Methods for identifying t-cell receptor (TCR) sequences

By comparing TCR chain sequences from subjects with different disease states and HLA alleles, the method identifies therapeutically relevant TCR chains with high specificity and accuracy, addressing the limitations of current TCR repertoire analysis.

WO2025213105A1PCT designated stage Publication Date: 2025-10-09BIONTECH US INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/023272
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-04
Filing Date
2025-04-04
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Current methods for analyzing T-cell receptor (TCR) repertoires are limited by access to large, real-world patient cohorts with genomic, TCR, and clinical profiling, hindering the characterization of TCR sequences and identifying therapeutically relevant TCR chains.

Method used

A method involving the comparison of TCR chain sequences from two sub-groups of subjects with different disease states, antigen profiles, or HLA alleles to identify TCR chains with statistically significant frequency differences, using statistical significance (p-value ≤ 0.1) and sequence similarity, and pairing alpha and beta chains for therapeutic relevance.

Benefits of technology

Enables the identification of therapeutically relevant TCR sequences with high specificity and accuracy, potentially applicable to diseases like HPV infection and cancer, by leveraging bulk and single-cell TCR sequencing data from diverse patient cohorts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025023272_09102025_PF_FP_ABST
    Figure US2025023272_09102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides methods for identifying T-cell receptors (TCRs) from sequencing data. The methods can comprise identifying a TCR alpha chain, a TCR beta chain, a TCR gamma chain, or a TCR delta chain from the sequencing data, and then identifying the corresponding paired chain. The "de novo" method can comprise the identification of statistically associated TCR chain sequences by assessing patients' antigenic status and / or HLA alleles. In the "bait" method, a TCR chain sequence of a TCR known to have some degree of antigen reactivity can be used to search for related sequences in the large database. Resulting sequences may or may not exhibit statistical enrichments that can be detected using the "de novo" approach. Various pairing methods are provided. Identified paired TCR chains can be used for cell-based therapies.
Need to check novelty before this filing date? Find Prior Art

Description

WSGR Docket No.: 50401-784.601 METHODS FOR IDENTIFYING T-CELL RECEPTOR (TCR) SEQUENCES CROSS-REFERENCE

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 574,530, filedon April 4, 2024, which is incorporated herein by reference in its entirety. BACKGROUND OF THE INVENTION

[0002] The T-cell receptor (TCR), located on the surface of T cells, is responsible for therecognition of the antigen-major histocompatibility complex, leading to the initiation of an inflammatory response. Analyzing the TCR repertoire may help to gain a better understanding of the immune system features and of the etiology and progression of diseases, in particular those with unknown antigenic triggers. The extreme diversity of the TCR repertoire represents a major analytical challenge; this has led to the development of specialized methods which aim to characterize the TCR repertoire in-depth. Currently, next generation sequencing based technologies are most widely employed for the high-throughput analysis of the immune cell repertoire. Despite recent advances in repertoire sequencing methods, the characterization of TCR repertoires has been limited by access to large, real-world patient cohorts with genomic, TCR, and clinical profiling. SUMMARY OF THE INVENTION

[0003] Recognized herein is a need to provide methods for analyzing the TCR chain sequencesand identifying therapeutically relevant TCR chain sequences from available clinicogenomic database from a large number of patients.

[0004] Provided herein is a method of identifying therapeutically relevant T-cell receptor (TCR)sequences, the method comprising: providing a plurality of libraries of TCR chain sequences having a first plurality of libraries of TCR chain sequences from a first sub-group of at least 10 subjects and a second plurality of libraries of TCR chain sequences from a second sub-group of at least 10 subjects, wherein each library of TCR chain sequences from a subject of the first sub- group of at least 10 subjects is from a single subject and each library of TCR chain sequences from a subject of the second sub-group of at least 10 subjects is from a single subject; wherein: each of the subjects of the first sub-group of at least 10 subjects has the same disease or condition, and each of the subjects of the second sub-group of at least 10 subjects does not have the same disease or condition as the first sub-group of at least 10 subjects; each of the subjects of the first sub-group of at least 10 subjects has the same antigen profile, and each of the subjects of the second sub-group of at least 10 subjects does not have the same antigen profile as the first sub-WSGR Docket No.: 50401-784.601 group of at least 10 subjects; and / or each of the subjects of the first sub-group of at least 10 subjects express a major histocompatibility complex (MHC) encoded by a same human leukocyte antigen (HLA) allele, and each of the subjects of the second sub-group of at least 10 subjects do not express an MHC encoded by the same HLA allele as the first sub-group of at least 10 subjects; identifying one or more TCR chain sequences based on an analysis of TCR chain sequences in the plurality of libraries of TCR chain sequences, wherein the analysis comprises comparing a frequency of one or more TCR chain sequences from the first sub-group to a frequency of the same one or more TCR chain sequences from the second sub-group, wherein the frequency of one or more TCR chain sequences from the first sub-group is the number of subjects of the first sub- group with the one or more TCR chain sequences over the total number of unique TCR chain sequences of the first sub-group, and the frequency of the same one or more TCR chain sequences from the second sub-group is the number of subjects of the second sub-group with the same one or more TCR chain sequence over the total number of unique TCR chain sequences of the second sub-group; determining a statistical significance of a difference in a frequency of a TCR chain sequence of the one or more TCR chain sequences from the first sub-group of at least 10 subjects to a frequency of the same TCR chain sequence of the one or more TCR chain sequences from the second sub-group of subjects based on the comparing in (b), thereby identifying a potentially therapeutic TCR chain sequence from the one or more TCR chain sequences obtained from the first sub-group of at least 10 subjects; and selecting a TCR chain sequence of the first sub-group of at least 10 subjects that (i) has a statistically significant difference in frequency to the frequency of the same TCR chain sequence in the second sub-group of subjects based on (c), wherein the statistically significant difference in the frequency is a p-value of at most 0.1, (ii) is present at a higher frequency in the first sub-group of at least 10 subjects relative to the frequency of the same TCR chain sequence in the second sub-group of subjects, and (iii) is a TCR chain sequence that is present in at least 2 subjects (e.g., at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 500, at least 1,000 or more subjects) from the first sub-group of at least 10 subjects.

[0005] In some embodiments, the number of subjects of the first sub-group of at least 10 subjectswith the TCR chain sequence selected in (d) over the total number of unique TCR chain sequences of the first sub-group of at least 10 subjects has a statistically significant difference compared to the number of subjects of the second sub-group of subjects with the TCR chain sequence selected in (d) over the total number of unique TCR chain sequences of the second sub-group of subjects. In some embodiments providing in (a) comprises: providing the plurality of libraries of TCR chain sequences from a population of subjects; identifying the first sub-group of at least 10 subjects asWSGR Docket No.: 50401-784.601 (i) having the same disease or condition, (ii) having the same antigen profile, (iii) expressing the MHC encoded by the same HLA allele, or (iv) the same combination thereof, and the second sub- group of subjects as (i) not having the same disease or condition as the first sub-group, (ii) not having the same antigen profile as the first sub-group, (iii) not expressing an MHC encoded by the same HLA allele as the first sub-group, or (iv) not having the same combination thereof as the first sub-group. In some embodiments, the statistically significant difference in the frequency is a p-value of at most 0.05, at most 0.001, or at most 0.0001. In some embodiments, the p-value is an adjusted p-value. In some embodiments, the selected TCR chain sequence is a TCR chain sequence that is present in at least 10 subjects from the first sub-group of at least 10 subjects. In some embodiments, the first sub-group of at least 10 subjects have the same disease or condition, and the second sub-group of at least 10 subjects do not have the same disease or condition as the first sub-group. In some embodiments, the method further comprises repeating (a)-(d) using the plurality of libraries of TCR chain sequences from a first sub-group of at least 10 subjects that express an MHC encoded by a same HLA allele, and a second sub-group of at least 10 subjects that do not express an MHC encoded by the same HLA allele as the first sub-group. In some embodiments, the first sub-group of at least 10 subjects have the same antigen profile, and the second sub-group of at least 10 subjects do not have the same antigen profile as the first sub- group. In some embodiments, the method further comprises repeating (a)-(d) using the plurality of libraries of TCR chain sequences from a first sub-group of at least 10 subjects that express an MHC encoded by a same HLA allele, and a second sub-group of at least 10 subjects that do not express an MHC encoded by the same HLA allele as the first sub-group. In some embodiments, the first sub-group of at least 10 subjects have the same disease or condition and the same antigen profile, and the second sub-group of at least 10 subjects do not have the same disease or condition as the first sub-group and does not have the same antigen profile as the first sub-group. In some embodiments, the method further comprises repeating (a)-(d) using the plurality of libraries of TCR chain sequences from a first sub-group of at least 10 subjects that express an MHC encoded by a same HLA allele, and a second sub-group of at least 10 subjects that do not express an MHC encoded by the same HLA allele as the first sub-group.

[0006] In some embodiments, the first sub-group of at least 10 subjects express an MHC encodedby a same HLA allele, and a second sub-group of at least 10 subjects do not express an MHC encoded by the same HLA allele as the first sub-group. In some embodiments, the method further comprises repeating (a)-(d) using the plurality of libraries of TCR chain sequences from a first sub-group of at least 10 subjects that have the same disease or condition, and the second sub-group of at least 10 subjects that do not have the same disease or condition as the first sub-group. In some embodiments, the method further comprises repeating (a)-(d) using the plurality of librariesWSGR Docket No.: 50401-784.601 of TCR chain sequences from a first sub-group of at least 10 subjects that have the same antigen profile, and the second sub-group of at least 10 subjects that do not have the same antigen profile as the first sub-group. In some embodiments, the first sub-group of at least 10 subjects comprises at least 100, at least 1,000, or more subjects. In some embodiments, the second sub-group of at least 10 subjects comprises at least 100, at least 1,000, or more subjects. In some embodiments, the selected TCR chain sequence is a TCR alpha chain sequence.

[0007] In some embodiments, the method further comprises selecting a TCR beta chain sequencethat is present at a higher frequency in the first sub-group of at least 10 subjects relative to the frequency of the same TCR beta chain sequence in the second sub-group of at least 10 subjects and is present in at least 2 subjects (e.g., at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 500, at least 1,000 or more) from the first sub-group of at least 10 subjects. In some embodiments, further comprises identifying a TCR beta chain sequence associated with or cognately paired to the selected TCR alpha chain. In some embodiments, the method further comprises (i) providing single-cell TCR sequencing data from a subject having the same disease or condition, having the same antigen profile, expressing an MHC encoded by the same HLA allele, or having the same combination thereof, (ii) determining if a TCR alpha chain sequence is present in the single-cell TCR sequencing data that is the same as or has at most 1, 2, 3, 4 or 5 amino acid differences with the selected TCR alpha chain sequence, and (iii) identifying a TCR beta chain sequence associated with or cognately paired to the selected TCR alpha chain sequence in the single-cell TCR sequencing data. In some embodiments, the selected TCR chain sequence is a TCR beta chain sequence. In some embodiments, the method further comprises selecting a TCR alpha chain sequence that is present at a higher frequency in the first sub-group of at least 10 subjects relative to the frequency of the same TCR alpha chain sequence in the second sub-group of at least 10 subjects, and is present in at least 2 subjects (e.g., at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 500, at least 1,000 or more) from the first sub-group of at least 10 subjects.

[0008] In some embodiments, the method further comprises identifying a TCR alpha chainsequence associated with or cognately paired to the selected TCR beta chain. In some embodiments, the method further comprises (i) providing a single-cell TCR sequencing data from a subject having the same disease or condition, having the same antigen profile, expressing an MHC encoded by the same HLA allele, or having the same combination thereof, (ii) determining if a TCR beta chain sequence is present in the single-cell TCR sequencing data that is the same as or has at most 1, 2, 3, 4 or 5 amino acid differences with the selected TCR beta chain sequence,WSGR Docket No.: 50401-784.601 and (iii) identifying a TCR alpha chain sequence associated with or cognately paired to the selected TCR beta chain sequence in the single-cell TCR sequencing data. In some embodiments, the method further comprises pairing the TCR alpha chain sequence and the TCR beta chain sequence to form a paired TCR chain sequences. In some embodiments, the selected TCR chain sequence is a TCR gamma chain sequence. In some embodiments, the method further comprises selecting a TCR delta chain sequence that is present at a higher frequency in the first sub-group of at least 10 subjects relative to the frequency of the same TCR delta chain sequence in the second sub-group of at least 10 subjects, and is present in at least 5 subjects from the first sub-group of at least 10 subjects. In some embodiments, the method further comprises identifying a TCR delta chain sequence associated with or cognately paired to the selected TCR gamma chain.

[0009] In some embodiments, the method further comprises (i) providing a single-cell TCRsequencing data from a subject having the same disease or condition, having the same antigen profile, expressing an MHC encoded by the same HLA allele, or having the same combination thereof, (ii) determining if a TCR gamma chain sequence is present in the single-cell TCR sequencing data that is the same as or has at most 1, 2, 3, 4 or 5 amino acid differences with the selected TCR gamma chain sequence, and (iii) identifying a TCR delta chain sequence associated with the selected TCR gamma chain sequence in the single-cell TCR sequencing data. In some embodiments, the selected TCR chain sequence is a TCR delta chain sequence. In some embodiments, the method further comprises selecting a TCR gamma chain sequence that is present at a higher frequency in the first sub-group of at least 10 subjects relative to the frequency of the same TCR gamma chain sequence in the second sub-group of at least 10 subjects, and is present in at least 5 subjects from the first sub-group of at least 10 subjects.

[0010] In some embodiments, the method further comprises identifying a TCR gamma chainsequence associated with or cognately paired to the selected TCR delta chain. In some embodiments, the method further comprises (i) providing a single-cell TCR sequencing data from a subject having the same disease or condition, having the same antigen profile, expressing an MHC encoded by the same HLA allele, or having the same combination thereof, (ii) determining if a TCR delta chain sequence is present in the single-cell TCR sequencing data that is the same as or has at most 1, 2, 3, 4 or 5 amino acid differences with the selected TCR delta chain sequence, and (iii) identifying a TCR gamma chain sequence associated with the selected TCR delta chain sequence in the single-cell TCR sequencing data. In some embodiments, the method further comprises pairing the TCR gamma chain sequence and the TCR delta chain sequence to form a paired TCR chain sequences. In some embodiments, a TCR comprising the paired TCR chain sequences is not a cognate TCR chain pair. In some embodiments, the same condition is HPV infection, and wherein a TCR having the paired TCR chain sequences binds to an epitopeWSGR Docket No.: 50401-784.601 associated with the HPV infection in complex with an HLA molecule. In some embodiments, the first sub-group of at least 10 subjects comprise a same HLA allele. In some embodiments, the same antigen profile comprises a same oncogenic driver mutation, and wherein a TCR having the paired TCR chain sequences binds to a neoepitope comprising the same oncogenic driver mutation in complex with an HLA molecule. In some embodiments, the first sub-group of at least 10 subjects comprise a same HLA allele. In some embodiments, the same condition comprises HPV infection. In some embodiments, the same condition comprises a cancer. In some embodiments, the cancer is associated with an antigen recognized by a TCR having the paired TCR chain sequences. In some embodiments, the cancer is not associated with an antigen recognized by a TCR having the paired TCR chain sequences. In some embodiments, the same condition comprises an autoimmune disorder. In some embodiments, the same condition comprises an infectious disease. In some embodiments, the same antigen profile comprises having a same cancer mutation, a same neoantigen, a same tumor-associated antigen, or any combination thereof.

[0011] In some embodiments, the same cancer mutation comprises KRAS-G12C, KRAS-G12D,or KRAS-G12V mutation. In some embodiments, the same HLA allele comprises an allele selected from the group consisting of HLA-A:01:01, HLA-A:02:01, HLA-A:03:01, HLA- A:11:01, HLA-B:07:02, HLA-B:08:01, HLA-B:44:02, HLA-B:44:03, HLA-C:04:01, HLA- C:05:01, HLA-C:06:02, HLA-C:07:02, and HLA-C:08:02. In some embodiments, each library of TCR chain sequences is a sequencing dataset obtained by bulk sequencing of TCR chain sequences from a subject. In some embodiments, an epitope or an HLA allele that a TCR comprising the paired TCR chain sequences recognizes is unknown. In some embodiments, the method further comprises assaying a TCR comprising the paired TCR chain sequences for a binding affinity against an epitope associated with the same disease or condition or associated with the same antigen profile in complex with the same HLA allele. In some embodiments, a TCR comprising the selected TCR chain sequence recognizes an epitope from HPV E2. In some embodiments, a TCR comprising the selected TCR chain sequence recognizes an epitope from HPV E2 in complex with HLA-A:01:01. In some embodiments, the frequency of one or more TCR chains in each library of TCR chain sequences from the first sub-group of at least 10 subjects is at least about 2-fold the frequency of one or more TCR chains in each library of TCR chain sequences from the second sub-group of at least 10 subjects. In some embodiments, the statistical significance of a TCR chain sequence is determined by a Fisher’s exact test. In some embodiments, the method further comprises, prior to determining the statistical significance, grouping two or more TCR chain sequences of the one or more TCR chain sequences based on sequence identity into a grouped TCR chain sequences, wherein the grouped TCR chain sequences share at least 70% sequence identity. In some embodiments, variable regions of the grouped TCRWSGR Docket No.: 50401-784.601 chain sequences share at least 70% sequence identity. In some embodiments, CDR3 sequences of the grouped TCR chain sequences share at least 70% sequence identity. In some embodiments, the one or more TCR chain sequences comprise two or more TCR chain sequences that have been grouped based on sequence identity, and wherein the grouped TCR chain sequences share at least 70% sequence identity. In some embodiments, variable regions of the grouped TCR chain sequences share at least 70% sequence identity. In some embodiments, CDR3 sequences of the grouped TCR chain sequences share at least 70% sequence identity. In some embodiments, selecting the TCR chain sequences comprises selecting a plurality of TCR chain sequences, and wherein the method further comprises aligning the plurality of TCR chain sequences to obtain conserved residues or sequence motifs.

[0012] In some embodiments, the method further comprises generating a paring score for a TCRhaving the paired TCR chain sequences, wherein the paring score predicts likelihood of the TCR to be an antigen-specific TCR to be validated experimentally. In some embodiments, the pairing score is calculated based on enrichment p-value, single-cell hit analysis, de novo match analysis, and / or enrichment match analysis, or combinations thereof. In some embodiments, the enrichment p-value is at most 1e-10. In some embodiments, the single-cell hit comprises identification of the TCR in a single-cell sample. In some embodiments, the de novo match or enrichment match comprises identification of the paired TCR chain sequences at a radius of a centroid of at least 5 of another identified antigen-specific TCR.

[0013] In some embodiments, the method further comprises generating an enrichment score forthe TCR chain sequence or the plurality of TCR chain sequences selected in (d), wherein the enrichment score predicts likelihood of the TCR chain sequences or the plurality of TCR chain sequences to be an antigen-specific TCR when paired with a corresponding TCR chain to be validated experimentally. In some embodiments, the enrichment score is calculated based on performing a singleton condition-enrichment analysis, and / or a cluster-based condition- enrichment analysis. In some embodiments, the cluster-based condition-enrichment analysis comprises identification of the TCR chain sequence or the plurality of TCR chain sequences at a radius of a centroid of at least 5, at least 10, and / or at least 20, or combinations thereof.

[0014] In some embodiments, the method further comprises analyzing the likelihood of theselected TCR chain sequence to be generated by thymic selection. In some embodiments, the first sub-group of at least 10 subjects have the same antigen profile as the second sub-group of at least 10 subjects, and wherein the first sub-group of at least 10 subjects express a protein with a same antigen or RNA encoding the protein with a same antigen at a higher level than the expression of the protein with a same antigen or RNA encoding the protein with a same antigen in the second sub-group of at least 10 subjects. In some embodiments, the same antigen comprises a mutation.WSGR Docket No.: 50401-784.601 In some embodiments, the first sub-group of at least 10 subjects have the same disease or condition and the second sub-group of at least 10 subjects do not have the same disease or condition.

[0015] Further provided herein is a method of identifying therapeutically relevant T-cell receptor(TCR) sequences, the method comprising: providing a plurality of libraries of TCR chain sequences from at least 10 subjects, each library of TCR chain sequences is from a single subject; selecting a TCR chain sequence that is present in at least 2 subjects (e.g., at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 500, at least 1,000 or more subjects) from the at least 10 subjects; based on the TCR chain sequence selected in (b), subgrouping the plurality of libraries of TCR chain sequences into a first plurality of libraries of TCR chain sequences from a first sub-group of subjects and a second plurality of libraries of TCR chain sequences from a second sub-group of subjects, wherein the first plurality of libraries of TCR chain sequences from the first sub-group of subjects comprises the TCR chain sequence selected in (b), and the second plurality of libraries of TCR chain sequences from the second sub- group of subjects does not comprise the TCR chain sequence selected in (b); and determining whether the first sub-group of subjects and the second sub-group of subjects have differential expression level of an antigen. In some embodiments, determining in (d) comprises analyzing sequencing data of the first sub-group and the second sub-group. In some embodiments, the first sub-group of subjects and the second sub-group of subjects have differential expression level of the antigen, and wherein the antigen has a probability of being recognized by the TCR chain sequence selected in (b) when being presented by an MHC.

[0016] Further provided herein is a method of identifying therapeutically relevant T-cell receptor(TCR) sequences, the method comprising: (a) providing a reference TCR chain sequence and libraries of TCR chain sequences from a population of subjects, wherein each library is from a single subject of the population of subjects; (b) analyzing one or more TCR chain sequences of the libraries of TCR chain sequences to determine sequence similarity between the one or more TCR chain sequences and the reference TCR chain sequence, wherein the reference TCR chain sequence is from a known antigen-specific TCR chain; and (c) identifying one or more TCR chain sequences from the libraries of TCR chain sequences that have a sequence similarity with the reference TCR chain sequence, wherein an identified TCR chain sequence a same amino acid sequence encoded by the V-gene as the reference TCR chain sequence and no more than three amino acid variations in a CDR3 sequence compared to a CDR3 sequence of the reference TCR chain sequence, wherein the amino acid variations comprises amino acids substitutions, deletions and / or insertions. In some embodiments, the reference TCR chain sequence comprises a reference TCR alpha chain and a reference TCR beta chain, and analyzing one or more TCR chain sequencesWSGR Docket No.: 50401-784.601 comprises analyzing each TCR alpha chain sequence and each TCR beta chain sequence separately. In some embodiments, analyzing one or more TCR chain sequences in (b) comprises analyzing each TCR alpha chain sequence of the libraries of TCR sequences to determine sequence similarity between each TCR alpha chain sequence and a reference TCR alpha chain sequence. In some embodiments, the method further comprises analyzing each TCR beta chain sequence of the libraries of TCR chain sequences to determine sequence similarity between each TCR beta chain sequence and a reference TCR beta chain sequence paired natively with the reference TCR alpha chain sequence. In some embodiments, the method further comprises, for each TCR chain of the one or more TCR chain sequences identified in (c), determining whether the TCR chain is likely to be restricted by an HLA allele by conducting a statistical test assessing whether a frequency of the TCR chain in a first subgroup of the population of subjects positive for the HLA allele is different from a frequency of the TCR chain in a second subgroup of the population of subjects negative for the HLA allele, and wherein the p-value for the statistical test is at most 0.1. In some embodiments, the p-value is at most 0.01. In some embodiments, the p- value is at most 0.001. In some embodiments, the p-value is at most 0.0001.

[0017] In some embodiments, the method further comprises determining the likelihood ofidentified TCR chain sequence specificity to the same pMHC complex as the reference TCR chain sequence. In some embodiments, determining the likelihood comprises aligning a CDR3 sequence of the identified TCR chain sequences to the CDR3 sequence of the reference TCR chain sequence. In some embodiments, a TCR comprising the reference TCR chain sequence recognizes HLA-A*02-presented epitopes, HLA-A*11-presented epitopes, HLA-A*03-presented epitopes, or HLA-C*08-presented epitopes. In some embodiments, the identified TCR chain sequence is a potentially therapeutic TCR chain sequence if the HLA allele is the same allele recognized by the reference TCR chain sequence. In some embodiments, the HLA allele recognized by the identified TCR chain sequence and the reference TCR chain sequence is selected from the group consisting of HLA-A*02, HLA-A*11, HLA-A*03, and HLA-C*08. In some embodiments, a TCR comprising the reference TCR chain sequence is specific for an epitope from HPV E7. In some embodiments, a TCR comprising the identified TCR chain sequence is specific for an epitope of HPV E7. In some embodiments, a TCR comprising the identified TCR chain sequence is specific for the epitope from HPV E7 in complex with HLA-A*02. In some embodiments, a TCR comprising the reference TCR chain sequence is specific for an epitope from KRAS-G12C. In some embodiments, a TCR comprising the reference TCR chain sequence is specific for the epitope from KRAS-G12C in complex with HLA-A*11. In some embodiments, a TCR comprising the identified TCR chain sequence is specific for the epitope from KRAS-G12C in complex with HL-A*11. In some embodiments, a TCR comprising the reference TCR chainWSGR Docket No.: 50401-784.601 sequence is specific for an epitope from KRAS-G12D. In some embodiments, a TCR comprising the reference TCR chain sequence is specific for the epitope from KRAS-G12D in complex with HLA-A*03, HLA-A*11, or HLA-C*08. In some embodiments, a TCR comprising the identified TCR chain sequence is specific for the epitope from KRAS-G12D in complex with HLA-A*03, HLA-A*11, or HLA-C*08. In some embodiments, a TCR comprising the reference TCR chain sequence is specific for an epitope from KRAS-G12V. In some embodiments, a TCR comprising the reference TCR chain sequence is specific for the epitope from KRAS-G12V in complex with HLA-A*03 or HLA-A*11. In some embodiments, a TCR comprising the identified TCR chain sequence is specific for the epitope from KRAS-G12V in complex with HLA-A*03 or HLA- A*11. In some embodiments, the method further comprises analyzing the likelihood for the identified TCR chain to be generated by thymic selection. In some embodiments, the reference TCR chain sequence is a TCR chain sequence selected by the methods described above.

[0018] Further provided herein is a method of making a TCR comprising preparing a TCRcomprising a TCR chain sequence selected according to the methods described above.

[0019] Further provided herein is a method of making a TCR comprising preparing a TCRcomprising a TCR chain sequence identified according to the methods described above and the reference TCR chain sequence.

[0020] Further provided herein is a pharmaceutical composition comprising a TCR comprising aTCR chain sequence selected according to the methods described above or a cell comprising the TCR comprising the TCR chain sequence selected according to the methods described above, and a pharmaceutically acceptable carrier.

[0021] Further provided herein is a pharmaceutical composition comprising a TCR comprising aTCR chain sequence selected according to the methods described above and the reference TCR chain sequence or a cell comprising the TCR comprising the TCR chain sequence selected according to the methods described above and the reference TCR chain sequence, and a pharmaceutically acceptable carrier.

[0022] Further provided herein is a method of treating a subject in need thereof comprisingadministering the pharmaceutical composition described above into the subject.

[0023] Further provided herein is a use of a TCR comprising a TCR chain sequence selectedaccording to the methods described above, a cell comprising the TCR comprising the TCR chain sequence selected according to the methods described above, a TCR comprising a TCR chain sequence selected according to the methods described above and the reference TCR chain sequence, a cell comprising the TCR comprising the TCR chain sequence selected according to the methods described above and the reference TCR chain sequence, or a pharmaceutical composition described above in the manufacture of a medicament for treating a disease.WSGR Docket No.: 50401-784.601 INCORPORATION BY REFERENCE

[0024] All publications, patents, and patent applications mentioned in this specification are hereinincorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The novel features of the present disclosure are set forth with particularity in the appendedclaims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:

[0026] FIG.1A depicts a schematic example of the de novo method of identifying T-cell receptors(TCRs) from a plurality of libraries of TCR chain sequences.

[0027] FIG. 1B depicts a schematic example of the bait method of identifying TCR chains froma plurality of libraries of TCR chain sequences.

[0028] FIG. 2 depicts the results of the de novo method for identifying statistically associatedTCR chain sequences for a particular condition.

[0029] FIG. 3A depicts the results of the bait method for identifying statistically associated TCRchain sequences for a particular condition.

[0030] FIG. 3B depicts sequence logo plots for identified neighbors of each bait (e.g., the knownreference TCR chain sequence(s)) to determine if particular amino acids are enriched in any CDR3 position relative to the bait sequence.

[0031] FIG. 4 depicts an association study using bulk TCR-Seq to identify candidate antigen-specific TCR chains.

[0032] FIG. 5A depicts sample counts for HPV-relevant cancers.

[0033] FIG. 5B depicts sample counts for KRAS-relevant cancers.

[0034] FIG. 6 depicts a method of identifying alpha and beta chains enriched in patientpopulations of interest.

[0035] FIG. 7A depicts the Fisher exact tests of association of one TCR chain by condition andby allele.

[0036] FIG. 7B depicts a plot of alpha and beta chains enriched in patient populations of interestthat were identified by statistical association.

[0037] FIG. 8A depicts the number of highest-tier candidate chains by across HPV and mutantKRAS targets.

[0038] FIG. 8B depicts the plotting and identification of dozens of candidate TCR alpha chainsfor KRAS-G12D including the highest-tier candidate chains.WSGR Docket No.: 50401-784.601

[0039] FIG. 9A depicts crystal structures of Ros9a and Ros9d TCRs in complex with HLA-C*08:02 / KRAS-G12D 9mer. See Sim et al.2020. High-affinity oligoclonal TCRs define effective adoptive T cell therapy targeting mutant KRAS-G12D. PNAS 117(23): 12826-12835.

[0040] FIG. 9B depicts the distribution of TCR-pMHC affinities. Data are obtained fromhttp: / / atlas.wenglab.org / .

[0041] FIG. 9C depicts T Cell Receptor Alpha Variable (TRAV) genes, T Cell Receptor Alphaof the Ros9 family of TCRs. Dots indicate amino acids matching those of the sequences in the first row of the table. Dashes indicate gaps relative to the sequences in the first row. Long dashes indicate the same V / J genes as the first row.

[0042] FIG. 10A depicts a method of identifying paired TCR chains for candidate alpha and betachains using an enrichment analysis.

[0043] FIG. 10B depicts a method of identifying paired TCR chains for candidate alpha and betachains using publicly available single-cell TCR-Seq (scTCR).

[0044] FIG. 11KDs in nM of the Ros9 family of TCR and selected new candidate TCR chains from the analysis and a portion of the Ros9a, Ros9d crystal structures. Dots indicate amino acids matching those of the sequences in the first row of the table. Dashes indicate gaps relative to the sequences in the first row. Long dashes indicate the same V / J genes as the first row.

[0045] FIG. 12A depicts the plotting and identification of candidate TCR alpha chains for HPVincluding the highest-tier candidate chains.

[0046] FIG. 12B depicts scTCR data coming from patients with HPV+ HNSC tumors. Dotsindicate amino acids matching those of the sequences in the first row of the table. Dashes indicate gaps relative to the sequences in the first row. Long dashes indicate the same V / J genes as the first row.

[0047] FIG. 12C depicts scTCR samples collected from HPV+ HNSC tumors.

[0048] FIG. 13A depicts a comparison of paired chains by condition identified using anenrichment analysis versus both the enrichment analysis and scTCR.

[0049] FIG. 13B depicts the number of candidate alpha-beta pairings by condition identified bythe de novo method.

[0050] FIG.14A depicts a method of validating candidate TCRs using surface plasmon resonance(SPR).

[0051] FIG. 14B depicts a method of validating candidate TCRs using cell lines.WSGR Docket No.: 50401-784.601

[0052] FIG. 15A depicts the plotting and identification of candidate TCR beta chains for HPVusing the “bait” method.

[0053] FIG. 15B depicts sequences of NCI TCR chains and close neighbors discovered from the"bait” method.

[0054] FIG. 15C depicts Binding affinities of NCI TCR chains and close neighbors measured bySPR.

[0055] FIG. 16 depicts the condition enrichment score, the pairing score, and the combined scoreof analyzing a test of 12 candidate HPV-specific TCR families. DETAILED DESCRIPTION Introduction

[0056] The present disclosure provides methods and systems for discovering therapeuticallyrelevant T-cell receptor (TCR) sequences from sequencing data, drawing upon two different methods. Both methods can depend on access to a database that includes bulk (or single-cell) TCR sequencing data from a large number of cancer patient tumor biopsies, where each patient / tumor can have an associated metadata (such cancer type / subtype, mutation profile, gene expression, HPV status, etc.). Utilizing an extensive plurality of patient libraries encompassing bulk TCR chain sequencing data from diverse cancer patient tumor biopsies, each enriched with relevant metadata, our methodology can provide a need in the field for identifying useful TCR chain sequences directly from sequencing data.

[0057] The “de novo” method can comprise the identification of statistically associated TCRchain sequences by assessing patients’ antigenic status and / or HLA alleles. For example, patients that do / don’t harbor an antigen or interest (e.g. HPV infection status or the presence / absence of a specific KRAS mutation) and patients that do / don’t express a specific HLA allele of interest can be identified and sub-grouped. Based on these criteria, whether a specific TCR chain sequence appears to be statistically associated with an antigen and / or HLA allele can be queried. Such TCR chains can be antigen-specific and can potentially be used as therapeutic molecules “as-is” or with further sequence optimization.

[0058] In the “bait” method, a TCR known to have some degree of antigen reactivity can be usedto search for related sequences in the large database. Resulting sequences may or may not exhibit statistical enrichments that can be detected using the “de novo” approach. Thus, the “bait” approach may be able to detect rarer TCR chains that lack sufficient counts to reach statistical significance.

[0059] To bolster the efficacy for the system, various additional methods can be performed beforeand / or after identifying one or more TCR chain sequence hits. For example, chain pairing can be performed since the results identify single chains (e.g., alpha, beta, gamma, or delta). There canWSGR Docket No.: 50401-784.601 be several approaches to find pairs for hits. If the hit is identified using the “bait” method, the corresponding chain of the bait TCR chain can be used to identify the corresponding pair of the hit to form the potential TCR pair. Auxiliary single-cell TCR (scTCR) sequencing databases can be used to see if they include the hit sequence or a close sequence neighbor thereof. Alternatively, alpha-beta sequences that frequently co-occur in the same patients in the original large bulk sequencing libraries can be identified.

[0060] In some cases, clustering can be performed. Rather than conducting the “de novo” analysisat the level of discrete TCR chains, it may be possible to first cluster the TCR chain sequences using a relevant similarity metric. This may improve the statistical power of the approach, especially for TCR chains that harbor rarer sequence motifs. For example, the method can comprise selecting a plurality of TCR chain sequences, and aligning the plurality of TCR chain sequences to obtain conserved residues or sequence motifs.

[0061] Motif-based triage can be performed. For example, hits from the “bait” and / or “de novo”methods may reveal sequence motifs such as residues that are conserved at certain positions. In this scenario, hits that conform to the overall motif may be considered as preferred candidates.

[0062] Generation probabilities can be performed. For example, the likelihood for a given TCRchain to be generated by thymic selection can be approximated using published techniques. For example to correct for the effect of thymic selection, a correction factor can be estimated. In some cases, a sequence-specific factor for each individual is learned. In some cases, all observed sequences passing thymic selection can be assumed. The correction factor can be a normalization factor accounting for the fact that just a fraction of sequences passes thymic selection. The correction factor can be determined for each VJ-combination. For example, the correction factor can be correction factor Q where the following factorized model for Q captures the main featuresof selection:where can be a CDR3 region, V and J can be a V and J gene choices, ( , , ) can be a full TCRchain sequence, (a1, …, aL) can be the amino acid sequence of the CDR3, and L is its length. The factors qL, qi;L(a), and qVJdenote selective pressures on the CDR3 length, its composition, and the associated VJ identities, respectively. (Pogorelyy et al. Elife. 2018; Elhanati et al. PNAS. 2014). Hits that are less probable may imply stronger selective forces to promote their frequency. Therefore, TCR chains with low generation probabilities may be considered as preferred candidates.

[0063] Iterative analysis can be performed. For example, hits from the “de novo” analysis maybe further used as starting sequences for the “bait” approach. “De novo” hits with marginalWSGR Docket No.: 50401-784.601 statistical significance may be rescued based on sequence similarity with a known TCR (e.g., a “bait”) or concordance with a known motif.

[0064] Chain fishing can be performed. For example, even if a known / putative alpha-beta pairingis available, additional potential partners for a given chain can be searched for using chain pairing approaches listed above (e.g., the single-cell approach and the database correlation approach). This can generate additional alpha-beta TCRs for consideration.

[0065] Continuous variables can be performed. For example, statistical analysis for the “de novo”method may not require that the antigen be strictly present / absent. Enrichment criteria can be devised for continuous variables, such as the expression of an antigenic gene (e.g., MAGEA4).

[0066] Antigen-agnostic analyses can be performed. For example, the “de novo” method maynot require the presence of an explicit antigen but may search for TCR chains enriched in a certain group of patients (e.g., a cancer subtype), possibly using healthy donors as controls. In these cases, additional analyses or experiments may be carried out to determine the antigen specificity of “hits”.

[0067] The methods and systems described herein are not limited to identifying alpha-beta TCRsand can be used to discover gamma-delta TCRs as well as BCRs and antibodies.

[0068] The methods and systems described herein are not limited to applications related to cancerand can be extended to discover TCRs relevant for infectious disease and autoimmunity. Definition

[0069] To facilitate an understanding of the present disclosure, a number of terms and phrases aredefined below. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0070] The term “about” or “approximately” means within an acceptable error range for theparticular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, “about” can mean within 1 or more than 1 standard deviation, per the practice in the art. Alternatively, “about” can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, preferably within 5-fold, and more preferably within 2- fold, of a value. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range for the particular value should be assumed.WSGR Docket No.: 50401-784.601

[0071] An “antigen” refers to a molecule, a moiety, a foreign substance, or an allergen that caninduce an immune response. A “antigen” can be a protein or fragment thereof that is encoded by a gene in a pathogen or a mutated gene. An antigen can be processed into peptide fragments and be bound by an MHC molecule, which can be recognized by a T-cell receptor.

[0072] An “epitope” is the collective features of a molecule (e.g., a peptide’s charge and primary,secondary and tertiary structure) that together form a site recognized by another molecule (e.g., an immunoglobulin, T-cell receptor, HLA molecule, or chimeric antigen receptor). For example, an epitope can be a set of amino acid residues involved in recognition by a particular immunoglobulin; a Major Histocompatibility Complex (MHC) receptor; or in the context of T cells, those residues recognized by a T-cell receptor protein and / or a chimeric antigen receptor. Epitopes can be prepared by isolation from a natural source, or they can be synthesized according to standard protocols in the art. Synthetic epitopes can comprise artificial amino acid residues, amino acid mimetics, (such as D isomers of naturally-occurring L amino acid residues or non- naturally-occurring amino acid residues). Throughout this disclosure, epitopes can be referred to in some cases as peptides or peptide epitopes. In certain embodiments, there is a limitation on the length of a peptide of the present disclosure. The embodiment that is length-limited occurs when the protein or peptide comprising an epitope described herein comprises a region (i.e., a contiguous series of amino acid residues) having 100% identity with a native sequence. In order to avoid the definition of epitope from reading, e.g., on whole natural molecules, there is a limitation on the length of any region that has 100% identity with a native peptide sequence. Thus, for a peptide comprising an epitope described herein and a region with 100% identity with a native peptide sequence, the region with 100% identity to a native sequence generally has a length of: less than or equal to 600 amino acid residues, less than or equal to 500 amino acid residues, less than or equal to 400 amino acid residues, less than or equal to 250 amino acid residues, less than or equal to 100 amino acid residues, less than or equal to 85 amino acid residues, less than or equal to 75 amino acid residues, less than or equal to 65 amino acid residues, and less than or equal to 50 amino acid residues. In certain embodiments, an “epitope” described herein is comprised by a peptide having a region with less than 51 amino acid residues that has 100% identity to a native peptide sequence, in any increment down to 5 amino acid residues; for example 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2 or 1 amino acid residues.

[0073] A “T-cell receptor” (“TCR”) refers to a molecule, whether natural or partly or whollysynthetically produced, found on the surface of T lymphocytes (T cells) that recognizes an antigen bound to a major histocompatibility complex (MHC) molecule. The ability of a T cells to recognize an antigen associated with various diseases (e.g., infectious diseases such as malaria)WSGR Docket No.: 50401-784.601 are encoded by DNA, which employs a unique mechanism for generating the tremendous diversity of the TCR. This multi-subunit immune recognition receptor associates with the CD3 complex and binds peptides presented by the MHC class I and II proteins on the surface of antigen- presenting cells (APCs). Binding of a TCR to a peptide on an APC is a central event in T cell activation.

[0074] “Major Histocompatibility Complex” or “MHC” is a cluster of genes or the proteinproducts thereof that plays a role in control of the cellular interactions responsible for physiologic immune responses. The terms “major histocompatibility complex” and the abbreviation “MHC” can include any class of MHC molecule, such as MHC class I and MHC class II molecules, and relate to a complex of genes which occurs in all vertebrates. In humans, the MHC complex is also known as the human leukocyte antigen (HLA) complex. Thus, a “Human Leukocyte Antigen” or “HLA” refers to a human Major Histocompatibility Complex (MHC) protein (see, e.g., Stites, et al., Immunology, 8TH Ed., Lange Publishing, Los Altos, Calif. (1994). For a detailed description of the MHC and HLA complexes, see, Paul, Fundamental Immunology, 3rd Ed., Raven Press, New York (1993).

[0075] The major histocompatibility complex in the genome comprises the genetic region whosegene products expressed on the cell surface are important for binding and presenting endogenous and / or foreign antigens and thus for regulating immunological processes. MHC proteins or molecules are important for signaling between lymphocytes and antigen-presenting cells or diseased cells in immune reactions. MHC proteins or molecules bind peptides and present them for recognition by T-cell receptors. The proteins encoded by the MHC can be expressed on the surface of cells and display both self-antigens (peptide fragments from the cell itself) and non- self-antigens (e.g., fragments of invading microorganisms) to a T-cell. MHC binding peptides can result from the proteolytic cleavage of protein antigens and represent potential lymphocyte epitopes. (e.g., T cell epitope and B cell epitope). MHCs can transport the peptides to the cell surface and present them there to specific cells, such as cytotoxic T-lymphocytes, T-helper cells, or B cells. The MHC region can be divided into three subgroups, class I, class II, and class III. by chromosome 15). They can present antigen fragments to cytotoxic T-cells. MHC class II MHC class III region can encode for other immune components, such as complement components and cytokines. The MHC can be both polygenic (there are several MHC class I and MHC class II genes) and polymorphic (there are multiple alleles of each gene).WSGR Docket No.: 50401-784.601

[0076] The term “motif” refers to a pattern of residues in an amino acid sequence of definedlength, for example, a peptide of less than about 15 amino acid residues in length, or less than about 13 amino acid residues in length, for example, from about 8 to about 13 amino acid residues (e.g., about 8, about 9, about 10, about 11, about 12, or about 13) for a class I HLA motif and from about 6 to about 25 amino acid residues (e.g., about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, or about 25) for a class II HLA motif, which is recognized by a particular HLA molecule. Motifs are typically different for each HLA protein encoded by a given human HLA allele. These motifs differ in their pattern of the primary and secondary anchor residues. In some embodiments, an MHC class I motif identifies a peptide of 7, 89, 10, 11, 12 or 13 amino acid residues in length. In some embodiments, an MHC class II motif identifies a peptide of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or 26 amino acid residues in length. A “cross-reactive binding” peptide refers to a peptide that binds to more than one member of a class of a binding pair members (e.g., a peptide bound by both a class I HLA molecule and a class II HLA molecule).

[0077] An “immune cell” refers to a cell that plays a role in the immune response. Immune cellsare of hematopoietic origin, and include lymphocytes, such as B cells and T cells; natural killer cells; myeloid cells, such as monocytes, macrophages, eosinophils, mast cells, basophils, and granulocytes.

[0078] “Pharmaceutically acceptable” refers to a generally non-toxic, inert, and / orphysiologically compatible composition or component of a composition. A “pharmaceutical excipient” or “excipient” comprises a material such as an adjuvant, a carrier, pH-adjusting and buffering agents, tonicity adjusting agents, wetting agents, preservatives, and the like. A “pharmaceutical excipient” is an excipient which is pharmaceutically acceptable.

[0079] The terms “identical” or percent “identity” in the context of two or more nucleic acids orpolypeptides, refer to two or more sequences or subsequences that are the same or have a specified percentage of nucleotides or amino acid residues that are the same, when compared and aligned (introducing gaps, if necessary) for maximum correspondence, not considering any conservative amino acid substitutions as part of the sequence identity. The percent identity can be measured using sequence comparison software or algorithms or by visual inspection. Various algorithms and software that can be used to obtain alignments of amino acid or nucleotide sequences are well- known in the art. These include, but are not limited to, BLAST, ALIGN, Megalign, BestFit, GCG Wisconsin Package, and variations thereof. In some embodiments, two nucleic acids or polypeptides described herein are substantially identical, meaning they have at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, and in some embodiments at least 95%, at least 96%, at least 97%, at least 98%, at least 99% nucleotide or amino acid residue identity, whenWSGR Docket No.: 50401-784.601 compared and aligned for maximum correspondence, as measured using a sequence comparison algorithm or by visual inspection. In some embodiments, identity exists over a region of the sequences that is at least about 10, at least about 20, at least about 40-60 residues, at least about 60-80 residues in length or any integral value there between. In some embodiments, identity exists over a longer region than 60-80 residues, such as at least about 80-100 residues, and in some embodiments the sequences are substantially identical over the full length of the sequences being compared, such as an amino acid sequence of a peptide or a coding region of a nucleotide sequence.

[0080] The term “subject” refers to any animal (e.g., a mammal), including, but not limited to,humans, non-human primates, canines, felines, rodents, and the like, which is to be the recipient of a particular treatment. Typically, the terms “subject” and “patient” are used interchangeably herein in reference to a human subject.

[0081] The terms “effective amount” or “therapeutically effective amount” or “therapeutic effect”refer to an amount of a therapeutic effective to “treat” a disease or disorder in a subject or mammal. The therapeutically effective amount of a drug has a therapeutic effect and as such can prevent the development of a disease or disorder; slow down the development of a disease or disorder; slow down the progression of a disease or disorder; relieve to some extent one or more of the symptoms associated with a disease or disorder; reduce morbidity and mortality; improve quality of life; or a combination of such effects.

[0082] The terms “treating” or “treatment” or “to treat” or “alleviating” or “to alleviate” refer toboth (1) therapeutic measures that cure, slow down, lessen symptoms of, and / or halt progression of a diagnosed pathologic condition or disorder. Thus, those in need of treatment include those already with the disorder. In some cases, treating may refer to reducing, or ameliorating a disorder and / or symptoms associated therewith (e.g., a neoplasia or tumor or infectious agent or an autoimmune disease). “Treating” can refer to administration of the therapy to a subject after the onset, or suspected onset, of a disease (e.g., cancer or infection by an infectious agent or an autoimmune disease). “Treating” includes the concepts of “alleviating”, which refers to lessening the frequency of occurrence or recurrence, or the severity, of any symptoms or other ill effects related to the disease and / or the side effects associated with therapy. The term “treating” may also encompass the concept of “managing” which refers to reducing the severity of a disease or disorder in a patient, e.g., extending the life or prolonging the survivability of a patient with the disease, or delaying its recurrence, e.g., lengthening the period of remission in a patient who had suffered from the disease. It is appreciated that, although not precluded, treating a disorder or condition does not require that the disorder, condition, or symptoms associated therewith be completely eliminated.WSGR Docket No.: 50401-784.601

[0083] The terms “prevent” or “prevention” refer to prophylactic or preventative measures thatslow down the development of a targeted pathologic condition or disorder. Thus, those in need of prevention include those prone to have the disorder or those in whom the disorder is to be prevented.

[0084] A “reference” can be used to correlate and / or compare the results obtained in the methodsof the present disclosure from a diseased specimen. Typically, a “reference” may be obtained on the basis of one or more normal specimens, in particular specimens which are not affected by a disease, either obtained from an individual or one or more different individuals (e.g., healthy individuals), such as individuals of the same species. A “reference” can be determined empirically by testing a sufficiently large number of normal specimens.

[0085] The term “Ros9 family of TCRs,” as used herein, refers to the known TCRs that recognizeKRAS G12D 9-mer presented by C*08:02 identified in Tran et al.2016. T-Cell Transfer Therapy Targeting Mutant KRAS in Cancer. New England Journal of Medicine 375:2255-2262; Leidner et al. 2022. Neoantigen T-Cell Receptor Gene Therapy in Pancreatic Cancer. New England Journal of Medicine 386: 2112-2119; and Sim et al. 2020. High-affinity oligoclonal TCRs define effective adoptive T cell therapy targeting mutant KRAS-G12D. PNAS 117(23): 12826-12835. The Ros9 family of TCRs can include Ros9a, Ros9b, Ros9c, and Ros9d. Methods of Identifying T-Cell Receptors (TCRs)

[0086] The present disclosure provides methods for discovering therapeutically relevant TCRs.The methods provided herein can depend on access to a database that includes bulk (or single- cell) TCR sequencing data from a large number of cancer patient tumor biopsies. Each patient / tumor can have an associated metadata (such cancer type / subtype, mutation profile, gene expression, HPV status, etc.).

[0087] In some cases, the methods are “de novo” methods, where patients that harbor or do notharbor an antigen or interest (e.g., HPV infection status or the presence / absence of a specific KRAS mutation) and patients that express or do not express a specific HLA allele of interest can be identified. Based on these criteria, whether a specific TCR chain sequence appears to be statistically associated with an antigen and / or HLA allele can be queried. Such TCR chains can be antigen-specific and can potentially be used as therapeutic molecules “as-is” or with further sequence optimization.

[0088] FIG.1A depicts a schematic example of the de novo method of identifying T-cell receptors(TCRs) from a plurality of libraries of TCR chain sequences. First, a plurality of libraries of TCR chain sequences can be sub-grouped libraries with and without a particular condition or set of conditions. (e.g., antigen expression, mutation profile, and / or HLA allele expression). The firstWSGR Docket No.: 50401-784.601 image depicts the result of sub-grouping a plurality of libraries of TCR chain sequences into “Sub- group A” and “Sub-group B.” Second, TCR alpha and / or beta chain sequences statistically associated with the condition or set of conditions are identified. The graph plots the -log10 of the p value of the condition against the -log10 of the p value of the allele to display statistically associated TCR chains. Third, and optionally, paired alpha / beta TCRs can be identified.

[0089] In some cases, the methods are “bait” methods. The “bait” method can start with a TCRknown to have some degree of antigen reactivity. Related sequences in the large database can then be searched for. Resulting sequences may or may not exhibit statistical enrichments that may have been detected using the “de novo” approach. Thus, the “bait” approach may be able to detect rarer TCRs that lack sufficient counts to reach statistical significance.

[0090] FIG. 1B depicts a schematic example of the bait method of identifying TCRs from aplurality of libraries of TCR chain sequences. First, TCR chain sequence(s) validated experimentally or from literature can be obtained. The image depicts TCR chain sequences validated experimentally or from literature. Second, similar sequences in a plurality of libraries of TCR chain sequences can be searched for. The image depicts an example sequence alignment of a plurality of libraries of TCR chain sequences with respect to the reference TCR chain sequence (e.g., alpha or beta chain from the reference TCR) validated experimentally or from literature. Third, TCR alpha and / or beta chain sequences statistically associated with the condition or set of conditions are identified. The graph plots the -log10 of the p value of the condition against the - log10 of the p value of the allele to display statistically associated TCR chains. Third, and optionally CDR3’s of the identified TCR chain sequences can be analyzed to obtain conserved amino acids in CDR3. The image is a sequence logo which is a graphical representation of the sequence conservation of nucleotides or amino acids. The image is a result of analyzing CDR3’s for conserved amino acids. Fourth, and optionally paired alpha / beta TCRs can be identified. De novo methods of identifying TCRs

[0091] The present disclosure provides a method of identifying therapeutically relevant T-cellreceptor (TCR) sequences. The method can comprise providing a plurality of libraries of TCR chain sequences. The plurality of libraries of TCR chain sequences can have a first plurality of libraries of TCR chain sequences from a first sub-group of at least 10 subjects (e.g., cancer patients) and a second plurality of libraries of TCR chain sequences from a second sub-group of at least 10 subjects. In some cases, the first sub-group or the second sub-group can comprise at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 140, at least 160, at least 180, at least 200, at least 240, at least 280, at least 320, at least 360, at least 400, at least 80, at least 560, at least 640, at least 720, at least 800, at least 920, at least 1040, at leastWSGR Docket No.: 50401-784.601 1160, at least 1280, at least 1400, at least 1600, at least 1800, at least 2000, at least 2200, at least 2400, at least 2800, at least 3200, at least 3600, at least 4000, at least 4400, at least 5000 or more subjects. Each library of TCR chain sequences from a subject of the first sub-group of at least 10 subjects can be from a single subject. Each library of TCR chain sequences from a subject of the second sub-group of at least 10 subjects can be from a single subject. Each of the subjects of the first sub-group of at least 10 subjects can have the same disease or condition. Each of the subjects of the second sub-group of at least 10 subjects may not have the same disease or condition as the first sub-group of at least 10 subjects. Each of the subjects of the first sub-group of at least 10 subjects can have the same antigen profile. Each of the subjects of the second sub-group of at least 10 subjects may not have the same antigen profile as the first sub-group of at least 10 subjects. Each of the subjects of the first sub-group of at least 10 subjects can express a major histocompatibility complex (MHC) encoded by a same human leukocyte antigen (HLA) allele. The HLA allele can be a class I HLA allele or a class II HLA allele. Each of the subjects of the second sub-group of at least 10 subjects may not express an MHC encoded by the same HLA allele as the first sub-group of at least 10 subjects. In some cases, each of the subjects of the first sub-group of at least 10 subjects can have the same disease or condition, the same antigen profile, and can express the MHC encoded by the same HLA allele. In some cases, each of the subjects of the second sub-group of at least 10 subjects may not have the same disease or condition, the same antigen profile, and may not express the MHC encoded by the same HLA allele. In some cases, the method can comprise providing a plurality of libraries of TCR chain sequences having a first plurality of libraries of TCR chain sequences from a first sub-group of at least 10 subjects and a second plurality of libraries of TCR chain sequences from a second sub-group of at least 10 subjects, where: (i) each of the subjects of the first sub-group of at least 10 subjects has the same disease or condition, and each of the subjects of the second sub-group of at least 10 subjects does not have the same disease or condition as the first sub-group of at least 10 subjects; (ii) each of the subjects of the first sub-group of at least 10 subjects has the same antigen profile, and each of the subjects of the second sub-group of at least 10 subjects does not have the same antigen profile as the first sub-group of at least 10 subjects; and / or (iii) each of the subjects of the first sub-group of at least 10 subjects express a major histocompatibility complex (MHC) encoded by a same human leukocyte antigen (HLA) allele, and each of the subjects of the second sub-group of at least 10 subjects do not express an MHC encoded by the same HLA allele as the first sub-group of at least 10 subjects.

[0092] The method can further comprise identifying one or more TCR chain sequences based onan analysis of TCR chain sequences in the plurality of libraries of TCR sequences. The analysis can comprise comparing a frequency of one or more TCR chain sequences from the first sub-WSGR Docket No.: 50401-784.601 group to a frequency of the same one or more TCR chain sequences from the second sub-group. The frequency of one or more TCR chain sequences from the first sub-group can be the number of subjects of the first sub-group with the one or more TCR chain sequences over the total number of unique TCR chain sequences of the first sub-group. The frequency of the same one or more TCR chain sequences from the second sub-group can be the number of subjects of the second sub- group with the same one or more TCR chain sequence over the total number of unique TCR chain sequences of the second sub-group. The total number of unique TCR chain sequences of the first sub-group or the second sub-group may comprise unique TCR chain sequences that are counted twice or more if the unique TCR chain sequences appear in two or more different subjects.

[0093] The method can further comprise determining a statistical significance of a difference in afrequency of a TCR chain sequence of the one or more TCR chain sequences from the first sub- group of at least 10 subjects to a frequency of the same TCR chain sequence of the one or more TCR chain sequences from the second sub-group of subjects based on the comparison described herein. A potentially therapeutic TCR chain sequence can be identified from the one or more TCR chain sequences obtained from the first sub-group of at least 10 subjects.

[0094] In some cases, determining a statistical significance may not comprise directly comparingthe frequency of a TCR chain sequence of the one or more TCR chain sequences from the first sub-group of at least 10 subjects to a frequency of the same TCR chain sequence of the one or more TCR chain sequences from the second sub-group of subjects. For example, when determining a statistical significance of a difference in a frequency of a TCR chain sequence of the one or more TCR chain sequences from the first sub-group of at least 10 subjects to a frequency of the same TCR chain sequence of the one or more TCR chain sequences from the second sub- group of subjects, four counts can be calculated, including counts of the candidate TCR chain of interest in the first sub-group, the total counts of all unique TCR chains in the first sub-group, counts of the candidate TCR chain of interest in the second sub-group, and the total counts of all unique TCR chains in the second sub-group. The total number of unique TCR chain sequences of the first sub-group or the second sub-group may comprise unique TCR chain sequences that are counted twice or more if the unique TCR chain sequences appear in two or more different subjects. The method can further comprise determining a statistical significance of a difference in a frequency of a TCR chain sequence of the one or more TCR chain sequences from the first sub-group of at least 10 subjects to a frequency of the same TCR chain sequence of the one or more TCR chain sequences from the second sub-group of subjects based on the calculated counts described herein.

[0095] Statistical significance refers to a mathematical tool used in hypothesis testing todetermine whether a particular result is beyond the realm of random chance. This concept isWSGR Docket No.: 50401-784.601 applied primarily when comparing two or more sets of data or testing the effect of certain variables within an experiment. The threshold of significance, often denoted by the p-value, is commonly set at 0.05, although this can be adjusted based on the circumstances or preferences of the researcher. If the calculated p-value falls below this threshold, it indicates that the likelihood of obtaining the observed data (or something more extreme) by mere chance is low, thereby suggesting that the observed results are statistically significant. This in turn provides evidence supporting the alternative hypothesis, or the assumption that some form of relationship exists between the variables being studied.

[0096] For example, statistical significance can be determined by the Fisher’s exact test, astatistical significance test used primarily in the analysis of small sample sizes. This test is employed when the numbers are too small for a Chi-square test to be applicable, to determine whether there are nonrandom associations between two categorical variables. An essential feature of the Fisher’s exact test is that it does not assume an equal distribution of probabilities for all possible outcomes, making it especially useful for studies with unequal, or 'non-standard', sample sizes. It calculates the exact probability of a specific distribution of outcomes, rather than approximating a probability based on a distribution, thus offering increased precision and reliability when dealing with small data sets. Some alternative tests to Fisher's exact test include the Chi-Square test, the Yates Correction test for continuity, or the Likelihood Ratio test.

[0097] FIG. 6 depicts a method of identifying TCR chains enriched in patient populations ofinterest and examples of the statistical analysis. First, all chains present in at least 5 patients (in some cases, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 or more patients) was identified. Second, for each TCR chain, a Fisher’s exact test is used to determine whether that TCR chain is over-represented in the HPV+ population. The table gives the input values for the Fisher’s exact test. Third, for each TCR chain, a Fisher’s exact test is used to determine whether that TCR chain is over-represented in any population defined by HLA class I alleles. The table give input values for the Fisher’s exact test. The results of the method are depicted by the two populations of circles- TCR chains from HPV+ patients and TCR chains from HPV- patients. Each circle represents an observation meaning 1 TCR chain in 1 patient. Dark circles represent TCR chains of interest. Light circles represent other TCR chains. Outline shade indicates Patient ID.

[0098] The method can further comprise selecting a TCR chain sequence of the first sub-group ofat least 10 subjects that has a statistically significant difference in frequency to the frequency of the same TCR chain sequence in the second sub-group of subjects as described herein. The statistically significant difference in the frequency can be a p-value of at most 0.1. In some cases, the statistically significant difference in the frequency can be a p-value of at most 0.1, at mostWSGR Docket No.: 50401-784.601 0.01, at most 0.001, at most 0.0001, at most 0.00001, at most 0.000001, at most 0.0000001, at most 0.00000001, at most 0.000000001, at most 0.0000000001, at most 0.00000000001, at most 0.000000000001, at most 0.0000000000001, at most 0.00000000000001, at most 0.000000000000001, at most 0.0000000000000001, at most 0.00000000000000001 or less. The TCR chain sequence can be present at a higher frequency in the first sub-group of at least 10 subjects relative to the frequency of the same TCR chain sequence in the second sub-group of subjects. The TCR chain sequence can be present in at least 5 subjects from the first sub-group of at least 10 subjects. In some cases, the TCR chain sequence can be present in at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 140, at least 160, at least 180, at least 200, at least 240, at least 280, at least 320, at least 360, at least 400, at least 480, at least 560, at least 640, at least 720, at least 800, at least 920, at least 1040, at least 1160, at least 1280, at least 1400, at least 1600, at least 1800, at least 2000, at least 2200, at least 2400, at least 2800, at least 3200, at least 3600, at least 4000, at least 4400, at least 5000 or more subjects from the first sub-group of at least 10 subjects.

[0099] In some cases, the method of identifying therapeutically relevant T-cell receptor (TCR)chain sequences can comprise (a) providing a plurality of libraries of TCR chain sequences having a first plurality of libraries of TCR chain sequences from a first sub-group of at least 10 subjects and a second plurality of libraries of TCR chain sequences from a second sub-group of at least 10 subjects, wherein each library of TCR chain sequences from a subject of the first sub-group of at least 10 subjects is from a single subject and each library of TCR chain sequences from a subject of the second sub-group of at least 10 subjects is from a single subject; wherein: each of the subjects of the first sub-group of at least 10 subjects has the same disease or condition, and each of the subjects of the second sub-group of at least 10 subjects does not have the same disease or condition as the first sub-group of at least 10 subjects; each of the subjects of the first sub-group of at least 10 subjects has the same antigen profile, and each of the subjects of the second sub- group of at least 10 subjects does not have the same antigen profile as the first sub-group of at least 10 subjects; and / or each of the subjects of the first sub-group of at least 10 subjects express a major histocompatibility complex (MHC) encoded by a same human leukocyte antigen (HLA) allele, and each of the subjects of the second sub-group of at least 10 subjects do not express an MHC encoded by the same HLA allele as the first sub-group of at least 10 subjects; (b) identifying one or more TCR chain sequences based on an analysis of TCR chain sequences in the plurality of libraries of TCR chain sequences, wherein the analysis comprises comparing a frequency of one or more TCR chain sequences from the first sub-group to a frequency of the same one or moreWSGR Docket No.: 50401-784.601 TCR chain sequences from the second sub-group , wherein the frequency of one or more TCR chain sequences from the first sub-group is the number of subjects of the first sub-group with the one or more TCR chain sequences over the total number of unique TCR chain sequences of the first sub-group, and the frequency of the same one or more TCR chain sequences from the second sub-group is the number of subjects of the second sub-group with the same one or more TCR chain sequence over the total number of unique TCR chain sequences of the second sub-group; (c) determining a statistical significance of a difference in a frequency of a TCR chain sequence of the one or more TCR chain sequences from the first sub-group of at least 10 subjects to a frequency of the same TCR chain sequence of the one or more TCR chain sequences from the second sub-group of subjects based on the comparing in (b), thereby identifying a potentially therapeutic TCR chain sequence from the one or more TCR chain sequences obtained from the first sub-group of at least 10 subjects; and (d) selecting a TCR chain sequence of the first sub- group of at least 10 subjects that (i) has a statistically significant difference in frequency to the frequency of the same TCR chain sequence in the second sub-group of subjects based on (c), wherein the statistically significant difference in the frequency is a p-value of at most 0.1, (ii) is present at a higher frequency in the first sub-group of at least 10 subjects relative to the frequency of the same TCR chain sequence in the second sub-group of subjects, and (iii) is a TCR chain sequence that is present in at least 5 subjects from the first sub-group of at least 10 subjects. In some cases, the TCR chain sequence can be present in at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 140, at least 160, at least 180, at least 200, at least 240, at least 280, at least 320, at least 360, at least 400, at least 480, at least 560, at least 640, at least 720, at least 800, at least 920, at least 1040, at least 1160, at least 1280, at least 1400, at least 1600, at least 1800, at least 2000, at least 2200, at least 2400, at least 2800, at least 3200, at least 3600, at least 4000, at least 4400, at least 5000 or more subjects from the first subgroup of at least 10 subjects.

[0100] In some cases, the number of subjects of the first sub-group of at least 10 subjects with theTCR chain sequence selected over the total number of unique TCR chain sequences of the first sub-group of at least 10 subjects can have a statistically significant difference compared to the number of subjects of the second sub-group of subjects with the TCR chain sequence selected over the total number of unique TCR chain sequences of the second sub-group of subjects.

[0101] In some cases, whether the subjects have the same disease or condition, whether thesubjects have the same antigen profile, and / or whether the subjects express the same MHC encoded by the same HLA allele may not be known. In such cases, the methods can further comprises, prior to identifying TCR chain sequences, identifying the first sub-group of at least 10WSGR Docket No.: 50401-784.601 subjects as (i) having the same disease or condition, (ii) having the same antigen profile, (iii) expressing the MHC encoded by the same HLA allele, or (iv) the same combination thereof, and the second sub-group of subjects as (i) not having the same disease or condition as the first sub- group, (ii) not having the same antigen profile as the first sub-group, (iii) not expressing an MHC encoded by the same HLA allele as the first sub-group, or (iv) not having the same combination thereof as the first sub-group.

[0102] The statistically significant difference in the frequency can be a p-value of at most 0.05, atmost 0.001, or at most 0.0001. In some cases the p value can be at most at most 0.01, at most 0.001, at most 0.0001, at most 0.00001, at most 0.000001, at most 0.0000001, at most 0.00000001, at most 0.000000001, at most 0.0000000001, at most 0.00000000001, at most 0.000000000001, at most 0.0000000000001, at most 0.00000000000001, at most 0.000000000000001, at most 0.0000000000000001, at most 0.00000000000000001 or less. The p-value can be an adjusted p- value. The selected TCR chain sequence can be a TCR chain sequence that is present in at least 10 subjects from the first sub-group of at least 10 subjects. The selected TCR chain sequence can be a TCR chain sequence that is present in at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 140, at least 160, at least 180, at least 200, at least 240, at least 280, at least 320, at least 360, at least 400, at least 480, at least 560, at least 640, at least 720, at least 800, at least 920, at least 1040, at least 1160, at least 1280, at least 1400, at least 1600, at least 1800, at least 2000, at least 2200, at least 2400, at least 2800, at least 3200, at least 3600, at least 4000, at least 4400, at least 5000 or more subjects from the first sub-group of at least 10 subjects.

[0103] In some cases, after selecting a candidate TCR chain sequence using the method (a)-(d)described above using the first sub-group and second sub-group separated based on one criterium (e.g., the same disease or condition, the method may be repeated using the first sub-group and the second sub-group separated based on other criteria (e.g., the HLA allele, the antigen, etc.). FIG. 4 depicts an example association study using bulk TCR-Seq to identify candidate antigen-specific TCR chains. Bulk TCR-seq data was segregated by TCR chain sequences statistically associated with and without HPV and with and without HLA-A*02:01 to produce 4 sub-groups. Dark gray figures represent patients with the query TCR chain. Light gray figures represent patients without query TCR chains. Candidate antigen-specific TCR chains can be enriched in the HPV+ HLA- A*02:01+ subgroup.

[0104] The first sub-group of at least 10 subjects can have the same disease or condition, and thesecond sub-group of at least 10 subjects may not have the same disease or condition as the first sub-group. The method can further comprise repeating the method (a)-(d) described above usingWSGR Docket No.: 50401-784.601 the plurality of libraries of TCR chain sequences from a first sub-group of at least 10 subjects that express an MHC encoded by a same HLA allele, and a second sub-group of at least 10 subjects that do not express an MHC encoded by the same HLA allele as the first sub-group.

[0105] The first sub-group of at least 10 subjects can have the same antigen profile. The secondsub-group of at least 10 subjects may not have the same antigen profile as the first sub-group. The method can further comprise repeating the method (a)-(d) as described above using the plurality of libraries of TCR chain sequences from a first sub-group of at least 10 subjects that express an MHC encoded by a same HLA allele, and a second sub-group of at least 10 subjects that do not express an MHC encoded by the same HLA allele as the first sub-group.

[0106] The first sub-group of at least 10 subjects can have the same disease or condition and thesame antigen profile, and the second sub-group of at least 10 subjects may not have the same disease or condition as the first sub-group and may not have the same antigen profile as the first sub-group. The method can further comprise repeating the method (a)-(d) as described above using the plurality of libraries of TCR chain sequences from a first sub-group of at least 10 subjects that express an MHC encoded by a same HLA allele, and a second sub-group of at least 10 subjects that do not express an MHC encoded by the same HLA allele as the first sub-group.

[0107] The first sub-group of at least 10 subjects may express an MHC encoded by a same HLAallele, and a second sub-group of at least 10 subjects may not express an MHC encoded by the same HLA allele as the first sub-group. The method can further comprise repeating the method (a)-(d) as described above using the plurality of libraries of TCR chain sequences from a first sub- group of at least 10 subjects that have the same disease or condition, and the second sub-group of at least 10 subjects that do not have the same disease or condition as the first sub-group. The method can further comprise repeating the method (a)-(d) as described above using the plurality of libraries of TCR chain sequences from a first sub-group of at least 10 subjects that have the same antigen profile, and the second sub-group of at least 10 subjects that do not have the same antigen profile as the first sub-group.

[0108] The first sub-group of at least 10 subjects can comprise at least 100, at least 1,000, or moresubjects. The first sub-group of at least 10 subjects can comprise at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 140, at least 160, at least 180, at least 200, at least 240, at least 280, at least 320, at least 360, at least 400, at least 480, at least 560, at least 640, at least 720, at least 800, at least 920, at least 960, at least 1000, at least 1040, at least 1160, at least 1280, at least 1400, at least 1600, at least 1800, at least 2000, at least 2200, at least 2400, at least 2800, at least 3200, at least 3600, at least 4000, at least 4400, at least 5000 or more subjects.WSGR Docket No.: 50401-784.601

[0109] The second sub-group of at least 10 subjects can comprise at least 100, at least 1,000, ormore subjects. The second sub-group of at least 10 subjects can comprise at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 140, at least 160, at least 180, at least 200, at least 240, at least 280, at least 320, at least 360, at least 400, at least 480, at least 560, at least 640, at least 720, at least 800, at least 920, at least 960, at least 1000, at least 1040, at least 1160, at least 1280, at least 1400, at least 1600, at least 1800, at least 2000, at least 2200, at least 2400, at least 2800, at least 3200, at least 3600, at least 4000, at least 4400, at least 5000 or more subjects.

[0110] The selected TCR chain sequence can be a TCR alpha chain sequence. The method canfurther comprise selecting a TCR beta chain sequence that is present at a higher frequency in the first sub-group of at least 10 subjects relative to the frequency of the same TCR beta chain sequence in the second sub-group of at least 10 subjects, and is present in at least 2 subjects (e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 140, at least 160, at least 180, at least 200, at least 240, at least 280, at least 320, at least 360, at least 400, at least 480, at least 560, at least 640, at least 720, at least 800, at least 920, at least 1040, at least 1160, at least 1280, at least 1400, at least 1600, at least 1800, at least 2000, at least 2200, at least 2400, at least 2800, at least 3200, at least 3600, at least 4000, at least 4400, at least 5000 or more subjects) from the first sub-group of at least 10 subjects. The method can further comprise identifying a TCR beta chain sequence associated with or cognately paired to the selected TCR alpha chain. The method can further comprise (i) providing single-cell TCR sequencing data from a subject having the same disease or condition, having the same antigen profile, expressing an MHC encoded by the same HLA allele, or having the same combination thereof, (ii) determining if a TCR alpha chain sequence is present in the single-cell TCR sequencing data that is the same as or has at most 1, 2, 3, 4 or 5 amino acid differences with the selected TCR alpha chain sequence, and (iii) identifying a TCR beta chain sequence associated with or cognately paired to the selected TCR alpha chain sequence in the single-cell TCR sequencing data.

[0111] The selected TCR chain sequence can be a TCR beta chain sequence. The method canfurther comprise selecting a TCR alpha chain sequence that is present at a higher frequency in the first sub-group of at least 10 subjects relative to the frequency of the same TCR alpha chain sequence in the second sub-group of at least 10 subjects, and is present in at least 2 subjects (e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 140, at least 160, at leastWSGR Docket No.: 50401-784.601 180, at least 200, at least 240, at least 280, at least 320, at least 360, at least 400, at least 480, at least 560, at least 640, at least 720, at least 800, at least 920, at least 1040, at least 1160, at least 1280, at least 1400, at least 1600, at least 1800, at least 2000, at least 2200, at least 2400, at least 2800, at least 3200, at least 3600, at least 4000, at least 4400, at least 5000 or more subjects) from the first sub-group of at least 10 subjects. The method can further comprise identifying a TCR alpha chain sequence associated with or cognately paired to the selected TCR beta chain. The method can further comprise (i) providing a single-cell TCR sequencing data from a subject having the same disease or condition, having the same antigen profile, expressing an MHC encoded by the same HLA allele, or having the same combination thereof, (ii) determining if a TCR beta chain sequence is present in the single-cell TCR sequencing data that is the same as or has at most 1, 2, 3, 4 or 5 amino acid differences with the selected TCR beta chain sequence, and (iii) identifying a TCR alpha chain sequence associated with or cognately paired to the selected TCR beta chain sequence in the single-cell TCR sequencing data.

[0112] The method can further comprise pairing the TCR alpha chain sequence and the TCR betachain sequence to form a paired TCR chain sequences.

[0113] The selected TCR chain sequence can be a TCR gamma chain sequence. The method canfurther comprise selecting a TCR delta chain sequence that is present at a higher frequency in the first sub-group of at least 10 subjects relative to the frequency of the same TCR delta chain sequence in the second sub-group of at least 10 subjects, and is present in at least 2 subjects (e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 140, at least 160, at least 180, at least 200, at least 240, at least 280, at least 320, at least 360, at least 400, at least 480, at least 560, at least 640, at least 720, at least 800, at least 920, at least 1040, at least 1160, at least 1280, at least 1400, at least 1600, at least 1800, at least 2000, at least 2200, at least 2400, at least 2800, at least 3200, at least 3600, at least 4000, at least 4400, at least 5000 or more subjects) from the first sub-group of at least 10 subjects. The method can further comprise identifying a TCR delta chain sequence associated with or cognately paired to the selected TCR gamma chain. The method can further comprise (i) providing a single-cell TCR sequencing data from a subject having the same disease or condition, having the same antigen profile, expressing an MHC encoded by the same HLA allele, or having the same combination thereof, (ii) determining if a TCR gamma chain sequence is present in the single-cell TCR sequencing data that is the same as or has at most 1, 2, 3, 4 or 5 amino acid differences with the selected TCR gamma chain sequence, and (iii) identifying a TCR delta chain sequence associated with the selected TCR gamma chain sequence in the single-cell TCR sequencing data.WSGR Docket No.: 50401-784.601

[0114] The selected TCR chain sequence can be a TCR delta chain sequence. The method canfurther comprise selecting a TCR gamma chain sequence that is present at a higher frequency in the first sub-group of at least 10 subjects relative to the frequency of the same TCR gamma chain sequence in the second sub-group of at least 10 subjects, and is present in at least 2 subjects (e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 140, at least 160, at least 180, at least 200, at least 240, at least 280, at least 320, at least 360, at least 400, at least 480, at least 560, at least 640, at least 720, at least 800, at least 920, at least 1040, at least 1160, at least 1280, at least 1400, at least 1600, at least 1800, at least 2000, at least 2200, at least 2400, at least 2800, at least 3200, at least 3600, at least 4000, at least 4400, at least 5000 or more subjects) from the first sub-group of at least 10 subjects. The method can further comprise identifying a TCR gamma chain sequence associated with or cognately paired to the selected TCR delta chain. The method can further comprise (i) providing a single-cell TCR sequencing data from a subject having the same disease or condition, having the same antigen profile, expressing an MHC encoded by the same HLA allele, or having the same combination thereof, (ii) determining if a TCR delta chain sequence is present in the single-cell TCR sequencing data that is the same as or has at most 1, 2, 3, 4 or 5 amino acid differences with the selected TCR delta chain sequence, and (iii) identifying a TCR gamma chain sequence associated with the selected TCR delta chain sequence in the single-cell TCR sequencing data.

[0115] The method can further comprise pairing the TCR gamma chain sequence and the TCRdelta chain sequence to form a paired TCR chain sequences.

[0116] The paired TCR chain sequences may not be a cognate TCR pair.

[0117] The same condition can be HPV infection. A TCR having the paired TCR chain sequencescan bind to an epitope associated with the HPV infection in complex with an HLA molecule. The first sub-group of at least 10 subjects can comprise a same HLA allele. The same antigen profile can comprise a same oncogenic driver mutation. A TCR having the paired TCR chain sequences can bind to a neoepitope comprising the same oncogenic driver mutation in complex with an HLA molecule. The first sub-group of at least 10 subjects can comprise a same HLA allele.

[0118] The same condition can comprise HPV infection. The same condition can comprise acancer. The cancer can be associated with an antigen recognized by a TCR having the paired TCR chain sequences. The cancer cannot be associated with an antigen recognized by a TCR having the paired TCR chain sequences. The same condition can comprise an autoimmune disorder. The same condition can comprise an infectious disease. The same antigen profile can comprise having a same cancer mutation, a same neoantigen, a same tumor-associated antigen, or any combinationWSGR Docket No.: 50401-784.601 thereof. The same cancer mutation can comprise KRAS-G12C, KRAS-G12D, or KRAS-G12V mutation. The same HLA allele can comprise an allele selected from the group consisting of HLA- A:01:01, HLA-A:02:01, HLA-A:03:01, HLA-A:11:01, HLA-B:07:02, HLA-B:08:01, HLA- B:44:02, HLA-B:44:03, HLA-C:04:01, HLA-C:05:01, HLA-C:06:02, HLA-C:07:02, and HLA- C:08:02.

[0119] Each library of TCR chain sequences can be a sequencing dataset obtained by bulksequencing of TCR chain sequences from a subject. The subject can be a cancer patient. The subject may also include a healthy subject.

[0120] An epitope or an HLA allele that a TCR comprising the paired TCR chain sequencesrecognizes can be unknown.

[0121] The method can further comprise assaying a TCR comprising the paired TCR chainsequences for a binding affinity against an epitope associated with the same disease or condition or associated with the same antigen profile in complex with the same HLA allele. For example, the binding affinity can be assayed by Surface Plasmon Resonance (SPR), isothermal titration calorimetry (ITC), or florescence anisotropy.

[0122] A TCR comprising the selected TCR chain sequence can recognize an epitope from HPVE2.

[0123] A TCR comprising the selected TCR chain sequence can recognize an epitope from HPVE2 in complex with HLA-A:01:01.

[0124] The frequency of one or more TCR chains in each library of TCR chain sequences fromthe first sub-group of at least 10 subjects can be at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 35-fold, at least about 40-fold, at least about 45-fold, at least about 50-fold, at least about 55-fold, at least about 60-fold, at least about 70-fold, at least about 80-fold, at least about 90-fold, at least about 100-fold or more the frequency of one or more TCR chains in each library of TCR chain sequences from the second sub-group of at least 10 subjects.

[0125] The statistical significance of a TCR chain sequence can be determined by a Fisher’s exacttest.

[0126] The method can further comprise, prior to determining the statistical significance,grouping two or more TCR chain sequences of the one or more TCR chain sequences based on sequence identity into a grouped TCR chain sequences. The grouped TCR chain sequences can share at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99% or more sequence identity. The variable regions of the groupedWSGR Docket No.: 50401-784.601 TCR chain sequences can share at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99% or more sequence identity. The CDR3 sequences of the grouped TCR chain sequences can share at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99% or more sequence identity.

[0127] The one or more TCR chain sequences can comprise two or more TCR chain sequencesthat have been grouped based on sequence identity. The grouped TCR chain sequences can share at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99% or more sequence identity. The variable regions of the grouped TCR chain sequences can share at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99% or more sequence identity. The CDR3 sequences of the grouped TCR chain sequences can share at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99% or more sequence identity.

[0128] Selecting the TCR chain sequences can comprise selecting a plurality of TCR chainsequences. The method can further comprise aligning the plurality of TCR chain sequences to obtain conserved residues or sequence motifs (e.g., clustering).

[0129] The method can further comprise generating a pairing score for a TCR having the pairedTCR chain sequences. The pairing score can predict likelihood of the TCR to be an antigen- specific TCR to be validated experimentally. In some cases, the pairing score can be calculated based on the metrics described in Table 6. In some cases, the pairing score can be calculated based on enrichment p-value analysis, single-cell hit analysis, de novo match analysis, and / or enrichment match analysis, or combinations thereof. In some cases, the enrichment p value can be at most 0.1, at most 0.01, at most 0.001, at most 0.0001, at most 0.00001, at most 0.000001, at most 0.0000001, at most 0.00000001, at most 0.000000001, at most 0.0000000001, at most 0.00000000001, at most 0.000000000001, at most 0.0000000000001, at most 0.00000000000001, at most 0.000000000000001, at most 0.0000000000000001, at most 0.00000000000000001 or less. In some cases, the enrichment p value can be at most 1e-10. In some cases, the single-cell hit analysis can comprise identification of the TCR in a single-cell sample. The single-cell sample can be derived from a sample from a disease type and / or an HLA allele of target. The de novo match analysis or enrichment match analysis can comprise identification of the paired TCR chain sequences at a radius of a centroid (e.g., TCRdist radius) of at least 2, at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, and / or at least 35, or combinations thereof, of another identified antigen-specific TCR. The radius of a centroid can refer to a TCR distance from a centroid sequence to other sequences in a given cluster.WSGR Docket No.: 50401-784.601

[0130] The method can further comprise generating an enrichment score for the TCR chainsequence or the plurality of TCR chain sequences selected, wherein the enrichment score predicts likelihood of the TCR chain sequences or the plurality of TCR chain sequences to be an antigen- specific TCR when paired with a corresponding TCR chain to be validated experimentally. The enrichment score can be calculated based on the metrics described in Table 7. In some cases, the enrichment score can be calculated based on performing a singleton condition-enrichment analysis, and / or a cluster-based condition-enrichment analysis. In some cases, the cluster-based condition-enrichment analysis comprises identification of the TCR chain sequence or the plurality of TCR chain sequences at a radius of a centroid (e.g., TCRdist radius) of at least 2, at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, and / or at least 35, or combinations thereof.

[0131] The method can further comprise analyzing the likelihood of the selected TCR chainsequence to be generated by thymic selection. For example to correct for the effect of thymic selection, a correction factor can be estimated. In some cases, a sequence-specific factor for each individual is learned. In some cases, all observed sequences passing thymic selection can be assumed. The correction factor can be a normalization factor accounting for the fact that just a fraction of sequences pass thymic selection. The correction factor can be determined for each VJ- combination. For example, the correction factor can be correction factor Q where the followingfactorized model for Q captures the main features of selection:where can be a CDR3 region, V and J can be a V and J gene choices, ( , , ) can be a full TCRchain sequence, (a1, …, aL) can be the amino acid sequence of the CDR3, and L is its length. The factors qL, qi;L(a), and qVJdenote selective pressures on the CDR3 length, its composition, and the associated VJ identities, respectively.

[0132] The first sub-group of at least 10 subjects can have the same antigen profile as the secondsub-group of at least 10 subjects. The first sub-group of at least 10 subjects may express a protein with a same antigen or RNA encoding the protein with a same antigen at a higher level than the expression of the protein with a same antigen or RNA encoding the protein with a same antigen in the second sub-group of at least 10 subjects. The same antigen can comprise a mutation.

[0133] The first sub-group of at least 10 subjects may have the same disease or condition and thesecond sub-group of at least 10 subjects may not have the same disease or condition.

[0134] An alternative method of the “de novo” methods can comprise analyzing differential targetantigen expression in two different groups of subjects (having a TCR chain or not having a TCR chain) given an arbitrary TCR chain sequence. The arbitrary TCR chain sequence can be selectedWSGR Docket No.: 50401-784.601 based on its presence in at least five subjects. For example, the method can comprise providing a plurality of libraries of TCR chain sequences from at least 10 subjects. Each library of TCR chain sequences can be from a single subject. The method can further comprise selecting a TCR chain sequence that is present in at least 5 subjects from the at least 10 subjects. In some cases, the TCR chain sequence can be present in at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 140, at least 160, at least 180, at least 200, at least 240, at least 280, at least 320, at least 360, at least 400, at least 480, at least 560, at least 640, at least 720, at least 800, at least 920, at least 1040, at least 1160, at least 1280, at least 1400, at least 1600, at least 1800, at least 2000, at least 2200, at least 2400, at least 2800, at least 3200, at least 3600, at least 4000, at least 4400, at least 5000 or more subjects from the first subgroup of at least 10 subjects. The method can further comprise, based on the TCR chain sequence selected, subgrouping the plurality of libraries of TCR chain sequences into a first plurality of libraries of TCR chain sequences from a first sub- group of subjects and a second plurality of libraries of TCR chain sequences from a second sub- group of subjects. The first plurality of libraries of TCR chain sequences from the first sub-group of subjects may comprise the TCR chain sequence selected, and the second plurality of libraries of TCR chain sequences from the second sub-group of subjects may not comprise the TCR chain sequence selected. The method can further comprise determining whether the first sub-group of subjects and the second sub-group of subjects may have differential expression level of an antigen.

[0135] Determining described herein can comprise analyzing sequencing data of the first sub-group and the second sub-group. The first sub-group of subjects and the second sub-group of subjects can have differential expression level of the antigen. The antigen can have a probability of being recognized by the TCR chain sequence selected when being presented by an MHC. The antigen described herein can be a neoantigen. The antigen can be a tumor associated antigen. The antigen may or may not contain a mutation. Bait method of identifying TCRs

[0136] The method of identifying therapeutically relevant T-cell receptor (TCR) sequences cancomprise providing a reference TCR chain sequence and libraries of TCR chain sequences from a population of subjects. Each library can be from a single subject of the population of subjects. The method can further comprise analyzing one or more TCR chain sequences of the libraries of TCR chain sequences to determine sequence similarity between the one or more TCR chain sequences and the reference TCR chain sequence. The reference TCR chain sequence can be from a known antigen-specific TCR. The method can further comprise identifying one or more TCR chain sequences from the libraries of TCR chain sequences that have a sequence similarity withWSGR Docket No.: 50401-784.601 the reference TCR chain sequence. An identified TCR chain sequence can have a same amino acid sequence encoded by the V-gene as the reference TCR chain sequence, and no more than three amino acid variations in a CDR3 sequence compared to a CDR3 sequence of the reference TCR chain sequence. The amino acid variations can comprise amino acids substitutions, deletions and / or insertions.

[0137] The reference TCR chain can comprise a reference TCR alpha chain sequence and areference TCR beta chain sequence from the known antigen-specific TCR. Analyzing one or more TCR chain sequences can comprise analyzing each TCR alpha chain sequence (or TCR gamma chain sequence) and each TCR beta chain sequence (or TCR delta chain sequence) separately.

[0138] Analyzing one or more TCR chain sequences can comprise analyzing each TCR alphachain sequence of the libraries of TCR chain sequences to determine sequence similarity between each TCR alpha chain sequence (or TCR gamma chain sequence) and a reference TCR alpha chain sequence (or a reference TCR gamma chain sequence).

[0139] The method can further comprise analyzing each TCR beta chain sequence (or TCR deltachain sequence) of the libraries of TCR chain sequences to determine sequence similarity between each TCR beta chain sequence (or TCR delta chain sequence) and a reference TCR beta chain sequence paired natively with the reference TCR alpha chain sequence (or TCR delta chain sequence paired natively with the reference TCR gamma chain sequence).

[0140] The method can further comprise, for each TCR chain of the one or more TCR chainsequences identified, determining whether the TCR chain is likely to be restricted by an HLA allele by conducting a statistical test assessing whether a frequency of the TCR chain in a first subgroup of the population of subjects positive for the HLA allele is different from a frequency of the TCR chain in a second subgroup of the population of subjects negative for the HLA allele. The p-value for the statistical test can be at most 0.1. In some cases, the p value can be at most 0.1, at most 0.01, at most 0.001, at most 0.0001, at most 0.00001, at most 0.000001, at most 0.0000001, at most 0.00000001, at most 0.000000001, at most 0.0000000001, at most 0.00000000001, at most 0.000000000001, at most 0.0000000000001, at most 0.00000000000001, at most 0.000000000000001, at most 0.0000000000000001, at most 0.00000000000000001 or less.

[0141] The p-value can be at most 0.01. The p-value can be at most 0.001. The p-value can beat most 0.0001.

[0142] The method can further comprise determining the likelihood of identified TCR chainsequence specificity to the same pMHC complex as the reference TCR chain sequence. Determining the likelihood can comprise aligning a CDR3 sequence of the identified TCR chainWSGR Docket No.: 50401-784.601 sequences to the CDR3 sequence of the reference TCR chain sequence. A TCR comprising the reference TCR chain sequence may recognize HLA-A*02-presented epitopes, HLA-A*11- presented epitopes, HLA-A*03-presented epitopes, or HLA-C*08-presented epitopes. The identified TCR chain sequence can be a potentially therapeutic TCR chain sequence if the HLA allele is the same allele recognized by the reference TCR chain sequence. The HLA allele recognized by the identified TCR chain sequence and the reference TCR chain sequence can be selected from the group consisting of HLA-A*02, HLA-A*11, HLA-A*03, and HLA-C*08.

[0143] A TCR comprising the reference TCR chain sequence can be specific for an epitope fromHPV E7.

[0144] A TCR comprising the identified TCR chain sequence can be specific for an epitope ofHPV E7.

[0145] A TCR comprising the identified TCR chain sequence can be specific for the epitope fromHPV E7 in complex with HLA-A*02.

[0146] A TCR comprising the reference TCR chain sequence can be specific for an epitope fromKRAS-G12C.

[0147] A TCR comprising the reference TCR chain sequence can be specific for the epitope fromKRAS-G12C in complex with HLA-A*11.

[0148] A TCR comprising the identified TCR chain sequence can be specific for the epitope fromKRAS-G12C in complex with HL-A*11.

[0149] A TCR comprising the reference TCR chain sequence can be specific for an epitope fromKRAS-G12D.

[0150] A TCR comprising the reference TCR chain sequence can be specific for the epitope fromKRAS-G12D in complex with HLA-A*03, HLA-A*11, or HLA-C*08.

[0151] A TCR comprising the identified TCR chain sequence can be specific for the epitope fromKRAS-G12D in complex with HLA-A*03, HLA-A*11, or HLA-C*08.

[0152] A TCR comprising the reference TCR chain sequence can be specific for an epitope fromKRAS-G12V.

[0153] A TCR comprising the reference TCR chain sequence can be specific for the epitope fromKRAS-G12V in complex with HLA-A*03 or HLA-A*11.

[0154] A TCR comprising the identified TCR chain sequence can be specific for the epitope fromKRAS-G12V in complex with HLA-A*03 or HLA-A*11.

[0155] The method can further comprise analyzing the likelihood for the identified TCR chain tobe generated by thymic selection. For example to correct for the effect of thymic selection, a correction factor can be estimated. In some cases, a sequence-specific factor for each individual is learned. In some cases, all observed sequences passing thymic selection can be assumed. TheWSGR Docket No.: 50401-784.601 correction factor can be a normalization factor accounting for the fact that just a fraction of sequences pass thymic selection. The correction factor can be determined for each VJ- combination. For example, the correction factor can be correction factor Q where the followingfactorized model for Q captures the main features of selection:where can be a CDR3 region, V and J can be a V and J gene choices, ( , , ) can be a full TCRchain sequence, (a1, …, aL) can be the amino acid sequence of the CDR3, and L is its length. The factors qL, qi;L(a), and qVJdenote selective pressures on the CDR3 length, its composition, and the associated VJ identities, respectively.

[0156] The reference TCR chain sequence can be a TCR chain sequence selected by the methoddescribed, for example, the de novo or bait methods described herein.

[0157] The present disclosure also provides methods of making the TCRs identified herein orcells comprising the TCRs identified. The method of making a TCR can comprise preparing a TCR comprising a TCR chain sequence selected according to the methods described herein. The method of making a cell comprising the identified TCR can comprise delivering a TCR sequence comprising a TCR chain sequence selected according to the methods described herein or a nucleic acid sequence encoding the TCR chain into the cell.

[0158] The method of making a TCR can comprise preparing a TCR comprising a TCR chainsequence identified according to the methods described herein and the reference TCR chain sequence.

[0159] The identified TCR chain can be used to prepare a pharmaceutical composition. Thepharmaceutical composition can comprise a TCR chain comprising a TCR chain sequence selected according to the method described, and a pharmaceutically acceptable carrier.

[0160] TCRs comprising the TCR chains identified by the methods described herein canrecognize epitopes in complex with an MHC molecule. The MHC molecules described herein can be encoded by a class I HLA allele, and / or a class II HLA allele. Identification of B-cell receptors or antibodies

[0161] The methods (e.g., de novo or bait methods) described herein are not limited to identifyingalpha-beta TCRs or gamma-delta TCRs and can be used to identify B-cell receptors (BCRs) and antibodies. For example, the method of identifying therapeutically relevant BCR or antibody sequences can comprise providing a plurality of libraries of BCR or antibody sequences having a first plurality of libraries of BCR or antibody sequences from a first sub-group of at least 10 subjects and a second plurality of libraries of BCR or antibody sequences from a second sub-group of at least 10 subjects. Each library of BCR or antibody sequences from a subject of the first sub-WSGR Docket No.: 50401-784.601 group of at least 10 subjects can be from a single subject and each library of BCR or antibody sequences from a subject of the second sub-group of at least 10 subjects can be from a single subject. In some cases, each of the subjects of the first sub-group of at least 10 subjects can have the same disease or condition, and each of the subjects of the second sub-group of at least 10 subjects may not have the same disease or condition as the first sub-group of at least 10 subjects. In some cases, each of the subjects of the first sub-group of at least 10 subjects can have the same antigen profile, and each of the subjects of the second sub-group of at least 10 subjects may not have the same antigen profile as the first sub-group of at least 10 subjects. The method can further comprise identifying one or more BCR or antibody sequences based on an analysis of BCR or antibody sequences in the plurality of libraries of BCR or antibody sequences. The analysis can comprise comparing a frequency of one or more BCR or antibody sequences from the first sub- group to a frequency of the same one or more BCR or antibody sequences from the second sub- group. The frequency of one or more BCR or antibody sequences from the first sub-group can be the number of subjects of the first sub-group with the one or more BCR or antibody sequences over the total number of unique BCR or antibody sequences of the first sub-group, and the frequency of the same one or more BCR or antibody sequences from the second sub-group can be the number of subjects of the second sub-group with the same one or more BCR or antibody sequence over the total number of unique BCR or antibody sequences of the second sub-group. The method can further comprise determining a statistical significance of a difference in a frequency of a BCR or antibody sequence of the one or more BCR or antibody sequences from the first sub-group of at least 10 subjects to a frequency of the same BCR or antibody sequence of the one or more BCR or antibody sequences from the second sub-group of subjects based on the comparison of the frequency above, thereby identifying a potentially therapeutic BCR or antibody sequence from the one or more BCR or antibody sequences obtained from the first sub-group of at least 10 subjects. The method can further comprise selecting a BCR or antibody sequence of the first sub-group of at least 10 subjects that has a statistically significant difference in frequency to the frequency of the same BCR or antibody sequence in the second sub-group of subjects based on the statistical significance, where the statistically significant difference in the frequency can be a p-value of at most 0.1. The method can further comprise selecting a BCR or antibody sequence of the first sub-group of at least 10 subjects that is present at a higher frequency in the first sub- group of at least 10 subjects relative to the frequency of the same BCR or antibody sequence in the second sub-group of subjects. The method can further comprise selecting a BCR or antibody sequence of the first sub-group of at least 10 subjects that is a BCR or antibody sequence that is present in at least 2 subjects (e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40,WSGR Docket No.: 50401-784.601 at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 140, at least 160, at least 180, at least 200, at least 240, at least 280, at least 320, at least 360, at least 400, at least 480, at least 560, at least 640, at least 720, at least 800, at least 920, at least 1040, at least 1160, at least 1280, at least 1400, at least 1600, at least 1800, at least 2000, at least 2200, at least 2400, at least 2800, at least 3200, at least 3600, at least 4000, at least 4400, at least 5000 or more subjects) from the first sub-group of at least 10 subjects.

[0162] The selected BCR or antibody sequence can be a heavy chain sequence. The method canfurther comprise selecting a light chain sequence that is present at a higher frequency in the first sub-group of at least 10 subjects relative to the frequency of the same light chain sequence in the second sub-group of at least 10 subjects, and is present in at least 2 subjects (e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 140, at least 160, at least 180, at least 200, at least 240, at least 280, at least 320, at least 360, at least 400, at least 480, at least 560, at least 640, at least 720, at least 800, at least 920, at least 1040, at least 1160, at least 1280, at least 1400, at least 1600, at least 1800, at least 2000, at least 2200, at least 2400, at least 2800, at least 3200, at least 3600, at least 4000, at least 4400, at least 5000 or more subjects) from the first sub-group of at least 10 subjects. The method can further comprise identifying a light chain sequence associated with or cognately paired to the selected heavy chain. The method can further comprise (i) providing single-cell BCR or antibody sequencing data from a subject having the same disease or condition and / or having the same antigen profile, (ii) determining if a heavy chain sequence is present in the single-cell BCR or antibody sequencing data that is the same as or has at most 1, 2, 3, 4 or 5 amino acid differences with the selected heavy chain sequence, and (iii) identifying a light chain sequence associated with or cognately paired to the selected heavy chain sequence in the single-cell BCR or antibody sequencing data. The same methods can be used to pair a heavy chain sequence to a selected light chain sequence. Pharmaceutical Compositions

[0163] Provided herein are compositions (e.g., pharmaceutical compositions) comprising a TCRchain sequence selected according to the methods described as, a nucleic acid molecule encoding the TCR chain, or an immune cell comprising the TCR chain or the nucleic acid molecule encoding the TCR chain. Pharmaceutical compositions can be formulated using one or more physiologically acceptable carriers including excipients and auxiliaries which facilitate processing of the active agents into preparations which can be used pharmaceutically. Proper formulation canWSGR Docket No.: 50401-784.601 be dependent upon the route of administration chosen. Any of the well-known techniques, carriers, and excipients can be used as suitable and as understood in the art.

[0164] In some cases, a pharmaceutical composition is formulated as cell based therapeutic, e.g.,a T cell therapeutic. In some embodiments, the pharmaceutical composition comprises a peptide- based therapy, a nucleic acid-based therapy, an antibody-based therapy, and / or a cell-based therapy. In some embodiments, a pharmaceutical composition comprises a peptide-based therapeutic, or nucleic acid based therapeutic in which the nucleic acid encodes the polypeptides. In some embodiments, a pharmaceutical composition comprises a peptide-based therapeutic, or nucleic acid based therapeutic in which the nucleic acid encodes the polypeptides; wherein the peptide-based therapeutic, or nucleic acid based therapeutic are comprised in a cell, wherein the cell is a T cell. In some embodiments, a pharmaceutical composition comprises as an antibody based therapeutic. A composition can comprise T cells specific for two or more immunogenic antigen or neoantigen peptides.

[0165] Pharmaceutical compositions can include, in addition to active ingredient, apharmaceutically acceptable excipient, carrier, buffer, stabilizer or other materials well known to those skilled in the art. Such materials may be non-toxic and may not interfere with the efficacy of the active ingredient. The precise nature of the carrier or other material will depend on the route of administration.

[0166] Acceptable carriers, excipients, or stabilizers are those that are non-toxic to recipients atthe dosages and concentrations employed, and include buffers such as phosphate, citrate, and other organic acids; antioxidants including ascorbic acid and methionine; preservatives (such as octadecyl dimethyl benzyl ammonium chloride; hexamethonium chloride; benzalkonium chloride, benzethonium chloride; phenol, butyl or benzyl alcohol; alkyl parabens such as methyl or propyl paraben; catechol; resorcinol; cyclohexanol; 3-pentanol; and m-cresol); low molecular weight (less than about 10 residues) polypeptides; proteins, such as serum albumin, gelatin, or immunoglobulins; hydrophilic polymers such as polyvinylpyrrolidone; amino acids such as glycine, glutamine, asparagine, histidine, arginine, or lysine; monosaccharides, disaccharides, and other carbohydrates including glucose, mannose, or dextrins; chelating agents such as EDTA; sugars such as sucrose, mannitol, trehalose or sorbitol; salt-forming counter-ions such as sodium; metal complexes (e.g., Zn-protein complexes); and / or non-ionic surfactants such as TWEEN®, PLURONICS® or polyethylene glycol (PEG).

[0167] Acceptable carriers are physiologically acceptable to the administered patient and retainthe therapeutic properties of the compounds with / in which it is administered. Acceptable carriers and their formulations are generally described in, for example, Remington’ pharmaceutical Sciences (18th ed. A. Gennaro, Mack Publishing Co., Easton, PA 1990). One example of carrierWSGR Docket No.: 50401-784.601 is physiological saline. A pharmaceutically acceptable carrier is a pharmaceutically acceptable material, composition or vehicle, such as a liquid or solid filler, diluent, excipient, solvent or encapsulating material, involved in carrying or transporting the subject compounds from the administration site of one organ, or portion of the body, to another organ, or portion of the body, or in an in vitro assay system. Acceptable carriers are compatible with the other ingredients of the formulation and not injurious to a subject to whom it is administered. Nor should an acceptable carrier alter the specific activity of the neoantigens.

[0168] In one aspect, provided herein are pharmaceutically acceptable or physiologicallyacceptable compositions including solvents (aqueous or non-aqueous), solutions, emulsions, dispersion media, coatings, isotonic and absorption promoting or delaying agents, compatible with pharmaceutical administration. Pharmaceutical compositions or pharmaceutical formulations therefore refer to a composition suitable for pharmaceutical use in a subject. Compositions can be formulated to be compatible with a particular route of administration (i.e., systemic or local). Thus, compositions include carriers, diluents, or excipients suitable for administration by various routes.

[0169] In some embodiments, a composition can further comprise an acceptable additive in orderto improve the stability of immune cells in the composition. Acceptable additives may not alter the specific activity of the immune cells. Examples of acceptable additives include, but are not limited to, a sugar such as mannitol, sorbitol, glucose, xylitol, trehalose, sorbose, sucrose, galactose, dextran, dextrose, fructose, lactose and mixtures thereof. Acceptable additives can be combined with acceptable carriers and / or excipients such as dextrose. Alternatively, examples of acceptable additives include, but are not limited to, a surfactant such as polysorbate 20 or polysorbate 80 to increase stability of the peptide and decrease gelling of the solution. The surfactant can be added to the composition in an amount of 0.01% to 5% of the solution. Addition of such acceptable additives increases the stability and half-life of the composition in storage.

[0170] The pharmaceutical composition can be administered, for example, by injection.Compositions for injection include aqueous solutions (where water soluble) or dispersions and sterile powders for the extemporaneous preparation of sterile injectable solutions or dispersion. For intravenous administration, suitable carriers include physiological saline, bacteriostatic water, or phosphate buffered saline (PBS). The carrier can be a solvent or dispersion medium containing, for example, water, ethanol, polyol (for example, glycerol, propylene glycol, and liquid polyethylene glycol, and the like), and suitable mixtures thereof. Fluidity can be maintained, for example, by the use of a coating such as lecithin, by the maintenance of the required particle size in the case of dispersion and by the use of surfactants. Antibacterial and antifungal agents include, for example, parabens, chlorobutanol, phenol, ascorbic acid and thimerosal. Isotonic agents, for example, sugars, polyalcohols such as mannitol, sorbitol, and sodium chloride can be included inWSGR Docket No.: 50401-784.601 the composition. The resulting solutions can be packaged for use as is, or lyophilized; the lyophilized preparation can later be combined with a sterile solution prior to administration. For intravenous, injection, or injection at the site of affliction, the active ingredient will be in the form of a parenterally acceptable aqueous solution which is pyrogen-free and has suitable pH, isotonicity and stability. Those of relevant skill in the art are well able to prepare suitable solutions using, for example, isotonic vehicles such as Sodium Chloride Injection, Ringer’s Injection, Lactated Ringer’s Injection. Preservatives, stabilizers, buffers, antioxidants and / or other additives can be included, as needed. Sterile injectable solutions can be prepared by incorporating an active ingredient in the required amount in an appropriate solvent with one or a combination of ingredients enumerated above, as required, followed by filtered sterilization. Generally, dispersions are prepared by incorporating the active ingredient into a sterile vehicle which contains a basic dispersion medium and the required other ingredients from those enumerated above. In the case of sterile powders for the preparation of sterile injectable solutions, the preferred methods of preparation can be vacuum drying and freeze drying which yields a powder of the active ingredient plus any additional desired ingredient from a previously sterile-filtered solution thereof.

[0171] Compositions can be conventionally administered intravenously, such as by injection of aunit dose, for example. For injection, an active ingredient can be in the form of a parenterally acceptable aqueous solution which is substantially pyrogen-free and has suitable pH, isotonicity and stability. One can prepare suitable solutions using, for example, isotonic vehicles such as Sodium Chloride Injection, Ringer’s Injection, Lactated Ringer’s Injection. Preservatives, stabilizers, buffers, antioxidants and / or other additives can be included, as required. Additionally, compositions can be administered via aerosolization.

[0172] When the compositions are considered for use in medicaments or any of the methodsprovided herein, it is contemplated that the composition can be substantially free of pyrogens such that the composition will not cause an inflammatory reaction or an unsafe allergic reaction when administered to a human patient. Testing compositions for pyrogens and preparing compositions substantially free of pyrogens are well understood to one or ordinary skill of the art and can be accomplished using commercially available kits.

[0173] Acceptable carriers can contain a compound that acts as a stabilizing agent, increases ordelays absorption, or increases or delays clearance. Such compounds include, for example, carbohydrates, such as glucose, sucrose, or dextrans; low molecular weight proteins; compositions that reduce the clearance or hydrolysis of peptides; or excipients or other stabilizers and / or buffers. Agents that delay absorption include, for example, aluminum monostearate and gelatin. Detergents can also be used to stabilize or to increase or decrease the absorption of the pharmaceutical composition, including liposomal carriers. To protect from digestion theWSGR Docket No.: 50401-784.601 compound can be complexed with a composition to render it resistant to acidic and enzymatic hydrolysis, or the compound can be complexed in an appropriately resistant carrier such as a liposome. Means of protecting compounds from digestion are known in the art (e.g., Fix (1996) Pharm Res. 13:17601764; Samanen (1996) J. Pharm. Pharmacol. 48:119135; and U.S. Pat. No. 5,391,377).

[0174] The compositions can be administered in a manner compatible with the dosageformulation, and in a therapeutically effective amount. The quantity to be administered depends on the subject to be treated, capacity of the subject’s immune system to utilize the active ingredient, and degree of binding capacity desired. Precise amounts of active ingredient required to be administered depend on the judgment of the practitioner and are peculiar to each individual. Suitable regimes for initial administration and booster shots are also variable, but are typified by an initial administration followed by repeated doses at one or more-hour intervals by a subsequent injection or other administration. Alternatively, continuous intravenous infusions sufficient to maintain concentrations in the blood are contemplated.

[0175] In some embodiments, the present invention is directed to an immunogenic composition,e.g., a pharmaceutical composition capable of raising a neoantigen-specific response (e.g., a humoral or cell-mediated immune response). In some embodiments, the immunogenic composition comprises neoantigen therapeutics (e.g., peptides, polynucleotides, TCR chains, cells containing TCR chains, etc.) described herein corresponding to a tumor specific antigen or neoantigen.

[0176] In some embodiments, a pharmaceutical composition described herein is capable of raisinga specific cytotoxic T cells response, specific helper T cell response, or a B cell response.

[0177] patient. The in vivo method can comprise targeting specific DC receptors using antibodiescoupled with the polypeptides described herein. The DC-based immunogenic pharmaceutical composition can further comprise DC activators such as TLR3, TLR-7-8, and CD40 agonists. The DC-based immunogenic pharmaceutical composition can further comprise adjuvants, and a pharmaceutically acceptable carrier.

[0178] An adjuvant can be used to enhance the immune response (humoral and / or cellular)elicited in a patient receiving the immunogenic pharmaceutical composition. Sometimes, adjuvants can elicit a Th1-type response. Other times, adjuvants can elicit a Th2-type response. A to a Th2-type response which can be characterized by the production of cytokines such as IL-4, IL-5 and IL-10.

[0179] In some aspects, lipid-based adjuvants, such as MPLA and MDP, can be used with theimmunogenic pharmaceutical compositions disclosed herein. Monophosphoryl lipid A (MPLA),WSGR Docket No.: 50401-784.601 for example, is an adjuvant that causes increased presentation of liposomal antigen to specific T Lymphocytes. In addition, a muramyl dipeptide (MDP) can also be used as a suitable adjuvant in conjunction with the immunogenic pharmaceutical formulations described herein.

[0180] Adjuvant can also comprise stimulatory molecules such as cytokines. Non-limitingcutaneous T cell-attracting chemokine (CTACK), epithelial thymus-expressed chemokine (TECK), mucosae-associated epithelial chemokine (MEC), IL-12, IL-15, IL-28, MHC, CD80, CD86, IL-1, IL-2, IL-4, IL-5, IL-6, IL-10, IL-18, MCP-1, MIP-la, MIP-1-, IL-8, L- selectin, P- selectin, E-selectin, CD34, GlyCAM-1, MadCAM-1, LFA-1, VLA-1, Mac-1, pl50.95, PECAM, ICAM-1, ICAM-2, ICAM-3, CD2, LFA-3, M-CSF, G-CSF, mutant forms of IL-18, CD40, CD40L, vascular growth factor, fibroblast growth factor, IL-7, nerve growth factor, vascular endothelial growth factor, Fas, TNF receptor, Fit, Apo-1, p55, WSL-1, DR3, TRAMP, Apo-3, AIR, LARD, NGRF, DR4, DRS, KILLER, TRAIL-R2, TRICK2, DR6, Caspase ICE, Fos, c-jun, R4, RANK, RANK LIGAND, Ox40, Ox40 LIGAND, NKG2D, MICA, MICB, NKG2A, NKG2B, NKG2C, NKG2E, NKG2F, TAPI, and TAP2.

[0181] Additional adjuvants include: MCP-1, MIP-la, MIP-lp, IL-8, RANTES, L-selectin, P-selectin, E-selectin, CD34, GlyCAM-1, MadCAM-1, LFA-1, VLA-1, Mac-1, pl50.95, PECAM, ICAM-1, ICAM-2, ICAM-3, CD2, LFA-3, M-CSF, G-CSF, IL-4, mutant forms of IL-18, CD40, CD40L, vascular growth factor, fibroblast growth factor, IL-7, IL-22, nerve growth factor, vascular endothelial growth factor, Fas, TNF receptor, Fit, Apo-1, p55, WSL-1, DR3, TRAMP, Apo-3, AIR, LARD, NGRF, DR4, DR5, KILLER, TRAIL-R2, TRICK2, DR6, Caspase ICE, Fos, TRAIL-R4, RANK, RANK LIGAND, Ox40, Ox40 LIGAND, NKG2D, MICA, MICB, NKG2A, NKG2B, NKG2C, NKG2E, NKG2F, TAP1, TAP2 and functional fragments thereof.

[0182] In some aspects, an adjuvant can be a modulator of a toll like receptor. Examples ofmodulators of toll-like receptors include TLR9 agonists and are not limited to small molecule modulators of toll-like receptors such as Imiquimod. Sometimes, an adjuvant is selected from bacteria toxoids, polyoxypropylene-polyoxyethylene block polymers, aluminum salts, liposomes, CpG polymers, oil-in-water emulsions, or a combination thereof. Sometimes, an adjuvant is an oil-in-water emulsion. The oil-in-water emulsion can include at least one oil and at least one surfactant, with the oil(s) and surfactant(s) being biodegradable (metabolizable) andWSGR Docket No.: 50401-784.601 have a sub-micron diameter, with these small sizes being achieved with a microfluidizer to provide stable emulsions. Droplets with a size less than 220 nm can be subjected to filter sterilization.

[0183] In some instances, an immunogenic pharmaceutical composition can include carriers andexcipients (including but not limited to buffers, carbohydrates, mannitol, proteins, polypeptides or amino acids such as glycine, antioxidants, bacteriostats, chelating agents, suspending agents, thickening agents and / or preservatives), water, oils including those of petroleum, animal, vegetable or synthetic origin, such as peanut oil, soybean oil, mineral oil, sesame oil and the like, saline solutions, aqueous dextrose and glycerol solutions, flavoring agents, coloring agents, detackifiers and other acceptable additives, adjuvants, or binders, other pharmaceutically acceptable auxiliary substances as required to approximate physiological conditions, such as pH buffering agents, tonicity adjusting agents, emulsifying agents, wetting agents and the like. Examples of excipients include starch, glucose, lactose, sucrose, gelatin, malt, rice, flour, chalk, silica gel, sodium stearate, glycerol monostearate, talc, sodium chloride, dried skim milk, glycerol, propylene, glycol, water, ethanol and the like. In another instances, the pharmaceutical preparation is substantially free of preservatives. In other instances, the pharmaceutical preparation can contain at least one preservative. It will be recognized that, while any suitable carrier known to those of ordinary skill in the art can be employed to administer the pharmaceutical compositions described herein, the type of carrier will vary depending on the mode of administration.

[0184] An immunogenic pharmaceutical composition can include preservatives such asthiomersal or 2-phenoxyethanol. In some instances, the immunogenic pharmaceutical

[0185] For controlling the tonicity, a physiological salt such as sodium salt can be included in theimmunogenic pharmaceutical composition. Other salts can include potassium chloride, potassium dihydrogen phosphate, disodium phosphate, and / or magnesium chloride, or the like.

[0186] An immunogenic pharmaceutical composition can have an osmolality of between 200mOsm / kg and 400 mOsm / kg, between 240-360 mOsm / kg, or within the range of 290-310 mOsm / kg.

[0187] An immunogenic pharmaceutical composition can comprise one or more buffers, such asa Tris buffer; a borate buffer; a succinate buffer; a histidine buffer (particularly with an aluminum hydroxide adjuvant); or a citrate buffer. Buffers, in some cases, are included in the 5-20 or 10-50 mM range.WSGR Docket No.: 50401-784.601

[0188] The pH of the immunogenic pharmaceutical composition can be between about 5.0 andabout 8.5, between about 6.0 and about 8.0, between about 6.5 and about 7.5, or between about 7.0 and about 7.8.

[0189] An immunogenic pharmaceutical composition can be sterile. The immunogenicpharmaceutical composition can be non-pyrogenic e.g. containing <1 EU (endotoxin unit, a standard measure) per dose, and can be <0.1 EU per dose. The composition can be gluten free.

[0190] An immunogenic pharmaceutical composition can include detergent e.g. apolyoxyethylene sorbitan ester surfactant (known as ‘Tweens’), or an octoxynol (such as octoxynol-9 (Triton X-100) or t-octylphenoxypolyethoxyethanol). The detergent can be present only at trace amounts. The immunogenic pharmaceutical composition can include less than 1 mg / mL of each of octoxynol-10 and polysorbate 80. Other residual components in trace amounts can be antibiotics (e.g. neomycin, kanamycin, polymyxin B).

[0191] An immunogenic pharmaceutical composition can be formulated as a sterile solution orsuspension, in suitable vehicles, well known in the art. The pharmaceutical compositions can be sterilized by conventional, well-known sterilization techniques, or can be sterile filtered. The resulting aqueous solutions can be packaged for use as is, or lyophilized, the lyophilized preparation being combined with a sterile solution prior to administration.

[0192] Pharmaceutical compositions comprising, for example, an active agent such as immunecells disclosed herein, in combination with one or more adjuvants can be formulated to comprise certain molar ratios. For example, molar ratios of about 99:1 to about 1:99 of an active agent such as an immune cell described herein, in combination with one or more adjuvants can be used. In some instances, the range of molar ratios of an active agent such as an immune cell described herein, in combination with one or more adjuvants can be selected from about 80:20 to about 20:80; about 75:25 to about 25:75, about 70:30 to about 30:70, about 66:33 to about 33:66, about 60:40 to about 40:60; about 50:50; and about 90:10 to about 10:90. The molar ratio of an active agent such as an immune cell described herein, in combination with one or more adjuvants can be about 1:9, and in some cases can be about 1:1. The active agent such as an immune cell described herein, in combination with one or more adjuvants can be formulated together, in the same dosage unit e.g., in one vial, suppository, tablet, capsule, an aerosol spray; or each agent, form, and / or compound can be formulated in separate units, e.g., two vials, suppositories, tablets, two capsules, a tablet and a vial, an aerosol spray, and the like.

[0193] In some instances, an immunogenic pharmaceutical composition can be administered withan additional agent. The choice of the additional agent can depend, at least in part, on the condition being treated. The additional agent can include, for example, a checkpoint inhibitor agent such as an anti-PD1, anti-CTLA4, anti-PD-L1, anti CD40, or anti-TIM3 agent (e.g., an anti-PD1, anti-WSGR Docket No.: 50401-784.601 CTLA4, anti-PD-L1, anti CD40, or anti-TIM3 antibody); or any agents having a therapeutic effect for a pathogen infection (e.g. viral infection), including, e.g., drugs used to treat inflammatory conditions such as an NSAID, e.g., ibuprofen, naproxen, acetaminophen, ketoprofen, or aspirin. For example, the checkpoint inhibitor can be a PD-1 / PD- L1 antagonist selected from the group consisting of: nivolumab (ONO-4538 / BMS-936558, MDX1 106, OPDIVO), pembrolizumab (MK-3475, KEYTRUDA), pidilizumab (CT-011), and MPDL328OA (ROCHE). As another example, formulations can additionally contain one or more supplements, such as vitamin C, E or other anti-oxidants.

[0194] A pharmaceutical composition comprising an active agent such as an immune celldescribed herein, in combination with one or more adjuvants can be formulated in conventional manner using one or more physiologically acceptable carriers, comprising excipients, diluents, and / or auxiliaries, e.g., which facilitate processing of the active agents into preparations that can be administered. Proper formulation can depend at least in part upon the route of administration chosen. The agent(s) described herein can be delivered to a patient using a number of routes or modes of administration, including oral, buccal, topical, rectal, transdermal, transmucosal, subcutaneous, intravenous, and intramuscular applications, as well as by inhalation.

[0195] The active agents can be formulated for parenteral administration (e.g., by injection, forexample bolus injection or continuous infusion) and can be presented in unit dose form in ampoules, pre-filled syringes, small volume infusion or in multi-dose containers with an added preservative. The compositions can take such forms as suspensions, solutions, or emulsions in oily or aqueous vehicles, for example solutions in aqueous polyethylene glycol.

[0196] In some embodiments, the pharmaceutical composition comprises a preservative orstabilizer. In some embodiments the preservative or stabilizer is selected from a cytokine, a growth factor or an adjuvant or a chemical substance. In some embodiments, the composition comprises at least one agent that helps preserve cell viability through at least one cycle of freeze-thaw. In some embodiments, the composition comprises at least one agent that helps preserve cell viability through at least more than one cycle of freeze-thaw.

[0197] For injectable formulations, the vehicle can be chosen from those known in art to besuitable, including aqueous solutions or oil suspensions, or emulsions, with sesame oil, corn oil, cottonseed oil, or peanut oil, as well as elixirs, mannitol, dextrose, or a sterile aqueous solution, and similar pharmaceutical vehicles. The formulation can also comprise polymer compositions which are biocompatible, biodegradable, such as poly(lactic-co-glycolic) acid. These materials can be made into micro or nanospheres, loaded with drug and further coated or derivatized to provide superior sustained release performance. Vehicles suitable for periocular or intraocular injection include, for example, suspensions of therapeutic agent in injection grade water,WSGR Docket No.: 50401-784.601 liposomes and vehicles suitable for lipophilic substances. Other vehicles for periocular or intraocular injection are well known in the art.

[0198] In some instances, pharmaceutical composition is formulated in accordance with routineprocedures as a pharmaceutical composition adapted for intravenous administration to human beings. Typically, compositions for intravenous administration are solutions in sterile isotonic aqueous buffer. Where necessary, the composition can also include a solubilizing agent and a local anesthetic such as lidocaine to ease pain at the site of the injection. Generally, the ingredients are supplied either separately or mixed together in unit dosage form, for example, as a dry lyophilized powder or water free concentrate in a hermetically sealed container such as an ampoule or sachette indicating the quantity of active agent. Where the composition is to be administered by infusion, it can be dispensed with an infusion bottle containing sterile pharmaceutical grade water or saline. Where the composition is administered by injection, an ampoule of sterile water for injection or saline can be provided so that the ingredients can be mixed prior to administration. Methods of Treatment

[0199] Provided herein is a method for treating a subject in need thereof comprising administeringthe pharmaceutical composition described herein into the subject. The pharmaceutical composition can be a therapeutic composition comprising a TCR comprising a TCR chain selected or identified using the methods described herein, a nucleic acid molecule encoding the TCR comprising the TCR chain selected or identified, an immune cell (e.g., T cells) comprising the TCR comprising the TCR chain or the nucleic acid molecule encoding the TCR chain selected or identified. In some embodiments, the therapeutic composition comprising T cells is administered by injection. In some embodiments, the therapeutic composition comprising T cells is administered by infusion. When administration is by injection, the active agent can be formulated in aqueous solutions, specifically in physiologically compatible buffers such as Hank’s solution, Ringer’s solution, or physiological saline buffer. The solution can contain formulator agents such as suspending, stabilizing and / or dispersing agents. In another embodiment, the pharmaceutical composition does not comprise an adjuvant or any other substance added to enhance the immune response stimulated by the peptide. In some embodiments, the method further comprises administering one or more of the at least one antigen specific T cell as a pharmaceutical composition described herein to a subject. In some embodiments, the pharmaceutical composition comprises a preservative or stabilizer. In some embodiments the preservative or stabilizer is selected from a cytokine, a growth factor or an adjuvant or a chemical substance. In some embodiments, the at least one antigen specific T cell is administered to a subject within 28 days from collecting a PBMC sample from the subject.WSGR Docket No.: 50401-784.601

[0200] In addition to the formulations described previously, the active agents can also beformulated as a depot preparation. Such long-acting formulations can be administered by implantation or transcutaneous delivery (for example subcutaneously or intramuscularly), intramuscular injection or use of a transdermal patch. Thus, for example, the agents can be formulated with suitable polymeric or hydrophobic materials (for example as an emulsion in an acceptable oil) or ion exchange resins, or as sparingly soluble derivatives, for example, as a sparingly soluble salt.

[0201] Also provided herein are methods of treating a subject with a disease, disorder orcondition. A method of treatment can comprise administering a composition or pharmaceutical composition disclosed herein to a subject with a disease, disorder or condition.

[0202] The present disclosure provides methods of treatment comprising an immunogenictherapy. Methods of treatment for a disease (such as cancer or a viral infection) are provided. A method can comprise administering to a subject an effective amount of a composition comprising an immunogenic antigen specific T cell according to the methods provided herein. In some embodiments, the antigen comprises a viral antigen. In some embodiments, the antigen comprises a tumor antigen.

[0203] Non-limiting examples of therapeutics that can be prepared include a peptide-basedtherapy, a nucleic acid-based therapy, an antibody-based therapy, a T cell-based therapy, and an antigen-presenting cell-based therapy.

[0204] In some other aspects, provided here is use of a composition or pharmaceuticalcomposition for the manufacture of a medicament for use in therapy. In some embodiments, a method of treatment comprises administering to a subject an effective amount of T cells specifically recognizing an immunogenic neoantigen peptide. In some embodiments, a method of treatment comprises administering to a subject an effective amount of a TCR that specifically recognizes an immunogenic neoantigen peptide, such as a TCR expressed in a T cell.

[0205] In some embodiments, the cancer is selected from the group consisting of carcinoma,lymphoma, blastoma, sarcoma, leukemia, squamous cell cancer, lung cancer (including small cell lung cancer, non-small cell lung cancer (NSCLC), adenocarcinoma of the lung, and squamous carcinoma of the lung), cancer of the peritoneum, hepatocellular cancer, gastric or stomach cancer (including gastrointestinal cancer), pancreatic cancer, glioblastoma, cervical cancer, ovarian cancer, liver cancer, bladder cancer, hepatoma, breast cancer, colon cancer, melanoma, endometrial or uterine carcinoma, salivary gland carcinoma, kidney or renal cancer, liver cancer, prostate cancer, vulval cancer, thyroid cancer, hepatic carcinoma, head and neck cancer, colorectal cancer, rectal cancer, soft-tissue sarcoma, Kaposi’s sarcoma, B-cell lymphoma (including low grade / follicular non-Hodgkin’s lymphoma (NHL), small lymphocytic (SL) NHL, intermediateWSGR Docket No.: 50401-784.601 grade / follicular NHL, intermediate grade diffuse NHL, high grade immunoblastic NHL, high grade lymphoblastic NHL, high grade small non-cleaved cell NHL, bulky disease NHL, mantle cell lymphoma, AIDS-related lymphoma, and Waldenstrom’s macroglobulinemia), chronic lymphocytic leukemia (CLL), acute lymphoblastic leukemia (ALL), myeloma, Hairy cell leukemia, chronic myeloblasts leukemia, and post-transplant lymphoproliferative disorder (PTLD), abnormal vascular proliferation associated with phakomatoses, edema, Meigs’ syndrome, and combinations thereof.

[0206] The methods described herein are particularly useful in the personalized medicine context,where immunogenic neoantigen peptides identified according to the methods described herein are used to develop therapeutics (such as vaccines or therapeutic antibodies) for the same individual. Thus, a method of treating a disease in a subject can comprise identifying an immunogenic neoantigen peptide in a subject according to the methods described herein; and synthesizing the peptide (or a precursor thereof, such as a polynucleotide (e.g., an mRNA) encoding the peptide); and manufacturing T cells specific for identified neoantigens; and administering the neoantigen specific T cells to the subject. In some embodiments, the method of treating a disease in a subject can comprise identifying an immunogenic neoantigen peptide in a subject according to the methods described herein; and synthesizing the polynucleotide, such as an mRNA, that encodes the immunogenic neoantigen peptide or a precursor thereof, and manufacturing T cells specific for identified neoantigens; and administering the neoantigen specific T cells to the subject.

[0207] The agents and compositions provided herein may be used alone or in combination withconventional therapeutic regimens such as surgery, irradiation, chemotherapy and / or bone marrow transplantation (autologous, syngeneic, allogeneic or unrelated). A set of tumor antigens can be identified using the methods described herein and are useful, e.g., in a large fraction of cancer patients.

[0208] In some embodiments, at least one or more chemotherapeutic agents may be administeredin addition to the composition comprising an immunogenic therapy. In some embodiments, the one or more chemotherapeutic agents may belong to different classes of chemotherapeutic agents.

[0209] In practicing the methods of treatment or use provided herein, therapeutically-effectiveamounts of the therapeutic agents can be administered to a subject having a disease or condition. A therapeutically-effective amount can vary widely depending on the severity of the disease, the age and relative health of the subject, the potency of the compounds used, and other factors.

[0210] Subjects can be, for example, mammal, humans, pregnant women, elderly adults, adults,adolescents, pre-adolescents, children, toddlers, infants, newborn, or neonates. A subject can be a patient. In some cases, a subject can be a human. In some cases, a subject can be a child (i.e. a young human being below the age of puberty). In some cases, a subject can be an infant. In someWSGR Docket No.: 50401-784.601 cases, the subject can be a formula-fed infant. In some cases, a subject can be an individual enrolled in a clinical study. In some cases, a subject can be a laboratory animal, for example, a mammal, or a rodent. In some cases, the subject can be a mouse. In some cases, the subject can be an obese or overweight subject.

[0211] In some embodiments, the subject has previously been treated with one or more differentcancer treatment modalities. In some embodiments, the subject has previously been treated with one or more of radiotherapy, chemotherapy, or immunotherapy. In some embodiments, the subject has been treated with one, two, three, four, or five lines of prior therapy. In some embodiments, the prior therapy is a cytotoxic therapy.

[0212] In some embodiments, the disease or condition that can be treated with the methodsdisclosed herein is cancer. Cancer is an abnormal growth of cells which tend to proliferate in an uncontrolled way and, in some cases, to metastasize (spread). A tumor can be cancerous or benign. A benign tumor means the tumor can grow but does not spread. A cancerous tumor is malignant, meaning it can grow and spread to other parts of the body. If a cancer spreads (metastasizes), the new tumor bears the same name as the original (primary) tumor.

[0213] Cancer refers to diseases in which abnormal cells divide out of control and are able toinvade other tissues. Cancer cells can spread to other parts of the body through the blood and lymph systems. Cancer can be characterized as a group of diseases involving abnormal cell growth that may begin in any tissue with the potential to invade or spread to other parts of the body. Some cancers can be characterized by their type, e.g., solid cancers, liquid cancers, or based on cellular origin such as hematopoietic cancers, osteosarcoma or lymphoma. Some cancers are known by the tissue of their origin or prevalence, e.g., endometrial cancers are characterized as cancers of the endometrial tissue. Some cancers are known by the organ or site of their origin or prevalence, e.g. lung cancer, head and neck cancer. Some cancers may be known by the overproductions of certain proteins, enzymes or biomarkers compared to their counterpart cells or tissues that are not cancerous. For example, certain proteins of viral origin may be associated with certain cancers, such as HPV-16 cancers, where certain proteins, for example HPV-16 E6 and E7 are overexpressed in cancer cells of this type. For example, certain antigens, such as KRAS may be highly expressed in certain cancer types, compared to non- cancer cells of the same type, and may be designated as KRAS overexpressing cancers. Typically, the overexpression of the antigen or the specific protein may be associated with or related to one or more mutations, and the cancer type may be associated with the mutation. For example mutation at the wild type G residue corresponding to position 12 in KRAS amino acid sequence may be mutated to V, D, C or other amino acids in KRAS-specific cancer cells. Certain specific antigens may be specifically expressed in cancer cells of certain cancer types,WSGR Docket No.: 50401-784.601 and not in other cancer types. Various cancer are contemplated herein that may not be restricted to a specific cell type, tissue type or organ, or even a certain stage of cancer. The TCRs of the present invention are directed to cancer cells that express a cancer antigen, that may be patient specific, which can be found during sequencing of a subject’s genome from biological sample obtained from a cancer cell, cancer site or cancer tissue and compared to a corresponding non- cancer sample from the same subject; wherein the patient-specific antigen may be expressed in the cancer cell, and not on the non-cancer cell of the subject. In some cases, cancer antigens may be cancer specific, where the antigen is reportedly present in the type of cancer observed in multiple patients in the human population, who have been diagnosed of the specific cancer. In some cases, certain types are cancers are associated with an antigen, a protein (e.g., a viral protein) a gene mutation. All forms of cancer are contemplated herein.

[0214] In some embodiments, the cancer is a solid cancer. In some cases, the cancer is a liquid / blood cancer. The cancer can express or be diagnosed as expressing a tumor antigen. The tumor antigen can be a tumor-associated antigen or a tumor-specific antigen. In some cases, the cancer expresses a tumor-associated antigen (TAA). In some cases, the cancer expresses a tumor-specific antigen (TSA).

[0215] In some embodiments, the cancer is a cancer expressing or diagnosed as expressing aTAA. In some embodiments, the cancer is a cancer expressing or diagnosed as expressing a TSA.

[0216] The current classification of TAA can include the following group:a) Cancer testis (CT) antigen: Since testis cells do not express HLA class I and class II molecules, these antigens may not be recognized by T cells in normal tissues and may therefore be immunologically considered tumor specific. Non-limiting examples of CT antigens include members of the MAGE family and NY-ESO-1; b) Differentiation antigen: both tumor and normal tissue (from which the tumor originates) may contain TAAs. Differentiation antigens may be found, for example, in melanoma and normal melanocytes. Many of these melanocyte lineage-associated proteins may be involved in melanin biosynthesis and therefore these proteins may not tumor-specific, but may still widely be used for immunotherapy of cancer. Examples include, but are not limited to, tyrosinase for melanoma and PSA for Melan-A / MART-1 or prostate cancer; c) Overexpressed TAA: gene-encoded widely expressed TAAs may be detected in histologically diverse tumors and in many normal tissues, with generally low expression levels. It is possible that many epitopes processed and potentially presented by normal tissues may be below the threshold level of T cell recognition, whereas their overexpression in tumor cells can triggerWSGR Docket No.: 50401-784.601 anticancer responses by breaking previously established tolerance. Non-limiting examples of such TAAs include Her-2 / neu, survivin, telomerase or WT1; d) tumor specific antigen can include unique TAAs resulted from mutations in normal genes (e.g., beta-catenin, CDK4). Some of these molecular changes can be associated with neoplastic transformation and / or progression. Tumor-specific antigens can generally induce strong immune responses without risking from the autoimmune response to normal tissue strips. On the other hand, these TAAs may only be associated with the exact tumor on which they are confirmed, and may not commonly shared among many individual tumors. In the case of tumor specific (related) isoform proteins, peptide tumor specificity (or relatedness) may also occur if the peptide is derived from tumor (related) exons; e) TAA resulting from aberrant post-translational modification: such TAAs may result from proteins in the tumor that are neither specific nor overexpressed, but which still have tumor relevance (this relevance is due to posttranslational processing that is primarily active on tumors). Such TAAs may result from an altered glycosylation pattern, resulting in a tumor producing a novel epitope for MUC1 or in an event such as protein splicing during degradation, which may or may not be tumor specific; and f) Tumor virus protein: these TTAs are viral proteins that may play a key role in the oncogenic process and, because they are foreign proteins (non-human proteins), may be able to trigger T cell responses. Non-limiting examples of such proteins include human papilloma type 16 viral proteins, E6 and E7, which are expressed in cervical cancer.

[0217] Examples of tumor antigens include, but not limited to new antigens expressed duringtumorigenesis, products of oncogenes and tumor suppressor genes, overexpressed or abnormally expressed intracellular proteins (e.g., HER2, MUC1, PSA, MUC1), carcinoembryonic antigen (CEA), tumor viruses (e.g., EBC, HPV, HBV, HCB, HTLV), cancer testis antigens (CTA) (e.g., MAGE family, NY-ESO), oncofetal antigens, altered surface glycolipids and glycoproteins, cell type-specific differentiation antigens (e.g., MART-1), or a derivative thereof. The tumor antigens can be selected from the group consisting of NY-ESO-1, Her2 / neu, SSX-2, MAGE-C2, MAGE-A1, M-2433-233, MAGE-A10254-262, KK-LC-1, p53, PRAME, Alpha fetoprotein, HPV6-E6, HPV16-E7, EBV-LMP1, RAS: G12D, RAS: G12C, RAS: G12A, RAS: G12S, RAS: G12R, RAS: G12R, RAS: G12R, RAS: G122 V, RAS: Q61H, RAS: Q61L, RAS: Q61R, RAS: G13D, TP53: V157G, TP53: V157F, TP53: R248Q, TP53: R248W, TP53: G245S, TP53: Y163C, TP53: G249S, TP53: Y240C, TP53: R175H, TP53: K132N, CDC73: Q254E, TPP2A6: N438Y, CTNN1: T41A, CTNNB1: S45P, CTNNB1: S37Y, CTNNB1: S33C, EGFR: L858R, EGFR: T790M, PIK3CA: E542K, PIK3CA: H1047R, GNAS: R201H, CDK4:R24, R24C H3.WSGR Docket No.: 50401-784.601 3:K28M, BRAF: V600E, CHD4 K73Rfs, NRAS Q61R, IDH1:R132H, TVP23C: C51Y, and any combination thereof. The RAS can be KRAS, HRAS, or NRAS.

[0218] Other non-limiting examples of tumor-associated antigen or tumor-specific antigenincludes antigens from Human Papilloma Virus, Epstein-Barr Virus, Merkel cell polyomavirus, Human Immunodeficiency Virus, Human T-cell Leukemia Virus, Human Herpes Virus 8, Hepatitis B virus, Hepatitis C virus, HCV, HBC, Cytomegalovirus, or from the group of single- point mutated antigens derived from the group consisting of the antigens of ctnnbl gene, casp8 gene, HER2 gene, p53 gene, KRAS gene, NRAS gene, or particular tumor antigens issued or derived from the group consisting of RAS oncogene, BCR-ABL tumor antigens, ETV6-AML1 tumor antigens, melanoma-antigen encoding genes (MAGE), BAGE antigens, GAGE antigens, ssx antigens, ny-eso-1 antigens, cyclin-A1 tumor antigens, MART-1 antigen, gp100 antigen, CD19 antigen, prostate specific antigen, prostatic acidic phosphatase antigen, carcinoembryonic antigen, alpha fetoprotein antigen, carcinoma antigen 125, mucin 16 antigen, mucin 1 antigen, human telomerase reverse transcriptase antigen, EGFR antigen, MOK antigen, RAGE-1 antigen, PRAME antigen, wild-type p53 antigen, oncogene ERBB2 antigen, sialyl-Tn tumor antigen, Wilms tumor 1 antigen, mesothelin antigen, carbohydrate antigens, B-catenin antigen, MUM-1 antigen, CDK4 antigen ERBB2IP antigen, and Melan-A melanoma tumor-associated antigen.

[0219] In some cases, the cancer cells express the tumor antigens, including and not limited to,NY-ESO-1, Her2 / neu, SSX-2, MAGE-C2, MAGE-A1, M-2433-233, MAGE-A10254-262, KK- LC-1, p53, PRAME, Alpha fetoprotein, HPV6-E6, HPV16-E7, EBV-LMP1, RAS: G12D, RAS: G12C, RAS: G12A, RAS: G12S, RAS: G12R, RAS: G12R, RAS: G12R, RAS: G122 V, RAS: Q61H, RAS: Q61L, RAS: Q61R, RAS: G13D, TP53: V157G, TP53: V157F, TP53: R248Q, TP53: R248W, TP53: G245S, TP53: Y163C, TP53: G249S, TP53: Y240C, TP53: R175H, TP53: K132N, CDC73: Q254E, TPP2A6: N438Y, CTNN1: T41A, CTNNB1: S45P, CTNNB1: S37Y, CTNNB1: S33C, EGFR: L858R, EGFR: T790M, PIK3CA: E542K, PIK3CA: H1047R, GNAS: R201H, CDK4:R24, R24C H3. 3:K28M, BRAF: V600E, CHD4 K73Rfs, NRAS Q61R, IDH1:R132H, or TVP23C: C51Y. The RAS can be KRAS, HRAS, or NRAS.

[0220] The methods of the disclosure can be used to treat any type of cancer known in the art.Non-limiting examples of cancers to be treated by the methods of the present disclosure can include melanoma (e.g., metastatic malignant melanoma), renal cancer (e.g., clear cell carcinoma), prostate cancer (e.g., hormone refractory prostate adenocarcinoma), pancreatic adenocarcinoma, breast cancer, colon cancer, lung cancer (e.g., non-small cell lung cancer), esophageal cancer, squamous cell carcinoma of the head and neck, liver cancer, ovarian cancer, cervical cancer, thyroid cancer, glioblastoma, glioma, leukemia, lymphoma, and other neoplastic malignancies.WSGR Docket No.: 50401-784.601

[0221] Additionally, the disease or condition provided herein includes refractory or recurrentmalignancies whose growth may be inhibited using the methods of treatment of the present disclosure. In some embodiments, a cancer to be treated by the methods of treatment of the present disclosure is selected from the group consisting of carcinoma, squamous carcinoma, adenocarcinoma, sarcomata, endometrial cancer, breast cancer, ovarian cancer, cervical cancer, fallopian tube cancer, primary peritoneal cancer, colon cancer, colorectal cancer, squamous cell carcinoma of the anogenital region, melanoma, renal cell carcinoma, lung cancer, non-small cell lung cancer, squamous cell carcinoma of the lung, stomach cancer, bladder cancer, gall bladder cancer, liver cancer, thyroid cancer, laryngeal cancer, salivary gland cancer, esophageal cancer, head and neck cancer, glioblastoma, glioma, squamous cell carcinoma of the head and neck, prostate cancer, pancreatic cancer, mesothelioma, sarcoma, hematological cancer, leukemia, lymphoma, neuroma, and combinations thereof. In some embodiments, a cancer to be treated by the methods of the present disclosure include, for example, carcinoma, squamous carcinoma (for example, cervical canal, eyelid, tunica conjunctiva, vagina, lung, oral cavity, skin, urinary bladder, tongue, larynx, and gullet), and adenocarcinoma (for example, prostate, small intestine, endometrium, cervical canal, large intestine, lung, pancreas, gullet, rectum, uterus, stomach, mammary gland, and ovary). In some embodiments, a cancer to be treated by the methods of the present disclosure further include sarcomata (for example, myogenic sarcoma), leukosis, neuroma, melanoma, and lymphoma. In some embodiments, a cancer to be treated by the methods of the present disclosure is breast cancer. In some embodiments, a cancer to be treated by the methods of treatment of the present disclosure is triple negative breast cancer (TNBC). In some embodiments, a cancer to be treated by the methods of treatment of the present disclosure is ovarian cancer. In some embodiments, a cancer to be treated by the methods of treatment of the present disclosure is colorectal cancer.

[0222] In some embodiments, a patient or population of patients to be treated with apharmaceutical composition of the present disclosure have a solid tumor. In some embodiments, a solid tumor is a melanoma, renal cell carcinoma, lung cancer, bladder cancer, breast cancer, cervical cancer, colon cancer, gall bladder cancer, laryngeal cancer, liver cancer, thyroid cancer, stomach cancer, salivary gland cancer, prostate cancer, pancreatic cancer, or Merkel cell carcinoma. In some embodiments, a patient or population of patients to be treated with a pharmaceutical composition of the present disclosure have a hematological cancer. In some embodiments, the patient has a hematological cancer such as Diffuse large B cell lymphoma (“DLBCL”), Hodgkin’s lymphoma (“HL”), Non-Hodgkin’s lymphoma (“NHL”), Follicular lymphoma (“FL”), acute myeloid leukemia (“AML”), or Multiple myeloma (“MM”). In someWSGR Docket No.: 50401-784.601 embodiments, a patient or population of patients to be treated having the cancer selected from the group consisting of ovarian cancer, lung cancer and melanoma.

[0223] Specific examples of cancers that can be prevented and / or treated in accordance withpresent disclosure include, but are not limited to, the following: renal cancer, kidney cancer, glioblastoma multiforme, metastatic breast cancer; breast carcinoma; breast sarcoma; neurofibroma; neurofibromatosis; pediatric tumors; neuroblastoma; malignant melanoma; carcinomas of the epidermis; leukemias such as but not limited to, acute leukemia, acute lymphocytic leukemia, acute myelocytic leukemias such as myeloblastic, promyelocytic, myelomonocytic, monocytic, erythroleukemia leukemias and myelodysplastic syndrome, chronic leukemias such as but not limited to, chronic myelocytic (granulocytic) leukemia, chronic lymphocytic leukemia, hairy cell leukemia; polycythemia vera; lymphomas such as but not limited to Hodgkin’s disease, non-Hodgkin’s disease; multiple myelomas such as but not limited to smoldering multiple myeloma, nonsecretory myeloma, osteosclerotic myeloma, plasma cell leukemia, solitary plasmacytoma and extramedullary plasmacytoma; Waldenstrom’s macroglobulinemia; monoclonal gammopathy of undetermined significance; benign monoclonal gammopathy; heavy chain disease; bone cancer and connective tissue sarcomas such as but not limited to bone sarcoma, myeloma bone disease, multiple myeloma, cholesteatoma-induced bone osteosarcoma, Paget’s disease of bone, osteosarcoma, chondrosarcoma, Ewing’s sarcoma, malignant giant cell tumor, fibrosarcoma of bone, chordoma, periosteal sarcoma, soft-tissue sarcomas, angiosarcoma (hemangiosarcoma), fibrosarcoma, Kaposi’s sarcoma, leiomyosarcoma, liposarcoma, lymphangiosarcoma, neurilemmoma, rhabdomyosarcoma, and synovial sarcoma; brain tumors such as but not limited to, glioma, astrocytoma, brain stem glioma, ependymoma, oligodendroglioma, nonglial tumor, acoustic neurinoma, craniopharyngioma, medulloblastoma, meningioma, pineocytoma, pineoblastoma, and primary brain lymphoma; breast cancer including but not limited to adenocarcinoma, lobular (small cell) carcinoma, intraductal carcinoma, medullary breast cancer, mucinous breast cancer, tubular breast cancer, papillary breast cancer, Paget’s disease (including juvenile Paget’s disease) and inflammatory breast cancer; adrenal cancer such as but not limited to pheochromocytoma and adrenocortical carcinoma; thyroid cancer such as but not limited to papillary or follicular thyroid cancer, medullary thyroid cancer and anaplastic thyroid cancer; pancreatic cancer such as but not limited to, insulinoma, gastrinoma, glucagonoma, vipoma, somatostatin-secreting tumor, and carcinoid or islet cell tumor; pituitary cancers such as but not limited to Cushing’s disease, prolactin-secreting tumor, acromegaly, and diabetes insipius; eye cancers such as but not limited to ocular melanoma such as iris melanoma, choroidal melanoma, and ciliary body melanoma, and retinoblastoma; vaginal cancers such as squamous cell carcinoma, adenocarcinoma, and melanoma; vulvar cancer such as squamous cellWSGR Docket No.: 50401-784.601 carcinoma, melanoma, adenocarcinoma, basal cell carcinoma, sarcoma, and Paget’s disease; cervical cancers such as but not limited to, squamous cell carcinoma, and adenocarcinoma; uterine cancers such as but not limited to endometrial carcinoma and uterine sarcoma; ovarian cancers such as but not limited to, ovarian epithelial carcinoma, borderline tumor, germ cell tumor, and stromal tumor; cervical carcinoma; esophageal cancers such as but not limited to, squamous cancer, adenocarcinoma, adenoid cyctic carcinoma, mucoepidermoid carcinoma, adenosquamous carcinoma, sarcoma, melanoma, plasmacytoma, verrucous carcinoma, and oat cell (small cell) carcinoma; stomach cancers such as but not limited to, adenocarcinoma, fungating (polypoid), ulcerating, superficial spreading, diffusely spreading, malignant lymphoma, liposarcoma, fibrosarcoma, and carcinosarcoma; colon cancers; colorectal cancer, KRAS mutated colorectal cancer; colon carcinoma; rectal cancers; liver cancers such as but not limited to hepatocellular carcinoma and hepatoblastoma, gallbladder cancers such as adenocarcinoma; cholangiocarcinomas such as but not limited to papillary, nodular, and diffuse; lung cancers such as KRAS-mutated non-small cell lung cancer, non-small cell lung cancer, squamous cell carcinoma (epidermoid carcinoma), adenocarcinoma, large-cell carcinoma and small-cell lung cancer; lung carcinoma; testicular cancers such as but not limited to germinal tumor, seminoma, anaplastic, classic (typical), spermatocytic, nonseminoma, embryonal carcinoma, teratoma carcinoma, choriocarcinoma (yolk-sac tumor), prostate cancers such as but not limited to, androgen-independent prostate cancer, androgen-dependent prostate cancer, adenocarcinoma, leiomyosarcoma, and rhabdomyosarcoma; penal cancers; oral cancers such as but not limited to squamous cell carcinoma; basal cancers; salivary gland cancers such as but not limited to adenocarcinoma, mucoepidermoid carcinoma, and adenoid cystic carcinoma; pharynx cancers such as but not limited to squamous cell cancer, and verrucous; skin cancers such as but not limited to, basal cell carcinoma, squamous cell carcinoma and melanoma, superficial spreading melanoma, nodular melanoma, lentigo malignant melanoma, acral lentiginous melanoma; kidney cancers such as but not limited to renal cell cancer, adenocarcinoma, hypernephroma, fibrosarcoma, transitional cell cancer (renal pelvis and / or ureter); renal carcinoma; Wilms’ tumor; bladder cancers such as but not limited to transitional cell carcinoma, squamous cell cancer, adenocarcinoma, carcinosarcoma. In addition, cancers include myosarcoma, osteogenic sarcoma, endotheliosarcoma, lymphangioendotheliosarcoma, mesothelioma, synovioma, hemangioblastoma, epithelial carcinoma, cystadenocarcinoma, bronchogenic carcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma and papillary adenocarcinomas.

[0224] In some embodiments, the treatment with adoptive T cells generated by the methoddescribed herein is directed to treatment of a specific patient population. In some embodiments,WSGR Docket No.: 50401-784.601 the adoptive T cells are directed to treatment of population of patients that are refractory to a certain therapy. For example, the T cells are directed to treatment of population of patients that are refractory to anti-checkpoint inhibitor therapy. In some embodiments, the patient is a melanoma patient. In some embodiments, the patient is a metastatic melanoma patient. In some embodiments, provided herein are methods of treating unresectable melanoma patient. In some embodiments, unresectable melanoma patients are selected for the T cell therapy described herein (such as NEO-PTC-01). Unresectable melanoma subjects may not be candidates for therapy with tumor infiltrating lymphocytes. In some embodiments, the treatment with adoptive T cells generated by the method described herein is directed to treatment of metastatic and unresectable melanoma patients. In some embodiments, the patient is refractory to anti-PD1 therapy. In some embodiments, the patient is refractory to anti-CTLA-4 therapy. In some embodiments, the patient is refractory to both anti-PD1 and anti-CTLA-4 therapy. In some embodiments, the therapy is administered by intravenously. In some embodiments, the therapy is administered by injection or infusion. In some embodiments the therapy is administered via a single dose, or 2, 3, 4, 5, 6, 7, 8, 9 or 10 doses. In some embodiments, the therapeutic or pharmaceutical composition comprises about 10^9 or higher total number of cells per dose. In some embodiments, the therapeutic or pharmaceutical composition comprises 10^10 or higher total number of cells per dose. In some embodiments, the therapeutic or pharmaceutical composition comprises 10^11 or higher total number of cells per dose. In some embodiments, the therapeutic or pharmaceutical composition comprises 10^12 or higher total number of cells per dose. In some embodiments, the subject is administered a therapeutic composition as described herein having about 10^10 to about 10^11 total cells per dose, wherein the cells have been validated for quality and have passed the release criteria. Kits

[0225] The methods and compositions described herein can be provided in kit form together withinstructions for administration. Typically, the kit can include the desired therapeutic compositions in a container, in unit dosage form and instructions for administration. Additional therapeutics, for example, cytokines, lymphokines, checkpoint inhibitors, antibodies, can also be included in the kit. Other kit components that can also be desirable include, for example, a sterile syringe, booster dosages, and other desired excipients.

[0226] Kits and articles of manufacture are also provided herein for use with one or more methodsdescribed herein. The kits can contain one or more types of immune cells. The kits can also contain reagents, peptides, and / or cells that are useful for antigen specific immune cell (e.g. neoantigenWSGR Docket No.: 50401-784.601 specific T cells) production as described herein. The kits can further contain adjuvants, reagents, and buffers necessary for the makeup and delivery of the antigen specific immune cells.

[0227] The kits can also include a carrier, package, or container that is compartmentalized toreceive one or more containers such as vials, tubes, and the like, each of the container(s) comprising one of the separate elements, such as the polypeptides and adjuvants, to be used in a method described herein. Suitable containers include, for example, bottles, vials, syringes, and test tubes. The containers can be formed from a variety of materials such as glass or plastic.

[0228] The articles of manufacture provided herein contain packaging materials. Examples ofpharmaceutical packaging materials include, but are not limited to, blister packs, bottles, tubes, bags, containers, bottles, and any packaging material suitable for a selected formulation and intended mode of administration and treatment. A kit typically includes labels listing contents and / or instructions for use, and package inserts with instructions for use. A set of instructions can also be included. EXAMPLES Example 1: A “de novo” method of identifying TCR chains from a plurality of patient libraries of TCR chain sequences

[0229] This example shows a method of identifying TCR chains from a plurality of patientlibraries of TCR chain sequences using “de novo” method. The general procedure is summarized in FIG. 1A. In order to identify T-cell receptor (TCR) chains recognizing specific antigens (e.g., HPV or mutant KRAS alleles) through de novo analysis, the patient population was separated into subgroups with or without a same condition or conditions (e.g., antigen expression, mutation profile, and / or HLA allele expression). Subgroups included HPV+ and HPV- patients as well as patients with mutant KRAS alleles (KRAS-G12C, KRAS-G12D, KRAS-G12V). For example, as shown in FIG. 4, TCR chains that recognize HPV were identified by searching for TCR chains enriched in that patient population (e.g. HPV+ and / or HLA-A*02:01+ patients) versus others (e.g. HPV- patients and / or HLA-A*02:01- patients). The relevant cohorts for HPV and KRAS were head / neck + cervical (FIG. 5A) and colorectal + pancreatic + lung (FIG. 5B), respectively. HPV status was defined by a threshold of 0 reads (negative) or > 100 reads (positive). Samples with intermediate values were excluded. Bulk TCR sequencing data (e.g., from a database) was analyzed to identify enriched TCR alpha and beta chains for the condition (e.g. HPV+ or positive for a specific KRAS mutation) vs control (e.g., HPV- or KRAS WT). Associations with specific HLA alleles (HLA-A:01:01, HLA-A:02:01, HLA-A:08:01, HLA-A:11:01, HLA-B:07:02, HLA-WSGR Docket No.: 50401-784.601 B:08:01, HLA-B:44:02, HLA-B:44:03, HLA-C:04:01, HLA-C:05:01, HLA-C:06:02, HLA- C:07:02, HLA-C:08:02) were statistically associated with enriched TCR chains.

[0230] In order to identify candidate TCR chains, all chains present in at least 5 patients wereanalyzed for enrichment by condition and HLA class I allele (FIG. 6). For each TCR chain a Fisher exact test was used to determine whether that TCR chain was over-represented in the patient population (e.g. HPV+ population) and any population defined by HLA class I alleles (e.g. HLA- A*02:01+ population). For example, FIG. 7A depicts the results of a Fisher exact test for one TCR chain by condition and by allele. The Fisher exact tests show association using patient counts. In addition, to control for the possibility that some patient populations might have different overall TCR chain counts, TCR chain detection events (e.g., one TCR chain detected in one patient) instead of patient counts were used. TCR chain enrichment by allele was plotted against TCR chain enrichment by condition to identify chains of interest (FIG. 7B). FIG. 8A summarizes the number of highest tier candidate chains identified by condition.

[0231] specific TCR chains of interest, two approaches were taken. First, paired TCR chains enriched in the same patient were searched for. For example, if a number of patients (e.g., 5 patients) had an patients (FIG. 10A). Second, publicly available scTCR datasets from these patient populations (e.g. HPV+ cancers for HPV and KRAS-mutation-enriched cancers for KRAS) were used to search for T cells contains the chains of interest plus paired chains (FIG. 10B). The number of paired chains by condition identified using enrichment analysis versus both enrichment analysis and scTCR shows that both methods can be paired to maximize paired chains uncovered (FIG. 13A). In total 689 candidate TCR chains across all four conditions were identified (FIG. 13B).

[0232] methods. In house avidity testing for the newly identified TCR chains has been initiated.

[0233] to a published TCR (Ros9a=Ros9b=Ros9d) that was used successfully in cell therapy against KRAS-G12D in two patients with pancreatic adenocarcinoma (PAAD) and colorectal adenocarcinoma (COAD) (Tran et al. 2016, NEJM; Sim et al. PNAS 2020; Leidner et al. 2022, Ros9d, at the same position that differs between Ros9a, Ros9b, Ros9d and Ros9c, along with 53 FIG 2.WSGR Docket No.: 50401-784.601

[0234] with Ros9a (they share the same TRBV gene, but a different TRBJ gene.) However, published residues critical for binding (CASSLG[E / R / Q]) are shared between our newly discovered paired betas and the original Ros9 TCRs (FIG. 11). Taken together, these results suggest that the de

[0235] discovered using the scTCR approach (using a dataset from HPV+ HNSC patients, disclosed in Eberhardt et al. 2021. Functional HPV-specific PD-1+ stem-like CD8 T cells in head and neck cancer. Nature 597: 279-284) (FIGs. 12A-C). Third, the de novo analysis for HPV returned an chain was enriched in patients with the same HLA allele (A:01:01) as the allele used for multimer sorting in Eberhardt et al, strengthening the case that it may be recognizing the same pMHC. This family showed no resemblance to the HPV TCR chains discovered in-house using NEO-STIM.

[0236] Avidity testing for the newly identified TCR chains is performed. For HPV TCR chains,the epitope they recognize is unknown; HPV TCR chains that do not recognize HPV E7 presented by HLA-A:02:01 (the combination being used for the avidity testing) may have other targets.

[0237] If any of our newly identified HPV and / or KRAS TCR chain s recognize the proposedtargets, the de novo method can be used to reveal additional TCR chains specific for these or other targets.Example 2: A “Bait” method of identifying TCR chains from a plurality of patient librariesof TCR chain sequences

[0238] This example shows a method of identifying TCR chains from a plurality of patientlibraries of TCR chain sequences using a known antigen specific TCR chain. The general procedure is summarized in FIG. 1B. In order to identify TCR chains based on known antigen specific TCR chain sequences, sequences of known HPV-specific and KRAS neoantigen-specific TCS were provided and queried against the plurality of libraries of TCR chain sequences forWSGR Docket No.: 50401-784.601 chains that are similar in sequence. Similarity was defined as matching V-gene and up to three amino acid substitutions / insertions / deletions in the CDR3. The following baits were included: -31 TCRs (18 identified in-house + 13 from the literature) specific for HPV E7. AllTCRs recognize HLA-A*02-presented epitopes. -3 TCRs specific for G12C KRAS. All HLA-A*11.- 10 TCRs specific for G12D KRAS. HLA-A*03, HLA-A*11, HLA-C*08.- 8 TCRs specific for G12V KRAS. HLA-A*03, HLA-A*11.

[0239] Each bait was independently queried against the plurality of libraries of TCR chainsequences. Table 1 summarizes the cohort size of each subgroup of subjects from whom the libraries of TCR chain sequences were obtained. Chains similar to the bait (e.g., bait neighbors) were identified. Bulk TCR sequencing data was analyzed to identify enriched TCR alpha and beta chains for the condition (e.g. HPV+ or positive for a specific KRAS mutation) vs control (e.g., HPV- or KRAS WT). Enrichment in patients that have the appropriate HLA allele was also tested as shown in FIG. 3A.

[0240] In addition, sequence logos made for identified bait neighbors were created as shown inFIG.3B and analyzed to determine if particular amino acids were enriched in any CDR3 position relative to the bait TCR chain sequence (e.g., the reference TCR chain sequence). While many TCR chain sequences have counts that are too low to determine statistical enrichment, logo analysis provided evidence that these TCR chain sequences were antigen-specific. Table 1. Cohort sizes.

[0241] The HPV cohorts included subjects with head and neck squamous cell carcinomas(HNSCC) and cervical squamous cell carcinoma (CESC), and KRAS cohorts included subjects with lung adenocarcinoma (LUAD), colon adenocarcinoma (COAD), and pancreatic adenocarcinoma (PAAD).

[0242] In Table 2 bait neighbor chains were classified into four tiers according to the level ofevidence of their specificity, with Tier 1 meaning very strong evidence and Tier 4 meaning low to no evidence. Table 2. Bait neighbor chains classified into four tiers.WSGR Docket No.: 50401-784.601Example 3: KRAS-specific TCRs identified from “de novo” and “bait” methods KRAS-specific TCRs identified from the “de novo” method

[0243] Using a clinically annotated database with T-cell repertoire data from patient tumors, ananalysis was performed to find TCR chains that were enriched with respect to allele and condition (KRAS-G12D). The data were from 5,714 KRAS-G12D samples and 16,181 KRAS WT samples, selected from patients with cancers known to have high rates of KRAS mutations: lung adenocarcinoma (LUAD), colon adenocarcinoma (COAD), and pancreatic adenocarcinoma (PAAD). Enrichment p-values were calculated using Fisher exact tests to compare the frequency of each TCR chain in the two sample groups (condition: KRAS-G12D vs WT; allele: allele+ vs allele-). Allele enrichments were performed for 22 Class I alleles, and the p-value from the strongest allele was used for plotting (FIG. 8B).

[0244] sequences (see arrows), two (indicated with grey boxes) were near-identical matches (within 1 2022). The Ros9 family comprises four closely related TCRs that recognize a peptide from mutant KRAS G12D presented by C*08:02, the same allele enriched in the analysis above (FIG. 8B). FIG.11). All four Ros9 acid of this motif is involved in recognizing the C:08:02 allele (Sim et al.2020). The Ros9 family has been used successfully in autologous T cell therapies for two patients with metastatic colon cancer (Tran et al. 2016) and metastatic pancreatic cancer (Leidner et al. 2022).To find paired were identified from the table above or a close variant (within 1 amino acid), then searched for publicly available single-cell TCR datasets from patient tumors from relevant cancer types (lung adenocarcinoma, pancreatic adenocarcinoma, and colorectal adenocarcinoma) were identified.WSGR Docket No.: 50401-784.601 FIG. 11 chains share the same TRBV gene as the four original Ros9 TCR beta chains, including the conserved 7 amino acid motif, strongly suggesting that they may recognize the same KRAS G12D epitope with a similar binding mechanism. The high sequence similarity of the alpha and beta chains discovered using the methods above to the original Ros9 family strongly suggests the same target specificity. KRAS G12D-specific Ros9 TCRs identified from “de novo” and “bait” methods

[0245] This example provides data analysis and summary of the KRAS G12D-specific TCRchains identified by “de novo” and “bait” methods as described in Example 1 and Example 2. The strongest statistical signal observed was for the alpha chains of ROS9a, ROS9b, ROS9c, ROS9dTCRs (Sim et al. PNAS 2020) (FIG. 2, FIG. 3A). ROS9a, ROS9b, ROS9d share the same alphachain (CDR3a: CLVGDMDQAGTALIF (SEQ ID NO: 16)) while ROS9c TRA differs in one amino acid (CLVGDRDQAGTALIF (SEQ ID NO: 17)) (FIG. 9C).

[0246] The top hit was an exact match of the ROS9a, ROS9b and ROS9d TRA which was stronglyassociated with the G12D mutation (p=2.6e-9, Fisher’s exact test for G12D vs WT) and with HLA-C*08 (p=1e-12, Fisher’s exact test for HLA-C*08+ vs HLA-C*08-). Restricting to HLA- C*08+ samples, this sequence appeared in 13 / 390 (3%) of KRAS G12D samples and in none of 1149 KRAS WT samples.

[0247] Another sequence (CLVGDIDQAGTALIF (SEQ ID NO: 49)) was also highly enrichedand differed from ROS9 CDR3s in the same position where ROS9c differs from ROS9a, ROS9b, and ROS9d (FIG. 8B). An exact match of the ROS9c TRA was found, although it was less enriched.

[0248] Overall six variants of ROS9 TCR chains that are likely specific for the antigen, and threemore that are potentially specific were found (FIGs. 8A and 8B). These sequences were additionally supported by x-ray structures for ROS9pMHC complexes (Protein Data Bank accession code 6ULR and 6ULN) (FIG. 9A).

[0249] Eight Beta chains were identified based on the shared motif CAS(S / T)(L / F / I)G(R / Q) thatappears in bait neighbors as well as in ROS9a,b,c,d beta chains. The importance of this motif for pMHC recognition was further supported by x-ray structures for ROS9a and ROS9d TCRpMHC complexes.

[0250] To conclude, multiple variants of Ros9 alpha and beta chains with strong evidence ofspecificity for HLA-C*08-presented KRAS G12D-derived peptide were identified.WSGR Docket No.: 50401-784.601 Example 4: HPV-specific TCRs identified from the “de novo” and “bait” methods HPV-specific TCRs identified from the “de novo” method

[0251] Using a clinically annotated database with T-cell repertoire data from patient tumors, ananalysis was performed to find TCR chains that were enriched with respect to allele and condition (HPV). The data were from 1755 HPV+ samples and 938 HPV- samples, all from patients with cervical cancer or head and neck cancer. Enrichment p-values were calculated using Fisher exact tests to compare the frequency of each TCR chain in the two sample groups (condition: HPV+ vs HPV-; allele: allele+ vs allele-). Allele enrichments were performed for 22 Class I alleles, and the p-value from the strongest allele was used for plotting (FIG. 12A).

[0252] from FIG. 12A these patients relative to all other patients (defined by an odds ratio >1 with detection in at least 2 patients) were searched for. For the “scTCR method,” publicly available single-cell TCR datasets from the tumors of patients with cancer types that are frequently driven by HPV infection (head and neck squamous cell carcinoma, cervical squamous cell carcinoma, cutaneous squamous cell carcinoma, and esophageal carcinoma) were identified. Clonotypes of at least 5 cells with our from these clonotypes were retrieved. Both methods identified cells from a family of TCRs with similar beta chains, defined by shared TRBV and TRBJ genes (FIG. 12B).

[0253] The scTCR data shown in FIG. 12B came from patients with HPV+ HNSC tumors,multimer-sorted cells specific for an epitope of HPV E2, QVDYYGLYY (SEQ ID NO: 50), presented by HLA-A:01:01 (FIG. 12C). Notably, HLA-A*01:01 was the same allele in which

[0254] The analyses above identified a family of TCRs (comprising the CDR3a shown in the FIG.12A, paired with the family of CDR3bs shown in FIG. 12B) that are either known to recognize HPV E2-QVDYYGLYY (SEQ ID NO: 50) / A*01:01 (for the five CDR3b chains identified from multimer-sorted cells in Eberhardt et al. 2021), or else share extremely high sequence similarity (for the two CDR3b chains identified using the enrichment approach, both within 2 amino acids of the most similar validated CDR3b from Eberhardt) (FIG.12B). These results indicated that the approach described above is capable of identifying complete TCRs, including both alpha and beta chains, that recognize targets of interest.WSGR Docket No.: 50401-784.601 HPV-specific TCRs identified from the “bait” method (I)

[0255] This example provides data analysis and summary of the HPV-specific TCR chainsidentified by the “bait” method as described in Example 2.

[0256] For HPV, neighbors were found with medium to strong evidence of specificity for threeout of 31 baits. Probably the strongest evidence is for the beta chain of the NCI TCR (or CRL4 TCR), for which two neighbors were found that appear in four and two HPV+ HLA-A*02+ samples respectively and do not appear in any HPV- or HLA-A*02- samples. The x-ray structure for the TCRpMHC for the NCI TCR chain suggested that the substitutions that these sequences have can be tolerated.

[0257] Among HPV E7-specific TCR chains, a cluster of beta chains stood out. They all haveTRBV2 and TRBJ2-7 and certain motifs in CDR3beta. The data currently includes eleven such TCR chains, and ten of them were included as baits. Some neighbors were found but none of them fully conformed to the CDR3 motif or were significantly enriched.

[0258] Further, TCR chain quality may not be the only parameter that affects TCR enrichment,the other being VDJ generation probability. For example, very strong enrichment for the alpha chains but not for the beta chains of ROS9 were found.

[0259] The TCR chain variants identified in the present disclosure will be tested experimentally.HPV-specific TCRs identified from the “bait” method (II)

[0260] The “bait approach” was used to find close neighbors of a published TCR that recognizedan epitope of HPV E7, YMLDLQPET (SEQ ID NO: 44), presented by HLA-A*02:01, referred to below as the NCI TCR chain (Jin et al. 2018. Engineered T cells targeting E7 mediate regression of human papillomavirus cancers in a murine model. JCI Insight 3(8):e99488; Nagarsheth et al. 2021. TCR-engineered T cells targeting E7 for patients with metastatic HPV-associated epithelial chain (within 2 amino acids) were significantly enriched in HPV+ patients compared to HPV- patients, as well as in patients with the A*02:01 allele compared to all other patients (FIG. 15A; see triangle indicated with an arrow).

[0261] The table below shows the CDR3 sequences and TRBV / TRBJ genes for the original NCITCR chain and two close neighbors discovered from the bait analysis (FIG. 15B). Positions that differ between the original CDR3b and the newly discovered CDR3b sequences are highlighted in red.WSGR Docket No.: 50401-784.601 Example 5: Experimental approaches for candidate TCR validation

[0262] This example provides two approaches to experimentally validating candidate TCRchains. The first approach is to use Surface plasmon resonance (SPR) (FIG. 14A). For SPR, pMHCs and soluble TCRs are produced, and affinity is measured using SPR. The second approach is to use cell lines (FIG. 14B) To validate using cell lines candidate TCR chains are expressed in Jurkat cells and then recognition against continuously infected HPV+, HLA-A:02:01 cell lines is tested. If recognition is seen, assays to reveal epitope and avidity are used. HPV-specific TCRs validated from the “bait” method (II)

[0263] Binding of the two variant TCRs plus the NCI TCR to HPV E7-YMLDLQPET (SEQ IDNO: 44) / HLA-A*02:01 was measured using surface plasmon resonance (SPR). The single- mutant hit (Hit 2) showed an improved Kd compared to the original NCI TCR, driven by an improved Koffvalue three times lower than that of the original (FIG. 15C).

[0264] Taken together, these results indicate that both variant TCRs identified using the baitapproach recognize the same target as the original NCI TCR. The single-point mutation (referred to above as Hit 2) moreover has improved binding compared to the original, which may make it a stronger therapeutic TCR. The NCI TCR has a high affinity of 290 nM (FIG. 15C) and has already been used successfully in clinical trials (Nagarsheth et al., 2021). These results suggest that the approach described above may be valuable for finding neighboring TCR chains with higher effectiveness than bait TCR chains, especially in cases where the initial bait has sub- optimal binding. Example 6: Example TCR sequences used for “bait” methods

[0265] Table 3 and Table 4 summarize example HPV and KRAS TCR chain sequences that canbe used as reference TCR chains or input TCR chains for “bait” methods described herein. Table 3. Summary of example TCR chain sequences (alpha chain sequences) used for “bait” methods, corresponding target genes, and alleles.WSGR Docket No.: 50401-784.601WSGR Docket No.: 50401-784.601Source Publications: 1https: / / doi.org / 10.4049 / jimmunol.202.Supp.131.4 Boutet et al, J Immunol 20192 https: / / doi.org / 10.1126 / sciadv.abf5835 Zhang et al, Sci Adv 20213 https: / / doi.org / 10.1172 / jci.insight.99488 Jin et al, JCI Insight 20184 https: / / doi.org / 10.1038 / s41591-020-01225-1 Nagarsheth, Nat Med 20215 https: / / doi.org / 10.1073 / pnas.1921964117 Sim et al, PNAS 20206 https: / / doi.org / 10.1038 / s41467-019-08304-z Cafri et al. Nat Comm 2019WSGR Docket No.: 50401-784.601 Table 4. Summary of example TCR chain sequences (corresponding beta chain sequences of the alpha chain sequences summarized in Table 3) used for “bait” methods, avidities, and epitope sequences.WSGR Docket No.: 50401-784.601WSGR Docket No.: 50401-784.601Example 7: Pan-Cancer approach to identifying TCRs

[0266] This example provides a method for identifying TCRs by adopting a pan-cancer approach.

[0267] When using condition association to find TCRs associated with a given antigen, it mayseem intuitive to restrict the analysis to cancer types relevant to the given antigen. For example, when analyzing for TCRs associated with HPV, HPV+ vs. HPV- patients within the population of patients with HPV-relevant cancer types (e.g., head and neck cancer, cervical cancer, analWSGR Docket No.: 50401-784.601 cancer, penile cancer, vulvar cancer, and vaginal cancer) may be compared. However, this example shows that significant statistical power may be achieved by adopting a pan-cancer approach wherein antigen-irrelevant cancer types are included.

[0268] TCR chain X (TRBV5-6, TRBJ2-1, CDR3: CASSLAWRGGSYNEQFF (SEQ ID NO:51)) was seen in 7 HPV+ samples and 0 HPV- samples among patients with an HPV-relevant cancer type. Considering the counts of all other TCR chains in those patients, the following contingency table was obtained:Fisher’s Exact Test p-value: 0.007676

[0269] Using a Bonferroni correction with alpha=0.05, a p-value of 0.05 / 1,485,605=3.4e-8 todetermine significance would be required. However, the p-value was not met (there were 1,485,605 unique TCR chains to be evaluated across those 6 cancer types); and thus, this TCR would not likely be selected for follow-up.

[0270] However, when all other cancer types were added into the analysis (lung cancer, prostatecancer, etc.), the following contingency table was obtained with much greater significance:Fisher’s Exact Test p-value: 2.295e-09

[0271] This p-value was less than the Bonferroni-corrected threshold of 3.4e-8, showing that thenew contingency was significant and would likely be selected for follow-up.

[0272] The finding suggests notable results because the updated analysis used no additionalobservations of TCR X. Rather, the improved significance was driven by the expanded population of antigen-negative samples. The absence of the TCRs in those samples implied that the TCR chain may be rare, making those 7 observations in HPV+ patient population surprising.

[0273] The particular chain in question in this example was found to be part of a potent TCRtargeting an HPV E7 epitope (https: / / pmc.ncbi.nlm.nih.gov / articles / PMC5931134 / ). Thus, the method described herein may provide the ability to pinpoint real, therapeutically TCRs, which otherwise may be ignored. Example 8: Approach to aid discovery of antigen-reactive TCRs

[0274] This example shows method of combining multiple sequence-similar TCRs into “meta-clonotypes” and analyzing antigen associations at the metaclonotype level to improve the statistical power. TCR ID Y (TRBV18, TRBJ2-7, CDR3: CASSPPEGSAYEQYF (SEQ ID NO:WSGR Docket No.: 50401-784.601 55)) was not significant after multiple hypothesis correction (e.g., via Benjamini-Hochberg method) in the pan-cancer analysis (Table 5). Table 5. Summary of analysis of example TCR chain sequences

[0275] However, this sequence was significant when considered as the centroid of a clusterdefined by including all TCR chains within radius=10 of the centroid, as shown in Table 5. Here, “distance” was measured according to a TCR-tailored sequence-similarity metric proposed in a publication (https: / / elifesciences.org / articles / 68605). When counting all samples that have at least one TCR chain within radius 10 of TCR ID Y, the p-value became extreme, and the FDR (e.g., per Benjamini-Hochberg) went below 1%. TCR ID Y was validated as part of a functional TCR (e.g., with TCRa: TRAV12-1, TRAJ9, CDR3: CVVPNIGGFKTIF (SEQ ID NO: 56)) recognizing an HPV E7 epitope presented on HLA-A*01:01, demonstrating the power of this approach to aid discovery of antigen-reactive TCRs. Example 9: Candidate TCR Scoring Method

[0276] This example shows a method of scoring identified candidate TCRs.

[0277] Enrichment pairing and single-cell pairing can return thousands of candidate TCRs for anantigen of interest, which may be a lot of TCRs to easily test experimentally. Applying a scoring system based on pairing and condition enrichment evidence can identify candidate antigen- specific TCRs that are most likely to be validated.

[0278] For HPV, candidate TCRs were scored as shown in Table 6.Table 6. Pairing score metric and points associated (maximum score 8)WSGR Docket No.: 50401-784.601

[0279] Further, condition enrichment score (applied separately to candidate alpha and beta chains)was assessed. TCR chains were scored based on the ‘minimal’ analysis in which they were significant, e.g., the analysis with the smallest cluster radius, as shown in Table 7. Table 7. Enrichment score metric and points associated (maximum score 8 (+4 each for alpha and beta))

[0280] The combined scores were highly predictive of which TCRs validated for HPV-specificityin a test of 12 candidate HPV-specific TCR families (where each family included 1-10 individual TCR variants). FIG. 16 shows the scores for the best-scoring member of each TCR family. Families were marked as ‘validated’ if at least one TCR was identified as HPV-specific in a cell recognition assay and as DNV (did not validate) if none were validated.

[0281] These results demonstrated the power of the scoring system to efficiently sort through ahigh volume of computationally predicted candidate alpha chain-beta chain pairings to identify the subset that may be most likely to validate as antigen-specific when tested experimentally.

[0282] While preferred embodiments of the present disclosure have been shown and describedherein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention.WSGR Docket No.: 50401-784.601 It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.

Claims

WSGR Docket No.: 50401-784.601 CLAIMS What is claimed is:

1. A method of identifying therapeutically relevant T-cell receptor (TCR) sequences, themethod comprising: (a) providing a plurality of libraries of TCR chain sequences having a first plurality oflibraries of TCR chain sequences from a first sub-group of at least 10 subjects and a second plurality of libraries of TCR chain sequences from a second sub-group of at least 10 subjects, wherein each library of TCR chain sequences from a subject of the first sub-group of at least 10 subjects is from a single subject and each library of TCR chain sequences from a subject of the second sub-group of at least 10 subjects is from a single subject; wherein: (i) each of the subjects of the first sub-group of at least 10 subjects has the samedisease or condition, and each of the subjects of the second sub-group of at least 10 subjects does not have the same disease or condition as the first sub- group of at least 10 subjects; (ii) each of the subjects of the first sub-group of at least 10 subjects has the sameantigen profile, and each of the subjects of the second sub-group of at least 10 subjects does not have the same antigen profile as the first sub-group of at least 10 subjects; and / or (iii) each of the subjects of the first sub-group of at least 10 subjects express amajor histocompatibility complex (MHC) encoded by a same human leukocyte antigen (HLA) allele, and each of the subjects of the second sub- group of at least 10 subjects do not express an MHC encoded by the same HLA allele as the first sub-group of at least 10 subjects; (b) identifying one or more TCR chain sequences based on an analysis of TCR chainsequences in the plurality of libraries of TCR chain sequences, wherein the analysis comprises comparing a frequency of one or more TCR chain sequences from the first sub-group to a frequency of the same one or more TCR chain sequences from the second sub-group, wherein the frequency of one or more TCR chain sequences from the first sub-group is the number of subjects of the first sub-group with the one or more TCR chain sequences over the total number of unique TCR chain sequences of the first sub-group, and the frequency of the same one or more TCR chain sequences from the second sub-group is the number of subjects of the second sub-group withWSGR Docket No.: 50401-784.601 the same one or more TCR chain sequence over the total number of unique TCR chain sequences of the second sub-group; (c) determining a statistical significance of a difference in a frequency of a TCR chainsequence of the one or more TCR chain sequences from the first sub-group of at least 10 subjects to a frequency of the same TCR chain sequence of the one or more TCR chain sequences from the second sub-group of subjects based on the comparing in (b), thereby identifying a potentially therapeutic TCR chain sequence from the one or more TCR chain sequences obtained from the first sub-group of at least 10 subjects; and (d) selecting a TCR chain sequence of the first sub-group of at least 10 subjects that (i)has a statistically significant difference in frequency to the frequency of the same TCR chain sequence in the second sub-group of subjects based on (c), wherein the statistically significant difference in the frequency is a p-value of at most 0.1, (ii) is present at a higher frequency in the first sub-group of at least 10 subjects relative to the frequency of the same TCR chain sequence in the second sub-group of subjects, and (iii) is a TCR chain sequence that is present in at least 2 subjects from the first sub-group of at least 10 subjects.

2. The method of claim 1, wherein the number of subjects of the first sub-group of at least10 subjects with the TCR chain sequence selected in (d) over the total number of unique TCR chain sequences of the first sub-group of at least 10 subjects has a statistically significant difference compared to the number of subjects of the second sub-group of subjects with the TCR chain sequence selected in (d) over the total number of unique TCR chain sequences of the second sub-group of subjects.

3. The method of claim 1 or 2, wherein providing in (a) comprises:(a) providing the plurality of libraries of TCR chain sequences from a population ofsubjects; (b) identifying the first sub-group of at least 10 subjects as (i) having the same disease orcondition, (ii) having the same antigen profile, (iii) expressing the MHC encoded by the same HLA allele, or (iv) the same combination thereof, and the second sub-group of subjects as (i) not having the same disease or condition as the first sub-group, (ii) not having the same antigen profile as the first sub-group, (iii) not expressing an MHC encoded by the same HLA allele as the first sub-group, or (iv) not having the same combination thereof as the first sub-group.

4. The method of any one of claims 1-3, wherein the statistically significant difference inthe frequency is a p-value of at most 0.05, at most 0.001, or at most 0.0001.WSGR Docket No.: 50401-784.6015. The method of any one of claims 1-4, wherein the p-value is an adjusted p-value.

6. The method of any one of claims 1-5, wherein the selected TCR chain sequence is aTCR chain sequence that is present in at least 10 subjects from the first sub-group of at least 10 subjects.

7. The method of any one of claims 1-6, wherein the first sub-group of at least 10 subjectshave the same disease or condition, and the second sub-group of at least 10 subjects do not have the same disease or condition as the first sub-group.

8. The method of claim 7, further comprising repeating (a)-(d) using the plurality oflibraries of TCR chain sequences from a first sub-group of at least 10 subjects that express a MHC encoded by a same HLA allele, and a second sub-group of at least 10 subjects that do not express an MHC encoded by the same HLA allele as the first sub- group.

9. The method of any one of claims 1-6, wherein the first sub-group of at least 10 subjectshave the same antigen profile, and the second sub-group of at least 10 subjects do not have the same antigen profile as the first sub-group.

10. The method of claim 9, further comprising repeating (a)-(d) using the plurality oflibraries of TCR chain sequences from a first sub-group of at least 10 subjects that express a MHC encoded by a same HLA allele, and a second sub-group of at least 10 subjects that do not express an MHC encoded by the same HLA allele as the first sub- group.

11. The method of any one of claims 1-6, wherein the first sub-group of at least 10 subjectshave the same disease or condition and the same antigen profile, and the second sub- group of at least 10 subjects do not have the same disease or condition as the first sub- group and does not have the same antigen profile as the first sub-group.

12. The method of claim 11, further comprising repeating (a)-(d) using the plurality oflibraries of TCR chain sequences from a first sub-group of at least 10 subjects that express a MHC encoded by a same HLA allele, and a second sub-group of at least 10 subjects that do not express an MHC encoded by the same HLA allele as the first sub- group.

13. The method of any one of claims 1-6, wherein the first sub-group of at least 10 subjectsexpress a MHC encoded by a same HLA allele, and a second sub-group of at least 10 subjects do not express an MHC encoded by the same HLA allele as the first sub- group.

14. The method of claim 13, further comprising repeating (a)-(d) using the plurality oflibraries of TCR chain sequences from a first sub-group of at least 10 subjects that haveWSGR Docket No.: 50401-784.601 the same disease or condition, and the second sub-group of at least 10 subjects that do not have the same disease or condition as the first sub-group.

15. The method of claim 13, further comprising repeating (a)-(d) using the plurality oflibraries of TCR chain sequences from a first sub-group of at least 10 subjects that have the same antigen profile, and the second sub-group of at least 10 subjects that do not have the same antigen profile as the first sub-group.

16. The method of any one of claims 1-15, wherein the first sub-group of at least 10subjects comprises at least 100, at least 1,000, or more subjects.

17. The method of any one of claims 1-16, wherein the second sub-group of at least 10subjects comprises at least 100, at least 1,000, or more subjects.

18. The method of any one of claim 1-17, wherein the selected TCR chain sequence is aTCR alpha chain sequence.

19. The method of claim 18, further comprising selecting a TCR beta chain sequence that ispresent at a higher frequency in the first sub-group of at least 10 subjects relative to the frequency of the same TCR beta chain sequence in the second sub-group of at least 10 subjects, and is present in at least 2 subjects from the first sub-group of at least 10 subjects.

20. The method of claim 18, further comprising identifying a TCR beta chain sequenceassociated with or cognately paired to the selected TCR alpha chain.

21. The method of claim 18, further comprising (i) providing single-cell TCR sequencingdata from a subject having the same disease or condition, having the same antigen profile, expressing an MHC encoded by the same HLA allele, or having the same combination thereof, (ii) determining if a TCR alpha chain sequence is present in the single-cell TCR sequencing data that is the same as or has at most 1, 2, 3, 4 or 5 amino acid differences with the selected TCR alpha chain sequence, and (iii) identifying a TCR beta chain sequence associated with or cognately paired to the selected TCR alpha chain sequence in the single-cell TCR sequencing data.

22. The method of any one of claim 1-17, wherein the selected TCR chain sequence is aTCR beta chain sequence.

23. The method of claim 22, further comprising selecting a TCR alpha chain sequence thatis present at a higher frequency in the first sub-group of at least 10 subjects relative to the frequency of the same TCR alpha chain sequence in the second sub-group of at least 10 subjects, and is present in at least 2 subjects from the first sub-group of at least 10 subjects.WSGR Docket No.: 50401-784.60124. The method of claim 22, further comprising identifying a TCR alpha chain sequenceassociated with or cognately paired to the selected TCR beta chain.

25. The method of claim 22, further comprising (i) providing a single-cell TCR sequencingdata from a subject having the same disease or condition, having the same antigen profile, expressing an MHC encoded by the same HLA allele, or having the same combination thereof, (ii) determining if a TCR beta chain sequence is present in the single-cell TCR sequencing data that is the same as or has at most 1, 2, 3, 4 or 5 amino acid differences with the selected TCR beta chain sequence, and (iii) identifying a TCR alpha chain sequence associated with or cognately paired to the selected TCR beta chain sequence in the single-cell TCR sequencing data.

26. The method of any one of claims 18-25, further comprising pairing the TCR alphachain sequence and the TCR beta chain sequence to form a paired TCR chain sequences.

27. The method of any one of claim 1-17, wherein the selected TCR chain sequence is aTCR gamma chain sequence.

28. The method of claim 27, further comprising selecting a TCR delta chain sequence thatis present at a higher frequency in the first sub-group of at least 10 subjects relative to the frequency of the same TCR delta chain sequence in the second sub-group of at least 10 subjects, and is present in at least 2 subjects from the first sub-group of at least 10 subjects.

29. The method of claim 27, further comprising identifying a TCR delta chain sequenceassociated with or cognately paired to the selected TCR gamma chain.

30. The method of claim 27, further comprising (i) providing a single-cell TCR sequencingdata from a subject having the same disease or condition, having the same antigen profile, expressing an MHC encoded by the same HLA allele, or having the same combination thereof, (ii) determining if a TCR gamma chain sequence is present in the single-cell TCR sequencing data that is the same as or has at most 1, 2, 3, 4 or 5 amino acid differences with the selected TCR gamma chain sequence, and (iii) identifying a TCR delta chain sequence associated with the selected TCR gamma chain sequence in the single-cell TCR sequencing data.

31. The method of any one of claim 1-17, wherein the selected TCR chain sequence is aTCR delta chain sequence.

32. The method of claim 31, further comprising selecting a TCR gamma chain sequencethat is present at a higher frequency in the first sub-group of at least 10 subjects relative to the frequency of the same TCR gamma chain sequence in the second sub-group of atWSGR Docket No.: 50401-784.601 least 10 subjects, and is present in at least 2 subjects from the first sub-group of at least 10 subjects.

33. The method of claim 31, further comprising identifying a TCR gamma chain sequenceassociated with or cognately paired to the selected TCR delta chain.

34. The method of claim 31, further comprising (i) providing a single-cell TCR sequencingdata from a subject having the same disease or condition, having the same antigen profile, expressing an MHC encoded by the same HLA allele, or having the same combination thereof, (ii) determining if a TCR delta chain sequence is present in the single-cell TCR sequencing data that is the same as or has at most 1, 2, 3, 4 or 5 amino acid differences with the selected TCR delta chain sequence, and (iii) identifying a TCR gamma chain sequence associated with the selected TCR delta chain sequence in the single-cell TCR sequencing data.

35. The method of any one of claims 27-34, further comprising pairing the TCR gammachain sequence and the TCR delta chain sequence to form a paired TCR chain sequences.

36. The method of claim 26 or 35, wherein a TCR having the paired TCR chain sequencesis not a cognate TCR pair.

37. The method of any one of claims 1-36, wherein the same condition is HPV infection,and wherein a TCR having the paired TCR chain sequences binds to an epitope associated with the HPV infection in complex with an HLA molecule.

38. The method of claim 37, wherein the first sub-group of at least 10 subjects comprise asame HLA allele.

39. The method of any one of claims 26-38, wherein the same antigen profile comprises asame oncogenic driver mutation, and wherein a TCR having the paired TCR chain sequences binds to a neoepitope comprising the same oncogenic driver mutation in complex with an HLA molecule.

40. The method of claim 39, wherein the first sub-group of at least 10 subjects comprise asame HLA allele.

41. The method of any one of claims 1-40, wherein the same condition comprises HPVinfection.

42. The method of any one of claims 1-40, wherein the same condition comprises a cancer.

43. The method of claim 42, wherein the cancer is associated with an antigen recognizedby a TCR having the paired TCR chain sequences.

44. The method of claim 42, wherein the cancer is not associated with an antigenrecognized by a TCR having the paired TCR chain sequences.WSGR Docket No.: 50401-784.60145. The method of any one of claims 1-40, wherein the same condition comprises anautoimmune disorder.

46. The method of any one of claims 1-40, wherein the same condition comprises aninfectious disease.

47. The method of any one of claims 1-46, wherein the same antigen profile compriseshaving a same cancer mutation, a same neoantigen, a same tumor-associated antigen, or any combination thereof.

48. The method of claim 47, wherein the same cancer mutation comprises KRAS-G12C,KRAS-G12D, or KRAS-G12V mutation.

49. The method of any one of claims 1-48, wherein the same HLA allele comprises anallele selected from the group consisting of HLA-A:01:01, HLA-A:02:01, HLA- A:03:01, HLA-A:11:01, HLA-B:07:02, HLA-B:08:01, HLA-B:44:02, HLA-B:44:03, HLA-C:04:01, HLA-C:05:01, HLA-C:06:02, HLA-C:07:02, and HLA-C:08:02.

50. The method of any one of claims 1-49, wherein each library of TCR chain sequences isa sequencing dataset obtained by bulk sequencing of TCR chain sequences from a subject.

51. The method of any one of claims 26-50, wherein an epitope or an HLA allele that aTCR comprising the paired TCR chain sequences recognizes is unknown.

52. The method of any one of claims 26-51, further comprising assaying a TCRcomprising the paired TCR chain sequences for a binding affinity against an epitope associated with the same disease or condition or associated with the same antigen profile in complex with the same HLA allele.

53. The method of any one of claims 1-52, wherein a TCR comprising the selected TCRchain sequence recognizes an epitope from HPV E2.

54. The method of claim 53, wherein a TCR comprising the selected TCR chain sequencerecognizes an epitope from HPV E2 in complex with HLA-A:01:01.

55. The method of any one of claims 1-54, wherein the frequency of one or more TCRchains in each library of TCR chain sequences from the first sub-group of at least 10 subjects is at least about 2-fold the frequency of one or more TCR chains in each library of TCR chain sequences from the second sub-group of at least 10 subjects.

56. The method of any one of claims 1-55, wherein the statistical significance of a TCRchain sequence is determined by a Fisher’s exact test.

57. The method of any one of claims 1-56, further comprising, prior to determining thestatistical significance, grouping two or more TCR chain sequences of the one or moreWSGR Docket No.: 50401-784.601 TCR chain sequences based on sequence identity into a grouped TCR chain sequences, wherein the grouped TCR chain sequences share at least 70% sequence identity.

58. The method of claim 57, wherein variable regions of the grouped TCR chain sequencesshare at least 70% sequence identity.

59. The method of claim 57 or 58, wherein CDR3 sequences of the grouped TCR chainsequences share at least 70% sequence identity.

60. The method of any one of claims 1-56, wherein the one or more TCR chain sequencescomprise two or more TCR chain sequences that have been grouped based on sequence identity, and wherein the grouped TCR chain sequences share at least 70% sequence identity.

61. The method of claim 60, wherein variable regions of the grouped TCR chain sequencesshare at least 70% sequence identity.

62. The method of claim 61, wherein CDR3 sequences of the grouped TCR chainsequences share at least 70% sequence identity.

63. The method of any one of claims 1-59, wherein selecting the TCR chain sequencescomprises selecting a plurality of TCR chain sequences, and wherein the method further comprises aligning the plurality of TCR chain sequences to obtain conserved residues or sequence motifs.

64. The method of any one of claims 26-63, further comprising generating a pairing scorefor a TCR having the paired TCR chain sequences, wherein the pairing score predicts likelihood of the TCR to be an antigen-specific TCR to be validated experimentally.

65. The method of claim 64, wherein the pairing score is calculated based on enrichment p-value, single-cell hit analysis, de novo match analysis, and / or enrichment match analysis, or combinations thereof.

66. The method of claim 65, wherein the enrichment p-value is at most 1e-10.

67. The method of claim 65, wherein the single-cell hit analysis comprises identification ofthe TCR in a single-cell sample.

68. The method of claim 65, wherein the de novo match analysis or enrichment matchanalysis comprises identification of the paired TCR chain sequences at a radius of a centroid of at least 5 of another identified antigen-specific TCR.

69. The method of any one of claims 1-68, further comprising generating an enrichmentscore for the TCR chain sequence or the plurality of TCR chain sequences selected in (d), wherein the enrichment score predicts likelihood of the TCR chain sequences or the plurality of TCR chain sequences to be an antigen-specific TCR when paired with a corresponding TCR chain to be validated experimentally.WSGR Docket No.: 50401-784.60170. The method of claim 69, wherein the enrichment score is calculated based onperforming a singleton condition-enrichment analysis, and / or a cluster-based condition- enrichment analysis.

71. The method of claim 70, wherein the cluster-based condition-enrichment analysiscomprises identification of the TCR chain sequence or the plurality of TCR chain sequences at a radius of a centroid of at least 5, at least 10, and / or at least 20, or combinations thereof.

72. The method of any one of claims 1-63, further comprising analyzing the likelihood ofthe selected TCR chain sequence to be generated by thymic selection.

73. The method of any one of claims 1-72, wherein the first sub-group of at least 10subjects have the same antigen profile as the second sub-group of at least 10 subjects, and wherein the first sub-group of at least 10 subjects express a protein with a same antigen or RNA encoding the protein with a same antigen at a higher level than the expression of the protein with a same antigen or RNA encoding the protein with a same antigen in the second sub-group of at least 10 subjects.

74. The method of claim 73, wherein the same antigen comprises a mutation.

75. The method of any one of claims 1-73, wherein the first sub-group of at least 10subjects have the same disease or condition and the second sub-group of at least 10 subjects do not have the same disease or condition.

76. A method of identifying therapeutically relevant T-cell receptor (TCR) sequences, themethod comprising: (a) providing a plurality of libraries of TCR chain sequences from at least 10 subjects,each library of TCR chain sequences is from a single subject; (b) selecting a TCR chain sequence that is present in at least 2 subjects from the at least10 subjects; (c) based on the TCR chain sequence selected in (b), subgrouping the plurality oflibraries of TCR chain sequences into a first plurality of libraries of TCR chain sequences from a first sub-group of subjects and a second plurality of libraries of TCR chain sequences from a second sub-group of subjects, wherein the first plurality of libraries of TCR chain sequences from the first sub-group of subjects comprises the TCR chain sequence selected in (b), and the second plurality of libraries of TCR chain sequences from the second sub-group of subjects does not comprise the TCR chain sequence selected in (b); and (d) determining whether the first sub-group of subjects and the second sub-group ofsubjects have differential expression level of an antigen.WSGR Docket No.: 50401-784.60177. The method of claim 76, wherein determining in (d) comprises analyzing sequencingdata of the first sub-group and the second sub-group.

78. The method of claim 76 or 77, wherein the first sub-group of subjects and the secondsub-group of subjects have differential expression level of the antigen, and wherein the antigen has a probability of being recognized by the TCR chain sequence selected in (b) when being presented by an MHC.

79. A method of identifying therapeutically relevant T-cell receptor (TCR) sequences, themethod comprising: (a) providing a reference TCR chain sequence and libraries of TCR chain sequencesfrom a population of subjects, wherein each library is from a single subject of the population of subjects; (b) analyzing one or more TCR chain sequences of the libraries of TCR chain sequencesto determine sequence similarity between the one or more TCR chain sequences and the reference TCR chain sequence, wherein the reference TCR chain sequence is from a known antigen-specific TCR; and (c) identifying one or more TCR chain sequences from the libraries of TCR chainsequences that have a sequence similarity with the reference TCR chain sequence, wherein an identified TCR chain sequence a same amino acid sequence encoded by the V-gene as the reference TCR chain sequence and no more than three amino acid variations in a CDR3 sequence compared to a CDR3 sequence of the reference TCR chain sequence, wherein the amino acid variations comprises amino acids substitutions, deletions and / or insertions.

80. The method of claim 79, wherein the reference TCR chain comprises a reference TCRalpha chain sequence and a reference TCR beta chain sequence from the known antigen-specific TCR, and wherein analyzing one or more TCR chain sequences in (b) comprises analyzing each TCR alpha chain sequence and each TCR beta chain sequence separately.

81. The method of claim 79 or 80, wherein analyzing one or more TCR chain sequences in(b) comprises analyzing each TCR alpha chain sequence of the libraries of TCR chain sequences to determine sequence similarity between each TCR alpha chain sequence and a reference TCR alpha chain sequence.

82. The method of claim 79 or 80, further comprising analyzing each TCR beta chainsequence of the libraries of TCR chain sequences to determine sequence similarity between each TCR beta chain sequence and a reference TCR beta chain sequence paired natively with the reference TCR alpha chain sequence.WSGR Docket No.: 50401-784.60183. The method of any one of claims 79-81, further comprising, for each TCR chain of theone or more TCR chain sequences identified in (c), determining whether the TCR chain is likely to be restricted by an HLA allele by conducting a statistical test assessing whether a frequency of the TCR chain in a first subgroup of the population of subjects positive for the HLA allele is different from a frequency of the TCR chain in a second subgroup of the population of subjects negative for the HLA allele, and wherein the p- value for the statistical test is at most 0.1.

84. The method of any one of claims 79-82, wherein the p-value is at most 0.01.

85. The method of any one of claims 79-85, wherein the p-value is at most 0.001.

86. The method of any one of claims 79-85, wherein the p-value is at most 0.0001.

87. The method of any one of claims 79-86, further comprising determining the likelihoodof identified TCR chain sequence specificity to the same pMHC complex as the reference TCR chain sequence.

88. The method of claim 87, wherein determining the likelihood comprises aligning aCDR3 sequence of the identified TCR chain sequences to the CDR3 sequence of the reference TCR chain sequence.

89. The method of any one of claims 79-88, wherein a TCR comprising the reference TCRchain sequence recognizes HLA-A*02-presented epitopes, HLA-A*11-presented epitopes, HLA-A*03-presented epitopes, or HLA-C*08-presented epitopes.

90. The method of any one of claims 79-89, wherein the identified TCR chain sequence is apotentially therapeutic TCR chain sequence if the HLA allele is the same allele recognized by the reference TCR chain sequence.

91. The method of any one of claims 79-90, wherein the HLA allele recognized by theidentified TCR chain sequence and the reference TCR chain sequence is selected from the group consisting of HLA-A*02, HLA-A*11, HLA-A*03, and HLA-C*08.

92. The method of any one of claims 79-91, wherein a TCR comprising the reference TCRchain sequence is specific for an epitope from HPV E7.

93. The method of claim 92, wherein a TCR comprising the identified TCR chain sequenceis specific for an epitope of HPV E7.

94. The method of claim 92 or 93, wherein a TCR comprising the identified TCR chainsequence is specific for the epitope from HPV E7 in complex with HLA-A*02.

95. The method of any one of claims 79-91, wherein a TCR comprising the reference TCRchain sequence is specific for an epitope from KRAS-G12C.

96. The method of claim 95, wherein a TCR comprising the reference TCR chain sequenceis specific for the epitope from KRAS-G12C in complex with HLA-A*11.WSGR Docket No.: 50401-784.60197. The method of claim 95 or 96, wherein a TCR comprising the identified TCR chainsequence is specific for the epitope from KRAS-G12C in complex with HL-A*11.

98. The method of any one of claims 79-91, wherein a TCR comprising the reference TCRchain sequence is specific for an epitope from KRAS-G12D.

99. The method of claim 98, wherein a TCR comprising the reference TCR chain sequenceis specific for the epitope from KRAS-G12D in complex with HLA-A*03, HLA-A*11, or HLA-C*08.

100. The method of claim 98 or 99, wherein a TCR comprising the identified TCR chainsequence is specific for the epitope from KRAS-G12D in complex with HLA-A*03, HLA-A*11, or HLA-C*08.

101. The method of any one of claims 79-91, wherein a TCR comprising the reference TCRchain sequence is specific for an epitope from KRAS-G12V.

102. The method of claim 101, wherein a TCR comprising the reference TCR chainsequence is specific for the epitope from KRAS-G12V in complex with HLA-A*03 or HLA-A*11.

103. The method of claim 101 or 102, wherein a TCR comprising the identified TCR chainsequence is specific for the epitope from KRAS-G12V in complex with HLA-A*03 or HLA-A*11.

104. The method of any one of claims 79-102, further comprising analyzing the likelihoodfor the identified TCR chain to be generated by thymic selection.

105. The method of any one of claims 79-104, wherein the reference TCR chain sequence isa TCR chain sequence selected by the method of any one of claims 1-78.

106. A method of making a TCR comprising preparing a TCR comprising a TCR chainsequence selected according to the method of any one of claims 1-105.

107. A method of making a TCR comprising preparing a TCR comprising a TCR chainsequence identified according to the method of any one of claims 79-105 and the reference TCR chain sequence.

108. A pharmaceutical composition comprising a TCR comprising a TCR chain sequenceselected according to the method of any one of claims 1-105 or a cell comprising the TCR comprising the TCR chain sequence selected according to the method of any one of claims 1-105, and a pharmaceutically acceptable carrier.

109. A pharmaceutical composition comprising a TCR comprising a TCR chain sequenceselected according to the method of any one of claims 79-105 and the reference TCR chain sequence of any one of claims 79-105 or a cell comprising the TCR comprising the TCR chain sequence selected according to the method of any one of claims 79-105WSGR Docket No.: 50401-784.601 and the reference TCR chain sequence of any one of claims 79-105, and a pharmaceutically acceptable carrier.

110. A method of treating a subject in need thereof comprising administering thepharmaceutical composition of claim 108 or 109 into the subject.

111. Use of a TCR comprising a TCR chain sequence selected according to the method ofany one of claims 1-105, a cell comprising the TCR comprising the TCR chain sequence selected according to the method of any one of claims 1-105, a TCR comprising a TCR chain sequence selected according to the method of any one of claims 79-105 and the reference TCR chain sequence of any one of claims 79-105, a cell comprising the TCR comprising the TCR chain sequence selected according to the method of any one of claims 79-105 and the reference TCR chain sequence of any one of claims 79-105, or a pharmaceutical composition of claim 108 or 109 in the manufacture of a medicament for treating a disease.

Citation Information

Patent Citations

  • Multigene analysis of tumor samples

    US20170356053A1

  • Methods of isolating neoantigen-specific t cell receptor sequences

    US20200056237A1

  • Method and systems for prediction of HLA class ii-specific epitopes and characterization of CD4+ t cells

    US20200279616A1

  • Method to isolate TCR genes

    US20210040558A1