Novel tumor-specific antigens for colorectal cancer and their uses
Patent Information
- Application Number
- JP2024516885
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-17
- Filing Date
- 2022-09-16
- Publication Date
- 2025-09-19
AI Technical Summary
Current immunotherapies for colorectal cancer, such as immune checkpoint inhibitors and tumor-associated antigen vaccines, have shown limited efficacy, particularly in microsatellite-stable tumors, and there is a need for more effective tumor-specific antigens that can elicit therapeutic immune responses.
Identification and utilization of tumor-specific antigens (TSAs) derived from non-coding regions of the genome, which are aberrantly expressed or mutated, and their use in vaccines or T-cell receptor-based therapies, potentially combined with immune checkpoint inhibitors.
The use of TSAs can enhance immune recognition and response against colorectal cancer, offering a more targeted and effective treatment approach across various tumor types, including microsatellite-stable tumors.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 261,315, filed September 17, 2021, which is incorporated herein by reference.
[0002] The present disclosure relates generally to the field of oncology, and more particularly to the treatment of cancers such as colorectal cancer. [Background technology]
[0003] Colorectal cancer (CRC) is the third most commonly diagnosed cancer and the second leading cause of cancer death worldwide, with over 1.8 million cases and an estimated 881,000 deaths in 2018 alone (1). The incidence of CRC is expected to increase as global socioeconomic changes occur, with a predicted 2.2 million cases and 1.1 million deaths per year by 2030 (1, 2). This significant disease burden highlights the need to develop novel and effective treatments for this disease.
[0004] The positive correlation between tumor-infiltrating lymphocyte (TIL) abundance and increased overall survival in both colon and rectal cancer suggests that T cells are able to recognize biologically relevant tumor antigens in these tumors (3, 4). The potential immunogenicity of these antigens has made immune checkpoint inhibitors (ICIs) a promising treatment for cancer patients. However, early clinical trials evaluating their efficacy in CRC have yielded mixed results. Colorectal tumors characterized by deficiencies in mismatch repair proteins, which cause the accumulation of repetitive DNA sequences (microsatellites) known as microsatellite instability (MSI), have shown relative success in phase II clinical trials using anti-PD1 therapy (5). In contrast, such treatments have shown little efficacy in clinical trials for microsatellite-stable (MSS) tumors, which do not have a high mutational burden and comprise approximately 80% of CRC cases (5, 6).
[0005] Given the importance of immune responses in CRC and the limited success of ICIs alone, promising avenues for research in recent years include neoantigen-based vaccines or T cell receptor-based therapies, which can be used with or without ICIs and ideally could bridge the gap in therapeutic efficacy across MSI and MSS tumors. For example, tumor-associated antigens (TAAs), antigens overexpressed in cancer cells compared with normal cells, have previously been identified in CRC (7, 8). Several TAAs have been tested in phase I trials for vaccines and CRC, but most have met with limited success, likely due to negative selection of such antigens by the thymus (9). In one study, treatment of metastatic CRC with genetically engineered anti-CEA T cells resulted in tumor regression in one patient but caused severe inflammatory bowel disease in all patients, indicating that adverse autoimmune responses are another possible negative consequence of using TAAs (10).
[0006] Due to the mixed response to TAAs, effective neoantigen-based therapies are more likely to utilize tumor-specific antigens (TSAs), which may be generated by genetic, epigenetic, and post-translational variations, including but not limited to single-nucleotide variants, aberrantly expressed transcripts, or novel splicing events, and are expressed exclusively by tumors (11). The high incidence of single-nucleotide variants, splice variants, and INDEL mutations in CRC suggests that unique antigens are likely presented by tumor major histocompatibility complexes (MHCs), potentially eliciting tumor-specific immune responses (12). TSAs have recently been identified in CRC and have shown some success in phase I and II vaccine trials. A 2015 vaccine trial using a frameshift antigen derived from MSI-high tumors demonstrated significant and specific immune responses in all patients (13). However, because this study used an antigen derived from a microsatellite instability frameshift, these findings may not be applicable to the majority of CRC patients. Other studies identifying TSAs in CRC to date have focused exclusively on mutant TSAs (mTSAs) derived from coding regions of the genome (13, 14). Examination of MSS CRC organoids revealed that only 0.5% of non-silent mutations were identified as mTSAs, a significantly lower percentage than predicted by HLA-binding prediction software (15). Additionally, previous studies have not used mass spectrometry (MS) techniques to quantify the expression of those TSAs on tumor cells, information that may impact the therapeutic potential of targeting a given TSA (13, 14).
[0007] Given this, there is an urgent need to identify tumor-specific antigens that can elicit therapeutic immune responses against CRC, which could be used as vaccines (± immune checkpoint inhibitors) or as targets for T cell receptor-based approaches (cell therapy, bispecific biologics).
[0008] This description makes reference to several documents, the contents of which are incorporated herein by reference in their entirety. Summary of the Invention
[0009] In various aspects and embodiments, the present disclosure provides the following items 1-62:
[0010] 1. The following amino acid sequence: [Table 1] A tumor antigen peptide (TAP) comprising or consisting of one of the following, or a nucleic acid encoding said TAP.
[0011] 2. The TAP or nucleic acid according to item 1, wherein the TAP comprises one of the sequences defined in SEQ ID NOs: 6, 1 to 5 and 6 to 17.
[0012] 3. A TAP or nucleic acid according to item 1 or 2, which binds to an HLA-A*02:01 molecule and comprises the sequence of SEQ ID NO: 6.
[0013] 4. A TAP or nucleic acid according to item 1 or 2, which binds to an HLA-A*03:01 molecule and comprises the sequence of SEQ ID NO: 1, 11 or 14.
[0014] 5. A TAP or nucleic acid according to item 1 or 2, which binds to an HLA-A*03:02 molecule and comprises the sequence of SEQ ID NO: 3, 5, 7, 16, or 23.
[0015] 6. A TAP or nucleic acid according to item 1 or 2, which binds to an HLA-A*11:01 molecule and comprises the sequence of SEQ ID NO: 9 or 18.
[0016] 7. A TAP or nucleic acid according to item 1 or 2, which binds to an HLA-A*30:01 molecule and comprises the sequence of SEQ ID NO: 19, 20 or 23.
[0017] 8. A TAP or nucleic acid according to item 1 or 2, which binds to an HLA-A*32:01 molecule and comprises the sequence of SEQ ID NO: 8.
[0018] 9. A TAP or nucleic acid according to item 1 or 2, which binds to an HLA-B*07:02 molecule and comprises the sequence of SEQ ID NO: 2 or 21.
[0019] 10. A TAP or nucleic acid according to item 1 or 2, which binds to an HLA-B*13:02 molecule and comprises the sequence of SEQ ID NO: 13.
[0020] 11. A TAP or nucleic acid according to item 1 or 2, which binds to an HLA-B*27:05 molecule and comprises the sequence SEQ ID NO: 4.
[0021] 12. A TAP or nucleic acid according to item 1 or 2, which binds to an HLA-B*52:01 molecule and comprises the sequence of SEQ ID NO: 10, 12 or 15.
[0022] 13. A TAP or nucleic acid according to item 1 or 2, which binds to an HLA-C*06:02 molecule and comprises the sequence SEQ ID NO: 17.
[0023] 14. TAP or nucleic acid according to any one of items 1 to 13, wherein TAP is encoded by a sequence located in a non-protein-coding region of the genome.
[0024] 15. The TAP or nucleic acid according to item 14, wherein the non-protein-coding region of the genome is an untranslated transcribed region (UTR).
[0025] 16. The TAP or nucleic acid according to item 14, wherein the non-protein-coding region of the genome is an intron.
[0026] 17. The TAP or nucleic acid according to item 14, wherein the non-protein-coding region of the genome is an intergenic region.
[0027] 18. The TAP or nucleic acid according to item 14, wherein the non-protein-coding region of the genome is a long non-coding RNA.
[0028] 19. The nucleic acid according to any one of items 1 to 18, wherein the nucleic acid is mRNA.
[0029] 20. The nucleic acid according to any one of items 1 to 18, wherein the nucleic acid is DNA.
[0030] 21. The nucleic acid according to any one of items 1 to 20, wherein the nucleic acid is a component of a viral vector.
[0031] 22. A combination comprising at least two of a TAP or a nucleic acid as defined in any one of items 1 to 21.
[0032] 23. A synthetic long peptide (SLP) comprising at least one of the amino acid sequences defined in item 1.
[0033] 24. A vesicle or particle comprising a TAP, a nucleic acid, a combination, or an SLP according to any one of items 1 to 23.
[0034] 25. The vesicle or particle according to item 24, wherein the vesicle is a lipid nanoparticle (LNP).
[0035] 26. A vesicle or particle according to item 24 or 25, comprising a cationic lipid.
[0036] 27. A composition comprising a TAP, nucleic acid, combination or SLP according to any one of items 1 to 23, or a vesicle or particle according to any one of items 24 to 26, and a pharmaceutically acceptable carrier.
[0037] 28. A vaccine comprising a TAP, nucleic acid, combination or SLP according to any one of items 1 to 23, a vesicle or particle according to any one of items 24 to 26, or a composition according to item 27, and an adjuvant.
[0038] 29. An isolated major histocompatibility complex (MHC) class I molecule, comprising within its peptide-binding groove a TAP according to any one of items 1 to 18.
[0039] 30. The isolated MHC class I molecule of item 29, which is in the form of a multimer.
[0040] 31. The isolated MHC class I molecule according to item 30, wherein the multimer is a tetramer.
[0041] 32. An isolated cell comprising (i) a TAP according to any one of items 1 to 18, (ii) the combination according to item 19, (iii) an SLP according to item 23, or (iv) a vector comprising a nucleotide sequence encoding a TAP according to any one of items 1 to 18, the combination according to item 19, or the SLP according to item 23.
[0042] 33. An isolated cell, which expresses on its surface major histocompatibility complex (MHC) class I molecules, and the MHC class I molecules contain in their peptide-binding groove a TAP according to any one of items 1 to 18 or a combination according to item 19.
[0043] 34. The cell according to item 33, which is an antigen-presenting cell (APC).
[0044] 35. The cell according to item 34, wherein the APC is a dendritic cell.
[0045] 36. A T cell receptor (TCR) that specifically recognizes the isolated MHC class I molecule according to any one of items 29 to 31 and / or the MHC class I molecule expressed on the surface of the cell according to any one of items 32 to 35.
[0046] 37. An antibody or antigen-binding fragment thereof that specifically binds to the isolated MHC class I molecule according to any one of items 29 to 31 and / or the MHC class I molecule expressed on the surface of the cell according to any one of items 33 to 35.
[0047] 38. The antibody or antigen-binding fragment thereof according to item 37, which is a bispecific antibody or antigen-binding fragment thereof.
[0048] 39. The antibody or antigen-binding fragment thereof according to item 38, wherein the bispecific antibody or antigen-binding fragment thereof is a single-chain diabody (scDb).
[0049] 40. The antibody or antigen-binding fragment thereof according to item 38 or 39, wherein the bispecific antibody or antigen-binding fragment thereof also specifically binds to a T cell signaling molecule.
[0050] 41. The antibody or antigen-binding fragment thereof according to item 40, wherein the T cell signaling molecule is a CD3 chain.
[0051] 42. An isolated cell that expresses on its cell surface the TCR according to item 36.
[0052] 43.CD8 + 43. The isolated cell of item 42, which is a T lymphocyte.
[0053] 44. A cell population comprising at least 0.5% isolated cells as defined in item 42 or 43.
[0054] 45. A method of treating cancer (e.g., colorectal cancer) in a subject, comprising administering to the subject an effective amount of: (a) a TAP comprising or consisting of one of the amino acid sequences defined in SEQ ID NOs: 1 to 23 and 25 to 50, or a synthetic long peptide (SLP) comprising at least one of the sequences set forth in SEQ ID NOs: 1 to 23 and 25 to 50; (b) at least one nucleic acid encoding a TAP, a combination thereof, or an SLP as defined in (a); (c) a vesicle or particle comprising a TAP, a combination thereof, or an SLP as defined in (a), or at least one nucleic acid as defined in (b); (d) a composition comprising a TAP, a combination thereof, or an SLP as defined in (a), at least one nucleic acid as defined in (b), or a vesicle or particle as defined in (c), and a pharmaceutically acceptable carrier; (e) a vaccine comprising a TAP, a combination thereof, or an SLP as defined in (a), at least one nucleic acid as defined in (b), a vesicle or particle as defined in (c), or a composition as defined in (d), and an adjuvant; (f) a cell expressing on its cell surface a major histocompatibility complex (MHC) class I molecule comprising a TAP as defined in (a) or a combination thereof in its peptide-binding groove; (g) a cell that expresses on its cell surface a T cell receptor (TCR) that specifically recognizes an MHC class I molecule expressed on the surface of a cell as defined in (f); or (h) A method comprising administering a soluble TCR, antibody, or antigen-binding fragment thereof that specifically binds to an MHC class I molecule expressed on the surface of a cell defined in (f).
[0055] 46. The method according to item 45, wherein the TAP or nucleic acid is as defined in any one of items 1 to 21, the combination is as defined in item 22, the SLP is as defined in item 23, the vesicle is as defined in any one of items 24 to 26, the composition is as defined in item 27, the vaccine is as defined in item 28, the cell is as defined in any one of items 32 to 35, 42 and 43, the cell population is as defined in item 44, and / or the antibody or antigen-binding fragment is as defined in any one of items 37 to 41.
[0056] 47. The method according to item 45 or 46, wherein the CRC is colon cancer.
[0057] 48. The method according to item 45 or 46, wherein the CRC is rectal cancer.
[0058] 49. The method of any one of items 45 to 48, further comprising administering to the subject at least one additional anti-tumor agent or therapy.
[0059] 50. The method of item 49, wherein the at least one additional anti-tumor agent or therapy is a chemotherapeutic agent, immunotherapy, immune checkpoint inhibitor, radiation therapy, or surgery.
[0060] 51. For treating cancer (e.g., colorectal cancer) in a subject, or for manufacturing a medicament for treating cancer (e.g., colorectal cancer) in a subject: (a) a TAP comprising or consisting of one of the amino acid sequences defined in SEQ ID NOs: 1 to 23 and 25 to 50, or a synthetic long peptide (SLP) comprising at least one of the sequences set forth in SEQ ID NOs: 1 to 23 and 25 to 50; (b) at least one nucleic acid encoding a TAP, a combination thereof, or an SLP as defined in (a); (c) a vesicle or particle comprising a TAP, a combination thereof, or an SLP as defined in (a), or at least one nucleic acid as defined in (b); (d) a composition comprising a TAP, a combination thereof, or an SLP as defined in (a), at least one nucleic acid as defined in (b), or a vesicle or particle as defined in (c), and a pharmaceutically acceptable carrier; (e) a vaccine comprising a TAP, a combination thereof, or an SLP as defined in (a), at least one nucleic acid as defined in (b), a vesicle or particle as defined in (c), or a composition as defined in (d), and an adjuvant; (f) a cell expressing on its cell surface a major histocompatibility complex (MHC) class I molecule comprising a TAP as defined in (a) or a combination thereof in its peptide-binding groove; (g) a cell that expresses on its cell surface a T cell receptor (TCR) that specifically recognizes an MHC class I molecule expressed on the surface of a cell as defined in (f); or (h) Use of a soluble TCR, antibody, or antigen-binding fragment thereof, which specifically binds to an MHC class I molecule expressed on the surface of the cell defined in (f).
[0061] 52. Use according to item 55, wherein the TAP or nucleic acid is as defined in any one of items 1 to 21, the combination is as defined in item 22, the SLP is as defined in item 23, the vesicle is as defined in any one of items 24 to 26, the composition is as defined in item 27, the vaccine is as defined in item 28, the cell is as defined in any one of items 32 to 35, 42 and 43, the cell population is as defined in item 44, and / or the antibody or antigen-binding fragment is as defined in any one of items 37 to 41.
[0062] 53. The use according to item 51 or 52, wherein the CRC is colon cancer.
[0063] 54. The use according to item 51 or 52, wherein the CRC is rectal cancer.
[0064] 55. The use according to any one of items 51 to 54, further comprising administering to the subject at least one additional anti-tumor agent or therapy.
[0065] 56. The use according to item 55, wherein the at least one additional anti-tumor agent or therapy is a chemotherapeutic agent, immunotherapy, immune checkpoint inhibitor, radiation therapy, or surgery.
[0066] 57. A drug for use in treating cancer (e.g., colorectal cancer) in a subject, the drug comprising: (a) a TAP comprising or consisting of one of the amino acid sequences defined in SEQ ID NOs: 1 to 23 and 25 to 50, or a synthetic long peptide (SLP) comprising at least one of the sequences set forth in SEQ ID NOs: 1 to 23 and 25 to 50; (b) at least one nucleic acid encoding a TAP, a combination thereof, or an SLP as defined in (a); (c) a vesicle or particle comprising a TAP, a combination thereof, or an SLP as defined in (a), or at least one nucleic acid as defined in (b); (d) a composition comprising a TAP, a combination thereof, or an SLP as defined in (a), at least one nucleic acid as defined in (b), or a vesicle or particle as defined in (c), and a pharmaceutically acceptable carrier; (e) a vaccine comprising a TAP, a combination thereof, or an SLP as defined in (a), at least one nucleic acid as defined in (b), a vesicle or particle as defined in (c), or a composition as defined in (d), and an adjuvant; (f) a cell expressing on its cell surface a major histocompatibility complex (MHC) class I molecule comprising a TAP as defined in (a) or a combination thereof in its peptide-binding groove; (g) a cell that expresses on its cell surface a T cell receptor (TCR) that specifically recognizes an MHC class I molecule expressed on the surface of a cell as defined in (f); or (h) A drug that is a soluble TCR, an antibody, or an antigen-binding fragment thereof that specifically binds to an MHC class I molecule expressed on the surface of a cell as defined in (f).
[0067] 58. The agent for use according to item 61, wherein the TAP or nucleic acid is as defined in any one of items 1 to 21, the combination is as defined in item 22, the SLP is as defined in item 23, the vesicle is as defined in any one of items 24 to 26, the composition is as defined in item 27, the vaccine is as defined in item 28, the cell is as defined in any one of items 32 to 35, 42 and 43, the cell population is as defined in item 44, and / or the antibody or antigen-binding fragment is as defined in any one of items 37 to 41.
[0068] 59. The agent for use according to item 57 or 58, wherein the CRC is colon cancer.
[0069] 60. The agent for use according to item 57 or 58, wherein the CRC is rectal cancer.
[0070] 61. The agent for use according to any one of items 57 to 60, further comprising administering to the subject at least one additional anti-tumor agent or therapy.
[0071] 62. The agent for use according to item 61, wherein the at least one additional anti-tumor agent or therapy is a chemotherapeutic agent, immunotherapy, immune checkpoint inhibitor, radiation therapy, or surgery.
[0072] Other objects, advantages and features of the present disclosure will become more apparent upon reading the following non-restrictive description of specific embodiments thereof, given by way of example only with reference to the accompanying drawings.
[0073] In the accompanying drawings: [Brief explanation of the drawings]
[0074] [Figure 1]This figure shows a proteogenomics workflow for tumor-specific antigen (TSA) discovery in both colorectal cancer (CRC)-derived cell lines and primary tumor samples. Samples generated from CRC- and normal intestinal-derived cell lines, as well as matched primary tumor / normal adjacent tissue (NAT) biopsies from six individuals, were all processed for both RNA sequencing and major histocompatibility complex class I (MHC-I) immunoprecipitation (IP). RNA sequencing data were used for both transcriptome characterization of the samples and generation of a customized global cancer proteome database. For each sample, MHC-I-associated peptides (MAPs) isolated by IP were identified by LC-MS / MS using the respective databases. After validating both the identification and tumor specificity of TSA candidates, their therapeutic potential was assessed by predicting both their immunogenicity and intertumor distribution. Figure created by BioRender.com. [Figure 2]Figures 2A-E show transcriptome profiles of CRC biopsies from primary tumors and normal adjacent tissues. Figure 2A: Principal component analysis (PCA) of the top 500 different genes in each tumor / NAT sample after paired-end RNA-seq and gene read count normalization by DESeq2. MSI tissues (determined by MSISensor) are circled. Figure 2B: GO term analysis of up- / down-regulated genes in CRC tissues compared to adjacent NATs. Genes submitted for GO term analysis were those with |log2FC| > 1 and found to be differentially regulated in all samples using TPM-normalized values. Figure 2C: Bar graph showing the mean ESTIMATE immunoscores for MSS NATs, MSI NATs, MSS CRCs, and MSI CRCs, with standard deviations indicated. Figure 2D: Stacked bar graph showing the average proportion of the transcriptome attributable to five distinct transcript biotypes in NAT versus CRC samples. The difference in the proportion of non-coding transcripts is statistically significant between NAT and CRC (p=0.0156). Figure 2E: Scatter plot showing non-coding RNA transcript expression (left), SNV counts (center), and INDEL counts (right) in MSS and MSI CRC tissues as determined by SNPEff genome annotation, with bar graphs of the mean and standard error. [Figure 3]Figures 3A-E show the results of immunopeptidome analysis of CRC-derived cell lines and tissues. Figure 3A: Upper panel: Stacked bar graph showing the number of unique peptides identified in CRC cell lines, with a horizontal line indicating the average number of MAPs per cell line. Lower panel: Scatter plot showing the correlation between the number of unique MAPs identified in each cell line and MHC I presentation on the cell surface (Pearson's r = 0.96). Figure 3B: Stacked bar graph showing the number of unique peptides identified in primary tissue samples, with a horizontal line indicating the average number of MAPs per tissue sample. In Figures 3A and 3B, "All peptides" indicates the number of peptides identified with a 5% FDR, while "MHC I peptides" indicates the number of peptides identified with a corresponding peptide score, 8-11 amino acid length, and a rank-eluted ligand threshold ≤ 2% using netpanMHC4.1b prediction. Figure 3C: Bar graph showing the percentage of unique MAPs predicted to bind a given HLA allele in each sample using elution-rank prediction. Figure 3D: GO term analysis of MAP source genes for CRC-derived cell lines and primary tissues. In this analysis, only source genes shared by four or more tissues were included for each tissue. Figure 3E, left panel: Stacked bar graphs showing the proportion of MAPs in each tissue sample derived from protein-coding, hypervariable genes (immunoglobulin or TCR), or non-coding transcripts, or unannotated transcripts. Right panel: Stacked bar graphs showing the proportion of non-coding MAPs derived from processed transcripts, retained introns, non-stop decay products, nonsense-mediated decay products, lncRNAs, or unannotated transcripts. [Figure 4]Figures 4A-E show that novel TSAs identified in CRCs are primarily derived from non-coding regions, while the majority of TAAs are derived from exons. Figure 4A: Bar graph showing the number of TSAs identified per sample, with the average number of TSAs per tissue sample indicated by a horizontal line. Figure 4B: Stacked pie chart identifying the genomic origin of TSAs in the inner circle and identifying what proportion of TSAs are mutated (mTSA) or aberrantly expressed (aeTSA). The outer circle indicates what proportion of TSAs are derived from coding or non-coding sequences. Figure 4C: Bar graph showing the number of TAAs identified per sample. Figure 4D: Stacked pie chart identifying the genomic origin of TAAs in the inner circle and identifying what proportion of TAAs are canonical or non-canonical. The outer circle indicates what proportion of TAAs are derived from coding or non-coding sequences. Figure 4E: Heatmap showing the presence or absence of putative TSAs and TAAs in two previous publications on immunopeptidomics of CRC (8, 15) as well as in the IEDB and HLA Ligand Atlas (all tissues and colon tissue only). [Figure 5] A-B show the RNA expression profiles of putative TSAs and TAAs. Figure 5A: MA plot showing the log2FC of transcripts in TPM, CRC, compared with the matched NAT on the y-axis, and the mean average expression in a given tissue sample (average of CRC and NAT). Highlighted points indicate the source transcripts of putative TAAs and TSAs. Both S4 and S5 plots have canonical TAA points that are not visible because they overlap with another canonical TAA source transcript. Figure 5B: Heatmap of the mean RNA expression in log(rphm+1) of aeTSA-coding sequences and TAA-coding sequences (separated as canonical TAAs (canTAAs) and non-canonical TAAs (non-canTAAs)) in normal tissues and pooled TEC samples from the Genotype Tissue Expression (GTEx) Portal. MHC-low tissues include those from brain, nerve, and testis, which have been shown to lowly express MHC I. Black outline indicates average RNA expression above 8.55 rphm. [Figure 6]Figures 6A-C show validation of TSA and TAA candidates. Figure 6A: Heatmap showing the log(rphm+1) mean RNA expression of TSAs and TAAs in 151 TCGA COAD samples. The percentage of TCGA COAD samples expressing TSA and TAA sequences at least 10-fold higher than the log-transformed (log(rphm+1)) mean expression of pooled GTEx and mTEC samples is shown on the left. Figure 6B: rEpitope immunogenicity scores for various groupings of validated TSAs and TAAs compared to putative non-immunogenic thymic peptides reported in Adamopoulou et al. 2013. rEpitope suggests a threshold of immunogenicity for MHC I peptides (0.36), indicated by the dashed line. Figure 6C: Predicted incidence of tumor antigen-binding MHC class I alleles in the US population (IEDB). [Figure 7] 1 shows an upset plot showing the number of HLA alleles unique to a sample, specifically a given intersection of MAPs that are unique to a given sample or uniquely shared by two samples. [Figure 8]Figures 8A-8D show transcriptome profiles of CRC-derived cell lines, ssGSEA analysis of immune infiltration in CRC tissues, and mutational profiles of all samples. Figure 8A: Principal component analysis (PCA) of the top 500 diverse genes in CRC-derived cell lines and one normal intestinal cell line (HIEC-6) after paired-end RNA-seq and gene read count normalization by DESeq2. Known MSI cell lines are circled. Figure 8B: ssGSEA analysis of immune infiltration in tumors and matched NATs using genes described in Danaher et al. 2017 and the GSVA R program (https: / / github.com / rcastelo / GSVA). Figure 8C: Scatter plots showing SNV and INDEL counts for MSS and MSI CRC-derived cell lines as determined by SNPEff genome annotation, with bar graphs of mean and standard error. (Figure 8D) Scatter plot showing SNV and INDEL counts for all MSS and MSI samples (cell lines and tissues) as determined by SNPEff genome annotation, with bar graphs of the mean and standard error. The difference in the number of INDEL mutations between MSS and MSI samples is statistically significant (p=0.00235). [Figure 9]Figures 9A and 9B show the results of GO term analysis of MSI and MSS primary tissue samples. Figure 9A: GO term analysis of up- / down-regulated genes in MSI tumors compared to adjacent NATs. Genes used for GO term analysis had a |log2FC| > 1 compared to the respective NATs using TPM-normalized values and were found to be uniquely differentially expressed in both MSI tumor samples (i.e., genes that were up- / down-regulated only in both MSI tumors but not in either MSS tumor). Figure 9B: GO term analysis of up- / down-regulated genes in MSS tumors compared to their NATs. Genes used for GO term analysis had a |log2FC| > 1 compared to the respective NATs using TPM-normalized values and were found to be uniquely differentially expressed in three or more MSI tumor samples (i.e., genes that were up- / down-regulated in at least three MSS tumors but not in the MSI tumor). [Figure 10] Figures 10A-D provide an overview of unique and shared MAPs in CRC-derived cell lines and CRC / NAT tissue samples. Figure 10A: Venn diagram showing MAP overlap in the MHC I immunopeptidome of four CRC-derived cell lines. Figure 10B: Venn diagram showing MAP overlap in the MHC I immunopeptidome of six primary tissue samples. Figure 10C: UpsetR plot showing the number of MAPs unique to a sample, specifically a given intersection of MAPs that are unique to a given sample or uniquely shared by two samples. Figure 10D: Heatmap showing the number of shared MAPs between any two cell lines or tissue samples. [Figure 11]Figures 11A-E provide an overview of unique and shared MAP source genes in CRC-derived cell lines and CRC / NAT tissue samples. Figure 11A: Top panel: Bar graph showing the number of unique source genes identified per sample. Bottom panel: Scatter plot showing the correlation between the number of unique MAPs identified in each sample and the corresponding number of unique source genes (Pearson's r = 0.99). Source genes were identified for peptides from the coding sequence, and any peptides that mapped to more than one source gene were excluded. Figure 11B: UpsetR plot showing the number of source genes unique to a given sample or intersection, specifically, uniquely shared by two samples. Figure 11C: Venn diagram showing source gene overlap in the MHC I immunopeptidomes of four CRC-derived cell lines. Figure 11D: Venn diagram showing source gene overlap in the MHC I immunopeptidomes of six primary tissue samples. Figure 11E: Heat map showing the number of shared source genes between any two cell lines or tissue samples. [Figure 12] Scatter plot showing the correlation between the number of unique MAPs identified in each sample and the number of TSAs identified and validated (Pearson's r=0.76). [Figure 13] Figure 6A shows the immunogenicity scores of the TSAs and TAAs described herein. The rEpitope immunogenicity scores of various groupings of validated TSAs and all TAAs are compared to the putative non-immunogenic thymic peptides reported in Adamopoulou et al. (2013). rEpitope suggests a threshold of immunogenicity for MHC I peptides (0.36), indicated by the dashed line. This figure differs from Figure 6A in that it includes all TAAs reported in this study, not just the nine selected for validation. DETAILED DESCRIPTION OF THE INVENTION
[0075] In the context of describing the technology (particularly in the context of the claims which follow), the use of the terms "a," "an," and "the" and similar referents are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context.
[0076] The terms "comprising," "having," "including," and "containing" are to be construed as open-ended terms (i.e., meaning "including, but not limited to") unless otherwise noted.
[0077] All methods described herein can be performed in any suitable order unless otherwise indicated herein or clearly contradicted by context.
[0078] Any and all examples provided herein, or the use of exemplary language ("for example," "etc.") are intended merely to better illustrate embodiments of the claimed technology and do not pose a limitation on scope unless otherwise claimed.
[0079] No language in the specification should be construed as indicating any non-claimed element as essential to the practice of embodiments of the claimed technology.
[0080] As used herein, the term "about" has its ordinary meaning. The term "about" is used to indicate that a value includes the inherent variation for error of the device or method being employed to determine the value, or encompasses values that approximate the recited value, e.g., within 10% of the recited value (or range of values).
[0081] The recitation of ranges of values herein, unless otherwise stated herein, is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, and each separate value is incorporated herein as if it were individually recited herein. Every subset of values within a range is also incorporated herein as if it were individually recited herein.
[0082] Where features or aspects of the present disclosure are described in terms of a Markush group or list of alternatives, those skilled in the art will recognize that the present disclosure is also thereby described in terms of any individual component or subgroup of components of the Markush group or list of alternatives.
[0083] Unless specifically defined otherwise, all technical and scientific terms used herein should be understood to have the same meaning as commonly understood by one of ordinary skill in the art (e.g., in stem cell biology, cell culture, molecular genetics, immunology, immunohistochemistry, protein chemistry, and biochemistry).
[0084] Unless otherwise indicated, the recombinant protein, cell culture, and immunological techniques utilized in this disclosure are standard procedures, well known to those skilled in the art. Such techniques are described in J. Perbal, A Practical Guide to Molecular Cloning, John Wiley and Sons (1984), J. Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Laboratory Press (1989), T.A. Brown (editor), Essential Molecular Biology: A Practical Approach, Volumes 1 and 2, IRL Press (1991), D.M.G. Lover and B.D.H. Memes (editors), DNA Cloning: A Practical Approach, Volumes 1-4, IRL Press (1995 and 1996), and F.M.A. Usubel et al. (editors), Current Protocols in Molecular Biology, Greene Pub. Associates and Wiley-Interscience (1988, including all revisions to date), Ed. Harlow and David Lane (editors), Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory, (1988), and J.E. Coligan et al. al. (eds.) Current Protocols in Immunology, John Wiley & Sons (including all current editions), and other sources.
[0085] In the study described herein, the inventors used a proteogenomics-based approach to identify TSA and TAA candidates from CRC cell lines and CRC specimens. Most of these TSAs were derived from aberrantly expressed, non-mutated genomic sequences that are not expressed in normal tissues, such as non-exonic sequences (e.g., intronic and intergenic sequences). The novel CRC TSA and TAA candidates identified herein may be useful, for example, for CRC T cell-based immunotherapy and vaccines.
[0086] Thus, in one aspect, the present disclosure relates to tumor antigen peptides (TAPs) (or tumor-specific peptides), more specifically, isolated TAPs, such as CRC TAPs, that comprise or consist of one of the following amino acid sequences: [Table 2]
[0087] In another aspect, the disclosure further relates to the use of a TAP (e.g., an isolated TAP) comprising or consisting of one of the following amino acid sequences for the treatment of cancer, more particularly CRC: [Table 3]
[0088] In one embodiment, TAP comprises or consists of one of the following amino acid sequences: SEQ ID NOs: 1-23. In one embodiment, TAP comprises or consists of one of the following amino acid sequences: SEQ ID NOs: 1-17, 22, 25, 27, 29, 31, 33, 38, 40, 41, and 46. In a further embodiment, TAP comprises or consists of one of the following amino acid sequences: SEQ ID NOs: 1-17, 22, and 25. In a further embodiment, TAP comprises or consists of one of the following amino acid sequences: SEQ ID NOs: 1-17 and 22.
[0089] Generally, peptides such as tumor antigen peptides (TAPs) presented in the context of HLA class I are about 7 or 8 to about 15, or preferably 8 to 14, amino acid residues in length. In some embodiments of the methods of the present disclosure, longer peptides comprising a TAP sequence as defined herein are artificially loaded into cells, such as antigen-presenting cells (APCs), where they are processed by the cells and the TAP is presented by MHC class I molecules on the surface of the APCs. In this method, peptides / polypeptides longer than 15 amino acid residues can be loaded into APCs and processed by proteases in the APC cytoplasm to provide the corresponding TAP as defined herein for presentation. In some embodiments, the precursor peptides / polypeptides used to generate the TAPs defined herein are, for example, 1000, 500, 400, 300, 200, 150, 100, 75, 50, 45, 40, 35, 30, 25, 20, or 15 amino acids or less. Thus, all methods and processes using TAPs described herein include the use of longer peptides or polypeptides (including naturally occurring proteins), i.e., tumor antigen precursor peptides / polypeptides, to induce the presentation of the "final" 8-14 TAPs after processing by cells (APCs). In some embodiments, the TAPs described herein are approximately 8-14, 8-13, or 8-12 amino acids in length (e.g., 8, 9, 10, 11, 12, or 13 amino acids in length), small enough to fit directly onto an HLA class I molecule. In one embodiment, the TAP contains 20 or fewer amino acids, preferably 15 or fewer amino acids, and more preferably 14 or fewer amino acids. In one embodiment, the TAP contains at least 7 amino acids, preferably at least 8 or fewer amino acids, and more preferably at least 9 amino acids.
[0090] As used herein, the term "amino acid" includes both L- and D-isomers of the naturally occurring amino acids as well as other amino acids (e.g., naturally occurring amino acids, non-naturally occurring amino acids, amino acids not encoded by nucleic acid sequences, etc.) used in peptide chemistry to prepare synthetic analogs of TAP. Examples of naturally occurring amino acids are glycine, alanine, valine, leucine, isoleucine, serine, threonine, etc. Other amino acids include, for example, non-genetically encoded forms of amino acids, as well as conservative substitutions for L-amino acids. Naturally occurring non-genetically encoded amino acids include, for example, beta-alanine, 3-aminopropionic acid, 2,3-diaminopropionic acid, alpha-aminoisobutyric acid (Aib), 4-amino-butyric acid, N-methylglycine (sarcosine), hydroxyproline, ornithine (e.g., L-ornithine), citrulline, t-butylalanine, t-butylglycine, N-methylisoleucine, phenylglycine, cyclohexylalanine, norleucine (Nle), norvaline, 2-naphthylalanine, pyridylalanine, 3-benzothienylalanine, 4-chlorophenylalanine, 2-fluorophenylalanine, and phenylalanine. Examples of amino acids include alanine, 3-fluorophenylalanine, 4-fluorophenylalanine, penicillamine, 1,2,3,4-tetrahydro-isoquinoline-3-carboxylic acid, beta-2-thienylalanine, methionine sulfoxide, L-homoarginine (Hoarg), N-acetyllysine, 2-aminobutyric acid, 2-aminobutyric acid, 2,4-diaminobutyric acid (D- or L-), p-aminophenylalanine, N-methylvaline, homocysteine, homoserine (HoSer), cysteic acid, epsilon-aminohexanoic acid, delta-aminovaleric acid, and 2,3-diaminobutyric acid (D- or L-). These amino acids are well known in the field of biochemistry / peptide chemistry. In one embodiment, TAP contains only naturally occurring amino acids.
[0091] In embodiments, the TAPs described herein include peptides with altered sequences containing functionally equivalent amino acid residue substitutions compared to the sequences described herein. For example, one or more amino acid residues within a sequence can be substituted with another amino acid of similar polarity (having similar physicochemical properties) that acts as a functional equivalent, resulting in a silent alteration. Substitutes for amino acids within a sequence can be selected from other members of the class to which the amino acid belongs. For example, positively charged (basic) amino acids include arginine, lysine, and histidine (as well as homoarginine and ornithine). Nonpolar (hydrophobic) amino acids include leucine, isoleucine, alanine, phenylalanine, valine, proline, tryptophan, and methionine. Uncharged polar amino acids include serine, threonine, cysteine, tyrosine, asparagine, and glutamine. Negatively charged (acidic) amino acids include glutamic acid and aspartic acid. The amino acid glycine can be included in either the nonpolar amino acid family or the uncharged (neutral) polar amino acid family. Substitutions made within a family of amino acids are generally understood to be conservative substitutions. The TAPs described herein can contain any L-amino acid, any D-amino acid, or a mixture of L- and D-amino acids. In one embodiment, the TAPs described herein contain all L-amino acids.
[0092] In one embodiment, in the sequence of TAP comprising or consisting of one of the sequences of SEQ ID NOs: 1 to 23 and 25 to 50, amino acid residues that do not substantially contribute to interaction with the T cell receptor can be modified by replacing them with other amino acids that do not substantially affect T cell responsiveness and do not eliminate binding to the relevant MHC.
[0093] TAP may also be N- and / or C-terminally capped or modified to prevent degradation, increase stability, affinity, and / or uptake. Thus, in another aspect, the present disclosure provides a compound of formula Z 1 -XZ 2wherein X is a TAP comprising or consisting of one of the amino acid sequences of SEQ ID NOs: 1 to 23 and 25 to 50.
[0094] In one embodiment, the amino terminal residue of TAP (i.e., the N-terminal free amino group) is, for example, a moiety / chemical group (Z 1 ) is modified (e.g., for protection against degradation). 1 may be a straight or branched chain alkyl group of 1 to 8 carbons, or an acyl group (R—CO—), where R is a hydrophobic moiety (e.g., acetyl, propionyl, butanyl, isopropionyl, or iso-butanyl), or an aroyl group (Ar—CO—), where Ar is an aryl group. In one embodiment, the acyl group is a C1-C 16 Or C3~C 16 It is an acyl group (linear or branched, saturated or unsaturated), and in a further embodiment is a saturated C1-C6 acyl group (linear or branched) or an unsaturated C3-C6 acyl group (linear or branched), for example, an acetyl group (CH3-CO-, Ac). In one embodiment, Z 1 The carboxy-terminal residue of TAP (i.e., the free carboxy group at the C-terminus of TAP) may be modified (e.g., for protection against degradation) by, for example, amidation (replacement of an OH group with an NH2 group), and thus, in such cases, Z 2 is an NH group. In one embodiment, Z 2 may be a hydroxamate group, a nitrile group, an amide (primary, secondary, or tertiary) group, an aliphatic amine of 1 to 10 carbons such as methylamine, iso-butylamine, iso-valerylamine, or cyclohexylamine, an aromatic or arylalkyl amine such as aniline, naphthylamine, benzylamine, cinnamylamine, or phenylethylamine, an alcohol, or CHOH. 2 In one embodiment, TAP comprises one of the amino acid sequences of SEQ ID NOs: 1 to 24 and 26 to 50. In one embodiment, TAP consists of one of the amino acid sequences of SEQ ID NOs: 1 to 24 and 26 to 50, i.e., Z1 and Z 2 does not exist.
[0095] In another aspect, the disclosure provides a TAP that binds to HLA-A*02:01 molecules, comprising or consisting of the sequence of SEQ ID NO: 6. Because HLA alleles exhibit promiscuity (particular HLA alleles present similar epitopes), the above-identified TAP may additionally bind to HLA-A*02:05, HLA-A*02:06, and / or HLA-A*02:07 molecules.
[0096] In another aspect, the disclosure provides a TAP that binds to HLA-A*03:01 molecules, comprising or consisting of the sequence of SEQ ID NO: 1, 11, 14, 38, 45, or 46, preferably SEQ ID NO: 1, 11, or 14. Due to HLA allele promiscuity (particular HLA alleles present similar epitopes), the above identified TAPs may further bind to HLA-A*03:02, HLA-A*11:01, or HLA-A*30:01 molecules.
[0097] In another aspect, the disclosure provides a TAP that binds to HLA-A*03:02 molecules, comprising or consisting of the sequence of SEQ ID NO: 3, 5, 7, 16, 23, 31, 32, 33, or 34, preferably SEQ ID NO: 3, 5, 7, 16, or 23. Due to HLA allele promiscuity (particular HLA alleles present similar epitopes), the above identified TAPs may additionally bind to HLA-A*03:01, HLA-A*11:01, or HLA-A*30:01 molecules.
[0098] In another aspect, the present disclosure provides a TAP that binds to HLA-A*11:01 molecules, comprising or consisting of the sequence of SEQ ID NO: 9, 18, 33, or 34, preferably SEQ ID NO: 9 or 18. Due to HLA allele promiscuity (particular HLA alleles present similar epitopes), the above identified TAPs may further bind to HLA-A*03:01, HLA-A*03:02, HLA-A*31:01, and / or HLA-A*68:01 molecules.
[0099] In another aspect, the present disclosure provides a TAP that binds to HLA-A*23:01 molecules, comprising or consisting of the sequence of SEQ ID NO: 29. Due to HLA allele promiscuity (particular HLA alleles present similar epitopes), the above identified TAP may also bind to HLA-A*24:02 molecules.
[0100] In another aspect, the disclosure provides a TAP that binds to HLA-A*24:02 molecules, comprising or consisting of the sequence of SEQ ID NO: 26, 27, or 29. Due to HLA allele promiscuity (particular HLA alleles present similar epitopes), the above identified TAPs may also bind to HLA-A*23:01 molecules.
[0101] In another aspect, the disclosure provides a TAP that binds to HLA-A*30:01 molecules, comprising or consisting of the sequence of SEQ ID NO: 19, 20, or 23. Due to HLA allele promiscuity (particular HLA alleles present similar epitopes), the above-identified TAPs may additionally bind to HLA-A*30:02 and / or HLA-B*15:02 molecules.
[0102] In another aspect, the present disclosure provides a TAP that binds to HLA-A*32:01 molecules, comprising or consisting of the sequence of SEQ ID NO: 8, 37, or 39, preferably SEQ ID NO: 8. Due to HLA allele promiscuity (particular HLA alleles present similar epitopes), the above identified TAPs may additionally bind to HLA-B*57:01 and / or HLA-B*58:01 molecules.
[0103] In another aspect, the present disclosure provides a TAP that binds to HLA-B*07:02 molecules, comprising or consisting of the sequence of SEQ ID NO: 2, 21, 24, or 50, preferably SEQ ID NO: 2 or 21. Due to HLA allele promiscuity (particular HLA alleles present similar epitopes), the above identified TAPs may additionally bind to HLA-B*35:02, HLA-B*35:03, HLA-B*55:01, and / or HLA-B*56:01 molecules.
[0104] In another aspect, the present disclosure provides a TAP that binds to HLA-B*13:02 molecules, comprising or consisting of the sequence of SEQ ID NO: 13 or 48, preferably SEQ ID NO: 13.
[0105] In another aspect, the disclosure provides a TAP that binds to HLA-B*18:01 molecules, comprising or consisting of the sequence of SEQ ID NO: 25. Due to HLA allele promiscuity (particular HLA alleles present similar epitopes), the above-identified TAP may further bind to HLA-B*40:01, HLA-B*44:02, HLA-B*44:03, and / or HLA-B*45:01 molecules.
[0106] In another aspect, the present disclosure provides a TAP that binds to HLA-B*27:05 molecules, comprising or consisting of the sequence of SEQ ID NO: 4. Due to HLA allele promiscuity (particular HLA alleles present similar epitopes), the above identified TAP may also bind to HLA-B*27:02.
[0107] In another aspect, the present disclosure provides a TAP that binds to HLA-B*52:01 molecules, comprising or consisting of the sequence of SEQ ID NO: 10, 12, 15, or 43, preferably SEQ ID NO: 10, 12, or 15. Due to HLA allele promiscuity (particular HLA alleles present similar epitopes), the above identified TAPs may further bind to HLA-B*51:01.
[0108] In another aspect, the present disclosure provides a TAP that binds to HLA-C*06:02 molecules, comprising or consisting of the sequence of SEQ ID NO: 17 or 44, preferably SEQ ID NO: 17. Due to HLA allele promiscuity (particular HLA alleles present similar epitopes), the above identified TAPs may further bind to HLA-B*27:02, HLA-C*07:01, and / or HLA-C*07:02 molecules.
[0109] In another aspect, the disclosure provides a TAP that binds to HLA-C*12:02 molecules, comprising or consisting of the sequence of SEQ ID NO: 47 or 49. Due to HLA allele promiscuity (particular HLA alleles present similar epitopes), the above-identified TAP may further bind to HLA-B*46:01, HLA-C*03:02, HLA-C*03:03, HLA-C*03:04, HLA-C*08:01, HLA-C*12:03, HLA-C*15:02, and / or HLA-C*16:01 molecules.
[0110] In one embodiment, TAP is encoded by a sequence located in a non-protein-coding region of the genome. In one embodiment, TAP is encoded by a sequence located in an untranslated transcribed region (UTR), i.e., the 3'-UTR or 5'-UTR region. In another embodiment, TAP is encoded by a sequence located in an intron. In another embodiment, TAP is encoded by a sequence located in an intergenic region. In another embodiment, TAP is encoded by a sequence located in an exon and results from a frameshift.
[0111] The TAP of the present disclosure can be produced by expression in a host cell containing a nucleic acid encoding TAP (recombinant expression) or by chemical synthesis (e.g., solid-phase peptide synthesis). Peptides can be readily synthesized by manual and / or automated solid-phase procedures well known in the art. Suitable synthesis can be carried out, for example, by utilizing the "T-boc" or "Fmoc" procedure. Techniques and procedures for solid-phase synthesis are described, for example, in "Solid Phase Peptide Synthesis: A Practical Approach," E. Atherton and R.C. Sheppard, IRL, Oxford University Press, 1989. Alternatively, the MiHA peptide can be used as described, for example, in Liu et al., Tetrahedron Lett.37:933-936,1996, Baca et al., J.Am.Chem.Soc.117:1881-1887,1995, Tam et al., Int.J.Peptide Protein Res.45:209-216,1995, Schnolzer and Kent,Science 256:221-225,1992, Liu and Tam,J.Am.Chem.Soc.116:4149-4153,1994,Liu and Tam,Proc.Natl.Acad.Sci.USA 91:6584-6588,1994, and Yamashiro and Li,Int.J.Peptide Protein Res. 31:322-334, 1988). Other methods useful for synthesizing TAP are described in Nakagawa et al., J. Am. Chem. Soc. 107:7087-7092, 1985. In one embodiment, TAP is chemically synthesized (synthetic peptide). Another embodiment of the present disclosure relates to non-naturally occurring peptides, which consist of or consist essentially of the amino acid sequences defined herein and have been synthetically produced (e.g., synthesized) as pharmaceutically acceptable salts. Salts of TAP according to the present disclosure are substantially different from the peptides in their state(s) in vivo, since peptides produced in vivo do not have salts.Non-natural salt forms of peptides can modulate the solubility of the peptides, particularly in the context of pharmaceutical compositions comprising the peptides, e.g., peptide vaccines disclosed herein. Preferably, the salts are pharmaceutically acceptable salts of the peptides.
[0112] In one embodiment, the TAPs described herein are isolated or substantially pure. A compound, such as a peptide or nucleic acid, is "isolated" or "substantially pure" when it is separated from components present in the molecule's natural environment or from naturally occurring source macromolecules (including, for example, other nucleic acids, proteins, lipids, sugars, etc.). Typically, a compound is substantially pure when it is at least 60%, more usually 75%, 80%, or 85%, preferably greater than 90%, and more preferably greater than 95%, by weight of the total material in a sample. Thus, for example, a polypeptide chemically synthesized or produced by recombinant techniques will generally be substantially free of its naturally associated components, e.g., components of its source macromolecules. A nucleic acid molecule is substantially pure when it is not immediately contiguous (i.e., not covalently linked) to coding sequences with which it is normally contiguous in the naturally occurring genome of the organism from which the nucleic acid is derived. A substantially pure compound can be obtained, for example, by extraction from a natural source, by expression of a recombinant nucleic acid molecule encoding the peptide compound, or by chemical synthesis. Purity can be measured using any suitable method, such as column chromatography, gel electrophoresis, HPLC, etc. In one embodiment, the TAP is in solution. In another embodiment, the TAP is in solid form, e.g., lyophilized.
[0113] In one embodiment, TAP is encoded by a sequence located in the non-protein-coding region of genome.In one embodiment, TAP is encoded by a sequence located in intergenic region.In another embodiment, TAP is encoded by non-coding RNA (ncRNA).
[0114] In another aspect, the present disclosure further provides synthetic long peptides (SLPs) comprising at least one of the TAPs described herein. In one embodiment, the SLP comprises at least two TAPs, at least one of which is a TAP described herein. In one embodiment, the SLP comprises at least two, three, four, or five of the TAPs described herein. In one embodiment, the SLP comprises at least one of the TAPs described herein linked to one or more amino acid sequences or domains that confer a desired property to the SLP, such as a sequence or domain that stabilizes the SLP and / or improves processing and presentation by MHC molecules, e.g., a sequence containing a motif cleavable by a cellular protease such as a cathepsin. In another embodiment, the SLP comprises at least one of the TAPs described herein and a TAP that binds to an MHC class II molecule. The TAPs may be directly connected to each other or indirectly connected via a linker, such as a short amino acid linker. In embodiments, the linker comprises about 4 to about 20 amino acids, or about 4 to about 15 amino acids, e.g., 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids. In one embodiment, the linker comprises glycine residues, serine residues, proline residues, threonine residues, or a mixture thereof. In one embodiment, the SLP has a length of 200, 150, 100, 90, 80, 70, 60, or 50 amino acids or less. In further embodiments, the SLP has a length of 20 to 50, 45, or 40 amino acids, e.g., 20 or 25 amino acids to 30, 35, or 40 amino acids. As used herein, "synthetic" refers to a peptide or nucleic molecule that has not been isolated from its natural source, e.g., produced by recombinant technology or using chemical synthesis.
[0115] In another aspect, the disclosure further provides (e.g., isolated) nucleic acids encoding a TAP or tumor antigen precursor peptide or SLP described herein. In one embodiment, the nucleic acid comprises about 21 nucleotides to about 45 nucleotides, about 24 to about 45 nucleotides, e.g., 24, 27, 30, 33, 36, 39, 42, or 45 nucleotides.
[0116] In one embodiment, a nucleic acid (DNA, RNA) encoding TAP of the present disclosure comprises any one of the sequences set forth in SEQ ID NOs: 51-73 and 75-100, or SEQ ID NOs: 51-67, 72, 75, 77, 79, 81, 83, 88, 90, 91, and 96, or SEQ ID NOs: 51-67, 72, and 75 (e.g., SEQ ID NOs: 51-67 and 72), or the corresponding RNA sequence (i.e., in which thymine nucleobases (T) are replaced by uracil nucleobases (U)). In one embodiment, the nucleic acid encoding TAP is an mRNA molecule. In one embodiment, the nucleic acid is in solution. In another embodiment, the nucleic acid is in solid form, e.g., lyophilized.
[0117] The nucleic acids of the present disclosure can be used for recombinant expression of the TAPs or SLPs of the present disclosure and can be contained within a vector or plasmid, such as a cloning or expression vector, that can be transfected into a host cell. In one embodiment, the present disclosure provides a cloning, expression, or viral vector, or a plasmid containing a nucleic acid sequence encoding a TAP of the present disclosure. Alternatively, the nucleic acid encoding a TAP of the present disclosure can be integrated into the genome of a host cell. In either case, the host cell expresses the TAP or protein encoded by the nucleic acid. As used herein, the term "host cell" refers not only to the particular subject cell but also to the progeny or potential progeny of such a cell. A host cell can be any prokaryotic (e.g., E. coli) or eukaryotic cell (e.g., insect, yeast, plant, or mammalian) capable of expressing a TAP described herein. A vector or plasmid contains elements necessary for the transcription and translation of an inserted coding sequence and may contain other components, such as resistance genes, cloning sites, etc. Methods well known to those skilled in the art may be used to construct expression vectors containing a peptide or polypeptide coding sequence and appropriate transcriptional and translational control / regulatory elements operably linked thereto. These methods include in vitro recombinant DNA techniques, synthetic techniques, and in vivo genetic recombination. Such techniques are described in Sambrook, et al. (1989) Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press, Plainview, NY, and Ausubel, FM et al. (1989) Current Protocols in Molecular Biology, John Wiley & Sons, New York, NY. "Operably linked" refers to the juxtaposition of components, particularly components that allow the nucleotide sequence to perform its normal function. Thus, a coding sequence operably linked to a regulatory sequence refers to a nucleotide sequence configuration in which the coding sequence can be expressed under the regulatory control, i.e., transcriptional and / or translational control, of the regulatory sequence.As used herein, "regulatory / control region" or "regulatory / control sequence" refers to a non-coding nucleotide sequence involved in regulating the expression of a coding nucleic acid. Thus, the term regulatory region includes promoter sequences, regulatory protein binding sites, upstream activator sequences, and the like. A vector (e.g., an expression vector) may have necessary 5' upstream and 3' downstream regulatory elements for efficient gene transcription and translation in a respective host cell, such as a promoter sequence (e.g., CMV, PGK, and EF-1α promoters), a ribosome recognition and binding TATA box, and a 3' UTR AAUAAA transcription termination sequence. Other suitable promoters include the simian vims 40 (SV40) early promoter, mouse mammary tumor virus (MMTV) promoter, HIV LTR promoter, MoMuLV promoter, avian leukosis virus promoter, EBV immediate early promoter, and constitutive promoter of the Rous sarcoma vims promoter. Human gene promoters may also be used, including, but not limited to, the actin promoter, myosin promoter, hemoglobin promoter, and creatine kinase promoter. In certain embodiments, an inducible promoter is also contemplated as part of a vector expressing TAP. This provides a molecular switch that can turn on or off expression of a polynucleotide sequence of interest. Examples of inducible promoters include, but are not limited to, metallothionine promoters, glucocorticoid promoters, progesterone promoters, or tetracycline promoters. Examples of vectors include plasmids, autonomously replicating sequences, and transposable elements. Additional exemplary vectors include, but are not limited to, plasmids, phagemids, cosmids, artificial chromosomes (e.g., yeast artificial chromosomes (YACs)), bacterial artificial chromosomes (BACs), or P1-derived artificial chromosomes (PACs), bacteriophages (e.g., lambda phage or M13 phage), and animal viruses.Examples of animal virus categories useful as vectors include, but are not limited to, retroviruses (including lentiviruses), adenoviruses, adeno-associated viruses, herpesviruses (e.g., herpes simplex viruses), poxviruses, baculoviruses, papillomaviruses, and papovaviruses (e.g., SV40). Examples of expression vectors include the Lenti-X™ Bicistronic Expression System (Neo) vector (Contech) for expression in mammalian cells, the pClneo vector (Promega), and pLenti4 / V5-DEST™, pLenti6 / V5-DEST™, and pLenti6.2N5-GW / lacZ (Invitrogen) for lentivirus-mediated gene transfer and expression in mammalian cells. The coding sequence for TAP disclosed herein can be ligated into such expression vectors for expression of TAP in mammalian cells.
[0118] In certain embodiments, a nucleic acid encoding a TAP of the present disclosure is provided in a viral vector. The viral vector may be derived from an adenovirus, vaccinia virus, retrovirus, lentivirus, or foamy virus. As used herein, the term "viral vector" refers to a nucleic acid vector construct that contains at least one element of viral origin and has the ability to be packaged into a viral vector particle. The viral vector may contain coding sequences for various proteins described herein in place of non-essential viral genes. In another embodiment, a nucleic acid encoding a TAP of the present disclosure is provided in a self-amplifying or self-replicating RNA (srRNA) vector. The srRNA is derived from a positive-strand RNA virus from which structural proteins have been removed and replaced with a heterologous gene of interest. srRNAs have been successfully derived from flaviviruses, nodamuraviruses, nidoviruses, and alphaviruses, with therapeutic applications providing structural proteins in trans to create single-cycle viral replicon particles (VRPs) (see, e.g., Aliahmad et al. Next generation self-replicating RNA vectors for vaccines and immunotherapies. Cancer Gene Ther (2022). https: / / doi.org / 10.1038 / s41417-022-00435-8). The vectors and / or particles can be used to transfer DNA, RNA, or other nucleic acids into cells, either in vitro or in vivo. Many forms of viral vectors are known in the art.
[0119] In embodiments, a nucleic acid (DNA, RNA) encoding a TAP of the present disclosure is contained within a vesicle or nanoparticle, such as a lipid vesicle (e.g., liposome) or lipid nanoparticle (LNP), or any other suitable vehicle. Thus, in another aspect, the present disclosure provides a vesicle or nanoparticle, such as a lipid vesicle or nanoparticle, comprising a nucleic acid, such as mRNA, encoding one or more of the TAPs described herein.
[0120] The term liposome, as used herein, refers in accordance with its ordinary meaning to a microscopic lipid vesicle composed of a bilayer of phospholipids or any similar amphiphilic lipid (e.g., sphingolipid) encapsulating an internal aqueous medium.
[0121] The term "lipid nanoparticle" refers to a liposome-like structure that may contain one or more lipid bilayer rings surrounding an internal aqueous medium, similar to liposomes, or a micelle-like structure that encapsulates molecules (e.g., nucleic acids) within a non-aqueous core. Lipid nanoparticles typically contain cationic lipids, such as ionizable cationic lipids. Examples of cationic lipids that can be used in LNPs include DOTMA, DOSPA, DOTAP, ePC, DLin-MC3-DMA, C12-200, ALC-0315, cKK-E12, Lipid H (SM-102), OF-Deg-Lin, A2-Iso5-2DC18, 306O i10 , BAME-O16B, TT3, 9A1P9, FTT5, COATSOME® SS-E, COATSOME® SS-EC, COATSOME® SS-OC, and COATSOME® SS-OP (see, e.g., Hou et al., Nature Reviews Materials, volume 6, pages 1078-1094 (2021); Tenchov et al., ACS Nano, 15, 16982-17015 (2021)).
[0122] Liposomes and lipid nanoparticles typically contain lipids, lipid-like materials, and other lipid components, such as polymers, which can improve liposome or nanoparticle properties such as stability, delivery efficacy, tolerability, and biodistribution. These include phospholipids (e.g., phosphatidylcholine, phosphatidylethanolamine, phosphatidylserine, and phosphatidylglycerol) such as 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC), as well as DOPE, sterols (e.g., cholesterol and cholesterol derivatives), 1,2-dimyristoyl-rac-glycero-3-methoxypolyethylene glycol-2000 (PEG), and other lipid-containing polymers.2000 -DMG) and 1,2-distearoyl-rac-glycero-3-methoxypolyethylene glycol-2000 (PEG 2000 PEGylated lipids (PEG lipids) such as PEG-DSG.
[0123] In one embodiment, the lipid nanoparticles according to the present disclosure comprise one or more cationic lipids, such as ionizable cationic lipids. Examples of ionizable cationic lipids include those listed in PCT Publication Nos. 2017 / 061150 and 2019 / 188867, including those commercially available under the trade names COATSOME® SS-E, COATSOME® SS-EC, COATSOME® SS-OC, and COATSOME® SS-OP.
[0124] Nucleic acids (e.g., mRNA) encoding one or more of the TAPs can be modified, for example, to increase stability and / or reduce immunogenicity. For example, the 5' end can be capped to stabilize the molecule and reduce immunogenicity (e.g., as described in U.S. Pat. Nos. 10,519,189 and 10,494,399). One or more nucleosides of the mRNA can be modified or substituted with 1-methylpseudouridine to either increase the stability of the molecule or reduce recognition of the molecule by the innate immune system. Modified nucleoside forms are described in U.S. Pat. No. 9,371,511. Other types of modifications that can be made to mRNA include anti-reverse cap analog (ARCA), 5'-methyl-cytidine triphosphate (m5CTP), N6-methyl-adenosine-5'-triphosphate (m6ATP), 2-thio-uridine triphosphate (s2UTP), pseudouridine triphosphate, N 1These modifications include the incorporation of methylpseudouridine triphosphate or 5-methoxyuridine triphosphate (5moUTP). The mRNA may also include additional modifications to the 5' and / or 3' untranslated regions (UTRs) and polyadenylation (polyA) tails (see, e.g., Kim et al., Molecular & Cellular Toxicology vol. 18, 1 (2022): 1-8). All of these and other modifications to nucleic acids (e.g., mRNAs) encoding TAP are encompassed by the present disclosure.
[0125] In another aspect, the present disclosure provides MHC class I molecules comprising (i.e., presenting or being bound to) one or more of the TAPs of SEQ ID NOs: 1-23 and 25-50, preferably SEQ ID NO: 1-23, e.g., SEQ ID NOs: 1-17 and 22.
[0126] In one embodiment, the MHC class I molecule is an HLA-A2 molecule, and in a further embodiment, an HLA-A*02:01 molecule. In one embodiment, the MHC class I molecule is an HLA-A3 molecule, and in a further embodiment, an HLA-A*03:01 or HLA-A*03:02 molecule. In another embodiment, the MHC class I molecule is an HLA-A11 molecule, and in a further embodiment, an HLA-A*11:01 molecule. In one embodiment, the MHC class I molecule is an HLA-A23 molecule, and in a further embodiment, an HLA-A*23:01 molecule. In one embodiment, the MHC class I molecule is an HLA-A24 molecule, and in a further embodiment, an HLA-A*24:02 molecule. In one embodiment, the MHC class I molecule is an HLA-A30 molecule, and in a further embodiment, an HLA-A*30:01 molecule. In one embodiment, the MHC class I molecule is an HLA-A32 molecule, and in a further embodiment, an HLA-A*32:01 molecule. In another embodiment, the MHC class I molecule is an HLA-B07 molecule, and in a further embodiment, an HLA-B*07:02 molecule. In another embodiment, the MHC class I molecule is an HLA-B13 molecule, and in a further embodiment, an HLA-B*13:02 molecule. In another embodiment, the MHC class I molecule is an HLA-B18 molecule, and in a further embodiment, an HLA-B*18:01 molecule. In another embodiment, the MHC class I molecule is an HLA-B27 molecule, and in a further embodiment, an HLA-B*27:05 molecule. In another embodiment, the MHC class I molecule is an HLA-B52 molecule, and in a further embodiment, an HLA-B*52:01 molecule. In another embodiment, the MHC class I molecule is an HLA-C06 molecule, and in a further embodiment, an HLA-C*06:02 molecule. In another embodiment, the MHC class I molecule is an HLA-C04 molecule, and in a further embodiment, an HLA-C*04:01 molecule. In another embodiment, the MHC class I molecule is an HLA-C12 molecule, and in a further embodiment, an HLA-C*12:02 molecule.
[0127] In one embodiment, TAP (e.g., SEQ ID NOS: 1-23 and 25-50, preferably SEQ ID NOS: 1-23) is non-covalently linked to an MHC class I molecule (i.e., TAP is loaded into the peptide-binding groove / pocket of the MHC class I molecule or non-covalently linked). In another embodiment, TAP is covalently linked / bound to an MHC class I molecule (alpha chain). In such constructs, TAP and the MHC class I molecule (alpha chain) are produced as a synthetic fusion protein, typically with a short (e.g., 5-20 residues, preferably about 8-12, e.g., 10) flexible linker or spacer (e.g., a polyglycine linker). In another aspect, the present disclosure provides nucleic acids encoding fusion proteins comprising TAP as defined herein fused to an MHC class I molecule (alpha chain). In one embodiment, the MHC class I molecule (alpha chain)-peptide complex is multimerized. Thus, in another aspect, the present disclosure provides multimers of MHC class I molecules loaded (covalently or non-covalently) with the TAPs described herein. Such multimers can be attached to a tag, e.g., a fluorescent tag, that allows for detection of the multimer. Numerous strategies have been developed for the production of MHC multimers, including MHC dimers, tetramers, pentamers, octamers, etc. (reviewed in Bakker and Schumacher, Current Opinion in Immunology 2005, 17:428-433). MHC multimers are useful, for example, for the detection and purification of antigen-specific T cells. Thus, in another aspect, the present disclosure provides multimers of CD8 specific MHC class I molecules loaded (covalently or non-covalently) with the TAPs defined herein. + A method for detecting or purifying (isolating, enriching) T lymphocytes is provided, comprising contacting a cell population with a multimer of MHC class I molecules loaded (covalently or non-covalently) with TAP, and detecting CD8 bound by the MHC class I multimer. + and detecting or isolating CD8 T lymphocytes bound by MHC class I multimers. + T lymphocytes can be isolated using known methods, for example, fluorescence activated cell sorting (FACS) or magnetic activated cell sorting (MACS).
[0128] In yet another aspect, the disclosure provides cells (e.g., host cells), and in one embodiment, an isolated cell comprising a nucleic acid, vector, or plasmid of the disclosure (i.e., a nucleic acid or vector encoding one or more TAPs) described herein. In another aspect, the disclosure provides a cell that expresses on its cell surface an MHC class I molecule (e.g., an MHC class I molecule of one of the alleles disclosed above) that is bound to or presents a TAP according to the disclosure. In one embodiment, the host cell is a eukaryotic cell, e.g., a mammalian cell, preferably a human cell, cell line, or immortalized cell. In another embodiment, the cell is an antigen-presenting cell (APC). In one embodiment, the host cell is a primary cell, cell line, or immortalized cell. In another embodiment, the cell is an antigen-presenting cell (APC). Nucleic acids and vectors can be introduced into cells via conventional transformation or transfection techniques. The terms "transformation" and "transfection" refer to techniques for introducing foreign nucleic acid into host cells, including calcium phosphate or calcium chloride co-precipitation, DEAE-dextran-mediated transfection, lipofection, electroporation, microinjection, and viral-mediated transfection. Suitable methods for transforming or transfecting host cells can be found, for example, in Sambrook et al. (supra), and other laboratory manuals. Methods for introducing nucleic acids into mammalian cells in vivo are also known and can be used to deliver the vectors or plasmids of the present disclosure to a subject for gene therapy.
[0129] Cells, such as APCs, can be loaded with one or more TAPs using various methods known in the art. As used herein, "loading cells" with TAP means that RNA or DNA encoding TAP or TAP is transfected into cells, or alternatively, APCs are transformed with a nucleic acid encoding TAP. Cells can also be loaded by contacting them with exogenous TAP, which can directly bind to MHC class I molecules present on the cell surface (e.g., cells pulsed with peptide). TAP can also be fused to a domain or motif (e.g., an endoplasmic reticulum (ER) retrieval signal, a C-terminal Lys-Asp-Glu-Leu sequence (see Wang et al., Eur J Immunol. 2004 Dec;34(12):3582-94)) that promotes its presentation by MHC class I molecules.
[0130] In another aspect, the disclosure provides compositions or combinations / pools of peptides comprising any one or any combination of the TAPs (or nucleic acids encoding said peptide(s)) defined herein. In one embodiment, the composition comprises any combination of the TAPs defined herein (any combination of 2, 3, 4, 5, 6, 7, 8, 9, 10, or more TAPs) or any combination of nucleic acids encoding said TAPs. Compositions comprising any combination / subcombination of the TAPs defined herein are encompassed by the disclosure. In another embodiment, the combination or pool may comprise one or more known tumor antigens.
[0131] Thus, in another aspect, the present disclosure provides compositions comprising any one or any combination of TAP or SLP(s) defined herein (e.g., comprising or consisting of the sequences of SEQ ID NOS: 1-23 and 25-50, preferably SEQ ID NOS: 1-23, e.g., SEQ ID NOS: 1-17 and 22) and cells expressing an MHC class I molecule (e.g., an MHC class I molecule of one of the alleles disclosed above). APCs for use in the present disclosure are not limited to a particular type of cell, and can be any type of cell, including CD8+ This includes professional APCs such as dendritic cells (DCs), Langerhans cells, macrophages, and B cells, which are known to present proteinaceous antigens on their cell surface for recognition by T lymphocytes. For example, APCs can be obtained by inducing DCs from peripheral blood monocytes and then contacting (stimulating) them with TAPs either in vitro, ex vivo, or in vivo. APCs can also be activated to present TAPs in vivo (one or more of the TAPs disclosed herein are administered to a subject), and APCs that present TAPs are induced within a subject's body. The phrases "inducing APCs" or "stimulating APCs" include contacting or loading cells with one or more TAPs or nucleic acids encoding TAPs, resulting in the presentation of TAPs on the cell surface by MHC class I molecules. As described herein, according to the present disclosure, TAP can be indirectly loaded using, for example, a longer peptide / polypeptide (including a natural protein) containing the sequence of TAP, which is then processed (e.g., by a protease) inside the APC to generate a TAP / MHC class I complex on the surface of the cell. After loading the APC with TAP and allowing the APC to present the TAP, the APC can be administered to a subject as a vaccine. For example, ex vivo administration can include (a) collecting APCs from a first subject; (b) contacting / loading the APCs of step (a) with TAP to form an MHC class I / TAP complex on the surface of the APCs; and (c) administering the peptide-loaded APCs to a second subject in need of treatment.
[0132] The first subject and the second subject may be the same subject (e.g., an autologous vaccine) or different subjects (e.g., an allogeneic vaccine). Alternatively, the present disclosure provides use of a TAP (or a combination thereof) described herein for manufacturing a composition (e.g., a pharmaceutical composition) for inducing antigen-presenting cells. In addition, the present disclosure provides a method or process for manufacturing a pharmaceutical composition for inducing antigen-presenting cells, the method or process comprising mixing or formulating a TAP, or a combination thereof, with a pharmaceutically acceptable carrier. Cells such as APCs that express MHC class I molecules (e.g., any of the HLA molecules described above) loaded with any one or any combination of the TAPs defined herein can induce CD8 + T lymphocytes, e.g., autologous CD8 + Thus, in another aspect, the present disclosure relates to a method for stimulating / proliferating T lymphocytes, more specifically CD8 T lymphocytes, by administering to a subject a TAP (or a nucleic acid or vector encoding the same) as defined herein, a cell expressing an MHC class I molecule, and a T lymphocyte, more specifically CD8 T lymphocytes. + T lymphocytes (e.g., CD8 + a cell population comprising T lymphocytes), or any combination thereof.
[0133] In one embodiment, the composition further comprises a buffer, excipient, carrier, diluent, and / or medium (e.g., culture medium). In a further embodiment, the buffer, excipient, carrier, diluent, and / or medium is a pharmaceutically acceptable buffer(s), excipient(s), carrier(s), diluent(s), and / or medium(s). As used herein, "pharmaceutically acceptable buffer, excipient, carrier, diluent, and / or medium" includes any and all solvents, buffers, binders, lubricants, fillers, thickeners, disintegrants, plasticizers, coatings, barrier layer formulations, lubricants, stabilizers, release retardants, dispersion media, coatings, antibacterial and antifungal agents, isotonic agents, etc. that are physiologically compatible, do not interfere with the effectiveness of the biological activity of the active ingredient(s), and are not toxic to the subject. The use of such media and agents for pharmaceutical active substances is well known in the art (Rowe et al., Handbook of Pharmaceutical Excipients, 2003, 4 th edition, Pharmaceutical Press, London UK). Use of any conventional media or agent in the compositions of the present disclosure is contemplated, except to the extent that the media or agent is incompatible with the active compound (peptide, cells). In one embodiment, the buffer, excipient, carrier, and / or medium is a non-naturally occurring buffer, excipient, carrier, and / or medium. In one embodiment, one or more of the TAPs defined herein, or nucleic acids (e.g., mRNA) encoding the one or more TAPs, are contained within or complexed with a vesicle, such as a lipid vesicle or liposome, e.g., a cationic lipid vesicle or liposome (see, e.g., Vitor MT et al., Recent Pat Drug Deliv Formul. 2013 Aug;7(2):99-110), or other suitable carrier.
[0134] In another aspect, the disclosure provides a composition comprising one of many, or any combination, of any one of the TAP or SLP(s) defined herein (e.g., comprising or consisting of the sequences of SEQ ID NOS: 1-23 and 25-50, preferably SEQ ID NOS: 1-23, e.g., SEQ ID NOS: 1-17 and 22) (or nucleic acid(s) encoding said peptide(s) or SLP(s)), and a buffer, excipient, carrier, diluent, and / or medium. For compositions comprising cells (e.g., APCs, T lymphocytes), the composition comprises a suitable medium that allows for the maintenance of viable cells. Representative examples of such medium include saline, Earl's buffered saline solution (Life Technologies®), or PlasmaLyte® (Baxter International®). In one embodiment, the composition (e.g., pharmaceutical composition) is an "immunogenic composition," "vaccine composition," or "vaccine." As used herein, the terms "immunogenic composition," "vaccine composition," or "vaccine" refer to a composition or formulation that contains one or more TAPs or vaccine vectors and that, when administered to a subject, can elicit an immune response against one or more TAPs present therein. The use of vaccines or vaccine vectors to induce an immune response in a mammal includes vaccines or vaccine vectors administered by any conventional route known in the vaccine art, for example, via a mucosal (e.g., ocular, intranasal, pulmonary, oral, gastric, intestinal, rectal, vaginal, or urinary) surface, via a parenteral (e.g., subcutaneous, intradermal, intramuscular, intravenous, or intraperitoneal) route, or by topical administration (e.g., via a transdermal delivery system such as a patch). In one embodiment, a TAP (or combination thereof) is conjugated to a carrier protein (conjugate vaccine) to increase the immunogenicity of the TAP(s). Accordingly, the present disclosure provides a composition (conjugate) comprising a TAP (or combination thereof) or a nucleic acid (or combination thereof) encoding a TAP and a carrier protein.For example, the TAP(s) or nucleic acid(s) may be conjugated or complexed to a Toll-like receptor (TLR) ligand (see, e.g., Zom et al., Adv Immunol. 2012, 114:177-201) or a polymer / dendrimer (see, e.g., Liu et al., Biomacromolecules. 2013 Aug 12;14(8):2798-806). In one embodiment, the immunogenic composition or vaccine further comprises an adjuvant. An "adjuvant" refers to a substance that, when added to an immunogenic agent, such as an antigen (a TAP, nucleic acid, and / or cell according to the present disclosure), nonspecifically enhances or potentiates the immune response to the agent in a host upon exposure to the mixture.Examples of adjuvants currently used in the field of vaccines include: (1) mineral salts (aluminum salts such as aluminum phosphate and aluminum hydroxide, calcium phosphate gel), squalene; (2) oil-based adjuvants (oil emulsions and surfactant-based formulations), such as MF59 (microfluidized detergent-stabilized oil-in-water emulsion), QS21 (purified saponin), AS02 [SBAS2] (oil-in-water emulsion + MPL + QS-21); (3) particulate adjuvants, such as virosomes (unilamellar liposomal vehicles incorporating influenza hemagglutinin), AS04 (aluminum salt of [SBAS4] with MPL), ISCOMS (structural complexes of saponin and lipids), polylactide-co-glycolide (PLG); (4) microbial derivatives (natural and (4) endogenous human immunomodulators, such as human GM-CSF or human IL-12 (cytokines that can be administered either as proteins or as encoded plasmids), Immudaptin (C3d tandem arrays), and / or (5) inert vehicles, such as gold particles.
[0135] In one embodiment, the TAP(s) or SLP(s) (e.g., comprising or consisting of the sequences of SEQ ID NOS: 1-23 and 25-50, preferably SEQ ID NOS: 1-23) (or nucleic acids such as mRNA encoding said peptide(s), or compositions comprising same) are in lyophilized form. In another embodiment, the TAP(s), SLP(s), nucleic acid(s), or compositions comprising same are liquid compositions. In a further embodiment, the TAP(s), SLP(s), or nucleic acid(s) are at a concentration of about 0.01 μg / mL to about 100 μg / mL in the composition. In further embodiments, the TAP(s), SLP(s), or nucleic acid(s) are at a concentration in the composition of from about 0.2 μg / mL to about 50 μg / mL, from about 0.5 μg / mL to about 10, 20, 30, 40, or 50 μg / mL, from about 1 μg / mL to about 10 μg / mL, or about 2 μg / mL.
[0136] As described herein, cells such as APCs expressing MHC class I molecules loaded with or bound to any one of the TAPs defined herein, or any combination thereof, can be used to induce CD8 + These molecules can be used to stimulate / proliferate T lymphocytes. Thus, in another aspect, the present disclosure provides T cell receptor (TCR) molecules capable of interacting with or binding to the MHC class I molecule / TAP complexes described herein, as well as nucleic acid molecules encoding such TCR molecules, and vectors comprising such nucleic acid molecules. TCRs according to the present disclosure can specifically interact with or bind to TAP loaded on or presented by MHC class I molecules, preferably on the surface of living cells in vitro or in vivo.
[0137] The term TCR, as used herein, refers to a member of the immunoglobulin superfamily that has a variable binding domain, a constant domain, a transmembrane region, and a short cytoplasmic tail (see, e.g., Janeway et al., Immunobiology: The Immune System in Health and Disease, 3rd Ed., Current Biology Publications, p. 4:33, 1997), and is capable of specifically binding to an antigenic peptide bound to an MHC receptor. TCRs can be found on the surface of cells and generally consist of a heterodimer having an α chain and a β chain (also known as TCRα and TCRβ, respectively). Similar to immunoglobulins, the extracellular portion of a TCR chain (e.g., α chain, β chain) contains two immunoglobulin regions: a variable region (e.g., a TCR variable α region or Vα, and a TCR variable β region or Vβ; typically, amino acids 1-116 according to Rabat numbering at the N-terminus) and one constant region adjacent to the cell membrane (e.g., a TCR constant domain α or Cα, typically, amino acids 117-259 according to Rabat; a TCR constant domain β or Cβ, typically, amino acids 117-295 according to Rabat). Also similar to immunoglobulins, the variable domains contain complementarity-determining regions (CDR3 in each chain) separated by framework regions (FR). In certain embodiments, TCRs are found on the surface of T cells (or T lymphocytes) and associate with the CD3 complex.
[0138] TCRs, and specifically nucleic acids encoding TCRs of the present disclosure, can be applied to, for example, T lymphocytes (e.g., CD8 + T lymphocytes) or other types of lymphocytes can be genetically transformed / modified to generate novel T lymphocyte clones that specifically recognize the MHC class I / TAP complex. In certain embodiments, T lymphocytes (e.g., CD8 + T lymphocytes (e.g., CD8 T lymphocytes) are transformed to express one or more TCRs that recognize TAP, and the transformed cells are administered to the patient (autologous cell transfusion). In certain embodiments, T lymphocytes (e.g., CD8 T lymphocytes) obtained from a donor are transfected with TCRs that recognize TAP. +In another embodiment, the present disclosure provides a method for administering T lymphocytes (e.g., CD8 T lymphocytes) transformed / transfected with a vector or plasmid encoding a TAP-specific TCR to a recipient (allogeneic cell transfusion). + In a further embodiment, the present disclosure provides a method of treating a patient with autologous or allogeneic cells transformed with a TAP-specific TCR. In certain embodiments, the TCR is expressed in primary T cells (e.g., cytotoxic T cells) by replacing the endogenous locus (e.g., the endogenous TRAC and / or TRBC locus) using, for example, CRISPR, TALEN, zinc finger, or other targeted disruption systems.
[0139] In another embodiment, the present disclosure provides a nucleic acid encoding the TCR described above. In a further embodiment, the nucleic acid is present in a vector, such as the vector described above. In a further embodiment, the nucleic acid is an mRNA molecule.
[0140] In yet a further embodiment, there is provided the use of a tumor antigen-specific TCR in the production of autologous or allogeneic cells for the treatment of cancer (e.g., colorectal cancer).
[0141] In some embodiments, patients treated with compositions (e.g., pharmaceutical compositions) of the present disclosure are treated prior to or after treatment with an anti-tumor agent and / or immunotherapy (e.g., CAR therapy). The compositions of the present disclosure may contain allogeneic T lymphocytes (e.g., CD8 + T lymphocytes), TAP-loaded allogeneic or autologous APC vaccines, TAP vaccines, and allogeneic or autologous T lymphocytes (e.g., CD8 +These include CD8 T lymphocytes (e.g., T lymphocytes), or lymphocytes transformed with a tumor antigen-specific TCR. Methods for providing T lymphocyte clones capable of recognizing TAP according to the present disclosure can be generated for and specifically target tumor cells expressing TAP in a subject (e.g., a transplant recipient), e.g., an ASCT and / or donor lymphocyte infusion (DLI) recipient. Thus, the present disclosure provides CD8 T lymphocytes that encode and express a T cell receptor that can specifically recognize or bind to a TAP / MHC class I molecule complex. + Providing T lymphocytes. The T lymphocytes (e.g., CD8 + The CD8 T lymphocytes may be recombinant (engineered) or naturally selected T lymphocytes. + At least two methods for producing T lymphocytes are provided, which involve contacting undifferentiated lymphocytes with TAP / MHC class I molecule complexes (typically expressed on the surface of cells such as APCs) under conditions conducive to T cell activation and induction of T cell proliferation, which can be done in vitro or in vivo (i.e., in patients administered an APC vaccine in which APCs are loaded with TAP, or in patients treated with a TAP vaccine). Combinations or pools of TAPs bound to MHC class I molecules can be used to generate CD8 T lymphocytes that can recognize multiple TAPs. + Alternatively, tumor antigen-specific or target T lymphocytes can be generated by integrating MHC class I molecule / TAP complexes (i.e., engineered or recombinant CD8 +TAP-specific TCRs can be produced / generated in vitro or ex vivo by cloning one or more nucleic acids (genes) encoding TCRs (more specifically, alpha and beta chains) that specifically bind to TAP (e.g., T lymphocytes). Nucleic acids encoding the TAP-specific TCRs of the present disclosure can be obtained from T lymphocytes activated against TAP ex vivo (e.g., by APCs loaded with TAP) or from an individual that exhibits an immune response to a peptide / MHC molecule complex, using methods known in the art. The TAP-specific TCRs of the present disclosure can be recombinantly expressed in host cells and / or host lymphocytes obtained from the graft recipient or graft donor and, optionally, differentiated in vitro to provide cytotoxic T lymphocytes (CTLs). The nucleic acid(s) (transgene(s)) encoding the TCR alpha and beta chains can be introduced into T cells (e.g., from the subject to be treated or another individual) using any suitable method, such as transfection (e.g., electroporation) or transduction (e.g., using a viral vector). Engineered CD8 expressing TCR specific for TAP + T lymphocytes can be expanded in vitro using well-known culture methods.
[0142] The present disclosure provides methods for generating immune effector cells that express a TCR described herein. In one embodiment, the method comprises transfecting or transducing immune effector cells, e.g., immune effector cells isolated from a subject, such as a subject with colorectal cancer (e.g., colon cancer, rectal cancer), such that the immune effector cells express one or more TCRs described herein. In certain embodiments, immune effector cells are isolated from an individual and genetically modified without further in vitro manipulation. Such cells can then be directly readministered to the individual. In a further embodiment, immune effector cells are first activated and stimulated to proliferate in vitro before being genetically modified to express a TCR. In this regard, the immune effector cells can be cultured before or after being genetically modified (i.e., transduced or transfected to express a TCR described herein).
[0143] Prior to in vitro manipulation or genetic modification of immune effector cells described herein, a cell source can be obtained from a subject. In particular, immune effector cells for use with the TCRs described herein include T cells. T cells can be obtained from several sources, including peripheral blood mononuclear cells (PBMCs), bone marrow, lymph node tissue, umbilical cord blood, thymus tissue, tissue from an infection site, ascites, pleural effusion, spleen tissue, and tumors. In certain embodiments, T cells can be obtained from a unit of blood collected from a subject using any number of techniques known to those skilled in the art, such as FICOLL™ separation. In one embodiment, cells from an individual's circulating blood can be obtained by apheresis. The apheresis product typically contains lymphocytes, including T cells, monocytes, granulocytes, B cells, other nucleated white blood cells, red blood cells, and platelets. In one embodiment, cells collected by apheresis can be washed to remove the plasma fraction and place in an appropriate buffer or medium for subsequent processing. In one embodiment of the present invention, cells are washed with PBS. In alternative embodiments, the wash solution may lack calcium, lack magnesium, or lack many, but not all, divalent cations. As will be appreciated by those skilled in the art, the wash step can be accomplished by methods known to those skilled in the art, for example, by using a semi-automated flow-through centrifuge. After washing, cells can be resuspended in various biocompatible buffers or other saline solutions, with or without buffers. In certain embodiments, undesirable components of the apheresis sample can be removed by resuspending cells directly in medium. In certain embodiments, T cells are isolated from peripheral blood mononuclear cells (PBMCs) by lysing red blood cells and depleting monocytes (e.g., by centrifugation through a PERCOLL™ gradient). Specific subpopulations of T cells, such as CD28+, CD4+, CD8+, CD45RA+, and CD45RO+ T cells, can be further isolated by positive or negative selection techniques. For example, enrichment of a T cell population by negative selection can be achieved by a combination of antibodies directed to surface markers unique to the negatively selected cells.One method for use herein is cell sorting and / or selection by negative magnetic immunoadhesion or flow cytometry, which uses a cocktail of monoclonal antibodies directed against cell surface markers present on the negatively selected cells. For example, to enrich for CD8+ cells by negative selection, the monoclonal antibody cocktail typically includes antibodies against CD14, CD20, CD11b, CD16, HLA-DR, and CD4. Flow cytometry and cell sorting can also be used to isolate cell populations of interest for use in the present disclosure. PBMCs can be used directly for TCR-mediated genetic modification using the methods described herein. In certain embodiments, after isolation of PBMCs, T lymphocytes are further isolated, and in certain embodiments, both cytotoxic and helper T lymphocytes can be sorted into naive, memory, and effector T cell subpopulations, either before or after genetic modification and / or expansion.
[0144] The present disclosure provides isolated immune cells (e.g., CD8 + The present disclosure also provides a method for detecting a TAP or a combination thereof (i.e., one or more TAPs bound to an MHC class I molecule) according to the present disclosure and a CD8 T lymphocyte capable of recognizing the TAP(s). + In another aspect, the present disclosure provides a composition comprising a CD8 T lymphocyte that specifically recognizes one or more MHC class I molecule / TAP complexes described herein. + Cell populations or cell cultures enriched in T lymphocytes (e.g., CD8 + Such enriched populations can be obtained by ex vivo expansion of specific T lymphocytes using cells such as APCs expressing MHC class I molecules that are loaded with (e.g., present) one or more of the TAPs disclosed herein. As used herein, "enriched" refers to the proliferation of tumor antigen-specific CD8 T lymphocytes in a population.+ In a further embodiment, the percentage of TAP-specific CD8 T lymphocytes in the cell population is significantly higher than in a natural cell population, i.e., a population that has not been subjected to the step of ex vivo expansion of specific T lymphocytes. + In some embodiments, the percentage of T lymphocytes is at least about 0.5%, e.g., at least about 1%, 1.5%, 2%, or 3%. In some embodiments, the percentage of TAP-specific CD8 T lymphocytes in the cell population is at least about 0.5%, e.g., at least about 1%, 1.5%, 2%, or 3%. + The percentage of T lymphocytes is about 0.5 to about 10%, about 0.5 to about 8%, about 0.5 to about 5%, about 0.5 to about 4%, about 0.5 to about 3%, about 1% to about 5%, about 1% to about 4%, about 1% to about 3%, about 2% to about 5%, about 2% to about 4%, about 2% to about 3%, about 3% to about 5%, or about 3% to about 4%. CD8 T lymphocytes that specifically recognize one or more MHC class I molecule / peptide (TAP) complexes of interest are also included. + Such cell populations or cultures enriched for T lymphocytes (e.g., CD8 + T lymphocyte populations) can be used in tumor antigen-based cancer immunotherapy, as described in detail below. In some embodiments, TAP-specific CD8 + The population of T lymphocytes may be further enriched using an affinity-based system, such as, for example, multimers of MHC class I molecules loaded (covalently or non-covalently) with TAP(s) as defined herein. Thus, the present disclosure provides methods for the detection of TAP-specific CD8 + A purified or isolated population of T lymphocytes is provided, e.g., TAP-specific CD8 + The percentage of T lymphocytes is at least about 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%.
[0145] In another aspect, the present disclosure provides an antibody or antigen-binding fragment thereof, or a soluble TCR, that specifically binds to a complex comprising a TAP described herein bound to an HLA molecule, such as an HLA molecule as defined herein. Such antibodies are generally referred to as TCR-like antibodies. As used herein, the term "antibody or antigen-binding fragment thereof" refers to any type of antibody / antibody fragment, including monoclonal antibodies (including full-length monoclonal antibodies), polyclonal antibodies, multispecific antibodies, humanized antibodies, CDR-grafted antibodies, chimeric antibodies, and antibody fragments, as long as they exhibit the desired antigen specificity / binding activity. An antibody fragment comprises a portion of a full-length antibody, generally the antigen-binding or variable region. Examples of antibody fragments include Fab, Fab', F(ab')2, and Fv fragments, diabodies, linear antibodies, single-chain antibody molecules (e.g., single-chain Fv, scFv), single-domain antibodies (e.g., from Camelidae), shark NAR single-domain antibodies, and multispecific antibodies formed from antibody fragments, single-chain diabodies (scDbs), bispecific T cell engagers (BiTEs), dual affinity retargeting molecules (DARTs), bivalent scFv-Fc, and trivalent scFv-Fc. Antibody fragments include V H Area (V H , V H -V H The term may also refer to a binding moiety comprising a CDR and an antigen-binding domain, such as, but not limited to, an anticalin, a pepbody, an antibody-T cell epitope fusion (troibody), or a peptibody. In one embodiment, the antibody or antigen-binding fragment thereof is a single-chain antibody, preferably a single-chain Fv (scFv). In one embodiment, the antibody or antigen-binding fragment thereof comprises at least one constant domain, e.g., a light chain and / or heavy chain constant domain, or a fragment thereof. In a further embodiment, the antibody or antigen-binding fragment thereof comprises a constant heavy chain fragment crystallizable (Fc) fragment of the antibody. In one embodiment, the antibody or antigen-binding fragment is an scFv comprising an Fc fragment (scFV-Fc). In one embodiment, the scFv component is connected to the Fc fragment by a linker, e.g., a hinge. The presence of an Fc region is useful for inducing complement-dependent cytotoxicity (CDC) or antibody-dependent cellular cytotoxicity (ADCC) responses against tumor cells.
[0146] In one embodiment, the antibody or antigen-binding fragment thereof is a multispecific antibody or antigen-binding fragment thereof, such as a bispecific antibody or antigen-binding fragment thereof, wherein at least one of the antigen-binding domains of the multispecific antibody or antibody fragment recognizes a complex comprising a TAP described herein bound to an HLA molecule. In one embodiment, at least one of the antigen-binding domains of the multispecific antibody or antibody fragment recognizes an immune cell effector molecule. The term "immune cell effector molecule" refers to a molecule (e.g., a protein) expressed by an immune cell, the engagement of which by the multispecific antibody or antibody fragment results in the activation of the immune cell. Examples of immune cell effector molecules include the CD3 signaling complex in T cells, such as CD8 T cells, and various activating receptors on NK cells (e.g., NKG2D, KIR2DS, NKp44, etc.). In a further embodiment, at least one of the antigen-binding domains of the multispecific antibody or antibody fragment recognizes and engages the CD3 signaling complex in a T cell (e.g., anti-CD3). In a further embodiment, the multispecific antibody or antibody fragment is a single-chain diabody (scDb). In further embodiments, the scDb comprises a first antibody fragment (e.g., scFv) that binds to a complex comprising a TAP described herein bound to an HLA molecule, and a second antibody fragment (e.g., scFv) that binds to and engages an immune cell effector molecule, such as a CD3 signaling complex (e.g., an anti-CD3 scFv) within a T cell. Such constructs can be used, for example, to induce cytotoxic T cell-mediated killing of tumor cells expressing a tumor antigen / MHC complex recognized by the multispecific antibody or antibody fragment. Antibodies or antigen-binding fragments thereof can also be used as chimeric antigen receptors (CARs) to generate CAR T cells, CAR NK cells, and the like. CARs combine a ligand-binding domain (e.g., an antibody or antibody fragment) that provides specificity for a desired antigen (e.g., an MHC / TAP complex) with an activating intracellular domain (or signaling domain) portion, such as a T cell or NK cell activation domain, to provide the primary activation signal.Antibody, more specifically, scFv, antigen-binding fragments capable of binding to molecules expressed by tumor cells are commonly used as the ligand-binding domain in CARs. Thus, in another aspect, the present disclosure provides host cells, preferably immune cells such as T cells or NK cells, that express the antibodies or antibody fragments (e.g., scFv) described herein.
[0147] In one embodiment, the soluble TCR is a soluble therapeutic bispecific TCR (see, e.g., Robinson et al., FEBS J. 2021 Nov;288(21):6159-6173; Dilchert et al., Antibodies (Basel). 2022 May 10;11(2):34).
[0148] The present disclosure relates to the immune cells (CD8 + T lymphocytes, CAR T cells) or TAP-specific CD8 + The present invention further relates to a pharmaceutical composition or vaccine comprising the population of T lymphocytes. Such a pharmaceutical composition or vaccine may comprise one or more pharmaceutically acceptable excipients and / or adjuvants, as described above.
[0149] The present disclosure further relates to the use of any TAP or SLP (e.g., comprising or consisting of any of the sequences of SEQ ID NOS: 1-23 and 25-50, preferably SEQ ID NOS: 1-23, e.g., SEQ ID NOS: 1-17 and 22), nucleic acid, expression vector, T cell receptor, antibody / antibody fragment, cell (e.g., T lymphocyte, APC CAR T cell), and / or composition according to the present disclosure, or any combination thereof, as a medicament or in the manufacture of a medicament. In one embodiment, the medicament is for the treatment of cancer, e.g., a cancer vaccine. The present disclosure relates to any TAP, SLP, nucleic acid, expression vector, T cell receptor, antibody / antibody fragment, cell (e.g., T lymphocyte, APC), and / or composition (e.g., vaccine composition) according to the present disclosure, or any combination thereof, for use in the treatment of cancer, e.g., as a cancer vaccine. The TAP sequences identified herein can be used for the production of synthetic peptides that are used i) for the in vitro stimulation and expansion of tumor antigen-specific T cells that are injected into tumor patients, and / or ii) for use as vaccines to induce or enhance anti-tumor T cell responses in cancer patients (e.g., CRC patients). In one embodiment, the cancer (e.g., CRC) expresses one or more of the TAPs described herein.
[0150] In another aspect, the disclosure provides for the use of a TAP described herein (e.g., SEQ ID NOS: 1-23 and 25-50, preferably SEQ ID NOS: 1-23, e.g., SEQ ID NOS: 1-17 and 22), or combinations thereof (e.g., peptide pools), or one or more nucleic acids encoding a TAP(s), as a vaccine for treating cancer, such as CRC, in a subject. The disclosure also provides for a TAP described herein, or combinations thereof (e.g., peptide pools), or one or more nucleic acids encoding a TAP(s), for use as a vaccine for treating cancer, such as CRC, in a subject. In one embodiment, the subject is immunized with TAP-specific T lymphocytes (e.g., CD8 +Thus, in another aspect, the present disclosure provides a method of treating (e.g., reducing the number of tumor cells, killing tumor cells) cancer, such as CRC, by administering to a subject in need thereof an effective amount of T lymphocytes (e.g., CD8 T cells) that recognize (i.e., express a TCR that binds to) one or more MHC class I molecule / TAP complexes (expressed on the surface of cells such as APCs). + In one embodiment, the method comprises administering (injecting) the CD8 + After administration / infusion of the T lymphocytes, the method further comprises administering to the subject an effective amount of a TAP or a combination thereof, or one or more nucleic acids encoding the TAP(s), and / or cells (e.g., APCs such as dendritic cells) expressing MHC class I molecule(s) loaded with the TAP(s). In yet a further embodiment, the method comprises administering to a subject in need thereof a therapeutically effective amount of dendritic cells loaded with one or more TAPs. In yet a further embodiment, the method comprises administering to a patient in need thereof a therapeutically effective amount of allogeneic or autologous cells expressing a recombinant TCR that binds to a TAP presented by an MHC class I molecule.
[0151] In another aspect, the present disclosure provides a method for treating cancer, such as CRC, in a subject by administering to a subject a T-lymphocyte (e.g., CD8 + In another aspect, the present disclosure provides the use of T lymphocytes (e.g., CD8 T lymphocytes) that recognize one or more MHC class I molecules loaded with (presenting) TAP for the preparation / manufacture of a medicament for treating cancer such as CRC (e.g., reducing the number of tumor cells, killing tumor cells) in a subject. +In another aspect, the present disclosure provides a method for treating cancer, such as CRC, in a subject by administering to a subject a TAP-loaded (presenting) TAP-containing T lymphocyte (e.g., CD8 T lymphocyte) that recognizes one or more MHC class I molecules, or a combination thereof. + In a further embodiment, the use provides the TAP-specific CD8 + After the use of T lymphocytes, the method further includes the use of an effective amount of TAP (or a combination thereof) and / or cells (e.g., APCs) expressing one or more MHC class I molecules that are loaded with (present) TAP.
[0152] The present disclosure also provides a method of generating an immune response in a subject against tumor cells (e.g., colorectal cancer cells) that express human class I MHC molecules loaded with any of the TAPs disclosed herein (e.g., SEQ ID NOS: 1-23 and 25-50, preferably SEQ ID NOS: 1-23, e.g., SEQ ID NOS: 1-17 and 22), or combinations thereof, the method comprising administering cytotoxic T lymphocytes that specifically recognize class I MHC molecules loaded with the TAP or combinations of TAPs. The present disclosure also provides use of cytotoxic T lymphocytes that specifically recognize class I MHC molecules loaded with any of the TAPs or combinations of TAPs disclosed herein to generate an immune response against tumor cells that express human class I MHC molecules loaded with the TAP or combinations thereof.
[0153] In one embodiment, the cancer is colorectal cancer. In a further embodiment, the cancer is colon cancer. In another embodiment, the cancer is rectal cancer. In one embodiment, the colorectal cancer is characterized by microsatellite instability (MSI). In one embodiment, the colorectal cancer is characterized by microsatellite stability (MSS). In one embodiment, the colorectal cancer is characterized by RAS (e.g., KRAS) and / or RAF mutations. In another embodiment, the colorectal cancer is resistant / refractory to chemotherapy. In another embodiment, the colorectal cancer is resistant / refractory to an EGFR inhibitor.
[0154] In one embodiment, the methods or uses described herein further comprise, prior to treatment / use, determining the HLA class I alleles expressed by the patient and administering or using a TAP that binds to one or more of the HLA class I alleles expressed by the patient. For example, if the patient is determined to express HLA-A2*01 and HLA-B07*02, a combination of TAP of SEQ ID NO: 6 (which binds to HLA-A2*01) and / or TAPs of SEQ ID NOs: 2, 21, 24, and / or 50 (which bind to HLA-B07*02) may be administered or used to the patient.
[0155] In one embodiment, the TAPs, SLPs, nucleic acids, expression vectors, T cell receptors, antibodies / antibody fragments, cells (e.g., T lymphocytes, CAR T or NK cells, APCs), and / or compositions according to the present disclosure, or any combination thereof, are administered in combination with one or more additional active agents or therapies for treating cancer (e.g., CRC), such as chemotherapy (e.g., vinca alkaloids, agents that interfere with microtubule formation (e.g., colchicine and its derivatives), anti-angiogenic agents, therapeutic antibodies, EGFR targeting agents, tyrosine kinase targeting agents (e.g., tyrosine kinase inhibitors), transition metal complexes, proteasome inhibitors, antimetabolites (e.g., nucleoside analogs), alkylating agents, platinum-based agents, anthracycline antibiotics, topoisomerase inhibitors, macrolides, retinoids (e.g., oncolytic agents, steroids, steroid drugs, steroid agonists ... all-trans retinoic acid or a derivative thereof), geldanamycin or a derivative thereof (e.g., 17-AAG), surgery, immune checkpoint inhibitors or immunotherapeutics (e.g., PD-1 / PD-L1 inhibitors such as anti-PD-1 / PD-L1 antibodies, CTLA-4 inhibitors such as anti-CTLA-4 antibodies, B7-1 / B7-2 inhibitors such as anti-B7-1 / B7-2 antibodies, TIM3 inhibitors such as anti-TIM3 antibodies, BTLA inhibitors such as anti-BTLA antibodies, CD47 inhibitors such as anti-CD47 antibodies, GITR inhibitors such as anti-GITR antibodies), antibodies against tumor antigens (e.g., anti-CD19, anti-CD22 antibodies), cell-based therapy (e.g., CAR In one embodiment, a TAP, nucleic acid, expression vector, T cell receptor, cell (e.g., T lymphocyte, APC), and / or composition according to the present disclosure is administered / used in combination with one or more therapies used to treat CRC (e.g., surgery, chemotherapy (e.g., using 5-fluorouracil, capecitabine, oxaliplatin, irinotecan, raltitrexed, trifluridine, tipiracil), radiation therapy, bevacizumab, cetuximab, panitumumab, regorafenib).
[0156] The additional therapy may be administered before, simultaneously with, or after administration of the TAP, SLP, nucleic acid, expression vector, T cell receptor, antibody / antibody fragment, cell (e.g., T lymphocyte, CAR T or NK cell, APC), and / or composition according to the present disclosure. [Example]
[0157] The present disclosure is illustrated in further detail by the following non-limiting examples.
[0158] Example 1: Experimental Procedure Cell Lines. Four colorectal cancer cell lines [COLO 205 (ATCC® CCL-222™), HCT 116 (ATCC® CCL-247™), RKO (ATCC® CRL-2577™), SW620 [SW-620] (ATCC® CCL-227™)] and one normal fetal small intestine cell line [HIEC6 (ATCC® CRL3266™)] were obtained from the American Type Culture Collection (ATCC). COLO205, HCT116, and SW620 were grown in RPMI-1640 (Gibco) supplemented with 10% fetal bovine serum (FBS). RKO was grown in Eagle's Minimum Essential Medium (EMEM) (ATCC®) supplemented with 10% FBS. HIEC-6 was grown in OptiMEM® 1 Reduced Serum Medium (Gibco) supplemented with 20 mM HEPES (Gibco), 10 mM GlutaMAX® (Gibco), 10 ng / mL epidermal growth factor (EGF) (Gibco), and FBS to a final concentration of 4%. All cells were maintained at 37°C and 5% CO2. For harvesting, cells were rinsed with warm PBS and then trypsinized with TrypLE™ Express Enzyme (1X) (Gibco) for 5–15 minutes at 37°C and 5% CO2. The harvested material was then spun at 1000 rpm for 5 min, rinsed once with warm PBS, and then resuspended in ice-cold PBS. After cell counting, 2 × 108 Replicates of CRC cells were pelleted and frozen at -80°C until further use. MHC class I surface density of CRC cell lines was determined by Qifikit™ (Agilent) using W6 / 32 anti-HLA class I antibody (BioXCell) according to the manufacturer's instructions.
[0159] Primary Tissues. Six paired primary human samples consisting of matched colon adenocarcinoma tumor and normal adjacent tissue (NAT) were purchased from Tissue Solutions. Tissue samples were collected from patients undergoing surgery as a first line of treatment and flash-frozen in liquid nitrogen. Further information about the primary tissue samples can be found in Table 2.
[0160] RNA extraction. For RNA extraction of cell lines, 1 to 2 million cells were collected and washed once with ice-cold PBS. Cells were then resuspended in Trizol™ (Invitrogen). For cell lines and primary tissue samples, total RNA was isolated using the AllPrep™ DNA / RNA / miRNA Universal Kit (Qiagen) or the RNeasy™ Mini Kit (Qiagen) as recommended by the manufacturer.
[0161] RNA sequencing. 500 ng of total RNA was used for library preparation. RNA quality control was assessed using the Bioanalyzer™ RNA 6000 Nano assay on a 2100 Bioanalyzer™ system (Agilent Technologies), and all samples had an RIN greater than 8. Libraries were prepared using the KAPA mRNAseq Hyperprep™ kit (Roche). Ligation was performed using an Illumina Dual Index UMI (IDT). After validation on a BioAnalyzer™ DNA1000 chip and quantification by QuBit and qPCR, libraries were pooled to equimolar concentrations and sequenced on an Illumina Nextseq500 using the Nextseq™ High Output 150 (2 × 75 bp) cycle kit. An average of 129 million and 95 million paired-end PF reads were generated for cell line and tissue samples, respectively. Library preparation and sequencing were performed at the Genomic Platform of the Institute for Research in Immunology and Cancer (IRIC).
[0162] Bioinformatics analysis. Sequences were trimmed using Trimmomatic version 0.35 (17) and aligned to the reference human genome version GRCh38 (gene annotation from Gencode version 33, based on Ensembl 99) using STAR version 2.7.1a (18). Gene expression was obtained as read counts directly from STAR and calculated using RSEM to obtain normalized gene and transcript-level expression in TPM values for these stranded RNA libraries (19).
[0163] HLA genotyping. HLA genotyping of cell lines and tissues was performed using OptiType (https: / / github.com / FRED-2 / OptiType) (20).
[0164] Microsatellite instability prediction. The MSI status of primary tumor samples was predicted using the MSIsensor program ( https: / / github.com / ding-lab / msisensor ) ( 21 ) using paired tumor and NAT.
[0165] Differential expression analysis. DESeq2 version 1.22.2 (22) was used to normalize gene read counts and calculate differential expression between tumor and normal samples. Principal component analysis (PCA) was generated using log read counts normalized for the first two most significant components. For cell line differential expression analysis, the fold change between the mean expression of four CRC cell lines compared to a normal cell line (HIEC-6) was calculated. Significantly differentially expressed genes (DEGs) were those with a padj of less than 0.05 and considered to have a |log2 fold change| > 1 for the GO term using the Metascape tool (23). For paired tissue differential expression analysis, TPM-normalized values were used to compare tumor / NAT pairs. Rather than filtering by adjusted p-values, only genes significantly differentially expressed in all six subjects for GO term analysis with a |log2 fold change| > 1 were selected because only a single replicate of tissue was sequenced. The same fold-change threshold was applied when examining differentially expressed genes between MSS and MSI tissues. For GO term analysis of MSI DEGs, genes that were exclusively differentially expressed in both MSI tissues (i.e., not considered DEGs in any MSS tissue) were selected. For GO term analysis of MSS DEGs, genes were considered if they were differentially expressed in three or more MSS tissues.
[0166] Transcriptome analysis of tissue samples. The proportions of various biotypes in the transcriptomes of tissue samples were determined as previously described (24). Briefly, after quantification and alignment of Ensembl-annotated transcripts by Kallisto (19), transcripts and repetitive elements were annotated using the Kallisto index, which includes Ensembl-annotated transcripts supplemented with gene repeat identification from the USCS Table Browser GRCh38 Repeat Masker database (25).
[0167] Mutational profile / "gene variant annotation." Gene variant calling was performed on both cell lines and primary biopsies using SNPEff (https: / / pcingola.github.io / SnpEff / #snpeff) (26).
[0168] Database generation. A general cancer database was constructed as previously described (16). Briefly, RNA-seq reads were trimmed using Trimmomatic version 0.35 (17) and aligned to the reference human genome version GRCh38 (based on Ensembl 99, with gene annotations from Gencode version 33) using STAR version 2.7.1a (18). Transcript expression at TPM was quantified using Kallisto (https: / / pachterlab.github.io / kallisto) (19). Sample-specific exomes were constructed by incorporating single-nucleotide variants (quality >20) identified in Freebayes (https: / / github.com / ekg / freebayes) into PyGeno (27). Annotated open reading frames with a TPM >0 were then translated in silico and added to the canonical proteome in fasta format. To generate cancer-specific proteomes, RNA-seq reads were cut into 33-nucleotide sequences known as kmers, preserving only kmers present in mTECs at less than two or matched NATs across cell lines and tissues, respectively. Overlapping kmers were assembled into contigs, and then the three frames were translated in silico. Notably, the short peptide sequences generated by the kmer approach were then concatenated into longer sequences of approximately 10,000 amino acids. These peptides were concatenated using a "JJ" sequence as a separator, which is internally recognized by the PeaksX+ software and splits the sequence upon the occurrence of this sequence. The canonical and cancer-specific proteomes were then concatenated to create a general cancer database. The cell line database contained an average of 3.38 × 10 6 It consisted of an array of pieces.
[0169] Isolation of MAPs. CRC cell line pellet samples (2 × 10 cells per replicate, 4 replicates per cell line) were resuspended in up to 2 mL of PBS and then solubilized by adding 2 mL of ice-cold 2x lysis buffer (1% w / v CHAPS). Tumor and normal adjacent tissue samples (455 mg–693 mg) were cut into small pieces (cubes, approximately 3 mm in size) and 5 mL of ice-cold PBS containing a protein inhibitor cocktail (Sigma, catalog no. P8340-5 ml) was added. The samples were homogenized twice, first for 20 seconds using an Ultra Turrax T25 homogenizer (IKA-Labortechnik) set at 20,000 rpm, and then for 20 seconds using an Ultra Turrax T8 homogenizer (IKA-Labortechnik) set at 25,000 rpm. Then, 550 μl of ice-cold 10x lysis buffer (5% w / v CHAPS) was added to each sample. After 60 min of incubation with tumbling at 4°C, the tissue samples and CRC cell line samples were centrifuged at 10,000 g for 30 min at 4°C. The supernatant was transferred to a new tube containing 1 mg of W6 / 32 antibody covalently cross-linked protein A magnetic beads, and MAPs were immunoprecipitated as previously described (28). The MAP extracts were then dried using a Speed-Vac and kept frozen prior to MS analysis.
[0170] TMT labeling. MAP extracts were resuspended in 200 mM Hepes buffer (pH 8.1). 50 μg of TMT reagent (Thermo Fisher Scientific) in anhydrous acetonitrile was added to the samples as follows: CRC cell line replicates were labeled with TMT6plex (Lot No. UG287166) channels TMT6-126-129, and tissue samples were labeled with TMT10plex (Lot No. UH285228) channels -126 (NAT) and -127N (tumor). Samples were gently vortexed and incubated at room temperature for 1.5 hours. Samples were then quenched with 50% hydroxylamine for 30 minutes at room temperature and then diluted with 4% FA / HO. CRC cell line replicates and individual NAT tumor pairs were combined. Samples were then desalted on a homemade C18 membrane (Empore) column and stored at -20°C until injection. To quantify MAPs of interest in primary tissue samples, synthetic peptides at concentrations ranging from 0.75 to 192 fmoles were labeled in the TMT 10plex channels -128N, -128C, -129N, -129C, -130N, -130C, and -131, while NAT and CRC tissues were labeled with -126 and 127N, respectively. Note that channel -127C was left empty to assess cross-channel contamination.
[0171] Liquid chromatography-tandem MS analysis. The dried peptide extract was resuspended in 4% FA and loaded onto a homemade C18 analytical column (20 cm x 150 μm i.d. packed with C18 Jupiter Phenomenex) on an EASY-nLC II system using a 106-minute gradient from 0% to 30% ACN (0.2% FA) at a flow rate of 600 nL / min. Samples were analyzed in positive ion mode on an Orbitrap Exploris 480 spectrometer (Thermo Fisher Scientific) with a Nanoflex source at 2.8 kV. Each full MS spectrum acquired at 240,000 resolution was followed by 20 MS / MS spectra, with the most abundant multiply charged ions selected for MS / MS sequencing using a resolution of 30,000, a 100% automatic gain control target, a 700 ms injection time, and 40% collision energy.
[0172] MAP Identification. Database searches were performed using PeaksX+ software (Bioinformatics Solutions Inc.) (29). The precursor mass and fragment ion error tolerances were set to 10.0 ppm and 0.01 Da, respectively. Nonspecific digest mode was used. TMT6plex or 10plex was set as the fixed PTM, and variable modifications included phosphorylation (STY), oxidation (M), deamidation (NQ), and TMT6plex or 10plex STY. The peak search was then loaded into MAPDP (30), which was used to apply the following filters: A 5% FDR was used to select peptides 8–11 amino acids long with a rank-eluted ligand threshold ≤2% based on NetMHCpan-4.1b predictions.
[0173] Quantification of MAP coding sequences in RNA-Seq data. MAP coding sequences (MCSs) were quantified in RNA-Seq data as previously described (31). Briefly, MCSs were reverse-translated into all possible nucleotide sequences using a Python script at our institution (deposited at Zenodo under DOI: 3739257). The nucleotide sequences were then mapped to the genome using GSNAP (32) to determine all possible genomic locations that could encode a given MAP. MCSs were also mapped to the transcriptome to account for splice sites where the MAPs overlap. The portions of the transcriptome corresponding to these MAPs were then also mapped to the reference genome using GSNAP. For the MAP of interest, genome alignment of all reads containing the MCS was performed. The GSNAP output was filtered to retain only exact matches between the sequence and the reference, resulting in a file containing all possible genomic regions that could encode a given MAP. The number of reads containing an MCS at each genomic location in each desired RNA-Seq sample (such as a CRC and NAT, GTEx, or TCGA sample) aligned to the reference genome by STAR was summed. Finally, all read counts for a given MAP were summed and normalized to the total number of reads sequenced in each sample of interest to obtain reads per hundred million (RPHM) counts.
[0174] Determination of MAP source transcripts. To determine what proportion of tissue sample MAPs originated from a particular transcript biotype, we determined the most abundant putative source transcript based on kmer per 100 million (KPHM) quantification. For peptides from cancer-specific (kmer) databases, we reverse-translated the MCS into all possible nucleotide sequences and identified all possible genomic regions that could encode a given MAP (see "Quantification of MAP-coding sequences in RNA-Seq data" above). Finally, we used Kallisto to determine the most abundant transcript at that location, which was then assigned as the most likely transcript for a given peptide. Peptides with more than one putative source transcript were excluded from the analysis.
[0175] Identification of TSA Candidates. TSA candidates were identified through a rigorous TSA identification pipeline. First, MAP underwent peptide classification, which retrieves peptide sequence accessions from protein databases and uses them to extract the nucleotide sequence of each peptide. RNA-Seq data from each cancer and normal tissue was converted into a 24-nucleotide-long k-mer database by Jellyfish 2.2.3 (using the -C option), which was then used to query the 24-nucleotide-long k-mer set of each TSA candidate coding sequence. The number of reads that completely overlap with a given peptide coding sequence was estimated using the minimum number of occurrences (kmin) of the k-mer set. In general, one k-mer always originates from a single RNA-Seq read. This kmin value was then calculated using the following formula: kphm = (kmin × 10 8 )rtot (where rtot represents the total number of reads sequenced in a given RNA-Seq experiment) 8The number of k-mers (kphm) detected per read was converted to the number of kphm detected per read. Peptides were retained only if their RNA coding sequences were expressed at least 10-fold higher in cancer than normal (pooled mTEC samples for cell lines, matched NATs for tissues) and expressed less than 2 kphm in normal. Subsequent filtering removed peptides with indistinguishable isoleucine / leucine variants. Peptides with IL variants were retained only if the most highly expressed variant met the criteria described above. The MCS of the remaining peptides were quantified in the RNA-seq data as described above, and they were retained only if their expression was less than 8.55 kphm in mTEC and other normal tissues (GTEx). The genomic location of each peptide was assigned by mapping reads containing each MCS to the reference genome (GRCh38.99) using BLAT (https: / / genome.ucsc.edu / cgi-bin / hgBlat, Kent WJ. BLAT - the BLAST-like alignment tool. Genome Res. 2002 Apr;12(4):656-64. PMID:11932250). Peptides were excluded if their genomic location was unknown or if they mapped to hypervariable regions (HLA, Ig, or T cell receptor (TCR) genes). Finally, MS / MS spectra of the remaining candidates were manually verified. Peptides were classified as mTSAs if their amino acid sequence differed from the reference and if the mutation was not a known germline polymorphism. Peptides were classified as aeTSAs if they were overexpressed 10-fold or more in tumors compared to normals and ≤0.2 kphm in mTECs (and NATs in the case of tissues), and as TAAs if they were overexpressed 10-fold or more in cancers but expression in mTECs and / or NATs was >0.2 kphm. Finally, the transcript of origin of each TSA / TAA was imputed by selecting the most highly expressed peptide overlap transcript from the kallisto quantification file (see database generation section).
[0176] Intertumor sharing. To examine the intertumor distribution of TSA and TAA sequences in other CRC tumors, we determined the log(RPHM+1) expression of peptide-coding sequences in 151 colon adenocarcinoma (COAD) samples from TCGA (see "Quantification of MAP-coding sequences in RNA-Seq data").
[0177] Immunogenicity prediction. The predicted immunogenicity of the MAPs of interest was determined using the R package Repitope v3.0.1 ( https: / / github.com / masato-ogishi / Repitope ) ( 34 ).
[0178] Validation of TSA peptide candidates. Synthetic peptides of TSA and selected TAA sequences were obtained from Genscript. Synthetic peptides were solubilized in DMSO to a concentration of 1 nmol / μL, and all synthetic peptides were combined in a stock solution at a concentration of 10 pmol / μL. The stock solution was desalted in 150 pmol aliquots on a homemade C18 membrane (Empore) column and dried using a Speed-Vac. The dried peptide extract was labeled with the TMT10plex channel as described (see the "TMT Labeling" section), desalted, and dried in a Speed-Vac. The labeled synthetic peptides were resuspended in 4% FA, and 1 pmol of each synthetic peptide was loaded onto a homemade C18 analytical column (20 cm × 150 μm i.d. packed with C18 Jupiter Phenomenex) on an EASY-nLC II system using a 76-minute gradient from 0% to 30% ACN (0.2% FA) at a flow rate of 600 nL / min. Samples were analyzed in positive ion mode on an Orbitrap Exploris 480 spectrometer (Thermo Fisher Scientific) with a Nanoflex source at 2.8 kV. Each full MS spectrum was acquired at 120,000 resolution, and an inclusion list was used to select ions for fragmentation with a collision energy of 40% and an isolation window of 1 m / z. MS / MS was acquired at 30,000 resolution. MS / MS correlations were calculated as previously described (17). Briefly, predicted peptide fragments were calculated using pyteomics v4.0.1 (https: / / bitbucket.org / levitsky / pyteomics) to identify reproducibly detected peptide fragments. The root-scale intensities of these fragments were correlated between endogenous and synthetic peptide scan pairs, and Pearson correlation coefficients, p-values, and confidence intervals were calculated using SciPy v1.2.1 (https: / / www.scipy.org / ). Mirror plots of the scan pair with the lowest p-value were generated for each peptide using spectrum_utils v0.2.1 (https: / / github.com / bittremieux / spectrum_utils).
[0179] Relative quantification. To relatively quantify MAPs of interest in primary tissue samples, synthetic peptides labeled with TMT 10plex-129N, 130N, and 131N, respectively, at concentrations of 10, 100, or 1,000 fmol were spiked into the remaining purified MAPs from NAT and CRC tissue samples labeled with TMT 10plex-126 and 127N, respectively. Note that channel TMT 10plex-127C was left empty to assess contamination. Samples were analyzed in positive ion mode with an Orbitrap Fusion Tribrid spectrometer (Thermo Fisher Scientific) at 2.4 kV and a Nanoflex source. For synchronous precursor selection MS3 (SPS-MS3), a full MS scan was acquired using an inclusive list of peptides of interest in the 300-1,000 m / z range, an Orbitrap resolution of 120,000, an automatic gain control (AGC) of 5.0e5, and a maximum injection time of 50 ms. A 3-second maximum speed approach for MS2 was used in the ion trap with a 0.4 m / z isolation window, 35% collision-induced dissociation, "normal" ion trap scan rate mode, an AGC target of 2.0e4, and a maximum injection time of 50 ms. This was followed by selection of eight synchronized precursor ions for MS3 acquisition, which was performed with a scan range of 110-500 m / z, an Orbitrap resolution of 50,000, an AGC of 1.0e5, a maximum injection time of 300 ms, a 2.0 m / z isolation window, and a HCD collision energy of 65%. The LC-MS instrument was controlled using Xcalibur version 4.4 (Thermo Fisher Scientific, Inc.). The error tolerances for precursor mass and fragment ions were set to 10.0 ppm and 0.5 Da, respectively. Nonspecific digest mode was used. TMT10plex was set as the fixed PTM, and variable modifications included phosphorylation (STY), oxidation (M), deamidation (NQ), and TMT10plex STY. For quantification, PSMs were filtered to exclude those with contamination in the TMT10plex-127C channel and select those within the 70th intensity percentile.Peptides were selected for quantification by manually inspecting the MS2 precursor and intensity profiles of all relevant channels. The intensity ratio of each peptide was calculated using the average TMT10plex-127N and TMT10plex-126 intensities of good-quality PSMs.
[0180] Data analysis and visualization. Figure 1 was generated using BioRender.com. Most other figures were created with Python v3.7.6, R v3.6.3, or Origin (Pro) 2019b. R packages included Repitope v3.0.1 (https: / / github.com / masato-ogishi / Repitope) (34), UpsetR v1.4.0 (https: / / github.com / hms-dbmi / UpSetR) (35), GSVA v1.38.2 (https: / / github.com / rcastelo / GSVA) (36), and ESTIMATE v1.0.13 (https: / / bioinformatics.mdanderson.org / estimate / ) (37).
[0181] Experimental Design and Statistical Rationale. To effectively elucidate the MHC I immunopeptidome of colorectal cancer, four CRC cell lines were selected and six samples from human subjects, consisting of both matched tumor and normal adjacent tissue (NAT), were obtained. NAT was used as a healthy tissue approximation, as it is the most effective control for each respective tumor. Because matched samples were not available for the cell lines, a pool of six mTEC samples was used to create a general cancer database to provide a broad range of approximate normal RNA expression. p-values in all cases were determined using two-sample t-tests, except for determining the significance of the immunogenicity score. In this case, a Mann-Whitney test was used because the data did not have a normal distribution, as determined by the Shapiro test. For t-tests, an f-test was performed to determine whether the dataset had significant variation. If there was significant variation, a t-test assuming variation was used; otherwise, a t-test assuming no variation was used. For CRC-derived cell lines, 2 × 10 8 Four technical replicates of cells were prepared, TMT-labeled, and multiplexed prior to injection. Due to limited tissue material, half of the purified MAP from the primary sample was injected to obtain global immunopeptidomic data, while the remaining sample was used for targeted analysis with synthetic peptides to confirm the sequence and abundance of putative TSAs and selected TAAs. To select high-quality PSMs for quantification, those with low intensity or contamination in the empty TMT channel were excluded. Furthermore, only peptides with favorable MS2 precursors and intensity profiles were quantified.
[0182] Example 2: Immunopeptidome analysis using a proteogenomics approach To determine the composition of the colorectal cancer immunopeptidome, we analyzed a sample set consisting of four colorectal cancer-derived cell lines and six sets of primary adenocarcinoma samples from matched tumor and normal adjacent tissues (Tables 1 and 2). Paired-end RNA sequencing (RNA-seq) enabled the creation of a general cancer database consisting of a canonical cancer proteome, created by generating cancer-specific kmers translated into three reading frames to encompass noncanonical sequences from any genomic origin, and a cancer-specific proteome for each sample (Figure 1). Medullary thymic epithelial cells (mTECs) present peripheral antigens within the thymus and mediate the negative selection of autoreactive T cells (38). For CRC-derived cell lines, cancer-specific kmers were obtained after subtracting mTEC-derived sequences that approximate the expression of these sequences in healthy tissues. For primary tissue samples, cancer-specific kmers were generated after subtracting sequences from matched NATs. This approach allowed the identification of sequences expressed in tumors that were not observed in healthy colon tissue from the same individuals. In addition to database construction, RNA-seq data were also used for transcriptome analysis, including GO term analysis, investigation of immune infiltration, mutation profiling, and transcript abundance determination (Figure 1). [Table 4] [Table 5] [Table 6] JPEG2024535853000008.jpg246130
[0183] MHC I:peptide complexes were isolated by immunoprecipitation, and the eluted MHC I-associated peptides (MAPs) were labeled with tandem mass tag (TMT) isobaric labeling reagents. TMT labeling has recently been shown to enhance the detection of MAPs by increasing their charge state and hydrophobicity (37). MAPs were then sequenced and analyzed by liquid chromatography-tandem mass spectrometry (LC-MS / MS) using a personalized cancer database generated by RNA-seq. The identified MAPs then underwent a rigorous series of classification and validation to identify putative TSAs and TAAs. The identified tumor antigens were then validated in CRC tissues and quantified with synthetic peptides to determine the extent to which they were overexpressed on the tumor cell surface. Their predicted immunogenicity and intertumor distribution were also examined to gain a sense of their clinical potential (Figure 1).
[0184] In this study, four CRC-derived cell lines with different alleles and characteristics were used, as summarized in Table 1. HCT116 and RKO were derived from primary tumors and characterized by microsatellite instability (MSI), while Colo205 and SW620 were derived from ascites and lymph node metastases, respectively, and both are microsatellite stable (MSS). These cell lines were 1.44 × 10 5 ~5.07×10 5 They have a wide range of MHC I surface expression, with 1.44 x 10 MHC I molecules / cell and a range of HLA allele diversity. Among the four cell lines, there are mutations in several key genes, including BRAF, RAS, SMAD4, TP53, and PI3CA. These cell lines have a 1.44 x 10 5 ~5.07×10 5 The diverse MHC I surface expression ranged from 100 MHC I molecules / cell, and the HLA allele diversity identified using OptiType, an HLA genotyping tool that uses RNA-Seq data to predict HLA alleles for samples, combined with HLA alleles for these cell lines shown in the literature (Table 1) (22).
[0185] All primary tumor samples were derived from stage II adenocarcinomas, with slight variations in tumor grade and tumor-node-metastasis (TNM) classification (Table 2). Primary samples contained 95-100% tumor content and had a mean mass of 0.6625 grams. All tumors originated from the sigmoid colon, except for S1 (cecum) and S5 (ascending colon). All patients except S2 were female, and patient ages ranged from 43 to 85 years, with a mean age of 62 years. Similar to cell lines, tissue samples also possess a variety of HLA alleles. A visualization of the number of unique or shared HLA alleles in cell lines and tissue samples is available in Figure 7. There are an average of 1.3 and 3.2 unique alleles per cell line and tissue, respectively.
[0186] Example 3: Transcriptome characterization of CRC samples Because outcomes for CRC patients within a given disease stage vary significantly based on the molecular characteristics of their tumors (40, 41), we used RNA sequencing data to characterize the molecular heterogeneity of the samples. After first examining the mutation status of key biomarkers (e.g., KRAS, NRAS, or BRAF) commonly used to guide therapeutic decisions and prognosis in the clinic (40, 41) (Table 1), we determined the microsatellite status of cell lines and primary samples, respectively, from the literature (42, 43). We then expanded our knowledge of the molecular features of these samples using the MSIsensor package (46). While MSI is found in a limited subset of CRC tumors (i.e., 15% of sporadic CRCs and 90% of nonpolyposis colorectal cancers) (47), in this study, 50% of tumorigenic cell lines and 33% of primary biopsies exhibited this phenotype (Table 3). Although several elements of the literature suggest that MSI and MSS tumors are immunologically distinct ( 5 , 11 , 48 , 49 ), the present study provides a comparison of MSI and MSS colorectal tumors at the immunopeptidome level. [Table 7]
[0187] Principal component analysis of the top 500 genes between normal and tumor biopsy samples (Figure 2A) or cell lines (Figure 9A) confirms their distinct transcriptome profiles. Accordingly, pathway and process enrichment analysis of both CRC-derived cell lines and biopsy samples revealed enriched transcriptome profiles related to their tumorigenesis status. The most significantly up- and down-regulated GO terms are associated with cell proliferation (Figure 2B, upper panel) and muscle phenotype and contractility (Figure 2B, lower panel), respectively. While enrichment of proliferation- and cell cycle-related GO terms is a general hallmark of cancer (50, 51), downregulation of muscle-related pathways is unique to CRC and arises from the functional dichotomy between poorly differentiated tumor regions and highly contractile NAT. Among tumor samples in the dataset used herein, MSI / MSS status appears to explain intertumor transcriptome differences (Figures 2A and 9A). While MSI samples tend to cluster tightly together, MSS tumors appear more dispersed and therefore transcriptomically more heterogeneous. Functionally, when analyzed separately, MSS and MSI CRC samples are enriched for very different gene sets. Compared to their NAT counterparts, MSI tumors are characterized by significant upregulation of various immune-related GO terms (Figure 9A), whereas MSS tumors are more associated with increased expression of genes related to both Wnt and PI3K-Akt signaling (Figure 9B). While the association of these two signaling pathways with CRC is known (52), we found no literature supporting the possibility that their contribution in CRC may differ between MSS and MSI tumors.
[0188] Next, the degree of immune infiltration in each sample was estimated by two approaches: using the immune infiltration score from the ESTIMATE package (37) (Figure 2C) and the enrichment score of known tumor-infiltrating leukocyte (TIL) markers (53) based on single-sample Gene Set Enrichment Analysis (ssGSEA) (54) (Figure 8B). All NAT samples showed similar levels of immune infiltration, whereas MSI and MSS tumors were characterized by increased and decreased immune infiltration scores, respectively (Figure 2C and Figure 8B). Such differences suggest that MSI tumors may be more immunogenic than their MSS homologs (48, 55-58).
[0189] Because TSAs can arise from a wide range of cancer-specific events / dysregulations (11, 12) and the immunopeptidome contribution of each antigen source varies significantly between malignancies (11), we also used RNA sequencing data to inform which TSA classes may be enriched in samples. By looking at the genomic origin of transcripts, we observed that both the proportion and absolute abundance of noncoding polyadenylated RNAs were significantly increased in tumors compared with NATs (Figure 2D). On average, the increase in absolute abundance was limited to 25%, but the data suggest that tumor-specific increases in noncoding transcripts may be higher in MSI tumors than in MSS tumors. While this comparison remains limited due to the small number of MSI samples (n = 2), it can be expected that the number of aeTSAs derived from noncoding transcripts in MSI samples will be higher than in MSS tumors. Similarly, while all CRC samples exhibit comparable single-nucleotide variant (SNV) burden and are therefore expected to have a similar number of mTSAs (Figure 2E), the insertion / deletion (indel) burden is significantly increased in MSI samples compared to MSS samples, an observation also noted for cell lines (Figure 8C). Indeed, when both cell lines and tissue samples are considered together, a statistically significant difference emerged in the number of INDEL mutations between MSS and MSI samples (p = 0.00235) (Figure 8D). Because both MSI and INDEL accumulation result from defects in the DNA mismatch repair (MMR) pathway (59), it can be hypothesized that the number of INDEL-derived TSAs (most likely frameshift-derived antigens) identified in a sample will be proportional to its MSI level.
[0190] Example 4: Immunopeptidome analysis highlights the diversity of CRC antigens To elucidate the MHC I immunopeptidome of CRC-derived cell lines, 2 × 10 8MAPs from four replicates of cells were immunoprecipitated for each line, and each replicate was derivatized with a separate TMT6plex channel (126, 127, 128, 129) for cell lines or with TMT10plex-126 and -127N for primary NAT and tissue samples, respectively. Four replicates of each cell line and half of each NAT and tumor MAP from each subject were multiplexed and analyzed by LC-MS / MS. The median labeling efficiency for cell lines and tissue samples was 72.4% or 87.8%, respectively. The low labeling efficiency in cell lines was attributed to the low MAP yield. 5281 and 27583 unique MAPs were identified in the cell line and tissue datasets, respectively, with an average of 1433 unique MAPs per cell line and 5855 unique MAPs per tissue (Figure 3A, upper panel, and Figure 3B). Although identification varied between strains, the number of identified MAPs correlated strongly with the abundance of MHC I molecules per cell (Fig. 3A, lower panel, Pearson's r = 0.96).
[0191] When cell lines and tissue samples were collected together, a total of 30,485 unique MAPs were identified. Within each sample's MAP repertoire, 32–68% of peptides were sample-specific, and even when comparing only cell lines or primary samples, very few MAPs were shared (Figure 10A–B). The majority of these unique MAPs can be attributed to the diversity of HLA alleles between samples, a major factor affecting which peptides can be presented on the cell surface (Figure 3C, Figure 7). On average, the number of MAPs shared by any two cell lines or any two tissue samples is 59 or 640, respectively. There are notable outliers. Tissue samples S1 and S6 share 2,079 MAPs, of which 1,673 are unique to these samples (Figure 10C), representing more than one-third of their respective MHC I immunopeptidomes (Figure 10D). The next closest similarity in the MAP repertoires between the two tissues is the 1,328 MAPs shared by the two MSI tissues (S5 and S6), corresponding to 21% of their repertoires. The reduction in MAP specificity in cell lines makes these comparisons less striking. For example, although HCT116 and RKO share most MAPs, these peptides represent less than 10% of their MAPs and are likely characteristic of their larger peptide repertoires (Figure 10C). In contrast, COLO205 and SW620 share 152 MAPs, nearly one-fifth of their immunopeptidomes.
[0192] To put these comparisons into context, we can again consider the HLA alleles of the samples. Of the 2079 MAPs shared by S1 and S6, 1542 are predicted to bind the same allele in both samples, which in more than 93% of these cases is HLA-B*07:02. Similarly, of the 152 MAPs shared by COLO205 and SW620, 136 are bound by the same allele, which in 113 of these cases is HLA-A*02:01. Thus, the MHC I immunopeptidome of a sample is primarily influenced by its HLA repertoire.
[0193] At the gene level, peptides derived from over 8,000 unique source genes were identified, with an average of 1,014 and 3,168 source genes per cell line and tissue sample, respectively (Figure 11A, upper panel). The number of source genes identified in each sample highly correlated with the number of identified MAPs (Figure 11A, lower panel). Approximately 6–14% of the source genes in a given immunopeptidome were sample-specific (Figure 11B), which may be due to sample-specific biological features or may reflect incomplete sampling of the immunopeptidome (Figure 11E). Because we do not expect to identify all MAPs displayed on the cell surface, and the majority of source genes in each sample are attributable to only a single MAP, it is almost certain that additional source genes contribute to the MAP repertoire and simply go undetected. When comparing any two tissue samples, they share an average of 48% of source genes, whereas when comparing any two cell lines, they share an average of 24% of genes (Figure 11C-E). Therefore, different cell lines appear less homogeneous at the source gene level. This likely reflects differences in sample composition, as tissue samples have source genes derived from natto, stromal, and infiltrating cells, while cell lines consist of only a single cell type. In addition, the reduced MHC I expression of cell lines, resulting in reduced specificity of MAPs, means that fewer source genes were present in the samples, reducing the likelihood of overlap. Nevertheless, all samples are more similar at the source gene level compared to the immunopeptidome level, and sample-specific MAPs are derived from shared source genes.
[0194] To obtain an overview of the genomic origin of the MHC I immunopeptidome and investigate the shared nature of source genes, we performed GO term analysis on all source genes identified in cell lines and those identified in four or more tissues. Several common features between cell lines and tissues are detectable at the immunopeptidome level, including significant enrichment for genes involved in RNA metabolism, ribonucleoprotein complex biogenesis, translation, and cellular responses to stress (Figure 3D). Thus, despite heterogeneity between both cell lines and tissue samples, including the large diversity of HLA alleles that significantly influences peptide repertoires and the low MAP specificity in cell lines, there is a striking similarity in which genes contribute to the MHC I immunopeptidome.
[0195] To determine what proportion of MAPs from tissue samples were derived from non-coding transcripts, we first determined the most abundant putative source transcript for each peptide (Ensembl Annotation 99). For peptides from cancer-specific databases, we mapped MCSs to the genome and determined the most abundant transcript at that location (see "Quantification of MAP-Coding Sequences in RNA-Seq Data" in the Methods section). Therefore, on average, we determined that 95.3% of MAPs from tissue samples were derived from protein-coding transcripts (i.e., UTRs or CDRs) (Figure 3E, left panel). These peptides are likely derived from intergenic sequences, so approximately 4.2% of peptides originate from non-coding regions, including the 2.8% of peptides derived from unannotated RNA transcripts. Approximately one-third of all non-coding MAPs (including those from unannotated transcripts) are derived from nonsense-mediated decay transcripts, and less than 1% of them are derived from lncRNAs, nonstop decay products, retained introns, or processed transcripts (transcripts without an open reading frame) ( Figure 3E , right panel).
[0196] Example 5: Identification of tumor-specific and tumor-associated antigens in CRC After identifying over 30,000 unique MAPs, we filtered peptide-coding sequences to select those that were overexpressed by at least 10-fold in cancer and expressed by less than 2 rphm in pooled mTEC samples or matched NATs for cell lines and primary samples, respectively. A recent immunopeptidome study in acute myeloid leukemia (AML) showed that MCSs with an RPHM <8.55 had a <5% probability of generating MAPs (18). We then quantified MAP-coding sequences in the RNA data and retained only those expressed at less than 8.55 rphm in mTECs and other normal tissues (GTEx). Following manual validation of the remaining peptides, peptides were classified as aberrantly expressed TSAs (aeTSAs) if they were overexpressed by at least 10-fold in tumors and expressed by less than 0.2 rphm in mTECs (and NATs, in the case of tissues). MAPs were also overexpressed at least 10-fold in cancers, but were classified as TAAs if their expression in mTECs and / or NATs was greater than 0.2 kphm.
[0197] Although TSA identification in CRC-derived cell lines was relatively poor, likely due in part to low MAP identification, an average of three TSAs were identified per primary tissue sample (Figure 4A). Overall, one putative TSA was identified in CRC-derived cell lines and 18 in primary tissues, with the TSA yield from each sample correlated with the number of identified MAPs (Pearson's r = 0.76) (Figure 12). Of these, approximately one-third were derived from coding regions, whereas the majority of identified putative TSAs were derived from non-coding regions (Figure 4B). Of the TSAs derived from coding regions, two were derived from non-canonical reading frames derived from exon frameshift sequences, and two were mutant TSAs identified in MSS tissues S2 and S3 (Figures 4A and 4B). Of the non-coding TSAs, all are aberrantly expressed, with the majority originating from intronic or intergenic regions, and fewer from 5'UTRs, 3'UTRs, or lncRNAs (Figure 4B). The sequences of nine aeTSAs (five intronic, three intergenic, and one lncRNA) overlap with ERE sequences (Table 4). Due to the ubiquitous nature of EREs, TSAs derived from aberrant ERE expression are potentially shared by tumors and have been shown to be immunogenic (61, 62). Notably, none of the putative TSAs were shared across multiple samples, even in samples with a high proportion of shared MAPs. However, two unique TSAs were identified in different tissues derived from the same transcript of COL11A1 (one exon frameshift and one 5'UTR), which were recently shown to play a role in the development and prognosis of CRC (63). Most of the other TSA source genes have also been shown to be biologically relevant in CRC (Table 5). [Table 8-1] [Table 8-2] [Table 8-3] [Table 9-1] [Table 9-2] [Table 9-3] [Table 10-1] JPEG2024535853000017.jpg34164 [Table 10-2]
[0198] Although the primary goal was to identify putative TSAs in CRC, an average of 6.33 TAAs were also identified in CRC tissue samples but not in CRC-derived cell lines (Figure 4C). In contrast to the primarily non-coding putative TSAs, the majority of identified TAAs were derived from canonical coding exon sequences, with only a minority derived from introns, intergenic sequences, or lncRNAs (Figure 4D). Two non-canonical TAAs were overlapped by ERE sequences (Table 4). Notably, four distinct TAAs were identified in more than one sample, and one TAA was identified in three tissues. All of these repeated TAAs were derived from canonical exons, and the source transcripts were ASPM, MKI67, DIAPH3, MMP12, NOS2, and SPC25, all of which have been shown to be associated with cancer (Table 6). [Table 11-1] [Table 11-2]
[0199] We initially expected to identify above-average numbers of both TSAs and TAAs in MSI tissues. This was true for S5 but not for the other MSI tissues (Figures 4A and 4C). This may be due to S6 having a lower "degree" of instability, as reflected in the MSIsensor results (Table 3). Furthermore, the sample with the highest number of identified TSAs was S2 (MSS tissue). Therefore, it appears that the TSA and TAA yield per sample may be due to other intrinsic biological features of the tumor, regardless of MSI status, that are outside the scope of this study.
[0200] To determine whether any of the putative TSAs or TAAs had been previously identified, we examined whether the peptide sequences were reported in the Immune Epitope Database (IEDB), the HLA Ligand Atlas (64), and two previous publications that attempted to identify tumor antigens in CRC: Loeffler et al. 2018 (8) and Newey et al. 2019 (15). Notably, none of the putative aeTSAs, mTSAs, or noncanonical TAAs had been previously reported in any of these sources. Of the 26 putative canonical TAAs identified, 24 had been reported in the IEDB, Loeffler et al. 2018 (PXD009602), Newey et al. 2019 (PXD014017), or some combination of these three (Figure 4E). Eight of these were also reported in the HLA Ligand Atlas, one of which was specifically represented in healthy colon tissue. Interestingly, none of the previously identified TAAs in these previous publications were reported as tumor antigens. Conversely, six of the 12 tumor antigens of interest reported in Loffler et al. were also identified in the immunopeptidome of this study; however, they did not exceed the threshold established in the identification pipeline for being considered TSAs or TAAs, in most cases due to their high expression in NAs (Table 7A). Thus, a selection of novel colorectal cancer TSAs, primarily derived from noncoding regions, and primarily coding TAAs (some of which have previously been reported as MAPs but not in the context of their biological relevance as TAAs), have been identified in this study. Tables 7B and 7C show the TSAs and TAAs identified in this study, respectively. [Table 12] [Table 13] [Table 14-1] [Table 14-2] [Table 14-3]
[0201] Example 6: Validation of putative tumor-specific and tumor-associated antigens Following the identification of putative TSAs and TAAs, validation of all TSAs and a subset of 11 TAAs was performed. These were selected based on favorable initial TMT intensity ratios and precursor ion proportions in cancers relative to matched NATs before validation with synthetic peptides. First, we studied the expression of their respective source transcripts in tumor samples compared to their matched NATs, and the mean average of their transcripts in CRC / NAT samples in TPM (Figure 5A). This analysis, of course, does not include peptides derived from intergenic regions. The mean log2FC of the source transcripts of putative TSAs and TAAs in the samples in which they were identified was 3.6 and 3.2, respectively. While there are several cases (S2, S6) in which the source transcript of an aeTSA is more abundant in NATs than in tumors, this reflects only the overall abundance of the entire transcript; the peptide-coding sequence is actually more abundant in cancers. This was also true for aeTSAs in which the peptide-coding region is either completely absent or is underexpressed in NATs but is more highly expressed in cancer tissues.
[0202] To assess the specificity of the putative tumor antigens, we determined the average expression of peptide-coding sequences in a large dataset of healthy tissues provided by the Genotype-Tissue Expression project (GTEx) (Figure 5B). No TSA sequences were expressed above 8.55 rphm in any healthy tissue, except for RIGGVGVEK, an aesthesia-associated TSA identified in S2, which is expressed above threshold in testis (Figure 5B). This suggests that this TSA could also be classified as a cancer-testis antigen (CTA), a type of aesthesia-associated TSA that is expressed in male germ cells but can also be aberrantly expressed in cancer. Due to the absence of MHC I in testis, these antigens are also promising candidates for cancer immunotherapy (65). This putative TSA is an LY6G6F-LY6G6D exon frameshift. These genes have not previously been reported as CTAs, although another member of the same gene family, LY6K, has been reported as a CTA in lung and esophageal cancer (66). TAA expression was below threshold in healthy tissues but tended to be higher in the esophagus and transverse colon. Seven of these peptides were also expressed above threshold in the testis.
[0203] Example 7: Cancer specificity and immunogenicity of TSAs and TAAs The putative TSAs and TAAs were validated by MS using their corresponding synthetic peptides. These TAAs were selected based on desirable initial TMT intensity ratios and precursor ion proportions in cancer relative to the matched NATs. All of these candidates had MS / MS that correlated well with the MS / MS of the synthetic peptides, with Pearson correlation scores of 0.6 or higher. The synthetic peptides were labeled with TMT10plex-129N, 130N, and 131 at concentrations of 10, 100, and 1000 fmol, respectively, and spiked into the remaining purified MAP from tissue samples labeled with TMT126 (NATs) and 127N (CRCs). SPS-MS3 was then used to quantify the peptides of interest in these samples. Despite the reduced sensitivity of SPS-MS3, it was possible to quantify seven TSAs and seven TAAs. Good-quality PSMs were selected for quantification, and all were more abundant in their respective CRCs compared to NATs (Table 8). Determining the ratio of the intensities of the TMT127N peptide compared to the TMT126 peptide revealed that TSAs had a median intensity fold change of 16.96 in CRC compared to NATs, and TAAs had a fold change of 6.93. Additionally, a TSA with the sequence RYLEKFYGL was also overexpressed in S1 tumors, but exceeded the transcriptome threshold only for S6. Therefore, it was possible to demonstrate that the TSA identification method used in this study successfully identified TSA and TAA sequences that were more highly present on the surface of cancer cells than those in NATs. The TSA and TAA candidates validated with synthetic peptides are listed in Tables 9A-9B. [Table 15-1] [Table 15-2]
[0204] Tables 9A and 9B show the TSA and TAA candidates that were validated with synthetic peptides. [Table 16] [Table 17]
[0205] To examine the intertumor distribution of these TSAs and TAAs in other CRC tumors, we plotted the log(RPHM+1) expression of peptide-coding sequences in 151 colon adenocarcinoma samples from The Cancer Genome Atlas (TCGA) (Figure 6A). To assess antigen sharing potential, we first calculated the average log-transformed (log(rphm+1)) values of pooled GTEx (n = 2442) and mTEC (n = 8) samples for each peptide of interest. Overall, nine TSAs (53%) and nine TAAs (100%) had expression 10-fold or higher than the corresponding average GTEx / mTEC values in at least 5% of TCGA COAD tumors. This indicates that TAAs are more frequently shared between COAD TCGA tumors than their TSA counterparts. However, this also implies that most TSAs are highly shared among these samples.
[0206] Another important consideration in identifying tumor antigens is whether these peptides can elicit effective antitumor immune responses. Immunogenic epitope prediction revealed that aeTSAs were predicted to be significantly more immunogenic than a set of presumed non-immunogenic thymic peptides (67) (Figure 6B). Additionally, aeTSAs had significantly higher immunogenicity scores compared to canonical TAAs and the entire coding TA (TSAs and TAAs derived from coding regions). Indeed, according to these predictions, TAAs derived from canonical regions were significantly less immunogenic than thymic peptides (p<0.01). This may be due in part to the small number of validated TAAs. This was no longer the case when these predictions were considered for the entire set of 31 TAAs (Figure 13). While MSI TAAs are predicted to be more immunogenic than thymic peptides when all 31 TAAs are considered, there is also a statistically significant increase in the predicted immunogenicity of TAs derived from MSI tissues compared to MSS.
[0207] Finally, we estimated the proportion of individuals with alleles predicted to bind and present tumor antigens (Figure 6C). Many of the antigens in the samples were highly abundant, and estimates using the IEDB population coverage tool predicted that 80.64% of the US population expresses at least one of the TA-associated alleles identified in this study.
[0208] While the present invention has been described herein with reference to specific embodiments set forth above, modifications can be made without departing from the spirit and nature of the subject invention as defined in the appended claims. In the claims, the word "comprising" is used as open-ended term substantially equivalent to the phrase "including, but not limited to." The singular forms "a," "an," and "the" include the corresponding plural referents unless the context clearly dictates otherwise.
[0209] References 1.Bray,F.,Ferlay,J.,Soerjomataram,I.,Siegel,R.L.,Torre,L.A.,and Jemal,A.(2018)Global cancer statistics 2018:GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries.CA Cancer J Clin 68,394-424 2.Arnold,M.,Sierra,M.S.,Laversanne,M.,Soerjomataram,I.,Jemal,A.,and Bray,F.(2017)Global patterns and trends in colorectal cancer incidence and mortality.Gut 66,683-691 3.Eriksen,A.C.,Sorensen,F.B.,Lindebjerg,J.,Hager,H.,dePont Christensen,R.,Kjaer-Frifeldt,S.,and Hansen,T.F.(2018)The Prognostic Value of Tumor-Infiltrating lymphocytes in Stage II Colon Cancer.A Nationwide Population-Based Study.Transl Oncol 11,979-987 4.Zhao,Y.,Ge,X.,He,J.,Cheng,Y.,Wang,Z.,Wang,J.,and Sun,L.(2019)The prognostic value of tumor-infiltrating lymphocytes in colorectal cancer differs by anatomical subsite:a systematic review and meta-analysis.World J Surg Oncol 17,85 5.Le,D.T.,Uram,J.N.,Wang,H.,Bartlett,B.R.,Kemberling,H.,Eyring,A.D.,Skora,A.D.,Luber,B.S.,Azad,N.S.,Laheru,D.,Biedrzycki,B.,Donehower,R.C.,Zaheer,A.,Fisher,G.A.,Crocenzi,T.S.,Lee,J.J.,Duffy,S.M.,Goldberg,R.M.,de la Chapelle,A.,Koshiji,M.,Bhaijee,F.,Huebner,T.,Hruban,R.H.,Wood,L.D.,Cuka,N.,Pardoll,D.M.,Papadopoulos,N.,Kinzler,K.W.,Zhou,S.,Cornish,T.C.,Taube,J.M.,Anders,R.A.,Eshleman,J.R.,Vogelstein,B.,and Diaz,L.A.,Jr.(2015)PD-1 Blockade in Tumors with Mismatch-Repair Deficiency.N Engl J Med 372,2509-2520 6.Fabrizio,D.A.,George,T.J.,Jr.,Dunne,R.F.,Frampton,G.,Sun,J.,Gowen,K.,Kennedy,M.,Greenbowe,J.,Schrock,A.B.,Hezel,A.F.,Ross,J.S.,Stephens,P.J.,Ali,S.M.,Miller,V.A.,Fakih,M.,and Klempner,S.J.(2018)Beyond microsatellite testing:assessment of tumor mutational burden identifies subsets of colorectal cancer who may respond to immune checkpoint inhibition.J Gastrointest Oncol 9,610-617 7.Wagner,S.,Mullins,C.S.,and Linnebacher,M.(2018)Colorectal cancer vaccines:Tumor-associated antigens vs neoantigens.World J Gastroenterol 24,5418-5432 8.Loffler,M.W.,Kowalewski,D.J.,Backert,L.,Bernhardt,J.,Adam,P.,Schuster,H.,Dengler,F.,Backes,D.,Kopp,H.G.,Beckert,S.,Wagner,S.,Konigsrainer,I.,Kohlbacher,O.,Kanz,L.,Konigsrainer,A.,Rammensee,H.G.,Stevanovic,S.,and Haen,S.P.(2018)Mapping the HLA Ligandome of Colorectal Cancer Reveals an Imprint of Malignant Cell Transformation.Cancer Res 78,4627-4641 9.Picard,E.,Verschoor,C.P.,Ma,G.W.,and Pawelec,G.(2020)Relationships Between Immune Landscapes,Genetic Subtypes and Responses to Immunotherapy in Colorectal Cancer.Front Immunol 11,369 10.Parkhurst,M.R.,Yang,J.C.,Langan,R.C.,Dudley,M.E.,Nathan,D.A.,Feldman,S.A.,Davis,J.L.,Morgan,R.A.,Merino,M.J.,Sherry,R.M.,Hughes,M.S.,Kammula,U.S.,Phan,G.Q.,Lim,R.M.,Wank,S.A.,Restifo,N.P.,Robbins,P.F.,Laurencot,C.M.,and Rosenberg,S.A.(2011)T cells targeting carcinoembryonic antigen can mediate regression of metastatic colorectal cancer but induce severe transient colitis.Mol Ther 19,620-626 11.Minati,R.,Perreault,C.,and Thibault,P.(2020)A Roadmap Toward the Definition of Actionable Tumor-Specific Antigens.Front Immunol 11,583287 12.Smith,C.C.,Selitsky,S.R.,Chai,S.,Armistead,P.M.,Vincent,B.G.,and Serody,J.S.(2019)Alternative tumour-specific antigens.Nat Rev Cancer 19,465-478 13.Kloor,M.,Reuschenbach,M.,Karbach,J.,Rafiyan,M.,Al-Batran,S.-E.,Pauligk,C.,Jaeger,E.,and Doeberitz,M.v.K.(2015)Vaccination of MSI-H colorectal cancer patients with frameshift peptide antigens:A phase I / IIa clinical trial.Journal of Clinical Oncology 33,3020-3020 14.van den Bulk,J.,Verdegaal,E.M.E.,Ruano,D.,Ijsselsteijn,M.E.,Visser,M.,van der Breggen,R.,Duhen,T.,van der Ploeg,M.,de Vries,N.L.,Oosting,J.,Peeters,K.,Weinberg,A.D.,Farina-Sarasqueta,A.,van der Burg,S.H.,and de Miranda,N.(2019)Neoantigen-specific immunity in low mutation burden colorectal cancers of the consensus molecular subtype 4.Genome Med 11,87 15.Newey,A.,Griffiths,B.,Michaux,J.,Pak,H.S.,Stevenson,B.J.,Woolston,A.,Semiannikova,M.,Spain,G.,Barber,L.J.,Matthews,N.,Rao,S.,Watkins,D.,Chau,I.,Coukos,G.,Racle,J.,Gfeller,D.,Starling,N.,Cunningham,D.,Bassani-Sternberg,M.,and Gerlinger,M.(2019)Immunopeptidomics of colorectal cancer organoids reveals a sparse HLA class I neoantigen landscape and no increase in neoantigens with interferon or MEK-inhibitor treatment.J Immunother Cancer 7,309 16.Laumont,C.M.,Vincent,K.,Hesnard,L.,Audemard,E.,Bonneil,E.,Laverdure,J.P.,Gendron,P.,Courcelles,M.,Hardy,M.P.,Cote,C.,Durette,C.,St-Pierre,C.,Benhammadi,M.,Lanoix,J.,Vobecky,S.,Haddad,E.,Lemieux,S.,Thibault,P.,and Perreault,C.(2018)Noncoding regions are the main source of targetable tumor-specific antigens.Sci Transl Med 10 17.Zhao,Q.,Laverdure,J.P.,Lanoix,J.,Durette,C.,Cote,C.,Bonneil,E.,Laumont,C.M.,Gendron,P.,Vincent,K.,Courcelles,M.,Lemieux,S.,Millar,D.G.,Ohashi,P.S.,Thibault,P.,and Perreault,C.(2020)Proteogenomics Uncovers a Vast Repertoire of Shared Tumor-Specific Antigens in Ovarian Cancer.Cancer Immunol Res 8,544-555 18.Ehx,G.,Larouche,J.D.,Durette,C.,Laverdure,J.P.,Hesnard,L.,Vincent,K.,Hardy,M.P.,Theriault,C.,Rulleau,C.,Lanoix,J.,Bonneil,E.,Feghaly,A.,Apavaloaei,A.,Noronha,N.,Laumont,C.M.,Delisle,J.S.,Vago,L.,Hebert,J.,Sauvageau,G.,Lemieux,S.,Thibault,P.,and Perreault,C.(2021)Atypical acute myeloid leukemia-specific transcripts generate shared and immunogenic MHC class-I-associated epitopes.Immunity 19.Bolger,A.M.,Lohse,M.,and Usadel,B.(2014)Trimmomatic:a flexible trimmer for Illumina sequence data.Bioinformatics 30,2114-2120 20.Dobin,A.,Davis,C.A.,Schlesinger,F.,Drenkow,J.,Zaleski,C.,Jha,S.,Batut,P.,Chaisson,M.,and Gingeras,T.R.(2013)STAR:ultrafast universal RNA-seq aligner.Bioinformatics 29,15-21 21.Bray,N.L.,Pimentel,H.,Melsted,P.,and Pachter,L.(2016)Near-optimal probabilistic RNA-seq quantification.Nat Biotechnol 34,525-527 22.Szolek,A.,Schubert,B.,Mohr,C.,Sturm,M.,Feldhahn,M.,and Kohlbacher,O.(2014)OptiType:precision HLA typing from next-generation sequencing data.Bioinformatics 30,3310-3316 23.Jia,P.,Yang,X.,Guo,L.,Liu,B.,Lin,J.,Liang,H.,Sun,J.,Zhang,C.,and Ye,K.(2020)MSIsensor-pro:Fast,Accurate,and Matched-normal-sample-free Detection of Microsatellite Instability.Genomics Proteomics Bioinformatics 18,65-71 24.Love,M.I.,Huber,W.,and Anders,S.(2014)Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2.Genome Biol 15,550 25.Zhou,Y.,Zhou,B.,Pache,L.,Chang,M.,Khodabakhshi,A.H.,Tanaseichuk,O.,Benner,C.,and Chanda,S.K.(2019)Metascape provides a biologist-oriented resource for the analysis of systems-level datasets.Nat Commun 10,1523 26.Hardy,M.P.,Audemard,E.,Migneault,F.,Feghaly,A.,Brochu,S.,Gendron,P.,Boilard,E.,Major,F.,Dieude,M.,Hebert,M.J.,and Perreault,C.(2019)Apoptotic endothelial cells release small extracellular vesicles loaded with immunostimulatory viral-like RNAs.Sci Rep 9,7203 27.Karolchik,D.,Hinrichs,A.S.,Furey,T.S.,Roskin,K.M.,Sugnet,C.W.,Haussler,D.,and Kent,W.J.(2004)The UCSC Table Browser data retrieval tool.Nucleic Acids Res 32,D493-496 28.Cingolani,P.,Platts,A.,Wang le,L.,Coon,M.,Nguyen,T.,Wang,L.,Land,S.J.,Lu,X.,and Ruden,D.M.(2012)A program for annotating and predicting the effects of single nucleotide polymorphisms,SnpEff:SNPs in the genome of Drosophila melanogaster strain w1118;iso-2;iso-3.Fly(Austin)6,80-92 29.Daouda,T.,Perreault,C.,and Lemieux,S.(2016)pyGeno:A Python package for precision medicine and proteogenomics.F1000Res 5,381 30.Lanoix,J.,Durette,C.,Courcelles,M.,Cossette,E.,Comtois-Marotte,S.,Hardy,M.P.,Cote,C.,Perreault,C.,and Thibault,P.(2018)Comparison of the MHC I immunopeptidome repertoir of B-cell lymphoblasts using two isolation methods.Proteomics 18,e1700251 31.Ma,B.,Zhang,K.,Hendrie,C.,Liang,C.,Li,M.,Doherty-Kirby,A.,and Lajoie,G.(2003)PEAKS:powerful software for peptide de novo sequencing by tandem mass spectrometry.Rapid Commun Mass Spectrom 17,2337-2342 32.Courcelles,M.,Durette,C.,Daouda,T.,Laverdure,J.P.,Vincent,K.,Lemieux,S.,Perreault,C.,and Thibault,P.(2020)MAPDP:A Cloud-Based Computational Platform for Immunopeptidomics Analyses.J Proteome Res 19,1873-1881 33.Wu,T.D.,Reeder,J.,Lawrence,M.,Becker,G.,and Brauer,M.J.(2016)GMAP and GSNAP for Genomic Sequence Alignment:Enhancements to Speed,Accuracy,and Functionality.Methods Mol Biol 1418,283-334 34.Ogishi,M.,and Yotsuyanagi,H.(2019)Quantitative Prediction of the Landscape of T Cell Epitope Immunogenicity in Sequence Space.Front Immunol 10,827 35.Conway,J.R.,Lex,A.,and Gehlenborg,N.(2017)UpSetR:an R package for the visualization of intersecting sets and their properties.Bioinformatics 33,2938-2940 36.Hanzelmann,S.,Castelo,R.,and Guinney,J.(2013)GSVA:gene set variation analysis for microarray and RNA-seq data.BMC Bioinformatics 14,7 37.Yoshihara,K.,Shahmoradgoli,M.,Martinez,E.,Vegesna,R.,Kim,H.,Torres-Garcia,W.,Trevino,V.,Shen,H.,Laird,P.W.,Levine,D.A.,Carter,S.L.,Getz,G.,Stemke-Hale,K.,Mills,G.B.,and Verhaak,R.G.(2013)Inferring tumour purity and stromal and immune cell admixture from expression data.Nat Commun 4,2612 38.Chan,A.Y.,and Anderson,M.S.(2015)Central tolerance to self revealed by the autoimmune regulator.Ann N Y Acad Sci 1356,80-89 39.Pfammatter,S.,Bonneil,E.,Lanoix,J.,Vincent,K.,Hardy,M.P.,Courcelles,M.,Perreault,C.,and Thibault,P.(2020)Extending the Comprehensiveness of Immunopeptidome Analyses Using Isobaric Peptide Labeling.Anal Chem 92,9194-9204 40.Pira,G.,Uva,P.,Scanu,A.M.,Rocca,P.C.,Murgia,L.,Uleri,E.,Piu,C.,Porcu,A.,Carru,C.,Manca,A.,Persico,I.,Muroni,M.R.,Sanges,F.,Serra,C.,Dolei,A.,Angius,A.,and De Miglio,M.R.(2020)Landscape of transcriptome variations uncovering known and novel driver events in colorectal carcinoma.Sci Rep 10,432 41.Kawakami,H.,Zaanan,A.,and Sinicrope,F.A.(2015)Microsatellite instability testing and its role in the management of colorectal cancer.Curr Treat Options Oncol 16,30 42.Jimeno,A.,Messersmith,W.A.,Hirsch,F.R.,Franklin,W.A.,and Eckhardt,S.G.(2009)KRAS mutations and sensitivity to epidermal growth factor receptor inhibitors in colorectal cancer:practical application of patient selection.J Clin Oncol 27,1130-1136 43.Van Cutsem,E.,Kohne,C.H.,Lang,I.,Folprecht,G.,Nowacki,M.P.,Cascinu,S.,Shchepotin,I.,Maurel,J.,Cunningham,D.,Tejpar,S.,Schlichting,M.,Zubel,A.,Celik,I.,Rougier,P.,and Ciardiello,F.(2011)Cetuximab plus irinotecan,fluorouracil,and leucovorin as first-line treatment for metastatic colorectal cancer:updated analysis of overall survival according to tumor KRAS and BRAF mutation status.J Clin Oncol 29,2011-2019 44.Ahmed,D.,Eide,P.W.,Eilertsen,I.A.,Danielsen,S.A.,Eknaes,M.,Hektoen,M.,Lind,G.E.,and Lothe,R.A.(2013)Epigenetic and genetic features of 24 colon cancer cell lines.Oncogenesis 2,e71 45.Berg,K.C.G.,Eide,P.W.,Eilertsen,I.A.,Johannessen,B.,Bruun,J.,Danielsen,S.A.,Bjornslett,M.,Meza-Zepeda,L.A.,Eknaes,M.,Lind,G.E.,Myklebost,O.,Skotheim,R.I.,Sveen,A.,and Lothe,R.A.(2017)Multi-omics of 34 colorectal cancer cell lines - a resource for biomedical studies.Mol Cancer 16,116 46.Niu,B.,Ye,K.,Zhang,Q.,Lu,C.,Xie,M.,McLellan,M.D.,Wendl,M.C.,and Ding,L.(2014)MSIsensor:microsatellite instability detection using paired tumor-normal sequence data.Bioinformatics 30,1015-1016 47.Aaltonen,L.A.,Peltomaki,P.,Mecklin,J.P.,Jarvinen,H.,Jass,J.R.,Green,J.S.,Lynch,H.T.,Watson,P.,Tallqvist,G.,Juhola,M.,and et al.(1994)Replication errors in benign and malignant tumors from hereditary nonpolyposis colorectal cancer patients.Cancer Res 54,1645-1648 48.Llosa,N.J.,Cruise,M.,Tam,A.,Wicks,E.C.,Hechenbleikner,E.M.,Taube,J.M.,Blosser,R.L.,Fan,H.,Wang,H.,Luber,B.S.,Zhang,M.,Papadopoulos,N.,Kinzler,K.W.,Vogelstein,B.,Sears,C.L.,Anders,R.A.,Pardoll,D.M.,and Housseau,F.(2015)The vigorous immune microenvironment of microsatellite instable colon cancer is balanced by multiple counter-inhibitory checkpoints.Cancer Discov 5,43-51 49.Le,D.T.,Durham,J.N.,Smith,K.N.,Wang,H.,Bartlett,B.R.,Aulakh,L.K.,Lu,S.,Kemberling,H.,Wilt,C.,Luber,B.S.,Wong,F.,Azad,N.S.,Rucki,A.A.,Laheru,D.,Donehower,R.,Zaheer,A.,Fisher,G.A.,Crocenzi,T.S.,Lee,J.J.,Greten,T.F.,Duffy,A.G.,Ciombor,K.K.,Eyring,A.D.,Lam,B.H.,Joe,A.,Kang,S.P.,Holdhoff,M.,Danilova,L.,Cope,L.,Meyer,C.,Zhou,S.,Goldberg,R.M.,Armstrong,D.K.,Bever,K.M.,Fader,A.N.,Taube,J.,Housseau,F.,Spetzler,D.,Xiao,N.,Pardoll,D.M.,Papadopoulos,N.,Kinzler,K.W.,Eshleman,J.R.,Vogelstein,B.,Anders,R.A.,and Diaz,L.A.,Jr.(2017)Mismatch repair deficiency predicts response of solid tumors to PD-1 blockade.Science 357,409-413 50.Hanahan,D.,and Weinberg,R.A.(2000)The hallmarks of cancer.Cell 100,57-70 51.Hanahan,D.,and Weinberg,R.A.(2011)Hallmarks of cancer:the next generation.Cell 144,646-674 52.Prossomariti,A.,Piazzi,G.,Alquati,C.,and Ricciardiello,L.(2020)Are Wnt / beta-Catenin and PI3K / AKT / mTORC1 Distinct Pathways in Colorectal Cancer? Cell Mol Gastroenterol Hepatol 10,491-506 53.Danaher,P.,Warren,S.,Dennis,L.,D’Amico,L.,White,A.,Disis,M.L.,Geller,M.A.,Odunsi,K.,Beechem,J.,and Fling,S.P.(2017)Gene expression markers of Tumor Infiltrating Leukocytes.J Immunother Cancer 5,18 54.Barbie,D.A.,Tamayo,P.,Boehm,J.S.,Kim,S.Y.,Moody,S.E.,Dunn,I.F.,Schinzel,A.C.,Sandy,P.,Meylan,E.,Scholl,C.,Frohling,S.,Chan,E.M.,Sos,M.L.,Michel,K.,Mermel,C.,Silver,S.J.,Weir,B.A.,Reiling,J.H.,Sheng,Q.,Gupta,P.B.,Wadlow,R.C.,Le,H.,Hoersch,S.,Wittner,B.S.,Ramaswamy,S.,Livingston,D.M.,Sabatini,D.M.,Meyerson,M.,Thomas,R.K.,Lander,E.S.,Mesirov,J.P.,Root,D.E.,Gilliland,D.G.,Jacks,T.,and Hahn,W.C.(2009)Systematic RNA interference reveals that oncogenic KRAS-driven cancers require TBK1.Nature 462,108-112 55.Kim,H.,Jen,J.,Vogelstein,B.,and Hamilton,S.R.(1994)Clinical and pathological characteristics of sporadic colorectal carcinomas with DNA replication errors in microsatellite sequences.Am J Pathol 145,148-156 56.Smyrk,T.C.,Watson,P.,Kaul,K.,and Lynch,H.T.(2001)Tumor-infiltrating lymphocytes are a marker for microsatellite instability in colorectal carcinoma.Cancer 91,2417-2422 57.Dolcetti,R.,Viel,A.,Doglioni,C.,Russo,A.,Guidoboni,M.,Capozzi,E.,Vecchiato,N.,Macri,E.,Fornasarig,M.,and Boiocchi,M.(1999)High prevalence of activated intraepithelial cytotoxic T lymphocytes and increased neoplastic cell apoptosis in colorectal carcinomas with microsatellite instability.Am J Pathol 154,1805-1813 58.Phillips,S.M.,Banerjea,A.,Feakins,R.,Li,S.R.,Bustin,S.A.,and Dorudi,S.(2004)Tumour-infiltrating lymphocytes in colorectal cancer with microsatellite instability are activated and cytotoxic.Br J Surg 91,469-475 59.Boland,C.R.,Koi,M.,Chang,D.K.,and Carethers,J.M.(2008)The biochemical basis of microsatellite instability and abnormal immunohistochemistry and clinical behavior in Lynch syndrome:from bench to bedside.Fam Cancer 7,41-52 60.Sarkizova,S.,Klaeger,S.,Le,P.M.,Li,L.W.,Oliveira,G.,Keshishian,H.,Hartigan,C.R.,Zhang,W.,Braun,D.A.,Ligon,K.L.,Bachireddy,P.,Zervantonakis,I.K.,Rosenbluth,J.M.,Ouspenskaia,T.,Law,T.,Justesen,S.,Stevens,J.,Lane,W.J.,Eisenhaure,T.,Lan Zhang,G.,Clauser,K.R.,Hacohen,N.,Carr,S.A.,Wu,C.J.,and Keskin,D.B.(2020)A large peptidome dataset improves HLA class I epitope prediction across most of the human population.Nat Biotechnol 38,199-209 61.Larouche,J.D.,Trofimov,A.,Hesnard,L.,Ehx,G.,Zhao,Q.,Vincent,K.,Durette,C.,Gendron,P.,Laverdure,J.P.,Bonneil,E.,Cote,C.,Lemieux,S.,Thibault,P.,and Perreault,C.(2020)Widespread and tissue-specific expression of endogenous retroelements in human somatic tissues.Genome Med 12,40 62.Cherkasova,E.,Scrivani,C.,Doh,S.,Weisman,Q.,Takahashi,Y.,Harashima,N.,Yokoyama,H.,Srinivasan,R.,Linehan,W.M.,Lerman,M.I.,and Childs,R.W.(2016)Detection of an Immunogenic HERV-E Envelope with Selective Expression in Clear Cell Kidney Cancer.Cancer Res 76,2177-2185 63.Patra,R.,Das,N.C.,and Mukherjee,S.(2021)Exploring the Differential Expression and Prognostic Significance of the COL11A1 Gene in Human Colorectal Carcinoma:An Integrated Bioinformatics Approach.Front Genet 12,608313 64.Marcu,A.,Bichmann,L.,Kuchenbecker,L.,Kowalewski,D.J.,Freudenmann,L.K.,Backert,L.,Muehlenbruch,L.,Szolek,A.,Luebke,M.,Wagner,P.,Engler,T.,Matovina,S.,Wang,J.,Hauri-Hohl,M.,Martin,R.,Kapolou,K.,Walz,J.S.,Velz,J.,Moch,H.,Regli,L.,Silginer,M.,Weller,M.,Loeffler,M.W.,Erhard,F.,Schlosser,A.,Kohlbacher,O.,Stevanovic,S.,Rammensee,H.-G.,and Neidert,M.C.(2020)The HLA Ligand Atlas - A resource of natural HLA ligands presented on benign tissues.bioRxiv,778944 65.Gjerstorff,MF,Andersen,MH,and Ditzel,HJ(2015)Oncogenic cancer / testis antigens:prime candidates for immunotherapy.Oncotarget 6,15772-15787 66. Ishikawa, N., Takano, A., Yasui, W., Inai, K., Nishimura, H., Ito, H., Miyagi, Y., Nakayama, H., Fujita, M., Hosokawa, M., Tsuchiya, E., Kohno, N., Nakamura, Y., and Daigo, Y. (2007)Cancer-testis Cancer Res 67,11601–11611: a serologic biomarker of the antigen lymphocyte antigen 6 complex locus K and a therapeutic target for lung and esophageal carcinomas 67.Adamopoulou, E., Tenzer, S., Hillen, N., Klug, P., Rota, IA, Tietz, S., Gebhardt, M., Stevanovic, S., Schild, H., Tolosa, E., Melms, A., and Stoeckle, C. (2013). human thymus.Nat Commun 4,2039 68. Kote, S., Pirog, A., Bedran, G., Alfaro, J., and Dapic, I. (2020)Mass Spectrometry-Based Identification of MHC-Associated Peptides.Cancers(Basel)1 69.Lin,A.,Zhang,J.,and Luo,P.(2020)Crosstalk Between the MSI Status and Tumor Microenvironment in Colorectal Cancer.Front Immunol 11,2039 70.Bonaventura,P.,Shekarian,T.,Alcazer,V.,Valladeau-Guilemond,J.,Valsesia-Wittmann,S.,Amigorena,S.,Caux,C.,and Depil,S.(2019)Cold Tumors:A Therapeutic Challenge for Immunotherapy.Front Immunol 10,168 71.Shihab,H.A.,Gough,J.,Cooper,D.N.,Stenson,P.D.,Barker,G.L.,Edwards,K.J.,Day,I.N.,and Gaunt,T.R.(2013)Predicting the functional,molecular,and phenotypic consequences of amino acid substitutions using hidden Markov models.Hum Mutat 34,57-65 72.Rentzsch,P.,Witten,D.,Cooper,G.M.,Shendure,J.,and Kircher,M.(2019)CADD:predicting the deleteriousness of variants throughout the human genome.Nucleic Acids Res 47,D886-D894 73.Sherry,S.T.,Ward,M.,and Sirotkin,K.(1999)dbSNP-database for single nucleotide polymorphisms and other classes of minor genetic variation.Genome Res 9,677-679 74.Tate,J.G.,Bamford,S.,Jubb,H.C.,Sondka,Z.,Beare,D.M.,Bindal,N.,Boutselakis,H.,Cole,C.G.,Creatore,C.,Dawson,E.,Fish,P.,Harsha,B.,Hathaway,C.,Jupe,S.C.,Kok,C.Y.,Noble,K.,Ponting,L.,Ramshaw,C.C.,Rye,C.E.,Speedy,H.E.,Stefancsik,R.,Thompson,S.L.,Wang,S.,Ward,S.,Campbell,P.J.,and Forbes,S.A.(2019)COSMIC:the Catalogue Of Somatic Mutations In Cancer.Nucleic Acids Res 47,D941-D947 75.Ehx,G.,and Perreault,C.(2019)Discovery and characterization of actionable tumor antigens.Genome Med 11,29 76.Vogel,C.,and Marcotte,E.M.(2012)Insights into the regulation of protein abundance from proteomic and transcriptomic analyses.Nat Rev Genet 13,227-232 77.Xiang,B.,Snook,A.E.,Magee,M.S.,and Waldman,S.A.(2013)Colorectal cancer immunotherapy.Discov Med 15,301-308 78.Gold,P.,and Freedman,S.O.(1965)Demonstration of Tumor-Specific Antigens in Human Colonic Carcinomata by Immunological Tolerance and Absorption Techniques.J Exp Med 121,439-462 79.Zhou,F.(2009)Molecular mechanisms of IFN-gamma to up-regulate MHC class I antigen processing and presentation.Int Rev Immunol 28,239-260 80.Deutsch,EW,Bandeira,N,Sharma,V,Perez-Riverol,Y,Carver,JJ,Kundu,DJ,Garcia-Sixfingers,D,Jarnuczak,AF,Hewapathirana ,S.,Pullman,BS,Wertz,J,Sun,Z,Kawano,S,Okuda,S,Watanabe,Y,Hermjakob,H,MacLean,B,MacCoss,MJ,Zhu,Y,Ishihama,Y,and Vizcaino,JA(2020)The ProteomeXchange Consortium in 2020:enabling 'big data' approaches in proteomics.Nucleic Acids Res 48,D1145-D1152 81.Perez-Riverol,Y.,Csordas,A.,Bai,J.,Bernal-Llinares,M.,Hewapathirana,S.,Kundu,DJ,Inuganti,A.,Griss,J.,Mayer,G.,Eisenacher,M.,Per ez , E. , Uszkoreit , J. , Pfeuffer , J. , Sachsenberg , T. , Yilmaz , S. , Tiwary , S. , Cox , J. , Audain , E. , Walzer , M. , Jarnuczak , AF , Ternent , T. , Brazma , A. , et al Vizcaino, JA (2019) Nucleic Acids Res 47, D442-D450: Improving Support for Quantification Data.
Claims
1. A tumor antigen peptide (TAP) comprising one of the amino acid sequences defined in SEQ ID NOs: 1 to 23.
2. 2. The TAP of claim 1, wherein the TAP comprises one of the sequences defined in SEQ ID NOs: 1-17.
3. (a) binds to HLA-A*02:01 molecules and comprises the sequence of SEQ ID NO: 6; (b) binds to HLA-A*03:01 molecules and comprises the sequence of SEQ ID NO: 1, 11, or 14; (c) binds to HLA-A*03:02 molecules and comprises the sequence of SEQ ID NO: 3, 5, 7, 16, or 23; (d) binds to HLA-A*11:01 molecules and comprises the sequence of SEQ ID NO: 9 or 18; (e) binds to HLA-A*30:01 molecules and comprises the sequence of SEQ ID NO: 19, 20, or 23; (f) binds to HLA-A*32:01 (g) binds to an HLA-B*07:02 molecule and comprises the sequence of SEQ ID NO: 2 or 21; (h) binds to an HLA-B*13:02 molecule and comprises the sequence of SEQ ID NO: 13; (i) binds to an HLA-B*27:05 molecule and comprises the sequence of SEQ ID NO: 4; (j) binds to an HLA-B*52:01 molecule and comprises the sequence of SEQ ID NO: 10, 12 or 15; or (k) binds to an HLA-C*06:02 molecule and comprises the sequence of SEQ ID NO:
17.
4. 2. The TAP of claim 1, wherein the TAP is encoded by a sequence or long non-coding RNA located in a non-protein-coding region of the genome.
5. 5. The TAP of claim 4, wherein the non-protein-coding region of the genome is an untranslated transcribed region (UTR), an intron, or an intergenic region.
6. A combination comprising at least two of the TAPs defined in claim 1.
7. A synthetic long peptide (SLP) comprising at least one of the amino acid sequences defined in claim 1.
8. A nucleic acid encoding one or more TAPs described in claim 1.
9. The nucleic acid described in claim 8, wherein the nucleic acid is mRNA.
10. The nucleic acid described in claim 8, wherein the nucleic acid is DNA.
11. A vesicle or particle comprising a TAP, a nucleic acid, a combination or an SLP according to any one of claims 1 to 10.
12. The vesicle or particle of claim 11 , wherein the vesicle is a lipid nanoparticle (LNP).
13. A composition comprising the TAP, nucleic acid, combination or SLP according to any one of claims 1 to 10 and a pharmaceutically acceptable carrier.
14. A vaccine comprising the TAP, nucleic acid, combination or SLP according to any one of claims 1 to 10 and an adjuvant.
15. An isolated major histocompatibility complex (MHC) class I molecule, comprising a TAP according to any one of claims 1 to 5 in its peptide-binding groove.
16. 10. An isolated cell expressing on its surface major histocompatibility complex (MHC) class I molecules, said MHC class I molecules comprising in their peptide-binding groove a TAP according to any one of claims 1 to 5 or a combination according to claim 6.
17. The cell of claim 16, which is an antigen-presenting cell (APC).
18. The cell of claim 17 , wherein the APC is a dendritic cell.
19. A T cell receptor (TCR), or an antibody or antigen-binding fragment thereof, that specifically recognizes an MHC class I molecule expressed on the surface of the cell described in claim 16.
20. 20. The antibody or antigen-binding fragment thereof of claim 19, which is a bispecific antibody or antigen-binding fragment thereof.
21. 21. The antibody or antigen-binding fragment thereof of claim 20, wherein the bispecific antibody or antigen-binding fragment thereof is a single-chain diabody (scDb).
22. 21. The antibody or antigen-binding fragment thereof of claim 20, wherein the bispecific antibody or antigen-binding fragment thereof also specifically binds to a T cell signaling molecule.
23. 23. The antibody or antigen-binding fragment thereof of claim 22, wherein the T cell signaling molecule is a CD3 chain.
24. 20. An isolated cell, which expresses the TCR of claim 19 on its cell surface.
25. CD8 + 25. The isolated cell of claim 24, which is a T lymphocyte.
26. 25. A cell population comprising at least 0.5% isolated cells as defined in claim 24.
27. 1. A medicament for use in treating colorectal cancer (CRC) in a subject, said medicament comprising: (a) a TAP comprising or consisting of one of the amino acid sequences defined in SEQ ID NOs: 1-23 and 25-50, or a synthetic long peptide (SLP) comprising at least one of the sequences set forth in SEQ ID NOs: 1-23 and 25-50; (b) at least one nucleic acid encoding the TAP, combination thereof, or SLP defined in (a); (c) a vesicle or particle comprising the TAP, a combination thereof, or an SLP as defined in (a), or the at least one nucleic acid as defined in (b); (d) a composition comprising the TAP, combination thereof, or SLP defined in (a), the at least one nucleic acid defined in (b), or the vesicle or particle defined in (c), and a pharmaceutically acceptable carrier; (e) a vaccine comprising the TAP, combination thereof, or SLP defined in (a), the at least one nucleic acid defined in (b), the vesicle or particle defined in (c), or the composition defined in (d), and an adjuvant; (f) a cell expressing at its cell surface a major histocompatibility complex (MHC) class I molecule comprising in its peptide-binding groove the TAP or a combination thereof as defined in (a); (g) a cell that expresses on its cell surface a T cell receptor (TCR) that specifically recognizes an MHC class I molecule expressed on the surface of the cell as defined in (f); or (h) An agent that is a soluble TCR, an antibody, or an antigen-binding fragment thereof that specifically binds to an MHC class I molecule expressed on the surface of the cell defined in (f).
28. 28. The method of claim 27, wherein the CRC is colon cancer.
29. 28. The method of claim 27, wherein the CRC is rectal cancer.
30. 28. The method of claim 27, further comprising administering to said subject at least one additional anti-tumor agent or therapy.
31. 31. The drug for use of claim 30, wherein the at least one additional anti-tumor agent or therapy is a chemotherapeutic agent, immunotherapy, immune checkpoint inhibitor, radiation therapy, or surgery.