Neoantigen compositions and uses thereof

The polypeptide, designed to enhance epitope processing and presentation by APCs, addresses the inefficiencies in current cancer therapeutic vaccines, resulting in improved immune responses and antitumor activity.

JP2025085674AInactive Publication Date: 2025-06-05BIONTECH US INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025038617
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-06-12
Filing Date
2025-03-11
Publication Date
2025-06-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Current cancer therapeutic vaccines face challenges in efficiently processing and presenting minimal epitopes for effective immune response generation.

Method used

A polypeptide comprising an epitope presented by class I or class II MHC of an antigen presenting cell (APC), with a specific formula and structure that enhances epitope processing and presentation, including the use of linkers and specific amino acid sequences.

Benefits of technology

The described polypeptide enhances epitope presentation and processing by APCs, leading to improved immune responses and antitumor activity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025085674000043
    Figure 2025085674000043
  • Figure 2025085674000044
    Figure 2025085674000044
  • Figure 2025085674000045
    Figure 2025085674000045
Patent Text Reader

Abstract

To provide neoantigen compositions and uses thereof.SOLUTION: The present disclosure relates to immunotherapeutic polypeptides comprising neoepitopes, antigen presenting cells comprising the immunotherapeutic polypeptides, and a pharmaceutical composition comprising the immunotherapeutic polypeptides. Also disclosed herein is use of the immunotherapeutic polypeptides in treating a disease or condition. The present disclosure provides a method for activating, promoting, enhancing, and / or augmenting an immune response using a polypeptide, a cell, or a pharmaceutical composition comprising a neoantigenic peptide or protein described herein.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 62 / 860,493, filed June 12, 2019, which is incorporated by reference in its entirety. This application is related to International Application No. PCT / US2020 / 031898, filed May 7, 2020, which is incorporated by reference in its entirety. [Background technology]

[0002] background Cancer immunotherapy is the use of the immune system to treat cancer. Immunotherapy exploits the fact that cancer cells often have molecules on their surface that can be detected by the immune system. These molecules are known as tumor antigens and are often proteins or other macromolecules (e.g. carbohydrates). Active immunotherapy directs the immune system to attack tumor cells by targeting tumor antigens. Passive immunotherapy enhances existing antitumor responses and includes the use of monoclonal antibodies, lymphocytes, and cytokines. Tumor vaccines are generally composed of tumor antigens and immunostimulatory molecules (e.g. adjuvants, cytokines, or Toll-like receptor (TLR) ligands) that work together to induce antigen-specific cytotoxic T cells (CTLs) that recognize and lyse tumor cells. Tumor neoantigens, which arise as a result of genetic alterations in malignant cells (e.g. inversions, translocations, deletions, missense mutations, splice site mutations, etc.), represent the most tumor-specific class of antigens and can be patient-specific or shared. Tumor neoantigens are unique to tumor cells because the mutations and their corresponding proteins are only present in tumors. Tumor neoantigens also circumvent central immune tolerance and are therefore more likely to be immunogenic. Thus, tumor neoantigens provide excellent targets for immune recognition, including both humoral and cellular immunity.

[0003] In order to elicit T cell response by vaccination, epitope-containing peptides must be processed by antigen-presenting cells (APCs) and epitopes must be presented on major histocompatibility complex (MHC) I or MHC II. One of the crucial barriers for the development of curative and tumor-specific immunotherapy is the insufficient processing and release of minimal epitopes for antigen presentation to generate proper immune response. Therefore, there is a need to develop additional cancer therapeutic vaccines to ensure efficient and sufficient epitope processing and presentation.

[0004] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. Summary of the Invention [Means for solving the problem]

[0005] overview In some embodiments, a polypeptide comprising an epitope presented by class I or class II MHC of an antigen presenting cell (APC) is provided, the polypeptide comprising the formula (I): Y n -B t -A r -X m -A s -C u -Z p Formula (I), or a pharma- ceutically acceptable salt thereof, (i)X m is an epitope, each X independently represents an amino acid in a contiguous amino acid sequence encoded by a nucleic acid sequence in the genome of a subject, and (a) the MHC is a class I MHC and m is an integer between 8 and 12; or (b) the MHC is a class II MHC and m is an integer from 9 to 25; (ii) each Y is independently an amino acid, an analog, or a derivative thereof; and (A) A of formula (I) r If the variable r in is 0, then Y n is B t -A r -X m is not encoded by a nucleic acid sequence immediately upstream of a nucleic acid sequence in the subject's genome encoding (B) A of formula (I) r The variable r is 1, and B in formula (I) t If the variable t in is 0, then Y n X m or (C) A of formula (I) r The variable r is 1, and B in formula (I) t If the variable t in is 1 or greater, then Y n is B t is not encoded by a nucleic acid sequence immediately upstream of a nucleic acid sequence in the subject's genome encoding further, n is an integer between 0 and 1000; (iii) each Z is independently an amino acid, an analog, or a derivative thereof; and (A) A of formula (I) s If the variable s in is 0, Z p X m -A s -C u is not encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the subject's genome encoding (B) A of formula (I) s The variable s in formula (I) is 1, u If the variable u in is 0, then Z p X m or (C) A of formula (I) s The variable s in formula (I) is 1, u If the variable u in is 1 or greater, then Z p is C uis not encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the subject's genome encoding Further, p is an integer from 0 to 1000; Further, when n is 0, p is an integer from 1 to 1000; When p is 0, n is an integer from 1 to 1000; (iv)A r is a linker and r is 0 or 1; (v)A s is a linker and s is 0 or 1; (vi) Each B independently m represents an amino acid encoded by a nucleic acid sequence in the subject's genome that is immediately upstream of a nucleic acid sequence in the subject's genome that encodes t is an integer between 0 and 1000; and (vii) Each C independently selects X m represents an amino acid encoded by a nucleic acid sequence in the subject's genome that is immediately downstream of a nucleic acid sequence in the subject's genome that encodes u is an integer between 0 and 1000; moreover, (a) the polypeptide does not consist of four distinct epitopes presented by class I MHC; (b) the polypeptide comprises at least two distinct polypeptide molecules; (c) the epitope comprises at least one mutant amino acid; and / or (d) When the polypeptide is processed by the APC, Y n and / or Z p is cleaved from the epitope, Polypeptides are provided herein.

[0006] In some embodiments, the epitope is presented by class II MHC. In some embodiments, m is an integer from 9 to 25. In some embodiments, t is 1, 2, 3, 4, or 5 or more and r is 0. In some embodiments, u is 1, 2, 3, 4, or 5 or more and s is 0. In some embodiments, t is 1 or more, r is 0 and n is 1-1000. In some embodiments, u is 1 or more, s is 0 and p is 1-1000. In some embodiments, t is 0. In some embodiments, u is 0. In some embodiments, t is at least 1 and B t In some embodiments, u is at least 1 and C u In some embodiments, the polypeptide, upon processing by the APC, becomes B t is cleaved from the epitope. In some embodiments, when the polypeptide is processed by the APC, it becomes u is cleaved from the epitope. In some embodiments, n is an integer from 1 to 5 or from 7 to 1000. In some embodiments, p is an integer from 1 to 4 or from 6 to 1000.

[0007] In some embodiments, the polypeptide does not consist of four different epitopes presented by class I MHC. In some embodiments, the polypeptide does not include four different epitopes presented by class I MHC. In some embodiments, the polypeptide comprises at least two different polypeptide molecules. In some embodiments, the epitope comprises at least one mutant amino acid. In some embodiments, the at least one mutant amino acid is encoded by an insertion, deletion, frameshift, neo-ORF, or point mutation in a nucleic acid sequence in the genome of the subject. In some embodiments, the polypeptide, upon processing by the APC, becomes a Y. n and / or Z p is cleaved from the epitope. m m is at least 8, and x mAA 1 AA 2 AA 3 AA 4 AA 5 AA 6 AA 7 AA 8 AA 9 AA 10 AA 11 AA 12 AA 13 AA 14 AA 15 AA 16 AA 17 AA 18 AA 19 AA 20 AA 21 AA 22 AA 23 AA 24 AA 25 where each AA is an amino acid, and 9 , A.A. 10 , A.A. 11 , A.A. 12 , A.A. 13 , A.A. 14 , A.A. 15 , A.A. 16 , A.A. 17 , A.A. 18 , A.A. 19 , A.A. 20 , A.A. 21 , A.A. 22 , A.A. 23 , A.A. 24 , and A.A. 25 are optionally present, and further wherein at least one AA is a mutated amino acid. In some embodiments, r is 1. In some embodiments, s is 1. In some embodiments, r is 1 and s is 1. In some embodiments, r is 0. In some embodiments, s is 0. In some embodiments, r is 0 and s is 0.

[0008] In some embodiments, A r and / or A s is a non-polypeptide linker. In some embodiments, A r and / or A sis a chemical linker. In some embodiments, A r and / or A s In some embodiments, A comprises a non-natural amino acid. r and / or A s In some embodiments, A does not contain any amino acid. r and / or A s does not contain a naturally occurring amino acid. r and / or A s In some embodiments, A contains bonds other than peptide bonds. r and / or A s In some embodiments, A r and A s In some embodiments, A r and A s is the same.

[0009] In some embodiments, the polypeptide comprises a hydrophilic tail. n -B t -A r and / or A s -C u -Z p The solubility of the polypeptide is increased by Y n -B t -A r and / or A s -Z p In some embodiments, the peptide is enhanced in comparison to a corresponding peptide that does not contain X. m Each X is a naturally occurring amino acid.

[0010] In some embodiments, the epitope becomes Y upon processing of the polypeptide by the APC. n -B t -A r and / or A s -C u -Z p In some embodiments, the polypeptide is released from A r and / or A sIn some embodiments, the polypeptide is cleaved at X m and X m and / or the polypeptide is cleaved at a rate greater than that of a corresponding polypeptide of the same length that contains at least one additional amino acid encoded by a nucleic acid sequence immediately upstream of the nucleic acid sequence in the subject's genome that encodes X m and X m The polypeptide is cleaved at a rate greater than that of a corresponding polypeptide of the same length that contains at least one additional amino acid encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the subject's genome that encodes the polypeptide.

[0011] In some embodiments, the polypeptide comprises a polypeptide having a structure similar to that of B, where n is an integer from 1 to 1000. t -X m wherein t is at least 1 and the variable A in formula (I) is r and / or, wherein r is 0; and / or, wherein p is an integer from 1 to 1000, X m -C u wherein u is at least 1 and the variable A in formula (I) is s The s is 0.

[0012] In some embodiments, the polypeptide comprises, where n is an integer from 1 to 1000, X m and X m A higher percentage of cleavage of a corresponding polypeptide of the same length that contains at least one additional amino acid encoded by a nucleic acid sequence immediately upstream of the nucleic acid sequence in the subject's genome that encodes A. r and / or the polypeptide is cleaved at X m and X mA higher rate of cleavage of a corresponding polypeptide of the same length that contains at least one additional amino acid encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the subject's genome that encodes A s is cut off at

[0013] In some embodiments, epitope presentation by APCs is determined by: X m and X m and / or, when p is an integer between 1 and 1000, epitope presentation by APCs is enhanced compared to epitope presentation of a corresponding polypeptide of the same length that contains at least one additional amino acid encoded by a nucleic acid sequence immediately upstream of the nucleic acid sequence in the subject's genome that encodes X. m and X m The epitope presentation is enhanced compared to that of a corresponding polypeptide of the same length that contains at least one additional amino acid encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the subject's genome that encodes the polypeptide.

[0014] In some embodiments, when n is an integer between 1 and 1000, epitope presentation by APCs is t -X m where t is at least 1 and the variable A in formula (I) is r and / or, epitope presentation by APCs is determined by: m -C u where u is at least 1 and the variable A in formula (I) is s The s is 0.

[0015] In some embodiments, the epitope is presented by the APC to immune cells. In some embodiments, the epitope is presented by the APC to phagocytes. In some embodiments, the epitope is presented by the APC to dendritic cells, macrophages, mast cells, neutrophils, or monocytes. In some embodiments, the epitope is presented by the APC preferentially or specifically to immune cells, phagocytes, dendritic cells, macrophages, mast cells, neutrophils, or monocytes.

[0016] In some embodiments, when n is an integer from 1 to 1000, the immunogenicity is m and X m and / or, where p is an integer between 1 and 1000, the immunogenicity is enhanced compared to the immunogenicity of a corresponding polypeptide of the same length that contains at least one additional amino acid encoded by a nucleic acid sequence immediately upstream of the nucleic acid sequence in the subject's genome that encodes X. m and X m The immunogenicity is enhanced compared to that of a corresponding polypeptide of the same length that contains at least one additional amino acid encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the subject's genome that encodes the polypeptide.

[0017] In some embodiments, when n is an integer from 1 to 1000, the immunogenicity is t -X m where t is at least 1 and the variable A in formula (I) is r When r is 0 and / or p is an integer from 1 to 1000, the immunogenicity is m -C u where u is at least 1 and the variable A in formula (I) is s The s is 0.

[0018] In some embodiments, the antitumor activity is greater than or equal to X, where n is an integer from 1 to 1000. m and X mand / or, when p is an integer between 1 and 1000, the antitumor activity is enhanced compared to the antitumor activity of a corresponding polypeptide of the same length that contains at least one additional amino acid encoded by a nucleic acid sequence immediately upstream of the nucleic acid sequence in the subject's genome that encodes X m and X m The anti-tumor activity of the polypeptide is enhanced compared to that of a corresponding polypeptide of the same length that contains at least one additional amino acid encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the subject's genome that encodes the polypeptide.

[0019] In some embodiments, when n is an integer from 1 to 1000, the antitumor activity is t -X m wherein t is at least 1 and the variable A in formula (I) is an antitumor activity of a polypeptide of the same length. r When r is 0 and / or p is an integer from 1 to 1000, the antitumor activity is m -C u wherein u is at least 1 and the variable A in formula (I) is s The s is 0.

[0020] In some embodiments, Y n and / or Z p comprises a sequence selected from the group consisting of poly-Lys (poly K) and poly-Arg (poly R). n and / or Z p comprises a sequence selected from the group consisting of polyK-AA-AA and polyR-AA-AA, where each AA is an amino acid or an analog or derivative thereof. In some embodiments, polyK comprises poly-L-Lys. In some embodiments, polyR comprises poly-L-Arg. In some embodiments, polyK or polyR comprises at least three or four consecutive lysine or arginine residues, respectively. In some embodiments, A r and / or A sis selected from the group consisting of disulfide; p-aminobenzyloxycarbonyl (PABC); and AA-AA-PABC, where each AA is an amino acid or an analog or derivative thereof. In some embodiments, AA-AA-PABC is selected from the group consisting of Ala-Lys-PABC, Val-Cit-PABC, and Phe-Lys-PABC.

[0021] In some embodiments, A r and / or A s teeth, [ka] It is.

[0022] In some embodiments, A r and / or A s teeth, [ka] where R 1 and R 2 are independently H or (C 1 ~C 6 ) alkyl; j is 1 or 2; G 1 is H or COOH; i is 1, 2, 3, 4, or 5.

[0023] In some embodiments, the polypeptide is ubiquitinated. In some embodiments, the polypeptide is ubiquitinated before cleavage. In some embodiments, the polypeptide is ubiquitinated at a lysine residue. In some embodiments, the polypeptide is not cleaved in the subject prior to processing by or internalization by APCs. In some embodiments, the polypeptide is not cleaved in the subject's blood prior to processing by or internalization by APCs. In some embodiments, the polypeptide is not cleaved by proteases in the blood. In some embodiments, the polypeptide is not cleaved by plasmin, plasma kallikrein, tissue kallikrein, thrombin, or clotting factors. In some embodiments, the polypeptide is stable in human plasma. In some embodiments, the polypeptide has a half-life in human plasma of 1 hour to 5 days. In some embodiments, the polypeptide is cleaved in a lysosome, endolysosome, endosome, or endoplasmic reticulum (ER). In some embodiments, the polypeptide is cleaved by aminopeptidases. In some embodiments, the aminopeptidase is insulin-regulated aminopeptidase (IRAP) or endoplasmic reticulum aminopeptidase (ERAP). In some embodiments, the polypeptide is processed by a trypsin-like domain of a proteasome and / or an immunoproteasome. In some embodiments, the trypsin-like domain comprises trypsin-like activity, chymotrypsin-like activity, or peptidylglutamyl-peptide hydrolase (PGPH) activity. In some embodiments, the polypeptide is cleaved by a protease. In some embodiments, the protease is a trypsin-like protease, a chymotrypsin-like protease, or a peptidylglutamyl-peptide hydrolase (PGPH). In some embodiments, the protease is selected from the group consisting of aspartic peptide lyase, aspartic acid protease, cysteine ​​protease, glutamic acid protease, metalloprotease, serine protease, and threonine protease.In some embodiments, the protease is a cysteine ​​protease selected from the group consisting of calpain, caspase, cathepsin B, cathepsin C, cathepsin F, cathepsin H, cathepsin K, cathepsin L1, cathepsin L2, cathepsin O, cathepsin S, cathepsin W, and cathepsin Z.

[0024] In some embodiments, the subject is a mammal, hi some embodiments, the subject is a human.

[0025] In some embodiments, the epitope binds to MHC class I HLA. In some embodiments, the epitope binds to MHC class I HLA with a stability of 10 minutes to 24 hours. In some embodiments, the epitope binds to MHC class I HLA with an affinity of 0.1 nM to 2000 nM. In some embodiments, the epitope binds to MHC class II HLA. In some embodiments, the epitope binds to MHC class II HLA with a stability of 10 minutes to 24 hours. In some embodiments, the epitope binds to MHC class II HLA with an affinity of 0.1 nM to 2000 nM, 1 nM to 1000 nM, 10 nM to 500 nM, or less than 1000 nM. In some embodiments, n is an integer from 1 to 20 or 5 to 12. In some embodiments, p is an integer from 1 to 20 or 5 to 12. In some embodiments, the epitope comprises a tumor-specific epitope.

[0026] In some embodiments, the polypeptide comprises at least two polypeptides, and two or more of the at least two polypeptides have the same formula Y n -B t -A r -X m -A s -C u -Z p In some embodiments, the polypeptide comprises at least two polypeptide molecules. In some embodiments, two or more of the at least two polypeptides or polypeptide molecules have X mIn some embodiments, two or more Y of the at least two polypeptides or polypeptide molecules are the same. n In some embodiments, two or more Z of the at least two polypeptides or polypeptide molecules are the same. p In some embodiments, two or more A of the at least two polypeptides or polypeptide molecules are the same. r and / or A s are different. In some embodiments, r=0 for a first polypeptide or polypeptide molecule of the at least two polypeptides or polypeptide molecules, and r=1 for a second polypeptide or polypeptide molecule of the at least two polypeptides or polypeptide molecules. In some embodiments, s=0 for a first polypeptide or polypeptide molecule of the at least two polypeptides or polypeptide molecules, and s=1 for a second polypeptide or polypeptide molecule of the at least two polypeptides or polypeptide molecules. In some embodiments, the polypeptide comprises at least 3, 4, 5, 6, 7, 8, 9, 10, or more polypeptides or polypeptide molecules.

[0027] In some embodiments, epitope is RAS epitope.In some embodiments, epitope comprises at least 8 consecutive amino acids of mutant RAS protein comprising mutation in G12, G13 or Q61 and mutant RAS peptide sequence comprising mutation in G12, G13 or Q61.In some embodiments, at least 8 consecutive amino acids of mutant RAS protein comprising mutation in G12, G13 or Q61 comprise G12A, G12C, G12D, G12R, G12S, G12V, G13A, G13C, G13D, G13R, G13S, G13V, Q61H, Q61L, Q61K or Q61R mutation. In some embodiments, the mutations at G12, G13, or Q61 include G12A, G12C, G12D, G12R, G12S, G12V, G13A, G13C, G13D, G13R, G13S, G13V, Q61H, Q61L, Q61K, or Q61R mutations. n and / or Z p comprises the amino acid sequence of a protein of pp65, human immunodeficiency virus (HIV), or cytomegalovirus (CMV), such as MART-1. In some embodiments, n and / or p are 1, 2, 3, or an integer greater than 3. In some embodiments, Y n and / or Z p comprises lysine or poly-lysine. In some embodiments, Y n and / or Z p includes K, KK, KKK, KKKK or KKKKK.

[0028] In some embodiments, the epitope binds to a protein encoded by an HLA allele with an affinity of less than 10 μM, less than 1 μM, less than 500 nM, less than 400 nM, less than 300 nM, less than 250 nM, less than 200 nM, less than 150 nM, less than 100 nM, or less than 50 nM. In some embodiments, the epitope binds to a protein encoded by an HLA allele with a stability of greater than 24 hours, greater than 12 hours, greater than 9 hours, greater than 6 hours, greater than 5 hours, greater than 4 hours, greater than 3 hours, greater than 2 hours, greater than 1 hour, greater than 45 minutes, greater than 30 minutes, greater than 15 minutes, or greater than 10 minutes. In some embodiments, the HLA allele is selected from the group consisting of an HLA-A02:01 allele, an HLA-A03:01 allele, an HLA-A11:01 allele, an HLA-A03:02 allele, an HLA-A30:01 allele, an HLA-A31:01 allele, an HLA-A33:01 allele, an HLA-A33:03 allele, an HLA-A68:01 allele, an HLA-A74:01 allele, and / or an HLA-C08:02 allele, and any combination thereof.

[0029] In some embodiments, the epitope is GADGVGKSAL, GACGVGKSAL, GAVGVGKSAL, GADGVGKSA, GACGVGKSA, GAVGVGKSA, KLVVVGACGV, FLVVVGACGL, FMVVVGACGI, FLVVVGACGI, FMVVVGACGV, FLVVVGACGV, MLVVVGACGV, FMVVVGACGL, YLVVVGACGV, KMVVVGACGV, YMVVVGACGV, MMVVVGACGV, DTAGHEEY, TAGHEEYSAM, DILDTAGHE, DILDTAG H, ILDTAGHEE, ILDTAGHE, DILDTAGHEEY, DTAGHEEYS, LLDILDTAGH, DILTAGRE, DILDTAGR, ILDTAGREE, ILDTAGRE, CLLDILDTAGR, TAGREEYSAM, REEYSAMRD, DTAGKEEYSAM, CLLDILDTAGK, DTAGKEEY, LLDILDTAGK, ILDTAGKE, ILDTAGKEE, DTAGLEEY, ILDTAGLE, DILDTAGL, ILDTAGLEE, GLEEYSAMRDQY, LLDILTAGLE, LDILDTAGL, DILTAGLE, DILDTAGLEEY, AGVGKSAL, GAAGVGKSAL, AAGVGKSAL, CGVGKSAL, ACGVGKSAL, DGVGKSAL, ADGVGKSAL, DGVGKSALTI, GARGVGKSA, KLVV VGARGV, VVVGARGV, SGVGKSAL, VVVGASGVGK, GASGVGKSAL, VGVGKSAL, VVVGAGCVGK, KLVVVGAGC, GDVGKSAL, DVGKSALTI, VVVGAGDVGK, TAGKEEYSAM, DTAGHEE YSAM, TAGHEEYSA, DTAGREEYSAM, TAGKEEYSA, AAGVGKSA, AGCVGKSAL, AGDVGKSAL, AGKEEYSAMR, AGVGKSALTI, ARGVGKSAL, ASGVGKSA, ASGVGKSAL, AVGVGKSA ,CVGKSALTI,DILDTAGK,DILDTAGREEY,DTAGHEEYSAMR,DTAGKEEYS,DTAGKEEYSAMR,DTAGLEEYS,DTAGLEEYSA,DTAGLEEYSAMR,DTAGREEYS,DTAGREEYSAMR,GAAGVGKSA, GACGVGKSA, GACGVGKSAL, GADGVGKS, GAGDVGKSA, GAGDVGKSAL, GASGVGKSA, GCVGKSAL, GCVGKSALTI, GHEEYSAM, GKEEYSAM, GLEEYSAMR, GREEYSAM, GREEYSAMR, HEEYSAMRD, KEEYSAMRD, KLVVVGASG, LDILDTAGR, LEEYSAMRD, LVVVGARGV, LVVVGASGV, REEYSAMRDQY, RGVGKSAL, TAGLEEYSA, TEYKLVVVGAA, VGAAGVGKSA, VGADGVGK, VGASGVGKSA, VGVGKSALTI, VVVGAAGV, VVVGAVGV, YKLVVVGAC, YKLVVVGAD, YKLVVVGAR, or DILDTAGKE.

[0030] In some embodiments, Y n comprises the amino acid sequence of IDIIMKIRNA, FFFFFFFFFFFFFFFFFFFFIIFFIFFWMC, FFFFFFFFFFFFFFFFFFFFFFFFFFAAFWFW, IFFIFFIIFFFFFFFFFFFFIIIIIIIWEC, FIFFFIIFFFFFFFFFFFIFIIIIIIFWEC, TEY, TEYKLV, WQAGILAR, HSYTTAE, PLTEEKIK, GALHFKPGSR, RRANKDATAE, KAFISHEEKR, TDLSSRFSKS, FDLGGGTFDV, CLLLHYSVSK, KKKKIIMKIRNA, or MTEYKLVVV. pcomprises the amino acid sequence of KKNKKDDI, KKNKKDDIKD, AGNDDDDDDDDDDDDDDDDDKKDKDDDDDD, AGNKKKKKKKNNNNNNNNNNNNNNNNNNNN, AGRDDDDDDDDDDDDDDDDDDDDDDDDDDDDD, SALTI, SALTIQL, GKSALTIQL, GKSALTI, QGQNLKYQ, ILGVLLLI, EKEGKISK, AASDFIFLVT, KELKQVASPF, KKKLINEKKE, KKCDISLQFF, KSTAGDTHLG, ATFYVAVTVP, LTIQLIQNHFVDEYDPTIEDSYRKQVVIDG, or TIQLIQNHFVDEYDPTIEDSYRKQVVIDGE.

[0031] In some embodiments, the epitope is not a RAS epitope. In some embodiments, the polypeptide is not KKKKPKRDGYMFLKAESKIMFAT, KKKKYMFLKAESKIMFATLQRSS, KKKKKAESKIMFATLQRSSLWCL, KKKKIMFATLQRSSLWCLCSNH, or KKKKMFATLQRSSLWCLCSNH.

[0032] In some embodiments, the epitope is a GATA3 epitope. In some embodiments, the GATA3 epitope comprises the amino acid sequence of MLTGPPARV, SMLTGPPARV, VLPEPHLAL, KPKRDGYMF, KPKRDGYMFL, ESKIMFATL, KRDGYMFL, PAVPFDLHF, AESKIMAFATL, FATLQRSSL, ARVPAVPFD, IMKPKRDGY, DGYMFLKA, MFLKAESKIMF, LTGPPARV, ARVPAVPF, SMLTGPPAR, RVPAVPFDL, or LTGPPARVP.

[0033] In some aspects, provided herein is a cell comprising a polypeptide as described herein. In some embodiments, the cell is an antigen-presenting cell. In some embodiments, the cell is a dendritic cell. In some embodiments, the cell is a mature antigen-presenting cell.

[0034] In some aspects, provided herein is a method of cleaving a polypeptide, the method comprising contacting a polypeptide described herein with an antigen-presenting cell (APC). In some embodiments, the method is performed in vivo. In some embodiments, the method is performed ex vivo.

[0035] In some aspects, a method of producing a polypeptide comprises the steps of: n -A r and / or A s -Z p to a sequence comprising an epitope sequence, wherein the epitope sequence is presented by class I MHC or class II MHC of an antigen presenting cell (APC); (i) each Y is independently an amino acid, an analogue or derivative thereof, and n is not encoded by a nucleic acid sequence immediately upstream of a nucleic acid sequence in the genome of the subject encoding the epitope, and n is an integer from 0 to 1000; (ii) each Z is independently an amino acid, an analog, or a derivative thereof, and Z p is not encoded by a nucleic acid sequence immediately downstream of a nucleic acid sequence in the subject's genome that encodes the epitope, and p is an integer from 0 to 1000; and r is a linker, and A s is a linker, and at least one of r and s is 1, and further wherein (a) the polypeptide does not consist of four distinct epitopes presented by class I MHC; (b) the polypeptide comprises at least two distinct polypeptide molecules; (c) the epitope comprises at least one mutant amino acid; and / or (d) the polypeptide, upon processing by the APC, becomes Y n and / or Z p is cleaved from the epitope.

[0036] In some aspects, a method of producing a polypeptide comprises the steps of: n B t -X m and / or Z p Xm -C u and X m is an epitope sequence presented by class I MHC or class II MHC of an antigen presenting cell (APC); and (i) each B independently is m (ii) each C independently represents an amino acid encoded by a nucleic acid sequence in the subject's genome immediately upstream of a nucleic acid sequence in the subject's genome encoding m (iii) each Y represents an amino acid encoded by a nucleic acid sequence in the subject's genome that is immediately downstream of a nucleic acid sequence in the subject's genome encoding n But, B t -X m and n is an integer from 0 to 1000; and (iv) each Z is independently an amino acid, an analog, or a derivative thereof, and Z p But X m -C u and p is an integer between 0 and 1000; further, (a) the polypeptide does not consist of four distinct epitopes presented by class I MHC; (b) the polypeptide comprises at least two distinct polypeptide molecules; (c) the epitope comprises at least one mutant amino acid; and / or (d) the polypeptide, upon processing by an APC, becomes Y n -B t and / or C. u -Z p is cleaved from the epitope.

[0037] In some embodiments, when n is 0, p is an integer from 1 to 1000, and when p is 0, n is an integer from 1 to 1000. In some embodiments, each X independently represents an amino acid of a peptide sequence comprising any contiguous amino acid sequence encoded by a nucleic acid sequence in the genome of the subject, and wherein (a) the MHC is a class I MHC and m is an integer from 8 to 12, or (b) the MHC is a class II MHC and m is an integer from 9 to 25.

[0038] In some aspects, provided herein is a pharmaceutical composition comprising a polypeptide as described herein and a pharma- ceutically acceptable excipient. In some embodiments, the pharmaceutical composition further comprises an immunomodulator or adjuvant. In some embodiments, the immunomodulator or adjuvant is Poly-ICLC, 1018 ISS, aluminum salt, Amplivax, AS15, BCG, CP-870,893, CpG7909, CyaA, ARNAX, STING agonist, dSLIM, GM-CSF, IC30, IC31, Imiquimod, ImuFact IMP321, IS Patch, ISS, ISCOMATRIX, Juvlmmune, LipoVac, MF59, monophosphoryl lipid A, Montanide IMS 1312, Montanide ISA 206, Montanide ISA 50V, Montanide The immunomodulatory agent is selected from the group consisting of ISA-51, OK-432, OM-174, OM-197-MP-EC, ONTAK, PepTel®, vector systems, PLGA microparticles, resiquimod, SRL172, virosomes and other virus-like particles, YF-17D, VEGF trap, R848, beta-glucan, Pam2Cys, Pam3Cys, Pam3CSK4, and Aquila's QS21 stimulon. In some embodiments, the immunomodulatory agent or adjuvant comprises poly-ICLC. In some embodiments, the pharmaceutical composition is a vaccine composition. In some embodiments, the pharmaceutical composition is aqueous or liquid.

[0039] In some embodiments, the epitope is present in the pharmaceutical composition in an amount of 1 ng to 10 mg, or 5 μg to 1.5 mg. In some embodiments, the pharmaceutical composition further comprises DMSO. In some embodiments, the pharma- ceutically acceptable excipient comprises water. In some embodiments, the pharmaceutical composition comprises a pH adjuster present at a concentration lower than 1 mM or higher than 1 mM. In some embodiments, the pH adjuster is a dicarboxylate or tricarboxylate. In some embodiments, the pH adjuster is a dicarboxylate or disuccinate of succinic acid. In some embodiments, the pH adjuster is a tricarboxylate or tricitrate of citric acid. In some embodiments, the pH adjuster is disodium succinate. In some embodiments, the dicarboxylate or disuccinate of succinic acid is present in the pharmaceutical composition at a concentration of 0.1 mM to 1 mM. In some embodiments, the dicarboxylate or disuccinate of succinic acid is present in the pharmaceutical composition at a concentration of 1 mM to 5 mM. In some embodiments, when administered to a subject, an immune response to the epitope is increased.

[0040] In some aspects, provided herein is a method for treating a disease or condition, comprising administering to a subject in need thereof a therapeutically effective amount of a pharmaceutical composition as described herein. In some embodiments, the disease or condition is cancer. In some embodiments, the cancer is selected from the group consisting of lung cancer, non-small cell lung cancer, pancreatic cancer, colorectal cancer, uterine cancer, and liver cancer. In some embodiments, the administering step comprises intradermal injection, intranasal spray application, intramuscular injection, intraperitoneal injection, intravenous injection, oral administration, or subcutaneous injection.

[0041] In some aspects, provided herein are prophylactic methods in a subject, the methods comprising contacting cells of the subject with a polypeptide, cell, or pharmaceutical composition described herein.

[0042] In some aspects, the method includes identifying an epitope expressed by a tumor cell of a subject and making a polypeptide comprising the epitope, the polypeptide having a structure represented by formula (I): Y n -B t -A r -X m -A s -C u -Z p Formula (I), or a pharma- ceutically acceptable salt thereof, (i)X m is an epitope, each X independently represents an amino acid in a contiguous amino acid sequence encoded by a nucleic acid sequence in the genome of a subject, and (a) the MHC is a class I MHC and m is an integer between 8 and 12; and (b) the MHC is a class II MHC and m is an integer from 9 to 25; (ii) each Y is independently an amino acid, an analog, or a derivative thereof; and (A) A of formula (I) r If the variable r in is 0, then Y n is B t -A r -X m is not encoded by a nucleic acid sequence immediately upstream of a nucleic acid sequence in the subject's genome encoding (B) A of formula (I) r The variable r is 1, and B in formula (I) t If the variable t in is 0, then Y n X m or (C) A of formula (I) r The variable r is 1, and B in formula (I) t If the variable t in is 1 or greater, then Y n is B t is not encoded by a nucleic acid sequence immediately upstream of a nucleic acid sequence in the subject's genome encoding further, n is an integer between 0 and 1000; (iii) each Z is independently an amino acid, an analog, or a derivative thereof; and (A) A of formula (I) s If the variable s in is 0, Z p Xm -A s -C u is not encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the subject's genome encoding (B) A of formula (I) s The variable s in formula (I) is 1, u If the variable u in is 0, then Z p X m or (C) A of formula (I) s The variable s in formula (I) is 1, u If the variable u in is 1 or greater, then Z p is C u is not encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the subject's genome encoding Further, p is an integer from 0 to 1000; Further, when n is 0, p is an integer from 1 to 1000; and When p is 0, n is an integer from 1 to 1000; (iv)A r is a linker and r is 0 or 1; (v)A s is a linker and s is 0 or 1; (vi) Each B independently m represents an amino acid encoded by a nucleic acid sequence in the subject's genome that is immediately upstream of a nucleic acid sequence in the subject's genome that encodes t is an integer from 0 to 1000; and (vii) Each C independently selects X m represents an amino acid encoded by a nucleic acid sequence in the subject's genome that is immediately downstream of a nucleic acid sequence in the subject's genome that encodes u is an integer between 0 and 1000; moreover, (a) the polypeptide does not consist of four distinct epitopes presented by class I MHC; (b) the polypeptide comprises at least two distinct polypeptide molecules; (c) the epitope comprises at least one mutant amino acid; and / or (d) When the polypeptide is processed by the APC, Y n and / or Z p is cleaved from the epitope, Provided herein is a method comprising the steps of:

[0043] In some embodiments, the identifying step comprises selecting from a pool of sequenced nucleic acid sequences from the subject's tumor cells that encode a plurality of candidate peptide sequences that contain one or more distinct mutations that are not present in a pool of sequenced nucleic acid sequences from the subject's non-tumor cells, where the pool of sequenced nucleic acid sequences from the subject's tumor cells and the pool of sequenced nucleic acid sequences from the subject's non-tumor cells are sequenced by whole genome or whole exome sequencing. In some embodiments, the identifying step further comprises predicting or determining which of the plurality of candidate peptide sequences form a complex with a protein encoded by an HLA allele of the same subject by HLA peptide binding analysis. In some embodiments, the identifying step further comprises selecting from the candidate peptide sequences a plurality of selected tumor-specific peptides or one or more polynucleotides that encode the plurality of selected tumor-specific peptides based on the HLA peptide binding analysis.

[0044] In some embodiments, the method further comprises administering the polypeptide to the subject. In some embodiments, the administering comprises intradermal injection, intranasal spray application, intramuscular injection, intraperitoneal injection, intravenous injection, oral administration, or subcutaneous injection. In some embodiments, an immune response is elicited in the subject. In some embodiments, the epitope expressed by tumor cells of the subject is a neoantigen, a tumor associated antigen, a mutated tumor associated antigen, and / or expression of the epitope in tumor cells of the subject is increased compared to expression of the epitope in cells of a normal subject.

[0045] A polypeptide comprising an epitope presented by class I or class II MHC of an antigen presenting cell (APC), comprising the formula (I): Y n -B t -A r -X m -A s -C u -Z p Formula (I), or a pharma- ceutically acceptable salt thereof, m is an epitope, each X independently represents an amino acid of a contiguous amino acid sequence encoded by a nucleic acid sequence in the genome of the subject, and (a) the MHC is a class I MHC and m is an integer from 8 to 12, or (b) the MHC is a class II MHC and m is an integer from 9 to 25; each Y independently is an amino acid, analog, or derivative thereof, and A of formula (I) r If the variable r in is 0, then Y n is B t -A r -X m or is not encoded by a nucleic acid sequence immediately upstream of a nucleic acid sequence in the genome of a subject encoding r The variable r is 1, and B in formula (I) t If the variable t in is 0, then Y n X m or is not encoded by a nucleic acid sequence immediately upstream of a nucleic acid sequence in the genome of a subject encoding r The variable r is 1, and B in formula (I) t If the variable t in is 1 or greater, then Y n is B t and n is an integer from 0 to 1000; each Z is independently an amino acid, an analog, or a derivative thereof, and is not encoded by a nucleic acid sequence immediately upstream of a nucleic acid sequence in the genome of the subject that encodes s If the variable s in is 0, Z p X m -A s -C uor is not encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the genome of the subject encoding s The variable s in formula (I) is 1, u If the variable u in is 0, then Z p X m or is not encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the genome of the subject encoding A s The variable s in formula (I) is 1, u If the variable u in is 1 or greater, then Z p is C u is not encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the genome of the subject encoding the nucleic acid sequence; further, p is an integer from 0 to 1000; further, when n is 0, p is an integer from 1 to 1000; when p is 0, n is an integer from 1 to 1000; A r is a linker, r is 0 or 1, and A s is a linker, s is 0 or 1, and each B is independently selected from X m represents an amino acid encoded by a nucleic acid sequence in the subject's genome that is immediately upstream of a nucleic acid sequence in the subject's genome that encodes m and u is an integer between 0 and 1000; further, the polypeptide does not consist of four distinct epitopes presented by class I MHC; the polypeptide comprises at least two distinct polypeptide molecules; the epitope comprises at least one mutant amino acid; and / or the polypeptide upon processing by an APC represents an amino acid sequence encoded by Y. n and / or Z p Provided herein are polypeptides in which the is cleaved from the epitope.

[0046] In some embodiments, the epitope is one presented by class II MHC and m is an integer from 9 to 25.

[0047] In some embodiments, Y n -B t -A r and / or A s -C u -Z p The solubility of the polypeptide is increased by Y n -B t -A r and / or A s -C u -Z p In some embodiments, the epitope is enhanced when the polypeptide is processed by the APC, compared to a corresponding peptide that does not contain Y. n -B t -A r and / or A s -C u -Z p In some embodiments, where n is an integer from 1 to 1000, the polypeptide is m and X m and / or, when p is an integer between 1 and 1000, the polypeptide is cleaved at a rate greater than that of a corresponding polypeptide of the same length that contains at least one additional amino acid encoded by a nucleic acid sequence immediately upstream of the nucleic acid sequence in the subject's genome that encodes X. m and X m In some embodiments, epitope presentation by APCs is enhanced by a rate of cleavage greater than that of a corresponding polypeptide of the same length that contains at least one additional amino acid encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the subject's genome that encodes X. m and X m and / or, when p is an integer between 1 and 1000, epitope presentation by APCs is enhanced compared to epitope presentation of a corresponding polypeptide of the same length that contains at least one additional amino acid encoded by a nucleic acid sequence immediately upstream of the nucleic acid sequence in the subject's genome that encodes X. m and X mIn some embodiments, the epitope is presented to an immune cell by an APC.

[0048] In some embodiments, when n is an integer from 1 to 1000, the immunogenicity is m and X m and / or, where p is an integer between 1 and 1000, the immunogenicity is enhanced compared to the immunogenicity of a corresponding polypeptide of the same length that contains at least one additional amino acid encoded by a nucleic acid sequence immediately upstream of the nucleic acid sequence in the subject's genome that encodes X. m and X m The immunogenicity is enhanced compared to that of a corresponding polypeptide of the same length that contains at least one additional amino acid encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the subject's genome that encodes the polypeptide.

[0049] In some embodiments, the antitumor activity is greater than or equal to X, where n is an integer from 1 to 1000. m and X m and / or, when p is an integer between 1 and 1000, the antitumor activity is enhanced compared to the antitumor activity of a corresponding polypeptide of the same length that contains at least one additional amino acid encoded by a nucleic acid sequence immediately upstream of the nucleic acid sequence in the subject's genome that encodes X m and X m The anti-tumor activity of the polypeptide is enhanced compared to that of a corresponding polypeptide of the same length that contains at least one additional amino acid encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the subject's genome that encodes the polypeptide.

[0050] In some embodiments, Y n and / or Z pcomprises a sequence selected from the group consisting of lysine (Lys), poly-Lys (polyK) and poly-Arg (polyR). In some embodiments, polyK comprises poly-L-Lys. In some embodiments, polyR comprises poly-L-Arg. In some embodiments, polyK or polyR comprises at least two, three or four consecutive lysine or arginine residues, respectively.

[0051] In some embodiments, the epitope binds to MHC II class HLA. In some embodiments, the epitope binds to MHC II class HLA with a stability of 10 minutes to 24 hours. In some embodiments, the epitope binds to MHC II class HLA with an affinity of 0.1 nM to 2000 nM, 1 nM to 1000 nM, 10 nM to 500 nM, or less than 1000 nM.

[0052] In some embodiments, the polypeptide is not cleaved before processing or internalization by APC in the subject.In some embodiments, the polypeptide is stable in human plasma.In some embodiments, the polypeptide has a half-life in human plasma of 1 hour to 5 days.In some embodiments, the subject is a human.

[0053] In some embodiments, the epitope binds to a protein encoded by an HLA allele with an affinity of less than 10 μM, less than 1 μM, less than 500 nM, less than 400 nM, less than 300 nM, less than 250 nM, less than 200 nM, less than 150 nM, less than 100 nM, or less than 50 nM. In some embodiments, the epitope binds to a protein encoded by an HLA allele with a stability of greater than 24 hours, greater than 12 hours, greater than 9 hours, greater than 6 hours, greater than 5 hours, greater than 4 hours, greater than 3 hours, greater than 2 hours, greater than 1 hour, greater than 45 minutes, greater than 30 minutes, greater than 15 minutes, or greater than 10 minutes. In some embodiments, the HLA allele is selected from the group consisting of HLA-A02:01 allele, HLA-A03:01 allele, HLA-A11:01 allele, HLA-A03:02 allele, HLA-A30:01 allele, HLA-A31:01 allele, HLA-A33:01 allele, HLA-A33:03 allele, HLA-A68:01 allele, HLA-A74:01 allele, and / or HLA-C08:02 allele, and any combination thereof. In some embodiments, the epitope comprises a tumor-specific epitope. In some embodiments, the epitope comprises at least one mutant amino acid. In some embodiments, the at least one mutant amino acid is encoded by an insertion, deletion, frameshift, neo-ORF, or point mutation in a nucleic acid sequence in the genome of the subject.

[0054] In some embodiments, epitope is RAS epitope.In some embodiments, epitope comprises at least 8 consecutive amino acids of mutant RAS protein comprising mutation in G12, G13 or Q61 and mutant RAS peptide sequence comprising mutation in G12, G13 or Q61.In some embodiments, at least 8 consecutive amino acids of mutant RAS protein comprising mutation in G12, G13 or Q61 comprise G12A, G12C, G12D, G12R, G12S, G12V, G13A, G13C, G13D, G13R, G13S, G13V, Q61H, Q61L, Q61K or Q61R mutation. In some embodiments, the mutations at G12, G13, or Q61 include G12A, G12C, G12D, G12R, G12S, G12V, G13A, G13C, G13D, G13R, G13S, G13V, Q61H, Q61L, Q61K, or Q61R mutations. In some embodiments, the RAS epitope is VVVGAAGVGK, VVVGAAGVG, VVVGAAGV, VVGAAGVGK, VVGAAGVG, VGAAGVGK, VVVGACGVG, VVVGACGV, VVGACGVGK, VVGACGVG, VGACGVGK, VVVGADGVGK, VVVGADGVG, VVVGADGV, VVGADGVGK, VVGADGVG, VGADGVG, VGADG In some embodiments, the amino acid sequence of Y is selected from the group consisting of YVGK, VVVGARGVGK, VVVGARGVG, VVVGARGV, VVGARGVGK, VVGARGVG, VGARGVGK, VVVGASGVGK, VVVGASGVG, VVVGASGV, VVGASGVGK, VVGASGVG, VGASGVGK, VVVGAVGVG, VVVGAVGV, VVGAVGVGK, VVGAVGVG, or VGAVGVGK. nK, KK, KKK, KKKK, KKKKK, KKKKKKK, KKKKKKKK, KTEY, KTEYK, KTEYKL, KTEYKLV, KTEYKLVV, KTEYKLVVV, KKTEY, KKTEYK, KKTEYKL, KKTEYKLV, KKTEYKLVV, KKTEYKLVVV, KKKT EY, KKKTEYK, KKKTEYKL, KKKTEYKLV, KKKTEYKLVV, KKKTEYKLVVV, KKKKTEY, KKKKTEYK, KKKKTEYKL, KKKKTEYKLV, KKKKTEYKLVV, KKKKTEYKLVVV, IDIIMKIRNA, FFFFFFFFFFFF In some embodiments, the amino acid sequence of Z is selected from the group consisting of FFFFFFFFIIFFIFFWMC, FFFFFFFFFFFFFFFFFFFFFFFFFFFFAAFWFW, IFFIFFIIFFFFFFFFFFFFIIIIIIIWEC, FIFFFIIFFFFFIFFFFFFFIFIIIIIIFWEC, TEY, TEYK, TEYKL, TEYKLV, TEYKLVV, TEYKLVVV, WQAGILAR, HSYTTAE, PLTEEKIK, GALHFKPGSR, RRANKDATAE, KAFISHEEKR, TDLSSRFSKS, FDLGGGTFDV, CLLLHYSVSK, KKKKIIMKIRNA, or MTEYKLVVV. pは、K、KK、KKK、KKKK、KKKKK、KKKKKKK、KKKKKKKKKKKKK、KKNKKDDI、KKNKKDDDIKD、AGNDDD DDDDDDDDDDDDDDKKDKDDDDDD、AGNKKKKKKNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN this DDDDDDDDDDDDDDDDDDDDD、JUMP、JUMPKL、GKJUMP、JUMP、JUMP IQLK、GKJUMPING、GKJUMPING、JUMPING、JUMPING、GKJUMPING、GKJUMPING、 JUMPKKK、JUMPQLKKK、GKJUMPKKKK、GKKKKK、JUMPKKKK 、GKSALTIQLKKKK、GKSALTI、KKKK、QGQNLKYQ、ILGVLLLI、EKEGKISK、AASDFIFLVT 、KELKQVASPF、KKKLINEKKE、KKCDISLQFF、KSTAGDTHLG、ATFYVAVTVP、LTIQLIQNH FVDEYDPTIEDSYRKQVVIDG is covered in TiQLIQNHFVDEYDPTIEDSYRKQVVIDG.In some embodiments, the polypeptide is KTEYKLVVVGAVGVGKSALTIQL, KTEYKLVVVGADGVGKSALTIQL, KTEYKLVVVGARGVGKSALTIQL, KTEYKLVVVGACGVGKSALTIQL, KKTEYKLVVVGAVGVGKSALTIQL, KKTEYKLVVVGADGVGKSALTIQL, KKTEYKLVVVGARGVGKSALTIQL, KKTEYKLVVVGACGVGKSALTIQL, KKKTEYKLVVVGAVGVGKSALTIQL , KKKTEYKLVVVGADGVGKSALTIQL, KKKTEYKLVVVGARGVGKSALTIQL, KKKTEYKLVVVGACGVGKSALTIQL, KKKKTEYKLVVVGAVGVGKSALTIQL, KKKKTEYKLVVVG ADGVGKSALTIQL, KKKKTEYKLVVVGARGVGKSALTIQL, KKKKTEYKLVVVGACGVGKSALTIQL, KKTEYKLVVVGAVGVGKSALTIQLKK, KKTEYKLVVVGADGVGKSALTIQLK K, KKTEYKLVVVGARGVGKSALTIQLKK, KKTEYKLVVVGACGVGKSALTIQLKK, TEYKLVVVGAVGVGKSALTIQLK, TEYKLVVVGADGVGKSALTIQLK, TEYKLVVVGARGVGK SALTIQLK, TEYKLVVVGACGVGKSALTIQLK, TEYKLVVVGAVGVGKSALTIQLKK, TEYKLVVVGADGVGKSALTIQLKK, TEYKLVVVGARGVGKSALTIQLKK, TEYKLVVVGAC In some embodiments, the epitope is not a RAS epitope.In some embodiments, the polypeptide is not KKKKPKRDGYMFLKAESKIMFAT, KKKKYMFLKAESKIMFATLQRSS, KKKKKAESKIMFATLQRSSLWCL, KKKKKIMFATLQRSSLWCLCSNH, or KKKKMFATLQRSSLWCLCSNH.

[0055] In some embodiments, Y n and / or Z p comprises the amino acid sequence of a protein that is different from the protein from which the epitope was derived. n and / or Z p comprises the amino acid sequence of a protein of CMV, such as pp65, HIV, or MART-1. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or an integer greater than 20. In some embodiments, p is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or an integer greater than 20.

[0056] In some embodiments, the epitope is the TMPRSS2:ERG epitope. In some embodiments, the TMPRSS2:ERG epitope comprises the amino acid sequence of ALNSEALSV.

[0057] Also provided herein is a polynucleotide comprising a sequence encoding a polypeptide described herein. In some embodiments, the polynucleotide is an mRNA.

[0058] Also provided herein is a pharmaceutical composition comprising a polypeptide described herein or a polynucleotide described herein; and a pharma- ceutically acceptable excipient.

[0059] Also provided herein is a method for treating a disease or condition, comprising administering to a subject in need thereof a therapeutically effective amount of a pharmaceutical composition described herein. In some embodiments, the disease or condition is a cancer selected from the group consisting of lung cancer, non-small cell lung cancer, pancreatic cancer, colorectal cancer, uterine cancer, prostate cancer, liver cancer, biliary malignancies, uterine cancer, cervical cancer, bladder cancer, liver cancer, myeloid leukemia, and breast cancer. In some embodiments, the administering step comprises intradermal injection, intranasal spray application, intramuscular injection, intraperitoneal injection, intravenous injection, oral administration, or subcutaneous injection.

[0060] Also provided herein are methods of preparing antigen-specific T cells, comprising stimulating the T cells with an antigen-presenting cell comprising a polypeptide described herein or a polynucleotide encoding a polypeptide described herein. In some embodiments, the method is performed ex vivo.

[0061] The features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings, in which: [Brief description of the drawings]

[0062] [Figure 1] Figure 1 shows a simplified exemplary epitope processing of epitope X by antigen presenting cells (APC) and presentation on HLA allele X. In a natural context, the peptide comprises amino acids or amino acid sequences that are naturally adjacent to the epitope sequence. In a rational context, the peptide comprises amino acids or amino acid sequences, and / or linkers, at the N-terminus and / or C-terminus of the epitope sequence that are not encoded by the genome encoding the epitope sequence.

[0063] [Diagram 2]FIG. 2 illustrates cathepsin B cleavage of an exemplary polypeptide containing a cathepsin B-cleavable linker.

[0064] [Diagram 3] FIG. 3 shows a diagram of the experimental design for screening polypeptides for epitope processing and presentation in vitro using T cell receptor (TCR)-transduced cells (results shown in FIGS. 4 and 5).

[0065] [Figure 4] FIG. 4 shows a graph demonstrating the levels of IL-2 (pg / mL) secreted from KRAS-specific Jurkat cells after they were co-cultured for 48 hours with peripheral blood mononuclear cells (PBMCs) loaded with equal amounts of either a peptide containing only the KRAS-G12V epitope or a peptide containing the KRAS-G12V epitope with additional amino acid sequences naturally adjacent to the N- and C-termini of the KRAS-G12V epitope.

[0066] [Diagram 5] FIG. 5 shows a graph demonstrating the levels of IL-2 (pg / mL) secreted from KRAS-specific Jurkat cells after 48 hours of co-culture with peripheral blood mononuclear cells (PBMCs) loaded with equal amounts of either a peptide containing only the KRAS-G12V epitope, a peptide containing the KRAS-G12V epitope and additional amino acid sequences that are naturally adjacent to the N-terminus and C-terminus of the KRAS-G12V epitope, or a peptide containing the KRAS-G12V epitope and additional amino acid sequences that are not naturally adjacent to the N-terminus and / or C-terminus of the KRAS-G12V epitope (rational situation).

[0067] [Figure 6]Figure 6 shows a diagram of the experimental design of the immunogenicity study. Mice were immunized with the various polypeptide designs on days 0, 7, and 14 and bled on days 7, 14, and 21 to evaluate antigen-specific CD8+ T cell responses (results are shown in Figures 7-9).

[0068] [Figure 7] FIG. 7 shows graphs demonstrating the total immune response (7A: H-2Kb, 7B: H-2Db, 7C: Total).

[0069] [Figure 8] FIG. 8 shows graphs demonstrating that immunization with the K4-epitope enhances the immune response against epitopes presented by H-2Kb (8A: Alg8, 8B: Lama4).

[0070] [Figure 9] FIG. 9 shows graphs demonstrating that immunization with the K4-epitope enhances the immune response against epitopes presented by the H-2Db (9A: Reps1, 9B: Adpgk, 9C: Irgq, 9D: Obsl1).

[0071] [Figure 10]FIG. 10 shows a graph demonstrating the levels of IL-2 (pg / mL) secreted by Jurkat cells after 24 hours of co-culture (5:1 Jurkat to 293T cell ratio) with 293T cells loaded with peptides containing the TMPRSS2::ERG epitope only, or with 293T cells transduced with a plasmid encoding a peptide containing the TMPRSS2::ERG epitope in a native context (i.e., the peptide further comprises amino acids or amino acid sequences that are not naturally adjacent to the N-terminus and / or C-terminus of the epitope sequence), with 293T cells transduced with a plasmid encoding a peptide containing the TMPRSS2::ERG epitope in a non-native context (i.e., the peptide further comprises amino acids or amino acid sequences that are not naturally adjacent to the epitope sequence), or with 293T cells transduced with a plasmid encoding an unrelated epitope in a non-native context (as a control).

[0072] [Figure 11] FIG. 11 shows a graph of IL-2 concentration (pg / mL) versus peptide concentration (nM) in FLT3L-treated PBMCs contacted with increasing amounts of the indicated RAS-G12V mutant peptides after co-culture with Jurkat cells transduced with a TCR that binds the underlined RAS-G12V epitope bound to MHC encoded by the HLA-A11:01 allele.

[0073] [Figure 12] FIG. 12 shows data illustrating the immunogenicity of the indicated RAS-G12V mutant peptides from FIG. 11 both in vitro using PBMCs from healthy donors (top) and in vivo using HLA-A11:01 transgenic mice immunized with the peptides (bottom).

[0074] [Figure 13A]FIG. 13A shows an exemplary schematic of the mRNA constructs used for expression in cells using short mers (9-10 amino acids, top) and long mers (25 amino acids, bottom).

[0075] [Figure 13B] Figure 13B shows an exemplary graph of the percentage of multimer-specific CD8+ cells relative to total CD8+ cells. The antigens used in the multimer assay are indicated.

[0076] [Figure 13C] FIG. 13C shows an exemplary flow cytometry analysis of detection of multimer-positive CD8+ T cells comparing APCs stimulated with a shortmer peptide (9-10 amino acids) and APCs stimulated with a longmer peptide (25 amino acids) as well as APCs containing RNA encoding the same shortmer peptide (9-10 amino acids) and APCs containing RNA encoding the longmer peptide (25 amino acids). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0077] Detailed Description Described herein are new immunotherapeutic compositions comprising individual tumor-specific antigens or neoepitopes and their uses based on the discovery of methods to enhance epitope processing and presentation to stimulate an immune response. Thus, the disclosure described herein provides peptides that can be used, for example, to stimulate an immune response against tumor-associated antigens or neoepitopes, to create immunogenic compositions or cancer vaccines for use in treating cancers, diseases or conditions.

[0078] The following description and examples are intended to illustrate the embodiments of the present disclosure in detail. It should be understood that the present disclosure is not limited to the specific embodiments described herein and may therefore vary. It will be understood by those skilled in the art that there are many variations and modifications of the present disclosure, which are encompassed within the scope of the present disclosure.

[0079] All terms are intended to be understood as understood by one of ordinary skill in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0080] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0081] Although various features of the disclosure may be described in the context of a single embodiment, the features may also be provided separately or in any suitable combination. Conversely, although the disclosure may for clarity be described herein in the context of separate embodiments, the disclosure may also be implemented in a single embodiment.

[0082] The following definitions are for the assistance of those skilled in the art, and are directed to this application, and are not to be attributed to any related or unrelated, for example, to any co-owned patent or application. Any method and material similar or equivalent to the methods and materials described herein can be used to carry out the testing of the present disclosure, but the preferred materials and methods are described herein. Therefore, the terms used herein are only to describe specific embodiments, and are not intended to be limiting. 1.Definition

[0083] The terms used herein are merely for the purpose of describing a particular instance and are not intended to be limiting. In this application, the use of the singular includes the plural unless otherwise specified. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0084] In this application, the use of "or" means "and / or" unless otherwise specified. The terms "and / or" and "any combination thereof" and their grammatical equivalents may be used interchangeably when used herein. These terms may indicate that any combination is specifically intended. For illustrative purposes only, the following phrase "A, B, and / or C" or "A, B, C, or any combination thereof" may mean "A individually; B individually; C individually; A and B; B and C; A and C; and A, B, and C." The term "or" may be used conjunctively or disjunctively, unless the context specifically dictates disjunctive use.

[0085] The term "about" or "approximately" may mean within an acceptable error range for a particular value as determined by one of ordinary skill in the art, and the acceptable error range will depend in part on how the value is measured or determined, i.e., on the limitations of the measurement system. For example, "about" may mean within 1 or more than 1 standard deviation, as is customary in the art. Alternatively, "about" may mean within 20%, 10%, 5%, or 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term may mean within an order of magnitude, within 5-fold, and more preferably within 2-fold of a value. When a particular value is described in this application and claims, the term "about" should be assumed to mean within an acceptable error range for the particular value, unless otherwise specified.

[0086] As used in the specification and claim(s), the word "comprising" (and any form of comprising, such as "comprise" and "comprises"), "having" (and any form of having, such as "have" and "has"), "including" (and any form of including, such as "includes" and "include") or "containing" (and any form of containing, such as "contains" and "contain"), etc. are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. It is intended that any embodiment discussed herein can be implemented with respect to any method or composition of the disclosure, and vice versa. Additionally, the compositions of the disclosure can be used to achieve the methods of the disclosure.

[0087] When reference is made herein to "some embodiments," "an embodiment," "one embodiment," or "other embodiments," it means that the particular feature, structure, or characteristic described in connection with an embodiment is included in at least some embodiments of the disclosure, but not necessarily in all embodiments. To facilitate understanding of this disclosure, certain terms and phrases are defined below.

[0088] The nomenclature used to describe peptides or proteins follows conventional practice, with the amino group of each amino acid residue shown on the left (amino-terminus or N-terminus) and the carboxyl group on the right (carboxyl-terminus or C-terminus). When reference is made to the position of an amino acid residue within a peptide epitope, it is numbered in the amino to carboxyl direction, with the residue located at the amino terminus of the epitope, or of the peptide or protein of which it may be a part, being numbered at position 1. In formulae representing selected particular embodiments of the present disclosure, the amino and carboxyl terminal groups are not specifically shown, but are in their assumed form at physiological pH values ​​unless otherwise indicated. In amino acid structural formulae, each residue is generally represented by a standard three-letter or one-letter name. Amino acid residues in the L-form are represented by a single capital letter or a three-letter symbol with the first letter capitalized, and the D-form of amino acid residues having the D-form are represented by a single lowercase letter or a three-letter symbol in lowercase. However, the three-letter symbol or full name may be used without a capital letter to refer to an L-amino acid residue. Glycine has no asymmetric carbon atom and is simply referred to as "Gly" or "G." The amino acid sequences of the peptides described herein are generally designated using the standard one-letter symbols (A, alanine; C, cysteine; D, aspartic acid; E, glutamic acid; F, phenylalanine; G, glycine; H, histidine; I, isoleucine; K, lysine; L, leucine; M, methionine; N, asparagine; P, proline; Q, glutamine; R, arginine; S, serine; T, threonine; V, valine; W, tryptophan; and Y, tyrosine).

[0089] The term "residue" refers to an amino acid residue or amino acid mimetic residue that is incorporated into a peptide or protein by an amide bond or amide bond mimetic, or a nucleic acid (DNA or RNA) that encodes an amino acid or amino acid mimetic.

[0090] "Polypeptide", "peptide" and their grammatical equivalents as used herein refer to a polymer of amino acid residues. A "mature protein" is a protein that is full-length and optionally includes glycosylation or other modifications typical of proteins in a given cellular environment. The polypeptides and proteins disclosed herein (including functional portions and functional variants thereof) may contain synthetic amino acids in place of one or more naturally occurring amino acids. Such synthetic amino acids are known in the art and include, for example, aminocyclohexanecarboxylic acid, norleucine, α-amino n-decanoic acid, homoserine, S-acetylaminomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine, β-phenylserine β-hydroxyphenylalanine, phenylglycine, α-naphthylalanine, cyclohexylalanine, cyclohexylglycine, indomethacin, cyclohexylglycine ... Examples of suitable lysine-2-carboxylic acids include lysine, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, aminomalonic acid monoamide, N'-benzyl-N'-methyl-lysine, N',N'-dibenzyl-lysine, 6-hydroxylysine, ornithine, α-aminocyclopentane carboxylic acid, α-aminocyclohexane carboxylic acid, α-aminocycloheptane carboxylic acid, α-(2-amino-2-norbornane)-carboxylic acid, α,γ-diaminobutyric acid, α,β-diaminopropionic acid, homophenylalanine, and α-tert-butylglycine. The present disclosure further contemplates that expression of the polypeptides described herein in engineered cells may be accompanied by post-translational modification of one or more amino acids of the polypeptide construct.Non-limiting examples of post-translational modifications include phosphorylation, acylation, including acetylation and formylation, glycosylation (including N-linked and O-linked), amidation, hydroxylation, alkylation, including methylation and ethylation, ubiquitination, addition of pyrrolidone carboxylic acid, formation of disulfide bridges, sulfation, myristoylation, palmitoylation, isoprenylation, farnesylation, geranylation, glypiation, lipoylation and iodination.

[0091] The term "peptide" refers to a stretch of amino acid residues in which one amino acid residue is joined to another, typically by a peptide bond between the α-amino and the carboxyl groups of the adjacent amino acid residues.

[0092] "Synthetic peptide" refers to a peptide obtained from a non-natural source, e.g., produced by man. Such peptides can be produced using methods such as chemical synthesis or recombinant DNA technology. "Synthetic peptide" includes "fusion proteins."

[0093] An "epitope" is a collective feature of a molecule, such as primary, secondary, and tertiary peptide structures and charges, which together form the site recognized by, for example, immunoglobulins, T cell receptors, HLA molecules, or chimeric antigen receptors. Alternatively, an epitope can be defined as a set of amino acid residues involved in recognition by a particular immunoglobulin, or, for T cells, a set of residues required for recognition by a T cell receptor protein, a chimeric antigen receptor, and / or a major histocompatibility complex (MHC) receptor. A "T cell epitope" should be understood to mean a peptide sequence to which a class I or II MHC molecule can bind in the form of a peptide-presenting MHC molecule or MHC complex, and which can then be recognized and bound by T cells, such as T lymphocytes or helper T cells, in this form. Epitopes can be prepared by isolation from natural sources or can be synthesized according to standard protocols in the art. Synthetic epitopes may include artificial amino acid residues, "amino acid mimetics," such as D-isomers of naturally occurring L-amino acid residues or non-naturally occurring amino acid residues such as cyclohexylalanine. Throughout this disclosure, epitopes may in some cases be referred to as peptides or peptide epitopes. It should be understood that proteins or peptides that include the epitopes or analogs described herein as well as additional amino acid(s) still fall within the scope of this disclosure. In certain embodiments, peptides include fragments of antigens. In certain embodiments, peptides of the present disclosure are limited in length. Limited-length embodiments occur when a protein or peptide that includes an epitope described herein includes a region (i.e., a stretch of contiguous amino acid residues) that has 100% identity to the native sequence. To avoid defining epitopes from read-through, the length of any region that has 100% identity to the native peptide sequence is limited, for example, to the entire natural molecule.Thus, for peptides that include an epitope described herein and a region having 100% identity to a native peptide sequence, the region having 100% identity to the native sequence generally has a length of less than or equal to 600 amino acid residues, less than or equal to 500 amino acid residues, less than or equal to 400 amino acid residues, less than or equal to 250 amino acid residues, less than or equal to 100 amino acid residues, less than or equal to 85 amino acid residues, less than or equal to 75 amino acid residues, less than or equal to 65 amino acid residues, and less than or equal to 50 amino acid residues. In certain embodiments, an "epitope" as described herein is comprised of a peptide having a region of less than 51 amino acid residues with 100% identity to the native peptide sequence, with any variation up to 5 amino acid residues; for example, 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 amino acid residue.

[0094] The term "derived" and its grammatical equivalents, when used in discussing epitopes, are synonymous with "prepared" and its grammatical equivalents. Derived epitopes can be isolated from natural sources or can be synthesized according to standard protocols in the art. Synthetic epitopes can include artificial amino acid residues, "amino acid mimetics," such as D-isomers of naturally occurring L-amino acid residues, or non-natural amino acid residues, such as cyclohexylalanine. Derived or prepared epitopes can be analogs of native epitopes.

[0095] An "immunogenic" peptide or "immunogenic" epitope or "peptide epitope" is a peptide that contains an allele-specific motif such that the peptide binds to an HLA molecule and elicits a cell-mediated or humoral response, such as the induction of cytotoxic T lymphocytes (CTLs (e.g., CD8 + )), helper T lymphocytes (Th (e.g., CD4 + )) and / or B lymphocyte responses. Thus, the immunogenic peptides described herein are capable of binding to the appropriate HLA molecule and subsequently inducing a CTL (cytotoxic) response, or an HTL (and humoral) response against the peptide.

[0096] "Neoantigen" refers to a class of tumor antigens that arise from tumor-specific alterations of proteins. Neoantigens include, but are not limited to, tumor antigens that arise from, for example, substitutions, frameshift mutations, fusion polypeptides, in-frame deletions, insertions, expression of endogenous retroviral polypeptides, and tumor-specific overexpression of polypeptides within a protein sequence.

[0097] The terms "mutant peptide", "tumor specific peptide", "neo-antigenic peptide", and "neo-antigenic peptide" are used interchangeably herein with "peptide" and refer to a stretch of one residue, typically an L-amino acid, connected by a peptide bond between another residue, typically an L-amino acid, and typically the α-amino and carboxyl group of the adjacent amino acid. Similarly, the term "polypeptide" is used interchangeably herein with "mutant polypeptide", "neo-antigenic polypeptide", and "neo-antigenic polypeptide" and refer to a stretch of one residue, e.g., an L-amino acid, connected by a peptide bond between another residue, e.g., an L-amino acid, and typically the α-amino and carboxyl group of the adjacent amino acid. Polypeptides or peptides may be of various lengths, may be in neutral (uncharged) or salt form, may be free of or contain modifications such as glycosylation, side chain oxidation, or phosphorylation, and may be subject to conditions under which the modifications do not impair the biological activity of the polypeptides described herein. A peptide or polypeptide, as used herein, comprises at least one flanking sequence. The term "flanking sequences" as used herein refers to fragments or regions of the neoantigenic peptide that are not part of the neoepitope.

[0098] "Neoepitope", "tumor specific neoepitope", "tumor specific epitope", or "tumor antigen" refers to an epitope or antigenic determining region that is not present in a reference non-disease cell, e.g., a non-cancerous cell or a germline cell, but is found in a disease cell, e.g., a cancer cell. This includes situations where the corresponding epitope is found in a normal non-disease cell or a germline cell, but due to one or more mutations in the disease cell, e.g., a cancer cell, the sequence of the epitope is altered, resulting in a neoepitope. The term "neoepitope" as used herein refers to an antigenic determining region within a peptide or neoantigenic peptide. A neoepitope may include at least one "anchor residue" and at least one "anchor residue adjacent region". A neoepitope may further include a "separation region". The term "anchor residue" refers to an amino acid residue that binds to a specific pocket on HLA and provides specificity of the interaction with HLA. In some cases, the anchor residue may be at a canonical anchor position. In other cases, the anchor residues may be at non-canonical anchor positions. Neoepitopes may bind to HLA molecules through primary and secondary anchor residues that protrude into pockets in the peptide-binding groove, where certain amino acids constitute pockets that accommodate the corresponding side chains of the anchor residues of the presented neoepitope. Peptide-binding preferences exist between different alleles of both HLA I and HLA II molecules. HLA class I molecules bind short neoepitopes, whose N- and C-termini are anchored in pockets located at the ends of the neoepitope-binding groove. The majority of HLA class I-binding neoepitopes are about 9 amino acids, but longer neoepitopes can also be accommodated in their central bulge, resulting in binding neoepitopes of about 8-12 amino acids. There is no size constraint on neoepitopes that bind to HLA class II proteins and can vary from about 16 amino acids to 25 amino acids. The neoepitope-binding groove of HLA class II molecules is open at both ends, allowing the binding of peptides of relatively long length.Although the core 9 amino acid residue long segment contributes most to neoepitope recognition, the anchor residue flanking region is also important for the specificity of the peptide to HLA class II alleles. In some cases, the anchor residue flanking region is the N-terminal residue. In other cases, the anchor residue flanking region is the C-terminal residue. In still other cases, the anchor residue flanking region is both the N-terminal and C-terminal residues. In some cases, the anchor residue flanking region is flanked by at least two anchor residues. The anchor residue flanking region that is flanked by anchor residues is a "separating region."

[0099] The "major histocompatibility complex" or "MHC" is a cluster of genes that plays a role in regulating the cellular interactions responsible for the physiological immune response. In humans, the MHC complex is also known as the human leukocyte antigen (HLA) complex. For a detailed description of the MHC and HLA complexes, see Paul, Fundamental Immunology, 3 rd Ed., Raven Press, New See York (1993). "Major histocompatibility complex (MHC) proteins or molecules", "MHC molecules", "MHC proteins" or "HLA proteins" should be understood to mean proteins that are capable of binding peptides resulting from proteolytic cleavage of protein antigens, presenting potential lymphocyte epitopes (e.g. T-cell and B-cell epitopes) and transporting them to the cell surface where they are presented to specific cells, in particular cytotoxic T-lymphocytes, helper T-cells or B-cells. The major histocompatibility complex in the genome comprises gene regions whose gene products expressed on the cell surface are important for binding and presenting endogenous and / or foreign antigens and therefore for regulating immunological processes. The major histocompatibility complex is divided into two groups of genes that code for different proteins, namely molecules of MHC class I and molecules of MHC class II. The cell biology and expression patterns of the two MHC classes are compatible with these different roles.

[0100] "Human leukocyte antigen" or "HLA" refers to a human class I or class II major histocompatibility complex (MHC) protein (see, e.g., Stites, et al., Immunology, 8 th Ed., Lange Publishing, Los Altos, Calif. (1994).

[0101] "Peptide-MHC (pMHC) stability" refers to the length of time it takes for half of a particular peptide to dissociate from its cognate HLA in a biochemical assay.

[0102] "Antigen-presenting cells" (APCs) are cells that present peptide fragments of protein antigens associated with MHC molecules on their cell surface. Some APCs can activate antigen-specific T cells. Mature professional antigen-presenting cells internalize antigens either by phagocytosis or by receptor-mediated endocytosis and then promote the expression of class II antigens. They are highly efficient at displaying fragments of antigens bound to MHC molecules on their membranes. T cells recognize and interact with the antigen-class II MHC molecule complex on the membrane of antigen-presenting cells. Additional costimulatory signals are then generated by the antigen-presenting cells, leading to T cell activation. The expression of costimulatory molecules is a defining feature of professional antigen-presenting cells. The major types of professional antigen-presenting cells are dendritic cells, macrophages, B cells, and certain activated epithelial cells, which have the broadest range of antigen presentation and are perhaps the most important antigen-presenting cells. "Dendritic cells (DCs)" are a population of leukocytes that present antigens captured in peripheral tissues to T cells via both the MHC class II and MHC class I antigen presentation pathways. It is well known that dendritic cells are potent inducers of immune responses and that activation of these cells is a crucial step for the induction of antitumor immunity. Dendritic cells are conveniently categorized into "immature" and "mature" cells, which can be used as a simple way to distinguish between two well-characterized phenotypes. However, this nomenclature should not be interpreted as excluding all possible intermediate stages of differentiation. Immature dendritic cells are characterized as antigen-presenting cells with high capacity for antigen uptake and processing, which correlates with high expression of Fc receptor (FcR) and mannose receptor. The mature phenotype is generally characterized by low expression of these markers, but high expression of cell surface molecules responsible for T cell activation, such as class I and class II MHC, adhesion molecules (e.g., CD54 and CD11), and costimulatory molecules (e.g., CD40, CD80, CD86, and 4-1BB).

[0103] The terms "polynucleotide", "nucleotide", "nucleic acid", "polynucleic acid", or "oligonucleotide" and their grammatical equivalents are used interchangeably herein and refer to polymers of nucleotides of any length, including DNA and RNA, such as mRNA. Thus, these terms encompass double- and single-stranded DNA, triple-stranded DNA, and double- and single-stranded RNA. The term also includes modified forms of polynucleotides, e.g., by methylation and / or by capping, as well as unmodified forms of polynucleotides. The term is also intended to include molecules that contain non-naturally occurring or synthetic nucleotides and nucleotide analogs. The nucleic acid sequences and vectors disclosed or contemplated herein can be introduced into cells, e.g., by transfection, transformation, or transduction. The nucleotides can be deoxyribonucleotides, ribonucleotides, modified nucleotides or bases, and / or their analogs, or any substrate that can be incorporated into a polymer by DNA or RNA polymerase. In some embodiments, the polynucleotides and nucleic acids can be in vitro transcribed mRNA. In some embodiments, the polynucleotide administered using the methods of the present disclosure is mRNA.

[0104] "Reference" is something that can be used to correlate and compare the results obtained from tumor specimen in the method of the present disclosure.Generally, "reference" can be obtained based on one or more normal specimens obtained from either a patient or one or more different individuals, such as healthy individuals, particularly individuals of the same species, particularly specimens that are not affected by cancer disease."Reference" can be empirically determined by testing a sufficient number of normal specimens.

[0105] The term "mutation" or "mutant" refers to a change or difference (nucleotide substitution, addition, insertion, or deletion) in a nucleic acid sequence compared to a reference. "Somatic mutations" can occur in any cell of the body other than germ cells (sperm and eggs) and therefore are not passed on to children. These changes can (but do not necessarily) cause cancer or other diseases. In some embodiments, the mutation is a nonsynonymous mutation. The term "nonsynonymous mutation" refers to a mutation that results in an amino acid change, such as an amino acid substitution in the translation product, e.g., a nucleotide substitution. A "frameshift" occurs when a mutation disrupts the normal phase of the codon periodicity (also known as the "reading frame") of a gene, resulting in the translation of a non-native protein sequence. Different mutations in a gene can achieve the same altered reading frame. When an open reading frame (ORF) is altered by various mutational events in the genome, such as missense mutations, fusion transcripts, frameshifts, and / or loss of a stop codon, a "neo-ORF" can be created. A neo-ORF can code for a novel amino acid sequence that is not present in the normal genome.

[0106] "Conservative amino acid substitution" refers to a substitution in which one amino acid residue is replaced with another amino acid residue having a similar side chain. Families of amino acid residues having similar side chains have been defined in the art, and include basic side chains (e.g., lysine, arginine, histidine), acidic side chains (e.g., aspartic acid, glutamic acid), uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine), non-polar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), beta-branched side chains (e.g., threonine, valine, isoleucine) and aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine). For example, the replacement of tyrosine with phenylalanine is a conservative substitution. Methods for identifying conservative substitutions of nucleotides and amino acids that do not eliminate peptide function are well known in the art.

[0107] "Native" or "wild-type" sequence refers to a sequence found in nature. Such a sequence may include a longer sequence found in nature.

[0108] As used herein, the term "affinity" refers to a measure of the strength of binding between two members of a binding pair, e.g., an HLA-binding peptide and class I or II HLA. D is the dissociation constant and has units of molar concentration. The affinity constant is the reciprocal of the dissociation constant. Sometimes affinity constant is used as a general term to describe this chemical entity. The affinity constant is a direct measure of the energy of binding. Affinity can be determined experimentally, for example, by surface plasmon resonance (SPR) using a commercially available Biacore SPR unit. Affinity is expressed as the inhibitory concentration 50 (IC 50 ), which is the concentration at which 50% of the peptide is displaced. Similarly, ln(IC 50 ) is IC 50 It refers to the natural logarithm of K off refers to, for example, the dissociation rate constant for the dissociation of an HLA-binding peptide with class I or II HLA. Throughout this disclosure, the results of "binding data" or "binding analysis" are referred to as "IC 50 " can be expressed as IC 50 is the concentration of the tested peptide at which 50% inhibition of binding of the labeled reference peptide is observed in the binding assay. Given the conditions under which the assay is performed (i.e., the limiting HLA protein and labeled reference peptide concentrations), these values ​​are D Assays for determining binding are well known in the art and are described, for example, in PCT Publication Nos. WO94 / 20127 and WO94 / 03205, and in Sidney et al., Current Protocols in Immunology 18.3.1 (1998); Sidney, et al., J. Immunol. 154:247 (1995); and Sette, et al., Mol. Immunol. 31:813 (1994). Alternatively, binding can be expressed relative to binding by a reference standard peptide. For example, the IC 50 Compared to IC 50 Binding can also be determined using other assay systems, including those using live cells (e.g., Ceppellini et al., Nature 339: 392 (1989); Christnick et al., Nature 352: 67 (1991); Busch et al., Int. Immunol. 2: 443 (1990); Hill et al., J. Immunol. 147: 189 (1991); del Guercio et al., J. Immunol. 154: 685 (1995)), non-cellular antibodies using detergent lysates, and antibodies against IgG. Cellular systems (e.g., Cerundolo et al., J. Immunol. 21: 2069 (1991)), immobilized purified MHC (e.g., Hill et al., J. Immunol. 152, 2890 (1994); Marshall et al., J. Immunol. 152, 2890 (1995)), and immunoglobulins (e.g., IgG1, IgG2, IgG3, IgG4, IgG5, IgG6, IgG7, IgG8, IgG9, IgG10, IgG11, IgG12, IgG13, IgG14, IgG15, IgG16, IgG17, IgG18, IgG19, IgG20, IgG21, IgG15, IgG16, IgG17, IgG18, IgG19, IgG20, IgG11, IgG12, et al., J. Immunol. 152: 4946 (1994)), ELISA systems (e.g., Reay et al., EMBO J. 11: 2829 (1992)), surface plasmon resonance (e.g., Khilko et al., J. Biol. Chem. 268: 15425 (1993)); high flow soluble phase assays (Hammer et al., J. Exp. Med. 180: 2353 (1994)), and measurements of class I MHC stabilization or assembly (e.g., Ljunggren et al., Nature 346: 476 (1990); Schumacher et al., Cell 62: 563 (1990); Townsend et al., Cell 62: 285 (1990); Parker et al., J. Immunol. 149: 1896 (1996)). (1992)). "Cross-reactive binding" indicates that a peptide is bound by more than one HLA molecule; a synonym is degenerate binding.

[0109] The term "naturally occurring" and its grammatical equivalents, as used herein, refers to the fact that an object can be found in nature. For example, a peptide or nucleic acid that is present in an organism (including viruses), can be isolated from a natural source, and has not been intentionally modified by humans in a laboratory, is naturally occurring.

[0110] "Antigen processing" or "processing" and its grammatical equivalents refer to the degradation of a polypeptide or antigen into processing products that are fragments of said polypeptide or antigen (e.g., degradation of a polypeptide into peptides), and the association (e.g., by binding) of these fragments to one or more MHC molecules for presentation by a cell, e.g., an antigen-presenting cell, to a specific T cell.

[0111] The term "subject" refers to any animal (e.g., mammal), including but not limited to humans, non-human primates, canines, felines, rodents, etc., that will be the recipient of a particular treatment. In general, the terms "subject" and "patient" are used interchangeably herein in reference to human subjects.

[0112] "Cell" and its grammatical equivalents refer to a cell of human or non-human animal origin.

[0113] "T cells" are CD4 + T cells and CD8 + Includes T cells. The term T cells also includes both T helper type 1 T cells and T helper type 2 T cells.

[0114] According to the present disclosure, the term "vaccine" refers to a pharmaceutical preparation (composition) or product that, when administered, induces an immune response, e.g., a cellular or humoral immune response, that recognizes and attacks pathogens or diseased cells, such as cancer cells. Vaccines can be used for the prevention or treatment of disease. The term "personalized cancer vaccine" or "individualized cancer vaccine" refers to a specific cancer patient and means that the cancer vaccine is adapted to the needs or special circumstances of the individual cancer patient.

[0115] The term "effective amount" or "therapeutically effective amount" or "therapeutic effect" refers to an amount that is therapeutically effective to "treat" a disease or disorder in a subject or mammal. A therapeutically effective amount of a drug has a therapeutic effect and may thus prevent the onset of a disease or disorder; slow the onset of a disease or disorder; slow the progression of a disease or disorder; relieve to some extent one or more of the symptoms associated with a disease or disorder; reduce morbidity and mortality; improve quality of life; or a combination of such effects.

[0116] The terms "treating" or "treatment" or "to treat" or "alleviating" or "to alleviate" refer to both (1) therapeutic measures that cure, slow, relieve the symptoms of, and / or halt the progression of a diagnosed pathological condition or disorder; and (2) prophylactic or preventative measures that prevent or slow the onset of the targeted pathological condition or disorder. Thus, those in need of treatment include those already with the disorder; those prone to having the disorder; and those in whom the disorder is to be prevented.

[0117] "Pharmaceutically acceptable" generally refers to a composition or component of a composition that is non-toxic, inert, and / or physiologically compatible.

[0118] A "pharmaceutical excipient" or "excipient" includes materials such as adjuvants, carriers, pH adjusting and buffering agents, tonicity adjusting agents, wetting agents, preservatives, etc. A "pharmaceutical excipient" is an excipient that is pharma- ceutically acceptable.

[0119] "Immunomodulators" or grammatical equivalents, as used herein, may refer to substances that can stimulate or suppress the immune system and help an individual's body fight disease, such as infection, cancer, etc. Examples of specific immunomodulators that affect specific parts of the immune system include, but are not limited to, monoclonal antibodies, cytokines, and vaccines. Non-specific immunomodulators affect the immune system generally, and non-limiting examples of these include Bacillus Calmette-Guerin (BCG) and levamisole.

[0120] The term "cancer" and its grammatical equivalents, as used herein, may refer to the hyperproliferation of cells, which is characterized by the loss of normal regulation, resulting in uncontrolled growth, lack of differentiation, local tissue invasion, and metastasis. In the context of the compositions and methods of the present invention, cancer includes acute lymphocytic cancer, acute myeloid leukemia, alveolar rhabdomyosarcoma, bladder cancer, bone cancer, brain cancer, breast cancer, anus, anal canal, rectal cancer, eye cancer, intrahepatic bile duct cancer, joint cancer, neck, gallbladder, or pleural cancer, nose, nasal cavity, or middle ear cancer, oral cavity cancer, vulva cancer, chronic lymphocytic leukemia, chronic myeloid cancer, colon cancer, esophageal cancer, cervical cancer, fibrosarcoma, gastrointestinal carcinoid tumor, Hodgkin's lymphoma, hypopharyngeal cancer, kidney cancer, laryngeal cancer, and the like. The tumor may be any cancer, including any of the following: leukemia, liquid tumors, liver cancer, lung cancer, lymphoma, malignant mesothelioma, mast cell tumor, melanoma, multiple myeloma, nasopharyngeal cancer, non-Hodgkin's lymphoma, ovarian cancer, pancreatic cancer, peritoneal, omental, and mesenteric cancer, pharyngeal cancer, prostate cancer, rectal cancer, renal cancer, skin cancer, small intestine cancer, soft tissue cancer, solid tumors, stomach cancer, testicular cancer, thyroid cancer, ureteral cancer, and / or urinary bladder cancer. As used herein, the term "tumor" refers to an abnormal growth of cells or tissue, e.g., of malignant or benign type.

[0121] The term "exome" refers to the portion of the genome that encodes a functional protein, or the sequence encompassing all the exons, or coding regions, of the protein-coding genes in the genome. The exome is approximately 1-2% of the whole genome, depending on the species.

[0122] "Diluents" include sterile liquids, such as water and oils, including those of petroleum, animal, vegetable or synthetic origin, for example, peanut oil, soybean oil, mineral oil, sesame oil, and the like. Water is also a diluent for pharmaceutical compositions. Saline solutions and aqueous dextrose and glycerol solutions can also be employed as diluents, for example, for injectable solutions.

[0123] "Receptor" should be understood to mean a biological molecule or group of molecules that can bind to a ligand. Receptors can be useful for transmitting information in a cell, cell formation, or organism. A receptor comprises at least one receptor unit, and each receptor unit can be, for example, a protein molecule. A receptor has a structure that complements the structure of a ligand and can form a complex with the ligand as a binding partner. Information is transmitted by a change in the conformation of the receptor after complexing with the ligand on the surface of a cell. In some embodiments, receptors should be understood to mean, in particular, proteins of MHC class I and II that can form a receptor / ligand complex with a ligand, in particular a peptide or peptide fragment of appropriate length.

[0124] A "ligand" should be understood to mean a molecule that has a structure complementary to that of a receptor and can form a complex with this receptor. In some embodiments, a ligand should be understood to mean a peptide or peptide fragment that has the appropriate length and the appropriate binding motif in the amino acid sequence, and thus can form a complex with a protein of MHC class I or MHC class II.

[0125] In some embodiments, "receptor / ligand complex" should also be understood to mean a "receptor / peptide complex" or a "receptor / peptide fragment complex" comprising a peptide-presenting or peptide fragment-presenting class I or class II MHC molecule.

[0126] The term "motif" refers to a pattern of residues within a defined length of amino acid sequence that is recognized by a particular HLA molecule, for example, a peptide less than about 15 amino acid residues in length, or less than about 13 amino acid residues in length, for example, about 8 to about 13 amino acid residues (e.g., 8, 9, 10, 11, 12, or 13) for class I HLA motifs, and about 6 to about 25 amino acid residues (e.g., 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25) for class II HLA motifs. Motifs are generally different for each HLA protein encoded by a given human HLA allele. These motifs differ in the pattern of primary and secondary anchor residues. In some embodiments, MHC class I motifs identify peptides that are 9 amino acid residues, 10 amino acid residues, or 11 amino acid residues in length.

[0127] As used herein, the term "identical" and its grammatical equivalents, or "sequence identity" in reference to the amino acid sequences of two nucleic acid sequences or polypeptides, refers to the residues in the two sequences being the same when aligned for maximum correspondence over a specified comparison window. A "comparison window," as used herein, refers to a segment of at least about 20, usually about 50 to about 200, more usually about 100 to about 150 contiguous positions in which a sequence can be compared to a reference sequence of the same number of contiguous positions after optimally aligning the two sequences. Methods for aligning sequences for comparison are well known in the art. Optimal alignment of sequences for comparison can be achieved by the local homology algorithm of Smith and Waterman, Adv. Appl. Math., 2: 482 (1981). Alignment of Needleman and Wunsch, J. Mol. Biol., 48: 443 (1970) by the similarity search method of Pearson and Lipman, Proc. Nat. Acad. Sci. USA, 85: 2444 (1988); The methods can be implemented using programs such as CLUSTAL, GAP, BESTFIT, BLAST, FASTA, and TFASTA in the PC / Gene program by Intelligentics, Mountain View Calif., in the Wisconsin Genetics Software Package, Genetics Computer Group (GCG), 575 Science Dr., Madison, Wis., USA; the CLUSTAL program is described in Higgins and Sharp, Gene, 73: 237-244 (1988) and Higgins and Sharp, CABIOS, 5: 151-153 (1989); Corpet et al., Nucleic Acids Res., 16: 10881-10890 (1988); Huang et al., Computer Applications in the Biosciences, 8: 155-165 (1992); and Pearson et al., Methods in Molecular Biology, 24: 307-331 (1994). Alignment is also often performed by inspection and manual alignment. In one class of embodiments, the polypeptides herein have at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a reference polypeptide or fragment thereof, as measured, for example, by BLASTP (or CLUSTAL, or any other available alignment software) using default parameters. Similarly, a nucleic acid can be described with reference to a starting nucleic acid, and the nucleic acid can have, for example, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 98%, 99%, or 100% sequence identity to a reference nucleic acid or a fragment thereof, as measured, for example, by BLASTN (or CLUSTAL, or any other available alignment software) using default parameters. When one molecule is said to have a certain percentage of sequence identity to a larger molecule, it means that when the two molecules are optimally aligned, said percentage of residues in the smaller molecule match the residues in the larger molecule according to the order in which the two molecules are optimally aligned.

[0128] The term "substantially identical" and its grammatical equivalents, when applied to a nucleic acid or amino acid sequence, means that the nucleic acid or amino acid sequence includes sequences having at least 90% or greater, at least 95%, at least 98%, and at least 99% sequence identity compared to a reference sequence using the above programs, e.g., BLAST, using standard parameters. For example, the BLASTN program (for nucleotide sequences) uses as defaults a word length (W) of 11, an expectation (E) of 10, M=5, N=-4, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a word length (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (Henikoff & Henikoff, (See Proc. Natl. Acad. Sci. USA 89: 10915 (1992)). The percentage of sequence identity is determined by comparing two optimally aligned sequences over a comparison window, where in order to optimally align the two sequences, a portion of the polynucleotide sequence within the comparison window may contain additions or deletions (i.e., gaps) compared to the reference sequence (which does not contain additions or deletions). The percentage is calculated by determining the number of positions at which identical nucleic acid bases or amino acid residues exist in both sequences to obtain the number of matching positions, dividing the number of matching positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity. In embodiments, substantial identity exists over a region of the sequences that is at least about 50 residues long, over a region of at least about 100 residues, and in embodiments, the sequences are substantially identical over at least about 150 residues. In embodiments, the sequences are substantially identical over the entire length of the coding regions.

[0129] The term "vector" as used herein refers to a construct that can deliver one or more genes or sequences of interest to host cell and usually express them there.Examples of vectors include, but are not limited to, virus vectors, naked DNA or RNA expression vectors, plasmids, cosmids or phage vectors, DNA or RNA expression vectors that are combined with cationic condensing agents, and DNA or RNA expression vectors that are encapsulated in liposomes.

[0130] An "isolated" polypeptide, antibody, polynucleotide, vector, cell, or composition is a polypeptide, antibody, polynucleotide, vector, cell, or composition in a form not found in nature. An isolated polypeptide, antibody, polynucleotide, vector, cell, or composition includes one that has been purified to the extent that it is no longer in a form found in nature. In some embodiments, an isolated polypeptide, antibody, polynucleotide, vector, cell, or composition is substantially pure. In some embodiments, an "isolated polynucleotide" encompasses a PCR or quantitative PCR reaction that includes a polynucleotide amplified in a PCR or quantitative PCR reaction.

[0131] The terms "isolated", "biologically pure" or their grammatical equivalents refer to a material that is substantially or essentially free from components that normally accompany the material when it is found in its native state. Thus, the isolated peptides described herein do not contain some or all of the materials that normally accompany them in the peptide's in situ environment. An "isolated" epitope refers to an epitope that does not contain the entire sequence of the antigen from which the epitope is derived. In general, an "isolated" epitope does not have additional amino acid residues attached that result in a sequence with 100% identity over the entire length of the native sequence. The native sequence can be a sequence such as a tumor-associated antigen from which the epitope is derived. Thus, the term "isolated" means that the material is removed from its original environment (e.g., the natural environment if it occurs in nature). An "isolated" nucleic acid is a nucleic acid that has been removed from its natural environment. For example, a naturally occurring polynucleotide or peptide present in a living animal is not isolated, but the same polynucleotide or peptide separated from some or all of the coexisting materials in the natural system is isolated. Such polynucleotides can be part of a vector, and / or such polynucleotides or peptides can be part of a composition, which is nevertheless "isolated" in that such vectors or compositions are not part of their natural environment. Isolated RNA molecules include in vivo or in vitro RNA transcripts of the DNA molecules described herein, and further include such molecules produced synthetically.

[0132] The term "substantially pure," as used herein, refers to a material that is at least 50% pure (i.e., free from contaminants), at least 90% pure, at least 95% pure, at least 98% pure, or at least 99% pure.

[0133] "Transfection", "transformation" or "transduction" as used herein refers to the introduction of one or more exogenous polynucleotides into a host cell by using physical or chemical methods. Many transfection techniques are known in the art, including, for example, calcium phosphate DNA co-precipitation (see, e.g., Murray EJ (ed.), Methods in Molecular Biology, Vol. 7, Gene Transfer and Expression Protocols, Humana Press (1991)); DEAE-dextran; electroporation; cationic liposome-mediated transfection; tungsten particle-facilitated microparticle bombardment (Johnston, Nature, 346: 776-777 (1990)); and strontium phosphate DNA coprecipitation (Brash et al., Mol. Cell Biol., 7: Phage vectors or viral vectors are suitable for use in infectious diseases. The chromogenic particles can be grown in suitable packaging cells and then introduced into a host cell, many of which are commercially available. 2. Enhanced cleavage and uses thereof

[0134] One of the crucial barriers for the development of curative and tumor-specific immunotherapy is the insufficient processing and release of minimal epitopes for antigen presentation to generate an adequate immune response. Antigen processing and presentation refers to the process that occurs within cells that results in protein fragmentation or proteolysis, association of protein fragments or peptides with major histocompatibility complex (MHC) molecules, and expression of peptide-MHC (pMHC) molecules on the cell surface for recognition by the T cell receptor (TCR) on T cells. Antigen presentation is mediated by MHC class I and MHC class II molecules found on the surface of antigen-presenting cells (APCs) and certain other cells. MHC class I and MHC class II molecules deliver short peptides to the cell surface, which then mediate their cytotoxic (CD8) and CD8+) effects, respectively. + ) T cells and helper (CD4 + ) can be recognized by T cells. TCRs can only recognize antigens in the form of peptides bound to MHC molecules on the cell surface, and the antigens recognized by T cells are peptides resulting from the degradation of the macromolecular structure, unfolding of individual proteins, and their cleavage into short fragments by antigen processing.

[0135] Antigen presentation on cell surface requires accurate processing of peptides by proteasomes to release minimal epitopes, cytosolic and endoplasmic reticulum (ER) aminopeptidases, efficient transporter associated with antigen processing (TAP) transport, and sufficient binding to MHC class I molecules. The efficiency of epitope generation depends not only on the epitope itself, but also on its adjacent regions or amino acid sequences adjacent to the epitope amino acid sequence. The efficiency of minimal epitope processing from peptides containing epitope sequences and amino acid sequences adjacent to the epitope sequence is not fully understood, but is known to be influenced by a number of factors, including specific amino acid residues on both sides of the cleavage site in the peptide and other competing cleavage sites nearby.

[0136] One way to address the problem of insufficient processing and release of minimal epitopes is to test and design specific amino acid residues or sequences that can be added to the N-terminus and / or C-terminus of the epitope sequence to enhance peptide cleavage and processing and epitope presentation. For example, amino acid residues or sequences from other epitopes known to be efficiently processed can be added to the epitope sequence. Another example is to use amino acid residues known to be commonly found around epitopes (Abelin, et al., 2017, Immunity 46, 315-326). This approach can confer additional benefits, including facilitating peptide production (e.g., synthesis, purification, and / or formulation) or easy downstream modification (e.g., conjugation with other molecules).

[0137] Another way to address the current barriers of efficient processing and release of minimal epitopes is to use protease-cleavable linkers to target epitope-containing peptides for site-specific protease processing to release epitopes.For example, specific linkers that can be easily cleaved inside dendritic cells (DCs) to release minimal epitope sequences can be used to enhance CD8-dependent immune responses after vaccination.These peptides also do not have non-selective binding to MHC class I molecules on the surface of non-professional APCs, but instead pass through specific (e.g., endocytosis) pathways to be appropriately processed and presented to T cells.Furthermore, another example of promoting sufficient epitope processing and presentation is to combine two strategies, namely, specific amino acid residues and specific linkers.

[0138] Provided herein is a polypeptide comprising an epitope sequence encoded by a subject's genome, an amino acid or amino acid sequence that may or may not be encoded by a nucleic acid sequence immediately upstream or downstream of the nucleic acid sequence encoding the epitope sequence in the subject's genome, an amino acid or amino acid sequence, and / or a linker. The addition of an amino acid, amino acid sequence, and / or a linker to the epitope sequence can enhance the processing and presentation of the epitope by APCs to generate an immune response. In one aspect, the amino acid or amino acid sequence is of an amino acid sequence or peptide sequence. In one embodiment, the amino acid sequence or peptide sequence is not encoded by a nucleic acid sequence immediately upstream or downstream of the nucleic acid sequence in the subject's genome that encodes the epitope sequence. In another embodiment, the amino acid or amino acid sequence is contiguous with the epitope sequence and is encoded by the subject's genome that encodes the epitope sequence. For example, the amino acid or amino acid sequence contiguous with the epitope sequence can include one or more amino acid residues (e.g., lysine) that enhance the cleavage of the polypeptide. In such embodiments, the polypeptide may comprise an amino acid or amino acid sequence that is contiguous with the epitope sequence and may further comprise an amino acid or amino acid sequence that is not encoded by a nucleic acid sequence immediately upstream or downstream of the nucleic acid sequence that encodes the epitope sequence in the genome of the subject.

[0139] In some embodiments, the epitope is presented by class I MHC of the APC. In some embodiments, the epitope is presented by class II MHC of the APC. In some embodiments, each amino acid of the epitope represents an amino acid of a peptide sequence that includes any contiguous amino acid sequence encoded by a nucleic acid sequence in the genome of the subject. In some embodiments, the epitope comprises 8-12 contiguous amino acid residues and is presented by class I MHC of the APC. In some embodiments, the epitope comprises 8, 9, 10, 11, or 12 contiguous amino acid residues and is presented by class I MHC of the APC. In some embodiments, the epitope comprises 9-25 contiguous amino acid residues and is presented by class II MHC of the APC. In some embodiments, the epitope comprises 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 consecutive amino acid residues and is presented by class II MHC of APC. In some embodiments, the epitope sequence comprises 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 consecutive amino acid residues, of which one or more of amino acids 13-25 are optionally present and at least one amino acid is a mutant amino acid. In some embodiments, the epitope sequence comprises AA 1 AA 2 AA 3 AA 4 AA 5 AA 6 AA 7 AA 8 AA 9 AA 10 AA 11 AA 12 AA 13 AA 14 AA 15 AA 16 AA 17 AA 18 AA 19 AA 20 AA 21 AA 22 AA23 AA 24 AA 25 where each AA is an amino acid, 9 , A.A. 10 , A.A. 11 , A.A. 12 , A.A. 13 , A.A. 14 , A.A. 15 , A.A. 16 , A.A. 17 , A.A. 18 , A.A. 19 , A.A. 20 , A.A. 21 , A.A. 22 , A.A. 23 , A.A. 24 , and A.A. 25 is optionally present, where at least one AA is a mutated amino acid.

[0140] In some embodiments, a polypeptide comprising an epitope sequence and an amino acid or amino acid sequence that is contiguous with the epitope sequence and is encoded by a nucleic acid sequence immediately upstream or downstream of a nucleic acid sequence that encodes the epitope in the genome of a subject may not include a linker. In some embodiments, a polypeptide comprising an epitope sequence and an amino acid or amino acid sequence that is contiguous with the epitope sequence and is encoded by a nucleic acid sequence immediately upstream or downstream of a nucleic acid sequence that encodes the epitope in the genome of a subject may include a linker. In some embodiments, a polypeptide comprising an epitope sequence and an amino acid or amino acid sequence that is not encoded by a nucleic acid sequence immediately upstream or downstream of a nucleic acid sequence in the genome of a subject that encodes the epitope sequence may further include a linker. In some embodiments, a polypeptide comprising an epitope sequence and an amino acid or amino acid sequence that is not encoded by a nucleic acid sequence immediately upstream or downstream of a nucleic acid sequence in the genome of a subject that encodes the epitope sequence may not include a linker.

[0141] In some embodiments, the amino acid or amino acid sequence comprises 0-1000 amino acid residues in length. In some embodiments, the amino acid or amino acid sequence encoded by a nucleic acid sequence immediately upstream of a nucleic acid sequence in a subject's genome encoding an epitope comprises 0-1000 amino acid residues in length. In some embodiments, the amino acid or amino acid sequence encoded by a nucleic acid sequence immediately downstream of a nucleic acid sequence in a subject's genome encoding an epitope comprises 0-1000 amino acid residues in length. In some embodiments, the amino acid or amino acid sequence is longer than 0 amino acid residues, longer than 1 amino acid residue, longer than 2 amino acid residues, longer than 3 amino acid residues, longer than 4 amino acid residues, longer than 5 amino acid residues, longer than 6 amino acid residues, longer than 7 amino acid residues, longer than 8 amino acid residues, longer than 9 amino acid residues, longer than 10 amino acid residues, longer than 15 amino acid residues, longer than 20 amino acid residues, longer than 25 amino acid residues, longer than 30 amino acid residues, longer than 35 amino acid residues, longer than 40 amino acid residues, longer than 45 amino acid residues, longer than 50 amino acid residues, longer than 55 amino acid residues, longer than 60 amino acid residues, longer than 65 amino acid residues, longer than 70 amino acid residues, longer than 75 amino acid residues. In some embodiments, the amino acid residue length may be longer than 80 amino acid residues, longer than 85 amino acid residues, longer than 90 amino acid residues, longer than 95 amino acid residues, longer than 100 amino acid residues, longer than 150 amino acid residues, longer than 200 amino acid residues, longer than 250 amino acid residues, longer than 300 amino acid residues, longer than 350 amino acid residues, longer than 400 amino acid residues, longer than 450 amino acid residues, longer than 500 amino acid residues, longer than 550 amino acid residues, longer than 600 amino acid residues, longer than 650 amino acid residues, longer than 700 amino acid residues, longer than 750 amino acid residues, longer than 800 amino acid residues, longer than 850 amino acid residues, longer than 900 amino acid residues, or longer than 950 amino acid residues.In some embodiments, the amino acid or amino acid sequence encoded by a nucleic acid sequence immediately upstream of a nucleic acid sequence in a subject's genome encoding an epitope is longer than 0 amino acid residues, longer than 1 amino acid residue, longer than 2 amino acid residues, longer than 3 amino acid residues, longer than 4 amino acid residues, longer than 5 amino acid residues, longer than 6 amino acid residues, longer than 7 amino acid residues, longer than 8 amino acid residues, longer than 9 amino acid residues, longer than 10 amino acid residues, longer than 15 amino acid residues, longer than 20 amino acid residues, longer than 25 amino acid residues, longer than 30 amino acid residues, longer than 35 amino acid residues, longer than 40 amino acid residues, longer than 45 amino acid residues, longer than 50 amino acid residues, longer than 55 amino acid residues, longer than 60 amino acid residues, longer than 65 ... The invention includes lengths of amino acid residues longer than 0 amino acid residues, longer than 75 amino acid residues, longer than 80 amino acid residues, longer than 85 amino acid residues, longer than 90 amino acid residues, longer than 95 amino acid residues, longer than 100 amino acid residues, longer than 150 amino acid residues, longer than 200 amino acid residues, longer than 250 amino acid residues, longer than 300 amino acid residues, longer than 350 amino acid residues, longer than 400 amino acid residues, longer than 450 amino acid residues, longer than 500 amino acid residues, longer than 550 amino acid residues, longer than 600 amino acid residues, longer than 650 amino acid residues, longer than 700 amino acid residues, longer than 750 amino acid residues, longer than 800 amino acid residues, longer than 850 amino acid residues, longer than 900 amino acid residues, or longer than 950 amino acid residues.In some embodiments, the amino acid or amino acid sequence encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the subject's genome encoding the epitope is longer than 0 amino acid residues, longer than 1 amino acid residue, longer than 2 amino acid residues, longer than 3 amino acid residues, longer than 4 amino acid residues, longer than 5 amino acid residues, longer than 6 amino acid residues, longer than 7 amino acid residues, longer than 8 amino acid residues, longer than 9 amino acid residues, longer than 10 amino acid residues, longer than 15 amino acid residues, longer than 20 amino acid residues, longer than 25 amino acid residues, longer than 30 amino acid residues, longer than 35 amino acid residues, longer than 40 amino acid residues, longer than 45 amino acid residues, longer than 50 amino acid residues, longer than 55 amino acid residues, longer than 60 amino acid residues, longer than 65 ... The invention includes lengths of amino acid residues longer than 0 amino acid residues, longer than 75 amino acid residues, longer than 80 amino acid residues, longer than 85 amino acid residues, longer than 90 amino acid residues, longer than 95 amino acid residues, longer than 100 amino acid residues, longer than 150 amino acid residues, longer than 200 amino acid residues, longer than 250 amino acid residues, longer than 300 amino acid residues, longer than 350 amino acid residues, longer than 400 amino acid residues, longer than 450 amino acid residues, longer than 500 amino acid residues, longer than 550 amino acid residues, longer than 600 amino acid residues, longer than 650 amino acid residues, longer than 700 amino acid residues, longer than 750 amino acid residues, longer than 800 amino acid residues, longer than 850 amino acid residues, longer than 900 amino acid residues, or longer than 950 amino acid residues.

[0142] In some embodiments, the amino acid or amino acid sequence comprises a length of 1-5 or 7-1000 amino acid residues. In some embodiments, the amino acid or amino acid sequence does not comprise a length of 6 amino acid residues. In some embodiments, the amino acid or amino acid sequence of the peptide sequence encoded by the nucleic acid sequence immediately upstream of the nucleic acid sequence in the subject's genome encoding the epitope comprises a length of 1-5 or 7-1000 amino acid residues. In some embodiments, the amino acid or amino acid sequence of the peptide sequence encoded by the nucleic acid sequence immediately upstream of the nucleic acid sequence in the subject's genome encoding the epitope does not comprise a length of 6 amino acid residues. In some embodiments, the amino acid or amino acid sequence of the peptide sequence encoded by the nucleic acid sequence immediately downstream of the nucleic acid sequence in the subject's genome encoding the epitope comprises a length of 1-4 or 6-1000 amino acid residues. In some embodiments, the amino acid or amino acid sequence of the peptide sequence encoded by the nucleic acid sequence immediately downstream of the nucleic acid sequence in the subject's genome encoding the epitope does not comprise a length of 5 amino acid residues.

[0143] In some embodiments, the polypeptide further comprises a linker. In some embodiments, the polypeptide does not consist of four different epitopes presented by class I MHC. In some embodiments, the polypeptide does not include four different epitopes presented by class I MHC. In some embodiments, the polypeptide comprises at least two different epitopes presented by class I MHC. In some embodiments, the polypeptide comprises at least three, at least five, or at least six different epitopes presented by class I MHC. In some embodiments, the epitope comprises at least one mutant amino acid. In some embodiments, the at least one mutant amino acid is encoded by an insertion, deletion, frameshift, neo-ORF, or point mutation in a nucleic acid sequence in the subject's genome. In some embodiments, an amino acid or amino acid sequence of the peptide sequence that is not encoded by a nucleic acid sequence immediately downstream or upstream of a nucleic acid sequence in the subject's genome that encodes the epitope is cleaved from the epitope when the polypeptide is processed by the APC. In some embodiments, the polypeptide comprises at least two different polypeptide molecules. In some embodiments, the polypeptide comprises at least three, at least four, or at least five different polypeptide molecules.

[0144] In some embodiments, the present disclosure includes polypeptides that include amino acids or amino acid sequences and / or linkers of peptide sequences that are not encoded by a nucleic acid sequence immediately downstream or upstream of a nucleic acid sequence in a subject's genome that encodes an epitope. The amino acids or amino acid sequences and / or linkers can provide the polypeptide with desired properties, such as solubility, stability, immunogenicity, antigen processing, or increased antigen presentation. In some embodiments, the polypeptides can include amino acids or amino acid sequences that enhance the processing and presentation of the epitope by APCs, for example, to generate an immune response. In some embodiments, the polypeptides can include amino acids or amino acid sequences at either the N-terminus and / or C-terminus of the epitope sequence. In some embodiments, the amino acids or amino acid sequences can include poly-lysine (poly-Lys or poly-K) or poly-arginine (poly-Arg or poly-R). In some embodiments, the amino acids or amino acid sequences can be of a polypeptide sequence of a protein that is not expressed in the subject that expresses the epitope (e.g., not encoded by the subject's genome that encodes the epitope sequence). In another embodiment, the polypeptide may include a linker cleavable by a protease. In some embodiments, the polypeptide may include both a linker cleavable by a protease and an amino acid or amino acid sequence. In some embodiments, provided herein are polypeptides of formula (I), (II), (III), and / or (IV), or pharma- ceutically acceptable salts of polypeptides of formula (I), (II), (III), and / or (IV), where the stereochemistry is indeterminate, e.g., a racemate or a mixture of diastereomers or individual diastereomers. It will be understood by those skilled in the art that at any stage in the preparation of the compounds of formula (I), (II), (III), and / or (IV), a mixture of isomers (e.g., a racemate) of the compounds corresponding to any of formulas (I), (II), (III), and / or may be utilized. At any stage in the preparation, a single stereoisomer may be obtained by isolation from a mixture of isomers (e.g., a racemate) using, for example, chiral chromatographic separation.

[0145] In some embodiments, the linker comprises a non-polypeptide linker. In some embodiments, the linker comprises a chemical linker. In some embodiments, the linker comprises a non-natural amino acid. In some embodiments, the non-natural amino acid comprises a β-γ-δ-amino acid. In some embodiments, the non-natural amino acid comprises a derivative of an L-α-amino acid. In some embodiments, the linker does not comprise an amino acid. In some embodiments, the linker does not comprise a natural amino acid. In some embodiments, the linker comprises a bond other than a peptide bond. In some embodiments, the linker comprises a disulfide bond. In some embodiments, the polypeptides described herein comprise more than one linker. In some embodiments, the polypeptides described herein comprise a first linker and a second linker, the first linker being at the N-terminus of the epitope and the second linker being at the C-terminus of the epitope. In some embodiments, the first linker and the second linker are different. In some embodiments, the first linker and the second linker are the same.

[0146] In some embodiments, the polypeptide comprises a hydrophilic tail. In some embodiments, a polypeptide comprising an epitope sequence, an amino acid or amino acid sequence of a peptide sequence not encoded by a nucleic acid sequence immediately downstream or upstream of a nucleic acid sequence in the genome of a subject encoding the epitope, and / or a linker has enhanced solubility compared to a polypeptide comprising the same epitope sequence but without the amino acid or amino acid sequence and / or linker. In some embodiments, a polypeptide comprising an epitope sequence and an amino acid or amino acid sequence contiguous with the epitope sequence encoded by a nucleic acid sequence in the genome of a subject has enhanced solubility compared to a polypeptide comprising the same epitope sequence but without the amino acid or amino acid sequence. For example, the amino acid or amino acid sequence contiguous with the epitope sequence may comprise one or more amino acid residues (e.g., lysine) that enhance the solubility of the polypeptide. In such embodiments, a polypeptide may comprise an amino acid or amino acid sequence contiguous with the epitope sequence and may further comprise an amino acid or amino acid sequence of a peptide sequence not encoded by a nucleic acid sequence immediately downstream or upstream of a nucleic acid sequence in the genome of a subject encoding the epitope.

[0147] In some embodiments, when the polypeptide is processed by APC, the epitope is released from the polypeptide comprising the epitope sequence. In some embodiments, when the polypeptide further comprises an amino acid or amino acid sequence that does not include at least one additional amino acid and / or a linker that is encoded by the nucleic acid sequence immediately upstream of the nucleic acid sequence in the genome of the subject that encodes the epitope, the epitope is released at a higher rate compared to the polypeptide comprising the same epitope but not including the amino acid or amino acid sequence that does not include at least one additional amino acid and / or a linker that is encoded by the nucleic acid sequence immediately upstream of the nucleic acid sequence in the genome of the subject that encodes the epitope. In some embodiments, when the polypeptide further comprises an amino acid or amino acid sequence that does not include at least one additional amino acid and / or a linker that is encoded by the nucleic acid sequence immediately downstream of the nucleic acid sequence in the genome of the subject that encodes the epitope, the epitope is released at a higher rate compared to the polypeptide comprising the same epitope but not including the amino acid or amino acid sequence that does not include at least one additional amino acid and / or a linker that is encoded by the nucleic acid sequence immediately downstream of the nucleic acid sequence in the genome of the subject that encodes the epitope. In some embodiments, when the amino acid or amino acid sequence is not that of the peptide sequence of the protein expressed in the subject, the epitope is released at a high rate.In some embodiments, when the polypeptide comprises a linker, the epitope is released at a high rate compared to the polypeptide that comprises the same epitope but does not comprise a linker.In some embodiments, when the polypeptide comprises a protease-cleavable linker, the epitope is released at a high rate compared to the polypeptide that comprises the same epitope but does not comprise a protease-cleavable linker.

[0148] In some embodiments, when a polypeptide that comprises an epitope and an amino acid or amino acid sequence that comprises at least one additional amino acid encoded by a nucleic acid sequence immediately upstream or downstream of the nucleic acid sequence in the subject's genome that encodes the epitope further comprises an amino acid or amino acid sequence and / or a linker not encoded by the nucleic acid sequence immediately upstream or downstream of the nucleic acid sequence in the subject's genome that encodes the epitope, a higher percentage of the epitope is released compared to a corresponding polypeptide that comprises the same epitope and an amino acid or amino acid sequence that comprises at least one additional amino acid encoded by a nucleic acid sequence immediately upstream or downstream of the nucleic acid sequence in the subject's genome that encodes the epitope, but does not comprise an amino acid or amino acid sequence and / or a linker not encoded by the nucleic acid sequence immediately upstream or downstream of the nucleic acid sequence in the subject's genome that encodes the epitope.

[0149] In some embodiments, when a polypeptide comprises an amino acid or amino acid sequence and / or a linker that does not include at least one additional amino acid encoded by a nucleic acid sequence immediately upstream of a nucleic acid sequence in the subject's genome that encodes an epitope, the polypeptide is cleaved at a higher rate compared to a corresponding polypeptide with the same length and epitope and an amino acid or amino acid sequence encoded by a nucleic acid sequence immediately upstream of a nucleic acid sequence in the subject's genome that encodes an epitope. In some embodiments, when a polypeptide comprises an amino acid or amino acid sequence and / or a linker that does not include at least one additional amino acid encoded by a nucleic acid sequence immediately downstream of a nucleic acid sequence in the subject's genome that encodes an epitope, the polypeptide is cleaved at a higher rate compared to a corresponding polypeptide with the same length and epitope and an amino acid or amino acid sequence encoded by a nucleic acid sequence immediately downstream of a nucleic acid sequence in the subject's genome that encodes an epitope. In some embodiments, when an amino acid or amino acid sequence is not of a peptide sequence of a protein expressed in the subject, the polypeptide is cleaved at a higher rate. In some embodiments, when a polypeptide comprises a linker, the polypeptide is cleaved at a higher rate compared to a polypeptide with the same epitope but without a linker. In some embodiments, when a polypeptide comprises a protease-cleavable linker, the polypeptide is cleaved at a higher rate compared to a polypeptide comprising the same epitope but without a protease-cleavable linker.

[0150] In some embodiments, when the polypeptide further comprises an amino acid or amino acid sequence that does not include at least one additional amino acid encoded by a nucleic acid sequence immediately upstream of the nucleic acid sequence in the subject's genome encoding the epitope, the polypeptide is cleaved at a higher rate compared to cleavage of a corresponding polypeptide of the same length that includes the epitope sequence and an amino acid or amino acid sequence that is contiguous with the epitope sequence encoded by the nucleic acid sequence and does not include a linker, in some embodiments, when the polypeptide further comprises an amino acid or amino acid sequence that does not include at least one additional amino acid encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the subject's genome encoding the epitope, the polypeptide is cleaved at a higher rate compared to cleavage of a corresponding polypeptide of the same length that includes the epitope sequence and an amino acid or amino acid sequence that is contiguous with the epitope sequence encoded by the nucleic acid sequence and does not include a linker.

[0151] In some embodiments, when a polypeptide comprises (i) an amino acid or amino acid sequence encoded by a nucleic acid sequence immediately upstream or downstream of a nucleic acid sequence in the subject's genome encoding an epitope, and (ii) an amino acid or amino acid sequence not encoded by a nucleic acid sequence immediately upstream or downstream of a nucleic acid sequence in the subject's genome encoding an epitope, and / or (iii) a linker, the polypeptide is cleaved at a higher rate compared to a corresponding polypeptide of the same length and epitope and an amino acid or amino acid sequence encoded by a nucleic acid sequence immediately upstream or downstream of a nucleic acid sequence in the subject's genome encoding an epitope.

[0152] In some embodiments, when the polypeptide is processed by APC, the polypeptide is cleaved at the linker region.In some embodiments, when the polypeptide further comprises an amino acid or amino acid sequence that does not include at least one additional amino acid encoded by the nucleic acid sequence immediately upstream of the nucleic acid sequence in the genome of the subject that encodes the epitope, and a linker, the polypeptide is cleaved at a higher rate at the linker region compared to the corresponding polypeptide with the same length and epitope, and the amino acid or amino acid sequence is encoded by the nucleic acid sequence immediately upstream of the nucleic acid sequence in the genome of the subject that encodes the epitope.In some embodiments, when the polypeptide further comprises an amino acid or amino acid sequence that does not include at least one additional amino acid encoded by the nucleic acid sequence immediately downstream of the nucleic acid sequence in the genome of the subject that encodes the epitope, and a linker, the polypeptide is cleaved at a higher rate at the linker region compared to the corresponding polypeptide with the same length and epitope, and the amino acid or amino acid sequence is encoded by the nucleic acid sequence immediately downstream of the nucleic acid sequence in the genome of the subject that encodes the epitope.In some embodiments, when the amino acid or amino acid sequence is not of the peptide sequence of the protein expressed in the subject, the polypeptide is cleaved at a higher rate at the linker region.

[0153] In some embodiments, when the polypeptide is processed by APC, epitope presentation by APC is enhanced.In some embodiments, when the polypeptide comprising epitope further comprises an amino acid or amino acid sequence and / or linker that does not include at least one additional amino acid encoded by the nucleic acid sequence immediately upstream of the nucleic acid sequence in the genome of the subject that encodes the epitope, epitope presentation by APC is enhanced compared to the corresponding polypeptide with the same length and epitope, and the amino acid or amino acid sequence is encoded by the nucleic acid sequence immediately upstream of the nucleic acid sequence in the genome of the subject that encodes the epitope.In some embodiments, when the polypeptide comprising epitope further comprises an amino acid or amino acid sequence and / or linker that does not include at least one additional amino acid encoded by the nucleic acid sequence immediately downstream of the nucleic acid sequence in the genome of the subject that encodes the epitope, epitope presentation by APC is enhanced compared to the corresponding polypeptide with the same length and epitope, and the amino acid or amino acid sequence is encoded by the nucleic acid sequence immediately downstream of the nucleic acid sequence in the genome of the subject that encodes the epitope. In some embodiments, when amino acid or amino acid sequence is not the peptide sequence of the protein expressed in subject, epitope presentation by APC is enhanced.In some embodiments, when polypeptide comprises linker, epitope presentation by APC is enhanced compared to the polypeptide that comprises the same epitope but does not comprise linker.In some embodiments, when polypeptide comprises protease cleavable linker, epitope presentation by APC is enhanced compared to the polypeptide that comprises the same epitope but does not comprise protease cleavable linker.

[0154] In some embodiments, when the polypeptide further comprises an amino acid or amino acid sequence that does not include at least one additional amino acid encoded by a nucleic acid sequence immediately upstream of the nucleic acid sequence in the genome of the subject that encodes the epitope, the epitope presentation by the APC is enhanced compared to the cleavage of the corresponding polypeptide of the same length that does not include the epitope sequence and the amino acid or amino acid sequence that is contiguous with the epitope sequence encoded by the nucleic acid sequence and does not include a linker. In some embodiments, when the polypeptide further comprises an amino acid or amino acid sequence that does not include at least one additional amino acid encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the genome of the subject that encodes the epitope, the epitope presentation by the APC is enhanced compared to the cleavage of the corresponding polypeptide of the same length that does not include the epitope sequence and the amino acid or amino acid sequence that is contiguous with the epitope sequence encoded by the nucleic acid sequence and does not include a linker.

[0155] In some embodiments, when a polypeptide comprises (i) an amino acid or amino acid sequence encoded by a nucleic acid sequence immediately upstream or downstream of a nucleic acid sequence in the subject's genome encoding the epitope, and (ii) an amino acid or amino acid sequence not encoded by a nucleic acid sequence immediately upstream or downstream of a nucleic acid sequence in the subject's genome encoding the epitope, and / or (iii) a linker, epitope presentation by APCs is enhanced compared to a corresponding polypeptide of the same length and epitope and amino acid or amino acid sequence encoded by a nucleic acid sequence immediately upstream or downstream of a nucleic acid sequence in the subject's genome encoding the epitope.

[0156] In some embodiments, when the polypeptide is processed by APC, immunogenicity is enhanced. In some embodiments, when the polypeptide comprising the epitope further comprises an amino acid or amino acid sequence and / or linker that does not include at least one additional amino acid encoded by the nucleic acid sequence immediately upstream of the nucleic acid sequence in the genome of the subject that encodes the epitope, immunogenicity is enhanced compared to the corresponding polypeptide with the same length and epitope, and the amino acid or amino acid sequence is encoded by the nucleic acid sequence immediately upstream of the nucleic acid sequence in the genome of the subject that encodes the epitope. In some embodiments, when the polypeptide comprising the epitope further comprises an amino acid or amino acid sequence and / or linker that does not include at least one additional amino acid encoded by the nucleic acid sequence immediately downstream of the nucleic acid sequence in the genome of the subject that encodes the epitope, immunogenicity is enhanced compared to the corresponding polypeptide with the same length and epitope, and the amino acid or amino acid sequence is encoded by the nucleic acid sequence immediately downstream of the nucleic acid sequence in the genome of the subject that encodes the epitope. In some embodiments, immunogenicity is enhanced when the amino acid or amino acid sequence is not of the peptide sequence of the protein expressed in the subject. In some embodiments, when a polypeptide includes a linker, the immunogenicity is enhanced compared to a polypeptide that includes the same epitope but does not include a linker. In some embodiments, when a polypeptide includes a linker that is cleavable by a protease, the immunogenicity is enhanced compared to a polypeptide that includes the same epitope but does not include a linker that is cleavable by a protease.

[0157] In some embodiments, when the polypeptide further comprises an amino acid or amino acid sequence that does not include at least one additional amino acid encoded by a nucleic acid sequence immediately upstream of the nucleic acid sequence in the genome of the subject that encodes the epitope, the immunogenicity is enhanced compared to the cleavage of the corresponding polypeptide of the same length that does not include the epitope sequence and the amino acid or amino acid sequence that is contiguous with the epitope sequence encoded by the nucleic acid sequence and does not include a linker. In some embodiments, when the polypeptide further comprises an amino acid or amino acid sequence that does not include at least one additional amino acid encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the genome of the subject that encodes the epitope, the immunogenicity is enhanced compared to the cleavage of the corresponding polypeptide of the same length that does not include the epitope sequence and the amino acid or amino acid sequence that is contiguous with the epitope sequence encoded by the nucleic acid sequence and does not include a linker.

[0158] In some embodiments, when a polypeptide comprises (i) an amino acid or amino acid sequence encoded by a nucleic acid sequence immediately upstream or downstream of a nucleic acid sequence in the subject's genome encoding an epitope, and (ii) an amino acid or amino acid sequence not encoded by a nucleic acid sequence immediately upstream or downstream of a nucleic acid sequence in the subject's genome encoding an epitope, and / or (iii) a linker, the immunogenicity is enhanced compared to a corresponding polypeptide of the same length and epitope and an amino acid or amino acid sequence encoded by a nucleic acid sequence immediately upstream or downstream of a nucleic acid sequence in the subject's genome encoding the epitope.

[0159] In some embodiments, when the polypeptide is processed by APC, the anti-tumor activity is enhanced. In some embodiments, when the polypeptide comprising the epitope further comprises an amino acid or amino acid sequence and / or linker that does not include at least one additional amino acid encoded by the nucleic acid sequence immediately upstream of the nucleic acid sequence in the genome of the subject that encodes the epitope, the anti-tumor activity by APC is enhanced compared to the corresponding polypeptide with the same length and epitope and the amino acid or amino acid sequence is encoded by the nucleic acid sequence immediately upstream of the nucleic acid sequence in the genome of the subject that encodes the epitope. In some embodiments, when the polypeptide comprising the epitope further comprises an amino acid or amino acid sequence and / or linker that does not include at least one additional amino acid encoded by the nucleic acid sequence immediately downstream of the nucleic acid sequence in the genome of the subject that encodes the epitope, the anti-tumor activity is enhanced compared to the corresponding polypeptide with the same length and epitope and the amino acid or amino acid sequence is encoded by the nucleic acid sequence immediately downstream of the nucleic acid sequence in the genome of the subject that encodes the epitope. In some embodiments, when the amino acid or amino acid sequence is not of the peptide sequence of the protein expressed in the subject, the anti-tumor activity is enhanced. In some embodiments, when a polypeptide includes a linker, the antitumor activity is enhanced compared to a polypeptide that includes the same epitope but does not include a linker. In some embodiments, when a polypeptide includes a linker that is cleavable by a protease, the antitumor activity is enhanced compared to a polypeptide that includes the same epitope but does not include a linker that is cleavable by a protease.

[0160] In some embodiments, when the polypeptide further comprises an amino acid or amino acid sequence that does not include at least one additional amino acid encoded by a nucleic acid sequence immediately upstream of the nucleic acid sequence in the genome of the subject that encodes the epitope, the anti-tumor activity is enhanced compared to the cleavage of the corresponding polypeptide of the same length that does not include the epitope sequence and the amino acid or amino acid sequence that is contiguous with the epitope sequence encoded by the nucleic acid sequence and does not include a linker. In some embodiments, when the polypeptide further comprises an amino acid or amino acid sequence that does not include at least one additional amino acid encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the genome of the subject that encodes the epitope, the anti-tumor activity is enhanced compared to the cleavage of the corresponding polypeptide of the same length that does not include the epitope sequence and the amino acid or amino acid sequence that is contiguous with the epitope sequence encoded by the nucleic acid sequence and does not include a linker.

[0161] In some embodiments, when a polypeptide comprises (i) an amino acid or amino acid sequence encoded by a nucleic acid sequence immediately upstream or downstream of a nucleic acid sequence in the subject's genome encoding an epitope, and (ii) an amino acid or amino acid sequence not encoded by a nucleic acid sequence immediately upstream or downstream of a nucleic acid sequence in the subject's genome encoding an epitope, and / or (iii) a linker, the anti-tumor activity is enhanced compared to a corresponding polypeptide of the same length and epitope and an amino acid or amino acid sequence encoded by a nucleic acid sequence immediately upstream or downstream of a nucleic acid sequence in the subject's genome encoding the epitope.

[0162] In some embodiments, when the polypeptide is processed by the APC, the epitope is presented by the APC to immune cells. In some embodiments, when the polypeptide is processed by the APC, the epitope is presented by the APC preferentially or specifically to immune cells. In some embodiments, when the polypeptide is processed by the APC, the epitope is presented by the APC to phagocytes. In some embodiments, when the polypeptide is processed by the APC, the epitope is presented by the APC preferentially or specifically to phagocytes. In some embodiments, when the polypeptide is processed by the APC, the epitope is presented by the APC to dendritic cells, macrophages, mast cells, neutrophils, or monocytes. In some embodiments, when the polypeptide is processed by the APC, the epitope is presented by the APC preferentially or specifically to dendritic cells, macrophages, mast cells, neutrophils, or monocytes.

[0163] In some embodiments, the polypeptide comprises an amino acid sequence selected from the group consisting of poly-Lys (poly-K) and poly-Arg (poly-R). In preferred embodiments, the polypeptide comprises a poly-K sequence. In some embodiments, the polypeptide comprises a sequence selected from the group consisting of poly-K-AA-AA and poly-R-AA-AA, where each AA is an amino acid or an analog or derivative thereof. In preferred embodiments, the polypeptide comprises poly-K-AA-AA. In some embodiments, the poly-K comprises poly-L-Lys. In some embodiments, the poly-K comprises at least two consecutive lysine residues. In some embodiments, the poly-K comprises at least three consecutive lysine residues, e.g., Lys-Lys-Lys. In preferred embodiments, the poly-K comprises at least four consecutive lysine residues, e.g., Lys-Lys-Lys-Lys, also known as K4. In some embodiments, the poly-K comprises at least five, at least six, at least seven, at least eight, at least nine, or at least ten consecutive lysine residues. In some embodiments, the poly-R comprises poly-L-Arg. In some embodiments, polyR comprises at least two consecutive arginine residues. In some embodiments, polyR comprises at least three consecutive arginine residues, e.g., Arg-Arg-Arg. In some embodiments, polyR comprises at least four, at least five, at least six, or at least seven consecutive arginine residues. In some embodiments, polyR comprises at least eight consecutive arginine residues, e.g., Arg-Arg-Arg-Arg-Arg-Arg-Arg, also known as R8. In some embodiments, polyR comprises at least five, at least six, at least seven, at least eight, at least nine, or at least ten consecutive arginine residues. In some embodiments, the lysine units in polyK and / or the arginine units in polyR may each have the (L) stereochemical configuration, the (D) stereochemical configuration, or any mixture of the (L) and (D) stereochemical configurations.

[0164] In some embodiments, the polypeptide comprises a linker selected from the group consisting of disulfide, p-aminobenzyloxycarbonyl (PABC), and AA-AA-PABC, where AA is an amino acid or an analog or derivative thereof. In some embodiments, AA-AA-PABC is selected from the group consisting of alanine-lysine-PABC (Ala-Lys-PABC), valine-citrulline-PABC (Val-Cit-PABC), and phenylalanine-lysine-PABC (Phe-Lys-PABC). In some embodiments, AA-AA-PABC is Ala-Lys-PABC. In some embodiments, AA-AA-PABC is Val-Cit-PABC. In some embodiments, AA-AA-PABC is Phe-Lys-PABC. In some embodiments, the valine and citrulline units in Val-Cit-PABC each have an (L) stereochemical configuration. In some embodiments, the phenylalanine and lysine units in Phe-Lys-PABC each have an (L) stereochemical configuration. In some embodiments, the valine and citrulline units in Val-Cit-PABC each have a (D) stereochemical configuration. In some embodiments, the phenylalanine and lysine units in Phe-Lys-PABC each have a (D) stereochemical configuration. In some embodiments, the valine and citrulline units in Val-Cit-PABC each have a mixture of (L) and (D) stereochemical configurations. In some embodiments, the phenylalanine and lysine units in Phe-Lys-PABC each have a mixture of (L) and (D) stereochemical configurations.

[0165] In some embodiments, the polypeptide comprises a linker having the following structure: [ka]

[0166] In some embodiments, the polypeptide is [ka] wherein R 1 and R 2 are independently H or (C 1 ~C 6 ) alkyl; j is 1 or 2; G 1 is H or COOH; i is 1, 2, 3, 4, or 5.

[0167] In some embodiments, A r and / or A s is of formula (III) or (IV), wherein R 1 and R 2 are independently H or (C 1 ~C 6 ) alkyl; j is 1 or 2; G 1 is H or COOH; i is 1, 2, 3, 4, or 5.

[0168] In some embodiments, the polypeptide comprises a linker that is Formula (III) or Formula (IV).

[0169] The disulfide linker of formula (IV) can be prepared by the methods described in Zhang, Donglu, et al., ACS Med. Chem. Lett. 2016, 7, 988-993; and Pillow, Thomas H., et al., Chem. Sci., 2017, 8, 366-370. PABC-containing peptides can be synthesized according to Laurent Ducry (ed.), Antibody-Drug Conjugates, Methods in Molecular Biology, vol. 1045, DOI 10.1007 / 978-1-62703-541-5_5, Springer Science+Business Media, LLC 2013. In some embodiments, any resin made for solid-phase peptide synthesis can be used. Antigen Processing Pathways

[0170] The polypeptides described herein can be processed by different pathways to release epitopes for epitope presentation. There are two important processing events in the antigen processing and presentation pathway to generate optimal peptide antigens. Cytoplasmic proteins are mainly processed by proteasomes. Short peptides are then transported into the endoplasmic reticulum (ER) by antigen processing-associated transporter (TAP) for subsequent assembly with MHC class I molecules. Exogenous proteins are mainly presented by MHC class II molecules. Antigens are internalized by several pathways, including phagocytosis, macropinocytosis, and endocytosis, and are finally transported to mature or late endosomal compartments, where they are processed and loaded onto MHC class II molecules. Cytoplasmic / nuclear antigens can also be transported into the endosomal network by autophagy for subsequent processing and presentation with MHC class II molecules.

[0171] Initial peptide proteolysis occurs in the cytosol of cells, where large protein fragments are degraded into small peptides by the proteasome or immunoproteasome. This processing event is often responsible for generating the final C-terminal residues of peptides that bind to class I MHC. The proteasome is a large proteolytic complex that contains multiple subunits, including two subunits, large multifunctional protease (LMP)2 and LMP7. Bound proteins for degradation are targeted to the proteasome by covalent linkage with ubiquitin. LMP2 and LMP7 induce the proteolytic complex to produce peptides that bind to class I MHC I. The peptides generated in the cytosol are then transported to the ER by TAP. Because TAP preferentially transports peptides of 11-14 amino acids, the peptides are often too long for stable class I MHC binding and require further processing upon entry into the ER. This processing involves trimming of the N-terminal region of antigenic peptides by endoplasmic reticulum aminopeptidases (ERAP) 1 and ERAP 2. This process creates a pool of peptides with high affinity for binding to class I MHC.

[0172] In the normal cellular environment, classical class II MHC molecules are expressed only on professional APCs such as dendritic cells (DCs) or macrophages. Exogenous or extracellular antigens internalized by phagocytosis, endocytosis, or pinocytosis are presented primarily on class II MHC to CD4+ T cells. However, a small subset of cytosolic antigens is also expressed on class II MHC as a result of autophagy. Briefly, antigens taken up by endocytosis are processed in a vesicular pathway consisting of compartments of increasingly more acidic and proteolytic activity, classically described as early endosomes (pH 6.0–pH 6.5), late endosomes or endolysosomes (pH 5.0–pH 6.0), and lysosomes (pH 4.5–pH 5.0). Antigens internalized by phagocytosis follow a similar pathway, terminating in phagolysosomes formed by fusion of phagosomes with lysosomes. Lysosomes and phagolysosomes (pH 4.0-pH 4.5) contain several acidic pH-optimal proteases commonly referred to as cathepsins. In highly degradative cells such as macrophages, sequential cleavages by these enzymes result in very short peptides and free amino acids that are translocated into the cytosol to recruit tRNAs for new protein synthesis. In APCs with low proteolytic activity, large intermediates form the predominant source of peptides for class II MHC binding, and these peptides usually consist of 13-18 amino acids.

[0173] Both class I MHC and class II MHC can access peptides processed from endogenous antigens and peptides processed from exogenous antigens.For example, class II MHC binds to peptides extracted from endogenous membrane proteins that are degraded in lysosomes.Similarly, class I MHC can bind to peptides extracted from exogenous proteins that are internalized by endocytosis or phagocytosis, a phenomenon called cross-presentation.Certain subsets of DCs are particularly skilled in mediating this process, which is crucial for the initiation of primary responses by naive CD8+ T cells.

[0174] In one aspect, provided herein is a method of cleaving a polypeptide, comprising contacting a polypeptide as described herein with an APC. In some embodiments, the method can be performed in vivo. In some embodiments, the method can be performed in vitro.

[0175] In some embodiments, the polypeptide is ubiquitinated. In some embodiments, the polypeptide is ubiquitinated before cleavage. In some embodiments, the polypeptide is ubiquitinated before processing by the proteasome and / or the immunoproteasome. In some embodiments, the polypeptide is ubiquitinated at a lysine residue. In some embodiments, the polypeptide is ubiquitinated at a lysine residue that is not on the epitope sequence. In some embodiments, the polypeptide is ubiquitinated at a lysine residue of polyK. In some embodiments, the polypeptide is ubiquitinated at the first lysine of polyK. In some embodiments, the polypeptide is ubiquitinated at the second lysine of polyK. In some embodiments, the polypeptide is ubiquitinated at the third lysine of polyK. In some embodiments, the polypeptide is ubiquitinated at the fourth lysine of polyK. In some embodiments, the polypeptide is ubiquitinated at the fifth, sixth, seventh, eighth, ninth, or tenth lysine of polyK. In some embodiments, the polypeptide is ubiquitinated at at least one lysine residue. In some embodiments, the polypeptide is ubiquitinated at more than one lysine residue. In some embodiments, the polypeptide is ubiquitinated at more than one lysine residue of polyK. In some embodiments, the polypeptide is ubiquitinated at each lysine residue. In some embodiments, the polypeptide is ubiquitinated at each lysine residue of polyK. In some embodiments, the polypeptide is ubiquitinated at two lysine residues of polyK. In some embodiments, the polypeptide is ubiquitinated at three lysine residues of polyK. In some embodiments, the polypeptide is ubiquitinated at four lysine residues of polyK. In some embodiments, the polypeptide is ubiquitinated at five, six, seven, eight, nine, or ten lysine residues of polyK. In some embodiments, the polypeptide is sequentially ubiquitinated at each lysine residue of polyK. In some embodiments, the polypeptide is not sequentially ubiquitinated at each lysine residue of polyK.

[0176] In some embodiments, the polypeptide is ubiquitinated at the lysine residue of Ala-Lys-PABC. In some embodiments, the polypeptide is ubiquitinated at the lysine residue of Phe-Lys-PABC. In some embodiments, the polypeptide comprises poly-K and AA-AA-PABC, where each AA is an amino acid or an analog or derivative thereof. In some embodiments, the polypeptide is ubiquitinated at at least one lysine residue of poly-K and AA-AA-PABC. In some embodiments, the polypeptide is ubiquitinated at one or more lysine residues of poly-K and AA-AA-PABC. In some embodiments, the polypeptide is ubiquitinated at one or more lysine residues of poly-K and Ala-Lys-PABC. In some embodiments, the polypeptide is ubiquitinated at one or more lysine residues of poly-K and Phe-Lys-PABC.

[0177] In some embodiments, the polypeptide is internalized by the APC. In some embodiments, the polypeptide is internalized by the APC via endocytosis. In some embodiments, the polypeptide is internalized by the APC via phagocytosis. In some embodiments, the polypeptide is internalized by the APC via pinocytosis. In some embodiments, the polypeptide is cleaved in the cytoplasm. In some embodiments, the polypeptide is cleaved in the endosome. In some embodiments, the polypeptide is cleaved in the endolysosome. In some embodiments, the polypeptide is cleaved in the lysosome. In some embodiments, the polypeptide is cleaved in the ER. In some embodiments, the polypeptide is cleaved by an aminopeptidase. In some embodiments, the aminopeptidase is an insulin-regulated aminopeptidase (IRAP). In some embodiments, the aminopeptidase is an endoplasmic reticulum aminopeptidase (ERAP). In some embodiments, the polypeptide is processed by a trypsin-like domain of the proteasome and / or the immunoproteasome. In some embodiments, the trypsin-like domain comprises trypsin-like activity. In some embodiments, the trypsin-like domain comprises a chymotrypsin-like activity. In some embodiments, the trypsin-like activity comprises a peptidylglutamyl-peptide hydrolase (PGPH) activity. In some embodiments, the polypeptide is cleaved by a protease. In some embodiments, the protease is a trypsin-like protease. In some embodiments, the protease is a chymotrypsin-like protease. In some embodiments, the protease is a peptidylglutamyl-peptide hydrolase (PGPH). In some embodiments, the protease is selected from the group consisting of an aspartic peptide lyase, an aspartic acid protease, a cysteine ​​protease, a glutamic acid protease, a metalloprotease, a serine protease, and a threonine protease. In a preferred embodiment, the protease is a cysteine ​​protease.In some embodiments, the cysteine ​​protease is selected from the group consisting of calpain, caspase, cathepsin B, cathepsin C, cathepsin F, cathepsin H, cathepsin K, cathepsin L1, cathepsin L2, cathepsin O, cathepsin S, cathepsin W, and cathepsin Z. In some embodiments, the protease is cathepsin B. In some embodiments, the protease is cathepsin C. In some embodiments, the protease is cathepsin F. In some embodiments, the protease is cathepsin Z.

[0178] In some embodiments, the polypeptide is cleaved at a lysine residue. In some embodiments, the polypeptide is cleaved at a lysine residue of poly-K. In some embodiments, the polypeptide is cleaved at the first lysine residue of poly-K. In some embodiments, the polypeptide is cleaved at the second lysine residue of poly-K. In some embodiments, the polypeptide is cleaved at the third lysine residue of poly-K. In some embodiments, the polypeptide is cleaved at the fourth lysine residue of poly-K. In some embodiments, the polypeptide is cleaved at the fifth lysine residue, the sixth lysine residue, the seventh lysine residue, the eighth lysine residue, the ninth lysine residue, or the tenth lysine residue of poly-K. In some embodiments, the polypeptide is cleaved at more than one lysine residue of poly-K. In some embodiments, the polypeptide is cleaved at each lysine residue of poly-K sequentially. In some embodiments, the polypeptide is not cleaved at each lysine residue of poly-K sequentially.

[0179] In some embodiments, the polypeptide is cleaved at AA-AA-PABC, where each AA is an amino acid or an analog or derivative thereof. In some embodiments, the polypeptide is cleaved at Ala-Lys-PABC. In some embodiments, the polypeptide is cleaved at the lysine residue of Ala-Lys-PABC. In some embodiments, the polypeptide is cleaved at Phe-Lys-PABC. In some embodiments, the polypeptide is cleaved at the lysine residue of Phe-Lys-PABC. In some embodiments, the polypeptide is cleaved at Val-Cit-PABC. In some embodiments, the polypeptide is cleaved at the citrulline (Cit) residue of Val-Cit-PABC. In some embodiments, the epitope is released when the polypeptide is cleaved.

[0180] One major drawback that limits the application of peptide-based drugs to systemic therapy is the proteolysis of peptides. Peptides administered by injection route reach the bloodstream. The bloodstream contains proteases that function in hemostasis, fibrinolysis, and tissue conversion, i.e., important processes in the case of injury. It is therefore important to stabilize the peptide against proteases present in blood, serum, or plasma. In one aspect, the polypeptides described herein are stable in plasma, blood, and / or serum. In some embodiments, the polypeptide is not cleaved prior to internalization by APCs in the subject. In some embodiments, the polypeptide is not cleaved prior to processing by APCs in the subject. In some embodiments, the polypeptide is not cleaved prior to internalization by APCs in the subject's blood. In some embodiments, the polypeptide is not cleaved prior to processing by APCs in the subject's blood. In some embodiments, the polypeptide is not cleaved by proteases in the blood. In some embodiments, the polypeptide is not cleaved by plasmin. In some embodiments, the polypeptide is not cleaved by plasma kallikrein. In some embodiments, the polypeptide is not cleaved by tissue kallikrein. In some embodiments, the polypeptide is not cleaved by thrombin. In some embodiments, the polypeptide is not cleaved by a clotting factor. In some embodiments, the polypeptide is not cleaved by clotting factor XII. In some embodiments, the polypeptide is stable in human plasma. In some embodiments, the polypeptide is stable in human blood. In some embodiments, the polypeptide is stable in human serum.

[0181] In some embodiments, the polypeptide has a half-life in human plasma of from 1 hour to 5 days. In some embodiments, the polypeptide has a half-life of about 1 hour to about 120 hours. In some embodiments, the polypeptide has a half-life of about 1 hour to about 5 hours, about 1 hour to about 10 hours, about 1 hour to about 12 hours, about 1 hour to about 24 hours, about 1 hour to about 36 hours, about 1 hour to about 48 hours, about 1 hour to about 60 hours, about 1 hour to about 72 hours, about 1 hour to about 84 hours, about 1 hour to about 96 hours, about 1 hour to about 120 hours, about 5 hours to about 10 hours, about 5 hours to about 12 hours, about 5 hours to about 24 hours, about 5 hours to about 36 hours, about 5 hours to about 48 hours, about 5 hours to about 60 hours. 5 hours to 72 hours, 5 hours to 84 hours, 5 hours to 96 hours, 5 hours to 120 hours, 10 hours to 12 hours, 10 hours to 24 hours, 10 hours to 36 hours, 10 hours to 48 hours, 10 hours to 60 hours, 10 hours to 72 hours, 10 hours to 84 hours, 10 hours to 96 hours, 10 hours to 120 hours, 12 hours to 24 hours, 12 hours to 36 hours, 12 hours to 48 hours, 12 hours to 60 hours between about 12 hours and about 72 hours, between about 12 hours and about 84 hours, between about 12 hours and about 96 hours, between about 12 hours and about 120 hours, between about 24 hours and about 36 hours, between about 24 hours and about 48 hours, between about 24 hours and about 60 hours, between about 24 hours and about 72 hours, between about 24 hours and about 84 hours, between about 24 hours and about 96 hours, between about 24 hours and about 120 hours, between about 36 hours and about 48 hours, between about 36 hours and about 60 hours, between about 36 hours and about 72 hours, between about 36 hours and about 84 hours, between about 36 hours and about 96 hours, between about 36 hours and about The polypeptide has a half-life of about 120 hours, about 48 hours to about 60 hours, about 48 hours to about 72 hours, about 48 hours to about 84 hours, about 48 hours to about 96 hours, about 48 hours to about 120 hours, about 60 hours to about 72 hours, about 60 hours to about 84 hours, about 60 hours to about 96 hours, about 60 hours to about 120 hours, about 72 hours to about 84 hours, about 72 hours to about 96 hours, about 72 hours to about 120 hours, about 84 hours to about 96 hours, about 84 hours to about 120 hours, or about 96 hours to about 120 hours. In some embodiments, the polypeptide has a half-life of about 1 hour, about 5 hours, about 10 hours, about 12 hours, about 24 hours, about 36 hours, about 48 hours, about 60 hours, about 72 hours, about 84 hours, about 96 hours, or about 120 hours.In some embodiments, the polypeptide has a half-life of at least about 1 hour, about 5 hours, about 10 hours, about 12 hours, about 24 hours, about 36 hours, about 48 hours, about 60 hours, about 72 hours, about 84 hours, or about 96 hours. In some embodiments, the polypeptide has a half-life of at most about 5 hours, about 10 hours, about 12 hours, about 24 hours, about 36 hours, about 48 hours, about 60 hours, about 72 hours, about 84 hours, about 96 hours, or about 120 hours. 3. Neoantigens and their uses

[0182] One of the crucial hurdles for the development of curative and tumor-specific immunotherapy is to identify and select highly specific and restricted tumor antigens to avoid autoimmunity. Tumor neo-antigens, which arise as a result of genetic alterations in malignant cells (e.g., inversions, translocations, deletions, missense mutations, splice site mutations, etc.), represent the most tumor-specific class of antigens. Neo-antigens have been used in cancer vaccines or immunogenic compositions only rarely due to the technical difficulties involved in identifying them, selecting optimized antigens, and generating neo-antigens for use in vaccines or immunogenic compositions. These problems can be addressed by identifying mutations in the neoplasia / tumor that are present at the DNA level in the tumor but not in the corresponding germline samples from a high percentage of subjects with cancer; analyzing the identified mutations with one or more peptide-MHC binding prediction algorithms to generate multiple neoantigenic T cell epitopes that are expressed in the neoplasia / tumor and bind to a high percentage of patient HLA alleles; and synthesizing multiple neoantigenic peptides selected from the set of all neoantigenic peptides and predicted binding peptides for use in a cancer vaccine or immunogenic composition suitable for treating a high percentage of subjects with cancer.

[0183] For example, translating peptide sequencing information into therapeutic vaccines may involve predicting which mutated peptides can bind to a high percentage of individuals' HLA molecules. Efficient selection of which specific mutations to utilize as immunogens requires the ability to predict which mutated peptides will efficiently bind to a high percentage of patients' HLA alleles. Recently, neural network-based learning approaches using validated binding and non-binding peptides have advanced the accuracy of prediction algorithms for major HLA-A and HLA-B alleles. However, even when advanced neural network-based algorithms are used to encode HLA-peptide binding rules, several factors limit the predictive power of peptides presented on HLA alleles.

[0184] Another example of translating peptide sequencing information into therapeutic vaccines may include formulating drugs as long peptide multi-epitope vaccines. Targeting as many mutated epitopes as practically possible takes advantage of the enormous capacity of the immune system, prevents the chance of immunological escape by down-modulation of immune-targeted gene products, and compensates for the known inaccuracies of epitope prediction techniques. Synthetic peptides provide a useful means for efficient preparation of large numbers of immunogens and for rapid translation of identification of mutated epitopes into effective vaccines. Peptides can be easily chemically synthesized and purified using reagents free of bacterial or animal contaminants. Their small size allows for a clear focus on the mutated regions of the protein and also reduces irrelevant antigenic competition from other components (non-mutated proteins or viral vector antigens).

[0185] Yet another example of the translation of peptide sequencing information into therapeutic vaccines may include combination with a strong vaccine adjuvant. Effective vaccines may require a strong adjuvant to initiate an immune response. For example, poly-ICLC, an agonist of TLR3 and the RNA helicase-domains of MDA5 and RIG3, has shown several desirable properties as a vaccine adjuvant. These properties include induction of local and systemic activation of immune cells in vivo, production of stimulatory chemokines and cytokines, and stimulation of antigen presentation by dendritic cells (DCs). In addition, poly-ICLC has been shown to induce long-lasting CD4+ expression in humans. + and CD8 + A long-lasting CD4 response can be induced in humans. + and CD8 + Importantly, striking similarities in upregulation of transcriptional and signal transduction pathways have been observed in subjects vaccinated with poly-ICLC and volunteers who received a highly effective replication-competent yellow fever vaccine. Furthermore, in a recent phase I study, CD4 + T cells and CD8 + Induction of T cell as well as antibody responses to the peptides has been demonstrated. At the same time, poly-ICLC has been extensively tested in more than 25 clinical trials to date, showing a relatively benign toxicity profile. peptide

[0186] In some aspects, the disclosure provides isolated peptides comprising tumor-specific mutations. These peptides and polypeptides are referred to herein as "neo-antigenic peptides" or "neo-antigenic polypeptides." The term "peptide" is used interchangeably herein with "mutated peptide," "neo-antigenic peptide," and "neo-antigenic peptide" to refer to a stretch of one residue, typically an L-amino acid, connected by a peptide bond between another residue, typically an L-amino acid, and typically the α-amino and carboxyl groups of the adjacent amino acid. Similarly, the term "polypeptide" is used interchangeably herein with "mutated polypeptide," "neo-antigenic polypeptide," and "neo-antigenic polypeptide" to refer to a stretch of one residue, e.g., an L-amino acid, connected by a peptide bond between another residue, e.g., an L-amino acid, and typically the α-amino and carboxyl groups of the adjacent amino acid. The polypeptides or peptides may be of various lengths, may be in neutral (uncharged) or salt form, and may be free or contain modifications such as glycosylation, side chain oxidation, or phosphorylation, provided that the modifications do not impair the biological activity of the polypeptides described herein.

[0187] In some embodiments, genome or exome sequencing methods are used to identify tumor-specific mutations. Any suitable sequencing method, such as next-generation sequencing (NGS) technology, can be used according to the present disclosure. In the future, third generation sequencing methods can be substituted for NGS technology to speed up the sequencing step of the method. For clarity, the term "next-generation sequencing" or "NGS" refers to all novel high-throughput sequencing technologies in the context of the present disclosure, as opposed to the "conventional" sequencing methodology known as Sanger chemistry, in which the nucleic acid template is read randomly in parallel along the entire genome by splitting the entire genome into small pieces. Such NGS technologies (also known as massively parallel sequencing technologies) can deliver nucleic acid sequence information of the whole genome, exome, transcriptome (all transcribed sequences of the genome), or methylome (all methylated sequences of the genome) in a very short time, e.g., within 1-2 weeks, e.g., within 1-7 days or less than 24 hours, and in principle also enable single-cell sequencing approaches. Many NGS platforms available commercially or referenced in the literature can be used in connection with the present disclosure, for example those described in detail in WO2012 / 159643.

[0188] In certain embodiments, the polypeptides described herein may comprise, but are not limited to, about 5 amino acids, about 6 amino acids, about 7 amino acids, about 8 amino acids, about 9 amino acids, about 10 amino acids, about 11 amino acids, about 12 amino acids, about 13 amino acids, about 14 amino acids, about 15 amino acids, about 16 amino acids, about 17 amino acids, about 18 amino acids, about 19 amino acids, about 20 amino acids, about 21 amino acids, about 22 amino acids, about 23 amino acids, about 24 amino acids, about 25 amino acids, about 26 amino acids, about 27 amino acids, about 28 amino acids, about 29 amino acids, about 30 amino acids, about 31 amino acids, about 32 amino acids, about 33 amino acids, about 34 amino acids, about 35 amino acids, about 36 amino acids, about 37 amino acids, about 38 amino acids, about 39 amino acids, about 40 amino acids, about 41 amino acids, about 42 amino acids, about 43 amino acids, about 44 amino acids, about 45 amino acids, about 46 amino acids, about 47 amino acids, about 48 amino acids, about 49 amino acids, about 50 amino acids, about 51 amino acids, about 52 amino acids, about 53 amino acids, about 54 amino acids, about 55 amino acids, about 56 amino acids, about 57 amino acids, about 58 amino acids, about 59 amino acids, about 60 amino acids, about 61 amino acids, about 62 amino acids, about 63 amino acids, about 64 amino acids, about 65 amino acids, about 66 amino acids, about 67 amino acids, about 68 amino acids, about 69 amino acids, about 70 amino acids, about 71 amino acids, about 72 amino acids, about 73 amino acids, about 74 amino acids, about 75 amino acids, about The amino acid sequence may comprise about 45 amino acids, about 46 amino acids, about 47 amino acids, about 48 amino acids, about 49 amino acids, about 50 amino acids, about 60 amino acids, about 70 amino acids, about 80 amino acids, about 90 amino acids, about 100 amino acids, about 110 amino acids, about 120 amino acids, about 150 amino acids, about 200 amino acids, about 300 amino acids, about 350 amino acids, about 400 amino acids, about 450 amino acids, about 500 amino acids, about 600 amino acids, about 700 amino acids, about 800 amino acids, about 900 amino acids, about 1,000 amino acids, about 1,500 amino acids, about 2,000 amino acids, about 2,500 amino acids, about 3,000 amino acids, about 4,000 amino acids, about 5,000 amino acids, about 7,500 amino acids, about 10,000 amino acids or more amino acid residues, and any range derivable therein. In certain embodiments, the neo-antigenic peptide molecule is equal to or less than 100 amino acids.

[0189] In some embodiments, the polypeptides may be from about 8 to about 50 amino acid residues in length, or from about 8 to about 30 amino acid residues, from about 8 to about 20 amino acid residues, from about 8 to about 18 amino acid residues, from about 8 to about 15 amino acid residues, or from about 8 to about 12 amino acid residues in length. In some embodiments, the peptides may be from about 8 to about 500 amino acid residues in length, or from about 8 to about 450 amino acid residues, from about 8 to about 400 amino acid residues, from about 8 to about 350 amino acid residues, from about 8 to about 300 amino acid residues, from about 8 to about 250 amino acid residues, from about 8 to about 200 amino acid residues, from about 8 to about 150 amino acid residues, from about 8 to about 100 amino acid residues, from about 8 to about 50 amino acid residues, or from about 8 to about 30 amino acid residues in length.

[0190] In some embodiments, a polypeptide may be at least 8 amino acid residues, 9 amino acid residues, 10 amino acid residues, 11 amino acid residues, 12 amino acid residues, 13 amino acid residues, 14 amino acid residues, 15 amino acid residues, 16 amino acid residues, 17 amino acid residues, 18 amino acid residues, 19 amino acid residues, 20 amino acid residues, 21 amino acid residues, 22 amino acid residues, 23 amino acid residues, 24 amino acid residues, 25 amino acid residues, 26 amino acid residues, 27 amino acid residues, 28 amino acid residues, 29 amino acid residues, 30 amino acid residues, 31 amino acid residues, 32 amino acid residues, 33 amino acid residues, 34 amino acid residues, 35 amino acid residues, 36 amino acid residues, 37 amino acid residues, 38 amino acid residues, 39 amino acid residues, 40 amino acid residues, 41 amino acid residues, 42 amino acid residues, 43 amino acid residues, 44 amino acid residues, 45 amino acid residues, 46 amino acid residues, 47 amino acid residues, 48 ​​amino acid residues, 49 amino acid residues, 50 amino acid residues, or more amino acid residues in length.In some embodiments, the polypeptide has at least 8 amino acid residues, 9 amino acid residues, 10 amino acid residues, 11 amino acid residues, 12 amino acid residues, 13 amino acid residues, 14 amino acid residues, 15 amino acid residues, 16 amino acid residues, 17 amino acid residues, 18 amino acid residues, 19 amino acid residues, 20 amino acid residues, 21 amino acid residues, 22 amino acid residues, 23 amino acid residues, 24 amino acid residues, 25 amino acid residues, 26 amino acid residues, 27 amino acid residues, 28 amino acid residues, 29 amino acid residues, 30 amino acid residues, 31 amino acid residues, 32 amino acid residues, 33 amino acid residues, 34 amino acid residues, 35 amino acid residues, 36 amino acid residues, 37 amino acid residues, 38 amino acid residues, 39 amino acid residues, 40 amino acid residues, 41 amino acid residues, 42 amino acid residues, 43 amino acid residues, 44 amino acid residues, 45 amino acid residues, 46 amino acid residues, 47 amino acid residues, 48 ​​amino acid residues, 49 amino acid residues, 50 amino acid residues, 51 amino acid residues, 52 amino acid residues, 53 amino acid residues, 54 amino acid residues, 55 amino acid residues, 56 amino acid residues, 57 amino acid residues, 58 amino acid residues, 59 amino acid residues, 60 amino acid residues, 61 amino acid residues, 62 amino acid residues, 63 amino acid residues, 64 amino acid residues, 65 amino acid residues, 66 amino acid residues, 67 amino acid residues, 68 amino acid residues, 69 amino acid residues, 70 amino acid residues, The length may be 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 55, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, or more amino acid residues. In some embodiments, a polypeptide may be at most 8 amino acid residues, 9 amino acid residues, 10 amino acid residues, 11 amino acid residues, 12 amino acid residues, 13 amino acid residues, 14 amino acid residues, 15 amino acid residues, 16 amino acid residues, 17 amino acid residues, 18 amino acid residues, 19 amino acid residues, 20 amino acid residues, 21 amino acid residues, 22 amino acid residues, 23 amino acid residues, 24 amino acid residues, 25 amino acid residues, 26 amino acid residues, 27 amino acid residues, 28 amino acid residues, 29 amino acid residues, 30 amino acid residues, 31 amino acid residues, 32 amino acid residues, 33 amino acid residues, 34 amino acid residues, 35 amino acid residues, 36 amino acid residues, 37 amino acid residues, 38 amino acid residues, 39 amino acid residues, 40 amino acid residues, 41 amino acid residues, 42 amino acid residues, 43 amino acid residues, 44 amino acid residues, 45 amino acid residues, 46 amino acid residues, 47 amino acid residues, 48 ​​amino acid residues, 49 amino acid residues, 50 amino acid residues, or fewer amino acid residues in length.In some embodiments, the polypeptide comprises at most 8 amino acid residues, 9 amino acid residues, 10 amino acid residues, 11 amino acid residues, 12 amino acid residues, 13 amino acid residues, 14 amino acid residues, 15 amino acid residues, 16 amino acid residues, 17 amino acid residues, 18 amino acid residues, 19 amino acid residues, 20 amino acid residues, 21 amino acid residues, 22 amino acid residues, 23 amino acid residues, 24 amino acid residues, 25 amino acid residues, 26 amino acid residues, 27 amino acid residues, 28 amino acid residues, 29 amino acid residues, 30 amino acid residues, 31 amino acid residues, 32 amino acid residues, 33 amino acid residues, 34 amino acid residues, 35 amino acid residues, 36 amino acid residues, 37 amino acid residues, 38 amino acid residues, 39 amino acid residues, 40 amino acid residues, 41 amino acid residues, 42 amino acid residues, 43 amino acid residues, 44 amino acid residues, 45 amino acid residues, 46 amino acid residues, 47 amino acid residues, 48 ​​amino acid residues, 49 amino acid residues, 50 amino acid residues, 51 amino acid residues, 52 amino acid residues, 53 amino acid residues, 54 amino acid residues, 55 amino acid residues, 56 amino acid residues, 57 amino acid residues, 58 amino acid residues, 59 amino acid residues, 60 amino acid residues, 61 amino acid residues, 62 amino acid residues, 63 amino acid residues, 64 amino acid residues, 65 amino acid residues, 66 amino acid residues, 67 amino acid residues, 68 amino acid residues, 69 amino acid residues, 70 amino acid residues, The length may be 7 amino acid residues, 38 amino acid residues, 39 amino acid residues, 40 amino acid residues, 41 amino acid residues, 42 amino acid residues, 43 amino acid residues, 44 amino acid residues, 45 amino acid residues, 46 amino acid residues, 47 amino acid residues, 48 ​​amino acid residues, 49 amino acid residues, 50 amino acid residues, 55 amino acid residues, 60 amino acid residues, 70 amino acid residues, 80 amino acid residues, 90 amino acid residues, 100 amino acid residues, 150 amino acid residues, 200 amino acid residues, 250 amino acid residues, 300 amino acid residues, 350 amino acid residues, 400 amino acid residues, 450 amino acid residues, 500 amino acid residues, or fewer amino acid residues.

[0191] In some embodiments, the polypeptide has a total length of at least 8 amino acids, at least 9 amino acids, at least 10 amino acids, at least 11 amino acids, at least 12 amino acids, at least 13 amino acids, at least 14 amino acids, at least 15 amino acids, at least 16 amino acids, at least 17 amino acids, at least 18 amino acids, at least 19 amino acids, at least 20 amino acids, at least 21 amino acids, at least 22 amino acids, at least 23 amino acids, at least 24 amino acids, at least 25 amino acids, at least 26 amino acids, at least 27 amino acids, at least 28 amino acids, at least 29 amino acids, at least 30 amino acids, at least 40 amino acids, at least 50 amino acids, at least 60 amino acids, at least 70 amino acids, at least 80 amino acids, at least 90 amino acids, at least 100 amino acids, at least 150 amino acids, at least 200 amino acids, at least 250 amino acids, at least 300 amino acids, at least 350 amino acids, at least 400 amino acids, at least 450 amino acids, at least 500 amino acids, at least 1000 amino acids, or at least 1500 amino acids.

[0192] In some embodiments, the polypeptide comprises at most 8 amino acids, at most 9 amino acids, at most 10 amino acids, at most 11 amino acids, at most 12 amino acids, at most 13 amino acids, at most 14 amino acids, at most 15 amino acids, at most 16 amino acids, at most 17 amino acids, at most 18 amino acids, at most 19 amino acids, at most 20 amino acids, at most 21 amino acids, at most 22 amino acids, at most 23 amino acids, at most 24 amino acids, at most 25 amino acids, at most 26 amino acids, at most 27 amino acids, at most 28 amino acids, at most 29 amino acids, at most 30 amino acids, at most 31 amino acids, at most 32 amino acids, at most 33 amino acids, at most 34 amino acids, at most 35 amino acids, at most 36 amino acids, at most 37 amino acids, at most 38 amino acids, at most 39 amino acids, at most 40 amino acids, at most 41 amino acids, at most 42 amino acids, at most 43 amino acids, at most 44 amino acids, at most 45 amino acids, at most 46 amino acids, at most 47 amino acids, at most 48 amino acids, at most 49 amino acids, at most 50 amino acids, at most 51 amino acids, at most 52 amino acids, at most 53 amino acids, at most 54 amino acids, at most 55 amino acids, at most 56 amino acids, at most 57 amino acids, at most 58 amino acids, at most 59 amino acids, at most 60 amino acids, at most 61 amino acids, at most 62 amino acids, at most 63 amino acids, at most 64 amino acids, at most 65 amino acids, at most 66 amino acids, at most 67 amino acids, at most 68 amino acids, at most 69 amino acids, at most 70 amino acids, at most It has a total length of 28 amino acids, at most 29 amino acids, at most 30 amino acids, at most 40 amino acids, at most 50 amino acids, at most 60 amino acids, at most 70 amino acids, at most 80 amino acids, at most 90 amino acids, at most 100 amino acids, at most 150 amino acids, at most 200 amino acids, at most 250 amino acids, at most 300 amino acids, at most 350 amino acids, at most 400 amino acids, at most 450 amino acids, at most 500 amino acids, at most 1000 amino acids, or at most 1500 amino acids.

[0193] In certain embodiments, the polypeptides described herein may comprise epitopes.In certain embodiments, epitopes may include, but are not limited to, about 5 amino acids, about 6 amino acids, about 7 amino acids, about 8 amino acids, about 9 amino acids, about 10 amino acids, about 11 amino acids, about 12 amino acids, about 13 amino acids, about 14 amino acids, about 15 amino acids, about 16 amino acids, about 17 amino acids, about 18 amino acids, about 19 amino acids, about 20 amino acids, about 21 amino acids, about 22 amino acids, about 23 amino acids, about 24 amino acids, about 25 amino acids, about 26 amino acids, about 27 amino acids, about 28 amino acids, about 29 amino acids, about 30 amino acids, about 31 amino acids, about 32 amino acids, about 33 amino acids, about 34 amino acids, about 35 amino acids, about 36 amino acids, about 37 amino acids, about 38 amino acids, about 39 amino acids, about 40 amino acids, about 41 amino acids, about 42 amino acids, about 43 amino acids, about 44 amino acids. , about 45 amino acids, about 46 amino acids, about 47 amino acids, about 48 amino acids, about 49 amino acids, about 50 amino acids, about 60 amino acids, about 70 amino acids, about 80 amino acids, about 90 amino acids, about 100 amino acids, about 110 amino acids, about 120 amino acids, about 150 amino acids, about 200 amino acids, about 300 amino acids, about 350 amino acids, about 400 amino acids, about 450 amino acids, about 500 amino acids, about 600 amino acids, about 700 amino acids, about 800 amino acids, about 900 amino acids, about 1,000 amino acids, about 1,500 amino acids, about 2,000 amino acids, about 2,500 amino acids, about 3,000 amino acids, about 4,000 amino acids, about 5,000 amino acids, about 7,500 amino acids, about 10,000 amino acids or more amino acid residues, and any range derivable therein.

[0194] In certain embodiments, the epitope may be from about 8 to about 50 amino acid residues in length, or from about 8 to about 30 amino acid residues, from about 8 to about 20 amino acid residues, from about 8 to about 18 amino acid residues, from about 8 to about 15 amino acid residues, or from about 8 to about 12 amino acid residues in length. In some embodiments, the peptide may be from about 8 to about 500 amino acid residues in length, or from about 8 to about 450 amino acid residues, from about 8 to about 400 amino acid residues, from about 8 to about 350 amino acid residues, from about 8 to about 300 amino acid residues, from about 8 to about 250 amino acid residues, from about 8 to about 200 amino acid residues, from about 8 to about 150 amino acid residues, from about 8 to about 100 amino acid residues, from about 8 to about 50 amino acid residues, or from about 8 to about 30 amino acid residues in length.

[0195] In certain embodiments, an epitope may be at least 8 amino acid residues, 9 amino acid residues, 10 amino acid residues, 11 amino acid residues, 12 amino acid residues, 13 amino acid residues, 14 amino acid residues, 15 amino acid residues, 16 amino acid residues, 17 amino acid residues, 18 amino acid residues, 19 amino acid residues, 20 amino acid residues, 21 amino acid residues, 22 amino acid residues, 23 amino acid residues, 24 amino acid residues, 25 amino acid residues, 26 amino acid residues, 27 amino acid residues, 28 amino acid residues, 29 amino acid residues, 30 amino acid residues, 31 amino acid residues, 32 amino acid residues, 33 amino acid residues, 34 amino acid residues, 35 amino acid residues, 36 amino acid residues, 37 amino acid residues, 38 amino acid residues, 39 amino acid residues, 40 amino acid residues, 41 amino acid residues, 42 amino acid residues, 43 amino acid residues, 44 amino acid residues, 45 amino acid residues, 46 amino acid residues, 47 amino acid residues, 48 ​​amino acid residues, 49 amino acid residues, 50 amino acid residues, or more amino acid residues in length.In some embodiments, the epitope comprises at least 8 amino acid residues, 9 amino acid residues, 10 amino acid residues, 11 amino acid residues, 12 amino acid residues, 13 amino acid residues, 14 amino acid residues, 15 amino acid residues, 16 amino acid residues, 17 amino acid residues, 18 amino acid residues, 19 amino acid residues, 20 amino acid residues, 21 amino acid residues, 22 amino acid residues, 23 amino acid residues, 24 amino acid residues, 25 amino acid residues, 26 amino acid residues, 27 amino acid residues, 28 amino acid residues, 29 amino acid residues, 30 amino acid residues, 31 amino acid residues, 32 amino acid residues, 33 amino acid residues, 34 amino acid residues, 35 amino acid residues, 36 amino acid residues, 37 amino acid residues, 38 amino acid residues, 39 amino acid residues, 40 amino acid residues, 41 amino acid residues, 42 amino acid residues, 43 amino acid residues, 44 amino acid residues, 45 amino acid residues, 46 amino acid residues, 47 amino acid residues, 48 ​​amino acid residues, 49 amino acid residues, 50 amino acid residues, 51 amino acid residues, 52 amino acid residues, 53 amino acid residues, 54 amino acid residues, 55 amino acid residues, 56 amino acid residues, 57 amino acid residues, 58 amino acid residues, 59 amino acid residues, 60 amino acid residues, 61 amino acid residues, 62 amino acid residues, 63 amino acid residues, 64 amino acid residues, 65 amino acid residues, 66 amino acid residues, 67 amino acid residues, 68 amino acid residues, 69 amino acid residues, 70 amino acid residue The length may be 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 55, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, or more amino acid residues. In some embodiments, an epitope may be at most 8 amino acid residues, 9 amino acid residues, 10 amino acid residues, 11 amino acid residues, 12 amino acid residues, 13 amino acid residues, 14 amino acid residues, 15 amino acid residues, 16 amino acid residues, 17 amino acid residues, 18 amino acid residues, 19 amino acid residues, 20 amino acid residues, 21 amino acid residues, 22 amino acid residues, 23 amino acid residues, 24 amino acid residues, 25 amino acid residues, 26 amino acid residues, 27 amino acid residues, 28 amino acid residues, 29 amino acid residues, 30 amino acid residues, 31 amino acid residues, 32 amino acid residues, 33 amino acid residues, 34 amino acid residues, 35 amino acid residues, 36 amino acid residues, 37 amino acid residues, 38 amino acid residues, 39 amino acid residues, 40 amino acid residues, 41 amino acid residues, 42 amino acid residues, 43 amino acid residues, 44 amino acid residues, 45 amino acid residues, 46 amino acid residues, 47 amino acid residues, 48 ​​amino acid residues, 49 amino acid residues, 50 amino acid residues, or fewer amino acid residues in length.In some embodiments, the epitope comprises at most 8 amino acid residues, 9 amino acid residues, 10 amino acid residues, 11 amino acid residues, 12 amino acid residues, 13 amino acid residues, 14 amino acid residues, 15 amino acid residues, 16 amino acid residues, 17 amino acid residues, 18 amino acid residues, 19 amino acid residues, 20 amino acid residues, 21 amino acid residues, 22 amino acid residues, 23 amino acid residues, 24 amino acid residues, 25 amino acid residues, 26 amino acid residues, 27 amino acid residues, 28 amino acid residues, 29 amino acid residues, 30 amino acid residues, 31 amino acid residues, 32 amino acid residues, 33 amino acid residues, 34 amino acid residues, 35 amino acid residues, 36 amino acid residues, 37 amino acid residues, 38 amino acid residues, 39 amino acid residues, 40 amino acid residues, 41 amino acid residues, 42 amino acid residues, 43 amino acid residues, 44 amino acid residues, 45 amino acid residues, 46 amino acid residues, 47 amino acid residues, 48 ​​amino acid residues, 49 amino acid residues, 50 amino acid residues, 51 amino acid residues, 52 amino acid residues, 53 amino acid residues, 54 amino acid residues, 55 amino acid residues, 56 amino acid residues, 57 amino acid residues, 58 amino acid residues, 59 amino acid residues, 60 amino acid residues, 61 amino acid residues, 62 amino acid residues, 63 amino acid residues, 64 amino acid residues, 65 amino acid residues, 66 amino acid residues, 67 amino acid residues, 68 amino acid residues, 69 amino acid residues, 70 amino acid residue The length may be 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 55, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, or fewer amino acid residues.

[0196] Longer peptides can be designed in several ways. In some embodiments, where HLA-binding peptides are predicted or known, the longer peptides include (1) individual binding peptides with 2-5 amino acid extensions toward the N-terminus and C-terminus of each of the corresponding gene products; or (2) a concatenation of some or all of the binding peptides with each extended sequence. In other embodiments, where sequencing reveals that long (>10 residues) neoepitope sequences are present in the tumor (e.g., due to frameshift, read-through or intron inclusion leading to a novel peptide sequence), the longer peptides can consist of the entire run of novel tumor-specific amino acids, either as a single longer peptide or as several overlapping longer peptides. In some embodiments, it is presumed that the use of longer peptides may allow endogenous processing by patient cells, leading to more effective antigen presentation and induction of T cell responses. In some embodiments, two or more peptides can be used, where the peptides are overlapping and arranged in a non-overlapping manner across the long neo-antigenic peptide.

[0197] In some embodiments, the MHC class I immunogenic antigen, neo-antigenic peptide, or epitope thereof is 12 amino acid residues or less in length, usually consisting of between about 8 and about 12 amino acid residues. In some embodiments, the MHC class I immunogenic antigen, neo-antigenic peptide, or epitope thereof is about 8, about 9, about 10, about 11, or about 12 amino acid residues. In some embodiments, the MHC class II immunogenic antigen, neo-antigenic peptide, or epitope thereof is 25 amino acid residues or less in length, usually consisting of between about 9 and about 25 amino acid residues. In some embodiments, the MHC class II immunogenic antigen, neo-antigenic peptide, or epitope thereof is about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, or about 25 amino acid residues.

[0198] In some embodiments, the antigen, neoantigenic peptide, or epitope binds to an HLA protein (e.g., an MHC class I HLA or an MHC class II HLA). In certain embodiments, the antigen, neoantigenic peptide, or epitope binds to an HLA protein with greater affinity than the corresponding wild-type peptide. In certain embodiments, the antigen, neoantigenic peptide, or epitope has an IC of at least 5000 nM or less, at least 500 nM or less, at least 100 nM or less, at least 50 nM or less. 50 or K DIn some embodiments, the antigen, neoantigenic peptide, or epitope binds to MHC class I HLA. In some embodiments, the antigen, neoantigenic peptide, or epitope binds to MHC class I HLA with an affinity of 0.1 nM to 2000 nM. In some embodiments, the antigen, neoantigenic peptide, or epitope binds to MHC class I HLA with an affinity of 0.1 nM to 2000 nM. 0.1nM, 0.2nM, 0.3nM, 0.4nM, 0.5nM, 0.6nM, 0.7nM, 0.8nM, 0.9nM, 1nM, 2nM, 3nM, 4nM, 5nM, 6nM, 7nM, 8nM, 9nM, 1 for HLA 0nM, 15nM, 20nM, 25nM, 30nM, 35nM, 40nM, 45nM, 50nM, 55nM, 60nM, 65nM, 70nM, 75nM, 80nM, 85nM, 90nM, 95nM, 100nM, In some embodiments, the antigen, neoantigen peptide, or epitope binds to an MHC class II HLA. In some embodiments, the antigen, neoantigenic peptide, or epitope binds to MHC class II HLA with an affinity of between 0.1 nM and 2000 nM, between 1 nM and 1000 nM, between 10 nM and 500 nM, or less than 1000 nM.In some embodiments, the antigen, neoantigen peptide, or epitope is administered to an MHC class II HLA at 0.1 nM, 0.2 nM, 0.3 nM, 0.4 nM, 0.5 nM, 0.6 nM, 0.7 nM, 0.8 nM, 0.9 nM, 1 nM, 2 nM, 3 nM, 4 nM, 5 nM, 6 nM, 7 nM, 8 nM, 9 nM, 10 nM, 15 nM, 20 nM, 25 nM, 30 nM, 35 nM, 40 nM, 45 nM, 50 nM, 55 nM, 60 nM, 65 nM, 70 nM, 75 nM, 80 nM, 85 nM, 90 nM, 95 nM, 100 nM, binds with an affinity of 150nM, 200nM, 250nM, 300nM, 350nM, 400nM, 450nM, 500nM, 550nM, 600nM, 650nM, 700nM, 750nM, 800nM, 850nM, 900nM, 950nM, 1000nM, 1100nM, 1200nM, 1300nM, 1400nM, 1500nM, 1600nM, 1700nM, 1800nM, 1900nM, or 2000nM.

[0199] In some embodiments, the antigen, neoantigenic peptide, or epitope binds to MHC class I HLA with a stability between 10 minutes and 24 hours. In some embodiments, the antigen, neoantigenic peptide, or epitope binds to MHC class I HLA with a stability of 10 minutes, 11 minutes, 12 minutes, 13 minutes, 14 minutes, 15 minutes, 16 minutes, 17 minutes, 18 minutes, 19 minutes, 20 minutes, 25 minutes, 30 minutes, 35 minutes, 40 minutes, 45 minutes, 50 minutes, 55 minutes, or 60 minutes. In some embodiments, the antigen, neoantigenic peptide, or epitope binds to MHC class I HLA with a stability of 1 hour, 1.5 hours, 2 hours, 2.5 hours, 3 hours, 3.5 hours, 4 hours, 4.5 hours, 5 hours, 5.5 hours, 6 hours, 6.5 hours, 7 hours, 7.5 hours, 8 hours, 8.5 hours, 9 hours, 9.5 hours, 10 hours, 11 hours, 12 hours, 13 hours, 14 hours, 15 hours, 16 hours, 17 hours, 18 hours, 19 hours, 20 hours, 21 hours, 22 hours, 23 hours, or 24 hours. In some embodiments, the antigen, neoantigenic peptide, or epitope binds to MHC class II HLA with a stability of 10 minutes to 24 hours. In some embodiments, the antigen, neoantigenic peptide, or epitope binds to MHC class II HLA with a stability of 10 minutes, 11 minutes, 12 minutes, 13 minutes, 14 minutes, 15 minutes, 16 minutes, 17 minutes, 18 minutes, 19 minutes, 20 minutes, 25 minutes, 30 minutes, 35 minutes, 40 minutes, 45 minutes, 50 minutes, 55 minutes, or 60 minutes. binds to HLA with stability for 1 hour, 1.5 hours, 2 hours, 2.5 hours, 3 hours, 3.5 hours, 4 hours, 4.5 hours, 5 hours, 5.5 hours, 6 hours, 6.5 hours, 7 hours, 7.5 hours, 8 hours, 8.5 hours, 9 hours, 9.5 hours, 10 hours, 11 hours, 12 hours, 13 hours, 14 hours, 15 hours, 16 hours, 17 hours, 18 hours, 19 hours, 20 hours, 21 hours, 22 hours, 23 hours, or 24 hours.

[0200] In some embodiments, the polypeptide may have a pI value of about 0.5 to about 12, about 2 to about 10, or about 4 to about 8. In some embodiments, the peptide may have a pI value of at least 4.5, 5, 5.5, 6, 6.5, 7, 7.5, or more. In some embodiments, the polypeptide may have a pI value of at most 4.5, 5, 5.5, 6, 6.5, 7, 7.5, or less.

[0201] In some embodiments, the polypeptides described herein comprise an amino acid or amino acid sequence of a peptide sequence that is not encoded by a nucleic acid sequence immediately upstream of a nucleic acid sequence in the subject's genome that encodes an epitope. In some embodiments, the polypeptides described herein comprise an amino acid or amino acid sequence of a peptide sequence that is not encoded by a nucleic acid sequence immediately downstream of a nucleic acid sequence in the subject's genome that encodes an epitope. In some embodiments, the amino acid or amino acid sequence comprises 0-1000, 1-900, 5-800, 10-700, 20-600, 30-500, 40-400, 50-300, 60-200, or 70-100 amino acid residues. In a preferred embodiment, the amino acid or amino acid sequence comprises from 1 to 20 amino acid residues. In another preferred embodiment, the amino acid or amino acid sequence comprises from 5 to 12 amino acid residues. In some embodiments, the amino acid or amino acid sequence is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, At least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 1000, or at least 1500 amino acid residues.In some embodiments, the amino acid or amino acid sequence is about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 101, about About 5, about 26, about 27, about 28, about 29, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 150, about 200, about 250, about 300, about 350, about 400, about 450, about 500, about 1000, or about 1500 amino acid residues.

[0202] In one aspect, provided herein is a method for producing a polypeptide, comprising linking an amino acid or amino acid sequence and / or a linker to the N-terminus and / or C-terminus of a sequence comprising an epitope sequence. In some embodiments, the polypeptides described herein may be in solution form, lyophilized form, or may be in crystalline form. In some embodiments, the polypeptides described herein may be synthetically prepared by recombinant DNA technology or chemical synthesis, or may be isolated from natural sources such as native tumors or pathogenic organisms. The epitopes or neoepitopes may be synthesized separately or may be directly or indirectly conjugated into the polypeptide. The polypeptides described herein may be substantially free of other naturally occurring host cell proteins and fragments thereof, although in some embodiments, the polypeptides may be synthetically conjugated to join native fragments or particles.

[0203] In some embodiments, the polypeptides described herein can be prepared in a variety of ways. In some embodiments, the polypeptides can be synthesized in solution or on solid support according to conventional techniques. A variety of automated synthesizers are commercially available and can be used according to known protocols. For example, see Stewart & Young, Solid Phase Peptide Synthesis, 2d. Ed., Pierce Chemical Co., 1984. In addition, individual polypeptides can be joined using chemical ligation to make larger polypeptides, which still falls within the scope of the present disclosure.

[0204] Alternatively, recombinant DNA technology can be used in which a nucleotide sequence encoding a polypeptide or a portion of a polypeptide is inserted into an expression vector, transformed or transfected into a suitable host cell, and cultivated under conditions suitable for expression. These procedures are described in Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press, Cold Spring Harbor, NY (1989). As described above, these methods are generally known in the art. Thus, recombinant peptides comprising one or more of the neoantigenic peptides described herein can be used to present appropriate T cell epitopes.

[0205] In some embodiments, the polypeptide comprises at least one mutant amino acid. In some embodiments, the at least one mutant amino acid is encoded by an insertion of one or more nucleotides in a nucleic acid sequence in the genome of the subject. In some embodiments, the at least one mutant amino acid is encoded by a deletion of one or more nucleotides in a nucleic acid sequence in the genome of the subject. In some embodiments, the at least one mutant amino acid is encoded by a frameshift in a nucleic acid sequence in the genome of the subject. A frameshift occurs when a mutation disrupts the normal phase of the codon periodicity (also known as "reading frame") of a gene, resulting in the translation of a non-native protein sequence. The same reading frame change can be achieved by different mutations in a gene. In some embodiments, the at least one mutant amino acid is encoded by a neo-ORF in a nucleic acid sequence in the genome of the subject. In some embodiments, the at least one mutant amino acid is encoded by a point mutation in a nucleic acid sequence in the genome of the subject. In some embodiments, the at least one mutant amino acid is encoded by a gene having a mutation that results in a fusion polypeptide, an in-frame deletion, an insertion, expression of an endogenous retroviral polypeptide, and tumor-specific overexpression of the polypeptide. In some embodiments, the at least one mutant amino acid is encoded by a fusion of a first gene and a second gene in the genome of the subject. In some embodiments, the at least one mutant amino acid is encoded by an in-frame fusion of a first gene and a second gene in the genome of the subject. In some embodiments, the at least one mutant amino acid is encoded by a fusion of an exon of a splice variant of the first gene in the genome of the subject. In some embodiments, the at least one mutant amino acid is encoded by a fusion of a cryptic exon of the first gene and a first gene in the genome of the subject.

[0206] In some aspects, the disclosure provides a polypeptide comprising at least two polypeptide molecules. In some embodiments, two or more of the at least two polypeptides or polypeptide molecules comprise an epitope. In some embodiments, two or more of the at least two polypeptides or polypeptide molecules comprise the same epitope. In some embodiments, two or more of the at least two polypeptides or polypeptide molecules comprise the same epitope of the same length. In some embodiments, two or more of the at least two polypeptides or polypeptide molecules comprise an amino acid or amino acid sequence that is a peptide sequence that is not encoded by a nucleic acid sequence immediately upstream or downstream of the nucleic acid sequence in the genome of the subject that encodes the epitope. In some embodiments, the amino acid or amino acid sequence that is a peptide sequence that is not encoded by a nucleic acid sequence immediately upstream of the nucleic acid sequence in the genome of the subject that encodes the epitope of two or more of the at least two polypeptides or polypeptide molecules is the same. In some embodiments, the amino acid or amino acid sequence that is a peptide sequence that is not encoded by a nucleic acid sequence immediately downstream of the nucleic acid sequence in the genome of the subject that encodes the epitope of two or more of the at least two polypeptides or polypeptide molecules is the same.

[0207] In some embodiments, two or more of the at least two polypeptides or polypeptide molecules include a linker. In some embodiments, two or more of the at least two polypeptides or polypeptide molecules include a linker at the N-terminus and / or C-terminus of the epitope. In some embodiments, two or more of the at least two polypeptides or polypeptide molecules include different linkers. In some embodiments, a first polypeptide or polypeptide molecule of the at least two polypeptides or polypeptide molecules does not include a linker, and a second polypeptide or polypeptide molecule of the at least two polypeptides or polypeptide molecules includes a linker. In some embodiments, a first polypeptide or polypeptide molecule of the at least two polypeptides or polypeptide molecules does not include a linker at the N-terminus of the epitope, and a second polypeptide or polypeptide molecule of the at least two polypeptides or polypeptide molecules includes a linker at the N-terminus of the epitope. In some embodiments, a first polypeptide or polypeptide molecule of the at least two polypeptides or polypeptide molecules does not include a linker at the C-terminus of the epitope, and a second polypeptide or polypeptide molecule of the at least two polypeptides or polypeptide molecules includes a linker at the C-terminus of the epitope. In some embodiments, a first polypeptide or polypeptide molecule of the at least two polypeptides or polypeptide molecules comprises a linker and a second polypeptide or polypeptide molecule of the at least two polypeptides or polypeptide molecules does not comprise a linker. In some embodiments, a first polypeptide or polypeptide molecule of the at least two polypeptides or polypeptide molecules comprises a linker at the N-terminus of the epitope and a second polypeptide or polypeptide molecule of the at least two polypeptides or polypeptide molecules does not comprise a linker at the N-terminus of the epitope.In some embodiments, a first polypeptide or polypeptide molecule of the at least two polypeptides or polypeptide molecules comprises a linker at the C-terminus of the epitope, and a second polypeptide or polypeptide molecule of the at least two polypeptides or polypeptide molecules does not comprise a linker at the C-terminus of the epitope.

[0208] Disulfide linkers can be synthesized using methods well known in the art. For example, disulfide linkers can be synthesized according to Zhang, Donglu, et al., ACS Med. Chem. Lett. 2016, 7, 988-993; and Pillow, Thomas H., et al., Chem. Sci., 2017, 8, 366-370. Examples of disulfide linker synthesis and disulfide-containing peptide synthesis are shown in Examples 3 and 4. PABC-containing peptides can be synthesized using methods well known in the art. For example, PABC-containing peptides can be synthesized according to Laurent Ducry (ed.), Antibody-Drug Conjugates, Methods in Molecular Biology, vol. 1045, DOI 10.1007 / 978-1-62703-541-5_5, Springer Science + Business Media, LLC 2013. PABC-containing An example of peptide synthesis is provided in Example 5. In some embodiments, any resin designed for solid phase peptide synthesis can be used.

[0209] In some embodiments, a polypeptide comprises at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or more polypeptides or polypeptide molecules. For example, a polypeptide may comprise 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 or more polypeptides or polypeptide molecules.

[0210] In some embodiments, the antigen, neoantigen peptide, or epitope-containing polypeptide comprises a RAS epitope. In some embodiments, the peptide can be derived from a protein having a substitution mutation, e.g., a KRAS G12C, G12D, G12V, Q61H, or Q61L mutation, or an NRAS Q61K or Q61R mutation. The substitution can be located anywhere along the length of the peptide. For example, the substitution can be located in the N-terminal third of the peptide, the central third of the peptide, or the C-terminal third of the peptide. In another embodiment, the substituted residue is located 2-5 residues away from the N-terminus or 2-5 residues away from the C-terminus. The peptide can similarly be derived from a tumor-specific insertion mutation, where the peptide comprises one or more or all of the inserted residues. In some embodiments, the epitope comprises at least 8 consecutive amino acids of a mutant RAS protein comprising a mutation at G12, G13, or Q61, and a mutant RAS sequence comprising a mutation at G12, G13, or Q61. In some embodiments, at least 8 consecutive amino acids of the mutant RAS protein comprising a mutation in G12, G13 or Q61 comprises a G12A, G12C, G12D, G12R, G12S, G12V, G13A, G13C, G13D, G13R, G13S, G13V, Q61H, Q61L, Q61K or Q61R mutation.In some embodiments, the mutation in G12, G13 or Q61 comprises a G12A, G12C, G12D, G12R, G12S, G12V, G13A, G13C, G13D, G13R, G13S, G13V, Q61H, Q61L, Q61K or Q61R mutation.

[0211] In some embodiments, the polypeptide comprising a RAS epitope further comprises an amino acid sequence. In some embodiments, the amino acid sequence is that of a cytomegalovirus (CMV) protein, such as pp65. In some embodiments, the amino acid sequence is that of a human immunodeficiency virus (HIV) protein. In some embodiments, the amino acid sequence is that of a MART-1 protein. In some embodiments, the amino acid sequence of a CMV protein, such as pp65, comprises one, two, three, or more than three amino acid residues. In some embodiments, the amino acid sequence of a CMV protein, such as pp65, comprises 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amino acid residues. In some embodiments, the amino acid sequence of an HIV protein comprises 1, 2, 3, or more than 3 amino acid residues. In some embodiments, the amino acid sequence of the HIV protein comprises 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amino acid residues. In some embodiments, the amino acid sequence of the MART-1 protein comprises 1, 2, 3, or more than 3 amino acid residues. In some embodiments, the amino acid sequence of the MART-1 protein comprises 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amino acid residues.

[0212] In some embodiments, the RAS epitope binds to a protein encoded by an HLA allele. In some embodiments, a RAS epitope binds to a protein encoded by an HLA allele with an affinity of less than 10 μM, less than 9 μM, less than 8 μM, less than 7 μM, less than 6 μM, less than 5 μM, less than 4 μM, less than 3 μM, less than 2 μM, less than 1 μM, less than 950 nM, less than 900 nM, less than 850 nM, less than 800 nM, less than 750 nM, less than 600 nM, less than 550 nM, less than 500 nM, less than 450 nM, less than 400 nM, less than 350 nM, less than 300 nM, less than 250 nM, less than 200 nM, less than 150 nM, less than 100 nM, less than 90 nM, less than 80 nM, less than 70 nM, less than 60 nM, less than 50 nM, less than 40 nM, less than 30 nM, less than 20 nM, or less than 10 nM. In some embodiments, the RAS epitope is associated with a protein encoded by an HLA allele for greater than 24 hours, greater than 23 hours, greater than 22 hours, greater than 21 hours, greater than 20 hours, greater than 19 hours, greater than 18 hours, greater than 17 hours, greater than 16 hours, greater than 15 hours, greater than 14 hours, greater than 13 hours, greater than 12 hours, greater than 11 hours, greater than 10 hours, greater than 9 hours, greater than 8 hours, greater than 7 hours, greater than 6 hours, greater than 5 hours In some embodiments, the binding occurs with a stability of greater than 1 hour, greater than 4 hours, greater than 3 hours, greater than 2 hours, greater than 1 hour, greater than 55 minutes, greater than 50 minutes, greater than 45 minutes, greater than 40 minutes, greater than 35 minutes, greater than 30 minutes, greater than 25 minutes, greater than 20 minutes, greater than 15 minutes, greater than 10 minutes, greater than 9 minutes, greater than 8 minutes, greater than 7 minutes, greater than 6 minutes, greater than 5 minutes, greater than 4 minutes, greater than 3 minutes, greater than 2 minutes, or greater than 1 minute.

[0213] In some embodiments, the HLA allele is selected from the group consisting of HLA-A02:01 allele, HLA-A03:01 allele, HLA-A11:01 allele, HLA-A03:02 allele, HLA-A30:01 allele, HLA-A31:01 allele, HLA-A33:01 allele, HLA-A33:03 allele, HLA-A68:01 allele, HLA-A74:01 allele, and / or HLA-C08:02 allele, and any combination thereof. In some embodiments, the HLA allele is HLA-A02:01. In some embodiments, the HLA allele is HLA-A03:01 allele. In some embodiments, the HLA allele is HLA-A11:01 allele. In some embodiments, the HLA allele is HLA-A03:02 allele. In some embodiments, the HLA allele is an HLA-A30:01 allele. In some embodiments, the HLA allele is an HLA-A31:01 allele. In some embodiments, the HLA allele is an HLA-A33:01 allele. In some embodiments, the HLA allele is an HLA-A33:03 allele. In some embodiments, the HLA allele is an HLA-A68:01 allele. In some embodiments, the HLA allele is an HLA-A74:01 allele. In some embodiments, the HLA allele is an HLA-C08:02 allele.

[0214] In some aspects, the present disclosure provides a composition comprising a single polypeptide comprising a first peptide and a second peptide, or a single polynucleotide encoding the first peptide and the second peptide. In some embodiments, the composition provided herein comprises one or more additional peptides, the one or more additional peptides comprising a third neoepitope. In some embodiments, the first peptide and the second peptide are encoded by a sequence transcribed from the same transcription start site. In some embodiments, the first peptide is encoded by a sequence transcribed from a first transcription start site, and the second peptide is encoded by a sequence transcribed from a second transcription start site. In some embodiments, the polypeptide has a length of at least 26 amino acids, 27 amino acids, 28 amino acids, 29 amino acids, 30 amino acids, 40 amino acids, 50 amino acids, 60 amino acids, 70 amino acids, 80 amino acids, 90 amino acids, 100 amino acids, 150 amino acids, 200 amino acids, 250 amino acids, 300 amino acids, 350 amino acids, 400 amino acids, 450 amino acids, 500 amino acids, 600 amino acids, 700 amino acids, 800 amino acids, 900 amino acids, 1,000 amino acids, 1,500 amino acids, 2,000 amino acids, 2,500 amino acids, 3,000 amino acids, 4,000 amino acids, 5,000 amino acids, 7,500 amino acids, or 10,000 amino acids.In some embodiments, the polypeptide has at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to the corresponding wild-type sequence. a first sequence; and a second sequence having at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to the corresponding wild-type sequence. In some embodiments, the polypeptide comprises a first repeat sequence of at least 8 or 9 contiguous amino acids that have at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to the corresponding wild-type sequence. 1 sequence; and a second sequence of at least 16 or 17 contiguous amino acids having at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to the corresponding wild-type sequence.

[0215] In some embodiments, the second peptide is longer than the first peptide. In some embodiments, the first peptide is longer than the second peptide. In some embodiments, the first peptide is at least 9 amino acids, 10 amino acids, 11 amino acids, 12 amino acids, 13 amino acids, 14 amino acids, 15 amino acids, 16 amino acids, 17 amino acids, 18 amino acids, 19 amino acids, 20 amino acids, 21 amino acids, 22 amino acids, 23 amino acids, 24 amino acids, 25 amino acids, 26 amino acids, 27 amino acids, 28 amino acids, 29 amino acids, 30 amino acids, 40 amino acids, 50 amino acids, 60 amino acids, 70 amino acids, 80 amino acids, The length may be 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 1,500, 2,000, 2,500, 3,000, 4,000, 5,000, 7,500, or 10,000 amino acids. In some embodiments, the second peptide has a length of at least 17 amino acids, 18 amino acids, 19 amino acids, 20 amino acids, 21 amino acids, 22 amino acids, 23 amino acids, 24 amino acids, 25 amino acids, 26 amino acids, 27 amino acids, 28 amino acids, 29 amino acids, 30 amino acids, 40 amino acids, 50 amino acids, 60 amino acids, 70 amino acids, 80 amino acids, 90 amino acids, 100 amino acids, 150 amino acids, 200 amino acids, 250 amino acids, 300 amino acids, 350 amino acids, 400 amino acids, 450 amino acids, 500 amino acids, 600 amino acids, 700 amino acids, 800 amino acids, 900 amino acids, 1,000 amino acids, 1,500 amino acids, 2,000 amino acids, 2,500 amino acids, 3,000 amino acids, 4,000 amino acids, 5,000 amino acids, 7,500 amino acids, or 10,000 amino acids.In some embodiments, the first peptide comprises a sequence of at least 9 contiguous amino acids that is at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to a corresponding wild-type sequence. In some embodiments, the second peptide comprises a sequence of at least 17 contiguous amino acids having at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to the corresponding wild-type sequence.

[0216] In some embodiments, the first peptide, the second peptide, or both, comprise at least one flanking sequence, and the at least one flanking sequence is upstream or downstream of the neoepitope.In some embodiments, the at least one flanking sequence has at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the corresponding wild-type sequence.In some embodiments, the at least one flanking sequence comprises a non-wild-type sequence. In some embodiments, at least one flanking sequence is an N-terminal flanking sequence. In some embodiments, at least one flanking sequence is a C-terminal flanking sequence. In some embodiments, at least one flanking sequence of the first peptide has at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to at least one flanking sequence of the second peptide. In some embodiments, at least one flanking region of the first peptide differs from at least one flanking region of the second peptide, hi some embodiments, at least one flanking residue comprises a mutation.

[0217] In some embodiments, the peptide comprises a neoepitope sequence that includes at least one mutant amino acid. In some embodiments, the peptide comprises a neoepitope sequence that includes at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more mutant amino acids. In some embodiments, the peptide comprises a neoepitope sequence derived from a protein that includes at least one mutated amino acid and at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more non-mutated amino acids. In some embodiments, the peptide comprises a neoepitope sequence derived from a protein that includes at least one mutant amino acid and at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more non-mutated amino acids upstream of the at least one mutant amino acid. In some embodiments, the peptide comprises a neoepitope sequence derived from a protein that includes at least one mutant amino acid and at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more non-mutated amino acids downstream of the at least one mutant amino acid.In some embodiments, the peptide comprises at least one mutated amino acid; at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more non-mutated amino acids upstream of the at least one mutated amino acid. and a neoepitope sequence derived from a protein comprising at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more non-mutated amino acids downstream of the at least one mutated amino acid.

[0218] In some embodiments, the peptide comprises a neoepitope sequence derived from a protein comprising at least one mutant amino acid and a sequence upstream of the at least one mutant amino acid that has at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the corresponding wild type sequence. In some embodiments, the peptide comprises a neoepitope sequence derived from a protein comprising at least one mutant amino acid and a sequence downstream of the at least one mutant amino acid and having at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the corresponding wild type sequence.In some embodiments, the peptide comprises at least one mutant amino acid, a sequence upstream of the at least one mutant amino acid that has at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the corresponding wild type sequence. and a neoepitope sequence derived from a protein comprising a sequence downstream of the at least one mutant amino acid and having at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the corresponding wild type sequence.

[0219] In some embodiments, the peptide has a sequence similar to that of the at least one mutant amino acid and at least one mutant amino acid upstream of the at least one mutant amino acid, relative to the corresponding wild type sequence by at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 140%, 141%, 142%, 143%, 144%, 145%, 146%, 147%, 148%, 149%, 150%, 151%, 152%, 153%, 154%, 155%, 156%, 157%, 158%, 159%, 160 and / or a neoepitope sequence derived from a protein comprising a sequence comprising at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more consecutive amino acids with 3%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity. In some embodiments, the peptide has a sequence similar to at least one mutant amino acid and at least one mutant amino acid downstream of the corresponding wild type sequence, with at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 140%, 141%, 142%, 143%, 144%, 145%, 146%, 147%, 148%, 149%, 150%, 151%, 152%, 153%, 154%, 155%, 156%, 157%, 158%, 159%, 160%, 161%, 162%, and / or a neoepitope sequence derived from a protein comprising a sequence comprising at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more consecutive amino acids with 3%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity.In some embodiments, the peptide comprises at least one mutant amino acid, at least one mutant amino acid upstream of the at least one mutant amino acid, and a sequence similar to that of the corresponding wild type sequence, the sequence similar to that of the corresponding wild type sequence, and the ... and the sequence similar to that of the corresponding wild type sequence, and the sequence similar to , 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the sequence of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more consecutive amino acids. , and downstream of at least one mutant amino acid that is at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, and neoepitope sequences derived from proteins that contain sequences that include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more consecutive amino acids with 97%, 98%, 99% or 100% sequence identity.

[0220] In some embodiments, the epitope is the TMPRSS2:ERG epitope. In some embodiments, the TMPRSS2:ERG epitope comprises the amino acid sequence of ALNSEALSV. In some embodiments, a polypeptide comprising a RAS epitope is selected from the group consisting of GADGVGKSAL, GACGVGKSAL, GAVGVGKSAL, GADGVGKSA, GACGVGKSA, GAVGVGKSA, KLVVVGACGV, FLVVVGACGL, FMVVVGACGI, FLVVVGACGI, FMVVVGACGV, FLVVVGACGV, MLVVVGACGV, FMVVVGACGL, YLVVVGACGV, KMVVVGACGV, YMVVVGACGV, MMVVVGACGV, DT AGHEEY, TAGEEYSAM, DILDTAGHE, DILDTAGH, ILDTAGHEE, ILDTAGHE, DILDTAGHEEY, DTAGHEEYS, LLDILDTAGH, DILTAGRE, DILDTAGR, ILDTA GREE, ILDTAGRE, CLLDILDTAGR, TAGREEYSAM, REEYSAMRD, DTAGKEEYSAM, CLLDILDTAGK, DTAGKEEY, LLDILDTAGK, ILDTAGKE, ILDTAGKEE, DTAG LEEY,ILDTAGLE,DILDTAGL,ILDTAGLEE,GLEEYSAMRDQY,LLDILDTAGLE,LDILDTAGL,DILDTAGLE,DILDTAGLEEY,AGVGKSAL,GAAGVGKSAL,AAG VGKSAL, CGVGKSAL, ACGVGKSAL, DGVGKSAL, ADGVGKSAL, DGVGKSALTI, GARGVGKSA, KLVVVGARGV, VVVGARGV, SGVGKSAL, VVVGASGVGK, GASGVGK SAL, VGVGKSAL, VVVGAGCVGK, KLVVVGAGC, GDVGKSAL, DVGKSALTI, VVVGAGDVGK, TAGKEEYSAM, DTAGEEYSAM, TAGHEEYSA, DTAGREEYSAM, TAGK EEYSA, AAGVGKSA, AGCVGKSAL, AGDVGKSAL, AGKEEYSAMR, AGVGKSALTI, ARGVGKSAL, ASGVGKSA, ASGVGKSAL, AVGVGKSA, CVGKSALTI, DILDTAGK,DILDTAGREEY, DTAGHEEYSAMR, DTAGKEEYS, DTAGKEEYSAMR, DTAGLEEYS, DTAGLEEYSA, DTAGLEEYSAMR, DTAGREEYS, DTAGREEYSAMR, GAAGVGKSA, GA CGVGKSA, GACGVGKSAL, GADGVGKS, GAGDVGKSA, GAGDVGKSAL, GASGVGKSA, GCVGKSAL, GCVGKSALTI, GHEEYSAM, GKEEYSAM, GLEEYSAMR, GREEYSAM, G The amino acid sequence of REEYSAMR, HEEYSAMRD, KEEYSAMRD, KLVVVGASG, LDILDTAGR, LEEYSAMRD, LVVVGARGV, LVVVGASGV, REEYSAMRDQY, RGVGKSAL, TAGLEEYSA, TEYKLVVVGAA, VGAAGVGKSA, VGADGVGK, VGASGVGKSA, VGVGKSALTI, VVVGAAGV, VVVGAVGV, YKLVVVGAC, YKLVVVGAD, YKLVVVGAR, or DILDTAGKE.

[0221] In some embodiments, a polypeptide comprising a RAS epitope can be, for example at the N-terminus, K, KK, KKK, KKKK, KKKKK, KKKKKK, KKKKKKKK, KTEY, KTEYK, KTEYKL, KTEYKLV, KTEYKLVV, KTEYKLVVV, KKTEY, KKTEYK, KKTEYKL, KKTEYKLV, KKTEYKLVV, KKTEYKLVVV, KKKTEY, KKKTEYK, KKKTEYKL, KKKTEYKLV, KKKTEYKLVV, KKKTEYKLVVV, KKKKTEY, KKKKTEYK, KKKKTEYKL, KKKKTEYKLV, KKKKTEYKLVV, KKKKTEYKLVVV, IDIIMKI and further comprising the amino acid sequence of RNA, FFFFFFFFFFFFFFFFFFFFIIFFIFFWMC, FFFFFFFFFFFFFFFFFFFFFFFFFFAAFWFW, IFFIFFIIFFFFFFFFFFFFIIIIIIIWEC, FIFFFIIFFFFFFFFFFFIFIFIIIIIIFWEC, TEY, TEYK, TEYKL, TEYKLV, TEYKLVV, TEYKLVVV, WQAGILAR, HSYTTAE, PLTEEKIK, GALHFKPGSR, RRANKDATAE, KAFISHEEKR, TDLSSRFSKS, FDLGGGTFDV, CLLLHYSVSK, KKKKIIMKIRNA, or MTEYKLVVV.

[0222] In some embodiments, a polypeptide comprising a RAS epitope has, e.g., at the C-terminus, K, KK, KKK, KKKK, KKKKK, KKKKKKK, KKKKKKKK, KKNKKDDI, KKNKKDDIKD, AGNDDDDDDDDDDDDDDDDDKKDKDDDDDD, AGNKKKKKKKNNNNNNNNNNNNNNNNNNNNNN, AGRDDDDDDDDDDDDDDDDDDDDDDDDDDDDDDD, SALTI, SALTIQL, GKSALTIQL, GKSALTI, SALTIK, SALTIQLK, GKSALTIQLK, GKSALTIK, SALTIKK, SALTIQLKK, GKSALT and further comprising the amino acid sequence of IQLKK, GKSALTIKK, SALTIKKK, SALTIQLKKK, GKSALTIQLKKK, GKSALTIKKK, SALTIKKKK, SALTIQLKKKK, GKSALTIQLKKKK, GKSALTI, KKKK, QGQNLKYQ, ILGVLLLI, EKEGKISK, AASDFIFLVT, KELKQVASPF, KKKLINEKKE, KKCDISLQFF, KSTAGDTHLG, ATFYVAVTVP, LTIQLIQNHFVDEYDPTIEDSYRKQVVIDG, or TIQLIQNHFVDEYDPTIEDSYRKQVVIDGE.

[0223] In some embodiments, a polypeptide comprising a RAS epitope is selected from the group consisting of KTEYKLVVVGAVGVGKSALTIQL, KTEYKLVVVGADGVGKSALTIQL, KTEYKLVVVGARGVGKSALTIQL, KTEYKLVVVGACGVGKSALTIQL, KKTEYKLVVVGAVGVGKSALTIQL, KKTEYKLVVVGADGVGKSALTIQL, KKTEYKLVVVGARGVGKSALTIQL, KKTEYKLVVVGACGVGKSALTIQL, KKKTEYKLVVVGAVGVG KSALTIQL, KKKTEYKLVVVGADGVGKSALTIQL, KKKTEYKLVVVGARGVGKSALTIQL, KKKTEYKLVVVGACGVGKSALTIQL, KKKKTEYKLVVVGAVGVGKSALTIQL, KKKKTEY KLVVVGADGVGKSALTIQL, KKKKTEYKLVVVGARGVGKSALTIQL, KKKKTEYKLVVVGACGVGKSALTIQL, KKTEYKLVVVGAVGVGKSALTIQLKK, KKTEYKLVVVGADGVGKSAL TIQLKK, KKTEYKLVVVGARGVGKSALTIQLKK, KKTEYKLVVVGACGVGKSALTIQLKK, TEYKLVVVGAVGVGKSALTIQLK, TEYKLVVVGADGVGKSALTIQLK, TEYKLVVVGARG VGKSALTIQLK, TEYKLVVVGACGVGKSALTIQLK, TEYKLVVVGAVGVGKSALTIQLKK, TEYKLVVVGADGVGKSALTIQLKK, TEYKLVVVGARGVGKSALTIQLKK, TEYKLVVVGA CGVGKSALTIQLKK, TEYKLVVVGAVGVGKSALTIQLKKK, TEYKLVVVGADGVGKSALTIQLKKK, TEYKLVVVGARGVGKSALTIQLKKKK, TEYKLVVVGACGVGKSALTIQLKKK, TEYKLVVVGAVGVGKSALTIQLKKKK, TEYKLVVVGADGVGKSALTIQLKKKK, and TEYKLVVVGARGVGKSALTIQLKKKK, TEYKLVVVGACGVGKSALTIQLKKKK.In some embodiments, the polypeptide comprising a RAS epitope is selected from the group consisting of KKKTEYKLVVVGADGVGKSALTIQL, KKKTEYKLVVVGARGVGKSALTIQL, KKKKTEYKLVVVGAVGVGKSALTIQL, and KKKKTEYKLVVVGACGVGKSALTIQL. In some embodiments, the polypeptide comprising a RAS epitope is KKKTEYKLVVVGADGVGKSALTIQL. In some embodiments, the polypeptide comprising a RAS epitope is KKKTEYKLVVVGARGVGKSALTIQL. In some embodiments, the polypeptide comprising a RAS epitope is KKKKTEYKLVVVGAVGVGKSALTIQL. In some embodiments, the polypeptide comprising a RAS epitope is KKKKTEYKLVVVGACGVGKSALTIQL. In some embodiments, the polypeptide comprising a RAS epitope is KKKKTEYKLVVVGACGVGKSALTIQL.

[0224] In some embodiments, the peptide comprising the KRAS G12C mutation comprises the sequence MTEYKLVVVGACGVGKSALTIQLIQNHFVDEYDPTIEDSYRKQVVIDGETC LLDILDTAGQE. In some embodiments, the peptide comprising the KRAS G12C mutation comprises the neoepitope sequence of KLVVVGACGV. In some embodiments, the peptide comprising the KRAS G12C mutation comprises the neoepitope sequence of LVVVGACGV. In some embodiments, the peptide comprising the KRAS G12C mutation comprises the neoepitope sequence of VVGACGVGK. In some embodiments, the peptide comprising the KRAS G12C mutation comprises the neoepitope sequence of VVVGACGVGK.

[0225] In some embodiments, the peptide comprising the KRAS G12D mutation comprises the sequence MTEYKLVVVGADGVGKSALTIQLIQNHFVDEYDPTIEDSYRKQVVIDGETCLLDILDTAGQE. In some embodiments, the peptide comprising the KRAS G12D mutation comprises the neoepitope sequence of VVGADGVGK. In some embodiments, the peptide comprising the KRAS G12D mutation comprises the neoepitope sequence of VVVGADGVGK. In some embodiments, the peptide comprising the KRAS G12D mutation comprises the neoepitope sequence of KLVVVGADGV. In some embodiments, the peptide comprising the KRAS G12D mutation comprises the neoepitope sequence of LVVVGADGV.

[0226] In some embodiments, the peptide comprising a KRAS G12V mutation comprises the sequence MTEYKLVVVGAVGVGKSALTIQLIQNHFVDEYDPTIEDSYRKQVVIDGETCLLDILDTAGQE. In some embodiments, the peptide comprising a KRAS G12V mutation comprises the neoepitope sequence of KLVVVGAVGV. In some embodiments, the peptide comprising a KRAS G12V mutation comprises the neoepitope sequence of LVVVGAVGV. In some embodiments, the peptide comprising a KRAS G12V mutation comprises the neoepitope sequence of VVGAVGVGK. In some embodiments, the peptide comprising a KRAS G12V mutation comprises the neoepitope sequence of VVVGAVGVGK.

[0227] In some embodiments, a peptide comprising a KRAS Q61H mutation comprises the sequence: AGGVGKSALTIQLIQNHFVDEYDPTIEDSYRKQVVIDGETCLLDILDTAGHEEYSAMRDQYMRTGEGFLCVFAINNTKSFEDIHHYREQIKRVKDSEDVPM In some embodiments, a peptide comprising a KRAS Q61H mutation comprises the neoepitope sequence of ILDTAGHEEY.

[0228] In some embodiments, the peptide comprising the KRAS Q61L mutation comprises the sequence AGGVGKSALTIQLIQNHFVDEYDPTIEDSYRKQVVIDGETCLLDILDTAGLEEYSAMRDQYMRTGEGFLCVFAINNTKSFEDIHHYREQIKRVKDSEDVPM. In some embodiments, the peptide comprising the KRAS Q61L mutation comprises the neoepitope sequence of ILDTAGLEEY. In some embodiments, the peptide comprising the KRAS Q61L mutation comprises the neoepitope sequence of LLDILDTAGL.

[0229] In some embodiments, the peptide comprising the NRAS Q61K mutation comprises the sequence AGGVGKSALTIQLIQNHFVDEYDPTIEDSYRKQVVIDGETCLLDILDTAGKEEYSAMRDQYMRTGEGFLCVFAINNSKSFADINLYREQIKRVKDSDDVPM In some embodiments, the peptide comprising the NRAS Q61K mutation comprises the neoepitope sequence of ILDTAGKEEY.

[0230] In some embodiments, the peptide comprising the NRAS Q61R mutation comprises the sequence AGGVGKSALTIQLIQNHFVDEYDPTIEDSYRKQVVIDGETCLLDILDTAGREEYSAMRDQYMRTGEGFLCVFAINNSKSFADINLYREQIKRVKDSDDVPM In some embodiments, the peptide comprising the NRAS Q61R mutation comprises the neoepitope sequence of ILDTAGREEY.

[0231] In some embodiments, the peptide comprising the RAS Q61H mutation comprises the sequence TCLLDILDTAGHEEYSAMRDQYM. In some embodiments, the peptide comprising the RAS Q61H mutation comprises a sequence presented in Table 1. In some embodiments, the peptide sequence presented in Table 1 is predicted to bind to or bind to a protein encoded by an HLA allele, which allele is presented in the corresponding row next to the peptide sequence in Table 1. Table 1. Peptide sequences containing the RAS Q61H mutation, corresponding HLA alleles, and ranking of binding potential. [Table 1-1] [Table 1-2]

[0232] In some embodiments, the peptide comprising the RAS Q61R mutation comprises the sequence TCLLDILDTAGREEYSAMRDQYM. In some embodiments, the peptide comprising the RAS Q61R mutation comprises a sequence provided in Table 2. In some embodiments, the peptide sequence provided in Table 2 is predicted to bind to or bind to a protein encoded by an HLA allele, which allele is provided in the corresponding row next to the peptide sequence in Table 2. Table 2. Peptide sequences containing the RAS Q61R mutation, corresponding HLA alleles, and ranking of binding potential. [Table 2-1] [Table 2-2] [Table 2-3]

[0233] In some embodiments, the peptide comprising the RAS Q61K mutation comprises the sequence TCLLDILDTAGKEEYSAMRDQYM. In some embodiments, the peptide comprising the RAS Q61K mutation comprises a sequence presented in Table 3. In some embodiments, the peptide sequence presented in Table 3 is predicted to bind to or bind to a protein encoded by an HLA allele, which allele is presented in the corresponding row next to the peptide sequence in Table 3. Table 3. Peptide sequences containing the RAS Q61K mutation, corresponding HLA alleles, and ranking of binding potential. [Table 3-1] [Table 3-2]

[0234] In some embodiments, the peptide comprising the RAS Q61L mutation comprises the sequence TCLLDILDTAGLEEYSAMRDQYM. In some embodiments, the peptide comprising the RAS Q61L mutation comprises a sequence provided in Table 4. In some embodiments, the peptide sequence provided in Table 4 is predicted to bind to or bind to a protein encoded by an HLA allele, which allele is provided in the corresponding row next to the peptide sequence in Table 4. Table 4. Peptide sequences containing the RAS Q61L mutation, corresponding HLA alleles, and ranking of binding potential. [Table 4-1] [Table 4-2] [Table 4-3]

[0235] In some embodiments, the peptide comprising the RAS G12A mutation comprises the sequence MTEYKLVVVGAAGVGKSALTIQL. In some embodiments, the peptide comprising the RAS G12A mutation comprises a sequence provided in Table 5. In some embodiments, the peptide sequence provided in Table 5 is predicted to bind to or bind to a protein encoded by an HLA allele, which allele is provided in the corresponding row next to the peptide sequence in Table 5. Table 5. Peptide sequences containing the RAS G12A mutation, corresponding HLA alleles, and ranking of binding potential. [Table 5-1] [Table 5-2] [Table 5-3]

[0236] In some embodiments, the peptide comprising the RAS G12C mutation comprises the sequence MTEYKLVVVGACGVGKSALTIQL. In some embodiments, the peptide comprising the RAS G12C mutation comprises a sequence provided in Table 6. In some embodiments, the peptide sequence provided in Table 6 is predicted to bind to or bind to a protein encoded by an HLA allele, which allele is provided in the corresponding row next to the peptide sequence in Table 6. Table 6. Peptide sequences containing the RAS G12C mutation, corresponding HLA alleles, and ranking of binding potential. [Table 6-1] [Table 6-2]

[0237] In some embodiments, the peptide comprising the RAS G12D mutation comprises the sequence MTEYKLVVVGADGVGKSALTIQL. In some embodiments, the peptide comprising the RAS G12D mutation comprises a sequence provided in Table 7. In some embodiments, the peptide sequence provided in Table 7 is predicted to bind to or bind to a protein encoded by an HLA allele, which allele is provided in the corresponding row next to the peptide sequence in Table 7. Table 7. Peptide sequences containing the RAS G12D mutation, corresponding HLA alleles, and ranking of binding potential. [Table 7-1] [Table 7-2]

[0238] In some embodiments, the peptide comprising the RAS G12R mutation comprises the sequence MTEYKLVVVGARGVGKSALTIQL. In some embodiments, the peptide comprising the RAS G12R mutation comprises a sequence provided in Table 8. In some embodiments, the peptide sequence provided in Table 8 is predicted to bind to or bind to a protein encoded by an HLA allele, which allele is provided in the corresponding row next to the peptide sequence in Table 8. Table 8. Peptide sequences containing the RAS G12R mutation, corresponding HLA alleles, and ranking of binding potential. [Table 8-1] [Table 8-2]

[0239] In some embodiments, the peptide comprising the RAS G12S mutation comprises the sequence MTEYKLVVVGASGVGKSALTIQL. In some embodiments, the peptide comprising the RAS G12S mutation comprises a sequence provided in Table 9. In some embodiments, the peptide sequence provided in Table 9 is predicted to bind to or bind to a protein encoded by an HLA allele, which allele is provided in the corresponding row next to the peptide sequence in Table 9. Table 9. Peptide sequences containing the RAS G12S mutation, corresponding HLA alleles, and ranking of binding potential. [Table 9-1] [Table 9-2]

[0240] In some embodiments, the peptide comprising the RAS G12V mutation comprises the sequence MTEYKLVVVGAVGVGKSALTIQL. In some embodiments, the peptide comprising the RAS G12V mutation comprises a sequence provided in Table 10. In some embodiments, the peptide sequence provided in Table 10 is predicted to bind to or bind to a protein encoded by an HLA allele, which allele is provided in the corresponding row adjacent to the peptide sequence in Table 10. Table 10. Peptide sequences containing the RAS G12V mutation, corresponding HLA alleles, and binding potential Presence ranking [Table 10-1] [Table 10-2] [Table 10-3]

[0241] In some embodiments, the peptide comprising the RAS G13C mutation comprises the sequence MTEYKLVVVGAGCVGKSALTIQL. In some embodiments, the peptide comprising the RAS G13C mutation comprises a sequence provided in Table 11. In some embodiments, the peptide sequence provided in Table 11 is predicted to bind to or bind to a protein encoded by an HLA allele, which allele is provided in the corresponding row next to the peptide sequence in Table 11. Table 11. Peptide sequences containing the RAS G13C mutation, corresponding HLA alleles, and binding potential Presence ranking [Table 11]

[0242] In some embodiments, the peptide comprising the RAS G13D mutation comprises the sequence MTEYKLVVVGAGDVGKSALTIQL. In some embodiments, the peptide comprising the RAS G13D mutation comprises a sequence provided in Table 12. In some embodiments, the peptide sequence provided in Table 12 is predicted to bind to or bind to a protein encoded by an HLA allele, which allele is provided in the corresponding row next to the peptide sequence in Table 12. Table 12. Peptide sequences containing the RAS G13D mutation, corresponding HLA alleles, and binding potential Presence ranking [Table 12-1] [Table 12-2]

[0243] In some embodiments, the polypeptides described herein do not include a RAS epitope. In some embodiments, the epitope is not a RAS epitope. In some embodiments, the polypeptides do not include KKKKKPKRDGYMFLKAESKIMFAT, KKKKYMFLKAESKIMFATLQRSS, KKKKKAESKIMFATLQRSSLWCL, KKKKIMFATLQRSSLWCLCSNH, or KKKKMFATLQRSSLWCLCSNH.

[0244] In some embodiments, the antigen, neoantigen peptide, or epitope-containing polypeptide comprises a GATA3 epitope, in some embodiments, the GATA3 epitope comprises the amino acid sequence of MLTGPPARV, SMLTGPPARV, VLPEPHLAL, KPKRDGYMF, KPKRDGYMFL, ESKIMFATL, KRDGYMFL, PAVPFDLHF, AESKIMAFATL, FATLQRSSL, ARVPAVPFD, IMKPKRDGY, DGYMFLKA, MFLKAESKIMF, LTGPPARV, ARVPAVPF, SMLTGPPAR, RVPAVPFDL, or LTGPPARVP. Peptide Modification

[0245] In some embodiments, the present disclosure includes modified peptides. Modifications can include covalent chemical modifications that do not change the primary amino acid sequence of the antigenic peptide itself. Modifications can result in peptides with desired properties, such as increased in vivo half-life, increased stability, reduced clearance, altered immunogenicity or allergenicity, enabling the generation of specific antibodies, cell targeting, antigen uptake, antigen processing, HLA affinity, HLA stability, or antigen presentation. In some embodiments, peptides can include one or more sequences that enhance the processing and presentation of epitopes by APCs, for example, to generate an immune response.

[0246] In some embodiments, the polypeptides can be modified to provide desired properties. For example, the ability of the peptide to induce cytotoxic T lymphocyte (CTL) activity can be enhanced by linking it to a sequence containing at least one epitope capable of inducing a helper T cell response. In some embodiments, the immunogenic peptide / helper T conjugate is linked by a spacer molecule. In some embodiments, the spacer comprises a relatively small neutral molecule, such as an amino acid or an amino acid mimetic, that is substantially uncharged under physiological conditions. The spacer can be selected, for example, from neutral spacers of Ala, Gly or other nonpolar amino acids or neutral polar amino acids. It will be understood that the optional spacer need not be comprised of the same residue and thus can be a hetero- or homo-oligomer. The neo-antigenic peptide can be linked to the helper T peptide either directly or via a spacer at either the amino or carboxy terminus of the peptide. The amino terminus of either the neo-antigenic peptide or the helper T peptide can be acylated. Examples of helper T peptides include tetanus toxoid residues 830-843, influenza residues 307-319, and malaria circumsporozoite residues 382-398 and residues 378-389.

[0247] The peptide sequences of the present disclosure can be altered, if desired, by making changes at the DNA level, in particular by mutating the DNA encoding the peptide at preselected bases so as to generate a codon that translates into the desired amino acid.

[0248] In some embodiments, the peptides described herein may contain substitutions to alter the physical properties (e.g., stability or solubility) of the resulting peptide. For example, the peptides can be modified by substituting cysteine ​​(C) with α-aminobutyric acid ("B"). Due to its chemical nature, cysteine ​​has a tendency to form disulfide bridges, structurally altering the peptide sufficiently to reduce its binding capacity. Substituting C with α-aminobutyric acid not only alleviates this problem, but in certain cases actually improves the binding and cross-linking capacity. Substitution of cysteine ​​with α-aminobutyric acid can be made at any residue of the neoantigenic peptide, for example at anchor or non-anchor positions of the epitope or analog within the peptide, or at other positions of the peptide.

[0249] Peptides can also be modified by, for example, adding or deleting amino acids to lengthen or shorten the amino acid sequence of the compound. Peptides or analogs can also be modified by changing the order or composition of certain residues. Those skilled in the art will understand that certain amino acid residues essential for biological activity, such as residues at critical contact sites or conserved residues, generally cannot be altered without adverse effects on biological activity. Non-critical amino acids need not be limited to amino acids naturally occurring in proteins, such as L-α-amino acids, or their D-isomers, but may include non-natural amino acids, such as β-γ-δ-amino acids, as well as many derivatives of L-α-amino acids.

[0250] In some embodiments, peptides can be modified using a series of peptides with single amino acid substitutions to determine the effect of electrostatic charge, hydrophobicity, etc. on HLA binding. For example, a series of positively charged amino acids (e.g., Lys or Arg) or negatively charged amino acids (e.g., Glu) can be substituted along the length of the peptide, thereby revealing different patterns of sensitivity to various HLA molecules and T cell receptors. In addition, multiple substitutions using small relatively neutral moieties such as Ala, Gly, Pro, or similar residues can be used. Substitutions can be homo- or hetero-oligomers. The number and type of residues substituted or added depend on the required spacing between essential contact points and on certain functional traits (e.g., hydrophobicity versus hydrophilicity) that are sought. Such substitutions can also achieve increased binding affinity to HLA molecules or T cell receptors compared to the affinity of the parent peptide. In any case, such substitutions should use amino acid residues or other molecular fragments selected to avoid, for example, steric and charge interferences that may interfere with binding. Amino acid substitutions are generally of single residues. Substitutions, deletions, insertions, or any combination thereof can be combined to arrive at the final peptide.

[0251] In some embodiments, the peptides described herein contain amino acid mimetics or unnatural amino acid residues, such as D- or L-naphthylalanine; D- or L-phenylglycine; D- or L-2-thienylalanine; D- or L-1, -2, 3-, or 4-pyrenylalanine; D- or L-3-thienylalanine; D- or L-(2-pyridinyl)-alanine; D- or L-(3-pyridinyl)-alanine; D- or L-(2-pyrazinyl)-alanine; D- or L-(4-isopropyl)-phenylglycine; D-(trifluoromethyl)-phenylglycine; D-(trifluoromethyl)-phenylalanine; Dp-fluorophenylalanine; D- or Lp-biphenyl-phenylalanine; D- or Lp-methoxybiphenylphenylalanine; D- or L-2-indole(allyl)alanine; and D- or L-alkylalanine, where the alkyl group can be substituted or unsubstituted methyl, ethyl, propyl, hexyl, butyl, pentyl, isopropyl, isobutyl, sec-isotyl, isopentyl, or non-acidic amino acid residue. Aromatic rings of unnatural amino acids include, for example, thiazolyl, thiophenyl, pyrazolyl, benzimidazolyl, naphthyl, furanyl, pyrrolyl, and pyridyl aromatic rings. Modified peptides with various amino acid mimics or unnatural amino acid residues can have increased stability in vivo. Such peptides can also have improved shelf life or manufacturing properties.

[0252] In some embodiments, the peptides described herein are provided with a terminal NH 2 By acylation, for example, alkanoyl (C 1 ~C 20) or thioglycolyl acetylation, terminal carboxylamidation, e.g., ammonia, methylamine, etc. In some embodiments, these modifications can provide sites for linkage to supports or other molecules. In some embodiments, the peptides described herein can contain modifications such as, but not limited to, glycosylation, oxidation of side chains, biotinylation, phosphorylation, addition of surface active materials, e.g., lipids, or can be chemically modified, e.g., acetylated. Additionally, bonds within a peptide can be bonds other than peptide bonds, e.g., covalent bonds, ester or ether bonds, disulfide bonds, hydrogen bonds, ionic bonds, etc.

[0253] In some embodiments, the peptides described herein may include a carrier, such as those known in the art, e.g., thyroglobulin, albumin such as human serum albumin, tetanus toxoid, polyamino acid residues such as poly-L-lysine and poly-L-glutamic acid, influenza virus proteins, hepatitis B virus core protein, and the like.

[0254] Peptides can be further modified to contain additional chemical moieties that are not normally part of a protein. These derivatized moieties can improve solubility, biological half-life, protein absorption, or binding affinity. The moieties can also reduce or eliminate any desirable side effects of the peptide, etc. A summary of these moieties can be found in Remington's Pharmaceutical Sciences, 20th ed., Mack Publishing Co., Easton, PA (2000). For example, a peptide having a desired activity can be modified to contain additional chemical moieties that are not normally part of a protein. Neo-antigenic peptides can be modified, if necessary, to provide certain desired properties, such as improved pharmacological characteristics, while increasing or at least retaining substantially all of the biological activity of the unmodified peptide to bind to desired HLA molecules and activate appropriate T cells. For example, peptides can be subjected to various changes, such as either conservative or non-conservative substitutions, which can provide certain advantages for their use, such as improved HLA binding. Such conservative substitutions can include replacing an amino acid residue with another that is biologically and / or chemically similar, such as replacing one hydrophobic residue with another, or replacing one polar residue with another. The effect of single amino acid substitutions can also be probed using D-amino acids. Such modifications are described, for example, in Merrifield, Science 232: 341-347 (1986), Barany & Merrifield, The Peptides, Gross & Meienhofer, eds. (NY, Academic Press), pp. 1-284 (1979); and Stewart & Young, Solid Phase Peptide Synthesis, (Rockford, III., Pierce), 2d Ed. (1984), using well-known peptide synthesis procedures. can be done.

[0255] In some embodiments, the peptides described herein may be conjugated to large, slowly metabolized macromolecules such as proteins; polysaccharides, e.g., sepharose, agarose, cellulose, cellulose beads, and the like; polymeric amino acids, e.g., polyglutamic acid, polylysine, and the like; amino acid copolymers; inactivated virus particles; inactivated bacterial toxins, e.g., toxoids derived from diphtheria, tetanus, cholera, leukotoxin molecules, and the like; inactivated bacteria; and dendritic cells.

[0256] Modifications to the peptide can include, but are not limited to, conjugation to a carrier protein, conjugation to a ligand, conjugation to an antibody, PEGylation, polysialylation, HESylation, recombinant PEG mimetics, Fc fusion, albumin fusion, nanoparticle attachment, nanoparticle encapsulation, cholesterol fusion, iron fusion, acylation, amidation, glycosylation, side chain oxidation, phosphorylation, biotinylation, addition of surface active materials, addition of amino acid mimetics, or addition of unnatural amino acids.

[0257] Glycosylation can affect the physical properties of proteins and can also be important for protein stability, secretion, and subcellular localization. Proper glycosylation can be important for biological activity. Indeed, some genes from eukaryotes, when expressed in bacteria (e.g., E. coli), which lack the cellular processes for glycosylation of proteins, result in proteins that are recovered with little or no activity due to lack of glycosylation. Addition of glycosylation sites can be achieved by altering the amino acid sequence. Modifications to the peptide or protein can be made, for example, by adding or substituting one or more serine or threonine residues (for O-linked glycosylation sites) or asparagine residues (for N-linked glycosylation sites). The structures of N-linked and O-linked oligosaccharides and the sugar residues found in each type can differ. One type of sugar commonly found in both is N-acetylneuraminic acid (hereinafter referred to as sialic acid). Sialic acid is usually the terminal residue of both N-linked and O-linked oligosaccharides and can confer acidic properties to glycoproteins due to a negative charge. Several embodiments of the present disclosure include the generation and use of N-glycosylation variants. Carbohydrate removal can be achieved chemically or enzymatically, or by substitution of the codons encoding the amino acid residues to be glycosylated. Chemical deglycosylation techniques are known, and enzymatic cleavage of carbohydrate moieties of polypeptides can be achieved by using a variety of endo- and exoglycosidases.

[0258] Additional suitable moieties and molecules for conjugation include, for example, molecules for targeting the lymphatic system, thyroglobulin; albumins such as human serum albumin (HAS); tetanus toxoid; diphtheria toxoid; polyamino acids such as poly(D-lysine:D-glutamic acid); VP6 polypeptide of rotavirus; influenza virus hemagglutinin, influenza virus nucleoprotein; keyhole limpet hemocyanin (KLH); and hepatitis B virus core protein and surface antigen; or any combination of the foregoing.

[0259] Another type of modification is to conjugate (e.g., link) one or more additional components or molecules, such as another protein (e.g., a protein having an amino acid sequence heterologous to the protein of interest) or a carrier molecule, to the N-terminus and / or C-terminus of the polypeptide sequence. Thus, an exemplary polypeptide sequence can be provided as a conjugate with another component or molecule. In some embodiments, fusion of albumin with a peptide or protein of the present disclosure can be achieved by genetic engineering, for example, joining DNA encoding HSA or a fragment thereof with DNA encoding one or more polypeptide sequences. A suitable host can then be transformed or transfected with the fused nucleotide sequence, for example in the form of a suitable plasmid, so that the fusion polypeptide is expressed. Expression can be performed in vitro, for example in prokaryotic or eukaryotic cells, or in vivo, for example in transgenic organisms. In some embodiments of the present disclosure, expression of the fusion protein is performed in a mammalian cell line, for example a CHO cell line. Furthermore, albumin itself can be modified to increase its circulating half-life. The fusion of modified albumin with one or more polypeptides can be achieved by the above-mentioned genetic engineering techniques or by chemical conjugation; the resulting fusion molecule has a half-life that exceeds that of fusion with unmodified albumin (see, for example, WO2011 / 051489).As an alternative to direct fusion, several albumin binding strategies have been developed, including albumin binding through conjugated fatty acid chains (acylation).Since serum albumin is a transport protein for fatty acids, these natural ligands with albumin binding activity are used to extend the half-life of small protein therapeutics.

[0260] Additional candidate components and molecules for conjugation include those suitable for isolation or purification.Non-limiting examples include molecules that include binding molecules such as biotin (biotin-avidin specific binding pair), antibodies, receptors, ligands, lectins, or solid supports, including, for example, plastic or polystyrene beads, plates or beads, magnetic beads, test strips, and membranes.Purification methods such as cation exchange chromatography can be used to separate conjugates by charge difference, which effectively separates conjugates into their various molecular weights.The contents of the fractions obtained by cation exchange chromatography can be identified by molecular weight using conventional methods, such as mass spectrometry, SDS-PAGE, or other known methods for separating molecular entities by molecular weight.

[0261] In some embodiments, the amino or carboxyl terminus of the peptide or protein sequence of the present disclosure can be fused with an immunoglobulin Fc region (e.g., human Fc) to form a fusion conjugate (or fusion molecule). Fc fusion conjugates have been shown to extend the systemic half-life of biologics, thus allowing less frequent administration of biologic products. Fc binds to neonatal Fc receptors (FcRn) in endothelial cells lining blood vessels, and upon binding, the Fc fusion molecule is protected from degradation and re-released into circulation, maintaining the molecule in circulation longer. This Fc binding is believed to be the mechanism by which endogenous IgG maintains its long plasma half-life. More recent Fc fusion technology has linked a single copy of the biologic to the Fc region of an antibody to optimize the pharmacokinetic and pharmacodynamic properties of the biologic compared to traditional Fc fusion conjugates.

[0262] The present disclosure contemplates the use of other currently known or developing modifications of peptides to improve one or more properties.One such method for extending the circulating half-life, increasing stability, reducing clearance, or modifying immunogenicity or allergenicity of the peptides of the present disclosure involves modifying peptide sequence by hesylation, which utilizes hydroxyethyl starch derivatives linked with other molecules to modify molecular characteristics.Various aspects of hesylation are described, for example, in US Patent Application Nos. 2007 / 0134197 and 2006 / 0258607.

[0263] The stability of peptides can be assayed in several ways. For example, peptidases and various biological media such as human plasma and serum have been used to test stability. See, for example, Verhoef, et al., Eur. J. Drug Metab. Pharmacokinetics 11: 291 (1986). The half-life of the peptides described herein is conveniently determined using a 25% human serum (v / v) assay. The protocol is as follows: pooled human serum (type AB, not heat-inactivated) is broken by centrifugation before use. The serum is then diluted to 25% with RPMI-1640 or another suitable tissue culture medium. At predetermined time intervals, small amounts of the reaction solution are removed and added to either 6% aqueous trichloroacetic acid (TCA) or ethanol. The cloudy reaction sample is cooled (4°C) for 15 minutes and then spun at high speed to pellet precipitated serum proteins. The presence of the peptide is then determined by reverse-phase HPLC using stability-specific chromatographic conditions.

[0264] Problems associated with short plasma half-life or susceptibility to protease degradation can be overcome by various modifications, including conjugating or linking the peptide or protein sequence to any of a variety of non-proteinaceous polymers, such as polyethylene glycol (PEG), polypropylene glycol, or polyoxyalkylene (e.g., typically via a linking moiety covalently attached to both the protein and the non-proteinaceous polymer, such as PEG). Such PEG-conjugated biomolecules have been shown to have clinically useful properties, including better physical and thermal stability, protection from susceptibility to enzymatic degradation, increased solubility, longer in vivo circulatory half-life and reduced clearance, reduced immunogenicity and antigenicity, and reduced toxicity.

[0265] PEGs suitable for conjugation to polypeptide or protein sequences are generally soluble in water at room temperature and have the general formula R-(O-CH 2 -CH 2 ) n-OR, where R is hydrogen or a protecting group such as an alkyl or alkanol group, and n is an integer between 1 and 1000. When R is a protecting group, it generally has 1 to 8 carbons. The PEG conjugated to the polypeptide sequence may be linear or branched. Branched PEG derivatives, "star PEGs" and multi-arm PEGs are contemplated by the present disclosure. The present disclosure also contemplates compositions of conjugates in which the PEG has different n values, and thus various different PEGs are present in specific ratios. For example, some compositions include a mixture of conjugates where n=1, 2, 3, and 4. In some compositions, the percentage of conjugates where n=1 is 18-25%, the percentage of conjugates where n=2 is 50-66%, the percentage of conjugates where n=3 is 12-16%, and the percentage of conjugates where n=4 is up to 5%. Such compositions can be made by reaction conditions and purification methods known in the art. For example, cation exchange chromatography can be used to separate the conjugates, and then fractions containing, for example, the conjugate with the desired number of PEG attached are identified and purified to be free of unmodified protein sequences and other numbers of PEG attached conjugates.

[0266] PEG can be attached to the peptide or protein of the present disclosure via a terminal reactive group ("spacer"). The spacer is, for example, a terminal reactive group that mediates the attachment of PEG to one or more free amino or carboxyl groups of a polypeptide sequence. PEG with a spacer that can be attached to a free amino group includes N-hydroxysuccinimide PEG, which can be prepared by activating the succinic acid ester of PEG with N-hydroxysuccinimide. Another activated PEG that can be attached to a free amino group is 2,4-bis(O-methoxypolyethylene glycol)-6-chloro-s-triazine, which can be prepared by reacting PEG monomethyl ether with cyanuric chloride. Activated PEG that is attached to a free carboxyl group includes polyoxyethylenediamine.

[0267] Conjugation of PEG with a spacer to one or more of the peptide or protein sequences of the present disclosure can be carried out by a variety of conventional methods. For example, the conjugation reaction can be carried out in solution at a pH of 5 to 10, at a temperature of 4° C. to room temperature, for 30 minutes to 20 hours, utilizing a molar ratio of reagent to peptide / protein of 4:1 to 30:1. Reaction conditions can be selected to direct the reaction to predominantly produce the desired degree of substitution. In general, low temperatures, low pH (e.g., pH=5), and short reaction times tend to decrease the number of PEGs attached, while high temperatures, neutral to high pH (e.g., pH>7), and longer reaction times tend to increase the number of PEGs attached. The reaction can be terminated using a variety of means known in the art. In some embodiments, the reaction is terminated by acidifying the reaction mixture and freezing, for example, at −20° C. Neoepitope

[0268] Neoepitopes include neoantigenic determinants of neoantigenic peptides or polypeptides that are recognized by the immune system. Neoepitopes refer to epitopes that are absent in reference non-disease cells, e.g., non-cancerous cells or germline cells, but are found in diseased cells, e.g., cancer cells. This includes situations where the corresponding epitope is found in normal non-disease cells or germline cells, but due to one or more mutations in diseased cells, e.g., cancer cells, the sequence of the epitope has changed, resulting in a neoepitope. The term "neoepitope" is used interchangeably herein with "tumor-specific epitope" or "tumor-specific neoepitope" and refers to a stretch of one residue, typically an L-amino acid, connected by a peptide bond to another residue, typically an L-amino acid, typically the α-amino and carboxyl group of the adjacent amino acid. Neoepitopes may be of various lengths, may be in neutral (uncharged) or salt form, may be free of or contain modifications such as glycosylation, side chain oxidation, or phosphorylation, and may be subject to conditions in which the modifications do not impair the biological activity of the polypeptides described herein. The present disclosure provides isolated neoepitopes that include the tumor-specific mutations of Tables 1-12.

[0269] In some embodiments, the neoepitopes described herein for MHC class I HLA are 12 amino acid residues or less in length, typically consisting of between about 8 and about 12 amino acid residues. In some embodiments, the neoepitopes described herein for MHC class I HLA are about 8, about 9, about 10, about 11, or about 12 amino acid residues. In some embodiments, the neoepitopes described herein for MHC class II HLA are 25 amino acid residues or less in length, typically consisting of between about 9 and about 25 amino acid residues. In some embodiments, the neoepitopes described herein for MHC class II HLA are about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, or about 25 amino acid residues.

[0270] In some embodiments, the compositions described herein comprise a first peptide comprising a first neoepitope of a protein and a second peptide comprising a second neoepitope of the same protein, where the first peptide is different from the second peptide, and where the first neoepitope comprises a mutation and the second neoepitope comprises the same mutation. In some embodiments, the compositions described herein comprise a first peptide comprising a first neoepitope of a first region of a protein and a second peptide comprising a second neoepitope of a second region of the same protein, where the first region comprises at least one amino acid of the second region, where the first peptide is different from the second peptide, and where the first neoepitope comprises a first mutation and the second neoepitope comprises a second mutation. In some embodiments, the first mutation and the second mutation are the same. In some embodiments, the mutation is selected from the group consisting of a point mutation, a splice site mutation, a frameshift mutation, a read-through mutation, a gene fusion mutation, and any combination thereof.

[0271] In some embodiments, the first neoepitope binds to a class I HLA protein to form a class I HLA-peptide complex. In some embodiments, the second neoepitope binds to a class II HLA protein to form a class II HLA-peptide complex. In some embodiments, the second neoepitope binds to a class I HLA protein to form a class I HLA-peptide complex. In some embodiments, the first neoepitope binds to a class II HLA protein to form a class II HLA-peptide complex. In some embodiments, the first neoepitope binds to a CD8 + Activates T cells. In some embodiments, the first neoepitope is a CD4 + In some embodiments, the second neoepitope activates CD4 T cells. + In some embodiments, the second neoepitope activates CD8 T cells. + In some embodiments, CD4 +The TCR of the T cell binds to a class II HLA-peptide complex. In some embodiments, the CD8 + The TCR of the T cell binds to a class II HLA-peptide complex. In some embodiments, the CD8 + The TCR of the T cell binds to a class I HLA-peptide complex. In some embodiments, CD4 + The TCR of the T cell binds to the class I HLA-peptide complex.

[0272] In some embodiments, the second neoepitope is longer than the first neoepitope. In some embodiments, the first neoepitope has a length of at least 8 amino acids. In some embodiments, the first neoepitope has a length of 8 to 12 amino acids. In some embodiments, the first neoepitope comprises a sequence of at least 8 consecutive amino acids, where at least one of the 8 consecutive amino acids differs from the corresponding position of the wild-type sequence. In some embodiments, the first neoepitope comprises a sequence of at least 8 consecutive amino acids, where at least two of the 8 consecutive amino acids differ from the corresponding position of the wild-type sequence. In some embodiments, the second neoepitope has a length of at least 16 amino acids. In some embodiments, the second neoepitope has a length of 16 to 25 amino acids. In some embodiments, the second neoepitope comprises a sequence of at least 16 consecutive amino acids, where at least one of the 16 consecutive amino acids differs from the corresponding position of the wild-type sequence. In some embodiments, the second neoepitope comprises a sequence of at least 16 contiguous amino acids, where at least two of the 16 contiguous amino acids differ from the corresponding position in the wild-type sequence.

[0273] In some embodiments, the neoepitope comprises at least one anchor residue. In some embodiments, the first neoepitope, the second neoepitope, or both comprise at least one anchor residue. In one embodiment, at least one anchor residue of the first neoepitope is at a canonical anchor position or a non-canonical anchor position. In another embodiment, at least one anchor residue of the second neoepitope is at a canonical anchor position or a non-canonical anchor position. In yet another embodiment, at least one anchor residue of the first neoepitope is different from at least one anchor residue of the second neoepitope.

[0274] In some embodiments, at least one anchor residue is a wild type residue. In some embodiments, at least one anchor residue is a substitution. In some embodiments, at least one anchor residue does not include a mutation.

[0275] In some embodiments, the second neoepitope or both comprise at least one anchor residue adjacent region. In some embodiments, the neoepitope comprises at least one anchor residue. In some embodiments, the at least one anchor residue comprises at least two anchor residues. In some embodiments, the at least two anchor residues are separated by a separation region comprising at least one amino acid. In some embodiments, the at least one anchor residue adjacent region is not within the separation region. In some embodiments, the at least one anchor residue adjacent region is (a) upstream of the N-terminal anchor residue of the at least two anchor residues; (b) downstream of the C-terminal anchor residue of the at least two anchor residues; or both (a) and (b).

[0276] In some embodiments, the neoepitope is associated with an HLA protein (e.g., MHC class I In some embodiments, the neoepitope binds to an HLA protein with greater affinity than the corresponding wild-type peptide. In some embodiments, the neoepitope binds to an HLA protein with an IC of less than 5,000 nM, less than 1,000 nM, less than 500 nM, less than 100 nM, less than 50 nM, or less than or equal to 5,000 nM. 50 In some embodiments, the neoepitope may have an HLA binding affinity of between about 1 pM and about 1 mM, between about 100 pM and about 500 μM, between about 500 pM and about 10 μM, between about 1 nM and about 1 μM, or between about 10 nM and about 1 μM. In some embodiments, the neoepitope may have an HLA binding affinity of at least 2 nM, 3 nM, 4 nM, 5 nM, 6 nM, 7 nM, 8 nM, 9 nM, 10 nM, 15 nM, 20 nM, 25 nM, 30 nM, 35 nM, 40 nM, 45 nM, 50 nM, 55 nM, 60 nM, 65 nM, 70 nM, 75 nM, 80 nM, 85 nM, 90 nM, It may have an HLA binding affinity of 95nM, 100nM, 150nM, 200nM, 250nM, 300nM, 350nM, 400nM, 450nM, 500nM, 550nM, 600nM, 700nM, 800nM, 900nM, 1,000nM, 1,500nM, or 2,000nM or greater. In some embodiments, the neoepitope is at most 2 nM, 3 nM, 4 nM, 5 nM, 6 nM, 7 nM, 8 nM, 9 nM, 10 nM, 15 nM, 20 nM, 25 nM, 30 nM, 35 nM, 40 nM, 45 nM, 50 nM, 55 nM, 60 nM, 65 nM, 70 nM, 75 nM, 80 nM, 85 nM, It may have an HLA binding affinity of 90nM, 95nM, 100nM, 150nM, 200nM, 250nM, 300nM, 350nM, 400nM, 450nM, 500nM, 550nM, 600nM, 700nM, 800nM, 900nM, 1,000nM, 1,500nM, or 2,000nM.

[0277] In some embodiments, the first neoepitope and / or the second neoepitope binds to the HLA protein with greater affinity than a wild-type neoepitope corresponding to the HLA protein, in some embodiments, the first neoepitope and / or the second neoepitope binds to the HLA protein with a K of less than 1,000 nM, less than 900 nM, less than 800 nM, less than 700 nM, less than 600 nM, less than 500 nM, less than 250 nM, less than 150 nM, less than 100 nM, less than 50 nM, less than 25 nM, or less than 10 nM. D or IC 50 In some embodiments, the first neoepitope and / or the second neoepitope binds to an HLA class I protein with a K of less than 1,000 nM, less than 900 nM, less than 800 nM, less than 700 nM, less than 600 nM, less than 500 nM, less than 250 nM, less than 150 nM, less than 100 nM, less than 50 nM, less than 25 nM, or less than 10 nM. D or IC 50 In some embodiments, the first neoepitope and / or the second neoepitope binds to an HLA class II protein with a K of less than 2,000 nM, less than 1,500 nM, less than 1,000 nM, less than 900 nM, less than 800 nM, less than 700 nM, less than 600 nM, less than 500 nM, less than 250 nM, less than 150 nM, less than 100 nM, less than 50 nM, less than 25 nM, or less than 10 nM. D or IC 50 Combine with.

[0278] In some embodiments, the neoepitope binds to MHC class I HLA. In some embodiments, the neoepitope binds to MHC class I HLA with an affinity of 0.1 nM to 2000 nM. In some embodiments, the neoepitope binds to MHC class I HLA with an affinity of 0.1 nM to 2000 nM. 0.1nM, 0.2nM, 0.3nM, 0.4nM, 0.5nM, 0.6nM, 0.7nM, 0.8nM, 0.9nM, 1nM, 2nM, 3nM, 4nM, 5nM, 6nM, 7nM, 8nM, 9nM, 1 for HLA 0nM, 15nM, 20nM, 25nM, 30nM, 35nM, 40nM, 45nM, 50nM, 55nM, 60nM, 65nM, 70nM, 75nM, 80nM, 85nM, 90nM, 95nM, 100nM, In some embodiments, the neoepitope binds to an MHC class II HLA with an affinity of 150nM, 200nM, 250nM, 300nM, 350nM, 400nM, 450nM, 500nM, 550nM, 600nM, 650nM, 700nM, 750nM, 800nM, 850nM, 900nM, 950nM, 1000nM, 1100nM, 1200nM, 1300nM, 1400nM, 1500nM, 1600nM, 1700nM, 1800nM, 1900nM, or 2000nM. In some embodiments, the neoepitope binds to an MHC class II HLA with an affinity of between 0.1 nM and 2000 nM, between 1 nM and 1000 nM, between 10 nM and 500 nM, or less than 1000 nM.In some embodiments, the neoepitope is present in an MHC class II HLA at 0.1 nM, 0.2 nM, 0.3 nM, 0.4 nM, 0.5 nM, 0.6 nM, 0.7 nM, 0.8 nM, 0.9 nM, 1 nM, 2 nM, 3 nM, 4 nM, 5 nM, 6 nM, 7 nM, 8 nM, 9 nM, 10 nM, 15 nM, 20 nM, 25 nM, 30 nM, 35 nM, 40 nM, 45 nM, 50 nM, 55 nM, 60 nM, 65 nM, 70 nM, 75 nM, 80 nM, 85 nM, 90 nM, 95 nM, 100 nM, binds with an affinity of 150nM, 200nM, 250nM, 300nM, 350nM, 400nM, 450nM, 500nM, 550nM, 600nM, 650nM, 700nM, 750nM, 800nM, 850nM, 900nM, 950nM, 1000nM, 1100nM, 1200nM, 1300nM, 1400nM, 1500nM, 1600nM, 1700nM, 1800nM, 1900nM, or 2000nM.

[0279] In some embodiments, the neoepitope binds to MHC class I HLA with a stability between 10 minutes and 24 hours. In some embodiments, the neoepitope binds to MHC class I HLA with a stability of 10 minutes, 11 minutes, 12 minutes, 13 minutes, 14 minutes, 15 minutes, 16 minutes, 17 minutes, 18 minutes, 19 minutes, 20 minutes, 25 minutes, 30 minutes, 35 minutes, 40 minutes, 45 minutes, 50 minutes, 55 minutes, or 60 minutes. In some embodiments, the neoepitope binds to MHC class I HLA with a stability of 1 hour, 1.5 hours, 2 hours, 2.5 hours, 3 hours, 3.5 hours, 4 hours, 4.5 hours, 5 hours, 5.5 hours, 6 hours, 6.5 hours, 7 hours, 7.5 hours, 8 hours, 8.5 hours, 9 hours, 9.5 hours, 10 hours, 11 hours, 12 hours, 13 hours, 14 hours, 15 hours, 16 hours, 17 hours, 18 hours, 19 hours, 20 hours, 21 hours, 22 hours, 23 hours, or 24 hours. In some embodiments, the neoepitope binds to MHC class II HLA with a stability of 10 minutes to 24 hours. In some embodiments, the neoepitope binds to an MHC class II HLA with stability for 10 minutes, 11 minutes, 12 minutes, 13 minutes, 14 minutes, 15 minutes, 16 minutes, 17 minutes, 18 minutes, 19 minutes, 20 minutes, 25 minutes, 30 minutes, 35 minutes, 40 minutes, 45 minutes, 50 minutes, 55 minutes, or 60 minutes. In some embodiments, the neoepitope binds to an MHC class II HLA with a stability of 1 hour, 1.5 hours, 2 hours, 2.5 hours, 3 hours, 3.5 hours, 4 hours, 4.5 hours, 5 hours, 5.5 hours, 6 hours, 6.5 hours, 7 hours, 7.5 hours, 8 hours, 8.5 hours, 9 hours, 9.5 hours, 10 hours, 11 hours, 12 hours, 13 hours, 14 hours, 15 hours, 16 hours, 17 hours, 18 hours, 19 hours, 20 hours, 21 hours, 22 hours, 23 hours, or 24 hours.

[0280] In some embodiments, the first neoepitope and / or the second neoepitope binds to a protein encoded by an HLA allele expressed by the subject. In another embodiment, the mutation is not present in the non-cancer cells of the subject. In yet another embodiment, the first neoepitope and / or the second neoepitope is a gene encoded or expressed by a gene in the cancer cells of the subject. In some embodiments, the first neoepitope comprises a mutation listed in column 1 of Tables 1-12. In some embodiments, the second neoepitope comprises a mutation listed in column 1 of Tables 1-12. For example, the first neoepitope and the second neoepitope can comprise the sequence ALNSEALSVV. For example, the first neoepitope and the second neoepitope can comprise the sequence MALNSEALSV.

[0281] In some embodiments, the first neoepitope and the second neoepitope are derived from KRAS protein. In some embodiments, the first neoepitope and the second neoepitope are derived from NRAS protein. In some embodiments, the first neoepitope and the second neoepitope are derived from KRAS protein comprising G12C, G12D, G12V, Q61H, or Q61L substitution mutation. In some embodiments, the first neoepitope and the second neoepitope are derived from NRAS protein comprising Q61K or Q61R substitution mutation. In some embodiments, the neoepitope comprises substitution mutation, for example, KRAS G12C, G12D, G12V, Q61H, or Q61L mutation, or NRAS Q61K or Q61R mutation. In some embodiments, the first neoepitope and the second neoepitope are derived from the KRAS or NRAS protein sequence of MTEYKLVVVGACGVGKSALTIQLIQNHFVDEYDPTIEDSYRKQVVIDGETCLLDILDTAGQE. For example, the first neoepitope and the second neoepitope can include the sequence KLVVVGACGV. For example, the first neoepitope and the second neoepitope can include the sequence LVVVGACGV. For example, the first neoepitope and the second neoepitope can include the sequence VVGACGVGK. For example, the first neoepitope and the second neoepitope can include the sequence VVVGACGVGK. In some embodiments, the first neoepitope and the second neoepitope are derived from the KRAS or NRAS protein sequence of MTEYKLVVVGADGVGKSALTIQLIQNHFVDEYDPTIEDSYRKQVVIDGETCLLDILDTAGQEVVGADGVGK. For example, the first neoepitope and the second neoepitope can include the sequence VVVGADGVGK. For example, the first neoepitope and the second neoepitope can include the sequence KLVVVGADGV. For example, the first neoepitope and the second neoepitope can include the sequence LVVVGADGV.

[0282] In some embodiments, the first neoepitope and the second neoepitope are derived from the KRAS or NRAS protein sequence of MTEYKLVVVGAVGVGKSALTIQLIQNHFVDEYDPTIEDSYRKQVVIDGETCLLDILDTAGQE. For example, the first neoepitope and the second neoepitope can include the sequence KLVVVGAVGV. For example, the first neoepitope and the second neoepitope can include the sequence LVVVGAVGV. For example, the first neoepitope and the second neoepitope can include the sequence VVGAVGVGK. For example, the first neoepitope and the second neoepitope can include the sequence VVVGAVGVGK.

[0283] In some embodiments, the first neoepitope and the second neoepitope are derived from the KRAS or NRAS protein sequence of AGGVGKSALTIQLIQNHFVDEYDPTIEDSYRKQVVIDGETCLLDILDTAGHEEYSAMRDQYMRTGEGFLCVFAINNTKSFEDIHHYREQIKRVKDSEDVPM. For example, the first neoepitope and the second neoepitope may include the sequence ILDTAGHEEY.

[0284] In some embodiments, the first neoepitope and the second neoepitope are derived from the KRAS or NRAS protein sequence of AGGVGKSALTIQLIQNHFVDEYDPTIEDSYRKQVVIDGETCLLDILDTAGLEEYSAMRDQYMRTGEGFLCVFAINNTKSFEDIHHYREQIKRVKDSEDVPM. For example, the first neoepitope and the second neoepitope can include the sequence ILDTAGLEEY. For example, the first neoepitope and the second neoepitope can include the sequence LLDILDTAGL.

[0285] In some embodiments, the first neoepitope and the second neoepitope are derived from the KRAS or NRAS protein sequence of AGGVGKSALTIQLIQNHFVDEYDPTIEDSYRKQVVIDGETCLLDILDTAGKEEYSAMRDQYMRTGEGFLCVFAINNSKSFADINLYREQIKRVKDSDDVPM. For example, the first neoepitope and the second neoepitope may comprise the sequence ILDTAGKEEY.

[0286] In some embodiments, the first neoepitope and the second neoepitope are derived from the KRAS or NRAS protein sequence AGGVGKSALTIQLIQNHFVDEYDPTIEDSYRKQVVIDGETCLLDILDTAGREEYSAMRDQYMRTGEGFLCVFAINNSKSFADINLYREQIKRVKDSDDVPM. For example, the first neoepitope and the second neoepitope may comprise the sequence ILDTAGREEY.

[0287] In some embodiments, the neoepitope is selected from the group consisting of DTAGHEEY, TAGHEEYSAM, DILDTAGHE, DILDTAGH, ILDTAGHEE, ILDTAGHE, DILDTAGHEEY, DTAGHEEYS, LLDILDTAGH, DILDTAGRE, DILDTAGR, ILDTAGREE, ILDTAGRE, CLLDILDTAGR, TAGREEYSAM, REEYSAMRD, DTAGKEEYSAM, CLLDILDTAGK, DTAGKEEY, LLDILDTAGK, ILDTAGKE, ILDTAGKEE, DTAG LEEY, ILDTAGLE, DILDTAGL, ILDTAGLEE, GLEEYSAMRDQY, LLDILDTAGLE, LDILDTAGL, DILTAGLE, DILDTAGLEEY, AGVGKSAL, GAAGVGKSAL, AAGVGKSAL, CGVG KSAL, ACGVGKSAL, DGVGKSAL, ADGVGKSAL, DGVGKSALTI, GARGVGKSA, KLVVVGARGV, VVVGARGV, SGVGKSAL, VVVGASGVGK, GASGVGKSAL, VGVGKSAL, VVVGAGCVGK , KLVVVGAGC, GDVGKSAL, DVGKSALTI, VVVGAGDVGK, TAGKEEYSAM, DTAGHEEYSAM, TAGHEEYSA, DTAGREEYSAM, TAGKEEYSA, AAGVGKSA, AGCVGKSAL, AGDVGKSAL ,AGKEEYSAMR,AGVGKSALTI,ARGVGKSAL,ASGVGKSA,ASGVGKSAL,AVGVGKSA,CVGKSALTI,DILDTAGK,DILDTAGREEY,DTAGHEEYSAMR,DTAGKEEYS,DTAGKEEYS AMR, DTAGLEEYS, DTAGLEEYSA, DTAGLEEYSAMR, DTAGREEYS, DTAGREEYSAMR, GAAGVGKSA, GACGVGKSA, GACGVGKSAL, GADGVGKS, GAGDVGKSA, GAGDVGKSAL, GA SGVGKSA, GCVGKSAL, GCVGKSALTI, GHEEYSAM, GKEEYSAM, GLEEYSAMR, GREEYSAM, GREEYSAMR, HEEYSAMRD, KEEYSAMRD, KLVVVGASG, LDILDTAGR, LEEYSAMRD,LVVVGARGV, LVVVGASGV, REEYSAMRDQY, RGVGKSAL, TAGLEEYSA, TEYKLVVVGAA, VGAAGVGKSA, VGADGVGK, VGASGVGKSA, VGVGKSALTI, VVVGAAGV, VVVGAVGV, YKLVVVGAC, YKLVVVGAD, YKLVVVGAR, and DILDTAGKE.

[0288] In some embodiments, the neoepitope comprises a RAS epitope. In some embodiments, the neoepitope comprises at least 8 consecutive amino acids of a mutant RAS protein comprising a mutation at G12, G13, or Q61 and a mutant RAS sequence comprising a mutation at G12, G13, or Q61. In some embodiments, the at least 8 consecutive amino acids of a mutant RAS protein comprising a mutation at G12, G13, or Q61 comprises a G12A, G12C, G12D, G12R, G12S, G12V, G13A, G13C, G13D, G13R, G13S, G13V, Q61H, Q61L, Q61K, or Q61R mutation. In some embodiments, mutations at G12, G13, or Q61 include G12A, G12C, G12D, G12R, G12S, G12V, G13A, G13C, G13D, G13R, G13S, G13V, Q61H, Q61L, Q61K, or Q61R mutations.

[0289] In some embodiments, the neoepitope comprising a mutant RAS sequence is GADGVGKSAL, GACGVGKSAL, GAVGVGKSAL, GADGVGKSA, GACGVGKSA, GAVGVGKSA, KLVVVGACGV, FLVVVGACGL, FMVVVGACGI, FLVVVGACGI, FMVVVGACGV, FLVVVGACGV, MLVVVGACGV, FMVVVGACGL, YLVVVGACGV, KMVVVGACGV, YMVVVGACGV, MMVVVGACGV, DTAGHEEY, TAGHEEYSAM, DIL DTAGHE, DILDTAGH, ILDTAGHEE, ILDTAGHE, DILDTAGHEEY, DTAGHEEYS, LLDILDTAGH, DILTAGRE, DILDTAGR, ILDTAGREE, ILDTAGRE, CLLDILDTAGR, TAGREEY SAM, REEYSAMRD, DTAGKEEYSAM, CLLDILDTAGK, DTAGKEEY, LLDILDTAGK, ILDTAGKE, ILDTAGKEE, DTAGLEEY, ILDTAGLE, DILDTAGL, ILDTAGLEE, GLEEYSAMRDQ Y,LLDILDTAGLE,LDILDTAGL,DILDTAGLE,DILDTAGLEEY,AGVGKSAL,GAAGVGKSAL,AAGVGKSAL,CGVGKSAL,ACGVGKSAL,DGVGKSAL,ADGVGKSAL,DGVGKSALTI, GARGVGKSA, KLVVVGARGV, VVVGARGV, SGVGKSAL, VVVGASGVGK, GASGVGKSAL, VGVGKSAL, VVVGAGCVGK, KLVVVGAGC, GDVGKSAL, DVGKSALTI, VVVGAGDVGK, TAGK EEYSAM, DTAGHEEYSAM, TAGHEEYSA, DTAGREEYSAM, TAGKEEYSA, AAGVGKSA, AGCVGKSAL, AGDVGKSAL, AGKEEYSAMR, AGVGKSALTI, ARGVGKSAL, ASGVGKSA, ASGV GKSAL, AVGVGKSA, CVGKSALTI, DILDTAGK, DILDTAGREEY, DTAGHEEYSAMR, DTAGKEEYS, DTAGKEEYSAMR, DTAGLEEYS, DTAGLEEYSA, DTAGLEEYSAMR, DTAGREEYS,DTAGREEYSAMR, GAAGVGKSA, GACGVGKSA, GACGVGKSAL, GADGVGKS, GAGDVGKSA, GAGDVGKSAL, GASGVGKSA, GCVGKSAL, GCVGKSALTI, GHEEYSAM, GKEEYSAM, GLEEYSAMR, GREEYSAM, GREEYSAMR, HEEYSAMRD, KEEYSAMRD, KLVVVGASG, LDILDTAGR, LEEYSAMRD, LVVVGARGV, LVVVGASGV, REEYSAMRDQY, RGVGKSAL, TAGLEEYSA, TEYKLVVVGAA, VGAAGVGKSA, VGADGVGK, VGASGVGKSA, VGVGKSALTI, VVVGAAGV, VVVGAVGV, YKLVVVGAC, YKLVVVGAD, YKLVVVGAR, or DILDTAGKE.

[0290] In some embodiments, the neoepitope comprising the mutant RAS sequence binds to a protein encoded by an HLA allele. In some embodiments, the neoepitope comprising the mutant RAS sequence binds to a protein encoded by an HLA allele at less than 10 μM, less than 9 μM, less than 8 μM, less than 7 μM, less than 6 μM, less than 5 μM, less than 4 μM, less than 3 μM, less than 2 μM, less than 1 μM, less than 950 nM, less than 900 nM, less than 850 nM, less than 800 nM, less than 750 nM, less than 600 nM, ...00 nM, less than 9 μM, less than 9 μM, less than 1 μM, less than 9 μM, less than 1 μM, less than 9 μM, less than 1 μM, less than 9 μM, less than 1 μM, less than 9 μM, less than 1 μM, less than 9 μM, less than 1 μM, less than 9 μM, less than 1 μM, less than 9 μM, less than 1 μM, less than 9 μM, less than 1 μM, less than 9 μM, less than 1 μM, less than 9 μM, less than 1 μM, less than 1 μM nM, less than 550 nM, less than 500 nM, less than 450 nM, less than 400 nM, less than 350 nM, less than 300 nM, less than 250 nM, less than 200 nM, less than 150 nM, less than 100 nM, less than 90 nM, less than 80 nM, less than 70 nM, less than 60 nM, less than 50 nM, less than 40 nM, less than 30 nM, less than 20 nM, or less than 10 nM. In some embodiments, the neoepitope comprising the mutant RAS sequence is associated with a protein encoded by an HLA allele for greater than 24 hours, greater than 23 hours, greater than 22 hours, greater than 21 hours, greater than 20 hours, greater than 19 hours, greater than 18 hours, greater than 17 hours, greater than 16 hours, greater than 15 hours, greater than 14 hours, greater than 13 hours, greater than 12 hours, greater than 11 hours, greater than 10 hours, greater than 9 hours, greater than 8 hours, greater than 7 hours, greater than 6 hours, greater than 8 hours, greater than 9 hours, greater than 10 ... The binding stability may be greater than 5 hours, greater than 4 hours, greater than 3 hours, greater than 2 hours, greater than 1 hour, greater than 55 minutes, greater than 50 minutes, greater than 45 minutes, greater than 40 minutes, greater than 35 minutes, greater than 30 minutes, greater than 25 minutes, greater than 20 minutes, greater than 15 minutes, greater than 10 minutes, greater than 9 minutes, greater than 8 minutes, greater than 7 minutes, greater than 6 minutes, greater than 5 minutes, greater than 4 minutes, greater than 3 minutes, greater than 2 minutes, or greater than 1 minute.

[0291] The substitutions may be located anywhere along the length of the neoepitope. For example, the substitutions may be located in the N-terminal third of the peptide, the central third of the peptide, or the C-terminal third of the peptide. In another embodiment, the substituted residues are located 2-5 residues away from the N-terminus or 2-5 residues away from the C-terminus. Peptides may similarly be derived from tumor-specific insertion mutations that include one or more or all of the peptide-inserted residues.

[0292] In some embodiments, the peptides described herein can be readily chemically synthesized utilizing reagents free of bacterial or animal contaminants (Merrifield RB: Solid phase peptide synthesis. I. The synthesis of a tetrapeptide. J. Am. Chem. Soc.85: 2149-54, 1963). In some embodiments, the peptides can be synthesized by: (1) homogeneous synthesis; (1) Parallel solid-phase synthesis on a multichannel instrument using synthesis and cleavage conditions; (2) purification on a RP-HPLC column with column stripping; and re-washing but no replacement between peptides; then (3) preparation by analysis with a limited set of the most informative assays. Good Manufacturing Practice (GMP) footprints can be defined for individual patients based on sets of peptides, and therefore a set of switchover procedures is only required between the synthesis of peptides for different patients. In some embodiments, any resin made for solid-phase peptide synthesis can be used. Polynucleotides

[0293] Alternatively, nucleic acids (e.g., polynucleotides) encoding the peptides of the present disclosure can be used to generate neo-antigenic peptides in vitro. The polynucleotides can be, for example, DNA, cDNA, RNA, single-stranded and / or double-stranded, or in native or stabilized form, such as, for example, polynucleotides with phosphorothioate backbones, or combinations thereof, and may or may not contain introns, so long as they encode the peptide. In some embodiments, in vitro translation is used to generate the peptides.

[0294] Provided herein are neo-antigenic polynucleotides encoding each of the neo-antigenic polypeptides described in this disclosure. The terms "polynucleotide", "nucleotide" or "nucleic acid" are used interchangeably in this disclosure with "mutated polynucleotide", "mutated nucleotide", "mutated nucleic acid", "neo-antigenic polynucleotide", "neo-antigenic nucleotide" or "neo-antigenic mutant nucleic acid". Due to redundancy in the genetic code, various nucleic acid sequences may encode the same peptide. Each of these nucleic ac...

Claims

[Claim 1] The invention described in the specification.