Methods and compositions for designing epitope strings with linkers

A computer-implemented algorithm optimizes T cell epitope ordering and linker insertion to enhance epitope cleavage and presentation, addressing the challenges in T cell vaccine design and improving vaccine efficacy.

WO2026013643A1PCT designated stage Publication Date: 2026-01-15BIONTECH SE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/057068
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-12
Filing Date
2025-07-11
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing methods for designing T cell vaccines face challenges in effectively ordering T cell epitopes and inserting linkers to ensure maximal cleavage and presentation of epitopes when expressed in cells, which is crucial for immune activation.

Method used

A method involving a computer-implemented algorithm that orders T cell epitope sequences into a defined N-terminus to C-terminus order and inserts amino acid linkers to optimize the fitness score, using a cleavage predictor to ensure efficient cleavage and presentation of epitopes.

Benefits of technology

The method enhances the probability of T cell epitope cleavage and presentation, improving the efficacy of T cell vaccines by maximizing the chances of epitope release and activation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025057068_15012026_PF_FP_ABST
    Figure IB2025057068_15012026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides computational methods for designing superior T cell epitopes followed by experimental verification, feedback and further improvement of cleavage linkers between certain epitopes.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND COMPOSITIONS FOR DESIGNING EPITOPE STRINGS WITH LINKERSCROSS REFERENCE

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 670,649, filed on July 12, 2024, which is incorporated herein by reference in its entirety.BACKGROUND

[0002] T cell vaccines are increasingly being considered as effective therapeutic modalities. Use of RNA for delivery of nucleic acid material to a subject with a disease, where the RNA is designed and optimized for encoding therapeutic proteins and peptides has advanced in recent years, thereby boosting the long-felt necessity to develop therapeutics that reprogram a subject’s own immune system to combat difficult diseases such as cancer and certain infectious diseases. However, major hurdles still exist for the developing immunotherapeutics using T cell activating peptides and generating effective T cell vaccines directed to treatment of cancer and infectious diseases such as HIV and EBV.SUMMARY

[0003] In one aspect, provided herein is a method for generating a multiepitopic polypeptide sequence from a plurality of T cell epitope containing sequences, each T cell epitope containing sequence of the plurality comprising or consisting of at least one T cell epitope sequence, wherein the plurality of T cell epitope containing sequences comprises at least two T cell epitope sequences, the method comprising: (a) ordering the plurality of T cell epitope containing sequences into a defined N-terminus to C-terminus order, thereby generating a starting polypeptide sequence, wherein the starting polypeptide sequence has a sequence comprising each T cell epitope containing sequence of the plurality directly linked to at least one other T cell epitope containing sequences of the plurality; (b) performing a first linker selection round, the first linker selection round comprising: (i) inserting an amino acid linker sequence into the starting polypeptide sequence between a junction of a first T cell epitope containing sequence and a second T cell epitope containing sequence, thereby generating a first modified starting polypeptide sequence, wherein the amino acid linker sequence consists of 1, 2, 3 or 4 amino acids, (ii) comparing a fitness score of the starting polypeptide sequence to a fitness score of the first modified starting polypeptide sequence; and (iii) selecting a first multiepitopic polypeptide sequence from the starting polypeptide sequence and the first modified starting polypeptide sequence based on the comparison of the fitness scores, thereby generating a first multiepitopic polypeptide sequence. In some embodiments, the method is performed in silico. In one embodiment, the method comprises optionally inserting no amino acid at a given junction of a first T cell epitope containing sequence and a second T cell epitope containing sequence of the plurality of T cell epitope containing sequences.

[0004] In one aspect, provided herein is a method for generating a multiepitopic polypeptide sequence from a plurality of T cell epitope containing sequences, each T cell epitope containingsequence of the plurality comprising or consisting of at least one T cell epitope sequence, wherein the plurality of T cell epitope containing sequences comprises at least three T cell epitope sequences, the method comprising: (a) ordering the plurality of T cell epitope containing sequences into a defined N-terminus to C-terminus order, thereby generating a starting polypeptide sequence, wherein the starting polypeptide sequence has a sequence comprising each T cell epitope containing sequence of the plurality directly linked to one or two other T cell epitope containing sequences of the plurality, ; (b) performing a first linker selection round, the first linker selection round comprising: (i) inserting an amino acid linker sequence into the starting polypeptide sequence between a junction of first T cell epitope containing sequence and a second T cell epitope containing sequence, thereby generating a first modified starting polypeptide sequence, wherein the amino acid linker sequence consists of 1, 2, 3 or 4 amino acids; (ii) comparing a fitness score of the starting polypeptide sequence to a fitness score of the first modified starting polypeptide sequence; and (iii) selecting a first multiepitopic polypeptide sequence from the starting polypeptide sequence and the first modified starting polypeptide sequence based on the comparison of the fitness scores, thereby generating a first multiepitopic polypeptide sequence. In some embodiments, the method is performed in silico.

[0005] In one aspect, provided herein is a method for generating a multiepitopic polypeptide sequence from a plurality of sequence fragments, each sequence fragment of the plurality of sequence fragments comprising at least one T cell epitope sequence, the method comprising: ordering the plurality of sequence fragments into a selected polypeptide sequence comprising a defined N-terminus to C- terminus order, thereby generating a starting polypeptide sequence, wherein the defined N-terminus to C-terminus order is a selected single possible order of all possible orders of the plurality of sequence fragments; comparing a fitness score of the starting polypeptide sequence to a fitness score of a modified polypeptide sequence, wherein the modified polypeptide sequence comprises (i) the defined N-terminus to C-terminus order of the selected single possible order of all possible orders that is in the starting polypeptide sequence and (ii) an amino acid linker sequence at one or more junctions of the plurality of sequence fragments of the starting polypeptide sequence, wherein the amino acid linker sequence consists of 1, 2, 3 or 4 amino acids; selecting a sequence between the sequence of the starting polypeptide sequence and the sequence of the modified polypeptide sequence based on the comparing of the fitness scores of the starting polypeptide sequence and the modified polypeptide sequence, thereby generating the multiepitopic polypeptide sequence, wherein the multiepitopic polypeptide sequence is the sequence that is selected.

[0006] In one aspect, provided herein is a method for generating a multiepitopic polypeptide sequence from a plurality of T cell epitope containing sequences, each T cell epitope containing sequence of the plurality comprising or consisting of at least one T cell epitope sequence, wherein the plurality of T cell epitope containing sequences comprises at least three T cell epitope sequences, the method comprising: (a) ordering the plurality of T cell epitope containing sequences into a defined N-terminus to C-terminus order, thereby generating a starting multiepitopic polypeptide sequence, wherein the starting multiepitopic polypeptide sequence has a sequence comprising each T cell epitope containing sequence of the plurality directly linked to one or two other T cell epitope containing sequences of the plurality; (b) performing a first linker selection round, the first linker selection round comprising: (i) inserting an amino acid linker sequence into the starting polypeptide sequence between a junction of a first T cell epitope containing sequence and a second T cell epitope containing sequence, thereby generating a first modified starting multiepitopic polypeptide sequence; (ii) comparing a fitness score of the starting polypeptide sequence to a fitness score of the first modified starting multiepitopic polypeptide sequence; and (iii) selecting the first multiepitopic polypeptide sequence from the starting multiepitopic polypeptide sequence and the first modified starting multiepitopic polypeptide sequence based on the comparison of the fitness scores, thereby generating a multiepitopic polypeptide sequence.

[0007] In some embodiments, the method is performed in silico.

[0008] In some embodiments, the amino acid linker sequence consists of 1 amino acid during the first linker selection round. In some embodiments, the first linker selection round comprises performing the inserting, comparing and selecting for each junction of two adjacent T cell epitope containing sequences. In some embodiments, the method further comprises performing a second linker selection round, wherein the amino acid linker sequence consists of 2 amino acids during the second linker selection round. In some embodiments, the second linker selection round comprises performing the inserting, comparing and selecting for each junction of two adjacent T cell epitope containing sequences.

[0009] In some embodiments, the method further comprises performing a third linker selection round, wherein the amino acid linker sequence consists of 3 amino acids during the third linker selection round.

[0010] In some embodiments, the third linker selection round comprises performing the inserting, comparing and selecting for each junction of two adjacent T cell epitope containing sequences.

[0011] In some embodiments, the method further comprises performing a fourth linker selection round, wherein the amino acid linker sequence consists of 4 amino acids during the fourth linker selection round.

[0012] In some embodiments, the selection linker selection round comprises performing the inserting, comparing and selecting for each junction of two adjacent T cell epitope containing sequences.

[0013] In some embodiments, the method comprises performing a number of additional linker selection rounds each additional selection round comprising: (i) inserting an amino acid linker sequence into the junction of two adjacent T cell epitope containing sequences in the selected multiepitopic polypeptide sequence from the previous linker selection round; (ii) comparing a fitnessscore of the selected multiepitopic polypeptide sequence from the previous linker selection round to a fitness score of the modified multiepitopic polypeptide sequence of the current additional linker selection round, and (iii) selecting a multiepitopic polypeptide sequence from the selected multiepitopic polypeptide sequence from the previous linker selection round and the modified multiepitopic polypeptide sequence of the current additional linker selection round based on the comparison of the fitness scores.

[0014] In some embodiments, each of the inserting steps of a given linker selection round comprises inserting, at a given junction of two adjacent T cell epitope containing sequences, an amino acid linker sequence that is one amino acid in length longer than the length of the amino acid linker sequence inserted during the previous linker selection round.

[0015] In some embodiments, each of the inserting steps of a first linker selection round comprises inserting, at a given junction of two adjacent T cell epitope containing sequences, an amino acid linker sequence that is one amino acid in length.

[0016] In some embodiments, each of the comparing steps of the first linker selection round comprises comparing a multiepitopic polypeptide sequence that lacks an amino acid linker sequence at the given junction of two adjacent T cell epitope containing sequences to a modified multiepitopic polypeptide sequence that has an amino acid linker sequence at the given junction of two adjacent T cell epitope containing sequence that is one amino acid in length.

[0017] In some embodiments, each of the inserting steps of a second linker selection round comprises inserting, at a given junction of two adjacent T cell epitope containing sequences, an amino acid linker sequence that is two amino acids in length.

[0018] In some embodiments, each of the comparing steps of the second linker selection round comprises comparing a multiepitopic polypeptide sequence that either lacks an amino acid linker sequence or has an amino acid linker sequence that is one amino acid in length at the given junction of two adjacent T cell epitope containing sequences to a modified multi epitopic polypeptide sequence that has an amino acid linker sequence at the given junction of two adjacent T cell epitope containing sequence that is two amino acid in length.

[0019] In some embodiments, each of the inserting steps of a third linker selection round comprises inserting, at a given junction of two adjacent T cell epitope containing sequences, an amino acid linker sequence that is three amino acids in length.

[0020] In some embodiments, of the comparing steps of the third linker selection round comprises comparing a multiepitopic polypeptide sequence that either lacks an amino acid linker sequence or has an amino acid linker sequence that is two amino acid in length at the given junction of two adjacent T cell epitope containing sequences to a modified multiepitopic polypeptide sequence that has an amino acid linker sequence at the given junction of two adjacent T cell epitope containing sequence that is three amino acid in length.

[0021] In some embodiments, each of the inserting steps of a fourth linker selection round comprises inserting, at a given junction of two adjacent T cell epitope containing sequences, an amino acid linker sequence that is four amino acids in length.

[0022] In some embodiments, each of the comparing steps of the fourth linker selection round comprises comparing a multiepitopic polypeptide sequence that either lacks an amino acid linker sequence or has an amino acid linker sequence that is three amino acid in length at the given junction of two adjacent T cell epitope containing sequences to a modified multi epitopic polypeptide sequence that has an amino acid linker sequence at the given junction of two adjacent T cell epitope containing sequence that is four amino acid in length.

[0023] In some embodiments, each of the comparing steps of a given linker selection round comprises comparing a fitness score of a given modified multiepitopic polypeptide comprising an amino acid linker sequence at a given junction of two adjacent T cell epitope containing sequences that is one amino acid in length longer than the length of the amino acid linker sequence inserted at the given junction of two adjacent T cell epitope containing sequences to a fitness score of a given multiepitopic polypeptide comprising no amino acid linker sequence or that has an amino acid linker sequence at the given junction of two adjacent T cell epitope containing sequences that is one amino acid in length shorter than the length of the amino acid linker sequence inserted at the given junction of two adjacent T cell epitope containing sequences of the given modified multi epitopic polypeptide.

[0024] In some embodiments, selecting comprises selecting a multiepitopic polypeptide sequence based a combined sequence of a multiepitopic polypeptide sequence not being present in the human proteome, wherein the combined sequence comprises a sequence of at least 8 amino acids, wherein: (A) when the multiepitopic polypeptide sequence comprises an amino acid linker sequence consisting of 1, 2, 3 or 4 amino acids: the sequence of at least 8 amino acids comprises at least one amino acid of the amino acid linker sequence and at least one amino acid of a T cell epitope containing sequence adjacent to the amino acid linker sequence, or (B) when the multi epitopic polypeptide sequence does not comprises an amino acid linker sequence: the sequence of at least 8 amino acids comprises at least one amino acid of each of the two adjacent T cell epitope containing sequences.

[0025] In some embodiments, the fitness score of a given multiepitopic polypeptide sequence is based on a cleavability score of each T cell epitope containing sequence, wherein the cleavability score of a T cell epitope containing sequence is indicative of the probability that a T cell epitope is cleaved from the given multiepitopic polypeptide sequence when the multiepitopic polypeptide sequence is expressed in a cell.

[0026] In some embodiments, the fitness score of a given multiepitopic polypeptide sequence is a function of the cleavability score of all T cell epitope containing sequences in the given multiepitopic polypeptide sequence, optionally, wherein the fitness score of a given multiepitopic polypeptidesequence is based on the average cleavability score of all T cell epitope containing sequences in the given multiepitopic polypeptide sequence.

[0027] In some embodiments, the cleavability score of a T cell epitope containing sequence is based on the T cell epitope containing sequence and the amino acid sequence upstream and / or downstream of the T cell epitope containing sequence.

[0028] In some embodiments, the T cell epitope containing sequence and the amino acid sequence upstream and / or downstream of the T cell epitope containing sequence is at least 20, 25, or 30 amino acids in length.

[0029] In some embodiments, the cleavability score and / or the fitness score is predicted by a trained prediction model implemented on a computer.

[0030] In some embodiments, comparing comprises inputting amino acid sequence information of a given multiepitopic polypeptide sequence using a computer processor into a trained prediction model to generate a plurality of cleavability predictions.

[0031] In some embodiments, the amino acid sequence information comprises the T cell epitope containing sequence and the amino acid sequence upstream and / or downstream of the T cell epitope containing sequence.

[0032] In some embodiments, the plurality of cleavability predictions comprises a cleavability prediction for each T cell epitope of the given multiepitopic polypeptide sequence.

[0033] In some embodiments, each cleavability prediction of the plurality of cleavability predictions is indicative of a probability that a given T cell epitope is cleaved from a given multiepitopic polypeptide sequence is expressed in a cell.

[0034] In some embodiments, the cleavability score is indicative of a probability that a given T cell epitope is cleaved from the multiepitopic polypeptide sequence when the multiepitopic polypeptide sequence is expressed in a cell.

[0035] In some embodiments, at least one T cell epitope containing sequence of the plurality comprises two or more T cell epitope sequences.

[0036] In some embodiments, at least one T cell epitope containing sequence of the plurality comprises two or more T cell epitope sequences that share at last one amino acid.

[0037] In some embodiments, the sequence of a first T cell epitope sequence of the two or more T cell epitope sequences starts at the N-terminus of the at least one T cell epitope containing sequence of the plurality that comprises two or more T cell epitope sequences, and wherein the sequence of a second T cell epitope sequence of the two or more T cell epitope sequences ends at the C-terminus of the at least one T cell epitope containing sequence of the plurality that comprises two or more T cell epitope sequences.

[0038] In some embodiments, inserting comprises randomly selecting two adjacent T cell epitope sequences of a multiepitopic polypeptide sequence and inserting an amino acid linker sequence intothe multi epitopic polypeptide sequence between a junction of the two adjacent T cell epitope sequences randomly selected.

[0039] In some embodiments, the method is performed in a computer implemented machine learning program having a framework that comprises an ordering algorithm, an objective function, and a sampling function.

[0040] In some embodiments, the ordering algorithm comprising a simulated annealing module.

[0041] In some embodiments, the objective function comprises a cleavage predictor algorithm.

[0042] In some embodiments, the sampling function comprises an epitope sampling function and linker sampling function.

[0043] In some embodiments, each of the junctions between two adjacent T cell epitope containing sequences are variable junctions.

[0044] In some embodiments, introducing comprises one or more rounds of introducing, each round comprising: (a) inputting a potential amino acid linker sequence from a set of potential amino acid linker sequences at a given variable junction; (b) determining whether at each given variable junction, a potential amino acid linker sequence together with a sequence that is immediately upstream or a sequence immediately downstream of the potential amino acid linker sequence comprises a sequence of at least 8 consecutive amino acids that is present in the human proteome; and (c) discarding the multiepitopic polypeptide sequence if the sequence of at least 8 consecutive amino acids is found to be present in the human proteome.

[0045] In some embodiments, the maximum number of amino acids for a given amino acid linker sequence is 4. In some embodiments, the minimum number of amino acids for a given amino acid linker sequence is 0.

[0046] In some embodiments, selecting comprises preferentially selecting a multiepitopic polypeptide sequence having an amino acid linker sequence with a minimal length at one or more or each junction between two adjacent T cell epitope containing sequences.

[0047] In one aspect, provided herein is a method for designing a multiepitopic polypeptide sequence from a plurality of T cell epitope containing sequences, each T cell epitope containing sequence of the plurality comprising or consisting of at least one T cell epitope sequence, wherein the plurality of T cell epitope containing sequences comprises at least three T cell epitope sequences, the method comprising performing a first T cell epitope containing sequence ordering round, the first T cell epitope containing sequence ordering round comprising: (a) ordering each T cell epitope containing sequence of the plurality into a first defined N to C terminal order, thereby generating a starting multiepitopic polypeptide sequence; (b) ordering each T cell epitope containing sequence of the plurality into a second defined N to C terminal order, thereby generating a first modified multiepitopic polypeptide sequence; (c) comparing a fitness score of the starting multiepitopic polypeptide sequence to a fitness score of the first modified starting multiepitopic polypeptidesequence; and (d) selecting a first multiepitopic polypeptide sequence from the starting multiepitopic polypeptide sequence and the first modified starting multiepitopic polypeptide sequence based on the comparison of the fitness scores, thereby generating a first multiepitopic polypeptide sequence.

[0048] In some embodiments, the starting polypeptide sequence has a sequence comprising each T cell epitope containing sequence of the plurality directly linked to one or two other T cell epitope containing sequences of the plurality.

[0049] In some embodiments, the modified polypeptide sequence has a sequence comprising each T cell epitope containing sequence of the plurality directly linked to one or two other T cell epitope containing sequences of the plurality.

[0050] In some embodiments, selecting comprises selecting the multiepitopic polypeptide sequence with the higher fitness score.

[0051] In some embodiments, selecting comprises selecting the multiepitopic polypeptide sequence with the lower fitness score ifwherein the second defined order score is the fitness score of the modified starting multiepitopic polypeptide sequence and the first defined order score is the fitness score of the starting multi epitopic polypeptide sequence, and the random sample value is a random number from a standard uniform distribution defined on a closed interval of [0,1],

[0052] In some embodiments, wherein the fitness score of a given multiepitopic polypeptide sequence is a fitness score computed for a predefined temperature parameter.

[0053] In some embodiments, the score of the second defined order is at least 1.1 -fold higher than the score of the first defined order.

[0054] In some embodiments, the method further comprises progressively repeating (a)-(d) thereby selecting an order that is the highest possible score for a given multiepitopic polypeptide sequence.

[0055] In some embodiments, a fitness score of a selected multiepitopic polypeptide sequence is at least 2.0.

[0056] In some embodiments, the fitness score of a multiepitopic polypeptide sequence ranges from (-)4.0 to +4.0. In some embodiments, the fitness score is a score that ranges from about (-)4.0 to about (+)4. In some embodiments, the fitness score is a score that ranges from about (-)4.0 to about (+)4 based on current scoring scheme.

[0057] In some embodiments, a cleavage predicting algorithm used to compute the fitness score has been trained with user-input data from experimentally verified T cell epitope sequences that are cleaved when expressed in a cell, as observed by mass spectrometry assay, an immunoassay or by a T cell activation assay.

[0058] In some embodiments, the immunoassay is a detection assay of an epitope that is bound to an MHC or a cell.

[0059] In some embodiments, the data from experimentally verified T cell epitope sequences comprises verified by a tetramer assay.

[0060] In some embodiments, the T cell activation assay comprises cytokine release assay by an activated T cell, as determined by ELISA or flow cytometry.

[0061] In some embodiments, the method comprises performing the method according to any one of the embodiments discussed above after performing the first T cell epitope containing sequence ordering round.

[0062] In some embodiments, the method comprises performing the method according to any one of any one of the embodiments discussed above prior to performing the first linker selection round.

[0063] In some embodiments, the method further comprises generating a therapeutic composition comprising the multiepitopic polypeptide sequence or a nucleic acid sequence encoding the multiepitopic polypeptide sequence.

[0064] In some embodiments, the therapeutic composition is a T cell vaccine composition.

[0065] In some embodiments, the multi epitopic polypeptide sequence comprises at least 10 T cell epitope sequences or from 10 to 100 T cell epitope sequences.

[0066] In one aspect, provided herein is a pharmaceutical composition comprising the multi epitopic polypeptide sequence or a nucleic acid sequence encoding the multi epitopic polypeptide sequence generated by a method of any one of the embodiments discussed above, for treating a disease in a subject.

[0067] In one aspect, provided herein is a method of treating a disease in a subject in need thereof, comprising administering the subject the pharmaceutical composition of the embodiment discussed above, wherein the subject is a human subject.

[0068] In one aspect, provided herein is a use of a composition comprising the multiepitopic polypeptide sequence or a nucleic acid sequence encoding the multiepitopic polypeptide sequence of the embodiment(s) presented above, or the multiepitopic polypeptide sequence or a nucleic acid sequence encoding the multiepitopic polypeptide sequence generated by a method of any one of embodiments described above in preparing a medicament for treating a disease in a subject in need thereof.

[0069] In some embodiments, the method in accordance to an embodiment discussed above, or the use in accordance to an embodiment discussed above, wherein the disease is a cancer or an infectious disease.

[0070] In one aspect, provided herein is a system for generating a multiepitopic polypeptide sequence from a plurality of T cell epitope sequences, the system comprising: (i) an input module for receiving a plurality of T cell epitope sequences; and (ii) framework module comprising one or moreprogrammed functions, wherein the one or more programmed functions operate on one or more trained algorithms comprising: (a) an ordering algorithm for transforming the plurality of T cell epitope sequences into a defined order of T cell epitope sequences from N terminus to C terminus, wherein the defined N-terminus to C-terminus order is a single possible order of all possible orders of the plurality of T cell epitope sequences, thereby forming a sequence having the defined N-terminus to C- terminus order; and (b) a cleavage predictor algorithm for predicting a probability that a T cell epitope sequence is cleaved and presented on a cell surface when the multiepitopic polypeptide sequence is expressed in a cell; and (c) a set of sampling rules for the sampling function comprises an epitope sampling and linker sampling functions.

[0071] In some embodiments, the further comprising an output module comprising a display that displays the multiepitopic polypeptide sequence and a fitness score for the multiepitopic polypeptide sequence representing a probability of cleavage of each epitope sequence within the multiepitopic polypeptide sequence when the multiepitopic polypeptide is expressed in a cell.

[0072] In some embodiments, the one or more trained algorithms are machine learning algorithms that have been trained with user-input data from experimentally verified T cell epitope sequences and linker sequences that are cleaved when expressed in a cell.

[0073] In some embodiments, the user-input data from experimentally verified T cell epitope sequences comprise a T cell epitope sequence that has been verified to (i) bind to an MHC molecule encoded by an HLA allele by a peptide-MHC binding assay and / or affinity assay; (ii) be presented by an APC as determined by a tetramer assay; (iii) bind to a T cell in vitro as determined by immunoassay, and / or (iv) activate a T cell upon contacting as determined by a cytokine release assay.

[0074] In some embodiments, user input data may comprise data generated in a computer implemented system, having appropriate statistical significance.

[0075] In some embodiments, provided herein is a system such that the framework module is set to generate a multiepitopic polypeptide, or a polynucleotide sequence encoding the multiepitopic polypeptide, containing epitopes and linkers, wherein an epitope encoded by the polynucleotide sequence is presented by a cell at an abundance that is at least 1.2 times the abundance with which the same epitope is presented that is encoded by (A) a control polynucleotide sequence; or (B) a polynucleotide sequence encoding a polypeptide (i) containing the same epitopes and with synthetic linkers not generated by the system, (ii) containing the same epitopes and linkers that are not in an order of epitopes and linkers as generated by the system, or (iii) containing a sequence that is not identical to the sequence of the multiepitopic polypeptide as generated by the framework module of the system. In some embodiments, the control polynucleotide is a polynucleotide encoding a control polypeptide lacking epitopes that are identical to epitopes of the multiepitopic polypeptide generated by the framework module of the system, or a control polypeptide that comprises the same epitopes but lacks any linker sequences. In some embodiments, the control polynucleotide is a polynucleotidehaving the same epitopes but lacking synthetic linkers. In some embodiments, an epitope encoded by the multiepitopic polynucleotide sequence is presented by a cell at an abundance that is at least 2 times the abundance with which the same epitope is presented that is encoded by (A) a control polynucleotide sequence; or (B) a polynucleotide sequence encoding a polypeptide (i) containing the same epitopes and with synthetic linkers not generated by the system, (ii) containing the same epitopes and linkers that are not in an order of epitopes and linkers as generated by the system, or (iii) containing a sequence that is not identical to the sequence of the multiepitopic polypeptide as generated by the framework module of the system.

[0076] Provided herein is a composition comprising a polynucleotide sequence encoding a multiepitopic polypeptide, the polynucleotide sequence having at least 90% sequence identity to a polynucleotide sequence set forth in any one of SEQ ID NOs: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, and 23.

[0077] Provided herein is a multiepitopic polypeptide encoded by a polynucleotide sequence having at least 90% sequence identity to a polynucleotide sequence set forth in any one of SEQ ID NOs: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21 and 23.

[0078] Provided herein is a multiepitopic polypeptide comprising an amino acid sequence that has at least 85% sequence identity to a polypeptide set forth in any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 and 24.

[0079] Provided herein is a multiepitopic polypeptide comprising at least 3 consecutive epitope sequences and linkers set forth in any one of Tables 2-11.

[0080] Provided herein is a multiepitopic polypeptide comprising at least 2, 3, 4, 5, 6 or more consecutive epitope sequences and linkers set forth in any one of Tables 2-11 and at least one GS linker. Said multiepitopic polypeptide comprising at least 2, 3, 4, 5, 6 or more consecutive epitope sequences and linkers may comprise a sequence in which at least two consecutive epitope sequences are intervened by a linker sequence, wherein the linker sequence is the one generated in an optimization run by the cleavage optimizer (CLEO) program for placement in between the two epitope sequences when placed in the N terminal to C order (i.e., consecutive). The linker sequence is at least one amino acid and up to four amino acids in length. The multiepitopic polypeptide may comprise at least three (3) epitope sequences in an N terminal to C order (e.g., consecutive epitope sequences that match with arrangement of epitopes of a multiepitopic polypeptide as generated by the framework module), in which at least two consecutive epitope sequences are intervened by at least one linker sequence in between the two consecutive epitope sequences wherein the linker sequence is the one generated in an optimization run by the cleavage optimizer program for placement in between the two consecutive epitope sequences. The linker sequence is at least one amino acid and up to four amino acids in length. Likewise, the multiepitopic polypeptide may comprise at least four (4) epitope sequences in the N terminal to C order (i.e., consecutive), in which at least two consecutive epitopesequences are intervened by at least one linker sequence in between the two consecutive epitope sequences wherein the linker sequence is the one generated in an optimization run by the cleavage optimizer program for placement in between the two epitope sequences. The linker sequence is at least one amino acid and up to four amino acids in length.

[0081] Provided herein is a multiepitopic polypeptide comprising at least 2, 3, 4, 5, 6 or more consecutive epitope sequences and linkers set forth in any one of Tables 2-11 and at least one helical linker. Said multiepitopic polypeptide comprising at least 2, 3, 4, 5, 6 or more consecutive epitope sequences and linkers may comprise a sequence in which two consecutive epitope sequences are intervened by a linker sequence, wherein the linker sequence was generated in an optimization run by the cleavage optimizer program for placement in between the two epitope sequences when placed in the N terminal to C order (consecutive); and at least one helical linker in between two other consecutive epitope sequences in the multiepitopic polypeptide. Similar to the above, also provided herein is a multiepitopic polypeptide comprising at least 2, 3, 4, 5, 6 or more consecutive epitope sequences and linkers set forth in any one of Tables 2-11 and at least one Furin linker. In some embodiments, the Furin linker is a viral Furin linker. In some embodiments, the Furin linker is a human Furin linker. In some embodiments, the multiepitopic polypeptide may further comprise a solubility enhancing modification. In some embodiments, the solubility enhancing modification is inclusion of a sequence for an N-terminal mannose binding (MBP) protein. In some embodiments, the multiepitopic polypeptide may further comprise one or more mutations that minimize ribosomal stalling motifs. A multiepitopic polypeptide described herein may comprise any one of the abovedescribed structures, including one or more conventional linkers between two consecutive epitope sequences, such linkers including, a GS linker, a helical linker, a Furin linker, or a combination of any of these.INCORPORATION BY REFERENCE

[0082] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWING

[0083] The salient features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “FIG.” herein), of which:

[0084] FIG. 1 is a schematic representation of a generalized and simplified mode of action of a computer implemented cleavage predictor: the algorithm uptakes as input an epitope along with “N” number of amino acids upstream and downstream. The output has a cleavage score for that epitope. A higher score corresponds to a better chance of being cleaved and vice versa.

[0085] FIG. 2A is a graphical representation of some drawbacks in existing methods of polyepitopic T cell vaccine design. Minimal epitopes or fragments are required to be arranged in certain order. Making an arbitrary ordering selection by hand creates a large chance that not all epitopes will be cleaved out of their respective sequence context.

[0086] FIG. 2B is a schematic representation showing the process of generating a multiepitopic polypeptide sequence from a plurality of T cell epitopes or fragments, where the multiepitopic polypeptide sequence has a defined N-terminal to C-terminal that allows optimal release of the epitopes when expressed in a cell, by using an in-house generated computer implemented program comprising a simulated annealing algorithm and a cleavage predictor as disclosed and described herein.

[0087] FIG. 2C shows a schematic representation of the generation of an exemplary T cell vaccine using the new multiepitopic polypeptide design incorporating a computer implemented cleavage predictor program as disclosed herein. The figure shows that the program can accept an input of a user-defined plurality of T cell epitope sequences or sequence fragments containing epitope sequences, and generate as an output a multiepitopic polypeptide sequence having a defined N- terminal to C-terminal order of the T cell epitopes, and also generates short linker sequences at variable T cell epitope junctions that are optimized for a high probability of cleavage and release of T cell epitope sequences from the polypeptide when expressed in a cell. Linkers are up to 4 amino acids long. Shorter linkers are prioritized. N terminal and C-terminal sequence are incorporated for efficient cleavage and / or immunogenicity.

[0088] FIG. 3A-FIG. 3C show a progressive optimizing scheme designed in the computer implemented cleavage optimizer (CLEO) program described herein, where FIG. 3 A shows a Round 0, FIG. 3B shows a Round 1, and FIG. 3C shows a Round 2-4 of optimization and generating a T cell vaccine product, comprising a multiepitopic polypeptide string.

[0089] FIG. 3 A: Cleavage optimizer starts at Round 0 to determine an optimal ordering of epitopes without any linkers. The goal is to arrange them on a string such that the chance of epitope release is maximized. The round is sub-divided into epochs wherein the epitope ordering changes. For each new string, a fitness score was calculated. If the score is better, the new ordering is accepted; if not, a random choice is made as to whether to accept or reject the new ordering. This process continues for a preset number of epochs after which the epitope ordering is fixed.

[0090] FIG. 3B: Round 1 is dedicated to adding the additional of optimal linkers of length 1. The task begins with linkers of length one under the assumption that shorter linkers are less likely tointroduce human proteome overlap at the junctions. The round is sub-divided into epochs. Within each epoch, a linker is placed in an epitope junction, a new fitness score is calculated; if it is better than previous one, the new candidate string is accepted, if not, a random choice is made whether to accept or rej ect the updated string candidate. This process continues for a preset number of epochs after which the string candidate contains optimal length-1 linkers. Some junctions may have no linker if an addition of the linker does not improve the fitness score.

[0091] FIG. 3C: Rounds 2 through 4 are dedicated to the additional of optimal length-n linkers, where n is the round number. The operation is continued under the assumption that shorter linkers are better. The same logic that applies to linker selection in round 1 is repeated for rounds 2, 3, and 4. In each round, linkers of length equal to the round number are selected for inclusion in the string candidate. The end of round 4 produces a fully optimized string in terms of epitope ordering, optimal linker selection, and minimized overlap with the human proteome.

[0092] FIG. 4 is a graph showing change of fitness score with the progress of comparing and selecting string candidates over the rounds of designing an exemplary multiepitopic polypeptide, in this case, for HIV vaccine, using the cleavage predictor software described herein. The graph shows progressive optimization and improving the average cleavage score of the polypeptide over several iterations of improvement cycles, till the score plateaus. Round zero, involves several iterations of ordering and reshuffling of the minimal T cell epitopes and sequence fragments containing T cell epitopes at the variable junctions. Round 1 onwards the optimization procedure involves progressive linker addition. The range of scores from bad to good are indicated in the index at the top right.

[0093] FIG. 5 shows proof of concept of the new multiepitopic polypeptide design incorporating the computer implemented cleavage predictor program where cleavage optimization was successfully applied in designing polyepitopic coronavirus mRNA vaccine string. As shown in the figure, epitopes from cleavage optimized regions were observed via targeted and discovery mode via mass spectrometry.

[0094] FIG. 6 shows design of two bivalent HIV vaccine strings, String 1 and String 2 with 33 / 34 epitope using a further improved HIV fragments for each string + C-terminal universal epitope domain (non-HIV specific). The length is drawn to scale; boxes of different shades are HIV fragments, grey is C-terminal fragments, and gaps between boxes represent linkers.

[0095] FIG. 7 shows graphical representation of the experimental validation and successful detection of the epitopes by mass spectrometry at the junctions (edge epitopes) of the fragments or epitopes designed by the cleavage predictor.

[0096] FIG. 8 shows experimental validation and detection of the HIV epitopes by mass spectrometry.

[0097] FIG. 9 shows data an exemplary multiepitopic polypeptide designed using the computer implemented program and tested to see if the epitopes are cleaved when the multiepitopic polypeptideexpressed in a cell. Detection of the epitopes at the junctions are specifically shown in the figure, where, the detection fairly correlates with the predicted cleavage score, and a score of about 2.0 corresponds to detection of the epitope.

[0098] FIG. 10 shows a graphical representation of the configuration of the newly developed machine learning cleavage predictor program (CLEO) disclosed herein. Based on the user’s specifications, this program generates numerous orderings of sequence fragments, generates short linker sequences and maximizes cleavage of defined T cell epitopes in a multiepitopic polypeptide sequence design, an generates sequences and cleavage scores. It takes advantage of the SLURM background and proceed with multi core processing . SLURM, Simple Linux Resource Management.

[0099] FIG. 11 shows expression of the mRNA strings depicted in Table 1 following transfection of each of the string mRNAs in Expi293 cells. This is a study with exemplary Lynch syndrome epitope strings with many different linkers designs, some designed in the cleavage optimizer (CLEO) platform and some with conventional linkers as described below. Expression was analyzed following cell lysis and protein quantitation, detection using LC-MS in the presence or absence of proteasomal inhibitor, Bortezomib. No PI, set without Bortezomib. With PI, set with Bortezomib. Control cells were cells transfection control, where no mRNA or scramble mRNA sequence was transfected. Control: negative experimental control. RNA44: as indicated in Table 1 and further illustrated in Table 2, the CLEO string with CLEO linkers. RNA51, RNA44 shuffled vl, this string has the same neoORFs as RNA44 but the neoepitope are different or are in a different order, and thus the chosen CLEO linkers are different as well. RNA52, RNA 44 shuffled v2: same as above. RNA53, RNA 44 shuffled v3: Same as above. RNA54, string with GS linkers. Instead of using CLEO, this string uses simple GS strings (“GGGGSGGGGS”) between epitopes. These are flexible linkers. RNA55 helical, string with helical linkers. Instead of using CLEO, this string uses helical linkers between epitopes. These are rigid, so it prevents the string from bending. RNA56 Furin RSV, this string has RSV Furin sites. Instead of using CLEO, this string uses a Furin cleavage site derived from RSV (respiratory syncytial virus). Furin sites are an alternative protease pathway to cleave the string. RNA57 Furin: Same as above, but the Furin cleavage site is derived from a naturally occurring human sequence. RNA58 OG, RNA44 string with a MBP (mannose binding protein) sequence at N terminus that increases the solubility of this string. RNA59 Min-ribo, RNA string with Minimal ribosomal stalling modification: RNA44 with 30 point mutations added that minimize ribosomal stalling motifs. RNA60, RNA44 using a codon optimization. RNA61, RNA44 using the different codon optimization.

[0100] FIG. 12 shows representative epitope presentation encoded by the transfected mRNA in a similar experiment as described in FIG. 11. Cell growth and transfection was done similarly to the experiment described for FIG. 11 except that an antibody pulldown was done post cell lysis to obtain A0201 -bound epitopes. The representative epitope sequence is VLDGTVSAV, encoded by the RNA strings.DETAILED DESCRIPTION

[0101] There exists a need to design T cell vaccines, comprising epitopes ordered in the form of amino acid sequences arranged as a continuous string, in such a way that the epitopes have a maximal chance of being cleaved out of the biomolecule produced from the design.

[0102] Accordingly, disclosed herein is a newly developed algorithm capable of handling this design optimization task.

[0103] All terms are intended to be understood as they would be understood by a person skilled in the art. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the disclosure pertains.

[0104] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0105] Although various features of the present disclosure can be described in the context of a single embodiment, the features can also be provided separately or in any suitable combination. Conversely, although the present disclosure can be described herein in the context of separate embodiments for clarity, the disclosure can also be implemented in a single embodiment.

[0106] As used in this specification and claim(s), the words “comprising” (and any form of comprising, such as “comprise” and “comprises”), “having” (and any form of having, such as “have” and “has”), “including” (and any form of including, such as “includes” and “include”) or “containing” (and any form of containing, such as “contains” and “contain”) may be inclusive or open-ended and do not exclude additional, unrecited elements or method steps. It is contemplated that any embodiment discussed in this specification can be implemented with respect to any method or composition of the disclosure, and vice versa. Furthermore, compositions of the disclosure can be used to achieve methods of the disclosure.

[0107] The term “about” or “approximately” as used herein when referring to a measurable value such as a parameter, an amount, a temporal duration, and the like, is meant to encompass variations of + / -20% or less, + / - 10% or less, + / -5% or less, or + / -!% or less of and from the specified value, insofar such variations are appropriate to perform in the present disclosure. It is to be understood that the value to which the modifier “about” or “approximately” refers is itself also specifically disclosed.

[0108] An “agent” can include any type of molecule and includes, but is not limited to, an antibody, a peptide, a protein, a polynucleotide (e.g., an oligonucleotide, RNA, or DNA), a small molecule, derivatives thereof and analogs thereof.

[0109] A “biologic sample” may be any tissue, cell, fluid, or other material derived from an organism. As used herein, the term “sample” may include a biologic sample such as any tissue, cell, fluid, or other material derived from an organism. The term "biological sample" may encompass a variety of sample types obtained from an organism and can be used in a diagnostic or monitoring assay. The term encompasses blood and other liquid samples of biological origin, solid tissue samples,such as a biopsy specimen or tissue cultures or cells derived therefrom and the progeny thereof. The term encompasses samples that have been manipulated in any way after their procurement, such as by treatment with reagents, solubilization, or enrichment for certain components. The term encompasses a clinical sample, and also includes cells in cell culture, cell supernatants, cell lysates, serum, plasma, biological fluids, and tissue samples.

[0110] “Specifically binds” may refer to a condition in which a compound (e.g., peptide) recognizes and binds to a molecule (e.g., peptide or polypeptide), but does not substantially recognize and bind other molecules in a sample, for example, a biological sample, that is, the compound exhibits a selective binding to a molecule. A “binder” as described herein includes, but is not limited to, a protein, a polypeptide or fragments thereof, that exhibits specific binding to a cognate molecule. A binder may refer to an antigen binding domain, such as the first binding domain of a bispecific or trispecific engager, or the second antigen binding domain of a bispecific or trispecific engager, and so on. In some cases, a binder may be any biomolecule or fragment thereof, such as a peptide or conjugated peptide or a ligand that can specifically bind to a receptor on a cell and therefore exhibits specific binding of one portion of an exemplary engager.[OHl] In several cases, an “immune response” may include T cell mediated and / or B cell mediated immune responses that are influenced by modulation of T cell co-stimulation. Exemplary immune responses include T cell responses, e.g., cytokine production, and cellular cytotoxicity. In addition, the term immune response includes immune responses that are indirectly affected by T cell activation, e.g., antibody production (humoral responses) and activation of cytokine responsive cells.

[0112] A "functional derivative" of a native sequence polypeptide may be a compound having a qualitative biological property in common with a native sequence polypeptide. "Functional derivatives" include, but are not limited to, fragments of a native sequence and derivatives of a native sequence polypeptide and its fragments, provided that they have a biological activity in common with a corresponding native sequence polypeptide. The term "derivative" may encompass both amino acid sequence variants of polypeptide and covalent modifications thereof.

[0113] In some cases, a “binder” may refer to a binding domain in a recombinant polypeptide generated by design, as described herein. A binding domain sequence may be derived from a naturally occurring protein and engineered into the recombinant polypeptide by recombinant DNA technology. A binding domain may be selected on the basis of its binding specificity to its target or cognate element. A binding domain may be derived from a protein that is an antibody or a functional fragment thereof, that binds to the target antigen or the cognate molecule. A desired characteristic of a binding domain may be high specificity, high binding affinity or both, towards its target. A binding domain may be termed an engager in that it engages to the target molecule it binds. Accordingly, in some cases a target for a binder may refer to the protein or the polypeptide or the biomolecule to which the binding domain binds. A target may be located on a different cell from the cell on which a binding domainmay be located. In some embodiments, the cell on which the target (e.g., the protein or the biomolecule to which the binder binds), may be referred to as a target cell. In some cases, a binder may not be located on a cell, e.g., a cellular, or may be referred to in such cases as a soluble binder. A binder may be an antibody, or any fragments thereof, an scFV, a sdAb, a VHH.

[0114] Often, “antibody” as used herein may refer to an antibody, an scFv, a VHH, single domain antibody (sdAb), or a protein or polypeptide that comprises an inactive antigen binding domain; wherein the antigen binding capability is designed to be blocked or inactive e.g., by binding a cleavable antigen domain binding polypeptide, until an active step is performed to convert the pro-antibody to its active form. In some embodiments, the active step involves a protease cleavage of the entity that block the antigen binding domain.

[0115] In many instances, the term “affinity” may refer to a quality of ability of one molecule (e.g., a protein molecule) to bind to another molecule or a ligand that binds with chemical specificity. Generally, a molecule a having higher affinity to bind another molecule b than to a third molecule c, would bind more strongly to b and to c. Chemical specificity is the ability of a protein's binding site to bind specific ligands. The fewer ligands a protein can bind, the greater its specificity. Specificity describes the strength of binding between a given protein and ligand. This relationship can be described by a first scFv specific to a cell surface component on a dissociation constant (KD), which characterizes the balance between bound and unbound states for the protein-ligand system.

[0116] In many instances, the term “antigen-presenting cell” or “antigen-presenting cells” or its abbreviation “APC” or “APCs” may refer to a cell or cells capable of endocytosis adsorption, processing and presenting of an antigen. The term includes professional antigen presenting cells, for example, B lymphocytes, monocytes, dendritic cells (DCs) and Langerhans cells, as well as other antigen presenting cells such as keratinocytes, endothelial cells, glial cells, fibroblasts and oligodendrocytes. The term “antigen presenting” may mean the display of antigen as peptide fragments bound to MHC molecules, on the cell surface. Many different kinds of cells may function as APCs including, for example, monocytes or macrophages, B cells, follicular dendritic cells and dendritic cells. APCs can also cross-present peptide antigens by processing exogenous antigens and presenting the processed antigens on class I MHC molecules. Antigens that give rise to proteins that are recognized in association with class I MHC molecules may generally be proteins that are produced within the cells, and these antigens are processed and associate with class I MHC molecules.

[0117] An “epitope” may refer to a portion of an antigen or other macromolecule capable of forming a binding interaction with the variable region binding pocket of an antibody or TCR. The term includes any protein determinant capable of specific binding to an antibody, antibody peptide, and / or antibody-like molecule (including but not limited to a T cell receptor) as defined herein. Epitopic determinants typically consist of chemically active surface groups of molecules such as amino acids or sugar side chains and generally have specific three-dimensional structural characteristics as well asspecific charge characteristics. In some cases, a minimal epitope is mentioned or indicated. A minimal epitope has a sequence which, from N-terminus to C-terminus is the sequence that is understood to bind to an MHC and is proteolytically cleaved inside a cell at these N-terminal and C-terminal ends for MHC binding. Epitopes longer than the minimal epitope may bind to MHCs. A minimal epitope for an MHC -I binding molecule may not promiscuously bind to, for example, MHC -II molecule, and epitope sequences containing the minimal sequence for the MHC -I binding may often bind to MHC- II molecule, thereby generating an opportunity for a CD4+ T cell activation. In some embodiments, minimal epitope may be designed as a candidate epitope or for specific CD8+ T cell response, or for specific binding to a known MHC -I molecule.

[0118] In many instances, the term “antigen” may be any organic or inorganic molecule capable of stimulating an immune response. The term “antigen” as used herein extends to any molecule such as, but not limited, to a peptide, polypeptide, protein, nucleic acid molecule, carbohydrate molecule, organic or inorganic molecule capable of stimulating an immune response.

[0119] In many instances, "antibody" or "antibody moiety" may include, but is not limited to any polypeptide chain-containing molecular structure that recognizes an epitope. Antibodies may utilize in the present invention may be polyclonal antibodies, although monoclonal antibodies are preferred because they may be reproduced by cell culture or recombinantly and can be modified to reduce their antigenicity. The term includes IgG (including IgGl, IgG2, IgG3, and IgG4), IgA (including IgAl and IgA2), IgD, IgE, IgM, andlgY, and is meant to include whole antibodies, including single-chain whole antibodies, and antigen-binding (Fab) fragments thereof. Antigen-binding antibody fragments include, but are not limited to, Fab, Fab' and F(ab')2, Fd (consisting of VH and CHI), single-chain variable fragment (scFv), single-chain antibodies, disulfide-linked variable fragment (dsFv) and fragments comprising either a VL or VH domain. The antibodies can be from any animal origin. Antigen-binding antibody fragments, including single-chain antibodies, can comprise the variable region(s) alone or in combination with the entire or partial of the following: hinge region, CHI, CH2, and CH3 domains. Also included are any combinations of variable region(s) and hinge region, CHI, CH2, and CH3 domains. Antibodies can be monoclonal, polyclonal, chimeric, humanized, and human monoclonal and polyclonal antibodies which, e.g., specifically bind an HLA-associated polypeptide or an HLA- peptide complex. A person of skill in the art will recognize that a variety of immunoaffinity techniques are suitable to enrich soluble proteins, such as soluble HLA-peptide complexes or membrane bound HLA-associated polypeptides, e.g., which have been proteolytically cleaved from the membrane. These include techniques in which, for example, (1) one or more antibodies capable of specifically binding to the soluble protein are immobilized to a fixed or mobile substrate (e.g., plastic wells or resin, latex or paramagnetic beads), and (2) a solution containing the soluble protein from a biological sample is passed over the antibody coated substrate, allowing the soluble protein to bind to the antibodies. The substrate with the antibody and bound soluble protein is separated from the solution,and optionally the antibody and soluble protein are disassociated, for example by varying the pH and / or the ionic strength and / or ionic composition of the solution bathing the antibodies. Alternatively, immunoprecipitation techniques in which the antibody and soluble protein are combined and allowed to form macromolecular aggregates can be used. The macromolecular aggregates can be separated from the solution by size exclusion techniques or by centrifugation.

[0120] The adaptive immune system reacts to molecular structures, referred to as antigens, of the intruding organism. Unlike the innate immune system, the adaptive immune system is highly specific to a pathogen. Adaptive immunity can also provide long-lasting protection; for example, someone who recovers from measles is now protected against measles for their lifetime. There are two types of adaptive immune reactions, which include the humoral immune reaction and the cell-mediated immune reaction. In the humoral immune reaction, antibodies secreted by B cells into bodily fluids bind to pathogen-derived antigens, leading to the elimination of the pathogen through a variety of mechanisms, e.g. complement-mediated lysis. In the cell-mediated immune reaction, T cells capable of destroying other cells are activated. For example, if proteins associated with a disease are present in a cell, they are fragmented proteolytically to peptides within the cell. Specific cell proteins then attach themselves to the antigen or peptide formed in this manner and transport them to the surface of the cell, where they are presented to the molecular defense mechanisms, in T cells, of the body. Cytotoxic T cells recognize these antigens and kill the cells that harbor the antigens.

[0121] In many instances, the term “major histocompatibility complex (MHC)”, “MHC molecules”, or “MHC proteins” may refer to proteins capable of binding antigenic peptides resulting from the proteolytic cleavage of protein antigens inside phagocytes or antigen presenting cells and for the purpose of presentation to and activation of T lymphocytes. Such antigenic peptides may represent T cell epitopes. The human MHC is also called the HLA complex. Thus, the term “human leukocyte antigen (HLA) system”, “HLA molecules” or “HLA proteins” may refer to a gene complex encoding the MHC proteins in humans. The term MHC may be referred to as the “H-2” complex in murine species. Those of ordinary skill in the art would recognize that the terms “major histocompatibility complex (MHC)”, “MHC molecules”, “MHC proteins” and “human leukocyte antigen (HLA) system”, “HLA molecules”, “HLA proteins” are used interchangeably herein.

[0122] HLA proteins are typically classified into two types, referred to as HLA class I and HLA class II. The structures of the proteins of the two HLA classes are very similar; however, they can have different functions. Class I HLA proteins are present on the surface of almost all cells of the body, including most tumor cells. Class I HLA proteins are loaded with antigens that usually originate from endogenous proteins or from pathogens present inside cells and are then presented to naive or cytotoxic T-lymphocytes (CTLs). HLA class II proteins are present on antigen presenting cells (APCs), including but not limited to dendritic cells, B cells, and monocytes or macrophages. They mainly present peptides, which are processed from external antigen sources, e.g., outside of the cells, to helperT cells. Most of the peptides bound by the HLA class I proteins originate from cytoplasmic proteins produced in the healthy host cells of an organism itself, and do not normally stimulate an immune reaction.

[0123] In HLA class II system, phagocytes such as monocytes or macrophages and immature dendritic cells take up entities by phagocytosis into phagosomes - though B cells exhibit the more general endocytosis into endosomes - which fuse with lysosomes whose acidic enzymes cleave the up taken protein into many different peptides. Autophagy is a source of HLA class II peptides. Via physicochemical dynamics in molecular interaction with the HLA class II variants bome by the host, encoded in the host's genome, a particular peptide exhibits immunodominance and loads onto HLA class II molecules. These are trafficked to and externalized on the cell surface. The most studied subclass II HLA genes are: HLA-DPA1, HLA-DPB1, HLA-DQA1, HLA-DQB1, HLA-DRA, and HLA-DRB1.

[0124] Presentation of peptides by HLA class II molecules to CD4+ helper T cells is required for immune responses to foreign antigens. Once activated, CD4+ T cells promote B cell differentiation and antibody production, as well as CD8+ T cell (CTL) responses. CD4+ T cells also secrete cytokines and chemokines that activate and induce differentiation of other immune cells. HLA class II molecules are heterodimers of a- and 0-chains that interact to form a peptide-binding groove that is more open than class I peptide-binding grooves. Peptides bound to HLA class II molecules are believed to have a 9-amino acid binding core with flanking residues on either N- or C-terminal side that overhang from the groove. These peptides are usually 12-16 amino acids in length and often contain 3-4 anchor residues at positions Pl, P4, P6 / 7 and P9 of the binding register (Rossjohn et al., 2015).

[0125] HLA alleles are expressed in codominant fashion, meaning that the alleles (variants) inherited from both parents are expressed equally. For example, each person carries 2 alleles of each of the 3 class I genes, (HLA- A, HLA-B and HLA-C) and so can express six different types of class II HLA. In the class II HLA locus, each person inherits a pair of HLA-DP genes (DPA1 and DPB1, which encode a and 0 chains), HLA-DQ (DQA1 and DQB1, for a and 0 chains), one gene HLA-DRa (DRA1), and one or more genes HLA-DR0 (DRB1 and DRB3, -4 or -5). HLA-DRB1, for example, has more than nearly 400 known alleles. That means that one heterozygous individual can inherit six or eight functioning class II HLA alleles: three or more from each parent. Thus, the HLA genes are highly polymorphic; many different alleles exist in the different individuals inside a population. Genes encoding HLA proteins have many possible variations, allowing each person’s immune system to react to a wide range of foreign invaders. Some HLA genes have hundreds of identified versions (alleles), each of which is given a particular number. In some embodiments, the class I HLA allele is an HLA- A*02:01, HLA-B*14:02, or HLA- A*23: 01, (classical), or HLA-E*01:01 (non-classical). In some embodiments, the class II HLA allele is HLA-DRB*01:01, HLA-DRB*01:02, HLA-DRB*l l:01, HLA-DRB*15:01, or HLA-DRB*07:01.

[0126] Subject specific HLA alleles or HLA genotype of a subject can be determined by any method known in the art In exemplary embodiments, the methods include determining polymorphic gene types that can comprise generating an alignment of reads extracted from a sequencing data set to a gene reference set comprising allele variants of the polymorphic gene, determining a first posterior probability or a posterior probability derived score for each allele variant in the alignment, identifying the allele variant with a maximum first posterior probability or posterior probability derived score as a first allele variant, identifying one or more overlapping reads that aligned with the first allele variant and one or more other allele variants, determining a second posterior probability or posterior probability derived score for the one or more other allele variants using a weighting factor, identifying a second allele variant by selecting the allele variant with a maximum second posterior probability or posterior probability derived score, the first and second allele variant defining the gene type for the polymorphic gene, and providing an output of the first and second allele variant.

[0127] ‘Amino acid” as used herein may be intended to include both natural and synthetic amino acids, and both D and L amino acids. A synthetic amino acid also encompasses chemically modified amino acids, including, but not limited to salts, and amino acid derivatives such as amides. Amino acids present within the polypeptides of the present invention can be modified by methylation, amidation, acetylation or substitution with other chemical groups which can change the circulating half-life without adversely affecting their biological activity.

[0128] In some instances, “operably linked,” “linked,” “fused” or “connected” may be used for interchangeably describing two structural units or subunits being structurally connected with each other. The connection may be direct, that is without any other components between the two, or it may be indirect, wherein one or more linkers as described elsewhere in the specification may be connecting the two units in the description.

[0129] In several occurrences throughout the document, the terms “peptide”, “polypeptide” and “protein” may be used herein interchangeably to describe a series of at least two amino acids covalently linked by peptide bonds or modified peptide bonds such as isosteres. No limitation is placed on the maximum number of amino acids which may comprise a peptide or protein. The terms “oligomer” and “oligopeptide” are also intended to mean a peptide as described herein. Furthermore, the term polypeptide extends to fragments, analogues and derivatives of a peptide, wherein said fragment, analogue or derivative retains the same biological functional activity as the peptide from which the fragment, derivative or analogue is derived.

[0130] A polypeptide as used herein may be a “protein”, including but not limited to a glycoprotein, a lipoprotein, a cellular protein or a membrane protein. A polypeptide may comprise one or more subunits of a protein. A polypeptide may be encoded by a recombinant nucleic acid. In some embodiments, a polypeptide described herein comprises one or more structurally distinct domains. In some embodiments, each domain of a polypeptide described herein may have a distinct function.Usually, a domain is a structural portion of a protein or polypeptide with a defined function. A moiety is a portion of polypeptide, a protein or a nucleic acid, having a specific structure or perform a specific function. For example, a signaling moiety is a specific unit within the larger structure of the polypeptide or protein or a recombinant nucleic acid, which (or the protein portion encoded by it in case of a nucleic acid) engages in a signal transduction process, for example a phosphorylation. In some examples, two or more domains may be required to accomplish a single function. In some instances, the two or more domains required to accomplish a single function may reside on two or more different polypeptides, and the function is accomplished only when a given three-dimensional structure is achieved, for example, by oligomerization and proper orientation of the two or more different polypeptides.

[0131] In several occurrences throughout the document, p-MHC may refer to a structural unit of a combination of an antigenic peptide and its cognate MHC. An MHC molecule is encoded by an HLA gene and may be an MHC class I allele and an MHC class II allele. In general, a peptide associated with an MHC class I molecule is shorter in length (9-11 amino acids) than that associated with an MHC class II molecule. A peptide presented in an MHC class I molecule may generally be presented to and activate a CD8+ T cell. A peptide presented in an MHC class II molecule may generally be presented to and activate a CD4+ T cell. An MHC class I peptide may refer to a peptide that is presented in association with an MHC class I molecule, and an MHC class I associated peptide may interchangeably be called a CD8 peptide or semantic variations thereof. An MHC class II peptide may likewise refer to a peptide that is presented in association with an MHC class II molecule, and an MHC class II associated peptide may interchangeably be called a CD4 peptide or semantic variations thereof.

[0132] As used herein, the term "recombinant nucleic acid molecule" may refer to a recombinant DNA molecule or a recombinant RNA molecule. A recombinant nucleic acid molecule may be any nucleic acid molecule containing joined nucleic acid molecules from different original sources and not naturally attached together. A recombinant nucleic acid may be synthesized in the laboratory. A recombinant nucleic acid can be prepared by using recombinant DNA technology by using enzymatic modification of DNA, such as enzymatic restriction digestion, ligation, and DNA cloning. A recombinant nucleic acid as used herein can be DNA, or RNA. A recombinant DNA may be transcribed in vitro, to generate a messenger RNA (mRNA), the recombinant mRNA may be isolated, purified and used to transfect a cell. A recombinant nucleic acid may encode a protein or a polypeptide. A recombinant nucleic acid, under suitable conditions, can be incorporated into a living cell, and can be expressed inside the living cell. As used herein, “expression” of a nucleic acid usually refers to transcription and / or translation of the nucleic acid. The product of a nucleic acid expression is usually a protein but can also be an mRNA. Detection of an mRNA encoded by a recombinant nucleic acid in a cell that has incorporated the recombinant nucleic acid, is considered positive proof that the nucleic acid is “expressed” in the cell.

[0133] The process of inserting or incorporating a nucleic acid into a cell can be via transformation, transfection or transduction. Transformation is the process of uptake of foreign nucleic acid by a bacterial cell. This process is adapted for propagation of plasmid DNA, protein production, and other applications. Transformation introduces recombinant plasmid DNA into competent bacterial cells that take up extracellular DNA from the environment. Some bacterial species are naturally competent under certain environmental conditions, but competence is artificially induced in a laboratory setting. Transfection is the forced introduction of small molecules such as DNA, RNA, or antibodies into eukaryotic cells. ‘Transfection’ may also refer to the introduction of bacteriophage into bacterial cells. ‘Transduction’ is mostly used to describe the introduction of recombinant viral vector particles into target cells, while ‘infection’ refers to natural infections of humans or animals with wildtype viruses.

[0134] As used herein, the term “vector” may mean any genetic construct, such as a plasmid, phage, transposon, cosmid, chromosome, virus, virion, etc., which is capable transferring nucleic acids between cells. Vectors may be capable of one or more of replication, expression, recombination, insertion or integration, but need not possess each of these capabilities. A plasmid is a species of the genus encompassed by the term “vector.” A vector typically refers to a nucleic acid sequence containing an origin of replication and other entities necessary for replication and / or maintenance in a host cell. Vectors capable of directing the expression of genes and / or nucleic acid sequence to which they are operatively linked are referred to herein as “expression vectors”. In general, expression vectors of utility are often in the form of “plasmids” which refer to circular double stranded DNA molecules which, in their vector form are not bound to the chromosome and typically comprise entities for stable or transient expression or the encoded DNA. Other expression vectors that can be used in the methods as disclosed herein include, but are not limited to plasmids, episomes, bacterial artificial chromosomes, yeast artificial chromosomes, bacteriophages or viral vectors, and such vectors can integrate into the host's genome or replicate autonomously in the cell. A vector can be a DNA or RNA vector. Other forms of expression vectors known by those skilled in the art which serve the equivalent functions can also be used, for example, self-replicating extrachromosomal vectors or vectors capable of integrating into a host genome. Exemplary vectors are those capable of autonomous replication and / or expression of nucleic acids to which they are linked.

[0135] The terms “spacer” or “linker” as used in reference to a fusion protein may refer to a peptide that joins the proteins comprising a fusion protein. In some embodiments, the constituent amino acids of a spacer can be selected to influence some property of the molecule such as the folding, net charge, or hydrophobicity of the molecule. Suitable linkers for use in an embodiment of the present disclosure are well known to those of skill in the art and include, but are not limited to, straight or branched-chain carbon linkers, heterocyclic carbon linkers, or peptide linkers. The linker is used to separate two antigenic peptides by a distance sufficient to ensure that, in some embodiments, eachantigenic peptide properly folds. Exemplary peptide linker sequences adopt a flexible extended conformation and do not exhibit a propensity for developing an ordered secondary structure. Typical amino acids in flexible protein regions include Gly, Asn and Ser. Virtually any permutation of amino acid sequences containing Gly, Asn and Ser would be expected to satisfy the above criteria for a linker sequence. Other near neutral amino acids, such as Thr and Ala, also can be used in the linker sequence.

[0136] In some embodiments, the peptide linkers have more than one functional properties, such as the ones described herein. For example, the peptide linker links two or more functional domains, such as binding domains. Additionally, the peptide linker may be a specific signal inducer when the linker contacts an extracellular portion of a cell, such as a receptor or a ligand binding protein.

[0137] The term “immunopurification (IP)” (or immunoaffinity purification or immunoprecipitation) may refer to a process well known in the art and is widely used for the isolation of a desired antigen from a sample. In general, the process involves contacting a sample containing a desired antigen with an affinity matrix comprising an antibody to the antigen covalently attached to a solid phase. The antigen in the sample becomes bound to the affinity matrix through an immunochemical bond. The affinity matrix is then washed to remove any unbound species. The antigen is removed from the affinity matrix by altering the chemical composition of a solution in contact with the affinity matrix. The immunopurification can be conducted on a column containing the affinity matrix, in which case the solution is an eluent. Alternatively, the immunopurification can be in a batch process, in which case the affinity matrix is maintained as a suspension in the solution. An important step in the process is the removal of antigen from the matrix. This is commonly achieved by increasing the ionic strength of the solution in contact with the affinity matrix, for example, by the addition of an inorganic salt. An alteration of pH can also be effective to dissociate the immunochemical bond between antigen and the affinity matrix.

[0138] As used herein, the terms “determining”, “assessing”, “assaying”, “measuring”, “detecting” and their grammatical equivalents refer to both quantitative and qualitative determinations, and as such, the term “determining” may be used interchangeably herein with “assaying,” “measuring,” and the like. Where a quantitative determination is intended, the phrase “determining an amount” of an analyte and the like is used. Where a qualitative and / or quantitative determination is intended, the phrase “determining a level” of an analyte or “detecting” an analyte may be used.

[0139] A “fragment” may be a portion of a protein or nucleic acid that is substantially identical to a reference protein or nucleic acid. In some embodiments, the portion retains at least 50%, 75%, or 80%, or 90%, 95%, or even 99% of the biological activity of the reference protein or nucleic acid described herein.

[0140] The terms “isolated,” “purified”, “biologically pure” and their grammatical equivalents refer to material that is free to varying degrees from components which normally accompany it as found in its native state. “Isolate” denotes a degree of separation from original source or surroundings.“Purify” denotes a degree of separation that is higher than isolation. A “purified” or “biologically pure” protein is sufficiently free of other materials such that any impurities do not materially affect the biological properties of the protein or cause other adverse consequences. That is, a nucleic acid or peptide of the present disclosure is purified if it is substantially free of cellular material, viral material, or culture medium when produced by recombinant DNA techniques, or chemical precursors or other chemicals when chemically synthesized. Purity and homogeneity are typically determined using analytical chemistry techniques, for example, polyacrylamide gel electrophoresis or high-performance liquid chromatography. The term “purified” can denote that a nucleic acid or protein gives rise to essentially one band in an electrophoretic gel. For a protein that can be subjected to modifications, for example, phosphorylation or glycosylation, different modifications can give rise to different isolated proteins, which can be separately purified.

[0141] A “cancer” may refer to any disease that is caused by or results in inappropriately high levels of cell division, inappropriately low levels of apoptosis, or both. Glioblastoma is one nonlimiting example of a neoplasia or cancer. The terms “cancer” or “tumor” or “hyperproliferative disorder” refer to the presence of cells possessing characteristics typical of cancer-causing cells, such as uncontrolled proliferation, immortality, metastatic potential, rapid growth and proliferation rate, and certain characteristic morphological features. Cancer cells are often in the form of a tumor, but such cells can exist alone within an animal, or can be anon-tumorigenic cancer cell, such as aleukemia cell.

[0142] As used herein, the term “pharmaceutically acceptable” may refer to approved or approvable by a regulatory agency of the Federal or a state government or listed in the U.S. Pharmacopeia or other generally recognized pharmacopeia for use in animals, including humans. A “pharmaceutically acceptable excipient, carrier or diluent” refers to an excipient, carrier or diluent that can be administered to a subject, together with an agent, and which does not destroy the pharmacological activity thereof and is nontoxic when administered in doses sufficient to deliver a therapeutic amount of the agent. A “pharmaceutically acceptable salt” of pooled disease specific antigens as recited herein can be an acid or base salt that is generally considered in the art to be suitable for use in contact with the tissues of human beings or animals without excessive toxicity, irritation, allergic response, or other problem or complication. Such salts include mineral and organic acid salts of basic residues such as amines, as well as alkali or organic salts of acidic residues such as carboxylic acids. Specific pharmaceutical salts include, but are not limited to, salts of acids such as hydrochloric, phosphoric, hydrobromic, malic, glycolic, fumaric, sulfuric, sulfamic, sulfanilic, formic, toluene sulfonic, methane sulfonic, benzene sulfonic, ethane disulfonic, 2-hydroxyethylsulfonic, nitric, benzoic, 2-acetoxybenzoic, citric, tartaric, lactic, stearic, salicylic, glutamic, ascorbic, pamoic, succinic, fumaric, maleic, propionic, hydroxymaleic, hydroiodic, phenylacetic, alkanoic such as acetic, HOOC-(CH2)n-COOH where n is 0-4, and the like. Similarly, pharmaceutically acceptablecations include, but are not limited to sodium, potassium, calcium, aluminum, lithium and ammonium. Those of ordinary skill in the art will recognize from this disclosure and the knowledge in the art that further pharmaceutically acceptable salts for the pooled disease specific antigens provided herein, including those listed by Remington's Pharmaceutical Sciences, 17th ed., Mack Publishing Company, Easton, PA, p. 1418 (1985).

[0143] Nucleic acid molecules useful in the methods of the disclosure may include any nucleic acid molecule that encodes a polypeptide of the disclosure or a fragment thereof. Such nucleic acid molecules need not be 100% identical with an endogenous nucleic acid sequence but will typically exhibit substantial identity. Polynucleotides having substantial identity to an endogenous sequence are typically capable of hybridizing with at least one strand of a double-stranded nucleic acid molecule. “Hybridize” refers to when nucleic acid molecules pair to form a double-stranded molecule between complementary polynucleotide sequences, or portions thereof, under various conditions of stringency. For example, stringent salt concentration can ordinarily be less than about 750 mM NaCl and 75 mM trisodium citrate, less than about 500 mM NaCl and 50 mM trisodium citrate, or less than about 250 mM NaCl and 25 mM trisodium citrate. Low stringency hybridization can be obtained in the absence of organic solvent, e.g., formamide, while high stringency hybridization can be obtained in the presence of at least about 35% formamide, or at least about 50% formamide. Stringent temperature conditions can ordinarily include temperatures of at least about 30° C, at least about 37°C, or at least about 42°C. Varying additional parameters, such as hybridization time, the concentration of detergent, e.g., sodium dodecyl sulfate (SDS), and the inclusion or exclusion of carrier DNA, are well known to those skilled in the art. Various levels of stringency are accomplished by combining these various conditions as needed. In an exemplary embodiment, hybridization can occur at 30° C in 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS. In another exemplary embodiment, hybridization can occur at 37° C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 pg / ml denatured salmon sperm DNA (ssDNA). In another exemplary embodiment, hybridization can occur at 42° C in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 pg / ml ssDNA. Useful variations on these conditions will be readily apparent to those skilled in the art. For most applications, washing steps that follow hybridization can also vary in stringency. Wash stringency conditions can be defined by salt concentration and by temperature. As above, wash stringency can be increased by decreasing salt concentration or by increasing temperature. For example, stringent salt concentration for the wash steps can be less than about 30 mM NaCl and 3 mM trisodium citrate, or less than about 15 mM NaCl and 1.5 mM trisodium citrate. Stringent temperature conditions for the wash steps can include a temperature of at least about 25°C, of at least about 42°C, or at least about 68°C. In exemplary embodiments, wash steps can occur at 25° C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In other exemplary embodiments, wash steps can occur at 42° C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In another exemplary embodiment, wash steps can occur at 68° C in15 mMNaCl, 1.5 mM trisodium citrate, and O.l% SDS. Additional variations on these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.

[0144] In some embodiments, the instant disclosure refers to cellular antigen processing machinery, which is a set of the intracellular functions or events leading up to antigen presentation and T cell activation. Antigen processing may refer to a proteolytic machinery inside a cell, comprising a network of cellular proteins including proteolytic enzymes that break down (e.g., cleave) an antigen into peptide fragments that are epitope sequences that can be loaded on to MHC molecules and presented on to T cells for activation and proliferation, an integral step for activating immune response against the specific antigen. Antigen presenting cells (APCs) present epitopes to T cells which comprise T cell receptors (TCRs), each TCR specific for recognition of an epitope sequence presented on a specific MHC molecule (forming an HLA-peptide pair), upon recognition of the HLA-peptide pair by the TCR a signaling cascade sets off via the TCR complex for T cell activation and proliferation. The proteolytic machinery may be referred to as a proteasome or proteasomal complex. Adequate cleavage of antigens into precise epitope sequence is a crucial first step for intracellular trafficking inside an APC and presentation on a cell surface being bound to an MHC molecule. Epitope sequences are presented in complex with MHC class I alleles to activate CD8+ T cells, thereby activating and generating cytotoxic cells. Epitope sequences presented on MHC class II epitopes generally activate CD4+ T cells, activating memory function.

[0145] “Substantially identical” may be used in reference to comparison between sequences of two or more polypeptide or nucleic acid molecules, for example, the sequence of one polypeptide may be described as exhibiting at least 50% identity to a reference amino acid sequence. A reference sequence may be a sequence disclosed in this specification or referred to as disclosed in another publicly available source. Such a sequence can be at least 60%, 80% or 85%, 90%, 95%, 96%, 97%, 98%, or even 99% or more identical at the amino acid level or nucleic acid level to the sequence used for comparison. Sequence identity is typically measured using sequence analysis software (for example, Sequence Analysis Software Package of the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. In an exemplary approach to determining the degree of identity, a BLAST program can be used, with a probability score between e-3 and e-m° indicating a closely related sequence. A “reference” is a standard of comparison.

[0146] In several occurrences, the term “scaffold” may refer to a molecular platform structure which may be further tweaked to render it suitable for a specific purpose, In the context described herein, a scaffold may refer to a multispecific T cell engager, that may comprise at least one arm that binds to a cell that presents an epitope of interest on its cell surface and at least a second arm that can bind to a T cell that can engage with the cell presenting the epitope and can do one or more of the following: destroy the cell, activate an immune response specific to the epitope, deactivate the cell.

[0147] The term “subject” or “patient” may refer to an animal which is the object of treatment, observation, or experiment. By way of example only, a subject includes, but is not limited to, a mammal, including, but not limited to, a human or a non-human mammal, such as a non-human primate, murine, bovine, equine, canine, ovine, or feline.

[0148] The terms “treat,” “treated,” “treating,” “treatment,” and the like may refer to reducing, preventing, or ameliorating a disorder and / or symptoms associated therewith (e.g., a neoplasia or tumor or infectious agent or an autoimmune disease). “Treating” can refer to administration of the therapy to a subject after the onset, or suspected onset, of a disease (e.g., cancer or infection by an infectious agent or an autoimmune disease). “Treating” includes the concepts of “alleviating”, which refers to lessening the frequency of occurrence or recurrence, or the severity, of any symptoms or other ill effects related to the disease and / or the side effects associated with therapy. The term “treating” also encompasses the concept of “managing” which refers to reducing the severity of a disease or disorder in a patient, e.g., extending the life or prolonging the survivability of a patient with the disease, or delaying its recurrence, e.g., lengthening the period of remission in a patient who had suffered from the disease. It is appreciated that, although not precluded, treating a disorder or condition does not require that the disorder, condition, or symptoms associated therewith be completely eliminated.

[0149] The term “therapeutic effect” refers to some extent of relief of one or more of the symptoms of a disorder (e.g., a neoplasia, tumor, or infection by an infectious agent or an autoimmune disease) or its associated pathology. “Therapeutically effective amount” as used herein may refer to an amount of an agent which is effective, upon single or multiple dose administration to the cell or subject, in prolonging the survivability of the patient with such a disorder, reducing one or more signs or symptoms of the disorder, preventing or delaying, and the like beyond that expected in the absence of such treatment. “Therapeutically effective amount” is intended to qualify the amount required to achieve a therapeutic effect. A physician or veterinarian having ordinary skill in the art can readily determine and prescribe the “therapeutically effective amount” (e.g., ED50) of the pharmaceutical composition required.

[0150] Reference in the specification to “some embodiments,” “an embodiment,” “one embodiment” or “other embodiments” means that a feature, structure, or characteristic described in connection with the embodiments is included in at least some embodiments, but not necessarily all embodiments, of the present disclosure.

[0151] Whereas a cancer cell or a tumor cell may be repeatedly referred here as the target cell, the concepts described here can be suitable for any type of a target cell, such as an infected cell, or a specific disease cell type that needs to be eliminated by the immune cells, as long as the binding domain for a cell surface component of a cancer cell is suitably replaced by a binding domain for a cell surface component specific for the target cell.Multiepitopic Polypeptide (“String”)

[0152] In one aspect, provided herein is a multiepitopic polypeptide that is designed and optimized for e.g., the most efficient epitope cleavage when the multiepitopic polypeptide is expressed in a mammalian cell, e.g., a human cell. For example, an ideal design of an optimized multiepitopic polypeptide is such that each epitope sequence in the multiepitopic polypeptide is cleaved and released for cell surface expression and is presented in an HLA-peptide complex for a chance to activate T cells. However, in reality, all epitope sequences may not be cleaved when the multiepitopic polypeptide is expressed in a mammalian cell. For example, presented herein are optimally designed multiepitopic polypeptides, where further optimization using an optimization scheme described herein leads to redundancy, or the effect of optimization plateaus off.

[0153] A multiepitopic polypeptide as described herein is also called a “string” in short throughout the document. In some embodiments, the multiepitopic polypeptide comprises an optimized ordering of epitopes on the polypeptide from N-terminus to C-terminus. In some embodiments, the multiepitopic polypeptide comprises linkers, wherein a linker may be a short peptide comprising 1-4 amino acids that comprise a cleavable site; wherein a linker may be present in between two adjacent epitopes in the multiepitopic polypeptide and may be absent in between adjacent epitopes in another position along the multiepitopic polypeptide. A string, therefore, is an optimized polypeptide comprising an ordered array of epitopes and with linkers between some adjacent epitopes.

[0154] In some embodiments, a multiepitopic polypeptide has been optimized for the sequence of epitopes in a particular defined order from N-terminus to C-terminus. In some embodiments, the defined order of epitopes in a multiepitopic polypeptide has been further optimized for the position of linkers between two adjacent epitopes and the sequence of amino acids in each linker in between each pair of adjacent epitopes that comprises a linker in between. In some embodiments, one or more adjacent epitopes do not comprise a linker. In some embodiments, provided herein is a multiepitopic polypeptide comprising a defined order of epitopes and linkers, and an N-terminal flanking sequence and a C-terminal flanking sequence. In some embodiments, a multiepitopic polypeptide may not comprise either an N-terminal flanking sequence, or a C-terminal flanking sequence; or may comprise neither an N-terminal flanking sequence, nor a C-terminal flanking sequence.

[0155] In some embodiments, provided herein is a multiepitopic polypeptide sequence that may be represented by a formula from N-terminus to C terminus,NH2-[A]-M-ai-N-a2-O... an-Xn+i-[B]-COOH ... (1)wherein M, N, O... and X represent T cell epitope sequences, ai, a2,... anrepresent linker sequences, n is an integer that is 3 or more; and[A] and [B] are optional N-terminal and C-terminal flanking sequences.

[0156] In some embodiments, there may not be any linkers at a junction of two T cell epitopes or T cell epitope-containing sequences. In some embodiments, the multiepitopic polypeptide comprises at least three epitopes. In some embodiments, the multi epitopic polypeptide comprises at most 100 epitopes. In some embodiments, the multi epitopic polypeptide comprises about 10 to about 100 epitopes. For example, the multiepitopic polypeptide comprises 50 epitopes, according to the Formula (1), n=50, and there are at the most 49 linkers. In some embodiments, there may be linkers with unique sequences between each adjacent epitope sequences. In some embodiments, the multi epitopic polypeptide comprising 50 epitopes may have about 15, about 20, about 30 or more linkers.

[0157] In some embodiments, the multi epitopic polypeptide comprises about 10 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 11 epitopes. In some embodiments, the multi epitopic polypeptide comprises about 12 epitopes. In some embodiments, the multi epitopic polypeptide comprises about 13 epitopes. In some embodiments, the multi epitopic polypeptide comprises about 14 epitopes. In some embodiments, the multi epitopic polypeptide comprises about 15 epitopes. In some embodiments, the multi epitopic polypeptide comprises about 16 epitopes. In some embodiments, the multi epitopic polypeptide comprises about 17 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 18 epitopes. In some embodiments, the multi epitopic polypeptide comprises about 19 epitopes. In some embodiments, the multi epitopic polypeptide comprises about 20 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 21 epitopes. In some embodiments, the multi epitopic polypeptide comprises about 22 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 23 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 24 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 25 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 26 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 27 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 28 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 29 epitopes. In some embodiments, the multi epitopic polypeptide comprises about 30 epitopes. In some embodiments, the multi epitopic polypeptide comprises about 31 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 32 epitopes. In some embodiments, the multi epitopic polypeptide comprises about 33 epitopes. In some embodiments, the multi epitopic polypeptide comprises about 34 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 35 epitopes. In some embodiments, the multi epitopic polypeptide comprises about 36 epitopes. In some embodiments, the multi epitopic polypeptide comprises about 37 epitopes. Insome embodiments, the multi epitopic polypeptide comprises about 38 epitopes. In some embodiments, the multi epitopic polypeptide comprises about 39 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 40 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 41 epitopes. In some embodiments, the multi epitopic polypeptide comprises about 42 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 43 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 44 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 45 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 46 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 47 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 48 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 49 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 50 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 55 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 60 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 70 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 80 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 90 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 95 epitopes. In some embodiments, the multiepitopic polypeptide comprises about 100 epitopes.

[0158] Each epitope is a T cell epitope. The T cell may be a CD8+ T cell. The T cell may be a CD4+ T cell. In some embodiments, a T cell epitope is a CD8+ T cell epitope comprising 8 amino acids. In some embodiments, a T cell epitope is a CD4+ T cell epitope comprising about 11-15 amino acids. Each epitope in a string is a prevalidated T cell epitope, that is a CD8+ T cell string or a CD4+ T cell epitope. A prevalidated T cell epitope may be an epitope that is predicted to bind to an MHC molecule, wherein, the predicted binding is according to a well- established peptide-MHC binding prediction algorithm (e.g., NetMHC-Pan 3.0 and higher versions; RECON, etc.). In some embodiments, the peptide-MHC binding prediction algorithm is one that functions reliably, for example, a prevalidated T cell epitope may be an epitope that is predicted to bind to an MHC by the NetMHC-Pan algorithm, or RECON algorithm. According to one objective of the present disclosure, an epitope in the multiepitopic polypeptide is designed for efficient cleavage of the epitope sequence when the multi epitopic polypeptide is expressed in a cell, the cleavage being by the cell’s proteolytic machinery, and the proteolytic cleavage inside the cell generates the epitope sequence, not with even one additional amino acid more or even one amino acid less, to ensure that the epitope is expressed on the cell surface, is presented in an HLA-peptide complex and activates a T cell when the cell contacts with the T cell, as determined by a T cell activation assay, e.g., a cytokine release assay. For example, a cytokine release assay may include release of IL2 by the activated T cell.

[0159] In some embodiments the multi epitopic polypeptide comprises from about 100 to about 10,000 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 9,000 amino acids. In some embodiments the multiepitopic polypeptide comprises from about 100 to about 8,000 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 7,000 amino acids. In some embodiments the multiepitopic polypeptide comprises from about 100 to about 6,000 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 5,000 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 4,000 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 3,000 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 2,900 amino acids. In some embodiments the multiepitopic polypeptide comprises from about 100 to about 2,800 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 2,700 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 2,600 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 2,500 amino acids. In some embodiments the multiepitopic polypeptide comprises from about 100 to about 2,400 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 2,300 amino acids. In some embodiments the multiepitopic polypeptide comprises from about 100 to about 2,200 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 2,100 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 2,000 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 1,900 amino acids. In some embodiments the multiepitopic polypeptide comprises from about 100 to about 1,800 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 1,700 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 1,600 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 1,500 amino acids. In some embodiments the multiepitopic polypeptide comprises from about 100 to about 1,400 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 1,300 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 1,200 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 1,100 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 1,000 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 900 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 800 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 700 amino acids. In some embodiments the multi epitopic polypeptide comprises from about 100 to about 600 amino acids. In some embodiments the multiepitopic polypeptide comprises about 400 amino acids. In someembodiments the multiepitopic polypeptide comprises about 420 amino acids. In some embodiments the multiepitopic polypeptide comprises about 440 amino acids. In some embodiments the multiepitopic polypeptide comprises about 460 amino acids. In some embodiments the multiepitopic polypeptide comprises about 480 amino acids. In some embodiments the multiepitopic polypeptide comprises about 500 amino acids. In some embodiments the multiepitopic polypeptide comprises about 520 amino acids. In some embodiments the multiepitopic polypeptide comprises about 540 amino acids. In some embodiments the multiepitopic polypeptide comprises about 560 amino acids. In some embodiments the multi epitopic polypeptide comprises about 580 amino acids. In some embodiments the multiepitopic polypeptide comprises about 600 amino acids. In some embodiments the multiepitopic polypeptide comprises about 620 amino acids. In some embodiments the multiepitopic polypeptide comprises about 640 amino acids. In some embodiments the multiepitopic polypeptide comprises about 660 amino acids. In some embodiments the multiepitopic polypeptide comprises about 680 amino acids, or about 700 amino acids.

[0160] In some embodiments, one objective of the work presented herein is to utilize the novel computer framework in the design and manufacture of a polyspecific and polyfunctional population of antigen specific T cells that have been stimulated and activated with multiepitopic polypeptide strings encoded by polynucleotide strings comprising nucleotide sequences encoding multiple T cell epitopes and linkers in a manner designed and / or selected by the framework program, such that all the T cell epitopes are presented by antigen presenting cells following endogenous processing of the epitopes in the antigen presenting cells, for effective stimulation that result in T cell responses spanning across the encoded epitopes including but not limited to exhibiting cytotoxic activity against candidate disease cells that present the antigens comprising the epitopes in a subject in need thereof. Accordingly, T cell epitope polypeptide strings are designed against HIV, EBV, Lynch, PRAME and Glioblastoma (GB) candidates till date, using the program. For example, HIV vaccine targeting both Clade B and Clade C HIV epitopes have been designed with a goal to generate CD8 T cell therapeutics. Exemplary 16 HIV strings have been generated. 56 strings have been generated for personalized cancer epitopes for CD8 / CD4 cells from 6 patients. With a goal to generate shared / off the shelf prophylactic vaccine, 22 exemplary strings have been generated spanning 250 candidates. 20 EBV specific mRNA vaccine strings were also generated.Linker

[0161] A linker sequence may be designed in between multiple peptides, wherein the peptides may represent sequence fragments containing at least one T cell epitope sequence, or a minimal epitope sequence. A linker comprises of amino acids, held by peptide linkages within the linker and with two adjacent T cell epitope sequences, one that is at the N-terminal side and one that is at the C- terminal side of the linker. The linker may enable or generate a higher likelihood of creation and presentation of the epitope while hinder creation of deleterious epitopes, such as those created at thejunction suture between adjacent peptides or between linker sequence and endogenous peptides.. It may be an objective of the design, for example by user intervention prior or during the process, that "junction" epitopes may not only compete with the intended epitopes to be presented on the cell surface, decreasing vaccine efficacy. In some embodiments, random cleavage could generate an unwanted auto-immune reaction. Thus, the designed linkers may be fine-tuned to be linker sequences that may a) avoid creating "junction" peptides from forming or may b) avoid proteasomal processing to create "junction" peptides. In some embodiments, it is desired to avoid creation of "junction" peptides that bind MHC molecules. Glycine, for example, inhibits strong binding in MHC binding groove positions. Multiple linker sequences and multiple linker lengths were designed and analyzed and calculated the number of "junction" peptides that bind MHC molecules. The software tools used herein comprise a variety of tools and databases, including those from the Immune Epitope Database (IEDB, http: / / www.immuneepitope.org / ) to calculate the likelihood that a given peptide sequence contains a ligand that will bind MHC Class I molecules. Such linker sequences as presented herein are distinct from commonly used GS linkers, for example, GS4 linkers, but designed specifically by computing and reserving the best fit according to a fitness score for any position of the linker between any two randomly selected adjacent T cell epitopes. An objective of the instant disclosure is the design of a polypeptide sequence with linkers between peptides, wherein the peptides comprise at least one T cell epitope sequence, and the linker facilitates or enhances intracellular epitope processing and presentation. As described herein, a linker sequence at a specific position juxtaposing two peptides is designed through the cleavage optimizer (CLEO) program and may be referred to simply as a linker or at times a cleavage optimizer linker (or, CLEO linker for the purposes of the disclosure herein) specifically differentiating it from a conventional linker, such as a GS linker. As is well understood by context, a linker described below is a CLEO linker.

[0162] A linker as provided herein is a short peptide comprising at most 4 amino acids, and is represented by the formulaX1X2X3X4 .... (2) wherein each X represents any amino acid, wherein each of Xi, X2, X3, and X4 is present or absent.

[0163] In some embodiments, two adjacent sequence fragments, or minimal T cell epitope sequences of the sequence having the defined N-terminus to C-terminus order may be adjacent to each other, and a selected linker sequence may be inserted at the junction that increases the probability of cleavage of the T cell epitopes that are immediately adjacent to / on either side of the linker, wherein the selected linker sequence comprises an amino acid Xi, wherein X2, X3, X4 are absent; wherein Xi is modified to a C-terminal amino acid of a T cell epitope sequence of the two adjacent T cell epitope sequences that is upstream of the linker, and is further connected to an N-terminal amino acid of a T cell epitope sequence of the two adjacent T cell epitope sequences that is downstream of the linker. In some embodiments, a proteolytic cleavage site may be designed to occur either before or after XI . Forthe purpose of the following few paragraphs on the linker, a proteolytic cleavage before or after ‘X’ may be understood as a cleavage between X and the amino acid immediately N-terminal to X; and cleavage between X and the amino acid immediately C-terminal to X respectively.

[0164] In some embodiments, two adjacent T cell epitope sequences of the sequence having the defined N-terminus to C-terminus order is further modified by introducing a selected linker sequence, wherein the selected linker sequence comprises an amino acid sequence according to a formula of X1X2; and wherein Xs and X4 are absent; wherein Xi is modified to a C-terminal amino acid of a T cell epitope sequence of the two adjacent T cell epitope sequences that is upstream of the linker, and X2 is modified to an N-terminal amino acid of a T cell epitope sequence of the two adjacent T cell epitope sequences that is downstream of the linker. In some embodiments, a proteolytic cleavage may be designed to occur either before or after XI or after X2.

[0165] In some embodiments, two adjacent T cell epitope sequences of the sequence having the defined N-terminus to C-terminus order is further modified by introducing a selected linker sequence, wherein the selected linker sequence comprises an amino acid sequence according to a formula of X1X2X3; and wherein X4 is absent; wherein Xi is modified to a C-terminal amino acid of the T cell epitope sequence of the two adjacent T cell epitope sequences that is upstream of the linker, and X3 is modified to an N-terminal amino acid of a T cell epitope sequence of the two adjacent T cell epitope sequences that is downstream of the linker. In some embodiments, a proteolytic cleavage may be designed to occur either before or after XI, before or after X2, or after X3.

[0166] In some embodiments, two adjacent T cell epitope sequences of the sequence having the defined N-terminus to C-terminus order is further modified by introducing a selected linker sequence, wherein the selected linker sequence comprises an amino acid sequence having the formula X1X2X3X4, wherein Xi is modified to a C-terminal amino acid of a T cell epitope sequence of the two adjacent T cell epitope sequences that is upstream of the linker, and X4 is modified to an N-terminal amino acid of a T cell epitope sequence of the two adjacent T cell epitope sequences that is downstream of the linker.

[0167] In some embodiments, the linker sequence comprises at the most four amino acids. In some embodiments, a proteolytic cleavage may be designed to occur either before or after XI, before or after X2, before or after X3, or after X4.

[0168] In some embodiments, two adjacent T cell epitope sequences may not have any linker sequences. For example, there may be a fragment of the multiepitopic polypeptide wherein 2, 3, 4, 5, 6, 7 or more adjacent epitopes lack any linker sequence and may be placed consecutively, each next to the other. Such positions between two epitopes lacking a linker have been determined to contain a proteolytic cleavage site of its own and are called fixed positions. Other positions in a given order of epitopes along early phase of a string design may comprise variable positions, such that the positionin between those two adjacent epitopes is not fixed, a linker sequence may be introduced, or the order of epitopes around that position may be altered or reshuffled.

[0169] In one aspect, a linker is selected for a position (e.g., a variable position) between two adjacent epitopes only if it satisfies the criteria that a combined sequence is not being present in the human proteome, wherein the combined sequence comprises: (i) the selected linker sequence and (ii) the sequence immediately upstream and / or downstream of the selected linker sequence. For example, a sequence upstream of the linker is a sequence that comprises the epitope sequence and beyond the N-terminus of the linker. For example, a sequence downstream of the linker is a sequence that comprises the epitope sequences and beyond the C-terminus of the linker. In some embodiments, the linker comprising a sequence of 4 amino acids, represented by the formula X1X2X3X4, may be taken together with, for example, up to 30 amino acids N-terminal to XI, and up to 30 amino acids to the C- terminus of X4 and is considered for the selection criteria, such that no stretch of 8 consecutive amino acids is found in the human proteome. Each 8 consecutive amino acid sequence from this combined sequence above is matched against the human proteome, and the selected linker at that position does not contain a match. In some embodiments, up to about 20 amino acids N-terminal to XI, and up to 20 amino acids to the C-terminus of X4 (or the C terminal amino acid of the linker, which could be any one of XI, X2, X3 or X4) and is considered for the selection criteria, such that no stretch of 8 consecutive amino acids is found in the human proteome. In some embodiments, up to about 15 amino acids N-terminal to XI, and up to 15 amino acids to the C-terminus of X4 (or the C terminal amino acid of the linker, which could be any one of XI, X2, X3 or X4) and is considered for the selection criteria, such that no stretch of 8 consecutive amino acids is found in the human proteome. In some embodiments, up to about 10 amino acids N-terminal to XI, and up to 10 amino acids to the C-terminus of X4 (or the C terminal amino acid of the linker, which could be any one of XI, X2, X3 or X4) and is considered for the selection criteria, such that no stretch of 8 consecutive amino acids is found in the human proteome. Each 8 consecutive amino acid sequence from this combined sequence above is matched against the human proteome, and the selected linker at that position is the linker that does not contain a match.

[0170] Exemplary cleavage optimizer (CLEO) linkers may contain an amino acid sequence of any one of the following from a non-exhaustive list: A, AA, AAR, RKY, LYM, RRY, AG, RAA, RMQ, GRTY, GRAA, AKN, HTY, KSA, RGA, ALNM, RGSY, RNS, HA, RNY, SM, AYN, AYH, ARA, RMR, AANA, AAAS, AMNM, GMNY, AM, KMH, HQY, S, ARGA, RHA, ARA, AAG, AAAG, ATN, KRSA, AKY, SRNY, KMK, ARN, GQAA, TL, RRRY, RSA, RYA, VRH, ARAA, RMYM, SRAY, SY A, YMA, AGA, NR, SGA, AG, KMA, SMY, KM, KMHH, HKY, RVY, SRRY, AAG, SQY, TYA, KA, AM, RKHY, HAG, GANA, MRA, RMAA, SMY, SR, ARNN, AQA, RMA, ASNY, RAG, RM, YYK, NMA, RMA, AMN, RQR, SG, RMS, NKMA, HRM, RK, AR, QMMA, SHR, SM, KMHK, MRQY, M, RRA, AG, AYGA, SAM, RHY, HAAA, AMG, AMA, SMA, SRY,RKA, SNGG, AMNG, GSN, RRM, GR, SKS, HRGY, AQMN, RRYR, MYS, RMA, AQRY, GRNY, ARAR, AMA, ST, SM, and KMHA. A cleavage optimizer linker may be as disclosed anywhere throughout the specification. A CLEO linker may comprise up to 4 amino acids. A CLEO linker may comprise any amino acid. A CLEO linker may comprise any naturally occurring amino acid. A CLEO linker may comprise any one or up to four of the 20 naturally occurring amino acids.

[0171] Other linkers may be used in addition to CLEO linkers. As described before, other linkers include conventional linkers and may include GS linkers, for example GGGGSGGGGS. Yet another linker may be a helical linker. Exemplary helical linker may comprise the sequence EKAAKAEEAAR. Yet other linkers may also comprise Furin site linkers of viral origin (e.g., respiratory syncytial virus, RSV), for example, a linker having a sequence GIRRKRSVSH. A linker with a Furin site may comprise human Furin site, for example, having a linker sequence RTKRELE. Exemplary epitopes and linkers may be found in Tables 1-11.

[0172] In some embodiments, a polypeptide may comprise other modifications, such as an N- terminal MBP (mannose binding protein) sequence. In some embodiments, addition of MBP sequence at the N terminus may increase solubility of the polypeptide.Polynucleotides

[0173] In one aspect, provided here is a method for preparing a polynucleotide encoding a polypeptide wherein the polypeptide is a multiepitopic polypeptide, wherein in some embodiments the polynucleotide is a ribonucleic acid (RNA). In some embodiments, the RNA is a messenger RNA (mRNA). The multiepitopic polypeptide described herein comprises a defined ordering of T cell epitope sequences and cleavable linkers and is often described as a string. The RNA encoding a string is designated as a string RNA or simply string where it is easily discerned whether the string refers to the RNA or the polypeptide. The RNA encoding a multiepitopic polypeptide (e.g., a string) comprises a long in vitro transcribed mRNA which consists of sequentially arranged sequences coding for the mutated peptides and one or more linker sequences at certain junctions of T cell epitopes or sequence fragments. The coding sequences are chosen from the non- synonymous mutations and are always built up of the codon for the mutated amino acid that may or may not be flanked by regions from the original sequence context. The linker sequence codes for amino acids that are preferentially not processed by the cellular antigen processing machinery. In some embodiments the string encodes a polypeptide comprises from about 100 to about 10,000 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 9,000 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 8,000 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 7,000 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 6,000 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 5,000 aminoacids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 4,000 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 3,000 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 2,900 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 2,800 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 2,700 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 2,600 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 2,500 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 2,400 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 2,300 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 2,200 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 2,100 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 2,000 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 1,900 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 1,800 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 1,700 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 1,600 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 1,500 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 1,400 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 1,300 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 1,200 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 1,100 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 1,000 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 900 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 800 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 700 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 600 amino acids. In some embodiments the string encodes a polypeptide comprises from about 100 to about 500 amino acids.

[0174] In some embodiments, the RNA is in vitro transcribed RNA (IVT-RNA) and may be obtained by in vitro transcription of an appropriate DNA template, wherein the in vitro transcription occurs with the help of a vector comprising a promoter. In some embodiments, the promoter for controlling transcription can be any promoter for any RNA polymerase. A DNA template for in vitro transcription may be obtained by cloning of a nucleic acid, in particular cDNA, and introducing it into an appropriate vector for in vitro transcription. The cDNA may be obtained by reverse transcriptionof RNA. In vitro transcription constructs are usually based on the pSTl-A120 vector containing a T7 promotor, a tandem beta-globin 3' UTR sequence and a 120-bp poly (A) tail, which have been shown to increase the stability and translational efficiency of the RNA thereby enhancing the T-cell stimulatory capacity of the encoded T cell epitope. In some embodiments, the poly-A tail comprises a poly-A sequence of at least 50 A residues, at least 80, or at least 100 and up to 500, up to 400, up to 300, up to 200, or up to 150 A nucleotides, and, in particular, about 120 A nucleotides. In some embodiments, the RNAs comprises a 5'-cap structure. In some embodiments, the RNAs comprises a modified cap. In some embodiments, the RNAs comprises a 5’ cap m27’2'° Gpp5p(5’)G. In some embodiments, the RNA comprises a cap analog anti-reverse cap (ARCA Cap(m27,2'°G(5’)ppp(5’)G), (CapO).

[0175] In some embodiments, the RNA may have modified nucleosides. In some embodiments, the RNA comprises a modified nucleoside in place of at least one (e.g., every) uridine, for example a pseudouridine. A pseudouridine (T) may be a modified nucleoside that is an isomer of uridine, where the uracil is attached to the pentose ring via a carbon-carbon bond instead of a nitrogen-carbon glycosidic bond. In some embodiments, the RNA comprises a modified nucleoside N1 -methylpseudouridine (m IT). In some embodiments, the RNA comprises a modified nucleoside 5-methyl- uridine m5U). In some embodiments, the modified nucleoside replacing one or more uridine in the RNA may be any one or more of 3-methyl-uridine (m3U), 5 -methoxy-uridine (mo5U), 5 -aza-uridine, 6-aza-uridine, 2-thio-5-aza-uridine, 2-thio-uridine (s2U), 4-thio-uridine (s4U), 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxy-uridine (ho5U), 5-aminoallyl-uridine, 5-halo uridine (e.g., 5-iodo- uridineor 5 -bromo-uridine), uridine 5-oxyacetic acid (cmo5U), uridine 5-oxyacetic acid methyl ester (mcmo5U), 5-carboxymethyl-uridine (cm5U), 1 -carboxymethyl- pseudouridine, 5- carboxyhydroxymethyl-uridine (chm5U), 5-carboxyhydroxymethyl-uridine methyl ester (mchm5U), 5 -methoxy carbonylmethyl-uri dine (mcm5U), 5-methoxycarbonylmethyl-2-thio-uridine (mcm5s2U), 5- aminomethyl-2-thio-uridine (nm5s2U), 5-methylaminomethyl-uridine (mnm5U), 1-ethyl- pseudouridine, 5-methylaminomethyl-2- thio-uridine (mnm5s2U), 5-methylaminomethyl-2-seleno- uridine (mnm5se2U), 5-carbamoylmethyl-uridine (ncm5U), 5-carboxymethylaminomethyl-uridine (cmnnfU), 5-carboxymethylaminomethyl-2-thio-uridine (cmnm5s2U), 5-propynyl-uridine, 1- propynyl-pseudouridine, 5-taurinomethyl-uridine (tm5U), 1 -taurinomethyl-pseudouridine, 5- taurinomethyl-2-thio-uridine(tm5s2U), 1 -taurinomethyl-4-thio-pseudouridine), 5-methyl-2-thio- uridine (m5s2U), 1 -methyl-4-thio-pseudouridine (m's4TI )), 4-thio-l-methyl-pseudouridine,3-methyl- pseudouridine (m3T), 2-thio-l-methyl-pseudouridine, 1 -methyl- 1 -deaza-pseudouri dine, 2-thio-l- methyl-l-deaza-pseudouridine, dihydrouridine (D), dihydropseudouridine, 5,6-dihydrouridine, 5- methyl-dihydrouridine (m5D), 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxy- uridine, 2-methoxy-4-thio-uridine,4-methoxy -pseudouridine, 4-methoxy-2-thio-pseudouridine, Nl- methyl-pseudouridine, 3-(3-amino-3-carboxypropyl)uridine (acp3U), l-methyl-3-(3-amino-3-carboxypropyl)pseudouridine (acp3vP), 5-(isopentenylaminomethyl)uridine (inm5U), 5- (isopentenylaminomethyl)-2-thio-uridine (inm5s2U), a-thio-uridine, 2'-0-methyl-uridine (Um), 5,2'- O-dimethyl-uridine (m5Um), 2'-O-methyl-pseudouridine ((Pm), 2-thio-2'-0-methyl-uridine (s2Um), 5-methoxycarbonylmethyl-2'-0-methyl-uridine (mcm5Um), 5-carbamoylmethyl-2'-O-methyl-uridine (ncm5Um), 5-carboxymethylaminomethyl-2'-O-methyl-uridine (cmnm5Um), 3,2'-0-dimethyl-uridine (m3Um), 5-(isopentenylaminomethyl)-2'-0-methyl-uridine (inm5Um), 1 -thio-uridine, deoxy thy mi dine, 2'-F-ara-uridine, 2'-F -uridine, 2' -OH-ara-uridine, 5-(2-carbomethoxyvinyl) uridine, 5-[3-(l-E-propenylamino)uridine, or any other modified uridine known in the art.

[0176] In some embodiments, at least one RNA comprises a modified nucleoside in place of at least one uridine. In some embodiments, at least one RNA comprises a modified nucleoside in place of each uridine. In some embodiments, each RNA comprises a modified nucleoside in place of at least one uridine. In some embodiments, each RNA comprises a modified nucleoside in place of each uridine. In some embodiments, the modified nucleoside is independently selected from pseudouridine ( ), Nl-methyl-pseudouridine (m l'P), and 5-methyl-uridine (m5U). In some embodiments, the modified nucleoside comprises pseudouridine (T). In some embodiments, the modified nucleoside comprises Nl-methyl-pseudouridine (mlT). In some embodiments, the modified nucleoside comprises 5-methyl-uridine (m5U). In some embodiments, at least one RNA may comprise more than one type of modified nucleoside, and the modified nucleosides are independently selected from pseudouridine Nl-methyl-pseudouridine (ml ), and 5- methyl-uridine (m5U). In some embodiments, the modified nucleosides comprise pseudouridine ('P) and Nl-methyl-pseudouridine (mlT). In some embodiments, the modified nucleosides comprise pseudouridine (VP) and 5-methyl-uridine (m5U). In some embodiments, the modified nucleosides comprise Nl-methyl-pseudouridine (mlT) and 5- methyl-uridine (m5U). In some embodiments, the modified nucleosides comprise pseudouridine (4)), Nl- methyl-pseudouridine (mlT), and 5-methyl uridine (m5U).

[0177] In some embodiments the RNA is codon optimized. Several codon optimized sequences are provided in the disclosure. Codon optimization may be done by available tools and methods such as provided by Linear Design (LD) and DERNA (J Comput Biol. 2024 Mar;31(3): 179-196. doi: 10.1089 / cmb.2023.0283).

[0178] An MHC class I signal peptide fragment and the transmembrane and cytosolic domains including the stop-codon is usually included in RNA encoding a multiepitopic polypeptide. In some embodiments, MHC class I trafficking signal (or MITD) sequence is included, flanking a poly-linker sequence for cloning the epitopes were inserted. The latter have been shown to increase the antigen presentation, thereby enhancing the expansion of antigen-specific CD8+ and CD4+ T cells and improving effector functions. In some embodiments, at least two of the peptide epitopes are separated from one another by a universal type II T-cell epitope. In one embodiment, all of the peptide epitopes are separated from one another by a universal type II T-cell epitope. In another embodiment, themRNA cancer vaccine encodes 1-20 universal type II T-cell epitopes. In some embodiments, the universal type II T- cell epitope comprises one or more tetanus toxoid-derived Helper sequences, for example, a sequence ILMQYIKANSKFIGI, or a sequence having 1, 2, or 3 amino acids differing from the sequence; a sequence FNNFTVSFWLRVPKVSASHLE, or a sequence having 1, 2, or 3 amino acids differing from the sequence; a Diptheria toxin epitope sequence QYIKANSKFIGITE, or QSIALSSLMVAQAIP, or a sequence having 1, 2, or 3 amino acids differing from the sequence; or a pan-DR epitope AKFVAAWTLKAAA. In some embodiments, the universal type II T-cell epitope is the same universal type II T-cell epitope throughout the mRNA. In some embodiments, the universal type II T-cell epitope is repeated 1-20 times in the mRNA. In another embodiment, the universal type II T-cell epitopes are different from one another throughout the mRNA. In some embodiments, the universal type II T-cell epitope is located between every other peptide epitope. In one embodiment, the universal type II T-cell epitope is located between every third peptide epitope. In some embodiments, the universal type II T-cell epitope is present only at the C terminal end upstream of the poly A sequence. In some embodiments, a universal epitope breaks immunological tolerance when administered to a subject, in combination with a T cell vaccine such as described herein, comprising the multiepitopic polypeptide or the polynucleic acid comprising a sequence encoding the multiepitopic polypeptide. In some embodiment, each RNA comprises a sequence encoding a tetanus toxin epitope sequence at a position that is 3 ’ of all sequences encoding epitopes of the multi epitopic polypeptide selected for treating a disease in a subject. The tetanus toxoid derived helper sequences comprise helper epitopes p2 (QYIKANSKFIGITEL; (tetanus toxoid TT) amino acids 830-844) and pl 6 (MTNSVDDALINSTKIYSYFPSVISKVNQGAQG; TT575 609) (p2pl6), which are incorporated in the RNA sequence at the position that is 3 ’ of all sequences encoding epitopes of the multi epitopic polypeptide selected for treating a disease in a subject. In some embodiments, the p2pl6 sequence is incorporated at a sequence upstream of the 3’UTR of the polynucleic acid.

[0179] In some embodiments, the RNA is formulated to be delivered via a lipid nanoparticle (LNP) delivery vehicle into a cell of the subject. In some embodiments the cell of the subject is in vivo. In some embodiments the RNA is formulated to administered into the bloodstream of a subject in an aqueous composition in which the RNA is encapsulated in an LNP delivery vehicle. In some embodiments the subject is a human.

[0180] Thus, in one aspect, provided herein is a composition comprising a vaccine, encapsulated by a lipid nanoparticle, wherein the vaccine is an RNA encoding a plurality of T cell epitopes that are a disease associated antigens, and further comprises a sequence that encodes an amino acid sequence that breaks immunological tolerance. In some embodiments, the sequence that encodes an amino acid sequence that breaks immunological tolerance comprises the p2pl6 tetanus toxoid-derived helper epitope sequences. In some embodiments, the plurality of T cell epitopes that are a disease associated-antigens comprise cancer antigens. In some embodiments, plurality of T cell epitopes that are a disease associated-antigens comprise antigens associated with a viral disease, e.g., HIV.LNP

[0181] In one aspect, provided herein is an RNA encoding a plurality of epitopes, e.g., T cell epitopes, formulated for administration to a subject in need thereof, in the form of an injectable or infusable composition, wherein the RNA is encapsulated in an LNP. In some embodiments, the LNP comprises a cationic lipid, and typically a non-cationic lipid or a neutral lipid. In some embodiments, the cationic lipid comprises l,2-di-O-octadecenyl-3 -trimethylammonium propane (DOTMA) and / or l,2-dioleoyl-3-trimethylammonium-propane (DOTAP). In some embodiments, the non-cationic lipid comprises l,2-di-(9Z-octadecenoyl)-sn-glycero-3 -phosphoethanolamine (DOPE), cholesterol (Choi) and / or l,2-dioleoyl-sn-glycero-3-phosphocholine (DOPC).

[0182] In some embodiments, the LNP in the composition generally comprises particles of uniform particle shape and diameter, generally ranging from an average diameter of about 200nm to 1 OOOnm. In some embodiments, the average LNP diameter is about 200 nm. . In some embodiments, the average LNP diameter is about 250 nm. In some embodiments, the LNP comprises a poly dexterity index of about 0.5 or less, about 0.4 or less, or about 0.3 or less and typically ranges between 0.1-0.3. Cleavage Predictor for Informed T Cell Epitope String Design: General Overview of the Method

[0183] In one aspect, provided herein is a method for designing, accurately predicting and optimizing a multiepitopic polypeptide (string) design for generating a polypeptide for a therapeutic vaccine product. In some embodiments, the method incorporates a new in-house generated, trained machine-learning algorithm that is capable of ordering a given set of epitopes in an N-terminal to C- terminal order and annealing in silico thereby generating a polypeptide having a defined N-terminal to C-terminal sequence. In addition, the trained machine-learning algorithm generates a prediction of a likelihood of cleavage of an epitope for antigen presentation from within a polypeptide sequence when inside a cell. In some embodiments, the likelihood of cleavage of an epitope is a measure of the efficiency of cleavage, or may interchangeably be understood as the degree of cleavage, the propensity of cleavage, or for example the chances of cleavage, and for which the machine learning algorithm is able to generate a quantitative score, wherein a higher score indicates the better changes / degree / likelihood of cleavage of the epitope from the polypeptide inside a cell. The method incorporates the algorithms to accurately predict an improved multiepitopic polypeptide sequence that is used as a T cell vaccine, wherein the polypeptide when expressed in vivo can activate T cells against said epitopes in the multiepitopic polypeptide.

[0184] The method, in some embodiments, further encompasses generating a multiepitopic polypeptide from the sequence generated by running the trained algorithm, and preparing a vaccine, and a method of treating a subject by administering the vaccine to a subject in need thereof. In someembodiments, the subject is a human subject. In some embodiments, the subject has a disease, and the multiepitopic polypeptide is a T cell vaccine to treat the disease.

[0185] Additionally, in some embodiments, the method incorporates providing a user defined set of epitopes as input for the algorithm. In some embodiments, the method incorporates user defined set of minimal epitopes and sequences comprising T cell epitopes (e.g., sequence fragments). In some embodiments, the method incorporates user defined input designating one or more T cell epitope junctions or sequence fragment junctions or edges as variable junctions; variable junctions are those designated junctions in a selection of T cell epitopes and sequences containing T cell epitopes at the time of input, at which junctions, a swapping of order of fragments or minimal epitopes may take place (e.g., ordering, reshuffling) to improve the score of the multiepitopic polypeptide during the progress of the method. Other than variable junctions, there may be one or more T cell epitope junctions where the order of the sequence from an N-terminal to C-terminal direction is invariable or fixed, where reordering of the junctions is not permissive.

[0186] In one embodiment, provided herein is a method of training a machine-learning algorithm by feeding the algorithm with a database of events, wherein each event may relate to an user generated data; providing a relationship between input and a possible outcome, wherein the possible outcome is epitope cleavage and release of the epitope sequence from a polypeptide, antigen presentation and T cell activation with the epitope when inside a cell. In some embodiments the data comprises mass spectrometry data. In some embodiments, the database comprises data obtained through large-scale immunopeptidomics profiling in-house, as also compilation from published database. In some embodiments, the database comprises ligandomics data.

[0187] A cell as discussed herein is a living cell. In some embodiments, the cell is a mammalian cell. In some embodiments the cell is a human cell. A release of an epitope sequence is understood as release of only the sequence that specifically binds to a cognate MHC molecule encoded by a specific allele. Typically, an MHC class I epitope sequence is an 8 amino acid sequence (8-mer). It may be understood that for a natural 8-mer epitope, the epitope sequence released or predicted to be released by the algorithm is the exact 8-mer epitope sequence, for efficient binding to the MHC and presentation to a T cell. In some embodiments, a typical T cell epitope sequence is a 9-mer sequence. It may be understood that for a natural 9-mer epitope the epitope sequence released or predicted to be released by the algorithm is the exact 9-mer epitope sequence, for efficient binding to the MHC and presentation to a T cell.

[0188] In some embodiments, the objective of the endeavor described herein is to maximize epitope release by maximizing predicted epitope cleavage through changing the peptide context with: (i) Linkers (1-4 amino acids) (ii) Changing the order of peptides in the string.

[0189] In some embodiments, the objective of the endeavor described herein is to implement utmost safety for a therapeutic design described herein, by providing in the design a reiterativechecking scheme for possible sequence match with the human proteome for preventing activation of the human proteasomal machinery and discarding a multiepitopic polypeptide sequence if one is found in during the design.

[0190] In some embodiments, the method begins by using a researcher-defined collection of epitopes, known as “a fragment file,” as input to generate a superstring for each unique set of fragments in the collection. Superstrings may be considered as primary phases of polypeptide designs comprising the researcher defined epitope and may contain high level information related to how the string should be designed, such as which positions should be fixed and which ones are variable, which positions can have linkers and which ones cannot, and so on. Thus, in some embodiments, an input file for the computer implemented machine learning program that comprises one or more algorithms, comprises a superstring comprising (a) a collection of two or more fragments, wherein each fragment comprises two or more T cell epitopes that are antigens for stimulating T cells against a disease protein or polypeptide with a pathogenic antigen, and (b) a designation of one or more variable epitope junctions, which may be at the terminal sides of the fragments. In some embodiments, to avoid generating fragment junctions that contain overlap with the human proteome, the procedure loads a reference FASTA file (gencode.v42.pc_translations.fa) containing human proteome sequences which are broken down into all possible sequences of length 8. This information may be encapsulated in a metadata object and used later in a checking schedule whether any human proteome fragments have been introduced into a string design. If it is introduced into a string design, the algorithm will attempt to adapt the design process to avoid such fragments in the final output.

[0191] In one aspect, provided herein is a method that generates a multiepitopic polypeptide sequence wherein each epitope within the multiepitopic polypeptide sequence has been designed and ordered in an N-terminal to C-terminal sequence within the multiepitopic polypeptide such that there is a high likelihood for each epitope being presented to a T cell by a cell expressing the multiepitopic polypeptide when incorporated in a cell. By incorporating in a cell, or expressing in a cell, it may be understood that a polynucleic acid encoding the polypeptide is introduced into a cell according to the methods described in the disclosure, or as is well known to one of skill in the art, e.g., transfection, transduction, electroporation, or any parallel method that would achieve the same. In some embodiments, ‘expressing in a cell,’ or ‘incorporating in a cell’ as described herein would encompass expressing in a cell in vivo, following administration of a composition comprising the polynucleic acid encoding the polypeptide in a composition to a subject, such that a cell in the body of the subject uptakes the polynucleic acid and translates the polynucleic acid thereby expressing the polypeptide encoded by the polynucleic acid. In some embodiments, the subject is a subject in need thereof, wherein the subject is a human subject. In some embodiments, the subject is a human subject with a disease, and the multiepitopic polypeptide comprises a vaccine, comprising T cell epitopes that can activate T cells in the subject against epitopes relevant for the disease, thereby generating an immuneresponse against the disease in the subject. In one embodiment, the output from the computer implemented machine learning program is an output comprising a defined sequence of a multiepitopic polypeptide from N-terminus to C-terminus, which is designed specifically for the instant purpose such that each epitope in the multiepitopic polypeptide has a highest likelihood of presentation to a T cell in vivo and activation of a T cell when expressed in a cell, that is a cell of a subject, wherein the subject is a human subject with a disease and is a recipient of a therapeutic comprising the multiepitopic polypeptide or a polynucleic acid encoding the multiepitopic polypeptide. In some embodiments the multiepitopic polypeptide comprises epitopes designed in a defined order such as to maximize the possibility of each epitope to be presented to T cells for activation by the cells of the subject, which may be determined by a suitable antigen presentation assay, such as a tetramer assay or a peptide-MHC binding assay, a determination of peptide MHC binding affinity in vitro or a T cell activation assay. In some embodiments the multiepitopic polypeptide comprises the epitopes designed in a way that each epitope can bind to and an MHC molecule encoded by the subject’s HLA. In some embodiments the multiepitopic polypeptide comprises the epitopes that can specifically bind to the subject’s MHC molecule with a high binding affinity. In some embodiments the multi epitopic polypeptide comprises the epitopes designed to be immunogenic against the disease of the subject, which may be measured by an immunogenicity assay. Exemplary immunogenicity assays may include a cytokine release assay when the MHC complex comprising the epitope is presented by an antigen presenting cell to a T cell, assayed in vitro by ELISA or flow cytometry. In some embodiments, the multi epitopic polypeptide comprises the epitopes that are specific for the subject, and for the disease of the subject. In some embodiments, the output comprises a cleavage score for the multi epitopic polypeptide sequence provided. A cleavage score is a score generated by the computer implemented machine learning program as further explained below.

[0192] In some embodiments, the framework for the computer implemented machine learning program comprises one or more algorithms. In some embodiments, the framework comprises an epitope ordering and simulated annealing algorithm, that can anneal fragments in defined orders and anneal them in silico. The algorithm further comprises an objective function of cleavage predictor, where the cleavage predictor predicts the likelihood of an epitope junction sequence to be cleaved by antigen presenting machinery inside a cell, thereby releasing the epitopes on either side of an epitope junction sequence. In some embodiments, the different components in the framework may be enumerated as follows:(I) Ordering optimization algorithm, comprising simulated annealing module;(II) Objective function, comprising a cleavage predictor;(III) Sampling function, comprising an epitope sampling and linker sampling functions.

[0193] In some embodiments, the method for designing a multiepitopic polypeptide as disclosed herein comprises a progressive designing and sampling process. In some embodiments, the frameworkdescribed herein is modular and allows for dynamic mixing- and-matching of components based on the design task. This includes selecting which optimization algorithm, objective function and sampling functions, along with their corresponding parameters, should be used for a given design task. In effect, the configuration file defines a meta-program that the cleavage optimization framework can assemble into a coherent, machine interpretable program.

[0194] In some embodiments, the framework is modular, thereby allowing one or more additional, later designed (or parallel designed) machine learning programs to be incorporated into the framework.

[0195] In some embodiments, the collection of superstrings, metadata and settings derived from the configuration file are used by a manager module that reasons about which optimization routine should be dispatched. The manager module calls the simulated annealing sub-manager module which is responsible for implementing the program defined by the configuration file. The sub-manager selects which objective function should be used to score the string designs and which sampling functions should be used to sample the string design space. The string design begins once the submanager has assembled the above components into a program.

[0196] In some embodiments, at the start of the procedure, the superstrings are used to generate string candidates which are randomly shuffled versions of the variable fragments defined by the superstrings. Therefore, in some embodiments, the input materials comprise a plurality of T cell epitope sequences, wherein the plurality of T cell epitope sequences comprises at least three T cell epitope sequences. The plurality of epitope sequences may comprise 4, 5, 6, 7, 8, 9. 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25. 26, 27, 28, 2, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 epitope sequences. In some embodiments, the plurality of T cell epitope sequences comprise about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 100, or about 200 epitope sequences. In some embodiments, the plurality of T cell epitope sequences comprise about 300, about 400, about 500, about 600, about 700, about 800, about 900, or about 1000 epitope sequences. In some embodiments the user defined superstring may comprise about 10 fragments to about 200 fragments. For example, in some embodiments there may be about 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 fragments. A fragment may thus comprise epitopes that are invariant with respect to each other, but owing to reordering of two or more fragments at one or more variable positions, epitopes within a fragment may have variable positions with epitopes on a different fragment on a string. The string candidates are subjected to optimization procedure. The superstrings are invariant and only serve as reference information. The set of possible linkers from which the procedure can choose is initially set to the empty linker.

[0197] A user defined input for designing a vaccine comprising a multiepitopic polypeptide comprising a plurality of T cell epitopes, the T cell epitopes being directed to a particular set of antigens relevant to the disease and that T cells in the subject are to be activated by the vaccine,comprise a set of sequence fragments, e.g., T cell epitope containing sequences, and / or minimal epitopes. Usually, a sequence fragment comprises one or more T cell epitopes. In some embodiments, a fragment usually comprises at least one epitope at each edge of the sequence fragment. In some embodiments the T cell epitopes may be partially overlapping or completely nested epitopes in a fragment. In some embodiments there may be 2, 3, 4, 5, 6, 7, 8, 9, 10 or more T cell epitopes within a fragment. In some embodiments, it may be possible that the sequence fragment (also designated as T cell containing sequences and grammatical equivalents thereof), may not have been all identified, other than the edge epitopes at the time the input is provided. In some embodiments, edge epitopes are minimal epitopes. In some embodiments, more than one epitope may reside in the edge, of which at least the minimal epitope is considered for scoring purposes. In some embodiments, only minimal epitopes at the edge of the sequence fragments may be considered for scoring purposes. In some embodiments, epitopes at the edge of the sequence fragments may be considered for scoring purposes, that may not be the minimal epitope. In some embodiments, at least the edge epitopes are identified and designated by the user at the time the input is provided.

[0198] In an exemplary sequence having about 30 fragments, the number of possible orderings may be 10A32. Thus, in some embodiments, the number of possible orderings may range from 10Al 0 to 10A50. For example, in some embodiments, the number of possible orderings may be 10A20, 10A30 or 10A40. In some embodiments, the number of possible orderings may be 10A12, 10A13, 10A14, 10A15, 10A16, 10A17, 10A18, 10A19, 10A20, 10A21, 10A22, 10A23, 10A24, 10A25, 10A26, 10A27, 10A28, 10A29, 10A30, 10A31, 10A32, 10A33, 10A34, 10A35, 10A36, 10A37, 10A38, 10A39, or 10A40.

[0199] In some embodiments, the procedure thereafter runs for the number of rounds specified in the configuration file, each of which is sub-divided into a specific number of epochs, beginning with round 0. Round 0 is specifically for the selection of an optimal epitope ordering. In each epoch of the round, the procedure proposes a new string candidate based on the current candidate by randomly shuffling all the current candidate's variable fragments. The new candidate string is passed to the objective function which assigns to it a score based on the string's fitness. With the current and new string candidates in hand, the string scores, along with a pre-defined temperature parameter, are used to determine whether to accept or reject the new string in a semi- stochastic fashion. If the new string's score is better than the current string, the new string is accepted and the current string is discarded. If the new string's score is worse than the current string’s score, then the two scores and temperature parameter along with the epoch number and a cooling factor, are used to compute an exponentiated value that, if greater than another randomly sampled value, decides whether to accept the new string.

[0200] In some embodiments, this selection process is represented as:

[0201] This stochastic selection process gives the optimization procedure a greater chance of not getting stuck in a local optimum. This procedure repeats for n epochs, after which the order of fragments in the string candidate is fixed.

[0202] A score as discussed above or elsewhere in the document, unless otherwise mentioned may be considered a fitness score. In some embodiments, a fitness score is a score given by the software as it is running, to each polypeptide, at each round or epoch of selection. A fitness score may be understood as a function of all the cleavability scores of all designated T cell epitopes at a given parameter, for example with a given set of polypeptide sequences, the temperature, taking into account the round or epoch of the progression of a run, whereas the run is for generating a multiepitopic polypeptide having a given set of sequences (e.g., fragments containing T cell epitopes and / or minimal epitopes). In some embodiments, as set by the machine learning software, a fitness score is within a range of (-4) to +4, and a score of 2 and above may be considered a good score. In experimental verifications, a score of 2 and above for a multiepitopic polypeptides highly correlates with identification of a given T cell epitope provided in the input sequences. In some embodiments, for each round of analysis (or Epoch, etc), each input epitope is considered in its sequence surrounding, considering sequences of up to about 30 amino acids on either side or more, and a cleavability score is assigned. For current practical purposes, a fitness score may be understood as an average of all cleavability scores of all epitope sequences for the polypeptide sequence.

[0203] For all subsequent rounds, starting with round 1 and proceeding to 4, the optimization procedure attempts to choose optimal linkers for every open position between fragments. The procedure begins with linkers of length 1 under the assumption that shorter linkers are preferable to longer linkers in that they have a smaller chance of introducing unwanted human proteome overlap into the design. Then, for a given round k, where k is between 1 and 4, all possible linkers of length k are generated for all open linker positions in the superstring. Each linker possibility at each position is then evaluated to determine whether it would generate overlap with the human proteome. If a linker causes overlap, that linker is excluded from the set of choices the optimizer can make during the next design round.

[0204] For example: If the linker choices for linker position 1 are R, H, and Q, during round 1 and Q is found to produce junctional epitopes that overlap with the human proteome, then Q is excluded from the list of options for position 1. This check is performed by taking the left and right flanking fragments and forming a junction of the form (left fragment + linker + right fragment) and asking if any of the sub-sequences of length 8 inside that junction occur in the human proteome. This step repeats for every linker position in every round. It is also possible that a given linker position will admit no linkers of any length, in which case the linker remains empty.

[0205] From this point onwards, the optimization procedure is largely the same as for round 0, with the primary difference being that, instead of swapping fragment orderings in each epoch, a linkercomponent at a single open linker position is randomly selected and set before scoring the new string candidate. That is, given some current string, generate a new string where a single, randomly chosen linker position is filled by selecting from the set of possible linkers components for that position at random. The new string is scored and accepted or rejected following the same heuristics described above.

[0206] This procedure repeats until all epochs in all rounds have been exhausted or an early stopping criterion has been reached, whichever happens first.

[0207] Hence, provided herein is a method for generating a multiepitopic polypeptide sequence from a plurality of T cell epitope sequences, wherein the plurality of T cell epitope sequences comprises at least three T cell epitope sequences, the method comprising: ordering the plurality of T cell epitope sequences in a defined N-terminus to C-terminus order, wherein the defined N-terminus to C-terminus order is a single possible order of all possible orders of the plurality of T cell epitope sequences, thereby forming a sequence having the defined N-terminus to C-terminus order; wherein two adjacent T cell epitope sequences of the sequence having the defined N-terminus to C-terminus order may be modified with a selected linker sequence, wherein each selected linker sequence is based on (a) a combined sequence not being present in the human proteome, wherein the combined sequence comprises: (i) the selected linker sequence and (ii) the sequence immediately upstream and / or downstream of the selected linker sequence; and / or (b) the selected linker sequence is a sequence predicted to be cleaved to a greater extent than another possible non-selected linker sequence when the multiepitopic polypeptide sequence is expressed in a cell; thereby generating the multiepitopic polypeptide sequence from the plurality of T cell epitope sequences.

[0208] In some embodiments, the selected linker sequence inserted in between the two adjacent T cell epitope sequences of the sequence having the defined N-terminus to C-terminus order is generated by an algorithm in a computer implemented process, wherein the algorithm is a prediction algorithm for predicting proteolytic cleavage of a linker sequence inserted in between two adjacent T cell epitope sequences in a multiepitopic polypeptide sequence, wherein the proteolytic cleavage releases each of the two adjacent T cell epitope sequences for antigen presentation and T cell activation when the multiepitopic polypeptide sequence is expressed in a cell. FIGs. 1A- 6 demonstrate various diagrammatic views exemplifying the objectives and functioning of the program.

[0209] In one aspect, therefore, provided herein is a method for designing a multiepitopic polypeptide sequence from a plurality of T cell epitope containing sequences, each T cell epitope containing sequence of the plurality comprising or consisting of at least one T cell epitope sequence, wherein the plurality of T cell epitope containing sequences comprises at least three T cell epitope sequences, the method comprising performing a first T cell epitope containing sequence ordering round, the first T cell epitope containing sequence ordering round comprising: (a) ordering each T cell epitope containing sequence of the plurality into a first defined N to C terminal order, therebygenerating a starting multiepitopic polypeptide sequence; (b) ordering each T cell epitope containing sequence of the plurality into a second defined N to C terminal order, thereby generating a first modified multiepitopic polypeptide sequence; (c) comparing a fitness score of the starting multiepitopic polypeptide sequence to a fitness score of the first modified starting multiepitopic polypeptide sequence; and (d) selecting a first multiepitopic polypeptide sequence from the starting multi epitopic polypeptide sequence and the first modified starting multiepitopic polypeptide sequence based on the comparison of the fitness scores, thereby generating a first multiepitopic polypeptide sequence.

[0210] In one aspect, provided herein is a method comprising, multiepitopic polypeptide sequence from a plurality of T cell epitope containing sequences, each T cell epitope containing sequence of the plurality comprising or consisting of at least one T cell epitope sequence, wherein the plurality of T cell epitope containing sequences comprises at least three T cell epitope sequences, the method comprising: (a) ordering the plurality of T cell epitope containing sequences into a defined N-terminus to C-terminus order, thereby generating a starting multiepitopic polypeptide sequence, wherein the starting multiepitopic polypeptide sequence has a sequence comprising each T cell epitope containing sequence of the plurality directly linked to one or two other T cell epitope containing sequences of the plurality; (b) performing a first linker selection round, the first linker selection round comprising: (i) inserting an amino acid linker sequence into the starting polypeptide sequence between a junction of a first T cell epitope containing sequence and a second T cell epitope containing sequence, thereby generating a first modified starting multiepitopic polypeptide sequence; (ii) comparing a fitness score of the starting polypeptide sequence to a fitness score of the first modified starting multiepitopic polypeptide sequence; and (iii) selecting the first multiepitopic polypeptide sequence from the starting multi epitopic polypeptide sequence and the first modified starting multiepitopic polypeptide sequence based on the comparison of the fitness scores, thereby generating a multiepitopic polypeptide sequence.

[0211] In one aspect, provided herein is a method for generating a multiepitopic polypeptide sequence from a plurality of T cell epitope sequences, wherein the plurality of T cell epitope sequences comprises at least three T cell epitope sequences, the method comprising: (i) randomly selecting two adjacent T cell epitope sequences of the plurality of T cell epitope sequences in a sequence having a defined N-terminus to C-terminus order; and (ii) generating a selected linker sequence between the randomly selected two adjacent T cell epitope sequences, thereby generating the multi epitopic polypeptide, wherein the selected linker sequence comprises a sequence according to a formula of X1X2X3X4, wherein each X represents any amino acid, wherein each of Xi, X2, X3, and X4 is present or absent, wherein generating the selected linker sequence comprises: (a) inputting a first potential linker sequence between the two randomly selected adjacent T cell epitope sequences, (b) determining whether the first potential linker sequence together with the sequence immediately upstream and / ordownstream of the first potential linker sequence is present in the human proteome, and (c) discarding the first potential linker sequence if the first potential linker sequence together with the sequence immediately upstream and / or downstream of the first potential linker sequence is present in the human proteome.

[0212] In some embodiments, generating the selected linker sequence is performed by a cleavage prediction algorithm for predicting proteolytic cleavage, wherein the cleavage prediction algorithm predicts the likelihood of proteolytic cleavage of a linker sequence between the two randomly selected adjacent T cell epitope sequences of the sequence having the defined N-terminus to C-terminus order, when the multiepitopic polypeptide sequence is expressed in a cell, as determined by a tetramer assay or mass spectrometry analysis. In some embodiments, the sequence immediately upstream and / or downstream of the selected linker sequence of a potential linker sequence is a sequence that is up to 30 amino acids upstream and / or downstream. In some embodiments, the sequence immediately upstream and / or downstream of the selected linker sequence of a potential linker sequence is a sequence that is up to 25 amino acids upstream and / or downstream. In some embodiments, the sequence immediately upstream and / or downstream of the selected linker sequence of a potential linker sequence is a sequence that is up to 20 amino acids upstream and / or downstream. In some embodiments, the sequence immediately upstream and / or downstream of the selected linker sequence of a potential linker sequence is a sequence that is at least 8 amino acids upstream and / or downstream.

[0213] In one aspect, the method of incorporating a linker is progressive with reiterative checking process for possible match with the human proteome database and discarding if the design has a match as discussed earlier. The progressive nature is also implemented by advancing linker rounds starting at 0 (empty linker), then adding a single amino acid sequence in round 1, adding 2 amino acids in round 2, and the maximum being four amino acids for a potential linker region.

[0214] In some embodiments, the number of rounds comprises at least 100, at least 1,000, at least 10,000 or more rounds.System, Neural Network and Training

[0215] The backbone of this algorithm is a tool that predicts the likelihood that a given HLA-I epitope will be processed by the proteasome; the algorithm is hereby referred to as CLEavage Predictor (CLEP in short) and considers the sequence of the epitope or fragment as well as its upstream and downstream sequence context (+-30AA). This further developed advanced proprietary algorithm is a neural network trained on a large in-house-curated database that draws on hundreds of internal and external HLA ligandomics datasets. The program is hereby referred to as the cleavage optimizer program (CLEO in short). For simplicity, the system used to perform the program’s functions may be referred to herein as the cleavage optimizer system or CLEO system, the tool may be referred to generally as the cleavage optimizer tool or CLEO tool, the platform as the cleavage optimizer platform or CLEO platform, or for example, a linker designed through the use of the program may be referredto as a CLEO linker for the purposes of this disclosure. The objective was to build a generalizable framework for cleavage optimization that is extensible to new methods and constraints. In one embodiment, it is modular and scalable. For example, the software framework is designed so that components can be easily added and swapped out. New optimization or objective functions can be added with minimal code. Incorporating new user-requested constraints is straightforward.

[0216] In some embodiments, cleavage optimizer program in-takes a user-defined set of T cell epitopes or fragments as input and seeks a final string design that optimizes the likelihood of successful cleavage of the epitopes according to CLEP (although CLEO is agnostic to the specific cleavage tool used). First, the algorithm performs multiple rounds of an iterative, stochastic selection process to determine optimal epitope / fragment ordering and junctional linkers. Each selection step leverages the in-house-designed cleavage predictor model as an objective to determine the fitness of any given string design. Once an optimal epitope / fragment ordering is set in the first round, the algorithm carries out multiple additional rounds of linker selection to further increase the chance that epitopes / fragments are cleaved from the biomolecule. Throughout this process, the cleavage optimizer has rules for determining whether to accept or reject new sequences and whether the process should be stopped due to a lack of improvement. Ultimately, each target epitope or fragment will have 0-4 amino acid(s) (AA(s)) of the linker sequence before and after it.

[0217] In one aspect, provided herein is a computer implemented cleavage optimization framework. In some embodiments, the framework is modular and allows for dynamic mixing-and- matching of components based on the design task. In some embodiments, this includes selecting which optimization algorithm, objective function and sampling functions, along with their corresponding parameters, should be used for a given design task. In some embodiments, the configuration file defines a meta-program that the cleavage optimization framework can assemble into a coherent, machine interpretable program.

[0218] In some embodiments, the computer implemented cleavage predictor comprises specialized optimization algorithms allow us to strategically decrease possible orderings to find best option.

[0219] In some embodiments, the computer implemented cleavage predictor outputs a score for a peptide given its surroundings. In some embodiments, the aim of this design endeavor is to maximize epitope release by maximizing predicted epitope cleavage through changing the peptide context with: Linkers (1-4 amino acids); changing the order of peptides in the string and wherein the higher the score, the higher the chance of being cleaved.

[0220] In some embodiments, the computer implemented cleavage predictor incorporates data obtained through large-scale immunopeptidomics profiling by in-house proteomics team and from external publications. In some embodiments, the processing rules are informed by greater than 300,000for example, for example, nearly 350,000 unique HLA-1 peptides representing 30 different cell types, tissues and tumors.

[0221] In some embodiments, the configuration of the cleavage predictor comprises a configuration of an optimizer; an objective function; and sampling rules. (FIG. 14). In some embodiments, the input data comprises user defined selection of epitope and larger fragment sequences comprising epitopes of interest, (hereafter fragments), and user defined designation of invariable and variable junctions between fragments and / or epitopes that can be developed further for maximizing the probability of cleavage using the cleavage predictor described herein. In some embodiments, the user defined fragments comprise 10, 20, 30, 40, 50, 60, 70 or more fragments. In some embodiments, the user defined epitopes comprise a minimal epitope, wherein a minimal epitope may comprise only the amino acids that bind to an HLA and can be determined by mass spectrometry after elusion from an HLA- peptide complex. For example, a minimal epitope for an HLA encoded Class I MHC molecule is an 8 amino acid long peptide epitope (8-mer). For example, a minimal epitope for an HLA encoded Class I MHC molecule is a 9 amino acid long peptide epitope (9-mer). In some embodiments, each fragment may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more individual T cell epitopes. In some embodiments, each fragment may have overlapping T cell epitopes. In some embodiments, each fragment may comprise one or more nested epitopes. In some embodiments, the epitopes in a fragment that are specifically defined for the cleavage optimizer to consider are the epitopes that are at the junctions of the fragments (edge epitopes) that are defined by user for specific interest in improving cleavage, for example, by introducing a linker, wherein the linker may be a 0, 1, 2, 3, or 4 amino acid residue linker.

[0222] In some embodiments, the objective of developing the cleavage predictor algorithm is aimed at building a generalizable framework for cleavage optimization that is extensible to new methods and constraints. For example, existing programs are inefficient and rudimentary for the purpose of therapeutic design. For example, an existing comparable program for the purpose may contain manual, ad hoc steps for every optimization run. For example, each run may contain one-off logic for handling specific cases, or inefficient logic may consume large amounts of memory at a run-time (e,g, > 200 GB per run). An existing procedure may not be generalizable to the constraints provided by individual requestors. For example, each request required a copy->paste->modify workflow that duplicated code. In some embodiments, generating the new program required unwinding the logic behind many one- off pieces of code that solved many different problems but didn’t generalize and reimplement it in a way that was flexible and free from errors. For example: a sampling function that acts on a set of allowed linkers per string, even when sets of allowed linkers differ between strings from the same input configuration. For example, an input file must follow a strict set of rules to be a valid input and requires extensive, automated checks. Examples (non-exhaustive) include:• First and last epitopes must be substrings of the sequence, up to being the sequence itself,• Epitope lengths must be within a particular size window,• At least one epitope must be cleavable,• Every sequence must conform to the amino acid alphabet,• Linkers cannot have positional conflicts,

[0223] For example, methods could include new optimization algorithms or objective functions. For example, constraints could be ways in which a requestor wants a particular optimization done. Therefore, provided herein is a new method wherein the generalized framework has been built and can perform constrain-free optimization.

[0224] The term “optimization” as used herein may comprise innovative design for improvement. In computer implemented program, an optimization is a new function, that has been objectively designed for solving a problem, that is neither routine nor ordinary. The instant programs, methods and systems described herein comprise the following and provide advantages and superiority over the difficulties in the existing programs discussed above.

[0225] Optimizer Program. Simulated Annealing: Stochastic algorithm that scores permutations of strings using the objective function and randomly selects the next candidate. Selection becomes less random as the optimization progresses.

[0226] Objective Function: The Objective function comprises the cleavage predictor program, that is a three-layer perceptron trained on protein cleavage data. Currently requires a context of 30 amino acids (AAs) on either side of an epitope.

[0227] Sampling Rules: sampling rules comprise a function to swap epitopes and / or linkers. Randomly exchanges epitopes or linkers to create a new string candidate for comparison with a previous candidate.

[0228] In some embodiments the program uses Python. In some embodiments, the components of the program are dynamically assembled.

[0229] In some embodiments, the program dynamically assembles these components into a coherent program based on the user’s specifications.

[0230] In some embodiments, the program is receptive to additional improvement and incorporation of new programs. In some embodiments, additional optimizers can be added to the framework with minimal code. In some embodiments, additional objective functions can be added to the framework with minimal code. In some embodiments, additional sampling rules can be added to the framework with minimal code.

[0231] For example, the new program comprises that the Code is contained in a single python package. Additionally, “.yaml” configuration files act as both the program and the documentation.

[0232] In some embodiments, extensive benchmarking of the system is undertaken. FIG. 14 shows an exemplary view of the organization of the system and framework.Pharmaceutical Composition

[0233] Provided herein is a pharmaceutical composition, comprising at least a first therapeutic agent which comprises multispecific molecule. The multispecific molecule in the composition may be in the form of peptides or polypeptides or a complex of multiple peptides. The multispecific molecule may be provided in a composition as purified recombinant proteins.

[0234] The multispecific molecule may be in the form of a polynucleotide encoding the recombinant multispecific molecule. In some embodiments, polynucleotide encoding the multispecific molecule may comprise DNA, mRNA or circRNA or a liposomal composition of any one of these. The liposome is an LNP.

[0235] Pharmaceutical compositions can include, in addition to active ingredient, a pharmaceutically acceptable excipient, carrier, buffer, stabilizer or other materials well known to those skilled in the art. Such materials should be non-toxic and should not interfere with the efficacy of the active ingredient. The precise nature of the carrier or other material will depend on the route of administration.

[0236] Acceptable carriers, excipients, or stabilizers are those that are non-toxic to recipients at the dosages and concentrations employed, and include buffers such as phosphate, citrate, and other organic acids; antioxidants including ascorbic acid and methionine; preservatives (such as octadecyldimethylbenzyl ammonium chloride; hexamethonium chloride; benzalkonium chloride, benzethonium chloride; phenol, butyl or benzyl alcohol; alkyl parabens such as methyl or propyl paraben; catechol; resorcinol; cyclohexanol; 3 -pentanol; and m-cresol); low molecular weight (less than about 10 residues) polypeptides; proteins, such as serum albumin, gelatin, or immunoglobulins; hydrophilic polymers such as polyvinylpyrrolidone; amino acids such as glycine, glutamine, asparagine, histidine, arginine, or lysine; monosaccharides, disaccharides, and other carbohydrates including glucose, mannose, or dextrins; chelating agents such as EDTA; sugars such as sucrose, mannitol, trehalose or sorbitol; salt-forming counter-ions such as sodium; metal complexes (e.g., Zn- protein complexes); and / or non-ionic surfactants such as TWEEN®, PLURONICS® or polyethylene glycol (PEG).

[0237] Acceptable carriers are physiologically acceptable to the administered patient and retain the therapeutic properties of the compounds with / in which it is administered. Acceptable carriers and their formulations are generally described in, for example, Remington’ pharmaceutical Sciences (18thed. A. Gennaro, Mack Publishing Co., Easton, PA 1990). One example of carrier is physiological saline. A pharmaceutically acceptable carrier is a pharmaceutically acceptable material, composition or vehicle, such as a liquid or solid fdler, diluent, excipient, solvent or encapsulating material, involved in carrying or transporting the subject compounds from the administration site of one organ, or portion of the body, to another organ, or portion of the body, or in an in vitro assay system. Acceptable carriers are compatible with the other ingredients of the formulation and not injurious to a subject to whom it is administered. Nor should an acceptable carrier alter the specific activity of the neoantigens.

[0238] In one aspect, provided herein are pharmaceutically acceptable or physiologically acceptable compositions including solvents (aqueous or non-aqueous), solutions, emulsions, dispersion media, coatings, isotonic and absorption promoting or delaying agents, compatible with pharmaceutical administration. Pharmaceutical compositions or pharmaceutical formulations therefore refer to a composition suitable for pharmaceutical use in a subject. Compositions can be formulated to be compatible with a particular route of administration (i.e., systemic or local). Thus, compositions include carriers, diluents, or excipients suitable for administration by various routes.

[0239] In some embodiments, a composition can further comprise an acceptable additive in order to improve the stability of immune cells in the composition. Acceptable additives may not alter the specific activity of the immune cells. Examples of acceptable additives include, but are not limited to, a sugar such as mannitol, sorbitol, glucose, xylitol, trehalose, sorbose, sucrose, galactose, dextran, dextrose, fructose, lactose and mixtures thereof. Acceptable additives can be combined with acceptable carriers and / or excipients such as dextrose. Alternatively, examples of acceptable additives include, but are not limited to, a surfactant such as polysorbate 20 or polysorbate 80 to increase stability of the peptide and decrease gelling of the solution. The surfactant can be added to the composition in an amount of 0.01% to 5% of the solution. Addition of such acceptable additives increases the stability and half-life of the composition in storage.

[0240] The pharmaceutical composition can be administered, for example, by injection. Compositions for injection include aqueous solutions (where water soluble) or dispersions and sterile powders for the extemporaneous preparation of sterile injectable solutions or dispersion. For intravenous administration, suitable carriers include physiological saline, bacteriostatic water, or phosphate buffered saline (PBS). The carrier can be a solvent or dispersion medium containing, for example, water, ethanol, polyol (for example, glycerol, propylene glycol, and liquid polyethylene glycol, and the like), and suitable mixtures thereof. Fluidity can be maintained, for example, by the use of a coating such as lecithin, by the maintenance of the required particle size in the case of dispersion and by the use of surfactants. Antibacterial and antifungal agents include, for example, parabens, chlorobutanol, phenol, ascorbic acid and thimerosal. Isotonic agents, for example, sugars, polyalcohols such as mannitol, sorbitol, and sodium chloride can be included in the composition. The resulting solutions can be packaged for use as is, or lyophilized; the lyophilized preparation can later be combined with a sterile solution prior to administration. For intravenous, injection, or injection at the site of affliction, the active ingredient will be in the form of a parenterally acceptable aqueous solution which is pyrogen-free and has suitable pH, isotonicity and stability. Those of relevant skill in the art are well able to prepare suitable solutions using, for example, isotonic vehicles such as Sodium Chloride Injection, Ringer’s Injection, Lactated Ringer’s Injection. Preservatives, stabilizers, buffers, antioxidants and / or other additives can be included, as needed. Sterile injectable solutions can be prepared by incorporating an active ingredient in the required amount in an appropriate solvent withone or a combination of ingredients enumerated above, as required, followed by fdtered sterilization. Generally, dispersions are prepared by incorporating the active ingredient into a sterile vehicle which contains a basic dispersion medium and the required other ingredients from those enumerated above. In the case of sterile powders for the preparation of sterile injectable solutions, the preferred methods of preparation can be vacuum drying and freeze drying which yields a powder of the active ingredient plus any additional desired ingredient from a previously sterile-filtered solution thereof.

[0241] Compositions can be conventionally administered intravenously, such as by injection of a unit dose, for example. For injection, an active ingredient can be in the form of a parenterally acceptable aqueous solution which is substantially pyrogen-free and has suitable pH, isotonicity and stability. One can prepare suitable solutions using, for example, isotonic vehicles such as Sodium Chloride Injection, Ringer’s Injection, Lactated Ringer’s Injection. Preservatives, stabilizers, buffers, antioxidants and / or other additives can be included, as required. Additionally, compositions can be administered via aerosolization.

[0242] When the compositions are considered for use in medicaments or any of the methods provided herein, it is contemplated that the composition can be substantially free of pyrogens such that the composition will not cause an inflammatory reaction or an unsafe allergic reaction when administered to a human patient. Testing compositions for pyrogens and preparing compositions substantially free of pyrogens are well understood to one or ordinary skill of the art and can be accomplished using commercially available kits.

[0243] Acceptable carriers can contain a compound that stabilizes, increases or delays absorption, or increases or delays clearance. Such compounds include, for example, carbohydrates, such as glucose, sucrose, or dextrans; low molecular weight proteins; compositions that reduce the clearance or hydrolysis of peptides; or excipients or other stabilizers and / or buffers. Agents that delay absorption include, for example, aluminum monostearate and gelatin. Detergents can also be used to stabilize or to increase or decrease the absorption of the pharmaceutical composition, including liposomal carriers. To protect from digestion the compound can be complexed with a composition to render it resistant to acidic and enzymatic hydrolysis, or the compound can be complexed in an appropriately resistant carrier such as a liposome. Means of protecting compounds from digestion are known in the art (e.g., Fix (1996) Pharm Res. 13:1760 1764; Samanen (1996) J. Pharm. Pharmacol. 48:119 135; and U.S. Pat. No. 5,391,377).

[0244] The compositions can be administered in a manner compatible with the dosage formulation, and in a therapeutically effective amount. The quantity to be administered depends on the subject to be treated, capacity of the subject’s immune system to utilize the active ingredient, and degree of binding capacity desired. Precise amounts of active ingredient required to be administered depend on the judgment of the practitioner and are peculiar to each individual. Suitable regimes for initial administration and booster shots are also variable but are typified by an initial administrationfollowed by repeated doses at one or more hourly intervals by a subsequent injection or other administration. Alternatively, continuous intravenous infusions sufficient to maintain concentrations in the blood are contemplated.Treatment Methods

[0245] Provided herein is a method for treating cancer or infectious disease in a subject, comprising administering to the subject in need thereof a composition comprising a T cell vaccine. In some embodiments, the T cell vaccine comprises a multiepitopic polypeptide or a polynucleotide string encoding the multiepitopic polypeptide. The multiepitopic polypeptide for example, directed towards treating a cancer comprises epitopes that are target epitopes derived from the cancer cells, cancer tissue, or cancer surrounding tissue. In some embodiments, for treating infectious diseases, multiepitopic polypeptide comprises epitopes that are target epitopes of the infectious entity, for example a virus, bacteria, etc. In some embodiments, the subject is human.

[0246] In some embodiments the T cell vaccine comprising a multiepitopic polypeptide or a polynucleotide string encoding the multiepitopic polypeptide is formulated for in vivo delivery via injection or infusion.

[0247] In some embodiments, the T cell vaccine comprises an RNA encoding the multiepitopic polypeptide in a lipid nanoparticle delivery vehicle.

[0248] Further treatment method may comprise cell therapy, comprising for example, contacting cancer neoantigen loaded antigen presenting cells (APCs) with isolated T cells ex vivo, wherein, the cancer neoantigen loaded antigen presenting cells (APCs) are CDl lb depleted; preparing cancer neoantigen primed T cells for a cellular composition for cancer immunotherapy ex vivo; and administering the cellular composition for cancer immunotherapy in the subject, wherein at least one or more conditions or symptoms related to the cancer are reduced or ameliorated by the administering, thereby treating the subject, wherein the cancer neoantigen loaded APCs and the cancer neoantigen primed T cells each express a protein encoded by an HLA allele that is expressed in the subject, and to which the neoantigen can specifically bind. Whereas cancer therapy is exemplified in the discussion more often, similar methods can be implemented in treatment of infectious diseases by one of skill in the art and is equally contemplated in the disclosure.

[0249] In some embodiments, the method further comprises administering one or more of the at least one antigen specific T cell to a subject. In some embodiments, the therapeutic composition comprising T cells is administered by injection. In some embodiments, the therapeutic composition comprising T cells is administered by infusion. When administration is by injection, the active agent can be formulated in aqueous solutions, specifically in physiologically compatible buffers such as Hanks solution, Ringer’s solution, or physiological saline buffer. The solution can contain formulator agents such as suspending, stabilizing and / or dispersing agents. In another embodiment, the pharmaceutical composition does not comprise an adjuvant or any other substance added to enhancethe immune response stimulated by the peptide. In some embodiments, the method further comprises administering one or more of the at least one antigen specific T cell as a pharmaceutical composition described herein to a subject. In some embodiments, the pharmaceutical composition comprises a preservative or stabilizer. In some embodiments the preservative or stabilizer is selected from a cytokine, a growth factor or an adjuvant or a chemical substance. In some embodiments, the at least one antigen specific T cell is administered to a subject within 28 days from collecting a PBMC sample from the subject.

[0250] In addition to the formulations described previously, the active agents can also be formulated as a depot preparation. Such long-acting formulations can be administered by implantation or transcutaneous delivery (for example subcutaneously or intramuscularly), intramuscular injection or use of a transdermal patch. Thus, for example, the agents can be formulated with suitable polymeric or hydrophobic materials (for example as an emulsion in an acceptable oil) or ion exchange resins, or as sparingly soluble derivatives, for example, as a sparingly soluble salt.

[0251] Also provided herein are methods of treating a subject with a disease, disorder or condition. A method of treatment can comprise administering a composition or pharmaceutical composition disclosed herein to a subject with a disease, disorder or condition.

[0252] The present disclosure provides methods of treatment comprising an immunogenic therapy. Methods of treatment for a disease (such as cancer or a viral infection) are provided. A method can comprise administering to a subject an effective amount of a composition comprising an immunogenic antigen specific T cell according to the methods provided herein. In some embodiments, the antigen comprises a viral antigen. In some embodiments, the antigen comprises a tumor antigen.

[0253] Non-limiting examples of therapeutics that can be prepared include a peptide-based therapy, a nucleic acid-based therapy, an antibody-based therapy, a T cell-based therapy, and an antigen-presenting cell-based therapy.

[0254] In some other aspects, provided herein is use of a composition or pharmaceutical composition for the manufacture of a medicament for use in therapy. In some embodiments, a method of treatment comprises administering to a subject an effective amount of T cells specifically recognizing an immunogenic neoantigen peptide. In some embodiments, a method of treatment comprises administering to a subject an effective amount of a TCR that specifically recognizes an immunogenic neoantigen peptide, such as a TCR expressed in a T cell.

[0255] In some embodiments, the cancer is selected from the group consisting of carcinoma, lymphoma, blastoma, sarcoma, leukemia, squamous cell cancer, lung cancer (including small cell lung cancer, non-small cell lung cancer (NSCLC), adenocarcinoma of the lung, and squamous carcinoma of the lung), cancer of the peritoneum, hepatocellular cancer, gastric or stomach cancer (including gastrointestinal cancer), pancreatic cancer, glioblastoma, cervical cancer, ovarian cancer, liver cancer, bladder cancer, hepatoma, breast cancer, colon cancer, melanoma, endometrial or uterine carcinoma,salivary gland carcinoma, kidney or renal cancer, liver cancer, prostate cancer, vulval cancer, thyroid cancer, hepatic carcinoma, head and neck cancer, colorectal cancer, rectal cancer, soft-tissue sarcoma, Kaposi’s sarcoma, B-cell lymphoma (including low grade / follicular non-Hodgkin’s lymphoma (NHL), small lymphocytic (SL) NHL, intermediate grade / follicular NHL, intermediate grade diffuse NHL, high grade immunoblastic NHL, high grade lymphoblastic NHL, high grade small non-cleaved cell NHL, bulky disease NHL, mantle cell lymphoma, AIDS-related lymphoma, and Waldenstrom’s macroglobulinemia), chronic lymphocytic leukemia (CLL), acute lymphoblastic leukemia (ALL), myeloma, Hairy cell leukemia, chronic myeloblasts leukemia, and post-transplant lymphoproliferative disorder (PTLD), abnormal vascular proliferation associated with phakomatoses, edema, Meigs’ syndrome, and combinations thereof.

[0256] The methods described herein are particularly useful in the personalized medicine context, where immunogenic neoantigen peptides identified according to the methods described herein are used to develop therapeutics (such as vaccines or therapeutic antibodies) for the same individual. Thus, a method of treating a disease in a subject can comprise identifying an immunogenic neoantigen peptide in a subject according to the methods described herein; and synthesizing the peptide (or a precursor thereof, such as a polynucleotide (e.g., an mRNA) encoding the peptide). In some embodiments the peptide is a multiepitopic polypeptide (wherein the therapeutic composition comprises an RNA, e.g., mRNA encoding the multiepitopic polypeptide as described herein, designed specifically for release and presentation of each epitope in the multiepitopic polypeptide, formulated in an LNP delivery vehicle.

[0257] In some embodiments, multiepitopic polypeptide may be used for manufacturing T cells specific for identified neoantigens; and administering the neoantigen specific T cells to the subject. In some embodiments, the method of treating a disease in a subject can comprise identifying an immunogenic neoantigen peptide; and synthesizing the polynucleotide, such as an mRNA, that encodes the immunogenic neoantigen peptide or polypeptide or a precursor thereof and administering the neoantigen polypeptide or a precursor thereof into the subject for activating antigen-specific T cells in the subject in vivo.

[0258] In some embodiments, the subject has previously been treated with one or more different cancer treatment modalities. In some embodiments, the subject has previously been treated with one or more of radiotherapy, chemotherapy, or immunotherapy. In some embodiments, the subject has been treated with one, two, three, four, or five lines of prior therapy. In some embodiments, the prior therapy is a cytotoxic therapy.EXAMPLESExample 1. MethodsCell Culture and Transfection for HLA Ligandomics MS Analysis

[0259] A375 cells were either engineered to stably express BAP-tagged alleles of interest or the allele of interest was overexpressed within the cell and used for transfection. 50e6 engineered cells transfected with unformulated MessengerMax Lipofectamine reagent prior to harvest.Sample Processing for HLA Ligandomics MS Analysis

[0260] Transfected cells were lysed and cleared before processing. For BAP-tagged cell lines the cleared lysate was biotinylated with biotin, ATP and BirA prior to incubation with NEUTRA VIDIN beads to affinity-enrich biotinylated-HLA-peptide complexes. For overexpressed cell lines, sepharose beads were changed with a pan-class-I antibody and incubated with the cleared lysate to isolate all HLA-peptide complexes. Peptides were washed and eluted from antibody-bound HLA complexes and molecular weight filtration was performed to isolate peptides. Isolated peptides were then labeled with TMTZERO, then reduced using TCEP, alkylated using IAA and desalted prior to analysis by nLC- MS / MS.HLA-Peptide Sequencing by nLC-MS / MS

[0261] Samples were resuspended in a 3% acetonitrile, 5% formic acid supplemented with 150 femtomoles of each TMT-131C labeled heavy synthetic peptide per injection. Peptides were chromatographically separated using a Vanquish Neo uHPLC fitted with an Aurora Ultimate packed emitter column and heated at 60 °C during separation. Peptides were eluted into an Orbitrap Ascend Tribrid Mass Spectrometer equipped with a Nanospray Flex Ion source. Data were acquired using internal standard triggered parallel reaction monitoring. Fast, low-resolution precursor scans were used to look for m / z values in an inclusion list associated with the TMT-131C labeled heavy synthetic peptide internal standards. When an m / z value from the inclusion list was observed, a fast, low- resolution tandem mass spectrum (MS / MS) survey scan was performed, and characteristic fragment ions associated with the peptide were monitored. If five or more monitored ions were observed, a second MS / MS scan was performed with a mass offset equal to the difference between the TMT-131C labeled heavy synthetic peptide and the TMTZERO labeled target HLA peptide.Targeted MS Data Analysis for HLA Ligandomics

[0262] Data analysis was performed using Skyline-daily software. Retention times and peptide fragments were identified by matching with the spiked-in heavy isotope labeled synthetic peptides. Relative abundance for each peptide was calculated by measuring the area under the curve (AUC) for the top 10 most abundant fragment ions.TCR Jurkat-NFAT assay

[0263] This is in vitro assay where it is known specific TCR that recognizes one of HIV epitopes on the strings (and some epitopes are edge epitopes). In the assay, Jurkat cells are transfected withTCR of interest, and presenting cells (cell lines or cells from the donors) are transfected with alleles of interest and one of the strings. If TCR will recognize the epitope being presented (meaning that it is processed and presented from the string), there is a luciferase readout.T-cell activation assay with K562

[0264] NFAT-TCR / CD3 effector cells were purchased from Promega as cryopreserved cells. These Jurkat T-cells express luciferase as a reporter, driven by an NFAT-response element (NFAT- RE). The endogenous TCR and B2M-gene have been removed in the Jurkat reporter cells by CRISPR- Cas9-mediated knockout. The alpha- and beta-chain of the CD8-coreceptor were stably inserted in the Jurkat reporter cells by via transposon. Reporter NF AT -luciferase cells were co-electroporated with two mRNAs, encoding for a TCR clones alpha and beta chains. Post transfection, 2 x 10A4 Jurkat cells were co-cultured with K562 cells at a 2: 1 ratio, in a 384-well-plate with 25 pL medium (RPMI1640 + 10% non-heat inactivated FBS) / well. Prior to co-culture, the K562 cells were transfected with an mRNA encoding for an HIV-derived polypeptide (string 1 or string 2) and mRNAs encoding for an HLA class I alleles. As a positive control for the specificity of the used TCRs, K562 cells only transfected with mRNAs encoding the HLA-I pulsed with a minimal HIV-epitope peptide target, were co-cultured with each TCR-encoding mRNA transfected Jurkat reporter cells. Moreover, stimulation with 2pg / ml Phytohemagglutinin-L (PHA-L) was used to corroborate TCR expression and downstream signaling. Transient expression of transfected HLA class-I was verified by flow cytometry after staining with HLA- A or HLA-B specific antibodies. After 16 h, an equal volume (15 pL) of luciferin (BIO-GLO, Promega) was added to each well and the luciferase activity was measured using a luminescence plate reader. The measured luminescence signal in the different wells corresponded to the level of TCR-mediated activation in the Jurkat cells. For each TCR, log2 fold change of luminescence compared to the “effectors only control” was calculated and a cut-off of twofold change was used to determine specific TCRs.T-cell activation assay with iDCs

[0265] Cells were used and prepared as described in example 4. Post transfection, 2 x 10A4 Jurkat cells were co-cultured with immature Dendritic cells (iDCs) cells at a 2:1 ratio, in a 384-well-plate with 25 pL medium (RPMI1640 + 10% non-heat inactivated FBS) / well. Prior to co-culture, iDCs were generated from donor PBMCs. Briefly, CD14+ monocytes were positively isolated from human PBMCs and cultivated for 5 days at 1x10A6 cells / mL in RPMI1640, 7 / 5% pooled Human Serum (PHS) / 1% Sodium pyruvate / 0,5% Penicillin-Streptomycin supplemented with lOOOU / mL IL-4 and 1600 U7mL GM-CSF to generate iDCs. iDCs were then transfected with mRNAs encoding for an HIV-derived polypeptide (string 1 and 2) and mRNAs encoding for HLA class I alleles before coincubation with the Jurkat reporter cells. As a positive control for the specificity of the used TCRs, iDCs only transfected with mRNAs encoding the HLA-I pulsed with a minimal HIV-epitope peptide target, were co-cultured with each TCR-encoding mRNA transfected Jurkat reporter cells. To evaluatethe impact of endogenous HLA-I alleles from the PBMCs donor, a control including iDCs only pulsed with peptide was included. Moreover, stimulation with 2pg / ml Phytohemagglutinin-L (PHA-L) was used to corroborate TCR expression and downstream signaling. After 16 h, an equal volume (15 L) of luciferin (BIO-GLO, Promega) was added to each well and the luciferase activity was measured using a luminescence plate reader. The measured luminescence signal in the different wells corresponded to the level of TCR-mediated activation in the Jurkat cells. For each TCR, log2 fold change of luminescence compared to the “effectors only control” was calculated and a cut-off of twofold change was used to determine specific TCRs. CLEO linkers are short (0-4aa) and therefore have better safety profile.Example 2. Experimental validation of T cell epitope cleavage from multiepitopic polypeptide

[0266] In this example, strings designed using the cleavage predictor software were tested for evidence of T cell epitopes being released by cleavage from a polypeptide, and expressed on a cell, when an RNA encoding the multiepitopic epitope is incorporated in the cell. For example, cleavage optimization was applied in designing coronavirus vaccine string (CorVac 2.0) as proof of concept. Epitopes from cleavage optimized regions were observed via targeted and discovery mode mass spectrometry. No junction-spanning epitopes were observed during discovery mode mass spectrometry. As shown in FIG. 5, mass spectrometry observed epitopes were generated from the sequence fragments.

[0267] The cleavage design program was further tested for a vaccine design against HIV. Two strings were designed. Each string contains 32-33 fragments / minimal epitopes, separated by cleavage optimized linkers (FIG. 6).

[0268] FIGs. 7 and 8 show snapshot summaries of all the detected epitopes mapping to the fragments originally designed for each string. The epitopes at the edge of the fragments were successfully detected. At least 20 fragments of 33 or 34 had epitopes that were detected. There were still a few fragments where no epitopes were detected. However, the prediction algorithm was a great success, and the detected epitopes correspond to epitopes that were predicted to be released and were at a score of 2.0 and above. (FIG. 9).

[0269] FIG. 6 is a closer view of the newly designed HIV vaccine string using the program. The arrangement of fragments on a strings and linkers are depicted. In this string CD8 T cell epitopes are incorporated, and longer epitopes to stimulate proliferation of CD4 T cells are discouraged. In these string designs P2P16 was inserted at the C-terminus and is intended to break immune tolerance. To validate that cleavage optimizer approach allows epitopes and fragments to be cleaved, ligandomics on 22 monoallelic / overexpressed cell lines were used including Jurkat-NFAT reporter assay with four (4) TCRs that recognize epitopes to show that minimal epitopes and epitopes on the edges of the fragment are cleaved out of the polypeptide. 24 edge epitopes (both minimal and from fragment) weredetected for string 1 (FIGs. 7 and 8). As shown in FIG. 8, a number of non-edge epitopes were also detected.Example 3. Experimental validation of CLEO linkers with Lynch syndrome epitopes, and comparison of different designs.

[0270] In this example, several string designs using the cleavage predictor software were tested for evidence of T cell epitopes being released by cleavage from a polypeptide. The strings designed are designated as follows: (a) control (b) RNA44, (c) RNA51, RNA44 shuffled vl, (d) RNA52, RNA44 shuffled v2, (e) RNA53, RNA44 shuffled v3, (f) RNA54, GS linkers: Instead of using CLEO, this string uses simple GS strings (“GGGGSGGGGS”) between epitopes. These are flexible linkers, (g) RNA55, Helical linkers: Instead of using cleavage optimizer CLEO linkers, this string uses helical linkers between epitopes. These are rigid, so it prevents the string from bending, (h) RNA56, RSV Furin sites: Instead of using CLEO linkers, this string uses a Furin cleavage site derived from RSV (respiratory syncytial virus). Furin sites are an alternative protease pathway to cleave the string, (i) RNA57, Furin sites: Same as above, but the Furin cleavage site is derived from a naturally occurring human sequence, (j) RNA58, Solubility + OG string: OG string here refers to RNA44. This string has a MBP (mannose binding protein) sequence in the front that theoretically drastically increases the solubility of this string, (k) RNA59, Min ribosomal stalling: This string design contains RNA44 with 30 point mutations added that minimize ribosomal stalling motifs. (1) RNA60,: RNA44 using a codon optimization scheme (optimization 2 or Opt 2) other than in the above strings (optimization 1 or Opt 1). (m) RNA61, using yet another codon optimization scheme on RNA44 (optimization 3 or Opt 3).

[0271] RNA 44 is the best scoring cleavage optimizer-designed (CLEO) string. RNA51 is RNA44 shuffled version l(vl) and is the second best-scoring CLEO string. This has the same epitopes as RNA44 but the epitopes are in a different order and thus the chosen CLEO linkers are different as well. RNA52, which is RNA44 shuffled v2 is the third best-scoring CLEO string. Similar to RNA51 or RNA44 shuffled vl conceptually, the epitopes are in a different order (and thus the chosen CLEO linkers are different as well). RNA53, or RNA44 shuffled v3 is the fourth best-scoring CLEO string. Similar to RNA44 shuffled vl conceptually, the epitopes here are in a different order (and thus the chosen CLEO linkers are different as well).

[0272] The sequences of each string are provided in the Table below.

[0273] Table 1. Lynch epitope string designs - Nucleotide and amino acid sequences

[0274] Individual epitopes and linkers for Lynch RNA44, Lynch RNA60 and Lynch RNA61 are shown in the Table below.

[0275] Table 2. Detailed sequences of linkers and epitopes of RNA 44, RNA 60, RNA 61

[0276] Individual epitopes and linkers for Lynch RNA51 are shown in the Table below.

[0277] Table 3. Detailed sequences of linkers and epitopes of RNA51

[0278] Individual epitopes and linkers for Lynch RNA52 are shown in the Table below.

[0279] Table 4. Detailed sequences of linkers and epitopes of RNA52

[0280] Individual epitopes and linkers for Lynch RNA53 are shown in the Table below.

[0281] Table 5. Detailed sequences of linkers and epitopes of RNA53

[0282] Individual epitopes and linkers for Lynch RNA54 are shown in the Table below.

[0283] Table 6. Detailed sequences of linkers and epitopes of RNA54

[0284] Individual epitopes and linkers for Lynch RNA55 are shown in the Table below.

[0285] Table 7. Detailed sequences of linkers and epitopes of RNA55

[0286] Individual epitopes and linkers for Lynch RNA56 are shown in the Table below.

[0287] Table 8. Detailed sequences of linkers and epitopes of RNA56

[0288] Individual epitopes and linkers for Lynch RNA57 are shown in the Table below.

[0289] Table 9. Detailed sequences of linkers and epitopes of RNA57

[0290] Individual epitopes and linkers for Lynch RNA58 are shown in the Table below.

[0291] Table 10. Detailed sequences of linkers and epitopes of RNA58

[0292] Individual epitopes and linkers for Lynch RNA59 are shown in the Table below.

[0293] Table 11. Detailed sequences of linkers and epitopes of RNA59

[0294] For each of these of the string constructs, expression as well as antigen presentation of the epitopes within are verified. Targeted mass spectrometry was used to assess the expression of all these string variants. Briefly, Expi293 cells were grown then transfected with the RNA string. The cells were then lysed and digested and peptide quantitation was done using targeted LC-MS. Each string was done in triplicate and in 2 conditions: with and without a proteasomal inhibitor (PI) (bortezomib). FIG. 11 shows data for expression of the strings. The results indicate that the strings that use CLEO generated linkers including the codon optimized and the MBP encoding variations of the RNA44 string show high expression. FIG. 12 shows antigen presentation. Antigen presentation is tested on a subset of the strings. The cell growth and transfection was done similarly to above except that an antibody pulldown was done post cell lysis to obtain A0201 -bound epitopes. A neoORF epitope having the sequence VLDGTVSAV was used to test epitope abundance. FIG. 12 confirms that RNA44 based strings with CLEO generated linkers, in particular RNA 44, RNA 51 (44 shuffled vl), RNA 52 (44 shuffled v2), show high epitope presentation based on the epitope abundance, which is relatively higher than the other string designs with GS linkers, furin linkers or other modifications. These data show the success of the CLEO generated strings for antigen presentation for either vaccine or T cell therapeutic developments.

Claims

CLAIMSWhat is claimed is:

1. A method for generating a multiepitopic polypeptide sequence from a plurality of T cell epitope containing sequences, each T cell epitope containing sequence of the plurality comprising or consisting of at least one T cell epitope sequence, wherein the plurality of T cell epitope containing sequences comprises at least three T cell epitope sequences, the method comprising:(a) ordering the plurality of T cell epitope containing sequences into a defined N- terminus to C-terminus order, thereby generating a starting multiepitopic polypeptide sequence, wherein the starting multiepitopic polypeptide sequence has a sequence comprising each T cell epitope containing sequence of the plurality directly linked to one or two other T cell epitope containing sequences of the plurality;(b) performing a first linker selection round, the first linker selection round comprising:(i) inserting an amino acid linker sequence into the starting polypeptide sequence between a junction of a first T cell epitope containing sequence and a second T cell epitope containing sequence, thereby generating a first modified starting multiepitopic polypeptide sequence;(ii) comparing a fitness score of the starting polypeptide sequence to a fitness score of the first modified starting multiepitopic polypeptide sequence; and(iii) selecting the first multiepitopic polypeptide sequence from the starting multiepitopic polypeptide sequence and the first modified starting multiepitopic polypeptide sequence based on the comparison of the fitness scores, thereby generating a multiepitopic polypeptide sequence.

2. The method of claim 1, wherein the method is performed in silico.

3. The method of claim 1 or 2, wherein the amino acid linker sequence consists of 1 amino acid during the first linker selection round.

4. The method of claim 3, wherein the first linker selection round comprises performing the inserting, comparing and selecting for each junction of two adjacent T cell epitope containing sequences.

5. The method of claim 4, wherein the method further comprises performing a second linker selection round, wherein the amino acid linker sequence consists of 2 amino acids during the second linker selection round.

6. The method of claim 5, wherein the second linker selection round comprises performing the inserting, comparing and selecting for each junction of two adjacent T cell epitope containing sequences.

7. The method of claim 6, wherein the method further comprises performing a third linker selection round, wherein the amino acid linker sequence consists of 3 amino acids during the third linker selection round.

8. The method of claim 7, wherein the third linker selection round comprises performing the inserting, comparing and selecting for each junction of two adjacent T cell epitope containing sequences.

9. The method of claim 8, wherein the method further comprises performing a fourth linker selection round, wherein the amino acid linker sequence consists of 4 amino acids during the fourth linker selection round.

10. The method of claim 9, wherein the selection linker selection round comprises performing the inserting, comparing and selecting for each junction of two adjacent T cell epitope containing sequences.

11. The method of claim 1 or 2, wherein the method comprises performing a number of additional linker selection rounds each additional selection round comprising(i) inserting an amino acid linker sequence into the junction of two adjacent T cell epitope containing sequences in the selected multiepitopic polypeptide sequence from the previous linker selection round;(ii) comparing a fitness score of the selected multiepitopic polypeptide sequence from the previous linker selection round to a fitness score of the modified multiepitopic polypeptide sequence of the current additional linker selection round, and(iii) selecting a multiepitopic polypeptide sequence from the selected multiepitopic polypeptide sequence from the previous linker selection round and the modified multiepitopic polypeptide sequence of the current additional linker selection round based on the comparison of the fitness scores.

12. The method of any one of claims 1-11, wherein each of the inserting steps of a given linker selection round comprises inserting, at a given junction of two adjacent T cell epitope containing sequences, an amino acid linker sequence that is one amino acid in length longer than the length of the amino acid linker sequence inserted during the previous linker selection round.

13. The method of claim 12, wherein each of the inserting steps of a first linker selection round comprises inserting, at a given junction of two adjacent T cell epitope containing sequences, an amino acid linker sequence that is one amino acid in length.

14. The method of claim 13 , wherein each of the comparing steps of the first linker selection round comprises comparing a multiepitopic polypeptide sequence that lacks an amino acid linker sequence at the given junction of two adjacent T cell epitope containing sequences to a modified multiepitopic polypeptide sequence that has an amino acid linker sequence at thegiven junction of two adjacent T cell epitope containing sequence that is one amino acid in length.

15. The method of any one of claims 12-14, wherein each of the inserting steps of a second linker selection round comprises inserting, at a given junction of two adjacent T cell epitope containing sequences, an amino acid linker sequence that is two amino acids in length.

16. The method of claim 15, wherein each of the comparing steps of the second linker selection round comprises comparing a multiepitopic polypeptide sequence that either lacks an amino acid linker sequence or has an amino acid linker sequence that is one amino acid in length at the given junction of two adjacent T cell epitope containing sequences to a modified multiepitopic polypeptide sequence that has an amino acid linker sequence at the given junction of two adjacent T cell epitope containing sequence that is two amino acid in length.

17. The method of any one of claims 12-16, wherein each of the inserting steps of a third linker selection round comprises inserting, at a given junction of two adjacent T cell epitope containing sequences, an amino acid linker sequence that is three amino acids in length.

18. The method of claim 17, wherein each of the comparing steps of the third linker selection round comprises comparing a multiepitopic polypeptide sequence that either lacks an amino acid linker sequence or has an amino acid linker sequence that is two amino acid in length at the given junction of two adjacent T cell epitope containing sequences to a modified multiepitopic polypeptide sequence that has an amino acid linker sequence at the given junction of two adjacent T cell epitope containing sequence that is three amino acid in length.

19. The method of any one of claims 12-18, wherein each of the inserting steps of a fourth linker selection round comprises inserting, at a given junction of two adjacent T cell epitope containing sequences, an amino acid linker sequence that is four amino acids in length.

20. The method of claim 19, wherein each of the comparing steps of the fourth linker selection round comprises comparing a multiepitopic polypeptide sequence that either lacks an amino acid linker sequence or has an amino acid linker sequence that is three amino acid in length at the given junction of two adjacent T cell epitope containing sequences to a modified multiepitopic polypeptide sequence that has an amino acid linker sequence at the given junction of two adjacent T cell epitope containing sequence that is four amino acid in length.

21. The method of any one of claims 1-20, wherein each of the comparing steps of a given linker selection round comprises comparing a fitness score of a given modified multiepitopic polypeptide comprising an amino acid linker sequence at a given junction of two adjacent T cell epitope containing sequences that is one amino acid in length longer than the length of the amino acid linker sequence inserted at the given junction of two adjacent T cell epitope containing sequences to a fitness score of a given multiepitopic polypeptide comprising no amino acid linker sequence or that has an amino acid linker sequence at the given junction of-I l l-two adjacent T cell epitope containing sequences that is one amino acid in length shorter than the length of the amino acid linker sequence inserted at the given junction of two adjacent T cell epitope containing sequences of the given modified multiepitopic polypeptide.

22. The method of any one of claims 1-21, wherein selecting comprises selecting a multi epitopic polypeptide sequence based a combined sequence of a multiepitopic polypeptide sequence not being present in the human proteome, wherein the combined sequence comprises a sequence of at least 8 amino acids, wherein:(A) when the multi epitopic polypeptide sequence comprises an amino acid linker sequence consisting of 1, 2, 3 or 4 amino acids: the sequence of at least 8 amino acids comprises at least one amino acid of the amino acid linker sequence and at least one amino acid of a T cell epitope containing sequence adjacent to the amino acid linker sequence, or(B) when the multiepitopic polypeptide sequence does not comprises an amino acid linker sequence: the sequence of at least 8 amino acids comprises at least one amino acid of each of the two adjacent T cell epitope containing sequences.

23. The method of any one of claims 1-22, wherein the fitness score of a given multi epitopic polypeptide sequence is based on a cleavability score of each T cell epitope containing sequence, wherein the cleavability score of a T cell epitope containing sequence is indicative of the probability that a T cell epitope is cleaved from the given multiepitopic polypeptide sequence when the multiepitopic polypeptide sequence is expressed in a cell.

24. The method of any one of claims 1-23, wherein the fitness score of a given multi epitopic polypeptide sequence is a function of the cleavability score of all T cell epitope containing sequences in the given multiepitopic polypeptide sequence, optionally, wherein the fitness score of a given multiepitopic polypeptide sequence is based on the average cleavability score of all T cell epitope containing sequences in the given multiepitopic polypeptide sequence.

25. The method of any one of claims 1-24, wherein the cleavability score of a T cell epitope containing sequence is based on the T cell epitope containing sequence and the amino acid sequence upstream and / or downstream of the T cell epitope containing sequence.

26. The method of any one of claims 1-25, wherein the T cell epitope containing sequence and the amino acid sequence upstream and / or downstream of the T cell epitope containing sequence is at least 20, 25, or 30 amino acids in length.

27. The method of any one of claims 23-26, wherein the cleavability score and / or the fitness score is predicted by a trained prediction model implemented on a computer.

28. The method of any one of claims 1-25, wherein comparing comprises inputting amino acid sequence information of a given multiepitopic polypeptide sequence using a computer processor into a trained prediction model to generate a plurality of cleavability predictions.

29. The method of claim 28, wherein the amino acid sequence information comprises the T cell epitope containing sequence and the amino acid sequence upstream and / or downstream of the T cell epitope containing sequence.

30. The method of claim 28 or 29, wherein the plurality of cleavability predictions comprises a cleavability prediction for each T cell epitope of the given multiepitopic polypeptide sequence.

31. The method of claim 30, wherein each cleavability prediction of the plurality of cleavability predictions is indicative of a probability that a given T cell epitope is cleaved from a given multiepitopic polypeptide sequence is expressed in a cell.

32. The method of any one of claims 23-31, wherein the cleavability score is indicative of a probability that a given T cell epitope is cleaved from the multiepitopic polypeptide sequence when the multiepitopic polypeptide sequence is expressed in a cell.

33. The method of any one of claims 1- 32, wherein at least one T cell epitope containing sequence of the plurality comprises two or more T cell epitope sequences.

34. The method of any one of claims 1- 33, wherein at least one T cell epitope containing sequence of the plurality comprises two or more T cell epitope sequences that share at last one amino acid.

35. The method of claim 33 or 34, wherein the sequence of a first T cell epitope sequence of the two or more T cell epitope sequences starts at the N-terminus of the at least one T cell epitope containing sequence of the plurality that comprises two or more T cell epitope sequences, and wherein the sequence of a second T cell epitope sequence of the two or more T cell epitope sequences ends at the C-terminus of the at least one T cell epitope containing sequence of the plurality that comprises two or more T cell epitope sequences.

36. The method of any one of claims 1- 35, wherein inserting comprises randomly selecting two adjacent T cell epitope sequences of a multi epitopic polypeptide sequence and inserting an amino acid linker sequence into the multiepitopic polypeptide sequence between a junction of the two adjacent T cell epitope sequences randomly selected.

37. The method of any one of claims 1- 36, wherein the method is performed in a computer implemented machine learning program having a framework that comprises an ordering algorithm, an objective function, and a sampling function.

38. The method of claim 37, wherein the ordering algorithm comprising a simulated annealing module.

39. The method of claim 37, wherein the objective function comprises a cleavage predictor algorithm.

40. The method of claim 37, wherein the sampling function comprises an epitope sampling function and linker sampling function.

41. The method of any one of claims 1-40, wherein each of the junctions between two adjacent T cell epitope containing sequences are variable junctions.

42. The method of any one of claims 1-41, wherein introducing comprises one or more rounds of introducing, each round comprising:(a) inputting a potential amino acid linker sequence from a set of potential amino acid linker sequences at a given variable junction;(b) determining whether at each given variable junction, a potential amino acid linker sequence together with a sequence that is immediately upstream or a sequence immediately downstream of the potential amino acid linker sequence comprises a sequence of at least 8 consecutive amino acids that is present in the human proteome; and(c) discarding the multiepitopic polypeptide sequence if the sequence of at least 8 consecutive amino acids is found to be present in the human proteome.

43. The method of any one of claims 1-42, wherein the maximum number of amino acids for a given amino acid linker sequence is 4.

44. The method of any one of claims 1-43, wherein the minimum number of amino acids for a given amino acid linker sequence is 0.

45. The method of any one of claims 1-44, wherein selecting comprises preferentially selecting a multiepitopic polypeptide sequence having an amino acid linker sequence with a minimal length at one or more or each junction between two adjacent T cell epitope containing sequences.

46. A method for designing a multiepitopic polypeptide sequence from a plurality of T cell epitope containing sequences, each T cell epitope containing sequence of the plurality comprising or consisting of at least one T cell epitope sequence, wherein the plurality of T cell epitope containing sequences comprises at least three T cell epitope sequences, the method comprising performing a first T cell epitope containing sequence ordering round, the first T cell epitope containing sequence ordering round comprising:(a) ordering each T cell epitope containing sequence of the plurality into a first defined N to C terminal order, thereby generating a starting multiepitopic polypeptide sequence;(b) ordering each T cell epitope containing sequence of the plurality into a second defined N to C terminal order, thereby generating a first modified multiepitopic polypeptide sequence;(c) comparing a fitness score of the starting multiepitopic polypeptide sequence to a fitness score of the first modified starting multiepitopic polypeptide sequence; and(d) selecting a first multiepitopic polypeptide sequence from the starting multiepitopic polypeptide sequence and the first modified starting multiepitopic polypeptide sequence basedon the comparison of the fitness scores, thereby generating a first multiepitopic polypeptide sequence.

47. The method of claim 46, wherein the starting polypeptide sequence has a sequence comprising each T cell epitope containing sequence of the plurality directly linked to one or two other T cell epitope containing sequences of the plurality.

48. The method of claim 46 or 47, wherein the modified polypeptide sequence has a sequence comprising each T cell epitope containing sequence of the plurality directly linked to one or two other T cell epitope containing sequences of the plurality.

49. The method of any one of claims 46-48, wherein selecting comprises selecting the multiepitopic polypeptide sequence with the higher fitness score.

50. The method of any one of claims 46-48, wherein selecting comprises selecting the multiepitopic polypeptide sequence with the lower fitness score ifwherein the second defined order score is the fitness score of the modified starting multiepitopic polypeptide sequence and the first defined order score is the fitness score of the starting multiepitopic polypeptide sequence, and the random sample value is a random number from a standard uniform distribution defined on a closed interval of [0,1],51. The method of any one of claims 1-50, wherein the fitness score of a given multiepitopic polypeptide sequence is a fitness score computed for a predefined temperature parameter.

52. The method of any one of claims 46-51, wherein the score of the second defined order is at least 1.1 -fold higher than the score of the first defined order.

53. The method of any one of claims 46-52, further comprising progressively repeating (a)-(d) thereby selecting an order that is the highest possible score for a given multiepitopic polypeptide sequence.

54. The method of any one of claim 1-53, wherein a fitness score of a selected multiepitopic polypeptide sequence is at least 2.0.

55. The method of any one of claims 1-54, wherein the fitness score of a multi epitopic polypeptide sequence ranges from (-)4.0 to +4.0.

56. The method of any one of claims 1-55, wherein a cleavage predicting algorithm used to compute the fitness score has been trained with user-input data from experimentally verified T cell epitope sequences that are cleaved when expressed in a cell, as observed by mass spectrometry assay, an immunoassay or by a T cell activation assay.

57. The method of claim 56, wherein the immunoassay is a detection assay of an epitope that is bound to an MHC or a cell.

58. The method of claim 56, wherein the data from experimentally verified T cell epitope sequences comprises verified by a tetramer assay.

59. The method of claim 57, wherein the T cell activation assay comprises cytokine release assay by an activated T cell, as determined by ELISA or flow cytometry.

60. The method of any one of claims 46-59, wherein the method comprises performing the method according to any one of claims 1-45 after performing the first T cell epitope containing sequence ordering round.

61. The method of any one of claims 1-45, wherein the method comprises performing the method according to any one of claims 46-59 prior to performing the first linker selection round.

62. The method of any one of claims 1-61, wherein the method further comprises generating a therapeutic composition comprising the multiepitopic polypeptide sequence or a nucleic acid sequence encoding the multiepitopic polypeptide sequence.

63. The method of claim 62, wherein the therapeutic composition is a T cell vaccine composition.

64. The method of claim 62 or 63, wherein the multi epitopic polypeptide sequence comprises at least 10 T cell epitope sequences or from 10 to 100 T cell epitope sequences.

65. A pharmaceutical composition comprising the multi epitopic polypeptide sequence generated by a method of any one of claims 1-64, for treating a disease in a subject.

66. A method of treating a disease in a subject in need thereof, comprising administering the subject the pharmaceutical composition of claim 65, wherein the subject is a human subject.

67. Use of a composition comprising the multiepitopic polypeptide sequence of claim 65, or the multi epitopic polypeptide sequence generated by a method of any one of claims 1-64 in preparing a medicament for treating a disease in a subject in need thereof.

68. The pharmaceutical composition of claim 65, the method in accordance to claim 66, or the use in accordance to claim 67, wherein the disease is a cancer or an infectious disease.

69. A system for generating a multiepitopic polypeptide sequence from a plurality of T cell epitope sequences, the system comprising:(i) an input module for receiving a plurality of T cell epitope sequences; and(ii) framework module comprising one or more programmed functions, wherein the one or more programmed functions operate on one or more trained algorithms comprising: a. an ordering algorithm for transforming the plurality of T cell epitope sequences into a defined order of T cell epitope sequences from N terminus to C terminus, wherein the defined N-terminus to C-terminus order is a single possible order of all possible orders of the plurality of T cell epitope sequences, thereby forming a sequence having the defined N-terminus to C-terminus order; b. a cleavage predictor algorithm for predicting a probability that a T cell epitope sequence is cleaved and presented on a cell surface when the multiepitopic polypeptide sequence is expressed in a cell; andc. a set of sampling rules for the sampling function comprises an epitope sampling and linker sampling functions.

70. The system of claim 69, further comprising an output module comprising a display that displays the multiepitopic polypeptide sequence and a fitness score for the multiepitopic polypeptide sequence representing a probability of cleavage of each epitope sequence within the multiepitopic polypeptide sequence when the multiepitopic polypeptide is expressed in a cell.

71. The system of claim 69 or 70, wherein the one or more trained algorithms are machine learning algorithms that have been trained with user-input data from experimentally verified T cell epitope sequences and linker sequences that are cleaved when expressed in a cell.

72. The system of claim 71, wherein the user-input data from experimentally verified T cell epitope sequences comprise a T cell epitope sequence that has been verified to (i) bind to an MHC molecule encoded by an HLA allele by a peptide-MHC binding assay and / or affinity assay; (ii) be presented by an APC as determined by a tetramer assay; (iii) bind to a T cell in vitro as determined by immunoassay, and / or (iv) activate a T cell upon contacting as determined by a cytokine release assay.

73. The system of any one of claims 69-72, wherein the framework module is set to generate a multiepitopic polypeptide, or a polynucleotide sequence encoding the multiepitopic polypeptide containing epitopes and linkers, wherein an epitope encoded by the polynucleotide sequence is presented by a cell at an abundance that is at least 1.2 times the abundance with which the same epitope is presented that is encoded by (A) a control polynucleotide sequence; or (B) a polynucleotide sequence encoding a polypeptide (i) containing the same epitopes and with synthetic linkers not generated by the system, (ii) containing the same epitopes and linkers that are not in an order of epitopes and linkers as generated by the system, or (iii) containing a sequence that is not identical to the sequence of the multiepitopic polypeptide as generated by the framework module of the system.

74. The system of claim 73, wherein the control polynucleotide is a polynucleotide encoding a control polypeptide lacking epitopes that are identical to epitopes of the multiepitopic polypeptide generated by the framework module, or a control polypeptide that comprises the same epitopes but lacks any linker sequences.

75. The system of claim 73 or 74, wherein an epitope encoded by the polynucleotide sequence is presented by a cell at an abundance that is at least 2 times the abundance with which the same epitope is presented that is encoded by (A) a control polynucleotide sequence; or (B) a polynucleotide sequence encoding a polypeptide (i) containing the same epitopes and with synthetic linkers not generated by the system, (ii) containing the same epitopes and linkers that are not in an order of epitopes and linkers as generated by the system, or (iii) containing asequence that is not identical to the sequence of the multiepitopic polypeptide as generated by the framework module of the system.

76. A composition comprising a polynucleotide sequence encoding a multiepitopic polypeptide, the polynucleotide sequence having at least 90% sequence identity to a polynucleotide sequence set forth in any one of SEQ ID NOs: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, and 23.

77. A multiepitopic polypeptide encoded by a polynucleotide sequence having at least 90% sequence identity to a polynucleotide sequence set forth in any one of SEQ ID NOs: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21 and 23.

78. A multiepitopic polypeptide comprising an amino acid sequence that has at least 85% sequence identity to a polypeptide set forth in any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22 and 24.

79. A multiepitopic polypeptide comprising at least 3 consecutive epitope sequences and linkers set forth in any one of Tables 2-11.

80. A multi epitopic polypeptide comprising at least 2, 3, 4, 5, 6 or more consecutive epitope sequences and linkers set forth in any one of Tables 2-11 and at least one GS linker.

81. A multiepitopic polypeptide comprising at least 2, 3, 4, 5, 6 or more consecutive epitope sequences and linkers set forth in any one of Tables 2-11 and at least one helical linker.

82. A multiepitopic polypeptide comprising at least 2, 3, 4, 5, 6 or more consecutive epitope sequences and linkers set forth in any one of Tables 2-11 and at least one Furin linker.

83. The multi epitopic polypeptide of claim 82, wherein the Furin linker is a viral Furin linker.

84. The multiepitopic polypeptide of claim 82, wherein the Furin linker is a human Furin linker.

85. The multiepitopic polypeptide of any one of claims 78-84, further comprising a solubility enhancing modification.

86. The multiepitopic polypeptide of claim 85, wherein the solubility enhancing modification is inclusion of a sequence for an N-terminal mannose binding (MBP) protein.

87. The multiepitopic polypeptide of any one of claims 78-84, further comprising one or more mutations that minimize ribosomal stalling motifs.

88. A method for designing a multi epitopic polypeptide sequence from a plurality of T cell epitope containing sequences comprising: (A) a first sequence comprising at least one neoORF sequence and (B) at least two sequences of the plurality of T cell epitope containing sequences each comprising at least one neoepitope comprising a point mutation; the method comprising performing a first T cell epitope containing sequence ordering round, the first T cell epitope containing sequence ordering round comprising:(a) ordering each of the at least two sequences of (B) keeping (A) constant into a first defined N to C terminal order, thereby generating a starting multiepitopic polypeptide sequence;(b) ordering each of the at least two sequences of (B) keeping (A) constant into a second defined N to C terminal order, thereby generating a first modified multiepitopic polypeptide sequence;(c) comparing a fitness score of the starting multiepitopic polypeptide sequence to a fitness score of the first modified starting multiepitopic polypeptide sequence; and(d) selecting a first multiepitopic polypeptide sequence from the starting multiepitopic polypeptide sequence and the first modified starting multiepitopic polypeptide sequence based on the comparison of the fitness scores, thereby generating a first multiepitopic polypeptide sequence, wherein the fitness score of a multi epitopic polypeptide sequence ranges from (-)4.0 to +4.0; and wherein a fitness score of a selected first multi epitopic polypeptide sequence is at least 2.0.

89. The method of claim 88, further comprising, generating a T cell epitope vaccine comprising the first selected polypeptide sequence.

90. The method of claim 88 or 89, wherein the T cell epitope vaccine is a polynucleotide comprising a sequence encoding the first selected polypeptide sequence; or a polypeptide comprising the first selected polypeptide sequence.

91. The method of any one of claims 89-90, wherein the plurality of T cell epitope containing sequences comprise at least one T cell epitope related to Lynch syndrome.

92. The method of any one of claims 89-91, wherein the vaccine is a Lynch syndrome vaccine.

93. The method of claim 88, further comprising, generating antigen-specific T cells, the method comprising: (I) (i) loading antigen presenting cells (APCs) with a polypeptide comprising the first selected polypeptide sequence or a part thereof; or (ii) expressing a polynucleic acid encoding a polypeptide comprising the first selected polypeptide sequence or a part thereof; and (II) contacting the APCs with T cells and incubating ex vivo, thereby generating antigenspecific T cells.

94. The method of claim 88, further comprising, generating a TCR vaccine using the first selected polypeptide sequence, the method comprising: (I) (i) loading antigen presenting cells (APCs) with a polypeptide comprising the first selected polypeptide sequence or a part thereof; or (ii) expressing a polynucleic acid encoding a polypeptide comprising the first selected polypeptide sequence or a part thereof; and (II) contacting the APCs with T cells and incubating ex vivo, thereby obtaining antigen specific T cells and (III) identifying TCR sequences from the antigen specific T cells.

95. The method of any one of claims 93 or 94, wherein the first selected polypeptide sequence or a part thereof comprises a Lynch syndrome epitope.

6. A method for treating Lynch syndrome or a hepatic cancer, the method comprising administering a subject in need thereof with one or more of the vaccine of any one of claims 89-91, the antigen-specific T cells of claim 93, or the TCR vaccine of claim 94.