Immunogenic epitope prediction

WO2026206880A1PCT designated stage Publication Date: 2026-10-01MODEX THERAPEUTICS INC +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/020437
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-08-27
Filing Date
2026-03-23
Publication Date
2026-10-01

Smart Images

  • Figure IMGF000079_0001
    Figure IMGF000079_0001
  • Figure IMGF000077_0001_TABLE
    Figure IMGF000077_0001_TABLE
  • Figure 00000101_0000
    Figure 00000101_0000
Patent Text Reader

Abstract

The disclosure provides a method for predicting immunogenicity based on the probability that the protein will trigger cytokine release. Predicting immunogenicity is of interest in biologics research and development due to the risk for biologic treatments to elicit an Anti-Drug Antibody (ADA) response. While existing methods such as MHCRoBERTa, CcBHLA, and AbImmPred utilize binding affinity prediction and / or T-cell receptor recognition to infer whether a protein will elicit an immune response based on primary peptide sequence, this approach is limited due to the variable nature of immune responses triggered by epitope recognition. Our method addresses this issue by training a Large Language Model (LLM) on a dataset of epitopes experimentally verified as immunogenic based on cytokine release assays. Since cytokine release assays are more directly associated with specific immune responses, our tool allows for more reliable prediction of immunogenic responses, which enables de-risking clinical antibody development.
Need to check novelty before this filing date? Find Prior Art

Description

IMMUNOGENIC EPITOPE PREDICTIONCROSS-REFERENCE TO RELATED APPLICATIONSAND INCORPORATION BY REFERENCE

[0001] This PCT application claims the priority benefit of U.S. Provisional Application No. 63 / 776,854, filed on March 24, 2025, and U.S. Provisional Application No. 63 / 871,488, filed on August 27, 2025, both of which are herein incorporated by reference in their entireties.REFERENCE TO SEQUENCE LISTING SUBMITTED ELECTRONICALLY

[0002] The content of the electronically submitted ST.26 sequence listing in XML format (Name 4850_021PC02_SequenceListing_ST26.xml; Size: 10,053 bytes; and Date of Creation: March 20, 2026) filed with the application is incorporated herein by reference in its entirety.FIELD

[0003] The present disclosure relates to methods for predicting protein immunogenicity using predictive models derived from epitopes experimentally verified to induce cytokine release.BACKGROUND

[0004] Immunogenicity is defined as the inherent ability of a substance, such as a drug, therapeutic antibody, or vaccine, to elicit an immune response within the body. Immunogenicity is a pivotal concept in drug development as the immune response can range from beneficial, as with vaccines, to potentially adverse and unwanted, as in the case with therapeutic proteins such as antibodies. In the latter, immunogenicity can be detrimental and lead to the development of antidrug antibodies (ADAs) that neutralize the effectiveness of a therapeutic drug or cause unwanted side effects. Immunogenicity is a critical consideration in the development and approval of biologic drugs, as it directly influences their efficacy and safety. In conclusion, the exploration of immunogenicity in drug and vaccine development is critical to ensure the safety and effectiveness of new therapeutics. This complex interplay between biologies and the immune system underscores the importance of meticulous research and development in the pharmaceutical industry. Byunderstanding the factors that influence immunogenicity and the nature of immune responses, scientists and clinicians can better predict, manage, and mitigate potential risks associated with biopharmaceuticals. This knowledge is instrumental in the creation of safer, more effective therapeutic agents, and in ensuring regulatory compliance.

[0005] Currently, most immunogenicity prediction research focuses on peptide processing and presentation, e.g. proteasomal cleavage, transporter associated with antigen processing (TAP) and major histocompatibility complex (MHC) combination. Most epitope prediction research focuses on peptide processing and presentation, e.g. proteasomal cleavage, transporter associated with antigen processing (TAP) and major histocompatibility complex (MHC) combination. However, not all peptides with a high (predicted) binding affinity also elicit an immune response. Using such binding prediction models, which predict whether a peptide is likely to be presented on the MHC, and assuming that those selected peptides are also immunogenic neoantigens, results in a large number of false positive predictions. The IEDB immunogenicity predictor and Ineo-Epp are two publicly available immunogenicity prediction tools. The IEDB immunogenicity predictor is a simple model that rates peptides based on the position and properties of their amino acids. It is based on the position-wise enrichment of amino acids in immunogenic versus non-immunogenic peptides. The resulting immunogenicity score of a peptide is the sum of its individual amino acid scores. Ineo-Epp is a random forest classifier trained on HLA-I immunogenic peptides of which several characteristics, like amino acid physicochemical properties and eluted ligand likelihood percentile rank (EL rank%) score, were extracted.BRIEF SUMMARY

[0006] The present disclosure provides a computer-implemented method for predicting the immunogenicity of a candidate protein, comprising (i) generating one or more training datasets derived from protein sequence data and associated in vitro cytokine release data from a population of training proteins; (ii) training one or more machine learning (ML) models using the one or more training datasets of (i), wherein the one or more ML models are capable of predicting the ability to induce the cytokine release and immunogenicity of the candidate protein based on the primary sequence of the candidate protein; (iii) preprocessing the primary sequence of the candidate protein generating a plurality of peptide subsequences derived from the primary sequence of the candidate protein; (iv) applying the one or more ML models to each of the plurality of peptide sequences to generate immunogenicity scores associated with each peptide subsequence; and, (v) generating animmunogenicity prediction for the candidate protein based on the immunogenicity scores associated with each peptide subsequence and / or generating a cytokine release prediction for the candidate protein based on the immunogenicity scores associated with each peptide subsequence.

[0007] The present disclosure also provides a computer-implemented method for predicting the immunogenicity of a candidate protein, comprising (i) training one or more ML models using the one or more training datasets wherein the one or more training datasets are derived from protein sequence data and associated in vitro cytokine release data from a population of training proteins, and wherein the one or more ML models are capable of predicting the immunogenicity of the candidate protein based on the primary sequence of the candidate protein; (ii) preprocessing the primary sequence of the candidate protein generating a plurality of peptide subsequences derived from the primary sequence of the candidate protein; (iii) applying the one or more ML models to each of the plurality of peptide sequences to generate immunogenicity scores associated with each peptide subsequence; and, (iv) generating an immunogenicity prediction for the candidate protein based on the immunogenicity scores associated with each peptide subsequence.

[0008] Also provided is a computer-implemented method for predicting the immunogenicity of a candidate protein, comprising: (i) preprocessing the primary sequence of the candidate protein generating a plurality of peptide subsequences derived from the primary sequence of the candidate protein; (ii) applying the one or more pretrained ML models to each of the plurality of peptide sequences to generate immunogenicity scores associated with each peptide subsequence, wherein the one or more ML models are trained using the one or more training datasets wherein the one or more training datasets are derived from protein sequence data and associated cytokine release data from a population of training proteins, and wherein the one or more ML models are capable of predicting the immunogenicity of the candidate protein based on the primary sequence of the candidate protein; and, (iii) generating an immunogenicity prediction for the candidate protein based on the immunogenicity scores associated with each peptide subsequence.

[0009] In some aspects, the one or more training datasets in the computer-implemented method disclosed above (i) do not use protein-binding data, and / or (ii) do not use protein-binding affinity data, and / or (iii) do not use protein-binding affinity prediction data. In some aspects, the computer-implemented method does not comprise a protein secondary or tertiary structure prediction. In some aspects, the computer-implemented method does not comprise a humanness assessment.

[0010] In some aspects, generating one or more training datasets comprises (a) preprocessing at least one population of training proteins; (b) optionally performing data augmentation; and (c) labeling. In some aspects, preprocessing comprises filling or removing missing data values, removing data redundancies, and handling data ambiguities associated with at least one population of training proteins. In some aspects, preprocessing the population of training proteins comprises (i) splitting the amino acid sequences of the training proteins into training peptide subsequences, and (ii) tokenizing the training peptide subsequences to generate embeddings. In some aspects, data augmentation comprises splitting the amino acid sequence of a training protein into one or more shorter training proteins. In some aspects, labeling comprises assigning a plurality of labels indicative of an immunogenicity or non-immunogenicity category of each training peptide subsequence derived from the at least one population of training proteins. In some aspects, labeling comprises converting a multi-class label associated with a training peptide subsequence located within the sequence of a training protein of the one or more training datasets to a binary-class label. In some aspects, the converting comprises assigning the binaryclass label, as positive based on determining that an IC50 value associated with the training peptide subsequence is larger than a limit threshold, or as negative based on determining that the IC50 value associated with the training peptide subsequence is smaller than the limit threshold.

[0011] In some aspects, the training peptide subsequences are overlapping. In some aspects, the population of training proteins comprises a set of immunogenic proteins. In some aspects, the population of training proteins comprises a plurality of protein subsequences from the set of immunogenic proteins and a plurality of data labels indicative of a categorized information of each protein sequence in the set of immunogenic proteins. In some aspects, the population of training proteins comprises a set of non-immunogenic proteins. In some aspects, the population of training proteins comprises a plurality of protein subsequences from the set of non-immunogenic proteins and a plurality of data labels indicative of a categorized information of each protein sequence in the set of non-immunogenic proteins.

[0012] In some aspects, the plurality of peptide subsequences derived from the amino acid sequences of the population of training proteins are generated using a sliding window and a sliding step. In some aspects, the sliding window has a variable length. In some aspects, the sliding window has a constant length. In some aspects, the sliding window constant length is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 amino acids, or the full protein length. In some aspects, the sliding window constant length is 15 amino acids. In some aspects, the sliding step has a variable length. In some aspects, the sliding step has a constant length. In some aspects,the sliding step length is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids or the full protein length. In some aspects, the sliding step length is 1 amino acid. In some aspects, the sliding window has a length of 15 amino acids, and the sliding step has a length of 1 amino acid.

[0013] In some aspects, the plurality of peptide subsequences derived from the amino acid sequences of the population of training proteins are overlapping. In some aspects, the plurality of peptide subsequences derived from the amino acid sequences of the population of training proteins are non-overlapping. In some aspects, the plurality of peptide subsequences derived from the amino acid sequences of the population of training proteins have the same length. In some aspects, the plurality of peptide subsequences derived from the amino acid sequences of the population of training proteins can have different lengths.

[0014] In some aspects, the training comprises fine-tuning a pre-trained LLM and / or finetuning a pre-trained LSTM. In some aspects, the protein sequence data comprises only primary sequence data and cytokine release data. In some aspects, the protein sequence data further comprised protein secondary structure data and / or protein tertiary structure data. In some aspects, the preprocessing of at least one population of training proteins protein comprises iterating the amino acid sequence of each training protein sequence using a sliding window and sliding step to generate a plurality of peptide subsequences. In some aspects, preprocessing further comprises tokenizing the plurality of peptide subsequences derived from the primary sequence of a training protein.

[0015] In some aspects, tokenizing a peptide subsequence derived from the primary sequence of a training protein comprises generating a numerical sequence associated with the peptide subsequence with a format applicable to the one or more ML models. In some aspects, tokenizing comprises using a pre-defined tokenizer to encode the amino acids within the peptide subsequence. In some aspects, tokenizing coverts each amino acid to a numeric and / or vector token. In some aspects, class weights in the loss function are scaled to account for class imbalance or class weights in the loss function are not scaled to account for class imbalance.

[0016] In some aspects, the ML model is a pretrained model or a newly trained model. In some aspects, the ML model is selected from the group consisting of Artificial Neural Network (ANN) model, Convolutional Neural Network (CNN) model, Recurrent Neural Networks (RNN) model, Large Language Model (LLM), or a combination thereof. In some aspects, the ML model comprises an RNN. In some aspects, the RNN is a Long Short-Term Memory (LSTM) model. In some aspects, the ML model comprises an LLM. In some aspects, the LLM is an EvolutionaryScale Modelling 2 (ESM-2) model. In some aspects, the LLM is a fine-tuned ESM-2 model. In some aspects, the fine-tuning of the fine-tuned ESM-2 model comprises further training a plurality of layers using the training data wherein the weights of the remaining layers are not updated / trained.

[0017] In some aspects, the ML model further comprises a classification layer. In some aspects, the binary classification layer outputs the probability of immunogenicity of a peptide. In some aspects, the binary classification layer outputs the probability of non-immunogenicity of a peptide. In some aspects, the plurality of peptide subsequences derived from the candidate protein sequence are generated using a sliding window and a sliding step. In some aspects, the sliding window applied to the candidate protein sequence has a variable length. In some aspects, the sliding window applied to the candidate protein sequence has a constant length. In some aspects, the sliding window constant length is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 amino acids, or the entire protein. In some aspects, the sliding window constant length is 15 amino acids. In some aspects, the sliding step has a variable length. In some aspects, the sliding step has a constant length. In some aspects, the sliding step length is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids. In some aspects, the sliding step length is 1 amino acid. In some aspects, the sliding window has a length of 15 amino acids, and the sliding step has a length of 1 amino acid.

[0018] In some aspects, the plurality of peptide subsequences derived from the candidate protein sequence are overlapping. In some aspects, the plurality of peptide subsequences derived from the candidate protein sequence are non-overlapping. In some aspects, the plurality of peptide subsequences derived from the candidate protein sequence have the same length. In some aspects, the plurality of peptide subsequences derived from the candidate protein sequence can have different lengths.

[0019] In some aspects, the immunogenicity prediction comprises a binary prediction of cytokine release induced by the candidate protein or a fragment thereof in a subject. In some aspects, the immunogenicity prediction comprises a probability of cytokine release induced by the candidate protein or a fragment thereof in a subject. In some aspects, cytokine release comprises the release a cytokine selected from the group consisting of interleukin 1 (IL-1), interleukin 2 (IL-2), interleukin 3 (IL-3), interleukin 4 (IL-4), interleukin 6 (IL-6), interleukin 8 (IL-8), interleukin 10 (IL-10), interleukin 13 (IL-13), interleukin 15 (IL-15), tumor necrosis factor alpha (TNF alpha, TNFa, TNFoc), transforming growth factor beta (TGFbeta, TGFb, TGF[3), cluster of differentiation40 ligand (CD40L, CD40 Ligand, CD 154), interferon alpha (IFNalpha, IFNa, IFNoc), interferon beta (IFNbeta, IFNb, IFN[3), interferon gamma (IFNgamma, IFNg, IFNy), and any combination thereof.

[0020] In some aspects, the predicted cytokine release comprises predicted severity of the cytokine release. In some aspects, generating an immunogenicity prediction for the candidate protein based comprises identifying one or more peptide subsequences having immunogenicity probabilities above a predetermined threshold. In some aspects, generating an immunogenicity prediction for the candidate protein comprises identifying one or more immunogenicity hot spots having immunogenicity probabilities above a predetermined threshold. In some aspects, generating an immunogenicity prediction for the candidate protein comprises (i) determining a global candidate protein immunogenicity score, (ii) determining number of immunogenicity hotspots on the candidate protein, (iii) determining the distance between immunogenicity hotspots on the candidate protein, (iv) determining median of the lowest quantile (MQl),(v) determining median of the top quantile (MQ3), (vi) determining distance of lowest to top quantile (MQ3-MQ1), (vii) determining dispersion of lowest and top quantile (MQ3-MQ1) / MQ3 ), (viii) determining the length of the immunogenicity hotspots on the candidate protein, or (ix) a combination thereof.

[0021] In some aspects, determining the global candidate protein immunogenicity score comprises integrating the immunogenicity probabilities of each of the peptide subsequences in the candidate protein. In some aspects, determining the global candidate protein immunogenicity score comprises calculating a Z-score. In some aspects, the immunogenicity prediction is outputted as a graphic representation showing the immunogenicity probabilities of each of the peptide subsequences in the candidate protein along the sequence of the candidate protein, wherein immunogenic regions or immunogenic hotspots correspond to segments of the candidate protein sequence having the immunogenicity probabilities above a predetermined threshold, such as 0.5 (50% probability) or MQ3.

[0022] The present disclosure also provides a method to determine the immunogenicity of a candidate protein comprising applying the computer-implemented methods disclosed herein to the candidate protein. In some aspects, the candidate protein is an immunogenic protein. In some aspects, the candidate protein is a non-immunogenic protein.

[0023] Also provided is a method to identify immunogenic hot spots in a candidate protein, wherein the method comprises applying a computer-implemented method disclosed herein to the candidate protein. In some aspects, the method comprises identifying (i) number of immunogenichot spots; (ii) location of the immunogenic hot spots; (iii) length of the immunogenic hot spots; (iv) degree of immunogenicity of the hots spots; (v) median of the lowest quantile (MQ1); (vi) median of the top quantile (MQ3); (vii) distance of lowest to top quantile (MQ3-MQ1); (viii) dispersion of lowest and top quantile (MQ3-MQ1) / MQ3 ); (ix) distance between immunogenic hot spots; (x) specific properties of the immunogenic hot spots; or, (xi) any combination thereof. In some aspects, the specific properties of the hotspots comprise charge, polarity, hydrophobicity, amino acid composition, surface / buried location, and any combination thereof.

[0024] The present disclosure also provides a method of predicting the immunogenicity of a peptide or protein, wherein the method comprises (i) inputting the amino acid sequence of the peptide or protein into a computer system comprising an immunogenicity prediction program comprising instructions for executing a computer-implemented method disclosed herein, wherein the peptide or protein is the candidate protein; (ii) executing in the computer system the computer-implemented method; and, (iii) receiving from the computer system the immunogenicity prediction resulting from executing the computer-implemented method. In some aspects, the code to execute the computer-implemented method, the training database of the computer-implemented method, the ML model of the computer-implemented method, the output of the ML model of the computer-implemented method, the immunogenicity prediction outputted by the computer-implemented method, or any combination thereof are hosted in a cloud computing environment.

[0025] The present disclosure also provides a non-transitory computer readable storage medium storing a computational module for predicting the immunogenicity of a protein, the computational module comprising code to execute a computer-implemented method disclosed herein or a portion thereof. Also provided is a computer readable storage medium having computer readable instructions to instruct a computer to perform a computer-implemented method disclosed herein or a portion thereof. The disclosure also provides a system for predicting the immunogenicity of a candidate protein according to a computer-implemented method disclosed herein, wherein the system is stored on computer-readable storage media.

[0026] Also provided in the present disclosure is a computer system for predicting the immunogenicity of a protein, the computer system comprising at least one processor and memory storing at least one program for execution by the at least one processor, the at least one program comprising instructions to execute a computer-implemented method disclosed herein. Also provided is a method to engineer a therapeutic protein comprising (i) predicting the immunogenicity of the therapeutic protein by applying a computer-implemented method disclosed herein to the therapeutic protein, wherein the therapeutic protein is the candidate protein; (ii)introducing one or more amino acid modifications to the therapeutic protein of step (i) to generate a modified therapeutic protein; (iii) predicting the immunogenicity of the modified protein by applying the computer-implemented to the modified therapeutic protein, wherein the modified therapeutic protein is the candidate protein; (iv) comparing the predicted immunogenicity of the therapeutic protein of (i) and the modified therapeutic protein of (iii); (v) optionally iterating (ii)-(iv) until the modified therapeutic protein has the desired degree of immunogenicity or lack thereof.

[0027] In some aspects, the amino acid modifications are selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof. In some aspects, the amino acid substitutions are conservative substitutions, non-conservative substitutions, or a combination thereof.

[0028] Also provided is a method to screen in silico a library of candidate proteins to determine their predicted immunogenicity, the method comprising applying the computer-implemented methods of immunogenicity prediction disclosed herein to the amino acid sequences of the library of candidate proteins. The present disclosure also provides a method of de-epitoping a therapeutic candidate protein comprising (i) applying a computer-implemented method of immunogenicity prediction disclosed herein to the therapeutic candidate protein or a modified variant thereof, (ii) applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof to the therapeutic candidate protein, (iii) applying the computer-implemented method to the therapeutic candidate protein or a modified variant thereof, and (iv) iterating steps (i) to (iii) to obtain a modified variant of the therapeutic candidate protein containing less antigenic epitopes than the therapeutic candidate protein. Also provided is a method to rescue a protein or peptide drug that triggers negative immunogenic effects by applying the methods of de-epitoping a therapeutic candidate protein disclosed herein.

[0029] The present disclosure also provides a method to predict negative responses to a drug comprising or consisting of a protein, the method comprising applying the computer-implemented method of immunogenicity predictions disclosed herein to the drug. In some aspects, the protein drug is used or is a candidate for use in a clinical trial.

[0030] The present disclosure also provides a method of vaccine design comprising applying the computer-implemented method of immunogenicity prediction to a vaccine candidate protein. In some aspects, vaccine design comprises selecting the vaccine candidate protein if thepredicted immunogenicity of the vaccine candidate protein is above a predetermined threshold value. In some aspects, vaccine design comprises modifying the vaccine candidate protein by applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof to the vaccine candidate protein, and iteratively executing the computer-implemented method of immunogenicity prediction.

[0031] Also provided is a method of modulating the immunogenicity of a candidate protein comprising applying the computer-implemented methods of immunogenicity prediction disclosed herein to the candidate protein. In some aspects, modulating comprises increasing the immunogenicity of the candidate protein by applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof to the candidate protein. In some aspects, modulating comprises decreasing the immunogenicity of the candidate protein by applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof to the candidate protein. In some aspects, modulating comprises introducing at least one immunogenic hot spot on the protein by applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof to the candidate protein. In some aspects, modulating comprises removing at least one immunogenic hot spot from the protein by applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof to the candidate protein.

[0032] The present disclosure provides a method of identifying an immunogenic epitope susceptible to antibody targeting in a candidate protein comprising applying the computer-implemented method disclosed herein to the candidate protein. Also provided is a method of generating an amino acid sequence predicted to have altered immunogenicity compared to the amino acid sequence a parent candidate protein comprising applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof to the amino acid sequence of parent candidate protein to generate theamino acid sequence of the modified candidate protein and determining the predicted immunogenicity of the modified candidate protein by applying the computer-implemented method disclosed herein to the amino acid sequence of modified candidate protein.

[0033] In some aspects of the computer-implemented methods disclosed herein the candidate protein is selected from the group consisting of a vaccine immunogen, an antibody, MSTAR, scFab, Fab, scFv, Fv, Fc, an enzyme, a growth factor, a chimeric antigen receptor (CAR), a T-cell receptor (TCR), a cytokine, chimeric cytokine or a cytokine receptor, glucagon-like peptide 1 (GLP1), glucagon-like peptide 2 (GLP2), and their close receptor and close analogues; and soluble and non-soluble forms of receptors mentioned above.

[0034] In some aspects, the computer-implemented methods disclosed herein further comprise evaluating the strength of interactions between (i) the candidate protein in a complex with the Major Histocompatibility Class I (MHCI) protein, (ii) the candidate protein in a complex with the Major Histocompatibility Class II (MHCII) protein, (iii) the candidate protein in a complex with the MHCI and MHCII proteins, (iv) the candidate protein and a T-cell receptor, or (v) a combination thereof. In some aspects, the computer-implemented methods disclosed herein further comprise evaluating the strength of interactions between (i) the candidate protein and (ii) the MHCII protein. In some aspects, the computer-implemented methods disclosed herein further comprise evaluating the strength of interactions between (i) the candidate protein and (ii) the protein-binding cavity of a complex comprising a MHCII protein and a T cell receptor. In some aspects, the computer-implemented methods disclosed herein further comprise inputting sequence information of the MHCII protein and / or the T cell receptor.

[0035] The present disclosure also provides a method to predict the immunogenicity of a plurality of fragments of a protein associated with cancer, and identifying a fragment of said protein that is predicted to be immunostimulatory comprising applying the computer-implemented method of the present disclosure to the plurality of fragments. Also provided is a method of producing a personalized cancer vaccine for a subject having a tumor, the method comprising the steps of (i) identifying a plurality of modified peptides expressed in the tumor, each comprising an amino acid substitution at a position, relative to a corresponding parent peptide expressed in the normal cells; (ii) determine the immunogenicity for each of the plurality of modified peptides, via the computer-implemented method of the present disclosure, wherein modified peptide is a candidate protein; and (iii) producing a personalized cancer vaccine for the subject, which comprises a peptide or polypeptide comprising the at least one modified peptide selected as immunogenic.

[0036] Also provided is a method of generating a library of immunogenic proteins and their predicted immunogenicity, wherein each immunogenic protein and / or their predicted immunogenicity is generated via the computer-implemented method of the present disclosure. Also provided is a method of selecting a therapeutic protein comprising applying the computer-implemented method of the present disclosure to a population of candidate proteins. Also provided is an immunogenicity-optimized protein sequence generated using the computer-implemented method of the present disclosure. Also provided is an immunogenic peptide identified based on a report generated by using the computer-implemented method of the present disclosure. Also provided is a vaccine protein sequence generated by using the computer-implemented method of the present disclosure.

[0037] Also provided is an amino acid sequence with desired immunogenic characteristics identified using the computer-implemented method of the present disclosure. Also provided is a protein having an amino acid sequence with desired immunogenic characteristics identified using the computer-implemented method of the present disclosure. Also provided is a polynucleotide encoding a protein having an amino acid sequence with desired immunogenic characteristics identified using the computer-implemented method of the present disclosure. Also provided is a vector comprising a polynucleotide encoding a protein having an amino acid sequence with desired immunogenic characteristics identified using the computer-implemented method of the present disclosure. Also provided is cell comprising a polynucleotide encoding a protein having an amino acid sequence with desired immunogenic characteristics identified using the computer-implemented method of the present disclosure. Also provided is a cell comprising a vector comprising a polynucleotide encoding a protein having an amino acid sequence with desired immunogenic characteristics identified using the computer-implemented method of the present disclosure. In some aspects, the cell is a T cell, a stem cell, a B cell, a NK cell, a bacterial cell, a dendritic cell, a mammalian cell, a yeast cell, or an insect cell. Also provided is a pharmaceutical composition comprising a protein, polynucleotide, vector, or the cell disclosed herein, and an excipient. The present disclosure also provides a delivery system comprising the protein, polynucleotide, or vector disclosed herein. In some aspects, the delivery system comprises a lipid nanoparticle, wherein the nanoparticle encapsulates the protein, polynucleotide, or vector.

[0038] The present disclosure provides a method to treat or prevent a disease or condition comprising administering the protein, polynucleotide, vector, cell, pharmaceutical composition, or delivery system disclosed herein to a subject. In some aspects, the disease or condition is selectedfrom the group consisting of cancer, infection, chronic inflammation, genetic disease, and autoimmune disease. In some aspects, the subject is a human subject.BRIEF DESCRIPTION OF THE DRAWINGS / FIGURES

[0039] FIG. 1 is a schematic representation of the immunogenicity prediction system of the present disclosure. Using primary protein sequence data and associated in vitro cytokine release data of thousands of proteins, and without using either experimentally determined or inferred / predicted MHC binding or affinity data, the encoder of a Protein Language Model (PLM) extracts relevant features, which are used by a classifier layer to predict immunogenicity. This model architecture has been shown to predict accurately the immunogenicity of peptide subsequences in protein candidate sequences such as antibodies.

[0040] FIG. 2 is an adalimumab performance comparison between the IFNy immunogenicity model (an implementation of the immunogenicity prediction system of the present disclosure trained using an IFNy release training set) and the IEDB CD4 Episcore method (Dhanda et al. (2018) “Predicting HLA CD4 Immunogenicity in Human Populations.” Frontiers in immunology 9: 1369. The IFNy Immunogenicity Model correctly predicted both epitopes, whereas the IEDB Method only predicted one of the epitopes.

[0041] FIG. 3 is an immunogenicity prediction comparison for heavy chain of Adalimumab using the True Negatives and True Positives models and the IEDB CD4 Episcore method. The True Negatives and True Positives models are different models from the IFNy Immunogenicity Model of FIG. 2. Whereas the IFNy Immunogenicity Model is trained only on epitopes that induce release of IFNy, the True Negative and True Positive models are trained on epitopes that induce the release of other cytokines such as IL-5, IL-2, IL-10, IL-4, IL-17, and relate to cytokine release-related parameters such as proliferation, cytotoxicity, activation, and qualitative binding.

[0042] FIGS. 4A and 4B show that the IFNy Immunogenicity Model correctly predicts the high immunogenicity potential for thrombopoietin and low immunogenic potential for follitropin-beta. FIG. 4A is a plot showing the probability of immunogenicity for thrombopoietin peptides.FIG. 4B is a plot showing the probability of immunogenicity probability for follitropin-beta peptides. Overlapping peptide subsequences are scanned for immunogenicity using a sliding window.

[0043] FIGS. 5A and 5B show that the True Negative Immunogenicity Model correctly predicts the high immunogenicity potential for thrombopoietin and low immunogenic potential for follitropin-beta. FIG. 5A is a plot showing the probability of immunogenicity for thrombopoietin peptides. FIG. 5B is a plot showing the probability of immunogenicity probability for follitropin-beta peptides. Overlapping peptide subsequences are scanned for immunogenicity using a sliding window.

[0044] FIGS. 6A and 6B show that the True Positive Immunogenicity Model correctly predicts the high immunogenicity potential for thrombopoietin and low immunogenic potential follitropin-beta. FIG. 6A is a plot showing the probability of immunogenicity for thrombopoietin peptides. FIG. 6B is a plot showing the probability of immunogenicity probability for follitropin-beta peptides. Overlapping peptide subsequences are scanned for immunogenicity using a sliding window.

[0045] FIG. 7 exemplifies the number of immunogenic and non-immunogenic peptides / protein included in the datasets used to train, validate, and test the different predictive models disclosed herein.

[0046] FIG. 8 is a scale showing proteins ranked based on their immunogenicity propensity determined experimentally.

[0047] FIGS. 9A and 9B shows results for the confusion matrix and the ROC curve illustrating the EPITOP™ model’s ability to distinguish between immunogenic and non-immunogenic peptides. FIG. 9A is a confusion matrix highlighting the model’s accuracy in predicting each immunogenic class. FIG. 9B is a ROC curve showing the model’s overall ability to distinguish between the 2 immunogenic classes.

[0048] FIG. 10A, 10B, 10C and 10D are t-SNE visualizations of the embedding space. Each point is a peptide color coded based on the class labels (orange for immunogenic and blue for non-immunogenic). Prior to fine tuning the model embeddings show poor separation as well as limited performance in predicting immunogenic peptides. After fine-tuning the model embeddings shows greatly improved class separation. FIG. 10A shows embedding space of final encoding layer before fine-tuning. FIG. 10B shows embedding space of prediction layer before fine-tuning. FIG.10C shows embedding space of final encoding layer for after fine-tuning. FIG. 10D shows embedding space of prediction layer after fine tuning.

[0049] FIGS. 11A and 11B show the predictions of the EPITOP™ method compared with with the IEDB tool and experimentally validated immunogenic epitopes for adalimumab. FIG.11A shows data corresponding to epitopes in adalimumab’s VH region, whereas FIG. 11B shows data corresponding to epitopes in adalimumab’s VL region.

[0050] FIG. 12 shows how the IMMUNOTOP™ tool ranks proteins based on their immunogenicity potential.DETAILED DESCRIPTION

[0051] The disclosure provides a method for predicting immunogenicity based on the probability that the protein will trigger cytokine release. Predicting immunogenicity is of interest in biologies research and development due to the risk for biologic treatments to elicit an Anti-Drug Antibody response. Other available methods have considerable drawbacks. EpiMatrix use statistical analysis to predict immunogenic propensity based on statistical sequence similarity and sequence consensus and is limited by the choice of the proper approximation function, which in is turn can catch only the basic trends. Most existing machine learning methods, such as IEDB (Vita et al. (2005) Nucleic Acids Res. 53(D1):D436-D443), MHCRoBERTa (Wang et al. (2022) Brief. Bioinform. 23(3):bbab595), CcBHLA (Wu et al. (2023) bioRxiv: 2023-04), AblmmPred (Wang et al. (2024) PLoS One 19(2):e0296737), TLimmuno (Wang et al. (2023) Brief. Bioinform. 24(3): bbadl 16) and similar, utilize secondary MHCI / MHCII binding affinity prediction tools (NetMHC / NetMHCpan, NetMHCII / NetMHCIIpan and similar) and / or T-cell receptor recognition, at least in training, to infer whether a protein will elicit an immune response based on primary peptide sequence, this approach is limited due to the variable nature of immune responses triggered by epitope recognition. See, e.g., NetMHC (Lundegaard et al. (2008) Nucleic Acids Research 36(S2): W509-W512), NetMHCII (Jensen et al. (2018) Immunology 154(3):394-406), NetMHCpan / NetMHCIIpan (Reynisson et al. (2020) Nucleic Acids Research 48(W1):W449-W454).

[0052] In contrast, our method addresses this issue by training a Large Language Model on a dataset of epitopes experimentally verified as immunogenic based on cytokine release assays. Since cytokine release assays are more directly associated with specific immune responses, our tool allows for more reliable prediction of immunogenic responses, which will enable de-risking clinical antibody development.

[0053] Methods in the art generally use experimental binding data, inferred binding data, or predicted binding or affinity data. In other words, predictive methods in the art generally determine first whether a candidate sequence will interact with MHC, and based on thosepredictions, the methods infer that an immune response will be elicited. The methods disclosed herein bypass the binding prediction step. Instead, by using highly curated and comprehensive training sets of protein sequences experimentally known to trigger the release of specific cytokines, the present disclosure provides methods that can predict immunogenicity based on the probability that a candidate protein will trigger the release of a specific cytokine or set thereof.

[0054] Various terms relating to aspects of disclosure are used throughout the specification and claims. Such terms are to be given their ordinary meaning in the art, unless otherwise indicated. Other specifically defined terms are to be construed in a manner consistent with the definitions provided herein.Definitions

[0055] In order that the present description can be more readily understood, certain terms are first defined. Additional definitions are set forth throughout the detailed description.

[0056] It is to be noted that the term "a" or "an" entity refers to one or more of that entity; for example, "a nucleotide sequence," is understood to represent one or more nucleotide sequences. As such, the terms "a" (or "an"), "one or more," and "at least one" can be used interchangeably herein.

[0057] Furthermore, "and / or" where used herein is to be taken as specific disclosure of each of the two specified features or components with or without the other. Thus, the term "and / or" as used in a phrase such as "A and / or B" herein is intended to include "A and B," "A or B," "A" (alone), and "B" (alone). Likewise, the term "and / or" as used in a phrase such as "A, B, and / or C" is intended to encompass each of the following aspects: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).

[0058] It is understood that wherever aspects are described herein with the language "comprising," otherwise analogous aspects described in terms of "consisting of and / or "consisting essentially of' are also provided.

[0059] The terms "about," "comprising essentially of," or "consisting essentially of," refer to a value or composition that is within an acceptable error range for the particular value or composition as determined by one of ordinary skill in the art, which will depend in part on how the value or composition is measured or determined, / .< ., the limitations of the measurement system. For example, "about," "comprising essentially of," or "consisting essentially of," can mean within 1 or more than 1 standard deviation per the practice in the art. Alternatively, "about," "comprising essentially of," or "consisting essentially of," can mean a range of up to 10%.Furthermore, particularly with respect to biological systems or processes, the terms can mean up to an order of magnitude or up to 5-fold of a value. When particular values or compositions are provided in the specification and claims, unless otherwise stated, the meaning of "about," "comprising essentially of," or "consisting essentially of," should be assumed to be within an acceptable error range for that particular value or composition.

[0060] As used herein, the term "approximately," as applied to one or more values of interest, refers to a value that is similar to a stated reference value. In certain aspects, the term "approximately" refers to a range of values that fall within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less in either direction (greater than or less than) of the stated reference value unless otherwise stated or otherwise evident from the context (except where such number would exceed 100% of a possible value).

[0061] As described herein, any concentration range, percentage range, ratio range or integer range is to be understood to include the value of any integer within the recited range and, when appropriate, fractions thereof (such as one tenth and one hundredth of an integer), unless otherwise indicated.

[0062] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure is related. For example, the Concise Dictionary of Biomedicine and Molecular Biology, Juo, Pei-Show, 2nd ed., 2002, CRC Press; The Dictionary of Cell and Molecular Biology, 3rd ed., 1999, Academic Press; and the Oxford Dictionary of Biochemistry And Molecular Biology, Revised, 2000, Oxford University Press, provide one of skill with a general dictionary of many of the terms used in this disclosure.

[0063] Units, prefixes, and symbols are denoted in their Systeme International de Unites (SI) accepted form. The headings provided herein are not limitations of the various aspects of the disclosure, which can be had by reference to the specification as a whole. Accordingly, the terms defined are more fully defined by reference to the specification in its entirety.

[0064] Abbreviations used herein are defined throughout the present disclosure. Various aspects of the disclosure are described in further detail in the following subsections.

[0065] Units, prefixes, and symbols are denoted in their Systeme International de Unites (SI) accepted form. Numeric ranges are inclusive of the numbers defining the range. Unless otherwise indicated, nucleotide sequences are written left to right in 5' to 3' orientation. Amino acid sequences are written left to right in amino to carboxy orientation. The headings provided herein are not limitations of the various aspects of the disclosure, which can be had by reference to thespecification as a whole. Accordingly, the terms defined immediately below are more fully defined by reference to the specification in its entirety.

[0066] The term "about" is used herein to mean approximately, roughly, around, or in the regions of. When the term "about" is used in conjunction with a numerical range, it modifies that range by extending the boundaries above and below the numerical values set forth. In general, the term "about" can modify a numerical value above and below the stated value by a variance of, e.g., 15 percent, up or down (higher or lower). Thus, in some aspects, about is interchangeable with ± 15%. When particular values or compositions are provided in the application and claims, unless otherwise stated, the meaning of "about" should be assumed to be within an acceptable error range for that particular value or composition.

[0067] As described herein, any numerical range, concentration range, percentage range, ratio range or integer range is to be understood to include the value of any integer within the recited range and, when appropriate, fractions thereof (such as one-tenth and one-hundredth of an integer), unless otherwise indicated.

[0068] As used herein, the term "antigen binding polypeptide" refers to a polypeptide having the ability to specifically bind to one or more substances that induce an immune response (i.e., one or more antigens or epitopes).

[0069] As used herein, the term "antigen binding polypeptide complex" refers to a group of two, three, four, or more associated polypeptides, wherein at least one polypeptide has the ability to specifically bind to one or more antigens. An antigen binding polypeptide complex, includes, but is not limited to, an antibody or antigen binding fragment thereof.

[0070] The term "antibody" includes, without limitation, a glycoprotein immunoglobulin that binds specifically to an antigen and comprises at least two heavy (H) chains and two light (L) chains interconnected by disulfide bonds. Each H chain comprises a heavy chain variable region (abbreviated herein as VH) and a heavy chain constant region. The heavy chain constant region comprises three constant domains, CHI, CH2 and CH3. Each light chain comprises a light chain variable region (abbreviated herein as VL) and a light chain constant region. The light chain constant region comprises one constant domain, CL. The VH and VL regions can be further subdivided into regions of hypervariability, termed complementarity determining regions (CDRs), interspersed with regions that are more conserved, termed framework regions (FR). Each VH and VL comprises three CDRs and four FRs, arranged from amino-terminus to carboxy-terminus in the following order: FR1, CDR1, FR2, CDR2, FR3, CDR3, and FR4. The variable regions of the heavy and light chains contain a binding domain that interacts with an antigen. The constantregions of the antibodies may mediate the binding of the immunoglobulin to host tissues or factors, including various cells of the immune system (e.g., effector cells) and the first component (Clq) of the classical complement system. A heavy chain may have the C-terminal lysine or not. Unless specified otherwise herein, the amino acids in the variable regions are numbered using the Kabat numbering system and those in the constant regions are numbered using the EU system.

[0071] The term "monoclonal antibody," as used herein, refers to an antibody that is produced by a single clone of B-cells and binds to the same epitope. In contrast, the term "polyclonal antibody" refers to a population of antibodies that are produced by different B-cells and bind to different epitopes of the same antigen. The term "antibody" includes, by way of example, monoclonal and polyclonal antibodies; chimeric and humanized antibodies; human or non-human antibodies; wholly synthetic antibodies; and single chain antibodies. A non-human antibody can be humanized by recombinant methods to reduce its immunogenicity in man.

[0072] The antibody can be an antibody that has been altered (e.g., by mutation, deletion, substitution, conjugation to a non-antibody moiety). For example, an antibody can include one or more variant amino acids (compared to a naturally occurring antibody) which change a property (e.g., a functional property) of the antibody. For example, several such alterations are known in the art, which affect, e.g., half-life, effector function, and / or immune responses to the antibody in a patient. The term antibody also includes artificial polypeptide constructs, which comprise at least one antibody-derived antigen binding site.

[0073] An "antigen binding fragment" of an antibody refers to one or more fragments or portions of an antibody that retain the ability to bind specifically to the antigen bound by the whole antibody. It has been shown that the antigen binding function of an antibody can be performed by fragments or portions of a full-length antibody. An antigen-binding fragment can contain the antigenic determining regions of an intact antibody (e.g., the complementarity determining regions (CDRs)). Examples of antigen binding fragments of antibodies that can be used in the targeted delivery systems of the present disclosure include, but are not limited to, Fab, Fab', F(ab')2, scFv, and Fv fragments, linear antibodies, and single chain antibodies. An antigen-binding fragment of an antibody can be derived from any animal species, such as rodents (e.g., mouse, rat, or hamster) and humans or can be artificially produced.

[0074] Furthermore, although the two domains of the Fv fragment, VL and VH, are coded for by separate genes, they can be joined, using recombinant methods, by a synthetic linker that enables them to be made as a single protein chain in which the VL and VH regions pair to form monovalent molecules (known as single chain Fv (scFv); see, e.g., Bird et al. (1988) Science242:423-426; and Huston et al. (1988) Proc. Natl. Acad. Sci. USA 85:5879-5883). Such single chain antibodies are also intended to be encompassed within the term "antigen-binding fragment" of an antibody.

[0075] Antigen binding fragments are obtained using conventional techniques known to those with skill in the art, and the fragments screened for utility in the same manner as are intact antibodies. Antigen binding fragments can be produced by recombinant DNA techniques, or by enzymatic or chemical cleavage of intact immunoglobulins.

[0076] As used herein, the term "variable region" typically refers to a portion of an antibody, generally, a portion of a light or heavy chain, typically about the amino-terminal 110 to 120 amino acids, or 110 to 125 amino acids in the mature heavy chain and about 90 to 115 amino acids in the mature light chain, which differ extensively in sequence among antibodies and are used in the binding and specificity of a particular antibody for its particular antigen. The variability in sequence is concentrated in those regions called Complementarity Determining Regions (CDRs) while the more highly conserved regions in the variable domain are called framework regions (FR). Without wishing to be bound by any particular mechanism or theory, it is believed that the CDRs of the light and heavy chains are primarily responsible for the interaction and specificity of an antibody with antigen. In some aspects, the variable region is a mammalian variable region, e.g., a human, mouse or rabbit variable region. In some aspects, the variable region comprises rodent or murine CDRs and human framework regions (FRs). In some aspects, the variable region is a primate (e.g., non-human primate) variable region. In some aspects, the variable region comprises rodent or murine CDRs and primate (e.g., non-human primate) framework regions (FRs).

[0077] The terms "complementarity determining region" or "CDR", as used herein, refer to each of the regions of an antibody variable domain which are hypervariable in sequence and / or form structurally defined loops (hypervariable loops) and / or contain the antigen-contacting residues. Antibodies can comprise six CDRs, e.g., three in the VH and three in the VL.

[0078] The terms "VL", "VL region," and "VL domain" are used herein interchangeably to refer to the light chain variable region of an antigen binding polypeptide, antigen binding polypeptide complex, antibody or antigen binding fragment thereof. In some aspects, a VL region is referred to herein as VL1 to denote a first light chain variable region, VL2 to denote a second light chain variable region, VL3 to denote a third light chain variable region, and VL4 to denote a fourth light chain variable region. An enumerated VL region (e.g., VL1) can have the same or different antigen binding properties and / or the same or different sequence as another enumerated VL region (e.g., VL2).

[0079] The terms "VH", "VH region," and "VH domain" are used herein interchangeably to refer to the heavy chain variable region of an antigen binding polypeptide, antigen binding polypeptide complex, antibody or antigen binding fragment thereof. In some aspects, a VH region is referred to herein as VH1 to denote a first heavy chain variable region, VH2 to denote a second heavy chain variable region, VH3 to denote a third heavy chain variable region, and VH4 to denote a fourth heavy chain variable region. An enumerated VH region (e.g., VH1) can have the same or different antigen binding properties and / or the same or different sequence as another enumerated VH region (e.g., VH2).

[0080] As used herein, "Kabat numbering" and like terms are recognized in the art and refer to a system of numbering amino acid residues in the heavy and light chain variable regions of an antibody or antigen binding fragment thereof. In some aspects, CDRs can be determined according to the Kabat numbering system (see, e.g., Kabat EA & Wu TT (1971) Ann NY Acad Sci 190: 382-391 and Kabat EA et al., (1991) Sequences of Proteins of Immunological Interest, Fifth Edition, U.S. Department of Health and Human Services, NIH Publication No. 91-3242). Using the Kabat numbering system, CDRs within an antibody heavy chain molecule are typically present at amino acid positions 31 to 35, which optionally can include one or two additional amino acids, following 35 (referred to in the Kabat numbering scheme as 35 A and 35B) (CDR1), amino acid positions 50 to 65 (CDR2), and amino acid positions 95 to 102 (CDR3). Using the Kabat numbering system, CDRs within an antibody light chain molecule are typically present at amino acid positions 24 to 34 (CDR1), amino acid positions 50 to 56 (CDR2), and amino acid positions 89 to 97 (CDR3).

[0081] As used herein, the terms "constant region" or "constant domain" are used interchangeably to refer to a portion of an antigen binding polypeptide, antigen binding polypeptide complex, antibody or antigen binding fragment thereof, e.g., a carboxyl terminal portion of a light and / or heavy chain which is not directly involved in binding of an antibody to antigen but which can exhibit various effector functions, such as interaction with the Fc region. The constant region generally has a more conserved amino acid sequence relative to a variable region. In some aspects, an antigen binding polypeptide, antigen binding polypeptide complex, antibody or antigen binding fragment thereof comprises a constant region or portion thereof that is sufficient for antibodydependent cell-mediated cytotoxicity (ADCC), antibody-dependent cellular phagocytosis (ADCP), and complement-dependent cytotoxicity (CDC). A constant region includes, but is not limited to, a light chain constant region (CL) or heavy chain constant region (CHI, CH2, and CH3).

[0082] As used herein, the terms "fragment crystallizable region," "Fc region," or "Fc domain" are used interchangeably herein to refer to the tail region of an antibody that interacts with cell surface receptors called Fc receptors and some proteins of the complement system. Fc regions typically comprise CH2 and CH3 regions, and, optionally, an immunoglobulin hinge.

[0083] As used herein, the terms "immunoglobulin hinge," "hinge," "hinge domain" or "hinge region" are used interchangeably to refer to a stretch of heavy chains between the Fab and Fc portions of an antigen binding polypeptide, antigen binding polypeptide complex, antibody or antigen binding fragment thereof. A hinge provides structure, position and flexibility, which assist with normal functioning of antibodies (e.g., for crosslinking two antigens or binding two antigenic determinants on the same antigen molecule). An immunoglobulin hinge is divided into upper, middle and lower hinge regions that can be separated based on structural and / or genetic components. An immunoglobulin hinge of the invention can contain one, two or all three of these regions. Structurally, the upper hinge region stretches from the C terminal end of CHI to the first hinge disulfide bond. The middle hinge region stretches from the first cysteine to the last cysteine in the hinge. The lower hinge region extends from the last cysteine to the glycine of CH2. The cysteines present in the hinge form interchain disulfide bonds that link the immunoglobulin monomers.

[0084] As used herein, the term "Fab" refers to a region of an antibody that binds to an antigen. It is typically composed of one constant and one variable domain of each of the heavy and the light chain.

[0085] As used herein, the term "heavy chain" refers to a portion of an antigen binding polypeptide, antigen binding polypeptide complex, antibody or antigen binding fragment thereof typically composed of a heavy chain variable region (VH), a heavy chain constant region 1 (CHI), a heavy chain constant region 2 (CH2), and a heavy chain constant region 3 (CH3). A typical antibody is composed of two heavy chains and two light chains. When used in reference to an antibody, a heavy chain can refer to any distinct type, e.g., alpha (a), delta (8), epsilon (a), gamma (y), and mu (p), based on the amino acid sequence of the constant region, which gives rise to IgA, IgD, IgE, IgG, and IgM classes of antibodies, respectively, including subclasses of IgG, e.g., IgGl, IgG2, IgG3, and IgG4. Heavy chain amino acid sequences are known in the art. In some aspects, the heavy chain is a human heavy chain.

[0086] As used herein, the term "light chain" refers to a portion of an antigen binding polypeptide, antigen binding polypeptide complex, antibody or antigen binding fragment thereof typically composed of a light chain variable region (VL) and a light chain constant region (CL). Atypical antibody is composed of two light chains and two heavy chains. When used in reference to an antibody, a light chain can refer to any distinct type, e.g., kappa (K) or lambda (1), based on the amino acid sequence of the constant region. Light chain amino acid sequences are known in the art. In some aspects, the light chain is a human light chain.

[0087] The term "chimeric" antibody or antigen-binding fragment thereof refers to an antibody or antigen binding fragments thereof wherein the amino acid sequence is derived from two or more species. Typically, the variable region of both light and heavy chains corresponds to the variable region of antibodies or antigen binding fragments thereof derived from one species of mammals (e.g., mouse, rat, rabbit, etc.) with the desired specificity, affinity and capability, while the constant regions are homologous to the sequences in antibodies or antigen binding fragments thereof derived from another (usually human) to avoid eliciting an immune response in that species.

[0088] The term "humanized" antibody or antigen binding fragment thereof refers to forms of non-human (e.g., murine) antibodies or antigen binding fragments that are specific immunoglobulin chains, chimeric immunoglobulins, or fragments thereof that contain minimal non-human (e.g., murine) sequences. Typically, humanized antibodies or antigen binding fragments thereof are human immunoglobulins in which residues from a complementary determining region (CDR) are replaced by residues from a CDR of a non-human species (e.g., mouse, rat, rabbit, hamster) that have the desired specificity, affinity, and capability (Jones et al., Nature 321:522-525 (1986); Riechmann et al., Nature 332:323-327 (1988); Verhoeyen et al., Science 239:1534-1536 (1988)). In some aspects, the Fv framework region (FR) residues of a human immunoglobulin are replaced with the corresponding residues in an antibody or fragment from a non-human species that has the desired specificity, affinity, and capability. The humanized antibody or antigen binding fragment thereof can be further modified by the substitution of additional residues either in the Fv framework region and / or within the replaced non-human residues to refine and optimize antibody or antigen-binding fragment thereof specificity, affinity, and / or capability. In general, a humanized antibody or antigen binding fragment thereof will comprise substantially all of at least one, and typically two or three, variable domains containing all or substantially all of the CDR regions that correspond to the non-human immunoglobulin whereas all or substantially all of the FR regions are those of a human immunoglobulin consensus sequence. A humanized antibody or antigen binding fragment thereof can also comprise at least a portion of a constant region, typically that of a human immunoglobulin. Examples of methods used to generate humanized antibodies are known and described, for example, in U.S. Pat. No.5,225,539; Roguska et al., Proc. Natl. Acad. Sci., USA, 91(3):969-973 (1994), and Roguska et al., Protein Eng. 9(10):895-904 (1996).

[0089] The term "human" antibody or antigen-binding fragment thereof, as used herein, means an antibody or antigen-binding fragment thereof having an amino acid sequence derived from a human immunoglobulin gene locus, where such antibody or antigen-binding fragment is made using recombinant techniques known in the art. This definition of a human antibody or antigen-binding fragment thereof includes intact or full-length antibodies and fragments thereof.

[0090] A polypeptide, polypeptide complex, antibody, antigen binding fragment thereof, polynucleotide, vector or host cell which is "isolated" is a polypeptide, polypeptide complex, antibody, antigen binding fragment thereof, polynucleotide, vector or host cell which is in a form not found in nature. Isolated polypeptides, polypeptide complexes, antibodies, antigen binding fragments thereof, polynucleotides, vectors or host cells include those which have been purified to a degree that they are no longer in a form in which they are found in nature. In some aspects, a polypeptide, polypeptide complex, antibody, antigen-binding fragment thereof, polynucleotide, vector or host cell which is isolated is substantially pure. As used herein, "substantially pure" refers to material which is at least 50% pure (i.e., free from contaminants), at least 90% pure, at least 95% pure, at least 98% pure, or at least 99% pure.

[0091] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acids of any length. The polymer can be linear or branched, it can comprise modified amino acids, and it can be interrupted by non-amino acids. The terms also encompass an amino acid polymer that has been modified naturally or by intervention; for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification, such as conjugation with a labeling component. Also included within the definition are, for example, polypeptides containing one or more analogs of an amino acid (including, for example, unnatural amino acids, etc.), as well as other modifications known in the art. It is understood that, because the polypeptides of this invention are based upon antibodies, in some aspects, the polypeptides can occur as single chains or associated chains.

[0092] As used herein, the term "identity" refers to the overall monomer conservation between polymeric molecules, e.g., between polypeptide molecules or polynucleotide molecules (e.g. DNA molecules and / or RNA molecules). The term "identical" without any additional qualifiers, e.g., protein A is identical to protein B, implies the sequences are 100% identical (100% sequence identity). Describing two sequences as, e.g., "70% identical," is equivalent to describing them as having, e.g., "70% sequence identity."

[0093] Calculation of the percent identity of two polypeptide sequences, for example, can be performed by aligning the two sequences for optimal comparison purposes (e.g., gaps can be introduced in one or both of a first and a second polypeptide sequences for optimal alignment and non-identical sequences can be disregarded for comparison purposes). In certain aspects, the length of a sequence aligned for comparison purposes is at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or about 100% of the length of the reference sequence. The amino acids at corresponding amino acid positions are then compared.

[0094] When a position in the first sequence is occupied by the same amino acid as the corresponding position in the second sequence, then the molecules are identical at that position. The percentage of sequence identity between the two sequences is a function of the number of identical positions shared by the sequences, taking into account the number of gaps, and the length of each gap, which needs to be introduced for optimal alignment of the two sequences. The comparison of sequences and determination of percent identity between two sequences can be accomplished using a mathematical algorithm.

[0095] Suitable software programs are available from various sources, and for alignment of both protein and nucleotide sequences. One suitable program to determine percent sequence identity is bl2seq, part of the BLAST suite of program available from the U.S. government's National Center for Biotechnology Information BLAST web site (blast.ncbi.nlm.nih.gov). B12seq performs a comparison between two sequences using either the BLASTN or BLASTP algorithm. BLASTN is used to compare nucleic acid sequences, while BLASTP is used to compare amino acid sequences. Other suitable programs are, e.g., Needle, Stretcher, Water, or Matcher, part of the EMBOSS suite of bioinformatics programs and also available from the European Bioinformatics Institute (EBI) at www.ebi.ac.uk / Tools / psa.

[0096] Sequence alignments can be conducted using methods known in the art such as MAFFT, Clustal (ClustalW, Clustal X or Clustal Omega), MUSCLE, etc.

[0097] Different regions within a single polynucleotide or polypeptide target sequence that aligns with a polynucleotide or polypeptide reference sequence can each have their own percent sequence identity. It is noted that the percent sequence identity value is rounded to the nearest tenth. For example, 80.11, 80.12, 80.13, and 80.14 are rounded down to 80.1, while 80.15, 80.16, 80.17, 80.18, and 80.19 are rounded up to 80.2. It also is noted that the length value will always be an integer.

[0098] In certain aspects, the percentage identity (%ID) or of a first amino acid sequence (or nucleic acid sequence) to a second amino acid sequence (or nucleic acid sequence) is calculated as %ID = 100 x (Y / Z), where Y is the number of amino acid residues (or nucleobases) scored as identical matches in the alignment of the first and second sequences (as aligned by visual inspection or a particular sequence alignment program) and Z is the total number of residues in the second sequence. If the length of a first sequence is longer than the second sequence, the percent identity of the first sequence to the second sequence will be higher than the percent identity of the second sequence to the first sequence.

[0099] One skilled in the art will appreciate that the generation of a sequence alignment for the calculation of a percent sequence identity is not limited to binary sequence-sequence comparisons exclusively driven by primary sequence data. It will also be appreciated that sequence alignments can be generated by integrating sequence data with data from heterogeneous sources such as structural data (e.g., crystallographic protein structures), functional data (e.g., location of mutations), or phylogenetic data. A suitable program that integrates heterogeneous data to generate a multiple sequence alignment is T-Coffee, available at www.tcoffee.org, and alternatively available, e.g., from the EBI. It will also be appreciated that the final alignment used to calculate percent sequence identity can be curated either automatically or manually.

[0100] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acids of any length. The polymer can comprise modified amino acids. The terms also encompass an amino acid polymer that has been modified naturally or by intervention; for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification, such as conjugation with a labeling component. Also included within the definition are, for example, polypeptides containing one or more analogs of an amino acid (including, for example, unnatural amino acids such as homocysteine, ornithine, p-acetylphenylalanine, D-amino acids, and creatine), as well as other modifications known in the art.

[0101] The term "polypeptide," as used herein, refers to proteins, polypeptides, and peptides of any size, structure, or function. Polypeptides include gene products, naturally occurring polypeptides, synthetic polypeptides, homologs, orthologs, paralogs, fragments and other equivalents, variants, and analogs of the foregoing. A polypeptide can be a single polypeptide or can be a multi-molecular complex such as a dimer, trimer or tetramer. They can also comprise single chain or multichain polypeptides. Most commonly disulfide linkages are found in multichain polypeptides. The term polypeptide can also apply to amino acid polymers in which one or moreamino acid residues are an artificial chemical analogue of a corresponding naturally occurring amino acid. In some aspects, a "peptide" can be less than or equal to 50 amino acids long, e.g., about 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 amino acids long.

[0102] As used herein, an "epitope" refers to a localized region of an antigen to which an antigen binding polypeptide or antigen binding polypeptide complex (e.g., antibody or antigen binding fragment thereof) can specifically bind. An epitope can be, for example, contiguous amino acids of a polypeptide (linear or contiguous epitope) or an epitope can, for example, come together from two or more non-contiguous regions of a polypeptide or polypeptides (conformational, nonlinear, discontinuous, or non-contiguous epitope). In some aspects, the epitope to which an antibody or antigen-binding fragment thereof binds can be determined by, e.g., NMR spectroscopy, X-ray diffraction crystallography studies, ELISA assays, hydrogen / deuterium exchange coupled with mass spectrometry (e.g., liquid chromatography electrospray mass spectrometry), array-based oligo-peptide scanning assays, and / or mutagenesis mapping (e.g., site-directed mutagenesis mapping). See, e.g., Giege R et al., (1994) Acta Crystallogr D Biol Crystallogr 50(Pt 4): 339-350; McPherson A (1990) Eur J Biochem 189: 1-23; Chayen NE (1997) Structure 5: 1269-1274; McPherson A (1976) J Biol Chem 251: 6300-6303; Meth Enzymol (1985) volumes 114 & 115, eds Wyckoff HW et al., U.S. Pub. No. 2004 / 0014194), Bricogne G (1993) Acta Crystallogr D Biol Crystallogr 49(Pt 1): 37-60, Bricogne G (1997) Meth Enzymol 276A: 361-423, ed Carter CW, and Roversi et al., (2000) Acta Crystallogr D Biol Crystallogr 56(Pt 10): 1316-1323 (X-ray diffraction crystallography studies); and Champe et al., (1995) J Biol Chem 270: 1388-1394 and Cunningham BC & Wells JA (1989) Science 244: 1081-1085 (mutagenesis mapping).

[0103] In some aspects, the term "epitope" refers to an antigenic determinant in a molecule such as an antigen, i.e., to a part in or fragment of the molecule that is recognized by the immune system, for example, that is recognized by a T cell, in particular when presented in the context of MHC molecules. An epitope of a protein such as a tumor antigen preferably comprises a continuous or discontinuous portion of said protein and is preferably between 5 and 100, preferably between 5 and 50, more preferably between 8 and 30, most preferably between 10 and 25 amino acids in length, for example, the epitope may be preferably 9, 10, 1 1, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids in length.

[0104] Specific binding can be represented by a "binding affinity." Binding affinity refers to an intrinsic binding affinity which reflects a 1 : 1 interaction between members of a binding pair (e.g., an antigen binding polypeptide complex and an antigen). Binding affinity can be measured and / or expressed in several ways known in the art, including, but not limited to, equilibriumdissociation constant (KD). KD is calculated from the quotient of koff / kon, where konrefers to the association rate constant of, e.g., an antigen binding polypeptide complex to an antigen, and kOff refers to the dissociation of, e.g., an antigen binding polypeptide complex from an antigen. The kon and koir can be determined by techniques known to one of ordinary skill in the art, such as Octet® BLI, BIAcore® or KinExA.

[0105] Accordingly, in some aspects, an antigen binding polypeptide complex provided herein is an antibody or antigen binding fragment thereof. In some aspects, an antigen binding polypeptide provided herein is part of an antibody or antigen-binding fragment thereof. In some aspects, the antibody or antigen binding fragment thereof specifically binds to an antigen with an equilibrium dissociation constant (KD) of from about 10 pM to about 1 pM.

[0106] The term "subsequence" or "portion" as applied to a protein refers to a fraction of a protein. With respect to a particular structure such as an amino acid sequence or protein the term "subsequence" or "portion" thereof may designate a continuous or a discontinuous fraction of said structure. In some aspects, a subsequence or portion of an amino acid sequence comprises at least about 1 %, at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95% of said amino acid sequence. In some aspects, a subsequence or portion of an amino acid sequence comprises about 1 %, about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, or about 95% of said amino acid sequence. In some aspects, a subsequence or portion of an amino acid sequence comprises at least about 10 continuous amino acids, at least about 20 continuous amino acids, at least about 30 continuous amino acids, at least about 40 continuous amino acids, at least about 50 continuous amino acids, at least about 60 continuous amino acids, at least about 70 continuous amino acids, at least about 80 continuous amino acids, at least about 90 continuous amino acids, at least about 100 continuous amino acids, at least about 120 continuous amino acids, at least about 140 continuous amino acids, at least about 160 continuous amino acids, at least about 180 continuous amino acids, at least about 200 continuous amino acids, at least about 220 continuous amino acids, at least about 240 continuous amino acids, at least about 260 continuous amino acids, at least about 280 continuous amino acids, at least about 300 continuous amino acids, at least about 325 continuous amino acids, at least about 350 continuous amino acids, at least about 375 continuous amino acids, at least about 400continuous amino acids, at least about 425 continuous amino acids, at least about 450 continuous amino acids, at least about 475 continuous amino acids, or at least about 500 continuous amino acids. In some aspects, a subsequence or portion of an amino acid sequence comprises about 10 continuous amino acids, about 20 continuous amino acids, about 30 continuous amino acids, about 40 continuous amino acids, about 50 continuous amino acids, about 60 continuous amino acids, about 70 continuous amino acids, about 80 continuous amino acids, about 90 continuous amino acids, about 100 continuous amino acids, about 120 continuous amino acids, about 140 continuous amino acids, about 160 continuous amino acids, about 180 continuous amino acids, about 200 continuous amino acids, about 220 continuous amino acids, about 240 continuous amino acids, about 260 continuous amino acids, about 280 continuous amino acids, about 300 continuous amino acids, about 325 continuous amino acids, about 350 continuous amino acids, about 375 continuous amino acids, about 400 continuous amino acids, about 425 continuous amino acids, about 450 continuous amino acids, about 475 continuous amino acids, or about 500 continuous amino acids.

[0107] The terms "subsequence" and "fragment" are used interchangeably herein and refer to a continuous element. For example, a part of the primary amino acid sequence of a protein and refers to a continuous element of said structure. For example, a portion, a part or a fragment of an epitope, peptide or protein is preferably immunologically equivalent to the epitope, peptide or protein it is derived from. In the context of the present invention, a "subsequence" of a structure such as an amino acid sequence preferably comprises or consists of about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, or about 99% of the primary amino acid sequence.

[0108] Terms such as "reducing" or "inhibiting" relate to the ability to cause an overall decrease, preferably of 5% or greater, 10% or greater, 20% or greater, more preferably of 50% or greater, and most preferably of 75% or greater, in the level. The term "inhibit" or similar phrases includes a complete or essentially complete inhibition, i.e. a reduction to zero or essentially to zero.

[0109] Terms such as "increasing", "enhancing", "promoting" or "prolonging" preferably relate to an increase, enhancement, promotion or prolongation by about at least 10%, preferably at least 20%, preferably at least 30%, preferably at least 40%, preferably at least 50%, preferably at least 80%, preferably at least 100%, preferably at least 200%, and in particular at least 300%. These terms may also relate to an increase, enhancement, promotion or prolongation from zero or a non-measurable or non-detectable level to a level of more than zero or a level, which is measurable or detectable.

[0110] The term “machine learning” as used herein generally refers to a type of Al that provides computers with the ability to learn without being explicitly programmed. Machine learning is a branch of Al focusing on systems that can learn from data, identify patterns, and make decisions with minimal human intervention.[OHl] As used herein, the terms “model” or “ML model” refer to the output of an ML algorithm that is trained with training data to learn from patterns and generalize to new data to make predictions.

[0112] As used herein, the term “tokenizing” and grammatical variants thereof refers to the process of breaking down a sequence or string into smaller entities or tokens so that they can be processed by machine learning models.

[0113] As used herein, the term “token” and grammatical variants thereof refers to the smallest meaningful units into which protein sequences are broken down for processing by machine learning models. In the context of this work, tokens represent individual amino acids.

[0114] As used herein, the term “embedding” and grammatical variants thereof refers to numerical vector representations that encode meaningful biological properties of proteins in a high dimensional space.

[0115] As used herein, the terms “SoftMax activation function” or “SoftMax function” refer to a mathematical function that converts a vector of raw prediction scores (often called logits) from the neural network into probabilities.

[0116] As used herein, the term “loss function” refers to a crucial component in machine learning that quantifies the difference between the predicted outputs of a machine learning algorithm and the actual target values. Optimizing a model entails adjusting model parameters to minimize the output of some loss function.

[0117] As used herein, the term “full length native protein” refers to a protein that is in its native or natural state and unaltered by any denaturing agent such as heat, chemical mutation or enzymatic reactions. A wild-type protein would be considered a full-length native protein. The term full length native protein sequence, as used herein, refers to the amino acid sequence found in the full-length native protein.

[0118] As used herein, “mutation” refers to a change in the amino acid sequence of a native protein. Mutations can be described by using the native sequence and then identifying the specific acid that have been changed. A “mutant” refers to the protein that contains the mutation. A full-length mutant sequence refers to the full amino acid sequence of the mutant protein, instead of describing the mutant as the amino acids that are different from the native protein.

[0119] The term “vaccine” as used herein refers to a pharmaceutical preparation (pharmaceutical composition) or product that upon administration induces an immune response, in particular a cellular immune response, which recognizes and attacks a pathogen or a diseased cell such as a cancer cell. A vaccine may be used for the prevention or treatment of a disease. The term "personalized cancer vaccine" or "individualized cancer vaccine" concerns a particular cancer patient and means that a cancer vaccine is adapted to the needs or special circumstances of an individual cancer patient. In one aspect, a vaccine provided according to the invention may comprise a peptide or polypeptide comprising one or more amino acid modifications or one or more modified peptides predicted as being immunogenic by the methods of the invention or a nucleic acid, preferably RNA, encoding said peptide or polypeptide.

[0120] As used herein, the terms “determining,” “assessing,” “measuring,” and their grammatical equivalents refer to both quantitative and qualitative determinations, and as such, the term “determining” is used interchangeably herein with“assaying,” “measuring,” and the like. Where a quantitative determination is intended, the phrase“determining an amount” of an analyte and the like is used. Where a qualitative and / or quantitative determination is intended, the phrase“determining a level” of an analyte or “detecting” an analyte is used.

[0121] As used herein, a “sequence” of a peptide or portion of a peptide refers to an aminoacid sequence that includes an ordered set of amino acid identifiers. The term sequence, in the context of the present disclosure, is equivalent to primary sequence when referring to proteins and peptides. A protein sequence is expressed as a series of consecutive amino acids in amino to carboxy orientation. Thus, the term sequence means the order in which amino acid residues, connected by peptide bonds, lie in the chains of peptides and proteins. A protein subsequence is a consecutive subset of amino acids from a protein sequence. A subsequence is typically defined as "a sequence that can be derived from another sequence by deleting some or no elements without changing the order of the remaining elements.I. Immunogenicity Prediction System

[0122] The present disclosure provides methods and systems for predicting the immunogenicity of a candidate protein, e.g., an antibody, comprising (i) generating one or more training datasets (e.g., an IFNy only dataset, an IFNy plus IL10 dataset, an IFNy, or any of the alternative training datasets disclosed more in detail below) derived from protein sequence data and associated in vitro cytokine release data from a population of training proteins; (ii) training one or more machine learning (ML) models using the one or more training datasets of (i), whereinthe one or more ML models are capable of predicting the ability to induce the cytokine release and / or immunogenicity of the candidate protein based on the primary sequence of the candidate protein; (iii) preprocessing the primary sequence of the candidate protein generating a plurality of peptide subsequences derived from the primary sequence of the candidate protein; (iv) applying the one or more ML models to each of the plurality of peptide sequences to generate immunogenicity scores associated with each peptide subsequence; and, (v) generating an immunogenicity prediction and / or a cytokine release prediction for the candidate protein based on the immunogenicity scores associated with each peptide subsequence.

[0123] The methods disclosed can be applied to a single candidate protein, to a subsequence or domain of a candidate protein, or to plurality or candidate proteins, e.g., a library of candidate proteins. In some aspects, the methods disclosed herein can be applied to a candidate protein that can be mutated either experimentally or in silico, and the mutated candidate sequence can be used as input to the system for example as a training sequence (inputting mutation data with associate immune response data, e.g., cytokine release induced by the protein following mutation) or as a candidate sequence to predict the immunogenicity of the mutated protein.

[0124] A general exemplary pipeline of the methods and systems of the present disclosure comprises (1) processing the training protein sequences in the training protein set by splitting them into overlapping frames (e.g., 15-mer peptide subsequences) and tokenizing them to generate embeddings (e.g., word embeddings) to convert them to the right format for an LLM or LSTM; (2) training the LLM or LSTM model using the converted input derived from the processed training protein set of (1); (3) processing a candidate protein by splitting it into overlapping frames (e.g., 15-mer peptide subsequences) and tokenizing it to generate embeddings (e.g., word embeddings) to convert them to the right format for the LLM or LSTM; (4) inputting the processed candidate protein of (3) into the trained LLM or LSTM of (2) to output predicted immunogenicity data; (5) processing the output from the LLM or LSTMs of (4) through a classifier layer to assign an immunogenic or non-immunogenic probability to each frame; (6) visualizing the immunogenicity prediction by plotting the immunogenicity probability of (5) for each frame (peptide subsequence) from the candidate with respect to a predetermined threshold value.

[0125] In one aspect of the immunogenicity predictor of the present disclosure, the input is a protein sequence, i.e., a linear arrangement of amino acids represented by a single letter notation. The protein sequence is first processed into overlapping peptides of 15 amino acids in length, where each peptide overlaps by 14 amino acids, i.e., it shifts by one amino acid. Each peptide is converted to a set of tokens, which is a vector representation of the text sequence. Thesetokens are used as an input to the protein language model (PLM), which generates embeddings that are numerical vector representations that encode meaningful biological properties of proteins learned by first training on unlabeled large-scale protein databases such as UniProt (www.uniprot.org), PRIDE Proteomics Identifications Database (www.ebi.ac.uk / pride / ), Peptide Atlas (peptideatlas.org), etc., and then fine-tuned on smaller labeled datasets. Therefore, these generated embeddings capture the structural, functional and evolutionary properties of the protein sequence in a high dimensional space. The generated embeddings are then used as input to a classification layer, which processes them to identify patterns. A Softmax activation function converts the raw predictions into probabilities, representing the likelihood of a peptide being immunogenic. A loss function then measures how different the predicted probabilities are from the true labels. The model adjusts itself using this loss to improve its predictions over time.

[0126] Immunogenicity generally refers to the ability of an antigen to elicit an immune response in the body of a subject. This response can be, for example, the development of anti-drug antibodies (AD As) that neutralize the effectiveness of a therapeutic drug or cause unwanted side effects, cytokine release, or a combination thereof. Immunogenicity is a critical consideration in the development and approval of biological drugs, as it directly impacts their efficacy and safety. The study of immunogenicity is a cornerstone in the development of drugs and vaccines, serving as a critical factor in determining their safety and efficacy. In the realm of biopharmaceuticals, particularly those involving therapeutic proteins, antibodies, and vaccines, understanding and managing immunogenicity is pivotal. The immune system’s response to these biologies can significantly impact their therapeutic effectiveness and patient safety.

[0127] Safety and Efficacy. Immunogenicity can lead to the production of AD As, which can neutralize the therapeutic effects of the compound or cause adverse immune reactions. In the case of vaccines, a robust immunogenic response is desired to provide protection against the target disease. Therefore, assessing immunogenicity helps in predicting and enhancing their clinical performance.

[0128] Regulatory Compliance'. Regulatory agencies like the FDA and EMA have set guidelines for immunogenicity assessment, making it a mandatory aspect of the drug and vaccine approval process. This assessment ensures that the benefits of the therapeutic outweigh its potential immunogenic risks to patients.

[0129] Personalized Medicine'. Understanding the immunogenic potential of drugs aids in the development of personalized medicine strategies. It allows for the prediction of patient-specific responses to biologies, leading to more tailored and efficacious treatments.

[0130] Advancements in Drug Design'. Studying immunogenicity guides scientists in modifying the molecular structure of biologies to diminish their immunogenic potential while preserving their therapeutic efficacy. This is particularly relevant in the development of biosimilars and next-generation biologies.

[0131] Definition of immunogenicity'. In the context of the predictions performed using the methods disclosed herein the term “immunogenicity” refers specifically to the ability of a protein (e.g., a candidate protein) to induce the release of cytokines or to the specific immune response associated with the set of immunogenic proteins used for the training of the machine learning (ML) models of the present invention. The specific variants or flavors of “immunogenicity” predicted by the ML models of the present invention will depend on the "immune response" or "immunological response" associated to the sequence data in the training set, e.g., in vitro cytokine release data. Thus, in the context of the present disclosure, a prediction of immunogenicity can refer to a prediction of release of cytokines (e.g., interferon gamma). In one particular aspect, e.g., when the IFNgamma Immunogenicity Model is applied, the model provides an “immunogenicity prediction” which is, e.g., the probability that the protein used as input to the Model or a subsequence thereof will elicit an immune response. Such probability can be referred at as immunogenic or immunogenicity potential, immunogenic or immunogenicity propensity, or immunogenic or immunogenicity probability. The immunogenicity propensity can lead to an immune response leading to neutralizing the effect of a drug, causing an adverse reaction or leading to a desired immune response as in the case of vaccines.

[0132] As used herein, the terms "immune response" and "immunological response" and grammatical equivalents herein refer to a response of the immune system to a molecule, including humoral or cellular immune responses. Non-limiting immunological responses include production of neutralizing and non-neutralizing antibodies, formation of immune complexes, complement activation, mast cell activation, inflammation, and anaphylaxis. In some specific aspects, the term immune response refers to cytokine release.

[0133] As used herein, the terms "predict," "predicting," "predictive," and grammatical variants thereof refer to the ability of a system or assay to act as a surrogate for in vivo immunogenicity and recapitulate or mimic the immunogenic outcome. That is, a system or assay is predictive of immunogenicity if the system can demonstrate with reasonable accuracy that the antigen would have or would not have elicited an immunogenic response had it been administered.

[0134] In some aspects, the present disclosure provides a computed-implemented method for predicting the immunogenicity of a candidate protein, comprising (i) training one or moremachine learning (ML) models using the one or more training datasets wherein the one or more training datasets are derived from protein sequence data and associated in vitro cytokine release data from a population of training proteins, and wherein the one or more ML models are capable of predicting the immunogenicity of the candidate protein based on the primary sequence of the candidate protein; (ii) preprocessing the primary sequence of the candidate protein generating a plurality of peptide subsequences derived from the primary sequence of the candidate protein; (iii) applying the one or more ML models to each of the plurality of peptide sequences to generate immunogenicity scores associated with each peptide subsequence; and, (iv) generating an immunogenicity prediction for the candidate protein based on the immunogenicity scores associated with each peptide subsequence.

[0135] Also provided is a computed-implemented method for predicting the immunogenicity of a candidate protein, comprising (i) preprocessing the primary sequence of the candidate protein generating a plurality of peptide subsequences derived from the primary sequence of the candidate protein; (ii) applying the one or more pretrained ML models to each of the plurality of peptide sequences to generate immunogenicity scores associated with each peptide subsequence, wherein the one or more machine learning (ML) models are trained using the one or more training datasets wherein the one or more training datasets are derived from protein sequence data and associated cytokine release data from a population of training proteins, and wherein the one or more ML models are capable of predicting the immunogenicity of the candidate protein based on the primary sequence of the candidate protein; and, (iii) generating an immunogenicity prediction for the candidate protein based on the immunogenicity scores associated with each peptide subsequence.

[0136] In some aspects, the methods disclosed herein possess characteristics that are advantageous with respect to other approaches to predict immunogenicity. Accordingly, in some aspects, the one or more training datasets used in the computer-implemented methods of the present disclosure do not use protein-binding data. In some aspects, the one or more training datasets used in the computer-implemented methods of the present disclosure do not use protein-binding affinity data. In some aspects, the one or more training datasets used in the computer-implemented methods of the present disclosure do not use protein-binding affinity prediction data. In some aspects, the one or more training datasets used in the computer-implemented methods of the present disclosure do not use (i) use protein-binding data, (ii) protein-binding affinity data, (iii) protein-binding affinity prediction data, or (iv) any combination thereof.

[0137] In some aspects, the computer-implemented methods disclosed herein do not comprise the use of protein secondary or tertiary structure data to predict immunogenicity. In some aspects, the computer-implemented methods disclosed herein do not use protein secondary structure data, e.g., secondary structure prediction data such as computational helicity predictions, or other computational structure predictions, or experimental secondary structure data, for example, circular dichroism measurements. In some aspects, the computer-implemented methods disclosed herein do not a comprising using protein tertiary structure prediction. In some aspects, the computer-implemented methods disclosed herein do not comprise using molecular dynamics predictions, or using three-dimensional data, for example, crystallographic data. In some aspects, the computer-implemented methods disclosed herein do not comprise a humanness assessment. For example, a crucial difference with respect to the prediction method disclosed in WO2012176756 is that the present method does not evaluate the strength of intermolecular interactions of a complex comprising a peptide, a MHCII protein, and a T cell receptor to provide a score that predicts the immunogenicity of the peptide.II. Training datasets

[0138] In some aspects, generating one or more training datasets comprises (a) preprocessing at least one population of training proteins; (b) optionally performing data augmentation; and (c) labeling. In some aspects, preprocessing comprises filling or removing missing data values, removing data redundancies, and handling data ambiguities associated with at least one population of training proteins. In some aspects, preprocessing the population of training proteins comprises (i) splitting the amino acid sequences of the training proteins into training peptide subsequences, and (ii) tokenizing the training peptide subsequences to generate embeddings.

[0139] In some aspects, data augmentation comprises splitting the amino acid sequence of a training protein into one or more shorter training proteins. In some aspects, labeling comprises assigning a plurality of labels indicative of an immunogenicity or non-immunogenicity category of each training peptide subsequence derived from the at least one population of training proteins. In some aspects, labeling comprises converting a multi-class label associated with a training peptide subsequence located within the sequence of a training protein of the one or more training datasets to a binary-class label. In some aspects, the converting comprises assigning the binaryclass label as positive or negative. For example, the converting can comprise assigning the binaryclass label as positive based on determining that a value, e.g., an ICso value, associated with thetraining peptide subsequence is larger than a limit threshold. Conversely, the converting can comprise assigning the binary-class label as negative based on determining that a value, e.g., an IC50 value, associated with the training peptide subsequence is smaller than the limit threshold. A person of ordinary skill in the art would appreciate that a variety of approaches can be used to convert a multi-class label to a binary-class label.

[0140] In some aspects of the methods disclosed herein, the training of the ML models can be conducted using full-length protein sequences. However, in other aspects, the training protein sequences can be split into subsequences, which can be of constant length or can be of variable length. In some aspects, the training peptide subsequences are overlapping. In other aspects, the training peptide subsequences are non-overlapping.

[0141] In the context of the present disclosure, a “population of training proteins” is defined as a combination of (i) a set of non-immunogenic protein (non-immunogenic subset) and (ii) a set of immunogenic protein capable of eliciting the release of a specific cytokine or combination thereof (immunogenic subset), wherein the release data associated with immunogenic proteins consists of experimental immunogenicity data (i.e., cytokine release data), and wherein the training set does not contain binding affinity prediction data release. In some aspects, the lengths of the proteins in the population of training proteins is between 13 amino acids and 26 amino acids.

[0142] In some aspects, the cytokine release data is selected from the group consisting of IFNy release data, proliferation data, IL-5 release, cytotoxicity data, activation data, qualitative binding data, IL-2 release data, IL- 10 release data, IL-4 release data, and IL- 17 release data. Therefore, models can be created to predict the release of each of these cytokines or a plurality of these cytokines such as IFNY plus IL- 10 plus IL-4 etc.

[0143] Proliferation quantitates or estimates the immunogenicity of a protein in the training set by measuring the increase in the number of T cells. T cells increase rapidly as a result of an immune response triggered by an immunogen, e.g., a protein such a therapeutic protein (e.g., an antibody or a vaccine) that is perceived as foreign or non-self by a human subject. Based on the level of observed proliferation, a protein in the training set can be classified as high proliferation or high stimulation index protein, or a low proliferation or low stimulation index protein. A high level of T-cell proliferation, indicated by a high Stimulation Index (SI) or a high percentage of proliferating cells indicates that the protein is eliciting a strong immune response, potentially leading to immunogenicity. Conversely, if the T-cell proliferation is low or the SI is below a certainthreshold (e.g., SI < 2.0), it indicates that the protein is less likely to cause a strong immune response and therefore has a lower risk of immunogenicity.

[0144] Cytotoxicity quantitates or estimates the immunogenicity of a protein in the training set by measuring its ability to trigger immune cells, such as CD8+ T cells or NK cells, to kill target cells presenting the protein. In these assays, target cells are loaded with protein and exposed to immune effector cells. If the protein is immunogenic, effector cells recognize the protein-MHC complex and initiate cell lysis through mechanisms like perforin-granzyme release or Fas-FasL signaling. The extent of cytotoxicity is then quantified using different methods, such as51Cr release, LDH leakage, flow cytometry-based viability stains, or caspase activity assays. A higher level of cell death indicates a strong cytotoxic immune response, suggesting that the protein is highly immunogenic. Conversely, low or no cytotoxicity suggests weak or non-immunogenic behavior.

[0145] Activation quantitates or estimates the immunogenicity of a protein in the training set by measuring cellular activation in response to the protein. These assays assess how immune cells, particularly T cells, B cells, or antigen-presenting cells (APCs), respond when exposed to the protein. Activation is determined by detecting changes in cell behavior, such as cytokine secretion (e.g., IFN-y, IL-2, TNF-a), upregulation of activation markers (CD69, CD25), proliferation, or metabolic activity. Techniques like ELISA, ELISPOT, flow cytometry, or luciferase-based reporter assays quantify these responses. A strong activation signal suggests the protein effectively stimulates the immune system and is likely immunogenic, whereas a weak or absent response indicates low immunogenicity.

[0146] Qualitative Binding quantitates or estimates the immunogenicity of a protein in the training set by assessing its ability to bind to immune system components, such as MHC molecules, antibodies, or immune receptors. These assays evaluate the strength, specificity, and stability of interactions using techniques like ELISA, surface plasmon resonance (SPR), isothermal titration calorimetry (ITC), or fluorescence polarization. For MHC binding assays, the proteins are tested for their ability to bind MHC class I or II molecules, which is a prerequisite for T cell activation. Antibody binding assays, such as ELISA or Western blot, detect whether antibodies in a sample recognize the protein, indicating prior immune exposure or potential immunogenicity.

[0147] In some aspects, the population of training proteins comprises a plurality of protein subsequences from the set of immunogenic proteins and a plurality of data labels indicative of a categorized information of each protein sequence in the set of immunogenic proteins. In some aspects, the population of training proteins comprises a set of non-immunogenic proteins. In someaspects, the population of training proteins comprises a plurality of protein subsequences from the set of non-immunogenic proteins and a plurality of data labels indicative of a categorized information of each protein sequence in the set of non-immunogenic proteins.

[0148] In some aspects, the total population of training proteins comprises at least about 1,000 proteins, at least about 2,000 proteins, at least about 3,000 proteins, at least about 4,000 proteins, at least about 5,000 proteins, at least about 6,000 proteins, at least about 7,000 proteins, at least about 8,000 proteins, at least about 9,000 proteins, at least about 10,000 proteins, at least about 11,000 proteins, at least about 12,000 proteins, at least about 13,000 proteins, at least about 14,000 proteins, at least about 15,000 proteins, at least about 16,000 proteins, at least about 17,000 proteins, at least about 18,000 proteins, at least about 19,000 proteins, at least about 20,000 proteins, at least about 21,000 proteins, at least about 22,000 proteins, at least about 23,000 proteins, at least about 24,000 proteins, at least about 25,000 proteins, at least about 26,000 proteins, at least about 27,000 proteins, at least about 28,000 proteins, at least about 29,000 proteins, at least about 30,000 proteins, at least about 31,000 proteins, at least about 32,000 proteins, at least about 33,000 proteins, at least about 34,000 proteins, at least about 35,000 proteins, at least about 36,000 proteins, at least about 37,000 proteins, at least about 38,000 proteins, at least about 39,000 proteins, at least about 40,000 proteins, at least about 41,000 proteins, at least about 42,000 proteins, at least about 43,000 proteins, at least about 44,000 proteins, at least about 45,000 proteins, at least about 46,000 proteins, at least about 47,000 proteins, at least about 48,000 proteins, at least about 49,000 proteins, at least about 50,000 proteins, at least about 51,000 proteins, at least about 55,000 proteins, at least about 60,000 proteins, at least about 65,000 proteins, at least about 70,000 proteins, at least about 75,000 proteins, at least about 80,000 proteins, at least about 85,000 proteins, at least about 90,000 proteins, at least about 95,000 proteins or at least about 10,000 proteins, or subsequences thereof.

[0149] In some aspects, the total population of training proteins comprises about 1,000 proteins, about 2,000 proteins, about 3,000 proteins, about 4,000 proteins, about 5,000 proteins, about 6,000 proteins, about 7,000 proteins, about 8,000 proteins, about 9,000 proteins, about 10,000 proteins, about 11,000 proteins, about 12,000 proteins, about 13,000 proteins, about 14,000 proteins, about 15,000 proteins, about 16,000 proteins, about 17,000 proteins, about 18,000 proteins, about 19,000 proteins, about 20,000 proteins, about 21,000 proteins, about 22,000 proteins, about 23,000 proteins, about 24,000 proteins, about 25,000 proteins, about 26,000 proteins, about 27,000 proteins, about 28,000 proteins, about 29,000 proteins, about 30,000 proteins, about 31,000 proteins, about 32,000 proteins, about 33,000 proteins, about 34,000proteins, about 35,000 proteins, about 36,000 proteins, about 37,000 proteins, about 38,000 proteins, about 39,000 proteins, about 40,000 proteins, about 41,000 proteins, about 42,000 proteins, about 43,000 proteins, about 44,000 proteins, about 45,000 proteins, about 46,000 proteins, about 47,000 proteins, about 48,000 proteins, about 49,000 proteins, about 50,000 proteins, about 51,000 proteins, about 55,000 proteins, about 60,000 proteins, about 65,000 proteins, about 70,000 proteins, about 75,000 proteins, about 80,000 proteins, about 85,000 proteins, about 90,000 proteins, about 95,000 or about 100,000 proteins, or subsequences thereof.

[0150] In some aspects, the total population of training proteins comprises between about 1,000 and about 5,000 proteins, between about 5,000 and about 10,000 proteins, between about 10,000 and about 20,000 proteins, between about 20,000 and about 30,000 proteins, between about 30,000 and about 40,000 proteins, between about 40,000 and about 50,000 proteins, between about 15,000 and about 25,000 proteins, between about 25,000 and about 35,000 proteins, between about 35,000 and about 45,000 proteins, between about 45,000 and about 55,000 proteins, between about 20,000 and about 60,000 proteins, between about 10,000 and about 50,000 proteins, or between about 30,000 and about 70,000 proteins, or subsequences thereof.

[0151] In some aspects, the set of immunogenic proteins comprises at least about 1,000 proteins, at least about 2,000 proteins, at least about 3,000 proteins, at least about 4,000 proteins, at least about 5,000 proteins, at least about 6,000 proteins, at least about 7,000 proteins, at least about 8,000 proteins, at least about 9,000 proteins, at least about 10,000 proteins, at least about 11,000 proteins, at least about 12,000 proteins, at least about 13,000 proteins, at least about 14,000 proteins, at least about 15,000 proteins, at least about 16,000 proteins, at least about 17,000 proteins, at least about 18,000 proteins, at least about 19,000 proteins, at least about 20,000 proteins, at least about 21,000 proteins, at least about 22,000 proteins, at least about 23,000 proteins, at least about 24,000 proteins, at least about 25,000 proteins, at least about 26,000 proteins, at least about 27,000 proteins, at least about 28,000 proteins, at least about 29,000 proteins, at least about 30,000 proteins, at least about 31,000 proteins, at least about 32,000 proteins, at least about 33,000 proteins, at least about 34,000 proteins, at least about 35,000 proteins, at least about 36,000 proteins, at least about 37,000 proteins, at least about 38,000 proteins, at least about 39,000 proteins, at least about 40,000 proteins, at least about 41,000 proteins, at least about 42,000 proteins, at least about 43,000 proteins, at least about 44,000 proteins, at least about 45,000 proteins, at least about 46,000 proteins, at least about 47,000 proteins, at least about 48,000 proteins, at least about 49,000 proteins, at least about 50,000 proteins, at least about 51,000 proteins, at least about 55,000 proteins, at least about 60,000proteins, at least about 65,000 proteins, at least about 70,000 proteins, at least about 75,000 proteins, at least about 80,000 proteins, at least about 85,000 proteins, at least about 90,000 proteins, at least about 95,000 proteins or at least about 100,000 proteins, or subsequences thereof.

[0152] In some aspects, the set of immunogenic proteins comprises about 1,000 proteins, about 2,000 proteins, about 3,000 proteins, about 4,000 proteins, about 5,000 proteins, about 6,000 proteins, about 7,000 proteins, about 8,000 proteins, about 9,000 proteins, about 10,000 proteins, about 11,000 proteins, about 12,000 proteins, about 13,000 proteins, about 14,000 proteins, about 15,000 proteins, about 16,000 proteins, about 17,000 proteins, about 18,000 proteins, about 19,000 proteins, about 20,000 proteins, about 21,000 proteins, about 22,000 proteins, about 23,000 proteins, about 24,000 proteins, about 25,000 proteins, about 26,000 proteins, about 27,000 proteins, about 28,000 proteins, about 29,000 proteins, about 30,000 proteins, about 31,000 proteins, about 32,000 proteins, about 33,000 proteins, about 34,000 proteins, about 35,000 proteins, about 36,000 proteins, about 37,000 proteins, about 38,000 proteins, about 39,000 proteins, about 40,000 proteins, about 41,000 proteins, about 42,000 proteins, about 43,000 proteins, about 44,000 proteins, about 45,000 proteins, about 46,000 proteins, about 47,000 proteins, about 48,000 proteins, about 49,000 proteins, about 50,000 proteins, about 51,000 proteins, about 55,000 proteins, about 60,000 proteins, about 65,000 proteins, about 70,000 proteins, about 75,000 proteins, about 80,000 proteins, about 85,000 proteins, about 90,000 proteins, about 95,000 or about 100,000 proteins, or subsequences thereof.

[0153] In some aspects, the set of immunogenic proteins comprises between about 1,000 and about 5,000 proteins, between about 5,000 and about 10,000 proteins, between about 10,000 and about 20,000 proteins, between about 20,000 and about 30,000 proteins, between about 30,000 and about 40,000 proteins, between about 40,000 and about 50,000 proteins, between about 15,000 and about 25,000 proteins, between about 25,000 and about 35,000 proteins, between about 35,000 and about 45,000 proteins, between about 45,000 and about 55,000 proteins, between about 20,000 and about 60,000 proteins, between about 10,000 and about 50,000 proteins, or between about 30,000 and about 70,000 proteins, or subsequences thereof.

[0154] In some aspects, the set of non-immunogenic proteins comprises at least about 1,000 proteins, at least about 2,000 proteins, at least about 3,000 proteins, at least about 4,000 proteins, at least about 5,000 proteins, at least about 6,000 proteins, at least about 7,000 proteins, at least about 8,000 proteins, at least about 9,000 proteins, at least about 10,000 proteins, at least about 11,000 proteins, at least about 12,000 proteins, at least about 13,000 proteins, at least about 14,000 proteins, at least about 15,000 proteins, at least about 16,000 proteins, at least about 17,000proteins, at least about 18,000 proteins, at least about 19,000 proteins, at least about 20,000 proteins, at least about 21,000 proteins, at least about 22,000 proteins, at least about 23,000 proteins, at least about 24,000 proteins, at least about 25,000 proteins, at least about 26,000 proteins, at least about 27,000 proteins, at least about 28,000 proteins, at least about 29,000 proteins, at least about 30,000 proteins, at least about 31,000 proteins, at least about 32,000 proteins, at least about 33,000 proteins, at least about 34,000 proteins, at least about 35,000 proteins, at least about 36,000 proteins, at least about 37,000 proteins, at least about 38,000 proteins, at least about 39,000 proteins, at least about 40,000 proteins, at least about 41,000 proteins, at least about 42,000 proteins, at least about 43,000 proteins, at least about 44,000 proteins, at least about 45,000 proteins, at least about 46,000 proteins, at least about 47,000 proteins, at least about 48,000 proteins, at least about 49,000 proteins, at least about 50,000 proteins, at least about 51,000 proteins, at least about 55,000 proteins, at least about 60,000 proteins, at least about 65,000 proteins, at least about 70,000 proteins, at least about 75,000 proteins, at least about 80,000 proteins, at least about 85,000 proteins, at least about 90,000 proteins, at least about 95,000 proteins or at least about 100,000 proteins, or subsequences thereof.

[0155] In some aspects, the set of non-immunogenic proteins comprises about 1,000 proteins, about 2,000 proteins, about 3,000 proteins, about 4,000 proteins, about 5,000 proteins, about 6,000 proteins, about 7,000 proteins, about 8,000 proteins, about 9,000 proteins, about 10,000 proteins, about 11,000 proteins, about 12,000 proteins, about 13,000 proteins, about 14,000 proteins, about 15,000 proteins, about 16,000 proteins, about 17,000 proteins, about 18,000 proteins, about 19,000 proteins, about 20,000 proteins, about 21,000 proteins, about 22,000 proteins, about 23,000 proteins, about 24,000 proteins, about 25,000 proteins, about 26,000 proteins, about 27,000 proteins, about 28,000 proteins, about 29,000 proteins, about 30,000 proteins, about 31,000 proteins, about 32,000 proteins, about 33,000 proteins, about 34,000 proteins, about 35,000 proteins, about 36,000 proteins, about 37,000 proteins, about 38,000 proteins, about 39,000 proteins, about 40,000 proteins, about 41,000 proteins, about 42,000 proteins, about 43,000 proteins, about 44,000 proteins, about 45,000 proteins, about 46,000 proteins, about 47,000 proteins, about 48,000 proteins, about 49,000 proteins, about 50,000 proteins, about 51,000 proteins, about 55,000 proteins, about 60,000 proteins, about 65,000 proteins, about 70,000 proteins, about 75,000 proteins, about 80,000 proteins, about 85,000 proteins, about 90,000 proteins, about 95,000 or about 100,000 proteins, or subsequences thereof.

[0156] In some aspects, the set of non-immunogenic proteins comprises between about 1,000 and about 5,000 proteins, between about 5,000 and about 10,000 proteins, between about10,000 and about 20,000 proteins, between about 20,000 and about 30,000 proteins, between about 30,000 and about 40,000 proteins, between about 40,000 and about 50,000 proteins, between about 15,000 and about 25,000 proteins, between about 25,000 and about 35,000 proteins, between about 35,000 and about 45,000 proteins, between about 45,000 and about 55,000 proteins, between about 20,000 and about 60,000 proteins, between about 10,000 and about 50,000 proteins, or between about 30,000 and about 70,000 proteins, or subsequences thereof.

[0157] In some aspects, the total population of training proteins comprises between about 20,000 and about 40,000 proteins, or subsequences thereof. In some aspects, the set of non-immunogenic proteins comprises between about 10,000 proteins and about 20,000 proteins, or subsequences thereof. In some aspects, the set of immunogenic proteins comprises between about 10,000 and about 20,000 protein proteins, or subsequences thereof. In some aspects, the total population of training proteins comprises 30,000 ± 10,000 proteins, or subsequences thereof. In some aspects, the set of immunogenic proteins comprises 20,000 ± 5,000 proteins, or subsequences thereof. In some aspects, the set of non-immunogenic proteins comprises 20,000 ± 5000 proteins, or subsequences thereof.

[0158] In one specific aspect, the total population of training proteins comprises approximately 36,000 proteins, or subsequences thereof. In another specific aspect, the set of immunogenic proteins comprises approximately 16,000 proteins, or subsequences thereof. In another specific aspect, the set of non-immunogenic proteins comprises approximately 20,000 proteins, or subsequences thereof. In one specific aspect, the total population of training proteins comprises approximately 23,000 proteins, or subsequences thereof. In one specific aspect, the set of immunogenic proteins comprises approximately 13,000 proteins, or subsequences thereof. In another specific aspect, the set of non-immunogenic protein comprises approximately 10,000 proteins, or subsequences thereof. In some aspects, the proteins are a curated subset of peptidic epitopes from the Immune Epitope Database (IEDB) available at www.iedb.org.

[0159] In some aspects, the plurality of peptide subsequences derived from the amino acid sequences of the population of training proteins are generated using a sliding window and a sliding step. In some aspects, the sliding window has a variable length. In some aspects, the sliding window has a constant length. In some aspects, the sliding window constant length is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids. In some aspects, the sliding window constant length is at least 1, at least 2, at least 3, at least 4, at least 5, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 12, at least 13, atleast 14 or at least 15 amino acids. In some aspects, the sliding window constant length is between I and 5, between 5 and 10, between 10 and 15, between 10 and 20, between 12 and 18, between I I and 16, between 12 and 17, or between 13 and 18. In some aspects, the sliding window length is 15 amino acids. In some aspects, the sliding window length is selected as a function of epitope class. For example, typical T cell epitopes presented by MHC class II molecules are 13 to 17 amino acids in length, and those presented by MHC class I molecules are 8 to 11 amino acids in length. In general, antibody epitopes are about 5 to 8 amino acids in length. Thus, in some aspects, the sliding window has a selectable variable length between 5 and 25 amino acids. In some aspects, the ‘sliding window’ is the whole length of the protein. In some aspects, the sliding step has a constant length. In some aspects, the sliding step length is 1 amino acid. In some aspects, wherein the sliding step length is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids. In some aspects, the sliding step has a variable length. In some aspects, the ‘sliding step’ is the full-length of the protein. In some aspects, the plurality of peptide subsequences derived from the amino acid sequences of the population of training proteins are overlapping. In some aspects, the plurality of peptide subsequences derived from the amino acid sequences of the population of training proteins are non-overlapping. In some aspects, the overlap between sequences is by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 amino acids. In some aspects, the plurality of peptide subsequences derived from the amino acid sequences of the population of training proteins have the same length. In some aspects, the plurality of peptide subsequences derived from the amino acid sequences of the population of training proteins can have different lengths. In some specific aspects, the sliding window has a length of 15 amino acids, and the sliding step has a length of 1 amino acid.

[0160] In some aspects of the methods disclosed herein, the training comprises fine-tuning a pre-trained large language model (LLM) and / or fine-tuning a pre-trained long short-term memory (LSTM) network.

[0161] In some aspects, the protein sequence data comprises only primary sequence data and associated cytokine release data. For example, whether the presence of a certain protein or epitope is associated with the release of a specific cytokine or set thereof. In some aspects, the associated cytokine release data comprises whether binary data indicating whether the cytokine is released or not. In other aspects, the data comprises, e.g., the duration of cytokine release, amount, severity of the immune reaction, nature of the immune reaction, etc. In some aspects, the cytokine is a proinflammatory cytokine, e.g., IL-1, IL-6, TNF alpha, or a chemokine such as IL-8. In some aspects, the cytokine is an anti-inflammatory cytokine, e.g., IL- 10 or TGFbeta. In some aspects,the cytokine is a B cell activating cytokine such as CD40L, IL-6, IL-3, or IL-4. In some aspects, the cytokine is a T cell activating cytokine such as IL-2, IL-4, IL- 10, IL-13, or IL-15. In some aspects, the cytokine is an anti-infectious cytokine such as IFNalpha, IFNbeta, IFNgamma, or TNFalpha. In some aspects, the cytokine is an anti-proliferative cytokine such as IFNalpha, IFNbeta, TNFalpha or TGFbeta. In some aspects, the cytokine is selected from the group consisting of IL-1, IL-2, IL-3, IL-4, IL-6, IL-8, IL-10, IL-13, IL-15, TNFalpha, TGFbeta, CD40L, IFNalpha, IFNbeta, IFNgamma, and any combination thereof.

[0162] In some aspects, the protein sequence data further comprises protein secondary structure data and / or protein tertiary structure data. Thus, in some aspects, the protein sequence data comprises alpha helical content, beta sheet content, beta turn content, omega loop content, random coil content or any combination thereof. In some aspects, the protein sequence secondary structure data is tokenized using, for example, the DSSP or SST dictionaries. The Dictionary of Protein Secondary Structure, in short DSSP, is commonly used to describe the protein secondary structure with single letter codes. DSSP defines eight types of secondary structure: G (3-turn helix; 310 helix; min length 3 residues), H (4-turn helix; a helix; minimum length 4 residues), I (5-tum helix; n helix; minimum length 5 residues), T (hydrogen bonded turn; 3, 4 or 5 turn), E (extended strand in parallel and / or anti-parallel P-sheet conformation; min length 2 residues), B (residue in isolated P-bridge; single pair P-sheet hydrogen bond formation), S (bend; the only non-hydrogen-bond based assignment), and C (coil; residues which are not in any of the above conformations). SST is a Bayesian method to assign secondary structure using the Shannon information criterion of Minimum Message Length (MML) inference. The core idea is that the best secondary structural assignment is the one that can explain (compress) the coordinates of a given protein coordinates in the most economical way, thus linking the inference of secondary structure to lossless data compression. SST accurately delineates any protein chain into regions associated with the following assignment types: E (extended; strand of a P-pleated sheet), G (right-handed 310 helix), H (right-handed a-helix), I (right-handed 7t-helix), g (left-handed 310 helix), h (left-handed a-helix), i (left-handed 7t-helix), 3 (310-like turn), 4 (a-like turn), 5 (jt-like turn), T (unspecified turn), C (coil), and - (unassigned residue).

[0163] In some aspects, the preprocessing of at least one population of training proteins protein comprises iterating the amino acid sequence of each training protein sequence using a sliding window and sliding step to generate a plurality of peptide subsequences. In some aspects, preprocessing further comprises tokenizing the plurality of peptide subsequences derived from the primary sequence of a training protein. In some aspects, tokenizing a peptide subsequence derivedfrom the primary sequence of a training protein comprises generating a numerical sequence associated with the peptide subsequence with a format applicable to the one or more ML models.

[0164] In some aspects, tokenizing comprises using a pre-defined tokenizer to encode the amino acids within the peptide subsequence. In some aspects, tokenizing coverts each amino acid to a numeric and / or vector token. For example, in some aspects, tokenization of an amino acid sequence comprises parsing the amino acid sequence into components (e.g., individual amino acids, n-mers, sub-words, etc.) and mapping each component to a value. For example, an amino acid sequence may be separated into individual amino acids, and each individual amino acid may be mapped to a numeric value. The output of the tokenizer module may be another data structure or structured data, e.g., comprising an array or matrix structure) with each entry corresponding to tokenized representation of an amino acid sequence. In some aspects, tokenization may be performed at the individual amino acid-level (individual amino acid-based tokenization), also referred to as single or individual amino acid-level tokenization, in which each individual amino acid is mapped to a different numeric value. In another aspect, n-mer tokenization may be performed. In this approach, short n-mers of adjacent amino acids (e.g., where n is a numeric value such as 2, 3, 4 or more to form respective strings of two amino acids, three amino acids, four amino acids, etc.), are each mapped to numeric values. In another aspect, sub-word tokenization may be performed. In this approach, sub-words or strings of amino acids of varying length are each mapped to particular numeric values. In some aspects, sub-words are determined based upon analysis of protein sequences or may be based upon knowledge from the literature and / or subject matter experts.

[0165] Most machine learning algorithms work best when the number of samples in each class is about equal. This is because most algorithms are designed to maximize accuracy and reduce errors. Class imbalance is a common problem in machine learning, especially in classification problems. To address the class imbalance, techniques like oversampling, under-sampling, and weighted loss functions are often used. Loss functions are mathematical equations that explain the deviation between the actual and prediction values. The loss function evaluates the performance of an algorithm on the dataset. The higher the loss values, the more significant the error rate. The aim is to minimize the loss function. The loss function helps in learning trainable parameters, weights and biases. In some aspects of the present disclosure, a loss function is application is application. In some aspects, the loss function is scaled or adjusted to account for class imbalance. In other aspects, the loss function is not scaled or adjusted to account for class imbalance.III. Machine Learning (ML) model and classification layer

[0166] The immunogenicity prediction methods disclosed herein comprise a machine learning (ML) model. In the context of the present disclosure, the terms machine learning and artificial intelligence can be used interchangeably. In machine learning, many machine learning algorithms have been developed regarding how to classify data. Representative examples include decision trees, Bayesian networks, support vector machines (SVMs) and artificial neural networks (ANNs).

[0167] As used herein, the terms “model,” “ML model,” and “machine learning model” and grammatical variants thereof are used interchangeably and refer to a data structure and / or set of rules that represents learned knowledge based on the training performed by the machine-learning algorithm. These models are computer models trained to perform one or more tasks by learning to approximate functions or parameters based on training input comprising massive amounts of data. Thus, these models are beyond mere calculations that can be performed manually by a human subject such a calculation of an average. Accordingly, a model according to the present disclose is not an arithmetic average of several inputs or a Z-score.

[0168] In some aspects, the ML model is a newly trained model. In other aspects, the ML model is a pretrained model. Pre-trained models are neural network architectures that have undergone a two-step process: pre-training and fine-tuning. Pre-trained models have garnered immense attention and have become a driving force in many machine-learning applications. Exemplary pre-trained Natural Language Processing (NLP) models are, e.g., BERT (Bidirectional Encoder Representations from Transformers), GPT-3 (Generative Pre-trained Transformer 3), or XLNet. In the first phase (pre-training), the model is exposed to vast data. This data is typically unstructured and unlabeled, such as a large text corpus for natural language processing (NLP) tasks. The model’s objective during pre-training is to learn the data’s underlying patterns, structures, and representations. This pre-training phase is achieved through deep neural network architectures. While the pre-trained model has gained substantial general knowledge during the pre-training phase, it is not yet task-specific. It goes through fine-tuning to make a valuable model for a particular task. During fine-tuning, the model is trained on a smaller, task-specific dataset. This dataset consists of labelled examples that are relevant to the specific task the model is intended to perform. For instance, if the pre-trained model was initially trained on general language understanding, it might be fine-tuned for a specific NLP task, like text classification, translation, or question answering. The fine-tuning process allows the model to adapt its general knowledge tothe nuances of the particular task. It learns how to utilize its pre-trained understanding to make predictions or generate accurate and relevant responses for the task at hand.

[0169] In some aspects, the ML model is selected from the group consisting of Artificial Neural Network (ANN) model, Convolutional Neural Network (CNN) model, Recurrent Neural Network (RNN) model, Large Language Model (LLM), or a combination thereof.

[0170] In some aspects, the ML model comprises, consists, or consists essentially of an ANN. As used herein, the terms “ANN” and “Artificial Neural Network” refer to an information processing system in which a plurality of neurons called nodes or processing elements are connected in the form of a layer structure by modeling the operating principle of biological neurons and connection relationship between neurons. An artificial neural network is an information processing system in which a plurality of neurons called nodes or processing elements are connected in the form of a layer structure by modeling the operating principle of biological neurons and the connection relationship between neurons. Specifically, an artificial neural network is an overall model that has problem-solving ability by changing synapse coupling strength through learning of artificial neurons (nodes) that form a network by synapse coupling. The term artificial neural network may be used interchangeably with the term neural network. An artificial neural network may include a plurality of layers, and each of the layers may include a plurality of neurons. In addition, the artificial neural network may include neurons and synapses connecting neurons. Artificial neural networks generally use the following three factors: (1) connection patterns between neurons in different layers, (2) a learning process that updates the weights of connections, and (3) an output value from the weighted sum of the inputs received from the previous layer. It can be defined by the activation function you create.

[0171] Artificial neural networks are classified into single-layer neural networks and multilayer neural networks according to the number of layers. A typical single-layer neural network consists of an input layer and an output layer. In addition, a general multilayer neural network is composed of an input layer, one or more hidden layers, and an output layer. The input layer is a layer that accepts external data. The number of neurons in the input layer is the same as the number of input variables. The hidden layer is located between the input layer and the output layer. The output layer receives a signal from the hidden layer and outputs an output value based on the received signal. The input signal between neurons is multiplied by each connection strength (weight) and then summed. If this sum is greater than the neuron's threshold, the neuron is activated and outputs the output value obtained through the activation function.

[0172] In some aspects, the ML model comprises, consists, or consists essentially of a CNN model. As used herein, the terms “CNN” and “Convolutional Neural Network” refers to a deep feed-forward artificial neural network. In some aspects, a convolutional neural network includes a plurality of convolutional layers, a plurality of up-sampling layers, and a plurality of downsampling layers. For example, a respective one of the plurality of convolutional layers can process an input. An up-sampling layer and a down-sampling layer can change a scale of an input to one corresponding to a certain convolutional layer. The output from the up-sampling layer or the downsampling layer can then be processed by a convolutional layer of a corresponding scale. This enables the convolutional layer to add or extract a feature having a scale different from that of the input. By pre-training, parameters include, but are not limited to, a convolutional kernel, a bias, and a weight of a convolutional layer of a convolutional neural network can be tuned. Accordingly, the convolutional neural network can be used in various applications. Biomedical application of convolutional neural networks are disclosed, for example, at U.S. Pat. Nos. 10,242,443B2 or 10,811,135B2.

[0173] In some aspects, the ML model comprises, consists, or consists essentially of an RNN. As used herein, the terms “RNN” and “Recurrent Neural Network” refer to neural networks which, in contrast to feedforward networks, are distinguished by links of neurons (i.e. nodes) of one layer to neurons of the same or a preceding layer. This is the preferred manner of interconnection of neural networks in the brain, in particular in the neocortex. In artificial neural networks, recurrent interconnections of model neurons are frequently used to discover time-encoded — i.e. dynamic — information in the data. Examples of such recurrent neural networks include the Elman network, the Jordan network, the Hopfield network and the fully connected neural network. Ordinary feed forward neural networks are only meant for data points, which are independent of each other. RNN method can be long short-term memory (LSTM) or Gated Recurrent Unit (GRU) network. RNNs can work in conjunction with CNNs to form networks, such as the CNN-LSTM. As used herein, the term “Gated Recurrent Unit (GRU)” refers to a type of Recurrent Neural Network (RNN) and uses less memory. It is a part of a specific model of recurrent neural network that intends to use connections through a sequence of nodes to perform machine learning tasks associated with memory and clustering. It has a gating mechanism in recurrent neural networks.

[0174] In some aspects, the RNN comprises, consists, or consists essentially of a Long Short-Term Memory (LSTM) model. As used herein, the terms “LSTM” and “Long Short-Term Memory” refers to a type of RNN aimed at mitigating the vanishing gradient problem commonlyencountered by traditional RNNs. Its relative insensitivity to gap length is its advantage over other RNNs, hidden Markov models, and other sequence learning methods. It aims to provide a shortterm memory for RNN that can last thousands of timesteps (thus "long short-term memory"). An LSTM unit is typically composed of a cell and three gates: an input gate, an output gate, and a forget gate. The cell remembers values over arbitrary time intervals, and the gates regulate the flow of information into and out of the cell. Forget gates decide what information to discard from the previous state, by mapping the previous state and the current input to a value between 0 and 1. A (rounded) value of 1 signifies retention of the information, and a value of 0 represents discarding. Input gates decide which pieces of new information to store in the current cell state, using the same system as forget gates. Output gates control which pieces of information in the current cell state to output, by assigning a value from 0 to 1 to the information, considering the previous and current states. Selectively outputting relevant information from the current state allows the LSTM network to maintain useful, long-term dependencies to make predictions, both in current and future timesteps.

[0175] In some aspects, the ML model comprises, consists, or consists essentially of an LLM. As used herein, the terms “LLM” and “Large Language Model” refers to a type of machine learning model designed for natural language processing tasks such as language generation. LLMs are language models with many parameters, and are trained with self-supervised learning on a vast amount of text. In the methods disclosed herein such vast amounts of text are primary sequences of immunogenic and non-immunogenic proteins. The largest and most capable LLMs are generative pretrained transformer (GPTs). Accordingly, in some implementations of the methods disclosed herein, the LLM is a GPT. As machine learning algorithms process numbers rather than text (i.e., protein sequences), the text must be converted to numbers through a tokenization process, i.e., the sequences must be tokenized. Thus, in the first step, a vocabulary is decided upon, then indices are arbitrarily but uniquely assigned to each vocabulary entry (e.g., each amino acid could be associated to a specific numerical token), and finally, an embedding is associated to the integer index. Tokenization also compresses the datasets. In the context of training LLMs, datasets are typically cleaned by removing low-quality or duplicated data. Cleaned datasets can increase training efficiency and lead to improved downstream performance. In some aspects, a trained LLM can be used to clean datasets for training a further LLM.

[0176] In some aspects, the methods of the present disclosure can be implemented using an LLM that comprises, consists, or consists essentially of a Protein Language Model (pLM). In some aspects, the pLM is an ESM model, e.g., ESM-2. ESM is a transformer-based proteinlanguage model which is presently considered the state-of-the-art in protein language modeling. ESM-2 has several different model sizes, ranging from eight million to 15 billion parameters. See, e.g., Rives et al. (2021) “Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences” Proc. Natl. Acad. Sci. 118(15):e2016239118, disclosing ESM-lb; and Lin et al. (2023) “Evolutionary-scale prediction of atomic-level protein structure with a language model” Science, 379(6637):1123-1130, disclosing ESM-2; both of which are incorporated by reference in their entireties.

[0177] In some aspects, the methods of the present disclosure can be implemented using a sequence-based ML model comprising ESM-2, AbLang, AntiBERTa, AntiBERTy, ESM-lb, IgLM, nanoBERT, Progen2-OAS, ProtBERT, AbDiffuser, DiffAb, EAGLE, FvHallucinator, RefineGNNor a combination thereof. See Joubbi et al. (2024) Briefings in Bioinformatics 25(4): bbae307, which is herein incorporated by reference in its entirety.

[0178] In some aspects, the immunogenicity predictions methods disclosed herein are combined with deep mutational scanning (DMS). DMS is a high-throughput approach that systematically assesses the effects of all possible single amino acid substitutions on protein function, e.g., protein immunogenicity, providing comprehensive mutational landscapes to guide engineering efforts. DMS involves creating large libraries of protein variants containing thousands to millions of mutations, followed by selection or screening and high throughput sequencing to quantify the functional effects of each mutation. DMS data can be used to train machine learning models for predicting the effect of mutations, further enhancing protein engineering capabilities.

[0179] In some aspects, the LLM is a fine-tuned ESM-2 model. In some aspects, finetuning of the fine-tuned ESM-2 model comprises further training a plurality of layers using the training data wherein the weights of the remaining layers are not updated / trained.

[0180] Classification layers are a crucial component in machine learning, particularly in the field of neural networks, serving as the final step in many classification tasks. Essentially, a classification layer is responsible for taking the output from the preceding layers and transforming it into probabilities for each possible class. Classification layers typically employ an activation function such as softmax or sigmoid to produce a probability distribution across different classes. Accordingly, in some aspects, the ML model further comprises a classification layer.

[0181] In some aspects, the classification layer is a binary classification layer. In some aspects, the classification layer is a two layer Multi-Layer perceptron neural network. In some aspects, the input to the classification layer are the embeddings generated using the Protein Language model. The model is trained using a cross entropy loss function to learn the patterns inthe embeddings and a SoftMax activation function is used to generate the output probability of a peptide being immunogenic.

[0182] In some aspects, the binary classification layer outputs the probability of immunogenicity of a peptide. In some aspects, the binary classification layer outputs the probability of non-immunogenicity of a peptide.IV. Processing of the candidate protein sequence into subsequences

[0183] The predictions of immunogenicity according to the methods disclosed herein comprise a step or combination thereof in which the candidate protein is processed before input in the ML model. Such processing generally requires splitting the candidate protein into subsequences, e.g., using a sliding window and a sliding step. However, in some aspects, the entire sequence of the candidate protein can be used as input.

[0184] In some aspects, the plurality of peptide subsequences derived from the candidate protein sequence are generated using a sliding window and a sliding step. In some aspects, the sliding window has a variable length. In some aspects, the sliding window has a constant length. In some aspects, the sliding window constant length is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids. In some aspects, the sliding window constant length is at least 1, at least 2, at least 3, at least 4, at least 5, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 12, at least 13, at least 14 or at least 15 amino acids. In some aspects, the sliding window constant length is between 1 and 5, between 5 and 10, between 10 and 15, between 10 and 20, between 12 and 18, between 11 and 16, between 12 and 17, or between 13 and 18. In some aspects, the sliding window length is 15 amino acids. In some aspects, the sliding window has a variable length. In some aspects, the variable length is between 5 and 25 amino acids, for examples, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25. The term “variable length” means that the sliding window for a certain candidate protein can be, for example, 5 amino acids, whereas for another candidate protein the sliding window could be, for example, 10 amino acids. In some aspects, the ‘sliding window’ is the whole length of the protein.

[0185] In some aspects, the sliding step has a constant length. In some aspects, the sliding step length is 1 amino acid. In some aspects, the sliding step length is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids. In some aspects, the sliding step has a variable length. In some aspects, the ‘sliding step’ is the full-length of the protein.

[0186] In some aspects, the plurality of peptide subsequences derived from the candidate protein sequence are overlapping. In some aspects, the plurality of peptide subsequences derived from the candidate protein sequence are non-overlapping. In some aspects, the overlap between sequences is by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 amino acids.

[0187] In some aspects, the plurality of peptide subsequences derived from the candidate protein sequence have the same length. In some aspects, the plurality of peptide subsequences derived from the candidate protein sequence can have different lengths.

[0188] In some specific aspects, the sliding window has a length of 15 amino acids, and the sliding step has a length of 1 amino acid.

[0189] Once the candidate protein sequence, or a population thereof, has been processed, the plurality of peptide subsequences if used as input to the ML model, and the output of the ML model is processed by the classification layers, and immunogenicity prediction is outputted.V. Immunogenicity prediction

[0190] The methods disclosed herein provide an immunogenicity prediction, which as discussed above is the output of the ML model after it has been processed by a classification layer. In some aspects, the immunogenicity prediction comprises a binary prediction of cytokine release induced by the candidate protein or a fragment thereof in a subject. In order words, in some aspects, the output indicates whether the candidate protein sequence or a subsequence thereof is immunogenic or non-immunogenic.

[0191] In general, using inputs of one or more amino acid sequences (i.e., a candidate protein) and a window size n, the ML model outputs a quantitative, 0-1 score corresponding to the degree of immunogenicity for each overlapping n-mer peptide. This is the primary output metric for the ML model, from which other metrics may be derived. The primary output metric numeric score can be used to classify the input peptide derived from the candidate protein as immunogenic / nonimmunogenic based on a threshold value, e.g., 0.5. This threshold value can be determined in a global manner (e.g., n-th percentile of values across all peptides in the training set), or a local manner (e.g., top n% of peptides in the input candidate protein).

[0192] Immunogenic hotspots on the candidate protein may be determined using the primary output metric values determined across multiple overlapping peptides. For example, a numeric value may be derived by averaging across n overlapping windows, which may be compared against a numeric threshold as described above to output a binary immunogenic / nonimmunogenic classification. A binary classification may also be determined bylooking at the number of peptides in a range above a certain threshold, i.e., at least k of n consecutive peptides must have an immunogenicity score above the threshold to be classified as immunogenic.

[0193] Additionally, a numeric or binary value for the immunogenicity of a full-length candidate protein may be derived from the scores (primary output metric values) generated for individual peptides as described above. For example, the sum, mean, median, interquartile range (IQR), or nth percentile of the immunogenicity scores for all overlapping peptides in a full-length candidate protein may be used as an immunogenicity score for the full-length candidate protein. A global or local threshold may be used to classify the candidate protein as immunogenic / non-immunogenic.

[0194] In some aspects, the immunogenicity prediction comprises a probability of cytokine release induced by the candidate protein or a fragment thereof in a subject.

[0195] In some aspects, the probability of cytokine release comprise a prediction of the release of a cytokine selected from the group consisting of IL-1, IL-2, IL-3, IL-4, IL-6, IL-8, IL-10, IL-13, IL-15, TNFalpha, TGFbeta, CD40L, IFNalpha, IFNbeta, IFNgamma, and any combination thereof. In one specific aspect, the cytokine is IFNgamma. In one specific aspect, the cytokine is IL-1. In one specific aspect, the cytokine is IL-2. In one specific aspect, the cytokine is IL-3. In one specific aspect, the cytokine is IL-4. In one specific aspect, the cytokine is IL-6. In one specific aspect, the cytokine is IL-8. In one specific aspect, the cytokine is IL-10. In one specific aspect, the cytokine is IL- 13. In one specific aspect, the cytokine is IL-15. In one specific aspect, the cytokine is TNFalpha. In one specific aspect, the cytokine is TGFbeta. In one specific aspect, the cytokine is CD40L. In one specific aspect, the cytokine is IFNalpha. In one specific aspect, the cytokine is IFNbeta.

[0196] In some aspects, the methods disclosed herein are fine-tuned to predict the release of a specific cytokine in the training data set. In some aspects, the methods disclosed herein are fine-tuned to predict Treg epitopes. Treg epitopes or tregitopes are natural regulatory T cell epitopes derived from immunoglobulin G (IgG) that stimulate CD25 (also known as IL2-RA or interleukin-2 receptor alpha chain), FoxP3 (forkhead box P3 protein, also known as scurfin), and T cells to expand.

[0197] In some aspects, the predicted cytokine release comprises a prediction of the severity of the cytokine release. For example, the predicted cytokine release can include a prediction of the likelihood of triggering Cytokine Release Syndrome (CRS). Cytokine release syndrome (CRS) is a form of systemic inflammatory response syndrome (SIRS) that can betriggered by a variety of factors such as infections and certain drugs. Thus, in some aspects, the predicted cytokine release can comprise a predicted grade of the CRS. In some aspects, the immunogenicity prediction comprises a CRS prediction, wherein the output is a CRS grade selected from the group consisting of Grade 1, Grade 2, Grade 3, Grade 4, and Grade. Grade 1 CRS comprises mild reaction, such that intervention is not indicated. Grade 2 CRS comprises a more severe reaction that requires symptomatic treatment, e.g., antihistamines, NSAIDS, narcotics, IV fluids, or prophylactic medication for less than 24 hours. Grade 2 CRS comprises prolonged recurrence of symptoms following initial improvement, and hospitalization is indicated for clinical sequelae such as renal impairment or pulmonary infiltrates. Grade 4 CRS comprises lifethreatening consequences, and pressor or ventilator support is indicated. Grade 5 CRS results in death.

[0198] In some aspects, generating an immunogenicity prediction for the candidate protein based comprises identifying one or more peptide subsequences having immunogenicity probabilities above a predetermined probability threshold. In some aspect, the immunogenicity probability threshold corresponds to probability percentage, for example, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70% about 75%, or about 80%. Thus, in some aspects, a candidate protein is considered immunogenic if contiguous subsequences corresponding to a certain length or a percentage of its length with respect to the total length of the candidate sequence have an immunogenicity probability above a predetermined probability threshold. Conversely, in some aspects, a candidate protein is considered non-immunogenic if a contiguous subsequence corresponding to a certain length or a percentage of its length with respect to the total length of the candidate sequence has an immunogenicity probability below a predetermined probability threshold.

[0199] In some aspects, the length of the contiguous sequence above or below the predetermined probability threshold must be at least about 10, at least about 15, at least about 20, or at least about 25 amino acids. In some aspects, the length of the contiguous sequence above or below the predetermined probability threshold must be no less than about 15, about 14, about 13, about 12, about 11, about 10, about 9, about 8, about 7, about 6, or about 5 amino acids. In some aspects, the percentage of the length of subsequences above or below an immunogenicity probability threshold with respect to the total length of the candidate sequence for the candidate protein to be considered immunogenic or non-immunogenic is about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%,about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 100%.

[0200] In some aspects, the output of the immunogenicity predictions is a combination of metrics. In some aspects, the output of the immunogenicity prediction comprises the calculation of (1) median of lowest quantile (MQ1), (2) median of the top quantile (MQ3), (3) distance between median of lowest quantile (MQ1) and median of the top quantile (MQ3) (MQ3-MQ1), (4) dispersion between median of lowest quantile (MQ1) and median of the top quantile (MQ3) (MQ3-MQ1) / MQ3 ).

[0201] In some aspects, the output of the immunogenicity prediction comprises using halfs, thirds, quantiles, pentiles, decimals or 5% fractions.

[0202] In some aspects, generating an immunogenicity prediction for the candidate protein (or a plurality thereof, e.g., a library of candidate proteins) comprises identifying one or more immunogenicity hot spots having immunogenicity probabilities above a predetermined threshold in the candidate protein (or in a plurality thereof, e.g., in a library of candidate proteins) or any other immunogenicity probabilities disclosed herein or combination thereof.

[0203] As used herein, the term “immunogenicity hot spot” refers to a contiguous amino acid subsequence in a candidate protein primary sequence having immunogenicity probabilities above a predetermined threshold value or matching any other metric that is considered to represent an immunogenic phenotype, i.e., the presence of such metric is associate with a specific immune response, for example, the release of a specific cytokine or combination thereof.

[0204] In some aspects, generating an immunogenicity prediction for the candidate protein comprises, for example:(i) determining a global candidate protein immunogenicity score,(ii) determining the number of immunogenicity hotspots on the candidate protein, (iii) determining the distance between immunogenicity hotspots on the candidate protein, i.e., number of amino acids between immunogenicity hotspots subsequences in the candidate protein primary sequence,(iv) determining the length (in number of amino acids) of the immunogenicity hotspots with respect to the length of primary amino acid sequence of the candidate protein, (v) determining the relative length (in percentage) of the immunogenicity hotspots with respect to the length of the primary amino acid sequence of the candidate protein, (vi) a combination thereof.

[0205] In some aspects, generating an immunogenicity prediction for the candidate protein comprises determining propensity to cause Anti-Drug Antibodies (ADA). For example, the predicted propensity rate of formation of anti-drug antibodies can be determined by using mean or median probability scores for the entire candidate protein, or by using any of the metrics described above, e.g., median or mean of different quantiles, thirds, decimals, halfs, etc.

[0206] In some aspects, determining the global candidate protein immunogenicity score comprising integrating the immunogenicity probabilities of each of the peptide subsequences in the candidate protein. In some aspects, integrating the immunogenicity probabilities comprises calculating an average score. In some aspects, determining the global candidate protein immunogenicity score comprises calculating a Z-score. A Z-score, also known as a standard score, measures how many standard deviations a data point is from the mean of a dataset.

[0207] In some aspects, the immunogenicity prediction is outputted as a graphic representation showing the immunogenicity probabilities of each of the peptide subsequences in the candidate protein along the sequence of the candidate protein, wherein immunogenic regions or immunogenic hotspots correspond to segments of the candidate protein sequence having the immunogenicity probabilities above a predetermined threshold. In some aspects, the threshold equals 0.5 (50% probability). In other aspects, the threshold is equal to MQ3.

[0208] The methods disclosed herein can be used to determine the immunogenicity of an immunogenic protein or the non-immunogenicity of a non-immunogenic protein. Thus, in some aspects, the candidate protein is an immunogenic protein. In other aspects, the candidate protein is a non-immunogenic protein. In some aspect, the input of the systems and methods disclosed herein is a population of proteins. In some aspect, the input of the systems and methods disclosed herein is a plurality of proteins. In some aspect, the input of the systems and methods disclosed herein is a library of proteins. In some aspect, the input of the systems and methods disclosed herein is a population of candidate proteins for vaccine development. In some aspect, the input of the systems and methods disclosed herein is a population of candidate epitopes for antibody development. In some aspect, the input of the systems and methods disclosed herein is a population of antibodies. In some aspect, the input of the systems and methods disclosed herein is a population of therapeutic proteins. In some aspect, the input of the systems and methods disclosed herein is a population of potential environmental allergens. In some aspect, the input of the systems and methods disclosed herein is a population of proteins that have been mutated in silico. In some aspect, the input of the systems and methods disclosed herein is a population of variants of a candidate protein.

[0209] In some aspects, the output of the systems and methods disclosed herein can identify immunogenic hot spots in a candidate protein. In some aspects, the identification of immunogenic hot spots comprises:(i) identifying the number of immunogenic hot spots;(ii) identifying the location of the immunogenic hot spots;(iii) determining the length of the immunogenic hot spots;(iv) determining the relative length of the hots spot with respect to the total length of the candidate protein;(v) determining the degree of immunogenicity of the hots spots;(vi) determining median of lowest quantile (MQ1);(vii) determining median of the top quantile (MQ3);(viii) determining distance between median of lowest quantile (MQ1) and median of the top quantile (MQ3) (MQ3-MQ1);(ix) determining dispersion between median of lowest quantile (MQ1) and median of the top quantile (MQ3) (MQ3-MQ1) / MQ3 );(x) determining mean or median distance between immunogenic hot spots;(xi) determining specific properties of the immunogenic hot spots; or,(xii) any combination thereof.

[0210] In some aspects, the specific properties of the immunogenic hotspots comprise, e.g., metrics (for example mean or median) related to amino acid charge; amino acid polarity; amino acid hydrophobicity; secondary structure propensity; aromaticity of amino acids; amino acid location propensity (e.g., tendency to be buried or exposed); presence / absence of specific functional sequences or motifs; presence / absence of MHC binding sites; presence / absence of know antibody epitopes; presence / absence of glycosylation target sequences; and any combinations thereof.VI. Computer-Based Implementations

[0211] The present disclosure provides a computer-implemented method of predicting the immunogenicity of a peptide, protein, or a plurality thereof comprising (i) inputting the amino acid sequence the peptide, protein, or plurality thereof into a computer system comprising a program comprising instructions for executing the immunogenicity prediction methods disclosed herein, wherein the peptide, protein, or plurality is the candidate protein; (ii) executing in the computersystem the immunogenicity prediction method; and, (iii) receiving from the computer system the immunogenicity prediction resulting from executing the immunogenicity prediction method.

[0212] In some aspects, the present disclosure provides a computer system for predicting the immunogenicity of a candidate protein or a plurality of candidate proteins, the computer system comprising at least one processor and memory storing at least one program for execution by the at least one processor, the at least one program comprising instructions to execute the computer-implemented methods for prediction of immunogenicity disclosed herein.

[0213] In some aspects, the present disclosure provides a system for predicting the immunogenicity of a candidate protein according to the computer-implemented methods of immunogenicity prediction disclosed herein, wherein the system is stored on computer-readable storage media.

[0214] In some aspects, the code to execute the computer-implemented method, the training database of the computer-implemented method, the ML model of the computer-implemented method, the output of the ML model of the computer-implemented method, the immunogenicity prediction outputted by the computer-implemented method, or any combination thereof is hosted in a cloud computing environment.

[0215] The present disclosure also provides a non-transitory computer readable storage medium storing a computational module for predicting the immunogenicity of a protein, the computational module comprising code to execute the computer-implemented methods of the present disclosure. The present disclosure also provides a computer readable storage medium having computer readable instructions to instruct a computer to perform the computer-implemented methods of the present disclosure.

[0216] The methods disclosed herein can be implemented by computer-executable instructions stored on one or more computer-readable media or conveyed by a signal of any suitable type. The steps of the methods disclosed herein can be implemented by software or combinations of software and hardware and in any of the ways described above. The computer-executable instructions can be the same process executing on a single or a plurality of microprocessors or multiple processes executing on a single or a plurality of microprocessors. The methods disclosed herein can be repeated any number of times as needed and the steps of the methods can be performed in any suitable order.

[0217] The subject matter described herein can operate in the general context of computerexecutable instructions, such as program modules, executed by one or more components. Generally, program modules include routines, programs, objects, data structures, etc., that performparticular tasks or implement particular abstract data types. Typically, the functionality of the program modules can be combined or distributed as desired. Although the description above relates generally to computer-executable instructions of a computer program that runs on a computer and / or computers, the user interfaces, methods and systems also can be implemented in combination with other program modules. Generally, program modules include routines, programs, components, data structures, etc. that perform particular tasks and / or implement particular abstract data types.

[0218] Moreover, the subject matter described herein can be practiced with most any suitable computer system configurations, including single-processor or multiprocessor computer systems, mini-computing devices, mainframe computers, personal computers, stand-alone computers, hand-held computing devices, wearable computing devices, microprocessor-based or programmable consumer electronics, and the like as well as distributed computing environments in which tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote memory storage devices. The methods and systems described herein can be embodied on a computer-readable medium having computer-executable instructions as well as signals (e.g., electronic signals) manufactured to transmit such information, for instance, on a network.

[0219] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing some of the claims.

[0220] It is, of course, not possible to describe every conceivable combination of components or methodologies that fall within the claimed subject matter, and many further combinations and permutations of the subject matter are possible. While a particular feature may have been disclosed with respect to only one of several implementations, such feature can be combined with one or more other features of the other implementations of the subject matter as may be desired and advantageous for any given or particular application.

[0221] Moreover, it is to be appreciated that various aspects as described herein can be implemented on portable computing devices (e.g., field medical device), and other aspects can be implemented across distributed computing platforms (e.g., remote medicine, or researchapplications). Likewise, various aspects as described herein can be implemented as a set of services (e.g., modeling, predicting, analytics, etc.).

[0222] Generally, program modules include routines, programs, components, data structures, etc., that perform particular tasks or implement particular abstract data types. Moreover, those skilled in the art will appreciate that the inventive methods can be practiced with other computer system configurations, including single-processor or multiprocessor computer systems, minicomputers, mainframe computers, as well as personal computers, hand-held computing devices, microprocessor-based or programmable consumer electronics, and the like, each of which can be operatively coupled to one or more associated devices.

[0223] The illustrated aspects of the specification may also be practiced in distributed computing environments where certain tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote memory storage devices.VII. Methods of Use

[0224] The immunogenicity prediction methods disclosed therein can be used in numerous applications. In general, the methods disclosed herein can be used to predict the immunogenicity on non-immunogenicity of a candidate protein, and based on those predictions, the candidate protein can be modified with the goal of increasing or decreasing its immunogenicity.

[0225] By "increased immunogenicity," "increasing immunogenicity," and grammatical equivalents herein is meant an increased ability to activate the immune system, when compared to a control, e.g., an unmodified protein. For example, a modified protein can be said to have "increased immunogenicity" if it elicits neutralizing or non-neutralizing antibodies in higher titer or in more subjects than an unmodified protein. In a particular embodiment, the probability of raising neutralizing antibodies is increased by at least 5%, e.g., at least 2-fold or at least 5-fold. For example, if an unmodified protein produces an immune response in 10% of subjects, a variant with enhanced immunogenicity would produce an immune response in more than 10% of subjects, e.g., more than 20% or more than 50% of subjects. A modified protein also can be said to have "increased immunogenicity" if it shows increased binding to one or more MHCI or MHCII alleles or if it induces T cell activation in an increased fraction of subjects relative to the parent protein. In some embodiments, the probability of T cell activation is increased by at least 5%, e.g., at least 2-fold or at least 5-fold.

[0226] By "reduced immunogenicity," "reducing immunogenicity," and grammatical equivalents herein is meant a decreased ability to activate the immune system, when compared to a control, e.g., an unmodified protein. For example, if the therapeutic agent is protein, a modified protein can be said to have "reduced immunogenicity" if it elicits neutralizing or non-neutralizing antibodies in lower titer or in fewer subjects than an unmodified protein. In some embodiments, the probability of raising neutralizing antibodies is decreased by, for example, at least 5%, e.g., at least 50% or at least 90%. For example, if a parent protein produces an immune response in 10% of subjects, a modified protein with reduced immunogenicity would produce an immune response in less than 10% of subjects, e.g., less than 5% or less than 1% of subjects. A modified protein also can be said to have "reduced immunogenicity" if it shows decreased binding to one or more MHCI or MHCII alleles or if it induces T cell activation in a decreased fraction of subjects relative to an unmodified protein. In some embodiments, the probability of T cell activation is decreased by at least 5%, e.g., by at least 50% or at least 90%.

[0227] The term “modulating immunogenicity” refers to altering the immunogenicity of a candidate protein, and encompasses both increasing and decreasing the immunogenicity of a protein, transforming an immunogenic protein into a non-immunogenic protein, or transforming a non-immunogenic protein into an immunogenic protein.

[0228] In some aspects, the methods disclosed herein can be used to generate an immunostimulatory protein using as an immunologically inert protein as starting point. Conversely, in some aspects, the methods disclosed herein can be used to generate an immunologically inert protein using an immunostimulatory protein as starting point. By "immunostimulatory" and grammatical equivalents herein is meant a part of a protein that stimulates an immune response that is greater than that generated in other, e.g., immunologically inert, parts of a protein. By "immunologically inert" and grammatical equivalents herein is meant a part of a protein that that does not stimulate an immune response.

[0229] Accordingly, in some aspect, the present disclosure provides methods for epitope prediction, i.e., the methods can be used to determine the immunogenicity of a subsequence of a target sequence (e.g., a spike protein of a virus), and such subsequence (i.e., an epitope) can be used to develop antibodies to treat a disease or condition or a vaccine. In some aspects, the methods disclosed herein can be used to engineer a protein, e.g., by increasing or decreasing its predicted immunogenicity, for example, above or below a certain threshold or with respect to a reference protein.

[0230] In some aspects, the present disclosure provides a method to engineer a therapeutic protein comprising (i) predicting the immunogenicity of the therapeutic protein by applying the computer-implemented methods of immunogenicity prediction disclosed herein to the therapeutic protein, wherein the therapeutic protein is the candidate protein; (ii) introducing one or more amino acid modifications to the therapeutic protein of step (i) to generate a modified therapeutic protein; (iii) predicting the immunogenicity of the modified protein by applying the computer-implemented methods of immunogenicity prediction disclosed herein to the modified therapeutic protein, wherein the modified therapeutic protein is the candidate protein; (iv) comparing the predicted immunogenicity of the therapeutic protein of (i) and the modified therapeutic protein of (iii); (v) optionally iterating (ii)-(iv) until the modified therapeutic protein has the desired degree of immunogenicity or lack thereof.

[0231] In some aspects, the immunogenicity prediction system disclosed herein can operate as a “Generative Al,” i.e., the ML model can suggest mutations, deletions, substitution, or other modifications to the sequence of the candidate protein to elicit changed in immunogenicity, e.g., increases, decreases, or modulation (e.g., variation in the degree of immunogenicity) of the candidate protein.

[0232] In some aspects, potential amino acid modifications to the sequence of the candidate protein can be selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof. In some aspects, the amino acid substitutions are conservative substitutions, non-conservative substitutions, or a combination thereof.

[0233] The present disclosure also provides a method to screen in silico a library of candidate proteins to determine their predicted immunogenicity, wherein the method comprises applying the computer-implemented method of immunogenicity prediction disclosed herein to the amino acid sequences of the library of candidate proteins.

[0234] Also provided is method of de-epitoping a therapeutic candidate protein comprising (i) applying the computer-implemented method of immunogenicity prediction disclosed herein to the therapeutic candidate protein or a modified variant thereof, (ii) applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof to the therapeutic candidate protein, (iii) applying the computer-implemented method of immunogenicity prediction disclosed herein to the therapeutic candidate protein or a modified variant thereof, and (iv) iterating steps (i) to (iii) to obtain a modified variantof the therapeutic candidate protein containing less antigenic epitopes than the therapeutic candidate protein. As use herein, the terms “de-epitoping” or “deepitoping” refers to removing an immunogenic epitope from an immunogenic protein, for example, by introducing mutations in the epitope or by excising the epitope.

[0235] In some aspects, the present disclosure provides a method to rescue a protein or peptide drug that triggers negative immunogenic effects by applying the computer-implemented methods of immunogenicity prediction disclosed herein. Also provided is a method to predict negative responses to a drug comprising or consisting of a protein, wherein the method comprises applying the computer-implemented methods of immunogenicity prediction disclosed herein to the drug. In some aspects, the protein drug is used as a candidate or is a candidate for use in a clinical trial. As used herein, the term “rescue” refers to making a candidate therapeutic protein previously discarded as unsuitable for therapeutic use due to undesirable immunogenic characteristics by modifying the immunogenic profile of the protein, e.g., by de-epitopizing the protein. Thus, the rescued protein becomes suitable for therapeutic use due to its reduced immunogenicity.

[0236] Also provided is a method of vaccine design comprising applying the computer-implemented methods of immunogenicity prediction disclosed herein to a vaccine candidate protein. In some aspects, vaccine design comprises selecting the vaccine candidate protein if the predicted immunogenicity of the vaccine candidate protein is above a predetermined threshold value. In some aspects, vaccine design comprises modifying the vaccine candidate protein by applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof to the vaccine candidate protein, and iteratively executing the computer-implemented methods of immunogenicity prediction disclosed herein.

[0237] Also provided is a method of modulating the immunogenicity of a candidate protein comprising applying the computer-implemented methods of immunogenicity prediction disclosed herein to the candidate protein. In some aspects, modulating comprises increasing the immunogenicity of the candidate protein by applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof to the candidate protein. In some aspects, modulating comprises decreasing the immunogenicity of the candidate protein by applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino aciddeletions, protein truncations, altering a post-translational modification, or any combination thereof to the candidate protein. In some aspects, modulating comprises introducing at least one immunogenic hot spot on the protein by applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof to the candidate protein. In some aspects, modulating comprises removing at least one immunogenic hot spot from the protein by applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof to the candidate protein.

[0238] Also provided is a method of identifying an immunogenic epitope susceptible to antibody targeting in a candidate protein comprising applying the computer-implemented methods of immunogenicity prediction disclosed herein to the candidate protein. Also provided is a method of generating an amino acid sequence predicted to have altered immunogenicity compared to the amino acid sequence a parent candidate protein comprising applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof to the amino acid sequence of parent candidate protein to generate the amino acid sequence of the modified candidate protein and determining the predicted immunogenicity of the modified candidate protein by applying the computer-implemented methods of immunogenicity prediction disclosed herein to the amino acid sequence of modified candidate protein.

[0239] In some aspects of the methods disclosed herein the candidate protein is selected from the group consisting of a vaccine immunogen, an antibody, MSTAR, scFab, Fab, scFv, Fv, Fc, an enzyme, a growth factor, a chimeric antigen receptor (CAR), a T-cell receptor (TCR), a cytokine, chimeric cytokine or a cytokine receptor, GLP1, GLP2, and their close receptor and close analogues; soluble and non-soluble forms of receptors mentioned above. In principle, the methods disclosed herein are applicable to any immunogenic or non-immunogenic protein or subset thereof. In some aspects, the computer-implemented methods of immunogenicity prediction disclosed herein further comprise evaluating the strength of interactions between (i) the candidate protein in a complex with the MHCI protein, (ii) and the candidate protein in a complex with the MHCII protein, (iii) the candidate protein in a complex with the MHC proteins, (iv) the candidate protein and a T-cell receptor. In some aspects, the computer-implemented methods of immunogenicityprediction disclosed herein further comprise evaluating the strength of interactions between (i) the candidate protein and (ii) the MHCII protein. In some aspects, the computer-implemented methods of immunogenicity prediction disclosed herein further comprise evaluating the strength of interactions between (i) the candidate protein and (ii) the protein-binding cavity of a complex comprising a MHCII protein and a T cell receptor. In some aspects, the computer-implemented methods of immunogenicity prediction disclosed herein further comprise inputting sequence information of the MHCII protein and / or the T cell receptor.

[0240] The present disclosure also provides a method to predict the immunogenicity of a plurality of fragments of a protein associated with cancer, and identifying a fragment of said protein that is predicted to be immunostimulatory comprising the computer-implemented methods of immunogenicity prediction disclosed herein to the plurality of fragments.

[0241] Also provided is a method of producing a personalized cancer vaccine for a subject having a tumor, the method comprising the steps of (i) identifying a plurality of modified peptides expressed in the tumor, each comprising an amino acid substitution at a position, relative to a corresponding parent peptide expressed in the normal cells; (ii) determine the immunogenicity for each of the plurality of modified peptides, via the computer-implemented methods of immunogenicity prediction disclosed herein, wherein modified peptide is a candidate protein; and (iii) producing a personalized cancer vaccine for the subject, which comprises a peptide or polypeptide comprising the at least one modified peptide selected as immunogenic.

[0242] Also provided is a method of generating a library of immunogenic proteins and their predicted immunogenicity, wherein each immunogenic protein and / or their predicted immunogenicity is generated via the computer-implemented methods of immunogenicity prediction disclosed herein.

[0243] The present disclosure also provides a method of selecting of therapeutic protein comprising applying the computer-implemented methods of immunogenicity prediction disclosed herein to a population of candidate proteins. In one particular aspect, the methods disclosed herein can be employed to identify an immunostimulatory fragment of a target protein (e.g., a therapeutic protein). In these aspects, the method can comprise identifying an immunostimulatory fragment of a protein, and altering that fragment, e.g., by altering the amino acid sequence of the fragment or by adding or removing a post-translational modification to or from the fragment, to decrease the immunogenicity of the protein. This method can be employed with any therapeutic protein, including but not limited to industrial, pharmaceutical, and agricultural proteins, including proteins that can be administered for the treatment of a blood disease or disorder, for example, anemia (e.g.,aplastic anemia, Fanconi anemia hemolytic anemia, sickle cell anemia hereditary spherocytosis, and thalassemia), hemoglobinuria, a blood coagulation disorder (including afibrinogenemia, factor V deficiency, factor VII deficiency, factor X deficiency, factor XI deficiency, factor XII deficiency, hemophilia A, hemophilia B, Von Willebrand disease, disseminated intravascular coagulation, antithrombin III deficiency, Bernard-Soulier syndrome, protein C deficiency, thrombasthenia, platelet storage pool deficiency, protein s deficiency), purpura (including Evans syndrome and thrombotic thrombocytopenic purpura), blood group incompatibility, a blood platelet disorder (e.g., thrombocytopenia), a blood protein disorder (e.g., cryoglobulinemia and Waldenstrom macro globulinemia), myelodysplasia syndrome, a myeloproliferative disorder, hemoglobinopathy, or a leukocyte disorder (e.g., eosinophilia, Kimura disease, leukopenia and neutropenia).

[0244] The present disclosure provides a method to treat or prevent a disease or condition comprising administering a protein, polynucleotide, vector, cell, pharmaceutical composition, or delivery system disclosed herein to a subject in need thereof. In some aspects, the disease or condition is selected from the group consisting of cancer, infection, chronic inflammation, genetic disease, and autoimmune disease. In some aspects, the subject is a human subject.

[0245] The method of immunogenicity prediction disclosed herein can be employed with proteins which are targets for the treatment of cancer and other diseases. Examples of such target proteins include and are not limited to ligands, cell surface receptors, antigens, antibodies, cytokines, hormones, transcription factors, signaling modules, cytoskeletal proteins, toxins and enzymes. Non-limiting examples of therapeutic proteins include, adenosine deamidase, arginase, asparaginase, bone morphogenic protein-7, ciliary neurotrophic factor, DNase, erythropoietin, factor IX, factor VIII, follicle stimulating hormone, glucocerebrocidase, gonadotrophin-releasing hormone, granulocyte-colony stimulating factor, granulocyte-macrophage-colony stimulating factor, growth hormone, growth hormone releasing hormone, human chorionic gonadotrophin, insulin, interferon alpha, interferon beta, interferon gamma, interleukin-2, interleukin-3, interleukin- 1, salmon calcitonin, staphylokinase, streptokinase, tissue plasminogen activator, and thrombopoietin. The parent protein can also comprise an extracellular domain of a receptor, including but not limited to CD4, interleukin- 1 receptor, tumor necrosis factor receptors, and antibodies (including a murine, chimeric, humanized, camelized, llamalized, single chain, or fully human antibodies). Proteinaceous therapeutic agents can be naturally occurring or synthetic.

[0246] In another aspect, the method of immunogenicity prediction disclosed herein can be employed to identify an immunologically inert fragment of a target protein. In these aspects,the method can comprise identifying an immunologically inert fragment of a target protein, and altering that fragment e.g., by altering the amino acid sequence of the fragment or by adding or removing a post-translational modification to or from the fragment, to increase the immunogenicity of the protein. The protein can be from any infectious disease, e.g., anthrax, chickenpox, diphtheria, hepatitis A, B or C, HIB, HPV, seasonal influenza, encephalitis, malaria, measles, meningitis, mumps, pertussis, polio, rabies, rubella, shingles, smallpox, tetanus, TB or yeller fever, etc. New targets for cancer therapy can be identified using this methodology, e.g. by identifying immunogenic epitopes associated with cancer. These epitopes can then be used to design vaccines or enhance a pre-existing immune response against the particular epitope.

[0247] In some aspects, the methods of immunogenicity prediction disclosed herein can also be used to identify an auto-antigen in a subject having an autoimmune disease. In this aspect, the T-cell repertoire of the individual can be sequenced, and protein sequences from the individual can be tested using the method described above to identify an immunodominant self-peptide, the sequence of which should allow the identification of the auto-antigen causing the autoimmune disease in the individual. In a similar aspect, the methods can be used to predict the immunogenicity of a plurality of fragments of a protein associated with a cancer, and to identify a fragment of the protein that is predicted to be immunostimulatory. The immunostimulatory fragment can be used as or developed into a cancer vaccine.

[0248] In some aspects, the computer-implemented methods of immunogenicity prediction disclosed herein can be used, for example, in methods to identify immunogenicity variants of a candidate protein or a population thereof, in methods to identify immunogenic hotspots in a candidate protein (e.g., a biologic) or in a library thereof, in methods to identify cancer neoantigens, in methods for vaccine epitope discovery, in methods for vaccine candidate selection and / or optimization and / or engineering, in methods for positive and / or negative immunogenicity detection, in methods of humanization and / or engineering of recombinant proteins to reduce, increase or modulate their immunogenicity, in methods of humanization and / or engineering of endogenous proteins to reduce, increase or modulate their immunogenicity, in methods to derisk biologies (for example, for development and / or candidate selection for clinical trials), in methods to rank candidate proteins based on their immunogenic profiles, in methods of treatment with specific biologies by selecting drugs capable of eliciting the appropriate response or avoiding a specific response, in methods of inducing or avoiding specific immune responses with different combinations of assays, in methods of preventing specific immune responses by selecting proteins with specific immunogenic profiles, in methods of eliciting specific immune responses by selectingproteins with specific immunogenic profiles, in methods of preventing rejection to transplantation by selecting proteins with specific immunogenic profiles, in methods of preventing specific immune responses by selecting proteins with specific immunogenic profiles, in methods of selecting specific antibody formats, in methods to select for or against Tregs, in methods to select for or against autoimmunity responses, in methods to select candidate proteins for development as vaccines (e.g., antiviral vaccines or anticancer vaccines) from a library of candidate proteins or via in silico iterative mutation, in methods to select candidate proteins for development as cancer therapeutics from a library of candidate proteins or via in silico iterative mutation, in methods to select proteins based on CD4 and or CD8 effector responses, in methods to select candidates proteins based on their stability (loss of structural integrity can cause loss of immunogenicity; see, e.g., Scheiblhofer et al. (2017) Expert Rev. Vaccines 16(5):479-489), or in methods to avoid adverse events caused my immune reactions triggered by a candidate protein.

[0249] In some aspects, the computer-implemented methods of immunogenicity prediction disclosed herein can be trained with specific sets of training proteins in order to identify targets with specific cytokine release profiled. For example, a training set that is IFNy+ (interferon gamma positive) but IL 10- (interleukin 10 negative) could be used to identify immunogenic proteins that do not induce Tregs.

[0250] In some aspects, the present disclosure provides methods to evaluate full-length candidate proteins. However, the methods disclosed herein can also be applied to evaluate the immunogenicity of a portion of a candidate protein, for example, a subunit of a multimeric protein, a domain of a multidomain protein, or a fragment or subsequence from the candidate protein.

[0251] In some aspects, the computer-implemented methods of immunogenicity prediction disclosed herein can provide an output based on a single, predetermined threshold. However, in other aspects, the output can contain protein-dependent thresholds. For example, if the candidate protein belongs to a specific protein class, e.g., an antibody, the threshold could be an antibodyspecific threshold. In some aspects, the threshold can be protein-dependent threshold based on the distribution of epitope values. In some aspects, the computer-implemented methods of immunogenicity prediction disclosed herein can provide a combined output that integrates the prediction of enhancing and suppressing epitopes. In some aspects, the computer-implemented methods of immunogenicity prediction disclosed herein can provide an evaluation of immunogenicity based on HLA profiles.

[0252] In some aspects, the training datasets used in the computer-implemented methods of immunogenicity prediction disclosed herein can comprise (i) CD4 epitopes, (ii) CD8 epitopes, (iii) CD4 and CD8 epitopes, (iv) epitopes clustered according to a narrow specific length, (v) epitopes with broad lengths, (vi) full-length protein training data to identify immunogenic sequences without relying on HLA binding data, (vii) proteins or epitopes belonging to specific allergen classes, (viii) proteins or epitopes linked to specific immune responses or inflammatory responses, or (ix) protein or epitopes linked to the release of specific cytokines.VIII. Compositions of Matter

[0253] The present disclosure also provides compositions of matter produced using the methods disclosed herein. In some aspects, the present disclosure provides an immunogenicity-optimized protein sequence generated using the computer-implemented methods of immunogenicity prediction disclosed herein. In some aspects, the present disclosure provides an immunogenic peptide identified based on a report generated by using the computer-implemented methods of immunogenicity prediction disclosed herein. In some aspects, the present disclosure provides vaccine protein sequence generated by using the computer-implemented methods of immunogenicity prediction disclosed herein. In some aspects, the present disclosure provides a vaccine protein generated by using the computer-implemented methods of immunogenicity prediction disclosed herein. In some aspects, the present disclosure provides a polynucleotide sequence, e.g., an mRNA encoding a vaccine protein sequence generated by using the computer-implemented methods of immunogenicity prediction disclosed herein.

[0254] In some aspects, the present disclosure provides an amino acid sequence with desired immunogenic characteristics identified using the computer-implemented methods of immunogenicity prediction disclosed herein. In some aspects, the present disclosure provides a protein having an amino acid sequence with desired immunogenic characteristics identified using the computer-implemented methods of immunogenicity prediction disclosed herein. In some aspects, the present disclosure provides a polynucleotide sequence encoding a protein having an amino acid sequence with desired immunogenic characteristics identified using the computer-implemented methods of immunogenicity prediction disclosed herein. In some aspects, the present disclosure provides a vector comprising a polynucleotide sequence encoding a protein having an amino acid sequence with desired immunogenic characteristics identified using the computer-implemented methods of immunogenicity prediction disclosed herein.

[0255] In some aspects, the present disclosure provides a cell comprising a vector comprising a polynucleotide sequence encoding a protein having an amino acid sequence with desired immunogenic characteristics identified using the computer-implemented methods of immunogenicity prediction disclosed herein. In some aspects, the present disclosure provides a cell comprising a polynucleotide sequence encoding a protein having an amino acid sequence with desired immunogenic characteristics identified using the computer-implemented methods of immunogenicity prediction disclosed herein. In some aspects, the present disclosure provides a cell comprising a protein, e.g., a recombinant protein, having an amino acid sequence with desired immunogenic characteristics identified using the computer-implemented methods of immunogenicity prediction disclosed herein. In some aspects, the cell is a host cell. In some aspects, the cell is a T cell, a stem cell, a B cell, a NK cell, a dendritic cell, a bacterial cell, a mammalian cell, a yeast cell, or an insect cell.

[0256] In some aspects, the present disclosure provides a pharmaceutical compositions comprising a protein, polynucleotide, or vector, or cell disclosed above, and an excipient.

[0257] In some aspects, the present disclosure provides a delivery system comprising a protein, polynucleotide, or vector disclosed above, e.g., a lipid nanoparticle, wherein the protein, polynucleotide, or vector are encapsulated in the lipid nanoparticle.

[0258] Although the foregoing embodiments have been described in some detail by way of illustration and example for purposes of clarity of understanding, it is readily apparent to those of ordinary skill in the art in light of the above teachings that certain changes and modifications can be made thereto without departing from the spirit or scope of the appended claims.EXAMPLESExample 1Predicting Immunogenicity Based on IFNy Release

[0259] Immunogenicity data was collected from the IEDB database and the following search criteria were used: linear peptide, T cell Assay, MHC Class II, Host Organism: Human, any disease, and Interferon-gamma release. Next, data cleaning was performed using the given conditions. First, all epitopes with an incorrect parent molecule accession number or epitopes with missing starting and ending positions were discarded. Furthermore, immunogenic epitopes between the lengths of 13 amino acids and 21 amino acids were included. Non-immunogenic epitopes between the lengths of 13 amino acids and 21 amino acids were included as is. However,for non-immunogenic epitopes between the lengths of 26 amino acids and 41 amino acids, they were split in half and included as non-immunogenic to add some negative controls. For epitopes that were determined experimentally in multiple studies with differing results the following criteria was used to determine immunogenicity: If the epitope is determined to be immunogenic in more than 70% of the experiments, it is immunogenic else it was excluded. Similarly, if the epitope is determined to be non-immunogenic more than 70% of the experiments, it is non-immunogenic else it was excluded. The final preprocessed dataset was split into 3 sub datasets for model training, model validation and model testing using a ratio of 65-15-20% respectively as shown in FIG. 7.

[0260] The ESM2 model, which is a Protein Language Model along with a classifier, was used to train on the curated dataset to predict immunogenicity. To train the model only peptides which are experimentally verified to release interferon-gamma (IFNy) were selected from the dataset. The ESM2 model was used to extract the embeddings which were the input to the classifier or neural network which generates the final probability of a peptide being immunogenic.

[0261] To assess the performance of the model on antibody sequences, predictions were made on the heavy chain sequence of adalimumab using the method described along with the tool from IEDB and compared with the experimentally found immunogenic epitopes. To make the predictions using this method the sequence was split into overlapping 15-mer peptides with an overlap of 14 amino acids between the peptides. Individual peptides with a predicted immunogenicity value greater than 0.5 were classified as immunogenic. For prediction using IEDB their recommended method was used. Comparing the predictions for the heavy region shown in FIG. 2 the model correctly classifies regions containing the protein’s 2 experimentally verified epitopes as immunogenic whereas the method from IEDB only predicts 1 out of the 2 experimentally verified immunogenic epitope. The model corrected classified immunogenic regions on the adalimumab VH region (SEQ ID NO:1). The experimentally confirmed Epitope 1 and Epitope 2 correspond, respectively, to SEQ ID NO: 2 and 3. The IFNy Immunogenicity Model correctly identified two immunogenic Epitopes: an Epitope 1 of SEQ ID NO:4 and an Epitope 2 of SEQ ID NO: 5. In contrast, the IEDB Method only identified an immunogenic epitope of SEQ ID vNO: 6 that corresponded to Epitope 2.

[0262] The performance of the model was also evaluated for general protein sequences by making predictions for thrombopoietin (high immunogenicity) and follitropin-beta (low immunogenicity) as shown in FIGS. 4A and 4B. To make the predictions using this method the sequence was split into overlapping 15-mer peptides with an overlap of 14 amino acids betweenthe peptides. Individual peptides with a predicted immunogenicity value greater than 0.5 were classified as immunogenic. As shown in FIG. 4A the model predicts 4 different regions as being immunogenic for thrombopoietin and as shown in FIG. 4B the model predicts only 1 region as being immunogenic for follitropin-beta. This aligns with the experimental observations for these 2 proteins shown in FIG. 8 as thrombopoietin has a much higher immunogenicity potential as compared to follitropin-beta.Example 2Predicting Immunogenicity Based on a Combination of Cytokine Release Assays, Weighted to Maximize True Negatives

[0263] Immunogenicity data was collected from public sources and the following search criteria were used: linear peptide, T cell Assay, MHC Class II, Host Organism: Human, any disease and cytokine release. Next, data cleaning was performed using the given conditions. First, all epitopes with an incorrect parent molecule accession number or epitopes with missing starting and ending positions were discarded. Furthermore, immunogenic epitopes between the lengths of 13 amino acids and 21 amino acids were included. Non-immunogenic epitopes between the lengths of 13 amino acids and 21 amino acids were included as is. However, for non-immunogenic epitopes between the lengths of 26 amino acids and 41 amino acids, they were split in half and included as non-immunogenic to add some negative controls. For epitopes that were determined experimentally in multiple studies with differing results the following criteria was used to determine immunogenicity: If the epitope is determined to be immunogenic in more than 70% of the experiments, it is immunogenic else it was excluded. Similarly, if the epitope is determined to be non-immunogenic more than 70% of the experiments, it is non-immunogenic else it was excluded. After preprocessing of the data, the dataset included the following cytokine release and cytokine release related factors: IFNy release, proliferation, IL-5 release, cytotoxicity, activation, qualitative binding, IL-2 release, IL- 10 release, IL-4 release, and IL- 17 release. Furthermore, the preprocessed dataset was split into 3 sub datasets for model training, model validation and model testing using a ratio of 65-15-20% respectively as shown in FIG. 7.

[0264] The ESM2 model, which is a Protein Language Model along with a neural network classifier, was used to train on the curated dataset to predict immunogenicity. The ESM2 model was used to extract the embeddings which were the input to the classifier or neural network which generates the final probability of a peptide being immunogenic. During training, greater weight was given to peptides correctly predicted as true negatives.

[0265] To assess the performance of the model on full-length protein sequences, predictions were made on the heavy chain sequence of adalimumab using the method described along with the tool from IEDB and compared with the experimentally found immunogenic epitopes. To make the predictions using this method the sequence was split into overlapping 15-mer peptides with an overlap of 14 amino acids between the peptides. Individual peptides with a predicted immunogenicity value greater than 0.5 were classified as immunogenic. For prediction using IEDB their recommended method was used. Comparing the predictions for the heavy region shown in FIG. 3 the model correctly classifies 1 region containing one of the protein’s 2 experimentally verified immunogenic epitopes. Similar performance is observed for IEDB as it only predicts 1 out of the 2 experimentally verified immunogenic epitopes. The experimentally identified epitopes on the VH domain of adalimumab (SEQ ID NO: 1) are Epitope 1, corresponding to SEQ ID NO:2, and Epitope 2, corresponding to SEQ ID NO:3. The True Negative Model predicted one Epitope 1 of SEQ ID NO: 7, whereas the True Positive Model predicted an Epitope 1 of SEQ ID NO: 8 and an Epitope 2 of SEQ ID NO:8. The IEDB CD4 Episcore Method, in contrast, predicted a single Epitope 2 of SEQ ID NO: 6.

[0266] The performance of the model was also evaluated for general protein sequences by making predictions for thrombopoietin (high immunogenicity) and follitropin-beta (low immunogenicity) as shown in FIGS. 5A and 5B. To make the predictions using this method the sequence was split into overlapping 15-mer peptides with an overlap of 14 amino acids between the peptides. Individual peptides with a predicted immunogenicity value greater than 0.5 were classified as immunogenic. As shown in FIG. 5A the model predicts 4 different regions as being immunogenic for thrombopoietin and as shown in FIG.5B no immunogenic peptides are predicted for follitropin-beta. This aligns with the experimental observations for these 2 proteins shown in FIG 8. as thrombopoietin has a much higher immunogenicity potential as compared to follitropin-beta.Example 3Predicting Immunogenicity Based on a Combination of Cytokine Release Assays, Weighted to Maximize True Positives

[0267] Immunogenicity data was collected from public sources and the following search criteria were used: linear peptide, T cell Assay, MHC Class II, Host Organism: Human, any disease and cytokine release. Next, data cleaning was performed using the given conditions. First, all epitopes with an incorrect parent molecule accession number or epitopes with missing starting andending positions were discarded. Furthermore, immunogenic epitopes between the lengths of 13 amino acids and 21 amino acids were included. Non-immunogenic epitopes between the lengths of 13 amino acids and 21 amino acids were included as is. However, for non-immunogenic epitopes between the lengths of 26 amino acids and 41 amino acids, they were split in half and included as non-immunogenic to add some negative controls. For epitopes that were determined experimentally in multiple studies with differing results the following criteria was used to determine immunogenicity: If the epitope is determined to be immunogenic in more than 70% of the experiments, it is immunogenic else it was excluded. Similarly, if the epitope is determined to be non-immunogenic more than 70% of the experiments, it is non-immunogenic else it was excluded. After preprocessing of the data, the dataset included the following cytokine release and cytokine release-related factors: IFNy release, proliferation, IL-5 release, cytotoxicity, activation, qualitative binding, IL-2 release, IL- 10 release, IL-4 release, and IL- 17 release. Furthermore, the preprocessed dataset was split into 3 sub datasets for model training, model validation and model testing using a ratio of 65-15-20% respectively as shown in FIG. 7.

[0268] The ESM2 model, which is a Protein Language Model along with a neural network classifier, was used to train on the curated dataset to predict immunogenicity. The ESM2 model was used to extract the embeddings which were the input to the classifier or neural network which generates the final probability of a peptide being immunogenic. During training, greater weight was given to peptides correctly predicted as true positives.

[0269] To assess the performance of the model on full-length protein sequences, predictions were made on the heavy chain sequence of adalimumab using the method described along with the tool from IEDB and compared with the experimentally found immunogenic epitopes. To make the predictions using this method the sequence was split into overlapping 15-mer peptides with an overlap of 14 amino acids between the peptides. Individual peptides with a predicted immunogenicity value greater than 0.5 were classified as immunogenic. For prediction using IEDB their recommended method was used. Comparing the predictions for the heavy region shown in FIG. 3 the model correctly classifies 1 region containing one of the protein’s 2 experimentally verified immunogenic epitopes. Similar performance is observed for IEDB as it only predicts 1 out of the 2 experimentally verified immunogenic epitopes.

[0270] The performance of the model was also evaluated for general protein sequences by making predictions for thrombopoietin (high immunogenicity) and follitropin-beta (low immunogenicity) as shown in FIGS. 6A and 6B. To make the predictions using this method thesequence was split into overlapping 15-mer peptides with an overlap of 14 amino acids between the peptides. Individual peptides with a predicted immunogenicity value greater than 0.5 were classified as immunogenic. As shown in FIG. 6A the model predicts 4 different regions as being immunogenic for thrombopoietin and as shown in FIG. 6B the model predicts only 1 region as being immunogenic for follitropin-beta. This aligns with the experimental observations for these 2 proteins shown in FIG 8 as thrombopoietin has a much higher immunogenicity potential as compared to follitropin-beta.Example 4Experimentally Anchored, MHC -Binding-Agnostic Prediction of Immunogenicity and Cytokine Release

[0271] Accurate prediction of anti-drug antibody (ADA) responses remains a challenge in the development of biologies, as current tools primarily rely on major histocompatibility complex (MHC) binding affinity, which does not capture the full complexity of immune activation. We have developed EPITOP™, a protein language model trained on epitopes experimentally verified via cytokine release assays, which enables direct immunogenicity prediction. Combined with IMMUNOTOP™, a tool that aggregates EPITOP™ outputs into protein-level scores, our approach outperforms existing models (AUC = 0.72), offering a more biologically grounded, scalable solution for immunogenicity risk assessment. Beyond identifying immunogenic hotspots, the model can also serve as a surrogate to approximate IFN-y cytokine release in scenarios where experimental measurement is impractical - for example, in modalities prone to systemic cytokine spikes, such as CAR-T therapies, and T-cell and NK-cell engagers.Methods

[0272] Peptide sequence dataset: All immunogenicity data was retrieved from IEDB available as of September 12, 2023. The following search terms were used to collect relevant data: linear peptide, T cell Assay, MHC Class II, Host: Human and any disease. This search yielded a total of 155,193 peptides belonging to 4 assay response categories Positive-High, Positive-Intermediate, Positive-Low and Negative. All epitopes marked as either Positive-High, Positive-Intermediate or Positive-Low were classified as immunogenic. Furthermore, only epitopes experimentally verified for the release of Interferon Gamma (IFNy) were included as it is commonly associated with inflammatory immune responses. There was further filtering to excludeepitopes that had an incorrect parent molecule accession number or the epitope starting and ending position was missing.

[0273] Next, immunogenic epitopes that were between the length of 13 and 21 were included as is. Similarly, non-immunogenic epitopes between the length of 13 and 21 were included as is and non-immunogenic epitopes between the length of 26 and 41 were split in half and both epitopes were classified as non-immunogenic. This was done to improve the class imbalance between immunogenic and non-immunogenic epitopes.

[0274] Lastly, there epitopes determined experimentally in multiple studies were considered. Therefore, to determine the qualitative label for these epitopes the following criteria was used: Epitopes determined to be immunogenic in more than 70% of the experiments were deemed to have sufficient evidence to be classified as immunogenic and likewise epitopes determined to be non-immunogenic in more than 70% of the experiments were deemed to have sufficient evidence to be classified as non-immunogenic. This resulted in a final dataset which contained a total of 15108 epitopes, in which 6320 were non-immunogenic and 8788 were immunogenic as shown in TABLE 1.TABLE 1: Test metrics results of EPITOP™ on the test dataset curated from IEDB.AUC Fl Precision Recall F0.5 F20.72 0.66 0.68 0.66 0.67 0.66

[0275] Training, validation and test datasets: To train, validate and test IMMUNOTOP™, the data was shuffled randomly and split into 3 datasets. 65% of the data was randomly selected to create the training dataset while maintaining the class balance. Similarly, 15% were selected to create the validation dataset and 20% of the data was randomly selected to create an independent test set to benchmark the performance of the models.

[0276] Implementation and training of EPITOP™: EPITOP™ was developed to predict immunogenic epitopes in protein sequences based on the release of IFNy cytokine, by finetuning the ESM-2 protein language model developed by Meta Al. See Lin, Z. et al. Science 379, 1123— 1130 (2023).

[0277] To optimize the performance of the model Partial Fine-Tuning was used along with population based hyperparameter-tuning and a linear learning rate scheduler. Partial Fine-Tuning is a parameter-efficient fine-tuning technique in which a subset of the model parameters is updatedby freezing the updates to selected model layers. This helps the model adapt to new tasks without requiring extensive computational resources as well as is useful in instances with limited training data.

[0278] There were several ESM-2 models with varying parameter sizes available. ESM-2 with 8M parameters was selected for training. The selected model consisted of 6 encoder layers and in EPITOP™ these were used to generate the embeddings which served as an input to a feedforward network consisting of 2 dense layers that used ReLu as an activation function to avoid vanishing gradients. A SoftMax function was used to generate the output probability of a peptide being immunogenic.

[0279] EPITOP™ was finetuned using Partial Fine-Tuning with the first 4 encoder layers of the ESM-2 model being frozen and the parameters for the last 2 encoder layers were only updated along with the feedforward network. Additionally, a linear learning rate scheduler with warmup was used to prevent unstable gradients in the initial updates and to avoid missing the optimal solution. To determine the optimal hyperparameters such as learning rate, weight decay, batch size and epochs population-based training was used with Accuracy as the metric to monitor to determine the optimal hyperparameters.

[0280] The following parameters were determined to provide the best accuracy learning: rate 8 - 10’5, weight decay 10'3, batch size 32, number of epochs 8 and 20,000 warmup steps for the learning rate scheduler. In addition to this, Adam optimizer was used to adjust the weights and to minimize the cross-entropy loss function.

[0281] The custom model was trained on subsequences between the lengths of 13 and 21 amino acids. To run predictions on full-length protein sequences a sliding window approach was used with the subsequence length being fixed to 15 and a sliding by 1 amino acid. Therefore, the input to EPITOP™ was a full-length sequence, and the sliding window method was used to predict epitopes in the sequence.

[0282] Immunogenicity score dataset: Different proteins and approved clinical molecules were used as benchmarks were used as a reference to benchmark the performance of IMMUNOTOP™, a downstream tool that aggregates EPITOP™ outputs to produce an overall immunogenicity score per protein. To further expand the dataset, FDA reports for several biologies were examined to identify antibodies with clinically reported high immunogenic responses, as well as biologies with low immunogenic responses. The latter classification was based on information provided in the FDA-approved labels of trastuzumab, romosozumab, bevacizumab, pembrolizumab, olaratumab, adalimumab, dupilumab, and golimumab.

[0283] IMMUNOTOP™ implementation: IMMUNOTOP™ was developed to extend interpretability by providing a qualitative metric to predict the immunogenic potential for a whole protein sequence and it accomplished so by performing further statistical analysis on the predictions generated by EPITOP™.

[0284] The input for IMMUNOTOP™ was the protein sequence. This sequence was split into overlapping 15-mer peptides with an overlap of 1 amino acid. For the resulting set of peptides, the underlying model in EPITOP™ was used to generate the immunogenic probability scores using the output of the SoftMax layer in the feedforward network.

[0285] Next, the following statistical values were calculated for the set of probabilities: the probability value below which 12.5% of the data falls (referred herein to as 01), and the probability value above which 12.5% of data lies (referred herein to as 07). Finally, the IMMUNOTOP™ score was determined as exp (2 * •Results

[0286] EPITOP™ Performance Evaluation on Test Set: To evaluate the prediction performance of the underlying model in EPITOP™ after training the corresponding test dataset created used data from IEDB was used. To assess the performance the AUC, Fl, F0.5, F2, precision and recall (TABLE 1) were calculated along with the confusion matrix. These metrics were selected as they provided a comprehensive measure of the performance of the model across thresholds and provided insights about how well the model distinguishes different classes given the imbalanced dataset.

[0287] EPITOP™ achieved an AUC score of 0.72 on the test set (FIG. 9B), indicating a good ability to distinguish between non-immunogenic and immunogenic peptides. In addition to this, precision and recall values were 0.68 and 0.66 respectively, which showed that the model made reliable predictions for the Immunogenic class while maintaining a balanced sensitivity. The Fl score demonstrated a balanced tradeoff between precision and recall. Furthermore, the F0.5 and F2 scores indicated that the performance of the model remained stable when the prediction was weighted towards either precision or recall. This was further supported by the results shown in the confusion matrix (FIG. 9A) as it showed that EPITOP™ had a 72% accuracy in predicting non-immunogenic epitopes and a 62% accuracy in predicting immunogenic epitopes. These metrics indicated that EPITOP™ performed consistently and maintained a balanced performance despiteclass imbalance, thus making it suitable for applications where an accurate performance is needed along with confidence in identifying true immunogenic peptides.

[0288] Analysis of embeddings learned by the models: The outputs from the final encoder layer and the final classification layer were extracted to understand how the model internally represented the peptides using the embeddings and to visualize the formation of a classification boundary.

[0289] T-SNE was used to reduce the outputs of the extracted layers to two dimensions as it captured both the local and non-linear relationships in the data. Specifically, this helped in understanding how the embeddings could be fine-tuned for the task of determining immunogenicity. The embeddings were compared for the ESM model before finetuning and after fine-tuning.

[0290] The output of the final encoder layer in the model captured the most relevant features for predicting immunogenicity. The results shown in FIG. 10A showed that before training the model had very few immunogenic features that were encoded or learned by the model and as the model was trained it learned many more relevant features related to predicting immunogenicity as shown by the embedding space in FIG. 10B. This was further validated by the Visualization of the classification layer in FIG. 10C and in FIG. 10D which showed that the model before fine-tuning predicted very few immunogenic peptides and as the model was trained the decision boundary was more clearly established with a higher number of immunogenic peptides being predicted. Therefore, it was observed that the model upon training learned the embeddings related to both immunogenic and non-immunogenic epitopes and it also created a separation between the two classes.

[0291] EPITOP™ outperforms existing methods: Next, the performance of EPITOP™ was evaluated in comparison to current publicly available methods. The current benchmark for MHCII dependent immunogenicity epitope prediction is CD4Episcore, created by IEDB, which generates allele independent predictions by using statistical methods. EPITOP™ has a slightly higher AUC score in comparison to the method by IEDB.

[0292] To further compare performance between the 2 methods, experimentally verified epitopes were determined for Adalimumab. Next, both the methods were used to predict epitopes on both the variable heavy and variable light chain for Adalimumab. As shown in FIGS. 11A and 11B there are 2 experimentally determined immunogenic epitopes each on the variable heavy and variable light chain. To make the prediction, the variable region was extracted using the AbNumber package in Python. For EPITOP™, the sequences were split into overlapping 15-mer peptides withan overlap of 14 amino acids, and for the IEDB CD4Episcore tool, the IEDB recommended method was used.

[0293] Comparing the predictions for the variable heavy region of adalimumab (SEQ ID NO:1), as shown in FIG. 11A, EPITOP™ (sequence labeled “B”) successfully predicted both epitopes (boxed sequences) correctly whereas the IEDB method (sequence labeled “C”) predicted only 1 epitope (boxed sequence). It was also observed that both methods also predicted 1 extra epitope which was a false positive (extended sequence located N-terminally to the N-terminal epitope detected by EPITOP™ and the corresponding epitope detected by IEDB).

[0294] The key difference lied in the prediction for the variable light region of adalimumab (SEQ ID NO: 10) shown in FIG. 11B, as EPITOP™ successfully predicted the 2 experimentally verified epitopes (the epitopes indicated by boxed sequences in the sequence labelled “B”) whereas the IEDB method predicted 2 epitopes both of which were false positives (the epitopes indicated by boxed sequences in the sequence labelled “C”).

[0295] Based on these results EPITOP™ outperformed the existing method in predicting the experimental epitopes correctly as well as on the overall AUC score.

[0296] IMMUNOTOP™ ranks proteins immunogenic potential: IMMUNOTOP™ utilizes statistical methods to aggregate the outputs from EPITOP™ to enable ranking of proteins based on the overall immunogenic potential. The input to IMMUNOTOP™ is a protein sequence, and it outputs an allele independent immunogenicity score along with a scale in reference to benchmark proteins. As seen in FIG. 12, a majority of the antibodies clustered between the range of 5.4 to 8 indicating a lower immunogenic potential as compared to other proteins, this aligned with expectations of antibodies generally being less immunogenic. Next, within the antibodies both Adalimumab and Golimumab had a higher score compared to other antibodies, which also matched clinical observations as both these antibodies have high immunogenicity levels as reported by FDA. Therefore, IMMUNOTOP™ provided a method to rank proteins independent of MHC and Treg concepts (with latter one playing a crucial role in statistical approaches, see Mattei et al. mAbs 16(1), 2333729 (2024)) as IMMUNOTOP™ predicts cytokine release as a measure of immunogenicity, and provides a method to manage clinical expectations by ranking proteins relative to their overall immunogenic potential.INCORPORATION BY REFERENCE

[0297] The contents of all cited references (including literature references, patents, patent applications, and websites) that may be cited throughout this application are hereby expresslyincorporated by reference in their entirety for any purpose, as are the references cited therein, in the versions publicly available on the date the present application was filed Protein and nucleic acid sequences identified by database accession number and other information contained in the subject database entries (e.g., non-sequence related content in database entries corresponding to specific Genbank accession numbers) are incorporated by reference, and correspond to the corresponding database release publicly available on the date the present application was filed.EQUIVALENTS

[0298] While various specific aspects have been illustrated and described, the above specification is not restrictive. It will be appreciated that various changes can be made without departing from the spirit and scope of the invention(s). Many variations will become apparent to those skilled in the art upon review of this specification.

Claims

WHAT IS CLAIMED IS:

1. A computer-implemented method for predicting the immunogenicity of a candidate protein, comprising:(i) generating one or more training datasets derived from protein sequence data and associated in vitro cytokine release data from a population of training proteins; (ii) training one or more machine learning (ML) models using the one or more training datasets of (i), wherein the one or more ML models are capable of predicting the ability to induce the cytokine release and immunogenicity of the candidate protein based on the primary sequence of the candidate protein;(iii) preprocessing the primary sequence of the candidate protein generating a plurality of peptide subsequences derived from the primary sequence of the candidate protein;(iv) applying the one or more ML models to each of the plurality of peptide sequences to generate immunogenicity scores associated with each peptide subsequence; and, (v) generating an immunogenicity prediction for the candidate protein based on the immunogenicity scores associated with each peptide subsequence and / or generating a cytokine release prediction for the candidate protein based on the immunogenicity scores associated with each peptide subsequence.

2. A computer-implemented method for predicting the immunogenicity of a candidate protein, comprising:(i) training one or more ML models using the one or more training datasets wherein the one or more training datasets are derived from protein sequence data and associated in vitro cytokine rlease data from a population of training proteins, and wherein the one or more ML models are capable of predicting the immunogenicity of the candidate protein based on the primary sequence of the candidate protein; (ii) preprocessing the primary sequence of the candidate protein generating a plurality of peptide subsequences derived from the primary sequence of the candidate protein;(iii) applying the one or more ML models to each of the plurality of peptide sequences to generate immunogenicity scores associated with each peptide subsequence; and,(iv) generating an immunogenicity prediction for the candidate protein based on the immunogenicity scores associated with each peptide subsequence.

3. A compute-implemented method for predicting the immunogenicity of a candidate protein, comprising:(i) preprocessing the primary sequence of the candidate protein generating a plurality of peptide subsequences derived from the primary sequence of the candidate protein;(ii) applying the one or more pretrained ML models to each of the plurality of peptide sequences to generate immunogenicity scores associated with each peptide subsequence, wherein the one or more ML models are trained using the one or more training datasets wherein the one or more training datasets are derived from protein sequence data and associated cytokine release data from a population of training proteins, and wherein the one or more ML models are capable of predicting the immunogenicity of the candidate protein based on the primary sequence of the candidate protein; and,(iii) generating an immunogenicity prediction for the candidate protein based on the immunogenicity scores associated with each peptide subsequence.

4. The computer-implemented method of any one of claims 1 to 3, wherein the one or more training datasets(i) do not use protein-binding data, and / or(ii) do not use protein-binding affinity data, and / or(iii) do not use protein-binding affinity prediction data.

5. The computer-implemented method of any one of claims 1 to 4, wherein the computer-implemented method does not comprise a protein secondary or tertiary structure prediction.

6. The computer-implemented method of any one of claims 1 to 5, wherein the computer-implemented method does not comprise a humanness assessment.

7. The computer-implemented method of any one of claims 1 to 6, wherein generating one or more training datasets comprises(a) preprocessing at least one population of training proteins;(b) optionally performing data augmentation; and(c) labeling.

8. The computer-implemented method claim 7, wherein preprocessing comprises filling or removing missing data values, removing data redundancies, and handling data ambiguities associated with at least one population of training proteins.

9. The computer-implemented method of claim 7, wherein preprocessing the population of training proteins comprises (i) splitting the amino acid sequences of the training proteins into training peptide subsequences, and (ii) tokenizing the training peptide subsequences to generate embeddings.

10. The computer-implemented method of any one of claims 7 to 9, wherein data augmentation comprises splitting the amino acid sequence of a training protein into one or more shorter training proteins.

11. The computer-implemented method of any one of claims 7 to 10, wherein labeling comprises assigning a plurality of labels indicative of an immunogenicity or non-immunogenicity category of each training peptide subsequence derived from the at least one population of training proteins.

12. The computer-implemented method of any one of claims 7 to 11, wherein labeling comprises converting a multi-class label associated with a training peptide subsequence located within the sequence of a training protein of the one or more training datasets to a binary-class label.

13. The computer-implemented method according to claim 12, wherein the converting comprises assigning the binary-class label, as positive based on determining that an ICso value associated with the training peptide subsequence is larger than a limit threshold, or as negative based on determining that the ICso value associated with the training peptide subsequence is smaller than the limit threshold.

14. The computer-implemented method of any one of claims 9 to 13, wherein the training peptide subsequences are overlapping.

15. The computer-implemental method of any one of claims 1 to 14, wherein the population of training proteins comprises a set of immunogenic proteins.

16. The computer-implemented method of claim 10, wherein the population of training proteins comprises a plurality of protein subsequences from the set of immunogenic proteins and a plurality of data labels indicative of a categorized information of each protein sequence in the set of immunogenic proteins.

17. The computer-implemented method of any one of claims 1 to 16, wherein the population of training proteins comprises a set of non-immunogenic proteins.

18. The computer-implemented method of claim 17, wherein the population of training proteins comprises a plurality of protein subsequences from the set of non-immunogenic proteins and a plurality of data labels indicative of a categorized information of each protein sequence in the set of non-immunogenic proteins.

19. The computer-implemented method of any one of claims 1 to 18, wherein the plurality of peptide subsequences derived from the amino acid sequences of the population of training proteins are generated using a sliding window and a sliding step.

20. The computer-implemented method of claim 19, wherein the sliding window has a variable length.

21. The computer-implemented method of claim 20, wherein the sliding window has a constant length.

22. The computer-implemented method of claim 21, wherein the sliding window constant length is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 amino acids, or the full protein length.

23. The computer-implemented method of claim 21, wherein the sliding window constant length is 15 amino acids.

24. The computer-implemented method of any one of claims 19 to 23, wherein the sliding step has a variable length.

25. The computer-implemented method of any one of claims 19 to 23, wherein the sliding step has a constant length.

26. The computer-implemented method of claim 25, wherein the sliding step length is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acidsm or the full protein length.

27. The computer-implemented method of claim 26, wherein the sliding step length is 1 amino acid.

28. The computer-implemented method of any one of claims 1 to 27, wherein the plurality of peptide subsequences derived from the amino acid sequences of the population of training proteins are overlapping.

29. The computer-implemented method of any one of claims 1 to 27, wherein the plurality of peptide subsequences derived from the amino acid sequences of the population of training proteins are non-overlapping.

30. The computer-implemented method of any one of claims 1 to 29, wherein the plurality of peptide subsequences derived from the amino acid sequences of the population of training proteins have the same length.

31. The computer-implemented method of any one of claims 1 to 29, wherein the plurality of peptide subsequences derived from the amino acid sequences of the population of training proteins can have different lengths.

32. The computer-implemented method of any one of claims 19 to 31, wherein the sliding window has a length of 15 amino acids, and the sliding step has a length of 1 amino acid.

33. The computer-implemented method of any one of claim 1 to 32, wherein the training of comprises fine-tuning a pre-trained LLM and / or fine-tuning a pre-trained LSTM.

34. The computer-implemented method of any one of claim 1 to 32, wherein the protein sequence data comprises only primary sequence data and cytokine release data.

35. The computer-implemented method of claim 34, wherein the protein sequence data further comprised protein secondary structure data and / or protein tertiary structure data.

36. The computer-implemented method of any one of claims 7 to 35, wherein the preprocessing of at least one population of training proteins protein comprises iterating the amino acid sequence of each training protein sequence using a sliding window and sliding step to generate a plurality of peptide subsequences.

37. The computer-implemented method of claim 36, wherein preprocessing further comprises tokenizing the plurality of peptide subsequences derived from the primary sequence of a training protein.

38. The computer-implemented method of claim 27, wherein tokenizing a peptide subsequence derived from the primary sequence of a training protein comprises generating a numerical sequence associated with the peptide subsequence with a format applicable to the one or more ML models.

39. The computer-implemented method of any one of claims 37 or 38, wherein tokenizing comprises using a pre-defined tokenizer to encode the amino acids within the peptide subsequence.

40. The computer-implemented method of any one of claims 37 to 39, wherein tokenizing coverts each amino acid to a numeric and / or vector token.

41. The computer-implemented method of any one of claims 1 to 40, wherein class weights in the loss function are scaled to account for class imbalance or class weights in the loss function are not scaled to account for class imbalance.

42. The computer-implemented method of any one of claims 1 to 41, wherein the ML model is a pretrained model or a newly trained model.

43. The computer-implemented method of any one of claims 1 to 42, wherein the ML model is selected from the group consisting of Artificial Neural Network (ANN) model, Convolutional Neural Network (CNN) model, Recurrent Neural Networks (RNN) model, large language model (LLM), or a combination thereof.

44. The computer-implemented method of any one of claims 1 to 43, wherein the ML model comprises an RNN.

45. The computer-implemented method of claim 44, wherein the RNN is a Long short-term memory (LSTM) model.

46. The computer-implemented method of any one of claims 1 to 45, wherein the ML model comprises an LLM.

47. The computer-implemented method of claim 46, wherein the LLM is an Evolutionary Scale Modelling 2 (ESM-2) model.

48. The computer-implemented method of claim 47, wherein the LLM is a fine-tuned ESM-2 model.

49. The computer-implemented method of claim 48, wherein fine-tuning of the fine-tuned ESM-2 model comprises further training a plurality of layers using the training data wherein the weights of the remaining layers are not updated / trained.

50. The computer-implemented method of any one of claims 1 to 49, wherein the ML model further comprises a classification layer.

51. The computer-implemented method of claim 50, wherein the binary classification layer outputs the probability of immunogenicity of a peptide.

52. The computer-implemented method of claim 50, wherein the binary classification layer outputs the probability of non-immunogenicity of a peptide.

53. The computer-implemented method of any one of claims 1 to 52, wherein the plurality of peptide subsequences derived from the candidate protein sequence are generated using a sliding window and a sliding step.

54. The computer-implemented method of claim 53, wherein the sliding window has a variable length.

55. The computer-implemented method of claim 53, wherein the sliding window has a constant length.

56. The computer-implemented method of claim 55, wherein the sliding window constant length is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 amino acids, or the entire protein.

57. The computer-implemented method of claim 55, wherein the sliding window constant length is 15 amino acids.

58. The computer-implemented method of any one of claims 53 to 57, wherein the sliding step has a variable length.

59. The computer-implemented method of any one of claims 54 to 58, wherein the sliding step has a constant length.

60. The computer-implemented method of claim 59, wherein the sliding step length is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids.

61. The computer-implemented method of claim 59, wherein the sliding step length is 1 amino acid.

62. The computer-implemented method of any one of claims 1 to 61, wherein the plurality of peptide subsequences derived from the candidate protein sequence are overlapping.

63. The computer-implemented method of any one of claims 1 to 61, wherein the plurality of peptide subsequences derived from the candidate protein sequence are non-overlapping.

64. The computer-implemented method of any one of claims 1 to 63, wherein the plurality of peptide subsequences derived from the candidate protein sequence have the same length.

65. The computer-implemented method of any one of claims 1 to 63, wherein the plurality of peptide subsequences derived from the candidate protein sequence can have different lengths.

66. The computer-implemented method of any one of claims 53 to 65, wherein the sliding window has a length of 15 amino acids, and the sliding step has a length of 1 amino acid.

67. The computer-implemented method of any one of claims 1 to 66, wherein the immunogenicity prediction comprises a binary prediction of cytokine release induced by the candidate protein or a fragment thereof in a subject.

68. The computer-implemented method of any one of claims 1 to 67, wherein the immunogenicity prediction comprises a probability of cytokine release induced by the candidate protein or a fragment thereof in a subject.

69. The computer-implemented method of claim 67 or claim 68, wherein cytokine release comprises the release a cytokine selected from the group consisting of IL-1, IL-2, IL-3, IL-4, IL-6, IL-8, IL-10, IL-13, IL-15, TNFalpha, TGFbeta, CD40L, IFNalpha, IFNbeta, IFNgamma, and any combination thereof.

70. The computer-implemented method of any one of claims 1 to 69, wherein the predicted cytokine release comprises predicted severity of the cytokine release.

71. The computer-implemented method of any one of claims 1 to 70, wherein generating an immunogenicity prediction for the candidate protein based comprises identifying one or more peptide subsequences having immunogenicity probabilities above a predetermined threshold.

72. The computer-implemented method of any one of claims 1 to 71, wherein generating an immunogenicity prediction for the candidate protein comprises identifying one or more immunogenicity hot spots having immunogenicity probabilities above a predetermined threshold.

73. The computer-implemented method of claim 72, wherein generating an immunogenicity prediction for the candidate protein comprises(i) determining a global candidate protein immunogenicity score,(ii) determining number of immunogenicity hotspots on the candidate protein, (iii) determining the distance between immunogenicity hotspots on the candidate protein,(iv) determining median of the lowest quantile (MQ1),(v) determining median of the top quantile (MQ3),(vi) determining distance of lowest to top quantile (MQ3-MQ1),(vii) determining dispersion of lowest and top quantile (MQ3-MQ1) / MQ3 ),(viii) determining the length of the immunogenicity hotspots on the candidate protein, or (ix) a combination thereof.

74. The computer-implemented method of claim 73, wherein determining the global candidate protein immunogenicity score comprises integrating the immunogenicity probabilities of each of the peptide subsequences in the candidate protein.

75. The computer-implemented method of claim 73, wherein determining the global candidate protein immunogenicity score comprises calculating a Z-score.

76. The computer-implemented method of any one of claims 1 to 75, wherein the immunogenicity prediction is outputted as a graphic representation showing the immunogenicity probabilities of each of the peptide subsequences in the candidate protein along the sequence of the candidate protein, wherein immunogenic regions or immunogenic hotspots correspond tosegments of the candidate protein sequence having the immunogenicity probabilities above a predetermined threshold, such as 0.5 (50% probability) or MQ3.

77. A method to determine the immunogenicity of a candidate protein comprising applying the computer-implemented method of any one of claims 1 to 76 to the candidate protein.

78. The method of claim 77, wherein the candidate protein is an immunogenic protein.

79. The method of claim 77, wherein the candidate protein is a non-immunogenic protein.

80. A method to identify immunogenic hot spots in a candidate protein comprising applying the computer-implemented method of any one of claims 1 to 79 to the candidate protein.

81. The method of claim 80, wherein the method comprises identifying:(i) number of immunogenic hot spots;(ii) location of the immunogenic hot spots;(iii) length of the immunogenic hot spots;(iv) degree of immunogenicity of the hots spots;(v) median of the lowest quantile (MQ1),(vi) median of the top quantile (MQ3),(vii) distance of lowest to top quantile (MQ3-MQ1),(viii) dispersion of lowest and top quantile (MQ3-MQ1) / MQ3 );(ix) distance between immunogenic hot spots;(x) specific properties of the immunogenic hot spots; or,(xi) any combination thereof.

82. The method of claim 81, wherein the specific properties of the hotspots comprise:(i) charge;(ii) polarity;(iii) hydrophobicity;(iv) amino acid composition;(v) surface / buried location; or,(vi) any combination thereof.

83. A method of predicting the immunogenicity of a peptide or protein comprising:(i) inputting the amino acid sequence of the peptide or protein into a computer system comprising an immunogenicity prediction program comprising instructions for executing the computer-implemented method of any one of claims 1 to 76, wherein the peptide or protein is the candidate protein;(ii) executing in the computer system the computer-implemented method of any one of claims 1 to 76; and,(iii) receiving from the computer system the immunogenicity prediction resulting from executing the computer-implemented method of any one of claims 1 to 76.

84. The method of any one of claims 1 to 83, wherein the code to execute the computer-implemented method of any one of claims 1 to 76, the training database of the computer-implemented method of any one of claims 1 to 76, the ML model of the computer-implemented method of any one of claims 1 to 76, the output of the ML model of the computer-implemented method of any one of claims 1 to 76, the immunogenicity prediction outputted by the computer-implemented method of any one of claims 1 to 76, or any combination thereof is hosted in a cloud computing environment.

85. A non-transitory computer readable storage medium storing a computational module for predicting the immunogenicity of a candidate protein, the computational module comprising code to execute the computer-implemented method of any one of claims 1 to 76 or a portion thereof.

86. A computer readable storage medium having computer readable instructions to instruct a computer to perform the computer-implemented method of any one of claims 1 to 76 or a portion thereof.

87. A system for predicting the immunogenicity of a candidate protein according to the computer-implemented method of any one of claims 1 to 76, wherein the system is stored on computer-readable storage media.

88. A computer system for predicting the immunogenicity of a candidate protein, the computer system comprising at least one processor and memory storing at least one program for executionby the at least one processor, the at least one program comprising instructions to execute the computer-implemented method of any one of claims 1 to 76.

89. A method to engineer a therapeutic protein comprising(i) predicting the immunogenicity of the therapeutic protein by applying the computer- implemented method of any one of claims 1 to 76 to the therapeutic protein, wherein the therapeutic protein is the candidate protein;(ii) introducing one or more amino acid modifications to the therapeutic protein of step (i) to generate a modified therapeutic protein;(iii) predicting the immunogenicity of the modified protein by applying the computer- implemented method of any one of claims 1 to 76 to the modified therapeutic protein, wherein the modified therapeutic protein is the candidate protein;(iv) comparing the predicted immunogenicity of the therapeutic protein of (i) and the modified therapeutic protein of (iii);(v) optionally iterating (ii)-(iv) until the modified therapeutic protein has the desired degree of immunogenicity or lack thereof.

90. The method of claim 89, wherein the amino acid modifications are selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof.

91. The method of claim 90, wherein the amino acid substitutions are conservative substitutions, non-conservative substitutions, or a combination thereof.

92. A method to screen in silico a library of candidate proteins to determine their predicted immunogenicity comprising applying the computer-implemented method of any one of claims 1 to 76 to the amino acid sequences of the library of candidate proteins.

93. A method of de-epitoping a therapeutic candidate protein comprising(i) applying the computer-implemented method of any one of claims 1 to 76 to the therapeutic candidate protein or a modified variant thereof,(ii) applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, proteintruncations, altering a post-translational modification, or any combination thereof to the therapeutic candidate protein,(iii) applying the computer-implemented method of any one of claims 1 to 76 to the therapeutic candidate protein or a modified variant thereof,(iv) iterating steps (i) to (iii) to obtain a modified variant of the therapeutic candidate protein containing less antigenic epitopes than the therapeutic candidate protein.

94. A method to rescue a protein or peptide drug that triggers negative immunogenic effects by applying the method of any one of claims 89 to 93.

95. A method to predict negative responses to a drug comprising or consisting of a protein, the method comprising applying the computer-implemented method of any one of claims 1 to 76 to the drug.

96. The method of claim 95, wherein the protein drug is used or is a candidate for use in a clinical trial.

97. A method of vaccine design comprising applying the computer-implemented method of any one of claims 1 to 76 to a vaccine candidate protein.

98. The method of claim 97, wherein vaccine design comprises selecting the vaccine candidate protein if the predicted immunogenicity of the vaccine candidate protein is above a predetermined threshold value.

99. The method of claims 97 or 98, wherein vaccine design comprises modifying the vaccine candidate protein by applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof to the vaccine candidate protein, and iteratively executing the computer-implemented method of any one of claims 1 to 76.

100. A method of modulating the immunogenicity of a candidate protein comprising applying the computer-implemented method of any one of claims 1 to 76 to the candidate protein.

101. The method of claim 100, wherein modulating comprises increasing the immunogenicity of the candidate protein by applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof to the candidate protein.

102. The method of claim 100, wherein modulating comprises decreasing the immunogenicity of the candidate protein by applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof to the candidate protein.

103. The method of claim 100, wherein modulating comprises introducing at least one immunogenic hot spot on the protein by applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof to the candidate protein.

104. The method of claim 102, wherein modulating comprises removing at least one immunogenic hot spot from the protein by applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification, or any combination thereof to the candidate protein.

105. A method of identifying an immunogenic epitope susceptible to antibody targeting in a candidate protein comprising applying the computer-implemented method of any one of claims 1 to 76 to the candidate protein.

106. A method of generating an amino acid sequence predicted to have altered immunogenicity compared to the amino acid sequence a parent candidate protein comprising applying one or more amino acid modifications selected from the group consisting of amino acid substitutions, amino acid insertions, amino acid deletions, protein truncations, altering a post-translational modification,or any combination thereof to the amino acid sequence of parent candidate protein to generate the amino acid sequence of the modified candidate protein and determining the predicted immunogenicity of the modified candidate protein by applying the computer-implemented method of any one of claims 1 to 76 to the amino acid sequence of modified candidate protein.

107. The computer-implemented method of any one of claims 1 to 76, wherein the candidate protein is selected from the group consisting of a vaccine immunogen, an antibody, MSTAR, scFab, Fab, scFv, Fv, Fc, an enzyme, a growth factor, a chimeric antigen receptor (CAR), a T-cell receptor (TCR), a cytokine, chimeric cytokine or a cytokine receptor, GLP1, GLP2, and their close receptor and close analogues; and soluble and non-soluble forms of receptors mentioned above.

108. The computer-implemented method of any one of claims 1 to 76, further comprising evaluating the strength of interactions between (i) the candidate protein in a complex with the MHCI protein, (ii) the candidate protein in a complex with the MHCII protein, (iii) the candidate protein in a complex with the MHCI and MHCII proteins, (iv) the candidate protein and a T-cell receptor, or (v) a combination thereof.

109. The computer-implemented method of any one of claims 1 to 76 or 108, further comprising evaluating the strength of interactions between (i) the candidate protein and (ii) the MHCII protein.

110. The computer-implemented method of any one of claims 1 to 76, 108 or 109, further comprising evaluating the strength of interactions between (i) the candidate protein and (ii) the protein-binding cavity of a complex comprising a MHCII protein and a T cell receptor.

111. The computer-implemented method of any one of claims 1 to 76, 108, 109 or 110, further comprising inputting sequence information of the MHCII protein and / or the T cell receptor.

112. A method to predict the immunogenicity of a plurality of fragments of a protein associated with cancer, and identifying a fragment of said protein that is predicted to be immunostimulatory comprising applying the computer-implemented method of any one of claims 1 to 76 to the plurality of fragments.

113. A method of producing a personalized cancer vaccine for a subject having a tumor, the method comprising the steps of (i) identifying a plurality of modified peptides expressed in the tumor, each comprising an amino acid substitution at a position, relative to a corresponding parent peptide expressed in the normal cells; (ii) determine the immunogenicity for each of the plurality of modified peptides, via the computer-implemented method of any one of claims 1 to 76, wherein modified peptide is a candidate protein; and (iii) producing a personalized cancer vaccine for the subject, which comprises a peptide or polypeptide comprising the at least one modified peptide selected as immunogenic.

114. A method of generating a library of immunogenic proteins and their predicted immunogenicity, wherein each immunogenic protein and / or their predicted immunogenicity is generated via the computer-implemented method of any one of claims 1 to 76.

115. A method of selecting of therapeutic protein comprising applying the computer-implemented method of any one of claims 1 to 76 to a population of candidate proteins.

116. An immunogenicity-optimized protein sequence generated using the computer-implemented method of any one of claims 1 to 76.

117. An immunogenic peptide identified based on a report generated by using the computer-implemented method of any one of claims 1 to 76.

118. A vaccine protein sequence generated by using the computer-implemented method of any one of claims 1 to 76.

119. An amino acid sequence with desired immunogenic characteristics identified using the computer-implemented method of any one of claims 1 to 76.

120. A protein having the amino acid sequence of claim 119.

121. A polynucleotide encoding the protein of claim 120.

122. A vector comprising the polynucleotide of claim 121.

123. A cell comprising the polynucleotide of claim 121 or the vector of claim 122.

124. The cell of claim 123, wherein the cell is a T cell, a stem cell, a B cell, a NK cell, a bacterial cell, a dendritic cell, a mammalian cell, a yeast cell, or an insect cell.

125. A pharmaceutical composition comprising the protein of claim 120, polynucleotide of claim 121, vector of claim 122, or the cell of claims 123 or 124, and an excipient.

126. A delivery system comprising the protein of claim 120, polynucleotide of claim 121, or vector of claim 122.

127. The delivery system of claim 126, comprising a lipid nanoparticle, wherein the nanoparticle encapsulated the protein of claim 120, polynucleotide of claim 121, or vector of claim 122.

128. A method to treat or prevent a disease or condition comprising administering the protein of claim 120, polynucleotide of claim 121, vector of claim 122, cell of claims 123 or 124, pharmaceutical composition of claim 125, or delivery system of claim 126 or 127 to a subject.

129. The method claim 128, wherein the disease or condition is selected from the group consisting of cancer, infection, chronic inflammation, genetic disease, and autoimmune disease.

130. The method of claim 128, wherein the subject is a human subject.