Systems and methods for machine learning model-based prediction of immune response

Machine learning models simulate immune responses to predict effective antigen targets, addressing the challenges of molecular heterogeneity in cancer and infectious diseases, thereby enhancing the development of personalized treatments.

WO2025257735A1PCT designated stage Publication Date: 2025-12-18BIONTECH SE

Patent Information

Application Number
PCT/IB2025/055939
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-05-08
Filing Date
2025-06-10
Publication Date
2025-12-18

AI Technical Summary

Technical Problem

Current cancer immunotherapy and infectious disease treatments face challenges due to the molecular heterogeneity of tumors and viruses, necessitating extensive and costly trial-and-error approaches to identify effective molecular targets for personalized treatments.

Method used

Utilizing machine learning models, particularly large language models (LLMs), to simulate immune system behavior and predict immune responses to candidate antigens, enabling the identification of effective molecular targets for cancer immunotherapies and vaccines through in-silico immune response prediction.

Benefits of technology

This approach reduces the time and cost of developing personalized treatments by accurately identifying antigen targets, facilitating the creation of effective cancer vaccines and infectious disease vaccines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025055939_18122025_PF_FP_ABST
    Figure IB2025055939_18122025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure, among other things, provides technologies for predicting immune response (e.g., in response to an antigen) and for utilizing these predictions to create personalized cancer vaccines and viral vaccines. Among other things, systems, methods, and architectures describe a use of large language models (LLMs) and / or multi-task models for immune response predictions. The present disclosure provides models that simultaneously predict multiple immune response tasks and / or benefit from training on multiple immune response tasks. The present disclosure also provides tools for identifying antigen fragments that will, among other things, be presented as ligands for T-cell recognition.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR MACHINE LEARNING MODEL-BASED PREDICTION OF IMMUNE RESPONSECROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and benefit from U.S. Provisional 63 / 658,210, filed June 10, 2024, and U.S. Provisional Application No. 63 / 802,233, filed May 8, 2025, the content of each of which is incorporated by reference herein in its entirety.BACKGROUND OF THE INVENTION

[0002] Immunotherapy is a promising avenue for cancer treatment that leverages the efficacy and specificity of a patient’s own immune system to identify and eliminate cancer cells in a targeted fashion. Yet, despite recent advances in cancer immunotherapy, underlying and sometimes fundamental features of cancer biology, such as the molecular heterogeneity of tumors, present significant challenges to development of effective and broadly applicable treatments. Moreover, owing to the biological complexity of the overall host immune response that immunotherapies aim to harness, identifying and developing treatments that will be successful in particular patient classes (for example, particular indications and genetic backgrounds) is a time consuming and costly endeavor, often involving extensive trial and error. These considerations, and the challenges that they present and costs they impose, are magnified further where personalized treatments, tailored to individual patients and their unique tumor biology and molecular profile, are the goal.

[0003] Treatment and prevention of infectious diseases, such as viral infections like SARS-CoV-2, influenza, respiratory syncytial virus (RSV), and the like, present similar challenges. Among other things, high mutation rates of viruses mean molecular targets of treatments, vaccines, and the like are often highly heterogeneous and vary over time and, as with cancer immunotherapy, vaccination efforts against infectious diseases rely on, and must account for the complexities of, host immune response.SUMMARY

[0004] The present disclosure, among other things, provides technologies for identifying and / or characterizing candidate antigens or antigen portions (e.g., antigen epitopecandidates) in-silico. In particular, in certain embodiments, systems, methods, and architectures of the present disclosure utilize machine learning models, such as language models (LLMs) that operate on and / or generate, as output, biological sequences of candidate antigens or portions thereof (e.g., epitopes). Among other things, machine learning models described herein may be trained to replicate immune system behavior, acting as a digital twin of a host immune system that performs functions such as evaluating, selecting, and generating sequences of candidate antigens or portions thereof based on their predicted ability to trigger a useful immune response. In this manner, in-silico immune response prediction technologies of the present disclosure can be used to identify (e.g., particular antigens or portions thereof) effective molecular targets for disease treatment and / or prevention, for example, for use in creation of cancer immunotherapies such as personalized cancer vaccines, vaccines for infectious diseases, and the like.

[0005] In particular, in certain embodiments, immune response prediction technologies of the present disclosure utilize machine learning techniques to model complex interactions between various steps in the biological pathway(s) that form a host’s immune response to a particular disease, such as a virus or tumor. In certain embodiments, particular steps in the immune response are designated as (e.g., distinct / discrete) tasks (immune response tasks) to be performed by a machine learning model. Each task may be associated with particular inputs and outputs. In certain embodiments, each task may be associated with particular training example datasets.

[0006] For example, antigen presentation plays a key role in immune response to an antigen, in particular triggering T- and B-cell responses. In the context of intracellular antigens, an antigen presentation process may include steps such as proteasomal cleavage, peptide-MHC binding, and surface presentation. As described in further detail herein, each of these steps may be formulated as an immune response task, to be modeled by a machine learning model. For example, in a proteasomal cleavage task, a machine learning model may evaluate candidate peptides and output scores (e.g., numerical values) that represent a likelihood and / or efficacy with which a given candidate peptide is cleaved by a proteasome. In a peptide-MHC binding task, a machine learning model may be used to predict binding affinity between candidate peptides and one or more MHC molecules. In certain embodiments, in addition to a representation (e.g., a biological sequence) of a given candidate peptide, a machine learning model may also receive, and base binding affinity predictions on, representations of particular MHC molecules, such as an identification orsequence of a particular HLA allele, thereby accounting for varying genetic makeup of patients. In certain embodiments, a machine learning model may be used to model a surface presentation task, whereby a particular peptide and MHC molecule are scored to measure an extent to which, upon forming a peptide-MHC complex, the particular peptide will be successfully presented at a cell surface. In certain embodiments, scores determined for each immune response task may be combined, for example, in a serial fashion, treating multiple steps in an antigen presentation biological pathway as a serial process, to identify candidate ligands that will be presented for T-cell regulation.

[0007] Additionally or alternatively, in certain embodiments, machine learning approaches may be used to model an entire immune response cascade end-to-end, for example, using a single machine learning model. For example, a machine learning model may receive, e.g., as input, a representation of an MHC molecule (e.g., a MHC-I allele) and a representation of an antigen sequence, such as an amino acid sequence of a particular protein or sub-unit thereof, and generate output representing a prediction of which portions of that protein will be cleaved into peptides and ultimately recognized as ligands by a host immune system, and, in turn, provoke an immune response. A machine learning model may, for example, score each amino acid or groups of amino acids in a protein sequence to reflect a likelihood of their being part of a presented ligand, or, additionally or alternatively, may output one or more sequences, representing those portion(s) of the protein that are predicted to ultimately be presented as ligands, e.g., via MHC molecules, for T-cell recognition.

[0008] Machine learning models utilized by immune response prediction technologies described herein may be trained to perform one or more of the above-described tasks, as well as other immune response prediction tasks. In certain embodiments, a single model may be trained to perform multiple tasks, for example, each task in the above-described antigen presentation pathway, as well as, additionally or alternatively, other tasks, such as direct ligand identification. Without wishing to be bound to any particular theory, this “multi-task” training approach, whereby a single machine learning model is trained to perform multiple distinct (e.g., but related) tasks, is believed to benefit from transfer learning between tasks and, as such, may allow for improved performance in comparison with dedicated, “taskspecific” models trained on one single task. Among other things, these benefits are believed to be particularly significant where data relating to a particular task is scarce. Accordingly, rather than limiting training to a small number of examples associated with that particulartask, a multi-task model can be trained using, and benefit from, additional examples from different but related tasks.

[0009] In certain embodiments, systems and methods of the present disclosure utilize language models (e.g., large language models (LLMs)) that receive textual representations of biological sequences - such as protein antigen sequences, candidate peptide sequences, HLA molecule sequences, as well as other data inputs - as input, and generate textual output. Among other things, this uniform textual representation approach employed by language models facilitates multi-task training and is well-suited to sequence data that can be used to represent (e.g., amino acid sequences of) proteins and / or peptides. Modeling immune responses, however, presents a challenge for language models, since predicted results of biological interactions, such as proteasomal cleavage likelihoods or efficacies, binding affinities, surface presentation efficacy, often are naturally captured as continuous variables, for example floating point number representing probabilities, relative rates, or real-world measurements, such as dissociation constants or other measures of affinity. Accordingly, to address this challenge, the present disclosure includes techniques for harmonizing input and outputs including biological sequences, categorical variables representing, for example, HLA allele types, and continuous numerical values, in textual form, thereby allowing for a single text-to-text language model to be trained and used to generate a wide range of predictions modeling expected immune response to various candidate antigens.

[0010] Accordingly, by providing technologies for accurately identifying and evaluating valuable antigen targets in-silico, methods and systems described herein can dramatically reduce the burden of extensive trial and error experimentation, allowing for improvements in efficacy with reduced costs and time to development.

[0011] In some aspects, the present disclosure provides methods for predicting an immune response to a prospective antigen (e.g., an antigen polypeptide; e.g., a source protein and / or portion thereof, such as a particular sub-unit) or portion thereof (e.g., a prospective epitope; e.g., a peptide) using a language model, in which provided methods comprise: (a) obtaining (e.g., receiving, accessing, generating, etc.), by a processor of a computing device, an input text string comprising a candidate peptide sequence representing a biological sequence (e.g., an amino acid sequence; e.g., a nucleic acid sequence) of at least a portion (e.g., a particular peptide corresponding to a prospective epitope) of the prospective antigen; (b) determining, by the processor, using the language model, an immune response prediction based on the candidate peptide sequence, wherein the language model receives, as input, theinput text string and generates, as output, the immune response prediction as an output text string encoding values of one or more immune response scores, each immune response score associated with and representing a predicted result of a particular immune response task (e.g., a particular step in a biological pathway of a host immune response to the prospective antigen); and (c) storing and / or providing, by the processor, the immune response prediction for display and / or further processing.

[0012] In certain embodiments, a prospective antigen is a neoantigen of a subject [e.g., a human protein resulting from one or more cancer-related somatic mutations (e.g., single nucleotide variants, insertions and deletions, etc.).

[0013] In certain embodiments, a prospective antigen is a shared tumor antigen [e.g., wherein the prospective antigen is or comprises a protein resulting from mutations that are associated with cancer and shared (e.g., having a high prevalence) amongst a population of individuals; e.g., mutations in stability genes, mutations linked to hereditary risk of cancer, mutations linked to particular cancers, such as breast cancer, colorectal cancer, endometrial cancer, ovarian cancer, gastric cancer, pancreatic cancer, melanoma, lung cancer, prostate cancer, etc.].

[0014] In certain embodiments, a candidate peptide sequence is an amino acid sequence encoded by a gene comprising one or more identified patient-specific tumor mutations (e.g., identified based on sequencing of a patient’s tumor genome).

[0015] In certain embodiments, a candidate peptide sequence corresponds to and / or comprises one or more neoantigen epitopes.

[0016] In certain embodiments, a candidate peptides sequence corresponds to and / or comprises one or more shared tumor antigen epitopes.

[0017] In certain embodiments, provided methods comprise obtaining, by the processor, a tumor sequence data (e.g., representing a sequence of at least a portion of a tumor genome and / or exome) for a patient and selecting and / or generating the candidate peptide sequence based on the tumor sequence data.

[0018] In certain embodiments, a candidate peptide sequence is a ribonucleic acid RNA transcribed from a gene encoding the candidate peptide.

[0019] In certain embodiments, a candidate peptide sequence is a deoxyribonucleic acid (DNA) sequence of a gene encoding the candidate peptide.

[0020] In certain embodiments, a prospective antigen is or comprises at least a portion of a viral protein.

[0021] In certain embodiments, a prospective antigen comprises at least a portion of a particular subunit or region (e.g., RBD and / or NTD) of a viral protein.

[0022] In certain embodiments, at least a portion of a prospective antigen is selected iteratively using a sliding window algorithm.

[0023] In certain embodiments, provided methods comprise performing step (b) for each of a plurality of candidate peptides of an initial set, thereby determining a plurality of particular immune response predictions, one for each of the plurality of candidate peptides; and determining, by the processor, the immune response prediction based at least in part on the plurality of candidate immune response predictions (e.g., average of all, maximum of all).

[0024] In certain embodiments, provided methods comprise determining the plurality of candidate peptides of the initial set by selecting peptides of one or more particular desired lengths from a region of a source protein sequence about a particular identified mutation [(e.g., a somatic mutation identified in a tumor cell; e.g., mutation giving rise to a neoantigen; e.g., a mutation giving rise to a shared tumor antigen); e.g., via a sliding window algorithm (e.g., thereby identifying peptides of a certain length, of multiple different lengths)].

[0025] In certain embodiments, at least a portion of a prospective antigen is selected based on known properties (e.g., known to be a problematic portion or unit, known to be a problematic mutation).

[0026] In certain embodiments, a language model is or comprises a transformer model (e.g., wherein the language model comprises one or more transformers).

[0027] In certain embodiments, a language model is a multi-task model, having been trained to determine values for at least two (e.g., a plurality) of different immune response score options, each associated with, and representing a predicted result of, a particular member of a set of immune response tasks (e.g., particular steps in a biological pathway of a host immune response to an antigen or portion thereof), and the one or more immune response scores whose value(s) is / are predicted by the language model at step (b) having been selected from the at least two different immune response score options.

[0028] In certain embodiments, a set of immune response tasks comprises one or both of (i) a peptide-MHC binding task and (ii) a peptide-MHC surface presentation task.

[0029] In certain embodiments, a set of immune response tasks comprises both the peptide-MHC binding task and the peptide-MHC surface presentation task.

[0030] In certain embodiments, a set of immune response tasks comprises a proteasomal cleavage task.

[0031] In certain embodiments, a set of immune response tasks comprises a ligand location task.

[0032] In certain embodiments, a set of immune response tasks comprises a ligand generation task.

[0033] In certain embodiments, a set of immune response tasks comprises two or more of (A), (B), and (C) [e.g., (A) and (B); e.g., (A) and (C); e.g., (B) and (C); e.g., all three of (A), (B), and (C)] as follows: (A) a binary classification task [e.g., where the output of the language model represents a binary value (e.g., a true or false)]; (B) a regression task [e.g., where the output of the language model represents a (e.g., continuously valued, discrete valued) numerical value]; and (C) a segmentation task [e.g., where the output of the language model represents a particular subsequence (e.g., representing a peptide fragment, epitope, etc.) of an input sequence (e.g., a protein sequence) or a location of the particular subsequence (e.g., indices (e.g., of letters) in the input sequence bounding the particular subsequence) in the input sequence] .

[0034] In certain embodiments, at least one of the immune response scores is an MHC binding affinity score representing a predicted binding affinity between the candidate peptide and a particular major histocompatibility complex (MHC) allele (e.g., a MHC class I allele).

[0035] In certain embodiments, an input text string comprises an identification of the particular MHC allele [e.g., one of a discrete set of text labels (e.g., alphanumeric strings) identifying a set of potential MHC alleles; e.g., a sequence (e.g., an amino acid sequence; e.g., a pseudosequence) of at least a portion of the particular MHC allele].

[0036] In certain embodiments, a value of the MHC binding affinity score is a continuous number [e.g., encoded in the output text string as a decimal string; e.g., encoded in the output text string via a token (e.g., an alphanumeric text label) identifying a particular one of a set of discrete bins, each representing a range of continuous values] .

[0037] In certain embodiments, at least one of the immune response scores is an MHC binding affinity classification label and / or value (e.g., a binary value; e.g., a likelihood value, which may be thresholded to produce a binary classification) indicative of whether or not the particular MHC allele is predicted to bind to the candidate peptide (e.g., wherein the machine learning model is trained using eluted ligand data).

[0038] In certain embodiments, at least one of the immune response scores is a surface presentation score representing a prediction of whether the candidate peptide will be presented at a surface via a particular (MHC) allele (e.g., a MHC class I allele) [e.g., a continuous value indicative of a predicted likelihood of presentation and / or efficacy thereof].

[0039] In certain embodiments, a machine learning model receives, as input, a gene bias value representing a background expression level of the candidate peptide [e.g., a continuous value computed using source transcripts of the peptide and measured number of peptides observed per gene relative to expected number based on gene length and expression] .

[0040] In certain embodiments, at least one of the immune response scores is a cleavage score representing a prediction of whether (e.g., a likelihood value, e.g., between 0 and 1) the candidate peptide is a fragment predicted to result from proteasomal cleavage of the antigen.

[0041] In certain embodiments, an input text string comprises an identification of a plurality of distinct MHC alleles and the output text string encodes, a plurality of distinct sets of allele-specific immune response score values, each corresponding to a particular one of the plurality of distinct MHC alleles.

[0042] In certain embodiments, provided methods comprise determining a set of overall immune response score values based on the plurality of distinct sets of allele-specific immune response score values.

[0043] In certain embodiments, provided methods comprise using the immune response prediction to create a personalized cancer vaccine for a patient (e.g., to select one or more neoantigens for inclusion in the personalized cancer vaccine).

[0044] In certain embodiments, provided methods comprise: performing steps (a) through (c) for each of a plurality of candidate peptides, each corresponding to a candidate epitope identified as associated with a tumor genome / exome of a patient (e.g., a candidateneoantigen epitope; e.g., a candidate shared tumor antigen epitope), thereby determining a plurality of immune response predictions, one for each candidate neoepitope; and selecting, by the processor, a subset of the candidate epitope based at least in part on the determined plurality of immune response predictions.

[0045] In certain embodiments, provided methods comprise including one or more biological sequences encoding the selected subset of candidate epitopes (e.g., a candidate neoantigen epitope; e.g., a candidate shared tumor antigen epitope) in an (e.g., individualized) cancer vaccine (e.g., creating one or more nucleic acids encoding the selected subset of candidate epitopes, e.g., via in-vitro transcription) (e.g., synthesizing one or more polypeptides encoding the selected subset of candidate epitopes).

[0046] In certain embodiments, provided methods comprise using the immune response prediction to create a personalized viral vaccine for a patient (e.g., to select one or more virus variants for inclusion in the personalized viral vaccine).

[0047] In certain embodiments, provided methods comprise: performing steps (a) through (c) for each of a plurality of candidate peptides, each corresponding to a candidate viral variant, thereby determining a plurality of immune response predictions, one for each candidate viral variant; and selecting, by the processor, a subset of the candidate viral variants based at least in part on the determined plurality of immune response predictions.

[0048] In some aspects, the present disclosure provides methods for identifying prospective antigen fragments (e.g., peptides; e.g., corresponding to or comprising epitopes) that will be presented as ligands for T-cell recognition, wherein, provided methods comprise: (a) obtaining (e.g., receiving, accessing, generating, etc.), by a processor of a computing device, sequence data comprising (i) a prospective antigen sequence representing a biological sequence of at least a portion of a target antigen and (ii) a major histocompatibility complex (MHC) [e.g., MHC class I (MHC-I) (e.g., a human leukocyte antigen (HLA-I) allele); e.g., MHC class II (MHC -II) (e.g., a human leukocyte antigen (HLA-II) allele)] sequence representing a biological sequence of at least a portion of a particular MHC allele; (b) identifying, by the processor, using a machine learning model (e.g., a language model), a subregion of the antigen sequence, thereby locating, within the antigen, a fragment to be presented as a ligand for T-cell recognition; and (c) storing and / or providing, by the processor, a representation of the located ligand for display and / or further processing.

[0049] In certain embodiments, a machine learning model:(i) receives a sequence data as input; and (ii) generates, as output, a mask or a set of indices identifying the subregion within the antigen sequence (e.g., an array of 1’s and 0’s identifying residues that are or are not, respectively, predicted to be the located ligand; e.g., a sequence of integers each corresponding to a position within the antigen sequence).

[0050] In certain embodiments, a machine learning model receives (i) an antigen sequence as a first channel of input and (ii) an MHC sequence as a second, separate channel of input.

[0051] In certain embodiments, a machine learning model is a multi-task model, having been trained to determine values for at least two (e.g., a plurality) of different immune response score options, each associated with, and representing a predicted result of, a particular member of a set of immune response tasks (e.g., particular steps in a biological pathway of a host immune response to an antigen or portion thereof).

[0052] In certain embodiments, a set of immune response tasks comprises a peptide- MHC binding task and a peptide-MHC surface presentation task.

[0053] In certain embodiments, provided method comprise: using the representation of the located ligand to create a personalized cancer vaccine for a patient (e.g., to select one or more neoantigens for inclusion in the personalized cancer vaccine) [e.g., encoding the located ligand in polynucleotide (e.g., via in-vitro transcription)].

[0054] In some aspects, the present disclosure provides methods for identifying antigen fragments that will be presented as ligands for T-cell recognition, wherein provided methods comprise: (a) obtaining (e.g., receiving, accessing, generating, etc.), by a processor of a computing device, sequence data comprising (i) an antigen sequence representing a biological sequence of at least a portion of a target antigen and (ii) a major histocompatibility complex (MHC) [e.g., MHC class I (MHC-I) (e.g., a human leukocyte antigen (HLA-I) allele); e.g., MHC class II (MHC -II) (e.g., a human leukocyte antigen (HLA-II) allele)] sequence representing a biological sequence of at least a portion of a particular MHC allele; (b) determining, by the processor, using a machine learning model, a ligand sequence corresponding to a subregion of the antigen sequence identified as (e.g., by the machine learning model) a fragment of the antigen predicted to be presented as a ligand for T-cell recognition; and (c) storing and / or providing, by the processor, the ligand sequence for display and / or further processing.

[0055] In certain embodiments, a machine learning model:(i) receives a sequence data as input; and (ii) generates, as output, a mask or a set of indices identifying the subregion within the antigen sequence (e.g., an array of 1’s and 0’s identifying residues that are or are not, respectively, predicted to be the located ligand; e.g., a sequence of integers each corresponding to a position within the antigen sequence).

[0056] In certain embodiments, a machine learning model receives (i) an antigen sequence as a first channel of input and (ii) an MHC sequence as a second, separate channel of input.

[0057] In certain embodiments, a machine learning model is a multi-task model, having been trained to determine values for at least two (e.g., a plurality) of different immune response score options, each associated with, and representing a predicted result of, a particular member of a set of immune response tasks (e.g., particular steps in a biological pathway of a host immune response to an antigen or portion thereof).

[0058] In certain embodiments, a set of immune response tasks comprises a peptide- MHC binding task and a peptide-MHC surface presentation task.

[0059] In certain embodiments, provided methods comprise: using the ligand sequence to create a personalized cancer vaccine for a patient (e.g., to select one or more neoantigens for inclusion in the personalized cancer vaccine) [e.g., encoding the ligand sequence in polynucleotide (e.g., via in-vitro transcription)].

[0060] In some aspects, the present disclosure provides methods for predicting an immune response to a prospective antigen (e.g., an antigen polypeptide; e.g., a source protein and / or portion thereof, such as a particular sub-unit) or portion thereof (e.g., a prospective epitope; e.g., a peptide), wherein provided methods comprise: (a) obtaining (e.g., receiving, accessing, generating, etc.), by a processor of a computing device, an input text string comprising a candidate peptide sequence representing a biological sequence of at least a portion of the prospective antigen; (b) determining, by the processor, a value of a selected immune response score based on the candidate peptide sequence using a machine learning model, wherein the machine learning model has been trained to determine values for a plurality of different immune response score options, each associated with, and representing a predicted result of, a particular member of a set of immune response tasks (e.g., particular steps in a biological pathway of a host immune response to an antigen), and wherein the selected immune response score is one of the plurality of different immune response scoreoptions; and (c) storing and / or providing, by the processor, the value of the selected immune response score for display and / or further processing.

[0061] In certain embodiments, provided methods comprise: repeating steps (a) - (c) for a plurality of candidate peptides, wherein each of the plurality of candidate peptides represents a biological sequence of at least a portion of the prospective antigen, thereby determining a plurality of values of the selected immune response score, one for each member of the plurality of candidate peptides; and selecting, by the processor, a subset of the plurality of candidate peptides for inclusion in composition (e.g., to be encoded by a polynucleotide composition).

[0062] In certain embodiments, a set of immune response tasks comprises a peptide- MHC binding task and a peptide-MHC surface presentation task.

[0063] In certain embodiments, a selected immune response score comprises an MHC binding affinity score representing a predicted binding affinity between the candidate peptide and a particular major histocompatibility complex (MHC) allele (e.g., a MHC class I allele).

[0064] In certain embodiments, an input text string comprises an identification of the particular MHC allele [e.g., one of a discrete set of text labels (e.g., alphanumeric strings) identifying a set of potential MHC alleles; e.g., a sequence (e.g., an amino acid sequence; e.g., a pseudosequence) of at least a portion of the particular MHC allele].

[0065] In certain embodiments, a selected immune response score comprises a surface presentation score representing a prediction of whether the candidate peptide will be presented at a surface via a particular (MHC) allele (e.g., a MHC class I allele) [e.g., a continuous value indicative of a predicted likelihood of presentation and / or efficacy thereof].

[0066] In certain embodiments, provided methods comprise: using the value of the selected immune response score to create a personalized cancer vaccine for a patient (e.g., to select one or more neoantigens for inclusion in the personalized cancer vaccine) [e.g., encoding the selected subset of candidate peptides determined via methods presented herein in polynucleotide (e.g., via in-vitro transcription)].

[0067] In some aspects, the present disclosure provides methods of training a machine learning model to predicting an immune response to a prospective antigen (e.g., an antigen polypeptide; e.g., a source protein and / or portion thereof, such as a particular sub-unit) or portion thereof (e.g., a prospective epitope; e.g., a peptide), wherein provided methodscomprise: (a) obtaining (e.g., receiving, accessing, generating, etc.), by a processor of a computing device, a plurality of datasets, each associated with a distinct particular immune response task (e.g., distinct particular steps in a biological pathway of a host immune response to an antigen) and comprising a plurality of examples, each example associated with and comprising sequence data representing a particular example peptide; (b) for each particular one of the plurality of datasets, repeatedly selecting, by the processor, examples from the particular dataset and using, by the processor, the selected examples to train the machine learning model to determine values of immune response scores, each immune response score associated with and representing a predicted result of the immune response task with which the particular dataset is associated, thereby generating a multi-task machine learning model trained to generate predictions for the at least two different immune response tasks; and (c) storing and / or providing, by the processor, the trained multi-task machine learning model for further processing (e.g., inference predictions based on new input source peptide sequences).

[0068] In certain embodiments, provided methods comprise: obtaining (e.g., receiving, accessing, generating), by the processor, an input text string comprising the sequence data; and determining, by the processor, using the machine learning model, the values of immune response scores, wherein the machine learning model receives, as input, the input text string and generates, as output, the values of immune response scores as an output text string.

[0069] In certain embodiments, at least two different immune response tasks comprise a peptide-MHC binding task and a peptide-MHC surface presentation task.

[0070] In certain embodiments, immune response scores comprise an MHC binding affinity score representing a predicted binding affinity between the example peptide and a particular major histocompatibility complex (MHC) allele (e.g., a MHC class I allele).

[0071] In certain embodiments, an input text string comprises an identification of the particular MHC allele [e.g., one of a discrete set of text labels (e.g., alphanumeric strings) identifying a set of potential MHC alleles; e.g., a sequence (e.g., an amino acid sequence; e.g., a pseudosequence) of at least a portion of the particular MHC allele].

[0072] In certain embodiments, immune response scores comprise a surface presentation score representing a prediction of whether the example peptide will be presentedat a surface via a particular (MHC) allele (e.g., a MHC class I allele) [e.g., a continuous value indicative of a predicted likelihood of presentation and / or efficacy thereof] .

[0073] In some aspects, the present disclosure provides methods for predicting an immune response to a prospective antigen (e.g., an antigen polypeptide; e.g., a source protein and / or portion thereof, such as a particular sub-unit) or portion thereof (e.g., a prospective epitope; e.g., a peptide) using a language model, wherein provided methods comprise: (a) obtaining (e.g., receiving, accessing, generating, etc.), by a processor of a computing device, an input text string comprising a candidate peptide sequence, wherein the candidate peptide sequence comprises a representation of a biological sequence (e.g., amino acid sequence of a candidate peptide, ribonucleic acid RNA transcribed from a gene encoding the candidate peptide, deoxyribonucleic acid (DNA) sequence of a gene encoding the candidate peptide) of at least a portion of the prospective antigen (e.g., an epitope); (b) determining, by the processor, an immune response prediction based on the candidate peptide sequence using the language model, wherein the language model receives, as input, the input text string and generates, as output, the immune response prediction as an output text string encoding values of one or more immune response scores, each immune response score associated with and representing a predicted result of a particular immune response task (e.g., a particular step in a biological pathway of a host immune response to the antigen); and (c) storing and / or providing, by the processor, the immune response prediction for display and / or further processing.

[0074] In certain embodiments, a language model comprises a transformer model (e.g., wherein the language model comprises one or more transformers).

[0075] In certain embodiments, a language model is a multi-task model, having been trained to determine values for at least two (e.g., a plurality) of different immune response score options, each associated with, and representing a predicted result of, a particular member of a set of immune response tasks (e.g., particular steps in a biological pathway of a host immune response to an antigen or portion thereof), and the one or more immune response scores whose value(s) is / are predicted by the language model at step (b) having been selected from the at least two different immune response score options.

[0076] In certain embodiments, at least two different immune response tasks comprise a peptide-MHC binding task and a peptide-MHC surface presentation task.

[0077] In certain embodiments, one or more immune response scores comprise an MHC binding affinity score representing a predicted binding affinity between the candidate peptide and a particular major histocompatibility complex (MHC) allele (e.g., a MHC class I allele).

[0078] In certain embodiments, an input text string comprises an identification of the particular MHC allele [e.g., one of a discrete set of text labels (e.g., alphanumeric strings) identifying a set of potential MHC alleles; e.g., a sequence (e.g., an amino acid sequence; e.g., a pseudosequence) of at least a portion of the particular MHC allele].

[0079] In certain embodiments, one or more immune response scores comprise a surface presentation score representing a prediction of whether the candidate peptide will be presented at a surface via a particular (MHC) allele (e.g., a MHC class I allele) [e.g., a continuous value indicative of a predicted likelihood of presentation and / or efficacy thereof].

[0080] In certain embodiments, presented method comprise using the immune response prediction to create a personalized cancer vaccine for a patient (e.g., to select one or more neoantigens for inclusion in the personalized cancer vaccine) [e.g., performing method presented herein for a plurality of candidate peptides, thereby determining a plurality of corresponding immune response predictions, selecting, based on the corresponding immune response predictions, a subset of the candidate peptides, and encoding the selected subset in polynucleotide (e.g., via in-vitro transcription)].

[0081] In some aspects, the present disclosure provides systems for predicting an immune response to a prospective antigen (e.g., an antigen polypeptide; e.g., a source protein and / or portion thereof, such as a particular sub-unit) or portion thereof (e.g., a prospective epitope; e.g., a peptide) using a language model, wherein provided systems comprise: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: (a) obtain (e.g., receiving, accessing, generating, etc.) an input text string comprising a candidate peptide sequence representing a biological sequence (e.g., an amino acid sequence; e.g., a nucleic acid sequence) of at least a portion (e.g., a particular peptide corresponding to a prospective epitope) of the prospective antigen; (b) determine, using the language model, an immune response prediction based on the candidate peptide sequence, wherein the language model receives, as input, the input text string and generates, as output, the immune response prediction as an output text string encoding values of one or more immune response scores,each immune response score associated with and representing a predicted result of a particular immune response task (e.g., a particular step in a biological pathway of a host immune response to the prospective antigen); and (c) store and / or provide the immune response prediction for display and / or further processing.

[0082] In some aspects, the present disclosure provides systems for identifying prospective antigen fragments (e.g., peptides; e.g., corresponding to or comprising epitopes) that will be presented as ligands for T-cell recognition, wherein provided systems comprise: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: (a) obtain (e.g., receiving, accessing, generating, etc.) sequence data comprising (i) a prospective antigen sequence representing a biological sequence of at least a portion of a target antigen and (ii) a major histocompatibility complex (MHC) [e.g., MHC class I (MHC-I) (e.g., a human leukocyte antigen (HLA-I) allele); e.g., MHC class II (MHC -II) (e.g., a human leukocyte antigen (HLA-II) allele)] sequence representing a biological sequence of at least a portion of a particular MHC allele; (b) identify, using a machine learning model (e.g., a language model), a subregion of the prospective antigen sequence, thereby locating, within the target antigen, a fragment to be presented as a ligand for T-cell recognition; and (c) store and / or provide a representation of the located ligand for display and / or further processing.

[0083] In some aspects, the present disclosure provides systems for identifying antigen fragments that will be presented as ligands for T-cell recognition, wherein provided systems comprise: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: (a) obtain (e.g., receiving, accessing, generating, etc.) sequence data comprising (i) a prospective antigen sequence representing a biological sequence of at least a portion of a target antigen and (ii) a major histocompatibility complex (MHC) [e.g., MHC class I (MHC-I) (e.g., a human leukocyte antigen (HLA-I) allele); e.g., MHC class II (MHC -II) (e.g., a human leukocyte antigen (HLA-II) allele)] sequence representing a biological sequence of at least a portion of a particular MHC allele; (b) determine, using a machine learning model, a ligand sequence corresponding to a subregion of the antigen sequence identified as (e.g., by the machine learning model, based on the sequence data) a fragment of the prospective antigen to be presented as a ligand for T-cell recognition; and (c) store and / or provide the ligand sequence for display and / or further processing.

[0084] In some aspects, the present disclosure provides systems for predicting an immune response to a prospective antigen (e.g., an antigen polypeptide; e.g., a source protein and / or portion thereof, such as a particular sub-unit) or portion thereof (e.g., a prospective epitope; e.g., a peptide), wherein provided systems comprise: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: (a) obtain (e.g., receiving, accessing, generating, etc.) an input text string comprising a candidate peptide sequence representing a biological sequence of at least a portion of the prospective antigen; (b) determine a value of a selected immune response score based on the candidate peptide sequence using a machine learning model, wherein the machine learning model has been trained to determine values for a plurality of different immune response score options, each associated with, and representing a predicted result of, a particular member of a set of immune response tasks (e.g., particular steps in a biological pathway of a host immune response to an antigen), and wherein the selected immune response score is one of the plurality of different immune response score options; and (c) store and / or provide the value of the selected immune response score for display and / or further processing.

[0085] In some aspects, the present disclosure provides systems of training a machine learning model to predicting an immune response to a prospective antigen (e.g., an antigen polypeptide; e.g., a source protein and / or portion thereof, such as a particular sub-unit) or portion thereof (e.g., a prospective epitope; e.g., a peptide), wherein provided systems comprise: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: (a) obtain (e.g., receiving, accessing, generating, etc.) a plurality of datasets, each associated with a distinct particular immune response task (e.g., distinct particular steps in a biological pathway of a host immune response to an antigen) and comprising a plurality of examples, each example associated with and comprising sequence data representing a particular example peptide; (b) for each particular one of the plurality of datasets, repeatedly select examples from the particular dataset and use the selected examples to train the machine learning model to determine values of immune response scores, each immune response score associated with and representing a predicted result of the immune response task with which the particular dataset is associated, thereby generating a multi-task machine learning model trained to generate predictions for the at least two different immune response tasks; and (c) store and / orprovide the trained multi-task machine learning model for further processing (e.g., inference predictions based on new input source peptide sequences).

[0086] In some aspects, the present disclosure provides systems for predicting an immune response to a prospective antigen (e.g., an antigen polypeptide; e.g., a source protein and / or portion thereof, such as a particular sub-unit) or portion thereof (e.g., a prospective epitope; e.g., a peptide) using a language model, wherein provided systems comprise: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: (a) obtain (e.g., receiving, accessing, generating, etc.) an input text string comprising a candidate peptide sequence, wherein the candidate peptide sequence comprises a representation of a biological sequence (e.g., amino acid sequence of the candidate peptide, ribonucleic acid RNA transcribed from a gene encoding the candidate peptide, deoxyribonucleic acid (DNA) sequence of a gene encoding the candidate peptide) of at least a portion of the prospective antigen; (b) determine an immune response prediction based on the candidate peptide sequence using the language model, wherein the language model receives, as input, the input text string and generates, as output, the immune response prediction as an output text string encoding values of one or more immune response scores, each immune response score associated with and representing a predicted result of a particular immune response task (e.g., a particular step in a biological pathway of a host immune response to the antigen); and (c) store and / or provide the immune response prediction for display and / or further processing.

[0087] Features of embodiments described with respect to one aspect of the invention may be applied with respect to another aspect of the invention.BRIEF DESCRIPTION OF THE DRAWING

[0088] The foregoing and other objects, aspects, features, and advantages of the present disclosure will become more apparent and better understood by referring to the following description taken in conjunction with the accompanying drawings, in which:

[0089] FIG. 1 is a schematic illustrating example inputs and outputs of a machine learning model for immune response prediction, according to an illustrative embodiment.

[0090] FIG. 2 is a schematic illustrating a peptide along with flanking residues, according to an illustrative embodiment.

[0091] FIG. 3A is a schematic showing an exemplary embodiment of a text representation of a single major histocompatibility (MHC) allele, according to an illustrative embodiment.

[0092] FIG. 3B is a schematic showing an exemplary embodiment of a text representation of multiple MHC alleles, according to an illustrative embodiment.

[0093] FIG. 4 is a block flow diagram of an example process for predicting an immune response to an antigen using a language model, according to an illustrative embodiment.

[0094] FIG. 5 is a diagram showing an example process for end-to-end modelling of antigen presentation via a multi-task model, according to an illustrative embodiment.

[0095] FIG. 6 is a block flow diagram of an example process for training a machine learning model to predicting an immune response to an antigen, according to an illustrative embodiment.

[0096] FIG. 7 is a block flow diagram of an example process for identifying antigen fragments that will be presented as ligands for T-cell recognition, according to an illustrative embodiment.

[0097] FIG. 8 is a block flow diagram of an example process for predicting antigen fragments that will be presented as ligands for T-cell recognition, according to an illustrative embodiment.

[0098] FIG. 9 is a block diagram of an exemplary cloud computing environment, used in certain embodiments.

[0099] FIG. 10 is a block diagram of an example computing device and an example mobile computing device used in certain embodiments.

[0100] FIG. 11 is a diagram showing an example task-specific pipeline for predicting antigen presentation, used in certain embodiments.

[0101] FIG. 12A is a diagram illustrating use of multiple public datasets for modelling immune responses in certain embodiments.

[0102] FIG. 12B is a diagram showing an approach for creating various datasets used in certain embodiments.

[0103] FIG. 13A is a bar chart showing an allele frequency distribution within an example dataset (the NetMHCpan-4.1 single-allelic (SA) Fold 0 Tune set described herein).

[0104] FIG. 13B is a bar chart showing allele frequency distribution within another example dataset (the mass spectrometry (MS) Ligand Test Set described herein).

[0105] FIG. 14 is a bar chart comparing a number of examples in the (Fold 0) train splits of the NetMHCpan-4.1 (“MHC-I”) and NetMHCIIpan-4.0 (“MHC-II”) eluted ligand (EL) datasets.

[0106] FIG. 15 is a histogram showing a distribution of binding affinities in a train split of the NetMHCpan-4.1 BA dataset (“BA Train Split 0”).

[0107] FIG. 16 is histogram showing a distribution of binding affinities in a tune split of the NetMHCpan-4.1 BA dataset (“BA Tune Split 0”).

[0108] FIG. 17 is a diagram showing an example text-to-text transformer (T5) language model architecture for modeling various immune response tasks, used in certain embodiments.

[0109] FIG. 18 is a schematic showing example text representations of a single major histocompatibility (MHC) allele, used in certain embodiments.

[0110] FIG. 19 is a schematic showing example text representations of multiple MHC alleles, used in certain embodiments.

[0111] FIG. 20 is a diagram showing example text representations of a continuous “gene bias” feature, used in certain embodiments.

[0112] FIG. 21 is a box plot showing Top-K values comparing performance of three example models on binding tasks, as evaluated on the MHC-I binding tune set.

[0113] FIG. 22 is a box plot showing Top-K values comparing performance of three example models on presentation tasks, as evaluated on the MHC-I presentation tune set.

[0114] FIG. 23A is a box plot showing Top-K values comparing performance of example multi-task models that were trained on different mixtures of (e.g., with varying ratios of) binding and presentation examples, evaluated on the MHC-I binding tune set.

[0115] FIG. 23B is a box plot showing Top-K values comparing performance of example multi-task models that were trained on different mixtures of (e.g., with varying ratios of) binding and presentation examples, evaluated on the MHC-I presentation tune set..

[0116] FIG. 24A is a box plot showing Top-K values comparing performance of different allele representations used in an example multi-task T5 model, evaluated per allele, on the MHC-I binding tune set.

[0117] FIG. 24B is a box plot showing Top-K values comparing performance of different allele representations used in an example multi-task T5 model, evaluated per allele, on the MHC-I presentation tune set.

[0118] FIG. 25A is a box plot showing Top-K values comparing performance of different allele representations used in an example multi-task T5 model, evaluated per allele, on a MHC-I MHCFovea unseen alleles set.

[0119] FIG. 25B is a box plot showing Top-K values comparing performance of different allele representations used in an example multi-task T5 model, evaluated per allele, on a MHC-I NetMHCPan4.1 unseen alleles set.

[0120] FIG. 26A is a box plot showing Top-K values measured on the MHC-II binding tune set per-allele for an example T5 model that used different allele representations.

[0121] FIG. 26B is a box plot showing Top-K values measured on the MHC-II presentation tune set per-allele for an example T5 model that used different allele representations.

[0122] FIG. 27A is a box plot showing Top-K values measured on the MHC-I binding tune set per-allele for an example T5 model that used different numbers of quantile bins.

[0123] FIG. 27B is a box plot showing Top-K values measured on the MHC-I presentation tune set per-allele for an example T5 model that used different numbers of quantile bins.

[0124] FIG. 28A is a box plot showing Top-K values measured on the MHC-I binding tune set per-allele for an example T5 model that used different methods of discretizing continuous presentation features.

[0125] FIG. 28B is a box plot showing Top-K values measured on the MHC-I presentation tune set per-allele for an example T5 model that used different methods of discretizing continuous presentation features.

[0126] FIG. 29A is a box plot showing Top-K values measured on the MHC-II binding tune set per-allele for an example T5 model that used different numbers of quantile bins.

[0127] FIG. 29B is a box plot showing Top-K values measured on the MHC-II presentation tune set per-allele for an example T5 model that used different numbers of quantile bins.

[0128] FIG. 30A is a box plot showing Top-K values measured on the MHC-II binding test set per-allele for various models.

[0129] FIG. 30B is a box plot showing Top-K values measured on the MHC-II presentation test set per-allele for various models.

[0130] FIG. 31A is a box plot showing Top-K values comparing performance of example multi-task models (T5 models) on the MHC-I binding tune set. The different multitask models evaluated were trained on different datasets, each comprising a different mixture of example tasks.

[0131] FIG. 31B is a box plot showing Top-K values comparing performance of example multi-task models (T5 models) on the MHC-I presentation tune set. The different multi-task models evaluated were trained on different datasets, each comprising a different mixture of example tasks.

[0132] FIG. 32A is a box plot showing Top-K values measured on the MHC-I binding test set per-allele for various models.

[0133] FIG. 32B is a box plot showing Top-K values measured on the MHC-I presentation test set per-allele for various models.

[0134] FIG. 33A is a box plot showing Top-K values for various models.

[0135] FIG. 33B is a box plot showing Top-K values for various models.

[0136] FIG. 34 is a box plot showing average precision scores of for several example multi-task models (T5 models) on the NetMHCpan-4.1 HLA single allele Fold 0 tune set. Different hyperparameters were used for the different models evaluated.

[0137] FIG. 35A is a box plot showing Top-K values comparing performance of an example multi-task model (T5 model) with performance of a task specific ensemble model, evaluated on a MS Ligands test set.

[0138] FIG. 35B is a box plot comparing average precision scores for an example multi-task model (T5 model) and an example task specific ensemble model, evaluated on the MS Ligands test set.

[0139] FIG. 36 is a box plot showing FRANK scores evaluating performance of an example multi-task model (T5 model) on a CD8 test set, where the example model was trained on the NetMHCpan-4.1 HLA SA set.

[0140] FIG. 37 is a bar chart showing classification performance the MS Ligands test set.

[0141] FIG. 38 is a bar chart showing mean squared distortion of NetMHCpan-4. 1 binding affinity values when using different quantization approaches as compared to original binding affinity data.

[0142] FIG. 39A is a box plot showing average precision scores on the NetMHCpan- 4.1 EL HLA single allele Fold 0 tune set for various example multi-task models (T5 models) trained with varying single allele to binding affinity mixture ratios.

[0143] FIG. 39B is a box plot showing Top-K values on the NetMHCpan-4. 1 EL HLA single allele Fold 0 tune set for various example multi-task models (T5 models) trained with varying single allele to binding affinity mixture ratios.

[0144] FIG. 40 is a bar chart showing R2-scores on the NetMHCpan-4. 1 BA tune set for example multi-task models (T5 models) trained with varying single allele to binding affinity mixture ratios.

[0145] FIG. 41A is a box plot showing Top-K values comparing performance of two different example multi-task models (T5 models) measured on the MS Ligands test set. One model (“T5 SA”) was trained on single allelic data only and another (“T5 3: 1”) trained on a three-to-one mixture of single allelic and binding affinity data.

[0146] FIG. 41B is a box plot showing average precision scores for the two example multi-task models of FIG. 41 A, as measured on the MS Ligands test set.

[0147] FIG. 42 is a box plot showing FRANK scores as measured per epitope for CD8 test results for an example T5 model trained on single allelic data only and an example T5 model trained on a mixture of single allelic and binding affinity data (“Binding Affinity” corresponds to a binding affinity task and “Binding” corresponds to a binding classification task) (with logarithmic scale for score above 10‘5).

[0148] FIG. 43 is a box plot showing average precision scores on a tune set for various example T5 models that used deconvolution and an end-to-end approach.

[0149] FIG. 44 is a box plot showing Top-K values on a tune set for various example T5 models that used deconvolution and an end-to-end approach.

[0150] FIG. 45 is bar chart showing R2-scores on the NetMHCpan-4.1 BA tune set for an example T5 model.

[0151] FIG. 46A is a box plot showing Top-K values on the MS Ligands test set for various models: an example T5 model, task specific (sub- and ensemble) models, and the NetMHCpan-4.1 model.

[0152] FIG. 46B is a box plot showing average precision scores measured per-allele on the MS Ligands test set for various models: an example T5 model, task specific (sub- and ensemble) models, and the NetMHCpan-4.1 model.

[0153] FIG. 47A is a plot showing precision-recall (PR) curves on the MS Ligands test set for various models: an example T5 model, task specific (sub- and ensemble) models, and the NetMHCpan-4.1 model.

[0154] FIG. 47B is a plot showing Receiver Operating Characteristic (ROC) curves on the MS Ligands test set for various models: an example T5 model, task specific (sub- and ensemble) models, and the NetMHCpan-4.1 model.

[0155] FIG. 48A is a box plot showing Top-K values on the MS Ligands test set for three example T5 models that used various single-allelic prefix notations.

[0156] FIG. 48B is a box plot showing average precision scores on the MS Ligands test set for three example T5 models that used various single-allelic prefix notations.

[0157] FIG. 48C is a box plot showing Top-K values on the MS Ligands test set for three example T5 models that used various multi-allelic prefix notations.

[0158] FIG. 48D is a box plot showing average precision scores on the MS Ligands test set for three example T5 models that used various multi-allelic prefix notations.

[0159] FIG. 49 is a box plot showing FRANK scores for the CD8 test results of an example T5 model that used single-allelic prefixes (with a logarithmic scale for scores above 10’5).

[0160] FIG. 50 is a box plot showing FRANK scores for the CD8 test results of an example T5 model that used multi-allelic prefixes (with a logarithmic scale for scores above 10’5).

[0161] FIG. 51A is a box plot showing Top-K values on the NetMHCpan-4.1 SA tune set for an example T5 Base Model and an example Fine-Tuned Model.

[0162] FIG. 51B is a box plot showing average precision scores on the NetMHCpan- 4.1 SA tune set for an example T5 Base Model and an example Fine-Tuned Model.

[0163] FIG. 51C is a box plot showing Top-K values on the MS Ligands test set for an example T5 Base Model and an example Fine-Tuned Model.

[0164] FIG. 51D is a box plot showing average precision scores on the MS Ligands test set for an example T5 Base Model and an example Fine-Tuned Model.

[0165] FIG. 51E is a box plot showing FRANK scores for CD8 test results for an example T5 Base Model and an example Fine-Tuned Model (with a logarithmic scale for scores above 10‘5).

[0166] FIG. 52A is a bar chart showing a distribution of allele sampling frequencies for a train Fold 0 of the NetMHCpan-4. 1 MHC-I binding data for 8 = 0.

[0167] FIG. 52B is a bar chart showing a distribution of allele sampling frequencies for the train Fold 0 of the MHC-I binding data for 8 = 0.6.

[0168] FIG. 52C is a bar chart showing a distribution of allele sampling frequencies for the train Fold 0 of the NetMHCpan-4.1 MHC-I binding data for 8 = 0.9.

[0169] FIG. 53A is a box plot showing Top-K values on the NetMHCpan-4.1 SA Fold 0 tune set for an example T5 model trained with allele frequencies resampled according to different 8 values.

[0170] FIG. 53B is a box plot showing average precision scores on the NetMHCpan- 4.1 SA Fold 0 tune set for an example T5 model trained with allele frequencies resampled according to different 8 values.

[0171] FIG. 53C is a box plot showing Top-K values on the MS Ligands test set for an example T5 model trained with allele frequencies resampled according to different 8 values.

[0172] FIG. 53D is a box plot showing average precision scores on the MS Ligands test set for an example T5 model trained with allele frequencies resampled according to different a values.

[0173] FIG. 54A is a box plot showing Top-K values on the NetMHCpan-4.1 SA Fold 0 tune set for an example T5-660k model that used pseudo-sequence-only allele representation and various hit to decoy ratios.

[0174] FIG. 54B is a box plot showing average precision scores on the NetMHCpan-4.1 SA Fold 0 tune set for an example T5-660k model that used pseudo-sequence-only allele representation and various hit to decoy ratios.

[0175] FIG. 54C is a box plot showing Top-K values on the MS Ligands test set for an example T5-660k model that used pseudo-sequence-only allele representation and various hit to decoy ratios.

[0176] FIG. 54D is a box plot showing average precision scores on the MS Ligands test set for an example T5-660k model that used pseudo-sequence-only allele representation and various hit to decoy ratios.

[0177] FIG. 55A is a box plot showing Top-K values on the NetMHCpan-4.1 SA Fold 0 tune set for an example T5 model that used single -allelic inversion.

[0178] FIG. 55B is a box plot showing average precision scores on the NetMHCpan-4.1 SA Fold 0 tune set for an example T5 model that used single-allelic inversion.

[0179] FIG. 55C is a box plot showing Top-K values on the MS Ligands test set for an example T5 model that used single -allelic inversion.

[0180] FIG. 55D is a box plot showing average precision scores on the MS Ligands test set for an example T5 model that used single-allelic inversion.

[0181] FIG. 55E is a box plot showing Top-K values on the NetMHCpan-4. 1 SA Fold 0 tune set for an example T5 model that used multi-allelic inversion.

[0182] FIG. 55F is a box plot showing average precision scores on the NetMHCpan-4.1 SA Fold 0 tune set for an example T5 model that used multi-allelic inversion.

[0183] FIG. 55G is a box plot showing Top-K values on the MS Ligands test set for an example T5 model that used multi-allelic inversion.

[0184] FIG. 55H is a box plot showing average precision scores on the MS Ligands test set for an example T5 model that used multi-allelic inversion.

[0185] FIG. 56A is a box plot showing Top-K values on the NetMHCpan-4.1 SA Fold 0 tune set for an example T5 model trained on various synthetic data.

[0186] FIG. 56B is a box plot showing average precision scores on the NetMHCpan- 4.1 SA Fold 0 tune set for an example T5 model trained on various synthetic data.

[0187] FIG. 56C is a box plot showing Top-K values on the MS Ligands test set for an example T5 model trained on various synthetic data.

[0188] FIG. 56D is a box plot showing average precision scores on the MS Ligands test set for an example T5 model trained on various synthetic data.

[0189] FIG. 57A is a box plot showing Top-K values on the NetMHCpan-4.1 SA Fold 0 tune set for an example T5 model that used various randomized allele representations.

[0190] FIG. 57B is a box plot showing average precision scores on the NetMHCpan- 4.1 SA Fold 0 tune set for an example T5 model that used various randomized allele representations.

[0191] FIG. 57C is a box plot showing Top-K values on the MS Ligands test set for an example T5 model that used various randomized allele representations.

[0192] FIG. 57D is a box plot showing average precision scores on the MS Ligands test set for an example T5 model that used various randomized allele representations.

[0193] FIG. 58 is a graph showing mean FRANK scores on the CD8 test set for an example T5 model trained with and without non-HLA (e.g., non-human) data and for the NetMHCpan-4. 1 model.

[0194] FIG. 59A is a box plot showing Top-K values on the NetMHCpan-4.1 SA Fold 0 tune set for an example T5-660k model trained on data with and without non-HLA (e.g., non-human) examples and that used pseudo-sequence allele representations.

[0195] FIG. 59B is a box plot showing average precision scores for two example multi-task models (T5-660k models), each trained on a different dataset, evaluated on the NetMHCpan-4. 1 SA Fold 0 tune set. Both models used pseudo-sequence allele representations; one was trained using data comprising non-HLA examples (e.g., non- human), while the other was trained using human only data.

[0196] FIG. 59C is a box plot showing Top-K values on the MS Ligands test set for an example T5-660k model trained on data with and without non-HLA examples and that used pseudo-sequence allele representations.

[0197] FIG. 59D is a box plot showing average precision scores on the MS Ligands test set for an example T5-660k models trained on data with and without non-HLA examples and that used pseudo-sequence allele representations.

[0198] FIG. 60A is a box plot showing Top-K values on the NetMHCpan-4.1 SA Fold 0 tune set for an example T5-25M model that used pseudo-sequence allele representations and various deconvolution methods.

[0199] FIG. 60B is a box plot showing average precision scores on the NetMHCpan- 4.1 SA Fold 0 tune set for an example T5-25M model that used pseudo-sequence allele representations and various deconvolution methods.

[0200] FIG. 60C is a box plot showing Top-K values on the MS Ligands test set for an example T5-25M model that used pseudo-sequence allele representations and various deconvolution methods.

[0201] FIG. 60D is a box plot showing average precision scores on the MS Ligands test set for an example T5-25M model that used pseudo-sequence allele representations and various deconvolution methods.

[0202] FIG. 61A is a box plot showing Top-K values on the NetMHCpan-4.1 SA Fold 0 tune set for an example T5-25M model that used allele name representations with various ways to shuffle an input order of alleles.

[0203] FIG. 61B is a box plot showing average precision scores on the NetMHCpan- 4.1 SA Fold 0 tune set for an example T5-25M model that used allele name representations with various ways to shuffle an input order of alleles.

[0204] FIG. 61C is a box plot showing Top-K values on the MS Ligands test set for an example T5-25M model that used allele name representations with various ways to shuffle an input order of alleles.

[0205] FIG. 61D is a box plot showing average precision scores on the MS Ligands test set for an example T5-25M model that used allele name representations with various ways to shuffle an input order of alleles.

[0206] FIG. 62 is a block diagram showing an example input and outputs of Ligand Location and Ligand Generation of an example T5 model.

[0207] FIG. 63 is a block diagram showing two ligand generation approaches of an example T5 model: single shot and multi-shot.

[0208] FIG. 64 is a block diagram showing ligand generation process for an example T5 model using a sequence and a structure.

[0209] FIG. 65 is a block diagram showing an architecture of an example T5 model, framing tasks as text input and output, according to an illustrative embodiment.

[0210] FIG. 66A is a box plot showing Top-K values on an in-house SA test dataset for an example T5 model and NetMHCpan-4.1 model.

[0211] FIG. 66B is a box plot showing Top-K values on an in-house MA test dataset for an example T5 model and NetMHCpan-4.1 model.

[0212] FIG. 67 is a box plot showing Top-K values on the NetMHCpan EL MS Ligands test set for an example T5 model and NetMHCpan-4. 1 model when both models are trained on a same data.

[0213] FIG. 68A is a bar chart showing AUC-ROC values on a p-TAP binding affinity prediction for an example T5 model and DeepTAP model.

[0214] FIG. 68B is a bar chart showing R2values on a p-MHC stability prediction for an example T5 model and NetMHCStabPan-1.0 model.

[0215] FIG. 69 is a bar chart showing Top-K values comparing impact of different training datasets on performance of an example multi-task (T5) model. As shown in the figure, the Top-K values evaluate performance of the different model versions on three different benchmarks.

[0216] The features and advantages of the present disclosure will become more apparent from the detailed description set forth below when taken in conjunction with the drawings, in which like reference characters identify corresponding elements throughout. In the drawings, like reference numbers generally indicate identical, functionally similar, and / or structurally similar elements.DEFINITIONS

[0217] About: The term “about”, when used herein in reference to a value, refers to a value that is similar, in context to the referenced value. In general, those skilled in the art, familiar with the context, will appreciate the relevant degree of variance encompassed by “about” in that context. For example, in some embodiments, the term “about” may encompass a range ofvalues that within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less of the referred value.

[0218] Affinity: As is known in the art, “affinity” is a measure of the tightness with which two or more binding partners associate with one another. Those skilled in the art are aware of a variety of assays that can be used to assess affinity, and will furthermore be aware of appropriate controls for such assays. In some embodiments, affinity is assessed in a quantitative assay. In some embodiments, affinity is assessed over a plurality of concentrations (e.g., of one binding partner at a time). In some embodiments, affinity is assessed in the presence of one or more potential competitor entities (e.g., that might be present in a relevant - e.g., physiological - setting). In some embodiments, affinity is assessed relative to a reference (e.g., that has a known affinity above a particular threshold [a “positive control” reference] or that has a known affinity below a particular threshold [ a “negative control” reference”]. In some embodiments, affinity may be assessed relative to a contemporaneous reference; in some embodiments, affinity may be assessed relative to a historical reference. Typically, when affinity is assessed relative to a reference, it is assessed under comparable conditions.

[0219] Amino acid: In its broadest sense, as used herein, the term “amino acid” refers to a compound and / or substance that can be, is, or has been incorporated into a polypeptide chain, e.g., through formation of one or more peptide bonds. In some embodiments, an amino acid has the general structure H2N-C(H)(R)-COOH. In some embodiments, an amino acid is a naturally-occurring amino acid. In some embodiments, an amino acid is a non-natural amino acid; in some embodiments, an amino acid is a D-amino acid; in some embodiments, an amino acid is an L-amino acid. “Standard amino acid” refers to any of the twenty standard L-amino acids commonly found in naturally occurring peptides. “Nonstandard amino acid” refers to any amino acid, other than the standard amino acids, regardless of whether it is prepared synthetically or obtained from a natural source. In some embodiments, an amino acid, including a carboxy- and / or amino-terminal amino acid in a polypeptide, can contain a structural modification as compared with the general structureabove. For example, in some embodiments, an amino acid may be modified by methylation, amidation, acetylation, pegylation, glycosylation, phosphorylation, and / or substitution (e.g., of the amino group, the carboxylic acid group, one or more protons, and / or the hydroxyl group) as compared with the general structure. In some embodiments, such modification may, for example, alter the circulating half-life of a polypeptide containing the modified amino acid as compared with one containing an otherwise identical unmodified amino acid. In some embodiments, such modification does not significantly alter a relevant activity of a polypeptide containing the modified amino acid, as compared with one containing an otherwise identical unmodified amino acid. As will be clear from context, in some embodiments, the term “amino acid” may be used to refer to a free amino acid; in some embodiments it may be used to refer to an amino acid residue of a polypeptide.

[0220] Antibody. As used herein, the term “antibody” refers to a polypeptide that includes canonical immunoglobulin sequence elements sufficient to confer specific binding to a particular target antigen. As is known in the art, intact antibodies as produced in nature are approximately 150 kD tetrameric agents comprised of two identical heavy chain polypeptides (about 50 kD each) and two identical light chain polypeptides (about 25 kD each) that associate with each other into what is commonly referred to as a “Y-shaped” structure. Each heavy chain is comprised of at least four domains (each about 110 amino acids long)- an amino-terminal variable (VH) domain (located at the tips of the Y structure), followed by three constant domains: CHI, CH2, and the carboxy-terminal CH3 (located at the base of the Y’s stem). A short region, known as the “switch”, connects the heavy chain variable and constant regions. The “hinge” connects CH2 and CH3 domains to the rest of the antibody. Two disulfide bonds in this hinge region connect the two heavy chain polypeptides to one another in an intact antibody. Each light chain is comprised of two domains - an aminoterminal variable (VL) domain, followed by a carboxy-terminal constant (CL) domain, separated from one another by another “switch”. Intact antibody tetramers are comprised of two heavy chain-light chain dimers in which the heavy and light chains are linked to one another by a single disulfide bond; two other disulfide bonds connect the heavy chain hinge regions to one another, so that the dimers are connected to one another and the tetramer is formed. Naturally-produced antibodies are also glycosylated, typically on the CH2 domain. Each domain in a natural antibody has a structure characterized by an “immunoglobulin fold” formed from two beta sheets (e.g., 3-, 4-, or 5-stranded sheets) packed against each other in a compressed antiparallel beta barrel. Each variable domain contains three hypervariable loopsknown as “complement determining regions” (CDR1, CDR2, and CDR3) and four somewhat invariant “framework” regions (FR1, FR2, FR3, and FR4). When natural antibodies fold, the FR regions form the beta sheets that provide the structural framework for the domains, and the CDR loop regions from both the heavy and light chains are brought together in three- dimensional space so that they create a single hypervariable antigen binding site located at the tip of the Y structure. The Fc region of naturally-occurring antibodies binds to elements of the complement system, and also to receptors on effector cells, including for example effector cells that mediate cytotoxicity. As is known in the art, affinity and / or other binding attributes of Fc regions for Fc receptors can be modulated through glycosylation or other modification. In some embodiments, antibodies produced and / or utilized in accordance with the present disclosure include glycosylated Fc domains, including Fc domains with modified or engineered such glycosylation. For purposes of the present disclosure, in certain embodiments, any polypeptide or complex of polypeptides that includes sufficient immunoglobulin domain sequences as found in natural antibodies can be referred to and / or used as an “antibody”, whether such polypeptide is naturally produced (e.g., generated by an organism reacting to an antigen), or produced by recombinant engineering, chemical synthesis, or other artificial system or methodology. In some embodiments, an antibody is polyclonal; in some embodiments, an antibody is monoclonal. In some embodiments, an antibody has constant region sequences that are characteristic of mouse, rabbit, primate, or human antibodies. In some embodiments, antibody sequence elements are humanized, primatized, chimeric, etc., as is known in the art. Moreover, the term “antibody” as used herein, can refer in appropriate embodiments (unless otherwise stated or clear from context) to any of the art-known or developed constructs or formats for utilizing antibody structural and functional features in alternative presentation. For example, in some embodiments, an antibody utilized in accordance with the present disclosure is in a format selected from, but not limited to, intact IgA, IgG, IgE or IgM antibodies; bi- or multi- specific antibodies (e.g., Zybodies®, etc.); antibody fragments such as Fab fragments, Fab’ fragments, F(ab’)2 fragments, Fd’ fragments, Fd fragments, and isolated CDRs or sets thereof; single chain Fvs; polypeptide-Fc fusions; single domain antibodies, alternative scaffolds or antibody mimetics (e.g., anticalins, FN3 monobodies, DARPins, Affibodies, Affilins, Affimers, Affitins, Alphabodies, Avimers, Fynomers, Im7, VLR, VNAR, Trimab, CrossMab, Trident); nanobodies, binanobodies, F(ab’)2, Fab’, di-sdFv, single domain antibodies, trifiinctional antibodies, diabodies, and minibodies, etc. In some embodiments, relevant formats may be orinclude: Adnectins®; Affibodies®; Affilins®; Anticalins®; Avimers®; BiTE®s; cameloid antibodies; Centyrins®; ankyrin repeat proteins or DARPINs®; dual-affinity re-targeting (DART) agents; Fynomers®; shark single domain antibodies such as IgNAR; immune mobilixing monoclonal T cell receptors against cancer (ImmTACs); KALBITOR®s; MicroProteins; Nanobodies® minibodies; masked antibodies (e.g., Probodies®); Small Modular ImmunoPharmaceuticals (“SMIPsTM”); single chain or Tandem diabodies (TandAb®); TCR-like antibodies;, Trans-bodies®; TrimerX®; VHHs. In some embodiments, an antibody may lack a covalent modification (e.g., attachment of a glycan) that it would have if produced naturally. In some embodiments, an antibody may contain a covalent modification (e.g., attachment of a glycan, a payload [e.g., a detectable moiety, a therapeutic moiety, a catalytic moiety, etc.], or other pendant group [e.g., poly-ethylene glycol, etc.]).

[0221] Antigen: term “antigen”, as used herein, refers to an agent that elicits an immune response; and / or an agent that binds to a T cell receptor (e.g., when presented by an MHC molecule) or to an antibody. In some embodiments, an antigen elicits a humoral response (e.g., including production of antigen-specific antibodies); in some embodiments, an antigen elicits a cellular response (e.g., involving T-cells whose receptors specifically interact with the antigen). In some embodiments, an antigen binds to an antibody and may or may not induce a particular physiological response in an organism. In general, an antigen may be or include any chemical entity such as, for example, a small molecule, a nucleic acid, a polypeptide, a carbohydrate, a lipid, a polymer (in some embodiments other than a biologic polymer [e.g., other than a nucleic acid or amino acid polymer]) etc. In some embodiments, an antigen is or comprises a polypeptide. In some embodiments, an antigen is or comprises a glycan. Those of ordinary skill in the art will appreciate that, in general, an antigen may be provided in isolated or pure form, or alternatively may be provided in crude form (e.g., together with other materials, for example in an extract such as a cellular extract or other relatively crude preparation of an antigen-containing source). In some embodiments, antigens utilized in accordance with the present invention are provided in a crude form. In some embodiments, an antigen is a recombinant antigen.

[0222] Antigen presenting cell. The phrase “antigen presenting cell” or “APC,” as used herein, has its art understood meaning referring to cells which process and present antigens to T-cells. Exemplary antigen cells include dendritic cells, macrophages and certain activated epithelial cells.

[0223] Cancer. The term “cancer” is used herein to generally refer to a disease or condition in which cells of a tissue of interest exhibit relatively abnormal, uncontrolled, and / or autonomous growth, so that they exhibit an aberrant growth phenotype characterized by a significant loss of control of cell proliferation. In some embodiments, cancer may comprise cells that are precancerous (e.g., benign), malignant, pre-metastatic, metastatic, and / or non-metastatic. In some embodiments, cancer may be characterized by a solid tumor. In some embodiments, cancer may be characterized by a hematologic tumor. In general, examples of different types of cancers known in the art include, for example, triple negative breast cancer (TNBC), hematopoietic cancers including leukemias, lymphomas (Hodgkin’s and non-Hodgkin’s), myelomas and myeloproliferative disorders; sarcomas, melanomas, adenomas, carcinomas of solid tissue, squamous cell carcinomas of the mouth, throat, larynx, and lung, liver cancer, genitourinary cancers such as prostate, cervical, bladder, uterine, and endometrial cancer and renal cell carcinomas, bone cancer, pancreatic cancer, skin cancer, cutaneous or intraocular melanoma, cancer of the endocrine system, cancer of the thyroid gland, cancer of the parathyroid gland, head and neck cancers, ovarian cancer, breast cancer, glioblastomas, colorectal cancer, gastro-intestinal cancers and nervous system cancers, benign lesions such as papillomas, and the like.

[0224] Epitope: As used herein, the term “epitope” refers to a moiety that is specifically recognized by an immune system (e.g., an immune system component) of a subject. For example, in some embodiments, an epitope may be a moiety that is specifically recognized by a T cell, a B cell, an immunoglobulin (e.g., antibody or receptor), binding component or an aptamer. In some embodiments, an epitope is comprised of a plurality of chemical atoms or groups on an antigen. In some embodiments, such chemical atoms or groups are surface-exposed when the antigen adopts a relevant three-dimensional conformation. In some embodiments, such chemical atoms or groups are physically near to each other in space when the antigen adopts such a conformation. In some embodiments, at least some such chemical atoms are groups are physically separated from one another when the antigen adopts an alternative conformation (e.g., is linearized).

[0225] Individualized shared tumor antigen : As used herein, an “individualized shared tumor antigen” is a shared tumor antigen that is expressed in an individual subject.

[0226] Individualized shared tumor antigen epitope: As used herein, an“individualized shared tumor antigen epitope” is a shared tumor antigen epitope that is expressed in an individual subject.

[0227] Machine learning module, machine learning model: As used herein, the terms “machine learning module” and “machine learning model” are used interchangeably and refer to a computer implemented process (e.g., a software function) that implements one or more particular machine learning algorithms, such as artificial neural networks (ANNs), random forests, decision trees, support vector machines, and the like, in order to determine, for a given input, one or more output values. In certain embodiments, machine learning models are deep learning models or deep neural networks - for example, ANNs that comprise, in addition to an input layer and an output layer, one or more hidden layers (e.g., in between). Examples of deep learning models include, without limitation, recurrent neural networks (RNNs) (e.g., long short-term memory networks (LSTMs), bi-directional LSTMs (biLSTMs)), attention-based networks, such as transformer models, and convolutional neural networks (CNNs). In some embodiments, machine learning modules implementing machine learning techniques are trained in a supervised manner, for example using curated and / or manually annotated datasets. In certain embodiments, machine learning models may be trained in an unsupervised manner, using unlabeled data. In certain embodiments, a machine learning model may be trained via a reinforcement approach, for example wherein a reward / penalty system is used to train a machine learning model to learn strategies for accomplishing specified tasks. Training a machine learning model may be used to determine various parameters of a model, such as weights associated with layers in neural networks. In some embodiments, once a machine learning module is trained, e.g. , to accomplish a specific task, such as predicting types of hidden amino acids within of polypeptide sequences based on their context, values of determined parameters are fixed and the machine learning module is used to process new data (e.g. , different from the training data), such as a new amino acid sequence. The process of presenting a machine learning model with multiple examples, comparing its output to known, ground truth values, and updating parameters to progressively improve performance may be referred to as training, while the use of a (e.g., previously trained) machine learning model to generate predictions about new data, for which ground truth values may be unknown, may be referred to as inference. In some embodiments, machine learning modules may receive feedback, e.g., based on user review of accuracy, and such feedback may be used as additional training data, for example to dynamically update the machine learning module. In some embodiments, a trained machine learning module is a classification algorithm with adjustable and / or fixed (e.g., locked) parameters, e.g., a random forest classifier. In some embodiments, two or more machine learning modules may becombined and implemented as a single module and / or a single software application. In some embodiments, two or more machine learning modules may also be implemented separately, e.g., as separate software applications. A machine learning module may be software and / or hardware. For example, a machine learning module may be implemented entirely as software, or certain functions of an ANN module may be carried out via specialized hardware (e.g., via an application specific integrated circuit (ASIC), field programmable gate arrays (FPGAs), and the like).

[0228] Model: As used herein, the term “model” is used to identify a computer representation of a particular physical object or quantity, e.g., that is accessed, displayed by, used as input to, generated as output of, etc., computer-implemented methods and systems and / or one or more steps and / or modules or functions thereof. For example, as described in further detail herein, various computer implemented systems and methods may operate on, process, and generate polypeptide models that represent physical polypeptides, such as particular proteins or portions thereof. Computer representations of polypeptides, e.g., polypeptide models, may be implemented in a variety of formats, such as a string of characters (e.g., letters, each representing a particular amino acid type) representing an amino acid sequence (e.g., a FASTA file), or a 3D structural model that includes information about a (e.g., relative) 3D location of amino acids and / or atoms thereof, such as a Protein Data Bank (PDB) file format which may be used to describe a 3D structure of a particular protein and includes, among other things, atomic coordinates of atoms of particular protein.

[0229] Neoantigen: As used herein, the term “neoantigen” refers to an antigen that is not present in a reference, such as a normal non-cancerous or germline cell, but is present in a cancer cell. In some embodiments, a neoantigen includes one or more mutations relative to a corresponding antigen present in a normal non-cancerous or germline cell.

[0230] Neoantigen epitope: As used herein, the term “neoantigen epitope” refers to an epitope that is not present in a reference, such as a normal non-cancerous or germline cell, but is present in a cancer cell.

[0231] Non-neoantigen: As used herein, the term “non-neoantigen” refers to a tumor antigen that is not a neoantigen. In some embodiments, a non-neoantigen is a shared tumor antigen. In some embodiments, a non-neoantigen is an individualized shared tumor antigen.

[0232] Non-neoantigen epitope: As used herein, the term “non-neoantigen epitope” refers to a tumor epitope that is not a neoantigen epitope. In some embodiments, a non-neoantigen epitope is a shared tumor antigen epitope. In some embodiments, a non- neoantigen epitope is an individualized shared tumor antigen epitope.

[0233] Nucleic acid. As used herein, the term “nucleic acid” in its broadest sense, refers to any compound and / or substance that is or can be incorporated into an oligonucleotide chain. In some embodiments, a nucleic acid is a compound and / or substance that is or can be incorporated into an oligonucleotide chain via a phosphodiester linkage. As will be clear from context, in some embodiments, "nucleic acid" refers to an individual nucleic acid residue (e.g., a nucleotide and / or nucleoside); in some embodiments, "nucleic acid" refers to an oligonucleotide chain comprising individual nucleic acid residues. In some embodiments, a "nucleic acid" is or comprises RNA; in some embodiments, a "nucleic acid" is or comprises DNA. In some embodiments, a nucleic acid is, comprises, or consists of one or more natural nucleic acid residues. In some embodiments, a nucleic acid is, comprises, or consists of one or more nucleic acid analogs. In some embodiments, a nucleic acid analog differs from a nucleic acid in that it does not utilize a phosphodiester backbone. For example, in some embodiments, a nucleic acid is, comprises, or consists of one or more "peptide nucleic acids", which are known in the art and have peptide bonds instead of phosphodiester bonds in the backbone, are considered within the scope of the present disclosure.Alternatively or additionally, in some embodiments, a nucleic acid has one or more phosphorothioate and / or 5'-N-phosphoramidite linkages rather than phosphodiester bonds. In some embodiments, a nucleic acid is, comprises, or consists of one or more natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxy guanosine, and deoxy cytidine). In some embodiments, a nucleic acid is, comprises, or consists of one or more nucleoside analogs (e.g., 2-aminoadenosine, 2- thiothymidine, inosine, pyrrolo-pyrimidine, 3 -methyl adenosine, 5-methylcytidine, C-5 propynyl-cytidine, C-5 propynyl-uridine, 2-aminoadenosine, C5 -bromouridine, C5- fluorouridine, C5 -iodouridine, C5-propynyl-uridine, C5 -propynyl-cytidine, C5- methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8- oxoguanosine, 0(6)-methylguanine, 2-thiocytidine, methylated bases, intercalated bases, and combinations thereof). In some embodiments, a nucleic acid comprises one or more modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose) as compared with those in natural nucleic acids. In some embodiments, a nucleic acid has a nucleotide sequence that encodes a functional gene product such as an RNA or protein. In some embodiments, a nucleic acid includes one or more introns. In some embodiments, nucleic acids are preparedby one or more of isolation from a natural source, enzymatic synthesis by polymerization based on a complementary template (in vivo or in vitro), reproduction in a recombinant cell or system, and chemical synthesis. In some embodiments, a nucleic acid is at least 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 1 10, 120, 130, 140, 150, 160, 170 180, 190, 20, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000 or more residues long. In some embodiments, a nucleic acid is partly or wholly single stranded; in some embodiments, a nucleic acid is partly or wholly double stranded. In some embodiments a nucleic acid has a nucleotide sequence comprising at least one element that encodes, or is the complement of a sequence that encodes, a polypeptide. In some embodiments, a nucleic acid has enzymatic activity.

[0234] Peptide. The term “peptide” as used herein refers to a polypeptide that is typically relatively short, for example having a length of less than about 100 amino acids, less than about 50 amino acids, less than about 40 amino acids less than about 30 amino acids, less than about 25 amino acids, less than about 20 amino acids, less than about 15 amino acids, or less than 10 amino acids.

[0235] Polypeptide: As used herein, the term “polypeptide” refers to a polymeric chain of amino acids. In some embodiments, a polypeptide has an amino acid sequence that occurs in nature. In some embodiments, a polypeptide has an amino acid sequence that does not occur in nature. In some embodiments, a polypeptide has an amino acid sequence that is engineered in that it is designed and / or produced through action of the hand of man. In some embodiments, a polypeptide may comprise or consist of natural amino acids, non-natural amino acids, or both. In some embodiments, a polypeptide may comprise or consist of only natural amino acids or only non-natural amino acids. In some embodiments, a polypeptide may comprise D-amino acids, L-amino acids, or both. In some embodiments, a polypeptide may comprise only D-amino acids. In some embodiments, a polypeptide may comprise only L-amino acids. In some embodiments, a polypeptide may include one or more pendant groups or other modifications, e.g., modifying or attached to one or more amino acid side chains, at the polypeptide’s N-terminus, at the polypeptide’s C-terminus, or any combination thereof. In some embodiments, such pendant groups or modifications comprise acetylation, amidation, lipidation, methylation, pegylation, etc., including combinations thereof. In some embodiments, a polypeptide may be cyclic, and / or may comprise a cyclic portion. In some embodiments, a polypeptide is not cyclic and / or does not comprise any cyclic portion. Insome embodiments, a polypeptide is linear. In some embodiments, a polypeptide may be or comprise a stapled polypeptide. In some embodiments, the term “polypeptide” may be appended to a name of a reference polypeptide, activity, or structure; in such instances it is used herein to refer to polypeptides that share the relevant activity or structure and thus can be considered to be members of the same class or family of polypeptides. For each such class, the present specification provides and / or those skilled in the art will be aware of exemplary polypeptides within the class whose amino acid sequences and / or functions are known; in some embodiments, such exemplary polypeptides are reference polypeptides for the polypeptide class or family. In some embodiments, a member of a polypeptide class or family shows significant sequence homology or identity with, shares a common sequence motif (e.g., a characteristic sequence element) with, and / or shares a common activity (in some embodiments at a comparable level or within a designated range) with a reference polypeptide of the class; in some embodiments with all polypeptides within the class). For example, in some embodiments, a member polypeptide shows an overall degree of sequence homology or identity with a reference polypeptide that is at least about 30-40%, and is often greater than about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more and / or includes at least one region (e.g., a conserved region that may in some embodiments be or comprise a characteristic sequence element) that shows very high sequence identity, often greater than 90% or even 95%, 96%, 97%, 98%, or 99%. Such a conserved region usually encompasses at least 3-4 and often up to 35 or more amino acids; in some embodiments, a conserved region encompasses at least one stretch of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 or more contiguous amino acids. In some embodiments, a relevant polypeptide may comprise or consist of a fragment of a parent polypeptide.

[0236] Ribonucleotide. As used herein, the term “ribonucleotide” encompasses unmodified ribonucleotides and modified ribonucleotides. For example, unmodified ribonucleotides include the purine bases adenine (A) and guanine (G), and the pyrimidine bases cytosine (C) and uracil (U). Modified ribonucleotides may include one or more modifications including, but not limited to, for example, (a) end modifications, e.g., 5' end modifications (e.g. , phosphorylation, dephosphorylation, conjugation, inverted linkages, etc.), 3' end modifications (e.g., conjugation, inverted linkages, etc.), (b) base modifications, e.g. , replacement with modified bases, stabilizing bases, destabilizing bases, or bases that base pair with an expanded repertoire of partners, or conjugated bases, (c) sugarmodifications (e.g., at the 2' position or 4' position) or replacement of the sugar, and (d) intemucleoside linkage modifications, including modification or replacement of the phosphodiester linkages. The term “ribonucleotide” also encompasses ribonucleotide triphosphates including modified and non-modified ribonucleotide triphosphates.

[0237] Ribonucleic acid (RNA): As used herein, the term “RNA” refers to a polymer of ribonucleotides. In some embodiments, an RNA is single stranded. In some embodiments, an RNA is double stranded. In some embodiments, an RNA comprises both single and double stranded portions. In some embodiments, an RNA can comprise a backbone structure as described in the definition of “Nucleic acid / Polynucleotide” above. An RNA can be a regulatory RNA (e.g., siRNA, microRNA, etc.), or a messenger RNA (mRNA). In some embodiments where an RNA is an mRNA. In some embodiments where an RNA is an mRNA, an RNA typically comprises at its 3 ’ end a poly(A) region. In some embodiments where an RNA is an mRNA, an RNA typically comprises at its 5’ end an art-recognized cap structure, e.g., for recognizing and attachment of an mRNA to a ribosome to initiate translation. In some embodiments, an RNA is a synthetic RNA. Synthetic RNAs include RNAs that are synthesized in vitro (e.g., by enzymatic synthesis methods and / or by chemical synthesis methods). In some embodiments, an RNA is a single-stranded RNA. In some embodiments, a single-stranded RNA may comprise self-complementary elements and / or may establish a secondary and / or tertiary structure. One of ordinary skill in the art will understand that when a single-stranded RNA is referred to as “encoding,” it can mean that it comprises a nucleic acid sequence that itself encodes or that it comprises a complement of the nucleic acid sequence that encodes. In some embodiments, a single-stranded RNA can be a self-amplifying RNA (also known as self-replicating RNA).

[0238] Sequence Data, Sequence Representation. As used herein, the terms“sequence data” and / or “sequence representation” are used to refer representations of biological sequences - such as amino acid sequences and nucleic acid sequences - for use in connection with computer modeling techniques - i.e., particular data representations of biological sequences. Sequence representations may take a variety of forms, including, for example, a text string comprising a plurality of individual characters (letters) and / or short (e.g., three letter code) sub-strings, each representing a particular amino acid and / or nucleic acid. For example, sequence data representing an amino acid sequence of a protein or peptide (e.g., an amino acid sequence representation) may be a text string comprising a plurality of individual characters (e.g., alphabetical characters, e.g., alphanumeric characters),each individual character representing a particular amino acid. In certain embodiments, each individual character of an amino acid sequence representation may correspond to and represent a particular one of the twenty standard amino acids. In certain embodiments, an amino acid sequence representation may use twenty-one types of individual characters, such as twenty for the twenty standard amino acids and a twenty-first, wildcard character (e.g., an “X”, other special letter or symbol, etc.) for an unknown and / or non-standard amino acid. In certain embodiments, capitalization may be used to differentiate between standard amino acids in their / .-configuration (e.g., via uppercase letters) and the corresponding amino acid in a D-configuration (e.g., via lowercase letters). For example, sequence data representing a nucleic acid sequence (e.g., a nucleic acid sequence representation) may be a text string comprising a plurality of individual characters (e.g., alphabetical characters, e.g., alphanumeric characters), each individual character representing a particular nucleic acid (e.g., a DNA sequence, an RNA sequence, etc.).

[0239] Shared tumor antigen . As used herein, the term “shared tumor antigen” refers to a tumor antigen expressed by a large fraction of cancers. In some embodiments, a “shared tumor antigen” is a tumor antigen expressed by a large fraction of cancers of the same type and / or a large fraction of cancers of different types. In some embodiments, a “shared tumor antigen” is a tumor antigen shared by a large fraction of different subjects having the same cancer type and / or different cancer types. With reference to “shared tumor antigen”, the term “large fraction” refers to at least 15%.

[0240] Shared tumor antigen epitope. As used herein, the term “shared tumor antigen epitope” refers to an epitope of and / or derived from a shared tumor antigen.DETAILED DESCRIPTION

[0241] It is contemplated that systems, architectures, devices, methods, and processes of the claimed invention encompass variations and adaptations developed using information from the embodiments described herein. Adaptation and / or modification of the systems, architectures, devices, methods, and processes described herein may be performed, as contemplated by this description.

[0242] Throughout the description, where articles, devices, systems, and architectures are described as having, including, or comprising specific components, or where processes and methods are described as having, including, or comprising specific steps, it iscontemplated that, additionally, there are articles, devices, systems, and architectures of the present invention that consist essentially of, or consist of, the recited components, and that there are processes and methods according to the present invention that consist essentially of, or consist of, the recited processing steps.

[0243] It should be understood that the order of steps or order for performing certain action is immaterial so long as the invention remains operable. Moreover, two or more steps or actions may be conducted simultaneously.

[0244] The mention herein of any publication, for example, in the Background and References sections, is not an admission that the publication serves as prior art with respect to any of the claims presented herein. The Background section is presented for purposes of clarity and is not meant as a description of prior art with respect to any claim.

[0245] Documents are incorporated herein by reference as noted. Where there is any discrepancy in the meaning of a particular term, the meaning provided in the Definition section above is controlling.

[0246] Headers are provided for the convenience of the reader - the presence and / or placement of a header is not intended to limit the scope of the subject matter described herein.A. Immune Response Predictions for Disease Prevention and / or Treatment

[0247] In certain embodiments, technologies of the present disclosure include in- silico techniques for predicting immune response. Immune response prediction technologies of the present disclosure may use and / or include machine learning models (e.g., language models) and / or various other computational models. Among other things, as described in further detail herein, these computational techniques may be used to evaluate (e.g., overall) immunogenicity of, and / or various characteristics of a host immune response to, particular pathogens, antigens, antigen portions - such as sub-regions / units, particular epitopes, etc. - and the like. In certain embodiments, immune response prediction technologies of the present disclosure may be used to determine scores that measure results of various steps in the biological pathways that constitute a (e.g., overall) host immune response (e.g., steps such as peptide cleavage, MHC binding, presentation, B- and T-cell responses thereto, etc.) to a particular antigen or portion thereof. These computational techniques, and / or scoresgenerated by them, may be used separately and / or in combination to accurately predict results (e.g., binding affinity, efficacy of immune response. Accurate immune response prediction may, for example, be used to build a digital twin of the immune system (e.g., to be used for monitoring, predictions, and assessment of the immune system), develop personalized treatments, such as cancer vaccines and viral vaccines, and others.

[0248] In the context of developing personalized treatments, the biological complexity of the overall host immune response, often coupled with heterogeneity and / or high evolution (e.g., mutation) rates associated with diseases such as cancer and viral infections, presents a major obstacle to identification and development of effective treatments, making them time consuming and costly process. By providing an accurate immune response prediction, the techniques of the present disclosure may improve efficacy, reduce time frames and associated costs of developing the related treatments.

[0249] In certain embodiments, immune response prediction comprises predicting particular steps in a biological pathway of a host immune response to an antigen. In certain embodiments, immune response prediction comprises predicting immune response using various tasks that are associated with processing via major histocompatibility complex (MHC) class I (MHC-I) and class II (MHC-II) antigen presentation pathways.

[0250] For example, immune response prediction technologies described herein may be used to model various steps along MHC-I antigen presentation pathways involved in presentation of portions of intracellular proteins, and involve steps such as cleavage (e.g., proteasomal cleavage; e.g., trimming via endoplasmic reticulum aminopeptidase (ERAP) enzymes) of intracellular proteins into peptides (e.g., 8 to 12 amino acids long), transport and / or binding of peptides to MHC-I molecules, thereby forming peptide-MHC complexes, and surface presentation of the resultant peptide-MHC complexes. Immunogenicity modelling may also include steps such as T-cell recognition of the presented peptide-MHC complexes, and T-cell immunogenic response (e.g., CD8+ T-cell response) induced by the recognition process. In certain embodiments, stability of peptide-MHC-I complexes may also be modeled and accounted for.

[0251] In certain embodiments, immune response prediction technologies described herein may be used to model various steps along MHC-II antigen processing and presentation pathways. MHC-II processing, and presentation is relevant for intercellular proteins recognition by the immune system and resultant responses thereto, for example via CD4+ T-cells. For example, MHC-II molecules are expressed on surfaces of antigen-presenting cells (APCs) such as dendritic cells. MHC-II pathways are responsible for presenting peptides to CD4+ cells (e.g., also known as T helper cells) and, without wishing to be bound to any particular theory, impact activation and regulation of other immune cells. MHC-II molecules are heterodimers comprising an a and a > chain, both of which take part in formation of the MHC-II binding groove that interacts with and influences binding by candidate peptides for presentation. In humans, these chains are encoded by three loci, referred to as HLA-DR, HLA-DP, and HLA-DQ, which exhibit a high degree of polymorphism and thereby producing a large spectrum of MHC-II peptide binding specificities. MHC-II molecules tend to accommodate binding to peptides with longer, and a greater degree in variability of, lengths than those bound by MHC-I molecules. For example, peptides bound by MHC-II molecules may range in length from twelve (12) to fifteen (15) amino acids, often comprising a continuous binding core of about nine amino acids and flanking residues. Accordingly, MHC-II antigen processing steps modeled via various approached described herein may include cleavage, MHC-II binding, surface presentation, and subsequent T-cell recognition. In certain embodiments, stability of peptide-MHC-II complexes may also be modeled and accounted for.

[0252] In certain embodiments, various steps described herein (e.g., above) may be modeled in combination, for example depending on and / or expressly to capture particular sources of experimental data. For example, as described in further detail herein, on the one hand, certain experiments, such as binding affinity measurements, may be used to produce data that isolates and represents a single step (e.g., peptide-MHC binding) associated with an MHC-I and / or MHC-II antigen presentation pathway. On the other hand, in certain cases, experimental data that results from an interplay between multiple antigen processing and presentation steps may be used. In particular, mass-spectrometry (MS) eluted ligand (EL) data is of particular relevance as data set of peptides that are presented by MHC-I and / or MHC-II molecules. MS data, among other things, identifies peptides that are successfully cleaved, bound, and presented by particular MHC-I and / or MHC-II alleles and, accordingly, captures an interaction / interplay between multiple (e.g., cleavage, binding, and presentation) steps along an MHC-I and / or MHC-II antigen processing and presentation pathway. In certain embodiments, as explained in further detail herein, such experimental data may be augmented via decoy generation processes that serve to create a modified dataset that isolates, and can be used to train machine learning models to predict, particular signalsassociated with particular individual tasks, such as cleavage, binding, and presentation, e.g., independently.

[0253] In this manner, technologies of the present disclosure aim, and may be used to model, various steps along MHC-I and / or MHC-II antigen processing pathways, as well as particular combinations thereof. These, provided, methods and systems may, accordingly, be used in the context of various disease diagnostics and / or treatment technologies, to predict various facets of, or overall, immunogenicity of candidate peptides, thereby allowing for those that are likely to induce an immune response to be identified and, for example, selected for use in vaccination strategies. Among other things, this ability to accurately model immunogenicity in-silico is of particular relevance, without limitation, to cancer treatments, such as targeted (e.g., individualized, tailored to particular patient cohorts, etc.) cancer vaccines, and vaccination against infectious diseases.A. i. Cancer Treatment

[0254] In certain embodiments, immune response prediction technologies of the present disclosure may be used to evaluate tumor antigens. For example, in-silico immune response prediction technologies described herein may be used to screen various candidate antigens and / or portions thereof, such as tumor antigen epitopes and neoantigens as targets for tailored treatments, such as (e.g., personalized) cancer vaccines. Antigen and / or antigen portions identified as targets in this manner may, for example, be used as targets for tailored / custom treatments, included in personalized cancer vaccines various forms. Tailored cancer treatment approaches that may be used in this manner, in combination with immune response prediction techniques of the present disclosure are described, for example, in U.S. Patent No. 10,738,355, entitled “Individualized Vaccines for Cancer,” and issued August 11, 2020, the content of which is incorporated by reference herein in its entirety.

[0255] In certain embodiments, tumor related antigens or portions thereof may be cancer-related epitopes (e.g., shared tumor antigen epitopes, neoepitopes). In certain embodiments, antigens may be neoantigens of a subject. Examples of neoantigens comprise human proteins resulting from one or more cancer-related genomic mutations (e.g., single nucleotide variants, insertions and deletions). In certain embodiments, mutations may occur in stability genes, mutations linked to hereditary risk of cancer, mutations linked to particularcancers, such as breast cancer, a colorectal cancer, an endometrial cancer, an ovarian cancer, a gastric cancer, a pancreatic cancer, a melanoma, a lung cancer, and a prostate cancer.

[0256] In certain embodiments, portions of antigens may comprise neoantigen epitopes and / or neoepitopes. In certain embodiments, neoantigen epitopes comprise cancerspecific somatic mutations of a cancer mutation signature [e.g., identified in a tumor specimen of a subject (e.g., using next generation sequencing)]. In certain embodiments, neoantigen epitopes may comprise one or more sequence differences identified between (1) a genome, exome, and / or transcriptome of a tumor specimen of a subject and (2) a genome, exome, and / or transcriptome of a non-tumor specimen. In certain embodiments, the nontumor specimen is from the subject.

[0257] In certain embodiments, portions of antigens may comprise shared tumor antigen epitopes (e.g., epitopes that are commonly expressed by tumors) and / or non- neoepitopes. In certain embodiments, shared tumor antigen epitopes comprise at least one shared tumor antigen epitope expressed by at least 15% of subjects having the same cancer type. In certain embodiments, shared tumor antigen epitopes comprise epitopes from one shared tumor antigen.

[0258] In certain embodiments, immune response predictions may be used to develop treatments (e.g., that induce engineered immune response). Using cancer as an example, a tumor sequence data (e.g., tumor genome and / or exome; e.g., comprising patient-specific tumor mutations) may be used to develop a treatment with a specific polypeptide sequence, where the treatment, in turn, induces a predicted immune response (e.g., to aid with specific disease or condition). Polypeptides encoding identified candidate peptides as described herein (e.g., corresponding to or comprising one or more epitopes) may be administered as peptides and / or as nucleic acids encoding candidate peptides, such as in RNA form.A. ii. Infectious Diseases

[0259] In certain embodiments, antigens may be a protein of (e.g., produced by) an infectious agent and / or a portion thereof. In certain embodiments, antigens may comprise a viral protein, a variant of a particular viral protein, or a portion thereof. In this context, the predicted immune responses using technologies of the present disclosure may be used to develop a (e.g., personalized) viral vaccine (e.g., to address rapid evolution of viral variants).

[0260] In certain embodiments, a viral protein may originate from (e.g., be a mutated version of) viruses with high mutation rates (e.g., RNA viruses). Such viruses include, without limitation, influenza, coronavirus (e.g., severe acute respiratory syndrome-related coronavirus), human immunodeficiency virus (HIV), Respiratory syncytial virus (RSV), and the like. For example, in the context of the recent SARS-CoV 2 pandemic, mutation of circulating virus has given rise to tens of thousands of viral variants, such as Omicron and recently emergent XBB (e.g., XBB.1.5).

[0261] In certain embodiments, an infectious agent is or comprises a coronavirus, such as SARS-CoV 2 and an antigen thereof is a SARS-CoV 2 Spike protein, or a portion of the SARS-CoV 2 Spike protein. In this context, an antigen may be, for example, an entire Spike protein, or a particular portion of the Spike protein, such as an N-terminal region (e.g., an N-Terminal Domain (NTD)) or a Receptor Binding Domain (RBD).

[0262] In certain embodiments, immune predictions of the present disclosure aim to design engineered versions of an antigen that, when introduced (e.g., administered) to a subject, encourage their immune system to generate new antibodies that are expressly tailored to particular epitopes of the antigen. In certain embodiments, this involves reducing a likelihood and / or an extent to which an engineered version of an antigen will trigger a memory immune response, thereby ameliorating certain obstacles immune imprinting phenomena can cause in regard to, e.g., vaccination.B. AI-Based Immune Response Prediction

[0263] Immune response prediction technologies of the present disclosure utilize machine learning models to generate predictions of various antigen processing, presentation, and recognition steps based on data provided as input, such as information about a candidate peptide, its context in an overall protein and / or expressed gene, MHC allele information (e.g., HLA types) associated with a particular patient and / or cohort, and the like.B.i. Model Inputs and Encoding Antigen Presentation Features

[0264] Turning to FIG. 1, machine learning models 102 may receive various inputs on which to base predictions of immune response behavior, for example depending on particular tasks and / or combinations of tasks being modeled.

[0265] For example, in certain embodiments, a machine learning model 102 may receive, as input, a candidate peptide representation 104 that represents a biological sequence of a particular candidate peptide to be evaluated (e.g., scored). The candidate peptide representation may be or comprise sequence data, such as a string of characters representing a sequence of amino acids that make up the candidate peptide (e.g., an amino acid sequence) and / or a string of characters representing a nucleic acid sequence (e.g., a DNA sequence; an RNA sequence) encoding the candidate peptide.

[0266] In certain embodiments, a machine learning model 102 may receive a representation of flank residues 106 (“flank representation(s)”) as input. As illustrated in FIG. 2A, in certain embodiments, flanks 202 and 204 of a candidate peptide 206 are those residues that lie immediately upstream (e.g., left of the A-terminus of the candidate peptide) and / or immediately downstream (e.g., right of the C-terminus of the candidate peptide), referred to herein as a left flank 202 and a right flank 204, respectively. Flanks of a candidate peptide 206 may be identified and extracted from a sequence of a source protein 208, from which the candidate peptide is known and / or believed to originate, and / or, additionally or alternatively, from a source gene whose expression results in production of the source protein. Flank representations 106, may be a string of characters representing a biological sequence, such as an amino acid sequence and / or a nucleic acid sequence encoding the flanks. Flanks represented and used as input to machine learning models described herein may be of a variety of lengths, i.e., wherein a flank of length x comprises (e.g., includes only) a first x amino acids upstream and / or downstream of an N and / or C-terminus of a candidate peptide. In certain embodiments, a flank may have a length of one or more amino acids, two or more amino acids, three or more amino acids, four or more amino acids, etc. In certain embodiments, a flank may have a length of ten or fewer amino acids, five or fewer amino acids, four or fewer amino acids, three or fewer amino acids, etc. In certain embodiments, a flank may have a length of one amino acid, two amino acids, three amino acids, four amino acids, etc. In certain embodiments, left and right flanks may have a same length. In certain embodiments, left and right flanks may have different lengths.

[0267] In certain embodiments, a machine learning model 102 may receive, as input, representations of one or more MHC alleles 108. MHC alleles may be represented using textlabels, each identifying a particular MHC allele, such as a particular HLA allele 222. For example, in the context of MHC -I antigen presentation modeling, each HLA-A, HLA-B, HLA-C allele type may be assigned a label, such as HLA-A*03:01, 222, for example as illustrated in FIG. 2B. Likewise, on the context of MHC-II antigen presentation modeling, various labels of HLA-DP, HLA-DM, HLA-DO, HLA-DQ, and HLA -DR types may be used to identify particular HLA alleles from which, for example, training and / or test data originates, and / or of particular individual subjects (e.g., human patients) for which immunogenicity predictions are being performed. In certain embodiments, MHC alleles may be represented via their sequences and / or portions thereof, such as sub-sequences comprising those amino acids that form a binding groove for binding with peptides to be presented. In certain embodiments, sub-sequences may be pseudo-sequences comprising only certain amino acids known and / or believed to be in contact (e.g., and / or close proximity) with peptides during peptide-MHC binding and / or associated with T-recognition / receptors. As illustrated in FIG. 2B, a pseudo-sequence may be represented using an amino acid sequence representation. In certain embodiments, a pseudo-sequence may be represented using sequence data representing a nucleic acid sequence (e.g., encoding a particular MHC chain portion). A variety of approaches may be used for selecting particular amino acids of an MHC -I molecule to include in a pseudo-sequence representation and have been used in connection with various MHC -I modeling approaches, such as NetMHCPan, MHCFlurry, BigMHC, and DePTH. For example, in certain embodiments a particular number - e.g., 30, 32, 34, 36, 38, 40, etc. - amino acids from or about a binding groove of an MHC protein may be used and represented as a pseudo-sequence. In certain embodiments, 34 amino acid residues exclusively from an MHC-I protein binding groove are selected and represented as a pseudo-sequence. In certain embodiments, two additional amino acids are included. In certain embodiments, residues that may interact with either an antigen (e.g., candidate peptide) and / or a T-Cell receptor are included. In certain embodiments, machine learning techniques, such as attention mechanisms, may be used to identify (e.g., score) a subset of amino acids for inclusion pseudo-sequences representing particular MHC alleles.

[0268] In certain embodiments, both an allele name and one or more pseudosequences may be used to represent a particular MHC allele, e.g., as a concatenation of two strings 226, as illustrated in FIG. 2C. In certain embodiments, representations of multiple alleles may be received as input, for example, to allow for multi -allelic modeling. Multipleallele representations may be received as allele names 242, pseudo-sequences 244, and / or combinations thereof 246.

[0269] In certain embodiments, a sequence representation of a protein and / or portion thereof 110 may be input to a machine learning model. For example, a source protein sequence representation may be received as input and, as described in further detail herein, machine learning model 102 may identify one or more portions thereof as peptides likely to be presented a host antigen processing pathway, e.g., given particular HLA allele type(s) of the host. In certain embodiments, a source protein or portion thereof is identified from genetic sequencing data obtained for a subject, such as a particular gene identified as having a targetable mutation, such potentially giving rise to a neoantigen or a shared tumor antigen. In certain embodiments, a source protein or portion thereof may be associated with a pathogen, such as a viral protein or portion thereof, such as a particular sub-unit (e.g., a receptor binding domain of a SARS-CoV-2 spike protein).

[0270] In certain embodiments, various continuous features may be used as input to a machine learning model. These may include, for example, numerical values indicative of expression levels, a gene bias value representing a ratio comparing the expression of proteins to their expected expression levels estimated from RNA-sequencing and gene length.

[0271] In certain embodiments, continuous features, such as numerical values, may be represented in a text format, e.g., for use in connection with language models. Various approaches for representing continuous, e.g., numerical, values may be used, including, without limitation, text representing a floating-point number, and quantization strategies whereby various ranges of values are assigned (e.g., distinct) bins and a bin label received as input. In certain embodiments, binned values are binned into 5 bins, 10 bins, 20 bins, or 50 bins. In certain embodiments, bins may be unique. In certain embodiments, bins may be shared / overlap.

[0272] In certain embodiments, additional features may be used as input, such as a “shadow” label indicates the presence or absence in the peptide of interest, of a 9-mer substring derived from presented peptides identified in other studies, a label that identifies a particular type of immune response prediction task to be performed, a cell or cancer type that specifies a particular cancer type, and others, such as structural features, which may be derived, e.g., from a model such as ESMFold.

[0273] For example, in certain embodiments, a multi-task model uses prefixes (e.g., as input strings) to differentiate between single-allelic and multi-allelic data. In certain embodiments, a multi-task model uses prefixes (e.g., as input strings) to differentiate between training datasets. In certain embodiments, a prefix comprises of multiple tokens. In certain embodiments, at least one of multiple tokens is shared across tasks (e.g., “Binding” followed by “SA” or “MA”). In certain embodiments, every token is unique (e.g., “SA-Binding” or “MA-Binding” tasks).

[0274] In the context of text-to-text language models described herein, which utilize all text input, these various features may, as with others, be represented via text strings for use as input to a machine learning model 102.B.ii. Immune Response Prediction Outputs

[0275] Machine learning models of the present disclosure may, based on received input, generate a variety of outputs, representing immune response scores and / or identified peptides that are likely to be immunogenic.

[0276] For example, in certain embodiments, a machine learning model 102 may generate, as output, a continuous value, such as a numerical value representing a binding affinity 116 or a likelihood (e.g., value between 0 and 1) of a particular peptide being cleaved 114, bound by a given MHC allele 116, or presented at a cell surface 118. In certain embodiments, for example in connection with text-to-text language models, continuous numerical values such as these may be represented via a text string, as described above with respect to continuous valued inputs. In certain embodiments, regression predictions may be converted from text back to continuous values for evaluation. For classification tasks, likelihood predictions may be estimated using the language model likelihood of the label sequence.

[0277] In certain embodiments, a binary label may be generated as output (e.g., and represented as text) indicating whether a particular candidate peptide, for example, is determined to be cleaved 114 for MHC presentation, bound to an MHC allele 116, presented at a cell surface 118, or likely immunogenic, or other immune response feature. Such binary labels may be represented via text in a variety of manners, such as a 0 or 1 (numericalcharacter), other binary set of characters (e.g., “T” or “F”, “P” or “N”, etc.) or strings (e.g., “positive” or “negative”, etc.).

[0278] In certain embodiments, a machine learning model 102 may generate, as output, peptide locations and / or sequences 119. For example, a peptide location may be a mask identifying various immunogenic peptides (as determined by the machine learning model) within a source protein sequence 110 received as input. This ‘peptide identification mask’ may, for example, have a same length as a source protein sequence 110 received as input and include 0’s and 1’s at various locations, with 0’s indicating amino acid residues not included in immunogenic peptides and 1’s identifying those amino acid residues making up immunogenic peptides (e.g., epitopes). In certain embodiments, a machine learning model 102 may generate peptide sequences - e.g., short strings - directly as output, given a particular source protein 110 received as input.B.iii. Large Language Models (LLM) for Immune Response Predictions

[0279] Machine learning models 102 used in connection with technologies of the present disclosure may utilize and implement a variety of machine learning techniques. For example, a machine learning model may be a deep learning model (e.g., an artificial neural network with one or more, e.g., plurality of, hidden layers), such as a language or large language model (LLM). I n certain embodiments, a machine learning model is or comprises one or more recurrent models, such long short-term memories (LSTMs), implemented alone or in combination, e.g., as in a bi-directional LSTM (bi-LSTM). In certain embodiments, a machine learning model comprises one or more transformer models. Examples of machine learning language models include, without limitation, evolutionary scale models (ESM), bidirectional encoder representations from transformers (BERT), and the like. In certain embodiments, LLMs may comprise one or more members selected from the group consisting of an autoregressive LLM, autoencoding LLM, encoder-decoder LLM, bidirectional LLM, fine-tuned LLMs, and multimodal LLMs.

[0280] In certain embodiments, a machine learning model of the present disclosure may be or comprise a model that processes input and generates output in a unified fashion, such as a text-based model, for example a Text-To-Text Transfer Transformer (T5) based model as described in (Raffel et al., 2020), the content of which is incorporated by reference herein in its entirety. Among other things, as described in further detail herein, the unifiedmanner in which a T5 or similar model process input and output facilitates multi-task learning, allowing for improved performance and / or simplified modeling of immune response tasks.

[0281] For example, FIG. 4 shows an example process 400 for predicting an immune response to a prospective antigen (e.g., an antigen polypeptide) using an LLM. At step 402, an input alphanumeric string of a candidate peptide is obtained. The candidate peptide may represent a biological sequence (e.g., RNA, DNA, amino acid) of at least a portion of the prospective antigen (e.g., a whole antigen; e.g., particular portions, such as individual epitopes). At step 404, an immune response prediction (e.g., as alphanumeric string) is determined based on the source peptide sequence using the language model. To this end, the model may receive the input alphanumeric string as input. The model may generate an output alphanumeric string. The output string may encode values of one or more immune response scores. Each immune response score may be associated with and / or represent a predicted result of a particular immune response task. At step 406, the immune response prediction is stored, provided for further processing, and / or displayed.

[0282] Machine learning language models may be trained, for example on protein data available from public and / or private repositories. Training may be accomplished by providing a machine learning model with input peptide sequence and tasking them to predict outputs associate with the immune response prediction and / or the related tasks. The outputs may be then evaluated against ground truth training data. Additional details of machine learning models and approaches for training them, such as particular techniques for training recurrent and / or transformer models, may be found, for example, in PCT publications WO 2022 / 235847 and WO 2022 / 235853, the content of each of which is incorporated by reference herein in its entirety.B. iv. Training Datasets and Decoy Generation

[0283] Machine learning models of the present disclosure may be trained for various immune response prediction tasks using a variety of data sources including datasets comprising examples from experimental sources as well as artificial training data, such as decoys, that can be designed to cause a model to learn certain types of signals or patterns.

[0284] Machine learning models used for immune response predictions may be trained, for example, using data available from public and / or private repositories, such as Internet Epitope Database (IEDB). Training may be accomplished by providing a machine learning model with input sequences and evaluating its output against ground truth (e.g., using a variety of metrics). In certain embodiments, training comprises hyperparameters tuning. Examples of hyperparameters are embedding size, MLP hidden size, layer number, heads, head size may be adjusted as well.

[0285] For example, in certain embodiments, immune response data may include binding affinity data, which results from binding assays and provides binding affinities for various peptides and MHC alleles. Binding affinity data may be obtained from public repositories, such as the immune epitope database (IEDB, https: / / www.iedb.org / ), as well as in-house (e.g., proprietary) experimental assays.

[0286] Additionally or alternatively, datasets used for training and evaluating machine learning models of the present disclosure may include eluted ligand mass spectrometry data, which is generated by extracting peptides bound to MHC molecules and sequencing them via mass spectrometry (MS). While binding affinity data isolates a particular single step in an antigen processing and presentation pathway - peptide-MHC binding - peptides identified as eluted ligands via MS have been expressed, cleaved, bound to MHC molecules, and presented. Accordingly, eluted ligand MS data captures interplay between multiple steps.

[0287] Due to natural presence of multiple MHC alleles, MS eluted ligand data may be multi -allelic, such that the particular MHC allele responsible for binding and presenting an identified peptide may not necessarily be known initially. In certain embodiments, specialized experimental techniques may be used to create single allele MS eluted ligand data in which identified peptides are known to have interacted with a particular, known, MHC allele. Examples of approaches for producing single-allelic data include, without limitation, using (i) allele-specific antibodies for pulling down pMHC complexes, (ii) cells that have been transduced with retroviral vectors to express only one HLA allele (Abelin et al., 2017), or (iii) a technique referred to as MAPTAC, wherein an allele is knocked in with special sequence at the C-terminus that allows it to be specifically pulled down (Abelin et al., 2019).

[0288] Multi -allelic data may be deconvolved, in order to identify a single allele responsible for peptide binding and presentation (Reynisson et al., 2020). Variousapproaches may be used for deconvolution. For example, in certain embodiments, the tool NNAlign_MA tackles this problem by iteratively (i.e., epoch after epoch, after a pre-training using SA data only) annotating the best single-allele to the MA data during the model training. (Alvarez et al., 2019). In certain embodiments, a multiple instance learning (MIL) approach may be used, where a particular sample is sample is a "bag" of instances, and each bag is associated with a bag label. The instances are unordered, independent and have unknown individual labels. BertMHC uses an instance-level classifier with a max-pooling operator to select the highest scoring allele as prediction. More generally, since a MIL model must be permutation-invariant, it can be approached using DeepSets (Zaheer et al., 2017), which allows the bag probability to be decomposed into instance representations and an appropriate pooling operator. The latter can be a max- or a mean-pooling. An attention-based pooling (Tarkhan et al., 2022) can also be applied to fully parameterize the bag probabilities. Here, MHC -peptide instances (the possible combinations of the MA sample) are first embedded and then pooled, effectively annotating a single MHC restriction while training, e.g., using an offline approach, iteratively, or via a multiple instance learning approach.

[0289] In certain embodiments, training datasets may be augmented by introducing artificial decoys. For example, MS eluted ligand data provides sequences of peptides that were identified as having been cleaved, bound, and presented successfully by MHC molecules. These experimentally determined example peptides are, accordingly, positive examples. Decoys, accordingly, may be generated to provide negative examples for purposes to of training a machine learning model, and to isolate particular steps - such as cleavage, binding, or presentation - from the multiple step interaction that results in positive ‘hits’ in MS data.

[0290] For example, in certain embodiments, distinct decoy generation processes may be used to create datasets with positive hits from MS data, and decoys that minimize correlations between cleavage and binding signals, thereby allowing a machine learning model to learn features indicative of cleavage or binding, separately.

[0291] For example, in certain embodiments, decoys for purposes of learning cleavage predictions may be generated to closely resemble positive hits in terms of certain features, determined to be relevant for binding, but differ primarily in those features determined to be relevant for cleavage. For example, a decoy may be generated from an identified hit by keeping certain amino acids the same, and varying others. In certain embodiments, for example, for a particular identified hit, a source protein may be identifiedand flanks for the particular identified hit determined. Decoys may, accordingly, be generated by altering flanks and holding at least a portion of other amino acids, e.g., associated with the peptide bound to the MHC molecule, constant. For example, in certain embodiments, one or more anchor amino acid positions may be identified and held constant. Anchor amino acids may include one or more of an A-tcnninus. JV+1, C-terminus, and C-l. In certain embodiments, flanks and / or other amino acids may be altered via random sampling to generate decoys. In certain embodiments, decoys may be extracted from a source protein and / or similar proteins (e.g., in related genes), for example by searching for subsequences with identical anchor amino acids, but differences elsewhere. In certain embodiments, decoys may be checked to ensure they do not correspond to peptides that also bind to the MHC allele. In certain embodiments, decoys may be selected based on expression, e.g., to prioritize more likely expressed decoys.

[0292] In certain embodiments, decoys may be generated using a dedicated binding predictor model, for example trained on binding affinity data. In this manner, peptides from a source protein associated with an identified hit peptide may be extracted and scored for binding, and those scored as strong binders (but not hits - hence, likely not cleaved). See, e.g., (O’Donnell at al., 2020), the content of which is incorporated by reference in its entirety.

[0293] In certain embodiments, decoys for binding prediction datasets based on eluted ligand data may be randomly extracted from source proteins that yield at least one identified hit peptide. In certain embodiments, decoys for presentation task datasets are randomly extracted from the proteins that are known to be expressed.

[0294] Accordingly, decoy generation procedures may be used to create training datasets that allow machine learning models to learn particular types of features and, accordingly, be trained on particular tasks.

[0295] In certain embodiments, a multi-task model is trained on a particular number / mixture of training samples. The particular number of training samples may comprise a set of training samples categories, wherein each training sample category corresponds to a particular immune response task. A set of training sample categories may have a specific number and / or ratio of samples (e.g., with respect to other categories) in every category.

[0296] In certain embodiments, a multi-task model is trained on a specific number and / or ratio (e.g., 1: 1, 1:2, 1 :5, 1: 10, 2: 1, 5: 1, 10: 1) of single-allelic and multi-allelic samples.

[0297] In certain embodiments, a multi-task model is trained only using human epitopes. In certain embodiments, a multi-task model is trained using human and non-human epitopes (e.g., viral) (e.g., using a specific ratio). In certain embodiments, a multi-task model is trained using non-human (e.g., viral) epitopes.

[0298] In certain embodiments, allele resampling (e.g., smoothing) is performed for a training dataset of a multi-task model. A more even distribution of allele frequencies during training may improve performance of a model on rare alleles.

[0299] Various metrics may be used to assess performance of a machine learning model and may, optionally, be used to select particular models having, e.g., particular hyperparameter values and / or particular input / output encoding strategies. For single -allelic data, where a binding score is produced by a machine learning model, such evaluation metrics as Top-K and Average Precision may be used. For multi-allelic data and ligand location, where a location and / or sequence of high-binding -likelihood epitopes with the protein are produced, such evaluation metrics as precision, recall, accuracy and Frank Score may be used.

[0300] Top-K scores may be determined as a fraction of true positives in the top k peptides as ranked by a model, where k is the number of known positive binders in the evaluation dataset (also known as R-Precision or Positive Predicted Value (PPV)). A Top-K score is therefore equivalent to both precision and recall at a particular k value.

[0301] Average Precision (AP) computes the average of the precision scores for every possible recall value (and hence every possible threshold). More formally, AP is defined as the area under the precision-recall curve obtained by varying the threshold (M. Zhu et al. 2004).

[0302] A metric called Frank score measures where the model ranks a known binding peptide amongst a set of decoy peptides generated from the source protein of the binder as defined in V. Jurtz et al. (2017). Formally, Frank score is defined as the fraction of the decoys which score higher than the binding peptide, where the decoys are generated by applying sliding windows of lengths 8 to 14 across the source protein (excluding the binding peptide itself). A “perfect” Frank score of zero indicates that the model ranks the binder above all of the decoys. Given a benchmark dataset of binding peptides and their associated source proteins, a Frank score is calculated for each binder and the mean and median scores as well as the number of perfect scores are reported.

[0303] The task of ligand prediction involves generating a set of ligands for T-cell recognition, after a cell that expresses a particular MHC-I allele processes an antigen sequence. The model’s predictive performance is evaluated by classifying each predicted ligand into one of three categories:• True Positive (TP): A predicted ligand that matches a target ligand.• False Positive (FP): A predicted ligand that does not match any target ligand.• False Negative (FN): A target ligand that is not predicted by the model.

[0304] To compare a predicted ligand with a target ligand, Windowed Exact Match (WEM) function may be used. Unlike directly comparing two ligands for equality, WEM checks if a target ligand falls within a window of ±5 amino acids around a predicted ligand in the antigen sequence. For example, for an antigen sequence and desired target ligand as shown below:Sequence: MYYKFSGFTQKLAGAWASEAYSPQGLKPVVS DSTVYDTarget: FTQKLAGAW on one hand, a predicted ligand (Prediction 1) FSGFTQKLAGAW would yield a WEM = 1.0 (since YKFSGFTQKLAGAWASEAY 9 FSGFTQKLAGAW. On the other hand, a predicted ligand (Prediction 2) of EAYSPQGLKPV would have a WEM = 0.0 (since YKFSGFTQKLAGAWASEAY 2 EAYSPQGLKPV).

[0305] After classifying each ligand, the number of TPs, FPs and FNs for each pair of allele and antigen sequence may be counted in a dataset and four classification metrics may be computed:• Precision = Z(TP) / (Z(TP) + Z(FP))• Recall = Z(TP) / (Z(TP) + Z(FN))• Accuracy = Z(TP) / (Z(TP) + Z(FN) + Z(FP))• Fl -score = 2 x (Precision x Recall) / (Precision + Recall).B.v. Multi-Task Machine Learning Models

[0306] In certain embodiments, machine learning model used to predict immune response to an antigen and / or a plurality of immune response tasks comprises a multi-task machine learning model. In certain embodiments, a multi-task machine learning model mayproduce values associated with predicting a plurality of immune response tasks (e.g., particular steps in a biological pathway of a host immune response to antigen and / or pathogen). In certain embodiments, a multi-task machine learning model is used for end-to- end prediction of immune response.

[0307] In certain embodiments, a multi-task machine learning model is trained on a plurality of immune response tasks. In certain embodiments, a multi-task model is trained on a plurality of immune response tasks, but may be used to predict scores associated with a single immune response task. For example, a multi-task trained model may benefit from transfer learning between tasks and, as such, may perform better on a single task as compared to models that were trained only a single task.

[0308] FIG. 5 shows an example process 500 for predicting an immune response to an antigen using a multi-task machine learning model. At step 502, an input alphanumeric string of a candidate peptide sequence is obtained. The candidate peptide sequence may represent a biological sequence (e.g., RNA, DNA, amino acid) of at least a portion of the antigen (e.g., a whole antigen, specific portions, individual epitopes). At step 504, a value of a selected immune response score based on the source peptide sequence is determined using the machine learning model. The machine learning model may be trained to determine values for a plurality of different immune responses scores, where each score is associated with and represents a predicted result of one of the immune response tasks. The selected immune response score may be one of the plurality of different immune response score options. At step 506, the value of the selected immune response score is stored, provided for further processing, and / or displayed.

[0309] FIG. 6 shows an exemplary method 600 of training a multi-task machine learning model for predicting immune responses. At step 602, multiple datasets are obtained. Each dataset may be associated with a particular immune response task from a set of at least two different immune response tasks, and comprise a plurality of examples, each example comprising at least an example peptide sequence. At step 604, examples from the datasets are repeatedly selected and used to train the machine learning model to determine values of immune response scores for each immune response task associated with a particular dataset. As a result of such training, a multi-task machine learning model trained to generate predictions for multiple different immune response tasks is generated. At step 606, the multitask machine learning model is stored and / or provided for further processing (e.g., inference predictions based on new input source peptide sequences).

[0310] In certain embodiments, a multi-task machine learning model may be or comprise a language model, such as Text-To-Text Transfer Transformer (T5). Use of a language model may include approaches for converting all tasks into “text-to-text” format, allowing the model to perform different tasks with a shared approach. The shared approach of a multi-task model may allow improvements on tasks where a limited training is performed (e.g., or available) by benefiting from knowledge acquired from tasks where a more substantial training is performed (e.g., or available). As such, multi-task modeling approaches can address challenges associated with limited training data availability. The available training data may be not uniformly distributed between various immune response tasks (e.g., some tasks may have limited training data available).

[0311] For example, single-task models may rely on specialized composite (e.g., multi-step) models, which may be challenging to scale and fail to comprehensively model intricate biological processes, such as T-cell recognition and response. Multi-task models may demonstrate better performance as compared to single-task models, due to, for example, transfer learning. Multi-task models may also simplify modeling process, introducing end-to- end modeling instead of multi-step process needed for single-task models. The unified task format may also enable the same model architecture, weights, loss function, hyper-parameters and training procedure to be used across tasks. Multi-task models may also lead to new insights in immune response predictions.

[0312] For instance, current technologies of HLA class I (HLA-I) and HLA class II (HLA-II) antigen presentation modeling are limited in their ability to capture information across the various steps in the HLA processing pathways. Generally, these technologies are trained to model peptide-MHC binding alone (B. Reynisson et al., 2020; X. Shao et al., 2020) or they are extended to model antigen processing and surface presentation using composite modelling (B. Bulik- Sullivan et al. 2019; S. Sarkizova et al. 2020; T. O’Donnell et al. 2020; R. Pyke et al. 2021). For example, Pyke et al. introduce SHERPA for modelling binding and presentation prediction using multi-stage gradient boosting decision trees: First a single- allelic binding predictor is trained, its outputs are used to deconvolute multi-allelic training data which is then used to train a second binding predictor for both single-allelic and multi- allelic inputs, and then finally integrating these models’ outputs with additional antigen processing features to train a presentation predictor.

[0313] These models are trained separately (z.e., not end-to-end), and may be implemented in a stacked / serial fashion, as a pipeline. A drawback of this approach,however, is that errors made early in the pipeline will couple, reducing downstream accuracy. Furthermore, in a single task, pipeline, format it is not possible for the training signal at each stage to benefit learning in other stages of the pipeline. Since peptide-MHC interaction is modelled by only the binding task, these features do not directly interact with expression and cleavage preference in the presentation task. Finally, limited work has been done to investigate powerful neural network algorithms for the complete immune response prediction pipeline due to low data availability for tasks, such as surface presentation and T-cell recognition. Allowing the various stages involved in predicting immune response to interact and share training signals can be important to reduce coupling errors, increase overall data availability and improve generalization.C. Example Immune Response Prediction Tasks

[0314] Machine learning models of the present disclosure may be used to predict a variety of immune response tasks, which may correspond to particular, individual steps in antigen processing and presentation pathways and / or combinations thereof. In certain embodiments, various immune response tasks described herein are designated to correspond to, and leverage, particular types of data, such as binding affinity data and MS eluted ligand data, which, in turn, may be single allelic or multi -allelic. Machine learning models used in connection with immune response prediction technologies of the present disclosure, as described above, may be used to generate various (e.g., a plurality) immune response scores, each immune response score is associated with a particular immune response task, or, additionally or alternatively, may be trained on one or more (e.g., a plurality of) immune response prediction tasks (e.g., any of those described herein), for example to optimize accuracy one or more particular tasks and / or an overall immune response prediction.C. i Peptide Cleavage

[0315] In certain embodiments, a machine learning model may be trained and / or used to generate (e.g., proteasomal) cleavage predictions. In this manner, an immune response score generated, e.g., as output, by a machine learning model may be or comprise a cleavage score that represents a prediction whether a particular candidate peptide is likely be produced result from proteasomal cleavage of an antigen, such as a source protein. In certainembodiments, to generate a cleavage prediction, a machine learning model may, accordingly, receive sequence representations of a candidate peptide and flanks (e.g., left and / or right flanks) as an input, and generate, as output, a cleavage score. A cleavage score may be represented as a binary, binned, tokenized, and / or continuous value.

[0316] Machine learning models may be trained and / or used to generate cleavage predictions for peptides associated with MHC-I presentation and / or MHC-II presentation.C.ii Peptide-MHC Binding

[0317] In certain embodiments, a machine learning model may be trained and / or used to generate peptide-MHC binding predictions, for example, such that an immune response score generated, as output, by a machine learning model may be or comprise a peptide-MHC binding score (e.g., an MHC binding affinity score) associated with the peptide-MHC binding task. A peptide-MHC binding score may represent a predicted binding affinity between the source peptide and a particular MHC allele. A peptide-MHC binding task may receive a peptide sequence and an MHC sequence as inputs, producing the peptide-MHC binding score as output. A peptide-MHC binding score may comprise a binary, binned, continuous, tokenized, and / or binding affinity value for a given allele of the MHC. A peptide-MHC binding score may comprise a ranked list of alleles of the MHC according to their probability of binding the peptide.C. Hi MHC Presentation

[0318] One of the pluralities of immune response scores may be a peptide-MHC presentation score (e.g., a surface presentation score) associated with the peptide-MHC presentation task. A peptide-MHC presentation score represents a prediction of whether the source peptide will be presented at a surface via a particular (MHC) allele (e.g., an MHC class I allele). The peptide-MHC presentation task may receive a peptide sequence and an MHC sequence as inputs, producing the peptide-MHC presentation score as output. The input may comprise additional parameters. The additional parameters may comprise flanking residues (e.g., amino acid sequences from the antigen in the local area to the left and the right of the peptide; e.g., binary, binned, continuous, tokenized), expression (e.g., values associated with expression levels of a protein that the peptide maps to; e.g., binary, binned, continuous,tokenized), gene bias (e.g., values associated with transcription of the peptide; e.g., a continuous value computed using source transcripts of the peptide and measured number of peptides observed per gene relative to expected number based on gene length and expression; e.g., binary, binned, continuous, tokenized), tissue (e.g., values associated with gene expression levels in specific tissues; e.g., binary, binned, continuous, tokenized), and / or shadow (e.g., values associated with a presence of preferred processing regions of proteins; e.g., a binary value checking whether the peptide contains any 9-mer in a list of “shadow” 9- mers which are found to indicate preferred processing regions of proteins; e.g., binary, binned, continuous, tokenized). The peptide-MHC presentation score may comprise a binary, binned, continuous, and / or tokenized, for a given allele of the MHC. The peptide-MHC presentation score may comprise a ranked list of alleles of the MHC according to their probability of being presented at a cell surface.

[0319] The MHC sequence may comprise identification of at least a portion of an allele or a plurality of distinct MHC alleles. The allele identification may comprise a label identifying an allele and / or a label identifying allele location. The output may comprise an output string that encodes a plurality of distinct sets of allele-specific immune response score values, each of corresponding to a particular one of the plurality of distinct MHC alleles. The immune response score values may be determined based on the plurality of distinct sets of allele-specific immune response score values.C. iv. Direct Ligand Identification and / or Generation

[0320] Machine learning models of the present disclosure may be used for identifying (e.g., locating) antigen fragments that will be presented as ligands, e.g., for T-cell recognition as shown in FIG. 7. At step 702, sequence data of an antigen or portion thereof is obtained. The sequence data may comprise (i) an antigen sequence representing a biological sequence (e.g., RNA, DNA, amino acid) of at least a portion of a target antigen (e.g., a whole antigen, specific portions, individual epitopes) and (ii) a MHC sequence representing a biological sequence (e.g., RNA, DNA, amino acid) of at least a portion of a particular MHC allele (e.g., a whole allele, specific portions). At step 704, at least one subregion of the antigen sequence is identified (e.g., via start and end locations, via a location mask) using a machine learning model, locating, within the antigen, fragment(s) to be presented as ligand(s) for T-cellrecognition. At step 706, a representation of the located ligand is stored, provided for further processing, and / or displayed.

[0321] Machine learning models of the present disclosure may be used for determining (e.g., generating) peptide fragments that will be presented as ligands, e.g., for T- cell recognition as shown in FIG. 8. At step 802, a sequence data is obtained. The sequence data may comprise (i) an antigen sequence representing a biological sequence (e.g., RNA, DNA, amino acid) of at least a portion of a target antigen (e.g., a whole antigen, specific portions, individual epitopes) and (ii) a MHC sequence representing a biological sequence (e.g., RNA, DNA, amino acid) of at least a portion of a particular MHC allele (e.g., a whole allele, specific portions). At step 804, at least one ligand sequence is determined (e.g., as an alphanumeric representation) using a machine learning model. The at least one ligand sequence corresponds to a subregion(s) of the antigen sequence identified as a fragment(s) of the antigen to be presented as a ligand(s) for T-cell recognition. At step 806, the ligand sequence is stored, provided for further processing, and / or displayed.C.v. Additional Immune Response Tasks

[0322] Examples of various immune response tasks, including those described above as well as others, are listed Table 1 along with their respective inputs and outputs.

[0323] In certain embodiments, a model input for one plurality of tasks may also be an output of the model for another plurality of tasks. For example, such values as at least a part of a peptide sequence, at least a part of a protein sequence, a protein structure information, a gene expression, a cleavage, a presentation, a binding, a binding affinity, an allele may serve both as inputs and outputs for various tasks in a given model. In certain embodiments, a model may be trained on a plurality of tasks, where at least one input in one plurality of tasks serves as an output in another plurality of tasks. For example, as shown in Table 1, while a binding result serves as an output for a binding classification task, it serves as an input for binding ligand and binding MHC tasks. Such task inversion capability may lead to better results as a model gets access to bigger training sets.Table 1: Examples of immune response tasks with their inputs and outputs for a multi-task model.D. Software, Computer System, and Network Environment

[0324] Certain embodiments described herein make use of computer algorithms in the form of software instructions executed by a computer processor. In certain embodiments, the software instructions include a machine learning module, also referred to herein as artificial intelligence software. As used herein, a machine learning module refers to a computer implemented process (e.g., a software function) that implements one or more specific machine learning algorithms, such as an artificial neural network (ANN), random forest, decision trees, support vector machines, and the like, in order to determine, for a given input, one or more output values. In certain embodiments, the input comprises alphanumeric data which can include numbers, words, phrases, or lengthier strings, for example. In certain embodiments, the one or more output values comprise values representing numeric values, words, phrases, or other alphanumeric strings. In certain embodiments, the one or more output values comprise an identification of one or more response strings (e.g., selected from a database).

[0325] In certain embodiments, machine learning modules implementing machine learning techniques are trained, for example using datasets that include categories of data described herein. Such training may be used to determine various parameters of machine learning algorithms implemented by a machine learning module, such as weights associated with layers in neural networks. In certain embodiments, once a machine learning module is trained, e.g., to accomplish a specific task such as identifying certain response strings, values of determined parameters are fixed and the (e.g., unchanging, static) machine learning module is used to process new data (e.g., different from the training data; e.g., infer a result) and accomplish its trained task without further updates to its parameters (e.g., the machine learning module does not receive feedback and / or updates). In certain embodiments,machine learning modules may receive feedback, e.g., based on automated review of accuracy or human user review of accuracy, and such feedback may be used as additional training data, to dynamically update the machine learning module. In certain embodiments, two or more machine learning modules may be combined and implemented as a single module and / or a single software application. In certain embodiments, two or more machine learning modules may also be implemented separately, e.g., as separate software applications. A machine learning module may be software and / or hardware. For example, a machine learning module may be implemented entirely as software, or certain functions of an ANN module may be carried out via specialized hardware (e.g., via an application specific integrated circuit (ASIC), field programmable gate arrays (FPGAs)).

[0326] In certain embodiments, machine learning modules implementing machine learning techniques may be composed of individual nodes (e.g., units, neurons). A node may receive a set of inputs that may include at least a portion of a given input data for the machine learning module and / or at least one output of another node. A node may have at least one parameter to apply and / or a set of instructions to perform (e.g., mathematical functions to execute) over the set of inputs. In certain embodiments, node instructions may include a step to provide various relative importance to the set of inputs using various parameters, such as weights. The weights may be applied by performing scalar multiplication (e.g., or other mathematical function) between a set of inputs values and the parameters, resulting in a set of weighted inputs. In certain embodiments, a node may have a transfer function to combine the set of weighted inputs into one output value. A transfer function may be implemented by a summation of all the weighted inputs and the addition of an offset (e.g., bias) value. In certain embodiments, a node may have an activation function to introduce non-linearity into the output value. Nonlimiting examples of the activation function include Rectified Linear Activation (ReLu), logistic (e.g., sigmoid), hyperbolic tangent (tanh), and softmax. In certain embodiments, a node may have a capability of remembering previous states (e.g., recurrent nodes). Previous states may be applied to the input and output values using a set of learning parameters.

[0327] A layer is a building block in a deep learning architecture composed of nodes. A layer is a set of nodes that receives data input (e.g., weighted or non-weighted input), transforms it (e.g., by carrying out instructions, e.g., applying a set of functions e.g., linear and / or non-linear functions), and passes transformed values as output (e.g., to the next layer). In certain embodiments, the set of nodes in a particular layer may share the same parametersand instructions without interacting with each other. A machine learning module may be composed of at least one layer (e.g., ordered). Examples of types of layers include convolutional layers (e.g., layers with a kernel, a matrix of parameters that is slid across an input to be multiplied with multiple input values to reduce them to a single output value); fully connected (FC) layers (e.g. all nodes are connected to all outputs of the previous layer); recurrent layers, long / short term memory (LSTM) layers, gated recurrent unit (GRU) layers (e.g., nodes with the various abilities to memorize and apply their previous inputs and / or outputs); batch normalization (BN) layers (e.g., layers that normalize a set of outputs from another layer, allowing for more independent learning of individual layers); activation layer (e.g., layers with nodes that only contain an activation function); (un)pooling layers [e.g., layers that reduce (increase) dimensions of an input by summarizing (splitting) input values in defined patches).

[0328] In certain embodiments, the performance of a machine learning module may be characterized by its ability to produce an output data that reproduces an input data with specific accuracy. To achieve specific accuracy, a training process is performed to find optimal parameters, such as weights, for every node in every layer of the machine learning module. In certain embodiments, the training process of a machine learning module may involve using output data to calculate an objective function (e.g., cost function, loss function, error function) that needs to be optimized (e.g., minimized, maximized). For example, a machine learning objective function may be a combination of a loss function and regularization parameter. The loss function is related to how well the output is able to predict the input. The loss function may take various forms, like mean squared error, mean absolute error, binary cross-entropy, categorical cross-entropy, for example. The regularization term may be needed to prevent overfitting and improve generalization of the training process. Typical regularization techniques include LI Regularization or Lasso Regression, L2 Regularization or Ridge Regression, and Dropout (e.g., dropping layer outputs at random during training process).

[0329] In certain embodiments, objective function optimization of a machine learning module may involve finding at least one (e.g., all) of the present global optima (e.g., as opposed to local optima). A typical algorithm for objective function optimization follows principles of mathematical optimization for a multi-variable function and relies on achieving specific accuracy of the process. Examples of objective function optimization algorithms include gradient descent, nonlinear conjugate gradient, random search, Levenberg-Marquardtalgorithm, limited-memory Broyden-Fietcher-Goldfarb-Shanno algorithm, pattern search, basin hopping method, Krylov method, Adam method, genetic algorithm, particle swarm optimization, surrogate optimization, and simulated annealing.

[0330] In certain embodiments, available input data includes training data and validation data, e.g., where the validation data is separate and non-overlapping with the training data. Training data is used during the training process to optimize a model, whereas validation data is used to check the accuracy of the model while operating on previously unseen data. In certain embodiments, training data is divided into batches (e.g., portions) that is sequentially used (e.g., in random order) as sets of inputs to train a model. In certain embodiments, a model is trained multiple times (e.g., epochs) on the entire set of training data.

[0331] As shown in FIG. 9, an implementation of a network environment 900 for use in providing systems, methods, and architectures as described herein is shown and described. In brief overview, referring now to FIG. 9, a block diagram of an exemplary cloud computing environment 900 is shown and described. The cloud computing environment 900 may include one or more resource providers 902a, 902b, 902c (collectively, 902). Each resource provider 902 may include computing resources. In some implementations, computing resources may include any hardware and / or software used to process data. For example, computing resources may include hardware and / or software capable of executing algorithms, computer programs, and / or computer applications. In some implementations, exemplary computing resources may include application servers and / or databases with storage and retrieval capabilities. Each resource provider 902 may be connected to any other resource provider 902 in the cloud computing environment 900. In some implementations, the resource providers 902 may be connected over a computer network 908. Each resource provider 902 may be connected to one or more computing device 904a, 904b, 904c (collectively, 904), over the computer network 908.

[0332] The cloud computing environment 900 may include a resource manager 906. The resource manager 906 may be connected to the resource providers 902 and the computing devices 904 over the computer network 908. In some implementations, the resource manager 906 may facilitate the provision of computing resources by one or more resource providers 902 to one or more computing devices 904. The resource manager 906 may receive a request for a computing resource from a particular computing device 904. The resource manager 906 may identify one or more resource providers 902 capable of providingthe computing resource requested by the computing device 904. The resource manager 906 may select a resource provider 902 to provide the computing resource. The resource manager 906 may facilitate a connection between the resource provider 902 and a particular computing device 904. In some implementations, the resource manager 906 may establish a connection between a particular resource provider 902 and a particular computing device 904. In some implementations, the resource manager 906 may redirect a particular computing device 904 to a particular resource provider 902 with the requested computing resource.

[0333] FIG. 10 shows an example of a computing device 1000 and a mobile computing device 1050 that can be used to implement the techniques described in this disclosure. The computing device 1000 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The mobile computing device 1050 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smart-phones, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to be limiting.

[0334] The computing device 1000 includes a processor 1002, a memory 1004, a storage device 1006, a high-speed interface 1008 connecting to the memory 1004 and multiple high-speed expansion ports 1010, and a low-speed interface 1012 connecting to a low-speed expansion port 1014 and the storage device 1006. Each of the processor 1002, the memory 1004, the storage device 1006, the high-speed interface 1008, the high-speed expansion ports 1010, and the low-speed interface 1012, are interconnected using various busses, and may be mounted on a common motherboard or in other manners as appropriate. The processor 1002 can process instructions for execution within the computing device 1000, including instructions stored in the memory 1004 or on the storage device 1006 to display graphical information for a GUI on an external input / output device, such as a display 1016 coupled to the high-speed interface 1008. In other implementations, multiple processors and / or multiple buses may be used, as appropriate, along with multiple memories and types of memory. Also, multiple computing devices may be connected, with each device providing portions of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi -processor system). Thus, as the term is used herein, where a plurality of functions are described as being performed by “a processor”, this encompasses embodiments wherein theplurality of functions are performed by any number of processors (one or more) of any number of computing devices (one or more). Furthermore, where a function is described as being performed by “a processor”, this encompasses embodiments wherein the function is performed by any number of processors (one or more) of any number of computing devices (one or more) (e.g., in a distributed computing system).

[0335] The memory 1004 stores information within the computing device 1000. In some implementations, the memory 1004 is a volatile memory unit or units. In some implementations, the memory 1004 is a non-volatile memory unit or units. The memory 1004 may also be another form of computer-readable medium, such as a magnetic or optical disk.

[0336] The storage device 1006 is capable of providing mass storage for the computing device 1000. In some implementations, the storage device 1006 may be or contain a computer-readable medium, such as a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid state memory device, or an array of devices, including devices in a storage area network or other configurations. Instructions can be stored in an information carrier. The instructions, when executed by one or more processing devices (for example, processor 1002), perform one or more methods, such as those described above. The instructions can also be stored by one or more storage devices such as computer- or machine-readable mediums (for example, the memory 1004, the storage device 1006, or memory on the processor 1002).

[0337] The high-speed interface 1008 manages bandwidth-intensive operations for the computing device 1000, while the low-speed interface 1012 manages lower bandwidthintensive operations. Such allocation of functions is an example only. In some implementations, the high-speed interface 1008 is coupled to the memory 1004, the display 1016 (e.g., through a graphics processor or accelerator), and to the high-speed expansion ports 1010, which may accept various expansion cards (not shown). In the implementation, the low-speed interface 1012 is coupled to the storage device 1006 and the low-speed expansion port 1014. The low-speed expansion port 1014, which may include various communication ports (e.g., USB, Bluetooth®, Ethernet, wireless Ethernet) may be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, e.g., through a network adapter.

[0338] The computing device 1000 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a standard server 1020, or multiple times in a group of such servers. In addition, it may be implemented in a personal computer such as a laptop computer 1022. It may also be implemented as part of a rack server system 1024. Alternatively, components from the computing device 1000 may be combined with other components in a mobile device (not shown), such as a mobile computing device 1050. Each of such devices may contain one or more of the computing device 1000 and the mobile computing device 1050, and an entire system may be made up of multiple computing devices communicating with each other.

[0339] The mobile computing device 1050 includes a processor 1052, a memory 1064, an input / output device such as a display 1054, a communication interface 1066, and a transceiver 1068, among other components. The mobile computing device 1050 may also be provided with a storage device, such as a micro-drive or other device, to provide additional storage. Each of the processor 1052, the memory 1064, the display 1054, the communication interface 1066, and the transceiver 1068, are interconnected using various buses, and several of the components may be mounted on a common motherboard or in other manners as appropriate.

[0340] The processor 1052 can execute instructions within the mobile computing device 1050, including instructions stored in the memory 1064. The processor 1052 may be implemented as a chipset of chips that include separate and multiple analog and digital processors. The processor 1052 may provide, for example, for coordination of the other components of the mobile computing device 1050, such as control of user interfaces, applications run by the mobile computing device 1050, and wireless communication by the mobile computing device 1050.

[0341] The processor 1052 may communicate with a user through a control interface 1058 and a display interface 1056 coupled to the display 1054. The display 1054 may be, for example, a TFT (Thin-Film-Transistor Liquid Crystal Display) display or an OLED (Organic Light Emitting Diode) display, or other appropriate display technology. The display interface 1056 may comprise appropriate circuitry for driving the display 1054 to present graphical and other information to a user. The control interface 1058 may receive commands from a user and convert them for submission to the processor 1052. In addition, an external interface 1062 may provide communication with the processor 1052, so as to enable near area communication of the mobile computing device 1050 with other devices. The externalinterface 1062 may provide, for example, for wired communication in some implementations, or for wireless communication in other implementations, and multiple interfaces may also be used.

[0342] The memory 1064 stores information within the mobile computing device 1050. The memory 1064 can be implemented as one or more of a computer-readable medium or media, a volatile memory unit or units, or a non-volatile memory unit or units. An expansion memory 1074 may also be provided and connected to the mobile computing device 1050 through an expansion interface 1072, which may include, for example, a SIMM (Single In Line Memory Module) card interface. The expansion memory 1074 may provide extra storage space for the mobile computing device 1050, or may also store applications or other information for the mobile computing device 1050. Specifically, the expansion memory 1074 may include instructions to carry out or supplement the processes described above, and may include secure information also. Thus, for example, the expansion memory 1074 may be provide as a security module for the mobile computing device 1050, and may be programmed with instructions that permit secure use of the mobile computing device 1050. In addition, secure applications may be provided via the SIMM cards, along with additional information, such as placing identifying information on the SIMM card in a non-hackable manner.

[0343] The memory may include, for example, flash memory and / or NVRAM memory (non-volatile random access memory), as discussed below. In some implementations, instructions are stored in an information carrier. The instructions, when executed by one or more processing devices (for example, processor 1052), perform one or more methods, such as those described above. The instructions can also be stored by one or more storage devices, such as one or more computer- or machine-readable mediums (for example, the memory 1064, the expansion memory 1074, or memory on the processor 1052). In some implementations, the instructions can be received in a propagated signal, for example, over the transceiver 1068 or the external interface 1062.

[0344] The mobile computing device 1050 may communicate wirelessly through the communication interface 1066, which may include digital signal processing circuitry where necessary. The communication interface 1066 may provide for communications under various modes or protocols, such as GSM voice calls (Global System for Mobile communications), SMS (Short Message Service), EMS (Enhanced Messaging Service), or MMS messaging (Multimedia Messaging Service), CDMA (code division multiple access),TDMA (time division multiple access), PDC (Personal Digital Cellular), WCDMA (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Packet Radio Service), among others. Such communication may occur, for example, through the transceiver 1068 using a radio-frequency. In addition, short-range communication may occur, such as using a Bluetooth®, Wi-Fi™, or other such transceiver (not shown). In addition, a GPS (Global Positioning System) receiver module 1070 may provide additional navigation- and location-related wireless data to the mobile computing device 1050, which may be used as appropriate by applications running on the mobile computing device 1050.

[0345] The mobile computing device 1050 may also communicate audibly using an audio codec 1060, which may receive spoken information from a user and convert it to usable digital information. The audio codec 1060 may likewise generate audible sound for a user, such as through a speaker, e.g., in a handset of the mobile computing device 1050. Such sound may include sound from voice telephone calls, may include recorded sound (e.g., voice messages, music fdes, etc.) and may also include sound generated by applications operating on the mobile computing device 1050.

[0346] The mobile computing device 1050 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a cellular telephone 1080. It may also be implemented as part of a smart-phone 1082, personal digital assistant, or other similar mobile device.

[0347] Various implementations of the systems and techniques described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0348] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms machine-readable medium and computer-readable medium refer to any computer program product, apparatus and / or device(e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine- readable medium that receives machine instructions as a machine-readable signal. The term machine-readable signal refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0349] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0350] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0351] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0352] In some implementations, various modules described herein can be separated, combined or incorporated into single or combined modules. Modules depicted in the figures are not intended to limit the systems described herein to the software architectures shown therein.

[0353] Elements of different implementations described herein may be combined to form other implementations not specifically set forth above. Elements may be left out of theprocesses, computer programs, databases, etc. described herein without adversely affecting their operation. In addition, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. Various separate elements may be combined into one or more individual elements to perform the functions described herein. Throughout the description, where apparatus and systems are described as having, including, or comprising specific components, or where processes and methods are described as having, including, or comprising specific steps, it is contemplated that, additionally, there are apparatus, and systems of the present invention that consist essentially of, or consist of, the recited components, and that there are processes and methods according to the present invention that consist essentially of, or consist of, the recited processing steps.

[0354] It should be understood that the order of steps or order for performing certain action is immaterial so long as the invention remains operable. Moreover, two or more steps or actions may be conducted simultaneously.

[0355] While the invention has been particularly shown and described with reference to specific preferred embodiments, it should be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the spirit and scope of the invention as defined by the appended claims.E. ExamplesE.i. Example 1: Immune Response Performance Benchmarks and Training Datasets

[0356] This example describes performance benchmarks and training datasets that can be used to evaluate performance and / or train machine learning models for immune response prediction in accordance with various embodiments described herein.(7) Task-Specific Pipeline Model

[0357] In certain cases, distinct, task-specific, machine learning models can be used to model particular steps in biological pathways that work together to form an overall immune response. For example, as shown in FIG. 11, a pipeline of three distinct taskspecific machine learning models can be used to generate immune response predictions. In the pipeline shown in FIG. 11, three different task-specific models - a binding predictor 1120, a cleavage predictor 1130, and presentation predictor 1110 are combined in acomposite (or stacked) modelling pipeline where the outputs of task-specific binding and cleavage prediction models are used as inputs to a presentation model. The task-specific binding 1120 and cleavage 1130 models are transformer-based while a logistic regression is used to integrate binding and cleavage model predictions and antigen processing features for surface presentation prediction 1110. In the figure, predicted values include a binding score 1102 (output by binding predictor model 1120), a cleavage score 1103 (output by cleavage predictor model 1130), and a presentation score 1101 that is output by presentation predictor model 1110. Antigen processing features used as input include representations of an allele 1104, a peptide 1106, and flanks 1106. The features gene bias 1108 and expression 1109 are used for HLA-I presentation and expression and shadow used for HLA-II presentation. An allele weighting 1107 feature is not used during training but is included when performing deconvolution.

[0358] Multi-Allelic Modelling. The task-specific pipeline shown in FIG. 11 used two approaches for multi -allelic modeling, referred to as (i) deconvolution and (ii) Multiple Instance Learning (MIL). The deconvolution approach uses a model, after training, to independently score each allele in a multi -allelic example (e.g., comprising two or more alleles) and returns the allele with the highest score among these scores. MIL generates scores or embeddings for each allele independently and in parallel, before aggregating them using a permutation invariant operation (typically an element-wise max) and uses this aggregate to make a final prediction Accordingly, on the one hand, the deconvolution approach, accordingly, introduces post-processing steps to combine multiple predictions from a model into a multi-allelic prediction, but does not alter the model itself or manner in which it is trained. On the other hand, the multi-allelic scoring and aggregation in MIL are incorporated during the training procedure (as well as at inference), such that in the MIL approach the machine learning model effectively learns to assign scores that will be aggregated across alleles.

[0359] When performing deconvolution, an allele weighting 1107 feature is used to weight individual allele binding scores. In particular the approach proceeds by multiplying per-allele binding scores with pre-computed weights, re-scoring the alleles in a multi-allelic sample and affecting the MHC restriction.

[0360] Binding Affinity Modelling. A task specific binding affinity model was also used. The binding affinity model was trained to model quantitative binding affinity data provided by NetMHCpan-4.1 (https: / / services.healthtech.dtu.dk / services / NetMHCIIpan-4.1 / ).Continuous-valued binding affinity values were modelled by adding a regression output head to the transformer-based peptide-MHC binding predictor 1120.(2) NetMHCpan-4.1 Framework

[0361] NetMHCpan-4.1 (B. Reynisson et al., 2020) is a tool used for predicting eluted ligand likelihood and binding affinity between peptides and MHC-I alleles. This model is used in immunology as a predictor of peptide-MHC binding and aids in understanding how the immune system recognizes and responds to pathogens and cancer cells.

[0362] NetMHCpan-4.1 Architecture and NNAlign MA Framework.NetMHCpan-4.1 represents MHC-I molecules using “pseudo-sequences” comprising the residues typically in contact with binding peptides. The amino acid sequence of training examples are encoded to a 20-dimensional vector corresponding to BLOSUM matrix scores. Peptides longer than 9 amino acids are reduced to a binding core of 9 amino acids by applying consecutive amino acid deletions. These include both deletions at the end terminals and consecutive deletions within the peptide. Additional features are included in training such as length of the deletion or insertion, length of the peptide flanking regions, and the length of the peptide (M. Nielsen et al., 2016). The peptide and allele pseudo-sequence are BLOSUM encoded and then flattened. This flattened representation is then concatenated with the additional features. The concatenated representation is then passed through dense layers. Finally, output neurons for binding affinity (BA) and eluted ligand (EL) respectively are used for prediction.

[0363] The NNAlign framework is a single-allelic framework permitting the integration of mixed data types (EL and BA) in the model training, which allows information to be leveraged across the different data types, resulting in a boosted predictive power (M. Nielsen et al., 2016). The NNAlign machine learning framework was updated to NNAlign_MA (B. Alvarez et al., 2019) to allow for effective handling of multi-allelic (MA) data. This was achieved by iteratively annotating the best single-allele to the MA data during the model training, effectively deconvoluting the MA binding motifs. An ensemble of neural networks described above are trained using the NNAlign_MA framework and the average prediction of the ensemble of models is used as the final output.(3) Tasks

[0364] Various tasks along HLA-I and HLA-II processing pathways were modeled using certain embodiments of immune response prediction systems and methods of the present disclosure, as described in Examples 2-4 below. Datasets for testing and training were obtained from in-house mass spectrometry data as well as public datasets released with the NetMHCpan-4.1 publication (B. Reynisson et al., 2020). FIGs. 12A and 12B illustrate a high-level categorization of the datasets in these collections.

[0365] Steps along the HLA-I and HLA-II processing pathways (e.g., antigen presentation) that were designated as tasks and modeled in the following examples include proteasomal cleavage, peptide-MHC binding, and peptide-MHC surface presentation.

[0366] Proteasomal Cleavage. The HLA-I immunopeptidome originates from a sequence of complex biochemical processes, beginning with proteasomal cleavage. Proteasomal cleavage is an initial step in antigen presentation, whereby proteasomes selectively degrade intracellular proteins, producing peptides typically 8 to 12 amino acids in size. Accordingly, proteasomal cleavage can be viewed as a first stage in modeling HLA-I antigen presentation and immune recognition. As described in further detail herein, peptides that are predicted to be generated via proteasomal cleavage eventually undergo testing for binding with HLA-I alleles expressed by a cell. The HLA-II processing pathway also includes a protein cleavage / processing step that also produces peptides, albeit typically of a slightly larger size, e.g., 10-30 residues, than those produced via the HLA-I pathway.Peptides produced in the HLA-II pathway, however, originate from cleavage of extracellular proteins that taken up by antigen presenting cells via a different set of mechanisms than the HLA-I cleavage step.

[0367] Peptide-MHC Binding. Peptides produced from cleavage steps in HLA-I and HLA-II pathways bind with MHC molecules to be transported and presented at cell surfaces. Accordingly, binding between a peptide and an MHC molecule can be designated as another step - a task to be modeled - for example, before surface presentation and T-cell recognition, in predicting antigen presentation. Data for peptide-MHC binding prediction typically originates from either MHC binding assays or MHC ligand elution data sequenced by mass-spectrometry. Binding assays yield affinity measurements - binding affinity data (BA) - which can be used for qualitative and quantitative modelling of peptide-MHC binding. Eluted ligands are modelled qualitatively as “hits”, corresponding to those ligands identifiedas binding to a particular tested MHC molecule. To train and test modeling approaches, artificial decoys, corresponding to peptides that presumably do not bind the tested MHC molecule are, accordingly, generated, in order to model peptide-MHC binding as a binary classification problem. Depending on the cell line used in ligand elution experiments, data may include either single-allelic samples, having a known MHC restriction since one MHC allele binds the peptide, or multi-allelic samples, for which MHC restriction is unknown since the cell line expresses multiple MHC alleles, any of which may bind to the peptide. For each of these (above-described) data sources, peptide-MHC binding is the task of predicting binding between a peptide and one or more MHC molecules. Both HLA-I and HLA-II peptide-MHC binding tasks were modeled and evaluated in Example 2.

[0368] Peptide-MHC Surface Presentation. Following successful proteasomal cleavage (e.g., in the case of the HLA-I antigen presentation pathway) of a peptide and subsequent binding to an MHC molecule, stable peptide-MHC complexes are transported to cell surfaces. These peptide-MHC complexes act as beacons, presenting a snapshot of the cellular proteome to surveillance T-cells. A number of factors impact likelihood of recognition by T-cells, such as abundance of the presented peptide-MHC. Accurate prediction of surface presentation taking into account peptide abundance, accordingly, facilitates antigen presentation modeling and evaluating potential of candidate antigens or portions thereof for T-cell recognition and providing immune response. Data for surface presentation prediction typically originates from MS-identified eluted ligands, the same as used for peptide-MHC binding. Since the hit data (i.e., positively presented peptides) are the same as for peptide-MHC binding, these tasks are distinguished by the manner in which a background distribution of decoy peptides is generated. Surface presentation tasks modeled herein also includes additional features, such as abundances of peptide source genes.(4) In-House Datasets

[0369] An in-house collection of MHC immunopeptidomics data identified by MSbased sequencing of eluted ligand assays was used to train and evaluate models for certain experiments described herein. This in-house MHC dataset included data generated from experiments using single-allelic and multi-allelic cell lines, ligands binding MHC -I and MHC -II molecules, and data collected from public and internal sources. From this database, three task datasets are formed for proteasomal cleavage prediction, peptide-MHC bindingprediction and peptide-MHC surface presentation prediction, differentiated in the manner in which negative decoys are generate. These datasets are all derived from the same immunopeptidome. In other words, the peptides which are cleaved, bind MHC molecules or present on the cell surface, are all taken from the same set of peptides observed in MHC-I or MHC -II presentation experiments.

[0370] While these tasks share the same process for collecting “hit” peptides - those which successfully cleave, bind or present - they differ in their decoy generation processes. The tasks also differ in their features. Cleavage prediction does not consider the MHC molecule but inputs the peptide flanking residues, for example.

[0371] Data Partitioning. A human proteome-level partitioning scheme was used to assign each catalogued human protein to either a train, tune, or test partition. When assigning peptides to a partition, an exact match search was performed over each partitioned human protein. Once a set of protein matches are found for a given peptide, a priority matching system is used to assign the peptides to one of the partitions.

[0372] A sub-partitioning scheme was used to further subdivide the train and tune partitions into five sub-partitions. In this manner, the full proteome was divided in the form {trainl, train5, tunel, tune5, test}. Subdividing the train and tune partitions was accomplished by assigning protein sub-clusters to a sub-partition by attempting balance values of a particular (e.g., pre-defined) quantity across the sub-partitions with the aim of creating an even distribution of peptides across the sub-partitions. In the examples described herein, in particular, the subdivision approach aimed to balance the number of protein transcripts that were assigned to each sub-partition. Other approaches, such as, for example, balancing a sum of representative protein sequence lengths from sub-clusters assigned to each sub-partition, may be used.

[0373] Peptide-MHC Binding Prediction. Datasets for training and evaluating peptide-MHC binding prediction models were constructed using peptides identified mono- allelic cell line experiments as positive examples and introducing decoys to produce hit: decoy ratios of 1 : 500 and 1: 19, respectively, for MHC-I and MHC-II peptide binding. The peptide- MHC binding prediction task involved training, and evaluating performance of, a machine learning model to generate, as output, a prediction of whether a particular candidate peptide was a hit or decoy, based on representations of the candidate peptide and a particular MHC allele or set of MHC alleles, received as input. For MHC-II binding predictions, two allelenames (such as HLA-DRA1 *01:01 and HLA-DRB 1*01:01) were used to represent the alpha and beta chains of an MHC-II molecule.

[0374] Surface Presentation Prediction. Surface presentation task datasets were constructed with hits and decoys at ratios of 1:5000 for MHC-I presentation and 1:500 for MHC-II presentation. Similar to peptide-MHC binding task, inputs for surface presentation tasks included representations of both a candidate peptide and a set of MHC alleles. The presentation task inputs also included additional inputs representing amino acid sequences from the source protein in the local vicinity to the left and right of the candidate peptide. These adjacent amino acids are referred to as the left and right flanking residues (or flanks). A numerical field, referred to as “expression,” was also used as input to represent the degree to which the candidate peptide’s source protein is expressed in the particular cell type used in the experiment. Where a candidate peptide mapped to multiple source proteins the maximum expression level was used.

[0375] Proteasomal Cleavage Prediction. In cleavage prediction tasks, models received representations of a candidate peptide and its flanking residues as input and generated, as output, a binary label indicating whether the candidate peptide is cleaved or not. The same mass spectrometry data used in the binding and presentation tasks was used to construct datasets used for proteasomal cleavage prediction. To control for the binding and presentation signals in the cleavage task, decoys were constructed to be similar to known presenting peptides.(5) Public, NetMHCpan, Datasets

[0376] Public datasets were also used to train and evaluate models. In particular, training and evaluation datasets were released with NetMHCpan-4.1 and NetMHCIIpan-4.0 (B. Reynisson et al., 2020). These datasets were created by querying the Internet Epitope Database (IEDB) and include binding affinity (BA), eluted ligand (EL) datasets with HLA, non-HLA, single-allelic (SA) and multi-allelic (MA) data. These datasets were each partitioned into train, tune, and test splits [referred to as “Fold 0 Train”, “Fold 0 Tune”, and “Fold 0 Test”, respectively, where the prefix “Fold 0” reflects only a single, as opposed to multiple different (e.g., as in w-fold cross-validation approaches), partition]. As described in further detail herein, these datasets were used to train a multi-task model in accordance with various embodiments described herein that was implemented using T5 framework in order tocompare example implementations of immune response prediction technologies described herein with strong baselines such as NetMHCpan.NetMHCpan-4.1 EL Data

[0377] Hit:Decoy Ratios. The hitdecoy ratios for the fold 0 training sets of the NetMHCpan data are shown in Table 2, below.Table 2: The hitdecoy ratios of the NetMHCpan-4. 1 (MHC-I) and NetMHCIIpan-4.0 (MHC-II) Fold 0 Train datasets.

[0378] Allele Frequencies. Since metrics are often reported “per-allele”, it is important to take into account the distribution of allele frequencies in the evaluation sets. Allele frequency distributions for NetMHCpan-4.1 SA Fold 0 Tune and Test sets (the test set is referred to as the “MS Ligands Test set”) are shown in FIGs. 13A and 13B, respectively.

[0379] As shown in FIGs. 13A and 13B, the MS Ligands Test set has a different distribution of allele frequencies than the NetMHCpan-4.1 SA Fold 0 Tune dataset that was used for selecting an EL model with optimal hyperparameters. In particular, the allele frequency distribution of the NetMHCpan-4. 1 SA Fold 0 Tune set is heavily skewed towards low numbers of ligands per allele. As a result, a small set of examples may have a large effect on aggregated per-allele metrics (such as mean-per-allele), making the aggregates noisy metrics. Accordingly, when using the NetMHCpan SA Fold 0 Tune set for model selection, global rather than per-allele metrics were used.

[0380] In the NetMHCpan SA Fold 0 Tune set, there are 12 MHC alleles with no hits and one MHC allele with no decoys. Metrics, such as Top-K and Average Precision, are undefined where there are no hits. For purposes of reporting in examples described herein, these undefined values were replaced with zeros. Alleles with no decoys always have perfect scores of 1 for both Top-K and Average Precision.

[0381] EL MA and SA Data. The NetMHCpan-4. 1 EL datasets include both MA and SA examples. FIG. 14 shows the number of SA and MA examples in the fold 0 train splits of the NetMHCpan-4.1 (MHC-I) and NetMHCIIpan-4.0 (MHC-II) EL datasets. As shown in FIG. 5, MA examples constitute a large fraction of the total available data.NetMHCpan-4.1 Binding Affinity Data

[0382] The NetMHCpan-4. 1 binding affinity (BA) dataset was obtained from the Immune Epitope Database (IEDB), which contains experimental data on peptide-MHC binding affinities. This dataset comprises of 186,684 peptide-MHC binding affinity measurements covering 172 MHC molecules from human, mouse, primates, cattle, and swine. Nielsen et al., 2016, introduced 25 random natural peptides for each of the lengths 8, 9, 10, and 11 as artificial negatives for each allele, to ensure a sufficiently diverse set of negative examples. These random sequences were only used for training and were excluded from all evaluations. The data were split into five partitions for cross-validation to ensure that no identical 8-mer segment was shared between partitions. The distribution of the binding affinity train and tune splits are shown below in FIGs. 15-16.E. ii. Example 2: Example Implementation of a Multi-Task LEM for Immune Response Prediction

[0383] This example describes an example implementation of a LLM in accordance with various embodiments of immune response methods and systems described herein. In particular, this example describes methods used to tailor a text-to-text LLM based on the model described in (Raffel et al., 2020), referred to as “T5,” for end-to-end learning of various immune response tasks, namely, proteasomal cleavage, peptide-MHC binding, and surface presentation, as described in Example 1. Among other things, this example includes descriptions of an overall modelling approach based on the T5 model, example approaches for representing biological features as text inputs, particular model settings that were used in experiments described in Examples 3-5 below, metrics for evaluating model performance and implementation approaches, including model development and model selection.

[0384] FIG. 17 is a schematic illustration of various immune response prediction tasks, such as tasks corresponding to particular steps in antigen presentation biologicalpathways as well as overall epitope identification and eluted ligand (EL) prediction tasks, in a “text-to-text” format that is compatible with the T5 framework and can, accordingly, be used for multi-task learning with an LLM. As shown in FIG. 17, tasks, including binding affinity, binding classification, proteasomal cleavage, surface presentation and direct ligand (or epitope) identification, are modelled as pairs of textual inputs and targets. Binding affinity and binding classification tasks are relevant for different types of data sources, and produce different outputs, with the binding affinity task involving predicting a quantitative binding affinity value as relevant for binding affinity data and the binding classification task referring to the binary classification task that is relevant for eluted ligand data. Although outputs are text targets, model predictions can still be used for task evaluation where results in datasets take the form of continuous numerical values, for example, by converting regression predictions from text to continuous values as described in further detail herein. For classification tasks, likelihood predictions are estimated using the language model likelihood of the label sequence. Natural language generations tasks like direct ligand identification are naturally suited for T5 and can be directly evaluated.( / ) Input Representations

[0385] Input features to T5 models are represented as text. For the immune response prediction tasks modeled using T5 in this Example and in Examples 3-5, peptides, MHC alleles, and various continuous features were encoded as textual features.

[0386] Peptides can be represented in text form via their amino acid sequence. Following the approach of NetMHCpan-4.1, this example represented peptides as sequences of twenty-one (21) possible letters, with 20 letters representing the twenty standard amino acid types and an “X” for unknown amino acids. All characters not in the set of 20 standard amino acids were replaced with an “X” character, including lowercase letters used in IEDB to denote amino acids in their D- configuration.

[0387] MHC alleles can be represented by their names, pseudo-sequences, or both. FIG. 18 shows examples whereby a particular HLA allele is represented using an alphanumeric textual string label for the particular allele, an amino acid pseudo sequence, and both used in combination, with the label and pseudo sequence concatenated. In some experiments, a particular allele representation (e.g., textual string label, pseudo sequence, concatenation of both) for each training example. Where training examples included data formultiple alleles, a list of alleles in a selected representation, as shown in FIG. 19, was used. During training, orders of alleles were shuffled to force the model to learn permutation invariance.

[0388] Continuous values can either be represented by their decimal strings or by grouping them into bins and representing each value using a single token denoting its bin, as illustrated in FIG. 20.(2) Model Sizes and Hyperparameters

[0389] Machine learning models, such as LLMs, typically include a number of hyperparameters that can be varied, such sizes of embeddings, hidden multi-layer perceptron’s (MLPs), and numbers of layers and heads. Several different model implementations, with variations in hyperparameters, were used in examples described herein, as listed in Table 3. All models used Gated Ge LU (Shazeer, 2020) as intermediate activations in MLPs. Model names reflect the number of trainable parameters in the resulting model.Table 3: Hyper-parameters of different model sizes used in experiments described in Examples 1-5.Name Legacy „ , . .. MLP,TEmbeddingT TT,TTcName „. Hidden Layer Heads Head SizeSizec•SizeT5-590k Tiny 128 512 2 2 64T5-660k Custom 128 128 4 4 48T5-2M Deep12s512 6432NarrowT5-4M Mini 256 1024 4 4 64T5-25M Small 512 2048 6 8 64T5-38M Small-Tall 512 2048 9 8 64(3) Metrics

[0390] Several metrics were used herein to evaluate two types of models. A first type of model developed and evaluated herein is an epitope ranking model, which takes a single epitope as input and outputs a binding score. These epitope ranking models were evaluated using the Top-K and average precision metrics described below. A second type of model that was developed and evaluated was a direct ligand identification (e.g., location or generation)model, which takes a full protein sequence as input, and outputs locations (e.g., via a binary mask) or sequences of high-binding -likelihood epitopes within the input protein sequence. These direct ligand identification models were evaluated using precision, recall, accuracy and Fl scores, described further in the following.

[0391] Top-K. A “Top-K” metric was defined as the fraction of true positives in the top k peptides as ranked by the model, where k is the number of known positive binders in the evaluation dataset. Top-K is, accordingly, equivalent to both precision and recall among these top- / : samples. This metric may also be referred to as Positive Predicted Value (PPV). Top-K can also be interpreted as the point on a precision-recall curve where precision is equal to recall. This metric was used for scoring epitope-ranking models.

[0392] Average Precision. Given a binding score output by a model for a particular epitope, to classify the particular epitope as either positive or negative (in a binary classification task, such as eluted ligand prediction, cleavage, binding, or surface presentation), a score threshold may be used, for example such that epitopes having a score above a particular score threshold value are classified as positive (for, e.g., cleavage, MHC binding, and / or surface presentation), while epitopes having scores below the score threshold are classified as negative. Modifying the score threshold allows for a trade-off between precision and recall to be adjusted. When calculating the Top-K metric a particular score threshold is selected such that precision and recall are equal on the benchmark dataset. Since the threshold is chosen using knowledge of the ground-truth labels, Top-K may be less desirable as an indicator of a model’s performance in a practical setting, where such labels are unavailable. Average Precision (AP) addresses this issue since it is does not choose a classification threshold. Instead, AP is computed as an average of the precision scores for every possible recall value (and hence every possible threshold). Formally, AP is defined as the area under the precision-recall curve obtained by varying the threshold. See, e.g., (Zhu, 2004).

[0393] Frank score. (Jurtz et al., 2017) introduce a metric called Frank score, which measures where a model ranks a known binding peptide amongst a set of decoy peptides generated from the source protein of the binder. Formally, Frank score is defined as the fraction of the decoys which score higher than the binding peptide, (see Dalia-Torre, 2023) where the decoys are generated by applying sliding windows of lengths 8 to 14 across the source protein (excluding the binding peptide itself). A ‘perfect’ Frank score of zero indicates that a model ranks the binder above all of the decoys. Given a benchmark dataset ofbinding peptides and their associated source proteins, a Frank score can be calculated for each binder, and mean and median scores, as well as the number of perfect scores, reported as overall measures of a model’s performance on the benchmark dataset.

[0394] As described herein, the task of direct ligand prediction involves generating one or more (e.g., a set of) ligands predicted to be likely resultant peptides to be presented for T-cell recognition, based on an antigen (e.g., protein) sequence and one or more MHC alleles received as input. Accordingly, a direct ligand prediction model’s predictive performance can be evaluated by classifying each predicted ligand into one of three categories:1. True Positive (TP): A predicted ligand that matches a target ligand.2. False Positive (FP): A predicted ligand that does not match any target ligand.3. False Negative (FN): A target ligand that is not predicted by the model.

[0395] Predicted ligands were compared with target ligands using a Windowed Exact Match (WEM) function. Unlike directly comparing two ligands for equality, WEM checks if a target ligand falls within a window of ±5 amino acids around a predicted ligand in the antigen sequence:Sequence: MYYKFSGFTQKLAGAWASEAYSPQGLKPVVS DSTVYDTarget: FTQKLAGAWPrediction 1: FSGFTQKLAGAW, WEM = 1.0 (YKFSGFTQKLAGAWASEAY 9 FSGFTQKLAGAWPrediction 2: EAYSPQGLKPV , WEM = 0.0 (YKFSGFTQKLAGAWASEAY 2 EAYSPQGLKPV)

[0396] After classifying each ligand, total numbers of TPs, FPs and FNs can be determined for each pair of allele and antigen sequences in a benchmark dataset and used to compute four classification metrics - precision, recall, accuracy, and Fl -score, defined the equations below. These metrics provide a comprehensive assessment of the model’s ability to accurately predict ligands that will be presented for T-cell recognition.Precision = Z(TP) / (Z(TP) + Z(FP))Recall = Z(TP) / (Z(TP) + Z(FN))Accuracy = Z(TP) / (Z(TP) + Z(FN) + Z(FP))Fl -score = 2 * (Precision x Recall) / (Precision + Recall)(4) Implementation Approaches

[0397] Example T5 LLMs for end-to-end immune response predictions as described in Examples 1-5 were implemented using a JAX-based framework, while Flaxformer was used for the transformer model implementation.

[0398] Several models were trained using various hyper parameters and input representations, as described herein. While training the models, checkpoints of model parameters were saved every fixed number of batches (typically 4096). Each time a model checkpoint was saved, the model was evaluated on the tune set of each task in the training mixture. When selecting a model within a given experiment, a single checkpoint per model was first chosen. Thereafter, the best checkpoints between different models were compared. Unless otherwise stated, the Global Top-K metric was used to select models.E. Hi. Example 3: LLM-Based Antigen Presentation Modeling Results for In-House Datasets

[0399] This Example describes experiments and results demonstrating use of a T5 implementation in accordance with various embodiments described herein to model in-house antigen presentation datasets. Results presented in this example demonstrate an end-to-end antigen presentation predictor using a single transformer-based neural network. Various experiments described herein test the effectiveness of a certain hyper-parameters or modelling decisions. Where relevant, parallel results for both MHC-I and MHC-II tasks are provided.

[0400] The single, T5, end-to-end model aimed to match performance of, and was compared with, the task-specific pipeline model approach described in Example 1, which utilizes binding and cleavage neural networks as input to a logistic regression presentation model. To achieve this the immune response prediction model was trained in a multi-task fashion, combining the proteasomal cleavage, peptide-MHC binding, and presentation tasks. Results for experiments on tune datasets were used to inform various modelling decisions made while creating an initial End-To-End (E2E) presentation predictor using binding and presentation single-allelic (SA) data. Additional results show how a T5 model implementedin accordance with various embodiments of the present disclosure can be used to model the cleavage task in both the single-task and multi-task settings. A final, single 38 million parameter T5 model was trained on binding and presentation multi-allelic (MA) data.(7) Modelling Single-Allelic Antigen Presentation

[0401] Results of single-allelic (SA) antigen presentation tasks presented herein show, among other things, impact of various allele input representations and approaches for tokenizing continuous values for the presentation task.

[0402] The inputs for modelling the presentation task were an allele, a peptide, and a set of continuous values, such as gene expression. Flanking residues were provided as inputs for the cleavage task, but not used in the presentation task.

[0403] Multi-task models for MHC-I binding and presentation tasks were evaluated on binding and presentation tune data. Details of the model and training parameters are shown in Table 4, below. FIGs. 21 and 22 show results for Top-K metrics calculated for single-task models trained on binding and presentation data only, along with a multi-task model trained on both binding and presentation data. FIG. 21 provides results for the MHC-I binding tune dataset, evaluating model performance on a binding task. FIG. 21 shows performance for the multi-task model trained with binding and presentation data to be similar to that of the single-task binding model on the binding task. FIG. 21 also shows that the model trained on only presentation data performs poorly on the binding task, however it is somewhat able to classify binders without presentation features. FIG. 22 provides results for the presentation task and shows the multi-task model to outperform the single-task presentation model. As shown in the figure, performance of the task-specific binding model on the presentation task is limited. Accordingly, the results shown in FIGs. 21 and 22 indicate that co-training binding and presentation in a multi-task fashion provides performance benefits, particularly in the context of the presentation task.Table 4: Details of model and training parameters for binding and presentation tasks.Model Details Training DetailsEmbed MLP Heads Layers Param Batch Dropout LR TrainDim Dim Steps

[0404] The influence of different mixture ratios of binding and presentation examples used during training was also evaluated. Parameters of the model used to evaluate impact of mixture ratios are shown in Table 5. FIG. 23 A shows that the binding task performance is barely affected by the mixture ratio. FIG. 23B shows that the presentation task performs best when using either a 4: 1 or 2: 1 binding to presentation ratio. Up-sampling the presentation task to an 8: 1 ratio appears to reduce presentation performance. Without wishing to be bound to any particular theory, this may be due to, for example, the presentation task being harder, and the binding task serving as an easier way to learn allele and peptide representations. Based on these results, 4: 1 was selected as the baseline mixture ratio between binding and presentation for subsequent experiments.Table 5: Details of the model and training parameters.Model Details Training DetailsEmbed MLP Heads Layers Params Batch Dropout LR TrainDim Dim Steps256 1024 4 4 4.4M 16K 0.1 le’4131K

[0405] Input representations for MHC-I and MHC-II alleles can be represented in a variety of fashions. Two approaches were evaluated in this example, namely an allele names and a pseudosequences approach. The allele names approach involves tokenizing the allele name text directly, to create a label that identifies a particular allele type, for example, [‘HLA-’, ‘A’, ‘01’, ‘:02’] or [‘HLA-’, DRA’, ‘ 1’, ‘02’, ‘:03’]. Additionally or alternatively, a pseudo-sequence approach can be used, whereby, rather than include an entire amino acid sequence of an MHC protein, a discontinuous subsequence is used, for example, including only those amino acids that are in contact with bound peptides and / or only those amino acids that are polymorphic between various MHC alleles. A variety of approaches for representing pseudosequences are possible.

[0406] Turning to FIGs. 24A and 24B, MHC-I binding and presentation prediction models using an allele names representation and various pseudosequence representation approaches were evaluated and compared. Four pseudosequence representation approaches were compared: NetMHCPan, MHCFlurry (O’Donnel, 2020), BigMHC, and DePTH. NetMHCPan selects 34 residues exclusively from the binding groove of the MHC protein. MHCFlurry pseudo-sequences are similar to NetMHCPan with 2 additional amino acids persequence. The DePTH method involves selecting residues which may interact with either an antigen or a T-Cell Receptor. BigMHC makes use of a learned approach whereby the attention scores of a model trained on full MHC protein sequences are used to infer the relative importance of different residues. The 30 residues selected are from both inside and outside of the binding groove. This approach differs in its ability to select different residues for different alleles.

[0407] The results in FIG. 25 A show that for the MHC -I binding task the Allele Names model performs best. For the presentation task, the results in FIG. 25B show that Allele Names again performs best, however the performance gap is narrower. The details of the models used are shown in Tables 6-7. While MHCFlurry underperforms on the binding task, on the presentation task it performs slightly better than other pseudo-sequence representation approaches.Table 6: Details for the model and training parameters used for MHC-I models exploring different allele representations and pseudo-sequences.Model Details Training DetailsEmbed MLP Heads Layers Params Task Batch Dropout LR TrainDim Dim Mixture Steps128 512 4 6 1.7M 4: 1 32K 0.1 le’4131KB+PTable 7: Details for the model and training parameters used.Model Details Training DetailsEmbed MLP Heads Layers Params Task Batch Dropout LR TrainDim Dim Mixture Steps256 1024 4 4 4.4M Single 32K 0.1 lc’413 IK task binding

[0408] MHC-I Unseen Allele Evaluation. Allele names and pseudo-sequence models were also compared in the unseen allele setting. This involves taking models already trained on test data and evaluating them in a new context. Unseen alleles were not in the models’ training data, and thus serve to evaluate how well a model generalizes. To this end,a dataset with data that originates from MHCFovea was used. The dataset has had all alleles found in the training data fdtered out.

[0409] The results in FIGs. 25A-B show that a pseudo-sequence model outperforms a model that uses allele names to identify MHC alleles, on unseen alleles. Parameters of these models are shown in Table 8. The models are tested on public data test sets collected from both MHCFovea and NetMHCPan4. 1 and have had all rows removed which contain alleles present in the training data. This holds true for both the MHCFovea and NetMHCPan4.1 unseen allele evaluation sets. Whether substituting unseen alleles with the most similar seen allele would improve allele name performance was also evaluated. This substitution was found to recover much of the lost performance when compared to pseudo-sequences. However, even with homologous allele substitution, the pseudo-sequence model still outperformed the allele names model. Accordingly, these results indicate that pseudosequence models generalize better to unseen alleles and likely would offer improved performance, e.g., in a clinical setting.Table 8: Details of the model and training parameters used for MHC -I models comparing allele names to the NetMHCIIPan4.0 pseudo-sequence representation when evaluated on unseen alleles using public data.Model Detail Training DetailsEmbed MLP Heads Layers Params Task Batch Dropout LR TrainDim Dim Mixture Steps128 512 4 6 1.7M 4: 1 B+P 16K 0.2 le’3131K

[0410] MHC-II alleles. Allele representations were also evaluated for MHC-II data. Unlike MHC-I, FIG.26A shows that the highest binding performance resulted from the NetMHCPan pseudo-sequences. Presentation tune results, shown in FIG. 26B, are less conclusive. However, since primarily a global Top-K was used as a model selection metric during model building, pseudo-sequence was selected as the better performing MHC-II allele representation.

[0411] Continuous variables. Various approaches for tokenizing continuous values for use in text-to-text models were also evaluated using the MHC-I and MHC-II presentation tune datasets. First, different strategies were explored in the MHC-II context since it wasfaster to train and evaluate. Then, the selected strategies were tuned for MHC-I and MHC-II separately.

[0412] Results for tuning various quantile bin parameters for an MHC-I model with parameters as in Table 9 are shown in FIGs. 27A and 27B. FIG. 27A shows similar binding results for different binning strategies. Results shown in FIG. 27B indicate that using 50 bins produces the best performance, however since global Top-K was used to perform model selection, 20 bins were selected. FIG. 27B also shows that shared bin tokens across fields perform better than using unique bin tokens per field. A label “unique” corresponds to a test of using distinct tokens per field for bins as compared to sharing bin tokens across fields (which is a default).Table 9: Details of the model and training parameters used for MHC-I models.Model Details Training DetailsEmbed MLP Heads Layers Params Task Batch Dropout LR TrainDim Dim Mixture Steps256 1024 4 4 4.4M 2: 1 B+P 16K 0.1 le’4131K

[0413] MHC-II. The MHC-II presentation task involved use of three continuous features as inputs, namely gene expression, gene bias, and shadow. These continuous input variables may be tokenized using binning strategies, or directly, e.g., as text representing a 16-bit floating point number. Results for various different representation strategies are shown for binding and presentation tasks in FIGs. 28A and 28B for models with parameters shown in Table 10. The binding performance results shown in FIG. 28A appear to indicate similar performance, while the presentation performance data shown in FIG. 28B points to best presentation performance being reached by tokenizing continuous values with a quantile bins strategy. Accordingly, the quantile binning strategy was selected for use in subsequent examples.Table 10: Details of the model and training parameters used for MHC-II models._ Model Details _ Training Details _ Embed MLP Heads Layers Params Task Batch Dropout LR Train Dim Dim Mixture Steps256 1024 4 4 4.4M 2: 1 16K 0.2 le’313 IKB+P

[0414] Results of a sweep over differing numbers of bins used in the quantile bins strategy are shown in FIGs. 29A and 29B. FIG. 29B shows that 5 and 10 bins appeared to perform best on the presentation task. Since model selection was performed on global Top-K only, 5 bins was chosen as a baseline number of bins in subsequent examples. However, FIGs. 29A and 29B show that when considering mean per-allele Top-K, it may be beneficial to proceed with 10 bins in MHC-II models.

[0415] FIGs. 29A-29B also show an entry for ’5 unique bins’. This is a variation of the bin tokenization strategy whereby the same bin tokens are utilized for each continuous field. The results show that in terms of global Top-K using distinct bins per field performs better than sharing bin tokens across the different continuous fields.

[0416] Turning to FIGs. 30A-30B, overall performance on MHC-II binding and presentation were compared for a T5 multi-task model trained on binding and presentation data with single task models, NetMHCPan, NeonMHC, and the task specific pipeline model as described in Example 1. FIG. 30A presents binding results and FIG. 30B shows presentation results. As shown in FIGs. 30A and 30B, a 4.4 million parameter T5 model outperforms the baseline task-specific pipeline model described in Example 1, which in turn outperforms both NeonMHC and NetMHCIIPan4.0. Since both NetMHCIIPan4.0 and NeonMHC are binding models, to evaluate them on the presentation task, their output was incorporated via a trained generalized linear model (GLM).

[0417] Accordingly, this example demonstrates an end-to-end presentation predictor built using a single T5 neural network. Co-training the presentation predictor with binding task data lead to significantly improved presentation performance. Various mixture ratios between binding and presentation were empirically evaluated and it was found that either a 2: 1 or a 4: 1 binding to presentation ratio produces the best performance. A ratio 4: 1 was selected for the subsequent models.

[0418] This example also demonstrates and evaluates performance of several options for tokenizing input representations and evaluated their impact on performance. Tokenizing MHC-II alleles as pseudo-sequences led to improved task performance for the binding and presentation multi-task predictors, while for MHC-I alleles the highest presentationperformance was produced using allele names. The later result was further investigated by testing pseudo-sequence and allele names binding models on unseen alleles. It was found that the pseudo-sequence models generalize beter to unseen alleles. Several strategies for discretizing continuous features in the presentation task were also investigated. It was found that quantile binning performs well.(2) Multi-Task Modelling of MHC-I Binding, Presentation and Cleavage

[0419] A T5 model was also developed and evaluated in predicting proteasomal cleavage. Results demonstrating performance and approaches for using a T5 model to predict proteasomal cleavage in a single-task seting are provided, followed by results demonstrating and evaluating incorporation of the cleavage task as a new task in multi-task binding and presentation models.Table 11: Details of the model and training parameters used to train the single-task T5 cleavage predictor.Model Details Training DetailsEmbed MLP Heads Layers Params Batch Dropout LR Label TrainDim Dim Smoothing Steps128 512 4 6 1.7M 32K 0.3 5e’40.3 13 IK

[0420] A T5 model with parameters as shown in Table 11, above, was developed and trained for comparison with the task-specific pipeline model described in Example 1. Results, shown in Table 12, indicate that the single-task T5 cleavage model is on par with the taskspecific pipeline model benchmark.Table 12: Comparing T5 to a Task-Specific Pipeline model in single-task cleavage prediction on Tune set.Model Global Top-K Global APT5 0.671 0.736Task-specific pipeline 0.674(Example 1)

[0421] Additional performance improvements were obtained by training the T5 model on an altered version of the cleavage dataset which contains additional decoys (1 :20 vs 1:3). Table 13 shows the improved performance on the external Wolf-Levy dataset (dataset described in Wolf-Levy et al., 2018). Accordingly, on this, the 1:20 dataset was used in the multi-task models described below.Table 13: Comparing Hit to Decoy ratio in single-task cleavage prediction on Wolf-Levy.HD Ratio Global Top-K Global AP1:3 0.263 0.2311:20 0.267 0.250

[0422] The cleavage task was included as an additional task in a multi-task training approach, whereby a mixture of 4: 1 binding and presentation data were used are shown. Performance results are shown in FIGs. 31A-B, indicating that in terms of both the binding and presentation performance of the model, adding the cleavage task results in comparable task performance. Details of the various model parameters are shown in Tables 14 and 15, below.Table 14: Details of the “Mini-Tall-Thin” model and training parameters.Model Details Training DetailsEmbed MLP Heads Layers Params Batch Dropout LR TrainDim Dim Steps128 512 4 6 1.7M 32K 0.2 le’413 IKTable 15: Details of the “Small” model and training parameters.Model Details Training DetailsEmbed MLP Heads Layers Params Batch Dropout LR TrainDim Dim Steps512 2048 8 6 26M 32K 0.2 le’413 IK

[0423] Performance of the T5 model trained on MHC-I binding and presentation data as well as 1:20 hit: decoy ratio cleavage data was compared with the stacked / pipeline model described in Example 1, as well as the NeonMHC and NetMHCPan4.1 models. The results,shown in in FIGs. 32A-32B indicate that on both the MHC-I binding and presentation tasks, the T5 model performs on par with the stacked / pipeline baseline, which, in turn, outperforms both NeonMHC and NetMHCPan4.1. Both NetMHCPan4.1 and NeonMHC are binding models. Accordingly, to evaluate these models on the presentation task, their output was incorporated into a trained generalized linear model (GLM).

[0424] An additional implementation of a single T5 model was trained on the combination of binding, presentation, and cleavage data in a second, updated, in-house, binding and presentation dataset containing roughly a billion training rows. In this implementation, a large model configuration was used — a 9 Layer, 512 Feature Transformer totaling around 38M parameters (see Table 16).Table 16: Details of the model and training parameters used to train a T5 model.Model Details Training DetailsEmbed MLP Heads Layers Params Task Batch Dropout LR TrainDim Dim Mixture Steps512 2048 8 9 38M 4: 1:0 1 32K 0.2 7e’5131KB+P+C

[0425] The second, updated, binding and presentation dataset included multi-allelic examples in all partitions (train / test / tune). In order to provide for a single end-to-end model training and evaluation pipeline each allele of an MA example was input into the model simultaneously. For example, the text “peptide: AABB allele: A10101;B10101” may directly be input to the T5 model used herein. This approach is simplified in comparison with deconvolution approaches and simplifies both training and evaluation of the model on MA data.

[0426] FIGs. 33A and 33B provide results for performance of a T5 model with a taskspecific pipeline model trained on SA and MA data on two different evaluation datasets.

[0427] Accordingly, this example demonstrates that a multi-task language model that uses a unified, all text, input and output format - i.e., the T5 model - is capable of modelling a cleavage task in a single-task setting, showing comparable results to a task-specific pipeline model benchmark. Performance on the cleavage task may be increased by including more decoys at a 1 :20 ratio. Co-training binding and presentation with a cleavage task was also evaluated and demonstrated.

[0428] This example also demonstrates performance of an end-to-end T5 multi- allelic presentation predictor. Notably, a T5 model is able to leverage MA training data to massively increase performance, with no architectural alterations, and only minor changes to how inputs are tokenized.E.iv. Example 4: LLM-Based Antigen Presentation Modeling Results for Public Datasets

[0429] This example describes implementation of a T5 model in accordance with various embodiments of the immune response prediction technologies of the present disclosure in order to model the NetMHCpan-4.1 dataset collection. The collection contains datasets for two tasks: Eluted Ligand (EL) and Binding Affinity (BA) prediction. First, different methods for multi-task training are evaluated across these tasks. Next, various methods for improving performance on the EL prediction task are explored through various resampling and data augmentation techniques.

[0430] The EL dataset was split into two parts, Multi-Allelic (MA) and Single Allelic (SA) and treated as separate tasks. MA data is more challenging to model, both in terms of the implementation complexity and computational intensity. Accordingly, experiments described herein began with developing a model on SA tasks and later extending the bestperforming approach to the MA task. In this example, multi-task models are denoted with labels identifying the tasks that they were trained on, separated by slashes. For example, a model labelled “SA / MA / BA” was trained on the “SA”, “MA” and “BA” tasks.(7) Multi-Task T5 Modeling of the NetMHCpan-4.1 DatasetsModelling Single-Allelic Eluted Ligand Prediction

[0431] Models were trained on the examples containing HLA alleles in the NetMHCpan-4. 1 SA data (referred to hereafter as “HLA SA”). As these were the first experiments on this dataset, hyper-parameter tuning was performed by training six different models. An ensemble of models using the best-performing set of hyper-parameters was then created and evaluated.

[0432] Experimental Design. The HLA SA dataset has five pre-defined train-tune splits (folds). This allows for 5-fold cross validation and model ensembling. The models were trained on all five train sets and with subsequent validation on the respective tune sets.The validation per-allele average precision (AP) scores for the different configurations are shown in FIG. 34 and the configuration can be found in Table 17. All models used pseudosequences to represent the alleles.Table 17: Hyper-Parameters of different model configurations used on the NP Split 0 Tune.Name „ . . .. „. MLP HiddenT TT> ■ ■Embedding SizeQ. Layers Heads Head Size izeConfig 1 128 256 4 4 64Config 2 256 512 3 4 48Config 3 256 256 1 3 64Config 4 256 256 4 3 64Config 5 256 256 5 4 48Config 6 256 128 4 4 48

[0433] Ensembling was performed by averaging the predictions from five models trained on each of the cross-validation folds. The ensemble was evaluated and compared to a task-specific ensemble model based on the pipeline approach described in Example 1 and trained on the same data.

[0434] Hyper-Parameter Tuning Results. Results for a hyper-parameter sweep are shown in FIG. 34. Configuration 6 provided the best Top-K and AP scores on average and was thus selected for further experiments. Configuration 6 six is hereafter referred to as “T5- 660k.”

[0435] Ensembling Results. The scores of the ensemble models on the MS Ligands and CD8 test sets are shown in FIGs. 35A-B and 36, respectively. The T5 Ensemble and Task-specific (Example 1) Ensembles showed comparable and consistent Top-K performance with a few outliers. The T5 Ensemble has a slightly higher mean and median AP with a similar number of outliers. Both ensembles demonstrated comparable performance, with a similar range in variance, on the MS Ligands test set.

[0436] FIG. 36 shows the performance of the HLA SA binding models on the CD8 test set. The T5 ensemble has a significantly lower mean and a higher number of perfect frank scores than the Task-specific (Example 1) model. However, the Task-specific (Example 1) ensemble has a slightly lower median than the T5 ensemble. The T5 ensemble outperformed the Task-specific (Example 1) ensemble on the CD8 test set.Threshold Classification and Model Selection

[0437] The evaluation framework was also used the test set logits to determine a threshold at which the model differentiates between a binder and non-binder. This threshold allowed to calculate metrics such as Fl score, precision, and recall. However, the test set should only be used as an out-of-distribution set. Determining the threshold on the test set also resulted in thresholds that would maximize performance on selected metrics.Accordingly, to ensure accurate and robust threshold determination and evaluation, instead of determining the threshold from the test set, five different methods of determining the threshold on the validation set were explored:1. Global Threshold• Maximize Fl score• Top-K (threshold at which precision equals recall)2. Per- Allele Thresholds• Maximize Fl score• Top-K (threshold at which precision equals recall)3. 50% Threshold

[0438] Experimental Design and Results. Since the NetMHCpan-4.1 dataset was split into five train and tune splits, a submodel on each of the five train splits was trained.For each submodel, the logits on the corresponding validation split were calculated. Using the logits, the threshold was calculated at which the enumerated metrics above are maximized. The thresholds across each validation set were then averaged. These thresholds were then applied to the MS Ligands test set to calculate threshold-dependent metrics, such as accuracy, balanced-accuracy, recall, precision, and Fl score. FIG. 37 shows the metrics on the MS ligands test set calculated using the five approaches for determining thresholds. Using the threshold that maximized the global Top-K generally resulted in the highest accuracy, balanced-accuracy, recall, and Fl score on the MS Ligands test set.

[0439] Threshold Independent Metrics. An additional challenge related to the threshold-dependent metrics is model selection. Given the selection of the classification threshold, the metrics may vary across models and testing datasets. Average precision offers an alternative route to model selection that provides an estimate of model performance that is threshold independent. Average precision is calculated as the area under the precision-recallcurve, essentially providing information on precision and recall at all thresholds. Using average precision as an additional metric for model selection provides a more robust estimate of model performance.Impact of Quantization on Binding Affinity Signal

[0440] Binding affinity (BA) is a continuous feature and is used as both an input and a target for the T5 models. The quantization was used as an approach for inputting and inferring continuous values in the T5 text-to-text model. Quantization is the mapping of continuous, infinite values to a smaller set of discrete, finite values. There were two quantization approaches explored to model continuous values which are shown below. In this section the impact of the different quantization approaches on the BA dataset is investigated.1. Binning: placing continuous values into a predefined number of buckets that are tokens in the vocabulary (GATO) (S. Reed et al., 2022).2. Float 16: converts continuous values into float 16 dtype, tokenizing each digit (PaLM and GPT) (S. Reed et al., 2022, T. Brown et al. 2020).

[0441] Measuring Impact of Quantization via Mean Squared Distortion. When the BA values are quantized, the signal is distorted. This distortion leads to lower precision when inputting BA as a feature or predicting BA as a target. The distortion can be theoretically calculated using a metric called mean squared distortion (MSD), which indicates the difference between the original signal, and the distorted signal. The MSD for each quantization approach is shown in FIG. 38.

[0442] The floatl6 approach can be viewed as a binning approach using 10,000+ bins, whereas the binning approaches use a maximum of 1,024 bins, following the GATO protocol.Modelling Binding Affinity Data

[0443] In this section, T5 models were trained and evaluated to predict BA values quantized using the different approaches. The performance of different quantization approaches on the NetMHCpan-4. 1 BA train and tune datasets was evaluated.

[0444] Experimental Design. The models shown in Table 18 were trained, with the same architecture, varying only the quantization approach. For the binning approaches a hyper-parameter sweep over the three binning strategies and the number of bins wasperformed. Lastly, the Floatl6 approach was explored. All models used the T5-4M config and used only pseudo-sequences to represent alleles.

[0445] Results. Table 18 below shows the regression metrics of T5 models trained on the NetMHCpan-4.1 train set and evaluated on the tune set. There is no test set for the BA data, so the tune set was used to estimate the performance of different quantization approaches. K-means quantization using 25 bins has the lowest MAE, with the uniform strategy with 10 bins having the lowest MSE and highest R2-score. However, the performance of the different approaches appeared to be comparable, with no significant change in performance metrics when the quantization strategy and corresponding number of bins was changed.Table 18: Metrics showing the quality of the binding affinity predictions made by T5 models using different quantization approaches on the NetMHCpan-4. 1 binding affinity tune set.Quantization Strategy Num. bins MSE MAE R2score10 0.0332 0.1194 0.5807K-means 25 0.0337 0.1090 0.573350 0.0353 0.1117 0.553210 0.0326 0.1186 0.5880Uniform 25 0.0342 0.1137 0.567050 0.0359 0.1123 0.5457Quantile 7 0.0334 0.1134 0.5767Multi-Task Training Eluted Ligand and Binding Affinity Prediction

[0446] This section presents the work of multi-task training BA and eluted ligand prediction using the aforementioned datasets. These models will be referred to as “T5 SA / BA” models.

[0447] Experimental Design. The impact of different “mixture ratios” between the HLA, SA and BA data was investigated, where the term “mixture ratio” refers to the relative frequencies at which the model is shown examples from each task during training. All models used the T5-660k config and used only pseudo-sequences to represent alleles.

[0448] Results on Tune Set. FIGs. 39A and 39B show average precision and Top-K performance, respectively, on the SA tune set of the T5 models trained at varying mixtureratios of BA and HLA SA data. FIG. 40 shows the performance of the same models on the BA tune set. The models in the figures are labelled according to their SA:BA mixture ratio.

[0449] Performance on the SA task tended to be higher where the mixture ratio favors the SA task and lower where the ratio favors the BA task. However, the performance on the BA task did not tend to improve as the proportion of BA examples in the mixture ratio was increased. Without wishing to be bound to any particular theory, it is hypothesized that the latter was due to the model overfitting the BA training data.

[0450] The T5 models trained with mixture ratios of 3: 1 and 1: 10 have the narrowest interquartile ranges (IQRs), indicating a higher consistency in model performance. The 1: 10 mixture ratio also shows the highest median and mean AP. The models with a higher proportion of BA samples showed lower mean and median AP scores and a larger IQR, indicating worse performance than the models trained on mixture ratios where there are more HLA SA samples than BA samples.

[0451] Similar to the AP metrics trend on the tune set, models trained with a mixture ratio where the HLA SA formed the majority of the training samples in a batch showed better Top-K scores on the tune set. There was a noticeable Top-K performance difference between the T5 10: 1 model and the remaining models. The best performing model based on the binding classification tune set uses a 10 : 1 ratio which is close to the natural ratio of 17 : 1. These results suggest that models should generally be trained with a balanced (1: 1) mixture ratio or a mixture ratio where the SA samples from the majority of the training samples.

[0452] In summary, the model trained at a ratio 3 : 1 was found to provide the second- best metrics on the SA tune set and the highest R2-score on the BA tune set. The T5 model trained at a 3: 1 mixture ratio and tested on the MS Ligands and CD8 test sets was therefore selected for further use.

[0453] Results on MS Ligands Test Set. FIGs. 41A and 4 IB show the Top-K and average precision performance, respectively, of the best model trained on SA data only, and the T5 3: 1 model trained on both HLA SA and BA data. It is important to note that the T5 SA:BA models are not ensemble models. This comparison was used to quantify the changes in performance on the MS Ligands and CD8 test sets after adding BA as an additional task.

[0454] FIGs. 41A-B show the T5 SA:BA model had a slightly higher mean and median Top-K and AP scores. However, the IQR is slightly larger on the T5 SA / BA model indicating that this model is less consistent, with a wider spread of per-allele Top-K and APscores. The addition of BA data resulted in better performance than an ensemble of T5 models trained on all HLA SA training data splits.

[0455] Evaluating CD8 Test Set via the Binding Affinity Task. Adding the BA regression task offers an additional method of evaluating EL tune or test sets. The CD8 test set can be evaluated as a regression task. When predicting binding affinity scores, a bin is predicted which indicates the range in which the predicted binding affinity value lies. The median value of the bin (determined from the training data) was then taken and outputted as the predicted binding affinity score. However, for evaluating the CD8 test set as a binding affinity task, the log-likelihood over all possible bins the model could predict was extracted. The log-likelihood score indicates the probability that the predicted value lies in a bin. A weighted average of the bins was then taken where the bins are weighted depending on the associated likelihood score for that particular bin. The weighted average was outputted as the prediction and then these predictions were ranked to calculate the FRANK scores on the CD8 test set.

[0456] Results on CD8 Test Set. The results on the CD8 test set are shown in FIG. 42. The T5 SA Binding model is the ensemble model. This model evaluated the CD8 test set as a binding classification task, indicated by the binding label. The T5 SA / BA model can evaluate the CD8 test as a BA task or a binding classification task. The Binding Affinity label indicates that the regression / BA task prefix was used to evaluate on the test set. Finally, the T5 SA / BA Binding label indicates that the binding classification task prefix to evaluate on the test set.

[0457] FIG. 42 shows the performance of the two models and evaluation methods on the CD8 test set. Adding the BA data and then evaluating the CD8 test set as a binding classification task yields the best results. This suggests that the BA distribution is like CD8 test set distribution. There is a large performance improvement on the CD 8 test when trained on BA data in addition to the HLA SA dataset. Finally, evaluating the CD 8 test set as a BA regression task resulted in the weakest performance. Without wishing to be bound to any particular theory, it is believed that this may be due to the quantization approach used.Multi-task Training T5 Models on Single / Multi-Allelic Eluted Ligand and Binding Affinity Data

[0458] This section presents the work on training T5 models on the entirety of the NetMHCpan-4. 1 dataset. Training on the entire dataset requires to handle EL / MA data. The two approaches explored to integrate the MA data are deconvolution, similar to the NNAlign_MA approach, and an end-to-end approach.

[0459] Experimental Design. The first approach followed the NetMHCpan-4.1 NNAlign approach, a T5 model was trained on the SA and BA data. This model was then used to deconvolute the multi-allelic data into single-allelic data. A new model was then trained on the SA, BA, and deconvoluted MA data. This deconvolution approach is abbreviated to DCNV.

[0460] The second approach is the end-to-end approach where the model does implicit deconvolution. Rather than deconvoluting the multi-allelic data explicitly, the peptide and all alleles constitutes a single MA training sample. This approach is significantly faster than the deconvolution approach that comprises training a deconvolution model and then using that model to deconvolute the MA data. The end-to-end approach is abbreviated to E2E.

[0461] For the E2E and DCNV approaches models were trained using the pseudosequence (PS) and allele name (AN) representations. For each of the E2E and DCNV approaches using the PS and AN representations, models were trained at two different ratios. The natural ratio (Nat) which is the natural ratio between the SA / MA and BA data. The second mixture ratio is the balance ratio (Bal), where the mixture ratio is roughly the same between the three data sources. The natural ratio was determined to be 19: 19: 1 (MA:SA:BA) using the UniMax sampler, whereas the balanced ratio was 2:2: 1 (MA:SA:BA). The E2E and DCNV models were trained on a single train split. All models were trained using the DeepNarrow T5-2M config.

[0462] Results on Tune Sets. The results of the models on the tune sets are shown in FIGs. 43-45. Models that use deconvolution are abbreviated to DCNV, whereas end-to-end models are abbreviated to E2E. The labels AN and PS refer to the allele name vs pseudosequence representations respectively.

[0463] The E2E model trained at the natural ratio using the allele-name representation has the highest global Top-K and AP. Generally, the models trained with a natural mixtureratio have higher AP and Top-K scores than their balanced mixture ratio counterparts. Similar performance between the E2E and DCNV models was observed. These results on the tune set indicated that the T5 model is capable of implicit deconvolution. The models trained at the natural ratio generally have lower R2-scores on the BA tune set, which is expected since the model sees fewer of the BA training samples. The E2E AN Nat model is an exception to this trend with the second highest R2-score after the DCNV PS Bal model. The E2E AN Nat model was selected as the final model to test on the MS Ligands and CD8 test set since it had the highest global Top-K and AP, along with the second highest R2-score on the BA tune set.

[0464] Results on MS Ligands Test Set. FIGs. 46A-46B show the performance of the penultimate T5 and version of the task-specific pipeline model (described Example 1) models alongside the NetMHCpan-4.1 model on the MS Ligands test set. It is important to note that the final task-specific pipeline model used for comparison is an ensemble of 15 submodels, with 3 models trained on each of the five training splits. The final T5 E2E model is a single model trained on a single train split. The performance of a single submodel in the task-specific pipeline model ensemble is also included for additional performance comparison.

[0465] From FIGs. 46A-46B, the T5 and task-specific pipeline models both outperform NetMHCpan-4.1 on the MS Ligands test set on both the Top-K and AP metrics. The T5 model has higher AP scores than the task-specific pipeline submodel trained on the same data, however, the Top-K scores between these two models are comparable. The taskspecific stacked / pipeline ensemble model and T5 show comparable performance on the test set. Since the T5 model is better than the task-specific pipeline model submodel, training T5 models on all training splits and ensembling these models will lead to comparable, if not better, performance than the task-specific pipeline model ensemble. The MS Ligands results are particularly promising as T5 integrated the multi-allelic data using the end-to-end approach which suggests the model is capable of implicit deconvolution. The task-specific pipeline submodel and ensemble were trained using deconvolution which is a complex and error-prone approach to integrating MA data. The T5 end-to-end approach significantly simplifies how multi-allelic data can be integrated, resulting in lower computational costs, shorter training times, and no error-propagation from one model to another as seen in deconvolution approaches.

[0466] The precision-recall curves and ROC curves are shown in FIGs. 47A and 4B8, respectively.

[0467] The task-specific pipeline model ensemble (15) has a PR-AUC of 0.87, which is the highest among the four. This suggests that on average, this ensemble model maintains a better balance between precision and recall across different thresholds, making it the superior model in this regard. T5 has a PR-AUC of 0.85, placing it second in this metric. It is slightly less precise or has a slightly lower recall than the task-specific pipeline ensemble, but it is better than the NetMHCpan-4. 1 and task-specific pipeline submodel.

[0468] FIGs. 47A and B shows that T5 outperforms the other models with a ROC- AUC of 0.96, suggesting it has the best discrimination ability between the positive and negative classes at various threshold settings. The task-specific pipeline ensemble (15) also has a ROC-AUC of 0.95, indicating that it is as effective as NetMHCpan-4. 1 in classifying the positive class correctly while keeping the false positives low. The task-specific pipeline submodel has a ROC-AUC of 0.93, which is still high but slightly lower than the other models, indicating that while it discriminates well, there may be a slightly higher rate of false positives or false negatives compared to the others.

[0469] In summary, when considering both PR and ROC AUC metrics, the taskspecific pipeline ensemble (15) and T5 models are superior, with T5 having a slight edge in ROC-AUC and task-specific pipeline ensemble (15) leading in PR-AUC.(2) Improvements to T5 Models on the NetMHCpan-4.1 Datasets

[0470] In this section, methods for improving performance on the EU prediction task are explored through various prefix-selection, resampling, data augmentation strategies.Impact of Task Prefixes on Performance

[0471] T5 uses input prefixes to differentiate between tasks during training and inference. Additional identifiers can be included to differentiate between datasets that fall under the same task. The elution assay data can either be single or multi -allelic. Different prefixes are used for the datasets in addition to the tasks. Three options for task and dataset prefixes were explored on the EE single and multi-allelic data:1. Same prefix - a single prefix token is used for all tasks (e.g., “Binding”).2. Common base prefix with unique identifier - prefix consists of multiple tokens where the some of the tokens are shared across tasks (e.g., “Binding” followed by “SA” or “MA”).3. Separate prefix - a single unique token is used for each task (e.g., “SA-Binding” or “MAbinding”).

[0472] Experimental Design. While these prefixes are used in training, any of these prefixes can also be used when evaluating on different test sets. Three separate models were trained using the three prefix notations enumerated above to investigate if different prefix notations resulted in performance differences. All the models were trained using the T5-660k config and used only pseudo-sequences to represent the alleles.

[0473] Results on MS Ligands Test Set. FIGs. 48A-48D shows the performance of the three models when evaluated using the single and multi-allelic prefixes on the MS Ligands test set. “Binding” represents the same prefix notation, “Binding MA” represents the common prefix notation, and “Binging-MA” represents the separate prefix notation. The model that uses the “Same prefix” notation is used as a baseline against which to compare the impact of using dataset identifiers during evaluation. FIGs. 48A-48B show the model that uses the same prefix notation i.e., “Binding” has comparable performance to the single-allelic prefix notations on the MS Ligands Test set. This indicates that there was no performance improvement by adding the single-allelic dataset identifier during training and evaluation on the MS Ligands test set. FIGs.48C-48D show the performance of the three models when evaluated using the different multi-allelic prefixes. The “same prefix” model results in the best average precision metrics on the MS Ligands. The Top-K metrics when evaluated via multi-allelic prefixes show worse performance on Top-K and AP metrics than the same prefix notation. The results indicate that training and evaluating using different prefix notations yields different results when evaluated on the MS Ligands test set. The common prefix notation with a unique identifier for the dataset that the training data is associated, in this case single-allelic, showed the best performance on the MS Ligands test set. This suggests that the distribution of the SA training data is similar to the MS Ligands test set, whereas the EL MA dataset has a slightly different distribution. In addition to different distributions, the quality of the SA data is considered higher than the EL MA dataset which provides additional evidence as to why training and evaluating via the multi -allelic prefix dataset identifiers yield worse performance on the MS Ligands test set.

[0474] Results on CD8 Test Set. FIGs. 49-50 show the results of having evaluated the three models using the three separate prefix notations on the CD8 test set. “Binding” represents the same prefix notation, “Binding MA” represents the common prefix notation, and “MA-Binding” or “SA-Binding” represent the separate prefix notation. The single -allelicprefix notations resulted in lower mean, median, and perfect number of Frank scores on the CD8 test set compared against the baseline common prefix model.

[0475] The multi-allelic prefix notations however resulted in higher mean, median, and perfect number of Frank scores on the CD8 test set compared against the baseline common prefix model. These results, similar to those shown on the MS Ligands test set, suggest that the SA training set has a distribution closer to that of the two test sets than the EL MA test set. The quantization approach does not effectively model BA as well as the task-specific stacked / pipeline approach, and as such, evaluations on the CD8 test set using the BA prefix always resulted in poor performance.Fine-tuning

[0476] Here the impact of fine-tuning a T5 SA / MA / BA model using the best checkpoint for each particular task is investigated. The model is fine-tuned on the BA and binding classification tasks with the aim of further improving performance on downstream tasks.

[0477] Experimental Design. A T5 base T5-660k model trained on SA, MA and BA data at a mixture ratio of 19: 19: 1 (SA: MA: BA) was fine-tuned on the SA data in an attempt to improve performance on downstream tasks like binding classification on the MS Ligands test set. The base model was fine-tuned using two different learning rate schedulers, the first used a constant learning rate whereas the second used warmup.

[0478] Results on Tune and MS Ligands Test Sets. The results on the tune set for the base model and fine-tuned models are shown in FIGs. 51A-D. “Base model” indicates the base model, “Fine-Tuned const Ir” indicates the fine-tuned model using a constant learning rate, and “Fine-Tuned wrmup Ir” indicated the fine-tuned model using a warmup learning rate scheduler. The results of the fine-tuned models on tune sets indicate that there were no performance gains as shown by comparable tune set AP and Top-K box plots. The fine-tuned models were still tested on the MS Ligands test set; however, no performance improvement was expected.

[0479] As expected, there was no performance improvement on the MS Ligands test set as a result of fine-tuning the base model. Fine-tuning the model on single-allelic data resulted in slightly lower performance than the base model. While this is not a large drop in performance, fine-tuning was not expected to significantly change performance on the MSligands test set as the single-allelic data used to fine-tune the base model was already seen in most training batches at a ratio of 19: 19: 1.

[0480] Results on CD8 Test Set. Multi-task training on the single-allelic and BA data resulted in improved results on the CD8 test when compared against the models trained on SA data alone. Since the base model was trained at the natural ratio of 19: 19: 1, finetuning the base model was investigated using the BA data in an attempt to further improve performance on the CD8 test set.

[0481] As shown in FIG. 5 IE, a slight performance improvement in the median and mean FRANK scores is observed along with a reduction in the number of perfect FRANK scores. These results suggest that fine-tuning on smaller datasets like BA can improve the performance of the base model on downstream tasks.Allele Resampling

[0482] In this section, the impact of smoothing the allele frequency distribution of the NetMHCpan-4. 1 data during training is investigated. The online resampling is performed to achieve the target allele distribution described below. Having a more even distribution of allele frequencies during training may improve performance on rare alleles thereby improving the aggregate per-allele metrics.

[0483] Allele Frequency Smoothing. A target allele distribution is chosen by linearly interpolating between the source distribution and the uniform distribution using the label smoothing formula. Let P(A) be the natural probability of sampling of allele A from the training set. In this experiment, A is sampled with probability P'(A). defined asP'(A) = (I - e) P(A) + e / K, where e G [0, 1] is a hyper-parameter and K is the number of alleles. Note how, if e = 0, then P'(A) = P(A), and if e = 1, then P'(A) = \!K. which is the uniform distribution. Intermediate values of e linearly interpolate between P(A) and l / K (see FIGs. 52A-52C).

[0484] Results. FIGs. 53A-D show the results for the allele resampling experiments. All models used the T5-660k config and pseudo-sequence allele representations and are trained on the NetMHCpan-4.1 SA-only data. The optimum values for r range from 0. 1 to 0.5 for the different evaluation sets and metrics.

[0485] Larger r values may improve the performance on rare alleles by oversampling them during training, at the expense of slightly degrading performance on more frequentalleles. For some values of e, the small performance losses in frequent alleles are outweighed by larger improvements on rare alleles.

[0486] When e is very large, performance degrades - without wishing to be bound to any particular theory, it is believed that this may be due to the reduced diversity in training examples, since the examples from rare alleles are repeated very frequently.Hit .Decoy Resampling

[0487] Online resampling during training was performed to test different hit decoy ratios.

[0488] Results and Discussion. The results are shown in FIGs. 54A-54D. All models are T5-660k with pseudo-sequence-only allele representations and are trained on the NetMHCpan-4. 1 SA-only data. The hit: decoy ratios of 1 : 10 and 1: 17 (the natural ratio occurring in the dataset) are the best performing on the tune and test sets respectively. Two potential explanations for this are considered. The first is that the NetMHCpan-4.1 SA training data already contains an optimal or near-optimal hit: decoy ratio for the current training setup, thus resampling is not required. The second is that resampling always reduces performance.

[0489] Preliminary experiments with hit: decoy resampling on the in-house dataset suggested that the former explanation is more likely. The proprietary dataset was resampled, which naturally had a hitdecoy ratio of 1 :500, and it was found that ratios closer to 1: 10 resulted in faster convergence and better final performance.Task Inversion

[0490] One technique in NLP is the use of “denoising” training objectives to improve the quality of learned representations. Typically, tokens or spans of tokens are randomly masked, and the model is trained to predict the masked spans. In this section, a similar method for Eluted Ligand (EL) binding prediction is tested. Specifically, new “inverted” forms of the binding prediction task were constructed and the models were trained in a multitask fashion on both the original binding prediction task and its inverted forms.

[0491] The inverted tasks are shown in Table 19. Only hits were used for auxiliary tasks. For the MA data, “presenting allele” refers to the allele that is expected to be most likely to bind. The presenting allele is chosen using annotation by SA data where possible, or by choosing the max-scoring allele using NetMHCpan-4. 1 if no SA annotation is available.Table 19: The input and output features for the inverted binding tasks. Task 1 reads, “given a hit peptide from the SA dataset, predict the presenting allele”.Task No. Source Data Input Target1 g peptide presenting allele2 presenting allele peptide3 phenotype, peptide presenting allele4 phenotype, presenting allele peptide

[0492] SA Results. FIGs. 55A-55D show that task inversion improves performance of the T5-25M model on the MS Ligands test set, but not on the NetMHCPan-4. 1 tune set. The T5-660k model does not appear to benefit from task inversion. This effect may be because it has insufficient model capacity to benefit from the effective increase in training dataset size. All models used pseudo-sequence allele representations.

[0493] MA Results. FIGs. 55E-55H show that the inclusion of the inverted tasks slightly improved the pseudo-sequence model performance on the NetMHCpan-4. 1 tune set, but not the MS Ligands test set. The performance of the allele-names model was not improved on either evaluation set. All models used the T5-25M config.Synthetic Multi-Allelic Data

[0494] In this section, synthetic multi -allelic samples were created fortraining models using existing SA and MA data. The synthetic data may improve the E2E model’s ability to implicitly deconvolute MA examples. Let ({A,5,C} A) denote an EL training example containing the phenotype (set of alleles) {A,B,C} and peptide X. The synthetic dataset is constructed by iterating over every existing phenotype, pairing it with every existing peptide and labelling the phenotype-peptide pair using the rules in Algorithm 1 :Algorithm 1. The synthetic MA data generation algorithm. if ({A,5,C}A), ({A, 5} A), ({5,C} A), ({A} A), ({5} A) or ({C}A) appear as a hit in some EL dataset then label ({A,5,C} A) as a hit. else if at least one of the above appear as a decoy in some EL dataset then label ({A,5,C} A) as a decoy.else exclude ({A,B,C},X) from the synthetic dataset. end if.

[0495] The process described above generates a dataset containing both the original examples from the source dataset and novel “synthetic” examples. Note that the phenotypes used in generating synthetic data can contain any number of unique alleles. Three unique alleles were used in the algorithm description above for illustration purposes only. After generating the synthetic data, the data was resampled such that the phenotype distribution matches that of the original data.

[0496] Results. The results are shown in FIGs. 56A-56D. All models were trained on NetMHCpan-4. 1 SA and MA data using the T5-25M config. The models labelled as “synthetic data” were trained on Synthetic MA data generated from the NetMHCpan-4.1 SA and MA in addition to the original data. The synthetic MA data reduced the variation in performance across alleles on the MS Ligands Test set for the pseudo-sequence model, but did not appear to have any beneficial effect on the performance on the NetMHCpan-4. 1 tune set. The synthetic data reduced the performance of the allele-names model.Randomized Allele Representation

[0497] Both the allele name and pseudo-sequence were provided as input to the model. For some alleles this performed worse than using pseudo-sequence alone. When provided with both representations, the model may ignore the pseudo-sequence and may learn to predict based on allele name alone since this may be a simpler task than making predictions based on pseudo-sequence.

[0498] In this section, randomly selecting different allele representations (allele names, pseudo-sequences or both) in an online fashion during training is explored. This procedure may force the model to learn to use both the pseudo-sequence and allele name features. During evaluation both representations are inputted.

[0499] The results are shown in FIGs. 57A-57D. All models used the T5-660k config and were trained on the HLA alleles of the NetMHCpan-4. 1 SA Fold 0 training set. On the tune set, the randomized representation model slightly outperforms the others in terms of average precision, but not Top-K. By contrast, on the test set, the randomized representations improve both the Top-K and Average Precision.Non-HLA EL Training Data

[0500] The NetMHCpan-4. 1 Eluted Ligand (EL) data contains examples with non- HLA (non-human) alleles. Thus far these were excluded from the training data of the models. This experiment investigates the effect of including this data on the MS Ligands and CD8 benchmark scores of the model. Without wishing to be bound to any particular theory, one possibility is that the CD8 benchmark includes more non-human peptides than the NetMHCpan-4. 1 EL training folds. The non-HLA NetMHCpan-4.1 data may include more non-human peptides than the HLA data, thus the inclusion of non-HLA may improve performance on the CD8 benchmark.

[0501] Results. The results are shown in EIGs. 58A-58C. The model trained with non-HLA examples seems to have slightly higher average precision scores, but lower Top-K scores on the NetMHCpan-4. 1 Eold 0 tune set and its best checkpoint appears to have a slightly higher number of perfect scores on the CD8 benchmark. FIGs. 59A-59D show performance results for models trained on data with and without non-HLA examples.Comparing E2E and Deconvolution MA Models

[0502] One approach for using multi-allelic data is first “deconvoluting” it with an existing binding prediction model. This experiment compares the use of a single-allelic (SA) T5 model and an SA TinyTrf ensemble for deconvolution, when training a multi-allelic T5 model.

[0503] Results. Tables 20 and 21 show the results for different deconvolution methods, as well as the SA and e2e-MA models for reference. Deconvolution models that use TinyTrf trained on proprietary data, were included since the checkpoints were available. For allele-names, the E2E model performance is similar to the other MA model, outperforming them on the tune set and scoring close to them on the test set. However, when using pseudo-sequence, the E2E model scores are lower than the other models on both the tune and test sets.Table 20: Results on NP fold 0 SA tune set. The “Deconvolution Model” column shows the model used to score alleles when performing explicit. Here “In-house” and “NP-SA” denote models trained on the in-house and NetMHCpan-4.1 SA training datasets respectively. “E2E”denotes that the model used all alleles as input in an end-to-end fashion, rather than performing an explicit deconvolution step.Table 21: Results on the MS Ligands test set. The “Deconvolution Model” column shows the model used to score alleles when performing explicit. Here “In-house” and “NP-SA” denote models trained on the in-house and NetMHCpan-4.1 SA training datasets respectively. “E2E” denotes that the model used all alleles as input in an end-to-end fashion, rather than performing an explicit deconvolution step.Alternate Deconvolution Algorithms

[0504] In certain examples, above, when testing “deconvolution” models, deconvolution was performed by selecting the allele with the highest score according to some trained binding prediction model. However, there are other possible methods and augmentations for this procedure which are explored in the experiments below.

[0505] Calibration. When using an SA binding model for deconvolution, inference was run on a set of reference peptides and then the mean and std statistics were calculated for the binding scores for each allele. Outliers (z-score < -3 or z-score > 3) were iteratively removed and the above statistics was recalculated until there were no remaining outliers.

[0506] Random-Decoy Deconvolution. In experiments labelled “random decoy deconvolution”, multiallelic decoys were deconvoluted by randomly selecting an allele.

[0507] Results. The experiments ran twice with different implementations to check for reproducibility. The results are shown in FIGs. 60A-60D. The “calibrated + random decoy” deconvolution strategy performed slightly better than the others. However, the difference was small and in later reproductions of this experiment the ranking of the models was different. Neither the calibration nor random decoy selection yielded a consistent advantage. These more advanced deconvolution methods do not improve performance.Shuffling the Allele Input Order for E2E MA models

[0508] This experiment tests the impact of shuffling the input order of the alleles to the e2e MA models during training. This was tested only on allele-names models, since they are less compute -intensive than pseudo-sequence models.

[0509] Results. The results are shown in FIGs. 61 A-6 ID. The use of allele order shuffling during training does not appear to affect the Top-K scores, but reduces the average precision.

[0510] Accordingly, it was found that using a separate prefix token for each task performed similarly to more complex prefix designs, and that fine-tuning multi-task models did not substantially improve performance on MS Ligands but did provide minor improvements on the CD8 test set. Similarly, allele and hitdecoy resampling, MA task inversion, the use of synthetic MA data and the use of randomized allele representations did not result in substantial performance improvements. However, SA Task Inversion did yield a substantial improvement on the MS Ligands test set.E.v. Example 5: Direct Ligand Identification

[0511] An antigen protein sequence undergoes multiple serial processes in a cell before fragments (ligands) from it are presented for T-cell recognition, including proteolysis, cleavage, binding and presentation. Each of these processes can be modelled interdependently in a serial manner to identify ligands that will be presented for T-Cell regulation. This approach may, however, be computationally intensive, in certain cases, suffer from compounding errors (e.g., with errors in each step compounding). In order to overcome these drawbacks, the entire cascading process end-to-end can be modelled using a single model, whereby, given an MHC-I allele and antigen sequence the MHC -peptide complexes are identified that result from their interaction.

[0512] Models receive, as input, an MHC-I allele and antigen sequence, and generate, as output, identifications of ligands from the antigen sequence that are predicted as likely to be presented for T-cell recognition. Two approaches for accomplishing this task - a binary classification and a generate approach were developed and evaluated. Since the ligands are a linear fragment of the antigen sequence, a binary classification approach can be used, whereby each amino acid in an antigen sequence is classified as belonging to a ligand or not. In a generative approach, a generative model is asked to directly identify - e.g., generate, as output - the ligand sequences. These two approaches are illustrated shown in FIG. 62 and are referred to herein as ligand location and ligand generation, respectively.( / ) Ligand Location

[0513] By posing the problem as a binary classification task, a significant class imbalance problem occurs because less than 10% of the amino acids in an antigen sequence belong to a ligand. Two approaches in overcoming this class imbalance problem is explored for the identification of ligands:

[0514] 1. Weighted Cross-Entropy: In the classic cross-entropy loss formulation each class has a default weight of 1. However, for the purpose of addressing class imbalance this weight can be skewed in favor of a class(es) with low sample counts. The weight for the positive class to a value of 5 is tuned.

[0515] 2. Random Decoy Sampling: Another approach used to address class imbalance is resampling so that all classes have equal sample proportions. In this case equal number of decoys as ligands are randomly sampled.(2) Ligand Generation

[0516] When posed as a generative problem, direct ligand identification formulated as a language modeling extractive question-answering task. In this task, the model is provided with the allele and antigen sequence and generates ligands that will be presented for T-cell recognition. In most cases, multiple ligands can be derived from a single antigen sequence. This has led to the exploration of two approaches as shown in FIG. 63 :• Single Shot Ligand Generation: In this approach, each trio of allele, antigen, and ligand is treated as a training sample, so that each example with multiple ligands is expanded into samples with multiple trios. The resulting model only learns to generate a single ligand during training. During inference, multiple inference steps are rolled out to recover the prediction (multiple ligands) for each allele-antigen pair.• Multi Shot Ligand Generation: In this approach, the model is trained to predict all possible ligands that can result from an allele-antigen pair.Table 22: Results ofNetMHCPan-4.1-Trained Ligand Location Models.Approach MetricsAccuracy Precision Recall FlCD8 Benchmark Weighted Cross Entropy 0.12 0.24 0.20 0.22Random Decoy Sampling 0.11 0.19 0.21 0.20MS Ligands Weighted Cross Entropy 0.16 0.27 0.28 0.27Random Decoy Sampling 0.13 0.23 0.22 0.23(3) Results and Discussion

[0517] Models were developed by initially fine tuning an Ankh-Large model (A.Elnaggar et al., 2023) using a maximum sequence length of 1024 on the NetMHCpan-4.1 and an in-house datasets described herein. The performance of our models was then evaluated using three distinct test sets: CD8 Benchmark, MS Ligands and the in-house dataset. This comprehensive evaluation allowed to thoroughly analyze the effectiveness of our approaches across various metrics and datasets.• NetMHCPan-Trained Models: Using the NetMHCPan-4. 1 dataset, models were trained using both the ligand location and ligand generation approaches. The results,detailed in Tables 22-23, reveal that the ligand generation approach outperforms the ligand location approach.• In-house dataset-Trained Models: Training only ligand generation models using the in-house dataset was performed.

[0518] As shown in Tables 23-24, the single-shot model demonstrated better performance compared to the multi-shot model in most evaluated metrics.Table 23: Results ofNetMHCPan-4.1-Trained Ligand Generation Models.Approach MetricsAccuracy Precision Recall Fl, Multi Shot 0.13 0.30 0.18 0.22CD8 BenchmarkSingle 0 15 0.26 0.25 0.26Shot, Multi Shot 0.18 0.36 0.27 0.31MS LigandsSingle 0.200.31 0.35 0.33ShotTable 24: Results of In-House-Trained Ligand Generation Models.Approach MetricsAccuracy Precision Recall FlCD8 Benchmark Multi Shot 0.05 0.13 0.08 0.10Single Shot 0.08 0.10 0.29 0.15MS Ligands Multi Shot 0.17 0.29 0.30 0.29Single Shot 0.17 0.19 0.61 0.29_ , Multi Shot 0.12 0.46 0.13 0.21In-house -Test singleShot 0 22 0.34 0.39 0.36Comparison

[0519] After experimentally identifying the most effective method for predicting ligands directly from protein sequences, the results were compared with an optimal protocol using the best task-specific binding models (TinyTrf). The results in Table 25 illustrate that when trained on NetMHCPan, the T5 models generally perform better than TinyTrf across various test benchmarks. However, as shown in Table 26, while the T5 models continue to show competitive performance against TinyTrf on the test set corresponding to the trainingdata, they appear exhibit slightly less robust results on the CD8 Epitope and MS Ligands Benchmark when trained on the in-house dataset.Table 25: Comparison between NetMHCPan-trained TinyTrf and T5.Approach MetricsAccuracy Precision Recall FlCD 8 Benchmark TinyTrf 0.09 0.14 0.20 0.16T5 0.15 0.26 0.25 0.26MS Ligands TinyTrf 0.13 0.17 0.34 0.23T5 0.20 0.31 0.35 0.33Table 26: Comparison between In-House-trained TinyTrf and T5.Approach MetricsAccuracy Precision Recall FlCD8 Benchmark TinyTrf 0.10 0.12 0.36 0.18T5 0.08 0.10 0.29 0.15MS Ligands TinyTrf 0.22 0.23 0.86 0.36T5 0.17 0.19 0.61 0.29 TinyTrf 0.10 0.19 0.17 0.18 T50.22Q J40.39 Q.36(4) Direct Ligand Identification from Protein Sequence & Structure

[0520] Having determined the best approach using only protein sequences, the model’s performance was enhanced by incorporating protein structure information. As illustrated in PIG. 64, a multi-modal T5 architecture was developed. This was initialized with pre-trained Ankh-Base (A. Elnaggar et al., 2023) weights and subsequently fine-tuned by injecting structure embeddings from ESMFold (Z. Lin et al., 2022) into the encoder. The structure embeddings were injected using an additive operation as shown in FIG. 64 and Multimodal Adaptation Gate (MAG) (W. Rahman et al., 2020). In both cases a maximum sequence length of 512 was used. The results of the experiments shown in Table 27 signify that the custom bottleneck approach is the best and it yields a marginal improvement on the test benchmarks.Table 27: Results of Single-Shot Ligand Generation Using Sequence and Structure Features.Approach MetricsAccuracy Precision Recall FlSequence Only 0.12 0.21 0.22 0.21CD8 Benchmark Sequence+Structure0 13 0 23 0 23 0 23(Bottleneck)E.vi. Example 6: Peptide-MEIC Presentation Predictions

[0521] This example demonstrates performance of a multimodal framework utilizing transformer-based neural networks that integrate data from multiple predictive tasks, in accordance with certain embodiments described herein. Among other things, without wishing to be bound to any particular theory, identifying patient-specific neoantigens that can elicit robust immune responses is believed to be an important component of personalized cancer immunotherapies, yet it remains challenging. For example, a significant obstacle is believed to be the complexity of accurately modeling the sequential processing steps involved in the biological pipeline that results in an immune response, including peptide-MHC binding, presentation, stability, and final immune response recognition across multiple HLA alleles. The complexity of these processes is further compounded by factors such as RNA abundance and subsequent protein expression. Among other things, the multimodal transformer-based model of the present example aims to address these challenges, and data presented in this example demonstrates enhanced prediction accuracy across these processing steps, thereby providing an important approach that can facilitate expediting the design of personalized cancer vaccines.

[0522] Methods. The model of the present example leveraged transfer learning between immunopeptidomics tasks. The model utilizes, in the context of immune response prediction, a text-to-text transfer transformer (T5) approach (Aribandi, Vamsi, et al. (2022); Raffel, Cohn et al. (2020)), framing every task’s input and output in a text format. This design is scalable to many different tasks, avoiding the need for, for example, loss balancing or the management of multiple output heads.

[0523] As shown in FIG. 65, the text-based format of the input and output allows for processing of single allelic (SA) and multi-allelic (MA) data in the same manner, without a need for explicit deconvolution. In this manner, the model can also learn interactions between alleles.

[0524] Results. FIGs. 66A-B show a comparison between the T5 model implementation of the present example and the NetMHCpan-4.1 predictor (see Reynisson, Birkir, et al. (2020)), evaluated on a large in-house eluted-ligand prediction benchmark. The T5 model substantially outperforms NetMHCpan-4. 1 for both SA and MA test datasets. Without wishing to be bound to any particular theory, this performance improvement is believed to largely be a result of the T5 model being a general-purpose model that is trained across many related tasks and datasets.

[0525] Integrating a multi-allelic (MA) dataset using the T5 end-to-end model without explicit deconvolution was also shown to yield performance improvements. This was demonstrated by training the T5 model on the same data as NetMHCpan4.1 model demonstrates and observing improved performance, as shown in FIG. 67.

[0526] To assess the benefit of transfer learning on low-data tasks, the T5 model was evaluated on transporters associated with antigen processing (TAP) tasks and its performance compared with the DeepTAP model (described in Zhang, Xue, et al. (2023)) as a reference. Results are shown in FIG. 68A. As partitioning methods are different for the two models, a conservative test set was chosen where the T5 model has never seen any of the test data, while the DeepTAP model has seen some of it. Stability prediction performance of the T5 model was also evaluated and compared with that ofNetMHCStabPan-1.0 model as shown in FIG. 68B. In both cases, the T5 model outperforms the task-specific models (e.g., DeepTAP and NetMHCStabPan-1.0) even in unfavorable setting, highlighting an impact of transfer learning forthose tasks.

[0527] Finally, the impact of including low-volume tasks in an already large pretraining mix was explored. To evaluate these effects, the T5 model was pre-trained on three different mixtures and evaluated on an in-house immune response benchmarks.• Mixture 1 : Data from the immune response tasks only;• Mixture 2: Pretraining mix containing 99.84% of the total data; and• Mixture 3 : Including the last two tasks to reach 100% of the data.

[0528] FIG. 69 shows that each task benefited from the transfer derived learning from the pretraining mix, with Mixture 2 outperforming Mixture 1 on all benchmarks. The last two tasks, despite representing only 0.16% of the total data, still significantly impact the performance. These results demonstrate that the T5 model may leverage data from all tasks, even those with very small volumes.REFERENCES

[0529] B. Alvarez, B. Reynisson, C. Barra, S. Buus, N. Temette, T. Connelley, M. Andreatta, and M. Nielsen. Nnalign_ma; mhc peptidome deconvolution for accurate mhc binding motif characterization and improved t-cell epitope predictions. Molecular & Cellular Proteomics, 18 (12):2459-2477, 2019.

[0530] R. Apweiler, A. Bairoch, C. H. Wu, W. C. Barker, B. Boeckmann, S. Ferro, E. Gasteiger, H. Huang, R. Lopez, M. Magrane, et al. Uniprot: the universal protein knowledgebase. Nucleic acids research, 32(suppl_l):D115-D119, 2004.

[0531] V. Aribandi, Y. Tay, T. Schuster, J. Rao, H. S. Zheng, S. V. Mehta, H. Zhuang, V. Q. Tran, D. Bahri, J. Ni, et al. Ext5: Towards extreme multi-task scaling for transfer learning. arXiv preprint.

[0532] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan,P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei. Language models are few-shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 1877-1901. Curran Associates, Inc., 2020.

[0533] T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei. Language models are few-shot learners, 2020.

[0534] B. Bulik-Sullivan, J. Busby, C. D. Palmer, M. J. Davis, T. Murphy, A. Clark, M. Busby, F. Duke, A. Yang, L. Young, et al. Deep learning using tumor hla peptide massspectrometry datasets improves neoantigen identification. Nature biotechnology, 37( 1): 55— 63, 2019.

[0535] R. Caruana. Multitask Learning, pages 95-133. Springer US, Boston, MA, 1998. ISBN 978-1-4615-5529-2.

[0536] B. Chen, X. Cheng, Y.-a. Geng, S. Li, X. Zeng, B. Wang, J. Gong, C. Liu, A. Zeng, Y. Dong, J. Tang, and L. Song, xtrimopghn: Unified lOOb-scale pre-trained transformer for deciphering the language of protein, 2023.

[0537] X. Chen, X. Wang, S. Changpinyo, A. Piergiovanni, P. Padlewski, D. Saiz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer, et al. Pali: A jointly-scaled multilingual language-image model. arXiv preprint, 2022.

[0538] A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Levskaya, S. Ghemawat, S. Dev, H. Michalewski, X. Garcia, V. Misra, K. Robinson, L. Fedus, D. Zhou, D. Ippolito, D. Luan, H. Lim, B. Zoph, A. Spiridonov, R. Sepassi, D. Dohan, S. Agrawal, M. Omemick, A. M. Dai, T. S. Pillai, M. Pellat, A. Lewkowycz, E. Moreira, R. Child, O. Polozov, K. Lee, Z. Zhou, X. Wang, B. Saeta, M. Diaz, O. Firat, M. Catasta, J. Wei, K. Meier-Hellstem, D. Eck, J. Dean, S. Petrov, and N. Fiedel. Palm: Scaling language modeling with pathways. arXiv:2204.02311, 2022.

[0539] A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240): 1-113, 2023.

[0540] H. Dalia-Torre, L. Gonzalez, J. Mendoza-Revilla, N. L. Carranza, A. H. Grzywaczewski, F. Oteri, C. Dallago, E. Trop, B. P. de Almeida, H. Sirelkhatim, et al. The nucleotide transformer: Building and evaluating robust foundation models for human genomics. bioRxiv, pages 2023-01, 2023.

[0541] D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al. Palm-e: An embodied multimodal language model. arXiv:2303.03378, 2023.

[0542] A. Elnaggar, M. Heinzinger, C. Dallago, G. Rehawi, Y. Wang, L. Jones, T. Gibbs, T. Feher, C. Angerer, M. Steinegger, D. Bhowmik, and B. Rost. Prottrans: Toward understanding the language of life through self-supervised learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44( 10):7112-7127, 2022. doi: 10.1109 / TP AMI.2021.3095381.

[0543] A. Elnaggar, H. Essam, W. Salah-Eldin, W. Moustafa, M. Elkerdawy, C. Rochereau, and B. Rost. Ankh: Optimized protein language model unlocks general -purpose modelling. bioRxiv, pages 2023-01, 2023.

[0544] R. D. Finn, A. Bateman, J. Clements, P. Coggill, R. Y. Eberhardt, S. R. Eddy, A. Heger, K. Hetherington, L. Holm, J. Mistry, E. L. L. Sonnhammer, J. Tate, and M. Punta. Pfam: the protein families database. Nucleic Acids Research, 42(Dl):D222-D230, 11 2013. ISSN 0305-1048.

[0545] A. Gane, M. Bileschi, D. Dohan, E. Speretta, A. Helion, L. Meng- Papaxanthos, H. Zellner, E. Brevdo, A. Parikh, M. Martin, et al. Protnlm: Model-based natural language protein annotation. Preprint, 2022.

[0546] S. Golkar, M. Pettee, M. Eickenberg, A. Bietti, M. Cranmer, G. Krawezik, F. Lanusse, M. Me- Cabe, R. Ghana, L. Parker, B. R.-S. Blancard, T. Tesileanu, K. Cho, and S. Ho. xval: Acontinuous number e...

Claims

What is claimed is:

1. A method for predicting an immune response to a prospective antigen or portion thereof using a language model, the method comprising:(a) obtaining, by a processor of a computing device, an input text string comprising a candidate peptide sequence representing a biological sequence of at least a portion of the prospective antigen;(b) determining, by the processor, using the language model, an immune response prediction based on the candidate peptide sequence, wherein the language model receives, as input, the input text string and generates, as output, the immune response prediction as an output text string encoding values of one or more immune response scores, each immune response score associated with and representing a predicted result of a particular immune response task; and(c) storing and / or providing, by the processor, the immune response prediction for display and / or further processing.

2. The method of claim 1, wherein the prospective antigen is a neoantigen of a subject.

3. The method of claim 1, wherein the prospective antigen is a shared tumor antigen.

4. The method of any one of the preceding claims, wherein the candidate peptide sequence is an amino acid sequence encoded by a gene comprising one or more identified patient-specific tumor mutations.

5. The method of any one of the preceding claims, wherein the candidate peptide sequence corresponds to and / or comprises one or more neoantigen epitopes.

6. The method of any one of the preceding claims, wherein the candidate peptides sequence corresponds to and / or comprises one or more shared tumor antigen epitopes.

7. The method of any one of the preceding claims, comprising obtaining, by the processor, a tumor sequence data for a patient and selecting and / or generating the candidate peptide sequence based on the tumor sequence data.

8. The method of any one of the preceding claims, wherein the candidate peptide sequence is a ribonucleic acid RNA transcribed from a gene encoding the candidate peptide.

9. The method of any one of the preceding claims, wherein the candidate peptide sequence is a deoxyribonucleic acid (DNA) sequence of a gene encoding the candidate peptide.

10. The method of any one of the preceding claims, wherein the prospective antigen is or comprises at least a portion of a viral protein.

11. The method of any of the preceding claims, wherein the prospective antigen comprises at least a portion of a particular subunit or region of a viral protein.

12. The method of any of the preceding claims, wherein the at least a portion of the prospective antigen is selected iteratively using a sliding window algorithm.

13. The method of any one of the preceding claims, comprising: performing step (b) for each of a plurality of candidate peptides of an initial set, thereby determining a plurality of particular immune response predictions, one for each of the plurality of candidate peptides; and determining, by the processor, the immune response prediction based at least in part on the plurality of candidate immune response predictions.

14. The method of claim 13, comprising determining the plurality of candidate peptides of the initial set by selecting peptides of one or more particular desired lengths from a region of a source protein sequence about a particular identified mutation.

15. The method of any of the preceding claims, wherein the portion of the prospective antigen is selected based on known properties.

16. The method of any one of the preceding claims, wherein the language model is or comprises a transformer model.

17. The method of any one of the preceding claims, wherein the language model is a multi-task model, having been trained to determine values for at least two of different immune response score options, each associated with, and representing a predicted result of, a particular member of a set of immune response tasks, and the one or more immune response scores whose value(s) is / are predicted by the language model at step (b) having been selected from the at least two different immune response score options.

18. The method of claim 17, wherein the set of immune response tasks comprises one or both of (i) a peptide-MHC binding task and (ii) a peptide-MHC surface presentation task.

19. The method of claim 18, wherein the set of immune response tasks comprises both the peptide-MHC binding task and the peptide-MHC surface presentation task.

20. The method of any one of claims 17 to 19, wherein the set of immune response tasks comprises a proteasomal cleavage task.

21. The method of any one of claims 17 to 20, wherein the set of immune response tasks comprises a ligand location task.

22. The method of any one of claims 17 to 21, wherein the set of immune response tasks comprises a ligand generation task.

23. The method of any one of claims 17 to 21, wherein the set of immune response tasks comprises two or more of (A), (B), and (C) as follows:(A) a binary classification task;(B) a regression task; and(C) a segmentation task.

24. The method of any one of the preceding claims, wherein at least one of the immune response scores is an MHC binding affinity score representing a predicted binding affinity between the candidate peptide and a particular major histocompatibility complex (MHC) allele.

25. The method of claim 24, wherein the input text string comprises an identification of the particular MHC allele.

26. The method of claim 24 or claim 25, wherein the value of the MHC binding affinity score is a continuous number.

27. The method of any one of the preceding claims, wherein at least one of the immune response scores is an MHC binding affinity classification label and / or value indicative of whether or not the particular MHC allele is predicted to bind to the candidate peptide.

28. The method of any one of the preceding claims, wherein at least one of the immune response scores is a surface presentation score representing a prediction of whether the candidate peptide will be presented at a surface via a particular (MHC) allele.

29. The method of claim 28, wherein the machine learning model receives, as input, a gene bias value representing a background expression level of the candidate peptide.

30. The method of any one of the preceding claims, wherein at least one of the immune response scores is a cleavage score representing a prediction of whether the candidate peptide is a fragment predicted to result from proteasomal cleavage of the antigen.

31. The method of any one of the preceding claims, wherein the input text string comprises an identification of a plurality of distinct MHC alleles and the output text string encodes a plurality of distinct sets of allele-specific immune response score values, each corresponding to a particular one of the plurality of distinct MHC alleles.

32. The method of claim 31, comprising determining a set of overall immune response score values based on the plurality of distinct sets of allele-specific immune response score values.

33. The method of any one of the preceding claims, comprising using the immune response prediction to create a personalized cancer vaccine for a patient.

34. The method of any one of the preceding claims, comprising: performing steps (a) through (c) for each of a plurality of candidate peptides, each corresponding to a candidate epitope identified as associated with a tumor genome / exome of a patient, thereby determining a plurality of immune response predictions, one for each candidate neoepitope; and selecting, by the processor, a subset of the candidate epitope based at least in part on the determined plurality of immune response predictions.

35. The method of claim 34, comprising including one or more biological sequences encoding the selected subset of candidate epitopes in a cancer vaccine.

36. The method of any one of the preceding claims, comprising using the immune response prediction to create a personalized viral vaccine for a patient.

37. The method of any one of the preceding claims, comprising: performing steps (a) through (c) for each of a plurality of candidate peptides, each corresponding to a candidate viral variant, thereby determining a plurality of immune response predictions, one for each candidate viral variant; and selecting, by the processor, a subset of the candidate viral variants based at least in part on the determined plurality of immune response predictions.

38. A method for identifying prospective antigen fragments that will be presented as ligands for T-cell recognition, the method comprising:(a) obtaining, by a processor of a computing device, sequence data comprising (i) a prospective antigen sequence representing a biological sequence of at least a portion of a target antigen and (ii) a major histocompatibility complex (MHC) sequence representing a biological sequence of at least a portion of a particular MHC allele;(b) identifying, by the processor, using a machine learning model, a subregion of the antigen sequence, thereby locating, within the antigen, a fragment to be presented as a ligand for T-cell recognition; and(c) storing and / or providing, by the processor, a representation of the located ligand for display and / or further processing.

39. The method of claim 38, wherein the machine learning model:(i) receives the sequence data as input; and(ii) generates, as output, a mask or a set of indices identifying the subregion within the antigen sequence.

40. The method of claim 38 or 39, wherein the machine learning model receives (i) the antigen sequence as a first channel of input and (ii) the MHC sequence as a second, separate channel of input.

41. The method of any one of claims 38-40, wherein the machine learning model is a multi-task model, having been trained to determine values for at least two of different immune response score options, each associated with, and representing a predicted result of, a particular member of a set of immune response tasks.

42. The method of claim 41, wherein the set of immune response tasks comprises a peptide-MHC binding task and a peptide-MHC surface presentation task.

43. The method of any one of claims 38-42, comprising using the representation of the located ligand to create a personalized cancer vaccine for a patient.

44. A method for identifying antigen fragments that will be presented as ligands for T-cell recognition, the method comprising:(a) obtaining, by a processor of a computing device, sequence data comprising (i) an antigen sequence representing a biological sequence of at least a portion of a target antigen and (ii) a major histocompatibility complex (MHC) sequence representing a biological sequence of at least a portion of a particular MHC allele;(b) determining, by the processor, using a machine learning model, a ligand sequence corresponding to a subregion of the antigen sequence identified as a fragment of the antigen predicted to be presented as a ligand for T-cell recognition; and(c) storing and / or providing, by the processor, the ligand sequence for display and / or further processing.

45. The method of claim 44, wherein the machine learning model:(i) receives the sequence data as input; and(ii) generates, as output, one or more strings representing the ligand sequence.

46. The method of claim 44 or 45, wherein the machine learning model receives (i) the antigen sequence as a first channel of input and (ii) the MHC sequence as a second, separate channel of input.

47. The method of any one of claims 44-46, wherein the machine learning model is a multi-task model, having been trained to determine values for at least two of different immune response score options, each associated with, and representing a predicted result of, a particular member of a set of immune response tasks.

48. The method of claim 47, wherein the set of immune response tasks comprises a peptide-MHC binding task and a peptide-MHC surface presentation task.

49. The method of any one of claims 44-48, comprising using the ligand sequence to create a personalized cancer vaccine for a patient.

50. A method for predicting an immune response to a prospective antigen or portion thereof, the method comprising:(a) obtaining, by a processor of a computing device, an input text string comprising a candidate peptide sequence representing a biological sequence of at least a portion of the prospective antigen;(b) determining, by the processor, a value of a selected immune response score, based on the candidate peptide sequence, using a machine learning model, wherein the machine learning model has been trained to determine values for a plurality of different immune response score options, each associated with, and representing a predicted result of, a particular member of a set of immune response tasks, and wherein the selected immune response score is one of the plurality of different immune response score options; and(c) storing and / or providing, by the processor, the value of the selected immune response score for display and / or further processing.

51. The method of claim 50, comprising: repeating steps (a) - (c) for a plurality of candidate peptides, wherein each of the plurality of candidate peptides represents a biological sequence of at least a portion of the prospective antigen, thereby determining a plurality of values of the selected immune response score, one for each member of the plurality of candidate peptides; and selecting, by the processor, a subset of the plurality of candidate peptides for inclusion in composition.

52. The method of claim 50 or 51, wherein the set of immune response tasks comprises a peptide-MHC binding task and a peptide-MHC surface presentation task.

53. The method of any one of claims 50-52, wherein the selected immune response score comprises an MHC binding affinity score representing a predicted binding affinity between the candidate peptide and a particular major histocompatibility complex (MHC) allele.

54. The method of claim 53, wherein the input text string comprises an identification of the particular MHC allele.

55. The method of any one of claims 50-54, wherein the selected immune response score comprises a surface presentation score representing a prediction of whether the candidate peptide will be presented at a surface via a particular (MHC) allele.

56. The method of any one of claims 50-55, comprising using the value of the selected immune response score to create a personalized cancer vaccine for a patient.

57. A method of training a machine learning model to predict an immune response to a prospective antigen or portion thereof, the method comprising:(a) obtaining, by a processor of a computing device, a plurality of datasets, each associated with a distinct particular immune response task and comprising a plurality ofexamples, each example associated with and comprising sequence data representing a particular example peptide;(b) for each particular one of the plurality of datasets, repeatedly selecting, by the processor, examples from the particular dataset and using, by the processor, the selected examples to train the machine learning model to determine values of immune response scores, each immune response score associated with and representing a predicted result of the immune response task with which the particular dataset is associated, thereby generating a multi-task machine learning model trained to generate predictions for the at least two different immune response tasks; and(c) storing and / or providing, by the processor, the trained multi-task machine learning model for further processing.

58. The method of claim 57, wherein step (b) comprises: obtaining, by the processor, an input text string comprising the sequence data; and determining, by the processor, using the machine learning model, the values of immune response scores, wherein the machine learning model receives, as input, the input text string and generates, as output, the values of immune response scores as an output text string.

59. The method of claim 57 or 58, wherein the at least two different immune response tasks comprise a peptide-MHC binding task and a peptide-MHC surface presentation task.

60. The method of any one of claim 57-59, wherein the immune response scores comprise an MHC binding affinity score representing a predicted binding affinity between the example peptide and a particular major histocompatibility complex (MHC) allele.

61. The method of claim 60, wherein the input text string comprises an identification of the particular MHC allele.

62. The method of any one of claims 57-61, wherein the immune response scores comprise a surface presentation score representing a prediction of whether the example peptide will be presented at a surface via a particular (MHC) allele.

63. A method for predicting an immune response to a prospective antigen or portion thereof using a language model, the method comprising:(a) obtaining, by a processor of a computing device, an input text string comprising a candidate peptide sequence, wherein the candidate peptide sequence comprises a representation of a biological sequence of at least a portion of the prospective antigen;(b) determining, by the processor, an immune response prediction based on the candidate peptide sequence using the language model, wherein the language model receives, as input, the input text string and generates, as output, the immune response prediction as an output text string encoding values of one or more immune response scores, each immune response score associated with and representing a predicted result of a particular immune response task; and(c) storing and / or providing, by the processor, the immune response prediction for display and / or further processing.

64. The method of claim 63, wherein the language model comprises a transformer model.

65. The method of claim 63 or 64, wherein the language model is a multi-task model, having been trained to determine values for at least two of different immune response score options, each associated with, and representing a predicted result of, a particular member of a set of immune response tasks, and the one or more immune response scores whose value(s) is / are predicted by the language model at step (b) having been selected from the at least two different immune response score options.

66. The method of claim 65, wherein the at least two different immune response tasks comprise a peptide-MHC binding task and a peptide-MHC surface presentation task.

67. The method of any one of claim 63-66, wherein the one or more immune response scores comprise an MHC binding affinity score representing a predicted binding affinity between the candidate peptide and a particular major histocompatibility complex (MHC) allele.

68. The method of claim 67, wherein the input text string comprises an identification of the particular MHC allele.

69. The method of any one of claims 63-68, wherein the one or more immune response scores comprise a surface presentation score representing a prediction of whether the candidate peptide will be presented at a surface via a particular (MHC) allele.

70. The method of any one of claims 63-69, comprising using the immune response prediction to create a personalized cancer vaccine for a patient.

71. A system for predicting an immune response to a prospective antigen or portion thereof using a language model, the system comprising: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to:(a) obtain an input text string comprising a candidate peptide sequence representing a biological sequence of at least a portion of the prospective antigen;(b) determine, using the language model, an immune response prediction based on the candidate peptide sequence, wherein the language model receives, as input, the input text string and generates, as output, the immune response prediction as an output text string encoding values of one or more immune response scores, each immune response score associated with and representing a predicted result of a particular immune response task; and(c) store and / or provide the immune response prediction for display and / or further processing.

72. A system for identifying prospective antigen fragments that will be presented as ligands for T-cell recognition, the system comprising: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to:(a) obtain sequence data comprising (i) a prospective antigen sequence representing a biological sequence of at least a portion of a target antigen and (ii) a major histocompatibility complex (MHC) sequence representing a biological sequence of at least a portion of a particular MHC allele;(b) identify, using a machine learning model, a subregion of the prospective antigen sequence, thereby locating, within the target antigen, a fragment to be presented as a ligand for T-cell recognition; and(c) store and / or provide a representation of the located ligand for display and / or further processing.

73. A system for identifying antigen fragments that will be presented as ligands for T-cell recognition, the system comprising: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to:(a) obtain sequence data comprising (i) a prospective antigen sequence representing a biological sequence of at least a portion of a target antigen and (ii) a major histocompatibility complex (MHC) sequence representing a biological sequence of at least a portion of a particular MHC allele;(b) determine, using a machine learning model, a ligand sequence corresponding to a subregion of the antigen sequence identified as a fragment of the prospective antigen to be presented as a ligand for T-cell recognition; and(c) store and / or provide the ligand sequence for display and / or further processing.

74. A system for predicting an immune response to a prospective antigen or portion thereof, the system comprising: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to:(a) obtain an input text string comprising a candidate peptide sequence representing a biological sequence of at least a portion of the prospective antigen;(b) determine a value of a selected immune response score based on the candidate peptide sequence using a machine learning model, wherein the machine learning model has been trained to determine values for a plurality of different immune response score options, each associated with, and representing a predicted result of, a particular member of a set of immune response tasks, and wherein the selected immune response score is one of the plurality of different immune response score options; and(c) store and / or provide the value of the selected immune response score for display and / or further processing.

75. A system for training a machine learning model to predicting an immune response to a prospective antigen or portion thereof, the system comprising: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to:(a) obtain a plurality of datasets, each associated with a distinct particular immune response task and comprising a plurality of examples, each example associated with and comprising sequence data representing a particular example peptide;(b) for each particular one of the plurality of datasets, repeatedly select examples from the particular dataset and use the selected examples to train the machine learning model to determine values of immune response scores, each immune response score associated with and representing a predicted result of theimmune response task with which the particular dataset is associated, thereby generating a multi-task machine learning model trained to generate predictions for the at least two different immune response tasks; and(c) store and / or provide the trained multi-task machine learning model for further processing.

76. A system for predicting an immune response to a prospective antigen or portion thereof using a language model, the system comprising: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to:(a) obtain an input text string comprising a candidate peptide sequence, wherein the candidate peptide sequence comprises a representation of a biological sequence of at least a portion of the prospective antigen;(b) determine an immune response prediction based on the candidate peptide sequence using the language model, wherein the language model receives, as input, the input text string and generates, as output, the immune response prediction as an output text string encoding values of one or more immune response scores, each immune response score associated with and representing a predicted result of a particular immune response task; and(c) store and / or provide the immune response prediction for display and / or further processing.

Citation Information

Patent Citations

  • Individualized vaccines for cancer

    US10738355B2

  • Technologies for early detection of variants of interest

    WO2022235847A1

  • Immunogen selection

    WO2022235853A1

Cited By

  • B cell immune repertoire antibody sequence characteristic analysis and specificity screening method and related equipment

    CN121789769A