Antigenic prediction of epitopes derived from infectious diseases

JP2025512989A5Pending Publication Date: 2026-04-07GRITSTONE BIO INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-04-07
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

The lack of effective models for predicting the probability of presentation of viral pellets in HLA alleles has led to increased difficulty in developing targeted therapeutic vaccines.

Method used

A multi-part presentation model is used to combine the pan-allele model and specific allele model to generate personalized granule presentation probability, and the training and prediction capabilities of the model are optimized by integrating mass spectrometry data and binding affinity data.

Benefits of technology

Improve the accuracy of prediction of viral particle presentation, help develop more effective personalized therapeutic vaccines, and enhance the antiviral immune response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Disclosed herein are systems and methods for determining allele, antigen, and infectious disease-based vaccine compositions that are determined based on the HLA alleles expressed in a patient. Specific infectious disease-derived vaccines are additionally described herein.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 329,259, filed April 8, 2022, the entire disclosure of which is incorporated herein by reference in its entirety for all purposes. [Background technology]

[0002] Therapeutic vaccines for viral infections hold great promise for personalized treatment. Accurate prediction of viral epitopes likely to be presented by HLA alleles may be useful for developing therapeutic vaccines. Therapeutic vaccines containing predicted viral epitopes may be more effective in eliciting antiviral immune responses. However, unlike previous efforts in cancer, where large mass spectrometry (MS) human immunopeptidomics datasets exist, such MS immunopeptidomics datasets are lacking for infectious disease-derived antigens. This represents a significant limitation in building models to predict likely presentation of viral epitopes. Thus, new models that can effectively leverage other types of data (e.g., in addition to MS immunopeptidomics datasets) are needed. Summary of the Invention

[0003] Disclosed herein is an optimized approach for identifying and selecting infectious disease-derived antigens for personalized infectious disease (ID) vaccines. In general, this approach involves applying a multi-part presentation model to generate a presentation likelihood for each allele that represents whether an individual HLA allele (e.g., an HLA allele expressed by a patient) is likely to present an epitope (e.g., an infectious disease-derived epitope). In various embodiments, the multi-part presentation model includes a first part that includes a pan-allele model portion (also referred to as a pan-specific model portion) and a second part that includes one or more allele-specific model portions (also referred to herein as per-allele models). In general, the pan-allele model portion allows for the sharing of information across similar alleles, thereby allowing the pan-allele model portion to generate presentation likelihoods across various HLA alleles, including HLA alleles that the pan-allele model portion has not previously encountered. Each of the allele-specific model portions represents a network model that generates a presentation likelihood for a particular HLA allele. Thus, both the pan-allelic model portion and the allele-specific model portion generate per-allele presentation likelihoods. By combining the outputs of the pan-allelic model portion and the allele-specific model portion, the multi-part presentation model outputs improved per-allele presentation likelihoods for individual peptide sequences (e.g., infectious disease-derived peptide sequences).

[0004] In various embodiments, the multipart presentation model is trained using training data derived from one of: 1) binding affinity data between training peptide sequences and HLA alleles; and 2) peptide data eluted from mass spectrometry representing the presentation of training peptide sequences and HLA alleles. In certain embodiments, the multipart presentation model is trained using training data derived from both: 1) binding affinity data between training peptide sequences and HLA alleles; and 2) peptide data eluted from mass spectrometry representing the presentation of training peptide sequences and HLA alleles. Here, it may be preferable to use both types of data to train the multipart presentation model, especially in situations where the amount of eluted peptide data of infectious disease-derived peptide sequences generated by mass spectrometry is limited.

[0005] Disclosed herein is a method for identifying one or more infectious disease derived antigens likely to be presented by cells of a subject, the method comprising: obtaining peptide sequences of a plurality of infectious disease derived antigens; obtaining sequences of one or more MHC alleles of the subject; inputting the peptide sequences of the plurality of infectious disease derived antigens and sequences of the one or more MHC alleles of the subject into a multipart presentation model to generate a set of numerical likelihoods that the plurality of infectious disease derived antigens will be presented by one or more MHC alleles expressed on the surface of cells of the subject, wherein a first part of the multipart presentation model includes a pan-allelic model portion that receives as input the peptide sequences of the one or more infectious disease derived antigens and sequences of the one or more MHC alleles of the subject or representations thereof, and a second part of the multipart presentation model includes a plurality of allele-specific models that each receive as input the peptide sequences of the plurality of infectious disease derived antigens or representations thereof; and selecting a subset of the plurality of infectious disease derived antigens based on the set of numerical likelihoods to generate a set of selected antigens.

[0006] In various embodiments, the multipart presentation model comprises a plurality of parameters generated using at least 1) mass spectrometry data and 2) binding affinity data determined from a plurality of samples. In various embodiments, the multipart presentation model comprises a plurality of parameters generated using a training dataset comprising training peptide sequences and, for one or more of the training peptide sequences, a label derived from mass spectrometry data indicating whether the training peptide sequence was presented by one or more class I MHC alleles present in the plurality of samples. In various embodiments, the training peptide sequences are identified by mass spectrometry on isolated peptides eluted from MHC alleles present in the plurality of samples. In various embodiments, the multipart presentation model comprises a plurality of parameters generated using a training dataset comprising a label derived from binding affinity data for one or more of the training peptide sequences indicating whether the training peptide sequence bound to one or more class I MHC alleles present in the plurality of samples.

[0007] In various embodiments, the training peptide sequences are in a k-mer range in length, where k is between 8 and 15, inclusive. In various embodiments, the training peptide sequences are in a k-mer range in length, where k is between 8 and 11, inclusive. In various embodiments, the training peptide sequences are training peptide sequences from an infectious disease.

[0008] In various embodiments, the pan-allele model portion includes a neural network. In various embodiments, a first layer set of the neural network of the pan-allele model portion performs dimensionality reduction of the sequences of one or more MHC alleles of the subject. In various embodiments, a second layer set of the neural network of the pan-allele model portion receives as inputs a representation of peptide sequences of the plurality of infectious disease-derived antigens and a dimensionality-reduced representation of the sequences of one or more MHC alleles of the subject. In various embodiments, the representation of peptide sequences of the plurality of infectious disease-derived antigens is generated by encoding the peptide sequences via a one-hot encoding scheme. In various embodiments, the second layer set of the neural network models interactions between the peptide sequences of the plurality of infectious disease-derived antigens and the sequences of one or more MHC alleles of the subject.

[0009] In various embodiments, one or more of the allele-specific models comprises a neural network. In various embodiments, the neural network of the allele-specific network receives as input a representation of peptide sequences of a plurality of infectious disease-derived antigens and outputs, for the alleles, a presentation likelihood for each allele. In various embodiments, the representation of peptide sequences of a plurality of infectious disease-derived antigens is generated by encoding the peptide sequences via a one-hot encoding scheme. In various embodiments, each of the allele-specific models comprises a neural network. In various embodiments, the second part of the multipart presentation model comprises 10 or more allele-specific models.

[0010] In various embodiments, the numerical likelihood of an antigen is a combination of the output of the pan-allelic model portion and the output of the multiple allele-specific models. In various embodiments, the set of numerical likelihoods is further specified by features including at least one of: (a) C-terminal sequences adjacent to the peptide sequences of the multiple infectious disease-derived antigens, and (b) N-terminal sequences adjacent to the peptide sequences of the multiple infectious disease-derived antigens.

[0011] In various embodiments, the plurality of samples comprises at least one of: (a) one or more cell lines engineered to express a single MHC class I allele; (b) one or more cell lines engineered to express multiple MHC class I alleles; (c) one or more human cell lines obtained or derived from multiple patients; (d) fresh or frozen samples obtained from multiple patients; and (e) fresh or frozen tissue samples obtained from multiple patients. In various embodiments, the subject's cells comprise cells infected with one of a pathogen, a virus, a bacterium, a fungus, or a parasite. In various embodiments, the infectious disease-derived antigen originates from one of a pathogen, a virus, a bacterium, a fungus, or a parasite. In various embodiments, the infectious disease derived antigen originates from an infectious disease organism selected from the group consisting of severe acute respiratory syndrome-associated coronavirus (SARS), severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), Ebola, HIV, hepatitis B virus (HBV), influenza, hepatitis C virus (HCV), human papillomavirus (HPV), cytomegalovirus (CMV), chikungunya virus, respiratory syncytial virus (RSV), dengue virus, orthomyxoviridae virus, tuberculosis, pancorona, herpes simplex virus infection (HSV), influenza, metapneumovirus (MPV), and parainfluenza virus (PIV).

[0012] Additionally disclosed herein is a method of treating a subject for an infectious disease comprising performing any of the methods disclosed herein and further comprising obtaining a vaccine comprising a set of selected antigens and administering the vaccine to the subject. In various embodiments, the vaccine is administered prophylactically to the subject. In various embodiments, the vaccine is administered therapeutically to the subject. Additionally disclosed herein is a method of manufacturing a vaccine comprising performing any of the methods disclosed herein and further comprising producing or having produced a vaccine comprising a set of selected antigens. Additionally disclosed herein is a vaccine comprising a set of selected antigens selected by performing any one of the methods disclosed herein. Additionally disclosed herein is a method of treating a subject for an infectious disease comprising performing any one of the methods disclosed herein and further comprising obtaining an isolated antigen binding protein exhibiting binding specificity for one or more of the selected antigens and administering the isolated antigen binding protein to the subject. Additionally disclosed herein is a method of treating a subject for an infectious disease, comprising carrying out any of the methods disclosed herein and further comprising obtaining an isolated antigen binding protein exhibiting binding specificity for one or more selected antigens, and administering the isolated antigen binding protein to the subject. In various embodiments, the antigen binding protein is an antibody or an antigen binding fragment. In various embodiments, the antigen binding protein is a T cell receptor or a chimeric antigen receptor.

[0013] Additionally disclosed herein is a non-transitory computer readable medium comprising instructions that, when executed by a processor, cause a processor to: obtain peptide sequences of a plurality of infectious disease derived antigens; obtain sequences of one or more MHC alleles of a subject; input the peptide sequences of the plurality of infectious disease derived antigens and sequences of the one or more MHC alleles of the subject into a multipart presentation model to generate a set of numerical likelihoods that the plurality of infectious disease derived antigens will be presented by one or more MHC alleles expressed on the surface of a cell of the subject, wherein a first part of the multipart presentation model includes a pan-allele model portion that receives as input the peptide sequences of the one or more infectious disease derived antigens and sequences of the one or more MHC alleles of the subject or representations thereof, and a second part of the multipart presentation model includes a plurality of allele-specific models that each receive as input the peptide sequences of the plurality of infectious disease derived antigens or representations thereof; and select a subset of the plurality of infectious disease derived antigens based on the set of numerical likelihoods to generate a set of selected antigens. In various embodiments, the multipart presentation model comprises a plurality of parameters generated using at least 1) mass spectrometry data and 2) binding affinity data determined from a plurality of samples. In various embodiments, the multipart presentation model comprises a plurality of parameters generated using a training dataset comprising training peptide sequences and, for one or more of the training peptide sequences, a label derived from the mass spectrometry data indicating whether the training peptide sequence was presented by one or more class I MHC alleles present in the plurality of samples. In various embodiments, the training peptide sequences are identified by mass spectrometry on isolated peptides eluted from MHC alleles present in the plurality of samples.

[0014] In various embodiments, the multipart presentation model comprises a plurality of parameters generated using a training dataset that includes labels derived from binding affinity data for one or more of the training peptide sequences that indicate whether the training peptide sequence bound to one or more class I MHC alleles present in the plurality of samples. In various embodiments, the training peptide sequences are in a k-mer range in length, where k is between 8 and 15, inclusive. In various embodiments, the training peptide sequences are in a k-mer range in length, where k is between 8 and 11, inclusive.

[0015] In various embodiments, the training peptide sequences are training peptide sequences from an infectious disease. In various embodiments, the pan-allele model portion comprises a neural network. In various embodiments, a first layer set of the neural network of the pan-allele model portion performs dimensionality reduction of sequences of one or more MHC alleles of the subject. In various embodiments, a second layer set of the neural network of the pan-allele model portion receives as inputs a representation of peptide sequences of antigens from a plurality of infectious diseases and a dimensionality reduced representation of sequences of one or more MHC alleles of the subject. In various embodiments, the representation of peptide sequences of antigens from a plurality of infectious diseases is generated by encoding the peptide sequences via a one-hot encoding scheme. In various embodiments, the second layer set of the neural network models interactions between peptide sequences of antigens from a plurality of infectious diseases and sequences of one or more MHC alleles of the subject. In various embodiments, one or more of the allele-specific models comprise a neural network. In various embodiments, the neural network of the allele-specific network receives as inputs representations of peptide sequences of antigens from multiple infectious diseases and outputs, for the alleles, the likelihood of presentation for each allele.

[0016] In various embodiments, the representations of the peptide sequences of the multiple infectious disease-derived antigens are generated by encoding the peptide sequences via a one-hot encoding scheme. In various embodiments, each of the allele-specific models comprises a neural network. In various embodiments, the second part of the multipart presentation model comprises 10 or more allele-specific models. In various embodiments, the numerical likelihood of the antigen is a combination of the output of the pan-allelic model portion and the output of the multiple allele-specific models. In various embodiments, the set of numerical likelihoods is further specified by features including at least one of: (a) a C-terminal sequence adjacent to the peptide sequences of the multiple infectious disease-derived antigens; and (b) an N-terminal sequence adjacent to the peptide sequences of the multiple infectious disease-derived antigens. In various embodiments, the plurality of samples includes at least one of: (a) one or more cell lines engineered to express a single MHC class I allele; (b) one or more cell lines engineered to express multiple MHC class I alleles; (c) one or more human cell lines obtained or derived from multiple patients; (d) fresh or frozen samples obtained from multiple patients; and (e) fresh or frozen tissue samples obtained from multiple patients.

[0017] In various embodiments, the cells of the subject include cells infected with one of a pathogen, a virus, a bacterium, a fungus, or a parasite. In various embodiments, the infectious disease derived antigen originates from one of a pathogen, a virus, a bacterium, a fungus, or a parasite. In various embodiments, the infectious disease derived antigen originates from an infectious disease organism selected from the group consisting of severe acute respiratory syndrome-related coronavirus (SARS), severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), Ebola, HIV, hepatitis B virus (HBV), influenza, hepatitis C virus (HCV), human papillomavirus (HPV), cytomegalovirus (CMV), chikungunya virus, respiratory syncytial virus (RSV), dengue virus, orthomyxoviridae virus, tuberculosis, pancorona, herpes simplex virus infection (HSV), influenza (flu), metapneumovirus (MPV), and parainfluenza virus (PIV). [Brief description of the drawings]

[0018] These and other features, aspects, and advantages of the present invention will become better understood with regard to the following description and accompanying drawings. [Figure 1A] 1 is an overview of an environment for identifying the likelihood of peptide presentation in a patient, according to an embodiment. [Figure 1B] According to an embodiment, an exemplary cassette design methodology is described. [Figure 1C] According to an embodiment, an exemplary cassette design methodology is described. [Diagram 2] 2A and 2B illustrate a method for obtaining presentation information according to one embodiment. [Figure 3A] FIG. 2 is a high-level block diagram illustrating computer logic components of a presentation specific system according to one embodiment. [Figure 3B] 1 illustrates an exemplary set of training data according to one embodiment. [Figure 4A]1 illustrates a flow process for implementing a multipart presentation model according to one embodiment. [Figure 4B] 1 illustrates a network architecture for a multipart presentation model according to one embodiment. [Figure 5A] 1 shows an implementation of an allele-specific model portion of a multi-part representation model, according to one embodiment. [Figure 5B] 1 shows an exemplary network model relating to MHC alleles. [Figure 5C] 1 shows an exemplary network model shared by MHC alleles. [Figure 5D] 1 shows the use of an exemplary network model to generate presentation likelihoods of peptides associated with MHC alleles. [Figure 5E] 1 shows the use of an exemplary network model to generate presentation likelihoods of peptides associated with MHC alleles. [Figure 5F] 1 shows the use of an exemplary network model to generate presentation likelihoods of peptides associated with MHC alleles. [Figure 5G] 1 shows the use of an exemplary network model to generate presentation likelihoods of peptides associated with MHC alleles. [Figure 5H] 1 shows the use of an exemplary network model to generate presentation likelihoods of peptides associated with MHC alleles. [Figure 5I] 1 shows the use of an exemplary network model to generate presentation likelihoods of peptides associated with MHC alleles. [Figure 6A] FIG. 1 shows an implementation of the pan-allele model portion of a multi-part presentation model, according to one embodiment. [Figure 6B] 1 illustrates an exemplary network model shared by MHC alleles, according to an embodiment. [Figure 6C] 1 shows an exemplary network model not linked to MHC alleles. [Figure 6D]1 shows an exemplary network model shared by MHC alleles can be used to generate presentation likelihoods for peptides associated with MHC alleles. [Figure 7] 1 illustrates an exemplary computer for implementing the entities illustrated in FIGS. 1A and 3A. [Figure 8A] 1 shows the performance of various models for predicting the presentation of HIV epitopes across the top five alleles. [Figure 8B] 1 shows the performance of various models for predicting presentation of influenza A epitopes across the top five alleles. [Figure 8C] Shows the performance of different models for predicting the presentation of SARS-CoV-2 epitopes across the top five alleles. [Figure 9A] Precision-recall curves of different models for predicting HIV epitope presentation across the top 25 alleles are shown. [Figure 9B] Precision-recall curves of different models for predicting influenza A epitope presentation across the top 25 alleles are shown. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0019] Detailed Description I. Definition In general, the terms used in the claims and this specification are intended to be interpreted as having the plain meaning understood by a person skilled in the art. To provide additional clarity, certain terms are defined below. In the event of a discrepancy between the plain meaning and the provided definition, the provided definition shall be used.

[0020] As used herein, the term "antigen" refers to a substance that induces an immune response. In various embodiments, "antigen" refers to an infectious disease-derived antigen that originates from one of the pathogens, viruses, bacteria, fungi, or parasites that can cause infectious disease. Antigens can include polypeptide sequences or nucleotide sequences. In various embodiments, antigens can include mutations such as frameshift or non-frameshift indels, missense or nonsense substitutions, splice site changes, splice variants, genomic rearrangements, or gene fusions.

[0021] As used herein, the term "antigen-based vaccine" is a vaccine construct that is based on one or more antigens, eg, multiple antigens.

[0022] As used herein, the term "coding region" is the portion or portions of a gene that encodes a protein.

[0023] As used herein, the term "coding mutation" is a mutation that occurs in a coding region.

[0024] As used herein, the term "indel" is an insertion or deletion of one or more nucleic acids.

[0025] As used herein, the term "epitope" is the particular portion of an antigen that is usually bound by an antibody or T-cell receptor.

[0026] As used herein, the term "immunogenic" is the ability to elicit an immune response, for example, via T cells, B cells, or both.

[0027] As used herein, the terms "HLA binding affinity," "MHC binding affinity," refer to the affinity of binding between a particular antigen and a particular MHC allele.

[0028] As used herein, the term "polymorphism" refers to a germline variant, i.e., a variant found in all DNA-carrying cells of an individual.

[0029] As used herein, the term "somatic variant" is a variant that occurs in the non-germline cells of an individual.

[0030] As used herein, the term "allele" is a version of a gene or a version of a gene sequence or a version of a protein.

[0031] As used herein, the term "HLA type" refers to the complement of HLA gene alleles.

[0032] As used herein, the term "nonsense-mediated decay" or "NMD" is the degradation of an mRNA by a cell due to a premature termination codon.

[0033] As used herein, the term "exome" is a subset of the genome that encodes proteins. The exome can be the comprehensive exons of the genome.

[0034] As used herein, the term "logistic regression" is a regression model for binary data from statistics in which the logit of the probability that the dependent variable equals 1 is modeled as a linear function of the dependent variable.

[0035] As used herein, the term "neural network" is a machine learning model for classification or regression that consists of multiple layers of linear transformations followed by element-wise nonlinearities typically trained via stochastic gradient descent and backpropagation.

[0036] As used herein, the term "proteome" is the set of all proteins expressed and / or translated by a cell, a population of cells, or an individual.

[0037] As used herein, the term "peptidome" refers to the set of all peptides presented by MHC-I or MHC-II on the cell surface. Peptidome may refer to the properties of a cell or a collection of cells.

[0038] As used herein, the term "ELISPOT" refers to enzyme-linked immunosorbent spot assay, a common method for monitoring immune responses in humans and animals.

[0039] As used herein, the term "tolerance or immune tolerance" is a state of immune non-responsiveness to one or more antigens, e.g., self-antigens.

[0040] As used herein, the term "central tolerance" is tolerance that is affected in the thymus, either by deleting autoreactive T cell clones or by promoting their differentiation into immunosuppressive regulatory T cells (Tregs).

[0041] As used herein, the term "peripheral tolerance" is tolerance that is affected in the periphery by downregulating or anergizing autoreactive T cells that survive central tolerance or promote these T cells to differentiate into Tregs.

[0042] The term "sample" may include a single cell or multiple cells or cell fragments or an aliquot of bodily fluid obtained from a subject by means including venipuncture, excretion, ejaculation, massage, biopsy, needle aspiration, lavage sample, scraping, surgical incision, or intervention, or other means known in the art.

[0043] The terms "subject" and "patient" are used interchangeably and include cells, tissues, or organisms, whether in vivo, ex vivo, or in vitro, male or female, human or non-human. The terms subject and patient include mammals, including humans.

[0044] The term "mammal" encompasses both humans and non-humans, and includes, but is not limited to, humans, non-human primates, canines, felines, murines, bovines, equines, and porcines.

[0045] The term "clinical factor" refers to a measure of a subject's condition, e.g., disease activity or severity. "Clinical factor" encompasses all markers of a subject's health status, including non-sample markers and / or other characteristics of the subject, such as, but not limited to, age and sex. A clinical factor can be a score, value, or set of values ​​that can be obtained from the evaluation of a subject or a sample (or a population of samples) from a subject under defined conditions. A clinical factor can also be predicted by other parameters, such as markers and / or surrogates of gene expression.

[0046] The term "antibody" herein is used in the broadest sense and includes polyclonal and monoclonal antibodies, including intact antibodies and functional (antigen-binding) antibody fragments, including fragment antigen-binding (Fab) fragments, F(ab')2 fragments, Fab' fragments, Fv fragments, recombinant IgG (rIgG) fragments, variable heavy chain (VH) regions capable of specific binding to antigen, single chain antibody fragments (including single chain variable fragments (scFv)), and single domain antibody (e.g., sdAb, sdFv, nanobodies, camelid VHH, engineered or evolved human VH such that pairing with a VL is not required for solubility or activity) fragments. The term encompasses genetically engineered and / or otherwise modified forms of immunoglobulins, such as intrabodies, peptibodies, chimeric antibodies, fully human antibodies, humanized antibodies, and heteroconjugate antibodies, multispecific, e.g., bispecific, antibodies, diabodies, triabodies, and tetrabodies, tandem di-scFvs, tandem tri-scFvs. Unless otherwise indicated, the term "antibody" should be understood to encompass functional antibody fragments thereof. The term also encompasses intact or full-length antibodies, including antibodies of any class or subclass, including IgG, and its subclasses IgM, IgE, IgA, and IgD.

[0047] Abbreviations: MHC: major histocompatibility complex; HLA: human leukocyte antigen, or human MHC locus; NGS: next generation sequencing; PPV: positive predictive value; FFPE: formalin-fixed, paraffin-embedded; NMD: nonsense-mediated decay; DC: dendritic cell.

[0048] Please note that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.

[0049] Any terms not directly defined herein should be understood to have the meanings generally associated with them as understood within the technical field of the present invention. Certain terms are discussed herein to provide additional guidance to the practitioner in describing the compositions, devices, methods, etc. of the present invention aspects and how to make or use them. It will be understood that the same thing may be referred to in more than one way. Thus, alternative language and synonyms may be used for any one or more of the terms discussed herein. No importance is placed on whether a term is detailed or discussed herein. Some synonyms or alternative methods, materials, etc. are provided. The listing of one or several synonyms or equivalents does not exclude the use of other synonyms or equivalents unless expressly stated. The use of examples, including examples of terms, is intended for illustrative purposes only and does not limit the scope and meaning of the present invention aspects herein.

[0050] All references, issued patents, and patent applications cited within the body of this specification are hereby incorporated by reference in their entirety for all purposes.

[0051] II. Methods for identifying antigens A method for identifying antigens (e.g., antigens derived from an infectious disease organism) includes identifying antigens from an infectious disease organism, an infection in a subject, or an infected cell of a subject that are likely to be presented on the cell surface of infected cells or immune cells, including professional antigen presenting cells such as dendritic cells, and / or likely to be immunogenic. As an example, one such method may include obtaining at least one of exome, transcriptome, or whole genome nucleotide sequencing and / or expression data of the infectious disease organism from an infected cell of the subject, using the nucleotide sequencing and / or expression data of the infectious disease organism to obtain data representing peptide sequences of each of a set of antigens (e.g., antigens derived from the infectious disease organism); inputting the peptide sequences of each antigen into one or more presentation models to generate a set of numerical likelihoods that each of the antigens will be presented by one or more MHC alleles on the cell surface of an infected cell of the subject or a cell present in the subject, the set of numerical likelihoods having been identified based at least on the received mass spectrometry data; and selecting a subset of the set of antigens based on the set of numerical likelihoods to generate a set of selected antigens.

[0052] A presentation model, such as a multipart presentation model, may include a statistical regression or machine learning (e.g., deep learning) model trained against a set of reference data (also referred to as a training dataset) that includes a corresponding set of labels, the set of reference data being optionally obtained from each of a plurality of different subjects, some of which may have an infectious disease, and the set of reference data includes at least one of data representing exome nucleotide sequences from infected tissue, data representing exome nucleotide sequences from normal tissue, data representing transcriptome nucleotide sequences from infected tissue, data representing proteome sequences from infected tissue, and data representing MHC peptidome sequences from infected tissue, and data representing MHC peptidome sequences from normal tissue. The reference data may further include mass spectrometry data, sequencing data, RNA sequencing data, expression profiling data, and proteomics data for synthetic proteins, normal human cell lines, and fresh and frozen primary samples, and monoallelic cell lines engineered to express a given MHC allele, which are then exposed to a T cell assay (e.g., ELISpot). In certain aspects, the set of reference data includes each form of reference data.

[0053] The proposed model may include a set of features derived at least in part from a set of reference data, the set of features including at least one of allele-dependent features and allele-independent features. In certain embodiments, each feature is included.

[0054] Methods for identifying shared antigens also include identifying one or more antigens from one or more cells of the subject that are likely to be presented on the surface of infected cells, thereby generating an output for constructing a personalized vaccine. As an example, one such method includes the steps of obtaining at least one of exome, transcriptome, or whole genome nucleotide sequencing and / or expression data from infected cells and normal cells of the subject, using the nucleotide sequencing and / or expression data to obtain data representing each peptide sequence of a set of identified antigens by comparing the nucleotide sequencing and / or expression data from the infected cells with the nucleotide sequencing and / or expression data from the normal cells and the peptide sequences identified from the normal cells of the subject; encoding each peptide sequence of the antigens into a corresponding numerical vector (each numerical vector includes information on a plurality of amino acids that make up the peptide sequence and a set of amino acid positions within the peptide sequence); and inputting the numerical vectors into a deep learning presentation model using a computer processor to generate a set of presentation likelihoods for the set of antigens (each presentation likelihood of the set indicates that the corresponding antigen is likely to be presented on the surface of infected cells of the subject in one or more class II or more classes II or more. The method may include: a deep learning presentation model, where the deep learning presentation model represents the likelihood of presentation by an MHC allele; selecting a subset of the set of antigens based on the set of presentation likelihoods to generate a set of selected antigens; and generating an output for constructing a personalized vaccine based on the set of selected antigens.

[0055] Specific methods for identifying antigens (e.g., antigens from infectious disease organisms) are known to those of skill in the art, such as those described in more detail in International Patent Application Publications WO / 2017 / 106638, WO / 2018 / 195357, and WO / 2018 / 208856, each of which is incorporated by reference in its entirety and for all purposes.

[0056] Disclosed herein is a method of treating a subject having an infectious disease comprising performing the steps of any of the antigen identification methods described herein and further comprising obtaining an infectious disease vaccine comprising the set of selected antigens and administering the infectious disease vaccine to the subject.

[0057] The methods disclosed herein may also include identifying one or more T cells that are antigen-specific for at least one of the antigens in the subset. In some embodiments, the identifying includes co-culturing one or more T cells with one or more of the antigens in the subset under conditions that expand the one or more antigen-specific T cells. In further embodiments, the identifying includes contacting one or more T cells with a tetramer that includes one or more of the antigens in the subset under conditions that allow binding between the T cells and the tetramer. In still further embodiments, the methods disclosed herein may also include identifying one or more T cell receptors (TCRs) of the one or more identified T cells. In certain embodiments, identifying the one or more T cell receptors includes sequencing the T cell receptor sequence of the one or more identified T cells. The methods disclosed herein may further include genetically engineering a plurality of T cells to express at least one of the one or more identified T cell receptors, culturing the plurality of T cells under conditions that expand the plurality of T cells, and injecting the expanded T cells into a subject. In some embodiments, engineering the plurality of T cells to express at least one of the one or more identified T cell receptors comprises cloning the T cell receptor sequence of the one or more identified T cells into an expression vector and transfecting each of the plurality of T cells with the expression vector. In some embodiments, the methods disclosed herein further comprise culturing the one or more identified T cells under conditions that expand the one or more identified T cells and injecting the expanded T cells into a subject.

[0058] Also disclosed herein are isolated T cells that are antigen-specific for at least one selected antigen in the subset.

[0059] Also disclosed herein is a method for producing an infectious disease vaccine comprising obtaining at least one of exome, transcriptome, or whole genome infectious disease organism nucleotide sequencing and / or expression data from infected cells of a subject, wherein the nucleotide sequencing and / or expression data of the infectious disease organism is used to obtain data representing peptide sequences for each of a set of antigens (e.g., where a peptide is derived from any polypeptide known or found to have altered expression in infected cells or tissues compared to normal cells or tissues); inputting the peptide sequences for each antigen into one or more presentation models to generate a set of numerical likelihoods that each of the antigens will be presented by one or more MHC alleles on the cell surface of infected cells of the subject (the set of numerical likelihoods having been identified based at least on the received mass spectrometry data); selecting a subset of the set of antigens based on the set of numerical likelihoods to generate a set of selected antigens; and producing, or has produced, an infectious disease vaccine comprising the set of selected antigens.

[0060] Also disclosed herein is an infectious disease vaccine comprising a set of selected antigens selected by performing a method comprising: obtaining at least one of exome, transcriptome, or whole genome nucleotide sequencing and / or expression data of the infectious disease organism from infected cells of the subject, wherein the nucleotide sequencing and / or expression data of the infectious disease organism is used to obtain data representing a peptide sequence for each of a set of antigens, the peptide sequence for each antigen (e.g., derived from any polypeptide known or found to have altered expression in infected cells or tissues compared to normal cells or tissues); inputting the peptide sequence for each antigen into one or more presentation models to generate a set of numerical likelihoods that each of the antigens will be presented by one or more MHC alleles on the cell surface of infected cells of the subject (the set of numerical likelihoods having been identified based at least on the received mass spectrometry data); selecting a subset of the set of antigens based on the set of numerical likelihoods to generate a set of selected antigens; and producing, or has produced, an infectious disease vaccine comprising the set of selected antigens.

[0061] The vaccine may include one or more of a nucleotide sequence, a polypeptide sequence, RNA, DNA, a cell, a plasmid, or a vector.

[0062] The vaccine may include one or more antigens that are displayed on the surface of an infected cell.

[0063] An infectious disease vaccine can include one or more antigens that are immunogenic in a subject.

[0064] An infectious disease vaccine may not include one or more antigens that induce an autoimmune response against normal tissues in a subject.

[0065] Infectious disease vaccines may include an adjuvant.

[0066] The infectious disease vaccine may include an excipient.

[0067] The methods disclosed herein may also include selecting antigens that have a higher likelihood of being presented on the infected cell surface compared to unselected antigens based on the presentation model.

[0068] The methods disclosed herein may also include selecting an antigen that has a high likelihood of being able to induce an infectious disease organism-specific immune response in a subject based on the presentation model, as compared to an unselected antigen.

[0069] The methods disclosed herein may also include selecting an antigen that has a higher likelihood of being presented to naive T cells by a professional antigen-presenting cell (APC) compared to an unselected antigen based on a presentation model, optionally the APC being a dendritic cell (DC).

[0070] The methods disclosed herein may also include selecting antigens that have a reduced likelihood of being subject to inhibition via central or peripheral tolerance compared to unselected antigens based on the presentation model.

[0071] The methods disclosed herein may also include selecting an antigen that has a reduced likelihood of being able to induce an autoimmune response against normal tissue in a subject, compared to an unselected antigen, based on the presentation model.

[0072] Exome or transcriptome nucleotide sequencing and / or expression data may be obtained by performing sequencing on infected tissue.

[0073] Sequencing may be next generation sequencing (NGS) or any massively parallel sequencing approach.

[0074] The set of numerical likelihoods may be further specified by at least one of the following characteristics of the MHC allele interaction: the predicted affinity with which the MHC allele and the antigen-encoded peptide bind; the predicted stability of the antigen-encoded peptide-MHC complex; the sequence and length of the antigen-encoded peptide; the probability of presentation of an antigen-encoded peptide with a similar sequence in cells from other individuals expressing the particular MHC allele as assessed by mass spectrometry proteomics or other means; the expression level of the particular MHC allele in the subject (e.g., as measured by RNA-seq or mass spectrometry); the overall antigen-encoded peptide sequence-independent probability of presentation by the particular MHC allele in other different subjects expressing the particular MHC allele; the overall antigen-encoded peptide sequence-independent probability of presentation by MHC alleles in the same molecular family (e.g., HLA-A, HLA-B, HLA-C, HLA-DQ, HLA-DR, HLA-DP) in other different subjects.

[0075] The set of numerical likelihoods is based on the C- and N-terminal sequences flanking the antigen-encoding peptide in the source protein sequence; optionally, the presence of protease cleavage motifs in the antigen-encoding peptide, weighted according to the expression of the corresponding protease in infected cells (as measured by RNA-seq or mass spectrometry); the conversion rate of the source protein, measured in the appropriate cell type; optionally, the degree of conversion of the source protein to the infected cell, as measured by RNA-seq or proteomic mass spectrometry, or as predicted from annotation of germline or somatic splicing mutations detected in DNA or RNA sequence data. the length of the source protein, taking into account the particular splice variants ("isoforms") that are most highly expressed in the infected cells; the level of expression of the proteasome, immunoproteasome, thymoproteasome, or other proteases in the infected cells (which may be measured by RNA-seq, proteomic mass spectrometry, or immunohistochemistry); the expression of the antigen-encoding peptide source gene (e.g., measured by RNA-seq or mass); the typical tissue-specific expression of the antigen-encoding peptide source gene during various phases of the cell cycle; a comprehensive catalog of features of the source protein and / or its domains that can be found at http: / / www.rcsb.org / pdb / home / home.do; features describing the properties of the domain of the source protein that contains the peptide, e.g., secondary or tertiary structure (e.g., alpha helix vs. beta sheet); alternative splicing; the probability of presentation of the peptide from the source protein of the peptide encoding the antigen in other different subjects; the probability that the peptide will not be detected and over-represented by mass spectrometry due to technical bias; the expression of various gene modules / pathways measured by RNASeq (which does not necessarily include the source protein of the peptide) that provides information about the status of the infected cell, stroma, or infected tissue; the copy number of the source gene of the peptide encoding the antigen in the infected cell; the probability that the peptide binds to TAP or the measured or predicted binding affinity of the peptide to TAP;The infection may be further identified by at least one of the following characteristics: MHC allele non-interaction, including the expression level of TAP in infected cells, which may be measured by RNA-seq, proteomic mass spectrometry, or immunohistochemistry. Peptide presentation can be achieved by the expression of any of the following proteins involved in the antigen presentation machinery (e.g., B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOB, HLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or genes encoding components of the proteasome or immunoproteasome. The presence or absence of functional germline polymorphisms may depend on the components of the antigen presentation machinery, including but not limited to the typical expression of the source gene of the peptide in the relevant infection type or clinical subtype; the gene encoding the peptide; the type of infection (e.g., pathogen infection, viral infection, bacterial infection, fungal infection, and parasitic infection); the clinical infection subtype (e.g., HIV infection, severe acute respiratory syndrome-related coronavirus (SARS) infection, severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infection, Ebola infection, hepatitis B virus (HBV) infection, influenza infection, hepatitis C virus (HCV) infection);

[0076] The methods disclosed herein may also include obtaining an infectious disease vaccine comprising a set of selected antigens (e.g., antigens from an infectious disease organism) or a subset thereof, and may optionally further include administering the infectious disease vaccine to a subject.

[0077] At least one of the antigens in the set of selected antigens (e.g., an antigen from an infectious disease organism) may include at least one of the following: a binding affinity to MHC when in polypeptide form with an IC50 value of less than 1000 nM; a length of 8-15, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids for an MHC class I polypeptide; a length of 6-30, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acids for an MHC class II polypeptide; the presence of a sequence motif within or near the polypeptide in the parent protein sequence that promotes proteasomal cleavage; and the presence of a sequence motif that promotes TAP transport. For MHC class II, the presence of sequence motifs within or near the peptide that promote cleavage by extracellular or lysosomal proteases (eg, cathepsins) or HLA-DM catalyzed HLA binding.

[0078] Disclosed herein is a method for identifying one or more antigens (e.g., antigens from an infectious disease organism) likely to be presented on the cell surface of an infected cell, comprising: receiving mass spectrometry data including data relating to a plurality of isolated peptides eluted from major histocompatibility complexes (MHC) derived from a plurality of fresh or frozen samples; obtaining a training dataset by identifying at least a set of training peptide sequences present in the sample and presented by one or more MHC alleles associated with each training peptide sequence; obtaining a set of training protein sequences based on the training peptide sequences; and training a set of numerical parameters of a presentation model using the training protein sequences and the training peptide sequences, the presentation model providing a plurality of numerical likelihoods that a peptide sequence from an infected cell will be presented by one or more MHC alleles on the infected cell surface.

[0079] The presentation model may represent the dependency between the presence of a particular one of the MHC alleles and a particular pair of amino acids at a particular position in a peptide sequence; and the likelihood that such a peptide sequence containing that particular amino acid at that particular position will be presented on the infected cell surface by that particular one of the MHC alleles of that pair.

[0080] The methods disclosed herein can also include selecting a subset of antigens (e.g., antigens from an infectious disease organism), where the subset of antigens is selected because each has a high likelihood of being presented on the cell surface of an infected cell compared to one or more different antigens.

[0081] The methods disclosed herein can also include selecting a subset of antigens (e.g., antigens from an infectious disease organism), where the subset of antigens is selected because each has a high likelihood of being able to induce a disease-specific immune response in a subject compared to one or more different antigens.

[0082] The methods disclosed herein may also include selecting a subset of antigens (e.g., antigens from an infectious disease organism), each selected because it has a high likelihood, relative to one or more distinct antigens, that it can be presented to a naive T cell by a professional antigen-presenting cell (APC), optionally the APC being a dendritic cell (DC).

[0083] The methods disclosed herein can also include selecting a subset of antigens (e.g., antigens from an infectious disease organism), where the subset of antigens is selected because each of the subsets of antigens has a low likelihood of being subject to inhibition via central or peripheral tolerance to one or more different antigens.

[0084] The methods disclosed herein can also include selecting a subset of antigens (e.g., antigens from an infectious disease organism), where the subset of antigens is selected because each has a lower likelihood of being able to induce an autoimmune response against normal tissue in a subject compared to one or more different antigens.

[0085] The methods disclosed herein may also include selecting a subset of antigens (e.g., antigens from an infectious disease organism), each selected because they have a lower likelihood of being differentially post-translationally modified in infected cells compared to APCs, optionally where the APCs are dendritic cells (DCs).

[0086] The practice of the methods herein employs, unless otherwise indicated, conventional methods of protein chemistry, biochemistry, recombinant DNA techniques, and pharmacology within the skill of the art. Such techniques are explained fully in the literature, see, for example, TECreighton, Proteins: Structures and Molecular Properties (WH Freeman and Company, 1993); A. L. Lehninger, Biochemistry (Worth Publishers, Inc., current addition); Sambrook, et al., Molecular Cloning: A Laboratory Manual (2nd Edition, 1989); Methods In Enzymology (S. Colowick, and N. Kaplan eds., Academic Press, Inc.); Remington's Pharmaceutical Sciences, 18th Edition (Easton, Pennsylvania: Mack Publishing Company, 1990); Carey and Sundberg Advanced Organic Chemistry 3rd Edition (Easton, Pennsylvania: Mack Publishing Company, 1992); rd Ed. (Plenum Press) Vols A and B (1992).

[0087] IV. Antigen Antigens can include nucleotides or polypeptides. For example, antigens can be RNA sequences that code for polypeptide sequences. Thus, antigens useful in vaccines can include nucleotide sequences or polypeptide sequences. Antigens can be derived from nucleotide sequences or polypeptide sequences of infectious disease organisms. Polypeptide sequences of infectious disease organisms include, but are not limited to, pathogen-derived peptides, virus-derived peptides, bacteria-derived peptides, fungi-derived peptides, and / or parasite-derived peptides. Infectious disease organisms include, but are not limited to, severe acute respiratory syndrome-associated coronavirus (SARS), severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), Ebola, HIV, hepatitis B virus (HBV), influenza, hepatitis C virus (HCV), human papillomavirus (HPV), cytomegalovirus (CMV), chikungunya virus, respiratory syncytial virus (RSV), dengue virus, orthomyxoviridae virus, tuberculosis, pancorona, herpes simplex virus infection (HSV), influenza, metapneumovirus (MPV), and parainfluenza virus (PIV).

[0088] Disclosed herein are isolated peptides comprising specific antigens or epitopes of infectious disease organisms identified by the methods disclosed herein, peptides comprising specific antigens or epitopes of known infectious disease organisms, and variant polypeptides or fragments thereof identified by the methods disclosed herein. Antigenic peptides may be described in the context of their coding sequences, where the antigen comprises a nucleotide sequence (e.g., DNA or RNA) that encodes the relevant polypeptide sequence.

[0089] The one or more polypeptides encoded by the antigen nucleotide sequence may include at least one of: a binding affinity to MHC with an IC50 value of less than 1000 nM, a length of 8-15, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids for MHC class I peptides, the presence of a sequence motif within or adjacent to the peptide that promotes proteasomal cleavage, and the presence of a sequence motif that promotes TAP transport; a length of 6-30, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acids for MHC class II peptides, the presence of a sequence motif within or adjacent to the peptide that promotes cleavage by extracellular or lysosomal proteases (e.g., cathepsins), or HLA-DM catalyzed HLA binding.

[0090] One or more antigens may be displayed on the surface of the infected cell.

[0091] The one or more antigens may be immunogenic in a subject with an infection, for example, capable of eliciting a T cell or B cell response in the subject.

[0092] One or more antigens that induce an autoimmune response in a subject may be excluded from consideration in the context of vaccine production for subjects with an infection.

[0093] The size of the at least one antigenic peptide molecule can include, but is not limited to, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120 or more amino acid residues, and any range derivable therein. In certain embodiments, the antigenic peptide molecule is equal to or less than 50 amino acids.

[0094] Antigenic peptides and polypeptides can be 15 residues or less in length, usually about 8 to about 11 residues, particularly 9 or 10 residues, for MHC class I, and 6 to 30 residues (inclusive) for MHC class II.

[0095] If desired, longer peptides can be designed in several ways. In some cases, where the presentation likelihood of a peptide to an HLA allele is predicted or known, the longer peptides can consist of either (1) individual presented peptides with extensions of 2-5 amino acids toward the N-terminus and C-terminus of each corresponding gene product, 2) for each, a concatenation of some or all of the presented peptides with the extended sequence. In another case, where sequencing reveals long (more than 10 residues) epitope sequences present in the infected cells, the longer peptides can consist of (3) the entire stretch of epitope sequence present in the infected cells - thus avoiding the need for computational or in vitro test-based selection of the strongest HLA-presented shorter peptides. In both cases, the use of longer peptides can allow for endogenous processing by patient cells, resulting in more effective antigen presentation and induction of T cell responses. The longer peptides can also be full-length proteins, protein subunits, protein domains, and combinations thereof of peptides expressed in the infectious disease organism.

[0096] Antigenic peptides and polypeptides can be presented by HLA proteins. In some embodiments, antigenic peptides and polypeptides are presented by HLA proteins with higher affinity than wild-type peptides. In some embodiments, antigenic peptides or polypeptides can have IC50 of at least 5000nM or less, at least 1000nM or less, at least 500nM or less, at least 250nM or less, at least 200nM or less, at least 150nM or less, at least 100nM or less, at least 50nM or less.

[0097] In some aspects, the antigenic peptides and polypeptides do not induce an autoimmune response and / or do not induce immune tolerance when administered to a subject.

[0098] Also provided is a composition comprising at least two or more antigenic peptides. In some embodiments, the composition comprises at least two different peptides. At least two different peptides can be derived from the same polypeptide. Different polypeptides means that the peptides vary in length, amino acid sequence, or both. The peptides are derived from any polypeptide that is known to be expressed or found to be expressed in infectious disease organisms.

[0099] Antigenic peptides and polypeptides with desired activities or properties can be modified to provide certain desired attributes, e.g., improved pharmacological properties, while increasing or at least substantially retaining all of the biological activity of the unmodified peptide to bind to the desired MHC molecule and activate the appropriate T cells. For example, antigenic peptides and polypeptides can undergo various changes, such as either conservative or non-conservative substitutions, which can provide certain advantages in their use, e.g., improved MHC binding, stability, or presentation. Conservative substitution means replacing an amino acid residue with another that is biologically and / or chemically similar, e.g., one hydrophobic residue with another, or one polar residue with another. Substitutions include combinations such as Gly, Ala; Val, Ile, Leu, Met; Asp, Glu; Asn, Gln; Ser, Thr; Lys, Arg; and Phe, Tyr. The effect of single amino acid substitutions can also be explored using D-amino acids. Such modifications can be made using well-known peptide synthesis procedures, as described, for example, in Merrifield, Science 232:341-347 (1986), Barany & Merrifield, The Peptides, Gross & Meienhofer, eds. (NY, Academic Press), pp. 1-284 (1979), and Stewart & Young, Solid Phase Peptide Synthesis, (Rockford, Ill., Pierce), 2d Ed. (1984).

[0100] Modification of peptides and polypeptides with various amino acid mimetics or unnatural amino acids can be particularly useful in increasing the stability of peptides and polypeptides in vivo. Stability can be assayed in several ways. For example, peptidases and various biological media, such as human plasma and serum, have been used to test stability. See, for example, Verhoef et al., Eur. J. Drug Metab Pharmacokin. 11:291-302 (1986). Peptide half-life can be conveniently determined using a 25% human serum (v / v) assay. The protocol is generally as follows: Pooled human serum (type AB, non-heat inactivated) is delipidated by centrifugation before use. The serum is then diluted to 25% with RPMI tissue culture medium and used to test peptide stability. At predetermined time intervals, a small amount of the reaction solution is removed and added to either 6% aqueous trichloroacetic acid or ethanol. The cloudy reaction samples are chilled for 15 minutes (4° C.) and then spun to pellet precipitated serum proteins. The presence of peptides is then determined by reverse-phase HPLC using stability-specific chromatographic conditions.

[0101] Peptides and polypeptides can be modified to provide desirable attributes other than improved serum half-life. For example, the ability of a peptide to induce CTL activity can be enhanced by linkage to a sequence containing at least one epitope capable of inducing a T helper cell response. The immunogenic peptide / T helper conjugate can be linked by a spacer molecule. The spacer is usually composed of relatively small neutral molecules, such as amino acids or amino acid mimetics, that are substantially uncharged under physiological conditions. The spacer is usually selected from, for example, Ala, Gly, or other neutral spacers of non-polar amino acids or neutral polar amino acids. It will be understood that the optionally present spacer need not be composed of the same residues and thus can be a hetero- or homo-oligomer. If present, the spacer is usually at least one or two residues, more usually three to six residues. Alternatively, the peptide can be linked to the T helper peptide without a spacer.

[0102] The antigenic peptide may be linked to a T helper peptide directly or via a spacer at either the amino or carboxy terminus of the peptide. The amino terminus of either the antigenic peptide or the T helper peptide may be acylated. Exemplary T helper peptides include tetanus toxoid 830-843, influenza 307-319, malaria spozois circumferential 382-398, and 378-389.

[0103] Proteins or peptides can be produced by any technique known to those skilled in the art, including expressing the protein, polypeptide, or peptide by standard molecular biology techniques, isolating the protein or peptide from a natural source, or chemically synthesizing the protein or peptide. Nucleotide and protein, polypeptide, and peptide sequences corresponding to various genes have been previously disclosed and can be found in computerized databases known to those skilled in the art. One such database is the Genbank and GenPept databases of the National Center for Biotechnology Information at the National Institutes of Health website. The coding regions of known genes can be amplified and / or expressed using techniques disclosed herein or that would be known to those skilled in the art. Alternatively, various commercial preparations of proteins, polypeptides, and peptides are known to those skilled in the art.

[0104] In a further aspect, the antigen includes a nucleic acid (e.g., a polynucleotide) encoding an antigenic peptide or a portion thereof. The polynucleotide may be, for example, DNA, cDNA, PNA, CNA, RNA (e.g., mRNA), either single-stranded and / or double-stranded, or a polynucleotide in a natural or stabilized form, e.g., a polynucleotide with a phosphorothioate backbone, or a combination thereof, with or without introns. Still further aspects provide expression vectors capable of expressing the polypeptide or a portion thereof. Expression vectors for various cell types are well known in the art and can be selected without undue experimentation. Generally, the DNA is inserted into an expression vector, such as a plasmid, in the appropriate orientation and correct reading frame for expression. If necessary, the DNA may be linked to appropriate transcriptional and translational regulatory control nucleotide sequences recognized by the desired host, although such controls are generally available in the expression vector. The vector is then introduced into the host via standard techniques. Guidance can be found, for example, in Sambrook et al. (1989) Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY.

[0105] IV. Exemplary Therapeutic Agents As disclosed herein, selected infectious disease-derived antigens (e.g., selected by predicting the likelihood of presentation of candidate antigens) can be used to develop therapeutics that, when administered to a patient, can induce an immune response against the infectious disease. Exemplary therapeutics include vaccine compositions, compositions comprising T cell receptors (TCRs) or chimeric antigen receptors (CARs) that bind to one or more selected infectious disease-derived antigens, and antibodies that exhibit binding specificity for one or more selected infectious disease-derived antigens. Further details of exemplary therapeutics are described herein.

[0106] IV.A. Vaccine Compositions Also disclosed herein are immunogenic compositions, e.g., vaccine compositions, that can generate a specific immune response, e.g., an immune response against an infectious disease. Vaccine compositions typically include multiple infectious disease-derived antigens, e.g., selected using the methods described herein. Vaccine compositions may also be referred to as vaccines.

[0107] The vaccine may include between 1 and 30 peptides, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 different peptides, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different peptides, or 12, 13, or 14 different peptides. The vaccine may contain 1-100 or more nucleotide sequences, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 109, 109, 109, 108, 109, 101, 102, 103, 104, 105, 106, 107, 2, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more different nucleotide sequences, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different nucleotide sequences, or 12, 13, or 14 different nucleotide sequences may be included. The vaccine contains antigen sequences 1-30, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, and 62. , 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more different antigen sequences, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different antigen sequences, or 12, 13, or 14 different antigen sequences.

[0108] In one embodiment, different peptides and / or polypeptides or nucleotide sequences encoding them are selected such that the peptides and / or polypeptides can associate with different MHC molecules, such as different MHC class I molecules. In some aspects, one vaccine composition comprises coding sequences for peptides and / or polypeptides that can associate with the most frequently occurring MHC class I molecules. Thus, a vaccine composition may comprise different fragments that can associate with at least two preferred, at least three preferred, or at least four preferred MHC class I molecules.

[0109] The vaccine composition may be capable of raising a specific cytotoxic T cell response and / or a specific helper T cell response.

[0110] The vaccine composition may further comprise an adjuvant and / or a carrier. Examples of useful adjuvants and carriers are provided herein below. The composition may be associated with a carrier, such as, for example, a protein, or an antigen-presenting cell, such as, for example, a dendritic cell (DC), capable of presenting peptides to T cells.

[0111] An adjuvant is any substance whose incorporation into a vaccine composition increases or otherwise modifies the immune response to an antigen. A carrier can be a scaffolding structure, such as a polypeptide or polysaccharide, to which an antigen can associate. Optionally, the adjuvant is covalently or non-covalently conjugated.

[0112] The ability of adjuvants to increase immune response to antigens is usually manifested by a significant or substantial increase in immune-mediated reactions or a reduction in disease symptoms.For example, an increase in humoral immunity is usually manifested by a significant increase in the titer of antibodies raised against antigens, and an increase in T cell activity is usually manifested by increased cell proliferation, or cytotoxicity, or cytokine secretion.Adjuvants can also change immune response, for example, by changing a predominantly humoral or Th response to a predominantly cellular or Th response.

[0113] Suitable adjuvants include 1018 ISS, alum, aluminum salts, Amplivax, AS15, BCG, CP-870, 893, CpG7909, CyaA, dSLIM, GM-CSF, IC30, IC31, Imiquimod, ImuFact IMP321, IS Patch, ISS, ISCOMATRIX, JuvImmune, LipoVac, MF59, monophosphoryl lipid A, Montanide IMS 1312, Montanide ISA 206, Montanide ISA 50V, Montanide Adjuvants include, but are not limited to, ISA-51, OK-432, OM-174, OM-197-MP-EC, ONTAK, PepTel vector system, PLG microparticles, resiquimod, SRL172, virosomes and other virus-like particles, YF-17D, VEGF trap, R848, beta-glucan, Pam3Cys, Aquila's QS21 Stimulon (Aquila Biotech, Worcester, Mass., USA) derived from saponin, mycobacterial extracts and synthetic bacterial cell wall mimics, as well as other proprietary adjuvants such as Ribi's Detox.Quil or Superfos. Adjuvants such as incomplete Freund's or GM-CSF are useful. Some immunological adjuvants specific for dendritic cells (e.g., MF59) and their preparation have been previously described (Dupuis M, et al., Cell Immunol. 1998; 186(1):18-27; Allison AC; Dev Biol Stand. 1998; 92:3-11). Cytokines can also be used. Some cytokines have been directly implicated in influencing dendritic cell migration to lymphoid tissues (e.g., TNF-alpha), promoting maturation of dendritic cells into efficient antigen-presenting cells for T lymphocytes (e.g., GM-CSF, IL-1, and IL-4) (U.S. Pat. No. 5,849,589, specifically incorporated herein by reference in its entirety), and acting as immune adjuvants (e.g., IL-12) (Gabrilovich DI, et al., J Immunother Emphasis Tumor Immunol. 1996 (6):414-418).

[0114] CpG immunostimulatory oligonucleotides have also been reported to enhance the effect of adjuvants in a vaccine setting. Other TLR binding molecules, such as RNA-binding TLR7, TLR8, and / or TLR9, can also be used.

[0115] Other examples of useful adjuvants include, but are not limited to, chemically modified CpG (e.g., CpR, Idera), poly(I:C) (e.g., poly i:CI2U), non-CpG bacterial DNA or RNA, as well as cyclophosphamide, sunitinib, bevacizumab, celebrex, NCX-4016, sildenafil, tadalafil, vardenafil, sorafenib, XL-999, CP-547632, pazopanib, ZD2171, AZD2171, ipilimumab, tremelimumab, and SC58175, which may act therapeutically and / or as an adjuvant. The amounts and concentrations of adjuvants and additives can be readily determined by those of skill in the art without undue experimentation. Additional adjuvants include colony stimulating factors, such as granulocyte-macrophage colony stimulating factor (GM-CSF, sargramostim).

[0116] Vaccine compositions may include two or more different adjuvants. Additionally, therapeutic compositions may include any adjuvant material, including any of the above or combinations thereof. It is also contemplated that the vaccine and adjuvant may be administered together or separately in any suitable order.

[0117] The carrier (or excipient) may be present independent of the adjuvant. The function of the carrier may be, for example, to increase the activity or immunogenicity, especially of the variant by increasing its molecular weight, to confer stability, to increase biological activity, or to extend serum half-life. Furthermore, the carrier may aid in the presentation of the peptide to T cells. The carrier may be any suitable carrier known to those skilled in the art, for example, a protein or an antigen-presenting cell. The carrier protein may be, but is not limited to, keyhole limpet hemocyanin, a serum protein, for example, transferrin, bovine serum albumin, human serum albumin, thyroglobulin or ovalbumin, an immunoglobulin, or a hormone, for example, insulin or palmitic acid. For human immunization, the carrier is generally a physiologically acceptable carrier that is acceptable and safe for humans. However, tetanus toxoid and / or diphtheria toxoid are suitable carriers. Alternatively, the carrier may be a dextran, for example, sepharose.

[0118] Cytotoxic T cells (CTLs) recognize antigens in the form of peptides bound to MHC molecules, not intact foreign antigens themselves. The MHC molecules themselves are located on the cell surface of antigen-presenting cells. Thus, when a trimeric complex of peptide antigens, MHC molecules, and APCs is present, activation of CTLs is possible. Correspondingly, not only peptides are used for activation of CTLs, but also when APCs are additionally added with the respective MHC molecules, immune responses can be enhanced. Thus, in some embodiments, the vaccine composition additionally comprises at least one antigen-presenting cell.

[0119] Infectious disease derived antigens may also be used in the treatment of infectious diseases, including but not limited to vaccinia, fowlpox, self-replicating alphaviruses, Maraba viruses, adenoviruses (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629), or lentiviruses, including but not limited to second, third, or hybrid second / third generation lentiviruses and recombinant lentiviruses of any generation, designed to target specific cell types or receptors (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1):45-61; Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3):603-18; Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human The ubiquitin C promoter can be included in a viral vector-based vaccine platform such as Nucl.Acids Res.(2015)43(1):682-690; Zufferey et al.,Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery,J.Virol.(1998)72(12):9873-9880). Depending on the packaging capacity of the viral vector-based vaccine platform, this approach can deliver one or more nucleotide sequences encoding one or more antigen peptides. A wide variety of other vaccine vectors useful for therapeutic administration or immunization of antigens, such as Salmonella typhi vectors, will be clear to those skilled in the art from the description herein.

[0120] IV.A.1. Considerations for Vaccine Design and Manufacturing Stem peptides, meaning those presented by all or most of the tumor subclones, are prioritized for inclusion in the vaccine. Optionally, if there are no stem peptides that are predicted to be highly presented and immunogenic, or if the number of stem peptides that are predicted to be highly presented and immunogenic is small enough that additional non-stem peptides can be included in the vaccine, additional peptides can be prioritized by estimating the number and identity of tumor subclones and selecting peptides to maximize the number of tumor subclones covered by the vaccine.

[0121] Additional candidate antigens may still be available for vaccine inclusion than vaccine technology can support. In addition, uncertainties regarding various aspects of antigen analysis may remain, and trade-offs may exist between different properties of candidate vaccine antigens. Therefore, instead of predefined filters at each step of the selection process, an integrated multidimensional model may be considered that places candidate antigens in a space with at least the following axes and optimizes the selection using an integrated approach: 1. Risk of autoimmunity or tolerance (germline risk) (lower autoimmunity risk is usually favorable) 2. Probability of sequencing artifacts (lower artifact probability is usually preferred) 3. Probability of immunogenicity (higher probability of immunogenicity is usually preferred) 4. Probability of presentation (higher probability of presentation is usually preferable) 5. Gene Expression (higher expression is usually preferred) 6. HLA gene coverage (a higher number of HLA molecules involved in the presentation of a set of antigens may decrease the probability that a tumor will escape immune attack via downregulation or mutation of HLA molecules)

[0122] In various embodiments, vaccine design includes the steps of: 1) identifying a set of epitopes; 2) aligning the epitopes to a reference proteome to define an initial footprint; 3) ranking the remaining epitopes according to coverage and cost relative to cassette size; 4) including the top ranked epitopes; and 5) repeating steps (3) and (4) until a maximal cassette design is reached or all epitopes are included.

[0123] With respect to step (1), identifying the set of epitopes can include a set of short amino acid sequences. In various embodiments, the short amino acid sequences have a length of 8-12 amino acids. In various embodiments, the short amino acid sequences have a length of 9-11 amino acids, and in various embodiments, the short amino acid sequences have a length of 8 amino acids, 9 amino acids, 10 amino acids, 11 amino acids, or 12 amino acids. In various embodiments, the set of epitopes includes epitopes identified using the methods disclosed herein (e.g., using the presentation models disclosed herein) as likely to be presented by MHC alleles. In various embodiments, the set of epitopes includes epitopes that have been publicly documented as likely to be presented by MHC alleles. Exemplary epitopes that have been publicly documented as likely to be presented can be found in the Immune Epitope Database (IEDB) and / or the HIV CD8 T-cell Epitope Database at Los Alamos National Laboratory. The Immune Epitope Database (IEDB) is described in further detail in Vita R, et al, The Immune Epitope Database (IEDB): 2018 update. Nucleic Acids Res. 2018 Oct 24, which is incorporated by reference in its entirety. The Los Alamos National Laboratory HIV CD8 T cell epitope database is described in further detail in Llano, A., et al., 2019 Optimal HIV CTL epitopes update: Growing diversity in epitope length and HLA restriction. HIV Molecular Immunology 2019, 3-27, which is incorporated by reference in its entirety.

[0124] In various embodiments, the set of epitopes includes both 1) epitopes identified using the methods disclosed herein (e.g., using the presentation models disclosed herein) as likely to be presented by MHC alleles, and 2) epitopes that have been publicly documented as likely to be presented by MHC alleles. In certain embodiments, the set of validated epitopes includes epitopes identified as likely to be presented by class I MHC alleles.

[0125] Step (2) involves aligning one or more epitopes of the set of epitopes to a reference proteome. See FIG. 1B, which illustrates an exemplary cassette design methodology according to certain embodiments. Here, one or more epitopes are aligned against the reference proteome to generate an initial footprint. In various embodiments, the one or more epitopes are validated epitopes (e.g., epitopes that have been publicly documented as likely to be presented by MHC alleles). In various embodiments, the initial footprint is constrained using two or more parameters that control the size of the initial footprint. For example, the parameters include a minimum epitope length (min 長さ ), the minimum required number of duplicates (min 重複 ) epitopes, and / or the minimum required number of overlaps (min 重複 ) epitope, the minimum epitope length (min 長さ ) percentage (minimum prop In various embodiments, a minimum epitope length (min 長さ ) is at least 2 amino acids. In various embodiments, the minimum epitope length is at least 3 amino acids, at least 4 amino acids, at least 5 amino acids, at least 6 amino acids, at least 7 amino acids, at least 8 amino acids, at least 9 amino acids, or at least 10 amino acids. In various embodiments, the minimum required number of overlaps (min 重複 An epitope is at least two overlapping epitopes. In various embodiments, the minimum required number of overlaps (min 重複) The epitopes are at least 3 overlapping epitopes, at least 4 overlapping epitopes, at least 5 overlapping epitopes, at least 6 overlapping epitopes, at least 7 overlapping epitopes, at least 8 overlapping epitopes, at least 9 overlapping epitopes, or at least 10 overlapping epitopes.

[0126] Step (3) involves ranking the remaining epitopes according to coverage and cost relative to cassette size. Specifically, coverage is denoted as c and refers to the additional population coverage provided by the epitope if it were included in the cassette. In various embodiments, the coverage provided by a given epitope is calculated according to the frequency of haplotypes in a reference population that cover at least one allele associated with the epitope. Exemplary reference populations may include any of African American (AFA), Hispanic (HIS), Asian Pacific Islander (API), and European (EUR) ancestry. Exemplary reference population data may be found in the U.S. National Bone Marrow Donor Program (e.g., Bioinformatics Be The Match®). See FIG. 1C, which shows an exemplary epitope sequence (AQTKILPR) with exemplary alleles. Here, coverage is calculated for two of the alleles (e.g., A*01:01 and B*08:01). Total coverage represents the sum of haplotypes that contain either or both of the two alleles (e.g., A*01:01 and B*08:01). The cost relative to cassette size is denoted as f and refers to the total increase in cassette size as a result of adding epitopes. In various embodiments, epitopes are ranked according to cost relative to coverage and cassette size, such that the best ranked epitopes provide the best improvement in coverage at the least cost relative to cassette size.

[0127] Step (4) involves including the top ranked epitopes in the cassette design. For example, additional epitopes, such as the top ranked epitopes, are added to expand the initial footprint, as shown under "Expansion - 1st iteration" in Figure 1B.

[0128] Step (5) involves a further repetition of step (3) ranking among the remaining epitopes in the set, and a further repetition of step (4) including the top ranked epitopes in the cassette design. Referring again to FIG. 1B "Expansion-Second Iteration," the footprint can be further expanded to include additional epitopes, such as the top ranked epitopes identified during this iteration. Steps (3) and (4) can be continued to be repeated until a maximal cassette design is reached, or until all epitopes are included in the cassette.

[0129] IV. BT Cell Receptor (TCR) and / or Chimeric Antigen Receptor (CAR) Additionally disclosed herein are T cell receptors (TCRs) designed to bind to antigens predicted to be presented on the surface of cells. The TCRs may be isolated or purified. In the majority of T cells, the TCR is a heterodimeric polypeptide with an alpha (α) chain and a beta (β) chain encoded by TRA and TRB, respectively. The alpha chain generally comprises an alpha variable region encoded by TRAV, an alpha junction region encoded by TRAJ, and an alpha constant region encoded by TRAC. The beta chain generally comprises a beta variable region encoded by TRBV, a beta diversity region encoded by TRBD, a beta junction region encoded by TRBJ, and a beta constant region encoded by TRBC. The TCR-alpha chain is generated by VJ recombination of alpha V and J segments, and the beta chain receptor is generated by V(D)J recombination of beta V, D, and J segments. Additional TCR diversity results from junction diversity. At each of the junctions, some bases may be deleted and others may be added (referred to as N and P nucleotides). In a minority of T cells, the TCR comprises a gamma chain and a delta chain. The TCR gamma chain is generated by VJ recombination, and the TCR delta chain is generated by V(D)J recombination (Kenneth Murphy, Paul Travers, and Mark Walport, Janeway's Immunology 7th edition, Garland Science, 2007, incorporated herein by reference in its entirety). The antigen-binding site of a TCR generally comprises six complementarity determining regions (CDRs). The alpha chain contributes three CDRs: alpha ("α") CDR1, αCDR2, and αCDR3. The beta chain also contributes three CDRs: beta ("β") CDR1, βCDR2, and βCDR3. In general, αCDR3 and βCDR3 are the regions most affected by V(D)J recombination and are responsible for most of the variation in the TCR repertoire.

[0130] The TCR can be designed to specifically recognize an antigen disclosed herein, such as an infectious disease-derived antigen predicted to be presented on the surface of a cell. In various embodiments, the TCR can also be membrane-bound on a cell, such as, for example, a T cell or a natural killer (NK) cell. Thus, the TCR can be used in the context of a soluble antibody and / or a membrane-bound CAR.

[0131] Any of the TCRs disclosed herein may include an alpha variable ("V") segment, an alpha joining ("J") segment, optionally an alpha constant region, a beta variable ("V") segment, optionally a beta diversity ("D") segment, a beta joining ("J") segment, and optionally a beta constant region.

[0132] In some embodiments, the TCR or CAR is a recombinant TCR or CAR. The recombinant TCR or CAR may include any of the TCRs identified herein, but may include one or more modifications. Exemplary modifications, such as amino acid substitutions, are described herein. The amino acid substitutions described herein may be made with reference to the IMGT nomenclature and amino acid numbering found at www.imgt.org.

[0133] The recombinant TCR or CAR may be a human TCR or CAR that includes a fully human sequence, e.g., a naturally occurring human sequence. The recombinant TCR or CAR may retain its naturally occurring human variable domain sequence, but may include modifications to the alpha constant region, the beta constant region, or both the alpha and beta constant regions. Such modifications to the TCR constant region may improve TCR assembly and expression for TCR gene therapy, for example, by driving preferential pairing of an exogenous TCR chain.

[0134] In some embodiments, the α and β constant regions are modified by replacing the mouse constant region sequences with the entire human constant region sequences. Such "murinized" TCRs and methods for their production are described in Cancer Res. 2006 Sep 1;66(17):8878-86, which is incorporated herein by reference in its entirety.

[0135] In some embodiments, the alpha and beta constant regions are modified by replacing certain human residues with mouse residues (human to mouse amino acid exchange), making one or more amino acid substitutions in the human TCR alpha constant (TRAC) region, the TCR beta constant (TRBC) region, or the TRAC and TRAB regions. The one or more amino acid substitutions in the TRAC region may include a Ser substitution at residue 90, an Asp substitution at residue 91, a Val substitution at residue 92, a Pro substitution at residue 93, or any combination thereof. The one or more amino acid substitutions in the human TRBC region may include a Lys substitution at residue 18, an Ala substitution at residue 22, an Ile substitution at residue 133, a His substitution at residue 139, or any combination of the above. Such targeted amino acid substitutions are described in J Immunol June 1, 2010, 184(11)6223-6231, which is incorporated herein by reference in its entirety.

[0136] In some embodiments, human TRAC contains an Asp substitution at residue 210 and human TRBC contains a Lys substitution at residue 134. Such substitutions may facilitate the formation of salt bridges between the alpha and beta chains and the formation of TCR interchain disulfide bonds. These targeted substitutions are described in J Immunol June 1, 2010, 184(11)6232-6241, which is incorporated herein by reference in its entirety.

[0137] In some embodiments, the human TRAC and human TRBC regions are modified to include an introduced cysteine ​​that may improve preferential pairing of exogenous TCR chains by forming additional disulfide bonds. For example, human TRAC may include a Cys substitution at residue 48 and human TRBC may include a Cys substitution at residue 57, as described in Cancer Res. 2007 Apr 15; 67(8): 3898-903 and Blood. 2007 Mar 15; 109(6): 2331-8, which are incorporated herein by reference in their entireties.

[0138] The recombinant TCR or CAR may contain other modifications to the α and β chains.

[0139] In some embodiments, the α and β chains are modified by linking the extracellular domains of the α and β chains to a fully human CD3ζ (CD3-zeta) molecule. Such modifications are described in J Immunol June 1, 2008, 180(11)7736-7746, Gene Ther. 2000 Aug;7(16):1369-77, and The Open Gene Therapy Journal, 2011, 4:11-22, which are incorporated herein by reference in their entireties.

[0140] In some embodiments, the alpha chain is modified by introducing hydrophobic amino acid substitutions into the transmembrane region of the alpha chain, as described in J Immunol June 1, 2012, 188(11)5538-5546, which is incorporated herein by reference in its entirety.

[0141] The alpha or beta chain may be modified by altering any one of the N-glycosylation sites in the amino acid sequence as described in J Exp Med. 2009 Feb 16;206(2):463-475, which is incorporated herein by reference in its entirety.

[0142] The alpha and beta chains may each include a dimerization domain, e.g., a heterologous dimerization domain. Such heterologous domains may be leucine zippers, 5H3 domains or hydrophobic proline-rich counter domains, or other similar modalities, as known in the art. In one example, the alpha and beta chains may be modified by introducing 30-mer segments at the carboxyl termini of the alpha and beta extracellular domains, which selectively associate to form stable leucine zippers. Such modifications are described in PNAS November 22, 1994.91(24)11408-11412, https: / / doi.org / 10.1073 / pnas.91.24.11408, which is incorporated herein by reference in its entirety.

[0143] The TCRs identified herein may be modified to include mutations that increase affinity or half-life, such as those described in WO2012 / 013913, which is incorporated herein by reference in its entirety.

[0144] The recombinant TCR or CAR may be a single chain TCR (scTCR). Such a scTCR may comprise an α chain variable region sequence fused to the N-terminus of a TCR α chain constant region extracellular sequence, a TCR β chain variable region fused to the N-terminus of a TCR β chain constant region extracellular sequence, and a linker sequence linking the C-terminus of the α segment to the N-terminus of the β segment, or vice versa. In some embodiments, the constant region extracellular sequences of the α and β segments of the scTCR are linked by a disulfide bond. In some embodiments, the length of the linker sequence and the position of the disulfide bond are such that the variable region sequences of the α and β segments are oriented with respect to each other substantially like a natural αβ T cell receptor. Exemplary scTCRs are described in U.S. Patent No. 7,569,664, which is incorporated by reference in its entirety.

[0145] In some cases, the variable regions of the scTCR may be covalently joined by a short peptide linker, such as those described in Gene Therapy volume 7, pages 1369-1377 (2000). The short peptide linker may be a serine-rich or glycine-rich linker. For example, the linker may be a (Gly)-rich linker, such as those described in Cancer Gene Therapy (2004) 11, 487-496, which is incorporated by reference in its entirety. 4 Ser) 3 may be also possible.

[0146] The recombinant TCR or antigen-binding fragment thereof may be expressed as a fusion protein. For example, the TCR or antigen-binding fragment thereof may be fused to a toxin. Such a fusion protein is described in Cancer Res. 2002 Mar 15; 62(6): 1757-60. The TCR or antigen-binding fragment thereof may be fused to an antibody Fc region. Such a fusion protein is described in J Immunol May 1, 2017, 198 (1 Supplement) 120.9.

[0147] The antigen recognition domain of a receptor, such as a TCR or a CAR, can be linked to one or more intracellular signaling components, such as signaling components that mimic activation through an antigen receptor complex, such as a TCR complex, and / or signal through another cell surface receptor. For example, a TCR or a CAR can be linked to one or more transmembrane and / or intracellular signaling domains. In some embodiments, the transmembrane domain is fused to the extracellular domain. In one embodiment, one of the domains in the receptor, such as a transmembrane domain that naturally associates with a CAR, is used. In some examples, the transmembrane domain is selected or modified by amino acid substitution to avoid binding of such domains to transmembrane domains of the same or different surface membrane proteins, minimizing interaction with other members of the receptor complex.

[0148] The transmembrane domain in some embodiments is derived from either natural or synthetic sources. When the source is natural, the domain in some aspects is derived from any membrane-bound or transmembrane protein. Transmembrane regions include those derived from (i.e., including at least the transmembrane region(s) of) the alpha, beta, or zeta chain of the T-cell receptor, CD28, CD3 epsilon, CD45, CD4, CD5, CDS, CD9, CD16, CD22, CD33, CD37, CD64, CD80, CD86, CD134, CD137, and / or CD154. Alternatively, the transmembrane domain in some embodiments is synthetic. In some aspects, the synthetic transmembrane domain comprises primarily hydrophobic residues such as leucine and valine. In some aspects, triplets of phenylalanine, tryptophan, and valine are found at each end of the synthetic transmembrane domain. In some embodiments, the linkage is by a linker, spacer, and / or transmembrane domain(s).

[0149] The intracellular signaling domain may mimic or approximate the signaling through a natural antigen receptor, through such a receptor in combination with a costimulatory receptor, and / or through a costimulatory receptor alone. In some embodiments, a short oligo- or polypeptide linker, e.g., a linker 2-10 amino acids in length, such as one that includes a glycine and a serine, e.g., a glycine-serine doublet, is present to form the bond between them. It forms the link between the transmembrane domain of the receptor and the cytoplasmic signaling domain.

[0150] A receptor, e.g., a TCR or a CAR, may include at least one intracellular signaling component(s). In some embodiments, the receptor includes an intracellular component of the TCR complex, such as the TCR CD3 chain, e.g., the CD3 zeta chain, which mediates T cell activation and cytotoxicity. For example, the HLA-PEPTIDE-binding ABP (e.g., a TCR or a CAR) is linked to one or more cell signaling modules. In some embodiments, the cell signaling module includes a CD3 transmembrane domain, a CD3 intracellular signaling domain, and / or other CD transmembrane domains. In some embodiments, the receptor, e.g., a TCR or a CAR, further includes a portion of one or more additional molecules, such as Fc receptor-gamma, CD8, CD4, CD25, or CD16. For example, in some aspects, the TCR or a CAR includes a chimeric molecule between CD3-zeta or Fc receptor-gamma and CD8, CD4, CD25, or CD16.

[0151] In some embodiments, upon ligation of the TCR or CAR, the cytoplasmic domain or intracellular signaling domain of the receptor activates at least one of the normal effector functions or responses of an immune cell, e.g., a T cell engineered to express the receptor. For example, in some situations, the receptor induces a T cell function, such as cytolytic activity or T helper activity, e.g., secretion of cytokines or other factors. In some embodiments, a truncated portion of the intracellular signaling domain of an antigen receptor component or a costimulatory molecule is used instead of an intact immunostimulatory chain, e.g., when transmitting an effector function signal. In some embodiments, the intracellular signaling domain(s) include the cytoplasmic sequence of a T cell receptor (TCR), and in some aspects, those of a co-receptor that cooperates with such receptor in a natural context to initiate signaling following antigen receptor engagement, and / or any derivative or variant of such molecule, and / or synthetic sequences with the same functional capabilities.

[0152] In the context of natural TCR, full activation generally requires not only signal transduction through TCR but also costimulatory signal.Thus, in some embodiments, the receptor also includes components for generating secondary or costimulatory signals to promote full activation.In other embodiments, the receptor does not include components for generating costimulatory signals.In some aspects, additional receptors are expressed in the same cell to provide components for generating secondary or costimulatory signals.

[0153] T cell activation is described in some embodiments as being mediated by two classes of cytoplasmic signaling sequences: those that initiate antigen-dependent primary activation through the TCR (primary cytoplasmic signaling sequences), and those that act in an antigen-independent manner to provide secondary or costimulatory signals (secondary cytoplasmic signaling sequences). In some embodiments, a receptor contains one or both of such signaling components.

[0154] In some aspects, the receptor comprises a primary cytoplasmic signaling sequence that regulates the primary activation of the TCR complex. The primary cytoplasmic signaling sequence that acts in a stimulatory manner may comprise a signaling motif known as an immunoreceptor tyrosine-based activation motif or ITAM. Examples of ITAM-containing primary cytoplasmic signaling sequences include those derived from TCR or CD3 zeta, FcR gamma, FcR beta, CD3 gamma, CD3 delta, CD3 epsilon, CDS, CD22, CD79a, CD79b, and CD66d. In some embodiments, the cytoplasmic signaling molecule(s) in the CAR comprises a cytoplasmic signaling domain, a portion thereof, or a sequence derived from CD3 zeta.

[0155] In some embodiments, the receptor comprises the signaling domain and / or transmembrane portion of a costimulatory receptor such as CD28, 4-1BB, OX40, DAP10, and ICOS, In some aspects, the same receptor comprises both an activating component and a costimulatory component.

[0156] In some embodiments, the activation domain is contained within one receptor, but the costimulatory component is provided by another receptor that recognizes a different antigen. In some embodiments, the receptor includes an activating or stimulatory receptor and a costimulatory receptor, both expressed on the same cell (see WO2014 / 055668). In some aspects, the HLA-PEPTIDE-targeted receptor is a stimulatory or activating receptor, and in other aspects, a costimulatory receptor. In some embodiments, the cell further includes an inhibitory receptor, such as a receptor that recognizes an antigen other than HLA-PEPTIDE (e.g., iCAR, see Fedorov et al., Sci.Transl.Medicine, 5(215) (December, 2013)), such that the activation signal delivered via the HLA-PEPTIDE-targeted receptor is attenuated or inhibited by the binding of the inhibitory receptor to its ligand, for example, reducing off-target effects.

[0157] In certain embodiments, the intracellular signaling domain comprises a CD28 transmembrane and signaling domain linked to a CD3 (e.g., CD3-zeta) intracellular domain. In some embodiments, the intracellular signaling domain comprises a chimeric CD28 and CD137 (4-1BB, TNFRSF9) costimulatory domain linked to a CD3 zeta intracellular domain.

[0158] In some embodiments, the receptor comprises one or more, e.g., two or more, costimulatory domains and an activation domain, e.g., a primary activation domain, in the cytoplasmic portion. Exemplary receptors include the intracellular components of CD3-zeta, CD28, and 4-1BB.

[0159] In some embodiments, the CAR (or other antigen receptor, such as TCR) further comprises a marker, such as a cell surface marker, which can be used to confirm the transduction or engineering of cells to express a receptor, e.g., a truncated version of a cell surface receptor, such as truncated EGFR (tEGFR). In some aspects, the marker comprises all or a portion (e.g., a truncated form) of CD34, nerve growth factor receptor (NGFR), or epidermal growth factor receptor (e.g., tEGFR). In some embodiments, the nucleic acid encoding the marker is operably linked to a polynucleotide encoding a cleavable linker sequence or a ribosomal skip sequence, e.g., a linker sequence such as T2A. See WO2014031687. In some embodiments, introduction of a construct encoding CAR and EGFRt separated by a T2A ribosomal switch allows the two proteins to be expressed from the same construct, so that EGFRt can be used as a marker to detect cells expressing such a construct. In some embodiments, the marker, and optionally the linker sequence, can be any of those disclosed in published patent application WO2014031687. For example, the marker can be a truncated EGFR (tEGFR), optionally linked to a linker sequence, such as a T2A ribosomal skip sequence.

[0160] In some embodiments, the marker is a molecule that is not naturally found on T cells or that is not naturally found on the surface of T cells, such as a cell surface protein, or portion thereof.

[0161] In some embodiments, the molecule is a non-self molecule, e.g., a non-self protein, ie, one that is not recognized as "self" by the immune system of the host into which the cells are adoptively transferred.

[0162] In some embodiments, the marker serves no therapeutic function and / or has no effect other than being used as a marker for genetic engineering, e.g., selection of successfully engineered cells. In other embodiments, the marker may be a therapeutic molecule or a molecule that exerts some other desired effect, e.g., a ligand that the cell encounters in vivo, such as a co-stimulatory or immune checkpoint molecule that enhances and / or suppresses the response of the cell upon adoptive transfer and encounter with the ligand.

[0163] The TCR or CAR may contain one or more modified synthetic amino acids in place of one or more naturally occurring amino acids. Exemplary modified amino acids include aminocyclohexane carboxylic acid, norleucine, α-amino n-decanoic acid, homoserine, S-acetylaminomethylcysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine, (3-phenylserine, (3-hydroxyphenylalanine, phenylglycine, α-naphthylalanine, cyclohexylalanine, cyclohexylglycine, indoline-2-carboxylic acid, 1,2 ,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, aminomalonic acid monoamide, N'-benzyl-N'-methyl-lysine, N',N'-dibenzyl-lysine, 6-hydroxylysine, ornithine, α-aminocyclopentanecarboxylic acid, α-aminocyclohexanecarboxylic acid, α-aminocycloheptanecarboxylic acid, α-(2-amino-2-norvomane)-carboxylic acid, α,γ-diaminobutyric acid, α,γ-diaminopropionic acid, homophenylalanine, and α-tertbutylglycine.

[0164] In some cases, CARs are referred to as first, second, and / or third generation CARs. In some embodiments, a first generation CAR is a CAR that provides only a CD3 chain-induced signal upon antigen binding, in some embodiments, a second generation CAR is a CAR that provides such a signal and a costimulatory signal such as one that includes an intracellular signaling domain from a costimulatory receptor such as CD28 or CD137, and in some embodiments, a third generation CAR in some embodiments is a CAR that includes multiple costimulatory domains of different costimulatory receptors.

[0165] In some embodiments, the chimeric antigen receptor comprises an extracellular portion comprising a TCR or fragment described herein. In some aspects, the chimeric antigen receptor comprises an extracellular portion comprising a TCR or fragment described herein and an intracellular signaling domain. In some embodiments, the intracellular domain comprises an ITAM. In some aspects, the intracellular signaling domain comprises the signaling domain of the zeta chain of the CD3-zeta (CD3) chain. In some embodiments, the chimeric antigen receptor comprises a transmembrane domain linking the extracellular domain and the intracellular signaling domain.

[0166] In some aspects, the transmembrane domain comprises the transmembrane portion of CD28. The extracellular domain and the transmembrane may be directly or indirectly linked. In some embodiments, the extracellular domain and the transmembrane are linked by a spacer, such as any of those described herein. In some embodiments, the chimeric antigen receptor comprises an intracellular domain of a T cell costimulatory molecule, such as between the transmembrane domain and the intracellular signaling domain. In some aspects, the T cell costimulatory molecule is CD28 or 41BB.

[0167] In some embodiments, the CAR comprises a TCR, e.g., a TCR fragment, a transmembrane domain that is or comprises a transmembrane portion of CD28 or a functional variant thereof, and an intracellular signaling domain that comprises a signaling portion of CD28 or a functional variant thereof and a signaling portion of CD3 zeta or a functional variant thereof. In some embodiments, the CAR comprises a TCR, e.g., a TCR fragment, a transmembrane domain that is or comprises a transmembrane portion of CD28 or a functional variant thereof, and an intracellular signaling domain that comprises a signaling portion of 4-1BB or a functional variant thereof and a signaling portion of CD3 zeta or a functional variant thereof. In some such embodiments, the receptor further comprises a spacer that comprises a portion of an Ig molecule, such as a human Ig molecule, such as an Ig hinge, e.g., an IgG4 hinge, such as a hinge-only spacer.

[0168] In some embodiments, the transmembrane domain of a receptor, e.g., a TCR or a CAR, is the transmembrane domain of human CD28 or a variant thereof, e.g., the 27 amino acid transmembrane domain of human CD28 (Accession Number: P10747.1).

[0169] In some embodiments, the chimeric antigen receptor comprises the intracellular domain of a T cell costimulatory molecule. In some aspects, the T cell costimulatory molecule is CD28 or 41BB.

[0170] In some embodiments, the intracellular signaling domain comprises an intracellular costimulatory signaling domain of human CD28 or a functional variant or portion thereof, e.g., the 41 amino acid domain thereof and / or a domain having an LL to GG substitution at positions 186-187 of the native CD28 protein. In some embodiments, the intracellular domain comprises an intracellular costimulatory signaling domain of 41BB or a functional variant or portion thereof, e.g., the 42 amino acid cytoplasmic domain of human 4-1BB (Accession No. Q07011.1) or a functional variant or portion thereof.

[0171] In some embodiments, the intracellular signaling domain comprises a human CD3 zeta stimulatory signaling domain, such as the 112 AA cytoplasmic domain of isoform 3 of human CD3 zeta or a functional variant thereof (Accession Number: P20963.2), or a CD3 zeta signaling domain described in U.S. Pat. No. 7,446,190 or U.S. Pat. No. 8,911,993.

[0172] In some aspects, the spacer comprises only the hinge region of an IgG, such as only an IgG4 or IgG1 hinge. In other embodiments, the spacer is an Ig hinge linked to the CH2 and / or CH3 domains, e.g., an IgG4 hinge. In some embodiments, the spacer is an Ig hinge linked to the CH2 and CH3 domains, e.g., an IgG4 hinge. In some embodiments, the spacer is an Ig hinge linked to the CH3 domain only, e.g., an IgG4 hinge. In some embodiments, the spacer is or comprises a glycine-serine rich sequence or other flexible linker, such as a known flexible linker.

[0173] For example, in some embodiments, a CAR includes a TCR or a fragment thereof, such as any of the HLA-PEPTIDE specific TCRs, a spacer, such as any of the Ig hinge-containing spacers, a CD28 transmembrane domain, a CD28 intracellular signaling domain, and a CD3 zeta signaling domain. In some embodiments, a CAR includes a TCR or a fragment thereof, such as any of the HLA-PEPTIDE specific TCRs, a spacer, such as any of the Ig hinge-containing spacers, a CD28 transmembrane domain, a CD28 intracellular signaling domain, and a CD3 zeta signaling domain.

[0174] IV.C. Methods for Engineering Cells with TCRs and / or CARs Also provided are methods, nucleic acids, compositions, and kits for expressing receptors, including TCRs, CARs, etc., and for producing engineered cells expressing such TCRs, CARs, etc. Genetic engineering generally involves introducing a nucleic acid encoding the recombinant or engineered component into a cell, such as by retroviral transduction, transfection, or transformation.

[0175] In some embodiments, gene transfer is achieved by first stimulating the cells, such as by combining the cells with a stimulus that induces a response such as proliferation, survival, and / or activation, e.g., as measured by expression of cytokines or activation markers, and then transducing the activated cells and expanding them in culture to sufficient numbers for clinical use.

[0176] In some situations, overexpression of stimulatory factors (e.g., lymphokines or cytokines) can be toxic to subjects. Thus, in some situations, engineered cells contain segments that make cells susceptible to negative selection in vivo, such as when administered in adoptive immunotherapy. For example, in some embodiments, cells are designed to be removed as a result of changes in the in vivo conditions of the patient to whom they are administered. A phenotype that can be negatively selected may result from the insertion of a gene that confers sensitivity to an administered agent, e.g., a compound. Negatively selectable genes include the herpes simplex virus type I thymidine kinase (HSV-I TK) gene, which confers sensitivity to ganciclovir (Wigler et al., Cell II:223, 1977), the cellular hypoxanthine phosphoribosyltransferase (HPRT) gene, the cellular adenine phosphoribosyltransferase (APRT) gene, and bacterial cytosine deaminase (Mullen et al., Proc. Natl. Acad. Sci. USA. 89:33 (1992)).

[0177] In some embodiments, the cells are further engineered to promote the expression of cytokines or other factors. A variety of methods for the introduction of engineered components, such as antigen receptors (e.g., TCRs), are well known and may be used with the provided methods and compositions. Exemplary methods include those for the transfer of nucleic acids encoding receptors, including via viral, e.g., retroviral or lentiviral transduction, transposons, nuclease-mediated gene editing (e.g., CRISPR, TALEN, meganuclease, or ZFN editing systems), and electroporation. For example, nuclease-mediated gene editing, particularly for editing T cells, is described in more detail in International Applications WO / 2018 / 232356 and PCT / US2018 / 058230, which are incorporated herein by reference for all purposes.

[0178] In some embodiments, recombinant nucleic acids are transferred to cells using recombinant infectious viral particles, such as vectors derived from Simian Virus 40 (SV40), adenovirus, or adeno-associated virus (AAV). In some embodiments, recombinant lentiviral or retroviral vectors, such as gamma retroviral vectors, are used to transfer recombinant nucleic acids to T cells (see, for example, Koste et al. (2014) Gene Therapy 2014 Apr.3.doi:10.1038 / gt.2014.25; Carlens et al. (2000) Exp Hematol 28(10):1137-46; Alonso-Camino et al. (2013) Mol Ther Nucl Acids 2,e93; Park et al., Trends Biotechnol. 2011 Nov.29(11):550-557).

[0179] In some embodiments, the retroviral vector has a long terminal repeat (LTR), for example, a retroviral vector derived from Moloney murine leukemia virus (MoMLV), myeloproliferative sarcoma virus (MPSV), mouse embryonic stem cell virus (MESV), murine stem cell virus (MSCV), spleen focus forming virus (SFFV), or adeno-associated virus (AAV). Most retroviral vectors are derived from murine retroviruses. In some embodiments, retroviruses include those derived from avian or mammalian cell sources. Retroviruses are typically amphotropic, meaning that they can infect host cells of several species, including humans. In one embodiment, the expressed gene replaces the retroviral gag, pol, and / or env sequences. Several exemplary retroviral systems have been described (e.g., U.S. Pat. Nos. 5,219,740, 6,207,453, 5,219,740; Miller and Rosman (1989) BioTechniques 7:980-990; Miller, AD (1990) Human Gene Therapy 1:5-14; Scarpa et al. (1991) Virology 180:849-852; Burns et al. (1993) Proc. Natl. Acad. Sci. USA 90:8033-8037; and Boris-Lawrie and Temin (1993) Cur. Opin. Genet. Develop. 3:102-109).

[0180] Methods for lentiviral transduction are known. Exemplary methods are described, for example, in Wang et al. (2012) J. Immunother. 35(9): 689-701, Cooper et al. (2003) Blood. 101: 1637-1644, Verhoeyen et al. (2009) Methods Mol Biol. 506: 97-114, and Cavalieri et al. (2003) Blood. 102(2): 497-505.

[0181] In some embodiments, the recombinant nucleic acid is transferred to the T cell via electroporation (see, e.g., Chicaybam et al, (2013) PLoS ONE 8(3):e60298, Van Tedeloo et al. (2000) Gene Therapy 7(16):1431-1437, and Roth et al. (2018) Nature 559:405-409). In some embodiments, the recombinant nucleic acid is transferred to the T cell via transposition (see, e.g., Manuri et al. (2010) Hum Gene Ther 21(4):427-437, Sharma et al. (2013) Molec Ther Nucl Acids 2,e74, and Huang et al. (2009) Methods Mol Biol 506:115-126). Other methods for introducing and expressing genetic material in immune cells include calcium phosphate transfection (e.g., as described in Current Protocols in Molecular Biology, John Wiley & Sons, New York, NY), protoplast fusion, cationic liposome-mediated transfection, tungsten particle-facilitated microparticle bombardment (Johnston, Nature, 346:776-777 (1990)), and strontium phosphate DNA co-precipitation (Brash et al., Mol. Cell Biol., 7:2031-2034 (1987)).

[0182] Other approaches and vectors for the transfer of nucleic acids encoding recombinant products are described, for example, in International Patent Application Publication No. WO2014055668 and U.S. Patent No. 7,446,190.

[0183] Among the additional nucleic acids, e.g., genes for transfer, are those that improve the efficacy of the treatment, such as by promoting the viability and / or function of the transferred cells; genes that provide genetic markers for the selection and / or evaluation of cells, such as to assess survival or localization in vivo; and genes that improve safety by rendering cells susceptible to negative selection in vivo, as described, for example, in SD et al., Mol. and Cell Biol., 11:6 (1991), and Riddell et al., Human Gene Therapy 3:319-338 (1992). See also PCT / US91 / 08442 and PCT / US94 / 05601 by Lupton et al., which describe the use of bifunctional selectable fusion genes resulting from fusing a dominant positive selectable marker with a negative selectable marker. See, e.g., U.S. Patent No. 6,040,177 to Riddell et al., columns 14-17.

[0184] IV.D. Antibodies or Antigen-Binding Fragments Additionally disclosed herein is an antibody or antigen-binding fragment that is designed to bind to an antigen that is predicted to be presented on the surface of a cell. In some embodiments, the antibody or antigen-binding fragment provided herein comprises a light chain. In some aspects, the light chain is a kappa light chain. In some aspects, the light chain is a lambda light chain.

[0185] In some embodiments, the antibodies or antigen-binding fragments provided herein comprise a heavy chain. In some aspects, the heavy chain is IgA. In some aspects, the heavy chain is IgD. In some aspects, the heavy chain is IgE. In some aspects, the heavy chain is IgG. In some aspects, the heavy chain is IgM. In some aspects, the heavy chain is IgG1. In some aspects, the heavy chain is IgG2. In some aspects, the heavy chain is IgG3. In some aspects, the heavy chain is IgG4. In some aspects, the heavy chain is IgA1. In some aspects, the heavy chain is IgA2.

[0186] In some embodiments, the antibodies or antigen-binding fragments provided herein comprise an antibody fragment. In some embodiments, the antibodies or antigen-binding fragments provided herein consist of an antibody fragment. In some embodiments, the antibodies or antigen-binding fragments provided herein consist essentially of an antibody fragment. In some aspects, the antibody fragment is an Fv fragment. In some aspects, the antibody fragment is a Fab fragment. In some aspects, the antibody fragment is an F(ab') 2 In some embodiments, the antibody fragment is a Fab' fragment. In some embodiments, the antibody fragment is an scFv (sFv) fragment. In some embodiments, the antibody fragment is an scFv-Fc fragment. In some embodiments, the antibody fragment is a fragment of a single domain ABP.

[0187] In some embodiments, the antibody fragments provided herein retain the ability to bind to a target, such as an infectious disease-derived antigen predicted to be presented by one or more HLA alleles on the surface of a cell, as measured by one or more assays or biological effects described herein. In some embodiments, the antibody fragments provided herein retain the ability to prevent an infectious disease-derived antigen from interacting with one or more of its ligands, as described herein.

[0188] Antibody fragments provided herein may be produced by any suitable method, including the exemplary methods described herein or known in the art. Suitable methods include recombinant techniques and proteolytic digestion of whole antibodies or antigen-binding fragments.

[0189] In some embodiments, the antibodies provided herein are monoclonal antibodies. Monoclonals may be obtained, for example, using hybridoma techniques or using phage or yeast-based libraries. DNA encoding monoclonal antibodies can be readily isolated and sequenced using conventional procedures. In some embodiments, the antibodies provided herein are polyclonal antibodies. In some embodiments, the antibodies provided herein comprise chimeric ABPs. In some embodiments, the antibodies provided herein consist essentially of chimeric antibodies. Chimeric antibodies can be made by any method known in the art. In some embodiments, chimeric antibodies are made by combining non-human variable regions (e.g., variable regions from mouse, rat, hamster, rabbit, or non-human primate, e.g., monkey) with human constant regions using recombinant techniques.

[0190] In some embodiments, the antibodies provided herein comprise humanized antibodies. In some embodiments, the antibodies provided herein consist of humanized antibodies. In some embodiments, the antibodies provided herein consist essentially of humanized antibodies. Humanized antibodies may be generated by replacing most or all of the structural parts of a non-human monoclonal with the corresponding human antibody sequences.

[0191] In some embodiments, the antibodies provided herein comprise human antibodies. In some embodiments, the antibodies provided herein consist of human antibodies. In some embodiments, the antibodies provided herein consist essentially of human antibodies. Human antibodies can be generated by a variety of techniques known in the art, for example, by using transgenic animals (e.g., humanized mice), can be derived from phage display libraries, can be generated by in vitro activated B cells, or can be derived from yeast-based libraries.

[0192] In some embodiments, the antibodies provided herein comprise an alternative scaffold. In some embodiments, the antibodies provided herein consist of an alternative scaffold. In some embodiments, the antibodies provided herein consist essentially of an alternative scaffold. Any suitable alternative scaffold may be used. In some aspects, the alternative scaffold is selected from Adnectin™, iMab, Anticalin®, EETI-II / AGRP, Kunitz domain, thioredoxin peptide aptamer, Affibody®, DARPin, Affilin, tetranectin, Finomer, and Avimer. The alternative scaffolds provided herein may be made by any suitable method, including the exemplary methods described herein or known in the art.

[0193] Also disclosed herein are isolated humanized, human, or chimeric antibodies that compete with the antibodies disclosed herein for binding to HLA-antigen complexes.

[0194] In certain aspects, the antibody may comprise a human Fc region that comprises at least one modification that reduces binding to a human Fc receptor. It is known that antibodies are post-translationally modified when expressed in cells. Examples of post-translational modifications include cleavage of lysine at the C-terminus of the heavy chain by carboxypeptidase; modification of glutamine or glutamic acid to pyroglutamic acid at the N-terminus of the heavy and light chains by pyroglutamylation; glycosylation; oxidation; deamidation; and glycation, which are known to occur in various ABPs (see Journal of Pharmaceutical Sciences, 2008, Vol. 97, p. 2426-2447, which is incorporated by reference in its entirety). In some embodiments, the antibody is a post-translationally modified antibody or antigen-binding fragment thereof. Examples of post-translationally modified antibodies or antigen-binding fragment thereof include antibodies or antigen-binding fragment thereof that have undergone pyroglutamylation at the N-terminus of the heavy chain variable region and / or deletion of lysine at the C-terminus of the heavy chain. It is known in the art that such post-translational modifications by pyroglutamylation at the N-terminus and deletion of lysine at the C-terminus have no effect on the activity of the antibody or fragment thereof (Analytical Biochemistry, 2006, Vol. 348, p. 24-39, incorporated by reference in its entirety).

[0195] In some embodiments, the antibodies provided herein are multispecific antibodies.

[0196] In some embodiments, the multispecific antibodies provided herein bind to two or more antigens. In some embodiments, the multispecific antibodies bind to two antigens (e.g., bispecific antibodies). In some embodiments, the multispecific antibodies bind to three antigens. In some embodiments, the multispecific antibodies bind to four antigens. In some embodiments, the multispecific antibodies bind to five antigens.

[0197] In some embodiments, the multispecific ABP comprises an antigen binding domain (ABD) that specifically binds to the HLA-PEPTIDE target disclosed herein, and an additional ABD that binds to an additional target antigen. Many multispecific antibody constructs are known in the art, and the antibodies provided herein can be provided in the form of any suitable multispecific construct. The multispecific antibodies provided herein can be made by any suitable method, including the exemplary methods described herein or methods known in the art.

[0198] In certain embodiments, the antibodies provided herein comprise an Fc region. The Fc region can be wild type or a variant thereof. In certain embodiments, the antibodies provided herein comprise an Fc region having one or more amino acid substitutions, insertions, or deletions compared to a naturally occurring Fc region. In some aspects, such substitutions, insertions, or deletions result in an antibody having altered stability, glycosylation, or other characteristics. In some aspects, such substitutions, insertions, or deletions result in a glycosylated antibody.

[0199] In some embodiments, the Fc region is a variant Fc region. A "variant Fc region" or "engineered Fc region" comprises an amino acid sequence that differs from that of a native sequence Fc region by at least one amino acid modification, preferably one or more amino acid substitution(s). Preferably, the variant Fc region has at least one amino acid substitution, e.g., about 1 to about 10 amino acid substitutions, preferably about 1 to about 5 amino acid substitutions, in the native sequence Fc region or the Fc region of the parent polypeptide compared to the native sequence Fc region or the Fc region of the parent polypeptide. The variant Fc region herein will preferably have at least about 80% homology with the native sequence Fc region and / or the Fc region of the parent polypeptide, most preferably at least about 90% homology therewith, more preferably at least about 95% homology therewith.

[0200] The term "antibody comprising an Fc region" refers to an antibody that comprises an Fc region. The C-terminal lysine (residue 447 according to the EU numbering system) of the Fc region may be removed, for example, during purification of the antibody or by recombinant engineering of the nucleic acid encoding the antibody. Thus, an antibody having an Fc region may include an antibody with or without K447.

[0201] In some aspects, the Fc region of the antibodies provided herein is modified to obtain antibodies with altered affinity for Fc receptors or that are more immunologically inactive. In some embodiments, the antibody variants provided herein retain some, but not all, effector functions. Such antibodies may be useful, for example, when antibody half-life is important in vivo, but certain effector functions (e.g., complement activation and ADCC) are unnecessary or deleterious.

[0202] In some embodiments, the antibodies provided herein contain one or more modifications that improve or reduce C1q binding and / or CDC.

[0203] In some embodiments, the antibodies provided herein comprise one or more modifications to enhance half-life, hi some embodiments, the antibodies comprise one or more non-Fc modifications that enhance half-life.

[0204] In some embodiments, the multispecific antibody comprises one or more Fc modifications that promote heteromultimerization, hi some embodiments, the Fc modifications comprise a set of mutations that electrostatically disfavor homodimerization but favor heterodimerization. In some embodiments, the Fc modification comprises a modification of the CH3 sequence that affects the ability of the CH3 domain to bind to an affinity agent, such as Protein A.

[0205] V. METHODS OF TREATMENT AND MANUFACTURING Also provided are methods of inducing an infectious disease organism-specific immune response in a subject, vaccinating against an infectious disease organism, and treating and / or alleviating symptoms of infection associated with an infectious disease organism in a subject by administering to the subject one or more antigens, such as a plurality of antigens identified using the methods disclosed herein.

[0206] In some embodiments, the subject has been diagnosed with or is at risk for an infectious disease. The subject may be a human, dog, cat, horse, or any animal in which an infectious disease organism-specific immune response is desired.

[0207] The antigen may be administered in an amount sufficient to induce a CTL response. The antigen may be administered in an amount sufficient to induce a T cell response. The antigen may be administered in an amount sufficient to induce a B cell response.

[0208] The antigen may be administered alone or in combination with other therapeutic agents. Any suitable treatment for the particular infectious disease may be administered.

[0209] The optimal amount and optimal dosing regimen of each antigen included in the vaccine composition can be determined. For example, the antigen or its variants can be prepared for intravenous (iv), subcutaneous (sc), intradermal (id), intraperitoneal (ip), or intramuscular (im) injection. Injection methods include sc, id, ip, im, and iv. Methods of DNA or RNA injection include id, im, sc, ip, and iv. Other methods of administration of the vaccine composition are known to those skilled in the art.

[0210] Vaccines can be edited such that the selection, number, and / or amount of antigens present in the composition are tissue, infectious disease, and / or patient specific. For example, the exact selection of peptides can be guided by the expression pattern of the parent protein in a given tissue or can be guided by the mutational status of the patient. Selection can depend on the specific type of infectious disease, the type of organism causing the infectious disease (e.g., pathogen, virus, bacteria, fungus, or parasite), the state of the disease, earlier treatment regimes, the immune status of the patient, and the patient's HLA haplotype. Additionally, vaccines can contain components that are personalized according to the personal needs of a particular patient. Examples include varying the selection of antigens according to the expression of antigens in a particular patient, or adjusting for secondary treatments after a first round or scheme of treatment.

[0211] Patients can be identified for administration of antigen vaccines through the use of various diagnostic methods, such as the patient selection methods described further below. Patient selection can involve identifying mutations in one or more genes, or the expression pattern of one or more genes. In some cases, patient selection involves identifying the patient's haplotype. Various patient selection methods can be performed in parallel, for example, sequencing diagnostics can identify both the patient's mutation and haplotype. Various patient selection methods can be performed sequentially, for example, one diagnostic test identifies mutations and a separate diagnostic test identifies the patient's haplotype, where each test can be the same (e.g., both high-throughput sequencing) or different (e.g., one high-throughput sequencing and the other Sanger sequencing) diagnostic method.

[0212] For compositions used as vaccines for infectious diseases, antigens with similar normal self peptides that are highly expressed in normal tissues may be avoided or present in low amounts in the compositions described herein. On the other hand, if it is known that infected cells of a patient express a certain antigen in high amounts, the respective pharmaceutical composition for the treatment of this infection may contain high amounts and / or more than one antigen specific to this antigen or pathway of this antigen that may be included.

[0213] Compositions containing antigens may be administered to individuals already suffering from infectious disease. In therapeutic applications, the compositions are administered to patients in an amount sufficient to induce an effective CTL response against the antigens of the infectious disease organism and cure or at least partially prevent symptoms and / or complications. An appropriate amount to accomplish this is defined as a "therapeutically effective dose." Amounts effective for this use will depend, for example, on the composition, the mode of administration, the stage and severity of the disease being treated, the weight and general health of the patient, and the judgment of the prescribing physician. It should be noted that the compositions may generally be used in severe infectious disease conditions, i.e., life-threatening or potentially life-threatening situations. In such cases, taking into account the minimization of foreign matter and the relatively non-toxic nature of the antigens, substantial excesses of these compositions may be administered and may be felt to be desirable by the treating physician.

[0214] For therapeutic use, administration may begin upon or prior to detection of infection, followed by increasing dosages at least until symptoms are substantially alleviated and for a period thereafter, or until immunity is deemed to be effected (e.g., memory B or T cell populations, or antigen-specific B cells or antibodies are produced).

[0215] Pharmaceutical compositions for therapeutic treatment (e.g., vaccine compositions) are intended for parenteral, topical, nasal, oral, or local administration. Pharmaceutical compositions can be administered parenterally, e.g., intravenously, subcutaneously, intradermally, or intramuscularly. The compositions can be administered to the site of infection to induce a local immune response against the infection. Compositions for parenteral administration are disclosed herein, which include a solution of the antigen, and the vaccine composition is dissolved or suspended in an acceptable carrier, e.g., an aqueous carrier. A variety of aqueous carriers can be used, e.g., water, buffered water, 0.9% saline, 0.3% glycine, hyaluronic acid, and the like. These compositions can be sterilized by conventional well-known sterilization techniques, or can be sterile filtered. The resulting aqueous solutions can be packaged for use as is, or lyophilized, and the lyophilized preparations are combined with a sterile solution before administration. The compositions may contain pharma- ceutically acceptable auxiliary substances required to approximate physiological conditions, such as, for example, pH adjusting and buffering agents, tonicity adjusting agents, wetting agents, and the like, e.g., sodium acetate, sodium lactate, sodium chloride, potassium chloride, calcium chloride, sorbitan monolaurate, triethanolamine oleate, and the like.

[0216] Antigens can also be administered via liposomes to target specific cellular tissues, such as lymphoid tissues. Liposomes are also useful for extending half-life. Liposomes include emulsions, foams, micelles, insoluble monolayers, liquid crystals, phospholipid dispersions, lamellar layers, and the like. In these preparations, the antigen to be delivered is incorporated as part of the liposome, either alone or in conjunction with a molecule that binds to a receptor prevalent among lymphoid cells, such as, for example, a monoclonal antibody that binds to the CD45 antigen, or in conjunction with other therapeutic or immunogenic compositions. Thus, liposomes filled with the desired antigen can be directed to the site of lymphoid cells, where the liposomes then deliver the selected therapeutic / immunogenic composition. Liposomes can be formed from standard vesicle-forming lipids, which generally include neutral and negatively charged phospholipids and sterols, such as cholesterol. The choice of lipid is generally guided by considerations, for example, of liposome size, acid instability, and stability of the liposomes in the bloodstream. A variety of methods for preparing liposomes are available, as described, for example, in Szoka et al., Ann. Rev. Biophys. Bioeng. 9;467 (1980), U.S. Pat. Nos. 4,235,871, 4,501,728, 4,501,728, 4,837,028, and 5,019,369.

[0217] To target immune cells, ligands incorporated into the liposomes can include, for example, antibodies or fragments thereof specific for cell surface determinants of the desired immune system cells. The liposomal suspensions can be administered intravenously, locally, topically, etc., at doses that vary according to, among other things, the mode of administration, the peptide being delivered, and the stage of the disease being treated.

[0218] For therapeutic or immunization purposes, the peptides described herein and optionally nucleic acids encoding one or more of the peptides may be administered to a patient. Several methods are conveniently used to deliver the nucleic acid to a patient. For example, the nucleic acid may be delivered directly as "naked DNA". This approach is described, for example, in Wolff et al., Science 247:1465-1468 (1990), and in U.S. Pat. Nos. 5,580,859 and 5,589,466. The nucleic acid may also be administered using ballistic delivery, for example, as described in U.S. Pat. No. 5,204,253. Particles containing only DNA may be administered. Alternatively, the DNA may be attached to particles such as gold particles. Approaches for delivering nucleic acid sequences may include viral vectors, mRNA vectors, and DNA vectors, with or without electroporation.

[0219] Nucleic acids can also be delivered by complexing with cationic compounds, such as cationic lipids. Lipid-mediated gene delivery methods are described, for example, in 9618372 WOAWO96 / 18372, 9324640 WOAWO93 / 24640, Mannino&Gould-Fogerite, BioTechniques 6(7):682-691(1988), U.S. Patent No. 5,279,833, Rose U.S. Patent No. 5,279,833, 9106309 WOAWO91 / 06309, and Felgner et al., Proc.Natl.Acad.Sci.USA 84:7413-7414(1987).

[0220] Antigens may also be derived from vaccinia, fowlpox, self-replicating alphaviruses, Maraba viruses, adenoviruses (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629), or lentiviruses, including but not limited to second, third, or hybrid second / third generation lentiviruses and recombinant lentiviruses of any generation, designed to target specific cell types or receptors (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1):45-61; Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3):603-18; Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human ubiquitin C (see, e.g., promoter, Nucl. Acids Res. (2015) 43 (1): 682-690; Zufferey et al., Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery, J. Virol. (1998) 72 (12): 9873-9880) can be included in a viral vector-based vaccine platform. Depending on the packaging capacity of the viral vector-based vaccine platform, this approach can deliver one or more nucleotide sequences encoding one or more antigen peptides.The sequences may be flanked by non-mutated sequences, separated by linkers, or preceded by one or more sequences that target subcellular compartments (see, e.g., Gros et al., Prospective identification of neoantigen-specific lymphocytes in the peripheral blood of melanoma patients, Nat Med. (2016) 22(4):433-8; Stronen et al., Targeting of cancer neoantigens with donor-derived T cell receptor repertoires, Science. (2016) 352(6291):1337-41; Lu et al., Efficient identification of mutated cancer antigens recognized by T cells associated with durable tumor regressions, Clin Cancer Res. (2014) 20(13):3401-10). Upon introduction into the host, the infected cells express the antigen, thereby eliciting a host immune (e.g., CTL) response against the infectious disease-derived peptide(s). Vaccinia vectors and methods useful for immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is BCG (Bacille Calmette Guerin). BCG vectors are described in Stover et al. (Nature 351:456-460 (1991)). A wide variety of other vaccine vectors useful for therapeutic administration of antigens or immunization, such as Salmonella typhi vectors, will be apparent to those skilled in the art from the disclosure herein.

[0221] The means of administering the nucleic acid uses a minigene construct that encodes one or more epitopes. The amino acid sequence of the epitope is reverse translated to generate a DNA sequence encoding the selected CTL epitope (minigene) for expression in human cells. A human codon usage table is used to guide the codon selection for each amino acid. The DNA sequences encoding these epitopes are directly adjacent to generate a contiguous polypeptide sequence. Additional elements can be incorporated into the minigene design to optimize expression and / or immunogenicity. Examples of amino acid sequences that can be reverse translated and included in the minigene sequence include helper T lymphocytes, epitopes, leader (signal) sequences, and endoplasmic reticulum retention signals. In addition, MHC presentation of the CTL epitopes can be improved by including synthetic (e.g., poly-alanine) or naturally occurring flanking sequences adjacent to the CTL epitopes. The minigene sequence is converted to DNA by assembling oligonucleotides that encode the plus and minus strands of the minigene. Overlapping oligonucleotides (30-100 bases long) are synthesized, phosphorylated, purified, and annealed under appropriate conditions using well-known techniques. The ends of the oligonucleotides are joined using T4 DNA ligase. This synthetic minigene encoding the CTL epitope polypeptide can then be cloned into a desired expression vector.

[0222] Purified plasmid DNA can be prepared for injection using a variety of formulations. The simplest of these is the reconstitution of lyophilized DNA in sterile phosphate-buffered saline (PBS). Various methods have been described, and new techniques may become available. As mentioned above, nucleic acids are conveniently formulated with cationic lipids. In addition, glycolipids, fusogenic liposomes, peptides, and compounds collectively referred to as protective, interactive, non-condensing (PINC) can also form complexes with purified plasmid DNA to affect variables such as stability, intramuscular distribution, or transport to specific organs or cell types.

[0223] Also disclosed is a method of manufacturing an infectious disease vaccine comprising carrying out the steps of the methods disclosed herein and producing an infectious disease vaccine comprising a plurality of antigens or a subset of a plurality of antigens.

[0224] The antigens disclosed herein can be produced using methods known in the art.For example, the method of producing the antigens or vectors disclosed herein (e.g., vectors that contain at least one sequence encoding one or more antigens) can include culturing host cells under suitable conditions for expressing the antigen or vector, where the host cells contain at least one polynucleotide encoding the antigen or vector, and purifying the antigen or vector.Standard purification methods include chromatographic techniques, electrophoretic techniques, immunological techniques, precipitation, dialysis, filtration, concentration, and chromatofocusing techniques.

[0225] The host cell may include Chinese hamster ovary (CHO) cells, NS0 cells, yeast, or HEK293 cells. The host cell may be transformed with one or more polynucleotides comprising at least one nucleic acid sequence encoding the antigen or vector disclosed herein, and optionally, the isolated polynucleotide further comprises a promoter sequence operably linked to at least one nucleic acid sequence encoding the antigen or vector. In certain embodiments, the isolated polynucleotide may be cDNA.

[0226] VII. Presentation Identification System VII.A. System Overview 1 is an overview of an environment for identifying the likelihood of peptide presentation in a patient, according to one embodiment. The environment 100 provides a context for implementing a presentation identification system 160, which itself includes a presentation information store 165.

[0227] The presentation identification system 160 is one or more computer models embodied in a computing system such as that discussed below with respect to FIG. 7 that receives a peptide sequence associated with a set of MHC alleles (e.g., an infectious disease-derived peptide sequence) and determines the likelihood that the infectious disease-derived peptide sequence will be presented by one or more of the set of associated MHC alleles. This is useful in a variety of contexts. One specific use of the presentation identification system 160 is to receive nucleotide sequences of candidate antigens (e.g., candidate antigen sequences 114) associated with a set of MHC alleles expressed by a patient 110 and determine the likelihood that the candidate antigen will be presented by one or more of the associated MHC alleles of the patient 110 and / or will induce an immunogenic response in the immune system of the patient 110. Those candidate antigens with a high likelihood, as determined by the system 160, can be selected for development of a therapeutic agent 118 (e.g., for inclusion in a vaccine, for development of a TCR specific for the selected antigen, and / or for development of an antibody exhibiting binding affinity to the selected antigen). Thus, when administered, the therapeutic agent can induce an anti-infectious disease immune response from the patient's 110 immune system.

[0228] The presentation identification system 160 determines the presentation likelihood through one or more presentation models, also referred to herein as multipart presentation models. Specifically, a presentation model generates a likelihood that a given peptide sequence is presented for a set of associated MHC alleles and is generated based on the presentation information stored in the store 165. For example, a presentation model may generate a likelihood that a peptide sequence "YVYVADVAAK" is presented for a set of alleles HLA-A*02:01, HLA-B*07:02, HLA-B*08:03, HLA-C*01:04, HLA-A*06:03, HLA-B*01:04 on the cell surface of a sample. The presentation information 165 includes information regarding whether peptides bind to different types of MHC alleles such that those peptides are presented by the MHC alleles, which is determined in the model depending on the position of the amino acid within the peptide sequence. The presentation model can predict whether an unrecognized peptide sequence will be presented in association with a set of associated MHC alleles based on the presentation information 165.

[0229] VII.B. Presentation information 2A and 2B illustrate a method for obtaining presentation information, according to an embodiment. In various embodiments, the presentation information 165 includes two general categories of information: allelic interaction information and allelic non-interaction information. Allelic interaction information includes information that affects presentation of a peptide sequence dependent on the type of MHC allele. Allelic non-interaction information includes information that affects presentation of a peptide sequence independent of the type of MHC allele.

[0230] VII.B.1. Allelic interaction information Allelic interaction information may include identified peptide sequences known to be presented by one or more identified MHC molecules from humans, mice, etc. In particular, this may or may not include data obtained from infectious disease samples. The presented peptide sequences may be identified from cells expressing a single MHC allele. In various embodiments, the presented peptide sequences are collected from a monoallelic cell line engineered to express a given MHC allele and then exposed to a synthetic protein. The peptides presented on the MHC allele are isolated by techniques such as acid elution and identified by mass spectrometry. FIG. 2A shows an example of this, where an exemplary peptide YEMFNDKS presented on a given MHC allele HLA-A*01:01 is isolated and identified by mass spectrometry. Because the peptides are identified through cells engineered to express a single, given MHC protein, a direct association between the presented peptide and the MHC protein to which it is bound is found categorically.

[0231] The presented peptide sequences may also be collected from cells expressing multiple MHC alleles. In humans, six different types of MHC molecules are expressed on a cell. Such presented peptide sequences may be identified from biallelic cell lines that have been engineered to express multiple predetermined MHC alleles. Such presented peptide sequences may also be identified from tissue samples, either from normal tissue samples or tissue samples exposed to one of the pathogens, viruses, bacteria, fungi, or parasites that can cause infectious disease. In this case, MHC molecules may be immunoprecipitated from normal or infectious disease tissue. Peptides presented on multiple MHC alleles may also be isolated by techniques such as acid elution and identified by mass spectrometry. 2B shows an example of this, where six exemplary peptides, YEMFNDKSF, HROEIFSHDFJ, FJIEJFOESS, NEIOREIREI, JFKSIFEMMSJDSSU, and KNFLENFIESOFI, are presented on identified MHC alleles HLA-A*01:01, HLA-A*02:01, HLA-B*07:02, HLA-B*08:01, HLA-C*01:03, and HLA-C*01:04, isolated, and identified by mass spectrometry. In contrast to monoallelic cell lines, the direct association between the presented peptide and the MHC protein to which it is bound may be unknown, since the bound peptide is isolated from the MHC molecule before being identified.

[0232] Allele interaction information can also include mass spectrometry ion current, which depends on both the concentration of peptide-MHC molecule complex and the ionization efficiency of peptide.Ionization efficiency varies from peptide to peptide in a sequence-dependent manner.In general, ionization efficiency varies over about two orders of magnitude from peptide to peptide, while the concentration of peptide-MHC complex varies over a larger range.

[0233] Allelic interaction information may also include measured or predicted binding affinities between a given MHC allele and a given peptide. One or more affinity models may generate such predictions. For example, presentation information 165 may include a predicted binding affinity of 1000 nM between peptide YEMFNDKSF and allele HLA-A*01:01. Most peptides with IC50>1000 nm will not be presented by the MHC, and lower IC50 values ​​increase the probability of presentation.

[0234] The allelic interaction information may also include measurements or predictions of stability of the MHC complex. One or more stability models capable of generating such predictions. For example, returning to the example shown in FIG. 2B, the display information 165 may include a stability prediction of a half-life of 1 hour for the molecule HLA-A*01:01.

[0235] Allelic interaction information can also include measured or predicted rates of peptide-MHC complex formation. Complexes that form at higher rates are more likely to be presented at high concentrations on the cell surface.

[0236] Allelic interaction information may also include peptide sequence and length. MHC class I molecules typically prefer to present peptides with a length of 8-15 peptides. 60-80% of presented peptides have a length of 9.

[0237] Allelic interaction information may also include the presence of kinase sequence motifs on the antigen-encoded peptide, and the absence or presence of certain post-translational modifications on the antigen-encoded peptide. The presence of a kinase motif influences the probability of post-translational modifications, which may enhance or interfere with MHC binding.

[0238] Allelic interaction information can also include expression or activity levels of proteins involved in post-translational modification processes, such as kinases (measured or predicted from RNAseq, mass spectrometry, or other methods).

[0239] Allelic interaction information can also include the probability of presentation of peptides with similar sequences in cells from other individuals expressing particular MHC alleles, as assessed by mass spectrometry proteomics or other means.

[0240] Allelic interaction information can also include the expression levels of particular MHC alleles in the individual (e.g., as measured by RNA-seq or mass spectrometry). Peptides that bind most strongly to MHC alleles expressed at high levels are more likely to be presented than peptides that bind most strongly to MHC alleles expressed at low levels.

[0241] Allelic interaction information can also include the overall antigen-encoded peptide sequence-independent probability of presentation by a particular MHC allele in other individuals expressing that allele.

[0242] Allelic interaction information may also include the overall peptide sequence-independent probability of presentation by MHC alleles of the same family of molecules (e.g., HLA-A, HLA-B, HLA-C, HLA-DQ, HLA-DR, HLA-DP) in other individuals. For example, HLA-C molecules are usually expressed at lower levels than HLA-A or HLA-B molecules, and thus presentation of a peptide by HLA-C is a priori less likely than presentation by HLA-A or HLA-B11.

[0243] The allelic interaction information may also include the protein sequence of a particular MHC allele.

[0244] Any of the MHC allele non-interacting information listed in the following section can also be modeled as MHC allele interacting information.

[0245] VII.B.2. Allelic non-interaction information The allele non-interacting information may include a C-terminal sequence adjacent to the peptide that encodes the antigen in its source protein sequence. The C-terminal flanking sequence may affect the proteasome processing of the peptide. However, the C-terminal flanking sequence is cleaved from the peptide by the proteasome before the peptide is transported to the endoplasmic reticulum and encounters the MHC allele on the surface of the cell. Therefore, the MHC molecule does not receive information about the C-terminal flanking sequence, and therefore the effect of the C-terminal flanking sequence cannot change depending on the MHC allele type. For example, returning to the example shown in FIG. 2B, the presentation information 165 may include the C-terminal flanking sequence FOEIFNDKSLDKFJI of the presented peptide FJIEJFOESS, which is identified from the source protein of the peptide.

[0246] Allele non-interaction information may also include mRNA quantification measurements. For example, mRNA quantification data may be obtained for the same samples that provide mass spectrometry training data. In one embodiment, mRNA quantification measurements are identified from the software tool RSEM. A detailed implementation of the RSEM software tool can be found in Bo Li and Colin N. Dewey. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics, 12:323, August 2011. In one embodiment, mRNA quantification is measured in fragments per kilobase of transcript per million mapped reads (FPKM).

[0247] The allelic non-interaction information may also include the N-terminal sequence adjacent to the peptide within the source protein sequence.

[0248] In certain embodiments, the allelic non-interaction information includes both the C-terminal sequence adjacent to the peptide in the source protein sequence and the N-terminal sequence adjacent to the peptide in the source protein sequence.

[0249] Allele non-interaction information can also include the presence of protease cleavage motifs in the peptides. Peptides containing protease cleavage motifs are less likely to be presented because they are more easily degraded by proteases, and therefore less stable within the cell.

[0250] Allelic non-interaction information can also include the turnover rate of the source protein measured in the appropriate cell type. A faster turnover rate (i.e., a lower half-life) increases the probability of presentation, but the predictive power of this feature is low when measured in different cell types.

[0251] The allelic non-interaction information can also include the length of the source protein.

[0252] The allele non-interaction information may also include the expression level of proteasomes, immunoproteasomes, thymoproteasomes, or other proteases. Different proteasomes have different cleavage site preferences. Proportionately to the expression level, more weight is given to the cleavage preference of each type of proteasome.

[0253] Allele non-interaction information can also include the expression of the peptide's source gene (e.g., measured by RNA-seq or mass spectrometry). Peptides from more highly expressed genes are more likely to be presented. Peptides from genes with undetectable expression levels can be excluded from consideration.

[0254] Allelic non-interaction information can also include the probability that the source mRNA of the peptide encoding the antigen is subject to nonsense-mediated decay, as predicted by models of nonsense-mediated decay, e.g., the model from Rivas et al., Science 2015.

[0255] Allelic non-interaction information may also include typical tissue-specific expression of the peptide source gene during various stages of the cell cycle: a gene that is expressed at low levels overall (as measured by RNA-seq or mass spectrometry proteomics), but known to be expressed at high levels during a particular stage of the cell cycle, is more likely to produce more presented peptides than a gene that is stably expressed at very low levels.

[0256] The allelic non-interaction information may also include a comprehensive catalog of the source protein's features, for example, as provided in uniProt or the PDB (http: / / www.rcsb.org / pdb / home / home.do). These features may include, among others, the protein's secondary and tertiary structure, subcellular localization,11 and Gene Ontology (GO) terms. Specifically, this information may include annotations that operate at the level of the protein (e.g., the length of the 5'UTR) and annotations that operate at the level of specific residues (e.g., a helix motif between residues 300 and 310). These features may also include turn motifs, sheet motifs, and disordered residues.

[0257] The allelic non-interaction information can also include features that describe the characteristics of the domain of the source protein that contains the peptide, such as secondary or tertiary structure (eg, alpha helices versus beta sheets), alternative splicing.

[0258] The allelic non-interaction information may also include features describing the presence or absence of presentation hotspots at the peptide's position in its source protein.

[0259] Allelic non-interaction information may also include the probability of presenting a peptide from its source protein in other individuals (after adjusting for the expression levels of the source protein in the individuals and the effects of different HLA types of those individuals).

[0260] Allelic non-interaction information can also include the probability that a peptide will not be detected or over-represented by mass spectrometry due to technical bias.

[0261] Allelic non-interaction information can also include the probability that a peptide will bind to TAP, or the measured or predicted binding affinity of the peptide to TAP. Peptides that are more likely to bind to TAP, or that bind with higher affinity to TAP, are more likely to be presented.

[0262] Allele non-interaction information may also include the known functionality of HLA alleles, for example, as reflected by HLA allele suffixes. For example, the N suffix in the allele name HLA-A*24:09N indicates a null allele that is not expressed and is therefore unlikely to present an epitope. The complete HLA allele suffix nomenclature is described at https: / / www.ebi.ac.uk / ipd / imgt / hla / nomenclature / suffixes.html.

[0263] VII.C. Presentation Identification System 3A is a high-level block diagram illustrating computer logic components of a presentation identification system 160 according to one embodiment. In this exemplary embodiment, the presentation identification system 160 includes a data management module 312, an encoding module 314, a training module 316, and a prediction module 320. The presentation identification system 160 is also comprised of a training data store 170A and a presentation model store 175. Some embodiments of the model management system 160 have different modules than those described herein. Similarly, functionality may be distributed among the modules in a different manner than that described herein.

[0264] VII.C.1. Data Management Module The data management module 312 generates sets of training data from the presentation information 165. Each set of training data includes a plurality of data instances, each data instance i including at least a presented or unpresented peptide sequence p i , peptide sequence p i One or more associated MHC alleles associated with a i Independent variables including z i , and a set of dependent variables y , which represent information of interest to the presentation identification system 160 in predicting new values ​​of the independent variables. i Includes:

[0265] In one particular embodiment, which will be mentioned throughout the remainder of this specification, the dependent variable y i is peptide p i but one or more associated MHC alleles a i However, in other implementations, the dependent variable y i is a function of the presentation identification system 160 for determining the independent variable z i It will be appreciated that the dependent variable y may represent any other type of information that is of interest in predicting the dependence of the dependent variable y i may also be a numerical value indicating the mass analysis ion current identified for the data instance.

[0266] Peptide sequence p of data instance i i is k i is a sequence of amino acids, k i can vary between data instances i within a range. For example, the range can be 8 to 15 for MHC class I, or 9 to 30 for MHC class II. In one specific embodiment of the system 160, all peptide sequences p in the training data set are i may have the same length, e.g., 9. The number of amino acids in a peptide sequence may vary depending on the type of MHC allele (e.g., MHC allele in humans, etc.). iWhich MHC allele corresponds to the corresponding peptide sequence p i Indicates whether the compound was present in association with

[0267] The data management module 312 also includes a set of peptide sequences p i and associated MHC alleles a i In conjunction with the binding affinity b i and stability i Additional allele interaction variables, such as predictions, may be included. For example, the training data 170 may include peptide p i And, a i The predicted binding affinity between each of the associated MHC molecules shown in i As another example, the training data 170 may include a i The stability predictions for each of the MHC alleles denoted by s i may include.

[0268] In certain embodiments, the data management module 312 includes: 1) peptide p, as determined by mass spectrometry; i but one or more associated MHC alleles a i and 2) a label indicating whether the peptide sequence p i and associated MHC alleles a i In conjunction with the additional binding affinity b i This includes prediction, which allows training of a presentation model (e.g., a multipart presentation model) using training data derived from both 1) binding affinity data between peptide sequences and HLA alleles, and 2) peptide data eluted from mass spectrometry representing the presentation of peptide sequences and HLA alleles. Here, it may be preferable to use both types of data to train the multipart presentation model, especially in situations where the amount of peptide data eluted from peptide sequences generated by mass spectrometry is limited.

[0269] In various embodiments, the peptide sequences (e.g., training peptide sequences) used to train the multipart presentation model are infectious disease-derived peptides. Thus, the multipart presentation model can be trained to accurately predict whether an infectious disease-derived peptide is likely to be presented by one or more HLA alleles. In various embodiments, the peptide sequences (e.g., training peptide sequences) used to train the multipart presentation model are human peptides. In various embodiments, the training peptide sequences used to train the multipart presentation model include both infectious disease-derived peptides and human peptides. In such embodiments, the human peptides may represent additional training peptide sequences to train the multipart presentation model when the amount of infectious disease-derived peptide sequences is limited.

[0270] In certain embodiments, the training data for training the multipart presentation model includes binding affinity data between infectious disease-derived peptide sequences and HLA alleles. In certain embodiments, the training data for training the multipart presentation model includes human peptide data (e.g., human immunopeptidemics) eluted from mass spectrometry representing the presentation of human peptide sequences and HLA alleles. In certain embodiments, the training data for training the multipart presentation model includes 1) binding affinity data between infectious disease-derived peptide sequences and HLA alleles, and 2) human peptide data (e.g., human immunopeptidemics) eluted from mass spectrometry representing the presentation of human peptide sequences and HLA alleles. This situation is beneficial when the eluted peptide data for infectious disease-derived peptide sequences is limited. In this way, the eluted human peptide data is used to complement the training of the multipart presentation model.

[0271] In various embodiments, the data management module 312 also includes the peptide sequence p i In conjunction with allele-free interaction variables such as C-terminal flanking sequences and mRNA quantification measurements, i may be included.

[0272] The data management module 312 may also identify peptide sequences that are not presented by MHC alleles to generate the training data 170. For example, this may include identifying a "longer" sequence of the source protein that includes the peptide sequence to be presented, prior to presentation. If the presentation information includes an engineered cell line, the data management module 312 may identify a set of peptide sequences in the synthetic protein to which the cells were exposed that were not presented to the MHC alleles of the cells. If the presentation information includes a tissue sample, the data management module 312 may identify the source protein from which the peptide sequence to be presented originates, and identify a set of peptide sequences in the source protein that were not presented to the MHC alleles of the tissue sample cells.

[0273] In various embodiments, the data management module 312 may also artificially generate peptides having random sequences of amino acids and identify the generated sequences as peptides that are not presented by MHC alleles. This may be accomplished by randomly generating peptide sequences, allowing the data management module 312 to easily generate large amounts of synthetic data of peptides that are not presented by MHC alleles. In practice, because a small percentage of peptide sequences are presented by MHC alleles, it is highly likely that synthetically generated peptide sequences were not presented by MHC alleles even if they were included in proteins processed by the cell.

[0274] In various embodiments, the data management module 312 artificially generates peptides to balance the training dataset to minimize bias resulting from an imbalanced training dataset. For example, the training dataset may include a set of peptides that are presented by some MHC alleles (M 1 ) and some peptides that were not presented by the MHC alleles (M 2In various embodiments, the training data set may have a smaller number of peptides not presented by the MHC alleles (e.g., M 2 <M 1 ), the data management module 312 artificially generates peptides that were not presented by the MHC allele to balance the training data set. In various embodiments, the data management module 312 calculates the ratio of peptides not presented by the MHC allele to peptides presented by the MHC allele (M 2 / M 1 If the ratio of peptides not presented by MHC alleles (M) is less than a threshold (N), then peptides not presented by MHC alleles are artificially generated. The threshold N can be any of 1, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100. In a specific embodiment, the threshold N is 50. Thus, the ratio of peptides not presented by MHC alleles to peptides presented by MHC alleles (M 2 / M 1 If (p < 0.05) is less than 50, the data management module 312 further balances the training dataset by artificially generating peptides that were not presented by MHC alleles. In such an embodiment, balancing the dataset ensures that there are significantly more peptides in the training dataset that are not presented by MHC alleles, such that the presentation model is properly trained to recognize peptides that are not presented by MHC alleles.

[0275] In various embodiments, the training data set includes binding affinity data between infectious disease-derived peptide sequences and HLA alleles. Thus, in such embodiments, the data management module 312 determines the ratio of peptides not presented by MHC alleles (as indicated by binding affinity values) to peptides presented by MHC alleles (as indicated by binding affinity values) (M 2 / M 1) is less than a threshold (N), then artificially generate peptides that were not presented by MHC alleles. In various embodiments, the training data set includes human peptide data eluted from mass spectrometry (e.g., human immunopeptidemics) representative of human peptide sequences and presentation of HLA alleles. Thus, in such embodiments, the data management module 312 determines the ratio of peptides not presented by MHC alleles to peptides presented by MHC alleles (M 2 / M 1 ) is less than a threshold (N), then artificially generate peptides that were not presented by the MHC allele.

[0276] In various embodiments, the training data set includes both 1) binding affinity data between infectious disease-derived peptide sequences and HLA alleles, and 2) human peptide data eluted from mass spectrometry (e.g., human immunopeptidemics) representing the presentation of human peptide sequences and HLA alleles. Thus, in such embodiments, the data management module 312 may artificially generate peptides from both the binding affinity data and / or the eluted human peptide data. For example, the ratio of peptides not presented by MHC alleles (as indicated by binding affinity values) to peptides presented by MHC alleles (as indicated by binding affinity values) (M 2 / M 1 ) is less than a threshold (N), the data management module 312 artificially generates peptides (as indicated by binding affinity values) that were not presented by the MHC allele to further balance the training dataset.

[0277] As another example, the ratio of peptides not presented by MHC alleles (present in the eluted human peptide data) to peptides presented by MHC alleles (present in the eluted human peptide data) (M 2 / M 1If (N) is less than a threshold (N), the data management module 312 artificially generates peptides that were not represented by the MHC alleles (in the eluted human peptide data) to further balance the training data set. The threshold N can be any of 1, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100. In a particular embodiment, the threshold N is 50.

[0278] In various embodiments, the data management module 312 artificially generates peptides of various lengths. In various embodiments, the data management module 312 artificially generates peptides by sampling from a length distribution that matches the distribution of peptide lengths in the source dataset. In various embodiments, the data management module 312 artificially generates peptides by sampling HLA genotype frequencies that are the same as or similar to the HLA genotype distribution in the source dataset. The source dataset may refer to the dataset from which the training dataset was originally obtained. After artificially generating peptide sequences, the generated peptides are combined with the original training dataset. For example, the data management module 312 adds the generated peptides to the original training dataset before shuffling the training dataset to ensure that the training data is randomly sampled when used to train the presented model.

[0279] 3B illustrates an exemplary set of training data 170, according to one embodiment. Specifically, the first three data instances in training data 170 represent peptide presentation information from a monoallelic cell line, including allele HLA-C*01:03 and three peptide sequences QCEIOWARE, FIEUHFWI, and FEWRHRJTRUJR. The fourth data instance in training data 170 represents peptide presentation information from a multiallelic cell line, including alleles HLA-B*07:02, HLA-C*01:03, HLA-A*01:01, and peptide sequence QIEJOEIJE. The first data instance indicates that peptide sequence QCEIOWARE was not presented by allele HLA-C*01:03. As discussed in the previous two paragraphs, peptide sequences may be randomly generated by data management module 312 or may be identified from the source protein of the presented peptide. The training data 170 also includes predicted binding affinities of 1000 nM and predicted half-lives of 1 hour for peptide sequence-allele pairs. 2 Included are allele-non-interacting variables such as mRNA quantification measurements of FPKM. A fourth data instance indicates that the peptide sequence QIEJOEIJE is presented by one of the alleles HLA-B*07:02, HLA-C*01:03, or HLA-A*01:01. Training data 170 also includes binding affinity and stability predictions for each of the alleles, as well as the C-flanking sequence of the peptide and mRNA quantification measurements for the peptide.

[0280] VII.C.2. Encoding Module In various embodiments, the encoding module 314 encodes information contained in the training data 170 into a representation (e.g., a numerical representation) that can be used to generate one or more presentation models. As one example, the encoding module 314 encodes peptide sequences of candidate infectious disease-derived peptides into a representation for analysis by one or more presentation models. As another example, the encoding module 314 encodes peptide sequences of HLA alleles into a representation for use in generating presentation likelihoods. In various embodiments, the encoding module 314 encodes information contained in the training data 170 into one or more numerical vectors that are input to one or more presentation models.

[0281] In one embodiment, the encoding module 314 one-hot encodes a sequence (e.g., a peptide sequence, a C-terminal flanking sequence, an N-terminal flanking sequence, or a peptide sequence of an MHC allele) through a predefined 20-letter amino acid alphabet. i A peptide sequence having amino acids p i is 20k i The jth position of the peptide sequence is represented as a row vector of p i 20(j-1)+1 , p i 20(j-1)+2 , …, p i 20j has a value of 1. Otherwise, the remaining elements have a value of 0. As an example, for a given alphabet {A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y}, the three amino acid peptide sequence EAF of data instance i is represented as a 60-element row vector p i =[00010000000000000000010000000000000000000000010000000000000000]. i is the protein sequence for the MHC allele in the presentation information h and other sequence data may be encoded in the same manner as above.

[0282] If the training data 170 includes sequences of amino acids of different lengths, the encoding module 314 may further encode the peptides into vectors of equal length by adding PAD characters to extend the predefined alphabet. For example, this may be performed by left-padding the peptide sequence with PAD characters until the length of the peptide sequence reaches the peptide sequence with the maximum length in the training data 170. Thus, if the peptide sequence with the maximum length is k max If the sequence has amino acids, the encoding module 314 encodes each sequence as (20+1) k max We represent it numerically as a row vector of elements. As an example, we have an extended alphabet {PAD, A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y} and a maximum amino acid length of k max For = 5, the same exemplary peptide sequence EAF of three amino acids is represented as a 105-element row vector p i =[100000000000000000000010000000000000000000000010000000000000000000000100000000000000000000000]. i or other sequence data can be encoded in the same manner as above. Thus, the peptide sequence p i or c i Each argument or column in represents the occurrence of a particular amino acid at a particular position in the sequence.

[0283] Although the method for encoding sequence data above has been described with reference to sequences having amino acid sequences, the method may be extended to other types of sequence data as well, such as DNA or RNA sequence data.

[0284] The encoding module 314 also encodes one or more MHC alleles a, b, c, and e of data instance i as a row vector of m elements. iwhere each element h=1, 2, ..., m corresponds to a uniquely specified MHC allele. The element corresponding to the MHC allele specified for data instance i has a value of 1. Otherwise, the remaining elements have a value of 0. As an example, alleles HLA-B*07:02 and HLA-C*01:03 of data instance i corresponding to a multi-allelic cell line in m=4 uniquely specified MHC allele types {HLA-A*01:01, HLA-C*01:08, HLA-B*07:02, HLA-C*01:03} are encoded as a four element row vector a i =

[0011] , where a 3 i =1 and a 4 i = 1. Although examples are described herein with four identified MHC allele types, the number of MHC allele types may in practice be hundreds or thousands. As mentioned above, each data instance i typically includes a peptide sequence p i It contains up to six different MHC allele types associated with

[0285] The encoding module 314 also encodes the label y i We encode,x,as a binary variable with values ​​from the set {0,1}, where a value of 1 represents the peptide x i However, the associated MHC allele a i A value of 0 indicates that the peptide was presented by one of the peptides x i However, the associated MHC allele a i The dependent variable y i If represents mass spectrometry ion current, the encoding module 314 may additionally scale the values ​​using various functions, such as logarithmic functions having a range of [-∞, ∞] for ion current values ​​between [0, ∞].

[0286] The encoding module 314 encodes the peptide p i Allele interaction variable x for hi The pair of x and the associated MHC allele h may be represented as a row vector, where the numerical representations of the allele interaction variables are concatenated one after the other. For example, the encoding module 314 may h i [p i ], [p i b h i ], [p i s h i ], or [p i b h i s h i ], where b h i is peptide p i and the binding affinity prediction for the associated MHC allele h, and s for stability h i Alternatively, one or more combinations of allele interaction variables may be stored individually (e.g., as individual vectors or matrices).

[0287] In one example, the encoding module 314 encodes an allele interaction variable x h i The binding affinity information is represented by incorporating measured or predicted binding affinities into the

[0288] In one example, the encoding module 314 encodes an allele interaction variable x h i The bond stability information is represented by incorporating measured or predicted bond stability into the

[0289] In one example, the encoding module 314 encodes an allele interaction variable x h i The binding on-rate information is represented by incorporating measured or predicted binding on-rates into

[0290] Additional examples of training data that the encoding module 314 may encode are described in WO2017106638 and WO2019168984, each of which is incorporated by reference in its entirety.

[0291] VIII. Training Module The training module 316 constructs one or more presentation models, e.g., one or more multipart presentation models, that generate a likelihood that a peptide sequence (e.g., an infectious disease-derived peptide sequence) will be presented by an MHC allele associated with the peptide sequence. k and MHC allele a k and / or a set of peptide sequences p k MHC allele sequence associated with k Given a peptide sequence p k However, the associated MHC allele a k The estimate u indicates the likelihood presented by one or more of k Generate.

[0292] Reference is now made to Figure 4A, which illustrates a flow process for implementing a multi-part presentation model, according to one embodiment. As shown in Figure 4A, the multi-part presentation model 430 receives as input one or more peptide sequences 410 (e.g., infectious disease-derived peptide sequences), HLA allele sequences 420 from a patient (e.g., sequences of HLA alleles expressed by the patient), and / or identified HLA alleles 425 of the patient. Although Figure 4A illustrates both the HLA allele sequences 420 from the patient and the identified HLA alleles 425 of the patient, in various embodiments, only the HLA allele sequences 420 from the patient are required. For example, the identified HLA alleles 425 may be derived from the HLA allele sequences, assuming that the HLA allele sequences are sufficiently distinct and attributable to only one HLA allele.

[0293] In general, the multipart presentation model 430 has been pre-trained using training data. In certain embodiments, the training data for training the multipart presentation model includes 1) binding affinity data between infectious disease-derived peptide sequences and HLA alleles, and 2) human peptide data eluted from mass spectrometry (e.g., human immunopeptidemics) representing the presentation of human peptide sequences and HLA alleles. This situation is beneficial when there is limited eluted peptide data for infectious disease-derived peptide sequences.

[0294] In general, the multipart presentation model 430 analyzes at least peptide sequences 410 and HLA allele sequences 420 from a patient to determine a set of presentation likelihoods 440, where each presentation likelihood may represent the likelihood that a particular peptide sequence will be presented by an HLA allele expressed by the patient. Based on the presentation likelihoods 440, antigens are identified to generate selected antigens 450. For example, a threshold number of peptide sequences having the highest presentation likelihood 440 values ​​are included as selected antigens, which may be used to develop therapeutics, as described in further detail herein.

[0295] VIII.A. Overview of the Multipart Presentation Model The training module 316 constructs one or more multi-part presentation models based on a training data set stored in store 170 that is generated from the presentation information stored in 165 .

[0296] Reference is now made to FIG. 4B, which illustrates in more detail the network architecture of the multipart presentation model, according to one embodiment. The multipart presentation model includes a first part including at least a pan-allele model 460 portion (also referred to as a pan-specific model portion) and a second part including multiple allele-specific models 470. In general, the pan-allele model 460 and the allele-specific model 470 analyze different information as input, but both output presentation likelihoods per allele. At a high level, the pan-allele model analyzes both peptide sequences and HLA sequences to generate presentation likelihoods per allele. Thus, the pan-allele model architecture represents alleles as HLA sequences instead of classification values, thereby allowing similar alleles to share information but blurring the distinction between alleles. In contrast, the allele-specific model analyzes peptide sequences without analyzing HLA sequences. In various embodiments, the allele-specific model may analyze additional information, such as allele non-interaction information, examples of which include C- or N-flanking sequences. Here, the allele-specific model treats each allele independently of all other alleles, and thus each instance of the allele-specific model learns from all alleles in the training dataset, but similar HLA alleles are learned independently of each other. By utilizing a multipart presentation model 430 that includes both a pan-allele model 460 and an allele-specific model 470, the multipart presentation model 430 can predict the presentation likelihood for each allele with improved accuracy.

[0297] Referring first to the pan-allele model 460 (also referred to as the pan-specific model), in various embodiments, the pan-allele model 460 may receive as input a representation of a peptide sequence 410, which is shown in FIG. 4B as a peptide representation 480. In various embodiments, the peptide sequence 410 is from an infectious disease-derived peptide. In various embodiments, the peptide sequence 410 is from a human peptide. As described herein, the peptide representation 480 may be a representation of the peptide sequence 410, generated by the encoding module 314. Thus, the pan-allele model 460 is trained to properly analyze the peptide representation 480. In some embodiments, the pan-allele model 460 receives the peptide sequence 410 directly as input. Thus, the peptide sequence 410 does not need to undergo conversion to a peptide representation before being input to the pan-allele model 460.

[0298] In various embodiments, the pan-allele model 460 may receive as input a representation of the HLA sequences 420 from the patient, which is shown in FIG. 4B as HLA sequence representation 485. As described herein, the HLA sequence representation 485 may be a representation of the HLA sequences 420 from the patient, generated by the encoding module 314. Thus, the pan-allele model 460 is trained to properly analyze the HLA sequence representation 485. In some embodiments, the pan-allele model 460 receives as input the HLA sequences 420 from the patient directly. Thus, the HLA sequences 420 from the patient do not need to undergo a conversion to a representation before being input to the pan-allele model 460.

[0299] In various embodiments, there may be multiple instances of the pan-allele model 460 in the multi-part presentation model 430. For example, as shown in FIG. 4B, there may be M instances of the pan-allele model 460. Each instance of the pan-allele model 460 may be trained on a random shuffle of the training data, so that each instance sees different combinations and learns different patterns in different training data. In various embodiments, M is any of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 instances of the pan-allele model 460. In a particular embodiment, M represents 10 instances of the pan-allele model 460.

[0300] With reference to the allele-specific model 470, in various embodiments, the allele-specific model 470 may receive as input a representation of the peptide sequence 410, which is shown in FIG. 4B as peptide representation 480. In various embodiments, as shown in FIG. 4B, the peptide representation input to the allele-specific model 470 may be the same peptide representation input to the pan-allele model 460. In other embodiments, different peptide representations are generated such that the pan-allele model 460 and the allele-specific model 470 analyze different peptide representations. In some embodiments, the allele-specific model 470 receives the peptide sequence 410 directly as input. Thus, the peptide sequence 410 does not need to undergo conversion to a peptide representation before being input to the allele-specific model 470.

[0301] Although not shown in FIG. 4B, in various embodiments, the allele-specific model may further receive additional allele non-interacting information as input, examples of which include the C-flanking or N-flanking sequences of the peptide sequence 410.

[0302] In various embodiments, there may be multiple instances of the allele-specific model 470 in the multi-part presentation model 430. For example, as shown in FIG. 4B, there may be N instances of the allele-specific model 470. In various embodiments, N is any of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 instances of the allele-specific model 470. In certain embodiments, N represents 10 instances of the allele-specific model 470.

[0303] 4B, the output of the allele-specific model 470 is combined with the HLA allele representation 490, which is a representation of the identified HLA alleles 425 expressed by the patient. The combination of the outputs of the allele-specific model 470 and the HLA allele representation 490 can be a per-allele likelihood of presentation. In certain embodiments, the HLA allele representation is a one-hot encoding of the identified HLA alleles 425. Thus, the combination of the outputs of the allele-specific model 470 and the HLA allele representation 490 can be a per-allele likelihood of presentation of the particular HLA alleles expressed by the patient.

[0304] In various embodiments, the per-allele likelihoods determined by the pan-allele model 460 and the per-allele likelihoods determined by the multiple allele-specific models 470 undergo a transformation, such as a sigmoid transformation as shown in Figure 4B. Other scaling transformations may be implemented. In various embodiments, the transformation of the per-allele likelihoods is optional and need not be performed.

[0305] The per allele likelihoods determined by the pan-allelic model 460 and the per allele likelihoods determined by the multiple allele-specific models 470 are combined to generate the final presented likelihood 440. In various embodiments, the combination is a statistical combination of the per allele likelihoods determined by the pan-allelic model 460 and the per allele likelihoods determined by the multiple allele-specific models 470. Exemplary statistical combinations include the mean, median, mode, variance, and standard deviation. In certain embodiments, the final presented likelihood 440 represents an average of the per allele likelihoods determined by the pan-allelic model 460 and the per allele likelihoods determined by the multiple allele-specific models 470. In various embodiments, the final presented likelihood 440 may represent a weighted average of the per allele likelihoods determined by the pan-allelic model 460 and the per allele likelihoods determined by the multiple allele-specific models 470. This allows the final presented likelihood 440 to be weighted more heavily to the contribution of one of the types of models.

[0306] In various embodiments, regardless of the particular type of multipart representation model, all multipart representation models capture the dependencies between independent and dependent variables in the training data 170 such that a loss function is minimized. Specifically, the loss function l(y i∈S ,u i∈S ;θ) is the dependent variable y for one or more data instances S in the training data 170. i∈S and the estimated likelihood u for a data instance S generated by the proposed model. i∈S In one particular implementation mentioned throughout the remainder of this specification, the loss function (y i∈S ,u i∈S ;θ) is the negative log-likelihood function given by equation (1a) as follows: In practice, however, a different loss function may be used. For example, if predictions are made for mass spectrometry ion currents, the loss function is the mean squared loss given by Equation 1b as follows: TIFF2025512989000003.tif11128

[0307] A multipart presentation model may be a parametric model in which one or more parameters θ mathematically specify the dependency between the independent and dependent variables. Typically, the loss function (y i∈S ,u i∈S The various parameters of the parametric type representation model that minimizes θ(θ;θ) are determined by a gradient-based numerical optimization algorithm such as a batch gradient algorithm, a stochastic gradient algorithm, etc. Alternatively, the multipart representation model may be a non-parametric model in which the model structure is determined from training data 170 and is not strictly based on a fixed set of parameters.

[0308] In various embodiments, the multipart presentation model yields a particular performance metric, such as the area under the precision-recall curve (PR-AUC).In various embodiments, the multipart presentation model is at least 0.05, at least 0.06, at least 0.07, at least 0.08, at least 0.09, at least 0.1, 0.11, at least 0.12, at least 0.13, at least 0.14, at least 0.15, at least 0.16, at least 0.17, at least 0.18, at least 0.19, at least 0.20, at least 0.21, at least 0.22, at least 0.23, at least 0.24, at least 0.25, at least 0.26, at least 0.27, at least 0.29, at least 0.30, at least 0.31, at least 0.32, at least 0.33, at least 0.34, at least 0.35, at least 0.36, at least 0.37, at least 0.38, at least 0.39, at least 0.40, at least 0.41, at least 0.42, at least 0.43, at least 0.44, at least 0.45, at least 0.46, at least 0.47, at least 0.48, at least 0.49, at least 0.50, at least 0.51, at least 0.52, at least 0.53, at least 0.54, at least 0.55, at least 0.56, at least 0.57, at least 0.58, at least 0.59, at least 0.60, at least 0.61, at least 0.62, at least 0.63, at least 0.64, at least 0.65, at least 0.66, at least 0.67, at least 0.68, 7, at least 0.28, at least 0.29, at least 0.30, at least 0.31, at least 0.32, at least 0.33, at least 0.34, at least 0.35, at least 0.36, at least 0.37, at least 0.38, at least 0.39, at least 0.40, at least 0.41, at least 0.42, at least 0.43, at least 0.44, at least 0.45, at least 0.46, at least 0.47, at least 0.48, at least 0.49, at least 0.50, at least 0.51, at least at least 0.52, at least 0.53, at least 0.54, at least 0.55, at least 0.56, at least 0.57, at least 0.58, at least 0.59, at least 0.60, at least 0.61, at least 0.62, at least 0.63, at least 0.64, at least 0.65, at least 0.66, at least 0.67, at least 0.68, at least 0.69, at least 0.70, at least 0.71, at least 0.72, at least 0.73, at least 0.74, at least 0.75, at least 0.76 , resulting in a PR-AUC of at least 0.77, at least 0.78, at least 0.79, at least 0.80, at least 0.81, at least 0.82, at least 0.83, at least 0.84, at least 0.85, at least 0.86, at least 0.87, at least 0.88, at least 0.89, at least 0.90, at least 0.91, at least 0.92, at least 0.93, at least 0.94, at least 0.95, at least 0.96, at least 0.97, at least 0.98, or at least 0.99.In various embodiments, the multipart display model provides a PR-AUC of at least 0.05. In various embodiments, the multipart display model provides a PR-AUC of at least 0.10. In various embodiments, the multipart display model provides a PR-AUC of at least 0.15. In various embodiments, the multipart display model provides a PR-AUC of at least 0.20. In various embodiments, the multipart display model provides a PR-AUC of at least 0.25. In various embodiments, the multipart display model provides a PR-AUC of at least 0.30. In various embodiments, the performance of the multipart display model is measured in relation to the presentation of viral epitopes by X alleles (e.g., class I or class II HLA alleles). In various embodiments, the X alleles include 1 allele, 2 alleles, 3 alleles, 4 alleles, 5 alleles, 6 alleles, 7 alleles, 8 alleles, 9 alleles, 10 alleles, 11 alleles, 12 alleles, 13 alleles, 14 alleles, 15 alleles, 16 alleles, 17 alleles, 18 alleles, 19 alleles, 20 alleles, 21 alleles, 22 alleles, 23 alleles, 24 alleles, 25 alleles, 26 alleles, 27 alleles, 28 alleles, 29 alleles, or 30 alleles. In certain embodiments, the X alleles include 5 alleles. In certain embodiments, the X alleles include 10 alleles. In certain embodiments, the X alleles include 15 alleles. In certain embodiments, the X alleles include 20 alleles. In certain embodiments, the X alleles include 25 alleles. In certain embodiments, the X alleles include 30 alleles.

[0309] VIII.B. Allele-by-Allele Model Reference is now made to FIG. 5A, which illustrates an implementation of an allele-specific model portion of a multi-part presentation model, according to one embodiment. In various embodiments, the allele-specific model 470 can be a neural network. In certain embodiments, the allele-specific model 470 is a multi-layer perceptron (MLP). Thus, the allele-specific model 470 can be composed of multiple layers 520. As described in more detail herein, the neural network can further include nodes in the layers 520, as well as connections between the nodes with associated parameters. In various embodiments, the allele-specific model 470 includes two or more layers 520. In various embodiments, the allele-specific model 470 includes three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more layers 520.

[0310] An initial layer of the allele-specific model 470 (hereafter referred to as the input layer) receives values ​​of peptide representations 480 derived from peptide sequences 410. In various embodiments, the peptide sequences 410 are from infectious disease derived peptides. In various embodiments, the peptide sequences 410 are from human peptides. The values ​​of peptide representations 480 are propagated through layer 520 to generate output values. Here, the output values ​​are combined with HLA allele representations 490, which are generated based on the identified HLA alleles (e.g., HLA allele 1 through HLA allele 6) expressed by the patient. As shown in FIG. 5A, the combination of the output of the allele-specific model 470 and the HLA allele representations 490 results in a value of likelihood per allele 550. Here, the combination may be a dot product between the output of the allele-specific model 470 and the HLA allele representations 490. Thus, the value of likelihood per allele 550 represents a value that is individualized for the patient based on the HLA alleles expressed. For example, the likelihood values ​​per allele 550 may include six likelihood values ​​corresponding to the six expressed HLA alleles, while other likelihood values ​​corresponding to non-expressed HLA alleles are equal to or near zero values.

[0311] The training module 316 may build an allele-specific model for each allele to predict the presentation likelihood of a peptide on an allele-specific basis, in which case the training module 316 may train the presentation model based on data instances S in the training data 170 generated from cells expressing a single MHC allele.

[0312] In one embodiment, for each allele model, the training module 316 determines the sequence of peptide p for a particular allele h by: k The estimated presentation likelihood u k To model: TIFF2025512989000004.tif8130 (where x h k is peptide p k and the encoded allele interaction variable for the corresponding MHC allele h, where f(·) is an arbitrary function, referred to as a transformation function throughout this specification for ease of explanation). h (·) is an arbitrary function, referred to throughout this specification as the dependence function for convenience of explanation, and the parameters θ h Based on the set of allele interaction variables x h k Generate a dependency score for each MHC allele h using the parameters θ h The set of values ​​is θ h where i is each instance in a subset S of training data 170 generated from cells expressing a single MHC allele h.

[0313] Dependent function g h (x h k ;θ h ) is the output of the MHC allele h that is associated with at least one allelic interaction feature x h k Based on this, in particular, the peptide p kThe dependency score for MHC allele h indicates whether or not the MHC allele h presents the corresponding antigen based on the amino acid position of the peptide sequence of p. k The transformation function f(·) transforms the input, more specifically, in this case, g h (x h k ;θ h ) is converted to an appropriate value to obtain the dependency score for peptide p k indicates the likelihood that a gene is presented by an MHC allele.

[0314] In one particular implementation mentioned throughout the remainder of this specification, f(·) is a function with range within [0,1] for the appropriate domain range. In one example, f(·) is the expit function given by: TIFF2025512989000005.tif10128As another example, f(·) can also be the hyperbolic tangent function given by f(z)=tanh(z) (4) (Here, values ​​in the domain z are equal to or exceed 0.) Alternatively, if predictions are made for mass spectrometry ion currents with values ​​outside the range [0,1], f(·) can be any function, such as the identity function, the exponential function, the logarithmic function, etc.

[0315] Therefore, the peptide sequence p k However, the likelihood of each allele presented by MHC allele h is a function of the dependence g h (·) is the peptide sequence p k to generate a corresponding dependency score. The dependency score can be transformed by a transformation function f(·) to generate a corresponding dependency score for the peptide sequence p k may generate the likelihood for each allele to be presented by MHC allele h.

[0316] VIII.B.1 Dependence functions for allelic interaction variables In one particular embodiment mentioned throughout this specification, the dependent function g h (·) is an affine function given by TIFF2025512989000006.tif6128This is x h k Each allele interaction variable in is expressed as a parameter θ determined for the associated MHC allele h. h , where .function..times ...

[0317] In another specific embodiment mentioned throughout this specification, the dependent function g h (·) is the network function given by TIFF2025512989000007.tif6128 A network model NN that has a set of nodes arranged in one or more layers h Each node is represented by a parameter θ h A particular node may be connected to other nodes via connections having associated parameters in a set of activation functions. The value at one particular node may be represented as the sum of the values ​​of the nodes connected to the particular node, weighted by the associated parameters that are mapped by the activation function associated with the particular node. In contrast to affine functions, network models are advantageous because the presentation model can incorporate nonlinearity and can handle data with amino acid sequences of different lengths. Specifically, through nonlinear modeling, the network model can capture the interactions between amino acids at different positions in a peptide sequence and how this interaction affects peptide presentation.

[0318] In general, the network model NN h(·) may be structured as a feedforward network, e.g., an artificial neural network (ANN), a convolutional neural network (CNN), a deep neural network (DNN), and / or a recurrent network, e.g., a long short-term memory network (LSTM), a bidirectional recurrent network, a deep bidirectional recurrent network, etc.

[0319] In one example, which will be mentioned throughout the remainder of this specification, each MHC allele, h=1,2,...,m, is associated with a separate network model, NN h (·) indicates the output(s) from the network model associated with MHC allele h.

[0320] 5B shows an exemplary network model associated with an MHC allele (e.g., any MHC allele h=3), where the exemplary network model may include a total of four layers and may represent layer 520 of the allele-specific model 470 described with reference to FIG. 5A. As shown in FIG. 5B, the network model NN for MHC allele h=3 3 (·) contains three input nodes at layer l=1, four nodes at layer l=2, two nodes at layer l=3, and one output node at layer l=4. The network model NN 3 (·) is the 10 parameters θ 3 (1), θ 3 (2), …, θ 3 (10) is related to the set of network models. 3 (·) is the three-allele interaction variable x for MHC allele h=3 3 k (1), x 3 k (2) and x 3 k (3) receives input values ​​(individual data instances including encoded polypeptide sequence data and any other training data used) for 3 (x 3 k) The network function may also include one or more network models, each of which takes different allele interaction variables as input.

[0321] In another example, the identified MHC alleles h = 1, 2, ..., m are modeled as a single network NN H (·), NN h (·) denotes one or more outputs of a single network model associated with MHC allele h. In such an example, the parameters θ h may correspond to a set of parameters for a single network model, and thus the parameters θ h A set of can be shared by all MHC alleles.

[0322] FIG. 5C illustrates an exemplary network model NN shared by MHC alleles h=1, 2, ..., m. H As shown in Figure 5C, the network model NN H (·) contains m output nodes, each corresponding to an MHC allele. 3 (·) is the allelic interaction variable x for MHC allele h=3 3 k , and the value NN 3 (x 3 k ) to output m values.

[0323] In yet another example, the dependent function g h (·) can be expressed as: TIFF2025512989000008.tif6128 (where g' h (x h k ;θ' h ) is the bias parameter θ in the set of parameters for allele interaction variables for MHC alleles, which represents the baseline probability of presentation of MHC allele h h 0 θ′, withh , which is an affine function with a set of,network functions, etc.

[0324] In another embodiment, the bias parameter θ h 0 can be shared according to the gene family of the MHC allele h. That is, the bias parameter θ h 0 θ 遺伝子(h) 0 where gene(h) is the gene family of MHC allele h. For example, class I MHC alleles HLA-A*02:01, HLA-A*02:02, and HLA-A*02:03 may be assigned to the gene family "HLA-A", and the bias parameter θ for each of these MHC alleles may be h 0 As another example, the class II MHC alleles HLA-DRB1:10:01, HLA-DRB1:11:01, and HLA-DRB3:01:01 may be assigned to the gene family "HLA-DRB", and the bias parameters θ for each of these MHC alleles may be h 0 may be shared.

[0325] Returning to equation (2), as an example, the affine dependence function g h (·) to identify peptide p among m = 4 different identified MHC alleles. k The likelihood that is presented by MHC allele h=3 can be generated by: TIFF2025512989000009.tif6128 (where x 3 k is the allelic interaction variable specified for MHC allele h=3, and θ 3 is the set of parameters determined for MHC allele h=3 through minimization of the loss function).

[0326] As another example, a separate network transformation function g h(·) to identify peptide p among m = 4 different identified MHC alleles. k The likelihood that is presented by MHC allele h=3 can be generated by: TIFF2025512989000010.tif6128 (where x 3 k is the allelic interaction variable specified for MHC allele h=3, and θ 3 is the network model NN associated with MHC allele h=3 3 (·) is the set of parameters to be determined.

[0327] FIG. 5D illustrates an exemplary network model NN 3 (·) to identify peptide p associated with MHC allele h=3 k As shown in Figure 5D, the network model NN 3 (·) is the allelic interaction variable x for MHC allele h=3 3 k Receives and outputs NN 3 (x 3 k ) The output is then mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.

[0328] VIII.B.2. Allele-by-allele using allele-noninteracting variables In one embodiment, the training module 316 incorporates allele-free interaction variables and trains peptide p by: k The estimated presentation likelihood u k To model: TIFF2025512989000011.tif8131 (where w k is peptide p k indicates the encoded allele non-interaction variable for g w (·) is the parameter θ determined for the allele non-interaction variables w Based on the set of allele non-interaction variables w kSpecifically, the parameters θ h and parameters θ for allele non-interacting variables w The set of values ​​is θ h and θ w where i is each instance in the subset S of training data 170 generated from cells expressing a single MHC allele.

[0329] Dependent function g w (w k ;θ w ) is the output of peptide p k represents a dependency score for the allele-non-interacting variables that indicates whether peptide p is presented by one or more MHC alleles based on the influence of the allele-non-interacting variables. k But peptide p k If the peptide p is associated with a C-terminal flanking sequence that is known to positively affect the presentation of the peptide, the dependence score for the allele-non-interacting variable may have a high value. k But peptide p k It may have a low value if it is associated with a C-terminal flanking sequence that is known to negatively affect the presentation of the protein.

[0330] According to formula (7), the peptide sequence p k However, the likelihood of each allele presented by MHC allele h is a function g h (·) is the peptide sequence p k to generate the corresponding dependency scores for the allele-interacting variables. w (·) is also applied to the encoded version of the allele-noninteracting variables to generate a dependency score for the allele-noninteracting variables. Both scores are combined and the combined score is transformed by a transformation function f(·) to obtain the dependency score for the peptide sequence p kgenerates the likelihood for each allele to be presented by MHC allele h.

[0331] Alternatively, the training module 316 may include the allele interaction variable x h k to the allele non-interaction variable w k By adding the allele non-interaction variable w k Thus, the presentation likelihood may be given by: TIFF2025512989000012.tif8128

[0332] VIII.B.3 Dependence functions for allele-noninteracting variables Dependence function g for allelic interaction variables h Similarly to (·), the dependence function g w (·) indicates that a separate network model is used to model the allele-noninteracting variables w k It may be an affine function or a network function related to

[0333] Specifically, the dependent function g w (·) is an affine function given by TIFF2025512989000013.tif6128This is w k The allele non-interaction variables in w , where .function..times ...

[0334] Dependent function g w (·) can also be a network function given by: TIFF2025512989000014.tif6128 Parameter θ w A network model NN having parameters associated in a set w The network function may also include one or more network models, each of which takes different allele-noninteracting variables as input.

[0335] In another example, the dependence function g for the allele non-interacting variables w (·) can be given by: TIFF2025512989000015.tif6128 (where g' w (w k ;θ' w ) is an affine function, and the allele non-interaction parameter θ' w and so on, with a set of m k is peptide p k is the mRNA quantification measurement for θ(·), h(·) is a function that transforms the quantification measurement, and θ w m is a parameter in the set of parameters for the allele-non-interacting variables that is combined with the mRNA quantification measure to generate a dependency score for the mRNA quantification measure.) In one particular embodiment mentioned throughout the remainder of this specification, h(·) is a logarithmic function, although in practice h(·) may be any one of a variety of different functions.

[0336] In yet another example, the dependence function g for the allele non-interacting variables w (·) can be given by: TIFF2025512989000016.tif6128 (where g' w (w k ;θ' w ) is an affine function, and the allele non-interaction parameter θ' w , and so on, with a set of k is peptide p k is a representation vector representing proteins and isoforms in the human proteome for w o is the set of parameters in the set of parameters for allele non-interaction variables that are combined with the representation vector). k Dimension and parameter θ w oIf the set of is significantly higher, then when determining the value of the parameter, TIFF2025512989000017.tif5128 (where A parameter regularization term, such as L1 norm, L2 norm, a combination, etc., may be added to the loss function. The optimal value of the hyperparameter λ may be determined by an appropriate method.

[0337] In yet another example, the dependence function g for the allele non-interacting variables w (·) can be given by: TIFF2025512989000019.tif14130 (where g' w (w k ;θ' w ) is an affine function, and the allele non-interaction parameter θ' w and so on, a network function with a set of TIFF2025512989000020.tif5128 is peptide p k is an indicator function that is equal to 1 if the allele-noninteracting variable is from source gene l, as described above for allele-noninteracting variables, and θ w l is a parameter that indicates the "antigenicity" of the source gene l. In one variation, L is significantly higher, and therefore the parameter θ w l=1、2、…、L If the number of is significantly higher, when determining the value of the parameter, TIFF2025512989000021.tif5128 (where A parameter regularization term, such as L1 norm, L2 norm, a combination, etc., may be added to the loss function. The optimal value of the hyperparameter λ may be determined by an appropriate method.

[0338] In yet another example, the dependence function g for the allele non-interacting variables w (·) can be given by: TIFF2025512989000023.tif14158 (where g' w (w k ;θ' w ) is an affine function, and the allele non-interaction parameter θ' w and so on, a network function with a set of TIFF2025512989000024.tif5128 is peptide p k is from source gene l and peptide p k is an indicator function equal to 1 if the allele is from tissue type m, as described above for allele non-interaction variables, and θ w lm is a parameter indicative of the antigenicity of the combination of source gene l and tissue type m. Specifically, the antigenicity of gene l to tissue type m may indicate the residual propensity of cells of tissue type m to present peptides from gene l after controlling for RNA expression and peptide sequence context).

[0339] In one variation, L or M is significantly higher, and therefore the parameter θ w lm=1、2、…、LM If the number of is significantly higher, when determining the value of the parameter, TIFF2025512989000025.tif5128 (where TIFF2025512989000026.tif5128 represents L1 norm, L2 norm, combination, etc.) may be added to the loss function. The optimal value of the hyperparameter λ may be determined by an appropriate method. In another variation, a parameter regularization term may be added to the loss function in determining the value of the parameter such that the parameters for the same source gene do not differ significantly between tissue types. For example, a penalty term such as: TIFF2025512989000027.tif19128 (where TIFF2025512989000028.tif5128 is the average antigenicity across tissue types for source gene l) can penalize the standard deviation of antigenicity across different tissue types in a loss function.

[0340] In yet another example, the dependence function g for the allele non-interacting variables w (·) can be given by: TIFF2025512989000029.tif31128 (where g' w (w k ;θ' w ) is an affine function, and the allele non-interaction parameter θ' w and so on, a network function with a set of TIFF2025512989000030.tif5128 is peptide p k is an indicator function that is equal to 1 if the allele-noninteracting variable is from source gene l, as described above for allele-noninteracting variables, and θ w l is a parameter indicating the "antigenicity" of the source gene l, TIFF2025512989000031.tif5128 is peptide p k is an indicator function that is equal to 1 if from proteomic position m, TIFF2025512989000032.tif4128 is a parameter indicating the degree to which proteomic position m is a presentation "hot spot." In one embodiment, a proteomic position can contain a block of n adjacent peptides from the same protein, where n is a hyperparameter of the model determined by a suitable method such as grid search cross-validation.

[0341] In practice, any of the additional terms in equations (9), (10), (11), (12a), and (12b) are combined to produce the dependence function g for the allele non-interaction variables. w For example, the term h(·), which represents the mRNA quantification measure in equation (9), and the terms representing the antigenicity of the source genes in equations (11), (12a), and (12b), may be summed together with any other affine or network functions to generate a dependence function for the allele non-interacting variables.

[0342] Returning to equation (7), as an example, the affine transformation function g h (·), g w (·) to identify peptide p among m = 4 different identified MHC alleles. k The likelihood that is presented by MHC allele h=3 can be generated by: TIFF2025512989000033.tif6128 (where w k is peptide p k is the allele-noninteraction variable identified for w is the set of parameters determined for the allele non-interacting variables).

[0343] As another example, the network transformation function g h (·), g w (·) to identify peptide p among m = 4 different identified MHC alleles. k The likelihood that is presented by MHC allele h=3 can be generated by: TIFF2025512989000034.tif6128 (where w k is peptide p k is the allele interaction variable specified for w is the set of parameters determined for the allele non-interacting variables).

[0344] FIG. 5E illustrates an exemplary network model NN 3 (·) and NN w (·) to identify peptide p associated with MHC allele h=3 k As shown in Figure 5E, the network model NN 3 (·) is the allelic interaction variable x for MHC allele h=3 3 k Receives and outputs NN 3 (x 3 k ) is generated. w (·) is peptide p kThe allele non-interaction variable w k Receives and outputs NN w (w k The outputs are combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.

[0345] VIII.C. Multi-allelic model The training module 316 may also build a presentation model to predict the presentation likelihood of a peptide in a multi-allelic setting when more than one MHC allele is present, in which case the training module 316 may train the presentation model based on data instances S in the training data 170 generated from cells expressing a single MHC allele, cells expressing multiple MHC alleles, or a combination thereof.

[0346] VIII.C.1. Example 1: Maximum of Models Per Allele In one embodiment, the training module 316 determines a presentation likelihood for each of the MHC alleles h in the set H determined based on cells expressing a single allele, as described above in conjunction with equations (2)-(10): Peptide p in association with a set of multiple MHC alleles H as a function of TIFF2025512989000035.tif4128 k The estimated presentation likelihood u k Specifically, we model the presentation likelihood u k teeth, In one embodiment, the function is a maximum function, as shown in equations (11), (12a), and (12b), where the presented likelihood u k may be determined as the maximum of the presentation likelihood of each MHC allele h in the set H. TIFF2025512989000037.tif6128

[0347] VIII.C.2. Example 2.1: Sum Function Model In one embodiment, the training module 316 trains peptide p by: k The estimated presentation likelihood u k Modeling TIFF2025512989000038.tif14128 (where element a h k is the peptide sequence p k 1 for multiple MHC alleles H associated with x h k is peptide p k and the encoded allele interaction variables for the corresponding MHC alleles). h The set of values ​​is θ h where i is each instance in the subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. h is the dependent function g introduced above. h The compound may be in any of the following forms:

[0348] According to formula (13), the peptide sequence p k The likelihood of presentation by one or more MHC alleles h is a function of the dependence g h (·) is the peptide sequence p k to generate a corresponding score for the allele interaction variables. The scores for each MHC allele h are combined and transformed by a transformation function f(·) to generate a corresponding score for the peptide sequence p k generates the likelihood of presentation given by the set H of MHC alleles.

[0349] The model presented in formula (13) is k This differs from the allele-by-allele model of equation (2) in that the number of alleles associated with a can exceed 1. In other words, h kTwo or more elements in k may have a value of 1 for multiple MHC alleles H associated with

[0350] As an example, the affine transformation function g h (·) to identify peptide p among m = 4 different identified MHC alleles. k The likelihood that is presented by MHC alleles h=2, h=3 can be generated by: TIFF2025512989000039.tif6128 (where x 2 k , x 3 k is the allelic interaction variable specified for MHC alleles h=2, h=3, and θ 2 , θ 3 is the set of parameters determined for MHC alleles h=2, h=3).

[0351] As another example, the network transformation function g h (·), g w (·) to identify peptide p among m = 4 different identified MHC alleles. k The likelihood that is presented by MHC alleles h=2, h=3 can be generated by: TIFF2025512989000040.tif6128 (where NN 2 (·), N.N. 3 (·) is the network model specified for MHC alleles h = 2, h = 3, and θ 2 , θ 3 is the set of parameters determined for MHC alleles h=2, h=3).

[0352] FIG. 5F illustrates an exemplary network model NN 2 (·) and NN 3 (·) to identify peptides p associated with MHC alleles h=2 and h=3. k As shown in Figure 5F, the network model NN 2(·) is the allelic interaction variable x for MHC allele h=2 2 k Receives and outputs NN 2 (x 2 k ) and generate a network model NN 3 (·) is the allelic interaction variable x for MHC allele h=3 3 k Receives and outputs NN 3 (x 3 k The outputs are combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.

[0353] VIII.C.3. Example 2.2: Sum Function Model with Allele Non-Interacting Variables In one embodiment, the training module 316 incorporates allele-free interaction variables and trains peptide p by: k The estimated presentation likelihood u k To model: TIFF2025512989000041.tif14142 (where w k is peptide p k (The encoded allele-noninteracting variables for each MHC allele h are shown in Fig. 1.) h and parameters θ for allele non-interacting variables w The set of values ​​is θ h and θ w where i is each instance in the subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. w is the dependent function g introduced above. w The compound may be in any of the following forms:

[0354] Thus, according to formula (14), the peptide sequence p kThe likelihood of presentation by one or more MHC alleles H is a function g h (·) is the peptide sequence p k to generate corresponding dependency scores for allele-interacting variables for each MHC allele h. w (·) is also applied to the encoded version of the allele-noninteracting variables to generate a dependency score for the allele-noninteracting variables. The scores are combined and the combined score is transformed by a transformation function f(·) to produce a dependency score for the peptide sequence p k generates the likelihood of presentation by the MHC allele H.

[0355] In the model presented in formula (14), each peptide p k The number of alleles associated with a can be more than one. h k Two or more elements in k may have a value of 1 for multiple MHC alleles H associated with

[0356] As an example, the affine transformation function g h (·), g w (·) to identify peptide p among m = 4 different identified MHC alleles. k The likelihood that is presented by MHC alleles h=2, h=3 can be generated by: TIFF2025512989000042.tif6128 (where w k is peptide p k is the allele-noninteraction variable identified for w is the set of parameters determined for the allele non-interacting variables).

[0357] As another example, the network transformation function g h (·), g w(·) to identify peptide p among m = 4 different identified MHC alleles. k The likelihood that is presented by MHC alleles h=2, h=3 can be generated by: TIFF2025512989000043.tif6128 (where w k is peptide p k is the allele interaction variable specified for w is the set of parameters determined for the allele non-interacting variables).

[0358] FIG. 5G illustrates an exemplary network model NN 2 (·), N.N. 3 (·), and N.N. w (·) is used to identify MHC alleles h=2, h= 3 Peptide p associated with k As shown in Figure 5G, the network model NN 2 (·) is the allelic interaction variable x for MHC allele h=2 2 k Receives and outputs NN 2 (x 2 k ) is generated. 3 (·) is the allelic interaction variable x for MHC allele h=3 3 k Receives and outputs NN 3 (x 3 k ) is generated. w (·) is peptide p k The allele non-interaction variable w k Receives and outputs NN w (w k The outputs are combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.

[0359] Alternatively, the training module 316 may include the allele interaction variable x hk to the allele non-interaction variable w k By adding the allele non-interaction variable w k Thus, the presentation likelihood may be given by: TIFF2025512989000044.tif14134

[0360] VIII.C.4. Example 3.1: Model using implicit allele-specific likelihood In another embodiment, the training module 316 trains peptide p by: k The estimated presentation likelihood u k To model: TIFF2025512989000045.tif8142 (where element a h k is the peptide sequence p k 1 for multiple MHC alleles h ∈ H associated with u' k h is the implied allele-wise presentation likelihood for MHC allele h, and vector v has elements v h A h k ·u' k h where v is a vector corresponding to v, s(·) is a function that maps the elements of v, and r(·) is a clipping function that clips the values ​​of the input to a given range. As explained in more detail below, s(·) may be an additive function or a quadratic function, although it will be appreciated that in other embodiments s(·) may be any function, such as a maximum function. Values ​​for a set of parameters θ for the implicit per-allele likelihood may be determined by minimizing a loss function with respect to θ, where i is each instance in a subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles.

[0361] The presentation likelihood in the presentation model of equation (16) is kThe implicit allele-wise presentation likelihood u' corresponds to the likelihood of being presented by an individual MHC allele h. k h The implicit per-allele likelihood differs from the per-allele presentation likelihood in that, in addition to the monoallelic setting, the parameters of the implicit per-allele likelihood can be learned from a biallelic setting, where the direct association between the presented peptide and the corresponding MHC allele is unknown. Thus, in the biallelic setting, the presentation model is k Not only can it be estimated whether a peptide is presented by the set of MHC alleles H as a whole, but also which MHC alleles h are most likely to present the peptide p k The individual likelihoods of presenting: TIFF2025512989000046.tif4128 can also be provided. The advantage of this is that the presented model can generate implicit likelihoods without using training data for cells expressing a single MHC allele.

[0362] In one particular implementation mentioned throughout the remainder of this specification, r(·) is a function with range [0,1]. For example, r(·) can be a clip function: r(z)=min(max(z,0),1) (where the minimum between z and 1 is the proposed likelihood u k In another embodiment, r(·) is the hyperbolic tangent function given by: r(z)=tanh(z) (where the values ​​of domain z are equal to or greater than 0).

[0363] VIII.C.5. Example 3.2: Sum of Functions Model In one particular embodiment, s(·) is an additive function and the representation likelihood is given by adding the representation likelihood for each implied allele. TIFF2025512989000047.tif16128

[0364] In one embodiment, the implied allele-by-allele presentation likelihood for MHC allele h is generated by: TIFF2025512989000048.tif8128As a result, the presentation likelihood is estimated by: TIFF2025512989000049.tif14129

[0365] According to formula (19), the peptide sequence p k The likelihood of presentation by one or more MHC alleles H is a function g h (·) is the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables. Each dependency score is first transformed by a function f(·) to produce an implicit per-allele presentation likelihood u' k h Generate the likelihood for each allele, u' k h are combined and a clipping function is applied to the combined likelihood to clip the values ​​to the range [0,1] to produce a peptide sequence p k can generate the likelihood of presentation given a set of MHC alleles H. The dependence function g h is the dependent function g introduced above. h The compound may be in any of the following forms:

[0366] As an example, the affine transformation function g h (·) to identify peptide p among m = 4 different identified MHC alleles. k The likelihood that is presented by MHC alleles h=2, h=3 can be generated by: TIFF2025512989000050.tif8128 (where x 2 k , x 3 k is the allelic interaction variable specified for MHC alleles h=2, h=3, and θ2 , θ 3 is the set of parameters determined for MHC alleles h=2, h=3).

[0367] As another example, the network transformation function g h (·), g w (·) to identify peptide p among m = 4 different identified MHC alleles. k The likelihood that is presented by MHC alleles h=2, h=3 can be generated by: TIFF2025512989000051.tif8128 (where NN 2 (·), N.N. 3 (·) is the network model specified for MHC alleles h = 2, h = 3, and θ 2 , θ 3 is the set of parameters determined for MHC alleles h=2, h=3).

[0368] FIG. 5H illustrates an exemplary network model NN 2 (·) and NN 3 (·) to identify peptides p associated with MHC alleles h=2 and h=3. k As shown in Figure 5H, the network model NN 2 (·) is the allelic interaction variable x for MHC allele h=2 2 k Receives and outputs NN 2 (x 2 k ) and generate a network model NN 3 (·) is the allelic interaction variable x for MHC allele h=3 3 k Receives and outputs NN 3 (x 3 k ) Each output is then mapped by a function f(·) and combined to produce an estimated presentation likelihood u k Generate.

[0369] In another embodiment, where the prediction is made on the logarithm of the mass analysis ion current, r(·) is a logarithmic function and f(·) is an exponential function.

[0370] VIII.C.6. Example 3.3: Sum of Functions Model with Allele Non-Interacting Variables In one embodiment, the implied allele-by-allele presentation likelihood for MHC allele h is generated by: TIFF2025512989000052.tif8128As a result, the presented likelihood is generated by: TIFF2025512989000053.tif14141 Incorporating the effects of allelic non-interacting variables on peptide presentation.

[0371] According to formula (21), the peptide sequence p k The likelihood of presentation by one or more MHC alleles H is a function g h (·) is the peptide sequence p k to generate corresponding dependency scores for allele-interacting variables for each MHC allele h. w (·) is also applied to the encoded version of the allele-non-interacting variables to generate a dependency score for the allele-non-interacting variables. The scores of the allele-non-interacting variables are combined into each of the dependency scores for the allele-interacting variables. Each of the combined scores is transformed by a function f(·) to generate an implicit per-allele presentation likelihood. The implicit likelihoods are combined and a clipping function is applied to the combined output to clip the values ​​to the range [0,1] to produce a representation of the peptide sequence p k may generate the likelihood of presentation by the MHC allele H. The dependence function g w is the dependent function g introduced above. w The compound may be in any of the following forms:

[0372] As an example, the affine transformation function g h (·), g w (·) to identify peptide p among m = 4 different identified MHC alleles. k The likelihood that is presented by MHC alleles h=2, h=3 can be generated by: TIFF2025512989000054.tif8128 (where w k is peptide p k is the allele-noninteraction variable identified for w is the set of parameters determined for the allele non-interacting variables).

[0373] As another example, the network transformation function g h (·), g w (·) to identify peptide p among m = 4 different identified MHC alleles. k The likelihood that is presented by MHC alleles h=2, h=3 can be generated by: TIFF2025512989000055.tif8139 (where w k is peptide p k is the allele interaction variable specified for w is the set of parameters determined for the allele non-interacting variables).

[0374] FIG. 5I shows an exemplary network model NN 2 (·), N.N. 3 (·), and N.N. w (·) to identify peptides p associated with MHC alleles h=2 and h=3. k As shown in Figure 5I, the network model NN 2 (·) is the allelic interaction variable x for MHC allele h=2 2 k Receives and outputs NN 2 (x 2 k ) is generated. w (·) is peptide pk The allele non-interaction variable w k Receives and outputs NN w (w k ) The outputs are combined and mapped by a function f(·). Network model NN 3 (·) is the allelic interaction variable x for MHC allele h=3 3 k Receives and outputs NN 3 (x 3 k ) which again uses the same network model NN w (·) Output NN w (w k ) and mapped by a function f(·). Both outputs are combined to produce an estimated presentation likelihood u k Generate.

[0375] In another embodiment, the implied allele-by-allele presentation likelihood for MHC allele h is generated by: TIFF2025512989000056.tif8128As a result, the presentation likelihood is generated by: TIFF2025512989000057.tif14128

[0376] VIII.C.7. Example 4: Second-Order Model In one embodiment, s(·) is a quadratic function, and the peptide p k The estimated presentation likelihood u k is given by: TIFF2025512989000058.tif14148 (where element u' k h(i is the implied per-allele presentation likelihood for MHC allele h). A set of values ​​for parameters θ for the implied per-allele likelihood may be determined by minimizing a loss function with respect to θ, where i is each instance in a subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. The implied per-allele presentation likelihood may be of any form shown in equations (18), (20), and (22) above.

[0377] In one embodiment, the model of formula (23) is a model for peptide p, whose presentation by the two HLA alleles is statistically independent. k However, this may suggest that there is a probability that a given antigen is simultaneously presented by two MHC alleles.

[0378] According to formula (23), the peptide sequence p k However, the presentation likelihood of being presented by one or more MHC alleles H is calculated by combining the implicit allele-wise presentation likelihoods, and assuming that each pair of MHC alleles simultaneously presents a peptide p k The likelihood of presenting the peptide sequence p is subtracted from the total. k can be generated by generating the presentation likelihood of being presented by the MHC allele H.

[0379] As an example, the affine transformation function g h (·) to identify peptide p among m = 4 different identified HLA alleles. k The likelihood that is represented by HLA alleles h=2, h=3 can be generated by: TIFF2025512989000059.tif6128 (where x 2 k , x 3 k is the allele interaction variable specified for HLA alleles h=2, h=3, and θ 2 , θ 3 is the set of parameters determined for HLA alleles h=2, h=3).

[0380] As another example, the network transformation function g h (·), g w (·) to identify peptide p among m = 4 different identified HLA alleles. k The likelihood that is represented by HLA alleles h=2, h=3 can be generated by: TIFF2025512989000060.tif6137 (where NN 2 (·), N.N. 3 (·) is the network model specified for HLA alleles h = 2, h = 3, and θ 2 , θ 3 is the set of parameters determined for HLA alleles h=2, h=3).

[0381] VIII.D. Pan-Allele Model In contrast to per-allele models, a pan-allelic model (e.g., pan-allelic model 460 depicted in FIG. 4B) is a model that can predict the presentation likelihood of a peptide on a pan-allelic basis. Specifically, unlike per-allele models that can predict the probability that a peptide will be presented by one or more known MHC alleles that were previously used to train the per-allele model, a pan-allelic model is a presentation model that can predict the probability that a peptide will be presented by any MHC allele, including unknown MHC alleles that the model did not previously encounter during training.

[0382] Briefly, the pan-allele model is trained by the training module 316. Similar to training the per-allele model, the training module 316 may train the pan-allele presentation model based on data instances S in the training data 170 generated from cells expressing a single MHC allele, cells expressing multiple MHC alleles, or a combination thereof. However, the training module 316 may train the pan-allele presentation model based on data instances S in the training data 170 generated from cells expressing a single MHC allele, cells expressing multiple MHC alleles, or a combination thereof. k hRather than training a pan-allele presentation model using all MHC allele peptide sequences d available in the training data 170, the training module 316 h In particular, the training module 316 trains the pan-allele presentation model based on the amino acid positions of the MHC alleles available in the training data 170.

[0383] After the pan-allele model is trained, when a peptide sequence and a known or unknown MHC allele peptide sequence are input into the model to determine the probability that a known or unknown MHC allele will present a peptide, the model can accurately predict this probability by using the information learned during training with similar MHC allele peptide sequences. For example, a pan-allele model trained using training data 170 that does not include occurrences of the A*02:07 allele can still accurately predict peptide presentation by the A*02:07 allele by applying the information learned during training with similar alleles (e.g., alleles of the A*02 gene family). In this way, the single presentation pan-allele model can predict the likelihood of presentation of a peptide for any MHC allele.

[0384] VIII.D.2. Advantages of pan-allelic models The main advantage of the pan-allele presentation model is that it has a higher generality than the per-allelic presentation model.As mentioned above, the per-allelic model can predict the probability that a peptide is presented by one or more specified MHC alleles that are used to train the per-allelic model.In other words, the per-allelic model is associated with a limited set of one or more known MHC alleles.

[0385] Thus, given a sample containing a particular set of one or more MHC alleles, a per-allele model trained using that particular set of MHC alleles is selected for use to determine the probability that a peptide will be presented by that particular set of MHC alleles. In other words, when relying on a per-allele model to predict the probability that a peptide will be presented by an MHC allele, predictions can only be made for MHC alleles that appeared in the training data 170. Due to the large number of MHC alleles (especially for small variations within the same gene family), a very large number of training samples is required to train a per-allele presentation model to make peptide presentation predictions for all MHC alleles.

[0386] In contrast, the pan-allele model is not limited to making predictions for the particular set of one or more MHC alleles on which it was trained. Instead, in use, the pan-allele model can accurately predict the probability that previously seen and / or previously unseen MHC alleles will present a given peptide by using information learned during training with similar MHC allele peptide sequences. As a result, the pan-allele model is not associated with a particular set of one or more MHC alleles and can predict the probability that a peptide will be presented by any MHC allele. This versatility of the pan-allele model means that a single model can be used to predict the likelihood that any peptide will be presented by any MHC allele. Thus, the use of the pan-allele model reduces the amount of training data required to maximize both individual and population HLA coverage.

[0387] VIII.D.3. Use of pan-allelic models Briefly, when using the pan-allele model to predict the likelihood that a peptide is presented by a single MHC allele, one set of inputs is provided to the pan-allele model, as described in detail below, and the pan-allele model generates a single output. On the other hand, when using the pan-allele model to predict the likelihood that a peptide is presented by multiple MHC alleles, the pan-allele model is used iteratively for each MHC allele of the multiple MHC alleles. Specifically, when using the pan-allele model to predict the likelihood that a peptide is presented by multiple MHC alleles, a first set of inputs associated with a first MHC allele of the multiple MHC alleles is provided to the pan-allele model, and the pan-allele model generates a first output for the first MHC allele. Then, a second set of inputs associated with a second MHC allele of the multiple MHC alleles is provided to the pan-allele model, and the pan-allele model generates a second output for the second MHC allele. This process is performed iteratively for each MHC allele of the multiple MHC alleles. Finally, the outputs generated by the pan-allelic model for each MHC allele of the multiple MHC alleles are combined to generate a single probability that the multiple MHC alleles will present a given peptide.

[0388] VIII.D.4. Overview of pan-allelic models Reference is now made to FIG. 6A, which illustrates an implementation of the pan-allele model portion 460 of a multi-part presentation model, according to one embodiment. As discussed herein, the multi-part presentation model may include multiple instances of the pan-allele model portion 460 (e.g., 10 instances of the pan-allele model portion 460). In various embodiments, the pan-allele model 460 is a neural network. In various embodiments, as shown in FIG. 6A, the pan-allele model 460 includes two layer sets (e.g., a first layer set 630 and a second layer set 640). In general, each of the first layer set 630 and the second layer set 640 may be composed of multiple layers, nodes within the layers, and connections between the nodes with associated parameters. In various embodiments, each of the first layer set 630 and the second layer set 640 includes two or more layers. In various embodiments, the first layer set 630 and the second layer set 640 each include three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more layers. In particular embodiments, the first layer set 630 and the second layer set 640 each are multi-layer perceptrons (MLPs). Generally, the first layer set 630 and the second layer set 640 perform different functions, as described in more detail herein.

[0389] The first layer set 630 receives as input the HLA sequence representation 485, which may be a representation of the sequence of individual HLA alleles expressed by the patient. In various embodiments, the sequence of the individual HLA alleles undergoes encoding (e.g., a one-hot encoding scheme) to generate the HLA sequence representation 485. The first layer set 630 reduces the HLA sequence representation 485 to a new representation with lower dimensions. For example, the new representation with lower dimensions may retain meaningful properties of the HLA sequence representation 485. In various embodiments, to reduce the dimension of the HLA sequence representation 485, the first layer set 630 includes an input layer having a number of nodes that is greater than the number of nodes in the final layer.

[0390] As shown in FIG. 6A, the second tier set 640 receives as input a peptide representation 480 derived from a peptide sequence 410. In various embodiments, the peptide sequence 410 is from an infectious disease-derived peptide. In various embodiments, the peptide sequence 410 is from a human peptide. Additionally, the second tier set 640 may receive as input an output of the first tier set 630. As discussed above, the output of the first tier set 630 is a dimensionality-reduced representation of an HLA allele sequence. The second tier set 640 models the interaction between the infectious disease-derived peptide and the HLA allele sequence to determine whether the HLA sequence is likely to present the peptide sequence. The output of the second tier set 640 is a per-allele likelihood 620 (e.g., the likelihood of presentation of the peptide sequence 410 for each of the six HLA alleles expressed by the patient). As discussed above with reference to FIG. 4B, the allele-by-allele likelihoods of the pan-allelic model 460 may be transformed and combined to generate the final presented likelihoods 440.

[0391] In one embodiment, a pan-allelic model is used to determine the identity of peptide p for allele h. k The presented likelihood u k In some embodiments, the pan-allelic model is represented by the following formula: TIFF2025512989000061.tif8128 (where p k indicates the peptide sequence, and d h denotes the peptide sequence of MHC allele h, f(·) is an arbitrary transformation function, and g H (·) is an arbitrary dependence function. The pan-allelic model uses a shared parameter θ H Based on the set of k and MHC allele peptide sequence d h Generate a dependency score for the shared parameter θ H The values ​​of the set of s are learned during training of the pan-allele model.

[0392] Dependent function gH ([p k d h ];θ H ) is the output of the MHC allele h that is associated with at least one peptide sequence p k Amino acid positions and MHC allele peptide sequences d h Based on the amino acid positions of peptide p k For example, the dependency score for MHC allele h indicates whether MHC allele h is present in the input MHC allele peptide sequence d h Given peptide p k The transformation function f(·) transforms the input, more specifically, in this case g H ([p k d h ];θ H ) is converted to an appropriate value to obtain the dependency score for peptide p k denotes the likelihood of being presented by MHC allele h.

[0393] In one particular implementation mentioned throughout the remainder of this specification, f(·) is a function having a range within [0,1] for the appropriate domain range. In one example, f(·) is an expit function. As another example, f(·) can also be a hyperbolic tangent function, where values ​​in the domain z are equal to or exceed 0. Alternatively, if predictions are made for mass spectrometry ion currents with values ​​outside the range [0,1], f(·) can be any function, such as an identity function, an exponential function, a logarithmic function, etc.

[0394] Therefore, the peptide sequence p k The likelihood of being presented by MHC allele h is governed by the dependence function g H (·) is the peptide sequence p k The encoded version of and MCH allele sequence d hto generate a corresponding dependency score. The dependency score can be transformed by a transformation function f(·) to generate a corresponding dependency score for the peptide sequence p k may generate the likelihood that is presented by MHC allele h.

[0395] VIII.D.5. Dependence Functions for Allelic Interaction Variables In one particular embodiment mentioned throughout this specification, the dependent function g H (·) is an affine function given by: TIFF2025512989000062.tif15158, where α is the intercept, TIFF2025512989000063.tif5128 is peptide p k indicates the residue at position i of hj denotes the residue at position j of MHC allele h, 1[] denotes an indicator variable that is 1 if the condition in brackets is true and 0 otherwise, TIFF2025512989000064.tif5128 is peptide p k is true if the amino acid at position i in is amino acid k, otherwise it is false, and d hj = l is true if the amino acid at position j of MHC allele h is amino acid l, and false otherwise; n pep denotes the length of the modeled peptide, and n MHC indicates the number of MHC residues considered in the modeling, and θ H,ijkl (where x is a coefficient describing the contribution to the likelihood of presentation of having residue k at position i of the peptide and residue l at position j of the MHC allele). This is a linear model of the one-hot encoded peptide sequence and the one-hot encoded MHC allele sequence, with peptide residue-to-MHC residue interactions for all peptide and MHC allele residues.

[0396] In another specific embodiment mentioned throughout this specification, the dependent function g H(·) is the network function given by TIFF2025512989000065.tif6128 A network model NN that has a set of nodes arranged in one or more layers H Each node is represented by a parameter θ H A particular node may be connected to other nodes via connections with associated parameters in a set of activation functions. The value at one particular node may be expressed as the sum of the values ​​of the nodes connected to the particular node, weighted by the associated parameters that are mapped by the activation function associated with the particular node. In contrast to affine functions, network models are advantageous because the presentation model can incorporate nonlinearity and can handle data with amino acid sequences of different lengths. Specifically, through nonlinear modeling, the network model can capture the interactions between amino acids at different positions in a peptide sequence and the interactions between amino acids at different positions in an MHC allele peptide sequence, and how these interactions affect peptide presentation.

[0397] In general, the network model NN H (·) may be structured as a feedforward network, e.g., an artificial neural network (ANN), a convolutional neural network (CNN), a deep neural network (DNN), and / or a recurrent network, e.g., a long short-term memory network (LSTM), a bidirectional recurrent network, a deep bidirectional recurrent network, etc.

[0398] In one example, a single network model NN H (·) is the encoded peptide sequence p k and the encoded protein sequence d of MHC allele h h In such an example, the parameter θ H may correspond to a set of parameters for a single network model, and thus the parameters θH A set of NN may be shared by all MHC alleles. H (·) is the arbitrary input [p k d h ], a single network model NN H As discussed above, such network models are advantageous because the peptide presentation probability of MHC alleles that were unknown in the training data can only be predicted by identifying the protein sequence of the MHC allele.

[0399] FIG. 6B shows an exemplary network model NN shared by MHC alleles. H As shown in Figure 6B, the network model NN H (·) is the peptide sequence p k and protein sequence d of MHC allele h h The input is a dependency score NN corresponding to the MHC allele h. H ([p k d h ]) is output.

[0400] FIG. 6C illustrates an exemplary network model NN H Here, the exemplary network model may represent the second layer set 640 of the pan-allele model 460 shown in FIG. 6A. As shown in FIG. 6C, the network model NN H (·) includes four input nodes at layer l=1, five nodes at layer l=2, two nodes at layer l=3, and one output node at layer l=4. In an alternative embodiment, the network model NN H (·) may contain any number of layers, and each layer may contain any number of nodes. H (·) is the 13 non-zero parameters θ H (1), θ H (2), …, θ H(13). These parameters serve to transform the values ​​that are propagated from node to node through the network model.

[0401] As shown in Figure 6C, the network model NN H The four input nodes at layer l=1 of (·) receive input values ​​including encoded polypeptide sequence data and encoded MHC allele peptide sequence data. The encoded polypeptide sequence data includes amino acid sequences of peptides, and the encoded MHC allele peptide sequence data includes amino acid sequences of MHC alleles that may (or may not) present peptides. In one particular embodiment, the network model NN is fed via the input nodes at layer l=1. H When input to (·), the encoded polypeptide sequence is fed to the network model NN H In the (·) layer, the encoded MHC allele peptide sequence is concatenated in front of it. These input values ​​are then fed into the network model NN according to the values ​​of the parameters. H In some embodiments, the network model NN H The layers of (·) include two fully connected dense network layers. In a further embodiment, a first layer of these two fully connected dense network layers includes 64-128 nodes with rectified linear unit activation functions. In yet a further embodiment, a second layer of these two fully connected dense network layers comprises a single node with a linear output. In such an embodiment, this single node is represented by the network model NN H (·) can be the output node. Finally, the network model NN H (·) is the value NN H ([p k d h ]) which outputs the MHC allele h associated with the peptide sequence p krepresents a dependency score for the MHC allele h, indicating whether the MHC allele h is present. The network function may also include one or more network models, each of which takes a different allele interaction variable (e.g., peptide sequence) as input.

[0402] In yet another example, the dependent function g H (·) can be expressed as: TIFF2025512989000066.tif6128 (where g' H ([p k d h ];θ' H ) is the shared parameter for allele interaction variables, θ, which represents the baseline probability of presentation for any MHC allele. H The bias parameter θ in the set H 0 With the parameter θ' H , which is an affine function involving a set of,network functions, etc.).

[0403] In another embodiment, the bias parameter θ H 0 may be shared according to the gene family of the MHC allele h. That is, the bias parameter θ H 0 θ 遺伝子(h) 0 where gene(h) is the gene family of MHC allele h. For example, class I MHC alleles HLA-A*02:01, HLA-A*02:02, and HLA-A*02:03 may be assigned to the gene family "HLA-A", and the bias parameter θ for each of these MHC alleles may be H 0 As another example, the class II MHC alleles HLA-DRB1:10:01, HLA-DRB1:11:01, and HLA-DRB3:01:01 may be assigned to the gene family "HLA-DRB", and the bias parameter θ for each of these MHC alleles may be H0 As discussed above, the gene family may be one of the allele interaction variables associated with the MHC allele h.

[0404] Returning to equation (23), as an example, the affine dependence function g H (·) to identify peptide p k The likelihood that is presented by MHC allele h can be generated by: TIFF2025512989000067.tif16128 (where α is the intercept, TIFF2025512989000068.tif5128 is peptide p k indicates the residue at position i of hj denotes the residue at position j of MHC allele h, 1[] denotes an indicator variable that is 1 if the condition in brackets is true and 0 otherwise, TIFF2025512989000069.tif5128 is peptide p k is true if the amino acid at position i in is amino acid k, otherwise it is false, and d hj = l is true if the amino acid at position j of MHC allele h is amino acid l, and false otherwise; n pep denotes the length of the modeled peptide, and n MHC indicates the number of MHC residues considered in the modeling, and θ H,ijkl (where x is a coefficient describing the contribution to the likelihood of presentation of having residue k at position i of the peptide and residue l at position j of the MHC allele). This is a linear model of the one-hot encoded peptide sequence and the one-hot encoded MHC allele sequence, with peptide residue-to-MHC residue interactions for all peptide and MHC allele residues.

[0405] As another example, the network transformation function g H (·) to identify peptide p k The likelihood that is presented by MHC allele h can be generated by: TIFF2025512989000070.tif6128 (where p k indicates the peptide sequence, and d h indicates the peptide sequence of MHC allele h, and θ H is a network model NN that associates all MHC alleles H (·) is the set of parameters to be determined.

[0406] FIG. 6D illustrates an exemplary shared network model NN H (·) is used to identify peptide p associated with MHC allele h k As shown in Figure 6D, the shared network model NN H (·) is the peptide sequence p k and MHC allele peptide sequence d h Receives and outputs NN H ([p k d h ]) The output is then mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.

[0407] VIII.D.6. Allelic non-interacting variables As discussed above, allele-free variables include information that influences the presentation of peptides that is independent of the type of MHC allele. For example, allele-free variables may include the N-terminal and C-terminal protein sequences of the peptide, the protein family of the presented peptide, the level of RNA expression of the source gene of the peptide, and any additional allele-free variables.

[0408] In one embodiment, the training module 316 incorporates allele-non-interacting variables into the pan-allele presentation model in a manner similar to that described for the allele-by-allele and multi-allelic models. For example, in some embodiments, the allele-non-interacting variables may be entered as inputs into a dependency function that is separate from the dependency function used for the allele-interacting variables. In such an embodiment, the outputs of the two separate dependency functions may be summed, and the resulting sum may be input into a transformation function to generate a presentation prediction.

[0409] VIII.D.7. Multi-allelic samples As mentioned above, the test sample may contain multiple MHC alleles, not just a single MHC allele. In fact, most samples taken from nature contain more than one MHC allele. For example, each human genome contains six MHC class I loci. Thus, a sample containing a human genome may contain up to six different MHC class I alleles. Thus, a sample containing multiple MHC alleles, not just a single MHC allele, is a typical sample for a real-life test case.

[0410] In embodiments where the test sample includes multiple MHC alleles, the pan-allelic model can be used to determine the probability that a given peptide from the test sample is presented by multiple MHC alleles. However, as briefly described above, when the pan-allelic model is used to predict the likelihood that a peptide is presented by multiple MHC alleles, the pan-allelic model described above is used iteratively for each MHC allele of the multiple MHC alleles. In other words, for each MHC allele of the multiple MHC alleles, the MHC allele peptide sequence and the peptide sequence are independently input into a dependency function shared by all MHC alleles. Based on these inputs, an output corresponding to the MHC allele is generated by the dependency function. This process is performed iteratively for each MHC allele of the multiple MHC alleles. Thus, each MHC allele of the multiple MHC alleles is independently associated with an output of the dependency function. The outputs associated with each MHC allele of the multiple MHC alleles are then combined.

[0411] The outputs of the dependency functions associated with each MHC allele of the plurality of MHC alleles may be combined. The manner in which the multiple outputs of the dependency functions are combined may vary. For example, in some embodiments, the outputs of the iterations of the dependency functions may be summed and the resulting sum may be input to a transformation function to generate the presentation prediction. The equation to obtain such an embodiment may be written as follows: TIFF2025512989000071.tif14153, where T is the total number of unique MHC alleles in the sample including multiple alleles. In an alternative embodiment, each individual output of the iterations of the dependency function may be input to a transformation function, and the resulting outputs from the transformation function may be summed to generate the presented prediction. The equation to obtain this alternative embodiment may be written as follows: TIFF2025512989000072.tif14156 Such embodiments, as well as other embodiments in which multiple outputs of the dependency functions are combined to predict the probability that a peptide is presented in a multi-allelic setting, are described herein.

[0412] VIII.D.8. Training the Pan-Allele Model Training the pan-allele model involves the use of parameters θ associated with the dependence function. H Specifically, the parameter θ H is optimized such that the dependency function can output a dependency score that indicates precisely whether a given MHC allele(s) presents a given peptide sequence.

[0413] Parameter θ H Training data 170 is used to optimize the value of . As mentioned above, the training data 170 used to train the model may include training samples including cells expressing a single MHC allele, training samples including cells expressing multiple MHC alleles, or training samples including cells expressing a combination of both a single MHC allele and multiple MHC alleles. Thus, each data instance i from the training data 170 is input to the pan-allele model, and more specifically, to the dependency function of the pan-allele model. For example, in certain embodiments, MHC allele peptide sequences and peptide sequences may be input to the pan-allele model. The pan-allele model then processes these inputs as if the model were in routine use. However, unlike during the operation of the pan-allele model, during training of the pan-allele model, the known outcomes of peptide presentation are also input to the model. In other words, the label y i is also input to the model. In embodiments where the training sample input to the pan-allelic model includes cells expressing multiple MHC alleles, y i is set to 1 for each allele of multiple MHC alleles in the sample.

[0414] After each iteration of the pan-allele model using data instance i, the model computes the predicted probability of an MHC allele presenting a peptide versus the known label y i Then, to minimize this difference, the pan-allele model uses the parameters θ H In other words, the pan-allele model modifies θ H By minimizing the loss function with respect to the parameters θ H Once the pan-allele model achieves a certain level of predictive accuracy, training is complete and the model is ready for use.

[0415] IX. Prediction Module The prediction module 320 receives sequence data (e.g., infectious disease-derived peptide sequence data and / or HLA sequence data for a patient) and uses a multipartite presentation model to select candidate antigens in the sequence data. Specifically, the sequence data can be DNA sequences, RNA sequences, and / or protein sequences. The prediction module 320 presents the sequence data as a set of multiple peptide sequences having 8-15 amino acids. k For example, the prediction module 320 may process a given sequence "IEFROEIFJEF" into three peptide sequences having the nine amino acids "IEFROEIFJ", "EFROEIFJE", and "FROEIFJEF".

[0416] The presentation module 320 applies one or more of the multipart presentation models to at least the processed peptide sequences to estimate the presentation likelihood of the peptide sequences (e.g., infectious disease-derived peptide sequences). Specifically, the prediction module 320 may select one or more candidate antigen peptide sequences that are likely to be presented on HLA molecules by applying the multipart presentation model to the candidate antigens. In one embodiment, the presentation module 320 selects the candidate antigen sequences that have an estimated presentation likelihood above a predetermined threshold. In another embodiment, the presentation model selects the N candidate antigen sequences with the highest estimated presentation likelihood.

[0417] In various embodiments, therapeutic agents can be provided to a patient to induce an immune response. Exemplary therapeutic agents include vaccine compositions comprising one or more selected antigens, engineered cells expressing TCRs or CARs specific for one or more selected antigens, or antibodies (e.g., bispecific antibodies) that exhibit binding specificity for one or more selected antigens.

[0418] XI. Exemplary Computer Figure 7 illustrates an exemplary computer for implementing the entities illustrated in Figures 1 and 3A. Computer 700 includes at least one processor 702 coupled to a chipset 704. Chipset 704 includes a memory controller hub 720 and an input / output (I / O) controller hub 722. Memory 706 and graphics adapter 712 are coupled to memory controller hub 720, and display 718 is coupled to graphics adapter 712. Storage device 708, input device 714, and network adapter 716 are coupled to I / O controller hub 722. Other embodiments of computer 700 have different architectures.

[0419] The storage device 708 is a non-transitory computer-readable storage medium, such as a hard drive, a compact disc read-only memory (CD-ROM), a DVD, or a solid-state memory device. The memory 706 holds instructions and data used by the processor 702. The input interface 714 is a touch screen interface, a mouse, a trackball, or other type of pointing device, a keyboard, or some combination thereof, and is used to input data into the computer 700. In some embodiments, the computer 700 may be configured to receive input (e.g., commands) from the input interface 714 via gestures from a user. The graphics adapter 712 displays images and other information on the display 718. The network adapter 716 couples the computer 700 to one or more computer networks.

[0420] The computer 700 is adapted to execute computer program modules for providing the functionality described herein. As used herein, the term "module" refers to computer program logic used to provide a specified functionality. Thus, a module may be implemented in hardware, firmware, and / or software. In one embodiment, the program module is stored in the storage device 708, loaded into the memory 706, and executed by the processor 702.

[0421] 1 may vary depending on the embodiment and the processing power required by the entities. For example, presentation specification system 160 may be implemented on a single computer 700 or on multiple computers 700 that communicate with each other over a network, such as in a server farm. Computer 700 may lack some of the components discussed above, such as graphics adapter 712 and display 718. EXAMPLES

[0422] Example 1: Multipart presentation models accurately predict the presentation of viral epitopes Various multipart presentation models were constructed and evaluated for their ability to predict the likely presentation of viral epitopes. Specifically, three viruses were selected, including human immunodeficiency virus (HIV), influenza A virus (IAV), and SARS-CoV-2. These three viruses are well described in the literature. Validated epitopes of T cells were aggregated from multiple sources, including the literature, the Immune Epitope Database (IEDB), as well as virus-specific databases such as the HIV CD8 T cell epitope database from Los Alamos National Laboratory. The Immune Epitope Database (IEDB) is described in further detail in Vita R, et al, The Immune Epitope Database (IEDB): 2018 update. Nucleic Acids Res. 2018 Oct 24, which is incorporated by reference in its entirety. The Los Alamos National Laboratory HIV CD8 T cell epitope database is described in further detail in Llano, A., et al., 2019 Optimal HIV CTL epitopes update: Growing diversity in epitope length and HLA restriction. HIV Molecular Immunology 2019, 3-27, which is incorporated by reference in its entirety.

[0423] The first multipart model, referred to herein as the "ViralEDGE" model, was a 10 pan-specific, 10 allele-specific model ensemble trained on all of the human mass spectrometry data from infectious diseases and binding affinity data from the IEDB. The architecture of the ViralEDGE model is described herein in FIG. 4B. For each selected virus, the ViralEDGE model was further retrained with all epitopes from these three viral taxa removed. Synthetic negatives per-positive were generated in the binding affinity training and validation datasets. Specifically, the retrained ViralEDGE model with no HIV, IAV, and SARS epitopes in its training set and with n=0 synthetic negatives per-positive is referred to herein as the "New Viral EDGE (0 Synth Neg)" model. A retrained ViralEDGE model that is free of HIV, IAV, and SARS epitopes in its training set and includes n=50 synthetic negatives for every positive is referred to herein as the "New Viral EDGE(50 Synth Neg)" model. A retrained ViralEDGE model that is free of HIV, IAV, and SARS epitopes in its training set and includes n=100 synthetic negatives for every positive is referred to herein as the "New Viral EDGE(100 Synth Neg)" model.

[0424] An additional previously published model was further constructed, including a presentation model previously trained and developed to predict the presentation of cancer epitopes (hereinafter referred to as the "EDGE" model). The EDGE model does not include a multipart presentation model and is trained on human immunopeptidemics. Further details of the EDGE model are described in U.S. Pat. No. 10,055,540 (see FIG. 13B showing the prediction of the "MS" model). An additional previously published model was further constructed, referred to as the MHCflurry model. The MHCflurry model incorporates both binding affinity and mass spectrometry data to train the model. Further details of the MHCflurry model are described in O'Donnell, TJ, et al., MHCflurry 2.0: Improved Pan-Allele Prediction of MHC Class I-Presented Peptides by Incorporating Antigen Processing. Cell Systems, 11(1), 42-48.e7, which is incorporated herein by reference in its entirety.

[0425] For each selected evaluated virus, a reference proteome for that virus was generated using the UniProt reference proteome (The UniProt Consortium. (2014). UniProt: a hub for protein information. Nucleic Acids Research, 43(D1), D204-212). Additionally, 8-11mer epitopes were generated for each protein sequence, removing any overlapping epitopes resulting from alternative polyprotein representation. Epitopes were analyzed by various models to predict likely presentation by various class I HLA alleles.

[0426] Each model (e.g., EDGE, ViralEDGE, New Viral EDGE(0 Synth Neg), New Viral EDGE(50 Synth Neg), New Viral EDGE(100 Synth Neg), and MHCFlurry) was deployed to make predictions for each allele. An epitope-HLA prediction was treated as true if that epitope-HLA pair was reported as CD8 positive in any of the collected validation sources, and all other predictions were assumed to be false.

[0427] Since each virus had a different number of proteins, validated epitopes, and HLA alleles with their associated validated epitopes, each benchmark was interpreted as the set of all HLA alleles with at least one validated epitope in that dataset, either allele-by-allele for the top 5 alleles ranked by the number of validated epitopes for that virus, or an aggregate of the top 5, 10, 20, or 25 alleles in total to detect any allele-specific bias in either prediction or validation. Models were evaluated according to precision-recall curves and AUC values, as well as the number of positives in the top 500 predictions ranked by model score.

[0428] Reference is now made to Figures 8A-8C, which show the performance of various models. For each of Figures 8A-8C, the bar plots for each allele show, from left to right, the performance of the MHC Flurry model, the EDGE model, the Viral EDGE model, the New Viral EDGE (50 Synth Neg) model, the New Viral EDGE (0 Synth Neg) model, and the New Viral EDGE (100 Synth Neg) model.

[0429] Specifically, Figure 8A shows the performance of various models for predicting the presentation of HIV epitopes across the top five class I HLA alleles, where the top five class I HLA alleles included HLA-A*02:01, HLA-B*07:02, HLA-A*03:01, HLA-A*11:01, and HLA-B*35:01. In general, the multipart presentation models (ViralEDGE, New Viral EDGE(0 Synth Neg), New Viral EDGE(50 Synth Neg), New Viral EDGE(100 Synth Neg)) achieved a larger area under the precision-recall curve (PR-AUC) across the five class I HLA alleles compared to two previously published models (MHCFlurry and EDGE).

[0430] Figure 8B shows the performance of various models for predicting presentation of influenza A epitopes across the top five alleles. Here, the top five class I HLA alleles included HLA-A*02:01, HLA-A*11:01, HLA-A*03:01, HLA-A*24:02, and HLA-A*68:01. Similar to the results for HIV epitopes, the multipart presentation models (ViralEDGE, New Viral EDGE(0 Synth Neg), New Viral EDGE(50 Synth Neg), New Viral EDGE(100 Synth Neg)) achieved a larger area under the precision-recall curve (PR-AUC) across the five class I HLA alleles compared to two previously published models (MHCFlurry and EDGE).

[0431] Figure 8C shows the performance of various models for predicting presentation of SARS-CoV-2 epitopes across the top five alleles, where the multipart presentation models (ViralEDGE, New Viral EDGE(0 Synth Neg), New Viral EDGE(50 Synth Neg), New Viral EDGE(100 Synth Neg)) achieved a larger area under the precision-recall curve (PR-AUC) compared to two previously published models (MHCFlurry and EDGE), at least for the HLA-A*01:01, HLA-A*03:01, and HLA-B*40:01 alleles. For two of the alleles (HLA-A*02:01 and HLA-A*24:02), the multipart presentation models (ViralEDGE, New Viral EDGE(0 Synth Neg), New Viral EDGE(50 Synth Neg), New Viral EDGE(100 Synth Neg)) achieved comparable PR-AUC values ​​compared to two previously published models (MHCFlurry and EDGE).

[0432] In general, Figures 8A-8C show that the various multipart presentation models generally outperform previously published models in predicting the presentation of different viral epitopes (e.g., HIV, Influenza A, and SARS-CoV-2).

[0433] See further FIG. 9A, which shows precision-recall curves for various models for predicting presentation of HIV epitopes across the top 25 alleles, as well as FIG. 9B, which shows precision-recall curves for various models for predicting presentation of Influenza A epitopes across the top 25 alleles. As shown in both FIG. 9A and 9B, the multipart presentation models (ViralEDGE, New Viral EDGE(0 Synth Neg), New Viral EDGE(50 Synth Neg), New Viral EDGE(100 Synth Neg)) achieved comparable or improved PR-AUC values ​​compared to the previously published MHCFlurry model.

[0434] While the present invention has been particularly shown and described with reference to preferred and various alternative embodiments, it will be understood by those skilled in the relevant art that various changes in form and detail may be made therein without departing from the spirit and scope of the invention.

[0435] All references, issued patents, and patent applications cited within the body of this specification are hereby incorporated by reference in their entirety for all purposes.

Claims

1. Methods for identifying one or more infectious disease-derived antigens likely to be presented by target cells, including the following: To obtain peptide sequences of antigens derived from multiple infectious diseases, To obtain the sequence of one or more MHC alleles of the aforementioned target, The process involves inputting the peptide sequences of the multiple infectious disease-derived antigens and the sequences of the one or more target MHC alleles into a multipart presentation model to generate a set of numerical likelihoods for the presentation of the multiple infectious disease-derived antigens by the one or more MHC alleles expressed on the surface of the target cells, The first part of the multipart presentation model includes a pan-allelic model portion that receives as input the peptide sequences of one or more infectious disease-derived antigens and the sequences or representations thereof of the one or more target MHC alleles, The second part of the multipart presentation model includes a plurality of allele-specific models, each receiving the peptide sequences or representations of the plurality of infectious disease-derived antigens as input. The matter, and To select a subset of the multiple infectious disease-derived antigens based on the set of numerical likelihoods and generate a set of selected antigens.

2. The aforementioned multipart presentation model, At least 1) mass spectrometry data and 2) binding affinity data are determined from multiple samples. The method according to claim 1, comprising a plurality of parameters generated using

3. The aforementioned multipart presentation model, Training peptide sequence and A label derived from mass spectrometry data indicating whether one or more of the aforementioned training peptide sequences were presented by one or more class I MHC alleles present in multiple samples, and The method according to claim 1, comprising a plurality of parameters generated using a training dataset including, optionally, the training peptide sequence having a length in the range of k-mer, where k is 8 to 15, including both ends.

4. The method according to claim 3, wherein the training peptide sequence is identified by mass spectrometry of isolated peptides eluted from MHC alleles present in the plurality of samples.

5. The aforementioned multipart presentation model, Labels derived from binding affinity data indicating whether one or more of the aforementioned training peptide sequences bound to one or more class I MHC alleles present in multiple samples. The method according to claim 3 or 4, comprising a plurality of parameters generated using a training dataset including, optionally, the training peptide sequence being a training peptide sequence derived from an infectious disease.

6. The method according to any one of claims 1 to 3, wherein the pan-allelic model portion includes a neural network.

7. The first set of layers of the neural network in the pan-allele model portion performs dimensionality reduction on the sequence of the one or more MHC alleles of interest, and / or The second set of layers of the neural network in the pan-allele model portion receives as input the representations of the peptide sequences of the plurality of infectious disease-derived antigens and the dimensionality-reduced representations of the sequences of one or more target MHC alleles. The method according to claim 6.

8. The method according to claim 7, wherein the representation of the peptide sequences of the plurality of infectious disease-derived antigens is generated by encoding the peptide sequences via a one-hot encoding scheme.

9. The method according to claim 7, wherein the second set of layers of the neural network models the interaction between the peptide sequences of the plurality of infectious disease-derived antigens and the sequences of the one or more target MHC alleles.

10. One or more of the aforementioned allele-specific models include a neural network, Optionally, the neural network of the allele-specific network receives the representations of the peptide sequences of the multiple infectious disease-derived antigens as input and outputs the presentation likelihood for each allele. The method according to claim 1.

11. The method according to claim 10, wherein the representation of the peptide sequences of the plurality of infectious disease-derived antigens is generated by encoding the peptide sequences via a one-hot encoding scheme.

12. Each of the allele-specific models includes and / or includes a neural network. The second part of the multipart presentation model includes 10 or more allele-specific models. The method according to claim 1.

13. The method according to claim 1, wherein the numerical likelihood of the antigen is a combination of the output of the pan-allelic model portion and the outputs of the plurality of allele-specific models.

14. The aforementioned set of numerical likelihoods is (a) C-terminal sequences adjacent to the peptide sequences of the plurality of infectious disease-derived antigens, and (b) N-terminal sequences adjacent to the peptide sequences of the plurality of infectious disease-derived antigens The method according to claim 1, further specified by a feature including at least one of the following.

15. The aforementioned multiple samples (a) One or more cell lines engineered to express a single MHC class I allele, (b) One or more cell lines engineered to express multiple MHC class I alleles, (c) One or more human cell lines obtained from or derived from multiple patients, (d) Fresh or frozen samples obtained from multiple patients, and (e) Fresh or frozen tissue samples obtained from multiple patients The method according to claim 2, comprising at least one of the following.

16. The method according to claim 1, wherein the target cells include cells infected with one of a pathogen, virus, bacterium, fungus, or parasite.

17. Infectious disease-derived antigens originate from one of the following: pathogens, viruses, bacteria, fungi, or parasites. Optionally, antigens derived from infectious diseases, Severe acute respiratory syndrome-associated coronavirus (SARS), severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), Ebola, HIV, hepatitis B virus (HBV), influenza (influenza), hepatitis C virus (HCV), human papillomavirus (HPV), cytomegalovirus (CMV), chikungunya virus, respiratory syncytial virus (RSV), dengue virus, orthomyxoviridae viruses, tuberculosis, pancoronavirus, herpes simplex virus infection (HSV), influenza (flu), metapneumovirus (MPV), and parainfluenza virus (PIV) Infectious disease organisms selected from the group consisting of, The method according to claim 1.

18. A method for producing a vaccine, comprising performing the method described in any one of claims 1 to 4 and 10 to 17, and further comprising producing or having produced a vaccine comprising the selected set of antigens.

19. A vaccine comprising a set of selected antigens, selected by performing the method described in any one of claims 1 to 4 and 10 to 17.

20. The vaccine according to claim 19, for the treatment of an infectious disease, and optionally administered to the subject prophylactically or therapeutically.