Method and system for predicting HLA epitopes
By using machine learning HLA-peptide presentation prediction model, identifying whether the peptide sequence expressed by cancer cells is presented by proteins encoded by HLA alleles, solving the problem of difficulty in identifying specific HLA-related peptides in the prior art, realizing the potential of personalized cancer treatment.
Patent Information
- Application Number
- CN202380070315.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-08-12
- Filing Date
- 2023-08-11
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has difficulty in effectively identifying and isolating specific HLA Class I or II related peptides, limiting the potential of developing immunotherapeutic agents for cancer or tumors.
Using a computer processor and a trained machine-learning HLA-peptide presentation prediction model, a presentation prediction was generated by inputting amino acid information of candidate peptide sequences expressed by cancer cells, identifying whether the peptide sequence is presented by proteins encoded by the subject's HLA allele.
The efficient identification of specific HLA class I or II-related peptides is achieved, providing potential methods for personalized cancer treatment and improving the efficiency of immunotherapeutic agent development for cancer or tumors.
Smart Images

Figure CN119998887A_ABST
Abstract
Description
[0001] Cross-references
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 397,669, filed on August 12, 2022, which is incorporated herein by reference in its entirety. Background Art
[0003] The major histocompatibility complex (MHC) is a gene complex encoding human leukocyte antigen (HLA) genes. HLA genes are expressed as protein heterodimers displayed on the surface of human cells to circulating T cells. HLA genes are highly polymorphic, allowing them to fine-tune the adaptive immune system. The adaptive immune response depends in part on the ability of T cells to identify and eliminate cells that display disease-associated peptide antigens that bind to human leukocyte antigen (HLA) heterodimers.
[0004] In humans, endogenous and exogenous proteins can be processed into peptides by proteasomes as well as by cytoplasmic and endosomal / lysosomal proteases and peptidases, and presented by two types of cell surface proteins encoded by MHC genes. These cell surface proteins are called human leukocyte antigens (HLA class I and class II), and a group of peptides that bind to them and trigger an immune response are called HLA epitopes. HLA epitopes are key components that enable the immune system to detect danger signals such as pathogen infection and self-transformation. Typically, CD8+T cells recognize MHC class I epitopes displayed on antigen presenting cells (APCs) such as dendritic cells and macrophages, and CD4+T cells recognize class II MHC (HLA-DR, HLA-DQ and HLA-DP) epitopes displayed on APCs. Endogenous processing and presentation of HLA epitopes is a complex process and involves a subset of multiple molecular chaperones and enzymes. HLA peptide presentation can activate cytotoxic T cells and helper T cells, subsequently promoting B cell differentiation and antibody production and CTL responses.
[0005] Understanding the peptide binding preferences of each HLA class I or class II molecule is key to successfully predicting which cancer or tumor-specific antigens are likely to elicit cancer or tumor-specific T cell responses. There is a need for methods of identifying and isolating specific HLA class I or class II associated peptides (e.g., neoantigenic peptides). Such methods and isolated molecules can be used, for example, to develop therapeutic agents, including but not limited to immune-based therapeutic agents. Summary of the invention
[0006] Provided herein is a method for identifying a peptide sequence as presented by at least one of one or more proteins encoded by an HLA allele of a cell of a subject, comprising: (a) using a computer processor to input amino acid sequence information of a set of candidate peptide sequences expressed by cancer cells of a single human subject into a trained machine learning HLA-peptide presentation prediction model to generate a plurality of presentation predictions, wherein each presentation prediction in the plurality of presentation predictions indicates a presentation likelihood of a peptide sequence in the set of candidate peptide sequences presented by an MHC protein of the single human subject; wherein the trained machine learning HLA-peptide presentation prediction model comprises: (i) a plurality of parameters, wherein the plurality of parameters are based on information from expression vectors; The invention relates to a method for preparing a subject's first peptide sequence according to the present invention and comprising: (i) preparing a plurality of training peptide sequences for a subject's first peptide sequence and a plurality of training peptide sequences for a subject's first peptide sequence; (ii) preparing a plurality of training peptide sequences for a subject's first peptide sequence; and (iii) preparing a plurality of training peptide sequences for a subject's first peptide sequence; and (iv) preparing a plurality of training peptide sequences for a subject's first peptide sequence; and (v) preparing a plurality of training peptide sequences for a subject's first peptide sequence; and (vi) preparing a plurality of training peptide sequences for a subject's first peptide sequence; and (vii ...
[0007] Provided herein is a method for selecting a peptide sequence, comprising: (a) using a computer processor to input amino acid sequence information of a set of candidate peptide sequences expressed by cancer cells of a single human subject into a trained machine learning HLA-peptide presentation prediction model to generate a plurality of presentation predictions, wherein each of the plurality of presentation predictions indicates a presentation likelihood of a peptide sequence in the set of candidate peptide sequences being presented by an MHC protein of the single human subject; wherein the trained machine learning HLA-peptide presentation prediction model comprises: (i) a plurality of parameters, wherein the plurality of parameters are based on training data from training cells expressing MHC proteins, wherein the training data comprises a plurality of training peptide sequences and epitope presentation quantitative information, wherein the epitope presentation quantitative information comprises the amount of one or more of the plurality of training peptide sequences presented by an MHC protein; and (ii) a function representing a relationship between the amino acid sequence information received as input and the presentation likelihood generated as output based on the amino acid sequence information and the plurality of parameters; and (b) selecting a subset of peptide sequences in the set of candidate peptide sequences based at least on the plurality of presentation predictions to generate a selected set of peptide sequences.
[0008] Provided herein is a method for treating cancer in a human subject in need thereof, comprising: (a) using a computer processor to input amino acid sequence information of a set of candidate peptide sequences expressed by cancer cells of a single human subject into a trained machine learning HLA-peptide presentation prediction model to generate a plurality of presentation predictions, wherein each presentation prediction in the plurality of presentation predictions indicates a presentation likelihood of a peptide sequence in the set of candidate peptide sequences presented by an MHC protein of the single human subject; wherein the trained machine learning HLA-peptide presentation prediction model comprises: (i) a plurality of parameters, wherein the plurality of parameters are based on training data from training cells expressing MHC proteins, wherein the training data comprises a plurality of training peptide sequences and epitope presentation quantitative information, wherein the epitope presentation quantitative information comprises one or more of the plurality of training peptide sequences presented by the MHC protein; The invention relates to a method for preparing a polypeptide according to the present invention and comprising the steps of: (i) preparing a polypeptide having one or more selected peptide sequences, and (ii) preparing a polypeptide having one or more selected peptide sequences; (iii) preparing a polypeptide having one or more selected peptide sequences; and (iv) preparing a polypeptide having one or more selected peptide sequences. The method comprises the steps of: (a) preparing a polypeptide having one or more selected peptide sequences, and (b) selecting or identifying a subset of peptide sequences in the set of candidate peptide sequences based at least on the plurality of presentation predictions to generate a set of selected or identified peptide sequences; and (c) administering a pharmaceutical composition to the single human subject, the pharmaceutical composition comprising: (i) a polypeptide having one or more selected peptide sequences, (ii) a polynucleotide encoding the polypeptide in (i); (iii) an APC comprising (i) or (ii), or (iv) a T cell comprising a T cell receptor (TCR) specific for an MHC protein of the single human subject complexed with one or more of the peptide sequences selected or identified in (b).
[0009] In some embodiments, the plurality of parameters are based on training data from training cells expressing the MHC proteins of the single human subject.
[0010] In some embodiments, each training peptide sequence in the plurality of training peptide sequences is associated with an MHC protein.
[0011] In some embodiments, the training data includes the identity of the MHC protein associated with each training peptide sequence in the plurality of training peptide sequences.
[0012] In some embodiments, the training data comprises observation of presentation of one or more of the plurality of training peptide sequences by an MHC protein by mass spectrometry.
[0013] In some embodiments, the MHC protein of the single human subject is a class I MHC protein.
[0014] In some embodiments, the plurality of candidate peptide sequences expressed by the cancer cells from the single human subject are identified by comparing the whole genome or whole exome sequence information of the cancer cells from the single human subject with the whole genome or whole exome sequence information of the non-cancerous cells from the single human subject, and identifying nucleic acid sequences that are unique to the cancer cells and not present in the non-cancerous cells.
[0015] In some embodiments, each candidate sequence in the plurality of candidate peptide sequences comprises a cancer-specific mutation.
[0016] In some embodiments, according to the presentation PPV determination method, the trained machine learning HLA-peptide presentation prediction model has a peptide presentation prediction value (PPV) of at least 0.2.
[0017] In some embodiments, the presentation PPV determination method includes inputting amino acid sequence information of multiple test peptide sequences into the trained machine learning HLA-peptide presentation prediction model to generate multiple test presentation predictions, each test presentation prediction indicating the likelihood that the one or more proteins encoded by the HLA alleles are capable of presenting a given test peptide sequence in the multiple test peptide sequences, wherein the multiple test peptide sequences comprise at least 500 test peptide sequences, the sequences comprising: (i) at least one hit peptide sequence identified by mass spectrometry as to be presented by an HLA protein expressed in a cell, and (ii) at least 499 decoy peptide sequences contained in a protein encoded by a genome of an organism, wherein the organism and the subject are the same species.
[0018] In some embodiments, the ratio of at least one hit peptide sequence among the multiple test peptide sequences to the at least 499 bait peptide sequences is 1:499, and based on the trained machine learning HLA-peptide presentation prediction model, the top 0.2% of the multiple test peptide sequences are predicted to be presented by the HLA protein expressed in the cell.
[0019] In some embodiments, (i) the at least one hit peptide sequence comprises at least 10 hit peptide sequences, and (ii) the at least 499 decoy peptide sequences comprise at least 4,990 decoy peptide sequences.
[0020] In some embodiments, the amount of one or more of the plurality of training peptide sequences presented by the MHC protein comprises the number of copies of one or more of the plurality of training peptide sequences presented by the MHC protein.
[0021] In some embodiments, the amount of one or more of the plurality of training peptide sequences presented by the MHC protein comprises the number of copies per cell of one or more of the plurality of training peptide sequences presented by the MHC protein.
[0022] In some embodiments, the amount of one or more of the multiple training peptide sequences presented by the MHC protein comprises the absolute amount, number of molecules, density, concentration, absolute amount per cell, number of molecules per cell, density per cell, or concentration in cells of one or more of the multiple training peptide sequences presented by the MHC protein.
[0023] In some embodiments, the amount of one or more of the plurality of training peptide sequences presented by the MHC protein is based on a large number of mass spectrometry observations, spectral counting, area under the curve (AUC), intensity-based absolute quantification (iBAQ), label-free quantification (LFQ), isotope dilution mass spectrometry, isobaric mass tags, stable isotope labels, and / or mass spectrometry peak intensities.
[0024] In some embodiments, the amount of one or more of the plurality of training peptide sequences presented by the MHC protein is obtained from quantitative mass spectrometry.
[0025] In some embodiments, the epitope presentation quantitative information is obtained from internal standard-parallel reaction monitoring (IS-PRM) mass spectrometry.
[0026] In some embodiments, the epitope presentation quantitative information is obtained from a xenograft sample.
[0027] In some embodiments, the xenograft sample is a patient-derived xenograft (PDX) sample.
[0028] Also provided herein is a method for selecting a peptide sequence, comprising: (a) using a computer processor to input amino acid sequence information of a set of candidate peptide sequences expressed by cancer cells of a single human subject into a trained machine learning HLA-peptide antigen-specific T cell prediction model to generate a plurality of antigen-specific T cell predictions, wherein each antigen-specific T cell prediction in the plurality of antigen-specific T cell predictions indicates a likelihood that an MHC complex comprising an MHC protein of the single human subject and a peptide sequence in the set of candidate peptide sequences will stimulate a T cell that is specific to a peptide sequence in the set of candidate peptide sequences; wherein the trained machine learning HLA-peptide cytotoxic T cell prediction model comprises: (i) a plurality of parameters; , wherein the plurality of parameters are based on training data from training cells expressing MHC proteins, wherein the training data includes a plurality of training peptide sequences and epitope presentation quantitative information, wherein the epitope presentation quantitative information includes the amount of one or more of the plurality of training peptide sequences presented by the MHC protein; and (ii) a function representing the relationship between the amino acid sequence information received as input and the likelihood that a T cell generated as output based on the amino acid sequence information and the plurality of parameters is specific for a peptide sequence in the set of candidate peptide sequences; and (b) selecting a subset of peptide sequences in the set of candidate peptide sequences based at least on the plurality of antigen-specific T cell predictions to generate a set of selected peptide sequences.
[0029] In some embodiments, each antigen-specific T cell prediction in the plurality of antigen-specific T cell predictions indicates the likelihood that an MHC complex comprising an MHC protein of the single human subject and a peptide sequence in the set of candidate peptide sequences will stimulate a T cell that is specific to a neo-antigen peptide sequence in the set of candidate peptide sequences.
[0030] In some embodiments, the function is a function that represents the relationship between the amino acid sequence information received as input and the likelihood that a T cell is specific to a neoantigen peptide sequence in the set of candidate peptide sequences generated as output based on the amino acid sequence information and the multiple parameters.
[0031] In some embodiments, each antigen-specific T cell prediction in the plurality of antigen-specific T cell predictions indicates a likelihood that an MHC complex comprising an MHC protein of the single human subject and a peptide sequence from the set of candidate peptide sequences will stimulate a T cell that is cytotoxic.
[0032] In some embodiments, the function is a function representing the relationship between the amino acid sequence information received as input and the likelihood of being a cytotoxic T cell generated as output based on the amino acid sequence information and the plurality of parameters.
[0033] Incorporation by reference
[0034] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. If the publications and patents or patent applications incorporated by reference conflict with the disclosure contained in this specification, this specification is intended to supersede and / or take precedence over any such conflicting material. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The novel features of the present invention are particularly set forth in the appended claims. The features and advantages of the present invention will be better understood by referring to the following detailed description of illustrative embodiments utilizing the principles of the present invention and the accompanying drawings (also referred to herein as "figures"), in which:
[0036] Figure 1A Data showing the evaluation results using a retained partition of the monoallelic data set are depicted.
[0037] Figure 1B Depicted are data showing predictor performance in ovarian tumors analyzed by MS.
[0038] Figure 1C RECON submission scores are shown.
[0039] Figure 2A A diagram showing an exemplary workflow for preparing patient-derived xenografts targeting MS is depicted.
[0040] Figure 2B A graphical representation of sequence overlap showing nonsynonymous mutations is depicted.
[0041] Figure 3A An exemplary workflow of the method for validating predicted neoantigens using internal standard-triggered parallel reaction monitoring is depicted.
[0042] Figure 3B Data from separate validation of predicted neoantigens by parallel reaction monitoring are shown.
[0043] Figure 4 Shown across epitopes targeted by MS Submit rating.
[0044] Figure 5A An exemplary workflow for quantitatively predicted neoantigens is shown.
[0045] Figure 5B Quantification of predicted neoantigens is shown.
[0046] Figure 5C Quantification of predicted neoantigens is shown.
[0047] Fig. 6A T cell responses to observed neoantigens are shown.
[0048] Figure 6B Data related to immune monitoring are shown.
[0049] Figure 7 Binding affinity related data are shown.
[0050] Figure 8 Clinical outcome data are shown. DETAILED DESCRIPTION
[0051] All terms should be understood as they would be understood by those skilled in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0052] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.
[0053] Although various features of the present disclosure may be described in the context of a single embodiment, these features may also be provided separately or in any suitable combination. Conversely, although for clarity, the present disclosure may be described herein in the context of separate embodiments, the present disclosure may also be implemented in a single embodiment.
[0054] The methods and compositions described herein can be used in a wide range of applications. For example, the methods and compositions described herein can be used to identify immunogenic antigenic peptides, and can be used to develop drugs, such as personalized medicine, and the isolation and characterization of antigen-specific T cells.
[0055] Methods disclosed herein may include generating LC-MS / MS allele data for training allele-specific machine learning methods for epitope prediction. Such methods may include using a set of quality metrics to improve LC-MS / MS data quality to rigorously remove false positives and increase prediction model performance; identifying allele-specific HLA class I or class II binding cores from HLA-ligand panel LC-MS / MS data sets; using machine learning algorithms to improve HLA class I or class II ligand and epitope predictions; and / or identifying biological variables that affect HLA class I or class II-ligand presentation and improve HLA class I or class II epitope predictions, such as gene expression, cleavage, gene preference, cellular localization, and secondary structure.
[0056] Provided herein is a method comprising: (a) processing amino acid information of a plurality of candidate peptide sequences using a machine learning HLA peptide presentation prediction model to generate a plurality of presentation predictions, wherein each candidate peptide sequence in the plurality of candidate peptide sequences is encoded by a genome or exome of a subject, wherein the plurality of presentation predictions comprises an HLA presentation prediction for each of the plurality of candidate peptide sequences, wherein each HLA presentation prediction indicates a likelihood that one or more proteins encoded by a class I or class II HLA allele of a cell of the subject is capable of presenting a given candidate peptide sequence in the plurality of candidate peptide sequences, wherein the machine learning HLA peptide presentation prediction model is trained using training data comprising sequence information of training peptide sequences, the training peptides being identified by mass spectrometry as being presented by an HLA protein expressed in a training cell; and (b) identifying a peptide sequence in the plurality of peptide sequences as presented by at least one of the one or more proteins encoded by a class I or class II HLA allele of a cell of the subject based at least on the plurality of presentation predictions; wherein the machine learning HLA peptide presentation prediction model has a PPV of at least 0.07 according to a presentation positive predictive value (PPV) determination method.
[0057] Provided herein is a method comprising: (a) processing amino acid information of a plurality of peptide sequences encoded by a genome or exome of a subject using a machine learning HLA peptide binding prediction model to generate a plurality of binding predictions, wherein the plurality of binding predictions comprises an HLA binding prediction for each of the plurality of candidate peptide sequences, each binding prediction indicating a likelihood that one or more proteins encoded by a class II HLA allele of a cell of the subject binds to a given candidate peptide sequence in the plurality of candidate peptide sequences, wherein the machine learning HLA peptide binding prediction model is trained using training data comprising sequence information of peptide sequences identified to bind to HLA class I or class II proteins or HLA class I or class II protein analogs; and (b) identifying, based at least on the plurality of binding predictions, a peptide sequence in the plurality of peptide sequences having a probability of binding to at least one of the one or more proteins encoded by a class I or class II HLA allele of a cell of the subject that is greater than a threshold binding prediction probability value; wherein the machine learning HLA peptide binding prediction model has a PPV of at least 0.1 according to a binding positive predictive value (PPV) determination method.
[0058] In some embodiments, the machine learning HLA peptide presentation prediction model is trained using training data comprising sequence information of training peptide sequences that are identified by mass spectrometry as being presented by HLA proteins expressed in training cells.
[0059] In some embodiments, the method comprises ranking at least two peptides presented by at least one of the one or more proteins identified as encoded by a class I or class II HLA allele of cells of the subject based on the presentation prediction.
[0060] In some embodiments, the method includes selecting one or more peptides from the two or more ranked peptides.
[0061] In some embodiments, the method includes selecting one or more peptides from the plurality of peptides that are identified as being presented by at least one of the one or more proteins encoded by a class I or class II HLA allele in cells of the subject.
[0062] In some embodiments, the method includes selecting one or more peptides from two or more peptides ranked based on the presentation prediction.
[0063] In some embodiments, the machine learning HLA peptide presentation prediction model has a positive predictive value (PPV) of at least 0.07 when processing amino acid information of a plurality of test peptide sequences to generate a plurality of test presentation predictions, each test presentation prediction indicating a likelihood that one or more proteins encoded by a class I or class II HLA allele of a cell of the subject is capable of presenting a given test peptide sequence in the plurality of test peptide sequences, wherein the plurality of test peptide sequences comprises at least 500 test peptide sequences, the sequences comprising (i) at least one hit peptide sequence identified by mass spectrometry as to be presented by an HLA protein expressed in a cell, and (ii) at least 499 decoy peptide sequences contained within a protein encoded by a genome of an organism, wherein the organism and the subject are of the same species, wherein the ratio of the at least one hit peptide sequence in the plurality of test peptide sequences to the at least 499 decoy peptide sequences is 1:499, and according to the machine learning HLA peptide presentation prediction model, a top percentage of the plurality of test peptide sequences are predicted to be presented by an HLA protein expressed in a cell.
[0064] In some embodiments, the machine learning HLA peptide presentation prediction model has a positive predictive value (PPV) of at least 0.1 when the amino acid information of a plurality of test peptide sequences is processed to generate a plurality of test binding predictions, each test binding prediction indicating the likelihood that one or more proteins encoded by a class I or class II HLA allele of a cell of the subject will bind to a given test peptide sequence in the plurality of test peptide sequences, wherein the plurality of test peptide sequences comprises at least 20 test peptide sequences comprising (i) at least one test peptide sequence identified by mass spectrometry as being expressed by an HLA protein expressed by a cell; The invention relates to a method of preparing a novel peptide sequence for a human chromatin protein comprising: (i) a hit peptide sequence presented by a human chromatin protein, and (ii) at least 19 decoy peptide sequences contained within a protein comprising at least one peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell, such as a single HLA protein expressed in a cell (e.g., a monoallelic cell), wherein the ratio of the at least one hit peptide sequence in the plurality of test peptide sequences to the at least 19 decoy peptide sequences is 1:19, and according to a machine learning HLA peptide presentation prediction model, a top percentage of the plurality of test peptide sequences are predicted to bind to an HLA protein expressed in a cell.
[0065] In some embodiments, there is no amino acid sequence overlap between the at least one hit peptide sequence and the bait peptide sequence.
[0066] In some embodiments, the machine learning HLA peptide presentation prediction model has a positive predictive value (PPV) of at least 0.08, 0.09, 0.1, 0.11, 0.12, 0.13, 0.14, 0.15, 0.16, 0.17, 0.18, 0.19, 0.2, 0.21, 0.22, 0.23, 0.24, 0.25, 0.26, 0.27, 0.28, 0.29, 0.3, 0.31, 0.32, 0.33, 0.34, 0.35, 0.36, 0.37, 0.38, 0.39, 0.4, 0.41, 0.42, 0.43, 0.44, 0.45, 0.46, 0.47, 0.48, 0.49, 0.5、0.51、0.52、0.53、0.54、0.55、0.56、0.57、0.58、0.59、0.6、0.61、0.62、0.63、0.64、0.65、0.66、0.67、0.68、0.69、0.7、0.71、0.72、0.73、0.74、 0.75, 0.76, 0.77, 0.78, 0.79, 0.8, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.9, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98 or 0.99.
[0067] In some embodiments, the at least one hit peptide sequence comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209,
[0068] In some embodiments, the at least 499 decoy peptide sequences comprise at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 500 0, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 75 00, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10 000、11000、12000、13000、14000、15000、16000、17000、18000、19000、20000、21000、22000、23000、24000、25000、26000、27000、28000、29000、30000、 31000、32000、33000、34000、35000、36000、37000、38000、39000、40000、41000、42000、43000、44000、45000、46000、47000、48000、49000、50000、52500 , 55000, 57500, 60000, 62500, 65000, 67500, 70000, 72500, 75000, 77500, 80000, 82500, 85000, 87500, 90000, 92500, 95000, 97500, 100000, 125000, 1 50000、175000、200000、225000、250000、275000、300000、325000、350000、375000、400000、425000、450000、475000、500000、600000、700000、800000、900,000 or 1,000,000 decoy peptide sequences. One skilled in the art will recognize that changing the hit: decoy ratio changes the PPV.
[0069] In some embodiments, the at least 500 test peptide sequences comprise at least 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300 100、5200、5300、5400、5500、5600、5700、5800、5900、6000、6100、6200、6300、6400、6500、6600、6700、6800、6900、7000、7100、7200、7300、7400、7500、 7600、7700、7800、7900、8000、8100、8200、8300、8400、8500、8600、8700、8800、8900、9000、9100、9200、9300、9400、9500、9600、9700、9800、9900、1000 0, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31 000、32000、33000、34000、35000、36000、37000、38000、39000、40000、41000、42000、43000、44000、45000、46000、47000、48000、49000、50000、52500、 55000、57500、60000、62500、65000、67500、70000、72500、75000、77500、80000、82500、85000、87500、90000、92500、95000、97500、100000、125000、15 0000、175000、200000、225000、250000、275000、300000、325000、350000、375000、400000、425000、450000、475000、500000、600000、700000、800000、900,000 or 1,000,000 test peptide sequences.
[0070] In some embodiments, the top percentage is the top 0.20%, 0.30%, 0.40%, 0.50%, 0.60%, 0.70%, 0.80%, 0.90%, 1.00%, 1.10%, 1.20%, 1.30%, 1.40%, 1.50%, 1.60%, 1.70%, 1.80%, 1.90%, 2.00%, 2.10%, 2.20%, 2.30%, 2.40%, 2.50%, %, 2.60%, 2.70%, 2.80%, 2.90%, 3.00%, 3.10%, 3.20%, 3.30%, 3.40%, 3.50%, 3.60%, 3.70%, 3.80%, 3.90%, 4.00%, 4.10%, 4.20%, 4.30%, 4.40%, 4.50%, 4.60%, 4.70%, 4.80%, 4.90%, 5.00%, 5.10%, 5.20% , 5.30%, 5.40%, 5.50%, 5.60%, 5.70%, 5.80%, 5.90%, 6.00%, 6.10%, 6.20%, 6.30%, 6.40%, 6.50%, 6.60%, 6.70%, 6.80%, 6.90%, 7.00%, 7.10%, 7.20%, 7.30%, 7.40%, 7.50%, 7.60%, 7.70%, 7.80%, 7.90%, 8.00%, 8.10%, 8.20%, 8.30%, 8.40%, 8.50%, 8.60%, 8.70%, 8.80%, 8.90%, 9.00%, 9.10%, 9.20%, 9.30%, 9.40%, 9.50%, 9.60%, 9.70%, 9.80%, 9.90%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19% or 20%.
[0071] In some embodiments, the at least one hit peptide sequence comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209,
[0072] In some embodiments, the at least 19 decoy peptide sequences comprise at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400 , 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900 ,4000,4100,4200,4300,4400,4500,4600,4700,4800,4900,5000,5100,5200,5300,5400,5500,5600,5700,5800,5900,6000,6100,6200,6300,6400 , 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900 , 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22 000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 41000, 42000, 4 3000、44000、45000、46000、47000、48000、49000、50000、52500、55000、57500、60000、62500、65000、67500、70000、72500、75000、77500、80000、82500、85,000, 87,500, 90,000, 92,500, 95,000, 97,500, 100,000, 125,000, 150,000, 175,000, 200,000, 225,000, 250,000, 275,000, 300,000, 325,000, 350,000, 375,000, 400,000, 425,000, 450,000, 475,000, 500,000, 600,000, 700,000, 800,000, 900,000, or 1,000,000 decoy peptide sequences.
[0073] In some embodiments, the at least 20 test peptide sequences comprise at least wherein the at least 500 test peptide sequences comprise at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500,600、700、800、900、1000、1100、1200、1300、1400、1500、1600、1700、1800、1900、2000、2100、2200、2300、2400、2500、2600、2700、2800、2900、3000、31 00、3200、3300、3400、3500、3600、3700、3800、3900、4000、4100、4200、4300、4400、4500、4600、4700、4800、4900、5000、5100、5200、5300、5400、5500、 5600、5700、5800、5900、6000、6100、6200、6300、6400、6500、6600、6700、6800、6900、7000、7100、7200、7300、7400、7500、7600、7700、7800、7900、800 0, 8100, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 140 00, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000、36000、37000、38000、39000、40000、41000、42000、43000、44000、45000、46000、47000、48000、49000、50000、52500、55000、57500、60000、6250 0, 65000, 67500, 70000, 72500, 75000, 77500, 80000, 82500, 85000, 87500, 90000, 92500, 95000, 97500, 100000, 125000, 150000, 175000, 200000, 22 or 1,000,000, 5,000, 2,500,000, 2,750,000, 3,000, 3,250,000, 3,500,000, 3,750,000, 4,000, 4,250,000, 4,500,000, 4,750,000, 5,000, 6,000,000, 7,000,000, 8,000,000, 9,000,000, or 1,000,000 test peptide sequences.
[0074] In some embodiments, the top percentage is the top 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, or 40%.
[0075] In some embodiments, the subject is a single subject.
[0076] In some embodiments, the subject is a mammal.
[0077] In some embodiments, the subject is a human.
[0078] In some embodiments, the training cell is a cell that expresses a single protein encoded by a class I or class II HLA allele of the subject cell.
[0079] In some embodiments, the training cell is a monoallelic HLA cell, or a cell expressing an HLA allele with an affinity tag.
[0080] In some embodiments, the subject's cells include cancer cells.
[0081] In some embodiments, the methods are used to identify peptide sequences.
[0082] In some embodiments, the methods are used to select peptide sequences.
[0083] In some embodiments, the methods are used to prepare for cancer treatment.
[0084] In some embodiments, the methods are used to prepare a subject for a specific cancer treatment.
[0085] In some embodiments, the methods are used to prepare cancer cell specific cancer treatments.
[0086] In some embodiments, each peptide sequence in the plurality of peptide sequences is associated with cancer.
[0087] In some embodiments, at least one peptide sequence in the plurality of peptide sequences is overexpressed by a cancer cell in the subject.
[0088] In some embodiments, each peptide sequence in the plurality of peptide sequences is overexpressed by a cancer cell in the subject.
[0089] In some embodiments, at least one peptide sequence in the plurality of peptide sequences is a cancer cell specific peptide.
[0090] In some embodiments, each peptide sequence in the plurality of peptide sequences is a cancer cell specific peptide.
[0091] In some embodiments, each peptide sequence in the plurality of peptide sequences is expressed by a cancer cell in the subject.
[0092] In some embodiments, at least one peptide sequence in the plurality of peptide sequences is not encoded by a non-cancer cell of the subject.
[0093] In some embodiments, each peptide sequence in the plurality of peptide sequences is not encoded by a non-cancer cell of the subject.
[0094] In some embodiments, at least one peptide sequence in the plurality of peptide sequences is not expressed by non-cancer cells of the subject.
[0095] In some embodiments, each peptide sequence in the plurality of peptide sequences is not expressed by non-cancer cells of the subject.
[0096] In some embodiments, the method includes obtaining the plurality of peptide sequences from a subject.
[0097] In some embodiments, the method includes obtaining a plurality of polynucleotide sequences from a subject.
[0098] In some embodiments, the method includes obtaining a plurality of polynucleotide sequences from a subject, the polynucleotide sequences encoding a plurality of peptide sequences encoded by the genome or exome of the subject or by a pathogen or virus in the subject.
[0099] In some embodiments, the method includes obtaining, by a computer processor, a plurality of polynucleotide sequences from a subject, the polynucleotide sequences encoding a plurality of peptide sequences encoded by the genome or exome of the subject.
[0100] In some embodiments, the method comprises obtaining a plurality of polynucleotide sequences of the subject by genome or exome sequencing.
[0101] In some embodiments, the method comprises obtaining a plurality of polynucleotide sequences of the subject by whole genome sequencing or whole exome sequencing.
[0102] In some embodiments, processing comprises processing with a computer processor.
[0103] In some embodiments, processing includes generating a plurality of predictor variables based at least on amino acid information of the plurality of peptide sequences.
[0104] In some embodiments, the plurality of predictor variables is processed using a machine learning HLA-peptide presentation prediction model.
[0105] In some embodiments, the one or more proteins encoded by a class I or class II HLA allele of a cell of the subject are one or more proteins encoded by a class I or class II HLA allele expressed by the subject.
[0106] In some embodiments, the one or more proteins encoded by the class I or class II HLA alleles of the subject's cells are one or more proteins encoded by the class I or class II HLA alleles expressed by cancer cells of the subject.
[0107] In some embodiments, the one or more proteins encoded by the class I or class II HLA alleles of the subject's cells is a single protein encoded by the class I or class II HLA alleles of the subject's cells.
[0108] In some embodiments, the one or more proteins encoded by the class II HLA alleles of the subject's cells are two, three, four, five, or six or more proteins encoded by class I or class II HLA alleles of the subject's cells.
[0109] In some embodiments, the one or more proteins encoded by a class I or class II HLA allele of a cell of the subject is every protein encoded by a class I or class II HLA allele of a cell of the subject.
[0110] In some embodiments, the method further comprises administering to the subject a composition comprising one or more peptide sequences from the selected subset of peptide sequences.
[0111] In some embodiments, identifying the plurality of peptide sequences comprises comparing a DNA, RNA, or protein sequence from a cancer cell of the subject to a DNA, RNA, or protein sequence from a normal cell of the subject, wherein each of the plurality of peptides comprises at least one mutation that is present in the cancer cell of the subject but not in the normal cell of the subject.
[0112] In some embodiments, the machine learning HLA-peptide presentation prediction model includes a plurality of predictor variables identified based at least on training data, wherein the training data includes training peptide sequence information comprising amino acid position information, wherein the training peptide sequence information is associated with HLA proteins expressed in cells; and a function representing a relationship between the amino acid position information and a presentation likelihood generated as an output based on the amino acid position information and the plurality of predictor variables.
[0113] In some embodiments, identifying includes identifying, based at least on the multiple presentation predictions, a peptide sequence from the multiple peptide sequences whose probability of being presented by at least one of the one or more proteins encoded by a class I or class II HLA allele of a cell of the subject is greater than a threshold presentation prediction probability value.
[0114] In some embodiments, the probability that one or more of 0.2% of the plurality of test peptide sequences predicted by the machine learning HLA peptide presentation prediction model to be presented is presented by at least one of the one or more proteins encoded by a class I or class II HLA allele of a cell of the subject is greater than a threshold presentation prediction probability value.
[0115] In some embodiments, each of 0.2% of the plurality of test peptide sequences predicted by the machine learning HLA peptide presentation prediction model to be presented has a probability of being presented by at least one of the one or more proteins encoded by a class I or class II HLA allele of a cell of the subject that is greater than a threshold presentation prediction probability value.
[0116] In some embodiments, the number of positives is limited to equal the number of hits.
[0117] In some embodiments, the mass spectrometry is monoallelic mass spectrometry.
[0118] In some embodiments, the peptide is presented by an HLA protein expressed in the cell via autophagy.
[0119] In some embodiments, the peptide is presented by an HLA protein expressed in a cell via phagocytosis.
[0120] In some embodiments, the plurality of predictor variables comprises predicted values for expression levels of source proteins comprising the peptides.
[0121] In some embodiments, the plurality of predictor variables comprises a stability prediction value for a source protein comprising the peptide.
[0122] In some embodiments, the plurality of predictor variables comprises a predicted value for a degradation rate of a source protein comprising the peptide.
[0123] In some embodiments, the plurality of predictor variables comprises a predicted value for protein cleavage of a source protein comprising the peptide.
[0124] In some embodiments, the plurality of predictor variables comprises a prediction of a cellular or tissue localization of a source protein comprising the peptide.
[0125] In some embodiments, the plurality of predictor variables comprises a predictor of an intracellular processing pattern of a source protein comprising the peptide, wherein the processing pattern of the source protein comprises, among other things, a predictor of whether the source protein undergoes autophagy, phagocytosis, and intracellular transport.
[0126] In some embodiments, the quality of the training data is improved by using multiple quality metrics.
[0127] In some embodiments, the plurality of quality metrics include common contaminant peptide removal, high scoring peak intensity, high score, and high mass accuracy.
[0128] In some embodiments, the scored peak intensity is at least 50%.
[0129] In some embodiments, the scored peak intensity is at least 60%.
[0130] In some embodiments, the score is at least 7.
[0131] In some embodiments, the mass accuracy is at most 5 ppm.
[0132] In some embodiments, the peptide presented by an HLA protein expressed in a cell is a peptide presented by a single immunoprecipitated HLA protein expressed in the cell.
[0133] In some embodiments, the peptide presented by the HLA protein expressed in the cell is a peptide presented by a single exogenous HLA protein expressed in the cell.
[0134] In some embodiments, the peptide presented by the HLA protein expressed in the cell is a peptide presented by a single recombinant HLA protein expressed in the cell.
[0135] In some embodiments, the plurality of predictor variables includes a peptide-HLA affinity predictor variable.
[0136] In some embodiments, peptides presented by HLA proteins include peptides identified by searching a database of non-enzyme-specific non-modified peptides.
[0137] In some embodiments, peptides presented by HLA proteins include peptides identified by searching a peptide database using an inverse database searching strategy.
[0138] In some embodiments, peptides presented by HLA proteins include peptides identified by comparing the MS / MS spectrum of the HLA-peptide to the MS / MS spectra of one or more peptides or proteins in a peptide or protein database.
[0139] In some embodiments, the mutation is selected from a point mutation, a splice site mutation, a frameshift mutation, a read-through mutation, and a gene fusion mutation.
[0140] In some embodiments, the peptide presented by the HLA protein has a length of 8-12 or 15-40 amino acids.
[0141] In some embodiments, peptides presented by HLA proteins include peptides identified by comparing the MS / MS spectrum of the HLA-peptide with the MS / MS spectra of one or more peptides or proteins in a peptide or protein database to identify peptides presented by HLA proteins.
[0142] In some embodiments, the personalized cancer therapy further comprises an adjuvant.
[0143] In some embodiments, the personalized cancer therapy further comprises an immune checkpoint inhibitor.
[0144] In some embodiments, the training data includes structured data, time series data, unstructured data, relational data, or any combination thereof.
[0145] In some embodiments, the unstructured data includes image data.
[0146] In some embodiments, the relationship data includes data from a client system, an enterprise system, an operating system, a website, a web-accessible application programming interface (API), or any combination thereof.
[0147] In some embodiments, the training data is uploaded to a cloud-based database.
[0148] In some embodiments, the training is performed using a convolutional neural network.
[0149] In some embodiments, the convolutional neural network comprises at least two convolutional layers.
[0150] In some embodiments, the convolutional neural network includes at least one batch normalization step.
[0151] In some embodiments, the convolutional neural network includes at least one spatial dropout step.
[0152] In some embodiments, the convolutional neural network includes at least one global maximum pooling step.
[0153] In some embodiments, the convolutional neural network comprises at least one dense layer.
[0154] In some embodiments, identifying a peptide sequence comprises identifying a peptide sequence having a mutation that is expressed in a cancer cell of the subject.
[0155] In some embodiments, identifying a peptide sequence comprises identifying a peptide sequence that is not expressed in normal cells of the subject.
[0156] In some embodiments, identifying a peptide sequence comprises identifying a viral peptide sequence.
[0157] In some embodiments, identifying a peptide sequence comprises identifying an overexpressed peptide sequence.
[0158] Provided herein is a method for identifying HLA class I or class II specific peptides for immunotherapy of a subject, comprising: obtaining, by a computer processor, a candidate peptide comprising an epitope and a plurality of peptide sequences, each of which comprises the epitope; using a machine learning HLA-peptide presentation prediction model, processing, by a computer processor, amino acid information of the plurality of peptide sequences to generate a presentation prediction for each of the plurality of peptide sequences to immune cells, each presentation prediction indicating the likelihood that one or more proteins encoded by an HLA class I or class II allele can present a given peptide sequence in the plurality of peptide sequences, wherein the machine learning HLA-peptide presentation prediction model is trained using training data, the training data comprising sequence information of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; obtaining a plurality of peptide sequences from the HLA of the subject's cells; Selecting a protein from one or more proteins encoded by a class I or class II allele, the protein being predicted to bind to the candidate peptide by a machine learning HLA-peptide presentation prediction model, wherein the probability of the protein presenting the candidate peptide to immune cells is greater than a threshold presentation prediction probability value; contacting the candidate peptide with the selected protein such that the candidate peptide competes for a placeholder peptide associated with the selected protein; and identifying the candidate peptide as a peptide for immunotherapy that is specific to the selected protein based on whether the candidate peptide replaces the placeholder peptide.
[0159] In some embodiments, obtaining includes identifying a candidate peptide, wherein identifying the candidate peptide includes comparing a DNA, RNA, or protein sequence from a cancer cell of the subject to a DNA, RNA, or protein sequence from a normal cell of the subject.
[0160] In some embodiments, processing includes determining a plurality of predictor variables based at least on amino acid information of the plurality of peptide sequences, and processing the plurality of predictor variables using a machine learning HLA-peptide presentation prediction model.
[0161] In some embodiments, the machine learning HLA-peptide presentation prediction model includes multiple predictor variables identified based on at least training data, wherein the training data includes: training peptide sequence information including amino acid position information, wherein the training peptide sequence information is associated with HLA proteins expressed in cells; and a function representing the relationship between the amino acid position information and the presentation likelihood generated as an output based on the amino acid position information and the multiple predictor variables.
[0162] In some embodiments, the number of positives is limited to equal the number of hits.
[0163] In some embodiments, the mass spectrometry is monoallelic mass spectrometry.
[0164] In some embodiments, the plurality of predictive variables include any one or more of: a predicted value for expression level of a source protein comprising the peptide, a predicted value for stability, a predicted value for degradation rate, a predicted value for cleavage, a predicted value for cell or tissue localization, and a predicted value for intracellular processing patterns including autophagy, phagocytosis, and intracellular transport.
[0165] In some embodiments, the quality of the training data is improved by using multiple quality metrics.
[0166] In some embodiments, the plurality of quality metrics include common contaminant peptide removal, high scoring peak intensity, high score, and high mass accuracy.
[0167] In some embodiments, the scored peak intensity is at least 50%.
[0168] In some embodiments, the scored peak intensity is at least 60%.
[0169] In some embodiments, the placeholder peptide is a CLIP peptide.
[0170] In some embodiments, the placeholder peptide is a CMV peptide.
[0171] In some embodiments, the method further comprises measuring the IC of the placeholder peptide replaced by the target peptide. 50 .
[0172] In some embodiments, the placeholder peptide is replaced by the target peptide. 50 Less than 500nM.
[0173] In some embodiments, the target peptide is further identified by mass spectrometry.
[0174] In some embodiments, the at least one protein encoded by an HLA class I or class II allele of a cell of the subject is a recombinant protein.
[0175] In some embodiments, the at least one protein encoded by an HLA class I or class II allele of a cell of the subject is expressed in a eukaryotic cell.
[0176] In some embodiments, the peptide is presented by an HLA protein expressed in the cell via autophagy.
[0177] In some embodiments, the peptide is presented by an HLA protein expressed in a cell via phagocytosis.
[0178] In some embodiments, the peptide presented by an HLA protein expressed in a cell is a peptide presented by a single immunoprecipitated HLA protein expressed in the cell.
[0179] In some embodiments, the peptide presented by the HLA protein expressed in the cell is a peptide presented by a single exogenous HLA protein expressed in the cell.
[0180] In some embodiments, the peptide presented by the HLA protein expressed in the cell is a peptide presented by a single recombinant HLA protein expressed in the cell.
[0181] In some embodiments, the plurality of predictor variables includes a peptide-HLA affinity predictor variable.
[0182] In some embodiments, peptides presented by HLA proteins include peptides identified by searching a database of non-enzyme-specific non-modified peptides.
[0183] In some embodiments, peptides presented by HLA proteins include peptides identified by searching a peptide database using an inverse database searching strategy.
[0184] In some embodiments, the immunotherapy is cancer immunotherapy.
[0185] In some embodiments, the epitope is a cancer-specific epitope.
[0186] In some embodiments, the identity of the peptide is known.
[0187] In some embodiments, the identity of the peptide is unknown.
[0188] In some embodiments, the identity of the peptide is determined by mass spectrometry.
[0189] In some embodiments, the peptide exchange assay comprises detecting a peptide fluorescent probe or tag.
[0190] In some embodiments, the placeholder peptide is a CLIP peptide. In some embodiments, the placeholder peptide has the amino acid sequence PVSKMRMATPLLMQA (SEQ ID NO: 1).
[0191] In some embodiments, the polynucleic acid construct comprises an expression vector, which further comprises one or more of the following: a promoter, a secretion signal, a dimerization factor, a ribosomal skipping sequence, and one or more tags for purification and / or detection.
[0192] In some embodiments, the placeholder peptide sequence is encoded by a nucleic acid sequence within a vector.
[0193] In some embodiments, the sequence encoding the cleavable domain is placed between the sequence encoding the placeholder peptide and the sequence encoding the HLAβ1 peptide.
[0194] Provided herein is a method for determining the immunogenicity of an MHC class I or class II binding peptide, comprising: selecting a protein encoded by an HLA class I or class II allele that is predicted to bind to an MHC class I or class II binding peptide by a machine learning HLA-peptide presentation prediction model, wherein the machine learning HLA-peptide presentation prediction model is configured to generate a presentation prediction for a given peptide sequence, the presentation prediction indicating the likelihood that one or more proteins encoded by the HLA class II allele can present the given peptide sequence, and wherein the probability that the protein presents the MHC class I or class II binding peptide is greater than a threshold presentation prediction probability value; contacting the peptide with the selected protein such that the peptide competes for a placeholder peptide associated with the selected protein and replaces the placeholder peptide, thereby forming a complex comprising the HLA class I or class II protein and the MHC class I or class II binding peptide; contacting the complex with a CD4+T cell, and determining one or more activation parameters of the CD4+T cell, the parameters being selected from: induction of cytokines, induction of chemokines, and expression of cell surface markers.
[0195] Provided herein is a method for inducing CD4+ T cell activation in a subject for cancer immunotherapy, the method comprising: identifying a peptide sequence associated with cancer and comprising a cancer mutation, wherein identifying the peptide sequence comprises comparing a DNA, RNA or protein sequence from a cancer cell of the subject with a DNA, RNA or protein sequence from a normal cell of the subject; selecting a protein encoded by an HLA class I or HLA class II allele, the protein being normally expressed by cells of the subject and predicted to bind to the peptide by a machine learning HLA-peptide presentation prediction model; wherein the prediction model has a positive predictive value of at least 0.1 at a recall rate of at least 0.1%, 0.1%-50%, or at most 50%, and wherein the probability of the protein presenting the identified peptide sequence is greater than a threshold presentation prediction probability value; contacting the identified peptide with the selected protein encoded by the HLA class I or HLA class II allele to verify whether the identified peptide competes with a placeholder peptide associated with the selected protein encoded by the HLA class I or HLA class II allele, thereby binding to the peptide with an IC value of less than 500 nM. 50 A value is used to replace the placeholder peptide; optionally, purifying the identified peptide; and administering to the subject an effective amount of a polypeptide comprising the sequence of the identified peptide or a polynucleotide encoding the polypeptide.
[0196] Provided herein is a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, comprising: obtaining, by a computer processor, a plurality of peptide sequences of the polypeptide sequence; processing, by a computer processor, the amino acid information of the plurality of peptide sequences using a machine learning HLA-peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences, each presentation prediction indicating a likelihood that one or more proteins encoded by a class I or class II MHC allele of a cell of the subject can present an epitope sequence of a given peptide sequence in the plurality of peptide sequences, wherein the machine learning HLA-peptide presentation prediction model is trained using training data comprising sequence information associated with HLA proteins expressed in cells; determining or predicting, based on the plurality of presentation predictions, that each of the plurality of peptide sequences of the polypeptide sequence is not immunogenic to the subject; and administering to the subject a composition comprising the drug.
[0197] Provided herein is a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, comprising: obtaining, by a computer processor, a plurality of peptide sequences of the polypeptide sequence; processing, by a computer processor, amino acid information of the plurality of peptide sequences using a machine learning HLA-peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences, each presentation prediction indicating a likelihood that one or more proteins encoded by a class I or class II MHC allele of a cell of the subject can present an epitope sequence of a given peptide sequence in the plurality of peptide sequences, wherein the machine learning HLA-peptide presentation prediction model is trained using training data, the training data comprising sequence information of sequences of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; and determining or predicting, based on the plurality of presentation predictions, that at least one of the plurality of peptide sequences of the polypeptide sequence is immunogenic to the subject.
[0198] Provided herein is a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: using a computer processor to input amino acid information of a peptide sequence of the polypeptide sequence into a machine learning HLA-peptide presentation prediction model to generate a set of presentation predictions for the peptide sequence, each presentation prediction representing a probability that an epitope sequence of a given peptide sequence is presented by one or more proteins encoded by a class I or class II MHC allele of a cell of the subject; wherein the machine learning HLA-peptide presentation prediction model comprises: a plurality of predictor variables identified at least based on training data; wherein the training data comprises: sequence information of sequences of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; training peptide sequence information comprising amino acid position information, wherein the training peptide sequence information is associated with HLA proteins expressed in cells; and a function representing a relationship between the amino acid position information received as input and the presentation likelihood generated as output based on the amino acid position information and the predictor variables; based on the set of presentation predictions, determining or predicting that each of the peptide sequences of the polypeptide sequence is not immunogenic to the subject; and administering a composition comprising the drug to the subject.
[0199] Provided herein is a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: using a computer processor to input amino acid information of a peptide sequence of the polypeptide sequence into a machine learning HLA-peptide presentation prediction model to generate a set of presentation predictions for the peptide sequence, each presentation prediction representing the probability that an epitope sequence of a given peptide sequence is presented by one or more proteins encoded by a class I or class II MHC allele of a cell of the subject; wherein the machine learning HLA-peptide presentation prediction model comprises: a plurality of predictor variables identified at least based on training data; wherein the training data comprises: sequence information of sequences of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; training peptide sequence information comprising amino acid position information, wherein the training peptide sequence information is associated with HLA proteins expressed in cells; and a function representing a relationship between the amino acid position information received as input and the presentation likelihood generated as output based on the amino acid position information and the predictor variables; based on the set of presentation predictions, determining or predicting that at least one of the peptide sequences of the polypeptide sequence is immunogenic to the subject.
[0200] Provided herein is a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, comprising: obtaining, by a computer processor, a plurality of peptide sequences of the polypeptide sequence; processing, by a computer processor, the amino acid information of the plurality of peptide sequences using a machine learning HLA-peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences, each presentation prediction indicating a likelihood that one or more proteins encoded by a class I or class II MHC allele of a cell of the subject can present an epitope sequence of a given peptide sequence in the plurality of peptide sequences, wherein the machine learning HLA-peptide presentation prediction model is trained using training data comprising sequence information associated with HLA proteins expressed in cells; determining or predicting, based on the plurality of presentation predictions, that each of the plurality of peptide sequences of the polypeptide sequence is not immunogenic to the subject; and administering to the subject a composition comprising the drug.
[0201] In some embodiments, the method further comprises deciding not to administer the drug to the subject.
[0202] In some embodiments, the medicament comprises an antibody or a binding fragment thereof.
[0203] In some embodiments, the peptide sequence of the polypeptide sequence has a length of 8, 9, 10, 11 or 12 amino acids, and wherein the protein encoded by a class I or class II MHC allele of a cell of the subject is a protein encoded by a class I MHC allele of a cell of the subject.
[0204] In some embodiments, the peptide sequence of the polypeptide sequence has a length of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids, and wherein the protein encoded by a class I or class II MHC allele of a cell of the subject is a protein encoded by a class II MHC allele of a cell of the subject.
[0205] Provided herein is a method of treating a subject having an autoimmune disease or condition, comprising: (a) identifying or predicting an epitope of an expressed protein presented by class I or class II MHC of a cell of the subject, wherein a complex comprising the identified or predicted epitope and class I or class II MHC is targeted by a CD8 or CD4 T cell of the subject; (b) identifying a T cell receptor (TCR) that binds to the complex; (c) expressing the TCR in a regulatory T cell or an allogeneic regulatory T cell from the subject; and (d) administering to the subject a regulatory T cell expressing the TCR.
[0206] Provided herein is a method of treating a subject having an autoimmune disease or condition, comprising administering to the subject a regulatory T cell expressing a T cell receptor (TCR) that binds to a complex comprising: (i) an epitope of an expressed protein identified or predicted to be presented by class I or class II MHC of a cell of the subject, and (ii) class I or class II MHC, wherein the complex is targeted by a CD8 or CD4 T cell of the subject.
[0207] Provided herein is a computer system for identifying peptide sequences for personalized cancer treatment of a subject, comprising: a database configured to store a plurality of peptide sequences of a subject; and one or more computer processors operably coupled to the database, wherein the one or more computer processors are individually and collectively programmed to: process amino acid information of the plurality of peptide sequences using a machine learning HLA-peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences, each presentation prediction indicating a likelihood that one or more proteins encoded by a class I or class II MHC allele of a cell of the subject is capable of presenting a given peptide sequence in the plurality of peptide sequences, wherein the machine learning HLA-peptide presentation prediction model is trained using training data, the training data comprising sequence information of sequences of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; and select a subset of the plurality of peptide sequences for personalized cancer treatment of the subject based at least on the plurality of presentation predictions.
[0208] Provided herein is a computer system for identifying HLA class I or HLA class II specific peptides for immunotherapy of a subject, comprising: a database configured to store candidate peptides comprising an epitope and a plurality of peptide sequences, each peptide sequence comprising the epitope; and one or more computer processors operably coupled to the database, wherein the one or more computer processors are individually and collectively programmed to: process amino acid information of the plurality of peptide sequences with a machine learning HLA-peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences to immune cells, each presentation prediction indicating a likelihood that one or more proteins encoded by an HLA class I or HLA class II allele is capable of presenting a given peptide sequence in the plurality of peptide sequences, wherein the machine learning HLA-peptide presentation prediction model is trained using training data, the training data comprising sequence information of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; and Selecting a protein from one or more proteins encoded by class II alleles, wherein the protein is predicted to bind to the candidate peptide by a machine learning HLA-peptide presentation prediction model, wherein the probability of the protein presenting the candidate peptide to immune cells is greater than a threshold presentation prediction probability value; and identifying the candidate peptide as a peptide for immunotherapy specific to the selected protein based on whether the candidate peptide replaces the placeholder peptide when the candidate peptide contacts the selected protein so that the candidate peptide competes for the placeholder peptide associated with the selected protein.
[0209] Provided herein is a computer system for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, comprising: a database configured to store a plurality of peptide sequences of the polypeptide sequence; and one or more computer processors operably coupled to the database, wherein the one or more computer processors are individually and collectively programmed to: process the amino acid information of the plurality of peptide sequences using a machine learning HLA-peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences, each presentation prediction indicating a likelihood that one or more proteins encoded by a class I or class II MHC allele of a cell of the subject is capable of presenting an epitope sequence of a given peptide sequence in the plurality of peptide sequences, wherein the machine learning HLA-peptide presentation prediction model is trained using training data, the training data comprising sequence information associated with HLA proteins expressed in cells; and based on the plurality of presentation predictions, determine or predict that each of the plurality of peptide sequences of the polypeptide sequence is not immunogenic to the subject, wherein a composition comprising the drug is administered to the subject.
[0210] Provided herein is a computer system for screening drugs comprising polypeptide sequences for immunogenicity in a subject, comprising: a database configured to store a plurality of peptide sequences of the polypeptide sequences; and one or more computer processors operably coupled to the database, wherein the one or more computer processors are individually and collectively programmed to: process amino acid information of the plurality of peptide sequences using a machine learning HLA-peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences, each presentation prediction indicating a likelihood that one or more proteins encoded by a class I or class II MHC allele of a cell of the subject is capable of presenting an epitope sequence of a given peptide sequence in the plurality of peptide sequences, wherein the machine learning HLA-peptide presentation prediction model is trained using training data, the training data comprising sequence information of sequences of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; and based on the plurality of presentation predictions, determine or predict that at least one of the plurality of peptide sequences of the polypeptide sequence is immunogenic to the subject.
[0211] Provided herein is a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements a method for identifying peptide sequences for personalized cancer treatment of a subject, the method comprising: obtaining a plurality of peptide sequences for the subject; processing amino acid information of the plurality of peptide sequences using a machine learning HLA-peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences, each presentation prediction indicating a likelihood that one or more proteins encoded by a class I or class II MHC allele of a cell of the subject is capable of presenting a given peptide sequence in the plurality of peptide sequences, wherein the machine learning HLA-peptide presentation prediction model is trained using training data comprising sequence information of sequences of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; and selecting a subset of the plurality of peptide sequences for personalized cancer treatment of the subject based at least on the plurality of presentation predictions.
[0212] Provided herein is a non-transitory computer-readable medium comprising machine executable code, which when executed by one or more computer processors implements a method for identifying HLA class II specific peptides for immunotherapy of a subject, the method comprising: obtaining a candidate peptide comprising an epitope and a plurality of peptide sequences, each peptide sequence comprising the epitope; processing amino acid information of the plurality of peptide sequences by a machine learning HLA-peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences to immune cells, each presentation prediction indicating a likelihood that one or more proteins encoded by an HLA class I or HLA class II allele is capable of presenting a given peptide sequence in the plurality of peptide sequences, wherein the machine learning HLA-peptide presentation prediction model is trained using training data, the training data comprising sequence information of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; obtaining amino acid information of the plurality of peptide sequences from the HLA of the subject's cells; Selecting a protein from one or more proteins encoded by class II alleles, wherein the protein is predicted to bind to the candidate peptide by a machine learning HLA-peptide presentation prediction model, wherein the probability of the protein presenting the candidate peptide to immune cells is greater than a threshold presentation prediction probability value; and identifying the candidate peptide as a peptide for immunotherapy specific to the selected protein based on whether the candidate peptide replaces the placeholder peptide when the candidate peptide contacts the selected protein so that the candidate peptide competes with the placeholder peptide.
[0213] Provided herein is a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: obtaining a plurality of peptide sequences of the polypeptide sequence; processing amino acid information of the plurality of peptide sequences using a machine learning HLA-peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences, each presentation prediction indicating a likelihood that one or more proteins encoded by a class I or class II MHC allele of a cell of the subject is capable of presenting an epitope sequence of a given peptide sequence in the plurality of peptide sequences, wherein the machine learning HLA-peptide presentation prediction model is trained using training data comprising sequence information associated with HLA proteins expressed in cells; and determining or predicting, based on the plurality of presentation predictions, that each of the plurality of peptide sequences of the polypeptide sequence is not immunogenic to the subject, wherein a composition comprising the drug is administered to the subject.
[0214] Provided herein is a non-transitory computer-readable medium comprising machine-executable code, which when executed by one or more computer processors implements a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: obtaining a plurality of peptide sequences of the polypeptide sequence; processing amino acid information of the plurality of peptide sequences using a machine learning HLA-peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences, each presentation prediction indicating a likelihood that one or more proteins encoded by a class I or class II MHC allele of a cell of the subject can present an epitope sequence of a given peptide sequence in the plurality of peptide sequences, wherein the machine learning HLA-peptide presentation prediction model is trained using training data, the training data comprising sequence information of sequences of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; and determining or predicting, based on the plurality of presentation predictions, that at least one of the plurality of peptide sequences of the polypeptide sequence is immunogenic to the subject.
[0215] Provided herein is a method comprising: processing amino acid information of a plurality of candidate peptide sequences using a machine learning HLA peptide presentation prediction model to generate a plurality of presentation predictions, wherein each candidate peptide sequence in the plurality of candidate peptide sequences is encoded by a genome or exome of a subject, wherein the plurality of presentation predictions comprises an HLA presentation prediction for each of the plurality of candidate peptide sequences, wherein each presentation prediction indicates a likelihood that one or more proteins encoded by an HLA class I or HLA class II allele of a cell of the subject is capable of presenting a given candidate peptide sequence in the plurality of candidate peptide sequences, wherein the machine learning HLA peptide presentation prediction model is trained using training data comprising sequence information of training peptide sequences, the training peptides being identified by mass spectrometry as to be presented by an HLA protein expressed in the training cells; and identifying, based at least on the plurality of presentation predictions, a peptide sequence in the plurality of peptide sequences that is presented by an HLA class I or HLA class II allele of a cell of the subject. wherein when the amino acid information of the plurality of test peptide sequences is processed to generate a plurality of test presentation predictions, each test presentation prediction indicates that the subject's HLA class I or HLA class II allele is expressed by the subject's cells; The machine learning HLA peptide presentation prediction model has a positive predictive value (PPV) of at least 0.07 when evaluating the likelihood that one or more proteins encoded by a class II allele are capable of presenting a given test peptide sequence in the multiple test peptide sequences, wherein the multiple test peptide sequences comprise at least 500 test peptide sequences, the sequences comprising (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell, and (ii) at least 499 bait peptide sequences contained in a protein encoded by a genome of an organism, wherein the organism and the subject are of the same species, wherein the ratio of the at least one hit peptide sequence in the multiple test peptide sequences to the at least 499 bait peptide sequences is 1:499, and according to the machine learning HLA peptide presentation prediction model, 0.2% of the multiple test peptide sequences are predicted to be presented by an HLA protein expressed in a cell.
[0216] Provided herein is a method comprising: processing amino acid information of a plurality of peptide sequences encoded by a genome or exome of a subject using a machine learning HLA peptide binding prediction model to generate a plurality of binding predictions, wherein the plurality of binding predictions comprises an HLA binding prediction for each of the plurality of candidate peptide sequences, each binding prediction indicating a likelihood that one or more proteins encoded by HLA class I or HLA class II of a cell of the subject binds to a given candidate peptide sequence in the plurality of candidate peptide sequences, wherein the machine learning HLA peptide binding prediction model is trained using training data, the training data comprising sequence information of peptide sequences identified as binding to HLA class I or HLA class II proteins or HLA class I or HLA class II protein analogs; and identifying, based at least on the plurality of binding predictions, a peptide sequence in the plurality of peptide sequences whose probability of binding to at least one of the one or more proteins encoded by an HLA class I or HLA class II allele of a cell of the subject is greater than a threshold binding prediction probability value; wherein when the amino acid information of a plurality of test peptide sequences is processed to generate a plurality of test binding predictions, each test binding prediction indicating a likelihood that one or more proteins encoded by an HLA class I or HLA class II of a cell of the subject binds to a given candidate peptide sequence in the plurality of candidate peptide sequences. When the machine learning HLA peptide binding prediction model is used to predict the likelihood that one or more proteins encoded by class II will bind to a given test peptide sequence in the multiple test peptide sequences, the machine learning HLA peptide binding prediction model has a positive predictive value (PPV) of at least 0.1, wherein the multiple test peptide sequences comprise at least 50 test peptide sequences, the sequences comprising (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell, and (ii) at least 19 bait peptide sequences contained in a protein, the protein comprising a peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell, wherein the organism and the subject are the same species, wherein the ratio of the at least one hit peptide sequence in the multiple test peptide sequences to the at least 19 bait peptide sequences is 1:19, and according to the machine learning HLA peptide presentation prediction model, 5% of the multiple test peptide sequences are predicted to bind to the HLA protein expressed in the cell.
[0217] In some embodiments, the machine learning HLA peptide presentation prediction model is trained using training data comprising sequence information of training peptide sequences that are identified by mass spectrometry as being presented by HLA proteins expressed in training cells.
[0218] In some embodiments, one or more of 0.2% of the plurality of test peptide sequences predicted to be presented by the machine learning HLA peptide presentation prediction model have a probability of being presented by at least one of the one or more proteins encoded by an HLA class I or HLA class II allele of a cell of the subject that is greater than a threshold presentation prediction probability value.
[0219] In some embodiments, each of 0.2% of the plurality of test peptide sequences predicted to be presented by a machine learning HLA peptide presentation prediction model has a probability of being presented by at least one of the one or more proteins encoded by an HLA class I or HLA class II allele of a cell of the subject that is greater than a threshold presentation prediction probability value.
[0220] Provided herein is a method for preparing a personalized cancer treatment, the method comprising: identifying a peptide sequence, wherein the peptide sequence is associated with cancer, wherein the identifying comprises comparing a DNA, RNA, or protein sequence from a cancer cell of a subject with a DNA, RNA, or protein sequence from a normal cell of the subject; inputting amino acid position information of the identified peptide sequence into a machine learning HLA-peptide presentation prediction model using a computer processor to generate a set of presentation predictions for the identified peptide sequence, each presentation prediction representing a peptide sequence presented by an HLA class I or HLA class II peptide of a cell of the subject; The method comprises the steps of: determining a probability that one or more proteins encoded by a class II allele present a given peptide sequence in the identified peptide sequences; wherein the machine learning HLA-peptide presentation prediction model comprises: a plurality of predictor variables identified at least based on training data; wherein the training data comprises: sequence information of sequences of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; training peptide sequence information comprising amino acid position information, wherein the training peptide sequence information is associated with HLA proteins expressed in cells; and a function representing a relationship between amino acid position information received as input and a presentation probability generated as output based on the amino acid position information and the predictor variables; and selecting a subset of the identified peptide sequences for preparing personalized cancer treatment based on the set of presentation predictions; wherein the prediction model has a positive predictive value of at least 0.1 at a recall rate of at least 0.1%, 0.1%-50%, or at most 50%.
[0221] Provided herein is a method comprising training a machine learning HLA-peptide presentation prediction model, wherein the training comprises inputting into the HLA-peptide presentation prediction model a sequence of amino acid position information of HLA-peptides separated from one or more HLA-peptide complexes from cells expressing HLA class II alleles using a computer processor; the machine learning HLA-peptide presentation prediction model comprises: a plurality of predictor variables identified based at least on training data, the training data comprising: sequence information of sequences of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; training peptide sequence information comprising amino acid position information of training peptides, wherein the training peptide sequence information is associated with HLA proteins expressed in cells; and a function representing a relationship between amino acid position information received as input and presentation likelihood generated as output based on the amino acid position information and the predictor variables.
[0222] In some embodiments, the presentation model has a positive predictive value of at least 0.25 at a recall of at least 0.1%, 0.1%-50%, or at most 50%.
[0223] In some embodiments, the presentation model has a positive predictive value of at least 0.4 at a recall of at least 0.1%, 0.1%-50%, or at most 50%.
[0224] In some embodiments, the presentation model has a positive predictive value of at least 0.6 at a recall of at least 0.1%, 0.1%-50%, or at most 50%.
[0225] In some embodiments, the mass spectrometry is monoallelic mass spectrometry.
[0226] In some embodiments, the peptide is presented by an HLA protein expressed in the cell via autophagy.
[0227] In some embodiments, the peptide is presented by an HLA protein expressed in a cell via phagocytosis.
[0228] In some embodiments, the quality of the training data is improved by using multiple quality metrics.
[0229] In some embodiments, the plurality of quality metrics include common contaminant peptide removal, high scoring peak intensity, high score, and high mass accuracy.
[0230] In some embodiments, the scored peak intensity is at least 50%.
[0231] In some embodiments, the scored peak intensity is at least 60%.
[0232] In some embodiments, the score is at least 7.
[0233] In some embodiments, the mass accuracy is at most 5 ppm.
[0234] In some embodiments, the mass accuracy is at most 2 ppm.
[0235] In some embodiments, the backbone cleavage score is at least 5.
[0236] In some embodiments, the backbone cleavage score is at least 8.
[0237] In some embodiments, the peptide presented by an HLA protein expressed in a cell is a peptide presented by a single immunoprecipitated HLA protein expressed in the cell.
[0238] In some embodiments, the peptide presented by the HLA protein expressed in the cell is a peptide presented by a single exogenous HLA protein expressed in the cell.
[0239] In some embodiments, the peptide presented by the HLA protein expressed in the cell is a peptide presented by a single recombinant HLA protein expressed in the cell.
[0240] In some embodiments, the plurality of predictor variables includes a peptide-HLA affinity predictor variable.
[0241] In some embodiments, the plurality of predictor variables includes a source protein expression level predictor variable.
[0242] In some embodiments, the plurality of predictor variables includes a peptide cleavage predictor variable.
[0243] In some embodiments, the training peptide sequence information includes sequences of peptides presented by HLA proteins, including peptides identified by searching a non-enzyme-specific non-modified peptide database. In some embodiments, peptides presented by HLA proteins include peptides identified by searching a de novo peptide sequencing tool.
[0244] In some embodiments, peptides presented by HLA proteins include peptides identified by searching a peptide database using an inverse database searching strategy.
[0245] In some embodiments, the HLA protein includes HLA-DR and HLA-DP or HLA-DQ protein. In some embodiments, the HLA protein includes an HLA-DR protein selected from HLA-DR and HLA-DP or HLA-DQ protein. In some embodiments, the HLA protein includes an HLA-DR protein selected from the group consisting of HLA-DPB1*01:01 / HLA-DPA1*01:03, HLA-DPB1*02:01 / HLA-DPA1*01:03, HLA-DPB1*03:01 / HLA-DPA1*01:03, HLA-DPB1*04:01 / HLA-DPA1*01:03, HLA-DPB1*04:02 / HLA-DPA1*01:03, HLA-DPB1*06:01 / HLA-DPA1*01:03, HLA- A-DQB1*02:01 / HLA-DQA1*05:01, HLA-DQB1*02:02 / HLA-DQA1*02:01, HLA-DQB1*06:02 / HLA-DQA1*01:02, HLA-DQB1*06:04 / HLA- DQA1*01:02, HLA-DRB1*01:01, HLA-DRB1*01:02, HLA-DRB1*03:01, HLA-DRB1*03:02, HLA-DRB1*04:01, HLA-DRB1*04:02, HLA-DR B1*04:03, HLA-DRB1*04:04, HLA-DRB1*04:05, HLA-DRB1*04:07, HLA-DRB1*07:01, HLA-DRB1*08:01, HLA-DRB1*08:02, HLA-DRB1 *08:03, HLA-DRB1*08:04, HLA-DRB1*09:01, HLA-DRB1*10:01, HLA-DRB1*11:01, HLA-DRB1*11:02, HLA-DRB1*11:04, HLA-DRB1*1 2:01, HLA-DRB1*12:02, HLA-DRB1*13:01, HLA-DRB1*13:02, HLA-DRB1*13:03, HLA-DRB1*14:01, HLA-DRB1*15:01, HLA-DRB1*15: 02, HLA-DRB1*15:03, HLA-DRB1*16:01, HLA-DRB3*01:01, HLA-DRB3*02:02, HLA-DRB3*03:01, HLA-DRB4*01:01 and HLA-DRB5*01:01.
[0246] In some embodiments, peptides presented by HLA proteins include peptides identified by comparing the MS / MS spectrum of the HLA-peptide to the MS / MS spectra of one or more HLA-peptides in a peptide database.
[0247] In some embodiments, the mutation is selected from a point mutation, a splice site mutation, a frameshift mutation, a read-through mutation, and a gene fusion mutation.
[0248] In some embodiments, the peptide presented by the HLA protein has a length of 8-12 or 15-40 amino acids.
[0249] In some embodiments, peptides presented by HLA proteins include peptides identified by: (a) isolating one or more HLA complexes from a cell line expressing a single HLA class I or HLA class II allele; (b) isolating one or more HLA-peptides from the one or more isolated HLA complexes; (c) obtaining MS / MS spectra of the one or more isolated HLA-peptides; and (d) obtaining peptide sequences corresponding to the MS / MS spectra of the one or more isolated HLA-peptides from a peptide database; wherein the sequences of the one or more isolated HLA-peptides are identified from the one or more sequences obtained in step (d).
[0250] In some embodiments, the personalized cancer therapy further comprises an adjuvant.
[0251] In some embodiments, the personalized cancer therapy further comprises an immune checkpoint inhibitor.
[0252] In some embodiments, the training data includes structured data, time series data, unstructured data, relational data, or any combination thereof.
[0253] In some embodiments, the unstructured data includes image data.
[0254] In some embodiments, the relationship data includes data from a client system, an enterprise system, an operating system, a website, a web-accessible application programming interface (API), or any combination thereof.
[0255] In some embodiments, the training data is uploaded to a cloud-based database.
[0256] In some embodiments, the training is performed using a convolutional neural network.
[0257] In some embodiments, the convolutional neural network comprises at least two convolutional layers.
[0258] In some embodiments, the convolutional neural network (CNN) includes at least one batch normalization step.
[0259] In some embodiments, the convolutional neural network includes at least one spatial dropout step.
[0260] In some embodiments, the convolutional neural network includes at least one global maximum pooling step.
[0261] In some embodiments, the convolutional neural network comprises at least one dense layer.
[0262] In some embodiments, identifying a peptide sequence comprises identifying a peptide sequence having a mutation that is expressed in a cancer cell of the subject.
[0263] In some embodiments, identifying a peptide sequence comprises identifying a peptide sequence that is not expressed in normal cells of the subject.
[0264] In some embodiments, identifying a peptide sequence comprises identifying an overexpressed peptide sequence.
[0265] In some embodiments, identifying a peptide sequence comprises identifying a viral peptide sequence. In one aspect, the present invention provides a method for identifying an HLA class I or HLA class II specific peptide for immunotherapy that is specific to a subject, the method comprising: identifying a candidate peptide comprising an epitope; inputting amino acid information of a plurality of peptide sequences (each comprising an epitope) into a machine learning HLA-peptide presentation prediction model using a computer processor to generate a set of HLA presentation predictions about the peptide sequence to immune cells, each presentation prediction representing a probability that a given peptide sequence comprising the epitope is presented by one or more proteins encoded by an HLA class I or HLA class II allele of a subject's cell; wherein the prediction model has a positive predictive value of at least 0.1 at a recall rate of at least 0.1%, 0.1%-50%, or at most 50%, selecting a protein from one or more proteins encoded by an HLA class I or HLA class II allele of a subject's cell, the protein being predicted by the prediction model to bind to the candidate peptide, wherein the probability that the protein presents the candidate peptide to an immune cell is greater than a threshold presentation prediction probability value; contacting the candidate peptide with a protein encoded by an HLA class I or HLA class II allele such that the candidate peptide competes with the HLA A placeholder peptide associated with a protein encoded by a class I or class II HLA allele; and identifying the candidate peptide as a peptide for immunotherapy that is specific to a protein encoded by a class II HLA allele based on whether the candidate peptide replaces the placeholder peptide.
[0266] In some embodiments, the immunotherapy is cancer immunotherapy.
[0267] In some embodiments, identifying comprises comparing a DNA, RNA, or protein sequence from a cancer cell of the subject to a DNA, RNA, or protein sequence from a normal cell of the subject.In some embodiments, the epitope is a cancer-specific epitope.
[0268] In some embodiments, the placeholder peptide is a CLIP peptide. In some embodiments, the placeholder peptide is a CMV peptide. In some embodiments, the method further comprises measuring the IC of the placeholder peptide replaced by the target peptide. 50 In some embodiments, the placeholder peptide is replaced by the target peptide 50 Less than 500 nM. In some embodiments, the target peptide is further identified by mass spectrometry. In some embodiments, the at least one protein encoded by the HLA class I or HLA class II allele of the subject's cells is a recombinant protein. In some embodiments, the at least one protein encoded by the HLA class I or HLA class II allele of the subject's cells is expressed in a eukaryotic cell.
[0269] In one aspect, provided herein is an assay method for verifying the specificity of a candidate peptide binding to an HLA class I or HLA class II protein, the method comprising: expressing in a eukaryotic cell a polynucleic acid construct comprising a nucleic acid sequence encoding an HLA class I or HLA class II protein, the protein comprising an α chain and a β chain or a portion thereof, capable of binding to a peptide comprising an MHC binding epitope, and wherein the expressed HLA class I or HLA class II protein or the portion thereof remains associated with a placeholder peptide; isolating the HLA class I or HLA class II protein or the portion thereof expressed in the eukaryotic cell; performing a peptide exchange assay by: (a) adding increasing amounts of a candidate peptide to determine whether the candidate peptide displaces the placeholder peptide associated with the HLA class I or HLA class II protein or the portion thereof; and (b) calculating the IC of the substitution reaction. 50 to determine the affinity of the candidate peptide to the HLA class I or HLA class II protein or a portion thereof relative to the placeholder peptide, thereby verifying the specificity of the candidate peptide in binding to the HLA class I or HLA class II protein.
[0270] In some embodiments, the identity of the peptide is known. In some embodiments, the identity of the peptide is unknown. In some embodiments, the identity of the peptide is determined by mass spectrometry.
[0271] In some embodiments, the peptide exchange assay comprises detecting a peptide fluorescent probe or tag. In some embodiments, the placeholder peptide is a CLIP peptide.
[0272] In some embodiments, the polynucleic acid construct comprises an expression vector, which further comprises one or more of the following: a promoter, a linker, one or more protease cleavage sites, a secretion signal, a dimerization factor, a ribosomal skipping sequence, one or more tags for purification and or detection.
[0273] In one aspect, the present invention provides a method for determining the immunogenicity of an MHC class II binding peptide, the method comprising: selecting a protein encoded by an HLA class II allele that is predicted to bind to a peptide by a machine learning HLA-peptide presentation prediction model; wherein the prediction model has a positive predictive value of at least 0.1 at a recall rate of at least 0.1%, 0.1%-50%, or at most 50%, and wherein the probability of the protein presenting the identified peptide sequence is greater than a threshold presentation prediction probability value; contacting the peptide with the selected protein encoded by the HLA class II allele such that the peptide competes with and replaces the placeholder peptide associated with the selected protein encoded by the HLA class II allele, thereby forming a complex comprising the HLA class II protein and the identified peptide; contacting the complex of the HLA class II protein and the identified peptide with CD4+T cells, and determining one or more activation parameters of the CD4+T cells, the parameters being selected from the induction of cytokines, the induction of chemokines, and the expression of cell surface markers.
[0274] In one aspect, the present invention provides a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: using a computer processor to input amino acid information of a peptide sequence of the polypeptide sequence into a machine learning HLA-peptide presentation prediction model to generate a set of presentation predictions for the peptide sequence, each presentation prediction representing the probability that an epitope sequence of a given peptide sequence is presented by one or more proteins encoded by an HLA class I or class II allele of a cell of the subject; wherein the machine learning HLA-peptide presentation prediction model comprises: a plurality of predictor variables identified based at least on training data; wherein the training data comprises: sequence information of sequences of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; training peptide sequence information comprising amino acid position information, wherein the training peptide sequence information is associated with HLA proteins expressed in cells; and a function representing a relationship between the amino acid position information received as input and the presentation likelihood generated as output based on the amino acid position information and the predictor variables; (b) determining or predicting, based on the set of presentation predictions, that each of the peptide sequences of the polypeptide sequence is not immunogenic to the subject; and (c) administering a composition comprising the drug to the subject.
[0275] In one aspect, the present invention provides a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: (a) using a computer processor to input amino acid information of a peptide sequence of the polypeptide sequence into a machine learning HLA-peptide presentation prediction model to generate a set of presentation predictions for the peptide sequence, each presentation prediction representing the probability that an epitope sequence of a given peptide sequence is presented by one or more proteins encoded by an HLA class I or class II allele of a cell of the subject; wherein the machine learning HLA-peptide presentation prediction model comprises: a plurality of predictor variables identified at least based on training data; wherein the training data comprises: sequence information of sequences of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; training peptide sequence information comprising amino acid position information, wherein the training peptide sequence information is associated with HLA proteins expressed in cells; and a function representing a relationship between the amino acid position information received as input and the presentation likelihood generated as output based on the amino acid position information and the predictor variables; (b) based on the set of presentation predictions, determining or predicting that at least one of the peptide sequences of the polypeptide sequence is immunogenic to the subject.
[0276] In one embodiment, the method further comprises deciding not to administer the drug to the subject.
[0277] In one embodiment, the medicament comprises an antibody or a binding fragment thereof.
[0278] In one embodiment, the peptide sequence of the polypeptide sequence comprises each contiguous peptide sequence of the polypeptide sequence, which has a length of 8, 9, 10, 11 or 12 amino acids, and wherein the protein encoded by an HLA class I or class II allele of a cell of the subject is a protein encoded by an HLA class I allele of a cell of the subject.
[0279] In one embodiment, the peptide sequence of the polypeptide sequence comprises each contiguous peptide sequence of the polypeptide sequence, which has a length of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 amino acids, and wherein the protein encoded by an HLA class I or class II allele of a subject's cell is a protein encoded by an MHC class II allele of a subject's cell.
[0280] In one aspect, provided herein is a method of treating a subject having an autoimmune disease or condition, comprising: (a) identifying or predicting an epitope of an expressed protein presented by HLA class I or class II of cells of the subject, wherein a complex comprising the identified or predicted epitope and HLA class I or class II is targeted by CD8 or CD4 T cells of the subject; (b) identifying a T cell receptor (TCR) that binds to the complex; (c) expressing the TCR in regulatory T cells or allogeneic regulatory T cells from the subject; and (d) administering regulatory T cells expressing the TCR to the subject.
[0281] In one embodiment, the autoimmune disease or condition is diabetes.
[0282] In one embodiment, the cells are pancreatic islet cells.
[0283] In one aspect, provided herein is a method of treating a subject having an autoimmune disease or condition, comprising administering to the subject a regulatory T cell expressing a T cell receptor (TCR) that binds to a complex comprising: (i) an epitope of an expressed protein identified or predicted to be presented by HLA class I or class II of a cell of the subject, and (ii) HLA class I or class II, wherein the complex is targeted by a CD8 or CD4 T cell of the subject.
[0284] Other aspects and advantages of the present disclosure will become apparent to those skilled in the art based on the following detailed description, which merely shows and describes illustrative embodiments of the present disclosure. It should be appreciated that the present disclosure is capable of other different embodiments, and that its several details are capable of modification in various obvious respects, all without departing from the present disclosure. Therefore, the drawings and description are to be regarded as illustrative in nature, and not restrictive.
[0285] In one aspect, the present invention provides a method for treating cancer in a subject, the method comprising: identifying a peptide sequence, wherein the peptide sequence is associated with cancer, wherein the identification comprises comparing a DNA, RNA or protein sequence from a cancer cell of the subject with a DNA, RNA or protein sequence from a normal cell of the subject; using a computer processor, inputting amino acid information of the identified peptide sequence into a machine learning HLA-peptide presentation prediction model to generate a set of presentation predictions for the identified peptide sequence, each presentation prediction representing a peptide sequence presented by an HLA class I or HLA class II peptide of a cell of the subject; The method comprises the steps of: determining a probability that one or more proteins encoded by a class II allele present a given sequence of the identified peptide sequence; wherein the machine learning HLA-peptide presentation prediction model comprises: a plurality of predictor variables identified at least based on training data, wherein the training data comprises: sequence information of sequences of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; training peptide sequence information comprising amino acid position information, wherein the training peptide sequence information is associated with HLA proteins expressed in cells; and a function representing a relationship between amino acid position information received as input and a presentation probability generated as output based on the amino acid position information and the predictor variables; and selecting a subset of the identified peptide sequences for preparing personalized cancer treatment based on the set of presentation predictions; and administering a composition comprising one or more of the peptides to the subject, wherein the prediction model has a positive predictive value of at least 0.1 at a recall of at least 0.1%, 0.1%-50%, or at most 50%.
[0286] In some embodiments, the machine learning HLA-peptide presentation prediction model comprises sequence information of the sequences of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry following reverse phase offline fractionation.
[0287] In some embodiments, the prediction model exhibits an improvement of 1.1x to 100x compared to NetMHCIIpan or NetMHCI. In some embodiments, the prediction model exhibits an improvement of 1.1, 2, 3, 4, 5, 6, 7, 7.4, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65 , 39, 40, 41, 42, 43, 44, 45, 50, 55, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 8, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 times or more.
[0288] In one aspect, the present disclosure provides methods for predicting peptides that can accurately pair or bind to specific HLA class I or class II molecules, such that high fidelity binding of the peptide to HLA class I or class II proteins (including α and β chain heterodimers) ensures that the specific peptide is presented to T lymphocytes, thereby eliciting a specific immune response and avoiding any cross-reaction or immune promiscuity. Several recent studies have shown that CD8+ or CD4+ T cells can also recognize ligands presented by HLA class I or class II and help control tumors. Ideally, cancer vaccines and other immunotherapies would utilize guided CD8+ or CD4+ T cell responses, but current efforts have completely abandoned HLA class I or class II antigen predictions because of the insufficient accuracy of current prediction tools.
[0289] In one aspect, the present disclosure provides methods for predicting peptides that can accurately bind to specific HLA class I or class II proteins, so that when the peptide is therapeutically administered to a subject expressing a specific homologous HLA class I or class II protein, the peptide can be used to activate a more sustained and potent immune response by virtue of the ability of the HLA class I or class II protein to activate CD8+ or CD4+ T cells and stimulate immune memory. In some embodiments, the methods provided herein show improvements in predictions of specific HLA class I or class II proteins relative to currently available predictors. In some embodiments, the methods provided herein show improvements to at least about 1.1 times in predictions of specific HLA class I or class II proteins relative to currently available predictors. In some embodiments, the methods provided herein show improvements to at least about 2 times in predictions of specific HLA class I or class II proteins relative to currently available predictors. In some embodiments, the methods provided herein show improvements to at least about 3 times in predictions of specific HLA class I or class II proteins relative to currently available predictors. In some embodiments, the methods provided herein show an improvement of at least about 4 times in the prediction of a specific HLA class I or class II protein relative to currently available predictors. In some embodiments, the methods provided herein show an improvement of at least about 5 times in the prediction of a specific HLA class I or class II protein relative to currently available predictors. In some embodiments, the methods provided herein show an improvement of at least about 6 times in the prediction of a specific HLA class I or class II protein relative to currently available predictors. In some embodiments, the methods provided herein show an improvement of at least about 7 times in the prediction of a specific HLA class I or class II protein relative to currently available predictors. In some embodiments, the methods provided herein show an improvement of at least about 8 times in the prediction of a specific HLA class I or class II protein relative to currently available predictors. In some embodiments, the methods provided herein show an improvement of at least about 9 times in the prediction of a specific HLA class I or class II protein relative to currently available predictors. In some embodiments, the methods provided herein show an improvement of at least about 10 times in the prediction of a specific HLA class I or class II protein relative to currently available predictors. In some embodiments, the methods provided herein show an improvement of at least about 15-fold in the prediction of a particular HLA class I or class II protein relative to currently available predictors. In some embodiments, the methods provided herein show an improvement of at least about 20-fold in the prediction of a particular HLA class I or class II protein relative to currently available predictors. In some embodiments, the methods provided herein show an improvement of at least about 30-fold in the prediction of a particular HLA class I or class II protein relative to currently available predictors.In some embodiments, the methods provided herein show an improvement of at least about 40-fold in the prediction of a particular HLA class I or class II protein relative to currently available predictors. In some embodiments, the methods provided herein show an improvement of at least about 50-fold in the prediction of a particular HLA class I or class II protein relative to currently available predictors. In some embodiments, the methods provided herein show an improvement of at least about 60-fold in the prediction of a particular HLA class I or class II protein relative to currently available predictors.
[0290] On the one hand, immunotherapy methods customized or personalized for a specific subject are proposed herein. Each subject or patient expresses a specific set of HLA class I and HLA class II proteins. HLA typing is a well-known technique that allows determination of a specific repertoire of HLA proteins expressed by a subject. Once the HLA heterodimers expressed by a specific subject are known, having an improved, sophisticated and reliable method as described herein for predicting with high fidelity peptides that can bind to a specific HLA class I or class II molecule or complex can ensure that a specific immune response customized specifically for the subject can be generated.
[0291] In this application, unless otherwise specifically stated, the use of the singular includes the plural. It must be noted that, as used in this specification, the singular forms "a", "an" and "the" include plural referents unless the context clearly indicates otherwise. In this application, the use of "or" means "and / or" unless otherwise stated. In addition, the use of the term "includes" and other forms such as "comprises", "contains" and "has" is not restrictive. The term "one or more" or "at least one", such as one or more or at least one member of a group of members, is clear in itself, and by further illustration, the term particularly includes reference to any one of the members, or any two or more of the members, for example, any ≥3, ≥4, ≥5, ≥6 or ≥7 of the members, etc., up to all the members.
[0292] References in this specification to "some embodiments," "embodiments," "one embodiment," or "other embodiments" mean that a feature, structure, or characteristic described in connection with the embodiment is included in at least some embodiments of the present disclosure, but not necessarily in all embodiments.
[0293] As used in this specification and claims, the words "comprising" (and any form of comprising), "having" (and any form of having), "including" (and any form of including), or "containing" (and any form of containing) are inclusive or open-ended and do not exclude other unrecited elements or method steps. It is contemplated that any embodiment discussed in this specification can be implemented using any method or composition of the present disclosure, and vice versa. In addition, the compositions of the present disclosure can be used to implement the methods of the present disclosure.
[0294] As used herein, the terms "about" or "approximately" when referring to measurable values such as parameters, amounts, durations, etc., are intended to encompass variations of + / -20% or less, + / -10% or less, + / -5% or less, or + / -1% or less of the specified value, as long as such variations are suitable for use in the present disclosure. It should be understood that the value to which the modifier "about" or "approximately" refers is itself also specifically disclosed.
[0295] The term "immune response" includes T cell-mediated and / or B cell-mediated immune responses that are affected by T cell costimulation regulation. Exemplary immune responses include T cell responses, such as the production of cytokines and the cytotoxicity of cells. In addition, the term immune response includes immune responses that are indirectly affected by T cell activation, such as antibody production (humoral response) and cytokine responsive cells such as the activation of macrophages.
[0296] "Receptor" should be understood to refer to a biological molecule or grouping of molecules that can bind to a ligand. Receptors can be used to transmit information in cells, cell formations or organisms. Receptors contain at least one receptor unit and may contain two or more receptor units, wherein each receptor unit may be composed of a protein molecule, such as a glycoprotein molecule. The receptor has a structure that is complementary to the structure of the ligand and can be complexed with the ligand as a binding partner. Signal information can be transmitted by conformational changes after the receptor binds to the ligand on the cell surface. According to the present disclosure, a receptor may refer to MHC class I and class II proteins that can form a receptor / ligand complex with a ligand, such as a peptide or peptide fragment of appropriate length. Class I and class II MHC peptides encoded by HLA class I and class II alleles are generally referred to herein as HLA class I and HLA class II peptides, or HLA class I and HLA class II peptides, or HLA class I and class II proteins, or HLA class I and class II proteins, or HLA class I and class II molecules, or such common variants thereof, as known to those of ordinary skill in the art in the context of the discussion.
[0297] A "ligand" is a molecule that is able to form a complex with a receptor. According to the present disclosure, a ligand is understood to refer to a peptide or peptide fragment having, for example, a suitable length and a suitable binding motif in its amino acid sequence, such that the peptide or peptide fragment is able to bind to and form a complex with an MHC class I or MHC class II protein (i.e., HLA class I and HLA class II protein).
[0298] "Antigen" is a molecule that can stimulate an immune response and can be produced by cancer cells or infectious agents or autoimmune diseases. Antigens recognized by T cells (whether helper T lymphocytes (T helper (TH) cells) or cytotoxic T lymphocytes (CTL)) are not recognized as complete proteins, but as small peptides associated with HLA class I or class II proteins on the cell surface. In the process of naturally occurring immune responses, antigens recognized by association with HLA class II molecules on antigen presenting cells (APCs) are obtained from outside the cell, internalized, and processed into small peptides associated with HLA class II molecules. APCs can also cross-present peptide antigens by processing exogenous antigens and presenting the processed antigens to HLA class I molecules. Antigens that produce peptides recognized by association with HLA class I MHC molecules are usually peptides produced intracellularly, and these antigens are processed and associated with class I MHC molecules. It is now understood that peptides that associate with a given HLA class I or class II molecule are characterized as having a common binding motif, and binding motifs have been determined for a large number of different HLA class I and class II molecules. Synthetic peptides that correspond to the amino acid sequence of a given antigen and contain a binding motif for a given HLA class I or class II molecule can also be synthesized. These peptides can then be added to appropriate APCs, and the APCs can be used to stimulate a T helper cell or CTL response in vitro or in vivo. The binding motifs, methods of synthesizing peptides, and methods of stimulating a T helper cell or CTL response are all known to those of ordinary skill in the art and are readily available.
[0299] In this specification, the term "peptide" is used interchangeably with "mutant peptide" and "neoantigenic peptide". Similarly, in this specification, the term "polypeptide" is used interchangeably with "mutant polypeptide" and "neoantigenic polypeptide". "Neoantigen" or "neo-epitope" refers to a class of tumor antigens or tumor epitopes produced by tumor-specific mutations in expressed proteins. The present disclosure further includes peptides comprising tumor-specific mutations, peptides comprising known tumor-specific mutations, and mutant polypeptides or fragments thereof identified by the methods of the present disclosure. These peptides and polypeptides are referred to herein as "neoantigenic peptides" or "neoantigenic polypeptides". These polypeptides or peptides may have a variety of lengths, may be in their neutral (uncharged) form, may be in salt form, and may not contain modifications such as glycosylation, side chain oxidation, phosphorylation, or any post-translational modification, or may contain these modifications, provided that the modification does not destroy the biological activity of the polypeptide described herein. In some embodiments, the neoantigenic peptides of the present disclosure may include: for HLA class I, a length of 22 or less residues, for example, about 8 to about 22 residues, about 8 to about 15 residues, or 9 or 10 residues; for HLA class II, a length of 40 or less residues, for example, a length of about 8 to about 40 residues, a length of about 8 to about 24 residues, about 12 to about 19 residues, or about 14 to about 18 residues. In some embodiments, the neoantigenic peptide or neoantigenic polypeptide comprises a neoepitope.
[0300] The term "epitope" includes any protein determinant that can specifically bind to an antibody, antibody peptide and / or antibody-like molecule (including but not limited to T cell receptor) as defined herein. Epitope determinants are usually composed of chemically active surface groups of molecules such as amino acids or sugar side chains, and usually have specific three-dimensional structural characteristics and specific charge characteristics.
[0301] A "T cell epitope" is a peptide sequence which can be bound by an MHC molecule of class I or II in the form of a peptide-presenting MHC molecule or MHC complex and which is then recognized and bound in this form by cytotoxic T lymphocytes or T helper cells, respectively.
[0302] The term "antibody" as used herein includes IgG (including IgG1, IgG2, IgG3 and IgG4), IgA (including IgA1 and IgA2), IgD, IgE, IgM and IgY, and is intended to include complete antibodies, including single-chain complete antibodies, and antigen-binding (Fab) fragments thereof. Antigen-binding antibody fragments include, but are not limited to, Fab, Fab' and F(ab')2, Fd (consisting of VH and CH1), single-chain variable fragments (scFv), single-chain antibodies, disulfide-linked variable fragments (dsFv) and fragments comprising VL or VH domains. Antibodies can be from any animal source. Antigen-binding antibody fragments, including single-chain antibodies, can contain variable regions alone or in combination with all or part of the following: hinge region, CH1, CH2 and CH3 domains. Any combination of variable regions and hinge regions, CH1, CH2 and CH3 domains is also included. The antibodies can be, for example, monoclonal antibodies, polyclonal antibodies, chimeric antibodies, humanized antibodies, and human monoclonal and polyclonal antibodies that specifically bind to HLA-associated polypeptides or HLA-HLA binding peptide (HLA-peptide) complexes. Those skilled in the art will recognize that a variety of immunoaffinity techniques are suitable for enriching soluble proteins, such as soluble HLA-peptide complexes or membrane-bound HLA-associated polypeptides, for example, which have been cleaved from a membrane by proteolysis. This includes techniques in which (1) one or more antibodies capable of specifically binding to soluble proteins are fixed to a fixed or removable substrate (e.g., a plastic well or resin, latex or paramagnetic beads), and (2) a solution containing soluble proteins from a biological sample is passed through the antibody-coated substrate, thereby binding the soluble proteins to the antibodies. The substrate with the antibodies and bound soluble proteins is separated from the solution, and the antibodies and soluble proteins are optionally dissociated, for example, by changing the pH and / or ionic strength and / or ionic composition of the solution in which the antibodies are bathed. Alternatively, immunoprecipitation techniques can be used, in which antibodies and soluble proteins are combined and form macromolecular aggregates. The macromolecular aggregates can be separated from the solution by size exclusion techniques or by centrifugation.
[0303] The term "immunopurification (IP)" (or immunoaffinity purification or immunoprecipitation) is a method well known in the art and is widely used to separate desired antigens from samples. Typically, the method includes contacting a sample containing the desired antigen with an affinity matrix, which comprises antibodies against the antigen covalently attached to a solid phase. The antigen in the sample is bound to the affinity matrix through an immunochemical bond. The affinity matrix is then washed to remove any unbound material. The antigen is removed from the affinity matrix by changing the chemical composition of the solution in contact with the affinity matrix. Immunopurification can be performed on a column containing an affinity matrix, in which case the solution is an eluent. Alternatively, immunopurification can be a batch process, in which case the affinity matrix is maintained as a suspension in a solution. An important step in the process is to remove the antigen from the matrix. This is typically achieved by increasing the ionic strength of the solution in contact with the affinity matrix, such as by adding an inorganic salt. Changes in pH can also effectively dissociate the immunochemical bonds between the antigen and the affinity matrix.
[0304] An "agent" is any small molecule compound, antibody, nucleic acid molecule or polypeptide or fragment thereof.
[0305] A "change" or "variation" is an increase or decrease. The change can be as little as 1%, 2%, 3%, 4%, 5%, 10%, 20%, 30%, or as much as 40%, 50%, 60%, or even as much as 70%, 75%, 80%, 90%, or 100%.
[0306] A "biological sample" is any tissue, cell, fluid or other substance derived from an organism. As used herein, the term "sample" includes biological samples, such as any tissue, cell, fluid or other substance derived from an organism. "Specific binding" refers to a compound (e.g., a peptide) that recognizes and binds to a molecule (e.g., a polypeptide) but does not substantially recognize and bind to other molecules in a sample (e.g., a biological sample).
[0307] "Capture reagent" refers to a reagent that specifically binds to a molecule (eg, a nucleic acid molecule or a polypeptide) to select or isolate the molecule (eg, a nucleic acid molecule or a polypeptide).
[0308] As used herein, the terms "determine," "assess," "determine," "measure," "detect," and their grammatical equivalents refer to both quantitative and qualitative determinations, and thus, the terms "determine" and "determine," "measure," and the like are used interchangeably herein. Where quantitative determination is intended, the phrase "determine the amount of an analyte, etc." is used. Where qualitative and / or quantitative determination is intended, the phrase "determine the level of an analyte" or "detect" an analyte is used.
[0309] A "fragment" is a portion of a protein or nucleic acid that is substantially identical to a reference protein or nucleic acid. In some embodiments, the portion retains at least 50%, 75%, or 80%, or 90%, 95%, or even 99% of the biological activity of the reference protein or nucleic acid described herein.
[0310] The terms "isolated," "purified," "biologically pure," and grammatical equivalents thereof refer to a substance that is freed to varying degrees from components that normally accompany it in its native state. "Isolated" means a degree of separation from an original source or environment. "Purified" means a degree of separation greater than isolated. A "purified" or "biologically pure" protein is sufficiently free of other substances such that any impurities do not materially affect the biological properties of the protein or cause other adverse consequences. That is, a nucleic acid or peptide of the present disclosure is purified if it is substantially free of cellular material, viral material, or culture medium when produced by recombinant DNA technology, or substantially free of chemical precursors or other chemicals when chemically synthesized. Purity and homogeneity are typically determined using analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high performance liquid chromatography. The term "purified" may mean that a nucleic acid or protein produces essentially one band in an electrophoretic gel. For proteins that can be modified, such as phosphorylation or glycosylation, different modifications may produce different isolated proteins that can be purified separately.
[0311] An "isolated" polypeptide (e.g., a peptide from an HLA-peptide complex) or polypeptide complex (e.g., an HLA-peptide complex) is a polypeptide or polypeptide complex of the present disclosure that has been separated from naturally associated components. Typically, a polypeptide or polypeptide complex is isolated when it is at least 60% by weight free of proteins and naturally occurring organic molecules with which it is naturally associated. The preparation may be at least 75%, at least 90%, or at least 99% by weight of a polypeptide or polypeptide complex of the present disclosure. An isolated polypeptide or polypeptide complex of the present disclosure may be obtained, for example, by extraction from a natural source, by expression of a recombinant nucleic acid encoding one or more components of the polypeptide or polypeptide complex, or by chemical synthesis of one or more components of the polypeptide or polypeptide complex. Purity may be measured by any appropriate method, such as column chromatography, polyacrylamide gel electrophoresis, or by HPLC analysis. In some cases, the MHC class II protein (i.e., MHC class II peptide) encoded by an HLA allele is referred to interchangeably in this document as an HLA class II protein (or HLA class II peptide).
[0312] The term "vector" refers to a nucleic acid molecule capable of transporting or mediating the expression of heterologous nucleic acids. Plasmid is one of the types covered by the term "vector". A vector generally refers to a nucleic acid sequence containing a replication origin and other entities necessary for replication and / or maintenance in a host cell. A vector capable of directing the expression of a gene and / or nucleic acid sequence operably connected thereto is referred to herein as an "expression vector". Typically, a useful expression vector is typically in the form of a "plasmid", which refers to a circular double-stranded DNA molecule that is not bound to a chromosome in the form of a vector and typically contains an entity or encoded DNA for stable or transient expression. Other expression vectors that can be used in the methods disclosed herein include, but are not limited to, plasmids, episomes, bacterial artificial chromosomes, yeast artificial chromosomes, phages or viral vectors, and such vectors can be integrated into the host's genome or replicate autonomously in a cell. The vector can be a DNA or RNA vector. Other forms of expression vectors known to those skilled in the art that function equivalently, for example, self-replicating extrachromosomal vectors or vectors capable of integrating into the host genome, can also be used. An exemplary vector is a vector capable of autonomous replication and / or expression of a nucleic acid connected thereto.
[0313] The term "spacer" or "connector" used for fusion proteins refers to peptides that connect proteins comprising fusion proteins. Generally, the spacer has no specific biological activity except for connecting or maintaining a certain minimum distance or other spatial relationship between proteins or RNA sequences. However, in some embodiments, the constituent amino acids of the spacer can be selected to affect certain properties of the molecule, such as the folding, net charge or hydrophobicity of the molecule. Suitable connectors for use in the embodiments of the present disclosure are well known to those skilled in the art, and include but are not limited to straight or branched carbon connectors, heterocyclic carbon connectors or peptide connectors. The connector is used to separate two antigenic peptides by a certain distance, which is sufficient to ensure that each antigenic peptide is correctly folded in some embodiments. Exemplary peptide connector sequences adopt a flexible extended conformation and do not show a tendency to develop an ordered secondary structure. Typical amino acids in flexible protein regions include Gly, Asn and Ser. In fact, it is expected that any arrangement of amino acid sequences containing Gly, Asn and Ser will meet the above-mentioned standards for connector sequences. Other near-neutral amino acids, such as Thr and Ala, can also be used in connector sequences. Other amino acid sequences that can be used as linkers are disclosed in Maratea et al. (1985), Gene 40:39-46, Murphy et al. (1986) Proc. Nat'l. Acad. Sci. USA 83:8258-62, US Pat. No. 4,935,233, and US Pat. No. 4,751,180.
[0314] The term "neoplasia" refers to any disease caused by or resulting in an inappropriately high level of cell division, an inappropriately low level of apoptosis, or both. Glioblastoma is a non-limiting example of neoplasia or cancer. The term "cancer" or "tumor" or "hyperproliferative disorder" refers to the presence of cells with typical characteristics of cancer cells such as uncontrolled proliferation, unlimited proliferation, metastatic potential, rapid growth and proliferation rate, and certain unique morphological characteristics. Cancer cells are usually in the form of tumors, but such cells can exist alone in an animal or can be non-tumorigenic cancer cells, such as leukemia cells. Cancers include, but are not limited to, B-cell cancers (e.g., multiple myeloma, Waldenstrom's macroglobulinemia), heavy chain diseases (e.g., alpha chain disease, gamma chain disease, and mu chain disease), benign monoclonal gammopathy and immune cell amyloidosis, melanoma, breast cancer, lung cancer, bronchial cancer, colorectal cancer, prostate cancer (e.g., metastatic, hormone-refractory prostate cancer), pancreatic cancer, gastric cancer, ovarian cancer, bladder cancer, brain or central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine cancer or endometrial cancer, cancer of the oral cavity or pharynx, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine or appendix cancer, salivary gland cancer, thyroid cancer, adrenal cancer, osteosarcoma, chondrosarcoma, cancer of blood tissue, etc. Other non-limiting examples of cancer types suitable for the methods encompassed by the present disclosure include human sarcomas and carcinomas, e.g., fibrosarcoma, myxosarcoma, liposarcoma, chondrosarcoma, osteosarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangioendotheliosarcoma, synovioma, mesothelioma, Ewing's tumor, tumor), leiomyosarcoma, rhabdomyosarcoma, colon cancer, colorectal cancer, pancreatic cancer, breast cancer, ovarian cancer, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland cancer, sebaceous gland cancer, papillary carcinoma, papillary adenocarcinoma, cystadenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct cancer, liver cancer, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, cervical cancer, bone cancer, brain tumor, testicular cancer, lung cancer, small cell lung cancer, bladder cancer, epithelial cancer, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pineal leukemias, e.g., acute lymphocytic leukemia and acute myeloid leukemia (myeloblastic, promyelocytic, myelomonocytic, monocytic and erythroleukemia); chronic leukemias (chronic myeloid (granulocytic) leukemia and chronic lymphocytic leukemia); as well as polycythemia vera, lymphomas (Hodgkin's disease and non-Hodgkin's disease), multiple myeloma, Waldenstrom's macroglobulinemia and heavy chain disease.In some embodiments, the cancer is an epithelial cancer, such as but not limited to bladder cancer, breast cancer, cervical cancer, colon cancer, gynecological cancer, kidney cancer, laryngeal cancer, lung cancer, oral cancer, head and neck cancer, ovarian cancer, pancreatic cancer, prostate cancer or skin cancer. In other embodiments, the cancer is breast cancer, prostate cancer, lung cancer or colon cancer. In other embodiments, the epithelial cancer is non-small cell lung cancer, non-papillary renal cell carcinoma, cervical cancer, ovarian cancer (e.g., serous ovarian cancer) or breast cancer. Epithelial cancer can be characterized in various other ways, including but not limited to serous, endometrioid, mucinous, clear cell, Brenner type (brenner) or undifferentiated. In some embodiments, the present disclosure is used for the treatment, diagnosis and / or prognosis of lymphoma or its subtypes (including but not limited to mantle cell lymphoma). Lymphoproliferative disorders are also considered to be proliferative diseases.
[0315] The term "vaccine" should be understood to refer to a composition used to generate immunity to prevent and / or treat a disease (e.g., neoplasia / tumor / infectious agent / autoimmune disease). Therefore, a vaccine is a drug containing an antigen and is intended to be used in humans or animals to produce specific defenses and protective substances by vaccination. A "vaccine composition" may include a pharmaceutically acceptable excipient, carrier, or diluent. Aspects of the present disclosure relate to the use of this technology in the preparation of antigen-based vaccines. In these embodiments, a vaccine refers to one or more disease-specific antigenic peptides (or their corresponding nucleic acids encoding them). In some embodiments, the antigen-based vaccine contains at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30 or more antigenic peptides. In some embodiments, the antigen-based vaccine contains 2 to 100, 2 to 75, 2 to 50, 2 to 25, 2 to 20, 2 to 19, 2 to 18, 2 to 17, 2 to 16, 2 to 15, 2 to 14, 2 to 13, 2 to 12, 2 to 10, 2 to 9, 2 to 8, 2 to 7, 2 to 6, 2 to 5, 2 to 4, 3 to 100, 3 to 75, 3 to 50, 3 to 25, 3 to 20, 3 to 19, 3 to 18, 3 to 17, 3 to 16, 3 to 15, 3 to 14, 3 to 13, 3 to 12, 3 to 10, 3 to 9, 3 to 8 4 to 15, 4 to 14, 4 to 13, 4 to 12, 4 to 10, 4 to 9, 4 to 8, 4 to 7, 4 to 6, 5 to 100, 5 to 75, 5 to 50, 5 to 25, 5 to 20, 5 to 19, 5 to 18, 4 to 17, 4 to 16, 4 to 15, 4 to 14, 4 to 13, 4 to 12, 4 to 10, 4 to 9, 4 to 8, 4 to 7, 4 to 6, 5 to 100, 5 to 75, 5 to 50, 5 to 25, 5 to 20, 5 to 19, 5 to 18, 5 to 17, 5 to 16, 5 to 15, 5 to 14, 5 to 13, 5 to 12, 5 to 10, 5 to 9, 5 to 8, or 5 to 7 antigenic peptides. In some embodiments, the antigen-based vaccine contains 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 antigenic peptides. In some cases, the antigenic peptide is a new antigenic peptide. In some cases, the antigenic peptide comprises one or more new epitopes.
[0316] The term "pharmaceutically acceptable" refers to approved or approvable by a regulatory agency of a federal or state government, or listed in the U.S. Pharmacopeia or other generally recognized pharmacopoeia for use in animals, including humans. "Pharmaceutically acceptable excipient, carrier or diluent" refers to an excipient, carrier or diluent that can be applied to a subject together with a medicament and does not destroy its pharmacological activity and is non-toxic when applied in a dose sufficient to deliver a therapeutic amount of the medicament. As described herein, a "pharmaceutically acceptable salt" of a combined disease-specific antigen can be an acid salt or basic salt that is generally considered in the art to be suitable for contact with human or animal tissues without excessive toxicity, irritation, allergic reactions or other problems or complications. Such salts include inorganic and organic acid salts of basic residues such as amines, and alkali metal or organic salts of acidic residues such as carboxylic acids. Specific pharmaceutical salts include, but are not limited to, salts of acids such as hydrochloric acid, phosphoric acid, hydrobromic acid, malic acid, glycolic acid, fumaric acid, sulfuric acid, sulfamic acid, aminobenzenesulfonic acid, formic acid, toluenesulfonic acid, methanesulfonic acid, benzenesulfonic acid, ethanedisulfonic acid, 2-hydroxyethylsulfonic acid, nitric acid, benzoic acid, 2-acetoxybenzoic acid, citric acid, tartaric acid, lactic acid, stearic acid, salicylic acid, glutamic acid, ascorbic acid, pamoic acid, succinic acid, fumaric acid, maleic acid, propionic acid, hydroxymaleic acid, hydroiodic acid, phenylacetic acid, alkanoic acids such as acetic acid, HOOC-(CH2)n-COOH, where n is 0-4, etc. Similarly, pharmaceutically acceptable cations include, but are not limited to, sodium, potassium, calcium, aluminum, lithium, and ammonium. Those of ordinary skill in the art will recognize from the present disclosure and the knowledge in the art that other pharmaceutically acceptable salts of the combined disease-specific antigens provided herein include those listed in Remington's Pharmaceutical Sciences, 17th edition, Mack Publishing Company, Easton, PA, p. 1418 (1985). Generally, pharmaceutically acceptable acid salts or basic salts can be synthesized from parent compounds containing basic or acidic moieties by any conventional chemical method. In short, such salts can be prepared by reacting the free acid or base form of these compounds with a stoichiometric amount of an appropriate base or acid in a suitable solvent.
[0317] Nucleic acid molecules that can be used in the methods of the present disclosure include any nucleic acid molecules encoding polypeptides of the present disclosure or fragments thereof. Such nucleic acid molecules do not have to be 100% identical to endogenous nucleic acid sequences, but generally show substantial identity. Polynucleotides having substantial identity to endogenous sequences are generally capable of hybridizing with at least one strand of a double-stranded nucleic acid molecule. "Hybridization" refers to the pairing of nucleic acid molecules under various stringent conditions between complementary polynucleotide sequences or portions thereof to form double-stranded molecules. (See, e.g., Wahl, GM and SL Berger (1987) Methods Enzymol. 152: 399; Kimmel, AR (1987) Methods Enzymol. 152: 507). For example, stringent salt concentrations may generally be less than about 750 mM NaCl and 75 mM trisodium citrate, less than about 500 mM NaCl and 50 mM trisodium citrate, or less than about 250 mM NaCl and 25 mM trisodium citrate. Low stringency hybridization can be obtained in the absence of organic solvents such as formamide, while high stringency hybridization can be obtained in the presence of at least about 35% formamide or at least about 50% formamide. Stringent temperature conditions can generally include a temperature of at least about 30°C, at least about 37°C, or at least about 42°C. It is well known to those skilled in the art to change other parameters, such as hybridization time, the concentration of detergents such as sodium dodecyl sulfate (SDS), and the inclusion or exclusion of carrier DNA. Various stringency levels are achieved by combining these various conditions as needed. In an exemplary embodiment, hybridization can occur at 30°C in 750mM NaCl, 75mM trisodium citrate, and 1% SDS. In another exemplary embodiment, hybridization can be carried out at 37°C in 500mM NaCl, 50mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA). In another exemplary embodiment, hybridization can occur at 42°C in 250mM NaCl, 25mM trisodium citrate, 1% SDS, 50% formamide, and 200 μg / ml ssDNA. Useful variations on these conditions will be apparent to those skilled in the art. For most applications, the washing steps after hybridization may also differ in stringency. Washing stringency conditions can be defined by salt concentration and by temperature. As described above, washing stringency can be increased by reducing salt concentration or by increasing temperature. For example, the stringent salt concentration of the washing step can be less than about 30mM NaCl and 3mM trisodium citrate, or less than about 15mM NaCl and 1.5mM trisodium citrate. The stringent temperature conditions of the washing step can include a temperature of at least about 25°C, at least about 42°C, or at least about 68°C.In an exemplary embodiment, the washing step can be performed at 25° C. in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In other exemplary embodiments, the washing step can be performed at 42° C. in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In another exemplary embodiment, the washing step can be performed at 68° C. in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Other variations of these conditions will be apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.
[0318] "Substantially identical" refers to a polypeptide or nucleic acid molecule that shows at least 50% identity with a reference amino acid sequence (e.g., any amino acid sequence described herein) or a nucleic acid sequence (e.g., any nucleic acid sequence described herein). Such a sequence can be at least 60%, 80% or 85%, 90%, 95%, 96%, 97%, 98% or even 99% or higher levels identical to the sequence used for comparison at the amino acid level or nucleic acid level. Sequence identity is usually measured using sequence analysis software (e.g., Sequence Analysis Software Package of the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions and / or other modifications. Conservative substitutions generally include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. In an exemplary method of determining the degree of identity, the BLAST program can be used, where probability scores between e-3 and em° represent closely related sequences. A "reference" is a comparison standard.
[0319] The term "subject" or "patient" refers to an animal that is the object of treatment, observation or experiment. By way of example only, a subject includes, but is not limited to, a mammal, including, but not limited to, a human or a non-human mammal, such as a non-human primate, mouse, cow, horse, dog, sheep or cat.
[0320] The terms "treatment", "treatment" and the like mean to reduce, prevent or improve a condition and / or symptoms associated therewith (e.g., neoplasia or tumor or infectious agent or autoimmune disease). "Treatment" may refer to administering treatment to a subject after an onset or suspected onset of a disease (e.g., cancer or infection with an infectious agent or autoimmune disease). "Treatment" includes the concept of "alleviation", which refers to reducing the frequency or severity of occurrence or recurrence of any symptom or other adverse effect associated with the disease and / or side effects associated with the treatment. The term "treatment" also encompasses the concept of "management", which refers to reducing the severity of a disease or condition in a patient, for example, prolonging the life span of a patient with the disease or prolonging their survival, or delaying their recurrence, for example, prolonging the remission period of a patient already suffering from the disease. It should be understood that, although not excluded, the treatment of a disease or condition does not require the complete elimination of the condition, condition or symptoms associated therewith.
[0321] As used herein, the terms "prevent," "preventing," and grammatical equivalents thereof refer to avoiding or delaying the onset of symptoms associated with a disease or condition in a subject who has not yet developed such symptoms when administration of an agent or compound is initiated.
[0322] The term "therapeutic effect" refers to a certain degree of alleviation of one or more symptoms of a disease (e.g., neoplasia, infection of a tumor or infectious agent or autoimmune disease) or its related pathology. "Therapeutically effective amount" as used herein refers to a dosage that effectively prolongs the survival of patients with such diseases, alleviates one or more signs or symptoms of the disease, prevents or delays, etc., beyond the expected degree in the absence of such treatments after a single or multiple dose administration to a cell or subject. "Therapeutically effective amount" is intended to limit the amount required for the therapeutic effect. A physician or veterinarian with ordinary skills in the art can easily determine and prescribe a "therapeutically effective amount" (e.g., ED50) of the desired pharmaceutical composition. For example, a physician or veterinarian can start with a level lower than the level required to obtain the desired therapeutic effect and gradually increase the dosage until the desired effect is obtained. The dosage of the compound of the present invention used in the pharmaceutical composition. Disease, condition and disease are used interchangeably herein.
[0323] Those of ordinary skill in the art will recognize that the terms "peptide tag," "affinity tag," "epitope tag," or "affinity receptor tag" are used interchangeably herein. As used herein, the term "affinity receptor tag" refers to an amino acid sequence that allows for easy detection or purification of the tagged protein, for example, by affinity purification. Affinity receptor tags are typically (but not necessarily) placed at or near the N- or C-terminus of the HLA allele. Various peptide tags are well known in the art. Non-limiting examples include polyhistidine tags (e.g., 4 to 15 consecutive His residues (SEQ ID NO: 4), such as 8 consecutive His residues (SEQ ID NO: 5)); polyhistidine-glycine tags; HA tags (e.g., Field et al., Mol. Cell. Biol., 8: 2159, 1988); c-myc tags (e.g., Evans et al., Mol. Cell. Biol., 5: 3610, 1985); herpes simplex virus glycoprotein D (gD) tags (e.g., Paborsky et al., Protein Engineering, 3:547, 1990); FLAG tag (e.g., Hopp et al., BioTechnology, 6:1204, 1988; U.S. Pat. Nos. 4,703,004 and 4,851,341); KT3 epitope tag (e.g., Martine et al., Science, 255:192, 1992); tubulin epitope tag (e.g., Skinner, Biol. Chem., 266:15173, 1991); T7 gene 10 protein peptide tag (e.g., Lutz-Freyemuth et al., Proc. Natl. Acad. Sci., 1996:114-117). c. Natl. Acad. Sci. USA, 87:6393, 1990); streptavidin tag (StrepTag.TM or StrepTagII.TM; see, e.g., Schmidt et al., J. Mol. Biol., 255(5):753-766, 1996 or U.S. Pat. No. 5,506,121; also commercially available from Sigma-Genosys); or a VSV-G epitope tag derived from the glycoprotein of vesicular stomatitis virus; or a V5 tag derived from a small epitope (Pk) found on the P and V proteins of the simian virus 5 (SV5) paramyxovirus. In some embodiments, the affinity receptor tag is an "epitope tag," a type of peptide tag that adds a recognizable epitope (antibody binding site) to an HLA protein to provide binding of a corresponding antibody, thereby allowing identification or affinity purification of the tagged protein. Non-limiting examples of epitope tags are protein A or protein G, which can bind to IgG. In some embodiments, the matrix of IgG Sepharose 6 Fast Flow chromatography resin is covalently coupled to human IgG. This resin allows for high flow rates, rapid and convenient purification of proteins labeled with Protein A.Many other tag moieties are known and contemplated by the skilled artisan and are contemplated herein.Any peptide tag may be used so long as it is capable of being expressed as a component of an affinity receptor-tagged HLA-peptide complex.
[0324] As used herein, the term "affinity molecule" refers to a molecule or ligand that binds to an affinity receptor peptide with chemical specificity. Chemical specificity is the ability of a protein binding site to bind a specific ligand. The fewer ligands a protein can bind, the higher its specificity. Specificity describes the strength of the binding between a given protein and a ligand. This relationship can be measured by the dissociation constant (K D ), which characterizes the equilibrium between the bound and unbound states of the protein-ligand system.
[0325] The term "affinity receptor-tagged HLA-peptide complex" refers to a complex comprising an HLA class I or class II associated peptide or a portion thereof that specifically binds to a monoallelic recombinant HLA class I or class II peptide comprising an affinity receptor peptide.
[0326] The terms "specific binding" or "specific binding" when applied to the interaction of affinity molecules and affinity receptor tags or epitopes with HLA peptides means that the interaction is dependent on the presence of a particular structure (e.g., an antigenic determinant or epitope) on the protein; in other words, the affinity molecule recognizes and binds to a specific affinity receptor peptide structure rather than binding to the protein in general.
[0327] As used herein, the term "affinity" refers to a measure of the strength of binding between two members of a binding pair (e.g., an "affinity receptor tag" and an "affinity molecule" and an HLA binding peptide and an HLA class I or class II molecule). D is the dissociation constant and has units of molar concentration. The affinity constant is the reciprocal of the dissociation constant. The affinity constant is sometimes used as a general term to describe the chemical entity. It is a direct measure of binding energy. Affinity can be determined experimentally, for example, by surface plasmon resonance (SPR) using a commercially available Biacore SPR unit. Affinity can also be expressed as an inhibitory concentration 50 (IC 50 ), which is the concentration at which 50% of the peptide is replaced. Similarly, lnIC 50 IC 50 The natural logarithm of K off Refers to the dissociation rate constant, for example, the dissociation rate constant of an affinity molecule from an affinity receptor-tagged HLA-peptide complex.
[0328] In some embodiments, the affinity receptor-tagged HLA-peptide complex comprises a biotin acceptor peptide (BAP) and is immunopurified from the complex cell mixture using streptavidin / NeutrAvidin beads. Biotin-avidin / streptavidin binding is the strongest non-covalent interaction known in nature. This property is used in a wide range of applications as a biological tool, such as immunopurification of proteins covalently linked to biotin. In an exemplary embodiment, the nucleic acid sequence encoding the HLA allele uses a biotin acceptor peptide (BAP) as an affinity receptor tag for immunopurification. BAP can be specifically biotinylated at a single lysine residue within the tag in vivo or in vitro (e.g., U.S. Patent Nos. 5,723,584; 5,874,239; and 5,932,433; and British Patent No. GB2370039). BAP is typically 15 amino acids long and contains one lysine as a biotin acceptor residue. In some embodiments, BAP is placed at or near the N- or C-terminus of a monoallelic HLA peptide. In some embodiments, the BAP is placed between the heavy chain domain and the β2 microglobulin domain of an HLA class I peptide. In some embodiments, the BAP is placed between the β chain domain and the α chain domain of an HLA class II peptide. In some embodiments, the BAP is placed in a loop region between the α1, α2, and α3 domains of an HLA class I heavy chain, or between the α1 and α2 and β1 and β2 domains of an HLA class II α chain and β chain, respectively.
[0329] As used herein, the term "biotin" refers to the compound biotin itself and its analogs, derivatives and variants. Therefore, the term "biotin" includes biotin (cis-hexahydro-2-oxo-1H-thieno[3,4]imidazole-4-pentanoic acid) and any derivatives and analogs thereof, including biotin-like compounds. Such compounds include, for example, amino or sulfhydryl derivatives of biotin-eN-lysine, biocytin hydrazide, 2-iminobiotin and biotinyl-E-aminocaproic acid-N-hydroxysuccinimide ester, sulfosuccinimidyl iminobiotin, biotin bromoacetyl hydrazide, p-diazobenzoyl biocytin, 3-(N-maleimidopropionyl) biocytin, desthiobiotin, etc. The term "biotin" also includes biotin variants that can specifically bind to one or more of Rhizavidin, avidin, streptavidin, tamavidin moieties or other avidin-like peptides.
[0330] As used herein, a "PPV determination method" may refer to a presentation PPV determination method. For example, a "PPV determination method" may refer to a method comprising the following steps: (a) using an HLA peptide presentation prediction model, such as a machine learning HLA peptide presentation prediction model, to process amino acid information of a plurality of test peptide sequences to generate a plurality of test presentation predictions, each test presentation prediction indicating the likelihood that one or more proteins encoded by a class II HLA allele of a cell (such as a class II HLA allele of a subject's cell) are capable of presenting a given test peptide sequence of the plurality of test peptide sequences, wherein the plurality of test peptide sequences comprises at least 500 test peptide sequences, the test peptide sequences comprising (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in the cell, and (ii) at least 499 hit peptide sequences contained in a hit peptide sequence expressed by an organism (e.g., a subject of the same species as the subject); (a) bait peptide sequences in proteins encoded by the genome of an organism of a species, wherein the ratio of the number of hit peptide sequences to the number of bait peptide sequences in the plurality of test peptide sequences is less than 1, for example, the ratio of the at least one hit peptide sequence to the at least 499 bait peptide sequences is 1:499; (b) identifying or determining the top percentage of the plurality of test peptide sequences, such as the top 0.2% of the plurality of test peptide sequences, as presented by the class II HLA alleles of the cell; and (c) calculating the PPV of the HLA peptide presentation prediction model, wherein the PPV is the fraction of the test peptide sequences identified or determined as presented by the class II HLA alleles of the cell in the plurality of test peptide sequences, which are peptides observed to be presented by the class II HLA alleles of the cell by mass spectrometry. In some embodiments, the bait peptides have the same length, i.e., contain the same number of amino acids as the hit peptides. In some embodiments, the bait peptides may contain one more or one less amino acid than the hit peptides. In some embodiments, the bait peptides are peptides that are endogenous peptides. In some embodiments, the bait peptide is a synthetic peptide. In some embodiments, the bait peptide is an endogenous peptide that has been identified by mass spectrometry as binding to a first MHC class I or class II protein, wherein the first MHC class I or class II protein is different from a second MHC class I or class II protein that binds to the hit peptide. In some embodiments, the bait peptide can be a scrambled peptide, for example, the bait peptide can comprise an amino acid sequence in which the amino acid positions are rearranged relative to the amino acid positions of the hit peptide within the length of the peptide. In some embodiments, the PPV determination method can be a presentation PPV determination method. In some embodiments, the ratio of the number of hit peptide sequences to the number of bait peptide sequences is about 1:10, 1:20, 1:50, 1:100, 1:250, 1:500, 1:1000, 1:1500, 1:2000, 1:2500, 1:5000, 1:7500, 1:10000, 1:25000,1:50000 or 1:100000. In some embodiments, the at least one hit peptide sequence comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, In some embodiments, the at least 499 decoy peptide sequences comprise at least 500 hit peptide sequences. 600、700、800、900、1000、1100、1200、1300、1400、1500、1600、1700、1800、1900、2000、2100、2200、2300、2400、2500、2600、2700、2800、2900、3000、3100、3200、3300、3400、3500、3600、3700、38 00, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7 000、7100、7200、7300、7400、7500、7600、7700、7800、7900、8000、8100、8200、8300、8400、8500、8600、8700、8800、8900、9000、9100、9200、9300、9400、9500、9600、9700、9800、9900、10000、110 00, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000,38000、39000、40000、41000、42000、43000、44000、45000、46000、47000、48000、49000、50000、52500、55000、57500、60000、62500、65000、67500、70000、72500、75000、77500、80000、82500、85000、87500、90000、92 700,000, 800,000, 900,000, or 1,000,000 decoy peptide sequences. In some embodiments, the at least 500 test peptide sequences comprise at least 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400、3500、3600、3700、3800、3900、4000、4100、4200、4300、4400、4500、4600、4700、4800、4900、5000、5100、5200、5300、5400、5500、5600、5700、5800、5900、6000、6100、6200、6300、6400、6500、66 00、6700、6800、6900、7000、7100、7200、7300、7400、7500、7600、7700、7800、7900、8000、8100、8200、8300、8400、8500、8600、8700、8800、8900、9000、9100、9200、9300、9400、9500、9600、9700、9800 、9900、10000、11000、12000、13000、14000、15000、16000、17000、18000、19000、20000、21000、22000、23000、24000、25000、26000、27000、28000、29000、30000、31000、32000、33000、34000、35000、36000、37000、38000、39000、40000、41000、42000、43000、44000、45000、46000、47000、48000、49000、50000、52500、55000、57500、60000、62500、65000、67500、70000、72500、75000、77500、80000、82500、85000、87500、90 0000, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000, 800000, 900000 or 1000000 test peptide sequences. In some embodiments, identifying or determining a top percentage of the plurality of test peptide sequences as presented by the class II HLA alleles of the cell comprises identifying or determining a top 0.20%, 0.30%, 0.40%, 0.50%, 0.60%, 0.70%, 0.80%, 0.90%, 1.00%, 1.10%, 1.20%, 1.30%, 1.40%, 1.50%, 1.60%, 1.70%, 1.80%, 1.90%, 1.01%, 1.12%, 1.13%, 1.14%, 1.15%, 1.16%, 1.17%, 1.18%, 1.19% or more of the plurality of test peptide sequences as presented by the class II HLA alleles of the cell. 0%, 2.00%, 2.10%, 2.20%, 2.30%, 2.40%, 2.50%, 2.60%, 2.70%, 2.80%, 2.90%, 3.00%, 3.10%, 3.20%, 3.30%, 3.40%, 3.50%, 3.60%, 3.70%, 3.80%, 3.90%, 4.00%, 4.10%, 4.20%, 4.30%, 4.40%, 4.50%, 4.60%, 4.70%, 4.80%, 4.90%, 5.00%, 5.10%, 5.20%, 5.30%, 5.40%, 5.50%, 5.60%, 5.70%, 5.80%, 5.90%, 6.00%, 6.10%, 6.20%, 6.30%, 6.40%, 6.50%, 6.60%, 6.70%, 6.80%, 6.90%, 7.00%, 7.10%, 7.2 0%, 7.30%, 7.40%, 7.50%, 7.60%, 7.70%, 7.80%, 7.90%, 8.00%, 8.10%, 8.20%, 8.30%, 8.40%, 8.50%, 8.60%, 8.70%, 8.80%, 8.90%, 9.00%, 9.10%, 9.20%, 9.30%, 9.40%, 9.50%, 9.60%, 9.70%, 9.80%,9. 90%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19% or 20% are identified or determined to be presented by a class II HLA allele of the cell. In some embodiments, the cell is a monoallelic cell.
[0331] As used herein, a "PPV determination method" may refer to a binding PPV determination method. For example, a "PPV determination method" may refer to a method comprising the following steps: (a) using an HLA peptide binding prediction model, such as a machine learning HLA peptide binding prediction model, to process amino acid information of a plurality of test peptide sequences to generate a plurality of test binding predictions, each test binding prediction indicating the likelihood that one or more proteins encoded by a class I or class II HLA allele of a cell (such as a class I or class II HLA allele of a subject's cell) bind to a given test peptide sequence of the plurality of test peptide sequences, wherein the plurality of test peptide sequences comprises at least 20 test peptide sequences, the test peptide sequences comprising (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in the cell, and (ii) at least 19 decoy peptide sequences contained within a protein, the protein The invention comprises at least one peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell, wherein the ratio of the number of hit peptide sequences to the number of bait peptide sequences in the plurality of test peptide sequences is less than 1, for example, the ratio of the at least one hit peptide sequence to the at least 19 bait peptide sequences is 1:19; (b) identifying or determining a top percentage of the plurality of test peptide sequences, such as the top 5% of the plurality of test peptide sequences, as binding to the HLA protein; and (c) calculating the PPV of the HLA peptide binding prediction model, wherein the PPV is the fraction of the test peptide sequences in the plurality of test peptide sequences identified or determined as binding to a class I or class II HLA allele of the cell, which peptides are peptides observed to be presented by a class I or class II HLA allele of the cell by mass spectrometry. In some embodiments, the ratio of the number of hit peptide sequences to the number of decoy peptide sequences is about 1:2, 1:3, 1:4, 1:5, 1:10, 1:20, 1:25, 1:30, 1:40, 1:50, 1:75, 1:100, 1:200, 1:250, 1:500, or 1:1000. In some embodiments, the at least one hit peptide sequence comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75 ,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99 or 100 hit peptide sequences. In some embodiments, the at least 19 decoy peptide sequences comprise at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400 400、1500、1600、1700、1800、1900、2000、2100、2200、2300、2400、2500、2600、2700、2800、2900、3000、3100、3200、3300、3400、3500、3600、3700、3800 , 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 630 0, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 8700, 8 800、8900、9000、9100、9200、9300、9400、9500、9600、9700、9800、9900、10000、11000、12000、13000、14000、15000、16000、17000、18000、19000、20000 , 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 41 000、42000、43000、44000、45000、46000、47000、48000、49000、50000、52500、55000、57500、60000、62500、65000、67500、70000、72500、75000、77500、425000, 450000, 475000, 500000, 600000, 700000, 800000, 90000 or 1000000 decoy peptide sequences. In some embodiments, the at least 20 test peptide sequences comprise at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470 , 480, 490, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300 , 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8 300、8400、8500、8600、8700、8800、8900、9000、9100、9200、9300、9400、9500、9600、9700、9800、9900、10000、11000、12000、13000、14000、15000、16000、17000、18000、19000、20000、21000、22000、23000、24000、25000、26000、27000、28000、29000、30000、31000、32000、33000、34000、35000、36000、37000、38000、39000、40000、41000、42000、43000、44000、45000、46000、47000、48000、49000、50000、52500、55000、57500、60000、62500、65000、67500、70000、72500、75000、77500、80000 , 82500, 85000, 87500, 90000, 92500, 95000, 97500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000, 800000, 900000 or 1000000 test peptide sequences. In some embodiments, identifying or determining a top percentage of the plurality of test peptide sequences as presented by a class II HLA allele of the cell comprises identifying or determining the top 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, or 40% as presented by a class II HLA allele of the cell. In some embodiments, the cell is a monoallelic cell.
[0332] Human Leukocyte Antigen (HLA) System
[0333] The immune system can be classified into two functional subsystems: the innate immune system and the adaptive immune system. The innate immune system is the first line of defense against infection, and most potential pathogens are rapidly neutralized by this system before they can cause, for example, an overt infection. The adaptive immune system reacts to molecular structures of invading organisms, known as antigens. Unlike the innate immune system, the adaptive immune system is highly specific to pathogens. Adaptive immunity can also provide long-lasting protection; for example, a person who has recovered from measles is now protected from measles for life. There are two types of adaptive immune responses, including humoral and cell-mediated. In humoral immune responses, antibodies secreted into body fluids by B cells bind to antigens derived from pathogens, resulting in the elimination of the pathogens through a variety of mechanisms, such as complement-mediated lysis. In cell-mediated immune responses, T cells that are capable of destroying other cells are activated. For example, if disease-associated proteins are present in cells, they are proteolytically broken into peptides within the cell. Specific cell proteins then attach themselves to the antigens or peptides formed in this way and transport them to the cell surface, where they are presented to the molecular defense mechanisms in the body's T cells. Cytotoxic T cells recognize these antigens and kill cells bearing the antigen.
[0334] The term "major histocompatibility complex (MHC)", "MHC molecule" or "MHC protein" refers to a protein that is capable of binding peptides generated by proteolytic cleavage of a protein antigen and representing a potential T-cell epitope, transporting it to the cell surface, and presenting the peptides to specific cells, such as in cytotoxic T lymphocytes or T helper cells. The human MHC is also known as the HLA complex. Thus, the term "human leukocyte antigen (HLA) system", "HLA molecule" or "HLA protein" refers to a gene complex encoding a human MHC protein. The term MHC is referred to as the "H-2" complex in murine species. One of ordinary skill in the art will recognize that the terms "major histocompatibility complex (MHC)", "MHC molecule", "MHC protein" and "human leukocyte antigen (HLA) system", "HLA molecule", "HLA protein" are used interchangeably herein.
[0335] HLA proteins are classified into two types, known as HLA class I and HLA class II. The structures of the proteins of both HLA classes are very similar; however, they have very different functions. HLA class I proteins are found on the surface of almost all cells in the body, including most tumor cells. HLA class I proteins are loaded with antigens, which usually originate from endogenous proteins or pathogens present within the cell, and are then presented to naive or cytotoxic T lymphocytes (CTLs). HLA class II proteins are found on antigen presenting cells (APCs), including but not limited to dendritic cells, B cells, and macrophages. They primarily present peptides processed from external antigen sources, such as outside the cell, to helper T cells. Most peptides bound by HLA class I proteins originate from cytoplasmic proteins produced in the organism's own healthy host cells and generally do not stimulate an immune response.
[0336] HLA class I molecules consist of two non-covalently linked polypeptide chains: an HLA-encoded alpha chain (heavy chain, 44 to 47 kD) and a non-HLA-encoded subunit, called beta-2 microglobulin (or beta-2m), (12 kD). The alpha chain has three extracellular domains, alpha 1, alpha 2, and alpha 3, and a transmembrane region, wherein the alpha 1 and alpha 2 regions are capable of binding peptides of about 7 to 13 amino acids (e.g., about 8 to 11 amino acids, or 9 or 10 amino acids). HLA class 1 molecules bind to peptides with appropriate binding motifs and present them to cytotoxic T lymphocytes. HLA class 1 heavy chains can be the protein product of the HLA-A allele, also known as the HLA-A monomer, or the protein product of the HLA-B allele (again, the HLA-B monomer) or the protein product of the HLA-C allele (HLA-C monomer), each of which is complexed with beta-2-microglobulin. α1 is dependent on the non-HLA protein β2m; β2m is encoded by the β-2-microglobulin gene located on chromosome 15 in humans. The α3 domain is linked to the transmembrane region, anchoring the HLA class I molecule to the cell membrane. The presented peptide is held in the central region of the α1 / α2 heterodimer (a molecule composed of two different subunits) by the bottom of the peptide-binding groove. HLA class IA, HLA class IB, or HLA class IC are highly polymorphic. The HLA class 1-A gene (referred to as the HLA-A gene), the HLA class 1-B gene (referred to as the HLA-B gene), and the HLA class 1-C gene (referred to as the HLA-C gene) each contain 8 exons, with exon 1 encoding the leader peptide, exons 2 and 3 encoding the α1 and α2 domains, exon 5 encoding the transmembrane region, and exons 6 and 7 encoding the cytoplasmic tail. Polymorphisms in exons 2 and 3 determine the peptide binding specificity of each class 1 molecule. HLA-B class genes (HLA-B) have many possible variations, expression patterns, and antigens presented. This group is subdivided into groups encoded within the HLA locus, such as HLA-E, HLA-F, HLA-G, and those not encoded therein, such as stress ligands, such as ULBP, Rae1, and H60. The antigen / ligand of many of these molecules remains unknown, but they can interact with CD8+T cells, NKT cells, and NK cells.
[0337] In some embodiments, the present disclosure utilizes non-classical HLA class IE alleles. HLA-E molecules are recognized by natural killer (NK) cells and CD8+T cells. HLA-E is expressed in almost all tissues, including lung, liver, skin and placental cells. HLA-E expression has also been detected in solid tumors (e.g., osteosarcoma and melanoma). HLA-E molecules bind to TCR expressed on CD8+T cells, thereby causing T cell activation. HLA-E is also known to bind to CD94 / NKG2 receptors expressed on NK cells and CD8+T cells. CD94 can be paired with several different isoforms of NKG2 to form receptors with the potential to inhibit (NKG2A, NKG2B) or promote (NKG2C) cell activation. HLA-E can bind to peptides derived from amino acid residues 3-11 of the leader sequences of most HLA-A, -B, -C and -G molecules, but cannot bind to its own leader peptide. It has also been shown that HLA-E presents peptides derived from endogenous proteins similar to HLA-A, -B and -C alleles. Under physiological conditions, CD94 / NKG2A and the engagement of HLA-E loaded with peptides from HLA class I leader sequences usually induce inhibitory signals. Cytomegalovirus (CMV) utilizes the mechanism of escaping NK cell immune surveillance by expressing UL40 glycoprotein (simulating HLA-A leader sequence). However, it has also been reported that CD8+T cells can identify HLA-E loaded with UL40 peptides derived from CMV Toledo strains, and work in defense CMV. A large number of studies have revealed several important functions of HLA-E in infectious diseases and cancer.
[0338] Peptide antigens attach themselves to HLA class I molecules by competitive affinity binding within the endoplasmic reticulum before presentation on the cell surface. Here, the affinity of an individual peptide antigen is directly related to its amino acid sequence and the presence of specific binding motifs at defined positions within the amino acid sequence. If the sequence of such a peptide is known, it is possible to manipulate the immune system against diseased cells using, for example, peptide vaccines.
[0339] MHC molecules are highly polymorphic, that is, there are many MHC variants. Each variant is encoded by a variant of the gene encoding the protein, and each such variant gene is called an allele. For humans, MHC is called human leukocyte antigen (HLA), which involves three types of HLA class II molecules: DP, DQ and DR. HLA class II peptides (Figure 1) have two chains: α and β, each of which has two domains-α1 and α2 and β1 and β2-each chain has a transmembrane domain: α2 and β2, anchoring HLA class II molecules to the cell membrane. The peptide binding groove is formed by heterodimers of α1 and β1. The most widely studied HLA-DR molecules have DRA and DRB, corresponding to the α and β domains, respectively. DRB is diverse, and DRA is almost the same. Therefore, the binding specificity of the DRB allele indicates the binding specificity of the corresponding HLA-DR. Each MHC protein has its own binding specificity, which means that a set of peptides bound to an MHC molecule may be different from the peptides bound to another MHC molecule. Classical molecules present peptides to CD4+ lymphocytes. Nonclassical molecules with intracellular functions, annexins, are not exposed on the cell membrane but in the inner membrane of lysosomes and usually load antigenic peptides onto classical HLA class II molecules.
[0340] In the HLA class II system, phagocytic cells such as macrophages and immature dendritic cells take up entities by phagocytosis into phagosomes - although B cells show more common endocytosis into endosomes - which fuse with lysosomes, whose acidic enzymes cleave the ingested proteins into many different peptides. Autophagy is another source of HLA class II peptides. Through physicochemical dynamics of molecular interactions with HLA class II variants carried by the host (encoded in the host genome), specific peptides become immunodominant and loaded onto HLA class II molecules. They are transported to and externalized on the cell surface. The most studied subclasses of HLA class II genes are: HLA-DPA1, HLA-DPB1, HLA-DQA1, HLA-DQB1, HLA-DRA, and HLA-DRB1.
[0341] Presentation of peptides by HLA class II molecules to CD4+ helper T cells is required for immune responses to foreign antigens (Roche and Furuta, 2015). Once activated, CD4+ T cells promote B cell differentiation and antibody production, as well as CD8+ T cell (CTL) responses. CD4+ T cells also secrete cytokines and chemokines that activate and induce differentiation of other immune cells. HLA class II molecules are heterodimers of α and β chains that interact to form a peptide binding groove that is more open than the HLA class I peptide binding groove (Unanue et al., 2016). Peptides that bind to HLA class II molecules are thought to have a 9-amino acid binding core with flanking residues protruding from the binding groove on the N-terminal or C-terminal side (Jardetzky et al., 1996; Stern et al., 1994). These peptides are typically 12-16 amino acids in length and usually contain 3-4 anchor residues at positions P1, P4, P6 / 7, and P9 of the binding moiety (Rossjohn et al., 2015).
[0342] HLA alleles are expressed in a codominant manner, which means that the alleles (variants) inherited from both parents are expressed equally. For example, each person carries 2 alleles of each gene in 3 class I genes (HLA-A, HLA-B and HLA-C), so six different types of HLA class II can be expressed. In the HLA class II locus, each person inherits a pair of HLA-DP genes (DPA1 and DPB1, encoding α and β chains), HLA-DQ (DQA1 and DQB1 for α and β chains), one HLA-DRα gene (DRA1) and one or more HLA-DRβ genes (DRB1 and DRB3, -4 or -5). For example, HLA-DRB1 has more than nearly 400 known alleles. This means that a heterozygous individual can inherit six or eight functional HLA class II alleles, three or more per parent. Therefore, HLA genes are highly polymorphic; there are many different alleles in different individuals within a population. The genes encoding the HLA proteins have many possible variations, allowing each person's immune system to respond to a wide variety of foreign invaders. Some HLA genes have hundreds of identified forms (alleles), each of which is given a specific number. In some embodiments, the HLA class I alleles are HLA-A*02:01, HLA-B*14:02, HLA-A*23:01, HLA-E*01:01 (non-classical). In some embodiments, the HLA class II alleles are HLA-DRB*01:01, HLA-DRB*01:02, HLA-DRB*11:01, HLA-DRB*15:01, and HLA-DRB*07:01.
[0343] The subject-specific HLA allele or HLA genotype of the subject can be determined by any method known in the art. In an exemplary embodiment, the HLA genotype is determined by any method described in International Patent Application No. PCT / US2014 / 068746 (published as WO2015085147 on June 11, 2015), which is incorporated herein by reference in its entirety. Briefly, the method includes determining a polymorphic gene type, which may include generating an alignment of reads extracted from a sequencing data set with a gene reference set comprising allelic variants of the polymorphic gene, determining a first posterior probability or a score derived from the posterior probability for each allelic variant in the alignment, identifying the allelic variant with the largest first posterior probability or a score derived from the posterior probability as a first allelic variant, identifying one or more overlapping reads aligned to the first allelic variant and one or more other allelic variants, determining a second posterior probability or a score derived from the posterior probability for the one or more other allelic variants using a weighting factor, determining a second allelic variant by selecting the allelic variant with the largest second posterior probability or a score derived from the posterior probability, the first and second allelic variants defining the gene type of the polymorphic gene, and providing an output of the first and second allelic variants.
[0344] In some embodiments, MHC class II peptides described herein: antigen peptides combine and present prediction methods have the ability of predicting binding substances from the large group library MHC class II peptides encoded by independent HLA alleles.In some embodiments, the large database of HLA matching peptides verified by mass spectrometry of MAPTAC technology is trained.In some embodiments, the large database of HLA matching peptides verified by mass spectrometry comprises more than 1.2x10^6 such HLA matching peptides.In some embodiments, the large database of HLA matching peptides verified by mass spectrometry covers more than 150 HLA alleles, including MHC class I and class II allele hypotypes.In some embodiments, the database covers at least 95% of the U.S. population for HLA-I and HLA-II (DR hypotype).
[0345] As described in this article, there is ample evidence in animals and humans that mutated epitopes can effectively induce immune responses, that spontaneous tumor regression or long-term survival is associated with CD8+ T cell responses to mutated epitopes, and that “immunoediting” can track changes in the expression of dominant mutant antigens in mice and humans.
[0346] Sequencing techniques reveal that each tumor contains multiple patient-specific mutations that alter the protein-coding content of genes. Such mutations produce altered proteins ranging from single amino acid changes (caused by missense mutations) to long regions that add new amino acid sequences due to frameshifts, readthrough of stop codons, or translation of intronic regions (new open reading frame mutations; neoORFs). These mutant proteins are valuable targets for the host's immune response to tumors because, unlike native proteins, they are not affected by the immunosuppressive effects of self-tolerance. Therefore, mutated proteins are more likely to be immunogenic and more specific to tumor cells than to normal cells of the patient. In essence, short peptides (8-24 amino acids long) containing cancer-associated mutations are candidates for cancer immunotherapy.
[0347] In some embodiments, the algorithm of the driver prediction method can be further used to determine the mutation of the peptide. In some embodiments, the prediction method can be used to determine the driver mutation status, and / or RNA expression status, and / or cleavage prediction within the peptide.
[0348] The term "T cell" includes CD4+T cells and CD8+T cells. The term T cell also includes T helper type 1 T cells and T helper type 2 T cells. T cells used herein are generally classified into two major categories: helper T (TH) cells and cytotoxic T lymphocytes (CTLs) by function and cell surface antigens (cluster differentiation antigens or CDs) that also help T cell receptors bind to antigens.
[0349] Mature helper T (TH) cells express surface protein CD4 and are referred to as CD4+T cells. After T cell development, mature immature T cells leave the thymus and begin to spread throughout the body, including lymph nodes. Immature T cells are those T cells that have never been exposed to the antigens to which they are programmed to respond. Like all T cells, they express T cell receptor-CD3 complexes. T cell receptor (TCR) consists of a constant region and a variable region. The variable region determines what antigens T cells can respond to. CD4+T cells have TCRs with affinity for MHC class II proteins, and CD4 participates in determining the MHC affinity during thymus maturation. MHC class II proteins are usually only found on the surface of specialized antigen presenting cells (APCs). Specialized antigen presenting cells (APCs) are mainly dendritic cells, macrophages and B cells, although dendritic cells are the only group of cells that constitutively (always) express MHC class II. Some APCs also bind native (or unprocessed) antigens to their surface, such as follicular dendritic cells, but unprocessed antigens do not interact with T cells and do not participate in their activation. Peptide antigens bound to HLA class I proteins are generally shorter than those bound to HLA class II proteins.
[0350] Cytotoxic T lymphocytes (CTL), also known as cytotoxic T cells, cytolytic T cells, CD8+T cells or killer T cells, refer to lymphocytes that induce apoptosis in targeted cells. CTL forms antigen-specific conjugates with target cells through the interaction of TCR and antigens (Ag) processed on the surface of target cells, thereby causing apoptosis of target cells. Apoptotic bodies are eliminated by macrophages. The term "CTL response" is used to refer to the primary immune response mediated by CTL cells. Cytotoxic T lymphocytes have both T cell receptors (TCR) and CD8 molecules on their surfaces. T cell receptors can recognize and bind peptides complexed with HLA class I molecules. Each cytotoxic T lymphocyte expresses a unique T cell receptor that can bind to a specific MHC / peptide complex. Most cytotoxic T cells express T cell receptors (TCRs) that can recognize specific antigens. In order for TCR to bind to HLA class I molecules, the former must be accompanied by a glycoprotein called CD8, which binds to the constant part of HLA class I molecules. Therefore, these T cells are called CD8+ T cells. The affinity between CD8 and MHC molecules brings T cells and target cells together during antigen-specific activation. Once activated, CD8+ T cells are recognized as T cells and are generally classified as having a predetermined cytotoxic role in the immune system. However, CD8+ T cells also have the ability to produce certain cytokines.
[0351] "T cell receptor (TCR)" is a cell surface receptor involved in the activation of T cells in response to antigen presentation. TCR is usually composed of two chains, α and β, which are assembled to form a heterodimer and associate with the CD3 transduction subunit to form a T cell receptor complex present on the cell surface. Each α and β chain of TCR consists of immunoglobulin-like N-terminal variable (V) and constant (C) regions, hydrophobic transmembrane domains, and short cytoplasmic regions. As for immunoglobulin molecules, the variable regions of α and β chains are produced by V(D)J recombination, thereby producing a great diversity of antigen specificity in T cell populations. However, compared with immunoglobulins that recognize complete antigens, T cells are activated by processed peptide fragments associated with MHC molecules, thereby introducing an additional dimension to T cell recognition of antigens, which is called MHC restriction. Recognition of MHC differences between donors and recipients by T cell receptors can lead to T cell proliferation and potential development of GVHD. It has been shown that normal surface expression of the TCR depends on the coordinated synthesis and assembly of all seven components of the complex (Ashwell and Klusner 1990). Inactivation of TCRα or TCRβ can lead to the elimination of the TCR from the T cell surface, thereby preventing recognition of alloantigens and thus preventing GVHD. However, TCR disruption usually leads to the elimination of CD3 signaling components and alters the pattern of further T cell expansion.
[0352] The term "HLA peptide group" refers to a group of peptides that specifically interact with a particular HLA class and can contain thousands of different sequences. The HLA peptide group includes diverse peptides that are derived from normal and abnormal proteins expressed in cells. Therefore, the HLA peptide group can be studied to identify cancer-specific peptides for the development of tumor immunotherapy and as a source of information about protein synthesis and degradation programs within cancer cells. In some embodiments, the HLA peptide group is a group of soluble HLA peptides (sHLA). In some embodiments, the HLA peptide group is a group of membrane-bound HLA (mHLA).
[0353] "Antigen presenting cells" or "APCs" include professional antigen presenting cells (e.g., B lymphocytes, macrophages, monocytes, dendritic cells, Langerhans cells), as well as other antigen presenting cells (e.g., keratinocytes, endothelial cells, astrocytes, fibroblasts, oligodendrocytes, thymic epithelial cells, thyroid epithelial cells, glial cells (brain), pancreatic β cells, and vascular endothelial cells). "Antigen presenting cells" or "APCs" are cells that express major histocompatibility complex (MHC) molecules and can display foreign antigens complexed with MHC on their surface.
[0354] Monoallelic HLA cell lines
[0355] A single HLA class I allele, a pair of HLA class II alleles or a single HLA class I allele and a pair of HLA class II alleles can be generated by transducing or transfecting a suitable cell colony with a polynucleic acid (e.g., vector) encoding a single HLA allele. A monoallelic cell line expressing a single HLA class I allele, a pair of HLA class II alleles or a single HLA class I allele and a pair of HLA class II alleles. Suitable cell colonies include, for example, an HLA class I defective cell line wherein a single HLA class I allele is exogenously expressed, an HLA class II defective cell line wherein a pair of single exogenous HLA class II alleles is expressed, or a class I and class II defective cell line wherein a single HLA class I and / or a pair of class II alleles are exogenously expressed. As an exemplary embodiment, the HLA class I defective B cell line is B721.221. However, it is clear to the technician that other cell colonies of HLA class I and / or HLA class II defects can be produced. Exemplary methods for deleting / inactivating endogenous HLA class I or HLA class II genes include CRISPR-Cas9-mediated genome editing, such as in THP-1 cells. In some embodiments, the cell colony is a professional antigen presenting cell, such as a macrophage, a B cell, and a dendritic cell. The cell may be a B cell or a dendritic cell. In some embodiments, the cell is a tumor cell or a cell from a tumor cell line. In some embodiments, the cell is isolated from a patient. In some embodiments, the cell contains an infectious agent or a portion thereof. In some embodiments, the cell colony comprises at least 107 cells. In some embodiments, the cell colony is further modified, for example, by increasing or decreasing the expression and / or activity of at least one gene. In some embodiments, the gene encodes a member of an immunoproteasome. It is known that the immunoproteasome is involved in the processing of HLA class I binding peptides and includes LMP2 (β1i), MECL-1 (β2i), and LMP7 (β5i) subunits. The immunoproteasome can also be induced by interferon-γ. Therefore, in some embodiments, the cell colony can be contacted with one or more cytokines, growth factors, or other proteins. The cells can be stimulated with inflammatory cytokines such as interferon-γ, IL-10, IL-6, and / or TNF-α. The cell population may also be subjected to various environmental conditions, such as stress (heat stress, hypoxia, glucose starvation, DNA damaging agents, etc.). In some embodiments, the cells are contacted with one or more of chemotherapeutic drugs, radiotherapy, targeted therapy, or immunotherapy. Thus, the methods disclosed herein can be used to study the effects of various genes or conditions on HLA peptide processing and presentation. In some embodiments, the conditions used are selected to match the condition of the patient whose HLA-peptide population is to be identified.
[0356] A virus-based system (e.g., an adenovirus system, an adeno-associated virus (AAV) vector, a poxvirus, or a lentivirus) can be used to encode and express a single HLA allele of the present disclosure. Plasmids that can be used for adeno-associated virus, adenovirus, and lentivirus delivery have been previously described (see, e.g., U.S. Patent Nos. 6,955,808 and 6,943,019, and U.S. Patent Application No. 20080254008, incorporated herein by reference). In the vectors that can be used in the practice of the present disclosure, integration into the host genome of the cell can be achieved using retroviral gene transfer methods, generally resulting in long-term expression of the inserted transgene. In an exemplary embodiment, the retrovirus is a lentivirus. In addition, high transduction efficiencies have been observed in many different cell types and target tissues. The tropism of retroviruses can be changed by incorporating foreign envelope proteins and expanding the potential target population of target cells. Retroviruses can also be engineered to allow conditional expression of the inserted transgene so that only certain cell types are infected by the lentivirus. Cell type-specific promoters can be used to target expression in specific cell types. Lentiviral vectors are retroviral vectors (thus both lentiviral and retroviral vectors can be used in the practice of the present disclosure). In addition, lentiviral vectors are able to transduce or infect non-dividing cells and typically produce high viral titers.
[0357] The choice of retroviral gene transfer system can depend on the target tissue. Retroviral vectors consist of cis-acting long terminal repeats with a packaging capacity of up to 6-10 kb of foreign sequence. Minimal cis-acting LTRs are sufficient to replicate and package the vector, which is then used to integrate the desired nucleic acid into the target cell to provide permanent expression. Widely used retroviral vectors that can be used in the practice of the present disclosure include those based on murine leukemia virus (MuLV), gibbon ape leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof (see, e.g., Buchscher et al. (1992) J. Virol. 66:2731-2739; Johann et al. (1992) J. Virol. 66:1635-1640; Sommnerfelt et al. (1990) Virol. 176:58-59; Wilson et al. (1998) J. Virol. 63:2374-2378; Miller et al. (1991) J. Virol. 65:2220-2224; PCT / US94 / 05700). Additionally, useful in the practice of the present disclosure are minimal non-primate lentiviral vectors, such as lentiviral vectors based on equine infectious anemia virus (EIAV) (see, e.g., Balagaan, (2006) J Gene Med; 8:275-285, published online on November 21, 2005 in Wiley InterScience DOI: 10.1002 / jgm.845). The vector may have a cytomegalovirus (CMV) promoter driving target gene expression. Thus, the present disclosure relates to one or more vectors that can be used to practice the present disclosure: viral vectors, including retroviral vectors and lentiviral vectors.
[0358] Any HLA allele can be expressed in a cell population. In an exemplary embodiment, the HLA allele is an HLA class I allele. In some embodiments, the HLA class I allele is an HLA-A allele or an HLA-B allele. In some embodiments, the HLA allele is an HLA class II allele. The sequences of HLA class I and class II alleles can be found in the IPD-IMGT / HLA database. Exemplary HLA alleles include, but are not limited to, HLA-A*02:01, HLA-B*14:02, HLA-A*23:01, HLA-E*01:01, HLA-DRB*01:01, HLA-DRB*01:02, HLA-DRB*11:01, HLA-DRB*15:01, and HLA-DRB*07:01.
[0359] In some embodiments, the HLA allele is selected to correspond to the target genotype. In some embodiments, the HLA allele is a mutated HLA allele, which can be a non-naturally occurring allele or a naturally occurring allele in a sick patient. The methods disclosed herein have the further advantage of identifying HLA binding peptides for HLA alleles associated with various diseases and alleles present at low frequencies. Therefore, in some embodiments, the methods provided herein can identify HLA alleles, even if they are present in a population such as a Caucasian population at a frequency of less than 1%.
[0360] In some embodiments, the nucleic acid sequence encoding the HLA allele further comprises an affinity receptor tag that can be used to immunopurify the HLA protein. Suitable tags are well known in the art. In some embodiments, the affinity receptor tag is a polyhistidine tag, a polyhistidine-glycine tag, a polyarginine tag, a polyaspartic acid tag, a polycysteine tag, a polyphenylalanine, a c-myc tag, a herpes simplex virus glycoprotein D (gD) tag, a FLAG tag, a KT3 epitope tag, a tubulin epitope tag, a T7 gene 10 protein peptide tag, a streptavidin tag, a streptavidin binding peptide (SPB) tag, a Strep-tag, a Strep-tag II, an albumin binding protein (ABP) tag, an alkaline phosphatase (AP) tag, a blue tongue virus tag (B-tag), a calmodulin binding peptide (CBP) tag, a chloramphenicol acetyltransferase (CAT) tag, a choline Binding domain (CBD) tag, chitin binding domain (CBD) tag, cellulose binding domain (CBP) tag, dihydrofolate reductase (DHFR) tag, galactose binding protein (GBP) tag, maltose binding protein (MBP), glutathione-S-transferase (GST), Glu-Glu (EE) tag, human influenza hemagglutinin (HA) tag, horseradish peroxidase (HRP) tag, NE-tag, HSV tag, ketosteroid isomerase (KSI) tag, KT3 tag, LacZ tag, luciferase tag, NusA tag, PDZ domain tag, AviTag, calmodulin tag, E-tag, S-tag, SBP-tag, Softag 1, Softag 3, TC tag, VSV-tag, Xpress tag, Isopeptag, SpyTag, SnoopTag, Profinity eXact tag, Protein C tag, S1-tag, S-tag, biotin-carboxyl carrier protein (BCCP) tag, green fluorescent protein (GFP) tag, small ubiquitin-like modifier (SUMO) tag, tandem affinity purification (TAP) tag, HaloTag, Nus-tag, thioredoxin tag, Fc-tag, CYD tag, HPC tag, TrpE tag, ubiquitin tag, VSV-G epitope tag derived from vesicular stomatitis virus glycoprotein, or V5 tag derived from a small epitope (Pk) found on the P and V proteins of simian virus 5 (SV5) paramyxovirus. In some embodiments, the affinity receptor tag is an "epitope tag", which is a type of peptide tag that adds a recognizable epitope (antibody binding site) to the HLA protein to provide binding of the corresponding antibody, thereby allowing identification or affinity purification of the tagged protein. Non-limiting examples of epitope tags are protein A or protein G, which bind to IgG. In some embodiments, the affinity receptor tag comprises a biotin acceptor peptide (BAP) or a human influenza hemagglutinin (HA) peptide sequence.Many other tag moieties are known and contemplated by the skilled artisan and are contemplated herein.Any peptide tag may be used so long as it is capable of being expressed as a component of an affinity receptor-tagged HLA-peptide complex.
[0361] The methods provided herein include isolating HLA-peptide complexes from cells transfected or transduced by affinity pull-down of HLA constructs. In some embodiments, standard immunoprecipitation techniques known in the art can be used to separate the complexes with commercially available antibodies. The cells can be lysed first. HLA class I-peptide complexes can be separated using HLA class I-specific antibodies such as W6 / 32 antibodies, while HLA class II-peptide complexes can be separated using HLA class II-specific antibodies such as M5 / 114.15.2 monoclonal antibodies. In some embodiments, a single (or a pair of) HLA alleles are expressed as fusion proteins with peptide tags, and HLA-peptide complexes are separated using binding molecules that recognize the peptide tags.
[0362] The method further comprises isolating the peptide from the HLA-peptide complex and sequencing the peptide. The peptide is isolated from the complex by any method known to those skilled in the art, such as acid elution. Although any sequencing method can be used, in some embodiments, a method using mass spectrometry, such as liquid chromatography-mass spectrometry (LC-MS or LC-MS / MS, or HPLC-MS or HPLC-MS / MS) is used. These sequencing methods are well known to the skilled person and are reviewed in Medzihradszky KF and Chalkley RJ. Mass Spectrom Rev. 2015 Jan-Feb; 34(1):43-63.
[0363] In some embodiments, the cell population expresses one or more endogenous HLA alleles. In some embodiments, the cell population is an engineered cell population lacking one or more endogenous HLA class I alleles. In some embodiments, the cell population is an engineered cell population lacking endogenous HLA class I alleles. In some embodiments, the cell population is an engineered cell population lacking one or more endogenous HLA class II alleles. In some embodiments, the cell population is an engineered cell population lacking endogenous HLA class II alleles or an engineered cell population lacking endogenous HLA class I alleles and endogenous HLA class II alleles. In some embodiments, the cell population comprises cells that have been enriched or sorted, such as by fluorescence activated cell sorting (FACS). In some embodiments, fluorescence activated cell sorting (FACS) is used to sort cell populations. In some embodiments, the cell population is pre-FACS sorted for cell surface expression of HLA class I or class II or both HLA class I and class II. For example, FACS can be used to sort a population of cells for cell surface expression of HLA class I alleles, HLA class II alleles, or a combination thereof.
[0364] Methods for preparing personalized cancer vaccines
[0365] Once a specific mutation for cancer is identified, such that the mutation is present in the DNA of cancer cells but not in normal cells of the same human subject, and the mutation causes a change in one or more amino acids in the protein encoded by the DNA, the mutation can be a target for the host immune response. The natural immune response can be directed against the mutated protein, resulting in the destruction of cancer cells expressing the protein. Due to the natural tolerance response and immunocompromised environment in cancer tissues, immunotherapy is a clinical approach that attempts to enhance this immune response to exceed the body's tolerance and immunosuppression. Therefore, proteins or peptides containing the above mutations are suitable candidates for immunotherapy.
[0366] The mutated protein is taken up by professional phagocytes acting as antigen presenting cells (APCs), chopped up, and displayed on the cell surface as an antigen for T cell activation in an antigen presentation complex containing major histocompatibility complex (MHC) proteins. Human MHC proteins are called human leukocyte antigens, HLA. MHC proteins can be MHC class I or class II proteins, and some functional differences are attributed to the presentation of peptides by class I or class II MHC proteins (HLA class I and HLA class II proteins). A significant difference is that HLA class I-peptide complexes present antigens to cytotoxic CD8+T cells, while HLA class II peptide complexes can also activate CD4+T cells, resulting in a prolonged immune response. CD8+T cells are essential in the task of clearing diseased cells (such as infected cells or tumor cells) cell by cell. CD4+T cells have a more lasting impact after activation, the most important of which is the generation of immune memory. CD4 subsets are recruited differently depending on the type of immune threat, and multiple subsets with overlapping or different functions can be recruited together. This helps to balance the immune response associated with pathogen threats. In these aspects, HLA class I or class II peptide-mediated antigen presentation enables sustained and tailored immune responses. On the other hand, HLA class I or class II binding to peptides can be promiscuous, so nonspecific peptide binding and presentation to the immune system can lead to abnormal immune responses, such as autoimmunity.
[0367] In one aspect, the present disclosure provides a method for predicting a peptide that can accurately pair or bind to a specific HLA class I or class II molecule such that high-fidelity binding of the peptide to the HLA class I or class II protein ensures that the specific peptide is presented to T lymphocytes, thereby eliciting a specific immune response and avoiding any cross-reaction or immune promiscuity.
[0368] In one aspect, the present disclosure provides a method for predicting a peptide that can accurately bind to a specific HLA class I or class II protein, so that when the peptide is therapeutically administered to a subject expressing a specific homologous HLA class I or class II protein, the peptide can be used to activate a more sustained and powerful immune response by virtue of the ability of the HLA class I or class II protein to activate CD4+T cells and stimulate immune memory. In some embodiments, a given peptide predicted to bind to an HLA class I or class II protein with high specificity is a peptide comprising a mutation, wherein the mutation is prevalent in a subject's cancer or tumor cells; and the same HLA class I or class II protein predicted to bind to the mutant peptide either (a) does not bind, or (b) binds to the corresponding non-mutated wild-type peptide with a significantly lower affinity than the mutant peptide of the subject. The preferential binding of HLA to mutant peptides is advantageous in the development of immunotherapeutics because cells expressing wild-type peptides will be protected from immune attacks by T cells reactive to HLA-presented peptides. In some embodiments, the predicted peptide that specifically binds to an HLA class I or class II protein is a peptide with a post-translational modification. Exemplary post-translational modifications include, but are not limited to, phosphorylation, ubiquitination, dephosphorylation, glycosylation, methylation, or acetylation. In some embodiments, the predicted peptide is post-translationally modified prior to use in immunotherapy.
[0369] In some embodiments, the immunotherapy methods and strategies disclosed herein can also be applied to inhibit unwanted immune activation, such as in autoimmune reactions. Specifically, peptides identified as potential binders to specific HLA subtypes can be tailored to bind to specific HLA molecules and induce tolerance rather than cause an immunogenic response.
[0370] On the one hand, the present invention proposes a customized or personalized immunotherapy method for a specific subject. Each subject or patient expresses a specific set of HLA class I and HLA class II proteins. HLA typing is a well-known technique that allows the determination of a specific set of HLA proteins expressed by a subject. Once the HLA heterodimers expressed by a specific subject are known, having an improved, sophisticated and reliable method as described herein for predicting peptides that can bind to a specific HLA class I or class II complex with high fidelity can ensure that a specific immune response customized specifically for the subject can be generated.
[0371] The genes encoding HLA heterodimers are highly polymorphic, with more than 4,000 HLA class II allelic variants identified throughout the human population. For each of the HLA class II loci, an individual can inherit different alleles from maternal and paternal HLA haplotypes, and each HLA class II heterodimer consists of an α chain and a β chain. Due to the large number of α chain and β chain pairing combinations, especially for HLA-DP and HLA-DQ alleles, the population of possible HLA heterodimers is very complex. HLA class II heterodimers are translated in the endoplasmic reticulum (ER) and assembled into a stable complex with the invariant chain (Ii) derived from the protein CD74. Ii stabilizes the class II complex by allowing correct protein folding and enables HLA class II heterodimers to be exported into endosomal / lysosomal compartments. Within these HLA class II loading compartments, Ii is proteolytically cleaved by cathepsins into placeholder peptides known as CLIPs. CLIP is then exchanged for higher affinity peptides in a low pH environment by the molecular chaperone HLA-DM (a non-classical HLA class II heterodimer). The HLA class II complex loaded with high affinity peptides then reaches the trans-Golgi apparatus and finally the cell surface for display by CD4+ T cells.
[0372] Each HLA heterodimer is estimated to bind thousands of peptides with allele-specific binding preferences. In fact, each HLA allele is estimated to bind and present approximately 1,000-10,000 unique peptides to T cells. Given this diversity in HLA binding, accurate prediction of whether a peptide is likely to bind to a specific HLA allele is very challenging. Little is known about the allele-specific peptide binding properties of HLA class II molecules because of the heterogeneity of α and β chain pairing, the complexity of the data that limits the ability to confidently assign core binding epitopes, and the lack of allele-specific antibodies of the immunoprecipitation grade required for high-resolution biochemical analysis. Furthermore, analysis of peptide epitopes derived from a given HLA allele introduces uncertainty when multiple HLA alleles are presented on the cell surface.
[0373] Disclosed herein is a method for preparing a personalized cancer vaccine. The method for preparing a personalized cancer vaccine may include identifying a peptide sequence with a mutation expressed in a subject's cancer cells; inputting the amino acid position information of the identified peptide sequence into a machine learning HLA-peptide presentation prediction model using a computer processor to generate a set of presentation predictions for the identified peptide sequence, each presentation prediction representing the probability that one or more proteins encoded by the class I or class II MHC alleles of the subject's cancer cells will present a given sequence of the identified peptide sequence; and selecting a subset of the peptide sequences identified based on the set of presentation predictions for use in preparing a personalized cancer vaccine.
[0374] In some embodiments, one or more results obtained from the methods described herein may provide one or more quantitative values indicating one or more of the following: the likelihood of diagnostic accuracy, the likelihood of the presence of a condition in a subject, the likelihood of a subject developing a condition, the likelihood of a particular treatment being successful, or any combination thereof. In some embodiments, the methods described herein may predict the risk or likelihood of developing a condition. In some embodiments, the methods described herein may be an early diagnostic indication of the occurrence of a condition. In some embodiments, the methods described herein may confirm the diagnosis or presence of a condition. In some embodiments, the methods described herein may monitor the progression of a condition. In some embodiments, the methods described herein may monitor the efficacy of a treatment on a subject's condition.
[0375] Methods for identification of MHC-presenting peptides
[0376] On the one hand, a method for identifying one or more peptides presented by MHC proteins for immune activation is provided herein. In some embodiments, the one or more peptides comprise an epitope. In some embodiments, the method relates to a computational prediction of the likelihood that a particular epitope is presented by an MHC protein. In some embodiments, the method relates to a computational prediction of the specificity of an epitope for MHC presentation. In some embodiments, the computational prediction method relates to an assessment of peptide-MHC interactions. In some embodiments, the computational prediction method relates to predicting the allele specificity of a peptide for antigen presentation.
[0377] In some embodiments, the computational prediction method involves bioinformatics information, such as nucleotide sequence, structural motifs of biomolecules, protein-protein interaction characteristics and functional efficacy such as immunogenicity. In some embodiments, the computational prediction method involves machine learning. Based on machine learning methods, such as simple pattern motifs, support vector machines (SVM), hidden Markov models (HMM), neural network (NN) models, quantitative structure-activity relationship (QSAR) analysis, structure-based methods and biophysical methods, many immunoinformatics methods for predicting peptide-MHC interactions have been developed for MHC class I and class II. These methods can be divided into two categories, i.e., intra-allelic (allele-specific) and trans-allelic (pan-specific) methods. The intra-allelic method is trained on a limited set of experimental peptide binding data for a specific MHC molecule and is applied to predict peptides bound to the molecule. Due to the extreme polymorphism of MHC molecules, the existence of thousands of allelic variants, and the lack of sufficient experimental binding data, it is impossible to establish a prediction model for each allele. Therefore, cross-allele and universal methods such as NetMHCIIpan (Karosiene E et al., NetMHCIIpan-3.0, a common pan-specific MHC classII prediction method including all three human MHC class II isotypes, HLA-DR, HLA-DP and HLADQ. Immunogenetics (2013) 65 (10): 711–24) and TEPITOPEpan (Zhang L et al., TEPITOPEpan: extending TEPITOPE for peptide binding prediction covering over 700 HLA-DR molecules. PLoS One (2012) 7 (2): e30483) have been developed using peptide binding data that extend over many alleles or across species. Similar methods for MHC-I are also available, such as NetMHCpan and KISS.
[0378] In some embodiments, the ahe peptide sequence may not be expressed in normal cells of the subject. In some embodiments, each and every cell of the subject may not be a cancer cell. The cancer cells may be generated by different cancers, including but not limited to thyroid cancer, adrenocortical cancer, anal cancer, aplastic anemia, bile duct cancer, bladder cancer, bone cancer, bone metastasis, central nervous system (CNS) cancer, peripheral nervous system (PNS) cancer, breast cancer, Castleman disease, cervical cancer, childhood non-Hodgkin lymphoma, lymphoma, colorectal cancer, endometrial cancer, esophageal cancer, Ewing tumor family (e.g., Ewing sarcoma), eye cancer, gallbladder cancer, gastrointestinal carcinoid tumors, gastrointestinal stromal tumors, gestational trophoblastic disease, hairy cell leukemia, Hodgkin's disease, Kaposi's sarcoma, kidney cancer, laryngeal cancer and hypopharyngeal cancer, acute lymphocytic leukemia, acute myeloid leukemia, childhood leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, liver cancer, lung cancer, lung carcinoid tumor, non-Hodgkin lymphoma, male breast cancer, malignant mesothelioma, multiple myeloma, myelodysplastic syndrome, myeloproliferative disorders, nasal and paranasal cancer, nasopharyngeal cancer, neuroblastoma, oral cavity and oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, penile cancer, pituitary tumor, prostate cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, sarcoma (adult soft tissue cancer), melanoma skin cancer, non-melanoma skin cancer, stomach cancer, testicular cancer, thymic cancer, uterine cancer (such as uterine sarcoma), vaginal cancer, vulvar cancer, or Waldenstrom's macroglobulinemia.
[0379] The identification may include comparing the DNA, RNA or protein sequence from the subject's cancer cell with the DNA, RNA or protein sequence from the subject's normal cell. The DNA, RNA or protein sequence from the subject's cancer cell may be different from the DNA, RNA or protein sequence from the subject's normal cell. The identification may identify nucleic acid variants with high sensitivity.
[0380] The machine learning HLA-peptide presentation prediction model may include a plurality of predictor variables identified based at least on training data. The training data may include sequence information of sequences of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; training peptide sequence information including amino acid position information, wherein the training peptide sequence information is associated with HLA proteins expressed in cells; and a function representing the relationship between the amino acid position information received as input and the presentation likelihood generated as output based on the amino acid position information and the predictor variables.
[0381] In some embodiments, the training data may further include structured data, time series data, unstructured data, and relational data. Unstructured data may include audio data, image data, video, mechanical data, electrical data, chemical data, and any combination thereof, for accurately simulating or training a robot or simulation. Time series data may include data from one or more of a smart meter, smart appliance, smart device, monitoring system, telemetry device, or sensor. Relational data includes data from a client system, an enterprise system, an operating system, a website, a network accessible application program interface (API), or any combination thereof. This may be accomplished by a user by any method in a file or other data format input into a software or system.
[0382] In some embodiments, the training data can be stored in a database. The database can be stored in a computer-readable format. The computer processor can be configured to access data stored in a computer-readable memory. In some embodiments, a computer system can be used to analyze data to obtain a result. The result can be stored remotely or internally on a storage medium and transmitted to personnel such as a drug expert. In some embodiments, a computer system can be operably coupled with a component for transmitting the result. The component for transmission can include wired and wireless components. Examples of wired communication components can include universal serial bus (USB) connections, coaxial cable connections, Ethernet cables such as Cat5 or Cat6 cables, fiber optic cables or telephone lines. Examples or wireless communication components can include Wi-Fi receivers, components for accessing mobile data standards such as 3G or 4GLTE data signals, or Bluetooth receivers. In some embodiments, all these data in the storage medium are collected and archived to build a data warehouse.
[0383] In some embodiments, the database includes an external database. The external database can be a medical database, such as, but not limited to, an adverse drug reaction database, an AHFS supplement file, an allergen selection list file, an average WAC pricing file, a brand probability file, a Canadian drug file v2, a comprehensive price history, a controlled substance file, a drug allergy cross reference file, a drug application file, a drug administration and administration database, a drug image database v2.0 / drug imprint database v2.0, a drug expiration date file, a drug indication database, a drug laboratory conflict database, a drug therapy monitoring system (DTMS) v2.2 / DTMS consumer monograph, a duplicate treatment database, a federal government pricing file, a healthcare common procedure coding system code (HCPCS) database, an ICD-10 mapping file, an immunization cross reference file, a comprehensive A to Z drug fact module, a comprehensive patient education, a master parameter database, a half span electronic drug file (MED-File) v2, a Medicaid Rebate file, health care plan file, medical condition selection list file, medical condition master database, medication order management database (MOMD), monitoring parameter database, patient safety program file, payment limits-part B (PAF-B) v2.0, precautionary measures database, RxNorm cross-reference file, standard drug identifier database, substitution group file, supplemental name file, uniform system of classification cross-reference file, or warning label database.
[0384] In some embodiments, training data may also be obtained through other data sources. Data sources may include sensors or smart devices, such as appliances, smart meters, wearable devices, monitoring systems, data stores, customer systems, billing systems, financial systems, crowd source data, weather data, social networks, or any other sensor, enterprise system, or data store. Examples of smart meters or sensors may include meters or sensors located at customer locations, or meters or sensors located between customers and generation or source locations. By integrating data from a wide range of sources, the system may be able to perform complex and detailed analysis. In some embodiments, data sources may include, but are not limited to, sensors or databases for other medical platforms.
[0385] HLA typing is conventionally performed by serological methods using antibodies or by PCR-based methods such as sequence-specific oligonucleotide probe hybridization (SSOP) or sequence-based typing (SBT). The first is hampered by a potentially high degree of cross-reactivity and limited resolution, while the second has difficulties associated with PCR efficiency, since the possibilities for positioning primers are very limited due to the polymorphic positions.
[0386] In some embodiments, sequence information is identified by a sequencing method or a method using mass spectrometry, such as liquid chromatography-mass spectrometry (LC-MS or LC-MS / MS, or HPLC-MS or HPLC-MS / MS). These sequencing methods may be well known to the skilled person and are reviewed in Medzihradszky KF and Chalkley RJ. Mass Spectrom Rev. 2015 Jan-Feb; 34(1):43-63. In some embodiments, the mass spectrometry is monoallelic mass spectrometry. In some embodiments, the mass spectrometry may be MS analysis, MS / MS analysis, LC-MS / MS analysis, or a combination thereof. In some embodiments, MS analysis may be used to determine the mass of the intact peptide. For example, the determination may include determining the mass of the intact peptide (e.g., MS analysis). In some embodiments, MS / MS analysis may be used to determine the mass of a peptide fragment. For example, the determination may include determining the mass of a peptide fragment, which may be used to determine the amino acid sequence of a peptide or a portion thereof (e.g., MS / MS analysis). In some embodiments, the mass of a peptide fragment may be used to determine the amino acid sequence within the peptide. In some embodiments, LC-MS / MS analysis can be used to separate complex peptide mixtures. For example, the determination can include, for example, separation of complex peptide mixtures by liquid chromatography, and determination of the mass of the complete peptide, the mass of the peptide fragments, or a combination thereof (e.g., LC-MS / MS analysis). The data can be used, for example, for peptide sequencing.
[0387] In some embodiments, the training peptide sequence information includes amino acid position information of the training peptide. In some embodiments, the training peptide sequence information includes at most about 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10% or less of the sequence of the peptide presented by the HLA protein expressed in the cell and identified by mass spectrometry. In some embodiments, the training peptide sequence information may include at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or more of the sequence of the peptide presented by the HLA protein expressed in the cell and identified by mass spectrometry.
[0388] Any information and data can be paired with the subject as the source of the information and data. The subject or medical professional can retrieve information and data from the storage or server through the subject identity. The subject identity can include a patient's photo, name, address, social security number, birthday, telephone number, zip code, or any combination thereof. The subject identity can be encrypted and encoded into a visible graphic code. The visible graphic code can be a one-time barcode that can be uniquely associated with the subject identity. The barcode can be a UPC barcode, an EAN barcode, a Code 39 barcode, a Code 128 barcode, an ITF barcode, a CodaBar barcode, a GS1 DataBar barcode, an MSI Plessey barcode, a QR barcode, a Datamatrix code, a PDF417 code, or an Aztec barcode. The visible graphic code can be configured to be displayed on a display screen. The barcode can include a QR that can be optically captured and read by a machine. The barcode can define elements such as the version, format, position, alignment, or timing of the barcode to enable the reading and decoding of the barcode. Barcodes can encode various types of information, such as binary or alphanumeric information, in any type of suitable format. QR codes can have various symbol sizes, as long as the QR code can be scanned from a reasonable distance by an imaging device. QR codes can be in any image file format (such as EPS or SVG vector graphics, PNG, TIF, GIF or JPEG raster graphics formats).
[0389] In some embodiments, the function representing the relationship between the amino acid position information received as input and the presentation probability generated as output based on the amino acid position information and the predictor variable comprises a linear or nonlinear function. The function can be, for example, a rectified linear unit (ReLU) activation function, a Leaky ReLu activation function, or other functions, such as saturated hyperbolic tangent, identity, binary step function, logistic function, arcTan, softsign, parametric rectified linear unit, exponential linear unit, softPlus, bent identity, softExponential, sinusoid, Sinc, Gaussian or sigmoid function, or any combination thereof.
[0390] In some embodiments, the linear function is obtained by linear regression. In some embodiments, linear regression is a method for predicting a target variable by fitting the best linear relationship between a dependent variable and an independent variable. The best fit may mean that the sum of all distances between the shape at each point and the actual observed value is the smallest. Linear regression may include simple linear regression or multiple linear regression. Simple linear regression can use a single independent variable to predict the dependent variable. Multiple linear regression can use more than one independent variable to predict the dependent variable by fitting the best linear relationship. Nonlinear function can be obtained by nonlinear regression. Nonlinear regression can be a form of regression analysis, in which observed data is modeled by a function that is a nonlinear combination of model parameters and depends on one or more independent variables. Nonlinear regression may include step functions, piecewise functions, splines, and generalized additive models.
[0391] In some embodiments, the presentation possibility is presented by a one-dimensional value (e.g., probability). In some embodiments, probability is configured to measure the possibility that an event may occur. In some embodiments, the probability range is about 0 to 1, 0.1 to 0.9, 0.2 to 0.8, 0.3 to 0.7 or 0.4 to 0.6. The higher the probability of an event occurring, the more likely the event is to occur. In some embodiments, an event includes any type of situation, including, as a non-limiting example, whether HLA-peptides will present certain peptides with specific amino acid position information, and whether people will get sick based on amino acid position information. In some embodiments, the possibility can be presented by a multidimensional value. Multidimensional values can be presented by multidimensional space, heat map or electronic table.
[0392] In one embodiment, a subset of peptide sequences identified based on a group of presentation predictions is selected to be configured to prepare a personalized cancer vaccine. In some embodiments, the subset comprises up to about 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10% or less of peptide sequences identified based on the group of presentation predictions. In other cases, the subset may comprise at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or more of peptide sequences identified based on the group of presentation predictions. A cancer vaccine may be a vaccine for treating an existing cancer or preventing cancer development. The vaccine may be prepared by a sample taken from a patient and may be specific to the patient.
[0393] In some embodiments, poxviruses are used in disease (e.g., cancer) vaccines or immunogenic compositions. These include orthopoxviruses, avipox, vaccinia, MVA, NYVAC, canarypox, ALVAC, fowlpox, TROVAC, and the like. Advantages of vectors may include simple construction, the ability to accommodate large amounts of foreign DNA, and high expression levels. Regarding poxviruses that can be used to practice the present disclosure, such as Chordopoxvirinae poxviruses (poxviruses of vertebrates), e.g., orthopoxviruses and avipox viruses, such as vaccinia viruses (e.g., Wyeth strain, WR strain (e.g., ), Copenhagen strain, NYVAC, NYVAC.1, NYVAC.2, MVA, MVA-BN), canarypox virus (e.g., Wheatley C93 strain, ALVAC), fowlpox virus (e.g., FP9 strain, Webster strain, TROVAC), dovepox, pigeonpox, quailpox and raccoonpox, especially their synthetic or non-naturally occurring recombinants, their uses and methods for making and using such recombinants can be found in the scientific and patent literature.
[0394] In some embodiments, vaccinia virus is used in disease vaccines or immunogenic compositions to express antigens. Recombinant vaccinia virus may be able to replicate in the cytoplasm of infected host cells, so the polypeptide of interest can induce an immune response.
[0395] In some embodiments, ALVAC is used as a vector in a disease vaccine or immunogenic composition. ALVAC may be a canarypox virus that can be modified to express foreign transgenes and has been used as a vaccination method against prokaryotic and eukaryotic antigens.
[0396] In some embodiments, a modified vaccinia Ankara (MVA) virus is used as a viral vector for an antigen vaccine or immunogenic composition. MVA can be a member of the Orthopoxviridae family and has been produced by approximately 570 serial passages of the Ankara strain of vaccinia virus (CVA) in chicken embryo fibroblasts. As a result of these passages, the resulting MVA virus contains 31 kilobases less genomic information than CVA and is highly host cell restricted. MVA can be characterized by its extreme attenuation, i.e., reduced virulence or infectivity, but still excellent immunogenicity. When tested in a variety of animal models, MVA can be shown to be non-toxic, even in immunosuppressed individuals. In addition, It may be a candidate immunotherapy for the treatment of HER-2 positive breast cancer and is currently undergoing clinical trials.
[0397] In some embodiments, a positive predictive value (PPV) is used as part of a prediction model. PPV, also referred to as a precision measurement, is the probability that an individual diagnosed with a disease or condition by, for example, a test or model actually suffers from the disease or condition. It can be calculated by dividing the number of true positive results by the total number of positive results (including false positive results) returned. PPV = true positive / (true positive + false positive). For example, if in a group of 100 patients, the model determines a positive result in 50 patients, 25 of which are true positives, then the PPV will be 25 / 50 = 0.5. The closer the PPV is to 1, the more accurate the diagnostic method, such as a test or model, is. PPV can be used to determine the accuracy of a prediction model. PPV can be used to adjust the prediction model to accommodate the false positive results that the model may generate.
[0398] Recall can be used as part of a predictive model. Recall can be thought of as the percentage of true positive results to the total number of positives in a sample set. Recall = true positives / (true positives + false negatives). For example, if in a group of 100 patients, the model identified positive results in 50 patients, 25 of which were true positives, and there were a total of 75 positives in the group of patients, then the recall is {25 / (25+25)}x100=50%. Recall can be used to determine the accuracy of a predictive model. Recall can be used to adjust a predictive model to accommodate false positives or false negatives that the model may generate.
[0399] In some embodiments, the prediction model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall of 0.1%-10%. In some embodiments, the prediction model may have a positive predictive value of at most 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1 or less at a recall of 0.1%-10%. The prediction model may have a positive predictive value of at least 0.05, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall of less than 0.1%. In some embodiments, the prediction model can have a positive predictive value of at most 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1 or less at a recall rate below 0.1%. The prediction model can have a positive predictive value of at least 0.05, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall rate above 10%. In some embodiments, the prediction model can have a positive predictive value of at most 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1 or less at a recall rate above 10%.
[0400] In some embodiments, the prediction model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall of 0.1% to 10%. In some embodiments, the prediction model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall of 0.1% to 10%. In some embodiments, the prediction model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall of 0.1% to 10%. .5% to 4%, 0.5% to 5%, 0.5% to 6%, 0.5% to 7%, 0.5% to 8%, 0.5% to 9%, 0.5% to 10%, 1% to 2%, 1% to 3%, 1% to 4%, 1% to 5%, 1% to 6%, 1% to 7%, 1% to 8%, 1% to 9%, 1% to 10%, 2% to 3%, 2% to 4%, 2% to 5%, 2% to 6 %, 2% to 7%, 2% to 8%, 2% to 9%, 2% to 10%, 3% to 4%, 3% to 5%, 3% to 6%, 3% to 7%, 3% to 8%, 3% to 9%, 3% to 10%, 4% to 5%, 4% to 6%, 4% to 7%, 4% to 8%, 4% to 9%, 4% to 10%, 5% to 6%, 5% to 7%, 5% to 8%, 5% to 9%, 5% to 10%, In some embodiments, the prediction model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall of 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9% or 10%. In some embodiments, the prediction model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall of at least 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8% or 9%. In some embodiments, the prediction model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall of at most 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9% or 10%.
[0401] In some embodiments, the prediction model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall of 10% to 20%. In some embodiments, the prediction model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall of 10% to 20%. In some embodiments, the prediction model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall of 10% to 20%. 1% to 16%, 11% to 17%, 11% to 18%, 11% to 19%, 11% to 20%, 12% to 13%, 12% to 14%, 12% to 15%, 12% to 16%, 12% to 17%, 12% to 18%, 12% to 19%, 12% to 20%, 13% to 14%, 13% to 15%, 13% to 16%, 13% to 17%, 13% to 18%, 13% to 19%, 13% to 20%, 14% to 15%, 14% to 16%, 14% to 17%, 14% to 18%, 14% to 19%, 14% to 20%, 15% to 16%, 15% to 17%, 15% to 18%, 15% to 19%, 15% to 20%, 16% to 17%, 16% to 1 In some embodiments, the prediction model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall of 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19% or 20%. In some embodiments, the prediction model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall of at least 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, or 19%. In some embodiments, the prediction model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall of at most 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19% or 20%.
[0402] In some embodiments, the prediction model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall rate of at least 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19% or 20%. For example, the prediction model can have a positive predictive value of at least 0.1 at a recall rate of at least 10%. For example, the prediction model can have a positive predictive value of at least 0.2 at a recall rate of at least 10%. For example, the prediction model can have a positive predictive value of at least 0.3 at a recall rate of at least 10%. For example, the prediction model may have a positive predictive value of at least 0.4 at a recall rate of at least 10%. For example, the prediction model may have a positive predictive value of at least 0.5 at a recall rate of at least 10%. For example, the prediction model may have a positive predictive value of at least 0.6 at a recall rate of at least 10%. For example, the prediction model may have a positive predictive value of at least 0.7 at a recall rate of at least 10%. For example, the prediction model may have a positive predictive value of at least 0.8 at a recall rate of at least 10%. For example, the prediction model may have a positive predictive value of at least 0.9 at a recall rate of at least 10%. For example, the prediction model may have a positive predictive value of at least 0.1 at a recall rate of at least 5%. For example, the prediction model may have a positive predictive value of at least 0.2 at a recall rate of at least 5%. For example, the prediction model may have a positive predictive value of at least 0.3 at a recall rate of at least 5%. For example, the prediction model may have a positive predictive value of at least 0.4 at a recall rate of at least 5%. For example, the prediction model may have a positive predictive value of at least 0.5 at a recall rate of at least 5%. For example, the prediction model may have a positive predictive value of at least 0.6 at a recall rate of at least 5%. For example, the prediction model may have a positive predictive value of at least 0.7 at a recall rate of at least 5%. For example, the prediction model may have a positive predictive value of at least 0.8 at a recall rate of at least 5%. For example, the prediction model may have a positive predictive value of at least 0.9 at a recall rate of at least 5%. For example, the prediction model may have a positive predictive value of at least 0.1 at a recall rate of at least 20%. For example, the prediction model may have a positive predictive value of at least 0.2 at a recall rate of at least 20%. For example, the prediction model may have a positive predictive value of at least 0.3 at a recall rate of at least 20%. For example, the prediction model may have a positive predictive value of at least 0.4 at a recall rate of at least 20%. For example, the prediction model may have a positive predictive value of at least 0.5 at a recall rate of at least 20%. For example, the predictive model may have a positive predictive value of at least 0.6 at a recall of at least 20%.For example, the prediction model can have a positive predictive value of at least 0.7 at a recall of at least 20%. For example, the prediction model can have a positive predictive value of at least 0.8 at a recall of at least 20%. For example, the prediction model can have a positive predictive value of at least 0.9 at a recall of at least 20%.
[0403] In some embodiments, the prediction model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall rate of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19% or 20%. For example, the prediction model can have a positive predictive value of at least 0.1 at a recall rate of about 10%. For example, the prediction model can have a positive predictive value of at least 0.2 at a recall rate of about 10%. For example, the prediction model can have a positive predictive value of at least 0.3 at a recall rate of about 10%. For example, the prediction model may have a positive predictive value of at least 0.4 at a recall rate of about 10%. For example, the prediction model may have a positive predictive value of at least 0.5 at a recall rate of about 10%. For example, the prediction model may have a positive predictive value of at least 0.6 at a recall rate of about 10%. For example, the prediction model may have a positive predictive value of at least 0.7 at a recall rate of about 10%. For example, the prediction model may have a positive predictive value of at least 0.8 at a recall rate of about 10%. For example, the prediction model may have a positive predictive value of at least 0.9 at a recall rate of about 10%. For example, the prediction model may have a positive predictive value of at least 0.1 at a recall rate of about 5%. For example, the prediction model may have a positive predictive value of at least 0.2 at a recall rate of about 5%.
[0404] For example, a predictive model may have a positive predictive value of at least 0.3 at a recall of about 5%.
[0405] For example, a predictive model may have a positive predictive value of at least 0.4 at a recall of about 5%.
[0406] For example, a predictive model may have a positive predictive value of at least 0.5 at a recall of about 5%.
[0407] For example, the predictive model may have a positive predictive value of at least 0.6 at a recall of about 5%.
[0408] For example, the predictive model may have a positive predictive value of at least 0.7 at a recall of about 5%.
[0409] For example, the predictive model may have a positive predictive value of at least 0.8 at a recall of about 5%.
[0410] For example, the predictive model may have a positive predictive value of at least 0.9 at a recall of about 5%.
[0411] For example, a predictive model may have a positive predictive value of at least 0.1 at a recall of about 20%.
[0412] For example, a predictive model may have a positive predictive value of at least 0.2 at a recall of about 20%.
[0413] For example, a predictive model may have a positive predictive value of at least 0.3 at a recall of about 20%.
[0414] For example, a predictive model may have a positive predictive value of at least 0.4 at a recall of about 20%.
[0415] For example, a predictive model may have a positive predictive value of at least 0.5 at a recall of about 20%.
[0416] For example, a predictive model may have a positive predictive value of at least 0.6 at a recall of about 20%.
[0417] For example, a predictive model may have a positive predictive value of at least 0.7 at a recall of about 20%.
[0418] For example, the prediction model may have a positive predictive value of at least 0.8 at a recall of about 20%.
[0419] For example, a predictive model may have a positive predictive value of at least 0.9 at a recall of about 20%.
[0420] In some embodiments, the prediction model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall rate of less than 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19% or 20%. For example, the prediction model can have a positive predictive value of at least 0.1 at a recall rate of at most 10%. For example, the prediction model can have a positive predictive value of at least 0.2 at a recall rate of at most 10%. For example, the prediction model can have a positive predictive value of at least 0.3 at a recall rate of at most 10%. For example, the prediction model may have a positive predictive value of at least 0.4 at a recall rate of at most 10%. For example, the prediction model may have a positive predictive value of at least 0.5 at a recall rate of at most 10%. For example, the prediction model may have a positive predictive value of at least 0.6 at a recall rate of at most 10%. For example, the prediction model may have a positive predictive value of at least 0.7 at a recall rate of at most 10%. For example, the prediction model may have a positive predictive value of at least 0.8 at a recall rate of at most 10%. For example, the prediction model may have a positive predictive value of at least 0.9 at a recall rate of at most 10%. For example, the prediction model may have a positive predictive value of at least 0.1 at a recall rate of at most 5%. For example, the prediction model may have a positive predictive value of at least 0.2 at a recall rate of at most 5%. For example, the prediction model may have a positive predictive value of at least 0.3 at a recall rate of at most 5%. For example, the prediction model may have a positive predictive value of at least 0.4 at a recall rate of at most 5%. For example, the prediction model may have a positive predictive value of at least 0.5 at a recall rate of at most 5%. For example, the prediction model may have a positive predictive value of at least 0.6 at a recall rate of at most 5%. For example, the prediction model may have a positive predictive value of at least 0.7 at a recall rate of at most 5%. For example, the prediction model may have a positive predictive value of at least 0.8 at a recall rate of at most 5%. For example, the prediction model may have a positive predictive value of at least 0.9 at a recall rate of at most 5%. For example, the prediction model may have a positive predictive value of at least 0.1 at a recall rate of at most 20%. For example, the prediction model may have a positive predictive value of at least 0.2 at a recall rate of at most 20%. For example, the prediction model may have a positive predictive value of at least 0.3 at a recall rate of at most 20%. For example, the prediction model can have a positive predictive value of at least 0.4 at a recall rate of at most 20%. For example, the prediction model can have a positive predictive value of at least 0.5 at a recall rate of at most 20%. For example, the prediction model can have a positive predictive value of at least 0.6 at a recall rate of at most 20%.For example, the prediction model can have a positive predictive value of at least 0.7 at a recall rate of at most 20%. For example, the prediction model can have a positive predictive value of at least 0.8 at a recall rate of at most 20%. For example, the prediction model can have a positive predictive value of at least 0.9 at a recall rate of at most 20%.
[0421] In some embodiments, the predictive model has a positive predictive value of 0.05% to 0.6% at a recall of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%. At a recall of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%, the predictive model can have a recall of 0.05% to 0.1%, 0.05% to 0.15%, 0.05% to 0.2%, 0.05% to 0.25%, 0.05% to 0.3%, 0.05% to 0.35%, 0.05% to 0.4%, 0.05% to 0.45%, 0.05% to 0.5%, 0.05% to 0.6%, 0.05% to 0.7%, 0.05% to 0.8%, 0.05% to 0.9%, 0.06% to 0.10%, 0.06% to 0.11%, 0.06% to 0.12%, 0.06% to 0.13%, 0.06% to 0.14%, 0.06% to 0.15%, 55%, 0.05% to 0.6%, 0.1% to 0.15%, 0.1% to 0.2%, 0.1% to 0.25%, 0.1% to 0.3%, 0.1% to 0.35%, 0.1% to 0.4%, 0.1% to 0.45%, 0.1% to 0.5%, 0.1% to 0.55%, 0.1% to 0.6%, 0.15% to 0.2%, 0.15% to 0.25%, 0.15% to 0.3%, 0.15% to 0.35%, 0.15% to 0.4%, 0.15% to 0.45%, 0.15% to 0.5%, 0.15% to 0 .55%, 0.15% to 0.6%, 0.2% to 0.25%, 0.2% to 0.3%, 0.2% to 0.35%, 0.2% to 0.4%, 0.2% to 0.45%, 0.2% to 0.5%, 0.2% to 0.55%, 0.2% to 0.6%, 0.25% to 0.3%, 0.25% to 0.35%, 0.25% to 0.4%, 0.25% to 0.45%, 0.25% to 0.5%, 0.25% to 0.55%, 0.25% to 0.6%, 0.3% to 0.35%, 0.3% to 0.4%, 0.3% to 0 .45%, 0.3% to 0.5%, 0.3% to 0.55%, 0.3% to 0.6%, 0.35% to 0.4%, 0.35% to 0.45%, 0.35% to 0.5%, 0.35% to 0.55%, 0.35% to 0.6%, 0.4% to 0.45%, 0.4% to 0.5%, 0.4% to 0.55%, 0.4% to 0.6%, 0.45% to 0.5%, 0.45% to 0.55%, 0.45% to 0.6%, 0.5% to 0.55%, 0.5% to 0.6%, or 0.55% to 0.6%.At a recall of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%, the predictive model can have a positive predictive value of 0.05%, 0.1%, 0.15%, 0.2%, 0.25%, 0.3%, 0.35%, 0.4%, 0.45%, 0.5%, 0.55%, or 0.6%. At a recall of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%, the predictive model can have a positive predictive value of at least 0.05%, 0.1%, 0.15%, 0.2%, 0.25%, 0.3%, 0.35%, 0.4%, 0.45%, 0.5%, or 0.55%. At a recall of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%, the predictive model can have a positive predictive value of at most 0.1%, 0.15%, 0.2%, 0.25%, 0.3%, 0.35%, 0.4%, 0.45%, 0.5%, 0.55%, or 0.6%.
[0422] At a recall of approximately 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%, the predictive model can have a positive predictive value of 0.45% to 0.98%. At a recall of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%, the predictive model can have a recall of 0.45% to 0.5%, 0.45% to 0.55%, 0.45% to 0.6%, 0.45% to 0.65%, 0.45% to 0.7%, 0.45% to 0.75%, 0.45% to 0.8%, 0.45% to 0.85%, 0.45% to 0.9%, 0.45% to 0.96%, or 0.5% to 0.5%. %, 0.45% to 0.98%, 0.5% to 0.55%, 0.5% to 0.6%, 0.5% to 0.65%, 0.5% to 0.7%, 0.5% to 0.75%, 0.5% to 0.8%, 0.5% to 0.85%, 0.5% to 0.9%, 0.5% to 0.96%, 0.5% to 0.98%, 0.55% to 0.6%, 0.55% to 0.65%, 0.55% to 0.7%, 0.55% to 0.75%, 0.55% to 0.8%, 0.55% to 0.85%, 0.55% to 0.9%, 0.55% to 0.96 %, 0.55% to 0.98%, 0.6% to 0.65%, 0.6% to 0.7%, 0.6% to 0.75%, 0.6% to 0.8%, 0.6% to 0.85%, 0.6% to 0.9%, 0.6% to 0.96%, 0.6% to 0.98%, 0.65% to 0.7%, 0.65% to 0.75%, 0.65% to 0.8%, 0.65% to 0.85%, 0.65% to 0.9%, 0.65% to 0.96%, 0.65% to 0.98%, 0.7% to 0.75%, 0.7% to 0.8%, 0.7% to 0.85 %, 0.7% to 0.9%, 0.7% to 0.96%, 0.7% to 0.98%, 0.75% to 0.8%, 0.75% to 0.85%, 0.75% to 0.9%, 0.75% to 0.96%, 0.75% to 0.98%, 0.8% to 0.85%, 0.8% to 0.9%, 0.8% to 0.96%, 0.8% to 0.98%, 0.85% to 0.9%, 0.85% to 0.96%, 0.85% to 0.98%, 0.9% to 0.96%, 0.9% to 0.98%, or 0.96% to 0.98%.At a recall of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%, the predictive model can have a positive predictive value of 0.45%, 0.5%, 0.55%, 0.6%, 0.65%, 0.7%, 0.75%, 0.8%, 0.85%, 0.9%, 0.96%, or 0.98%. At a recall of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%, the predictive model can have a positive predictive value of at least 0.45%, 0.5%, 0.55%, 0.6%, 0.65%, 0.7%, 0.75%, 0.8%, 0.85%, 0.9%, or 0.96%. At a recall of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%, the predictive model can have a positive predictive value of at most 0.5%, 0.55%, 0.6%, 0.65%, 0.7%, 0.75%, 0.8%, 0.85%, 0.9%, 0.96%, or 0.98%.
[0423] Methods for training machine learning HLA-peptide presentation prediction models
[0424] In one aspect, a method for training a machine learning HLA-peptide presentation prediction model may include using a computer processor to input into the HLA-peptide presentation prediction model a sequence of amino acid position information of an HLA-peptide isolated from one or more HLA-peptide complexes of a cell expressing an HLA class I or class II allele; training the machine learning HLA-peptide presentation prediction model may include adjusting weight values on the neural network nodes to best match the provided training data.
[0425] The training data may include sequence information of the sequence of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; training peptide sequence information including amino acid position information of training peptides, wherein the training peptide sequence information is associated with HLA proteins expressed in cells; and a function representing the relationship between amino acid position information received as input and presentation probability generated as output based on the amino acid position information and predictor variables. Training data, training peptide sequence information, function, and presentation probability are disclosed elsewhere herein.
[0426] The trained algorithm may include one or more neural networks. A neural network may be a type of computing system based on a graph of several connected neurons (or nodes) in a series of layers. A neural network may include an input layer to which data is presented; one or more internal and / or "hidden" layers; and an output layer from which results are presented. A neural network may learn the relationship between an input data set and a target data set by adjusting a series of connection weights. Neurons may be connected to neurons in other layers through connections with weights, which are parameters that control the strength of the connections. The number of neurons in each layer may be related to the complexity of the problem to be solved. The minimum number of neurons required in a layer may be determined by the complexity of the problem, and the maximum number may be limited by the generalization ability of the neural network. The input neurons may receive the presented data and then transmit the data to the nodes in the first hidden layer through the connection weights, which are modified during training. The result node may sum the products of all input pairs and their associated weights. The weighted sum may be offset by a bias to adjust the value of the result node. The output of a node or neuron may be gated using a threshold or an activation function. The activation function may be a linear or nonlinear function. The activation function can be, for example, a rectified linear unit (ReLU) activation function, a leaky ReLu activation function, or other functions such as saturated hyperbolic tangent, identity, binary step function, logistic function, arcTan, softsign, parametric rectified linear unit, exponential linear unit, softPlus, bent identity, softExponential, sinusoid, Sinc, Gaussian or sigmoid function, or any combination thereof.
[0427] Hidden layers in a neural network can process data and transmit their results to the next layer through a second set of weighted connections. Each subsequent layer can "pool" the results from previous layers into more complex relationships. Neural networks can be trained using known sets of examples of training data (data collected from one or more sensors) by causing them to modify themselves during (and after) training to provide a desired output, such as an output value, from a given set of inputs. Trained algorithms can include convolutional neural networks, recurrent neural networks, dilated convolutional neural networks, fully connected neural networks, deep generative models, and Boltzmann machines.
[0428] One or more sets of training data may be used during a training phase to "teach" or "learn" weighting factors, preference values, and thresholds or other computational parameters of a neural network. For example, input data from a training data set and a gradient descent or back-propagation method may be used to train the parameters so that the output values from the neural network are consistent with the examples included in the training data set.
[0429] The number of nodes used in the input layer of the neural network can be at least about 10, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000 or more. In other cases, the number of nodes used in the input layer can be at most about 100,000, 90,000, 80,000, 70,000, 60,000, 50,000, 40,000, 30,000, 20,000, 10,000, 9000, 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 50, or 10 less. In some cases, the total number of layers used in the neural network (including input and output layers) can be at least about 3, 4, 5, 10, 15, 20, or more. In other cases, the total number of layers can be at most about 20, 15, 10, 5, 4, 3, or less.
[0430] In some cases, the total number of learnable or trainable parameters (e.g., weighting factors, preferences, or thresholds) used in a neural network can be at least about 10, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, or more. In other cases, the number of learnable parameters may be at most about 100,000, 90,000, 80,000, 70,000, 60,000, 50,000, 40,000, 30,000, 20,000, 10,000, 9000, 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1000, 900, 800, 700, 600, 500, 400, 300, 200, 1000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 50, or 10 or less.
[0431] The neural network may include a convolutional neural network. The convolutional neural network may include one or more convolutional layers, expansion layers, or fully connected layers. The number of convolutional layers may be 1-10, and the number of expansion layers may be 0-10. The total number of convolutional layers (including input layers and output layers) may be at least about 1, 2, 3, 4, 5, 10, 15, 20 or more, and the total number of expansion layers may be at least about 1, 2, 3, 4, 5, 10, 15, 20 or more. The total number of convolutional layers may be at most about 20, 15, 10, 5, 4, 3 or less, and the total number of expansion layers may be at most about 20, 15, 10, 5, 4, 3 or less. In some embodiments, the number of convolutional layers is 1-10, and the number of fully connected layers is 0-10. The total number of convolutional layers (including input and output layers) can be at least about 1, 2, 3, 4, 5, 10, 15, 20, or more, and the total number of fully connected layers can be at least about 1, 2, 3, 4, 5, 10, 15, 20, or more. The total number of convolutional layers can be at most about 20, 15, 10, 5, 4, 3, or less, and the total number of fully connected layers can be at most about 20, 15, 10, 5, 4, 3, or less.
[0432] A convolutional neural network (CNN) can be a deep and feed-forward artificial neural network. CNN can be suitable for analyzing visual images. CNN can contain an input layer, an output layer, and multiple hidden layers. The hidden layers of a CNN can include convolutional layers, pooling layers, fully connected layers, and normalization layers. Layers can be organized in 3 dimensions: width, height, and depth.
[0433] A convolutional layer can apply a convolution operation to the input and pass the result of the convolution operation to the next layer. For processing images, the convolution operation can reduce the number of free parameters, allowing the network to become deeper with fewer parameters. In a convolutional layer, neurons can receive input only from restricted partitions of the previous layer. The parameters of a convolutional layer can include a set of learnable filters (or kernels). Learnable filters can have small receptive fields and extend to the entire depth of the input volume. During the forward pass, each filter can be convolved over the width and height of the input volume, calculate the dot product between the filter entries and the input, and generate a two-dimensional activation map for that filter. As a result, the network can learn filters that activate when a certain type of feature is detected at a certain spatial location in the input.
[0434] The pooling layer may include a global pooling layer. The global pooling layer may combine the outputs of a layer of neuron clusters into a single neuron in the next layer. For example, a max pooling layer may use the maximum value from each of the neuron clusters in the previous layer; and an average pooling layer may use the average value from each of the neuron clusters in the previous layer. A fully connected layer may connect each neuron in one layer to each neuron in another layer. In a fully connected layer, each neuron may receive input from each element of the previous layer. The normalization layer may be a batch normalization layer. The batch normalization layer may improve the performance and stability of the neural network. The batch normalization layer may provide any layer in the neural network with input as zero mean / unit variance. The advantages of using a batch normalization layer may include faster training networks, higher learning rates, easier initialization of weights, more feasible activation functions, and a simpler process for creating deep networks.
[0435] Neural networks may include recurrent neural networks. Recurrent neural networks may be configured to receive sequential data as input, such as continuous data input, and the recurrent neural network software module may update internal states at each time step. Recurrent neural networks may use internal states (memory) to process input sequences. Recurrent neural networks may be suitable for tasks such as handwriting recognition or speech recognition, next word prediction, music composition, image description generation, time series anomaly detection, machine translation, scene tagging, and stock market prediction. Recurrent neural networks may include fully recurrent neural networks, independent recurrent neural networks, Elman networks, Jordan networks, echo states, neural history compressors, long short-term memory, gated recurrent units, multi-timescale models, neural Turing machines, differentiable neural computers, neural network push-down automata, or any combination thereof.
[0436] The trained algorithm may include supervised or unsupervised learning methods, such as SVM, random forest, clustering algorithm (or software module), gradient boosting, logistic regression and / or decision tree. A supervised learning algorithm may be an algorithm that relies on using a set of labeled, paired training data examples to infer the relationship between input data and output data. An unsupervised learning algorithm may be an algorithm for inferring output data from a training data set. An unsupervised learning algorithm may include cluster analysis, which may be used for exploratory data analysis to discover hidden patterns or groupings in process data. An example of an unsupervised learning method may include principal component analysis. Principal component analysis may include reducing the dimensionality of one or more variables. The dimensionality of a given variable may be at least 1, 5, 10, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200 1300, 1400, 1500, 1600, 1700, 1800 or more. The dimensionality of a given variable may be at most 1800, 1600, 1500, 1400, 1300, 1200, 1100, 1000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 50, 10, or less.
[0437] The training algorithm can be obtained by statistical techniques. In some embodiments, the statistical techniques may include linear regression, classification, resampling methods, subset selection, shrinkage, dimensionality reduction, nonlinear models, tree-based methods, support vector machines, unsupervised learning, or any combination thereof.
[0438] Linear regression can be a method of predicting a target variable by fitting the best linear relationship between the dependent variable and the independent variables. The best fit may mean that the sum of all distances between the shape at each point and the actual observed value is minimized. Linear regression can include simple linear regression and multiple linear regression. Simple linear regression can use a single independent variable to predict the dependent variable. Multiple linear regression can use more than one independent variable to predict the dependent variable by fitting the best linear relationship.
[0439] Classification can be a data mining technique that assigns classes to a data set to enable accurate prediction and analysis. Classification techniques can include logistic regression and discriminant analysis. Logistic regression can be used when the dependent variable is dichotomous (binary). Logistic regression can be used to discover and describe the relationship between a binary dependent variable and one or more independent variables at the nominal, ordinal, interval, or ratio level. Resampling can be a method that includes drawing repeated samples from an original data sample. Resampling may not involve using a common distribution table to calculate approximate probability values. Resampling can generate a unique sampling distribution based on the actual data. In some embodiments, resampling can use experimental methods rather than analytical methods to generate a unique sampling distribution. Resampling techniques can include bootstrapping and cross-validation. Bootstrapping can be performed by sampling with replacement from the original data and using the "not selected" data points as test cases. Cross-validation can be performed by dividing the training data into multiple parts.
[0440] Subset selection can identify a subset of predictors that are related to the response. Subset selection can include best subset selection, forward stepwise selection, backward stepwise selection, hybrid methods, or any combination thereof. In some embodiments, shrinkage fitting involves a model with all the predictor variables, but the estimated coefficients shrink towards zero relative to the least squares estimates. This shrinkage can reduce the variance. Shrinkage can include ridge regression and lasso. Dimensionality reduction can simplify the problem of estimating n + 1 coefficients to a simpler problem of estimating m + 1 coefficients, where m < n. It can be obtained by computing n different linear combinations or projections of the variables. Then these n projections are used as predictor variables to fit a linear regression model by least squares. Dimensionality reduction can include principal component regression and partial least squares. Principal component regression can be used to derive a set of low-dimensional features from a large set of variables. The principal components used in principal component regression can capture the maximum variance in the data using linear combinations of the data in subsequent orthogonal directions. Partial least squares can be a supervised alternative to principal component regression because partial least squares can utilize the response variable to identify new features.
[0441] Nonlinear regression can be a form of regression analysis in which the observed data are modeled by a function that is a nonlinear combination of the model parameters and depends on one or more independent variables. Nonlinear regression can include step functions, piecewise functions, splines, generalized additive models, or any combination thereof.
[0442] Tree-based methods can be used for both regression and classification problems. Regression and classification problems may involve stratifying or partitioning the predictor space into multiple simple regions. Tree-based methods may include bagging, boosting, random forests, or any combination thereof. Bagging can reduce the variance of predictions by generating additional data for training from the original data set using repeated combinations to produce multiple steps with the same carnality / size as the original data. Boosting can calculate the output using several different models and then average the results using a weighted average method. The random forest algorithm can draw random bootstrap samples of the training set. Support vector machines can be a classification technique. Support vector machines can include finding a hyperplane that best separates two classes of points with the largest margin. Support vector machines can constrain the optimization problem so that the margin is maximized, subject to the constraint that it perfectly classifies the data.
[0443] An unsupervised method may be a method that draws inferences from a data set containing input data without labeled responses. Unsupervised methods may include clustering, principal component analysis, k-means clustering, hierarchical clustering, or any combination thereof.
[0444] The mass spectrometry can be a monoallelic mass spectrometry. In some embodiments, the mass spectrometry can be MS analysis, MS / MS analysis, LC-MS / MS analysis or a combination thereof. In some embodiments, the mass of the complete peptide can be determined using MS analysis. For example, the determination can include determining the mass of the complete peptide (for example, MS analysis). In some embodiments, the mass of the peptide fragment can be determined using MS / MS analysis. For example, the determination can include determining the mass of the peptide fragment, which can be used to determine the amino acid sequence of the peptide or its part (for example, MS / MS analysis). In some embodiments, the mass of the peptide fragment can be used to determine the amino acid sequence in the peptide. In some embodiments, LC-MS / MS analysis can be used to separate complex peptide mixtures. For example, the determination can include, for example, separating complex peptide mixtures by liquid chromatography, and determining the mass of the complete peptide, the mass of the peptide fragment or a combination thereof (for example, LC-MS / MS analysis). The data can be used for example peptide sequencing.
[0445] The peptides can be presented by HLA proteins expressed in cells through autophagy. Autophagy allows orderly degradation and recycling of cellular components. Autophagy can include macroautophagy, microautophagy and chaperone-mediated autophagy. The peptides can be presented by HLA proteins expressed in cells through phagocytosis. Phagocytosis can be the main mechanism for removing pathogens and cell debris. For example, when macrophages ingest pathogenic microorganisms, the pathogen is trapped in the phagosome, and then the phagosome fuses with the lysosome to form a phagolysosome. In HLA class II, phagocytic cells such as macrophages and immature dendritic cells can ingest entities through phagocytosis in the phagosome-although B cells show more general endocytosis in the endosome-endosome fuses with the lysosome, and the acidic enzymes of the lysosomes break the ingested proteins into many different peptides.
[0446] The quality of training data can be improved by using multiple quality metrics. The multiple quality metrics may include common pollutant peptide removal, high scoring peak intensity, high scoring and high mass accuracy. The peak intensity of the score can be used before scoring. The MS / MS search first uses a simple filter to screen the MS / MS spectrum for the candidate sequence. The filter can be the minimum scoring peak intensity. Once a sufficient number of spectral peaks are checked and it is found that it does not meet the threshold established by the filter, the search speed can be improved by quickly and generally rejecting the candidate sequence using the scoring peak intensity. The scoring peak intensity can be at least 50%. The scoring peak intensity can be at least 70%. The scoring peak intensity can be at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or more. In some cases, the scoring peak intensity can be at most 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10% or less. The score can be at least 7. The score can be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or more. In some cases, the score can be at most about 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1 or less. The mass accuracy can be at most 5 ppm. The mass accuracy can be at most 10 ppm, 9 ppm, 8 ppm, 7 ppm, 6 ppm, 5 ppm, 4 ppm, 3 ppm, 2 ppm, 1 ppm or less. The mass accuracy can be at least 1 ppm, 2 ppm, 3 ppm, 4 ppm, 5 ppm, 6 ppm, 7 ppm, 8 ppm, 9 ppm, 10 ppm or more.
[0447] In some embodiments, the mass accuracy is at most 2 ppm. In some embodiments, the backbone cleavage score is at least 5. In some embodiments, the backbone cleavage score is at least 8.
[0448] The peptides presented by the HLA proteins expressed in the cell can be peptides presented by a single immunoprecipitated HLA protein expressed in the cell. Immunoprecipitation (IP) can be a technique for precipitating protein antigens from a solution using antibodies that specifically bind to a particular protein. This process can be used to separate and concentrate a particular protein from a sample containing thousands of different proteins. Immunoprecipitation may require coupling the antibody to a solid substrate at a certain point in the process.
[0449] The peptide presented by the HLA protein expressed in the cell can be a peptide presented by a single exogenous HLA protein expressed in the cell. A single exogenous HLA protein can be produced by introducing one or more exogenous peptides into a cell population. In some embodiments, the introduction includes contacting the cell population with the one or more exogenous peptides or expressing the one or more exogenous peptides in the cell population. In some embodiments, the introduction includes contacting the cell population with one or more nucleic acids encoding the one or more exogenous peptides. In some embodiments, the one or more nucleic acids encoding the one or more peptides are DNA. In some embodiments, the one or more nucleic acids encoding the one or more peptides are RNA, optionally wherein the RNA is mRNA. In some embodiments, the enrichment does not include the use of tetramer (or polymer) reagents.
[0450] The peptide presented by the HLA protein expressed in the cell can be a peptide presented by a single recombinant HLA protein expressed in the cell. The recombinant HLA protein can be encoded by a recombinant HLA class I or HLA class II allele. The HLA class I can be selected from HLA-A, HLA-B, HLA-C. The HLA class I can be a non-classical class Ib group. The HLA class I can be selected from HLA-E, HLA-F and HLA-G. The HLA class I can be a non-classical class Ib group selected from HLA-E, HLA-F and HLA-G. In some embodiments, the HLA class II comprises an HLA class II alpha chain, an HLA class II beta chain, or a combination thereof.
[0451] The multiple predictors may include peptide-HLA affinity predictors. The multiple predictors may include source protein expression level predictors. The source protein expression level may be the expression level of the source protein of the intracellular peptide. In some embodiments, the expression level may be determined by measuring the amount of the source protein or the amount of the RNA encoding the source protein. The multiple predictors may include peptide sequence, amino acid physical properties, peptide physical properties, the expression level of the source protein of the intracellular peptide, protein stability, protein translation rate, ubiquitination site, protein degradation rate, translation efficiency from ribosome profiling, protein cleavage, protein localization, the motif of the host protein that promotes TAP transport, the host protein that experiences autophagy, the motif that is conducive to ribosome pause (e.g., polyproline or polylysine segment), the protein feature that is conducive to NMD (e.g., long 3'UTR, the last exon: exon connection upstream>50nt stop codon and peptide cleavage).
[0452] The multiple predictors may include peptide cleavage predictors. Peptide cleavage may be associated with a cleavable linker or cleavage sequence. In some embodiments, the cleavable linker is a ribosome jump site or an internal ribosome entry site (IRES) element. In some embodiments, when expressed in a cell, the ribosome jump site or IRES is cut. In some embodiments, the ribosome jump site is selected from F2A, T2A, P2A and E2A. In some embodiments, the IRES element is selected from common cell or viral IRES sequences. A cleavage sequence, such as F2A, or an internal ribosome entry site (IRES), may be placed between the α chain and the β2-microglobulin (HLA class I), or between the α chain and the β chain (HLA class II). In some embodiments, the single HLA class I allele is HLA-A*02:01, HLA-A*23:01 and HLA-B*14:02 or HLA-E*01:01, and the HLA class II allele is HLA-DRB*01:01, HLA-DRB*01:02 and HLA-DRB*11:01, HLA-DRB*15:01 or HLA-DRB*07:01. In some embodiments, the cleavage sequence is a T2A, P2A, E2A or F2A sequence. For example, the cleavage sequence can be EGRGSLTCG DV ENPGP (SEQ ID NO:6) (T2A), ATNF SL KQAGDV EN PGP (SEQ ID NO:7) (P2A), QCTNYALKLAGDVE SN PGP (SEQ ID NO:8) (E2A), or VKQTLNFDLKLAGD VE SNPGP (SEQ ID NO:9) (F2A).
[0453] In some embodiments, the cleavage sequence can be a thrombin cleavage site CLIP.
[0454] The peptides presented by the HLA protein may include peptides identified by searching a non-enzyme-specific non-modified peptide database. The peptide database may be a non-enzyme-specific peptide database, such as a non-modified database or a modified (e.g., phosphorylated or cysteine) database. In some embodiments, the peptide database is a polypeptide database. In some embodiments, the polypeptide database may be a protein database. In some embodiments, the method further comprises searching the peptide database using a reverse database search strategy. In some embodiments, the method further comprises searching the protein database using a reverse database search strategy. In some embodiments, a de novo search is performed, for example, to find new peptides not included in a normal peptide or protein database. The peptide database may be generated by the following method: providing a first and a second cell population, each cell population comprising one or more cells comprising an affinity receptor-tagged HLA, wherein the sequence affinity receptor-tagged HLA comprises different recombinant polypeptides encoded by different HLA alleles operably linked to an affinity receptor peptide; enriching affinity receptor-tagged HLA-peptide complexes; characterizing peptides or portions thereof bound to the affinity receptor-tagged HLA-peptide complexes from the enrichment; and generating an HLA allele-specific peptide database.
[0455] Peptides presented by HLA proteins may include peptides identified by comparing the MS / MS spectrum of the HLA-peptide with the MS / MS spectra of one or more HLA-peptides in a peptide database.
[0456] Mutation may exist on the nucleic acid of peptide or coding peptide.The mutation may be selected from point mutation, splice site mutation, frameshift mutation, read-through mutation and gene fusion mutation.The point mutation may be a gene mutation in which a single nucleotide base is changed, inserted or deleted from a DNA or RNA sequence.The splice site mutation may be a gene mutation in which a specific site where splicing occurs during the processing of precursor messenger RNA into mature messenger RNA is inserted, deleted or changed many nucleotides.The frameshift mutation may be a gene mutation caused by the insertion or deletion (insertion or deletion) of many nucleotides in a DNA sequence that is not divisible by three.The mutation may also include insertion, deletion, substitution mutation, gene duplication, chromosome translocation and chromosome inversion.
[0457] In some embodiments, the HLA class II protein comprises an HLA-DR protein.
[0458] In some embodiments, the HLA class II protein comprises an HLA-DP protein.
[0459] In some embodiments, the HLA class II protein comprises an HLA-DQ protein.
[0460] In some embodiments, the HLA class II protein can be selected from HLA-DR and HLA-DP or HLA-DQ proteins. In some embodiments, the HLA protein is an HLA class II protein selected from the group consisting of HLA-DPB1*01:01 / HLA-DPA1*01:03, HLA-DPB1*02:01 / HLA-DPA1*01:03, HLA-DPB1*03:01 / HLA-DPA1*01:03, HLA-DPB1*04:01 / HLA-DPA1*01:03, HLA-DPB1*04:02 / HLA-DPA1*01:03, HLA-DPB1*06:01 / HLA-DPA1*01:03, HLA-DQB1*02:01 / HLA-DQA 1*05:01, HLA-DQB1*02:02 / HLA-DQA1*02:01, HLA-DQB1*06:02 / HLA-DQA1*01:02, HLA-DQB1*06:04 / HLA-DQA1*01:02, HLA-DR B1*01:01, HLA-DRB1*01:02, HLA-DRB1*03:01, HLA-DRB1*03:02, HLA-DRB1*04:01, HLA-DRB1*04:02, HLA-DRB1*04:03, HLA-D RB1*04:04, HLA-DRB1*04:05, HLA-DRB1*04:07, HLA-DRB1*07:01, HLA-DRB1*08:01, HLA-DRB1*08:02, HLA-DRB1*08:03, HLA- DRB1*08:04, HLA-DRB1*09:01, HLA-DRB1*10:01, HLA-DRB1*11:01, HLA-DRB1*11:02, HLA-DRB1*11:04, HLA-DRB1*12:01, HLA -DRB1*12:02, HLA-DRB1*13:01, HLA-DRB1*13:02, HLA-DRB1*13:03, HLA-DRB1*14:01, HLA-DRB1*15:01, HLA-DRB1*15:02, HLA-DRB1*15:03, HLA-DRB1*16:01, HLA-DRB3*01:01, HLA-DRB3*02:02, HLA-DRB3*03:01, HLA-DRB4*01:01, HLA-DRB5*01:01). The peptide presented by the HLA protein can have a length of 15-40 amino acids.The peptides presented by the HLA proteins may have a length of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27 or more amino acids. In some embodiments, the peptides presented by the HLA proteins may have a length of at most 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11 or fewer amino acids.
[0461] The peptides presented by HLA proteins may include peptides identified by: (a) isolating one or more HLA complexes from a cell line expressing a single HLA class II allele; (b) isolating one or more HLA-peptides from the one or more isolated HLA complexes; (c) obtaining MS / MS spectra of the one or more isolated HLA-peptides; and (d) obtaining peptide sequences corresponding to the MS / MS spectra of the one or more isolated HLA-peptides from a peptide database; wherein the sequences of the one or more isolated HLA-peptides are identified from the one or more sequences obtained in steps (a, b, c) and (d).
[0462] The separation may include isolating the HLA-peptide complex from cells transfected or transduced with an affinity-tagged HLA construct. In some embodiments, the complex may be isolated using standard immunoprecipitation techniques known in the art with commercially available antibodies. The cells may be lysed first. HLA class II-peptide complexes may be isolated using HLA class II-specific antibodies, such as the M5 / 114.15.2 monoclonal antibody. In some embodiments, a single (or a pair of) HLA alleles are expressed as a fusion protein with a peptide tag, and the HLA-peptide complex is isolated using a binding molecule that recognizes the peptide tag.
[0463] The separation may include separating the peptides from the HLA-peptide complex and sequencing the peptides. The peptides are separated from the complex by any method known to those skilled in the art, such as acid elution. Although any sequencing method may be used, in some embodiments, a method using mass spectrometry is used, such as liquid chromatography-mass spectrometry (LC-MS or LC-MS / MS, or HPLC-MS or HPLC-MS / MS). These sequencing methods may be well known to the skilled person and are reviewed in Medzihradszky KF and Chalkley RJ. Mass Spectrom Rev. 2015 Jan-Feb; 34(1):43-63.
[0464] Other candidate components and molecules suitable for separation or purification may include binding molecules, such as biotin (biotin-avidin specific binding pair), antibodies, receptors, ligands, lectins or molecules comprising solid supports, including, for example, plastics or polystyrene beads, plates or beads, magnetic beads, test strips and membranes. Purification methods such as cation exchange chromatography can be used to separate conjugates by charge differences, which effectively separate conjugates into their various molecular weights. The content of the fractions obtained by cation exchange chromatography can be identified by conventional methods, for example, mass spectrometry, SDS-PAGE or other known methods for separating molecular entities by molecular weight.
[0465] In some embodiments, the method further includes separating peptides from HLA-peptide complexes labeled with affinity receptors before characterization. In some embodiments, HLA-peptide complexes are separated using anti-HLA antibodies. In some cases, HLA-peptide complexes with or without affinity tags are separated using anti-HLA antibodies. In some cases, soluble HLA (sHLA) with or without affinity tags are separated from the culture medium of cell cultures. In some cases, soluble HLA (sHLA) with or without affinity tags are separated using anti-HLA antibodies. For example, beads or columns containing anti-HLA antibodies can be used to separate HLA, such as soluble HLA (sHLA) with or without affinity tags. In some embodiments, peptides are separated using anti-HLA antibodies. In some cases, soluble HLA (sHLA) with or without affinity tags are separated using anti-HLA antibodies. In some cases, soluble HLA (sHLA) with or without affinity tags are separated using columns containing anti-HLA antibodies. In some embodiments, the method further comprises removing one or more amino acids from the terminus of the peptide that binds to the affinity receptor-tagged HLA-peptide complex.
[0466] The personalized cancer vaccine may further comprise an adjuvant. For example, poly-ICLC, an agonist of the RNA helicase domains of TLR3 and MDA5 and RIG3, has shown several desirable properties for vaccine adjuvants. These properties may include inducing local and systemic activation of immune cells in vivo, producing stimulatory chemokines and cytokines, and stimulating antigen presentation by DCs. In addition, poly-ICLC can induce persistent CD4+ and CD8+ responses in humans. Importantly, striking similarities in upregulation of transcriptional and signal transduction pathways were observed in subjects vaccinated with poly-ICLC and volunteers who received a highly effective, replication-competent yellow fever vaccine. In addition, in a recent Phase 1 study, >90% of ovarian cancer patients immunized with poly-ICLC combined with NYESO-1 peptide vaccine (in addition to Montanide) showed induction of CD4+ and CD8+ T cells and antibody responses to the peptide.
[0467] The personalized cancer vaccine may further comprise an immune checkpoint inhibitor. Immune checkpoint inhibitors may include a type of drug that blocks certain proteins produced by certain types of immune system cells, such as T cells, and certain cancer cells. These proteins help control the immune response and can prevent T cells from killing cancer cells. When these proteins are blocked, the "brakes" on the immune system are released and T cells are better able to kill cancer cells. Examples of checkpoint proteins found on T cells or cancer cells include PD-1 / PD-L1 and CTLA-4 / B7-1 / B7-2. Some immune checkpoint inhibitors are used to treat cancer.
[0468] Training data may also include structured data, time series data, unstructured data, and relational data. Unstructured data may include audio data, image data, video, mechanical data, electrical data, chemical data, and any combination thereof, for accurately simulating or training a robot or simulation. Time series data may include data from one or more of smart meters, smart appliances, smart devices, monitoring systems, telemetry devices, or sensors. Relational data includes data from customer systems, enterprise systems, operating systems, websites, web-accessible application programming interfaces (APIs), or any combination thereof. This may be accomplished by a user through any method of inputting files or other data formats into the software or system.
[0469] Training data can be uploaded to a cloud-based database. The cloud-based database can be accessed from local and / or remote computer systems running machine learning-based sensor signal processing algorithms. Cloud-based databases and associated software can be used to archive electronic data, share electronic data, and analyze electronic data. Locally generated data or data sets can be uploaded to a cloud-based database from which they can be accessed and used to train other machine learning-based detection systems at the same site or at different sites. Locally generated sensor device and system test results can be uploaded to a cloud-based database and used to update training data sets in real time to continuously improve the test performance of sensor devices and detection systems.
[0470] Convolutional neural networks can be used to perform training. Convolutional neural networks (CNNs) are described elsewhere herein. Convolutional neural networks can include at least two convolutional layers. The number of convolutional layers can be 1-10, and the number of expansion layers can be 0-10. The total number of convolutional layers (including input layers and output layers) can be at least about 1, 2, 3, 4, 5, 10, 15, 20 or more, and the total number of expansion layers can be at least about 1, 2, 3, 4, 5, 10, 15, 20 or more. The total number of convolutional layers can be at most about 20, 15, 10, 5, 4, 3 or less, and the total number of expansion layers can be at most about 20, 15, 10, 5, 4, 3 or less. In some embodiments, the number of convolutional layers is 1-10, and the number of fully connected layers is 0-10. The total number of convolutional layers (including input and output layers) can be at least about 1, 2, 3, 4, 5, 10, 15, 20, or more, and the total number of fully connected layers can be at least about 1, 2, 3, 4, 5, 10, 15, 20, or more. The total number of convolutional layers can be at most about 20, 15, 10, 5, 4, 3, or less, and the total number of fully connected layers can be at most about 20, 15, 10, 5, 4, 3, or less.
[0471] The convolutional neural network can include at least one batch normalization step. The batch normalization layer can improve the performance and stability of the neural network. The batch normalization layer can provide zero mean / unit variance input to any layer in the neural network. The total number of batch normalization layers can be at least about 3, 4, 5, 10, 15, 20 or more. The total number of batch normalization layers can be at most about 20, 15, 10, 5, 4, 3 or less.
[0472] The convolutional neural network may include at least one spatial dropout step. The total number of spatial dropout steps may be at least about 3, 4, 5, 10, 15, 20 or more, and the total number of spatial dropout steps may be at most about 20, 15, 10, 5, 4, 3 or less.
[0473] The convolutional neural network may include at least one global maxi...
Claims
1. A method of identifying a peptide sequence as presented by at least one of one or more proteins encoded by an HLA allele of a cell of a subject, comprising: (a) using a computer processor to input amino acid sequence information of a set of candidate peptide sequences expressed by cancer cells of a single human subject into a trained machine learning HLA-peptide presentation prediction model to generate a plurality of presentation predictions, wherein each presentation prediction in the plurality of presentation predictions indicates a presentation likelihood of a peptide sequence in the set of candidate peptide sequences being presented by an MHC protein of the single human subject; Wherein the trained machine learning HLA-peptide presentation prediction model comprises: (i) a plurality of parameters, wherein the plurality of parameters are based on training data from training cells expressing MHC proteins, wherein the training data comprises a plurality of training peptide sequences and epitope presentation quantitative information, wherein the epitope presentation quantitative information comprises an amount of one or more of the plurality of training peptide sequences presented by the MHC protein; and (ii) a function representing the relationship between the amino acid sequence information received as input and the presentation likelihood generated as output based on the amino acid sequence information and the plurality of parameters; and (b) identifying, based at least on the plurality of presentation predictions, a peptide sequence from the plurality of peptide sequences in the set of candidate peptide sequences as being presented by at least one of the one or more proteins encoded by the HLA alleles of cells of the subject.
2. A method for selecting a peptide sequence, comprising: (a) using a computer processor to input amino acid sequence information of a set of candidate peptide sequences expressed by cancer cells of a single human subject into a trained machine learning HLA-peptide presentation prediction model to generate a plurality of presentation predictions, wherein each presentation prediction in the plurality of presentation predictions indicates a presentation likelihood of a peptide sequence in the set of candidate peptide sequences being presented by an MHC protein of the single human subject; Wherein the trained machine learning HLA-peptide presentation prediction model comprises: (i) a plurality of parameters, wherein the plurality of parameters are based on training data from training cells expressing MHC proteins, wherein the training data comprises a plurality of training peptide sequences and epitope presentation quantitative information, wherein the epitope presentation quantitative information comprises an amount of one or more of the plurality of training peptide sequences presented by the MHC protein; and (ii) a function representing the relationship between the amino acid sequence information received as input and the presentation likelihood generated as output based on the amino acid sequence information and the plurality of parameters; and (b) selecting a subset of peptide sequences in the set of candidate peptide sequences based at least on the plurality of presentation predictions to generate a set of selected peptide sequences.
3. A method of treating cancer in a human subject in need thereof, comprising: (a) using a computer processor to input amino acid sequence information of a set of candidate peptide sequences expressed by cancer cells of a single human subject into a trained machine learning HLA-peptide presentation prediction model to generate a plurality of presentation predictions, wherein each presentation prediction in the plurality of presentation predictions indicates a presentation likelihood of a peptide sequence in the set of candidate peptide sequences being presented by an MHC protein of the single human subject; Wherein the trained machine learning HLA-peptide presentation prediction model comprises: (i) a plurality of parameters, wherein the plurality of parameters are based on training data from training cells expressing MHC proteins, wherein the training data comprises a plurality of training peptide sequences and epitope presentation quantitative information, wherein the epitope presentation quantitative information comprises an amount of one or more of the plurality of training peptide sequences presented by the MHC protein; and (ii) a function representing the relationship between the amino acid sequence information received as input and the presentation likelihood generated as output based on the amino acid sequence information and the plurality of parameters; (b) selecting or identifying a subset of peptide sequences in the set of candidate peptide sequences based at least on the plurality of presentation predictions to generate a set of selected or identified peptide sequences; and (c) administering to the single human subject a pharmaceutical composition comprising: (i) a polypeptide having one or more selected peptide sequences, (ii) a polynucleotide encoding the polypeptide described in (i); (iii) an APC comprising (i) or (ii), or (iv) a T cell comprising a T cell receptor (TCR) specific for an MHC protein of said single human subject complexed with one or more of said peptide sequences selected or identified in (b).
4. The method of any one of claims 1-3, wherein the plurality of parameters are based on training data from training cells expressing the MHC proteins of the single human subject. The method of claim 4 , wherein each of the plurality of training peptide sequences is associated with an MHC protein.
6. The method of claim 5, wherein the training data comprises the identity of the MHC protein associated with each training peptide sequence in the plurality of training peptide sequences.
7. The method of claim 6, wherein the training data comprises observation of presentation of one or more of the plurality of training peptide sequences by MHC proteins by mass spectrometry.
8. The method of any one of claims 1-7, wherein the MHC protein of the single human subject is a class I MHC protein.
9. The method according to any one of claims 1-8, wherein the multiple candidate peptide sequences expressed by the cancer cells of the single human subject are identified by comparing the whole genome or whole exome sequence information of the cancer cells from the single human subject with the whole genome or whole exome sequence information of the non-cancerous cells from the single human subject, and identifying nucleic acid sequences that are unique to the cancer cells and not present in the non-cancerous cells.
10. The method of any one of claims 1-9, wherein each candidate sequence in the plurality of candidate peptide sequences comprises a cancer-specific mutation.
11. The method of any one of claims 1-10, wherein the trained machine learning HLA-peptide presentation prediction model has a peptide presentation prediction value (PPV) of at least 0.2 according to the presentation PPV determination method.
12. The method according to any one of claims 1-11, wherein the presentation PPV determination method comprises inputting amino acid sequence information of a plurality of test peptide sequences into the trained machine learning HLA-peptide presentation prediction model to generate a plurality of test presentation predictions, each test presentation prediction indicating the likelihood that the one or more proteins encoded by the HLA alleles are capable of presenting a given test peptide sequence in the plurality of test peptide sequences, wherein the plurality of test peptide sequences comprises at least 500 test peptide sequences, the sequences comprising: (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell, and (ii) at least 499 decoy peptide sequences contained within a protein encoded by the genome of an organism, wherein the organism and the subject are of the same species.
13. A method according to claim 12, wherein the ratio of at least one hit peptide sequence among the multiple test peptide sequences to the at least 499 bait peptide sequences is 1:499, and according to the trained machine learning HLA-peptide presentation prediction model, the top 0.2% of the multiple test peptide sequences are predicted to be presented by the HLA protein expressed in the cell.
14. The method of claim 13, wherein (i) the at least one hit peptide sequence comprises at least 10 hit peptide sequences, and (ii) the at least 499 decoy peptide sequences comprise at least 4,990 decoy peptide sequences.
15. The method according to any one of claims 1-14, wherein the amount of one or more of the plurality of training peptide sequences presented by the MHC protein comprises the number of copies of one or more of the plurality of training peptide sequences presented by the MHC protein.
16. The method according to any one of claims 1-15, wherein the amount of one or more of the plurality of training peptide sequences presented by the MHC protein comprises the number of copies per cell of one or more of the plurality of training peptide sequences presented by the MHC protein.
17. The method according to any one of claims 1-16, wherein the amount of one or more of the multiple training peptide sequences presented by the MHC protein comprises the absolute amount, number of molecules, density, concentration, absolute amount per cell, number of molecules per cell, density per cell or concentration in cells of one or more of the multiple training peptide sequences presented by the MHC protein.
18. The method of any one of claims 1-17, wherein the amount of one or more of the plurality of training peptide sequences presented by the MHC protein is based on a large number of mass spectrometry observations, spectral counting, area under the curve (AUC), intensity-based absolute quantification (iBAQ), label-free quantification (LFQ), isotope dilution mass spectrometry, isobaric mass tags, stable isotope labels and / or mass spectrometry peak intensities.
19. The method of any one of claims 1-18, wherein the amount of one or more of the plurality of training peptide sequences presented by an MHC protein is obtained from quantitative mass spectrometry.
20. The method of any one of claims 1-19, wherein the epitope presentation quantitative information is obtained from internal standard-parallel reaction monitoring (IS-PRM) mass spectrometry.
21. The method of any one of claims 1-20, wherein the epitope presentation quantitative information is obtained from a xenograft sample.
22. The method of claim 21, wherein the xenograft sample is a patient-derived xenograft (PDX) sample.
23. A method for selecting a peptide sequence, comprising: (a) using a computer processor to input amino acid sequence information of a set of candidate peptide sequences expressed by cancer cells of a single human subject into a trained machine learning HLA-peptide antigen-specific T cell prediction model to generate a plurality of antigen-specific T cell predictions, wherein each antigen-specific T cell prediction in the plurality of antigen-specific T cell predictions indicates a likelihood that an MHC complex comprising an MHC protein of the single human subject and a peptide sequence in the set of candidate peptide sequences will stimulate a T cell that is specific to a peptide sequence in the set of candidate peptide sequences; The trained machine learning HLA-peptide cytotoxic T cell prediction model comprises: (i) a plurality of parameters, wherein the plurality of parameters are based on training data from training cells expressing MHC proteins, wherein the training data comprises a plurality of training peptide sequences and epitope presentation quantitative information, wherein the epitope presentation quantitative information comprises an amount of one or more of the plurality of training peptide sequences presented by the MHC protein; and (ii) a function representing the relationship between the amino acid sequence information received as input and the likelihood that a T cell is specific for a peptide sequence in the set of candidate peptide sequences generated as output based on the amino acid sequence information and the plurality of parameters; and (b) selecting a subset of peptide sequences in the set of candidate peptide sequences based at least on the plurality of antigen-specific T cell predictions to generate a set of selected peptide sequences.
24. The method of claim 23, wherein each of the plurality of antigen-specific T cell predictions indicates the likelihood that an MHC complex comprising an MHC protein of the single human subject and a peptide sequence in the set of candidate peptide sequences will stimulate a T cell that is specific to a neoantigen peptide sequence in the set of candidate peptide sequences.
25. The method of claim 24, wherein the function is a function representing the relationship between the amino acid sequence information received as input and the likelihood that a T cell is specific to a neoantigen peptide sequence in the set of candidate peptide sequences generated as output based on the amino acid sequence information and the plurality of parameters.
26. The method of claim 23, wherein each of the plurality of antigen-specific T cell predictions indicates a likelihood that an MHC complex comprising an MHC protein of the single human subject and a peptide sequence from the set of candidate peptide sequences will stimulate a T cell that is cytotoxic.
27. The method according to claim 26, wherein the function is a function representing the relationship between the amino acid sequence information received as input and the possibility of becoming a cytotoxic T cell generated as output based on the amino acid sequence information and the plurality of parameters.
Citation Information
Patent Citations
Producing protein arrays and fusion protein for use therein
GB2370039A
Protective face mask for attachment to protective eye-ware
US12041988B2
Lentiviral Vectors and Their Use
US20080254008A1
Engineered CD19-specific t lymphocytes that coexpress il-15 and an inducible caspase-9 based suicide gene for the treatment of b-cell malignancies
US20130071414A1
Synthesis of protein with an identification peptide
US4703004A