Machine learning techniques for predicting surface-presented peptides

A machine learning model using genomic and transcriptome data predicts MHC-binding peptides' surface presentation, overcoming conventional limitations, enhancing the accuracy of immune response prediction and personalized cancer therapies.

JP2026071377APending Publication Date: 2026-04-28PERSONALIS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PERSONALIS INC
Filing Date
2026-02-10
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Conventional techniques for predicting MHC-binding peptides are inadequate as they fail to accurately determine whether a peptide is presented on the cell surface, leading to incomplete understanding of immune system responses to tumor cells.

Method used

A machine learning model is trained using genomic and transcriptome data, including MHC molecule identification and peptide presentation metrics, to predict the degree of peptide presentation on the cell surface, utilizing single-allelic and multi-allelic immunopeptidomics data to enhance accuracy.

Benefits of technology

The model achieves significantly higher positive predictive values compared to conventional methods like NetMHCPan, enabling precise identification of MHC-binding peptides presented on the cell surface, facilitating personalized immunotherapies and biomarker development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071377000001_ABST
    Figure 2026071377000001_ABST
Patent Text Reader

Abstract

This provides machine learning techniques for predicting surface-presented peptides. [Solution] This disclosure provides a method for predicting surface-presented peptides using binding and surface-presentation properties. The method may include accessing a trained machine learning model configured to produce an output indicating the degree to which one or more expression levels and one or more peptide presentation metrics are related according to population-level relationships between expression and presentation. For each peptide in a set of peptides from a tissue sample, a score can be determined using the machine learning model and corresponding genomic and transcriptome data. The score predicts whether the corresponding peptide is a surface-presented peptide that binds to an MHC molecule and is presented on the cell surface.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications This application claims priority based on U.S. Provisional Patent Application No. 63 / 040,943, titled "Composite Biomarkers for Immunotherapy for Cancer," filed on June 18, 2020, and U.S. Provisional Patent Application No. 63 / 111,007, titled "Machine - Learning Techniques For Predicting Surface - Presenting Peptides," filed on November 7, 2020, the entire contents of which are hereby incorporated by reference in their entirety for all purposes.

[0002] This disclosure relates to machine - learning techniques for predicting surface - presenting peptides.

Background Art

[0003] Cancer contains mutations, which can be somatic and / or tumor - specific. The immune system detects cancer - based mutations by identifying peptides derived from these mutations. Peptides bind to proteins encoded by major histocompatibility complex (MHC) genes and are recognized by the immune system when presented on the cell surface. For example, peptides corresponding to mutated genes can bind to specific MHC molecules (e.g., human leukocyte antigen (HLA) proteins) and be presented on the cell surface. Predicting the peptides expressed on the surface of tumor cells can provide information for the development of precision cancer treatment and diagnosis. For example, identifying the genomic variants corresponding to these peptides can analyze the complex system responses and resistances to certain cancer immunotherapies. As another example, analyzing the peptides presented on the surface of tumor cells can create individualized immuno - oncology (I - O) therapies and / or neoantigen cancer vaccines.

[0004] Such peptides expressed on the surface of tumor cells are also known as "neoantigens," and techniques for predicting them require a thorough analysis of many technical factors, including, but not limited to, the quality of peptide sequencing data, the availability of tumor and normal sample pairs, HLA typing, and the identification of other peptide characteristics. For example, neoantigens can be identified based on the prediction of peptides that bind to MHC molecules and are presented on the cell surface. To identify neoantigens, determining the peptide encoded by the somatic variant and identifying the HLA molecule that binds to the peptide is only the first step in a very complex process. This is because each peptide identified from sequence data may or may not be processed by the proteasome; transported for MHC binding; presented on the tumor cell surface; and finally recognized by the immune system. Due to this complex process, many peptides that bind to HLA molecules (for example) may not be expressed on the cell surface.

[0005] Furthermore, by identifying one or more binding motifs of MHC molecules, it is possible to determine whether a given peptide binds to an MHC molecule. While the binding motifs of some MHC molecules (e.g., HLA-A molecules) are known, many MHC molecules still lack identified binding motifs. For example, the binding motifs of MHC class II molecules are relatively unknown due to the limited availability of experimental data. Without this information, it would be difficult to determine whether a peptide binds to the corresponding MHC molecule. Conventional techniques have attempted to address this problem by training machine learning models using known MHC binding motifs to predict whether a peptide binds to one of various types of MHC molecules. However, even if such peptides are identified, some peptides are not present on the cell surface. In other words, while conventional techniques can identify MHC-binding peptides, only a small fraction of them can be successfully presented on the cell surface. Since the immune system response is triggered when MHC-binding peptides are presented on the cell surface, simply identifying MHC-binding peptides does not fully explain how the immune system responds to tumor cells, foreign proteins, and other factors.

[0006] Thus, conventional techniques for predicting MHC-binding peptides do not address whether the peptide is actually presented and expressed on the cell surface. Furthermore, conventional techniques are insufficient to identify the peptide characteristics that indicate a given peptide is presented on the cell surface. Therefore, it is necessary to accurately predict the peptide that binds to the corresponding MHC molecule and is presented on the cell surface. [Overview of the Initiative] [Problems that the invention aims to solve]

[0007] In some embodiments, a method for predicting surface-presented peptides is provided. This method may include accessing a trained machine learning model that has been trained using a training dataset and, for each of a plurality of peptides identified by the training dataset, includes protein properties of the major histocompatibility complex (MHC) molecule that binds and presents the peptide, one or more expression levels representing the expression level of the gene encoding the peptide, and one or more peptide presentation metrics representing the amount of peptide detected as being presented by the MHC molecule. The machine learning model may be configured to produce an output indicating the degree to which one or more expression levels and one or more peptide presentation metrics are related according to a population-level relationship between expression and presentation. [Means for solving the problem]

[0008] This method may also include accessing genomic and transcriptome data corresponding to the biological sample under consideration. The genomic and transcriptome data may include the identification of one or more MHC molecules from the biological sample and, for each peptide in a set of peptides identified from the cell line or tissue sample, one or more values ​​representing that peptide. The one or more values ​​may be determined based on the processing of the tissue sample. This method may also include determining a score for each peptide in the set of peptides using a machine learning model, one or more MHC molecules identified from the biological sample, and one or more values ​​representing the peptide. This method may include generating results based on the scores and outputting the results.

[0009] Some embodiments of the present disclosure include a system comprising one or more data processors. In some embodiments, the system includes a non-temporary computer-readable storage medium containing instructions that cause one or more data processors to perform some or all of the methods and / or some or all of the processes disclosed herein when the data processors are executed. Some embodiments of the present disclosure include a computer program product materialized on a non-temporary machine-readable storage medium, which includes instructions configured to cause one or more data processors to perform some or all of the methods and / or some or all of the processes disclosed herein.

[0010] The terms and expressions used are descriptive, not restrictive, and in using such terms and expressions, there is no intention to exclude any equivalent or part thereof of the features shown and described, but it is recognized that various modifications are possible within the scope of the claimed invention. Accordingly, although the claimed invention is specifically disclosed by embodiments and optional features, it should be understood that modifications and variations of the concepts disclosed herein may be available to those skilled in the art, and that such modifications and variations are considered to fall within the scope of the invention as defined by the appended claims. [Brief explanation of the drawing]

[0011] This disclosure will be explained in conjunction with the attached diagram. [Figure 1] Figure 1 shows a schematic diagram of a peptide that binds to an MHC molecule and is presented on the cell surface. [Figure 2] Figure 2 shows a schematic diagram representing peptides that may be presented on the cell surface in response to gene therapy. [Figure 3] Figure 3 shows a schematic diagram illustrating the identification of single-allelic immunopeptidomics data that can be used to train machine learning models, according to several embodiments. [Figure 4]Figure 4 shows allelic diversity data corresponding to MHC-binding peptides in several embodiments. [Figure 5] Figure 5 shows source diversity data identified from tissue and cell line samples for training machine learning models to predict surface-presented peptides, according to several embodiments. [Figure 6] Figure 6 shows plots of comparative data between the predicted number of peptides based on gene expression levels and the number of peptides actually observed, for several embodiments. [Figure 7] Figure 7 shows the process for determining the genetic propensity scores used to train a machine learning model, according to several embodiments. [Figure 8] Figure 8 shows plots comparing the expected number of peptides for one or more regions within a gene with the number of peptides actually observed for those regions, according to several embodiments. [Figure 9] Figure 9 shows the process for determining the hotspot score used to train a machine learning model, according to several embodiments. [Figure 10] Figure 10 shows examples of features used in combined and presented models according to several embodiments. [Figure 11] Figure 11 shows exemplary model architectures for training machine learning models to predict surface-presented peptides, according to several embodiments. [Figure 12] Figure 12 shows the performance levels of trained combined models and trained presentation models, measured by positive predictive value based on 10% holdout data, in several embodiments. [Figure 13] Figure 13 shows a comparison of the performance levels of a trained machine learning model compared to conventional techniques for predicting surface-presented peptides. [Figure 14] Figure 14 shows a comparison of the performance levels of trained presentation models across various alleles compared to conventional techniques for predicting surface-presented peptides. [Figure 15]FIG. 15 shows the results of a leave-one-out analysis of a trained presentation model according to some embodiments. [Figure 16] FIG. 16 shows a graph showing the precision and recall for evaluating a trained machine learning model according to some embodiments. [Figure 17] FIG. 17 shows a box plot representing the performance levels of a trained machine learning model across various tissue samples according to some embodiments. [Figure 18] FIG. 18 shows a graph comparing the performance levels of a trained machine learning model with those of other prior arts according to some embodiments. [Figure 19] FIG. 19 includes a flowchart showing an example of a method for predicting surface presentation peptides according to a particular embodiment. [Figure 20] FIG. 20 shows an example of a computer system for implementing some of the embodiments disclosed herein.

DETAILED DESCRIPTION OF THE INVENTION

[0012] I. Overview To address at least the aforementioned shortcomings of conventional systems, this technique can be used to predict surface-presented peptides. As used herein, “surface-presented peptide” may mean a peptide that binds to an MHC molecule (e.g., HLA-A protein) and is presented on the corresponding cell surface. One or more somatic variants can be identified by determining the DNA sequences from normal and tumor samples. Somatic variants include one or more gene mutations present in tumor and normal samples. Somatic variants from tumor samples can be processed using a trained machine learning model to predict whether the peptide encoded by the somatic variant will bind to an MHC molecule (e.g., MHC class 1) and be presented on the cell surface. The machine learning model may include a binding model that predicts whether the peptide encoded by the somatic variant will bind to an MHC molecule. In some embodiments, the machine learning model includes a presentation model that predicts whether the peptide encoded by the somatic variant will be expressed on the cell surface.

[0013] A machine learning model can be trained using (i) a genetically engineered single allele cell line and (ii) a training data set obtained from multi-allele data of other tissue samples. In some cases, the machine learning model is trained using binding array data (e.g., IEDB data). The training data set can include, for each peptide identified by the training data set, one or more expression levels representing the expression level of the gene encoding the peptide and one or more peptide presentation metrics representing the amount of peptide detected as being presented by MHC molecules. The training data set can include immunopeptidomic data of peptides generated from a plurality of genetically engineered cell lines (e.g., K562 cells) expressing a single allele of interest (e.g., HLA-A). In particular, MHC-peptide complexes in these cell lines can be immunoprecipitated using the W6 / 32 antibody, followed by peptide elution and peptide sequencing using tandem mass spectrometry. The training data set corresponding to multi-allele data from other tissue samples can be obtained using curated public data.

[0014] Prediction of surface-presented peptides can be performed by biasing the selection of peptides towards those associated with a score predicting a more certain presentation compared to the probability expected by the population-level relationship between peptide expression and presentation. Additionally or alternatively, prediction of surface-presented peptides is performed by biasing the selection of peptides towards those associated with a region in space, the region being associated with outlier peptides in a training data set where expression levels and peptide presentation metrics are related in a way that deviates from the population-level relationship.

[0015] Accordingly, embodiments of this disclosure offer technical advantages over conventional systems by accurately predicting peptides that bind to their corresponding MHC molecules and are presented on the cell surface. As previously stated, the binding and expression of peptides on the surface of tumor cells can predict how the immune system will respond to neoantigens and / or certain cancer immunotherapies. Therefore, accurate prediction of surface-presented peptides facilitates the selection or development of immunotherapies that will be most effective for a given target. Furthermore, based on model evaluation, embodiments exhibit significantly higher positive predictive values ​​compared to conventional technologies such as NetMHCPan 4.0. Thus, the high sensitivity and specificity of embodiments enable accurate identification of MHC-binding peptides presented on the cell surface, thereby facilitating their application to the development of personalized immunotherapies and biomarkers.

[0016] The following examples are provided to illustrate a particular embodiment. Specific details are included in the following description for illustrative purposes and to ensure a full understanding of the examples disclosed. However, it will be apparent that various examples can be performed without these specific details. For example, devices, systems, structures, assemblies, methods, and other components may be shown as components in block diagrams to avoid obscuring the examples with unnecessary details. In other cases, well-known devices, processes, systems, structures, and techniques may be shown without necessary details to avoid obscuring the examples. The figures and descriptions are not intended to be limiting. The terms and expressions used in this disclosure are for illustrative purposes only, not limiting purposes, and the use of such terms and expressions is not intended to exclude any equivalent or part of any of the features shown and described. The term “example” as used herein means “to function as an example, instance, or illustration.” No embodiment or design described herein as an “example” should necessarily be construed as preferable or advantageous to any other embodiment or design.

[0017] II. Surface presentation of peptides 1. Neoantigens in tumor samples Neoantigens can be found in tumor samples and refer to one or more peptides that are presented on the surface of tumor cells, thereby triggering an immune system response. The immune system can be tuned to search for pathogens, including cancer, and therefore has the ability to cure cancer. The immune system can distinguish between self and non-self antigens. Since tumors are caused by genetic mutations (e.g., somatic variants), peptides that correspond to these genetic mutations and are expressed on the cell surface can be considered neoantigens. Because these peptides are considered "novel" to the immune system, ideally, the immune system can recognize and eliminate tumor cells based on the detection of neoantigens presented on the tumor cell surface. As described above, tumor samples can be analyzed to reveal sequence data, which can then be compared to that of normal samples to identify somatic variants. Further analysis of somatic variants can determine which subsets of variants will manifest as peptides. Neoantigens can be predicted by identifying peptides that bind to MHC molecules and are presented on the cell surface. Therefore, the ability of peptides to be presented on the cell surface can be a crucial factor in developing immunotherapies against cancer.

[0018] 2. Peptides that respond to the treatment of certain autoimmune diseases Surface-presented peptides (SPRs) can be identified in relation to autoimmune diseases, and these peptides are encoded based on genetic alterations resulting from specific immunotherapies. Figure 2 shows a schematic diagram illustrating a SPRS in response to gene therapy. Figure 2 shows a mutation in the dystrophin gene, which typically causes debilitating muscular dystrophy. The dystrophin gene encodes the dystrophin protein molecule, which acts as a cushion, a shock-absorbing substance in muscle cells. The absence of a fully functional dystrophin protein can lead to muscle degeneration. Typically, muscular dystrophy can be treated with exome skipping therapy, which skips the exome (e.g., exon 52) that causes the dystrophin gene mutation, generating a semi-functional dystrophin protein for the target. While exome skipping therapy can be effective, the intentional skipping of exomes through genetic modification can induce the generation of new types of peptides. These new peptides can bind to MHC molecules and be presented on the cell surface, potentially triggering a destructive immune response.

[0019] III. Example of a training dataset Machine learning models for predicting surface-presented peptides can be trained using supervised training algorithms. Machine learning models can be trained using training datasets. Training datasets for training machine learning models can include sequence data from a variety of sources: (i) peptides identified as binding to HLA molecules based on in vitro experiments, (ii) peptides identified by mass spectrometry from tumor samples, (iii) HLA alleles, and (iv) non-tumor samples. However, some training sequence data may be inaccurate for training machine learning models. For example, training sequence data generated from tissue samples would require the challenging process of mapping peptides to one of several types of HLA proteins (e.g., HLA-A, HLA-B) that are co-expressed on the cell surface. In another example, sequence data generated using in vitro methods may not mimic surface presentation. Embodiments of this disclosure for systematically addressing training dataset inconsistencies are used to train a machine learning model that predicts peptides likely to be "shuttled" to the cell surface from somatic cell variants called from sequence data.

[0020] As an addition or alternative, the training dataset may further include data corresponding to somatic variants, each somatic variant being labeled to indicate whether the peptide encoded by the somatic variant binds to an MHC molecule (e.g., HLA-A protein) and is presented on the cell surface. The training dataset may also include one or more features derived from the somatic variant (e.g., peptide sequence, peptide length, peptide expression in tumor samples).

[0021] To prepare the training dataset, tumor samples and corresponding normal control samples can be sequenced to generate tumor-normal paired sequence data. The tumor-normal paired sequence data can be compared to identify somatic variants, including altered genes containing single nucleotide variants (SNVs), indels, and / or copy number variations. In some cases, machine learning models can be used to process the tumor-normal paired sequence data and identify somatic variants in tumor samples.

[0022] 1. Training data source a) Single-allele immunopeptidomics data In some cases, at least a portion of the training data corresponds to peptides identified from genetically engineered single-allelic cell lines. Figure 3 shows a schematic diagram of identifying single-allelic immunopeptide mix data that can be used to train machine learning models according to several embodiments. As shown in Figure 3, a genetically engineered single-allelic K562 cell line can be constructed and then transfected with a specific HLA molecule of interest (e.g., HLA-B) (step 305). As previously mentioned, the HLA complex is a group of related proteins encoded by the MHC gene complex in humans. These cell surface proteins are responsible for regulating the immune system. From the cell line, HLA-binding peptides can be identified by immunoprecipitation of the HLA-peptide complex using a W6 / 32 antibody (step 310), applying peptide elution (step 315), and performing peptide sequencing on the eluted peptide using mass spectrometry (e.g., liquid chromatography-mass spectrometry, mass spectrometry) (step 320). Thus, HLA-binding peptides can be identified for that specific HLA molecule of interest (step 325).

[0023] Single-allelic immunopeptide mix data that identifies various features of HLA-binding peptides can be revealed and included as part of the training dataset. Examples of training data from single-allelic immunopeptide mix data may include, for a given HLA-binding peptide, the peptide type, peptide length, peptide amino acid sequence, HLA allele that binds to the peptide, the number of transcripts corresponding to the peptide, and the expression of the gene region encoding the peptide. To optimize the performance of the machine learning model, training data representing HLA genotypes of the general population was generated. For example, Figure 4 shows allele diversity data corresponding to HLA-binding peptides in several embodiments. To determine allele diversity data, identified peptides can be clustered based on their similarity to peptide sequences corresponding to all known alleles of the target HLA molecule (e.g., HLA alleles identified from the IMGT database). Thus, identified peptides can be clustered based on the similarity of their respective binding pockets. In some cases, identified peptides can be clustered using a BLOSUM similarity matrix. Based on these clusters, one or more alleles encoding the HLA-binding peptide can also be identified. In some cases, peptide clusters are visualized on a heatmap. For example, Figure 4 shows a first heatmap identifying allele diversity of HLA-A molecules and a second heatmap identifying allele diversity of HLA-B molecules. Additionally or alternatively, the training dataset can be enhanced with training data corresponding to allele frequency data of alleles encoding HLA-binding proteins, where the allele frequency data is categorized into different segments of the world population.

[0024] Training data corresponding to HLA-binding peptides can facilitate the training of machine learning models by using single-allelic immunopeptide mix data expressing one specific type of HLA at once. Furthermore, allelic diversity in single-allelic immunopeptide mix training data allows machine learning models to predict surface-presented peptides derived from various alleles that may not be present in the training data.

[0025] b) Multi-allele immunopeptidomics data In some cases, at least a portion of the training data corresponds to peptides identified from sequencing other target tissue samples. Multiple peptides that bind to different types of HLA molecules (e.g., HLA-A, HLA-B, HLA-C) can be identified by sequencing various tissue samples or cell lines of the target tissue sample. In some cases, cell lines and tissue samples are processed using mass spectrometry. Multiallelic immunopeptide mix data obtained from the identified peptides can be used as part of the training data. Multiallelic immunopeptide mix data can include various features corresponding to the identified peptides, including peptide length and allele diversity. Figure 5 shows source diversity data identified from target tissue samples for training a machine learning model to predict surface-presented peptides, according to several embodiments. Figure 5 shows the amount of each peptide type for both single and multiallelic samples. Additionally or alternatively, multiallelic data can also be obtained from publicly available data sources.

[0026] Multiallelic immunopeptide mix data generated from diverse tissues and cell lines can be integrated into training datasets to improve the performance of trained machine learning models. In particular, training machine learning models with multiallelic immunopeptide mix data can reduce overfitting and / or underfitting. For example, both single-allelic and multiallelic immunopeptides from several publicly available data sources can be added to training datasets. Single-allelic immunopeptide mix data from genetically engineered cell lines, as well as single-allelic and multiallelic immunopeptide mix data from tissue samples, can all be integrated into training datasets to scale them (e.g., by increasing the number of unique peptides).

[0027] 2. Additional enhancement features As described above, immunopeptidemix data from training datasets identify various features of HLA-binding peptides, including peptide sequence, peptide length, binding pocket sequence, left flanking region, and right flanking region. In some cases, training datasets also include antigen presentation features, such as peptide expression levels measured by DPM. In addition to the above, two additional features can be generated from immunopeptidemix data and used to enhance training datasets.

[0028] a) Comparative data between predicted peptide counts based on gene expression levels and actually observed peptide counts. A first feature generated from immunopeptide mix data may include comparative data between the predicted number of peptides based on gene expression levels and the number of peptides actually observed. By including a training dataset with this first feature, a trained machine learning model trained on the above training data can improve its prediction of surface-presented peptides, resulting in predictions that are biased towards peptides with scores that predict presentations that are more certain than the probabilities predicted by the population-level relationship between peptide expression and presentation. Furthermore, a trained machine learning model trained on the above training data can facilitate the prediction of surface-presented peptides, resulting in predictions that are performed in a way that biases the selection to regions in space and to associated peptides, where these regions are associated with outlier peptides in the training dataset where expression levels and peptide presentation metrics relate in a way that deviates from population-level relationships.

[0029] Figure 6 shows plots of comparative data between the predicted number of peptides based on gene expression levels and the number of peptides actually observed, according to several embodiments. To generate comparative data, all transcripts corresponding to HLA-binding peptides can be identified from a training dataset and organized into sets of bins based on their respective gene expression levels. For example, as shown in Figure 6, the x-axis represents 10 sections (e.g., deciles) into which transcripts can be grouped based on their respective gene expression levels. The bars in each section represent the measured gene expression levels corresponding to the transcripts grouped in the section. The Y-axis of the plots shown in Figure 6 represents the number of peptides, and the diamond-shaped dots represent the amount of peptides counted from the cell lines of the training sample.

[0030] The initial hypothesis in Figure 6, without comparative data, appears to suggest that the expected amount of peptide is directly proportional to the measured gene expression level. However, using the comparative data in Figure 6, one or more outliers that deviate from the initial hypothesis can be identified. The first outlier, 605, contains a large amount of peptide observed in bin "1" and indicates a very low gene expression level. The second outlier, 610, contains almost no peptide observed in bin "10" and indicates a high gene expression level. Therefore, the comparative data in Figure 6 suggests that measuring the gene expression level of HLA-binding peptides may not be sufficient to predict whether the HLA-binding peptides are actually presented on the cell surface. Using the comparative data, gene propensity scores ("GPS") can be calculated for the genes encoding HLA-binding peptides, and these gene propensity scores predict whether the peptides are presented on the cell surface. In some cases, the calculated gene propensity scores can be added as additional features to the training dataset, and as a result, machine learning models can be further trained based on the gene propensity scores.

[0031] Figure 7 shows the process for determining gene propensity scores used to train machine learning models in several embodiments. In block 705, immunopeptidomics data is obtained, which includes the expression levels of genes encoding HLA-binding peptides. For example, immunopeptidomics data can be obtained by reprocessing existing mass spectrometry (MS) data or by directly accessing immunopeptidomics data from a database (e.g., an immunopeptide database).

[0032] In block 710, the predicted peptide count for genes identified in immunopeptidomics data is calculated. Specifically, the predicted peptide count is calculated based on the number of transcripts (e.g., TPMs) and the gene sequence length. In block 715, the ratio between the predicted peptide count and the observed peptide count is calculated to generate a gene propensity score (e.g., log10(observed / predicted)). In some cases, the gene propensity score can be added as an additional feature to the training dataset.

[0033] b) Comparison of the predicted number of peptides per gene region with the actual number of peptides observed. A second feature generated from immunopeptide mix data may include comparative data between the predicted number of peptides based on the expression levels within one or more regions of a given gene, and the actual observed number of peptides corresponding to those one or more regions. In contrast to the first feature, which identifies gene expression levels across various genes, the second feature identifies the expression level of a region within a single gene. Based on the identified expression level, a predicted amount of peptide can be generated. By comparing the predicted amount with the observed amount of peptide, the second feature of the training dataset can be identified, where the second feature indicates the surface presentation characteristics of one or more regions within the corresponding gene.

[0034] In some cases, the first and second features are integrated into the training dataset. A trained machine learning model trained on training data with integrated features can facilitate prediction of surface-presented peptides, and as a result, the prediction is biased towards peptides associated with scores that predict presentation with greater certainty compared to the probability expected by the population-level relationship between peptide expression and presentation. Furthermore, a trained machine learning model trained on the above training data can facilitate prediction of surface-presented peptides, and as a result, the prediction is performed in a way that biases the selection to regions in space and associated peptides, where the regions are associated with outlier peptides in the training dataset whose expression levels and peptide presentation metrics relate in a way that deviates from population-level relationships.

[0035] Figure 8 shows plots of comparative data between the expected number of peptides for one or more regions within a gene and the actually observed number of peptides for those regions, according to several embodiments. For each genomic region of a given gene, the gene expression level can be calculated and the expected amount of peptides can be measured. For example, the expected amount of peptides for the genomic region of the ACTB gene is shown by the black plot line 805. The expected amount of peptides can then be compared with the observed amount of peptides for each genomic region, which is shown by the gray region 810. Similar to Figure 6, several outliers can be identified within various genomic regions, and the observed amount of peptides does not necessarily correlate with the measured gene expression level. For example, region 815 of the ACTC1 gene (e.g., region number 230) may show a very high expected amount of peptides (e.g., >3000 peptides), but the observed amount of peptides in the same region 815 is actually much less than the expected amount (e.g., about 1000 peptides). Therefore, the comparative data in Figure 8 suggests that measuring the gene expression level at the region level of HLA-binding peptides may not be sufficient to predict whether the HLA-binding peptides are actually presented on the cell surface. Using the comparative data shown in Figure 8, a hotspot score ("hhs") can be calculated for the genes encoding the HLA-binding peptides, and the hotspot score predicts whether the peptide corresponding to the region of the gene is presented on the cell surface.

[0036] Figure 9 shows the process for determining hotspot scores used to train a machine learning model in several embodiments. In block 905, immunopeptidomics data is obtained, which includes the expression levels of genes encoding HLA-binding peptides. For example, immunopeptidomics data can be obtained by reprocessing existing mass spectrometry (MS) data or by directly accessing immunopeptidomics data from a database (e.g., an immunopeptide database).

[0037] In step 910, the predicted number of peptides in each region of a particular gene is compared with the actual number of peptides in that region. In step 915, a hotspot score is calculated for the particular gene, and the hotspot score identifies the distribution of peptide numbers observed across the gene region (e.g., ACTB gene, ACTC1 gene).

[0038] IV. Example of a model architecture for predicting MHC-binding peptides presented on the cell surface A machine learning model can be trained to predict surface-presented peptides using a training dataset. The machine learning model includes one or more submodels configured to identify the binding and surface-presentation properties of peptides in a sample. These submodels can be trained separately on corresponding subsets of the training dataset, so that each submodel can predict surface-presented peptides based on parameters learned from the features corresponding to that subset.

[0039] 1. Combined Model and Presentation Model In some cases, machine learning models include coupled models and presentation models, each trained to process various features of the input data. Figure 10 shows examples of features used by coupled model 1005 and presentation model 1010 in several embodiments. Coupled model 1005 can be trained using a training dataset that includes information related to a set of peptides (e.g., sequences of MHC molecules that bind the peptides, peptide lengths). In some cases, coupled model 1005 includes one or more trained gradient boosting algorithms. Gradient boosting refers to a machine learning technique for regression and classification problems that creates a predictive model in the form of an ensemble of weak predictive models. The technique can generalize the model by building the model step by step and allowing the optimization of any differentiable loss function. Gradient boosting combines weak learners into a single strong learner in an iterative manner. As each weak learner is added, a new model is fitted to yield a more accurate estimate of the response variable. The new weak learners can be maximally correlated with the negative gradient of the loss function and associated with the entire ensemble. Examples of gradient boosting machines include XGBoost and LightGBM. Additionally or alternatively, combined models can be constructed using other types of machine learning techniques, including bagging methods, boosting methods, and / or random forest algorithms.

[0040] The presentation model 1010 can be trained using information related not only to the peptide (e.g., peptide sequence, sequence of MHC molecule binding to the peptide, peptide length) but also to information related to the expression level of the source protein from which the peptide originates, the peptide's surface presentation characteristics, gene propensity score, and hotspot score. Therefore, the trained presentation model 1010 can identify the binding characteristics and surface presentation characteristics of a given peptide, i.e., whether the peptide is presented on the cell surface. Similar to the binding model 1005, the presentation model 1010 may include one or more trained gradient boosting algorithms.

[0041] 2. Model Architecture Figure 11 shows exemplary model architectures for training machine learning models to predict surface-presented peptides, according to several embodiments. As shown in Figure 11, the training database is represented by cylinders containing various types of information and includes allele data obtained from publicly available sources. For example, the dark gray cylinder contains immunopeptide mix data corresponding to genetically engineered single-allelic cell lines (see Figure 4). In another example, the training database may also include in vitro binding data from publicly available data sources (e.g., the IEDB database, represented by the white cylinder). In some cases, training datasets from each training database are used to train the corresponding binding and presentation models individually. Furthermore, training databases can be integrated into a larger training database to train its corresponding binding and presentation models (e.g., the "ALL(MONO)" light gray cylinder in Figure 11).

[0042] Figure 11 further shows multiple sets of binding and presentation models trained to predict surface-presented peptides. Each set of binding and presentation models is shown as being trained with various training datasets. In some cases, the outputs generated from a first set of models 1105 ("initial models") are used as input features to train a second set of models 1110 ("intermediate models"). For example, the outputs generated by the initial models corresponding to in vitro binding, single-allelic data from genetically engineered single-allelic cell lines, can be used as input features to train the intermediate models. The intermediate models can also be trained individually with training data obtained from all single-allelic data 1115. Furthermore, the outputs from the intermediate models can be deconvoluted and added to another training database containing both single-allelic and multi-allelic data. The outputs can be deconvoluted using one or more base mono-allelic base models or unsupervised clustering and alignment algorithms such as GibbsCluster.

[0043] A third set of presentation and binding models 1120 ("final models") can be trained using training datasets from a database containing deconvoluted sets of peptides for each HLA allele. The trained final models 1120 can be deployed to predict surface-presented peptides. The purpose of building a training database to train the final models is to obtain as much allele diversity as possible to avoid problems caused by underfitting or overfitting. Additionally or alternatively, the performance level of the trained final models tends to be better than that of the trained intermediate models, although the trained intermediate models can also be deployed to predict surface-presented peptides.

[0044] V. Evaluation of the performance level of machine learning models To evaluate the performance of the trained machine learning model, a test dataset is generated containing several experimentally observed peptides and synthetic decoys that are not part of the training process. The trained machine learning model processes these candidate test peptides and outputs scores predicting MHC class I binding and cell surface presentation, and the machine learning model is trained using the large immunopeptide training dataset as described above. The scores are then compared to corresponding data obtained from validated MHC-binding peptides presented on the cell surface to reveal the performance level of the trained machine learning model. The output scores are also evaluated against NetMHCpan 4.0 (a known platform for predicting peptide binding to MHC molecules), and the trained machine learning algorithm shows higher overall sensitivity and specificity. Based on the output scores, the predicted antigen loading score of the peptide can be calculated using peptides with output scores above a confidence threshold.

[0045] In another example, peptides experimentally generated from tissue samples using a mass spectrometry-based immunopeptide approach were mixed with decoys in a 1:999 ratio, and a trained machine learning model was tested and evaluated. Compared to NetMHCPan 4.0, a publicly available tool considered the gold standard for MHC-binding peptide prediction, the trained machine learning model showed significantly higher positive predictive value for the top 0.1% of predicted ligands. In yet another example, the trained machine learning model was evaluated using a slip-out analysis, revealing a high degree of agreement between the motifs in the raw data and the motifs predicted by the trained machine learning model.

[0046] 1. Model evaluation of single-allele data a) Positive predictive value Figure 12 shows the performance levels of trained binding and trained presentation models, measured in terms of positive predictive value (PPV) based on 10% holdout data, for several embodiments. The evaluation data is based on single-allelic immunopeptide mix data. Positive predictive value (PPV) is defined as the proportion of predicted positivity by the trained machine learning model that were actually positive. Therefore, PPV reflects the probability that a predicted positivity is a true positivity. In the evaluation dataset, the positive rate, which represents the ratio of positives to negatives, is 1:999.

[0047] As shown in Figure 12, the median PPV for NetMHCpan is approximately 0.4. In contrast, the trained combined models performed relatively better than NetMHCpan, with the combined model trained on single-allelic data having a median PPV of approximately 0.6, and the combined models trained on single-allelic and multi-allelic data having a similar median PPV of approximately 0.6. The first trained presentation model trained on single-allelic data and the second trained presentation model trained on single-allelic and multi-allelic data performed significantly better than NetMHCpan, with a median PPV of approximately 0.7. The performance difference of 0.1 in PPV values ​​may be due to the evaluation data originating from single-allelic data.

[0048] Figure 13 shows a comparison of the performance levels of a trained machine learning model compared to conventional techniques for predicting MHC-bound peptides. As shown in Figure 13, other conventional techniques show a median PPV of approximately 0.6, which is comparable to the median PPV corresponding to the trained binding model. Compared to the above models, the trained presentation model tends to perform better, with a median PPV of approximately 0.7.

[0049] Figure 14 shows a comparison of the performance levels of a trained presentation model for various alleles compared to conventional techniques for predicting MHC-binding peptides. The PPV values ​​for each allele are shown for both NetMHCpan and the trained presentation model. As shown in Figure 14, the PPV values ​​corresponding to the trained presentation model are significantly higher than those of NetMHCpan across all single alleles. Therefore, the trained presentation model shows a significant improvement over NetMHCpan in predicting MHC-binding peptides presented and expressed on the cell surface.

[0050] b) Single-item analysis Figure 15 shows the results of a miss-one analysis of a trained presentation model in several embodiments. To demonstrate the performance of a trained presentation model in discovering unknown types of MHC-binding peptides that may be presented on the cell surface, a miss-one analysis can be used to evaluate whether the trained presentation model can predict surface-presented peptides corresponding to alleles that are not present in any of the training data. To perform a miss-one analysis, the presentation model was trained on a training dataset with the training data corresponding to one specific allele excluded. After training, the trained machine learning model was evaluated by processing 500,000 random peptides and predicting surface-presented peptides that at least some MHC-binding peptides are encoded by the excluded alleles. To evaluate the accuracy of the peptide predictions by the trained machine learning model, the predicted MHC-binding peptide motifs were compared to motifs obtained from the raw data for which the specific allele was available.

[0051] As shown in Figure 15, the motifs corresponding to the predicted surface-presented peptides substantially coincide with the motifs of the peptides corresponding to the excluded alleles, and these motifs exhibit equivalent amino acid expression levels across nine positions of the target peptide. For example, HLA-B *The second position of the peptide corresponding to 44.03 shows a high expression level of glutamate ("E") in the raw data. The predicted MHC-binding peptide presented on the cell surface also shows a high expression level of glutamate at the same second position. Therefore, the trained machine learning model can accurately predict the MHC-binding peptide presented on the cell surface even when the corresponding allele is not part of the training data.

[0052] c) Precision and recall Figure 16 shows graphs illustrating precision and recall values ​​for evaluating trained machine learning models in several embodiments. Precision-recall can be a useful indicator of predictive success. In information retrieval, precision is an indicator of result relevance, while recall is an indicator of how many truly relevant results were returned. High precision is associated with a low false positive rate, and high recall is associated with a low false negative rate. High scores in both precision and recall can indicate that a given classifier is returning accurate results (high precision) and returning a large proportion of positive results (high recall). The performance of the trained machine learning models was evaluated based on holdout single-allelic data using 10% of the immunopeptide mix data from training mixed with synthetic negative cases at a ratio of 1:999. The X-axis of the graph corresponds to a set of rank percentile thresholds ranging from 0.02 to 1.0, identifying surface-presented peptides that fall within a specific rank percentile threshold considered for either binding or presentation.

[0053] As shown in Figure 16, the trained machine learning model responds with higher precision for all recall values ​​compared to NetMHCpan. This difference is further highlighted for the top 1% peptides in the test data, where the median precision to recall is approximately 0.8 / 0.6 for the trained machine learning model, compared to approximately 0.5 / 0.2 for NetMHCpan. Therefore, the trained machine learning model can demonstrate improved prediction of surface-presented peptides compared to NetMHCpan, returning accurate results and the majority of all positive results.

[0054] 2. Model evaluation using polyallele (tissue) samples Furthermore, the performance level of machine learning models trained using multi-allelic samples shows improved prediction of surface-presented peptides compared to conventional techniques such as NetMHCpan. Figure 17 shows box plots representing the performance levels of trained machine learning models across various tissue samples in several embodiments. In Figure 17, three types of tissue samples were processed with the trained machine learning model to obtain the fraction of true candidates corresponding to surface-presented peptides. Therefore, a higher fraction may suggest that the trained machine learning model can demonstrate a high level of performance in accurately identifying surface-presented peptides across various tissue samples.

[0055] For example, the percentage value corresponding to NetMHCpan is approximately 0.65. This percentage value indicates that NetMHCpan was able to predict approximately 65% ​​of the surface-presented peptides actually present in the tissue sample. In contrast, the trained binding models performed better than NetMHCpan, with percentage values ​​of approximately 0.81 for the binding model trained on single-allele data and approximately 0.85 for the binding models trained on single and multi-allele data. The first trained presentation model trained on single-allele data and the second trained presentation model trained on single and multi-allele data performed even better, both corresponding to percentage values ​​of approximately 0.9. Thus, the trained presentation models revealed approximately 90% of the surface-presented peptides experimentally identified in the tissue sample. Similar improvements in the prediction of surface-presented peptides were shown in other tissue samples. Figure 18 shows graphs comparing the performance levels of the trained machine learning models with other prior arts in several embodiments.

[0056] VI. Example of a process for predicting MHC-binding peptides presented on the cell surface Figure 19 includes a flowchart 1900 illustrating an example of a method for predicting a surface-presented peptide according to a particular embodiment. The operations described in flowchart 1900 can be performed by a computer system implementing a trained machine learning model, such as a trained binding and presentation model. While flowchart 1900 can describe the operations as a sequential process, in various embodiments, many of the operations can be performed in parallel or simultaneously. Furthermore, the order of the operations can be rearranged. The operations may have additional steps not shown in the figure. Moreover, embodiments of the method can be implemented by hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof. When implemented by software, firmware, middleware, or microcode, the program code or code segments that perform the relevant tasks can be stored on a computer-readable medium such as a storage medium.

[0057] In operation 1910, the computer system accesses a machine learning model. The machine learning model is trained using a training dataset that, for each of several peptides identified by the training dataset, includes the protein properties of the MHC molecule (e.g., HLA allele) that binds to and presents the peptide on the cell surface, one or more expression levels representing the expression level of the gene encoding the peptide, and one or more peptide presentation metrics representing the amount of the peptide detected as being presented by the MHC molecule. The machine learning model is configured to produce an output indicating the degree to which one or more expression levels and one or more peptide presentation metrics are related according to the population-level relationship between expression and presentation.

[0058] In operation 1920, the computer system accesses genomic and transcriptome data corresponding to the biological sample of interest. The genomic and transcriptome data of the biological sample are processed to identify candidate neoantigens (peptides). The genomic and transcriptome data identify one or more MHC molecules from the biological sample and, for each peptide in the set of peptides identified from the tissue sample (e.g., candidate neoantigens), include one or more values ​​representing that peptide. At least one of the one or more values ​​can be determined based on the processing of the tissue sample. The one or more values ​​can correspond to the type of peptide, the length of the peptide, the allele to which the peptide binds, and the expression of the gene region encoding the peptide.

[0059] In operation 1930, the computer system determines a score for each peptide in the set of peptides using a machine learning model, one or more MHC molecules identified from a biological sample, and one or more values ​​representing the peptide in genomic and transcriptome data. In some cases, the computer system uses a trained machine learning model to process one or more values ​​and output a score for a given peptide that predicts MHC molecule binding and presentation.

[0060] In operation 1940, the computer system generates results based on the scores. The results may include an incomplete subset of peptides where the peptide subset exceeds a predetermined threshold expected to be a surface-presented peptide. In some cases, the results may include motifs corresponding to each of the peptide subsets. Additionally or alternatively, the results may include a subset of peptides with scores above a particular ranking percentile (e.g., 0.02). In some cases, for each peptide in the set of peptides, the results indicate whether that peptide is a surface-presented peptide, i.e., whether it binds to the corresponding MHC molecule and is presented on the cell surface.

[0061] In some cases, the computer system selects an incomplete subset of the peptide set, and the identification of the incomplete subset is performed in a way that biases the selection to peptides associated with regions in space, where the expression levels and peptide presentation metrics relate to outlier peptides in the training dataset in a way that deviates from population-level relationships.

[0062] In operation 1950, the computer system outputs the result. Afterward, process 1900 terminates.

[0063] VII. Computer Environment Figure 20 shows an example of a computer system 2000 for carrying out some embodiments of the embodiments disclosed herein. The computer system 2000 may have a distributed architecture in which some components (e.g., memory and processors) are part of an end-user device and some other similar components (e.g., memory and processors) are part of a computer server. The computer system 2000 includes at least a processor 2002, memory 2004, storage device 2006, input / output (I / O) peripherals 2008, communication peripherals 2010, and an interface bus 2012. The interface bus 2012 is configured to communicate, transmit, and transfer data, control, and commands between various components of the computer system 2000. The processor 2002 may include one or more processing units such as a CPU, GPU, TPU, systolic array, or SIMD processor. Memory 2004 and storage devices 2006 include computer-readable storage media, such as RAM, ROM, electrically erasable programmable read-only memory (EEPROM), hard drives, CD-ROMs, optical storage devices, magnetic storage devices, electronically non-volatile computer storage devices, such as Flash® memory, and other tangible storage media. Any such computer-readable storage media can be configured to store instructions or program code that embody aspects of the present disclosure. Memory 2004 and storage devices 2006 also include computer-readable signal media. Computer-readable signal media include propagating data signals in which computer-readable program code is embodied. Such propagating signals may take any of various forms, including but not limited to electromagnetic, optical, or any combination thereof. Computer-readable signal media include any computer-readable medium that is not a computer-readable storage medium and can communicate, propagate, or transmit programs for use in connection with a computer system 2000.

[0064] Furthermore, memory 2004 includes the operating system, programs, and applications. The processor 2002 is configured to execute stored instructions and includes, for example, logic processing units, microprocessors, digital signal processors, and other processors. Memory 2004 and / or processor 2002 can be virtualized and hosted, for example, in another computing system in a cloud network or data center. I / O peripherals 2008 include user interfaces such as keyboards, screens (e.g., touchscreens), microphones, speakers, and other input / output devices, and computing components such as graphical processing units, serial ports, parallel ports, universal serial buses, and other input / output peripherals. I / O peripherals 2008 are connected to processor 2002 through any of the ports connected to interface bus 2012. Communication peripherals 2010 are configured to facilitate communication between the computer system 2000 and other computing devices via a communication network and include, for example, network interface controllers, modems, wireless and wired interface cards, antennas, and other communication peripherals.

[0065] While the subject matter of the present invention is described in detail with respect to specific embodiments thereof, it will be understood that those skilled in the art, having obtained the foregoing understanding, can readily create modifications, variations, and equivalents to such embodiments. Therefore, it should be understood that this disclosure is presented for illustrative purposes only, and not as an limitation, and does not preclude the inclusion of modifications, variations, and / or additions to the subject matter which would be readily apparent to those skilled in the art. Indeed, the methods and systems described herein can be embodied in a variety of other forms, and furthermore, various omissions, substitutions, and modifications in the forms of the methods and systems described herein can be made without departing from the spirit of this disclosure. The appended claims and their equivalents are intended to cover forms or modifications which fall within the scope and spirit of this disclosure.

[0066] Unless otherwise specifically stated, throughout this specification, any use of terms or similar expressions such as “processing,” “computing,” “calculating,” “determining,” and “identifying” is understood to mean the operation or process of a computing device, such as one or more computers or similar electronic computing devices or devices, that manipulate or transform data represented as physical electronic or magnetic quantities in the memory, registers, or other information storage devices, transmission devices, or display devices of a computing platform.

[0067] The systems or systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditioned on one or more inputs. Suitable computing devices include multipurpose microprocessor-based computing systems that access stored software that programs or configures a computing system, ranging from general-purpose computing devices to specialized computing devices that implement one or more embodiments of the subject matter of the present invention. The teachings contained herein may be implemented in the software used to program or configure a computing device using any suitable programming, scripting, or other type of language or combination of languages.

[0068] Embodiments of the methods disclosed herein can be implemented in the operation of such computing devices. The order of the blocks presented in the above examples can be changed, for example, by rearranging, combining, and / or dividing the blocks into subblocks. Certain blocks or processes can be executed in parallel.

[0069] The conditional language used herein, particularly "can," "could," "might," "may," "eg," and similar expressions, is generally intended to convey that certain examples include certain features, elements, and / or steps, while other examples do not, unless otherwise specifically stated or understood within the context in which they are used. Therefore, such conditional language is generally not intended to imply that features, elements, and / or steps are somehow required in one or more examples, nor to necessarily include logic for determining, with or without author input or prompting, whether one or more examples include these features, elements, and / or steps, or whether they are performed in a particular example.

[0070] The terms “comprising,” “including,” and “having,” and similar terms, are synonymous and are used in an open-ended and inclusive manner, without excluding additional elements, features, actions, operations, etc. The term “or” is also used in an inclusive (and not exclusive) sense, and as a result, when used, for example, to connect an enumeration of elements, “or” means one, some, or all of the elements in the enumeration. The use of “adapted to” or “configured to” herein means an open and inclusive language that does not exclude devices adapted to or configured to perform additional tasks or steps. Furthermore, the use of “based on” means that a process, step, calculation, or other action “based on” one or more described conditions or values ​​may actually be based on additional conditions or values ​​beyond those described. Similarly, the use of “based at least in part on” means open and inclusive, in that a process, step, calculation, or other action “based at least in part on” one or more of the stated conditions or values ​​may actually be based on additional conditions or values ​​beyond those stated. The headings, lists, and numbering included herein are for illustrative purposes only and not to limit.

[0071] The various functions and processes described above can be used independently of each other or in various combinations. All possible combinations and sub-combinations are intended to fall within the scope of this disclosure. Furthermore, some implementations may omit certain methods or process blocks. The methods and processes described herein are also not limited to a specific order, and the associated blocks or states may be executed in any other appropriate order. For example, the described blocks or states may be executed in an order other than that specifically disclosed, or multiple blocks or states may be combined into a single block or state. Examples of blocks or states may be performed sequentially, in parallel, or in any other way. Blocks or states may be added to or removed from the examples of disclosure. Similarly, the examples of systems and components described herein may have different configurations than those described. For example, elements may be added, removed, or rearranged compared to the examples of disclosure.

Claims

1. A method of executing a program stored in the memory of a computer system using the processor of that system, (a) Accessing machine learning models stored in a memory device, The aforementioned machine learning model, (i) A training dataset as an attribute is used to train each of the multiple peptides identified by the training dataset, the training dataset being labeled to include the protein properties of the major histocompatibility complex (MHC) molecule to which the peptides bind and present: (1) Protein properties of the MHC molecule to which the peptide is bound and presented; (2) One or more expression levels representing the expression level of the gene encoding the peptide; and (3) One or more peptide presentation metrics representing the amount of peptide detected as presented by the MHC molecule; and, (ii) Accessing the machine learning model configured to generate a score indicating the degree to which the one or more expression levels of a tissue sample and the one or more peptide presentation metrics of the tissue sample are related according to the relationship between the expression levels of the training dataset and the peptide presentation metrics of the training dataset, wherein the relationship between the expression levels and the peptide presentation metrics of the training dataset differs depending on the population level of the training dataset, (b) Accessing genome and transcriptome data corresponding to a target tissue sample stored in the storage device, wherein the genome and transcriptome data identifies one or more MHC molecules from the tissue sample and includes one or more values ​​representing each peptide in the set of peptides identified from the tissue sample, and at least one of the one or more values ​​is determined based on processing of the tissue sample. (c) For each peptide in the set of peptides, a score is determined using the machine learning model, the one or more MHC molecules identified from the tissue sample, and the one or more values ​​representing the peptide. (d) generating results based on the score, (e) Outputting the above result using an input / output (I / O) peripheral device, Includes, (i) The computer system includes the processor, the memory, the storage device, the input / output (I / O) peripherals, the communication peripherals and the interface bus, (ii) The training dataset is obtained from single-allelic data corresponding to peptides derived from single-allelic cell lines and / or multi-allelic data corresponding to peptides derived from other tissue samples. (iii) The attributes of the training data vectors are data corresponding to somatic cell mutants that include one or more features including peptide sequences, (iv) The labels of the data vectors of the training data are peptides encoded by somatic mutants that bind to MHC molecules and are present on the cell surface, (v) The training dataset for training the machine learning model includes sequence data from various sources: (1) peptides identified as binding to HLA molecules based on in vitro experiments, (2) peptides identified by mass spectrometry from tumor samples, (3) HLA alleles, and (4) non-tumor samples. method.

2. The further includes selecting an incomplete subset of the set of peptides based on the score, The method according to claim 1, wherein the identification of the incomplete subset is performed in a manner that biases the selection to peptides associated with scores that are expected to be presented more certainly than the probabilities expected by the corresponding population-level relationships of the training dataset, and the result includes the incomplete subset of the set of peptides.

3. The further includes selecting an incomplete subset of the set of peptides based on the score, The method according to claim 1, wherein the identification of the incomplete subset is performed in a manner that biases the selection to peptides associated with regions in space, the regions being associated with outlier peptides in the training dataset whose expression levels and peptide presentation metrics are related in a manner that deviates from the population-level relationships.

4. The method according to claim 1, wherein the results include, for each of the one or more peptides in the set of peptides, the identification of the peptide and the score.

5. The method according to claim 1, wherein, for each peptide in the set of peptides, the one or more values ​​representing the peptide are generated based on the amino acid sequence of the peptide, an indicator of whether the peptide binds to one or more binding pockets of the MHC molecule, the expression level of the peptide in the tissue sample, and / or the length of the peptide.

6. The method according to claim 1, wherein the score corresponding to a peptide in the set of peptides corresponds to a predicted probability of whether the peptide binds to the MHC molecule and is presented on the cell surface.

7. The method according to claim 1, wherein the machine learning model includes one or more trained gradient boosting algorithms.

8. The method according to claim 1, wherein the machine learning model includes, for each of the plurality of peptides, a first submodel trained on a first subset of the training dataset which includes a sequence corresponding to the peptide, a sequence of an MHC molecule bound to the peptide, and / or the length of the peptide.

9. The method according to claim 8, wherein the machine learning model includes, for each of the plurality of peptides, a second submodel trained on a second subset of the training dataset, which includes the expression levels of one or more source proteins from which the peptide originates and the surface presentation characteristics of the peptide.

10. The method according to claim 9, wherein the first submodel and the second submodel are each trained based on one or more scores generated by another set of submodels.

11. (a) One or more data processors, (b) A non-temporary computer-readable storage medium that includes an instruction causing the one or more data processors to execute the method according to claim 1 when the one or more data processors are executed, A system that includes this.

12. A non-temporary, machine-readable storage medium comprising instructions for causing one or more data processors to perform the method according to claim 1.