Machine learning technique for predicting surface presentation peptide
Patent Information
- Application Number
- JP2024143327
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-11-07
- Filing Date
- 2024-08-23
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2041-06-17
AI Technical Summary
Conventional techniques for predicting MHC-binding peptides fail to accurately determine which peptides will be displayed on the cell surface, leading to incomplete understanding of immune system responses to tumor cells and limitations in developing personalized cancer therapies.
A machine learning model trained using genomic and transcriptomic data, including monoallelic and multi-allelic immunopeptidomic data, to predict peptides that bind to MHC molecules and are presented on the cell surface, utilizing features like genetic propensity scores and hotspot scores to enhance prediction accuracy.
The model achieves significantly higher positive predictive value and specificity in identifying surface-presented peptides, facilitating the development of personalized immunotherapies and biomarkers by accurately predicting immune system responses.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 040,943, filed June 18, 2020, entitled “Composite Biomarkers for Immunotherapy for Cancer,” and U.S. Provisional Patent Application No. 63 / 111,007, filed November 7, 2020, entitled “Machine-Learning Techniques For Predicting Surface-Presenting Peptides,” the entire contents of which are incorporated herein by reference in their entireties for all purposes.
[0002] The present disclosure relates to machine learning techniques for predicting surface-displayed peptides. [Background technology]
[0003] Cancers contain mutations, which can be somatic or tumor-specific. The immune system detects these cancer-based mutations by identifying peptides derived from these mutations. Peptides are identified by the immune system when they bind to proteins encoded by major histocompatibility complex (MHC) genes and are presented on the surface of cells. For example, peptides corresponding to mutated genes can bind to specific MHC molecules (e.g., human leukocyte antigen (HLA) proteins) and be presented on the cell surface. Predicting peptides expressed on tumor cell surfaces can inform the development of precision cancer treatments and diagnostics. For example, genomic variants corresponding to these peptides can be identified to analyze the response and resistance of complex systems to certain cancer immunotherapies. As another example, peptides presented on tumor cell surfaces can be analyzed to generate personalized immuno-oncology (IO) therapies and / or neoantigen cancer vaccines.
[0004] Such peptides expressed on the tumor cell surface are also known as "neoantigens", and techniques to predict them require thorough analysis of many technical factors, including but not limited to the quality of peptide sequencing data, availability of tumor and normal sample pairs, HLA typing, and identification of other peptide characteristics. For example, neoantigens can be identified based on prediction of peptides that bind to MHC molecules and are presented on the cell surface. To identify neoantigens, determining the peptides encoded by somatic variants and identifying the HLA molecules that bind to the peptides is only the first step of a very complex process. The reason is that each peptide identified from the sequence data may or may not be processed by the proteasome; transported for MHC binding; presented on the tumor cell surface; and finally recognized by the immune system. Due to this complex process, many peptides that bind to HLA molecules (for example) may not be expressed on the cell surface.
[0005] Furthermore, one or more binding motifs of an MHC molecule can be identified to determine whether a given peptide binds to the MHC molecule. Although the binding motifs of some MHC molecules (e.g., HLA-A molecules) are known, there are many MHC molecules for which the binding motif has not yet been identified. For example, the binding motifs of MHC class II molecules are relatively unknown due to limited availability of experimental data. Without that information, it would be difficult to determine whether a peptide binds to a corresponding MHC molecule. Conventional techniques have attempted to address this problem by using known MHC binding motifs to train machine learning models to predict whether a peptide will bind to one of the various types of MHC molecules. However, even if such peptides are identified, some peptides are not present on the cell surface. In other words, conventional techniques can identify MHC-binding peptides, but only a small portion of them can be successfully presented on the cell surface. Because the immune system response is triggered when the MHC-binding peptide is presented on the cell surface, identifying the MHC-binding peptide alone cannot provide all the details of how the immune system responds to tumor cells, foreign proteins, etc.
[0006] Thus, conventional techniques for predicting MHC-binding peptides do not address whether a peptide is actually presented and expressed on the cell surface. Furthermore, conventional techniques are insufficient to identify peptide characteristics that indicate that a given peptide is presented on the cell surface. Therefore, there is a need to accurately predict which peptides bind to their corresponding MHC molecules and are presented on the cell surface. Summary of the Invention [Problem to be solved by the invention]
[0007] In some embodiments, a method of predicting surface-presented peptides is provided. The method can include accessing a trained machine learning model that has been trained using a training dataset and that includes, for each peptide of a plurality of peptides identified by the training dataset, a protein signature of a major histocompatibility complex (MHC) molecule that binds and presents the peptide, one or more expression levels representing the expression level of a gene encoding the peptide, and one or more peptide presentation metrics representing the amount of the peptide detected as being presented by the MHC molecule. The machine learning model can be configured to generate an output that indicates the degree to which the one or more expression levels and the one or more peptide presentation metrics are related according to a population-level relationship between expression and presentation. [Means for solving the problem]
[0008] The method may also include accessing genomic and transcriptomic data corresponding to the subject's biological sample. The genomic and transcriptomic data may identify one or more MHC molecules from the biological sample and, for each peptide of the set of peptides identified from the cell line or tissue sample, one or more values representing the peptide. The one or more values may be determined based on processing of the tissue sample. The method may also include determining a score for each peptide of the set of peptides using the machine learning model, the one or more MHC molecules identified from the biological sample, and the one or more values representing the peptide. The method may include generating a result based on the score and outputting the result.
[0009] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed by the one or more data processors, cause the one or more data processors to perform some or all of the methods and / or some or all of the processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium, the product including instructions configured to cause the one or more data processors to perform some or all of the methods and / or some or all of the processes disclosed herein.
[0010] The terms and expressions which have been employed are used as terms of description, not of limitation, and in the use of such terms and expressions there is no intention to exclude any equivalents or portions thereof of the features shown and described, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, although the invention as claimed has been specifically disclosed by embodiments and optional features, it is to be understood that modifications and variations of the concepts disclosed herein may be utilized by those skilled in the art, and that such modifications and variations are considered to be within the scope of the invention as defined by the appended claims. [Brief description of the drawings]
[0011] The present disclosure will be described in conjunction with the accompanying figures. [Figure 1] FIG. 1 shows a schematic diagram of peptides binding to MHC molecules and being presented on the cell surface. [Diagram 2] FIG. 2 shows a schematic diagram depicting peptides that can be displayed on the cell surface for gene therapy. [Diagram 3] FIG. 3 shows a schematic diagram of identifying single-allele immunopeptidomic data that can be used to train a machine learning model, according to some embodiments. [Figure 4]FIG. 4 shows allelic diversity data corresponding to MHC binding peptides, according to some embodiments. [Diagram 5] FIG. 5 shows source diversity data identified from tissue and cell line samples for training a machine learning model for predicting surface-displayed peptides, according to some embodiments. [Figure 6] FIG. 6 shows a plot of comparative data between expected peptide numbers based on gene expression levels and actual observed peptide numbers, according to some embodiments. [Figure 7] FIG. 7 shows a process for determining genetic propensity scores used to train a machine learning model, according to some embodiments. [Figure 8] FIG. 8 shows a plot of comparative data between the number of peptides expected for one or more regions within a gene and the number of peptides actually observed for that region, according to some embodiments. [Figure 9] FIG. 9 illustrates a process for determining hotspot scores used to train a machine learning model, according to some embodiments. [Figure 10] FIG. 10 shows examples of features used in the binding model and the presentation model, according to some embodiments. [Figure 11] FIG. 11 shows an exemplary model architecture for training a machine learning model for predicting surface-displayed peptides, according to some embodiments. [Figure 12] FIG. 12 shows the performance levels of the trained combined model and the trained presentation model, measured by positive predictive value, based on 10% holdout data, according to some embodiments. [Figure 13] FIG. 13 shows a comparison of the performance levels of the trained machine learning models compared to conventional techniques for predicting surface-displayed peptides. [Figure 14] FIG. 14 shows a comparison of the performance levels of the trained display models across different alleles compared to conventional techniques for predicting surface-displayed peptides. [Figure 15]FIG. 15 shows the results of a leave-one-out analysis of the trained representation model, according to some embodiments. [Figure 16] FIG. 16 illustrates a graph showing precision and recall for evaluating a trained machine learning model, according to some embodiments. [Figure 17] FIG. 17 shows a box plot depicting the performance level of a trained machine learning model across various tissue samples, according to some embodiments. [Figure 18] FIG. 18 shows a graph comparing the performance levels of a trained machine learning model in accordance with some embodiments with other prior art. [Figure 19] FIG. 19 includes a flow chart illustrating an example of a method for predicting surface-displayed peptides, according to certain embodiments. [Figure 20] FIG. 20 illustrates an example of a computer system for implementing some of the embodiments disclosed herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] I. Overview To address at least the above-mentioned shortcomings of conventional systems, the present technology can be used to predict surface-presented peptides. As used herein, "surface-presented peptide" can refer to a peptide that binds to an MHC molecule (e.g., HLA-A protein) and is presented on the corresponding cell surface. One or more somatic variants can be identified by sequencing DNA from normal and tumor samples. The somatic variants include one or more genetic mutations present in the tumor and normal samples. The somatic variants of the tumor sample can be processed using a trained machine learning model to predict whether a peptide encoded by the somatic variant binds to an MHC molecule (e.g., MHC class 1) and is presented on the cell surface. The machine learning model can include a binding model that predicts whether a peptide encoded by the somatic variant binds to an MHC molecule. In some embodiments, the machine learning model includes a presentation model that predicts whether a peptide encoded by the somatic variant is expressed on the cell surface.
[0013] The machine learning model can be trained using training datasets obtained from (i) engineered monoallelic cell lines and (ii) biallelic data of tissue samples of other subjects. In some cases, the machine learning model has been trained using binding array data (e.g., IEDB data). The training dataset can include, for each peptide identified by the training dataset, one or more expression levels representing the expression level of the gene encoding said peptide and one or more peptide presentation metrics representing the amount of the peptide detected as presented by the MHC molecule. The training dataset can include immunopeptidomic data of peptides generated from multiple engineered cell lines (e.g., K562 cells) expressing a single allele of interest (e.g., HLA-A). In particular, MHC-peptide complexes in these cell lines can be immunoprecipitated using W6 / 32 antibody, followed by peptide elution and peptide sequencing using tandem mass spectrometry. Training datasets corresponding to biallelic data from other tissue samples can be obtained using curated public data.
[0014] The prediction of surface-displayed peptides can be performed in a manner that biases selection towards peptides associated with scores that predict presentation that is more likely than would be expected by a population-level relationship between peptide expression and presentation. Additionally or alternatively, the prediction of surface-displayed peptides can be performed in a manner that biases selection towards peptides associated with regions in space that are associated with outlier peptides in the training dataset where expression levels and peptide presentation metrics are related in a way that deviates from the population-level relationship.
[0015] Thus, the embodiments of the present disclosure provide technical advantages over conventional systems by accurately predicting peptides that bind to their corresponding MHC molecules and are presented on the cell surface. As previously discussed, tumor cell surface peptide binding and expression can predict how the immune system will respond to neoantigens and / or certain cancer immunotherapies. Thus, accurate prediction of surface-presented peptides facilitates the selection or development of immunotherapies that will be most effective for a given subject. Furthermore, based on model evaluation, the embodiments show significantly higher positive predictive value compared to conventional techniques such as NetMHCPan 4.0. Thus, the high sensitivity and specificity of the embodiments allows for accurate identification of MHC-bound peptides presented on the cell surface, thereby facilitating application in the development of personalized immunotherapies and biomarker development.
[0016] The following examples are provided to introduce certain embodiments. In the following description, for purposes of explanation, specific details are set forth to provide a thorough understanding of the disclosed examples. However, it will be apparent that various examples may be practiced without these specific details. For example, devices, systems, structures, assemblies, methods, and other components may be shown as components in block diagram form so as not to obscure the examples in unnecessary detail. In other cases, well-known devices, processes, systems, structures, and techniques may be shown without necessary details so as to avoid obscuring the examples. The figures and descriptions are not intended to be limiting. The terms and expressions used in this disclosure are used as terms of description rather than limitation, and there is no intention in the use of such terms and expressions to exclude any equivalents of the features shown and described, or portions thereof. The term "example" is used herein to mean "to serve as an example, instance, or illustration." Any embodiment or design described herein as an "example" is not necessarily to be construed as preferred or advantageous over other embodiments or designs.
[0017] II. Surface display of peptides 1. Neoantigens in tumor samples Neoantigens can be found in tumor samples, where a neoantigen refers to one or more peptides that are presented on the tumor cell surface, thereby triggering an immune system response. The immune system can be tuned to look for pathogens, including cancer, and thus has the ability to cure cancer. The immune system can distinguish between self and non-self antigens. Because tumors are caused by genetic mutations (e.g., somatic variants), peptides that correspond to these genetic mutations and are expressed on the cell surface can be considered neoantigens. These peptides are considered "new" to the immune system, so ideally, the immune system can recognize and eliminate tumor cells based on detecting neoantigens presented on the tumor cell surface. As explained above, tumor samples can be analyzed to reveal sequence data, which can be compared to those of normal samples to identify somatic variants. The somatic variants can be further analyzed to determine which subsets of variants will appear as peptides. Neoantigens can be predicted by identifying peptides that bind to MHC molecules and are presented on the cell surface. Thus, the ability of peptides to be presented on the cell surface can be a key factor for developing immunotherapies against cancer.
[0018] 2. Peptides that respond to the treatment of certain autoimmune diseases Surface-presented peptides can be identified in the context of autoimmune diseases, where the peptides are encoded based on genetic alterations resulting from certain immunotherapies. Figure 2 shows a schematic diagram illustrating surface-presented peptides that respond to gene therapy. In Figure 2, a mutation in the dystrophin gene is shown, which typically causes debilitating muscular dystrophies. The dystrophin gene encodes the dystrophin protein molecule, which acts as a cushioning agent for muscle cells as a shock absorber. The lack of a fully functional dystrophin protein can lead to muscle degeneration. Typically, muscular dystrophies can be treated with exome skipping therapy, which skips exomes (e.g., exon 52) responsible for the dystrophin gene mutation and generates semi-functional dystrophin protein for the subject. Although exome skipping therapy can be effective, it can induce the generation of new types of peptides due to the intentional skipping of exomes through genetic modification. The new peptides can bind to MHC molecules and be presented on the cell surface, which can trigger destructive immune system responses.
[0019] III. Example of training dataset The machine learning model for predicting surface-presented peptides can be trained using a supervised training algorithm. The machine learning model can be trained using a training dataset. The training dataset for training the machine learning model can include sequence data from various sources: (i) peptides identified as binding to HLA molecules based on in vitro experiments, (ii) peptides identified by performing mass spectrometry on tumor samples, (iii) HLA alleles, and (iv) non-tumor samples. However, some training sequence data may be inaccurate for training the machine learning model. For example, training sequence data generated from tissue samples will require the difficult process of mapping peptides to one of several types of HLA proteins (e.g., HLA-A, HLA-B) that are simultaneously expressed on the cell surface. In another example, sequence data generated using in vitro methods may not mimic surface presentation. An embodiment of the present disclosure for systematically resolving inconsistencies in the training dataset is used to train a machine learning model that predicts peptides that are likely to be "shuttled" to the cell surface from somatic variants called from the sequence data.
[0020] Additionally or alternatively, the training dataset can further include data corresponding to somatic variants, each somatic variant being labeled to indicate whether a peptide encoded by the somatic variant binds to an MHC molecule (e.g., an HLA-A protein) and is presented on the cell surface. The training dataset can also include one or more features derived from the somatic variants (e.g., peptide sequence, peptide length, expression of the peptide in a tumor sample).
[0021] To prepare a training dataset, tumor samples and corresponding normal control samples can be sequenced to generate tumor-normal paired sequence data. The tumor-normal paired sequence data are compared to identify somatic variants, including single nucleotide variants (SNVs), indels, and / or altered genes including copy number variations. In some cases, a machine learning model is used to process the tumor-normal paired sequence data to identify somatic variants in the tumor sample.
[0022] 1. Training Data Source a) Single allele immunopeptidomic data In some cases, at least a portion of the training data corresponds to peptides identified from engineered monoallelic cell lines. FIG. 3 shows a schematic diagram of identifying monoallelic immunopeptidomic data that can be used to train a machine learning model, according to some embodiments. As shown in FIG. 3, a monoallelic engineered K562 cell line can be created and then transfected with a specific HLA molecule of interest (e.g., HLA-B) (step 305). As previously mentioned, the HLA complex is a group of related proteins that are encoded by the MHC gene complex in humans. These cell surface proteins are responsible for regulating the immune system. From the cell line, HLA-binding peptides can be identified by immunoprecipitating the HLA-peptide complexes using the W6 / 32 antibody (step 310), applying peptide elution (step 315), and performing peptide sequencing on the eluted peptides using mass spectrometry (e.g., liquid chromatography-mass spectrometry, mass spectrometry) (step 320). Thus, HLA-binding peptides can be identified for that specific HLA molecule of interest (step 325).
[0023] Single-allelic immunopeptidomic data identifying various features of HLA-binding peptides can be revealed and included as part of the training data set. Examples of training data from single-allelic immunopeptidomic data can include, for a given HLA-binding peptide, the type of peptide, the length of the peptide, the amino acid sequence of the peptide, the HLA alleles that bind to the peptide, the number of transcripts that correspond to the peptide, and the expression of the gene region that codes for the peptide. To optimize the performance of the machine learning model, training data was generated to represent the HLA genotypes of the general population. For example, FIG. 4 shows allele diversity data corresponding to HLA-binding peptides according to some embodiments. To determine the allele diversity data, the identified peptides can be clustered based on similarity to peptide sequences corresponding to all known alleles of the HLA molecule of interest (e.g., HLA alleles identified from the IMGT database). Thus, the identified peptides can be clustered based on the similarity of their respective binding pockets. In some cases, the identified peptides can be clustered using a BLOSUM similarity matrix. Based on these clusters, one or more alleles that code for the HLA-binding peptide can also be identified. In some cases, the peptide clusters are visualized on a heat map. For example, Figure 4 shows a first heat map identifying allelic diversity of HLA-A molecules and a second heat map identifying allelic diversity of HLA-B molecules. Additionally or alternatively, the training data set can be augmented with training data corresponding to allele frequency data for alleles encoding HLA binding proteins, where the allele frequency data is classified for different portions of the world's population.
[0024] Training data corresponding to HLA-binding peptides can facilitate training of machine learning models by using single-allelic immunopeptidomic data expressing one particular type of HLA at a time. Furthermore, the allelic diversity in the single-allelic immunopeptidomic training data allows the machine learning model to predict surface-presented peptides derived from various alleles that may not be present in the training data.
[0025] b) Multi-allele immunopeptidomic data In some cases, at least a portion of the training data corresponds to peptides identified from sequencing tissue samples of other subjects. Various tissue samples or cell lines of the tissue samples of the subject can be sequenced to identify multiple peptides that bind to different types of HLA molecules (e.g., HLA-A, HLA-B, HLA-C). In some cases, the cell lines and tissue samples are processed using mass spectrometry. Multi-allelic immunopeptidomic data obtained from the multiple identified peptides can be used as part of the training data. The multi-allelic immunopeptidomic data can include various features corresponding to the identified peptides, including peptide length and allelic diversity. Figure 5 shows source diversity data identified from tissue samples of the subject for training a machine learning model for predicting surface-displayed peptides according to some embodiments. In Figure 5, the amount of each type of peptide is shown for each of the single and multi-allelic samples. Additionally or alternatively, the multi-allelic data can be obtained from public data sources.
[0026] Multi-allelic immunopeptidomic data generated from various tissues and cell lines can be integrated into the training dataset to improve the performance of the trained machine learning model. In particular, overfitting and / or underfitting can be reduced by training the machine learning model with multi-allelic immunopeptidomic data. For example, both mono-allelic and bi-allelic immunopeptide peptides from several publicly available data sources can be added to the training dataset. Mono-allelic immunopeptidomic data from genetically engineered cell lines and mono-allelic and bi-allelic immunopeptidomic data from tissue samples can all be integrated into the training dataset to expand its size (e.g., a larger number of unique peptides).
[0027] 2. Additional enhancing features As explained above, the immunopeptidomic data from the training dataset identifies various features of HLA-bound peptides, including peptide sequence, peptide length, binding pocket sequence, left flanking region, and right flanking region. In some cases, the training dataset also includes antigen presentation features, such as expression levels of peptides as measured by DPM. In addition to the above, two additional features can be generated from the immunopeptidomic data and used to enhance the training dataset.
[0028] a) Comparison data between predicted peptide numbers based on gene expression levels and actual observed peptide numbers The first features generated from the immunopeptidomic data can include comparison data between expected peptide counts based on gene expression levels and actual observed peptide counts. By including a training dataset with the first features, a trained machine learning model trained from the training data can improve prediction of surface-displayed peptides, such that predictions are biased towards peptides associated with scores predicting more certain presentation compared to the probability expected by the population-level relationship between peptide expression and presentation. Additionally, a trained machine learning model trained from the training data can facilitate prediction of surface-displayed peptides, such that predictions are performed in a manner that biases selection towards peptides associated with regions in space associated with outlier peptides in the training dataset where expression levels and peptide presentation metrics are related in a manner that deviates from the population-level relationship.
[0029] FIG. 6 shows a plot of comparative data between expected peptide counts based on gene expression levels and actual observed peptide counts, according to some embodiments. To generate the comparative data, all transcripts corresponding to HLA-binding peptides from a training dataset can be identified and organized into a set of bins based on their respective gene expression levels. For example, as shown in FIG. 6, the x-axis shows 10 sections (e.g., deciles) into which transcripts can be grouped based on their respective gene expression levels. The bars in each section show the measured amount of gene expression levels corresponding to the transcripts grouped in the section. The y-axis of the plot shown in FIG. 6 shows the number of peptides, and the diamond points show the amount of peptides counted from the training sample cell lines.
[0030] The initial hypothesis without the comparative data in FIG. 6 appears to indicate that the expected peptide abundance is directly proportional to the measured gene expression level. However, with the comparative data in FIG. 6, one or more outliers can be identified that deviate from the initial hypothesis. A first outlier 605 includes a large amount of peptide observed in bin "1", indicating a very low amount of gene expression level. A second outlier 610 includes very little amount of peptide observed in bin "10", indicating a high amount of gene expression level. Thus, the comparative data in FIG. 6 indicates that measuring the gene expression level of an HLA-binding peptide may not be sufficient to predict whether the HLA-binding peptide will actually be presented on the cell surface. Using the comparative data, a genetic propensity score ("gps") can be calculated for the gene encoding the HLA-binding peptide, which genetic propensity score predicts whether the peptide will be presented on the cell surface. In some cases, the calculated genetic propensity score can be added as an additional feature of the training dataset, so that a machine learning model can be further trained based on the genetic propensity scores.
[0031] 7 illustrates a process for determining gene propensity scores used to train a machine learning model, according to some embodiments. In block 705, immunopeptidomics data is obtained, the immunopeptidomics data including expression levels of genes encoding HLA-binding peptides. For example, the immunopeptidomics data can be obtained by reprocessing existing mass spectrometry (MS) data or by accessing the immunopeptidomics data directly from a database (e.g., an immunopeptide database).
[0032] At block 710, the expected peptide counts for the genes identified in the immunopeptidomic data are calculated. In particular, the expected peptide counts are calculated based on the number of transcripts (e.g., TPMs) and the sequence length of the genes. At block 715, a ratio between the expected peptide counts and the observed peptide counts can be calculated to generate a gene propensity score (e.g., log10(observed / expected)). In some cases, the gene propensity score is added as an additional feature to the training dataset.
[0033] b) Comparison of predicted and observed peptide numbers per gene region The second features generated from the immunopeptidomic data can include a comparison of the expected number of peptides based on the expression levels in one or more regions of a given gene to the number of actually observed peptides corresponding to the one or more regions. In contrast to the first features that identify gene expression levels across a range of genes, the second features identify the expression levels of a region within a single gene. Based on the identified expression levels, an expected amount of peptides can be generated. The expected amount can be compared to the observed amount of peptides to identify a training dataset second feature, where the second feature is indicative of one or more surface display properties of the region within the corresponding gene.
[0034] In some cases, the first feature and the second feature are combined in a training dataset. A trained machine learning model trained from training data with the combined features can facilitate prediction of surface-displayed peptide predictions, such that predictions are biased towards peptides associated with scores predicting more certain presentation compared to the probability expected by a population-level relationship between peptide expression and presentation. Additionally, a trained machine learning model trained from the training data can facilitate prediction of surface-displayed peptide predictions, such that predictions are performed in a manner that biases selection towards peptides associated with regions in space associated with outlier peptides in the training dataset whose expression levels and peptide presentation metrics are related in a manner that deviates from the population-level relationship.
[0035] FIG. 8 shows a plot of comparative data between expected peptide counts for one or more regions in a gene and the actual observed peptide counts for that region, according to some embodiments. For each genomic region of a given gene, gene expression levels can be calculated and the expected peptide amounts can be measured. For example, the expected peptide amounts for the genomic region of the ACTB gene are shown in black plot line 805. The expected peptide amounts can then be compared to the observed peptide amounts for each genomic region, with the observed peptide amounts shown in grey areas 810. As in FIG. 6, some outliers can be identified within various genomic regions, where the observed peptide amounts are not proportional to the measured gene expression levels. For example, a region 815 (e.g., region number 230) of the ACTC1 gene may show a very high expected peptide amount (e.g., >3000 peptides), but the observed amount of peptides for the same region 815 is actually much less than the expected amount (e.g., about 1000 peptides). Thus, the comparative data in Figure 8 indicates that measuring the region-level gene expression levels of an HLA-binding peptide may not be sufficient to predict whether the HLA-binding peptide will actually be presented on the cell surface. Using the comparative data shown in Figure 8, a hotspot score ("hhs") for a gene encoding an HLA-binding peptide can be calculated, which predicts whether a peptide corresponding to a region of the gene will be presented on the cell surface.
[0036] 9 illustrates a process for determining hotspot scores used to train a machine learning model, according to some embodiments. In block 905, immunopeptidomics data is obtained, the immunopeptidomics data including expression levels of genes encoding HLA-binding peptides. For example, the immunopeptidomics data can be obtained by reprocessing existing mass spectrometry (MS) data or by accessing the immunopeptidomics data directly from a database (e.g., an immunopeptide database).
[0037] In step 910, the predicted peptide counts for each region of a particular gene are compared to the actual peptide counts for that region. In step 915, a hotspot score is calculated for the particular gene, which specifies the distribution of peptide counts observed across regions of the gene (e.g., ACTB gene, ACTC1 gene).
[0038] IV. Example of a model architecture for predicting MHC-binding peptides presented on the cell surface The training data set can be used to train a machine learning model for predicting surface-displayed peptides. The machine learning model includes one or more sub-models configured to identify binding and surface-display characteristics of peptides in a sample. These sub-models can be trained separately with corresponding subsets of the training data set, such that each sub-model is capable of predicting surface-displayed peptides based on parameters learned from features corresponding to the subset.
[0039] 1. Binding model and presentation model In some cases, the machine learning models include binding models and presentation models, each trained to handle different features of the input data. FIG. 10 shows examples of features used by the binding model 1005 and presentation model 1010, according to some embodiments. The binding model 1005 can be trained using a training dataset that includes information related to a set of peptides (e.g., the sequence of the MHC molecule that binds the peptide, the length of the peptide). In some cases, the binding model 1005 includes one or more trained gradient boosting algorithms. Gradient boosting refers to a machine learning technique for regression and classification problems that creates a predictive model in the form of an ensemble of weak predictive models. The technique can generalize the model by building the model incrementally and allowing the optimization of any differentiable loss function. Gradient boosting combines the weak learners into a single strong learner in an iterative manner. As each weak learner is added, a new model is fitted, resulting in a more accurate estimate of the response variable. The new weak learner can be maximally correlated with the negative gradient of the loss function, which is associated with the entire ensemble. Examples of gradient boosting machines include XGBoost and LightGBM. Additionally or alternatively, the combined model can be constructed using other types of machine learning techniques, including bagging techniques, boosting techniques, and / or random forest algorithms.
[0040] The presentation model 1010 can be trained using information related to the peptide (e.g., the peptide sequence, the sequence of the MHC molecule that binds the peptide, the length of the peptide), as well as information related to the expression level of the source protein from which the peptide is derived, the surface presentation properties of the peptide, gene propensity scores, and hotspot scores. Thus, the trained presentation model 1010 can identify the binding properties of a given peptide and its surface presentation properties, i.e., whether the peptide is presented on the cell surface. Similar to the binding model 1005, the presentation model 1010 can include one or more trained gradient boosting algorithms.
[0041] 2. Model Architecture FIG. 11 illustrates an exemplary model architecture for training a machine learning model for predicting surface-displayed peptides, according to some embodiments. As illustrated in FIG. 11, the training database is represented by cylinders containing various types of information and includes allele data obtained from publicly available sources. For example, the dark grey cylinders include immunopeptidomics data corresponding to engineered monoallelic cell lines (see FIG. 4). In another example, the training database can also include in vitro binding data from publicly available data sources (e.g., the IEDB database represented by the white cylinders). In some cases, the training datasets from each training database are used to train the corresponding binding and presentation models individually. Additionally, the training databases can be merged into a larger training database to train its corresponding binding and presentation models (e.g., the "ALL(MONO)" light grey cylinders in FIG. 11).
[0042] FIG. 11 further illustrates multiple sets of binding and presentation models trained to predict surface-displayed peptides. Each set of binding and presentation models is shown as being trained with a different training data set. In some cases, the output generated from a first set of models 1105 ("initial models") is used as input features to train a second set of models 1110 ("intermediate models"). For example, the output generated by the initial models corresponding to in vitro binding, single-allelic data from an engineered mono-allelic cell line can be used as input features to train the intermediate models. The intermediate models can also be trained separately with training data obtained from all the single-allelic data 1115. Additionally, the output from the intermediate models can be deconvolved and added to another training database that includes both single-allelic and multi-allelic data. The output can be deconvolved using one or more base mono-allelic base models, or unsupervised clustering and alignment algorithms such as GibbsCluster.
[0043] A training dataset from the database containing the deconvoluted set of peptides for each HLA allele can be used to train a third set of presentation and binding models 1120 ("final models"). The trained final models 1120 can be deployed to predict surface-presented peptides. The goal in building the training database and training the final models is to obtain as much allelic diversity as possible and avoid problems caused by underfitting and overfitting. Additionally or alternatively, the intermediate trained models can also be deployed to predict surface-presented peptides, although the performance level of the trained final models tends to be superior to that of the intermediate trained models.
[0044] V. Evaluating the performance level of machine learning models To evaluate the performance of the trained machine learning model, a test dataset is generated that includes several experimentally observed peptides and synthetic decoys that are not part of the training process. The trained machine learning model processes these candidate test peptides to output scores predicting MHC class I binding and cell surface presentation, the machine learning model being trained using a large immunopeptidome training dataset as described above. The scores are then compared to corresponding data obtained from validated MHC-binding peptides presented on the cell surface to reveal the performance level of the trained machine learning model. The output scores are also evaluated against NetMHCpan 4.0, a known platform that predicts binding of peptides to MHC molecules, and the trained machine learning algorithm shows higher overall sensitivity and specificity. Based on the output scores, the antigen burden score of the predicted peptides can be calculated using peptides with output scores above a confidence threshold.
[0045] In another example, the trained machine learning model was tested and evaluated using experimentally generated peptides from tissue samples using a mass spectrometry-based immunopeptide approach, mixed with decoys at a ratio of 1:999. The positive predictive value of the trained machine learning model in the top 0.1% predicted ligands is significantly higher than that of NetMHCPan 4.0, a publicly available tool considered the gold standard for MHC-binding peptide prediction. In yet another example, the trained machine learning model was evaluated using leave-one-out analysis, which showed high concordance between motifs in the raw data and those predicted by the trained machine learning model.
[0046] 1. Model evaluation for single allele data a) Positive predictive value Figure 12 shows the performance level of the trained binding model and the trained presentation model measured in terms of positive predictive value based on 10% holdout data according to some embodiments. The evaluation data is based on single allele immunopeptidomic data. The positive predictive value (PPV) is defined as the proportion of predicted positives of the trained machine learning model that are actually positive. Thus, PPV reflects the probability that a predicted positive is a true positive. In the evaluation dataset, the positive rate, which indicates the ratio of positives to negatives, is 1:999.
[0047] As shown in Figure 12, the median PPV corresponding to NetMHCpan is about 0.4. In contrast, the trained combined models perform relatively better than NetMHCpan, with the combined models trained on single-allelic data having a median PPV of about 0.6, and the combined models trained on single-allelic and multi-allelic data having comparable median PPVs of about 0.6. The first trained presentation model trained on single-allelic data and the second trained presentation model trained on single-allelic and multi-allelic data perform significantly better than NetMHCpan, with a median PPV of about 0.7. The performance difference of the PPV value of 0.1 may be due to the fact that the evaluation data is derived from single-allelic data.
[0048] Figure 13 shows a comparison of the performance levels of the trained machine learning model compared to conventional techniques for predicting MHC binding peptides. As shown in Figure 13, other conventional techniques show a median PPV of about 0.6, which is comparable to the median PPV of the trained binding model. Compared to the above models, the trained proposed model tends to perform better, with a median PPV of about 0.7.
[0049] FIG. 14 shows a comparison of the performance levels of the trained display model for various alleles compared to conventional techniques for predicting MHC-binding peptides. The PPV values for each allele are shown for NetMHCpan and the trained display model. As shown in FIG. 14, the PPV values corresponding to the trained display model are significantly higher than those of NetMHCpan across all single alleles. Thus, the trained display model shows a significant improvement over NetMHCpan in predicting MHC-binding peptides presented and expressed on the cell surface.
[0050] b) Leave-one-out analysis FIG. 15 shows the results of a leave-one-out analysis of a trained display model according to some embodiments. To show the performance of a trained display model in discovering unknown types of MHC-binding peptides that may be displayed on the cell surface, a leave-one-out analysis can be used to evaluate whether the trained display model can predict surface-displayed peptides corresponding to alleles that are not present in any training data. To perform the leave-one-out analysis, the display model was trained on a training data set that excluded training data corresponding to one specific allele. After training, the trained machine learning model was evaluated by processing 500,000 random peptides to predict surface-displayed peptides that are encoded by alleles from which at least some MHC-binding peptides were excluded. To evaluate the accuracy of peptide prediction by the trained machine learning model, the motifs of the predicted MHC-binding peptides were compared with the motifs obtained from the raw data in which the specific alleles were available.
[0051] As shown in Figure 15, the motifs corresponding to the predicted surface-displayed peptides are substantially identical to the motifs of the peptides corresponding to the excluded alleles, which show comparable amino acid expression levels across the nine positions of the subject peptide. *The second position of the peptide corresponding to 44.03 shows high expression levels of glutamic acid ("E") in the raw data. The predicted MHC binding peptides presented on the cell surface also show high expression levels of glutamic acid at the same second position. Thus, the trained machine learning model is able to accurately predict MHC binding peptides presented on the cell surface even when the corresponding allele is not part of the training data.
[0052] c) Precision and recall FIG. 16 shows a graph depicting precision and recall values for evaluating trained machine learning models, according to some embodiments. Precision-recall can be a useful indicator of prediction success. In information retrieval, precision is an indicator of result relevancy, while recall is an indicator of how many truly relevant results are returned. High precision is related to low false positive rates, and high recall is related to low false negative rates. High scores for both precision and recall can indicate that a given classifier is returning accurate results (high precision) and returning a majority of positive results (high recall). The performance of the trained machine learning models was evaluated based on held-out single allele data, using 10% of the immunopeptidomics data from training mixed with synthetic negative examples in a 1:999 ratio. The X-axis of the graph corresponds to a set of rank percentile thresholds ranging from 0.02 to 1.0, which will identify surface-displayed peptides that are within a particular rank percentile threshold to be considered for either binding or presentation.
[0053] As shown in Figure 16, the trained machine learning model corresponds with a higher precision for all recall values compared to NetMHCpan. The difference is further accentuated for the top 1% peptides in the test data, where the trained machine learning model has a median precision to recall ratio of about 0.8 / 0.6, whereas NetMHCpan has a median precision to recall ratio of about 0.5 / 0.2. Thus, the trained machine learning model can show improved prediction of surface-displayed peptides over NetMHCpan, returning accurate results and returning a large proportion of all positive results.
[0054] 2. Model evaluation with multi-allele (tissue) samples Furthermore, the performance level of the machine learning model trained using multi-allelic samples shows improved prediction of surface-displayed peptides compared to conventional techniques such as NetMHCpan. Figure 17 shows a box plot representing the performance level of the trained machine learning model across various tissue samples according to some embodiments. In Figure 17, three types of tissue samples were processed with the trained machine learning model to generate the fraction of recovery of true candidates corresponding to surface-displayed peptides. Thus, a higher fraction may suggest that the trained machine learning model can show a high performance level in accurately identifying surface-displayed peptides across various tissue samples.
[0055] For example, the ratio value corresponding to NetMHCpan is about 0.65. This ratio value indicates that NetMHCpan was able to predict about 65% of the surface-displayed peptides actually present in the tissue sample. In contrast, the trained binding model performed better than NetMHCpan, with ratio values of about 0.81 for the binding model trained on single-allelic data and about 0.85 for the binding model trained on single- and multi-allelic data. The first trained display model trained on single-allelic data and the second trained display model trained on single- and multi-allelic data performed even better, both corresponding to ratio values of about 0.9. Thus, the trained display model revealed about 90% of the surface-displayed peptides experimentally identified in the tissue sample. Similar improvements in predicting surface-displayed peptides were also shown in other tissue samples. Figure 18 shows a graph comparing the performance levels of the trained machine learning model according to some embodiments with other conventional techniques.
[0056] VI. An Example of a Process for Predicting MHC-Binding Peptides Presented on the Cell Surface FIG. 19 includes a flowchart 1900 illustrating an example of a method for predicting surface-displayed peptides, according to certain embodiments. The operations described in the flowchart 1900 can be performed by a computer system implementing a trained machine learning model, such as a trained binding and presentation model. Although the flowchart 1900 may describe the operations as a sequential process, in various embodiments, many of the operations can be performed in parallel or simultaneously. Furthermore, the operations can be rearranged. The operations may have additional steps not shown in the figure. Furthermore, embodiments of the method can be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments performing the associated tasks can be stored on a computer-readable medium, such as a storage medium.
[0057] In operation 1910, the computer system accesses a machine learning model that has been trained using a training dataset that includes, for each peptide of a plurality of peptides identified by the training dataset, a protein signature of an MHC molecule (e.g., an HLA allele) that binds and presents the peptide on a cell surface, one or more expression levels representing the expression level of a gene encoding the peptide, and one or more peptide presentation metrics representing the amount of the peptide detected as being presented by the MHC molecule. The machine learning model is configured to generate an output indicative of the degree to which the one or more expression levels and the one or more peptide presentation metrics are related according to a population-level relationship between expression and presentation.
[0058] In operation 1920, the computer system accesses genomic and transcriptomic data corresponding to the subject's biological sample. The genomic and transcriptomic data of the biological sample is processed to identify candidate neoantigens (peptides). The genomic and transcriptomic data identifies one or more MHC molecules from the biological sample, and for each peptide of a set of peptides identified from the tissue sample (e.g., candidate neoantigens), includes one or more values representing the peptide. At least one of the one or more values can be determined based on processing of the tissue sample. The one or more values can correspond to a type of peptide, a length of the peptide, an allele that binds the peptide, and expression of a gene region that encodes the peptide.
[0059] In operation 1930, the computer system determines a score for each peptide in the set of peptides using the machine learning model, the one or more MHC molecules identified from the biological sample, and one or more values representing the peptide in the genomic and transcriptomic data. In some cases, the computer system uses a trained machine learning model to process the one or more values to output a score predictive of MHC molecule binding and presentation for a given peptide.
[0060] In operation 1940, the computer system generates results based on the scores. The results may include an incomplete subset of peptides where the subset of peptides exceeds a predetermined threshold predicted to be surface-presented peptides. In some cases, the results may include motifs corresponding to each of the subset of peptides. Additionally or alternatively, the results may include a subset of peptides having a score above a particular ranking percentile (e.g., 0.02). In some cases, the results indicate, for each peptide in the set of peptides, whether the peptide is a surface-presented peptide, i.e., whether it is a peptide that binds to a corresponding MHC molecule and is presented on the cell surface.
[0061] In some cases, the computer system selects an incomplete subset of the set of peptides, where the identification of the incomplete subset is performed in a manner that biases the selection towards peptides associated with regions in space that are associated with outlier peptides in the training dataset whose expression levels and peptide presentation metrics are related in a manner that deviates from population-level relationships.
[0062] In operation 1950, the computer system outputs the results, after which process 1900 ends.
[0063] VII. Computer Environment 20 illustrates an example of a computer system 2000 for implementing some of the embodiments disclosed herein. The computer system 2000 may have a distributed architecture, where some components (e.g., memory and processor) are part of an end-user device and some other similar components (e.g., memory and processor) are part of a computer server. The computer system 2000 includes at least a processor 2002, a memory 2004, a storage device 2006, an input / output (I / O) peripherals 2008, a communication peripherals 2010, and an interface bus 2012. The interface bus 2012 is configured to communicate, transmit, and transfer data, control, and commands between various components of the computer system 2000. The processor 2002 may include one or more processing units, such as a CPU, a GPU, a TPU, a systolic array, or a SIMD processor. Memory 2004 and storage 2006 include computer readable storage media, such as RAM, ROM, electrically erasable programmable read only memory (EEPROM), hard drives, CD-ROM, optical storage, magnetic storage, electronic non-volatile computer storage, such as Flash memory, and other tangible storage media. Any such computer readable storage media can be configured to store instructions or program code embodying aspects of the present disclosure. Memory 2004 and storage 2006 also include computer readable signal media. Computer readable signal media include propagated data signals having computer readable program code embodied therein. Such propagated signals may take any of a variety of forms, including, but not limited to, electromagnetic, optical, or any combination thereof. Computer readable signal media includes any computer readable medium that can communicate, propagate, or transmit a program for use in connection with computer system 2000, other than computer readable storage media.
[0064] Additionally, memory 2004 includes an operating system, programs, and applications. Processor 2002 is configured to execute stored instructions and includes, for example, logic processing units, microprocessors, digital signal processors, and other processors. Memory 2004 and / or processor 2002 may be virtualized and hosted within another computing system, for example, in a cloud network or data center. I / O peripherals 2008 include computing components such as user interfaces and graphical processing units, such as keyboards, screens (e.g., touch screens), microphones, speakers, other input / output devices, serial ports, parallel ports, universal serial buses, and other input / output peripherals. I / O peripherals 2008 are connected to processor 2002 through any of the ports connected to interface bus 2012. Communication peripherals 2010 are configured to facilitate communication between computer system 2000 and other computing devices over a communication network and include, for example, network interface controllers, modems, wireless and wired interface cards, antennas, and other communication peripherals.
[0065] Although the subject matter of the present invention has been described in detail with respect to specific embodiments thereof, it will be understood that those skilled in the art, upon gaining the above understanding, can easily make modifications, variations, and equivalents to such embodiments. Thus, it should be understood that the present disclosure is presented for purposes of illustration and not limitation, and does not preclude the inclusion of modifications, variations, and / or additions to the subject matter as would be readily apparent to one skilled in the art. Indeed, the methods and systems described herein may be embodied in other various forms, and further, various omissions, substitutions, and changes in the forms of the methods and systems described herein may be made without departing from the spirit of the present disclosure. The appended claims and their equivalents are intended to cover such forms or modifications as fall within the scope and spirit of the present disclosure.
[0066] Unless specifically stated otherwise, throughout this specification, discussions using terms such as "processing," "computing," "calculating," "determining," and "identifying" or the like are understood to refer to operations or processes of a computing device, such as one or more computers or similar electronic computing device or devices, that manipulate or transform data represented as physical electronic or magnetic quantities in the memory, registers, or other information storage, transmission, or display devices of a computing platform.
[0067] The system or systems discussed herein are not limited to a particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide a result conditioned on one or more inputs. Suitable computing devices include general-purpose microprocessor-based computing systems that access stored software that programs or configures the computing system, ranging from general-purpose computing devices to specialized computing devices that implement one or more embodiments of the inventive subject matter. Any suitable programming, scripting, or other type of language or combination of languages may be used to implement the teachings contained herein in the software used to program or configure the computing device.
[0068] The method embodiments disclosed herein may be performed in operation of such a computing device. The order of the blocks presented in the above examples may be changed, e.g., the blocks may be rearranged, combined, and / or divided into sub-blocks. Certain blocks or processes may be performed in parallel.
[0069] Conditional language as used herein, particularly "can," "could," "might," "may," "eg," and the like, is generally intended to convey that certain examples include certain features, elements, and / or steps, while other examples do not, unless specifically stated otherwise or understood otherwise within the context in which it is used. Thus, such conditional language is generally not intended to imply that features, elements, and / or steps are in any way required by one or more examples, or that one or more examples necessarily include logic for determining whether those features, elements, and / or steps are included or performed in a particular example, with or without author input or prompting.
[0070] The terms "comprising," "including," "having," and the like are synonymous and are used in an open-ended, inclusive manner and do not exclude additional elements, features, acts, operations, etc. Additionally, the term "or" is used in an inclusive (and not exclusive) sense, so that, for example, when used connecting a list of elements, the term "or" means one, some, or all of the elements in the list. Use of "adapted to" or "configured to" herein means open, inclusive language that does not exclude devices adapted or configured to perform additional tasks or steps. Additionally, use of "based on" means that a process, step, calculation, or other action that is "based on" one or more recited conditions or values may, in fact, be based on additional conditions or values beyond those recited. Similarly, the use of "based at least in part on" is meant to be open and inclusive, in that a process, step, calculation, or other act that is "based at least in part on" one or more recited conditions or values may in fact be based on additional conditions or values beyond those recited. The headings, lists, and numbering contained herein are for ease of description and are not meant to be limiting.
[0071] The various functions and processes described above may be used independently of one another or may be combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of the present disclosure. In addition, in some implementations, certain method or process blocks may be omitted. The methods and processes described herein are also not limited to a particular order, and the blocks or states associated therewith may be performed in other orders as appropriate. For example, the blocks or states described may be performed in an order other than that specifically disclosed, or multiple blocks or states may be combined into one block or state. The examples of blocks or states may be performed sequentially, in parallel, or in other ways. Blocks or states may be added to or removed from the disclosed examples. Similarly, the examples of systems and components described herein may be configured differently than described. For example, elements may be added, removed, or rearranged as compared to the disclosed examples.
Claims
1. 1. A method of using a processor of a computer system to execute a program stored in its memory, comprising: (a) accessing a machine learning model stored in a storage device, The machine learning model is (i) trained for each peptide of a plurality of peptides identified by said training dataset using as attributes a training dataset that includes: (1) the protein characteristics of the MHC molecule that bind and present the peptide; (2) one or more expression levels representing the expression level of the gene encoding the peptide; and (3) one or more peptide presentation metrics representing the amount of peptide detected as presented by the MHC molecule; and (ii) accessing the machine learning model, the machine learning model being configured to generate a score indicating the degree to which the one or more expression levels of a tissue sample and the one or more peptide presentation metrics of the tissue sample are related according to a relationship between the expression levels of the training dataset and the peptide presentation metrics of the training dataset, wherein the relationship between the expression levels and the peptide presentation metrics of the training dataset varies depending on the population level of the training dataset; (b) accessing genomic and transcriptomic data stored in the storage device corresponding to a tissue sample of the subject, the genomic and transcriptomic data identifying one or more MHC molecules from the tissue sample and, for each peptide of a set of peptides identified from the tissue sample, including one or more values representing the peptide, at least one of the one or more values being determined based on processing of the tissue sample; (c) for each peptide in the set of peptides, determining a score using the machine learning model, the one or more MHC molecules identified from the tissue sample, and the one or more values representing the peptide; (d) generating a result based on the score; and (e) outputting the results using input / output (I / O) peripherals; Including, (i) the computer system includes the processor, the memory, the storage device, the input / output (I / O) peripherals, communication peripherals, and an interface bus; (ii) the training dataset is derived from monoallelic data corresponding to peptides derived from monoallelic cell lines and / or biallelic data corresponding to peptides derived from other tissue samples; (iii) attributes of the training data vector are data corresponding to somatic variants, including one or more features including peptide sequences; (iv) the labels of the data vectors of the training data are peptides encoded by somatic variants that bind to MHC molecules and are present on the cell surface; (v) the training dataset for training the machine learning model includes sequence data from various sources: (1) peptides identified as binding to HLA molecules based on in vitro experiments, (2) peptides identified by mass spectrometry analysis of tumor samples, (3) HLA alleles, and (4) non-tumor samples. method.
2. selecting an incomplete subset of the set of peptides based on the score; 2. The method of claim 1, wherein the identification of the incomplete subset is performed in a manner that biases the selection toward peptides associated with scores that predict a more likely presentation compared to the probability predicted by the corresponding population-level relationship in the training dataset, and the result comprises the incomplete subset of the set of peptides.
3. selecting an incomplete subset of the set of peptides based on the score; 2. The method of claim 1, wherein the identification of the incomplete subset is performed in a manner that biases the selection toward peptides associated with regions in space, the regions being associated with outlier peptides in the training dataset whose expression levels and peptide presentation metrics are related in a manner that deviates from the population-level relationships.
4. The method of claim 1 , wherein the results include, for each of one or more peptides in the set of peptides, the peptide identification and the score.
5. 2. The method of claim 1, wherein for each peptide in the set of peptides, the one or more values representing the peptide are generated based on the amino acid sequence of the peptide, an indication of whether the peptide binds to one or more binding pockets of the MHC molecule, the expression level of the peptide in the tissue sample, and / or the length of the peptide.
6. The method of claim 1 , wherein the score corresponding to a peptide of the set of peptides corresponds to a predicted probability as to whether the peptide will bind to the MHC molecule and be presented on a cell surface.
7. The method of claim 1 , wherein the machine learning models include one or more trained gradient boosting algorithms.
8. 2. The method of claim 1, wherein the machine learning model includes, for each peptide of the plurality of peptides, a first sub-model trained on a first subset of the training dataset that includes a sequence corresponding to the peptide, a sequence of an MHC molecule that binds to the peptide, and / or a length of the peptide.
9. 9. The method of claim 8, wherein the machine learning model comprises, for each peptide of the plurality of peptides, a second sub-model trained on a second subset of the training dataset comprising one or more expression levels of a source protein from which the peptide is derived and surface presentation properties of the peptide.
10. The method of claim 9 , wherein the first sub-model and the second sub-model are each trained based on one or more scores generated by another set of sub-models. (a) one or more data processors; (b) a non-transitory computer-readable storage medium containing instructions that, when executed by said one or more data processors, cause said one or more data processors to perform the method of claim 1; Including, the system.
12. A non-transitory machine-readable storage medium containing instructions for causing one or more data processors to perform the method of claim 1.