Machine learning techniques for predicting surface-displayed peptides
A machine learning model using genomic and transcriptomic data predicts MHC-binding peptides that are presented on the cell surface, improving immune response prediction and enabling personalized cancer therapies.
Patent Information
- Application Number
- JP2024143327
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-07
- Filing Date
- 2024-08-23
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2041-06-17
AI Technical Summary
Conventional techniques for predicting MHC-binding peptides fail to accurately determine which peptides are presented on the cell surface, leading to incomplete understanding of immune system responses to tumor cells.
A machine learning model trained using genomic and transcriptomic data, including MHC molecule binding and expression levels, to predict peptides that bind and are presented on the cell surface, utilizing genetically engineered cell lines and biallelic data to enhance accuracy.
The model achieves significantly higher positive predictive value compared to conventional methods, enabling precise identification of MHC-bound peptides on the cell surface, facilitating personalized immunotherapies and biomarker development.
Smart Images

Figure 00000019_0000 
Figure 00000019_0001 
Figure 00000020_0000
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 040,943, filed June 18, 2020, entitled "Composite Biomarkers for Immunotherapy for Cancer," and U.S. Provisional Patent Application No. 63 / 111,007, filed November 7, 2020, entitled "Machine-Learning Techniques For Predicting Surface-Presenting Peptides," the entire contents of which are incorporated herein by reference in their entireties for all purposes.
[0002] The present disclosure relates to machine learning techniques for predicting surface-displayed peptides. [Background technology]
[0003] Cancers contain mutations, which can be somatic or tumor-specific. The immune system detects these cancer-based mutations by identifying peptides derived from these mutations. Peptides are recognized by the immune system when they bind to proteins encoded by major histocompatibility complex (MHC) genes and are presented on the surface of cells. For example, peptides corresponding to mutated genes can bind to specific MHC molecules (e.g., human leukocyte antigen (HLA) proteins) and be presented on the cell surface. Predicting peptides expressed on tumor cell surfaces can inform the development of precision cancer treatments and diagnostics. For example, identifying genomic variants corresponding to these peptides can analyze the response and resistance of this complex system to specific cancer immunotherapies. As another example, analyzing peptides presented on tumor cell surfaces can be used to develop personalized immuno-oncology (IO) therapies and / or neoantigen cancer vaccines.
[0004] Such peptides expressed on the surface of tumor cells are also known as "neoantigens," and their prediction requires thorough analysis of many technical factors, including, but not limited to, the quality of peptide sequencing data, the availability of tumor and normal sample pairs, HLA typing, and the identification of other peptide characteristics. For example, neoantigens can be identified based on the prediction of peptides that bind to MHC molecules and are presented on the cell surface. To identify neoantigens, determining the peptides encoded by somatic variants and identifying the HLA molecules that bind to the peptides is only the first step in a highly complex process. This is because each peptide identified from the sequence data must be processed by the proteasome, transported for MHC binding, presented on the tumor cell surface, and ultimately may or may not be recognized by the immune system. Due to this complex process, many peptides that bind to HLA molecules (for example) may not be expressed on the cell surface.
[0005] Furthermore, one or more binding motifs of an MHC molecule can be identified to determine whether a given peptide binds to the MHC molecule. While the binding motifs of some MHC molecules (e.g., HLA-A molecules) are known, there are many MHC molecules for which the binding motifs have not yet been identified. For example, the binding motifs of MHC class II molecules are relatively unknown due to limited experimental data available. Without this information, it would be difficult to determine whether a peptide binds to a corresponding MHC molecule. Previous techniques have attempted to address this problem by training machine learning models using known MHC-binding motifs to predict whether a peptide will bind to one of various types of MHC molecules. However, even if such peptides are identified, some peptides may not be present on the cell surface. In other words, while previous techniques can identify MHC-binding peptides, only a small proportion of them can be successfully presented on the cell surface. Because immune system responses are triggered when MHC-binding peptides are presented on the cell surface, simply identifying MHC-binding peptides does not fully reveal how the immune system responds to tumor cells, foreign proteins, etc.
[0006] Thus, conventional techniques for predicting MHC-binding peptides do not address whether a peptide is actually presented and expressed on the cell surface. Furthermore, conventional techniques are insufficient to identify peptide characteristics that indicate that a given peptide is presented on the cell surface. Therefore, there is a need to accurately predict peptides that bind to their corresponding MHC molecules and are presented on the cell surface. Summary of the Invention [Problem to be solved by the invention]
[0007] In some embodiments, a method for predicting surface-presented peptides is provided. The method can include accessing a trained machine learning model that has been trained using a training dataset and that includes, for each peptide of a plurality of peptides identified by the training dataset, a protein signature of a major histocompatibility complex (MHC) molecule that binds and presents the peptide, one or more expression levels representing the expression level of a gene encoding the peptide, and one or more peptide presentation metrics representing the amount of the peptide detected as being presented by the MHC molecule. The machine learning model can be configured to generate an output that indicates the degree to which the one or more expression levels and the one or more peptide presentation metrics are related according to a population-level relationship between expression and presentation. [Means for solving the problem]
[0008] The method may also include accessing genomic and transcriptomic data corresponding to the subject's biological sample. The genomic and transcriptomic data may identify one or more MHC molecules from the biological sample and, for each peptide in a set of peptides identified from the cell line or tissue sample, one or more values representing the peptide. The one or more values may be determined based on processing of the tissue sample. The method may also include determining a score for each peptide in the set of peptides using a machine learning model, the one or more MHC molecules identified from the biological sample, and the one or more values representing the peptide. The method may include generating a result based on the score and outputting the result.
[0009] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed by the one or more data processors, cause the one or more data processors to perform some or all of one or more of the methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium, the product including instructions configured to cause one or more data processors to perform some or all of one or more of the methods and / or some or all of one or more processes disclosed herein.
[0010] The terms and expressions which have been employed are used as terms of description rather than of limitation, and there is no intention in the use of such terms and expressions to exclude any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, while the invention as claimed has been specifically disclosed by embodiments and optional features, it is to be understood that modifications and variations of the concepts disclosed herein may be available to those skilled in the art, and that such modifications and variations are considered to be within the scope of the invention as defined by the appended claims. [Brief explanation of the drawings]
[0011] The present disclosure will be described in conjunction with the accompanying drawings. [Figure 1] FIG. 1 shows a schematic diagram of peptides bound to MHC molecules and presented on the cell surface. [Figure 2] FIG. 2 shows a schematic diagram depicting peptides that can be presented on the cell surface in response to gene therapy. [Figure 3] FIG. 3 shows a schematic diagram for identifying single-allele immunopeptidomic data that can be used to train a machine learning model, according to some embodiments. [Figure 4]FIG. 4 shows allelic diversity data corresponding to MHC-binding peptides, according to some embodiments. [Figure 5] FIG. 5 shows source diversity data identified from tissue and cell line samples for training machine learning models to predict surface-displayed peptides, according to some embodiments. [Figure 6] FIG. 6 shows a plot of comparative data between the expected number of peptides based on gene expression levels and the actual number of peptides observed, according to some embodiments. [Figure 7] FIG. 7 shows a process for determining genetic propensity scores used to train machine learning models, according to some embodiments. [Figure 8] FIG. 8 shows a plot of comparative data between the number of peptides expected for one or more regions within a gene and the number of peptides actually observed for that region, according to some embodiments. [Figure 9] FIG. 9 illustrates a process for determining hotspot scores used to train machine learning models, according to some embodiments. [Figure 10] FIG. 10 shows examples of features used in the combined model and the presented model, according to some embodiments. [Figure 11] FIG. 11 shows an exemplary model architecture for training a machine learning model for predicting surface-displayed peptides, according to some embodiments. [Figure 12] FIG. 12 shows the performance levels of the trained combined model and the trained presentation model, measured by positive predictive value, based on 10% holdout data, according to some embodiments. [Figure 13] FIG. 13 shows a comparison of the performance levels of the trained machine learning models compared to conventional techniques for predicting surface-displayed peptides. [Figure 14] FIG. 14 shows a comparison of the performance levels of the trained display models across different alleles compared to conventional techniques for predicting surface-displayed peptides. [Figure 15]FIG. 15 shows the results of a leave-one-out analysis of the trained presentation model, according to some embodiments. [Figure 16] FIG. 16 shows a graph illustrating precision and recall for evaluating a trained machine learning model, according to some embodiments. [Figure 17] FIG. 17 shows a box plot depicting the performance level of a trained machine learning model across various tissue samples, according to some embodiments. [Figure 18] FIG. 18 shows a graph comparing the performance levels of a trained machine learning model with other prior art, according to some embodiments. [Figure 19] FIG. 19 includes a flowchart illustrating an example of a method for predicting surface-displayed peptides, according to certain embodiments. [Figure 20] FIG. 20 illustrates an example computer system for implementing some of the embodiments disclosed herein. DETAILED DESCRIPTION OF THE INVENTION
[0012] I. Overview To address at least the above-mentioned shortcomings of conventional systems, the present technology can be used to predict surface-presented peptides. As used herein, "surface-presented peptide" can refer to a peptide that binds to an MHC molecule (e.g., an HLA-A protein) and is presented on the surface of a corresponding cell. One or more somatic variants can be identified by sequencing DNA from a normal sample and a tumor sample. The somatic variants include one or more genetic mutations present in the tumor sample and the normal sample. The somatic variants of the tumor sample can be processed using a trained machine learning model to predict whether a peptide encoded by the somatic variant will bind to an MHC molecule (e.g., MHC class 1) and be presented on the cell surface. The machine learning model can include a binding model that predicts whether a peptide encoded by the somatic variant will bind to an MHC molecule. In some embodiments, the machine learning model includes a presentation model that predicts whether a peptide encoded by the somatic variant will be expressed on the cell surface.
[0013] The machine learning model can be trained using training datasets obtained from (i) genetically engineered monoallelic cell lines and (ii) biallelic data from tissue samples from other subjects. In some cases, the machine learning model has been trained using binding array data (e.g., IEDB data). The training dataset can include, for each peptide identified by the training dataset, one or more expression levels representing the expression level of the gene encoding the peptide and one or more peptide presentation metrics representing the amount of the peptide detected as presented by an MHC molecule. The training dataset can include immunopeptidomic data of peptides generated from multiple genetically engineered cell lines (e.g., K562 cells) expressing a single allele of interest (e.g., HLA-A). In particular, MHC-peptide complexes in these cell lines can be immunoprecipitated using the W6 / 32 antibody, followed by peptide elution and peptide sequencing using tandem mass spectrometry. Training datasets corresponding to biallelic data from other tissue samples can be obtained using curated public data.
[0014] The prediction of surface-displayed peptides can be performed in a manner that biases selection toward peptides associated with scores that predict more certain presentation than would be expected by the population-level relationship between peptide expression and presentation. Additionally or alternatively, the prediction of surface-displayed peptides can be performed in a manner that biases selection toward peptides associated with regions in space that are associated with outlier peptides in the training dataset where expression levels and peptide presentation metrics are related in a way that deviates from the population-level relationship.
[0015] Thus, embodiments of the present disclosure provide technical advantages over conventional systems by accurately predicting peptides that bind to their corresponding MHC molecules and are presented on the cell surface. As previously discussed, tumor cell surface peptide binding and expression can predict how the immune system will respond to neoantigens and / or certain cancer immunotherapies. Accurate prediction of surface-presented peptides therefore facilitates the selection or development of immunotherapies that will be most effective for a given subject. Furthermore, based on model evaluation, embodiments demonstrate significantly higher positive predictive value compared to conventional techniques, such as NetMHCPan 4.0. Thus, the high sensitivity and specificity of embodiments enable accurate identification of MHC-bound peptides presented on the cell surface, thereby facilitating applications in the development of personalized immunotherapies and biomarkers.
[0016] The following examples are provided to introduce certain particular embodiments. In the following description, for purposes of explanation, specific details are set forth to provide a thorough understanding of the disclosed examples. However, it will be apparent that various examples can be practiced without these specific details. For example, devices, systems, structures, assemblies, methods, and other components may be shown as components in block diagram form so as not to obscure the examples in unnecessary detail. In other instances, well-known devices, processes, systems, structures, and techniques may be shown without necessary detail so as to avoid obscuring the examples. The figures and descriptions are not intended to be limiting. The terms and expressions used in this disclosure are used in terms of description rather than limitation, and the use of such terms and expressions is not intended to exclude any equivalents of the features shown and described, or portions thereof. The word "example" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as an "example" is not necessarily to be construed as preferred or advantageous over other embodiments or designs.
[0017] II. Surface display of peptides 1. Neoantigens in tumor samples Neoantigens can be found in tumor samples. A neoantigen refers to one or more peptides presented on the surface of tumor cells, thereby eliciting an immune system response. The immune system can be tuned to seek out pathogens, including cancer, and thus has the potential to cure cancer. The immune system can distinguish between self and non-self antigens. Because tumors are caused by genetic mutations (e.g., somatic variants), peptides expressed on the cell surface that correspond to these genetic mutations can be considered neoantigens. These peptides are considered "new" to the immune system, so ideally, the immune system can recognize and eliminate tumor cells based on detecting neoantigens presented on the tumor cell surface. As explained above, tumor samples can be analyzed to reveal sequence data, which can be compared with those of normal samples to identify somatic variants. Somatic variants can be further analyzed to determine which subsets of variants will appear as peptides. Neoantigens can be predicted by identifying peptides that bind to MHC molecules and are presented on the cell surface. Therefore, the ability of peptides to be presented on the cell surface can be a key factor in developing immunotherapies for cancer.
[0018] 2. Peptides that are effective in treating certain autoimmune diseases Surface-presented peptides can be identified in the context of autoimmune diseases, and these peptides are encoded based on genetic alterations resulting from specific immunotherapies. Figure 2 shows a schematic diagram illustrating surface-presented peptides that respond to gene therapy. Figure 2 shows a mutation in the dystrophin gene, which typically causes debilitating muscular dystrophies. The dystrophin gene encodes the dystrophin protein molecule, which acts as a shock absorber for muscle cells. The lack of fully functional dystrophin protein can lead to muscle degeneration. Muscular dystrophies can typically be treated with exome skipping therapy, which skips exomes (e.g., exon 52) responsible for the dystrophin gene mutation and generates semi-functional dystrophin protein for the subject. While exome skipping therapy can be effective, it can also induce the production of novel peptides due to the intentional skipping of exomes through genetic modification. These novel peptides can bind to MHC molecules and be presented on the cell surface, potentially triggering a destructive immune system response.
[0019] III. Example of a training dataset Machine learning models for predicting surface-presented peptides can be trained using supervised training algorithms. Machine learning models can be trained using training datasets. Training datasets for training machine learning models can include sequence data from a variety of sources: (i) peptides identified as binding to HLA molecules based on in vitro experiments, (ii) peptides identified by mass spectrometry analysis of tumor samples, (iii) HLA alleles, and (iv) non-tumor samples. However, some training sequence data may be inaccurate for training machine learning models. For example, training sequence data generated from tissue samples requires the difficult process of mapping peptides to one of several types of HLA proteins (e.g., HLA-A, HLA-B) co-expressed on the cell surface. In another example, sequence data generated using in vitro methods may not mimic surface presentation. An embodiment of the present disclosure for systematically resolving inconsistencies in training datasets is used to train machine learning models that predict peptides likely to be "shuttled" to the cell surface from somatic variants called from the sequence data.
[0020] Additionally or alternatively, the training dataset can further include data corresponding to somatic variants, each labeled to indicate whether the peptide encoded by the somatic variant binds to an MHC molecule (e.g., an HLA-A protein) and is presented on the cell surface. The training dataset can also include one or more features derived from the somatic variants (e.g., peptide sequence, peptide length, expression of the peptide in tumor samples).
[0021] To prepare a training dataset, tumor samples and matched normal control samples can be sequenced to generate tumor-normal paired sequence data. The tumor-normal paired sequence data are compared to identify somatic variants, including single nucleotide variants (SNVs), indels, and / or altered genes containing copy number variations. In some cases, machine learning models are used to process the tumor-normal paired sequence data and identify somatic variants in the tumor sample.
[0022] 1. Training Data Source a) Single-allele immunopeptidomics data In some cases, at least a portion of the training data corresponds to peptides identified from a genetically engineered, monoallelic cell line. Figure 3 shows a schematic diagram for identifying monoallelic immunopeptidomic data that can be used to train a machine learning model, according to some embodiments. As shown in Figure 3, a genetically engineered, monoallelic K562 cell line can be generated and then transfected with a specific HLA molecule of interest (e.g., HLA-B) (step 305). As previously described, the HLA complex is a group of related proteins encoded by the MHC gene complex in humans. These cell surface proteins are responsible for regulating the immune system. From the cell line, HLA-binding peptides can be identified by immunoprecipitating the HLA-peptide complex using the W6 / 32 antibody (step 310), applying peptide elution (step 315), and performing peptide sequencing on the eluted peptides using mass spectrometry (e.g., liquid chromatography-mass spectrometry, mass spectrometry) (step 320). Thus, HLA-binding peptides can be identified for the specific HLA molecule of interest (step 325).
[0023] Single-allele immunopeptidomics data identifying various characteristics of HLA-binding peptides can be identified and included as part of a training dataset. Examples of training data from single-allele immunopeptidomics data can include, for a given HLA-binding peptide, the type of peptide, the length of the peptide, the amino acid sequence of the peptide, the HLA alleles that bind to the peptide, the number of transcripts corresponding to the peptide, and the expression of the gene region encoding the peptide. To optimize the performance of the machine learning model, training data was generated to represent the HLA genotypes of the general population. For example, FIG. 4 shows allele diversity data corresponding to HLA-binding peptides according to some embodiments. To determine allele diversity data, the identified peptides can be clustered based on their similarity to peptide sequences corresponding to all known alleles of the HLA molecule of interest (e.g., HLA alleles identified from the IMGT database). Thus, the identified peptides can be clustered based on the similarity of their respective binding pockets. In some cases, the identified peptides can be clustered using a BLOSUM similarity matrix. Based on these clusters, one or more alleles encoding the HLA-binding peptide can also be identified. In some cases, peptide clusters are visualized on a heat map. For example, Figure 4 shows a first heatmap identifying allelic diversity of HLA-A molecules and a second heatmap identifying allelic diversity of HLA-B molecules. Additionally or alternatively, the training dataset can be augmented using training data corresponding to allele frequency data for alleles encoding HLA binding proteins, where the allele frequency data is classified across different portions of the world's population.
[0024] Training data corresponding to HLA-binding peptides can facilitate training of machine learning models by using single-allelic immunopeptidomic data expressing one specific type of HLA at a time. Furthermore, the allelic diversity in the single-allelic immunopeptidomic training data allows machine learning models to predict surface-presented peptides derived from various alleles that may not be present in the training data.
[0025] b) Multi-allele immunopeptidomic data In some cases, at least a portion of the training data corresponds to peptides identified from sequencing tissue samples from other subjects. Various tissue samples or cell lines from a subject's tissue sample can be sequenced to identify multiple peptides that bind to different types of HLA molecules (e.g., HLA-A, HLA-B, HLA-C). In some cases, the cell lines and tissue samples are processed using mass spectrometry. Multi-allelic immunopeptidomic data obtained from the multiple identified peptides can be used as part of the training data. The multi-allelic immunopeptidomic data can include various features corresponding to the identified peptides, including peptide length and allelic diversity. Figure 5 shows source diversity data identified from a subject's tissue sample for training a machine learning model for predicting surface-displayed peptides, according to some embodiments. Figure 5 shows the abundance of each type of peptide for each of the single and multi-allelic samples. Additionally or alternatively, multi-allelic data can be obtained from public data sources.
[0026] Biallelic immunopeptidomic data generated from diverse tissues and cell lines can be integrated into a training dataset to improve the performance of the trained machine learning model. In particular, training a machine learning model using biallelic immunopeptidomic data can reduce overfitting and / or underfitting. For example, both monoallelic and biallelic immunopeptide peptides from several publicly available data sources can be added to the training dataset. Monoallelic immunopeptidomic data from genetically engineered cell lines and monoallelic and biallelic immunopeptidomic data from tissue samples can all be integrated into the training dataset to expand its size (e.g., a larger number of unique peptides).
[0027] 2. Additional enhancing features As explained above, immunopeptidomic data from a training dataset identifies various features of HLA-bound peptides, including peptide sequence, peptide length, binding pocket sequence, left flanking region, and right flanking region. In some cases, the training dataset also includes antigen presentation features, such as peptide expression levels measured by DPM. In addition to the above, two additional features can be generated from the immunopeptidomic data and used to enhance the training dataset.
[0028] a) Comparison data between the predicted number of peptides based on gene expression levels and the number of peptides actually observed The first features generated from the immunopeptidomics data can include comparison data between the expected number of peptides based on gene expression levels and the number of peptides actually observed. By including a training dataset with the first features, a trained machine learning model trained from the training data can improve prediction of surface-displayed peptides, such that predictions are biased toward peptides associated with scores that predict presentation with greater certainty compared to the probability expected by the population-level relationship between peptide expression and presentation. Furthermore, a trained machine learning model trained from the training data can facilitate prediction of surface-displayed peptides, such that predictions are performed in a manner that biases selection toward peptides associated with regions in space associated with outlier peptides in the training dataset where expression levels and peptide presentation metrics are related in a way that deviates from the population-level relationship.
[0029] Figure 6 shows a plot of comparative data between the expected number of peptides based on gene expression levels and the number of actually observed peptides, according to some embodiments. To generate the comparative data, all transcripts corresponding to HLA-binding peptides from a training dataset can be identified and organized into a set of bins based on their respective gene expression levels. For example, as shown in Figure 6, the x-axis shows 10 sections (e.g., deciles) into which transcripts can be grouped based on their respective gene expression levels. The bars in each section indicate the measured amount of gene expression levels corresponding to the transcripts grouped in the section. The y-axis of the plot shown in Figure 6 indicates the number of peptides, and the diamonds indicate the amount of peptides counted from the training sample cell lines.
[0030] The initial hypothesis in Figure 6 without comparative data appears to indicate that the expected peptide abundance is directly proportional to the measured gene expression level. However, using the comparative data in Figure 6, one or more outliers that deviate from the initial hypothesis can be identified. The first outlier 605 contains a large amount of peptides observed in bin "1," indicating a very low amount of gene expression. The second outlier 610 contains almost no amount of peptides observed in bin "10," indicating a high amount of gene expression. Therefore, the comparative data in Figure 6 indicates that measuring the gene expression level of an HLA-binding peptide may not be sufficient to predict whether the HLA-binding peptide will actually be presented on the cell surface. Using the comparative data, gene propensity scores ("GPS") for genes encoding HLA-binding peptides can be calculated, and the gene propensity scores predict whether the peptide will be presented on the cell surface. In some cases, the calculated gene propensity scores can be added as additional features to the training dataset, allowing further training of a machine learning model based on the gene propensity scores.
[0031] 7 illustrates a process for determining gene propensity scores used to train a machine learning model, according to some embodiments. At block 705, immunopeptidomics data is obtained, the immunopeptidomics data including expression levels of genes encoding HLA-binding peptides. For example, the immunopeptidomics data can be obtained by reprocessing existing mass spectrometry (MS) data or by accessing the immunopeptidomics data directly from a database (e.g., an immunopeptide database).
[0032] At block 710, the expected peptide counts for the genes identified in the immunopeptidomic data are calculated. In particular, the expected peptide counts are calculated based on the number of transcripts (e.g., TPMs) and the sequence length of the gene. At block 715, the ratio between the expected peptide counts and the observed peptide counts can be calculated to generate a gene propensity score (e.g., log10(observed / expected)). In some cases, the gene propensity score is added as an additional feature to the training dataset.
[0033] b) Comparison of the predicted number of peptides per gene region with the number of peptides actually observed The second features generated from the immunopeptidomic data can include a comparison of the expected number of peptides based on expression levels within one or more regions of a given gene with the actual observed number of peptides corresponding to the one or more regions. In contrast to the first features, which identify gene expression levels across various genes, the second features identify the expression levels of regions within a single gene. Based on the identified expression levels, predicted amounts of peptides can be generated. The predicted amounts can be compared to the observed amounts of peptides to identify training dataset second features, where the second features are indicative of one or more surface display properties of the regions within the corresponding genes.
[0034] In some cases, the first feature and the second feature are combined in a training dataset. A trained machine learning model trained from training data with the combined features can facilitate prediction of surface-displayed peptide predictions, such that predictions are biased toward peptides associated with scores that predict presentation with greater certainty compared to the probability expected by the population-level relationship between peptide expression and presentation. Furthermore, a trained machine learning model trained from the training data can facilitate prediction of surface-displayed peptide predictions, such that predictions are performed in a manner that biases selection toward peptides associated with regions in space associated with outlier peptides in the training dataset where expression levels and peptide presentation metrics are related in a way that deviates from the population-level relationship.
[0035] FIG. 8 shows a plot of comparative data between the expected peptide counts for one or more regions within a gene and the actual observed peptide counts for that region, according to some embodiments. For each genomic region of a given gene, gene expression levels can be calculated and the expected peptide abundance can be measured. For example, the expected peptide abundance for the genomic region of the ACTB gene is shown by the black plot line 805. The expected peptide abundance can then be compared to the observed peptide abundance for each genomic region, with the observed peptide abundance shown by the gray area 810. As in FIG. 6, some outliers can be identified within various genomic regions, where the observed peptide abundance is not proportional to the measured gene expression level. For example, region 815 (e.g., region number 230) of the ACTC1 gene may exhibit a very high expected peptide abundance (e.g., >3,000 peptides), yet the observed peptide abundance for the same region 815 is actually much lower than the expected amount (e.g., approximately 1,000 peptides). Therefore, the comparative data in Figure 8 indicates that measuring the gene expression level at the region level of an HLA-binding peptide may not be sufficient to predict whether the HLA-binding peptide will actually be presented on the cell surface. Using the comparative data shown in Figure 8, a hot spot score ("hhs") can be calculated for a gene encoding an HLA-binding peptide, and the hot spot score predicts whether a peptide corresponding to a region of the gene will be presented on the cell surface.
[0036] 9 illustrates a process for determining hotspot scores used to train a machine learning model, according to some embodiments. At block 905, immunopeptidomics data is obtained, the immunopeptidomics data including expression levels of genes encoding HLA-binding peptides. For example, the immunopeptidomics data can be obtained by reprocessing existing mass spectrometry (MS) data or by accessing the immunopeptidomics data directly from a database (e.g., an immunopeptide database).
[0037] In step 910, the predicted peptide counts for each region of a particular gene are compared to the actual peptide counts for that region. In step 915, a hotspot score is calculated for the particular gene, which specifies the distribution of observed peptide counts across regions of the gene (e.g., ACTB gene, ACTC1 gene).
[0038] IV. Example of a model architecture for predicting MHC-binding peptides presented on the cell surface The training dataset can be used to train a machine learning model for predicting surface-displayed peptides. The machine learning model includes one or more submodels configured to identify binding and surface-display characteristics of peptides in a sample. These submodels can be trained separately on corresponding subsets of the training dataset, such that each submodel is capable of predicting surface-displayed peptides based on parameters learned from features corresponding to the subset.
[0039] 1. Binding model and presentation model In some cases, the machine learning models include a binding model and a presentation model, each trained to process different features of the input data. FIG. 10 shows example features used by the binding model 1005 and the presentation model 1010, according to some embodiments. The binding model 1005 can be trained using a training dataset containing information related to a set of peptides (e.g., the sequence of the MHC molecule that binds the peptide, the length of the peptide). In some cases, the binding model 1005 includes one or more trained gradient boosting algorithms. Gradient boosting refers to a machine learning technique for regression and classification problems that creates a predictive model in the form of an ensemble of weak predictive models. The technique can generalize the model by building models incrementally and enabling the optimization of any differentiable loss function. Gradient boosting combines weak learners into a single strong learner in an iterative manner. As each weak learner is added, a new model is fitted, resulting in a more accurate estimate of the response variable. The new weak learner can be maximally correlated with the negative gradient of the loss function, which is associated with the entire ensemble. Examples of gradient boosting machines include XGBoost and LightGBM. Additionally or alternatively, other types of machine learning techniques can be used to build the combined model, including bagging, boosting, and / or random forest algorithms.
[0040] The presentation model 1010 can be trained using information related to the peptide (e.g., peptide sequence, sequence of the MHC molecule that binds the peptide, length of the peptide), as well as information related to the expression level of the source protein from which the peptide is derived, surface presentation properties of the peptide, gene propensity score, and hotspot score. Thus, the trained presentation model 1010 can identify the binding properties of a given peptide and its surface presentation properties, i.e., whether the peptide is presented on the cell surface. Similar to the binding model 1005, the presentation model 1010 can include one or more trained gradient boosting algorithms.
[0041] 2. Model Architecture FIG. 11 illustrates an exemplary model architecture for training a machine learning model to predict surface-displayed peptides, according to some embodiments. As shown in FIG. 11, training databases are represented by columns containing various types of information, including allele data obtained from publicly available sources. For example, the dark gray columns contain immunopeptidomics data corresponding to engineered monoallelic cell lines (see FIG. 4). In another example, the training databases can also include in vitro binding data from publicly available data sources (e.g., the IEDB database, represented by the white columns). In some cases, training datasets from each training database are used to train corresponding binding and presentation models individually. Furthermore, training databases can be combined into a larger training database to train its corresponding binding and presentation models (e.g., the "ALL(MONO)" light gray columns in FIG. 11).
[0042] FIG. 11 further illustrates multiple sets of binding and presentation models trained to predict surface-displayed peptides. Each set of binding and presentation models is shown as being trained with a different training dataset. In some cases, output generated from a first set of models 1105 ("initial models") is used as input features to train a second set of models 1110 ("intermediate models"). For example, output generated by an initial model corresponding to in vitro binding, single-allelic data from a genetically engineered monoallelic cell line, can be used as input features to train the intermediate models. The intermediate models can also be trained separately with training data obtained from all the single-allelic data 1115. Furthermore, output from the intermediate models can be deconvolved and added to a separate training database containing both single-allelic and multi-allelic data. Output can be deconvolved using one or more base monoallelic base models or unsupervised clustering and alignment algorithms such as GibbsCluster.
[0043] A training dataset from the database containing the deconvoluted set of peptides for each HLA allele can be used to train a third set of presentation and binding models 1120 ("final models"). The trained final models 1120 can be deployed to predict surface-presented peptides. The goal in building the training database and training the final models is to capture as much allelic diversity as possible and avoid problems caused by underfitting and overfitting. Additionally or alternatively, the intermediate trained models can also be deployed to predict surface-presented peptides, although the performance level of the final trained models tends to be superior to that of the intermediate trained models.
[0044] V. Evaluating the performance level of machine learning models To evaluate the performance of the trained machine learning model, a test dataset is generated containing several experimentally observed peptides and synthetic decoys that were not part of the training process. The trained machine learning model processes these candidate test peptides and outputs scores predicting MHC class I binding and cell surface presentation. The machine learning model was trained using the large-scale immunopeptidome training dataset described above. The scores are then compared to corresponding data obtained from validated MHC-binding peptides presented on the cell surface to reveal the performance level of the trained machine learning model. The output scores are also evaluated against NetMHCpan 4.0 (a known platform for predicting peptide binding to MHC molecules), and the trained machine learning algorithm shows higher overall sensitivity and specificity. Based on the output scores, the antigen burden scores of the predicted peptides can be calculated using peptides with output scores above a confidence threshold.
[0045] In another example, we tested and evaluated the trained machine learning model using experimentally generated peptides from tissue samples using a mass spectrometry-based immunopeptide approach, mixed with decoys at a 1:999 ratio. Compared to NetMHCPan 4.0, a publicly available tool considered the gold standard for MHC-binding peptide prediction, the trained machine learning model showed significantly higher positive predictive value in the top 0.1% of predicted ligands. In yet another example, we evaluated the trained machine learning model using leave-one-out analysis, demonstrating high concordance between motifs in the raw data and those predicted by the trained machine learning model.
[0046] 1. Model evaluation for single-allele data a) Positive predictive value 12 shows the performance levels of the trained binding model and the trained display model measured in terms of positive predictive value based on 10% holdout data, according to some embodiments. The evaluation data is based on single-allele immunopeptidomics data. The positive predictive value (PPV) is defined as the proportion of predicted positives of the trained machine learning model that are actually positive. Thus, PPV reflects the probability that a predicted positive is a true positive. In the evaluation dataset, the positive rate, which indicates the ratio of positives to negatives, is 1:999.
[0047] As shown in Figure 12, the median PPV corresponding to NetMHCpan is approximately 0.4. In contrast, the trained combined models performed relatively better than NetMHCpan, with combined models trained on single-allelic data having a median PPV of approximately 0.6, and combined models trained on single-allelic and multi-allelic data having comparable median PPVs of approximately 0.6. The first trained representation model trained on single-allelic data and the second trained representation model trained on single-allelic and multi-allelic data performed significantly better than NetMHCpan, with a median PPV of approximately 0.7. The performance difference of 0.1 PPV value may be due to the fact that the evaluation data was derived from single-allelic data.
[0048] Figure 13 shows a comparison of the performance levels of the trained machine learning models compared to conventional techniques for predicting MHC-binding peptides. As shown in Figure 13, the other conventional techniques exhibit a median PPV of approximately 0.6, which is comparable to the median PPV corresponding to the trained binding model. Compared to the above models, the trained proposed model tends to perform better, with a median PPV of approximately 0.7.
[0049] Figure 14 shows a comparison of the performance levels of the trained display model for various alleles compared to conventional techniques for predicting MHC-binding peptides. The PPV values for each allele are shown for NetMHCpan and the trained display model. As shown in Figure 14, the PPV values corresponding to the trained display model are significantly higher than those of NetMHCpan across all single alleles. Thus, the trained display model demonstrates a significant improvement over NetMHCpan in predicting MHC-binding peptides displayed and expressed on the cell surface.
[0050] b) Leave-one-out analysis Figure 15 shows the results of a leave-one-out analysis of a trained display model, according to some embodiments. To demonstrate the performance of a trained display model in discovering unknown types of MHC-binding peptides that may be displayed on cell surfaces, a leave-one-out analysis can be used to evaluate whether the trained display model can predict surface-displayed peptides corresponding to alleles that are not present in any of the training data. To perform the leave-one-out analysis, the display model was trained on a training data set that excluded training data corresponding to one specific allele. After training, the trained machine learning model was evaluated by processing 500,000 random peptides and predicting surface-displayed peptides encoded by alleles from which at least some MHC-binding peptides were excluded. To evaluate the accuracy of peptide predictions by the trained machine learning model, the motifs of the predicted MHC-binding peptides were compared with motifs obtained from raw data for which specific alleles were available.
[0051] As shown in Figure 15, the motifs corresponding to the predicted surface-presented peptides substantially match the motifs of the peptides corresponding to the excluded alleles, which show comparable amino acid expression levels across the nine positions of the subject peptide. *The second position of the peptide corresponding to 44.03 shows a high expression level of glutamic acid ("E") in the raw data. The predicted MHC-binding peptides presented on the cell surface also show a high expression level of glutamic acid at the same second position. Therefore, the trained machine learning model can accurately predict MHC-binding peptides presented on the cell surface even when the corresponding allele is not part of the training data.
[0052] c) Precision and recall Figure 16 shows a graph illustrating precision and recall values for evaluating trained machine learning models, according to some embodiments. Precision-recall can be a useful indicator of prediction success. In information retrieval, precision is a measure of result relevance, while recall is a measure of how many truly relevant results are returned. High precision is related to a low false positive rate, and high recall is related to a low false negative rate. High scores for both precision and recall can indicate that a given classifier is returning accurate results (high precision) and a large proportion of positive results (high recall). The performance of the trained machine learning model was evaluated based on held-out single-allele data, using 10% of the immunopeptidomics data from training mixed with synthetic negative examples at a 1:999 ratio. The x-axis of the graph corresponds to a set of rank percentile thresholds ranging from 0.02 to 1.0, which will identify surface-displayed peptides within a particular rank percentile threshold to be considered for either binding or presentation.
[0053] As shown in Figure 16, the trained machine learning model corresponds to a higher precision for all recall values compared to NetMHCpan. The difference is further accentuated for the top 1% peptides in the test data, where the median precision to recall ratio for the trained machine learning model is approximately 0.8 / 0.6, while the median precision to recall ratio for NetMHCpan is approximately 0.5 / 0.2. Therefore, the trained machine learning model can demonstrate improved prediction of surface-displayed peptides compared to NetMHCpan, returning accurate results and a larger proportion of all positive results.
[0054] 2. Model evaluation with multi-allele (tissue) samples Furthermore, the performance level of the machine learning model trained using multi-allelic samples demonstrates improved prediction of surface-displayed peptides compared to conventional techniques such as NetMHCpan. Figure 17 shows a box plot depicting the performance level of a trained machine learning model across various tissue samples, according to some embodiments. In Figure 17, three types of tissue samples were processed with the trained machine learning model to generate the fraction of true candidates corresponding to surface-displayed peptides recovered. Therefore, a higher fraction may suggest that the trained machine learning model can demonstrate a high performance level in accurately identifying surface-displayed peptides across various tissue samples.
[0055] For example, the ratio value corresponding to NetMHCpan is approximately 0.65. This ratio value indicates that NetMHCpan was able to predict approximately 65% of the surface-displayed peptides actually present in the tissue sample. In contrast, the trained binding model performed better than NetMHCpan, with ratio values of approximately 0.81 for the binding model trained with single-allelic data and approximately 0.85 for the binding model trained with single- and multi-allelic data. The first trained display model trained with single-allelic data and the second trained display model trained with single- and multi-allelic data performed even better, both corresponding to ratio values of approximately 0.9. Thus, the trained display models revealed approximately 90% of the surface-displayed peptides experimentally identified in the tissue sample. Similar improvements in predicting surface-displayed peptides were also demonstrated for other tissue samples. Figure 18 shows a graph comparing the performance levels of a trained machine learning model according to some embodiments with other conventional techniques.
[0056] VI. Example of a process for predicting MHC-binding peptides presented on the cell surface FIG. 19 includes a flowchart 1900 illustrating an example of a method for predicting surface-displayed peptides, according to certain embodiments. The operations described in flowchart 1900 can be performed by a computer system implementing a trained machine learning model, such as a trained binding and display model. While flowchart 1900 may describe the operations as a sequential process, in various embodiments, many of the operations can be performed in parallel or simultaneously. Furthermore, the order of the operations can be rearranged. The operations may include additional steps not shown in the figure. Furthermore, embodiments of the method can be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented by software, firmware, middleware, or microcode, the program code or code segments performing the associated tasks can be stored on a computer-readable medium, such as a storage medium.
[0057] In operation 1910, the computer system accesses a machine learning model trained using a training dataset that includes, for each peptide of a plurality of peptides identified by the training dataset, a protein signature of an MHC molecule (e.g., an HLA allele) that binds and presents the peptide on a cell surface, one or more expression levels representing the expression level of a gene encoding the peptide, and one or more peptide presentation metrics representing the amount of the peptide detected as being presented by the MHC molecule. The machine learning model is configured to generate an output indicating the degree to which the one or more expression levels and the one or more peptide presentation metrics are related according to a population-level relationship between expression and presentation.
[0058] In operation 1920, the computer system accesses genomic and transcriptomic data corresponding to a biological sample of a subject. The genomic and transcriptomic data of the biological sample is processed to identify candidate neoantigens (peptides). The genomic and transcriptomic data identifies one or more MHC molecules from the biological sample, and for each peptide (e.g., candidate neoantigens) in a set of peptides identified from a tissue sample, includes one or more values representing the peptide. At least one of the one or more values can be determined based on processing of the tissue sample. The one or more values can correspond to the type of peptide, the length of the peptide, the allele that binds the peptide, and the expression of a gene region encoding the peptide.
[0059] In operation 1930, the computer system determines a score for each peptide in the set of peptides using the machine learning model, one or more MHC molecules identified from the biological sample, and one or more values representing the peptide in the genomic and transcriptomic data. In some cases, the computer system uses a trained machine learning model to process the one or more values and output a score predicting MHC molecule binding and presentation for a given peptide.
[0060] In operation 1940, the computer system generates results based on the scores. The results may include an incomplete subset of peptides where the subset of peptides exceeds a predetermined threshold that is predicted to be surface-presented peptides. In some cases, the results may include motifs corresponding to each of the subset of peptides. Additionally or alternatively, the results may include a subset of peptides having a score above a particular ranking percentile (e.g., 0.02). In some cases, the results indicate, for each peptide in the set of peptides, whether the peptide is a surface-presented peptide, i.e., whether the peptide binds to a corresponding MHC molecule and is presented on the cell surface.
[0061] In some cases, the computer system selects an incomplete subset of the set of peptides, and the identification of the incomplete subset is performed in a manner that biases the selection towards peptides associated with regions in space that are associated with outlier peptides in the training dataset whose expression levels and peptide presentation metrics are related in a way that deviates from population-level relationships.
[0062] In operation 1950, the computer system outputs the results, after which the process 1900 ends.
[0063] VII. Computer Environment 20 illustrates an example of a computer system 2000 for implementing some of the embodiments disclosed herein. The computer system 2000 may have a distributed architecture, with some components (e.g., memory and processor) being part of an end-user device and some other similar components (e.g., memory and processor) being part of a computer server. The computer system 2000 includes at least a processor 2002, a memory 2004, a storage device 2006, input / output (I / O) peripherals 2008, communication peripherals 2010, and an interface bus 2012. The interface bus 2012 is configured to communicate, transmit, and transfer data, control, and commands between various components of the computer system 2000. The processor 2002 may include one or more processing units, such as a CPU, a GPU, a TPU, a systolic array, or a SIMD processor. Memory 2004 and storage 2006 include computer-readable storage media, such as RAM, ROM, electrically erasable programmable read-only memory (EEPROM), hard drives, CD-ROMs, optical storage devices, magnetic storage devices, electronic non-volatile computer storage devices, such as Flash memory, and other tangible storage media. Any such computer-readable storage media can be configured to store instructions or program code embodying aspects of the present disclosure. Memory 2004 and storage 2006 also include computer-readable signal media. Computer-readable signal media include a propagated data signal having computer-readable program code embodied therein. Such a propagated signal may take any of a variety of forms, including, but not limited to, electromagnetic, optical, or any combination thereof. Computer-readable signal media is not a computer-readable storage medium, but includes any computer-readable medium capable of communicating, propagating, or transmitting a program for use in connection with computer system 2000.
[0064] Additionally, memory 2004 includes an operating system, programs, and applications. Processor 2002 is configured to execute stored instructions and includes, for example, a logic processing unit, microprocessor, digital signal processor, and other processors. Memory 2004 and / or processor 2002 can be virtualized and / or hosted within another computing system, for example, in a cloud network or data center. I / O peripherals 2008 include computing components such as user interfaces and graphical processing units, such as keyboards, screens (e.g., touchscreens), microphones, speakers, and other input / output devices, as well as serial ports, parallel ports, universal serial buses, and other input / output peripherals. I / O peripherals 2008 are connected to processor 2002 through any of the ports connected to interface bus 2012. Communications peripherals 2010 are configured to facilitate communications between computer system 2000 and other computing devices over a communications network and include, for example, network interface controllers, modems, wireless and wired interface cards, antennas, and other communications peripherals.
[0065] While the subject matter of the present invention has been described in detail with reference to specific embodiments thereof, it will be understood that those skilled in the art, upon gaining the foregoing understanding, will be able to readily make modifications, variations, and equivalents to such embodiments. Accordingly, it should be understood that the present disclosure is presented for purposes of illustration and not limitation, and does not preclude the inclusion of modifications, variations, and / or additions to the subject matter that would be readily apparent to those skilled in the art. Indeed, the methods and systems described herein may be embodied in a variety of other forms, and further, various omissions, substitutions, and changes in the form of the methods and systems described herein may be made without departing from the spirit of the present disclosure. The appended claims and their equivalents are intended to cover such forms or modifications as fall within the scope and spirit of the present disclosure.
[0066] Unless specifically stated otherwise, throughout this specification, discussions using terms such as "processing," "computing," "calculating," "determining," and "identifying," or the like, are understood to refer to the operations or processes of a computing device, such as one or more computers or similar electronic computing device or devices, that manipulate or transform data represented as physical electronic or magnetic quantities in the memory, registers, or other information storage, transmission, or display devices of the computing platform.
[0067] The system or systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device can include any suitable arrangement of components that provides a result conditioned on one or more inputs. Suitable computing devices range from general-purpose computing devices to specialized computing devices that implement one or more embodiments of the present subject matter, including general-purpose microprocessor-based computing systems that access stored software that programs or configures the computing system. Any suitable programming, scripting, or other type or combination of languages can be used to implement the teachings contained herein in software used to program or configure a computing device.
[0068] Embodiments of the methods disclosed herein may be performed in operation of such a computing device. The order of the blocks presented in the above examples may be changed, e.g., the blocks may be rearranged, combined, and / or divided into sub-blocks. Certain blocks or processes may be performed in parallel.
[0069] Conditional language used herein, particularly "can," "could," "might," "may," "e.g.," and the like, is generally intended to convey that certain examples include certain features, elements, and / or steps, while other examples do not, unless specifically stated otherwise or understood otherwise within the context in which it is used. Thus, such conditional language is generally not intended to imply that features, elements, and / or steps are in any way required by one or more examples, or that one or more examples necessarily include logic for determining whether those features, elements, and / or steps are included or performed in a particular example, with or without author input or prompting.
[0070] The terms "comprising," "including," "having," and the like are synonymous and are used in an open-ended, inclusive manner and do not exclude additional elements, features, acts, operations, etc. Also, the term "or" is used in an inclusive (and not exclusive) sense, so that, for example, when used connecting a list of elements, the term "or" means one, some, or all of the elements in the list. The use of "adapted to" or "configured to" herein means open, inclusive language that does not exclude devices adapted or configured to perform additional tasks or steps. Furthermore, the use of "based on" means open and inclusive in that a process, step, calculation, or other act "based on" one or more stated conditions or values may, in fact, be based on additional conditions or values beyond those stated. Similarly, the use of "based at least in part on" is meant to be open and inclusive, in that a process, step, calculation, or other act that is "based at least in part on" one or more stated conditions or values may, in fact, be based on additional conditions or values beyond those stated. The headings, lists, and numbering contained herein are for ease of description and are not meant to be limiting.
[0071] The various functions and processes described above can be used independently of one another or can be combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of the present disclosure. Furthermore, in some implementations, certain method or process blocks may be omitted. The methods and processes described herein are also not limited to a particular order, and the associated blocks or states may be performed in other orders as appropriate. For example, the described blocks or states may be performed in an order other than that specifically disclosed, or multiple blocks or states may be combined into a single block or state. Instances of blocks or states may be performed sequentially, in parallel, or in other ways. Blocks or states may be added to or deleted from the disclosed examples. Similarly, the example systems and components described herein may be configured differently than described. For example, elements may be added, deleted, or rearranged compared to the disclosed examples.
Claims
1. 1. A method of using a processor of a computer system to execute a program stored in its memory, comprising: (a) accessing a machine learning model stored in a storage device, The machine learning model is (i) trained for each peptide of a plurality of peptides identified by said training dataset using as attributes a training dataset that includes: (1) the protein characteristics of the MHC molecule that bind and present the peptide; (2) one or more expression levels representing the expression level of the gene encoding the peptide; and (3) one or more peptide presentation metrics representing the amount of peptide detected as presented by the MHC molecule; and (ii) accessing the machine learning model, the machine learning model being configured to generate a score indicating the degree to which the one or more expression levels of a tissue sample and the one or more peptide presentation metrics of the tissue sample are related according to a relationship between the expression levels of the training dataset and the peptide presentation metrics of the training dataset, wherein the relationship between the expression levels and the peptide presentation metrics of the training dataset varies depending on the population level of the training dataset; (b) accessing genomic and transcriptomic data stored in the storage device corresponding to a tissue sample of the subject, the genomic and transcriptomic data identifying one or more MHC molecules from the tissue sample and, for each peptide of a set of peptides identified from the tissue sample, including one or more values representing the peptide, at least one of the one or more values being determined based on processing of the tissue sample; (c) for each peptide in the set of peptides, determining a score using the machine learning model, the one or more MHC molecules identified from the tissue sample, and the one or more values representing the peptide; (d) generating a result based on the score; and (e) outputting the results using input / output (I / O) peripherals; Including, (i) the computer system includes the processor, the memory, the storage device, the input / output (I / O) peripherals, communication peripherals, and an interface bus; (ii) the training dataset is derived from monoallelic data corresponding to peptides derived from monoallelic cell lines and / or biallelic data corresponding to peptides derived from other tissue samples; (iii) attributes of the training data vector are data corresponding to somatic variants, including one or more features including peptide sequences; (iv) the labels of the data vectors of the training data are peptides encoded by somatic variants that bind to MHC molecules and are present on the cell surface; (v) the training dataset for training the machine learning model includes sequence data from various sources: (1) peptides identified as binding to HLA molecules based on in vitro experiments, (2) peptides identified by mass spectrometry analysis of tumor samples, (3) HLA alleles, and (4) non-tumor samples. method.
2. selecting an incomplete subset of the set of peptides based on the score; 2. The method of claim 1, wherein the identification of the incomplete subset is performed in a manner that biases the selection toward peptides associated with scores that predict a more likely presentation compared to the probability predicted by the corresponding population-level relationship in the training dataset, and the result comprises the incomplete subset of the set of peptides.
3. selecting an incomplete subset of the set of peptides based on the score; 2. The method of claim 1, wherein the identification of the incomplete subset is performed in a manner that biases the selection toward peptides associated with regions in space, the regions being associated with outlier peptides in the training dataset whose expression levels and peptide presentation metrics are related in a manner that deviates from the population-level relationships.
4. The method of claim 1 , wherein the results include, for each of one or more peptides in the set of peptides, the peptide identification and the score.
5. 2. The method of claim 1, wherein for each peptide in the set of peptides, the one or more values representing the peptide are generated based on the amino acid sequence of the peptide, an indication of whether the peptide binds to one or more binding pockets of the MHC molecule, the expression level of the peptide in the tissue sample, and / or the length of the peptide.
6. The method of claim 1 , wherein the score corresponding to a peptide of the set of peptides corresponds to a predicted probability as to whether the peptide will bind to the MHC molecule and be presented on a cell surface.
7. The method of claim 1 , wherein the machine learning models include one or more trained gradient boosting algorithms.
8. 2. The method of claim 1, wherein the machine learning model includes, for each peptide of the plurality of peptides, a first sub-model trained on a first subset of the training dataset that includes a sequence corresponding to the peptide, a sequence of an MHC molecule that binds to the peptide, and / or a length of the peptide.
9. 9. The method of claim 8, wherein the machine learning model comprises, for each peptide of the plurality of peptides, a second sub-model trained on a second subset of the training dataset comprising one or more expression levels of a source protein from which the peptide is derived and surface presentation properties of the peptide.
10. The method of claim 9 , wherein the first sub-model and the second sub-model are each trained based on one or more scores generated by another set of sub-models. (a) one or more data processors; (b) a non-transitory computer-readable storage medium containing instructions that, when executed by said one or more data processors, cause said one or more data processors to perform the method of claim 1; Including, the system.
12. A non-transitory machine-readable storage medium containing instructions for causing one or more data processors to perform the method of claim 1.
Citation Information
Patent Citations
Improved HLA epitope prediction
US20190346442A1
Neoantigen identification, manufacture, and use
WO2018195357A1