Data-independent acquisition (DIA) for mass spectrometric analysis of MHC peptides

DIA and PASEF with machine learning-based libraries improve the detection and quantification of MHC peptides, addressing the challenge of low abundance and facilitating personalized cancer treatments.

WO2025166132A1PCT designated stage Publication Date: 2025-08-07GENENTECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/013990
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-26
Filing Date
2025-01-31
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Current mass spectrometry methods face challenges in accurately detecting and quantifying tumor-specific epitopes due to the low abundance of Major Histocompatibility Complex (MHC) peptides, particularly HLA-I peptides, which are crucial for personalized cancer treatments.

Method used

Utilizing data-independent acquisition (DIA) mass spectrometry combined with parallel accumulation serial fragmentation (PASEF) and machine learning-based methods to generate predicted patient-specific mass spectral libraries, enabling the detection and quantification of tumor-specific epitopes.

Benefits of technology

Enhances the sensitivity and accuracy of identifying MHC peptides, including patient-specific neoantigen peptides, facilitating personalized cancer treatments such as vaccines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025013990_07082025_PF_FP_ABST
    Figure US2025013990_07082025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the use of data-independent acquisition (DIA), optionally in combination with ion mobility (diaPASEF), for high-sensitivity mass spectrometry-based identification of MHC presented peptides in a sample (e.g., a patient sample). In some instances, for example, the methods can comprise: predicting, using one or more machine-learning models, an MHC allele-specific mass spectral library for the sample; performing a mass spectrometric analysis of peptides extracted from the sample using data-independent acquisition (DIA) to generate mass spectral data for the peptides; and performing a spectral library search using the mass spectral data for the peptides and the predicted MHC allele-specific mass spectral library for the sample to identify at least one MHC peptide present in the sample.
Need to check novelty before this filing date? Find Prior Art

Description

DATA-INDEPENDENT ACQUISITION (DIA) FOR MASS SPECTROMETRIC ANALYSIS OF MHC PEPTIDESCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the priority benefit of United States Provisional Patent Application Serial No. 63 / 548,610, filed February 1, 2024, and to United States Provisional Patent Application Serial No. 63 / 676,256, filed July 26, 2024, the contents of each of which are incorporated herein by reference in their entirety.FIELD

[0002] This disclosure relates generally to mass spectrometric analysis of MHC peptides, and more specifically to the use of data-independent acquisition (DIA), alone or in combination with ion mobility (diaPASEF), for high sensitivity mass spectrometry-based identification of MHC presented peptides.BACKGROUND

[0003] Major histocompatibility complex (MHC) molecules are cell surface proteins that bind peptide fragments derived from endogenous or foreign proteins and display them on the cell surface for recognition by T-cells in vertebrate organisms. Human leukocyte antigen class I (HLA-I) molecules (human-specific MHC molecules) present short peptide sequences from endogenous or foreign proteins to cytotoxic T cells. The low abundance of HLA-I peptides poses significant technical challenges for their identification and accurate quantification. While mass spectrometry (MS) is currently a method of choice for direct system-wide identification of cellular immunopeptidome, there is still a need for enhanced sensitivity in detecting and quantifying tumor specific epitopes.SUMMARY

[0004] Disclosed herein are sensitive mass spectrometry-based methods and systems for analyzing peptides that bind to major histocompatibility complex (MHC) proteins. The methods comprise the use of data-independent acquisition (DIA) mass spectrometry (optionally in combination with parallel accumulation serial fragmentation (PASEF)) to detect and quantify tumor specific epitopes. In some embodiments, the methods further comprise the use of machine learning (ML)-based methods for generating predicted (or “synthetic”) patient-specific mass spectral libraries to facilitate the detection and quantification of tumor specific epitopes.

[0005] The disclosed methods and systems enable more accurate and sensitive detection of MHC peptides (e.g., HLA-I peptides) that are present in samples, including patient-specific neoantigen peptides present in samples from patients diagnosed with a cancer. The improved detection sensitivity, quantification, and accurate identification of MHC peptides, e.g., patientspecific neoantigen peptides, enabled by the disclosed methods can, in turn, facilitate the development of personalized cancer treatments, e.g., personalized cancer vaccines.

[0006] Disclosed herein are methods for identifying MHC peptides present in a sample, the method comprising: predicting, using one or more machine-learning models, an MHC allelespecific mass spectral library for the sample; performing a mass spectrometric analysis of peptides extracted from the sample using data-independent acquisition (DIA) to generate mass spectral data for the peptides; and performing a spectral library search using the mass spectral data for the peptides and the predicted MHC allele-specific mass spectral library for the sample to identify at least one MHC peptide present in the sample.

[0007] In some embodiments, predicting the MHC allele-specific mass spectral library comprises: generating a plurality of candidate peptide sequences from a reference proteome having a sequence length within a specified range of sequence lengths; predicting binding of the candidate peptide sequences of the plurality to at least one MHC allele using a machine learning-based MHC binding prediction model to generate a plurality of candidate MHC peptide sequences; and predicting mass spectral characteristics for each candidate MHC peptide sequence of the plurality using a machine learning-based mass spectra prediction model to generate the predicted MHC allele-specific mass spectral library.

[0008] In some embodiments, the specified range of sequence lengths is from 8 to 13 amino acid residues.

[0009] In some embodiments, the machine learning-based MHC binding prediction model is HLApollo.

[0010] In some embodiments, the machine learning-based mass spectra prediction model is AlphaPeptDeep, directDIA™, DIA-NN, Prosit, or a proprietary mass spectra prediction model.

[0011] In some embodiments, the mass spectral characteristics for each candidate MHC peptide sequence comprise a peptide retention time, a peptide fragmentation pattern, an ion mobility prediction, or any combination thereof.

[0012] In some embodiments, the at least one MHC peptide present in the sample comprises at least one neoantigen.

[0013] In some embodiments, the method further comprises extracting and / or purifying the peptides from the sample.

[0014] In some embodiments, performing the mass spectrometric analysis of peptides extracted from the sample comprises performing parallel accumulation serial fragmentation (PASEF) in combination with data-independent acquisition (DIA).

[0015] In some embodiments, the sample is derived from a human subject, and the at least one MHC peptide is an HLA peptide. In some embodiments, the human subject is a patient. In some embodiments, the sample is derived from a primate, a mouse, a bacterial culture, a viral culture, or a cultured cell line.

[0016] In some embodiments, the reference proteome is a human reference proteome, a primate reference proteome, a mouse reference proteome, a bacterial reference proteome, or a viral reference proteome.

[0017] Also disclosed herein are systems comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform any of the methods described herein.

[0018] Disclosed herein are non-transitory computer-readable storage media storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to perform any of the methods described herein.

[0019] The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention as claimedhas been disclosed by specific, exemplary implementations and optional features, modification and variation of the concepts herein disclosed can be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of the invention as defined by the appended claims.INCORPORATION BY REFERENCE

[0020] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety. In the event of a conflict between a term herein and a term in an incorporated reference, the term herein controls.BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0022] The present disclosure is described in conjunction with the appended figures:

[0023] FIG. 1 provides a non-limiting example of a process flowchart for identifying MHC peptides (e.g., HLA-1 peptides) present in a sample, in accordance with one implementation or the disclosed methods.

[0024] FIG. 2 provides a non-limiting schematic illustration of a method for identifying HLA peptides present in a sample, in accordance with one implementation of the disclosed methods.

[0025] FIG. 3 provides a non-limiting example of a block diagram of a computer system, in accordance with some implementations of the methods and systems disclosed herein.

[0026] FIG. 4 provides a non-limiting example of a block diagram of an artificial intelligence (Al) architecture included as part of an example spectral library prediction system, in accordance with some implementations of the methods and systems disclosed herein.

[0027] FIGS. 5A-F provide an overview of an HLA peptide enrichment workflow and nonlimiting examples of diaPASEF data that demonstrates identification of HLA peptides with high reproducibility. FIG. 5A: Overview of HLA peptide enrichment workflow. FIG. 5B: preparation of a sample specific spectral library from HLA peptides. FIG. 5C: Number ofunique peptide identifications. FIG. 5D: Length distribution of peptide IDs. FIG. 5E: Charge distribution. FIG. 5F: Intensity correlation between 2 replicates. FIG. 5G: Fraction of observed peptides assigned to alleles in DDA library and DIA peptide identifications.

[0028] FIGS. 6A-F provide non-limiting examples of data comparing the performance of DIA to that of DDA. DIA outperforms DDA in half the analysis time. FIG. 6A: HLA peptide identifications in DDA and DIA measurements of peptides enriched from decreasing number of A375 cells. FIG. 6B: Percent overlap of peptide identifications for DDA and DIA for A375. FIG. 6C: Intensity distribution for precursors identified in both DDA and DIA (overlap) and precursors unique to either acquisition type for A375. FIG. 6D: Same plot description as that for FIG. 6A for peptides isolated from C1R cell line monoallelic for HLA-A* 11:01. FIG. 6E: Same plot description as that for FIG. 6B for peptides isolated from C1R cell line monoallelic for HLA-A* 11 :01. FIG. 6F: Same plot description as that for FIG. 6C for peptides isolated from C1R cell line monoallelic for HLA-A*l l :01.

[0029] FIGS. 7A-E provide non-limiting examples of data comparing the use of predicted and empirical spectral libraries for HLA peptides to query diaPASEF data. FIG. 7A: Comparison of empirically observed retention times to predicted retention times of peptides in the spectral library. FIG. 7B: Difference of DIA-measured RT apex between both libraries. FIG. 7C: Ratio of quantified fragment ions compared to available fragment ions in empirical and predicted libraries. The symbol *** denotes p value <0.01. FIG. 7D: Spectral angle of identifications in empirical and predicted libraries. The dashed line indicates the median spectral angle. FIG. 7E: Overlap of identifications when using empirical or predicted libraries.

[0030] FIGS. 8A-F provide non-limiting examples of data illustrating that the predicted spectral library for all A375 HLA alleles generally outperformed empirical library. FIG. 8A: In silico library creation by predicting all possible HLA- peptide combinations for A375 cells and spectral feature prediction with Salud. FIG. 8B: Number of unique HLA peptides identified with empirical (green / light) or predicted library (red / dark). FIG. 8C: Coefficient of variation of identifications using different spectral libraries. FIG. 8D: Venn diagrams of identified peptides for 0.5e6(top) and 25e6cells (bottom). FIG. 8E: Peptide sequence motifs for identifications overlapping between libraries or uniquely identified with either library. FIG. 8F: Intensity and spectral angle distribution for identifications overlapping between libraries or uniquely identified with either library for peptides enriched from 0.5e6cells.

[0031] FIGS. 9A-B provide a non-limiting example of diaPASEF quantification benchmarking with SILAC. FIG. 9A: SILAC labeling strategy for HLA peptide analysis. FIG. 9B: Violin and density plot for log2 H / L ratios across mixing ratio experiments. Box plots show median ratio, dotted lines indicate expected ratio.

[0032] FIGS. 10A-G provide a non-limiting example of diaPASEF identification and quantification of A* 11 :01 neoantigens. FIG. 10A: HLA-I null C1R cell lines were transfected with a neoantigen cassette containing 47 common cancer mutations and an HLA-I allele of interest and HLA ligands were eluted from increasing cell amounts. FIG. 10B: HLA peptide identifications in DDA and DIA measurements of peptides enriched from increasing cell amounts. FIG. 10C: MSI area of endogenous presented peptides increased while synthetic spike in areas were stable. FIG. 10D-1: Extracted Ion Chromatograms (XIC) of endogenous (top) and heavy labeled neoantigen peptides (le6cell input). FIG. 10D-2: Extracted Ion Chromatograms (XIC) of endogenous (top) and heavy labeled neoantigen peptides (50e6cell input). FIG. 10E: The same as FIGS. 10D-1 and 10D-2, but quantified using DIANN MSI Area. FIG. 10F: Intensity of neoantigen identifications in DDA and DIA acquisition increases with increasing sample input. FIG. 10G: Abundance of all neoantigens identified and quantified by DIA in A*l l :01 C1R cells normalized to lOOfmol of heavy synthetic spike in peptides.

[0033] FIGS. 11A-C provide a non-limiting example of the impact of variable window design on peptide identification. FIG. 11 A: Variable window design for diaPASEF with HLA Class I peptides. FIG. 11B: Variable window design enables equal feature distribution per cycle. FIG. 11C: Number of peptides missing in replicate injections.

[0034] FIGS. 12A-E provide a non-limiting example of data for identification of HLA peptides. FIG. 12A: Precursor quantity in DIA across samples. FIG. 12B: Upset plot for DIA identifications. Horizontal bars at bottom left indicate the number of total unique peptides per acquisition; dots and lines under the vertical bars identify peptides identified in only one (single dot) or multiple acquisitions (dots with lines). FIG. 12C: Same description as that for FIG. 12B for DDA identifications. FIG. 12D: Fraction of observed peptides assigned to HLA alleles in A375, separated by identifications found in both DDA and DIA or unique to either acquisition. FIG. 12E: Hexbin plot for precursor m / z distribution and intensity values in 10e6 sample separated by identifications found in both DDA and DIA or unique to either acquisition.

[0035] FIGS. 13A-B provide non-limiting examples of the correlation between predicted retention time as calculated by DIA-NN (iRT) and measured RT for Salud predicted libraries (FIG. 13A) and an empirical library (FIG. 13B).

[0036] FIGS. 14A-C provide a non-limiting example of the identification of precursors using predicted or empirical spectral libraries. FIG. 14A: Overlap of peptide sequences in predicted library and empirical library (left) and properties of peptides unique to empirical library (right). FIG. 14B-1: Low abundant precursors are preferentially identified with the empirical library only (plots of kernel density versus log2(intensity) for cell inputs of 0.5e6, le6, and 2.5e6). FIG. 14B-2 : Low abundant precursors are preferentially identified with the empirical library only (plots of kernel density versus log2(intensity) for cell inputs of 5e6, 10e6, and 25e6). FIG. 14C-1: Spectral angle for identifications overlapping between libraries or uniquely identified with either library (plots of kernel density versus spectral angle for cell inputs of 0.5e6, le6, and 2.5e6). FIG. 14C-2: Spectral angle for identifications overlapping between libraries or uniquely identified with either library (plots of kernel density versus spectral angle for cell inputs of 5e6, 10e6, and 25e6).

[0037] FIGS. 15A-D provide non-limiting examples of data for predicted ion mobilities and fragment ion distributions. FIG. 15A: Ion mobility reduces co-eluting theoretical precursors. FIG. 15B: Predicted Ion Mobility for theoretically presented peptides, squares indicate elution window during acquisition. FIG. 15C: Number of theoretical eluting precursors per ion mobility window and instrument acquisition cycle. FIGS. 15D-1 to 15D-4: Fragment ion distribution (b-ions: FIGS. 15D-1 and 15D-3, y-ions: FIGS. 15D-2 and 15D-4) of all possible coeluting predicted precursors in a large window of mostly singly charged precursors (window 1.2: FIGS. 15D-1 and 15D-2) and a narrow window of doubly charged precursors (window 8.2: FIGS. 15D-3 and 15D-4). Color indicates fragment ion occurrences in multiple precursors (blue / dark) or in a unique precursor (orange / light) in a particular window / cycle combination.

[0038] FIGS. 16A-D provide non-limiting examples of density plot and C score data for heavy and light chain peptides. FIG. 16A: 2D density plots showing log2 H / L ratios for peptides according to their summed H+L intensity. FIG. 16B: C-Score distribution of heavy and light peptide identifications for each mixing ratio. FIG. 16C: Same as FIG. 16B, but after filtering for channel q-vale <0.01 and translated q-value <0.01. FIG. 16D: H / L density basedon summed intensity across mixing ratios after filtering for channel q-vale <0.01 and translated q-value <0.01. Dotted lines indicate expected ratio.

[0039] In the appended figures, similar components and / or features can have the same reference label. Further, various components of the same type can be distinguished by following the reference label by a dash and a second label that distinguishes among the similar components. If only the first reference label is used in the specification, the description is applicable to any one of the similar components having the same first reference label irrespective of the second reference label.DETAILED DESCRIPTION

[0040] Mass spectrometry -based methods and systems for analyzing peptides that bind to major histocompatibility complex (MHC) proteins are described. The methods comprise the use of data-independent acquisition (DIA) mass spectrometry (optionally in combination with parallel accumulation serial fragmentation (PASEF)) to detect and quantify tumor specific epitopes. In some instances, the methods can further comprise the use of machine learning (ML)-based methods for generating predicted (or “synthetic”) patient-specific mass spectral libraries to facilitate the detection and quantification of tumor specific epitopes.

[0041] The disclosed methods and systems provide more accurate and sensitive detection of MHC peptides (e.g., HLA- 1 peptides) that are present in samples, including patient-specific neoantigen peptides present in samples from patients diagnosed with a cancer.

[0042] In some instances, for example, the disclosed methods for identifying MHC (e.g., HLA) peptides present in a sample may comprise predicting, using one or more machinelearning models, an MHC allele-specific (e.g., HLA allele-specific) mass spectral library for the sample; performing a mass spectrometric analysis of peptides extracted from the sample using data-independent acquisition (DIA) to generate mass spectral data for the peptides; and performing a spectral library search using the mass spectral data for the peptides and the predicted MHC allele-specific (e.g., HLA allele-specific) mass spectral library for the sample to identify at least one MHC peptide (e.g., HLA peptide) present in the sample.

[0043] In some instances, predicting the MHC allele-specific mass spectral library can comprise: generating a plurality of candidate peptide sequences from a reference proteome having a sequence length within a specified range of sequence lengths; predicting binding of the candidate peptide sequences of the plurality to at least one MHC allele using a machinelearning-based MHC binding prediction model to generate a plurality of candidate MHC peptide sequences; and predicting mass spectral characteristics for each candidate MHC peptide sequence of the plurality using a machine learning-based mass spectra prediction model to generate the predicted MHC allele-specific mass spectral library.

[0044] In some instances, the mass spectral characteristics for each candidate MHC peptide sequence comprise a peptide retention time, a peptide fragmentation pattern, an ion mobility prediction, or any combination thereof.

[0045] In some instances, the mass spectrometric analysis of peptides extracted from the sample can comprise using data-independent acquisition (DIA) in combination with parallel accumulation serial fragmentation (PASEF).

[0046] In some instances, the sample can be derived from a human subject (e.g., a patient), and the at least one MHC peptide comprises an HLA peptide. In some instances, the sample can be derived from a primate, a mouse, a bacterial culture, a viral culture, or a cultured cell line.

[0047] In some instances, the reference proteome can be a human reference proteome, a primate reference proteome, a mouse reference proteome, a bacterial reference proteome, or a viral reference proteome.Example Description of Terms

[0048] Unless otherwise defined, all of the technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art in the field to which this disclosure belongs.

[0049] As used in this specification and the appended claims, the singular forms “a”, “an”, and “the” include plural references unless the context clearly dictates otherwise. Any reference to “or” herein is intended to encompass “and / or” unless otherwise stated.

[0050] “About” and “approximately” shall generally mean an acceptable degree of error for the quantity measured given the nature or precision of the measurements. Examples of acceptable degrees of error are typically within 20 percent (%), within 10%, or within 5% of a given value or range of values.

[0051] As used herein, the terms "comprising" (and any form or variant of comprising, such as "comprise" and "comprises"), "having" (and any form or variant of having, such as "have"and "has"), "including" (and any form or variant of including, such as "includes" and "include"), or "containing" (and any form or variant of containing, such as "contains" and "contain"), are inclusive or open-ended and do not exclude additional, un-recited additives, components, integers, elements, or method steps.

[0052] As used herein, the terms “individual”, “patient”, or “subject” are used interchangeably and refer to any single being, e.g., a human being or a non-human mammal (e.g. , a dog, a cat, a horse, a cow, a pig, a sheep, a rabbit, a mouse, or a non-human primate) for which diagnosis and / or treatment is desired. In particular implementations, the individual, patient, or subject herein is a human.

[0053] The terms “cancer” and “tumor” are used interchangeably herein. These terms refer to the presence of cells possessing characteristics typical of cancer-causing cells, such as uncontrolled proliferation, immortality, metastatic potential, rapid growth and proliferation rate, and certain characteristic morphological features. Cancer cells are often found in the form of a tumor, but such cells can exist alone within an animal, or can be a non-tumorigenic cancer cell, such as a leukemia cell. These terms include a solid tumor, a soft tissue tumor, or a metastatic lesion. As used herein, the term “cancer” includes premalignant, as well as malignant cancers.

[0054] As used herein, “therapy” and “treatment” (and grammatical variations thereof, such as “treat” or “treating”) are used interchangeably and refer to clinical intervention (e.g., administration of an anti -cancer agent or anti-cancer therapy) in an attempt to alter the natural course of disease in the individual being treated, and can be performed either for prophylaxis or during the course of clinical pathology. Desirable effects of treatment include, but are not limited to, preventing occurrence or recurrence of disease, alleviation of symptoms, diminishment of any direct or indirect pathological consequences of the disease, preventing metastasis, decreasing the rate of disease progression, amelioration or palliation of the disease state, and remission or improved prognosis.

[0055] As used herein, the term “HLA peptide” refers to a peptide that is predicted or known to bind to an HLA protein (z.e., the expressed protein corresponding to an HLA gene allele).

[0056] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.Data-independent Acquisition (DIA) for Mass Spectrometric Analysis of MHC Peptide

[0057] While DIA for immunopeptidomics is not yet widely adopted, more groups have been exploring this type of acquisition for HLA peptides. For example, DIA has been used for identification of neoantigens (Pak et al. (2021) “Sensitive Immunopeptidomics by Leveraging Available Large-Scale Multi-HLA Spectral Libraries, Data-independent Acquisition, and MS / MS Prediction”, Mol. Cell. Proteomics 20, 100080) or for identification of peptides bound to soluble HLA (sHLA) (Wahle et al. (2024) “IMBAS-MS Discovers Organ-Specific HLA Peptide Patterns in Plasma”, Mol. Cell. Proteomics 23, 100689). A significant drawback for use of DIA in immunopeptidomic applications stems from the lack of clear digestion rules for HLA-presented peptides which makes the generation of predicted spectral library a computationally challenging task. Therefore, immunopeptidomics mainly relies on generation of experimental spectral libraries. This can be disadvantageous if analyzing samples across multiple cell lines or patients with varying HLA types as each of these will present peptides with different binding rules and varying amino acid anchor residues. Moreover, the nonspecific nature of these peptides can also complicate data analysis as coeluting precursors might lead to false identifications.

[0058] To address these shortcoming, we have evaluated the use of machine learning (ML)- based methods for generating predicted (or “synthetic”) patient-specific mass spectral libraries to facilitate the detection and quantification of tumor specific epitopes using data-independent acquisition (DIA) mass spectrometry (optionally in combination with parallel accumulation serial fragmentation (PASEF)).

[0059] FIG. 1 provides a non-limiting example of a flowchart for a process 200 for identifying MHC peptides (e.g., HLA-1 peptides) present in a patient sample. Process 100 can be performed, for example, as a computer-implemented method using software running on one or more processors of one or more electronic devices, computers, or computing platforms. In some examples, process 100 is performed using a client-server system, and the blocks of process 100 are divided up in any manner between the server and a client device. In other examples, the blocks of process 100 are divided up between the server and multiple client devices. Thus, while portions of process 100 are described herein as being performed by particular devices of a client-server system, it will be appreciated that process 100 is not so limited. In other examples, process 100 is performed using only a client device or only multiple client devices. In process 100, some blocks are, optionally, combined, the order of some blocksis, optionally, changed, and some blocks are, optionally, omitted. In some examples, additional steps may be performed in combination with the process 100. Accordingly, the operations as illustrated (and described in greater detail below) are exemplary by nature and, as such, should not be viewed as limiting.

[0060] At step 102 in FIG. 1, an MHC allele-specific (e.g., an HLA allele-specific) mass spectral library for a sample (e.g. , a sample of peptides extracted from a sample such as a patient tumor specimen) is predicted using one or more (e.g., 1, 2, 3, 4,5 or more than 5) machine learning models.

[0061] In some instances, for example, predicting the MHC allele-specific mass spectral library can comprise: generating a plurality of candidate peptide sequences from a reference proteome, where each candidate peptide has a sequence length within a specified range of sequence lengths; predicting binding of each candidate peptide sequences of the plurality of candidate peptide sequences to at least one MHC allele (z.e., the expressed protein corresponding to the at least one MHC gene allele) using a machine learning-based MHC binding prediction model to generate a plurality of candidate MHC peptide sequences; and predicting mass spectral characteristics for each candidate MHC peptide sequence of the plurality of candidate MHC peptide sequences using a machine learning-based mass spectra prediction model to generate the predicted MHC allele-specific mass spectral library.

[0062] In some instances, for example, the specified range of sequence lengths for candidate peptide sequences can be 6 to 13 amino acid residues, 6 to 15 amino acid residues, 8 to 13 amino acid residues, or 8 to 15 amino acid residues, including all possible subranges encompassed by these example ranges (e.g., from 7 to 12 amino acid residues in length).

[0063] In some instances, the plurality of candidate peptide sequences may comprise all possible peptide sequences within the specified range of lengths that can be generated from a reference proteome, e.g., a reference human proteome, a primate reference proteome, a mouse reference proteome, a bacterial reference proteome, or a viral reference proteome. In some instances, the plurality of candidate peptide sequences may comprise at least 106, 5 x 106, 107, 5 x 107, 108, 5 x 108, 109, or more than 109candidate peptide sequences. In some instances, the reference genome may comprise a human reference proteome as provided in, e.g., Swiss-PROT (Bairoch et al. (2000), “The SWISS-PROT Protein Sequence Database and its Supplement TrEMBL in 2000”, Nucleic Acids Research 28(l):45-48) or UniProt (The UniProt Consortium,“UniProt: the Universal Protein Knowledgebase in 2023”, Nucleic Acids Research 51(D1):D523-D531).

[0064] In some instances, at least one MHC allele (e.g., an HL A allele) present in the sample can be determined based on MHC typing (e.g., HLA typing; see, e.g., Bravo-Egana et al. (2021), “New challenges, new opportunities: Next generation sequencing and its place in the advancement of HLA typing”, Human Immunology 82:478-487), and may comprise, for example, an allele associated with the HLA Class I (HLA-I) genes ( . e. , the HLA- A, HLA-B, and HLA-C genes) or HLA Class II (HLA-II) genes (z.e., the HLA-DR, HLA-DQ, and HLA- DP genes). In some instances, the binding of each candidate peptide sequence to at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 MHC alleles / HLA alleles (i.e., the expressed proteins corresponding to the at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 MHC gene alleles / HLA gene alleles) is predicted. Candidate peptides that are predicted to bind to the at least one MHC / HLA allele (i.e., the expressed protein corresponding to the at least one MHC / HLA gene allele) can be referred to as candidate MHC (e.g., HLA) peptides.

[0065] In some instances, the machine learning-based MHC binding prediction model can be, e.g., HLApollo, a pan-allelic, transformer-based model for predicting peptide presentation by MHC-I (e.g. HLA-I). See, for example, Thrift et al. (2022), “HLApollo: A superior transformer model for pan-allelic peptide-MHC-I presentation prediction, with diverse negative coverage, deconvolution and protein language features”, bioRxiv, 2022.12.08.519673).

[0066] In some instances, the machine learning-based mass spectra prediction model can be, for example, AlphaPeptDeep (see, e.g., Zeng et al. (2022), “ AlphaP eptDeep: a modular deep learning framework to predict peptide properties for proteomics”, Nature Communications 13:7238), directDIA™ (Biognosys AG, Zurich, CH), DIA-NN (see, e.g., Demichev et al. (2020), “DIA-NN: neural networks and interference correction enable deep proteome coverage in high throughput”, Nature Methods 17:41-44), Prosit (see, e.g, Gessulat et al. (2019), “Prosit: proteome-wide prediction of peptide tandem mass spectra by deep learning”, Nature Methods 16:509-518), or a proprietary mass spectra prediction model.

[0067] Examples of mass spectral characteristics that may be predicted for each candidate MHC (e.g, HLA) peptide sequence include, but are not limited to, a peptide retention time, a peptide fragmentation pattern, an ion mobility prediction, or any combination thereof.

[0068] In some instances, the generation of the predicted mass spectral library may include or require the use of additional data.

[0069] At step 104 in FIG. 1, a mass spectrometric analysis of peptides extracted from the sample using data-independent acquisition (DIA) is performed to generate mass spectral data for the peptides.

[0070] In some instances, the sample may be derived from a human subject (e.g., a patient) or from a cultured cell line. In some instances, a sample derived from a human subject (e.g., a patient) may comprise a tumor specimen, e.g., a surgical resection sample, a tissue biopsy sample, or a formalin-fixed, paraffin-embedded (FFPE) tissue sample. Peptides can be subjected to enzymatic digestion and / or extracted from the tumor specimen using any of a variety of techniques known to those of skill in the art. Examples of extraction techniques include, but are not limited to, organic solvent precipitation, differential solubilization, centrifugal ultrafiltration, and solid phase extraction (see, for example, Peng et al. (2020), “Peptidomic analyses: The progress in enrichment and identification of endogenous peptides”, Trends in Analytical Chemistry 125: 115835).

[0071] In some instances, the peptide extraction method may be coupled directly to the mass spectrometric analysis, e.g., through the use of a liquid chromatography-mass spectrometric (LC-MS) or liquid chromatography-mass spectrometry-mass spectrometry (LC- MS-MS) technique. In some instances, for example, the LC-MS or LC-MS-MS technique may comprise the use of a reverse phase liquid chromatography (RPLC) technique.

[0072] In some instances, mass spectrometric analysis of the extracted peptides can be performed using data-independent acquisition (DIA) - a data acquisition mode that isolates and concurrently fragments populations of different precursors by cycling through segments of a predefined precursor m / z range (see, e.g., Meier et al. (2020), “diaPASEF: parallel accumulation- serial fragmentation combined with data-independent acquisition”, Nature Methods 17(12): 1229-1236).

[0073] In some instances, mass spectrometric analysis of the extracted peptides can be performed using parallel accumulation serial fragmentation (PASEF) in combination with data- independent acquisition (DIA) - a data acquisition mode that makes use of the correlation of molecular weight and ion mobility in a trapped ion mobility device (e.g. , a trapped ion mobility spectrometry (TIMS)-time-of-flight (TOF) mass spectrometer) that samples up to 100% of thepeptide precursor ion current in m / z and mobility windows (see, e.g., Meier et al. (2020), ibid.) and increases the reproducibility and quantitative accuracy of peptide identification.

[0074] At step 106 in FIG. 1, a spectral library search using the mass spectral data for the peptides and the predicted MHC allele-specific (e.g., HLA allele-specific) mass spectral library for the sample is performed to identify at least one MHC peptide (e.g., HLA peptide) present in the sample.

[0075] In some instances, the at least one MHC peptide (e.g., HLA peptide) identified as being present in the sample may include at least 1, at least 10, at least 100, at least 103, at least 104, at least 105, at least 106, or more than 106peptides. In some instances, the at least one MHC peptide (e.g., HLA peptide) present in the sample (e.g., a patient sample) can comprise at least one neoantigen peptide.

[0076] In some instances, the method can further comprise extracting and / or purifying the peptides from the sample.

[0077] Also described herein are systems (and associated computer-readable storage media) configured to perform process 100 (as illustrated in FIG. 1 and described in more detail below in Appendices I - III). For example, a system as described herein can comprise: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform any of the methods described herein. Also disclosed are non-transitory computer-readable storage media storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to perform any of the methods described herein.

[0078] FIG. 2 provides a schematic illustration of one implementation of the disclosed methods for identifying HLA peptides present in a sample. A conventional approach to identification of HLA peptides is illustrated in steps 202 to 212. Bound HLA peptides can be isolated from a sample (e.g., a patient sample) using an HLA immunoprecipitation technique, 202. The bound peptides can then be eluted at step 204 (e.g., by changing the assay buffer) and subjected to analysis using a direct data acquisition (DDA) approach to performing tandem mass spectrometry at step 206. Liquid chromatography-tandem mass spectrometry (LC- MS / MS) is a widely used technique for identifying unknown ions (e.g., ionized metabolites or peptide fragments) in untargeted metabolomics and proteomics studies. Data-dependentacquisition (DDA) chooses which ions to fragment for the second MS scan based upon intensities observed in the first MS survey scan, and typically only fragments a small subset of the ions present. The data acquired using DDA can then be used to build an empirical mass spectral library, 208. However, the generation of an empirical mass spectral library can be prohibitively burdensome in terms of sample consumption and the analysis time required.

[0079] The eluted peptides can also be subjected to analysis using a data independent acquisition (DIA) approach to performing tandem mass spectrometry at step 210. As noted above, DIA is a data acquisition mode that isolates and concurrently fragments populations of different precursors by cycling through segments of a predefined precursor m / z range, and in combination with parallel accumulation serial fragmentation (PASEF) provides significantly improved coverage of ionized peptide fragments with increased reproducibility and quantitation. The empirical library 208 and DIA mass spectral data 210 may then be utilized to perform spectral library searching and HLA peptide identification at step 212.

[0080] The approach disclosed herein provides an alternative approach to identifying HLA peptides present in a sample, and is summarized in steps 202, 204, 210, 214 (including 214a - 214d), and 212 in FIG. 2. Peptides may be isolated from a sample (e.g., a patient sample) using HLA immunoprecipitation, 202, and elution, 204, prior to being subjected to mass spectral analysis using a data independent acquisition (DIA) approach, 210. However, rather than generating an empirical spectral library, 208, the disclosed methods can rely on generating a predicted spectral library, 214, for example, by identifying all possible 8-12mer peptides in the human proteome (z.e., 154,423,819 peptide sequences in this example) at step 214a, identifying all possible paired combinations of the peptides and HLA proteins present in the patient (z.e., 926,542,914 pairwise combinations in this example) at step 214b, predicting which peptides bind to the HLA proteins present in the patient (z.e., 382,537 predicted binding peptides) using a machine learning-based binding prediction model at step 214c, and predicting the mass spectral characteristics for the predicted binding peptides (e.g., peptide retention time, peptide fragmentation pattern, ion mobility prediction, or any combination thereof) using a machine learning-based mass spectral characteristic prediction model at step 214d. The predicted spectral library 214 and DIA mass spectral data 210 may then be utilized to perform spectral library searching and HLA peptide identification at step 212. In some instances, the predicted mass spectral library 214 may be used alone, or in combination with an empirically-generated mass spectral library 208, in performing the spectral library search to identify HLA peptides.Example Computer Systems

[0081] FIG. 3 provides a non-limiting example of a block diagram for a computer system, in accordance with some implementations of the disclosed systems and methods.

[0082] Computer system 300 can be a host computer connected to a network. Computer system 300 can be a client computer or a server. As shown in FIG. 3, computer system 300 can be any suitable type of microprocessor-based device, such as a personal computer, workstation, server, or handheld computing device (portable electronic device), such as a phone or tablet. The device can include, for example, one or more of processor 310, input device 320, output device 330, storage 340, and communication device 360. Input device 320 and output device 330 can generally correspond to those described elsewhere herein, and they can either be connectable or integrated with the computer.

[0083] Input device 320 can be any suitable device that provides input, such as a touch screen, keyboard or keypad, mouse, or voice-recognition device. Output device 330 can be any suitable device that provides output, such as a touch screen, haptics device, or speaker.

[0084] Storage 340 can be any suitable device that provides storage, such as an electrical, magnetic, or optical memory including a RAM, cache, hard drive, or removable storage disk. Communication device 360 can include any suitable device capable of transmitting and receiving signals over a network, such as a network interface chip or device. The components of the computer can be connected in any suitable manner, such as via a physical bus 370 or wirelessly.

[0085] Software 350, which can be stored in memory / storage 340 and executed by processor 310, can include, for example, the programming that embodies the functionality of the present disclosure (e.g., as embodied in the methods described above).

[0086] Software 350 can also be stored and / or transported within any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context disclosure, a computer-readable storage medium can be any medium, such as storage 340, that can contain or store programming for use by or in connection with an instruction execution system, apparatus, or device.

[0087] Software 350 can also be propagated within any transport medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of the present disclosure, a transport medium can be any medium that can communicate, propagate, or transport programming for use by or in connection with an instruction execution system, apparatus, or device. The transport readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, or infrared wired or wireless propagation medium.

[0088] Computer system 300 may be connected to a network, which can be any suitable type of interconnected communication system. The network can implement any suitable communications protocol and can be secured by any suitable security protocol. The network can comprise network links of any suitable arrangement that can implement the transmission and reception of network signals, such as wireless network connections, T1 or T3 lines, cable networks, DSL, or telephone lines.

[0089] Computer system 300 can implement any operating system suitable for operating on the network. Software 350 can be written in any suitable programming language, such as C, C++, Java, or Python. In various instances, application software embodying the functionality of the present disclosure can be deployed in different configurations, such as in a client / server arrangement or through a web browser as a web-based application or web service, for example.

[0090] FIG. 4 illustrates a diagram 400 of an example artificial intelligence (Al) architecture 402 (which can be included as part of the one or more computing device(s) 300 as discussed above with respect to FIG. 3) that can be utilized to implement the methods described herein. In certain instances, the Al architecture 402 can be implemented utilizing, for example, one or more processing devices that may include hardware (e.g., a general purpose processor, a graphic processing unit (GPU), an application-specific integrated circuit (ASIC), a system- on-chip (SoC), a microcontroller, a field-programmable gate array (FPGA), a central processing unit (CPU), an application processor (AP), a visual processing unit (VPU), a neural processing unit (NPU), a neural decision processor (NDP), a deep learning processor (DLP), a tensor processing unit (TPU), a neuromorphic processing unit (NPU), and / or other processing device(s) that can be suitable for processing various molecular data and making one or moredecisions based thereon), software (e.g., instructions running / executing on one or more processing devices), firmware (e.g., microcode), or some combination thereof.

[0091] In certain instances, as depicted by FIG. 4, the Al architecture 402 may include machine learning (ML) algorithms and functions 404, natural language processing (NLP) algorithms and functions 406, expert systems 408, computer-based vision algorithms and functions 410, speech recognition algorithms and functions 412, planning algorithms and functions 414, and robotics algorithms and functions 416. In certain instances, the ML algorithms and functions 404 may include any statistics-based algorithms that can be suitable for finding patterns across large amounts of data (e.g., “Big Data” such as genomics data, proteomics data, metabolomics data, metagenomics data, transcriptomics data, or other omics data). For example, in certain instances, the ML algorithms and functions 404 may include deep learning algorithms 418, supervised learning algorithms 420, and unsupervised learning algorithms 422.

[0092] In certain instances, the deep learning algorithms 418 may include any artificial neural networks (ANNs) that can be utilized to learn deep levels of representations and abstractions from large amounts of data. For example, the deep learning algorithms 418 may include ANNs, such as a perceptron, a multilayer perceptron (MLP), an autoencoder (AE), a convolution neural network (CNN), a recurrent neural network (RNN), long short term memory (LSTM), a grated recurrent unit (GRU), a restricted Boltzmann Machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a generative adversarial network (GAN), and deep Q-networks, a neural autoregressive distribution estimation (NADE), an adversarial network (AN), attentional models (AM), a spiking neural network (SNN), deep reinforcement learning, and so forth.

[0093] In certain instances, the supervised learning algorithms 420 may include any algorithms that can be utilized to apply, for example, what has been learned in the past to new data using labeled examples for predicting future events. For example, starting from the analysis of a known training data set, the supervised learning algorithms 420 may produce an inferred function to make predictions about the output values. The supervised learning algorithms 420 may also compare its output with the correct and intended output and find errors in order to modify the supervised learning algorithms 420 accordingly. On the other hand, the unsupervised learning algorithms 422 may include any algorithms that may applied, for example, when the data used to train the unsupervised learning algorithms 422 are neitherclassified nor labeled. For example, the unsupervised learning algorithms 422 may study and analyze how systems may infer a function to describe a hidden structure from unlabeled data.

[0094] In certain instances, the NLP algorithms and functions 406 may include any algorithms or functions that can be suitable for automatically manipulating natural language, such as speech and / or text. For example, the NLP algorithms and functions 406 may include content extraction algorithms or functions 424, classification algorithms or functions 426, machine translation algorithms or functions 428, question answering (QA) algorithms or functions 430, and text generation algorithms or functions 432. In certain instances, the content extraction algorithms or functions 424 may include a means for extracting text or images from electronic documents (e.g., webpages, text editor documents, and so forth) to be utilized, for example, in other applications.

[0095] In certain instances, the classification algorithms or functions 426 may include any algorithms that may utilize a supervised learning model (e.g., logistic regression, naive Bayes, stochastic gradient descent (SGD), k-nearest neighbors, decision trees, random forests, support vector machine (SVM), and so forth) to learn from the data input to the supervised learning model and to make new observations or classifications based thereon. The machine translation algorithms or functions 428 may include any algorithms or functions that can be suitable for automatically converting source text in one language, for example, into text in another language. The QA algorithms or functions 2130 may include any algorithms or functions that can be suitable for automatically answering questions posed by humans in, for example, a natural language, such as that performed by voice-controlled personal assistant devices. The text generation algorithms or functions 432 may include any algorithms or functions that can be suitable for automatically generating natural language texts.

[0096] In certain instances, the expert systems 408 may include any algorithms or functions that can be suitable for simulating the judgment and behavior of a human or an organization that has expert knowledge and experience in a particular field (e.g., stock trading, medicine, sports statistics, and so forth). The computer-based vision algorithms and functions 410 may include any algorithms or functions that can be suitable for automatically extracting information from images (e.g., photo images, video images). For example, the computer-based vision algorithms and functions 410 may include image recognition algorithms 434 and machine vision algorithms 436. The image recognition algorithms 434 may include any algorithms that can be suitable for automatically identifying and / or classifying objects, places,people, and so forth that can be included in, for example, one or more image frames or other displayed data. The machine vision algorithms 436 may include any algorithms that can be suitable for allowing computers to “see”, or, for example, to rely on image sensors cameras with specialized optics to acquire images for processing, analyzing, and / or measuring various data characteristics for decision making purposes.

[0097] In certain instances, the speech recognition algorithms and functions 412 may include any algorithms or functions that can be suitable for recognizing and translating spoken language into text, such as through automatic speech recognition (ASR), computer speech recognition, speech -to-text (STT) 438, or text-to-speech (TTS) 440 in order for the computing to communicate via speech with one or more users, for example. In certain instances, the planning algorithms and functions 414 may include any algorithms or functions that can be suitable for generating a sequence of actions, in which each action may include its own set of preconditions to be satisfied before performing the action. Examples of Al planning may include classical planning, reduction to other problems, temporal planning, probabilistic planning, preference-based planning, conditional planning, and so forth. Lastly, the robotics algorithms and functions 416 may include any algorithms, functions, or systems that may enable one or more devices to replicate human behavior through, for example, motions, gestures, performance tasks, decision-making, emotions, and so forth.EXAMPLESExample 1 - Real-Time Prediction of Fragmentation and Retention Time Combined with a Real-Time Search Improves Multiplexed Immunopeptidome Quantification

[0098] Introduction: Immunopeptidomics, the analysis of peptides from endogenous or foreign proteins presented on the cell surface, is a vital part of the development of cancer vaccines and targeted cell therapies. Recent advances in machine learning (ML) prediction of peptide fragmentation and retention time have dramatically improved the identification rates for non-tryptic peptides enabling deeper and more confident analysis of immunopeptidomes. Here we introduce Salud, a collection of ML models specifically developed to predict fragmentation and retention time in real-time, during mass spectrometry analysis. We combine Salud with a rapid real-time MSFragger search within the instrument application programming interface (iAPI) program inSeqAPI to enable real-time rescoring of spectra and improve theconfidence of identification and depth of multiplexed quantification for immunopeptidomic analyses.

[0099] Methods: Salud and RT-MSFragger were incorporated into inSeqAPI through creation of custom C# “APIFilter” class objects. Similar to Prosit, Salud is built from an encoder / decoder architecture, but lacking components that prohibited conversion to a common model format that could be called from inSeqAPI. During real-time analysis a “SaludFragmentFilter” predicts the theoretical spectrum for a given peptide while a “SaludRTFilter” predicts the normalized retention time. Both values are then used within a real-time support vector machine (SVM) false discovery rate filter that performs real-time peptide FDR. For “MSFragger”, inSeqAPI utilizes OpenJDK to start an MSFragger service based on a predefined parameter file. MSFragger is then queried during the analysis to provide a rapid search of large non-tryptic databases.

[0100] Preliminary Data: We sought to improve the depth of multiplexed immunopeptidome quantification through improvements to real-time search methods within inSeqAPI. Real-time searches have not yet been applied to non-tryptic samples due to the size of the database and corresponding long search time per spectrum. Furthermore, large databases can lead to decreased identifications, a phenomenon remedied by machine learning (ML) prediction algorithms that allow rescoring of spectra following fragmentation and retention time prediction.

[0101] Here we describe Salud, a collection of ML models that enable real-time fragment and retention time prediction within inSeqAPI. We benchmarked Salud to determine the amount of time needed to perform real-time predictions. For a single peptide we found that Salud fragmentation prediction is performed in ~3 ms (median of 27,400 predictions) and retention time prediction is performed in ~1.8 ms (median of 27,400 predictions). Furthermore, utilization of these ML models during FDR determination improves identification of HLA peptides by 25-45%, similar to other described ML models.

[0102] To address the challenge of lengthy real-time non-tryptic searches, we introduce Real-Time MSFragger (RT-MSFragger) an implementation of the MSFragger algorithm that runs as a service and accepts spectra in real-time from the C# code of inSeqAPI. We compared Comet and MSFragger search times for a non-tryptic database of 7-12 mer peptides (z.e., common lengths for peptides identified in immunopeptidome analysis). Here we found thatMSFragger improved search times nearly ~80x compared to Comet (0.55 ms for MSFragger vs. 43 ms for Comet). Lastly, we demonstrate the ability of RT -MSFragger to search 7-25 mer non-tryptic peptides, a real-time search the current version of comet cannot perform.

[0103] The results indicate that machine learning algorithms executed in real-time for fragment and retention time prediction improve identification confidence. The RT -MSFragger search dramatically reduced the search time required for non-tryptic peptides.

[0104] Future work will benchmark the improvement in quantified peptides from the analysis of peptidomes and immunopeptidomes in a cell line that expresses mutation bearing regions from 47 common cancer mutations.Example 2 - Improving Multiplexed Single Cell Proteomics Through Combination of Real- Time Fragmentation and Retention Time Prediction and Real-Time Search

[0105] Introduction: Recent advances in data acquisition have improved the depth of multiplexed single cell proteomics (SCP). In particular, the utilization of a real-time search dedicates more instrument time to the quantitative analysis of peptides that will be identified in post-run analysis. Separately, the use of machine learning algorithms has improved peptide identification for challenging samples by utilizing predicted fragmentation intensities and retention times to aid in false discovery rate (FDR) determination. Here we describe two additions to inSeqAPI that improve the number of peptides and proteins quantified for SCP experiments: 1) Salud-a collection of ML algorithms that can be executed in real-time to enable rescoring during real-time FDR, and 2) Real-time MSFragger which performs a real-time faster than the current standard.

[0106] Methods: Salud is a collection of ML algorithms that have an encoder / decoder architecture similar to Prosit, but lack particular layers that prohibit conversion to a platform agnostic file. RT -MSFragger is written in Java and enables real-time search through creation of a service that can be queried from C# code. During real-time inSeqAPI analysis a “SaludFragmentFilter” predicts the theoretical spectrum for a given peptide while a “SaludRTFilter” predicts the normalized retention time. Both values are then used within a real-time support vector machine (SVM) false discovery rate filter that performs real-time peptide FDR. For “MSFragger”, inSeqAPI utilizes OpenJDK to start an MSFragger service based on a predefined parameter file. MSFragger is queried during the analysis to provide a rapid real-time search.

[0107] Preliminary Data: Previously, we described our real-time data analysis and instrument control software, inSeqAPI - a program that enables flexible method generation and rapid incorporation of new types of analysis modules. Multiple groups have demonstrated that a real-time database search and peptide identification improves the depth of single cell proteomics by dedicating more instrument time to the quantification of identifiable peptides. Separately, it has been demonstrated that utilizing fragment intensity and retention time prediction during pot-analysis FDR calculations improves the number of peptides and proteins that can be quantified within an experiment. However, since these prediction algorithms are not available during single cell analysis there is an inherent disconnect between results generated online and those generated post-acquisition.

[0108] Here, we bridge this gap by introducing Salud - a collection of ML models that enable real-time fragmentation and retention time prediction. We combine Salud predictions with a rapid MSFragger search to improve the number of peptides and proteins that can be quantified within a single cell proteomics experiment.

[0109] To date we have constructed the Salud ML models and demonstrated that real-time spectral predictions take 3 ms per spectrum while retention time prediction takes 1.8 ms per spectrum (median values from 27,400 single cell spectra). While these times are relatively short, inSeqAPI will also store previously predicted spectra for use as a spectral library. When using a stored library the fragmentation and retention time prediction filters take only 0.12 ms and 0.014 ms, respectively.

[0110] We combined Salud prediction with MSFragger and within an example run we find a 20% increase in total peptide identifications (7,177 without prediction and 8,604 with prediction), 19% increase in unique peptide identifications (6,964 without prediction and 8,320 with prediction), and 12% increase in protein identifications (1,807 without prediction and 2,026 with prediction).[OHl] Real-time prediction of peptide fragmentation and retention time in combination with real-time MSFragger search improved detection and quantification of peptides for single cell proteomics, implementation of real-time MSFragger search.Example 3 - diaPASEF Analysis for HLA-I Peptides Enables Quantification of Common Cancer Neoantigens

[0112] Sensitive mass spectrometry (MS) methods have greatly advanced the discovery of peptides presented by major histocompatibility complex (MHC) or human leukocyte antigen (HLA) molecules and helped to inform cancer immunotherapies (Yewdell et al. (2022), “MHC Class I Immunopeptidome: Past, Present, and Future”, Mol Cell Proteomics 21(7): 100230; B as sani- Sternberg et al. (2016) “Direct Identification of Clinically Relevant Neoepitopes Presented on Native Human Melanoma Tissue by Mass Spectrometry”, Nature Communication 7: 13404; Yadav et al. (2014) “Predicting Immunogenic Tumour Mutations by Combining Mass Spectrometry and Exome Sequencing”, Nature 515:572-576). However, low abundance, poor recovery yields, and lack of clear digestion rules pose significant challenges for efficient identification of HLA-I peptidome.

[0113] As of today, data-dependent acquisition (DDA) remains a method of choice for MSbased immunopeptidomics. The high quality of MS / MS spectra acquired in DDA mode can be readily translated into high confidence determination of peptide sequences that makes DDA ideal for discovery immunopeptidomics. Nevertheless, DDA inherently suffers from stochastic intensity-based selection of precursor ions that often leads to low sensitivity and low data completeness.

[0114] Data-independent acquisition (DIA), in contrast to DDA, employs fragmentation of all precursor ions that fall into predefined mass windows, generating highly complex fragment spectra. Fragmentation of all available precursor ions without prior selection significantly enhances sensitivity and data completeness of an analysis. In recent years, DIA has proven to be an attractive acquisition method for label free quantification in global proteomics. Analysis of highly complex DIA fragment spectra is a non-trivial task, which traditionally employs time and resource intensive DDA-based spectral libraries. Development and rapid evolution of advanced computational tools enabled an accurate prediction of peptide retention time, ion mobility, and fragment ion intensity (Adams et al. (2023) “Fragment ion intensity prediction improves the identification rate of non- tryptic peptides in TimsTOF”, bioRxiv, 2023.07.17.549401; Bouwmeester et al. (2021) “DeepLC can predict retention times for peptides that carry as-yet unseen modifications”, Nat Methods 18, 1363-1369; Declercq et al. (2022) “MS2Rescore: Data-driven rescoring dramatically boosts immunopeptide identification rates”, Mol Cell Proteomics, 100266; Gessulat et al. (2019) “Prosit: proteome-wide prediction of peptide tandem mass spectra by deep learning”, Nat Methods 16, 509-518; Teschner et al. (2023) “lonmob: a Python package for prediction of peptide collisional crosssection values”, Bioinformatics 39, btad486). This information can then be parsed into in-silico predicted spectral libraries that can serve as an attractive alternative to empirical spectral libraries.

[0115] We have recently shown that ion mobility separation, provided by high field asymmetric waveform ion mobility (FAIMS, Thermo Fisher) (Klaeger et al. (2021) “Optimized Liquid and Gas Phase Fractionation Increases HLA-Peptidome Coverage for Primary Cell and Tissue Samples”, Mol. Cell. Proteomics 20, 100133) or trapped ion mobility spectrometry (TIMS, Bruker Daltonics) (Phulphagar et al. (2023) “Sensitive, High-Throughput HLA-I and HLA-II Immunopeptidomics Using Parallel Accumulation-Serial Fragmentation Mass Spectrometry”, Mol. Cell. Proteomics 22, 100563), helped to increase HLA-I and HLA-II peptide identification starting from more reasonable sample input amounts. Increased sensitivity due to gas phase separation provided by these approaches facilitate detection of sub- stoichiometric post-translational modifications (Oliinyk et al. (2022) “Ion mobility - resolved phosphoproteomics with dia - PASEF and short gradients”, Proteomics, 2200032) or neoantigens.

[0116] Here, we evaluated the quantitative accuracy of DIA for HLA Class I peptides in combination with ion mobility on the timsTOFscp and propose an in silico library approach using a combination of prediction tools. We then used DIA for identification and quantification of previously identified shared neoantigens in an engineered monoallelic cell line model.

[0117] Materials and Methods

[0118] I. Cell culture and SILAC labeling

[0119] A375 cells (multiallelic: A*01:01, A*02:01, B*44:03, B*57:01, C*06:02, C*16:01) were cultured in DMEM supplemented with 10% FBS, 2 mM L- glutamine and 1% Penicillin- Streptomycin (Gibco). For SILAC experiments, A375 cells were passaged at least 7 times in powdered DMEM for SILAC (Thermo Scientific) prepared with either light or heavy amino acid isotopes (84 mg / L L-arginine:HCl [13C6,15N4], 175 mg / L L- lysine:2HCl [13C6,15N2], 100 mg / L L-leucine [13C6]; Cambridge Isotope Laboratories, Inc.), 44 mM sodium bicarbonate, 10% dialyzed FBS (Thermo Scientific), 200 mg / L L-proline, 2 mM L-glutamine,and 1% Penicillin-Streptomycin (Gibco). Cells were grown to 90% confluency, harvested with Trypsin-EDTA, washed three times with ice cold PBS and snap frozen.

[0120] C1R HLA-A*11 :O1 monoallelic line containing 47 common cancer mutations was engineered as previously described (Gurung et al. (2023) “Systematic discovery of neoepitope- HLA pairs for neoantigens shared among patients and tumor types”, Nat. BiotechnoL, 1-11). Briefly, HLA-I null C1R cells were electroporated with a piggyBac neoantigen expression plasmid system and transduced with lentiviral HLA expression vector. Cells were grown in IMDM supplemented with 10% FBS and lug / ml puromycin (Gibco).

[0121] HLA-I Peptide Enrichment and Peptide Elution

[0122] A375 cell pellets were lysed in lysis buffer containing 1% CHAPS, 20mM Tris, pH8.0, 15mM NaCl, 2mM MgC12, 0.2 mM iodoacetamide, ImM EDTA, lx Complete Protease Inhibitor Tablet-EDTA free, 2 pl benzonase, 0.2 mM PMSF in total of 2 ml lysis buffer per 100 million cells. Each lysate was incubated on ice for 30 min and vortexed every 5 min. Lysates were then centrifuged at 20,000 ref for 15 min at 4°C and supernatants were transferred to 1 ml 96 deep-well plates. For neoantigen detection, a 500 million cell pellet of a C1R A* 11 :01 monoallelic cell line containing stably transfected 47-mer cassette without GS linkers was lysed in 5 mL 1% CHAPS lysis buffer and appropriate volume was taken for the respective dilution point for HLA Class I enrichment.

[0123] Immunoprecipitation (IP) of HLA-I peptides and sample preclear was performed on AssayMAP Bravo Sample Prep Platform (Agilent), using Affinity Purification v.4 Protocol. Sample preclear was performed by dispensing lysates through 5 pl Protein-A cartridges (Agilent), previously primed and equilibrated with PBS. For enrichment of HLA-I complexes, flow through was loaded on 5 pl W6 / 32 crosslinked Protein-A cartridges. Cartridges were primed and equilibrated with 100 pl of 20mM Tris pH 8.0 and 150mM NaCl in water. After sample loading, cartridges were washed with 250 pl of 20 mM Tris pH 8.0 and 400mM NaCl in water followed by final wash with 100 pl of 20mM Tris pH 8.0 in water. The antibodybound HLA-peptide complexes were eluted with 60 pl of 0.1M acetic acid in 0.15% trifluoroacetic acid (TFA) and collected to 96 well PCR, Full Skirt, PolyPro plates (Eppendorf). 2 pl of 30% NH40H was added to the eluate to neutralize the pH. Samples were then reduced with 5 mM Dithiothreitol (DTT) at 56°C for 15 min, and alkylated with 10 mM lodoacetamide(IAA) in the dark at RT for 20 min. The samples were then acidified by adding 10% TFA to pH ~ 3.

[0124] Next, samples were loaded on the Assay MAP Bravo platform for Cl 8-based desalting. 5 pl C18 cartridges (Agilent) were primed with 80% acetonitrile (ACN) 0.1% TFA and equilibrated with 0.1% TFA. The samples were afterwards loaded through the cartridges, washed with 0.1% TFA and eluted with 30% ACN in 0.1% TFA and dried by vacuum centrifugation.

[0125] LC-MSMS Analysis

[0126] Purified and desalted peptides were reconstituted in 4 pl of solvent A (0.1% formic acid) and separated via nanoflow reversed-phase liquid chromatography (Vanquish Neo, Thermo) with 60 min or 120 min methods at flow rate of 0.3 pl / min on an Aurora Ultimate nanoflow UHPLC column with CSI fitting (25 cm x 75 pm ID, 1.7 pm C18) (lonOpticks) at 50°C. Biognosys iRT peptides were spiked into each sample for both DDA and DIA acquisitions. For HLA-A* 11 :01 dilution DIA runs 100 firnol of previously targeted neoantigens (Gurung 2023) were spiked into the samples. Mobile phase A was water with 0.1 vol% formic acid and B was ACN with 0.1 vol% formic acid. Peptides eluting from the column were electrosprayed (Captive Spray) into a TIMS quadrupole time-of-flight mass spectrometer (Bruker timsTOF SCP). When operated in dda-PASEF, a ten or three PASEF / MSMS scan per topN acquisition method was utilized. An adapted polygon in the m / z and ion mobility space was employed (Phulphagar et al. (2023) “Sensitive, High-Throughput HLA-I and HLA-II Immunopeptidomics Using Parallel Accumulation-Serial Fragmentation Mass Spectrometry”, Mol. Cell. Proteomics 22, 100563; Gomez-Zepeda et al. (2024) “Thunder-DDA-PASEF enables high-coverage immunopeptidomics and is boosted by MS2Rescore with MS2PIP timsTOF fragmentation prediction model”, Nat. Commun. 15, 2288). The mass spectrometer was operated with an accumulation and ramp time of 166 ms or 300 ms. Precursors were isolated with a 2 Th window below m / z 700 and 3 Th above and actively excluded for 0.4 min with a threshold of 10,000 arbitrary units. A range from 100 to 2000 m / z and 0.65 to 1.67 Vs cm-2 was covered with a collision energy 20 eV at 0.6 Vs cm-2 with a linear increase to 59 eV at 1.6 Vs cm-2. When operated in dia-PASEF mode, an optimized isolation window scheme in the m / z vs ion mobility plane was designed using pyDIAid (Skowronek et al. (2022) “Rapid and In-Depth Coverage of the (Phospho-)Proteome With Deep Libraries and Optimal Window Design for dia-PASEF”, Mol. Cell. Proteomics 21, 100279), covering >99.5% of all precursorions including singly charged species. The method covered precursors within 300-1200 Da (Table 1). The mass spectrometer was operated in sensitivity mode with an accumulation and ramp time of 100 ms and a cycle time of 1.17 s.Table 1. DIA window settings.

[0127] Raw data analysis

[0128] DDA data was analyzed using FragPipe v20 with a nonspecific HLA workflow, peptide length was set to 7-15 mers against a human database (Uniprot 08 / 2023, 104,452 protein sequences) including swissprot and trembl entries and 48 common contaminants. Precursor and fragment mass tolerance were both set to 20 ppm. Carbamidomethylation of cysteines was set as static modification and methionine oxidation, cysteinylation, and pyro-Glu or acetylation of the N terminus were set as variable modification. We allowed a maximum of three variable modifications per peptide. For validation, MSBooster in combination with Percolator was enabled and protein FDR was disabled. A spectral library was built with the spectral library module in FragPipe using standard parameters and retention time alignment to iRT standards (Biognosys).

[0129] DIA data was analyzed using DIA-NN version 1.8.2 beta 11 with standard settings with Mass accuracy = 10 ppm and MSI accuracy = 5 ppm, no match between runs, precursor FDR filtering at 1% searching against either sample specific spectral library, generated by FragPipe, or a predicted library.

[0130] SILAC DIA data was analyzed with DIA-NN version 1.8.2 beta 11 with following search settings: Library Generation was set to “IDs, RT and IM Profiling”, Quantification Strategy was set to “Peak height”, Scan window = 1, Mass accuracy = 10 p.p.m. and MSI accuracy = 5 p.p.m. Additional commands, entered in DIA-NN command window: {—peaktranslation}, {-fixed-mod SILAC, 0.0, LKR, label}, {—lib -fixed-mod SILAC}, {—channels SILAC, L,LKR, 0:0:0; SILAC, H, LKR, 6.020129:8.014199: 10.008269}. Unless otherwise noted, identifications were further filtered for length 8-11, tryptic contamination and contaminants.

[0131] HLA Peptide Binding Prediction

[0132] HLA peptide to allele binding was predicted with HL Apollo vl (Thrift et al. (2024), “Towards designing improved cancer immunotherapy targets with a peptide-MHC-Ipresentation model, HLApollo”, Nat. Commun. 15, 10752). To evaluate MS peptide identifications, identified peptide sequences of length 8-13 were evaluated for binding on the cell lines' respective alleles. Peptides were assigned as binders with an HLApollo score >0 (Thrift et al. (2022) “HLApollo: A superior transformer model for pan-allelic peptide-MHC-I presentation prediction, with diverse negative coverage, deconvolution and protein language features”, bioRxiv 2022, 12.08.519673), peptides predicted to bind to multiple alleles were assigned to every possible allele. For spectral library generation, all possible 8-13mers from the entire human proteome (swissprot: 62,266 protein sequences, 57,172,493 possible 8-13mer peptides including sequence context for 4 residues at the N and C terminus) were generated and predicted for presentation on A375 cells. Peptides were only retained if they would bind to any of the six alleles with a score >0.

[0133] Peptide Retention Time, Fragmentation and Ion Mobility Prediction with Salud and lonmob

[0134] For the predicted A375 spectral library (predicted library), HLA peptides identified as binders by HLApollo were considered for charge states 1-3 per peptide with fixed cysteine alkylation and variable methionine oxidation as possible modifications. Spectral features were predicted using Salud, an in-house collection of ML models based on Prosit (Gessulat et al. (2019), ibid.) to predict retention time and fragmentation. The peptide fragmentation model was based on the Prosit nontryptic HCD model (Wilhelm et al. (2021) “Deep learning boosts sensitivity of mass spectrometry-based immunopeptidomics”, Nat Commun 12, 3346) and the collision energies used for prediction were 15, 20, and 25 for charge states 1, 2, and 3, respectively. For ion mobility, collisional cross sections were predicted using lonmob (Teschner et al. (2023), ibid.) and converted to 1 / Ko values.

[0135] Neoantigen identification and quantification

[0136] In order to systematically identify HLA A* 11 :01 neoantigens, a two-step approach was used to build a spectral library (Al lOl-combo; see Table 2) in Fragpipe v20. First, a spectral library was created from all A* 11 :01 DDA runs using a human proteome fasta file containing the neoantigen sequences along with BFP reporter sequences and iRT sequences (Biognosys). In order to maximize chances of identifying the putative neoantigens, three replicate DDA injections of an equimolar mix of 456 unique synthetic heavy 8-l lmer neoantigens that potentially resulted from 47 common cancer mutations (Gurung et al. (2023)“Systematic discovery of neoepitope- HLA pairs for neoantigens shared among patients and tumor types”, Nat. BiotechnoL, 1-11) were performed as described above using an effective gradient of 51 minutes. Since the synthetic neoantigen pool contained one heavy amino acid per peptide, mass delta for lysine (K: 8.01499), arginine (R: 10.008269), valine / proline (V / R: 6.0138), leucine / isoleucine (L / I: 7.0171), phenylalanine / tryptophan (F / Y: 10.0272), and alanine (A:4.007) along with methionine oxidation were set as variable modifications. We allowed a maximum of two variable modifications per peptide and carbamidomethylation of cysteine residue was set as a static modification. We used a fasta file containing only the neoantigen sequences along with BFP reporter sequences and iRT sequences (Biognosys) to build the spectral library containing the heavy peptide sequences. Ion mobility calibration was performed with default selection of Automatic selection of a run as reference IM and only b and y ions were selected for library generation. Following heavy spectral library generation in Fragpipe, a custom R script was used to add light labeled transitions to the library for every neoantigen peptide sequence. Then, interference correction was performed by removing identical fragment ion pairs from a combined library of heavy and light neoantigens. Finally, for the library used for searching (Al 101-combo), all peptides present in the synthetic peptide library were added to the DDA library and duplicates were removed. The data were analyzed in DIA-NN version 1.8.2 beta 11 at 5% precursor FDR by importing the spectral library created with Fragpipe v20. Quantification strategy was set as Robust LC (high precision), cross-run normalization as RT-dependent, and library generation as smart profiling. All of the neoantigen identifications by DIANN were cross-validated by analyzing them with Skyline (v23.1.0.268) whenever synthetic heavy peptides were available. 100 firnol of each of the previously predicted neoantigens for A* 11 :01 allele (total 57) were spiked into each dilution sample before data collection. The ratio of endogenous light peptide MSI Area signal to spiked in synthetic heavy was then used to back calculate attomole amount for each neoantigen peptide.Table 2: Overview of Spectral Library Characteristics

[0137] Statistical analysis

[0138] All data analysis was performed using custom scripts in python (3.9.4) with packages pandas (1.1.5), numpy (1.22.2), plotly (5.4.0) and scipy (1.7.3) or R (4.3.1) with packages dplyr (1.1.3), data.table (1.14.8), stringr (1.5.1), and ggplot2 (3.4.4) Statistical analysis parameters, if applicable, are described in the respective paragraph and / or figure legend.

[0139] Experimental Design and Statistical Rationale

[0140] Experiments were performed using A375 cells or engineered C1R cell lines (see respective sections). A375 DIA experiments were performed using three technical replicates. We used the same batch of cultured cells for library creation and DIA injections. Statistical analysis parameters, if applicable, are described in the respective paragraph and / or figure legend. For dilution samples, injections were run from low input to high input to reduce carry over effects.

[0141] Results

[0142] HI A peptide analysis with diaPASEF workflow

[0143] The analytical depth and sensitivity of DIA make it attractive for low-yield samples, such as HLA-peptides. For the initial assessment of a diaPASEF immunopeptidomics workflow, we applied our automated immunoaffinity purification protocol to enrich HLA-I peptides from A375 multiallelic melanoma cells (FIG. 5A). In short, we lysed 125 e6A375 cells and enriched HLA-I peptides. Peptides from an equivalent of 75 e6cells were further split into 3 replicates, fractionated, and analyzed in DDA mode. To create an experimental DDA- based spectral library, data was analyzed using Fragpipe resulting in a final library with 20,178 sequences. The remaining amount was injected in technical quadruplicates (12.5 e6cells each) in diaPASEF mode (FIG. 5B; FIGS. HA and 11B) and analyzed by DIA-NN using the sample matched spectral library. This led to identification of 10,297 HLA-I peptides from the equivalent of 12.5 e6cells in just 60 min of LC gradient time (FIGS. 5C-D). Notably, weobserved a high degree of data completeness with >9,500 peptides present in all 4 replicates (FIG. 11C). Interestingly, > 35% of identified peptides are singly charged, which emphasizes the importance of isolation polygon optimization (FIG. 5E). Furthermore, all replicates exhibited similar quantitative behavior with a median R2 of 0.96 and a median coefficient of variation of 12% (FIG. 5F). Based on their length distribution and presence of A375 alleletypical anchor residues, we could further verify that identified peptides are indeed HLA-bound with > 92% of all identified peptides being 8-13-mers, > 70% are predicted to bind to a specific allele, and ~5% could be assigned to multiple alleles in line with observations made for DDA based experiments (FIG. 5G).

[0144] DIA offers high immunopeptidome coverage in half the analysis time

[0145] We then enriched HLA peptides from reduced cell numbers (500,000 - 10e6A375 cells) in triplicate and found that DIA also increased immunopeptidome coverage when starting with fewer cells (FIG. 6A). Notably, DIA identified more precursors than a previously acquired DDA dataset for all cell inputs analyzed using only a 60 min gradient compared to a 120 min gradient for our standard DDA method. Median precursor quantity increased when isolating peptides from a larger number of cells, reflecting the expected abundance changes of HLA ligands (FIG. 12A). Furthermore, DIA allowed the identification of >3,000 precursors in all conditions (FIG. 12B), while only >1,000 peptides were found in all samples with DDA (FIG. 12C), likely due to stochastic intensity-based selection of precursors for MS / MS. For both methods, the majority of peptides are shared between the two higher input conditions (7,388 for DIA, 4702 for DDA). We found that the overlap between DDA and DIA identifications for this experiment is 20-26% for increasing cell amounts (FIG. 6B). Unique DDA identifications were measured with higher intensity values as compared to unique DIA identifications (FIG. 6C). These peptides differ slightly in their assigned HLA allele (FIG. 12D), with a slightly larger fraction of unique DDA identifications matching to HLA A*02:01. This can be partially explained by the m / z space of identifications. DDA acquired more singly charged precursors (800-1200 m / z) with higher intensity (FIG. 12E) than DIA and A*02:01 often produces singly charged precursors. To estimate whether these differences are related to DDA and DIA acquisition at different time points, we performed a similar experiment for peptides isolated from A* 11 :01 monoallelic cells where DDA and DIA runs and the spectral library were generated from the same cell lysate and HLA peptide pool. Again, we found that DIA analysis identified more unique peptides compared to DDA (FIG. 6D). We observed a larger overlapfor DDA and DIA identifications (-30-40%, FIG. 6E) and peptide intensity distributions are more similar (FIG. 6F) when samples are generated and acquired from the same batch.

[0146] Predicted spectral features are a good estimate of empirically determined parameters

[0147] Experiment-specific spectral libraries facilitate a sensitive matching of HLA peptides, acquired with DIA. However, creation of such a library comes at the cost of vast material expenses and typically up to 80% of all available starting material is used to generate a deep experiment-specific spectral library. Alas, often sample specific libraries have a limited coverage due to low quantity of eluted HLA peptides and inherent disadvantages of the DDA approach. Deep learning algorithms for prediction of MS / MS spectra for existing peptide sequences offer an alternative approach for creation of experiment-specific sample libraries. However, due to the non-tryptic nature of HLA peptides, the number of theoretical 8-13-mers from the human proteome exceeds 50 million variants, which complicates directDIA searches.

[0148] We, therefore, compared a library containing predicted spectral features of sequences in the library and the empirically generated spectral library. First, to verify the accuracy of retention time prediction, we inspected the correlation between predicted retention time (RT), calculated by DIA-NN, and measured RT (FIGS. 13Aand 13B). Our predicted library showed a slightly higher deviation in comparison with the empirical library that might point to bigger differences in mean RT prediction between libraries. Indeed, we observed that standard deviation for difference between measured RT and predicted RT in predicted library is ~3-fold higher than in empirical (FIG. 7A). However, the difference of DIA-measured RT apex of identified elution profiles between two libraries was within 3 seconds meaning that both libraries managed to identify the same elution profiles (FIG. 7B). We next asked if the predicted spectral library contains sufficient information for quantifying fragment ions in DIA files. Reassuringly, the predicted library yielded a similar fraction of quantified fragment ions hinting on a good quality of fragment intensity predictions by Salud (FIG. 7C). Moreover, the overall quality of predictions (median spectral angle = 0.68) was on par with the experimentspecific library (spectral angle = 0.7) (FIG. 7D). In concordance with previous studies (Pak et al. (2021), ibid.), we found a drop of -15% in the number of identified HLA peptides with a predicted library in comparison with empirical library (8,424 vs 10,597). More than 85% of peptides, identified with predicted spectral library, were shared with the empirical one.Interestingly, we found -15% (1,062 peptides) to be found only with predicted library while 38% of peptides were unique to the empirical library (FIG. 7E).

[0149] High -sensitivity immunopeptidomics with predicted libraries

[0150] Encouraged by the correlation of empirical and predicted spectral library we generated a two-fold prediction strategy to create a synthetic spectral library for A375 that comprised all possible 8-13-mer binders. First, human canonical proteins were predicted for presentation (343,034,958 peptide-HLA allele combinations) using the state of the art predictor HLApollo (Thrift et al. (2022) “HLApollo: A superior transformer model for pan-allelic peptide-MHC-I presentation prediction, with diverse negative coverage, deconvolution and protein language features”, bioRxiv, 2022.12.08.519673), reducing the possible search space to only 382,537 sequences that could possibly be presented. Retention time, fragment ions, and ion mobility were then predicted for this reduced set of possible precursors using Salud to generate an allele-specific, sample matched spectral library (FIG. 8A). 26% of precursors in the empirically generated library were not contained in this predicted library, the majority of those peptides were not predicted to be binders and hence were removed with this proposed approach (FIG. 14A). While we identified slightly more peptides for lower load samples (5e5- 2.5e6cells) with the empirical library, the predicted library was able to identify more peptides for samples enriched from 5e6- 25e6cells (summarized across triplicates, FIG. 8B). Both approaches resulted in high degree of reproducibility with the median coefficient of variation < 0.2 across all conditions (FIG. 8C). For example, at the 25e6input level we identified > 12,643 peptides with in-silico library versus -8,288 peptides with the DDA-based library (FIG. 8D). While we observed - 35% overlap between libraries, - 42% of peptide IDs found in the empirical library were also identified by the predicted library at 25e6input while at 5e5input -61% of empirical library peptides were also identified with the predicted approach. Interestingly, at an input amount of <5e6cells, we identified more peptides with the empirical library than in the in silico library analysis. However, when comparing motifs of overlapping and unique peptides to either library for the 0.5e6experiment, we found that all peptides show characteristic sequences for A375 presented peptides (FIG. 8E). Upon examination of additional quality criteria such as intensity and spectral angle, peptides identified by the empirical library but not the predicted library show lower intensity and lower spectral angles and might be less confident identifications to begin with (FIGS. 8F; FIGS. 14B-1 and 14B-2;FIGS. 14C-1 and 14C-2)

[0151] Ion mobility reduces co eluting peptides

[0152] Using the predicted spectral features of over 300,000 possibly presented peptides, we further evaluated the ion distribution across DIA windows and LC gradient (FIG. 15A-B, Table 1). Of note, ion mobility already decreased the number of co-eluting precursors per scan from a median of 10 precursors to 1. The majority of predicted precursors fell into windows 1.2 (l / k0:0.4-0.87; m / z: 300.53 mz- 430.23 mz) and window 12.1 (1 / kO: 1.09 - 1.7; m / z: 975.48-1340.62) with >200 precursors eluting in the middle of the LC gradient (FIG. 15C). However, these are deliberately large windows and not many actual features are observed for these mass / ion mobility combinations. All other DIA window and PASEF cycle combinations contain fewer than 50 precursors that are predicted to be analyzed at a given time in the same window. To characterize the extent of co- fragmentation on HLA peptide sequences, we calculated the occurrence of shared peptide fragments in a PASEF cycle window at a given predicted LC elution point. We compared fragment ions in a wide m / z window and found that many of the predicted, coeluting ions share fragment ion masses wouldn’t be uniquely assigned to a particular peptide unless the ions are of longer length (b / y9 or higher, FIGS. 15D-1 to 15D4). Reassuringly, this effect is drastically reduced for windows with fewer co-eluting features (such as DIA window 8.2), where only lower m / z fragment ions (<b3 / y3) would be shared among coeluting precursors and the majority of fragment ions can be uniquely assigned to a single precursor (FIGS. 15D-1 to 15D-4). Overall this analysis shows that the addition of ion mobility significantly reduces the problem of indistinguishable, coeluting HLA peptides in DIA.

[0153] DIA quantification benchmarking with SILAC

[0154] Next, we derived a SILAC HLA peptide DIA strategy to evaluate quantitative accuracy and precision. A375 cells grown in heavy SILAC (K, R, L) and light medium were mixed to a total of 20e6cells each at ratios 1 : 1, 1 :2, 1 :4, 1 :9 and 1 : 19 prior to lysis, HLA peptide enrichment and DIA (FIG. 9A). DIA data was analyzed using an A375 spectral library and DIANN with SILAC specific parameters (Methods). Not all presented peptides contained an amino acid used in the SILAC mix. Therefore, H / L ratios could be calculated for 70-80 % of peptides identified in each experiment. Median H / L ratios for all peptides identified per sample closely matched expected H / L ratios for 1 : 1, 1 :2 and 1 :4 dilutions with only 0-4% relative error (FIG. 9B). For dilutions 1 :9 and 1 : 19, measured ratios had 16% or 46% relative error compared to the expected ratio. This is likely due to identification of lower intensity precursors, mis-identifying light labeled peptides as heavy labeled and vice versa (FIG. 16A). Indeed, when examining the quality of the spectral match, heavy identifications have a greater C-Score distribution than light labeled peptide identifications, with a median of 0.5 C-score for heavy peptides in the 1 : 19 dilution (FIG. 16B). However, while additional filtering for channel and translated q values (Bortegen et al. (2023) “An integrated workflow for quantitative analysis of the newly synthesized proteome”, Nat. Commun. 14, 8237) improved C-scores of remaining peptides (FIG. 16C), it did not affect median ratio or remove peptides with extreme ratios (FIG. 16D) This indicates that peptides are either misidentified or suffer from high signal to noise at extreme ratios and need to be evaluated carefully. Overall, quantitation is accurate up to 5-fold dilution, in line with similar studies on low input samples (Petrosius et al. (2023) “Exploration of cell state heterogeneity using single-cell proteomics through sensitivity- tailored data-independent acquisition”, Nat. Commun. 14, 5910).

[0155] Identification and quantification of common cancer neoantigens

[0156] Next, we evaluated how DIA analysis could benefit neoantigen identification and quantification. We used a mono-allelic C1R cell line expressing HLA-A*l l :01 that was transfected with a neoantigen cassette encoding 47 common cancer mutations (FIG. 10A) (Gurung et al. (2023), ibid.). HLA complexes were enriched from le6- 50e6cells and eluted, followed by DDA for library generation and then DIA. Prior to DIA analysis, 100 fmol of synthetic, heavy standard peptides were spiked into the sample. This allowed for an internal control and helped eliminate false positive identification as spiked in peptides co- elute and cofragment. We observed a steady increase in identifications up to 10e6and 25e6cells. Interestingly, identifications decreased for 50e6cells in DIA runs compared to DDA, likely due to space charging, gradient or library constraints (FIG. 10B). We observed an increase in abundance of neoantigens while the spike- in MSI area stayed relatively constant (FIG. 10C). For example, while a neoantigen from EGFR G719A was low abundant at le6cell input, it could be detected across all input ranges tested with increasing presentation on the surface (FIGS. 10D-1, 10D-2, E). We identified 16 peptide sequences from the neoantigen construct, the majority of sequences from KRAS mutant sequences. As expected, DIA provided a more complete coverage across the dilution series and detected 14 / 16 neoantigens in as low as le6cell input (FIG. 10F). Using the signal of heavy spike-in peptides, we quantified the relative abundance of 13 neoantigens on the surface starting at 33 attomoles for le6cells (FIG. 10G). For comparison, 20 neoantigens were found in a targeted PRM experiment and 9 with DDAfrom 166 e6cells (Gurung et al. (2023), ibid.). This highlights the general utility of DIA acquisition for quantification in immunopeptidomics experiments when following selected peptides of interest across samples.

[0157] Collectively, these results point to high analytical depth and exceptional quantitative reproducibility of DIA in the context of immunopeptidomics.

[0158] Discussion

[0159] Advancements in instrumentation have further increased the application of DIA in numerous proteomic studies. Key advantages include reproducibility between replicates and relative quantification across a large number of samples. For DDA based immunopeptidomic measurements, reproducibility between replicates historically ranges between 60-70%. With DIA we achieved over 99% overlap of identified peptides in replicates. The integration of ion mobility, such as FAIMS or TIMS, decreases the number of co eluting precursors and fewer overlapping or shared fragment ions allow for higher confidence in identifications.

[0160] Though DIA resulted in great depth of coverage with reduced acquisition time of an immunopeptidome sample of interest, additional sample material is needed to create a spectral library. Due to the biological complexity of HLA peptidome samples, library based approaches are necessary to keep FDR low. Other studies have predicted a smaller set of peptides of interest to be included in a spectral library, or generated a large library of available public HLA peptide data on either sequence or spectral level independent of sample context and HLA types present in a sample (Pak et al. (2021), ibid.; Wahle et al. (2024), ibid.). We propose a two-step strategy that first employs state-of-the-art HLA binding predictors to reduce the possible number of peptide sequences in the human proteome and tailor them to the sample of investigation. The second step includes the prediction of spectral features. This combined prediction approach leads to similar or more identifications depending on the amount of peptide loaded on column. Of note, such a predicted library is highly dependent on the prediction algorithm used and might miss novel peptide sequences not included in the binding prediction or sequences predicted as non-binders. It can be further refined by adding contaminant species and non- canonical peptide sequences. Prediction results can also be used to filter peptides further, for example, peptides uniquely identified in experimentally determined spectral libraries are of lower intensity and show poorer spectral angle, indicating that these might be lower confidence hits.

[0161] Using all possible presented precursors, we could also evaluate our data acquisition approach. We found that ion mobility in general reduces the number of co-eluting and cofragmenting peptides and even for highly similar peptide sequences like HLA presented peptides, most DIA spectra contain <50 features with distinct fragment ions. Additional data acquisition approaches further improving window design, as well as novel approaches such as sliceDIA or synchroPASEF (Skowronek et al. (2023) “Synchro-PASEF Allows Precursor- Specific Fragment Ion Extraction and Interference Removal in Data-independent Acquisition”, Mol. Cell. Proteomics 22, 100489; Szyrwiel et al. (2022) “Slice-PASEF: fragmenting all ions for maximum sensitivity in proteomics”, bioRxiv, 2022.10.31.514544), midiaPASEF (Distler et al. (2023) “midiaPASEF maximizes information content in data-independent acquisition proteomics”, bioRxiv, 2023.01.30.526204 ), and narrow window acquisition (Guzman et al. (2024) “Ultra-fast label-free quantification and comprehensive proteome coverage with narrow-window data-independent acquisition”, Nat. BiotechnoL , 1-12) can further decrease co-elution and co-fragmentation leading to confident identifications.

[0162] DIA acquisition is quite powerful for quantification across many samples compared to DDA-based label free or TMT-based quantification. For instance, assessing peptide abundance changes due to varying treatment conditions or input amounts in a consistent sample background is likely executed more efficiently with DIA. But due to the inherent diversity of the HLA peptidome, stratification by respective HLA types is pivotal before conducting such experiments. We evaluated quantitative accuracy using a SIL AC approach and found that DIA analysis can accurately recover known quantities up to a 5-fold difference. Quantification of larger ratios suffers from background noise contamination or incomplete fragmentation and wrong assignment of spectral matches. This is in line with reports in other studies including single cell analyses (Petrosius et al. (20203, ibid.; Pino et al. (2021) “Improved SIL AC Quantification with Data-independent Acquisition to Investigate Bortezomib-Induced Protein Degradation”, J. Proteome Res. 20, 1918-1927). In cell line samples carrying common tumor mutations, we were able to detect neoantigen peptides as low as le6cells input into the IP and quantify increasing amounts of peptides with increasing amounts of cells. If coupled with a hip-MHC approach (Stopfer et al. (2021) “Absolute quantification of tumor antigens using embedded MHC-I isotopologue calibrants”, Proc National Acad Sci 118, e2111173118), ratios of light to heavy peptides can be used to determine the copy number of peptides presented on the surface, while still gathering data on all other peptides present in the sample, and hence notlosing additional peptidome information discarded in PRM acquisitions (Martinez-Vai et al. (2023) “Hybrid-DIA: intelligent data acquisition integrates targeted and discovery proteomics to analyze phospho-signaling in single spheroids”, Nat. Commun. 14, 3599; Goetze et al. (2024) “Simultaneous targeted and discovery-driven clinical proteotyping using hybrid- PRM / DIA”, Clin. Proteom. 21, 26).

[0163] Collectively, the integration of DIA with ion mobility emerges as a potent data acquisition approach. It enhances reproducibility and facilitates effortless quantification of peptides within the same HLA allele type background. Additionally, the library-free strategy we proposed for HLA immunopeptidome measurements can be conveniently constructed using existing resources. This approach, coupled with DIA measurements, ensures robust quantification which can significantly increase throughput, proving advantageous to studies appraising immunotherapy mode of action.

[0164] In summary, gas phase separation in data-dependent MS data acquisition (DDA) increased HLA- 1 peptide detection by up to 50%. We have evaluated the performance of data- independent acquisition (DIA) in combination with ion mobility (diaPASEF) for high- sensitivity identification of HLA presented peptides. A streamlined diaPASEF workflow enabled identification of 11,412 unique peptides from 12.5 million A375 cells and 3,426 8-11 mers from as few as 500,000 cells with high reproducibility. By taking advantage of HLA binder-specific in-silico predicted spectral libraries, we were able to further increase the number of identified HLA-I peptides. We applied SILAC-DIA to a mixture of labeled HLA-I peptides, calculated heavy-to-light ratios for 7,742 peptides across 5 conditions and demonstrated that diaPASEF achieves high quantitative accuracy up to 4-fold dilution. Finally, we identified and quantified shared neoantigens in a monoallelic C1R cell line model. By spiking in heavy synthetic peptides, we verified the identification of the peptide sequences and calculated relative abundances for 13 neoantigens. Taken together, diaPASEF analysis workflows for HLA-I peptides can increase the peptidome coverage for lower sample amounts. The sensitivity and quantitative precision provided by DIA can enable the detection and quantification of less abundant peptide species such as neoantigens across samples from the same background.EXAMPLE EMBODIMENTS

[0165] Embodiments disclosed herein can include:1. A method for identifying MHC peptides present in a sample, the method comprising: predicting, using one or more machine-learning models, an MHC allele-specific mass spectral library for the sample; performing a mass spectrometric analysis of peptides extracted from the sample using data-independent acquisition (DIA) to generate mass spectral data for the peptides; and performing a spectral library search using the mass spectral data for the peptides and the predicted MHC allele-specific mass spectral library for the sample to identify at least one MHC peptide present in the sample.2. The method of embodiment 1, wherein predicting the MHC allele-specific mass spectral library comprises: generating a plurality of candidate peptide sequences from a reference proteome having a sequence length within a specified range of sequence lengths; predicting binding of the candidate peptide sequences of the plurality to at least one MHC allele using a machine learning-based MHC binding prediction model to generate a plurality of candidate MHC peptide sequences; and predicting mass spectral characteristics for each candidate MHC peptide sequence of the plurality using a machine learning-based mass spectra prediction model to generate the predicted MHC allele-specific mass spectral library.3. The method of embodiment 1 or embodiment 2, wherein the mass spectral characteristics for each candidate MHC peptide sequence comprise a peptide retention time, a peptide fragmentation pattern, an ion mobility prediction, or any combination thereof.4. The method of embodiment 2 or embodiment 3, wherein the specified range of sequence lengths is from 8 to 13 amino acid residues.5. The method of any one of embodiments 2 to 4, wherein the machine learning-based MHC binding prediction model is HLApollo.6. The method of any one of embodiments 2 to 5, wherein the machine learning-based mass spectra prediction model is AlphaPeptDeep, directDIA™, DIA-NN, Prosit, or a proprietary mass spectra prediction model.7. The method of any one of embodiments 1 to 6, wherein the at least one MHC peptide present in the sample comprises at least one neoantigen.8. The method of any one of embodiments 1 to 7, further comprising extracting and / or purifying the peptides from the sample.9. The method of any one of embodiments 1 to 8, wherein performing the mass spectrometric analysis of peptides extracted from the sample comprises performing parallel accumulation serial fragmentation (PASEF) in combination with data-independent acquisition (DIA).10. The method of any one of embodiments 1 to 9, wherein the sample is derived from a human subject, and the at least one MHC peptide is an HLA peptide.11. The method of embodiment 10, wherein the human subject is a patient.12. The method of any one of embodiments 1 to 9, wherein the sample is derived from a primate, a mouse, a bacterial culture, a viral culture, or a cultured cell line.13. The method of any one of embodiments 2 to 12, wherein the reference proteome is a human reference proteome, a primate reference proteome, a mouse reference proteome, a bacterial reference proteome, or a viral reference proteome.14. The method of any one of embodiments 1 to 13, wherein the at least one MHC peptide is used in the development of a cancer vaccine or targeted cell therapy.15. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform the method of any one of embodiments 1 to 14.16. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to perform the method of any one of embodiments 1 to 14.

[0166] The description provides preferred example embodiments only, and is not intended to limit the scope, applicability or configuration of the disclosure. Rather, the description of the preferred example embodiments will provide those skilled in the art with an enablingdescription for implementing various embodiments. It is understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope as set forth in the appended claims.

Claims

CLAIMSWhat is claimed is:

1. A method for identifying MHC peptides present in a sample, the method comprising: predicting, using one or more machine-learning models, an MHC allele-specific mass spectral library for the sample; performing a mass spectrometric analysis of peptides extracted from the sample using data-independent acquisition (DIA) to generate mass spectral data for the peptides; and performing a spectral library search using the mass spectral data for the peptides and the predicted MHC allele-specific mass spectral library for the sample to identify at least one MHC peptide present in the sample.

2. The method of claim 1, wherein predicting the MHC allele-specific mass spectral library comprises: generating a plurality of candidate peptide sequences from a reference proteome having a sequence length within a specified range of sequence lengths; predicting binding of the candidate peptide sequences of the plurality to at least one MHC allele using a machine learning-based MHC binding prediction model to generate a plurality of candidate MHC peptide sequences; and predicting mass spectral characteristics for each candidate MHC peptide sequence of the plurality using a machine learning-based mass spectra prediction model to generate the predicted MHC allele-specific mass spectral library.

3. The method of claim 1 or claim 2, wherein the mass spectral characteristics for each candidate MHC peptide sequence comprise a peptide retention time, a peptide fragmentation pattern, an ion mobility prediction, or any combination thereof.

4. The method of claim 2 or claim 3, wherein the specified range of sequence lengths is from 8 to 13 amino acid residues.

5. The method of any one of claims 2 to 4, wherein the machine learning -based MHC binding prediction model is HLApollo.

6. The method of any one of claims 2 to 5, wherein the machine learning-based mass spectra prediction model is AlphaPeptDeep, directDIA™, DIA-NN, Prosit, or a proprietary mass spectra prediction model.

7. The method of any one of claims 1 to 6, wherein the at least one MHC peptide present in the sample comprises at least one neoantigen.

8. The method of any one of claims 1 to 7, further comprising extracting and / or purifying the peptides from the sample.

9. The method of any one of claims 1 to 8, wherein performing the mass spectrometric analysis of peptides extracted from the sample comprises performing parallel accumulation serial fragmentation (PASEF) in combination with data-independent acquisition (DIA).

10. The method of any one of claims 1 to 9, wherein the sample is derived from a human subject, and the at least one MHC peptide is an HLA peptide.

11. The method of claim 10, wherein the human subject is a patient.

12. The method of any one of claims 1 to 9, wherein the sample is derived from a primate, a mouse, a bacterial culture, a viral culture, or a cultured cell line.

13. The method of any one of claims 2 to 12, wherein the reference proteome is a human reference proteome, a primate reference proteome, a mouse reference proteome, a bacterial reference proteome, or a viral reference proteome.

14. The method of any one of claims 1 to 13, wherein the at least one MHC peptide is used in the development of a cancer vaccine or targeted cell therapy.

15. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform the method of any one of claims 1 to 14.

16. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to perform the method of any one of claims 1 to 14.