Data-independent acquisition (dia) for mass spectrometric analysis of mhc peptides

By combining DIA mass spectrometry and PASEF technology with machine learning to generate patient-specific mass spectrometry libraries, the problem of insufficient sensitivity in HLA-I peptide identification has been solved, enabling high-precision detection and quantification of HLA-I peptides and supporting personalized cancer treatment.

CN122641894APending Publication Date: 2026-08-25GENENTECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202580012207.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-07-26
Filing Date
2025-01-31
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing mass spectrometry techniques lack sufficient sensitivity for detecting and quantifying HLA-I peptides, making it difficult to accurately identify low-abundance tumor-specific epitopes, which poses a challenge, especially in personalized cancer treatment.

Method used

We employed data-independent acquisition (DIA) mass spectrometry combined with parallel cumulative serial fragmentation (PASEF) and used machine learning-based methods to generate predicted patient-specific mass spectrometry libraries. We then used machine learning models to predict MHC allele-specific mass spectrometry libraries and combined mass spectrometry features to perform library searches, thereby improving identification accuracy.

Benefits of technology

This enables more accurate and sensitive detection of HLA-I peptides, promotes the development of personalized cancer treatments, and improves the ability to detect and quantify tumor-specific epitopes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122641894A_ABST
    Figure CN122641894A_ABST
Patent Text Reader

Abstract

The present disclosure relates to data independent acquisition (DIA), optionally in combination with ion mobility (diaPASEF), for high-sensitivity identification of MHC-presented peptides in a sample For example of mass spectrometry in some examples, for example, the method can comprise: predicting, using one or more machine learning models, a MHC allele-specific mass spectral library for the sample; performing mass spectrometric analysis of peptides extracted from the sample using data independent acquisition (DIA) to generate mass spectrometric data for the peptides; and performing library searching using the mass spectrometric data for the peptides and the predicted MHC allele-specific mass spectral library for the sample to identify at least one MHC peptide present in the sample.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 548,610, filed February 1, 2024, and U.S. Provisional Patent Application No. 63 / 676,256, filed July 26, 2024, the contents of each of which are incorporated herein by reference in their entirety. Technical Field

[0002] This disclosure generally relates to the mass spectrometry analysis of MHC peptides, and more specifically to the identification of MHC-presenting peptides based on high-sensitivity mass spectrometry using data-independent acquisition (DIA), alone or in combination with ion mobility (diaPASEF). Background Technology

[0003] Major histocompatibility complex (MHC) molecules are cell surface proteins that bind peptide fragments derived from endogenous or exogenous proteins and display these peptide fragments on the cell surface for recognition by T cells in vertebrate organisms. Human leukocyte antigen class I (HLA-I) molecules (human-specific MHC molecules) present short peptide sequences derived from endogenous or exogenous proteins to cytotoxic T cells. The low abundance of HLA-I peptides poses a significant technical challenge to their identification and accurate quantification. Although mass spectrometry (MS) is currently the preferred method for directly identifying the entire cellular immune peptidomome, its sensitivity in detecting and quantifying tumor-specific epitopes still needs improvement. Summary of the Invention

[0004] This article discloses mass spectrometry-based sensitive methods and systems for analyzing peptides that bind to major histocompatibility complex (MHC) proteins. These methods include the use of data-independent acquisition (DIA) mass spectrometry (optionally coupled with parallel cumulative serial fragmentation (PASEF)) to detect and quantify tumor-specific epitopes. In some embodiments, these methods further include using machine learning (ML)-based approaches to generate a predicted (or “synthetic”) patient-specific mass spectrometry library to facilitate the detection and quantification of tumor-specific epitopes.

[0005] These disclosed methods and systems enable more accurate and sensitive detection of MHC peptides (e.g., HLA-I peptides) present in samples, including patient-specific neoantigen peptides present in samples from patients diagnosed with cancer. The improved detection sensitivity, quantification, and accurate identification of MHC peptides (e.g., patient-specific neoantigen peptides) achieved through these disclosed methods can further facilitate the development of personalized cancer therapies (e.g., personalized cancer vaccines).

[0006] This article discloses a method for identifying MHC peptides present in a sample, the method comprising: using one or more machine learning models to predict an MHC allele-specific mass spectrometry library for the sample; performing mass spectrometry analysis on peptides extracted from the sample using data-independent acquisition (DIA) to generate mass spectrometry data of the peptides; and performing a library search using the mass spectrometry data of the peptides and the predicted MHC allele-specific mass spectrometry library for the sample to identify at least one MHC peptide present in the sample.

[0007] In some embodiments, predicting an MHC allele-specific mass spectrometry library includes: generating a plurality of candidate peptide sequences having sequence lengths within a specified range from a reference proteome; using a machine learning-based MHC binding prediction model to predict the binding of candidate peptide sequences among the plurality of candidate peptide sequences to at least one MHC allele, thereby generating a plurality of candidate MHC peptide sequences; and using a machine learning-based mass spectrometry prediction model to predict the mass spectrometry characteristics of each candidate MHC peptide sequence among the plurality of candidate MHC peptide sequences, thereby generating the predicted MHC allele-specific mass spectrometry library.

[0008] In some embodiments, the specified sequence length is in the range of 8 to 13 amino acid residues.

[0009] In some embodiments, the machine learning-based MHC combined prediction model is HLApollo.

[0010] In some embodiments, the machine learning-based mass spectrometry prediction model is AlphaPeptDeep or directDIA. TM DIA-NN, Prosit, or proprietary mass spectrometry prediction models.

[0011] In some embodiments, the mass spectrometry features of each candidate MHC peptide sequence include peptide retention time, peptide fragmentation mode, ion mobility prediction, or any combination thereof.

[0012] In some embodiments, at least one MHC peptide present in the sample contains at least one neoantigen.

[0013] In some embodiments, the method further includes extracting and / or purifying peptides from a sample.

[0014] In some embodiments, mass spectrometry analysis of peptides extracted from a sample includes combining parallel cumulative serial fragmentation (PASEF) with data-independent acquisition (DIA).

[0015] In some embodiments, the sample is derived from a human subject, and at least one MHC peptide is an HLA peptide. In some embodiments, the human subject is a patient. In some embodiments, the sample is derived from a primate, mouse, bacterial culture, viral culture, or cultured cell line.

[0016] In some embodiments, the reference proteome is a human reference proteome, a primate reference proteome, a mouse reference proteome, a bacterial reference proteome, or a viral reference proteome.

[0017] This document also discloses systems comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform any of the methods described herein.

[0018] This document discloses a non-transitory computer-readable storage medium that stores one or more programs, including instructions that, when executed by one or more processors of a system, cause the system to perform any of the methods described herein.

[0019] The terms and expressions used are used as descriptive rather than restrictive terms, and in using such terms and expressions, no equivalent of the features shown and described or any part thereof is intended to be excluded. However, it should be recognized that various modifications are possible within the scope of the claimed invention. Therefore, it should be understood that although the claimed invention has been disclosed through specific exemplary embodiments and optional features, those skilled in the art can adopt modifications and variations of the concepts disclosed herein, and such modifications and variations are considered to be within the scope of the invention as defined by the appended claims.

[0020] Incorporate by reference All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference in their entirety, to the extent that each individual publication, patent, or patent application is specifically and individually indicated as being incorporated herein by reference in its entirety. In the event of any conflict between terminology used herein and terminology found in the incorporated references, the terminology used herein shall prevail. Attached Figure Description

[0021] This patent or application contains at least one color drawing. Upon request and payment of the necessary fees, the official authority will provide a copy of this patent or application publication with one or more color drawings.

[0022] This disclosure is described in conjunction with the accompanying drawings: Figure 1 Non-limiting examples of process flow diagrams for identifying MHC peptides (e.g., HLA-1 peptides) present in a sample according to one embodiment or the disclosed method are provided.

[0023] Figure 2 A non-limiting schematic diagram of an embodiment of the disclosed method for identifying HLA peptides present in a sample is provided.

[0024] Figure 3 Non-limiting examples of block diagrams of computer systems according to some implementations of the methods and systems disclosed herein are provided.

[0025] Figure 4 Non-limiting examples of block diagrams of artificial intelligence (AI) architectures (included as part of an instance spectral library prediction system) based on some implementations of the methods and systems disclosed herein are provided.

[0026] Figures 5A to 5F An overview of the HLA peptide enrichment workflow is provided, along with non-limiting examples of diaPASEF data demonstrating highly reproducible identification of HLA peptides. Figure 5A Overview of the HLA peptide enrichment workflow. Figure 5B : Preparation of sample-specific spectral libraries from HLA peptides. Figure 5C The number of unique peptides identified. Figure 5D : Length distribution of peptide IDs. Figure 5E Charge distribution. Figure 5F The strength of the correlation between two repetitions. Figure 5G : The proportion of observed peptides assigned to alleles in DDA library and DIA peptide identification.

[0027] Figures 6A to 6F Unrestricted examples of data comparing the performance of DIA to that of DDA are provided. DIA outperforms DDA in half the analysis time. Figure 6A HLA peptide identification in DDA and DIA measurements of peptides enriched from A375 cells in decreasing numbers. Figure 6B : Percentage overlap in peptide identification of DDA and DIA for A375. Figure 6CIntensity distributions (overlap) of precursors identified in both DDA and DIA, and intensity distributions of precursors unique to any acquisition type for A375. Figure 6D :and Figure 6A The same caption description applies to HLA-A 11:01 Peptides isolated from the monoallelic C1R cell line. Figure 6E :and Figure 6B The same caption description applies to HLA-A 11:01 Peptides isolated from the monoallelic C1R cell line. Figure 6F :and Figure 6C The same caption description applies to HLA-A 11:01 Peptides isolated from the monoallelic C1R cell line.

[0028] Figures 7A to 7E Non-limiting examples of data for querying diaPASEF data are provided, comparing the use of predictive and empirical spectral libraries for HLA peptides. Figure 7A Comparison of empirically observed retention times with predicted retention times of peptides in the spectral library. Figure 7B The difference in peak RT measured by DIA between the two libraries. Figure 7C : The ratio of quantified fragment ions in the experience library and prediction library to available fragment ions. Symbol This indicates that the p-value is <0.01. Figure 7D Spectral angles identified in the experience and prediction databases. Dashed lines represent median spectral angles. Figure 7E Overlap identified when using an experience library or prediction library.

[0029] Figures 8A to 8F Non-restrictive examples are provided illustrating that predictive libraries for all A375 HLA alleles are generally superior to empirical libraries. Figure 8A A computer simulation library was created by predicting all possible HLA-peptide combinations in A375 cells and using Salud for spectral feature prediction. Figure 8B Identify the number of unique HLA peptides using an empirical library (green / light) or a prediction library (red / dark). Figure 8C : Coefficient of variation for identification using different spectral libraries. Figure 8D 0.5e 6 Cells (top) and 25e 6 Venn diagram of identified peptides in individual cells (bottom). Figure 8E Identification of overlapping peptide sequences between libraries or identification using peptide sequence motifs unique to any one library. Figure 8F For from 0.5e 6 The intensity and spectral distribution of peptides enriched in individual cells, identified by overlap between libraries or by identification unique to any one library.

[0030] Figures 9A to 9B A non-limiting example of using SILAC for quantitative benchmarking of diaPASEF is provided. Figure 9A SILAC labeling strategy for HLA peptide analysis. Figure 9B Violin plot and density plot of log2 H / L ratio across mixing ratio experiments. Box plot shows the median ratio, and dashed lines represent the expected ratio.

[0031] Figures 10A to 10G Provided A 11:01 Non-limiting examples of diaPASEF identification and quantification of neoantigens. Figure 10A HLA-I-deficient C1R cell lines were transfected with a neoantigen cassette containing 47 common cancer mutations and HLA-I alleles of interest, and HLA ligands were eluted from the increasing cell volume. Figure 10B HLA peptide identification in DDA and DIA measurements of peptides enriched from increasing cell mass. Figure 10C The MS1 area of ​​endogenously presented peptides increased, while the area of ​​synthetically spiked peptides remained stable. Figure 10D-1 Endogenous (top) and heavy isotope-labeled neoantigen peptides (1e 6 Extraction ion chromatogram (XIC) of individual cell inputs. Figure 10D-2 Endogenous (top) and heavy isotope-labeled neoantigen peptides (50e 6 Extraction ion chromatogram (XIC) of individual cell inputs. Figure 10E :and Figure 10D-1 and Figure 10D-2 The same, but quantification was performed using DIANN MS1 area. Figure 10F The intensity of neoantigen identification in DDA and DIA collections increases with increasing sample input. Figure 10G In A 11:01 Abundance of all neoantigens identified and quantified by DIA in C1R cells, normalized to 100 fmol of heavy isotope-synthesized spiked peptides.

[0032] Figures 11A to 11C Non-limiting examples of the impact of variable window design on peptide identification are provided. Figure 11A : Variable window design for diaPASEF using HLA class I peptides. Figure 11BThe variable window design enables a uniform feature distribution in each iteration. Figure 11C The number of peptides missing in repeated injections.

[0033] Figures 12A to 12E Non-limiting examples of data for the identification of HLA peptides are provided. Figure 12A Number of precursors in DIA across samples. Figure 12B Upset plot used for DIA identification. The horizontal bar in the lower left corner represents the total number of unique peptides in each sample; the dots and lines below the vertical bar represent peptides identified only in a single sample (single point) or in multiple samples (dots with lines). Figure 12C Regarding DDA identification, and Figure 12B Same description. Figure 12D The proportions of peptides allocated to HLA alleles observed in A375 were isolated according to the identification found in both DDA and DIA or the identification unique to either collection. Figure 12E Hexagonal box plots of precursor m / z distribution and intensity values ​​in 10e6 samples, separated according to identification found in both DDA and DIA or identification unique to either collection.

[0034] Figures 13A to 13B Provides the prediction retention time calculated by DIA-NN (iRT) and the prediction library for Salud ( Figure 13A ) and experience base ( Figure 13B (A non-limiting example of the correlation between measured RTs.)

[0035] Figures 14A to 1 4C provides non-limiting examples of using predictive or empirical spectral libraries to identify precursors. Figure 14A The overlap of peptide sequences in the prediction library and the experience library (left) and the properties of peptides unique to the experience library (right). Figure 14B-1 When using the experience library (alone), low-abundance precursors were preferentially identified (plots of nuclear density versus log2 (intensity) for 0.5e6, 1e6, and 2.5e6 cell inputs). Figure 14B-2 When using the experience library (alone), low-abundance precursors were preferentially identified (plots of nuclear density versus log2 (intensity) for 5e6, 10e6, and 25e6 cell inputs). Figure 14C-1 : Identification overlapping between libraries or spectral angles for identification unique to any library (plots of nuclear density versus spectral angles for inputs of 0.5e6, 1e6, and 2.5e6 cells). Figure 14C-2: Identification overlapping between libraries or spectral angles for identification unique to any library (plots of nuclear density versus spectral angles for inputs of 5e6, 10e6, and 25e6 cells).

[0036] Figures 15A to 1 5D provides a non-limiting instance of data for predicting ion mobility and fragment ion distribution. Figure 15A The reduced ion mobility decreases the theoretical precursors for co-elution. Figure 15B : Predicted ion mobility of the theoretically presented peptide; squares represent the elution window during the collection process. Figure 15C : The theoretical amount of eluted precursors per ion mobility window and per instrument acquisition cycle. Figures 15D-1 to 15D-4 : In the large window that mainly contains single-charge precursors (window 1.2: Figure 15D-1 and Figure 15D-2 ) and the narrow window of the double-charged precursor (window 8.2: Figure 15D-3 and Figure 15D-4 In the sample, the fragment ion distribution of all possible co-eluted predicted precursors (b ions: Figure 15D-1 and Figure 15D-3 γ ions: Figure 15D-2 and Figure 15D-4 The color indicates whether, in a specific window / cycle combination, the fragment ion appears in multiple precursors (blue / dark) or in a unique precursor (orange / light).

[0037] Figures 16A to 16D Non-limiting examples of density maps and C-score data for heavy chain and light chain peptides are provided. Figure 16A : 2D density plot, which shows the log2 H / L ratio of peptides based on the sum of their H+L intensities. Figure 16B C-score distribution for heavy and light peptides for each mixing ratio. Figure 16C :and Figure 16B The same, but after filtering for channel q values ​​< 0.01 and converted q values ​​< 0.01. Figure 16D H / L density based on total intensity across the mixing ratio after filtering for channel q values ​​< 0.01 and converted q values ​​< 0.01. The dashed line represents the expected ratio.

[0038] In the accompanying drawings, similar components and / or features may have the same reference numerals. Furthermore, various parts of the same type can be distinguished by adding a dash after the reference numerals and a second reference numeral to differentiate similar parts. If only the first reference numeral is used in the description, the description applies to any similar parts having the same first reference numeral, regardless of the second reference numeral. Detailed Implementation

[0039] Mass spectrometry-based methods and systems for analyzing peptides that bind to major histocompatibility complex (MHC) proteins are described. These methods include the use of data-independent acquisition (DIA) mass spectrometry (optionally coupled with parallel cumulative serial fragmentation (PASEF)) to detect and quantify tumor-specific epitopes. In some instances, these methods may further include the use of machine learning (ML)-based methods to generate predicted (or “synthetic”) patient-specific mass spectrometry libraries to facilitate the detection and quantification of tumor-specific epitopes.

[0040] The disclosed methods and systems enable more accurate and sensitive detection of MHC peptides (such as HLA-I peptides) present in samples, including patient-specific neoantigen peptides present in samples from patients diagnosed with cancer.

[0041] In some instances, for example, the disclosed methods for identifying MHC (e.g., HLA) peptides present in a sample may include: using one or more machine learning models to predict a sample-specific (e.g., HLA allele-specific) mass spectrometry library; performing mass spectrometry analysis on peptides extracted from the sample using data-independent acquisition (DIA) to generate mass spectrometry data of the peptides; and performing a library search using the peptide mass spectrometry data and the predicted sample-specific (e.g., HLA allele-specific) mass spectrometry library to identify at least one MHC peptide (e.g., an HLA peptide) present in the sample.

[0042] In some instances, predicting an MHC allele-specific mass spectrometry library may include: generating multiple candidate peptide sequences with sequence lengths within a specified range from a reference proteome; using a machine learning-based MHC binding prediction model to predict the binding of candidate peptide sequences among the multiple candidate peptide sequences to at least one MHC allele, thereby generating multiple candidate MHC peptide sequences; and using a machine learning-based mass spectrometry prediction model to predict the mass spectrometry characteristics of each candidate MHC peptide sequence among the multiple candidate MHC peptide sequences, thereby generating the predicted MHC allele-specific mass spectrometry library.

[0043] In some instances, the mass spectrometry features of each candidate MHC peptide sequence include peptide retention time, peptide fragmentation mode, ion mobility prediction, or any combination thereof.

[0044] In some instances, mass spectrometry analysis of peptides extracted from a sample may include combining data-independent acquisition (DIA) with parallel cumulative serial fragmentation (PASEF).

[0045] In some instances, the sample may be derived from a human subject (e.g., a patient), and at least one MHC peptide contains an HLA peptide. In some instances, the sample may be derived from a primate, mouse, bacterial culture, viral culture, or cultured cell line.

[0046] In some instances, the reference proteome can be the human reference proteome, the primate reference proteome, the mouse reference proteome, the bacterial reference proteome, or the viral reference proteome.

[0047] Examples of terms Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0048] As used in this specification and the appended claims, unless the context clearly indicates otherwise, the singular forms “a” and “the” include plural references. Unless otherwise stated, any reference to “or” herein is intended to cover “and / or”.

[0049] "Approximately" and "about" should generally refer to the acceptable degree of error for the quantity being measured, taking into account the nature or precision of the measurement. Examples of acceptable degrees of error are typically within 20 percent, 10 percent, or 5 percent of a given value or range of values.

[0050] As used herein, the terms “comprising” (and any form or variation of “comprising”, such as “comprising multiple” and “comprising one”), “having” (and any form or variation of “having”, such as “having multiple” and “having one”), “including” (and any form or variation of “including”, such as “including multiple” and “including one”), or “containing” (and any form or variation of “containing”, such as “containing multiple” and “containing one”) are non-exhaustive or open-ended and do not exclude additional, unlisted additives, components, integers, elements, or method steps.

[0051] As used herein, the terms “individual,” “patient,” or “subject” are used interchangeably and refer to any single organism, such as a human or non-human mammal (e.g., a dog, cat, horse, cow, pig, sheep, rabbit, mouse, or non-human primate), that requires diagnosis and / or treatment. In a particular embodiment, the individual, patient, or subject herein is a human.

[0052] The terms “cancer” and “tumor” are used interchangeably in this document. These terms refer to cells exhibiting typical characteristics of cancerous cells, such as uncontrolled proliferation, immortality, metastatic potential, rapid growth and proliferation rates, and certain characteristic morphological features. Cancer cells typically exist in the form of tumors, but such cells can exist alone in an animal or can be non-tumorigenic cancer cells, such as leukemia cells. These terms include solid tumors, soft tissue tumors, or metastatic lesions. As used herein, the term “cancer” includes both precancerous lesions and malignant cancers.

[0053] As used herein, “therapeutic” and “treatment” (and their grammatical variations, such as “treatment” or “being treated”) are used interchangeably and refer to a clinical intervention (e.g., administration of an anticancer agent or anticancer therapy) that attempts to alter the natural course of a disease in an individual receiving treatment, and may be performed for preventative purposes or in the course of clinicopathology. The desired effects of treatment include, but are not limited to, preventing the onset or recurrence of disease, alleviating symptoms, attenuating any direct or indirect pathological consequences of the disease, preventing metastasis, slowing the rate of disease progression, improving or alleviating the disease state, and mitigating or improving prognosis.

[0054] As used herein, the term "HLA peptide" refers to a peptide that is predicted or known to bind to HLA proteins (i.e., expressed proteins corresponding to HLA gene alleles).

[0055] The chapter headings used here are for organizational purposes only and should not be construed as limiting the topics described.

[0056] Data-independent acquisition (DIA) for mass spectrometry analysis of MHC peptides Although DIA for immunopeptidomics is not yet widely adopted, a growing number of research teams are exploring this type of acquisition for HLA peptides. For example, DIA has been used for the identification of neoantigens (Pak et al. (2021) "Sensitive Immunopeptidomics by Leveraging Available Large-Scale Multi-HLA Spectral Libraries, Data-Independent Acquisition, and MS / MS Prediction", Mol. Cell. Proteomics 20, 100080) or for the identification of peptides that bind to soluble HLA (sHLA) (Wahle et al. (2024) "IMBAS-MS Discovers Organ-Specific HLA Peptide Patterns in Plasma", Mol. Cell. Proteomics 23, 100689). A major drawback of using DIA in immunopeptidomics applications stems from the lack of well-defined digestion rules for HLA-presenting peptides, making the prediction of spectral library generation a computationally challenging task. Therefore, immunopeptidomics relies heavily on the generation of experimental libraries. This can be disadvantageous when analyzing samples from multiple cell lines or patients with different HLA types, as each of these will present peptides with different binding rules and different amino acid anchoring residues. Furthermore, the non-specific nature of these peptides can complicate data analysis, as co-eluted precursors may lead to erroneous identification.

[0057] To address these shortcomings, we have evaluated the use of machine learning (ML)-based approaches to generate predictive (or “synthetic”) patient-specific mass spectrometry libraries to facilitate the detection and quantification of tumor-specific epitopes using data-independent acquisition (DIA) mass spectrometry (optionally coupled with parallel cumulative serial fragmentation (PASEF)).

[0058] Figure 1A non-limiting example of a flowchart of process 100 for identifying MHC peptides (e.g., HLA-1 peptides) present in a patient sample is provided. Process 100 can be performed as a computer-implemented method, for example, using software running on one or more processors of one or more electronic devices, computers, or computing platforms. In some instances, a client-server system is used to perform process 100, and the boxes of process 100 are divided in any way between server and client devices. In other instances, the boxes of process 100 are divided between a server and multiple client devices. Therefore, although parts of process 100 are described herein as being performed by a specific device of a client-server system, it should be understood that process 100 is not so limited. In other instances, process 100 is performed using only client devices or only multiple client devices. In process 100, some boxes may be combined, some boxes may be rearranged, and some boxes may be omitted. In some instances, additional steps may be performed in combination with process 100. Therefore, the operations shown (and described in more detail below) are exemplary in nature and should not be considered restrictive.

[0059] exist Figure 1 In step 102, one or more (e.g., 1, 2, 3, 4, 5 or more) machine learning models are used to predict the MHC allele-specific (e.g. HLA allele-specific) mass spectrometry library for the sample (e.g., a sample of peptides extracted from a sample such as a patient's tumor specimen).

[0060] In some instances, for example, predicting an MHC allele-specific mass spectrometry library may include: generating multiple candidate peptide sequences from a reference proteome, wherein each candidate peptide has a sequence length within a specified range; using a machine learning-based MHC binding prediction model to predict the binding of each of the multiple candidate peptide sequences to at least one MHC allele (i.e., an expressed protein corresponding to at least one MHC gene allele), to generate multiple candidate MHC peptide sequences; and using a machine learning-based mass spectrometry prediction model to predict the mass spectrometry characteristics of each of the multiple candidate MHC peptide sequences, to generate the predicted MHC allele-specific mass spectrometry library.

[0061] In some instances, for example, the sequence length of a specified range of candidate peptide sequences may be 6 to 13 amino acid residues, 6 to 15 amino acid residues, 8 to 13 amino acid residues, or 8 to 15 amino acid residues, including all possible subranges (e.g., 7 to 12 amino acid residues in length) covered by these exemplary ranges.

[0062] In some instances, the multiple candidate peptide sequences may comprise all possible peptide sequences within a specified length range that can be generated from a reference proteome, such as a human reference proteome, primate reference proteome, mouse reference proteome, bacterial reference proteome, or viral reference proteome. In some instances, the multiple candidate peptide sequences may comprise at least 10 6 5 × 10 6 10 7 5 × 10 7 10 8 5 × 10 8 10 9 or more than 10 9 A candidate peptide sequence. In some instances, the reference genome may contain, for example, the human reference proteome provided in Swiss-PROT (Bairoch et al. (2000), “The SWISS-PROT Protein Sequence Database and its Supplement TrEMBL in 2000”, Nucleic Acids Research 28(1):45-48) or UniProt (The UniProt Consortium, “UniProt: the Universal Protein Knowledgebase in 2023”, Nucleic Acids Research 51(D1):D523-D531).

[0063] In some instances, at least one MHC allele (e.g., an HLA allele) present in a sample can be determined based on MHC typing (e.g., HLA typing; see, for example, Bravo-Egana et al. (2021), “New challenges, new opportunities: Next generation sequencing and its place in the advancement of HLA typing”, Human Immunology 82:478-487), and may include, for example, alleles associated with HLA class I (HLA-I) genes (i.e., HLA-A, HLA-B, and HLA-C genes) or HLA class II (HLA-II) genes (i.e., HLA-DR, HLA-DQ, and HLA-DP genes). In some instances, the binding of each candidate peptide sequence to at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 MHC / HLA alleles (i.e., expressed proteins corresponding to at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 MHC / HLA alleles) is predicted. A candidate peptide predicted to bind to at least one MHC / HLA allele (i.e., expressed protein corresponding to at least one MHC / HLA allele) can be referred to as a candidate MHC (e.g., HLA) peptide.

[0064] In some instances, machine learning-based MHC binding prediction models can be, for example, HLApollo, a pan-allelic, Transformer-based model for predicting peptide presentation of MHC-I (e.g., HLA-I). See, for example, Thrift et al. (2022), “HLApollo: A superior transformer model for pan-allelicpeptide-MHC-I presentation prediction, with diverse negative coverage, deconvolution and protein language features”, bioRxiv, 2022.12.08.519673).

[0065] In some instances, machine learning-based mass spectrometry prediction models can be, for example, AlphaPeptDeep (see, for example, Zeng et al. (2022), “AlphaPeptDeep: a modular deep learning framework to predict peptide properties for proteomics”, Nature Communications 13:7238), directDIA, etc. TM (Biognosys AG, Zurich, CH), DIA-NN (see, for example, Demichev et al. (2020), “DIA-NN: neural networks and interference correction enable deep proteome coverage in high throughput”, Nature Methods 17:41-44), Prosit (see, for example, Gessulat et al. (2019), “Prosit: proteome-wide prediction of peptide tandem mass spectra by deep learning”, Nature Methods 16:509-518) or proprietary mass spectrometry prediction models.

[0066] Examples of mass spectrometry features that can be predicted for each candidate MHC (e.g., HLA) peptide sequence include, but are not limited to, peptide retention time, peptide fragmentation mode, ion mobility prediction, or any combination thereof.

[0067] In some instances, the generation of a predictive mass spectrometry library may include or require the use of additional data.

[0068] exist Figure 1 In step 104, data-independent acquisition (DIA) is used to perform mass spectrometry analysis on the peptides extracted from the sample to generate mass spectrometry data of the peptides.

[0069] In some instances, the sample may be derived from a human subject (e.g., a patient) or from a cultured cell line. In some instances, samples derived from a human subject (e.g., a patient) may include tumor specimens, such as surgically resected samples, tissue biopsy samples, or formalin-fixed paraffin-embedded (FFPE) tissue samples. The peptides can be enzymatically digested and / or extracted from the tumor specimen using any of a variety of techniques known to those skilled in the art. Examples of extraction techniques include, but are not limited to, organic solvent precipitation, differential solubilization, centrifugal ultrafiltration, and solid-phase extraction (see, for example, Peng et al. (2020), “Peptidomic analyses: The progress in enrichment and identification of endogenous peptides”, Trends in Analytical Chemistry 125:115835).

[0070] In some instances, peptide extraction methods can be directly coupled with mass spectrometry analysis, such as through the use of liquid chromatography-mass spectrometry (LC-MS) or liquid chromatography-tandem mass spectrometry (LC-MS-MS) techniques. In some instances, for example, LC-MS or LC-MS-MS techniques may include the use of reversed-phase liquid chromatography (RPLC) techniques.

[0071] In some instances, data-independent acquisition (DIA) can be used to perform mass spectrometry analysis on extracted peptides. DIA is a data acquisition mode that separates and simultaneously fragments populations of different precursors by iteratively traversing predefined m / z ranges of precursors (see, for example, Meier et al. (2020), “diaPASEF: parallel accumulation-serial fragmentation combined with data-independent acquisition”, Nature Methods 17(12):1229-1236).

[0072] In some instances, parallel cumulative serial fragmentation (PASEF) can be coupled with data-independent acquisition (DIA) for mass spectrometry analysis of extracted peptides. DIA is a data acquisition mode that utilizes the correlation between molecular weight and ion mobility in a captured ion mobility spectrometer (e.g., a captured ion mobility spectroscopy (TIMS)-time-of-flight (TOF) mass spectrometer), which samples up to 100% of the peptide precursor ion current in both m / z and mobility windows (see, for example, Meier et al. (2020), ibid.), and improves the reproducibility and quantitative accuracy of peptide identification.

[0073] exist Figure 1 In step 106, a spectral library search is performed using the peptide's mass spectrometry data and a sample-specific (e.g., HLA allele-specific) mass spectrometry library to identify at least one MHC peptide (e.g., HLA peptide) present in the sample.

[0074] In some instances, the presence of at least one MHC peptide (e.g., an HLA peptide) in the sample may include at least 1, at least 10, at least 100, or at least 10 3 At least 10 4 At least 10 5 At least 10 6 One or more 10 6 Each peptide. In some instances, at least one MHC peptide (e.g., an HLA peptide) present in a sample (e.g., a patient sample) may contain at least one neoantigen peptide.

[0075] In some instances, the method may further include extracting and / or purifying peptides from a sample.

[0076] This document also describes a system (and associated computer-readable storage media) configured to execute process 100. Figure 1 As shown, and described in more detail in Appendices I through III below. For example, a system as described herein may include: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform any of the methods described herein. This document also discloses a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions that, when executed by the one or more processors of the system, cause the system to perform any of the methods described herein.

[0077] Figure 2A schematic diagram of one embodiment of the disclosed method for identifying HLA peptides present in a sample is provided. A conventional method for identifying HLA peptides is shown in steps 202 through 212. The bound HLA peptide can be separated from the sample (e.g., a patient sample) using HLA immunoprecipitation, 202. The bound peptide can then be eluted at step 204 (e.g., by changing the assay buffer), and tandem mass spectrometry analysis can be performed at step 206 using a direct data acquisition (DDA) method. Liquid chromatography-tandem mass spectrometry (LC-MS / MS) is a technique widely used in untargeted metabolomics and proteomics studies to identify unknown ions (e.g., ionized metabolites or peptide fragments). Data-dependent acquisition (DDA) selects which ions to fragment for a second MS scan based on the intensity observed in a first full-spectrum MS scan, and typically only fragments a small subset of the present ions. An empirical mass spectrometric library can then be constructed using the data acquired using DDA, 208. However, generating an empirical mass spectrometry library can be extremely cumbersome in terms of sample consumption and required analysis time.

[0078] At step 210, the eluted peptides can also be analyzed by tandem mass spectrometry using a data-independent acquisition (DIA) method. As described above, DIA is a data acquisition mode that separates and simultaneously fragments populations of different precursors by cyclically traversing predefined precursor m / z ranges, and, when coupled with Parallel Accumulated Serial Fragmentation (PASEF), provides significantly improved coverage of ionized peptide fragments and enhances reproducibility and quantification. Then, at step 212, library search and HLA peptide identification can be performed using the empirical library 208 and the DIA mass spectrometry data 210.

[0079] The method disclosed in this paper provides an alternative method for identifying HLA peptides present in a sample, and... Figure 2Steps 202, 204, 210, 214 (including 214a to 214d) and 212 are summarized in the table. Before performing mass spectrometry analysis 210 using a data-independent acquisition (DIA) method, peptides can be isolated from a sample (e.g., a patient sample) using HLA immunoprecipitation 202 and elution 204. However, unlike generating the empirical spectral library 208, the disclosed method can rely on generating a predictive spectral library 214, for example, by identifying all possible 8-12 polymeric peptides in the human proteome at step 214a (i.e., 154,423,819 peptide sequences in this example), identifying all possible pairwise combinations of peptides and HLA proteins present in the patient at step 214b (i.e., 926,542,914 pairwise combinations in this example), predicting which peptides bind to HLA proteins present in the patient using a machine learning-based binding prediction model at step 214c (i.e., 382,537 predicted binding peptides), and predicting the mass spectrometric characteristics of these predicted binding peptides (e.g., peptide retention time, peptide fragmentation pattern, ion mobility prediction, or any combination thereof) using a machine learning-based mass spectrometry feature prediction model at step 214d. Then, at step 212, the predictive spectral library 214 and DIA mass spectrometry data 210 can be used to perform a library search and HLA peptide identification. In some instances, the predicted mass spectrum library 214 can be used alone or in combination with the empirically generated mass spectrum library 208 to identify HLA peptides during a library search.

[0080] Example computer system Figure 3 Non-limiting examples of block diagrams of computer systems based on some implementations of the disclosed systems and methods are provided.

[0081] Computer system 300 can be a host computer connected to a network. Computer system 300 can be a client computer or a server. For example... Figure 3 As shown, computer system 300 can be any suitable type of microprocessor-based device, such as a personal computer, workstation, server, or handheld computing device (portable electronic device), such as a telephone or tablet. The device may include, for example, one or more of processor 310, input device 320, output device 330, storage 340, and communication device 360. Input device 320 and output device 330 may generally correspond to devices described elsewhere herein, and they may be connectable or integrated with a computer.

[0082] Input device 320 can be any suitable device that provides input, such as a touchscreen, keyboard or keypad, mouse, or voice recognition device. Output device 330 can be any suitable device that provides output, such as a touchscreen, haptic feedback device, or speaker.

[0083] Storage 340 can be any suitable device providing storage, such as electrical memory, magnetic memory, or optical memory including RAM, cache, hard disk drive, or removable storage disk. Communication device 360 ​​can include any suitable device capable of transmitting and receiving signals over a network, such as a network interface chip or device. Computer components can be connected in any suitable manner, such as via physical bus 370 or wireless connection.

[0084] The software 350, which can be stored in memory / storage 340 and executed by processor 310, may include, for example, programming that embodies the functionality of this disclosure (e.g., as embodied in the methods described above).

[0085] Software 350 may also be stored and / or transferred to any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, device, or apparatus (such as those described above), from which the instruction execution system, device, or apparatus may retrieve and execute instructions associated with the software. In the context of this disclosure, the computer-readable storage medium may be any medium (such as storage 340) that may contain or store programming for use by or in connection with an instruction execution system, device, or apparatus.

[0086] Software 350 can also be propagated within any transmission medium for use by or in conjunction with an instruction execution system, device, or apparatus (such as those described above), from which the instruction execution system, device, or apparatus can retrieve and execute instructions associated with the software. In the context of this disclosure, the transmission medium can be any medium capable of communicating, propagating, or transmitting programming for use by or in conjunction with an instruction execution system, device, or apparatus. Transmission readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, or infrared wired or wireless propagation media.

[0087] Computer system 300 can be connected to a network, which can be any suitable type of interconnected communication system. The network can implement any suitable communication protocol and can be protected by any suitable security protocol. The network can include any suitable arrangement of network links that enables the transmission and reception of network signals, such as wireless network connections, T1 or T3 lines, cable networks, DSL, or telephone lines.

[0088] Computer system 300 can implement any operating system suitable for operation on a network. Software 350 can be written in any suitable programming language, such as C, C++, Java, or Python. In various instances, application software embodying the functionality of this disclosure can be deployed in different configurations, such as, for example, in a client / server setup, or as a web-based application or web service via a web browser.

[0089] Figure 4 Figure 400 illustrates the example artificial intelligence (AI) architecture 402 (this architecture can be used as a reference in the above discussion). Figure 3 (Including a portion of the one or more computing devices 300 discussed herein), the architecture can be used to implement the methods described herein. In some instances, the AI ​​architecture 402 can be implemented using, for example, one or more processing devices, which may include hardware (e.g., general-purpose processors, graphics processing units (GPUs), application-specific integrated circuits (ASICs), system-on-a-chip (SoCs), microcontrollers, field-programmable gate arrays (FPGAs), central processing units (CPUs), application processors (APs), vision processing units (VPUs), neural processing units (NPUs), neural decision processors (NDPs), deep learning processors (DLPs), tensor processing units (TPUs), neuromorphic processing units (NPUs), and / or other processing devices that may be adapted to process various molecular data and make one or more decisions thereon), software (e.g., instructions that run / execute on one or more processing devices), firmware (e.g., microcode), or some combination thereof.

[0090] In some instances, such as Figure 4 The AI ​​architecture 402 described may include machine learning (ML) algorithms and functions 404, natural language processing (NLP) algorithms and functions 406, expert systems 408, computer-based vision algorithms and functions 410, speech recognition algorithms and functions 412, planning algorithms and functions 414, and robotics algorithms and functions 416. In some instances, ML algorithms and functions 404 may include any statistically based algorithms applicable to finding patterns in large amounts of data (e.g., “big data,” such as genomic data, proteomic data, metabolomic data, metagenomic data, transcriptomic data, or other omics data). For example, in some instances, ML algorithms and functions 404 may include deep learning algorithms 418, supervised learning algorithms 420, and unsupervised learning algorithms 422.

[0091] In some instances, deep learning algorithm 418 may include any artificial neural network (ANN) that can be used to learn deep representations and abstractions from large amounts of data. For example, deep learning algorithm 418 may include ANNs such as perceptrons, multilayer perceptrons (MLPs), autoencoders (AEs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory (LSTMs), gated recurrent units (GRUs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs) and deep Q-networks, neural autoregressive distribution estimation (NADEs), adversarial networks (ANs), attention models (AMs), spiking neural networks (SNNs), deep reinforcement learning, etc.

[0092] In some instances, supervised learning algorithm 420 can include any algorithm that can be used to apply, for example, previously learned content to new data using labeled examples to predict future events. For example, starting with the analysis of a known training dataset, supervised learning algorithm 420 can produce an inference function to predict output values. Supervised learning algorithm 420 can also compare its output with correct and expected outputs and identify errors so that it can be modified accordingly. On the other hand, unsupervised learning algorithm 422 can include, for example, any algorithm that can be applied when the data used to train unsupervised learning algorithm 422 is neither classified nor labeled. For example, unsupervised learning algorithm 422 can study and analyze how the system infers functions describing the hidden structure from unlabeled data.

[0093] In some instances, NLP algorithms and functions 406 may include any algorithms or functions applicable to the automatic processing of natural language, such as speech and / or text. For example, NLP algorithms and functions 406 may include content extraction algorithms or functions 424, classification algorithms or functions 426, machine translation algorithms or functions 428, question answering (QA) algorithms or functions 430, and text generation algorithms or functions 432. In some instances, content extraction algorithms or functions 424 may include means for extracting text or images from electronic documents (e.g., web pages, text editor documents, etc.) for use, for example, in other applications.

[0094] In some instances, classification algorithm or function 426 may include any algorithm that can learn from data input to a supervised learning model using a supervised learning model (e.g., logistic regression, Naive Bayes, stochastic gradient descent (SGD), k-nearest neighbors, decision trees, random forests, support vector machines (SVMs), etc.) and make new observations or classifications accordingly. Machine translation algorithm or function 428 may include any algorithm or function that can be adapted to automatically convert source text in one language into text in another language. QA algorithm or function 2130 may include any algorithm or function that can be adapted to automatically answer questions posed by humans in natural language, such as algorithms or functions performed by voice-controlled personal assistant devices. Text generation algorithm or function 432 may include any algorithm or function that can be adapted to automatically generate natural language text.

[0095] In some instances, expert system 408 may include any algorithms or functions applicable to simulating the judgment and behavior of a person or organization with expertise and experience in a specific domain (e.g., stock trading, medicine, sports statistics, etc.). Computer-based vision algorithms and functions 410 may include any algorithms or functions applicable to automatically extracting information from images (e.g., photographic images, video images). For example, computer-based vision algorithms and functions 410 may include image recognition algorithm 434 and machine vision algorithm 436. Image recognition algorithm 434 may include any algorithms applicable to automatically identifying and / or classifying objects, locations, people, etc., that may be included in one or more image frames or other display data. Machine vision algorithm 436 may include any algorithms applicable to allowing a computer to "see," or, for example, relying on an image sensor camera with specialized optics to acquire images for processing, analyzing, and / or measuring various data features for decision-making purposes.

[0096] In some instances, speech recognition algorithms and functions 412 may include any algorithms or functions applicable to recognizing spoken language and translating it into text, such as through automatic speech recognition (ASR), computer speech recognition, speech-to-text (STT) 438, or text-to-speech (TTS) 440, for example, to enable a computing device to communicate with one or more users via voice. In some instances, planning algorithms and functions 414 may include any algorithms or functions applicable to generating sequences of actions, where each action may include a set of prerequisites that must be satisfied before it is performed. Examples of AI planning may include classical planning, simplified to other problems, time planning, probabilistic planning, preference-based planning, conditional planning, etc. Finally, robotics algorithms and functions 416 may include any algorithm, function, or system that enables one or more devices to replicate human behavior through, for example, movement, posture, performance of tasks, decision-making, emotions, etc.

[0097] Example Example 1 - Real-time prediction and real-time search of fragmentation and retention times improve multiplex immunopeptidomics quantification Introduction: Immunopeptidomics analyzes peptides presented on cell surfaces, whether derived from endogenous or exogenous proteins, and is a crucial component of cancer vaccine and targeted cell therapy development. Recent advances in machine learning (ML) prediction of peptide fragmentation and retention times have significantly improved the identification rate of non-trypsin peptides, enabling deeper and more reliable analysis of the immunopeptidome. Here, we introduce Salud, a set of specially developed ML models for real-time prediction of fragmentation and retention times during mass spectrometry analysis. We combine Salud with a fast, real-time MSFragger search within the instrument application programming interface (iAPI) program inSeqAPI to enable real-time re-scoring of spectra and improve the confidence level of identification and the depth of multiplex quantification in immunopeptidomics analysis.

[0098] Method: Salud and RT-MSFragger are integrated into inSeqAPI by creating a custom C# "APIFilter" class object. Similar to Prosit, Salud is built on an encoder / decoder architecture but lacks components that prevent conversion to a common model format that can be called from inSeqAPI. During real-time analysis, the "SaludFragmentFilter" predicts the theoretical spectrum for a given peptide, while the "SaludRTFilter" predicts the normalized retention time. These two values ​​are then used within a real-time support vector machine (SVM) false discovery rate filter, which performs real-time peptide FDR. For "MSFragger," inSeqAPI leverages OpenJDK to launch the MSFragger service based on a predefined parameter file. MSFragger is then queried during analysis to provide fast searching of a large non-trypsin database.

[0099] Preliminary data: We attempted to improve the depth of multiplex immunopeptidome quantification by refining the real-time search method within inSeqAPI. Due to the size of the database and the correspondingly long search time for each spectrum, real-time search has not yet been applied to non-trypsin samples. Furthermore, the large database may lead to reduced identification, a phenomenon that can be mitigated by machine learning (ML) prediction algorithms that allow for re-scoring of spectra after fragmentation and retention time prediction.

[0100] Here, we describe Salud, a set of ML models capable of real-time fragmentation and retention time prediction within inSeqAPI. We benchmarked Salud to determine the amount of time required to perform real-time predictions. For a single peptide, we found that Salud fragmentation prediction was performed within approximately 3 ms (median of 27,400 predictions), and retention time prediction within approximately 1.8 ms (median of 27,400 predictions). Furthermore, using these ML models improved the identification of HLA peptides by 25%–45% during FDR assays, similar to other described ML models.

[0101] To address the challenge of time-consuming real-time non-trypsin searches, we introduce Real-Time MSFragger (RT-MSFragger), an implementation of the MSFragger algorithm that runs as a service and accepts spectra from C# code in the inSeqAPI in real time. We compare the search times of Comet and MSFragger on a non-trypsin database of 7-12 polymeric peptides (i.e., the common length of peptides identified in immunopeptidomics analysis). Here, we find that MSFragger reduces the search time by nearly 80 times compared to Comet (MSFragger is 0.55 ms, while Comet is 43 ms). Finally, we demonstrate RT-MSFragger's ability to search for 7-25 polymeric non-trypsin peptides, a real-time search capability that the current version of Comet cannot perform.

[0102] The results show that real-time execution of machine learning algorithms for fragmentation and retention time prediction improves identification confidence. RT-MSFragger search significantly reduces the search time required for non-trypsin peptides.

[0103] Future work will benchmark improvements in the quantification of peptides in the analysis of peptidomics and immune peptidomics from cell lines expressing mutant regions derived from 47 common cancer mutations.

[0104] Example 2 - Improving multivariate single-cell proteomics by combining real-time fragmentation and retention time prediction with real-time search Preface: Recent advances in data acquisition have enhanced the depth of multiplex single-cell proteomics (SCP). In particular, the use of real-time searches dedicates more instrument time to the quantification of peptides identified in post-run analyses. Furthermore, the use of machine learning algorithms improves peptide identification in challenging samples by utilizing predicted fragmentation intensity and retention time to aid in the determination of false discovery rate (FDR). Here, we describe two new features for inSeqAPI that increase the number of peptides and proteins quantified in SCP experiments: 1) Salud—a set of ML algorithms that can be executed in real-time to enable re-scoring during real-time FDR; and 2) Real-time MSFragger, which performs real-time searches faster than current standards.

[0105] Method: Salud is a set of ML algorithms with an encoder / decoder architecture similar to Prosit, but lacks a specific layer to prevent conversion to platform-independent files. RT-MSFragger is written in Java and enables real-time searching by creating a service that can be queried from C# code. During real-time inSeqAPI analysis, the "SaludFragmentFilter" predicts the theoretical spectrum for a given peptide, while the "SaludRTFilter" predicts the normalized retention time. These two values ​​are then used within a real-time support vector machine (SVM) false discovery rate filter, which performs real-time peptide FDR. For "MSFragger," inSeqAPI leverages OpenJDK to launch the MSFragger service based on a predefined parameter file. MSFragger is queried during analysis to provide fast real-time searching.

[0106] Preliminary Data: Previously, we described our real-time data analysis and instrumentation control software, inSeqAPI—a program that enables flexible method generation and rapid integration of novel analytical modules. Multiple research teams have demonstrated that real-time database searching and peptide identification enhance the depth of single-cell proteomics by dedicating more instrumentation time to the quantification of identifiable peptides. Furthermore, it has been shown that using fragment intensity and retention time predictions during post-analysis FDR calculations increases the number of quantifiable peptides and proteins in experiments. However, because these prediction algorithms are not available during single-cell analysis, there is an inherent disconnect between online-generated results and those generated after acquisition.

[0107] Here, we bridge this gap by introducing Salud, a suite of ML models that support real-time fragmentation and retention time prediction. We combine Salud predictions with fast MSFragger search to increase the number of peptides and proteins that can be quantified in single-cell proteomics experiments.

[0108] To date, we have built the Salud ML model and demonstrated that real-time spectral prediction takes 3 ms per spectrum, while retention time prediction takes 1.8 ms per spectrum (median value obtained from 27,400 single-cell spectra). While these times are relatively short, inSeqAPI also stores previously predicted spectra for use as a spectral library. When using the stored library, fragmentation and retention time prediction filters take only 0.12 ms and 0.014 ms, respectively.

[0109] We combined Salud prediction with MSFragger, and within a single example run, we found a 20% increase in total peptide identification (7,177 without prediction, 8,604 with prediction), a 19% increase in unique peptide identification (6,964 without prediction, 8,320 with prediction), and a 12% increase in protein identification (1,807 without prediction, 2,026 with prediction).

[0110] The combination of real-time prediction of peptide fragmentation and retention time with real-time MSFragger search improves the detection and quantification of peptides for single-cell proteomics, as well as the implementation of real-time MSFragger search.

[0111] Example 3 - diaPASEF analysis for HLA-I peptides enables the quantification of common cancer neoantigens. Sensitive mass spectrometry (MS) methods have greatly facilitated the discovery of peptides presented by major histocompatibility complex (MHC) or human leukocyte antigen (HLA) molecules and have helped to provide important information for cancer immunotherapy (Yewdell et al. (2022), “MHC Class I Immunopeptidome: Past, Present, and Future”, Mol CellProteomics21 (7): 100230; Bassani-Sternberg et al. (2016) “Direct Identification of Clinically Relevant Neoepitopes Presented on Native Human Melanoma Tissue by Mass Spectrometry”, Nature Communication7: 13404; Yadav et al. (2014) “Predicting Immunogenic Tumour Mutations by Combining Mass Spectrometry and Exome Sequencing”, Nature515: 572-576). However, low abundance, poor recovery, and lack of clear digestion rules pose significant challenges to the efficient identification of HLA-I peptide groups.

[0112] To date, data-dependent acquisition (DDA) remains the preferred method for MS-based immunopeptidomics. The high-quality MS / MS spectra acquired in DDA mode can be readily converted into high-confidence peptide sequencing, making DDA ideal for immunopeptidomics discovery. However, DDA is inherently affected by intensity-based random precursor ion selection, which typically leads to low sensitivity and low data integrity.

[0113] In contrast to DDA, data-independent acquisition (DIA) fragments all precursor ions falling within a predefined quality window, generating highly complex fragment spectra. The ability to fragment all available precursor ions without pre-selection significantly improves analytical sensitivity and data integrity. In recent years, DIA has proven to be a highly attractive acquisition method for label-free quantification in global proteomics. The analysis of highly complex DIA fragment spectra is a non-trivial task traditionally undertaken using time-consuming and resource-intensive DDA-based spectral libraries. The development and rapid evolution of advanced computational tools have enabled accurate prediction of peptide retention time, ion mobility, and fragment ion intensity (Adams et al. (2023) "Fragment ion intensity prediction improves the identification rate of non-tryptic peptides in TimsTOF", bioRxiv, 2023.07.17.549401; Bouwmeester et al. (2021) "DeepLC can predict retention times for peptides that carry as-yetunseen modifications", Nat Methods 18, 1363–1369; Declercq et al. (2022) "MS2Rescore: Data-driven rescoring dramatically boosts immunopeptide identification rates", Mol Cell Proteomics, 100266; Gessulat et al. (2019) "Prosit: proteome-wide prediction of peptide tandem mass spectra by deep learning", Nat Methods 16, 509–518; Teschner et al. (2023) "Ionmob: a Python package for prediction of peptide collisional cross-section values", Bioinformatics 39, btad486. This information can then be parsed into a computer simulation prediction spectral library, which can serve as a highly attractive alternative to empirical spectral libraries.

[0114] We have recently demonstrated that ion mobility separation provided by high-field asymmetric waveform ion mobility spectrometry (FAIMS, Thermo Fisher Scientific) (Klaeger et al. (2021) “Optimized Liquid and Gas Phase Fractionation Increases HLA-Peptidome Coverage for Primary Cell and Tissue Samples”, Mol. Cell. Proteomics 20, 100133) or trapped ion mobility spectrometry (TIMS, Bruker Daltonics) (Phulphagar et al. (2023) “Sensitive, High-Throughput HLA-I and HLA-II Immunopeptidomics Using Parallel Accumulation-Serial Fragmentation Mass Spectrometry”, Mol. Cell. Proteomics 22, 100563) can help increase the identification of HLA-I and HLA-II peptides starting from a more reasonable sample input starting amount. The improved sensitivity provided by gas-phase separation by these methods has facilitated the detection of substoichiometric post-translational modifications (Oliinyk et al. (2022) "Ion mobility‐resolved phosphoproteomics with dia‐PASEF and short gradients", Proteomics, 2200032) or neoantigens.

[0115] Here, we evaluate the accuracy of DIA combined with ion mobility for the quantification of HLA class I peptides on timsTOFscp and present a computer-simulated spectral library approach using a combination of predictive tools. We then use DIA to identify and quantify previously identified shared neoantigens in engineered monoallelic cell line models.

[0116] Materials and methods I. Cell Culture and SILAC Labeling A375 cells (multiple alleles: A) 01:01, A 02:01, B 44:03, B 57:01, C 06:02, C 16:01) Cells were cultured in DMEM supplemented with 10% FBS, 2 mM L-glutamine, and 1% penicillin-streptomycin (Gibco). For SILAC experiments, A375 cells were passaged at least 7 times in powdered DMEM (Thermo Fisher Scientific) used for SILAC, prepared using any light or heavy amino acid isotope (84 mg / L L-arginine: HCl [13C6,15N4], 175 mg / L L-lysine: 2HCl [13C6,15N2], 100 mg / L L-leucine [13C6]; Cambridge Isotope Laboratories, Inc.), 44 mM sodium bicarbonate, 10% dialyzed FBS (Thermo Fisher Scientific), 200 mg / L L-proline, 2 mM L-glutamine, and 1% penicillin-streptomycin (Gibco). Cells were cultured to 90% confluence, harvested using trypsin-EDTA, washed three times with ice-cold PBS, and then flash-frozen.

[0117] As mentioned earlier, C1R HLA-A containing 47 common cancer mutations... 11:01 Engineering of monoallelic cell lines (Gurung et al. (2023) “Systematic discovery of neoepitope–HLA pairs for neoantigens shared among patients and tumor types”, Nat. Biotechnol., 1–11). Briefly, HLA-I-deficient C1R cells were electroporated using the piggyBac neoantigen expression plasmid system and transduced using a lentiviral HLA expression vector. Cells were cultured in IMDM supplemented with 10% FBS and 1 μg / ml puromycin (Gibco).

[0118] HLA-I peptide enrichment and peptide elution A375 cell pellets were lysed in lysis buffer containing 1% CHAPS, 20 mM Tris (pH 8.0), 15 mM NaCl, 2 mM MgCl2, 0.2 mM iodoacetamide, 1 mM EDTA, 1x EDTA-free complete protease inhibitor tablet, 2 μl benzoxoside, and 0.2 mM PMSF, for a total of 2 ml of lysis buffer per 100 million cells. Each lysate was incubated on ice for 30 min, vortexed every 5 min. The lysate was then centrifuged at 20,000 rcf for 15 min at 4 °C, and the supernatant was transferred to 1 ml of 96-well plates. For neoantigen detection, C1RA containing a stable transfected 47-mer cassette without GS adapter was used. At 11:01, 500 million cells from a monoallelic cell line were lysed in 5 mL of 1% CHAPS lysis buffer, and an appropriate volume was taken for the corresponding dilution point for HLA class I enrichment.

[0119] Immunoprecipitation (IP) of HLA-I peptides and sample pre-cleaning were performed on the AssayMAP Bravo sample preparation platform (Agilent) using the affinity purification v.4 protocol. Sample pre-cleaning was performed by passing the lysate through a 5 μl Protein-A column (Agilent), which was pre-wetted and equilibrated with PBS. To enrich the HLA-I complex, the eluent was loaded onto a 5 μl W6 / 32 crosslinked Protein-A column. These columns were wetted and equilibrated with 100 μl of an aqueous solution containing 20 mM Tris pH 8.0 and 150 mM NaCl. After sample loading, the columns were washed with 250 μl of an aqueous solution containing 20 mM Tris pH 8.0 and 400 mM NaCl, followed by a final wash with 100 μl of an aqueous solution containing 20 mM Tris pH 8.0. The HLA-peptide complex bound to the antibody was eluted with 60 μl of 0.1 M acetic acid containing 0.15% trifluoroacetic acid (TFA) and collected into a 96-well Eppendorf PCR plate. 2 μl of 30% NH4OH was added to the eluent to neutralize the pH. The samples were then reduced with 5 mM dithiothreitol (DTT) at 56 °C for 15 min and alkylated with 10 mM iodoacetamide (IAA) at RT in the dark for 20 min. The samples were then acidified to approximately pH 3 by adding 10% TFA.

[0120] Next, the samples were loaded onto the AssayMAP Bravo platform for C18-based desalting. Five μl of an Agilent C18 column was wetted with 80% acetonitrile (ACN) containing 0.1% TFA, and the columns were equilibrated with 0.1% TFA. The samples were then loaded through these columns, washed with 0.1% TFA, eluted with 30% ACN containing 0.1% TFA, and dried by vacuum centrifugation.

[0121] LC-MS / MS analysis The purified and desalted peptides were resuspended in 4 μl of solvent A (0.1% formic acid) and separated at 50 °C using a CSI-connected Aurora Ultimate nanoflow UHPLC column (25 cm × 75 μm ID, 1.7 μm C18) (IonOpticks) via nanoflow reversed-phase liquid chromatography (VanquishNeo, Thermo) at a flow rate of 0.3 μl / min for 60 min or 120 min. Biognosys iRT peptides were spiked into each sample for both DDA and DIA acquisitions. For HLA-A... At 11:01, the DIA was diluted, and these samples were spiked with 100 fmol of the previously targeted neoantigen (Gurung 2023). Mobile phase A was water containing 0.1 vol% formic acid, and mobile phase B was ACN containing 0.1 vol% formic acid. Peptides eluted from the column were electrosprayed into a TIMS quadrupole time-of-flight mass spectrometer (Bruker timsTOF SCP). When operating in dda-PASEF mode, an acquisition method of ten or three PASEF / MSMS scans per topN was used. Adjusted polygons were used in the m / z and ion mobility spaces (Phulphagar et al. (2023) "Sensitive, High-Throughput HLA-I and HLA-II Immunopeptidomics Using Parallel Accumulation-Serial Fragmentation Mass Spectrometry", Mol. Cell. Proteomics 22, 100563; Gomez-Zepeda et al. (2024) "Thunder-DDA-PASEF enables high-coverage immunoopeptidomics and is boosted by MS2Rescore with MS2PIP timsTOFfragmentation prediction model", Nat. Commun. 15, 2288). The mass spectrometer was operated at accumulation and ramp times of 166 ms or 300 ms. The isolation window for precursors below m / z 700 is 2 Th, and the isolation window for precursors above m / z 700 is 3 Th, with active exclusion for 0.4 min at an arbitrary threshold of 10,000 units. It covers the range from 100 m / z to 2000 m / z and from 0.65 Vs cm⁻² to 1.67 Vs cm⁻², with a collision energy of 20 eV at 0.6 Vs cm⁻², linearly increasing to 59 eV at 1.6 Vs cm⁻².When operating in dia-PASEF mode, an optimized isolation window scheme with m / z relative to the ion mobility plane was designed using pyDIAid (Skowronek et al. (2022) “Rapid and In-Depth Coverage of the (Phospho-)Proteome With DeepLibraries and Optimal Window Design for dia-PASEF”, Mol. Cell. Proteomics 21, 100279), covering >99.5% of all precursor ions (including single-charged species). This method covered precursors in the 300–1200 Da range (Table 1). The mass spectrometer was operated in sensitivity mode with a cumulative time and ramp time of 100 ms and a cycle time of 1.17 s.

[0122] Table 1. DIA Window Settings.

[0123] Raw data analysis DDA data were analyzed using FragPipe v20 and a nonspecific HLA workflow, targeting a human database (Uniprot 08 / 2023, 104,452 protein sequences) including Swissprot and Trembl entries and 48 common contaminants, with peptide lengths set to 7–15 units. Both precursor and fragment quality tolerances were set to 20 ppm. Cysteine ​​ureomethylation was set as a static modification, while methionine oxidation, cysteylation, and N-terminal pyroglutamate or acetylation were set as variable modifications. Up to three variable modifications were allowed per peptide. For validation, MSBooster in conjunction with Percolator was enabled, and protein FDR was disabled. A library was constructed using the library module in FragPipe with standard parameters, and retention times were calibrated against iRT standards (Biognosys).

[0124] The DIA data was analyzed using DIA-NN version 1.8.2 beta 11 with the following standard settings: quality precision = 10 ppm and MS1 precision = 5 ppm, no inter-run matching, searching against sample-specific spectral libraries or prediction libraries generated by FragPipe, and precursor FDR filtering set to 1%.

[0125] SILAC DIA data were analyzed using DIA-NN version 1.8.2 beta 11 with the following search settings: library generation set to "IDs, RT and IM Profiling", quantification strategy set to "Peak height", scan window = 1, quality precision = 10 ppm, and MS1 precision = 5 p.pm. Additional commands entered in the DIA-NN command window were: {--peak-translation}, {--fixed-modSILAC,0.0,LKR,label}, {--lib-fixed-mod SILAC}, {--channels SILAC,L,LKR,0:0:0;SILAC,H,LKR,6.020129:8.014199:10.008269}. Unless otherwise specified, identification was further filtered to lengths 8-11, and trypsin contamination and other contaminants were excluded.

[0126] HLA peptide binding prediction HLApollo v1 was used to predict the binding of HLA peptides to alleles (Thrift et al. (2024), "Towards designing improved cancer immunotherapy targets with a peptide-MHC-I presentation model, HLApollo", Nat. Commun. 15, 10752). To evaluate MS peptide identification, identified peptide sequences of length 8–13 were evaluated to examine their binding to their respective alleles in cell lines. Peptides with an HLApollo score >0 were designated as binding peptides (Thrift et al. (2022), "HLApollo: A superior transformer model for pan-allelic peptide-MHC-I presentation prediction, with diverse negative coverage, deconvolution and protein language features", bioRxiv2022, 12.08.519673), and peptides predicted to bind to multiple alleles were assigned to all possible alleles. For library generation, all possible 8-13 polymers from the entire human proteome (swissprot: 62,266 protein sequences, 57,172,493 possible 8-13 polymers, including the sequence context of the four residues at the N-terminus and C-terminus) were generated, and their presentation on A375 cells was predicted. Only peptides that bind to any one of the six alleles and have a score >0 were retained.

[0127] Predicting peptide retention time, fragmentation, and ion mobility using Salud and Ionmob For the predicted A375 spectral library (prediction library), HLA peptides identified as binding peptides by HLApollo were taken into consideration, with charge states 1–3 considered for each peptide, and cysteine ​​immobilization alkylation and methionine variable oxidation considered as possible modifications. Spectral features were predicted using Salud, an internal ML model set based on Prosit (Gessulat et al. (2019), ibid.) for predicting retention time and fragmentation. This peptide fragmentation model is based on the Prosit non-trypsin digestion HCD model (Wilhelm et al. (2021) “Deep learning boosts sensitivity of mass spectrometry-based immunopeptidomics”, Nat Commun 12, 3346), with collision energies of 15, 20, and 25 for charge states 1, 2, and 3, respectively. For ion mobility, collision cross-sections were predicted using Ionmob (Teschner et al. (2023), ibid.) and converted to 1 / K0 values.

[0128] Neoantigen identification and quantification In order to systematically identify HLA A 11:01 Neoantigen, a spectral library was constructed in Fragpipe v20 using a two-step method (A1101-combo; see Table 2). First, using a human proteome fasta file containing the neoantigen sequence as well as the BFP reporter sequence and iRT sequence (Biognosys), the spectral library was constructed from all A... 11:01 DDA run data creation library. To maximize the chance of identifying putative neoantigens, three repeated DDA injections were performed on an equimolar mixture of 456 unique synthetic relabeled 8-11 polymeric neoantigens, which may be generated by mutations in 47 common cancers, using an effective gradient of 51 minutes as described above. Since each peptide in the synthesized neoantigen pool contains one heavy amino acid, the mass difference of lysine (K: 8.01499), arginine (R: 10.008269), valine / proline (V / R: 6.0138), leucine / isoleucine (L / I: 7.0171), phenylalanine / tryptophan (F / Y: 10.0272), and alanine (A: 4.007), as well as methionine oxidation, are set as variable modifications. We allow a maximum of two variable modifications per peptide and set urease methylation of cysteine ​​residues as a static modification. We constructed a spectral library containing heavily isotopically labeled peptide sequences using a fasta file containing only the neoantigen sequence, BFP reporter sequence, and iRT sequence (Biognosys). Ion mobility calibration was performed using the default settings, i.e., automatically selecting one run as the reference IM and selecting only b and y ions for library generation. After generating the heavy isotope-labeled spectral library in Fragpipe, a custom R script was used to add the light isotope-labeled transitions for each neoantigen peptide sequence to the library. Interference correction was then performed by removing identical fragment ion pairs from the combined library of heavy and light isotope-labeled neoantigens. Finally, for the library used for searching (A1101-combo), all peptides present in the synthetic peptide library were added to the DDA library, and duplicates were removed. In DIA-NN version 1.8.2 beta 11, the precursor FDR was set to 5%, and the data was analyzed by importing the spectral library created with Fragpipe v20. The quantification strategy was set to Robust LC (high precision), the inter-run normalization was set to RT-dependent, and the library generation was set to smart profiling. Whenever a synthetic heavy isotope-labeled peptide is available, it will be analyzed using Skyline (v23.1.0.268) to cross-validate all neoantigen identifications of DIANN.Prior to data acquisition, 100 fmol of the previously predicted A was spiked into each diluted sample. 11:01 Neoantigens of alleles (57 in total). Then, the amount of atomolar amount of each neoantigen peptide was calculated by using the ratio of the MS1 area signal of the endogenous light isotope-labeled peptide to the signal of the spiked synthetic heavy isotope-labeled peptide.

[0129] Table 2: Overview of Spectral Library Characteristics Statistical analysis All data analyses were performed using custom scripts in Python (3.9.4) with packages including pandas (1.1.5), numpy (1.22.2), plotly (5.4.0), and scipy (1.7.3), or R (4.3.1) with packages dplyr (1.1.3), data.table (1.14.8), stringr (1.5.1), and ggplot2 (3.4.4). Statistical analysis parameters (where applicable) are described in the relevant paragraphs and / or figure captions.

[0130] Experimental design and statistical basis Experiments were performed using A375 cells or engineered C1R cell lines (see relevant sections). A375 DIA experiments were repeated using a triple technique. We used the same batch of cultured cells for library creation and DIA injection. Statistical analysis parameters (if applicable) are described in the relevant paragraphs and / or figure captions. For diluted samples, injections were performed in order from low input to high input to minimize residual effects.

[0131] result HLA peptide analysis using the diaPASEF workflow DIA's analytical depth and sensitivity make it attractive for low-yield samples, such as HLA peptides. To preliminarily evaluate the diaPASEF immunopeptidomics workflow, we applied our automated immunoaffinity purification protocol to enrich HLA-I peptides ( ) from A375 multi-allelic melanoma cells. Figure 5A In short, we cleaved 125 e. 6 A375 cells were collected and enriched with HLA-I peptide. This is equivalent to 75 e 6The peptides from each cell were further fractionated into three replicates, separated, and analyzed in DDA mode. To create a DDA-based experimental library, data analysis was performed using Fragpipe, resulting in a final library containing 20,178 sequences. The remaining amount was processed in quadruplicates (each equivalent to 12.5 e^(-1 / 2)). 6 Inject (number of cells) and run in diaPASEF mode ( ). Figure 5B ; Figure 11A and Figure 11B Furthermore, analysis was performed using a sample-matched spectral library via DIA-NN. This enabled analysis from the equivalent of 12.5 e spectral density within a short 60 min LC gradient time. 6 10,297 HLA-I peptides were identified in the cells. Figure 5C -D). Notably, we observed high data integrity, with >9,500 peptides in all four replicates. Figure 11C Interestingly, >35% of the identified peptides were single-charged peptides, highlighting the importance of isolation polygon optimization. Figure 5E Furthermore, all replicates exhibited similar quantitative behavior, with a median R² of 0.96 and a median coefficient of variation of 12%. Figure 5F Based on the peptide length distribution and the presence of typical anchoring residues in the A375 allele, we can further verify that the identified peptides do indeed bind to HLA: >92% of the identified peptides are 8-13 polymers, >70% are predicted to bind to specific alleles, and approximately 5% can be assigned to multiple alleles, which is consistent with experimental observations based on DDA. Figure 5G ).

[0132] DIA provides high immunopeptide coverage with half the analysis time. Then, we started with the reduced cell number (500,000-10e) 6 HLA peptides were enriched in A375 cells (triple replicates), and DIA was found to improve immune peptidomome coverage when starting with fewer cells. Figure 6A Notably, compared to our standard DDA method (using a 120-minute gradient), DIA, using only a 60-minute gradient, identified a greater number of precursors across all analyzed cell inputs than previously acquired DDA datasets. The median number of precursors increased as peptides were isolated from a larger number of cells, reflecting the expected changes in HLA ligand abundance. Figure 12A Furthermore, DIA allows for the identification of >3,000 precursors under all conditions. Figure 12B ), while using DDA, only >1,000 peptides were found in all samples ( Figure 12C This may be due to the intensity-based random precursor selection mechanism of DDA in MS / MS analysis. For both methods, most peptides were shared between the two higher input conditions (7,388 for DIA and 4,702 for DDA). We found that the overlap between DDA and DIA identification in this experiment was 20%-26% with increasing cell quantity. Figure 6B The measured unique DDA assay has a higher intensity value compared to the proprietary DIA assay. Figure 6C These peptides differ slightly in their assigned HLA alleles. Figure 12D Among them, the unique DDA identification is related to HLA A. 02:01 The matching ratio is slightly larger. This can be partly explained by the identified m / z space. Compared to DIA, DDA collected more single-charge precursors at a higher intensity (800-1200 m / z). Figure 12E ), and A 02:01 typically produces a single-charge precursor. To assess whether these differences are related to the collection of DDA and DIA at different time points, we analyzed data from A... At 11:01, similar experiments were performed on peptides isolated from monoallelic cells, where DDA and DIA runs and libraries were generated from the same cell lysate and HLA peptide pool. Similarly, we found that DIA analysis identified more unique peptides compared to DDA. Figure 6D When samples were generated and collected from the same batch, we observed a greater overlap in DDA and DIA identification (approximately 30%-40%). Figure 6E ), and the peptide intensity distribution is more similar ( Figure 6F ).

[0133] The predicted spectral characteristics are good estimates of the empirically determined parameters. Experiment-specific spectral libraries facilitate sensitive matching of HLA peptides obtained via direct-DIA (DDA). However, creating such libraries is resource-intensive, with up to 80% of available starting materials typically used to generate deep experiment-specific libraries. Unfortunately, sample-specific libraries often have limited coverage due to the limited number of eluted HLA peptides and inherent limitations of the DDA method. Deep learning algorithms for predicting MS / MS spectra of existing peptide sequences offer an alternative approach for creating experiment-specific sample libraries. However, the non-trypsinic nature of HLA peptides and the theoretical number of over 50 million variants of 8-13 polymers from the human proteome complicate direct-DIA searches.

[0134] Therefore, we compared a library containing predicted spectral features of sequences from the library with an empirically generated spectral library. First, to verify the accuracy of the retention time prediction, we examined the correlation between the predicted retention time (RT) calculated by DIA-NN and the measured RT. Figure 13A and Figure 13B Compared to the experience database, our prediction database shows a slightly higher bias, which may indicate a significant difference in the average RT predictions between the two databases. In fact, we observed that the standard deviation of the difference between the measured RT and the predicted RT in the prediction database was approximately 3 times higher than that in the experience database. Figure 7A However, the difference between the measured RT peak values ​​of the DIA of the elution curves identified between the two libraries was within 3 seconds, which means that the same elution curves were identified for both libraries. Figure 7B Next, we investigated whether the prediction library contained sufficient information for quantifying fragment ions in the DIA file. Encouragingly, the prediction library produced a similar proportion of quantified fragment ions, indicating that Salud's fragment intensity predictions are of good quality. Figure 7C Furthermore, the overall predicted quality (median spectral angle = 0.68) was comparable to that of the experimentally specific library (spectral angle = 0.7). Figure 7D Consistent with previous studies (Pak et al. (2021), ibid.), we found that the number of HLA peptides identified using the prediction library decreased by approximately 15% compared to the empirical library (8,424 vs. 10,597). Of the peptides identified using the prediction library, over 85% were shared with the empirical library. Interestingly, we found that approximately 15% (1,062 peptides) were found exclusively in the prediction library, while 38% were unique to the empirical library. Figure 7E ).

[0135] High-sensitivity immune peptidomics using prediction libraries Inspired by the correlation between empirical and predicted libraries, we employed a dual prediction strategy to create a synthetic library for A375 containing all possible 8-13 polymer-binding peptides. First, the presentation of classic human proteins (343,034,958 peptide-HLA allele combinations) was predicted using the state-of-the-art predictor HLApollo (Thrift et al. (2022) “HLApollo: A superior transformer model for pan-allelic peptide-MHC-I presentation prediction, with diverse negative coverage, deconvolution and protein language features”, bioRxiv, 2022.12.08.519673), narrowing the search space to only 382,537 potentially presentable sequences. Then, Salud was used to predict retention time, fragment ions, and ion mobilities of this narrowed set of possible precursors to generate an allele-specific library matched to the sample. Figure 8A ). 26% of the precursors in the empirically generated library were not included in the prediction library, and most of these peptides were not predicted as binding peptides. Therefore, these peptides are removed using the proposed method. Figure 14A Although we used an experience library for samples with lower injection volumes (5e) 5 -2.5e 6 A slightly larger number of peptides were identified (from 5e cells), but the prediction library was able to identify peptides from 5e cells. 6 -25e 6 Samples enriched in each cell identified more peptides (summary results from three replicates). Figure 8B Both methods demonstrated high reproducibility, with a median coefficient of variation <0.2 under all conditions. Figure 8C For example, in 25e 6 At the input level, we identified >12,643 peptides using a computer-simulated library, while approximately 8,288 peptides were identified using a DDA-based library. Figure 8D Although we observed approximately 35% overlap between the two libraries, in 25e 6 Under input, approximately 42% of the peptide IDs found in the experience library were also identified by the prediction library; while in 5e 5 With this input, approximately 61% of the empirical library peptides were also identified by predictive methods. Interestingly, this was also true when the input amount was <5e 6 When analyzing 100 cells, we identified more peptides using the empirical library than in the computer simulation library. However, when comparing 0.5e...6 When examining the motifs of overlapping and unique peptides in a cell experiment in relation to either library, we found that all peptides showed the characteristic sequence of A375-presented peptides ( Figure 8E ). When examining additional quality criteria such as intensity and spectral angle, peptides identified by the empirical library rather than the predicted library showed lower intensity and lower spectral angle, and the confidence in their identification results may have been relatively low ( Figure 8F ; Figure 14B-1 and Figure 14B-2 ; Figure 14C-1 and Figure 14C-2 ).

[0136] Peptides with reduced co-elution by ion mobility Using the predicted spectral features of over 300,000 peptides that may be presented, we further evaluated the ion distribution on the DIA window and LC gradient ( Figure 15A -B, Table 1). Notably, ion mobility has reduced the median number of co-eluting precursors per scan from 10 to 1. Most predicted precursors fell into window 1.2 (1 / k0: 0.4 - 0.87; m / z: 300.53 mz - 430.23 mz) and window 12.1 (1 / k0: 1.09 - 1.7; m / z: 975.48 - 1340.62), where >200 precursors eluted in the middle of the LC gradient ( Figure 15C ). However, these are deliberately set large windows, and there are not many actual observed features for these mass / ion mobility combinations. All other DIA window and PASEF cycle combinations contained fewer than 50 precursors, and the analysis of these predicted precursors was performed within the same window at a given time. To characterize the degree of co-fragmentation on HLA peptide sequences, we calculated the occurrence of shared peptide fragments in the PASEF cycle window at a given predicted LC elution point. We compared fragment ions in a wider m / z window and found that many predicted co-eluting ions shared fragment ion masses, and unless the ion length was long (b / y9 or higher, Figures 15D-1 to 1 5D4), it could not be uniquely assigned to a specific peptide. Fortunately, for windows with fewer co-elution features (such as DIA window 8.2), this effect was significantly reduced, where only fragment ions with lower m / z (<b3 / y3) were shared between co-eluting precursors, and most fragment ions could be uniquely assigned to a single precursor ( Figures 15D-1 to 15D-4 ). Overall, this analysis indicates that the addition of ion mobility significantly reduces the problem of co-eluting HLA peptides that are difficult to distinguish in DIA.

[0137] Quantitative benchmarking of DIA using SILAC Next, we developed a SILAC HLA peptide DIA strategy to evaluate the accuracy and precision of quantification. Prior to lysis, HLA peptide enrichment, and DIA, A375 cells grown in heavy SILAC (K, R, L) and light SILAC media were mixed at ratios of 1:1, 1:2, 1:4, 1:9, and 1:19, respectively, to achieve a total cell count of 20e for each medium. 6 Cells ( Figure 9A The DIA data were analyzed using the A375 library and DIANN (with SILAC-specific parameters) (Methods). Not all presented peptides contained the amino acids used in the SILAC mixture. Therefore, 70%–80% of the peptides identified in each experiment had a calculable H / L ratio. The median H / L ratio of all peptides identified in each sample was in high agreement with the expected H / L ratios for 1:1, 1:2, and 1:4 dilutions, with a relative error of only 0–4%. Figure 9B For dilutions of 1:9 and 1:19, the measured ratios showed a relative error of 16% and 46% respectively compared to the expected ratios. This may be due to the identification of lower-intensity precursors, which incorrectly identify lightly isotopically labeled peptides as heavily isotopically labeled peptides, and vice versa. Figure 16A In fact, when checking spectral matching quality, heavy isotope-labeled peptide identification showed a larger C0 score distribution than light isotope-labeled peptide identification; at a 1:19 dilution, the median C0 score for heavy isotope-labeled peptides was 0.5 (…). Figure 16B However, while additional filtering of the channels and the converted q-value (Borteçen et al. (2023) “An integrated workflow for quantitative analysis of the newly synthesized proteome”, Nat. Commun. 14, 8237) improved the C-score of the remaining peptides ( Figure 16C However, this did not affect the median ratio or remove peptides with extreme ratios ( Figure 16DThis suggests that these peptides were either misidentified or subjected to high signal-to-noise ratio interference at extreme ratios, and require careful evaluation. Overall, quantification was accurate up to 5-fold dilution, consistent with similar studies for low-input samples (Petrosius et al. (2023) “Exploration of cell state heterogeneity using single-cell proteomics through sensitivity-tailored data-independent acquisition”, Nat. Commun. 14, 5910).

[0138] Identification and quantification of common cancer neoantigens Next, we evaluated how DIA analysis could benefit neoantigen identification and quantification. We used HLA-A expression... A monoallelic C1R cell line with a mutation rate of 11:01 was transfected with a neoantigen cassette encoding 47 common cancer mutations. Figure 10A (Gurung et al. (2023), ibid.). From 1e 6 -50e 6 HLA complexes were enriched in cells and eluted, followed by DDA to generate a library, and then DIA was performed. Prior to DIA analysis, 100 fmol of synthetic heavy isotope-labeled standard peptides were spiked into the samples. This allowed for internal control and, due to the co-elution and co-fragmentation of the spiked peptides, helped eliminate false positives. We observed that as the cell number increased to 10e... 6 and 25e 6 The number of identified [items] has steadily increased. Interestingly, compared to DDA, in 50e [items / items]... 6 The number of DIA identified per cell decreased during the run, which may be due to space charge effects, gradient limitations, or library limitations. Figure 10B We observed an increase in neoantigen abundance, while the spiked MS1 area remained relatively constant. Figure 10C For example, although neoantigens from EGFR G719A are present in 1e 6 The abundance was low when individual cells were introduced, but it could be detected across all test input ranges as the amount presented on the surface increased. Figure 10D-1 , Figure 10D-2 , Figure 10E We identified 16 peptide sequences from the neoantigen construct, most of which were derived from KRAS mutant sequences. As expected, DIA provided more complete coverage across the entire dilution series and down to 1 e6 14 / 16 of the new antigens were detected in the cell input ( Figure 10F Using the signal of a spiked peptide labeled with a heavy isotope, we obtain the signal from the equivalent of 1e. 6 Starting with a base dose of 33 atomoles per cell, the relative abundance of 13 neoantigens on the cell surface was quantified. Figure 10G In contrast, from 166 e 6 In 100 cells, 20 new antigens were identified in the targeted PRM assay, while 9 were identified by DDA (Gurung et al. (2023), ibid.). This highlights the broad applicability of DDA collection in immunopeptidomics assays for quantification when tracking selected target peptides across samples.

[0139] In summary, these results demonstrate that DIA offers high analytical depth and excellent quantitative reproducibility in the context of immunopeptidomics.

[0140] discuss Advances in instrumentation have further expanded the application of DIA in numerous proteomics studies. Key advantages include reproducibility between replicates and relative quantification across large sample volumes. For DIA-based immunopeptidomics measurements, reproducibility between replicates has historically ranged from 60% to 70%. With DIA, we achieved over 99% overlap of identified peptides in replicates. Integration of ion mobility metrics such as FAIMS or TIMS reduces the amount of co-eluted precursors, while fewer overlapping or shared fragment ions result in higher confidence levels for identification.

[0141] While DIA achieves extremely high coverage depth while shortening the acquisition time of target immune peptide samples, additional sample material is still required to create a spectral library. Due to the biological complexity of HLA peptide samples, a library-based approach must be adopted to maintain a low FDR. Other studies have predicted a small set of target peptides and incorporated them into spectral libraries, or generated large libraries containing a wealth of available public HLA peptide data at the sequence or spectral level, independent of sample background and the HLA types present in the sample (Pak et al. (2021), ibid.; Wahle et al. (2024), ibid.). We propose a two-step strategy: first, utilizing a state-of-the-art HLA binding predictor to reduce the number of possible peptide sequences in the human proteome and tailor it to the sample under study. The second step involves predicting spectral features. This joint prediction approach leads to similar or more identifications, depending on the amount of peptide loaded onto the column. It is noteworthy that this prediction library is highly dependent on the prediction algorithm used and may miss novel peptide sequences not included in the binding predictions or sequences predicted as non-bindings. It can be further optimized by adding contaminant species and non-classical peptide sequences. The prediction results can also be used to further filter peptides. For example, peptides identified that are unique to the experimentally determined spectral library tend to have lower intensity and show worse spectral angles, which suggests that the confidence level of these identifications may be low.

[0142] Using all possible presentation precursors, we were also able to evaluate our data acquisition methods. We found that ion mobility generally reduced the number of co-eluted and co-fragmented peptides; and even for highly similar peptide sequences (such as HLA-presented peptides), most DIA spectra contained fewer than 50 features of ions with different fragments. Further improvements to window design can be made through other data acquisition methods, such as sliceDIA or synchronizePASEF (Skowronek et al. (2023) “Synchro-PASEF Allows Precursor-Specific Fragment Ion Extraction and Interference Removal in Data-Independent Acquisition”, Mol. Cell. Proteomics 22, 100489; Szyrwiel et al. (2022) “Slice-PASEF: fragmenting all ions for maximum sensitivity in proteomics”, bioRxiv, 2022.10.31.514544), midiaPASEF (Distler et al. (2023) “midiaPASEF maximizes information content in data-independent acquisition proteomics”, bioRxiv, 2023.01.30.526204), and narrow window acquisition (Guzman et al. (2024) “Ultra-fast label-free”). Novel methods such as “quantification and comprehensive proteome coverage with narrow-window data-independent acquisition” (Nat. Biotechnol., 1–12) can further reduce co-elution and co-fragmentation, thereby achieving high-confidence identification.

[0143] Compared to label-free quantification based on DDA or TMT-based quantification, DIA acquisition offers significant advantages for multi-sample quantification. For example, in a consistent sample background, assessing changes in peptide abundance caused by different treatment conditions or input amounts can be performed more efficiently using DIA. However, due to the inherent diversity of the HLA peptidomome, stratification by corresponding HLA type is crucial before performing such experiments. We evaluated the quantification accuracy using the SILAC method and found that DIA analysis could accurately recover known quantities with differences up to 5-fold. Quantification over a larger ratio range can be affected by background noise contamination, incomplete fragmentation, and misassignment of spectral matching. This is consistent with reports in other studies, including single-cell analyses (Petrosius et al. (20203), ibid.; Pino et al. (2021) “Improved SILAC Quantification with Data-Independent Acquisition to Investigate Bortezomib-Induced Protein Degradation”, J. Proteome Res. 20, 1918–1927). In cell line samples carrying common tumor mutations, we were able to detect levels as low as 1e. 6The novel antigen peptides are introduced into cells into the IP, and the incremental peptides are quantified as the cell count increases. When combined with the hip-MHC method (Stopfer et al. (2021) "Absolute quantification of tumor antigens using embedded MHC-I isotopologue calibrants", Proc National Acad Sci 118, e2111173118), the copy number of peptides presented on the surface can be determined using the ratio of light to heavy peptides, while still collecting data on all other peptides present in the sample, and thus without losing additional peptidominant information discarded in PRM acquisition (Martínez-Val et al. (2023) "Hybrid-DIA: intelligent data acquisition integrates targeted and discovery proteomics to analyze phospho-signaling in single spheroids", Nat. Commun. 14, 3599; Goetze et al. (2024) "Simultaneous targeted and discovery-driven clinical proteotyping using hybrid-PRM / DIA", Clin. Proteom. 21, 26).

[0144] In summary, the integration of DIA with ion mobility provides a powerful data acquisition method. This enhances reproducibility and facilitates convenient peptide quantification within the same HLA allele type context. Furthermore, our proposed library-free strategy for HLA immunopeptidome measurements can be readily constructed using existing resources. The combination of this method with DIA measurements ensures robust quantification, which can significantly increase throughput, thus demonstrating its advantage for studies evaluating the mechanisms of action of immunotherapies.

[0145] In summary, gas phase separation in data-dependent MS data acquisition (DDA) improved HLA-I peptide detection by up to 50%. We evaluated the performance of data-independent acquisition (DIA) combined with ion mobility spectrometry (diaPASEF) in the high-sensitivity identification of HLA-presenting peptides. The simplified diaPASEF workflow enabled the identification of 11,412 unique peptides from 12.5 million A375 cells and the identification of 3,426 8-11 polymers with high reproducibility from as few as 500,000 cells. By utilizing computer simulations to predict spectral libraries specific to HLA conjugates, we were able to further increase the number of identified HLA-I peptides. We applied SILAC-DIA to mixtures of labeled HLA-I peptides, calculated the heavy peptide to light peptide ratios of 7,742 peptides under five conditions, and demonstrated that diaPASEF achieved high quantitative accuracy at dilutions up to 4-fold. Finally, we identified and quantified shared neoantigens in a monoallelic C1R cell line model. We validated peptide sequence identification by spiking and relabeling synthetic peptides and calculated the relative abundance of 13 neoantigens. In summary, the diaPASEF analytical workflow for HLA-I peptides can increase peptide coverage with low sample volumes. The sensitivity and quantitative precision provided by DIA enable the detection and quantification of low-abundance peptides, such as neoantigens, from samples with similar backgrounds.

[0146] Exemplary embodiments The embodiments disclosed herein may include: 1. A method for identifying MHC peptides present in a sample, the method comprising: Use one or more machine learning models to predict MHC allele-specific mass spectrometry libraries for a sample. Mass spectrometry analysis of peptides extracted from samples was performed using data-independent acquisition (DIA) to generate mass spectrometry data of the peptides; and Mass spectrometry data of the peptides and a sample-specific MHC allele-specific mass spectrometry library were used to identify at least one MHC peptide present in the sample.

[0147] 2. The method according to Example 1, wherein predicting the MHC allele-specific mass spectrometry library includes: Generate multiple candidate peptide sequences with sequence lengths within a specified range from a reference proteome; A machine learning-based MHC binding prediction model is used to predict the binding of a candidate peptide sequence to at least one MHC allele from a pool of candidate peptide sequences, thereby generating multiple candidate MHC peptide sequences; and A machine learning-based mass spectrometry prediction model is used to predict the mass spectrometry features of each candidate MHC peptide sequence among multiple candidate MHC peptide sequences, in order to generate a predicted MHC allele-specific mass spectrometry library.

[0148] 3. The method according to Example 1 or Example 2, wherein the mass spectrometry features of each candidate MHC peptide sequence include peptide retention time, peptide fragmentation mode, ion mobility prediction, or any combination thereof.

[0149] 4. The method according to Example 2 or Example 3, wherein the specified range of sequence length is 8 to 13 amino acid residues.

[0150] 5. The method according to any one of Examples 2 to 4, wherein the machine learning-based MHC combined prediction model is HLApollo.

[0151] 6. The method according to any one of Examples 2 to 5, wherein the machine learning-based mass spectrometry prediction model is AlphaPeptDeep. TM DirectDIA, DIA-NN, Prosit, or proprietary mass spectrometry prediction models.

[0152] 7. The method according to any one of Examples 1 to 6, wherein at least one MHC peptide present in the sample contains at least one neoantigen.

[0153] 8. The method according to any one of Examples 1 to 7, further comprising extracting and / or purifying peptides from a sample.

[0154] 9. The method according to any one of Examples 1 to 8, wherein mass spectrometry analysis of peptides extracted from a sample comprises combining parallel cumulative serial fragmentation (PASEF) with data-independent acquisition (DIA).

[0155] 10. The method according to any one of Examples 1 to 9, wherein the sample is derived from a human subject, and at least one MHC peptide is an HLA peptide.

[0156] 11. The method according to Example 10, wherein the human subject is a patient.

[0157] 12. The method according to any one of Examples 1 to 9, wherein the sample is derived from a primate, mouse, bacterial culture, viral culture or cultured cell line.

[0158] 13. The method according to any one of Examples 2 to 12, wherein the reference proteome is a human reference proteome, a primate reference proteome, a mouse reference proteome, a bacterial reference proteome, or a viral reference proteome.

[0159] 14. The method according to any one of Examples 1 to 13, wherein at least one MHC peptide is used to develop cancer vaccines or targeted cell therapies.

[0160] 15. A system comprising: One or more processors; and A memory, communicatively coupled to one or more processors, and configured to store instructions that, when executed by one or more processors, cause the system to perform the method according to any one of embodiments 1 to 14.

[0161] 16. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of the system, cause the system to perform the method according to any one of embodiments 1 to 14.

[0162] This description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the description of preferred exemplary embodiments will provide those skilled in the art with a feasible description for implementing various embodiments. It should be understood that various changes can be made to the function and arrangement of the elements without departing from the spirit and scope set forth in the appended claims.

Claims

1. A method for identifying MHC peptides present in a sample, the method comprising: Use one or more machine learning models to predict the MHC allele-specific mass spectrometry library for the sample. Mass spectrometry analysis of peptides extracted from the sample was performed using data-independent acquisition (DIA) to generate mass spectrometry data of the peptides. as well as The mass spectrometry data of the peptide and the predicted MHC allele-specific mass spectrometry library for the sample were used to perform a library search to identify at least one MHC peptide present in the sample.

2. The method according to claim 1, wherein predicting the MHC allele-specific mass spectrometry library comprises: Generate multiple candidate peptide sequences with sequence lengths within a specified range from a reference proteome; A machine learning-based MHC binding prediction model is used to predict the binding of candidate peptide sequences among the plurality of candidate peptide sequences to at least one MHC allele, thereby generating a plurality of candidate MHC peptide sequences. as well as A machine learning-based mass spectrometry prediction model is used to predict the mass spectrometry characteristics of each candidate MHC peptide sequence among the multiple candidate MHC peptide sequences, in order to generate a predicted MHC allele-specific mass spectrometry library.

3. The method according to claim 1 or claim 2, wherein the mass spectrometry features of each candidate MHC peptide sequence include peptide retention time, peptide fragmentation mode, ion mobility prediction, or any combination thereof.

4. The method according to claim 2 or claim 3, wherein the specified range of sequence length is 8 to 13 amino acid residues.

5. The method according to any one of claims 2 to 4, wherein the machine learning-based MHC combined prediction model is HLApollo.

6. The method according to any one of claims 2 to 5, wherein the machine learning-based mass spectrometry prediction model is AlphaPeptDeep or directDIA. TM DIA-NN, Prosit, or proprietary mass spectrometry prediction models.

7. The method according to any one of claims 1 to 6, wherein the at least one MHC peptide present in the sample comprises at least one neoantigen.

8. The method according to any one of claims 1 to 7, the method further comprising extracting and / or purifying the peptide from the sample.

9. The method according to any one of claims 1 to 8, wherein the mass spectrometry analysis of the peptides extracted from the sample comprises combining parallel cumulative serial fragmentation (PASEF) with data-independent acquisition (DIA).

10. The method according to any one of claims 1 to 9, wherein the sample is derived from a human subject, and the at least one MHC peptide is an HLA peptide.

11. The method of claim 10, wherein the human subject is a patient.

12. The method according to any one of claims 1 to 9, wherein the sample is derived from a primate, mouse, bacterial culture, viral culture, or cultured cell line.

13. The method according to any one of claims 2 to 12, wherein the reference proteome is a human reference proteome, a primate reference proteome, a mouse reference proteome, a bacterial reference proteome, or a viral reference proteome.

14. The method according to any one of claims 1 to 13, wherein the at least one MHC peptide is used for the development of cancer vaccines or targeted cell therapies.

15. A system comprising: One or more processors; as well as A memory, communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform the method according to any one of claims 1 to 14.

16. A non-transitory computer-readable storage medium storing one or more programs, said one or more programs including instructions that, when executed by one or more processors of the system, cause the system to perform the method according to any one of claims 1 to 14.