Differential population-scale immunopeptidomics

By integrating immunopeptidomic analysis with orthogonal omics data and a population-scale reference dataset, the method effectively identifies differential MHC-presented peptides, addressing the limitations of existing technologies and enhancing therapeutic targeting.

WO2025125535A1PCT designated stage expired Publication Date: 2025-06-19IMMATICS BIOTECHNOLOGIES GMBH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/086145
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-12-13
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Current methods for identifying differential MHC-presented peptides are limited by their reliance on epitope prediction and screening in complex peptide mixtures, which lack sensitivity and specificity, especially in distinguishing between healthy and diseased states.

Method used

The method combines immunopeptidomic analysis with orthogonal measurements of RNA, protein, and DNA, using a population-scale reference dataset to identify differential MHC-presented peptides by matching paired data against quantitative human reference data based on at least 25 donors.

Benefits of technology

This approach enhances the sensitivity and specificity of identifying differential MHC-presented peptides, providing robust evidence of their clinical relevance and therapeutic potential.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000031_0001
    Figure IMGF000031_0001
  • Figure 00000043_0000
    Figure 00000043_0000
  • Figure 00000043_0001
    Figure 00000043_0001
Patent Text Reader

Abstract

The present invention relates to a method for identifying differential MHC-presented peptides.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Differential population-scale immunopeptidomics

[0002] The present invention relates to a method for identifying differential MHC-presented peptides.

[0003] Background of the Invention

[0004] The major histocompatibility complex (MHC) is an essential component of the adaptive immune system responsible for antigen presentation at the cell surface. In a cell, proteins are constantly synthesized and proteasomally degraded. Short peptide fragments of these degraded proteins are presented by the MHC molecules on the cell surface. In case a cell suffers from a disease or other pathological condition such as e.g. viral infection, intracellular microorganism infection, cancerous transformation, etc., the resultant altered protein repertoire is presented on the cell surface through MHC. This enables T lymphocytes to surveil cellular health by screening MHC-presented peptides with their T cell receptor (TCR) and initiate immune reactions in case of a peptide-MHC complex being recognized.

[0005] This mechanism makes it possible to target differentially presented, disease-associated MHC peptides with immunotherapies, including the use of binding moieties such as TCRs, antibodies, and antigen binding fragments, as well as vaccination, as therapeutics for various diseases. Such a treatment would be based on qualitative or quantitative differences for specific peptides between cells affected by the disease or clinical condition and unaffected / benign cells. There is therefore a great need for the identification of such MHC-presented peptides in general and in particular of those being indicative or even specific for a certain disease.

[0006] Methods known in the art rely on the combination of epitope prediction and screening in complex peptide mixtures derived from tissue samples of clinical relevance and comparative benign tissue samples. The detection of the peptides may include highly sensitive capillary liquid chromatography mass spectrometry (LC-MS) (see Schirle et al 2000, Eur. J. Immunol. 30:2216-2225). Other known methods rely on comparative expression profiling of cancerous and corresponding normal tissue to identify differential peptides (see Young et al 2001, Am. J. Pathol. 158: 1639-1651) or combine expression analysis with epitope prediction (Mathiassen et al 2001, Eur. J. Immuno. 31 : 1239-1246). The present invention relies on the combination of immunop eptidomic analysis in combination with orthogonal measurements of other parameters (such as RNA, protein and DNA) resulting in paired data. The differential analysis is then empowered by the use of a population-scale reference dataset that is used to determine which peptides show pronounced differences between a sample type of clinical interest (e.g. various tumor types or autoimmune diseases) and the reference dataset.

[0007] Summary of the Invention

[0008] In a first aspect, the present invention provides an in vitro method for identifying a differential MHC-presented peptide comprising the following steps:

[0009] (a) providing a sample or group of samples of clinically relevant tissue,

[0010] (b) isolation and analysis of RNA, proteins, and / or DNA in the sample or group of samples of clinically relevant tissue,

[0011] (c) isolation and analysis of MHC-presented peptides in the sample or group of samples used in step (b),

[0012] (d) matching the paired data obtained in steps b) and c) against quantitative human reference data generated as described in steps b) and c), wherein the quantitative human reference data is based on at least 25 donors; and

[0013] (e) identification of the differential MHC-presented peptide on the basis of the matched data of step (d).

[0014] In a second aspect the present invention provides a method for the analysis of a data- independent acquisition (DIA) dataset obtained by mass spectrometry analysis of MHC- presented peptides in a biological sample comprising the steps of:

[0015] (i) providing at least one spectral library based on predicted MHC-presented peptides specific for a single MHC allotype;

[0016] (ii) analyzing the DIA-dataset based on the at least one spectral library of step (ii).

[0017] In particular the invention relates to the following items:

[0018] 1. An in vitro method for identifying a differential MHC-presented peptide comprising the following steps:

[0019] (a) providing a sample or group of samples of clinically relevant tissue,

[0020] (b) isolation and analysis of RNA, proteins, and / or DNA in the sample or group of samples of clinically relevant tissue, (c) isolation and analysis of MHC-presented peptides in the sample or group of samples used in step (b),

[0021] (d) matching the paired data obtained in steps b) and c) against quantitative human reference data generated as described in steps b) and c), wherein the quantitative human reference data is based on at least 25 donors; and

[0022] (e) identification of the differential MHC-presented peptide on the basis of the matched data of step (d).

[0023] 2. The method according to item 1, wherein the quantitative human reference data comprises data on at least one MHC-allotype and wherein each MHC-allotype in the quantitative human reference data is based on the analysis of at least 20 samples of this MHC-allotype.

[0024] 3. The method according to any one of the preceding items, wherein the analysis of MHC- presented peptides in step (c) comprises mass spectrometry analysis using data- independent acquisition (DIA), wherein analysis of DIA data comprises the use of at least one peptide spectral library based on predicted MHC-presented peptides specific for a single MHC allotype.

[0025] 4. The method according to any one of the preceding items, wherein the isolation of MHC- presented peptides of the sample or group of samples of clinically relevant tissue in step (c) comprises: affinity purification with one or more binders specific for a class of MHC- molecules or a specific MHC-allotypes; or acidic elution of MHC peptides.

[0026] 5. The method according to any one of the preceding items, wherein the quantitative human reference data is univariate or multivariate reference data, preferably wherein the human reference data comprises data on MHC-presented peptide expression, nucleic acid expression, protein expression, and / or DNA variation.

[0027] 6. The method according to any one of the preceding items, wherein the quantitative human reference data is combined with further datasets obtained with other methods, preferably wherein the further datasets comprise gene expression data. 7. The method according to any one of the preceding items, wherein the analysis in step

[0028] (b) is performed by methods that allow the analysis of the complete set of analytes in the sample of clinically relevant tissue, preferably by RN A- sequencing, DNA- sequencing and LC-MS / MS.

[0029] 8. The method according to any one of the preceding items, wherein the clinically relevant issue is tissues affected by cancer, inflammation, auto-immune disease.

[0030] 9. The method according to any one of items 3 to8, wherein the peptide spectral library is obtained by a method comprising the following steps:

[0031] (I) obtaining a protein database representing the proteome of the biological sample to be analyzed;

[0032] (II) applying a prediction algorithm to the protein database of step (I) for generating a database of potential MHC-presented peptides,

[0033] (III) generating a peptide spectral library based on the database of potential MHC- presented peptides.

[0034] 10. The method according to item 9, wherein the generation of the peptide spectral library further includes empirical optimization of the spectral library.

[0035] 11. The method according to any one of items 3 to 10, wherein the spectral library comprises of target and decoy spectra and associated retention times.

[0036] 12. The method according to any one of items 9 to 11, wherein the peptide spectral library is generated by a method comprising the steps:

[0037] (a) based on the database of potential MHC-presented peptides of step (II) a database of negative controls is generated;

[0038] (b) a chromatogram for each MHC-presented peptide and negative control is generated and elution peaks are identified therefrom and scored;

[0039] (c) the elution peaks with the highest scores are selected;

[0040] (d) a deep neural network algorithm is applied to the selected elution peaks to distinguish between the MHC-presented peptides and negative controls resulting in a peptide spectral library. 13. The method according to any one of items 9 to 12, wherein the spectral library is generated using DIA-NN, preferably in library free mode.

[0041] 14. A peptide as identified by the method according to any one of items 1 to 13.

[0042] 15. An (in vitro) method for identifying an antigen-binding protein binding to a peptide (preferably bound to a MHC protein) identified by the method according to any one of items 1 to 13 comprising: a) contacting a plurality of antigen-binding proteins with said peptide (preferably bound to a MHC protein), b) identifying an antigen-binding protein that binds to said peptide (preferably bound to a MHC protein), and c) selecting the antigen-binding protein identified in b).

[0043] 16. An antigen-binding protein as obtained by the method according to item 15.

[0044] 17. A host cell comprising the antigen-binding protein according to item 16.

[0045] 18. A pharmaceutical composition comprising the peptide according to item 14, the antigenbinding protein according to item 16 or the host cell according to item 17.

[0046] 19. The peptide according to item 14, the antigen-binding protein according to item 16, the host cell according to item 17 or the pharmaceutical composition according to item 18 for use as a medicament.

[0047] 20. The peptide according to item 14, the antigen-binding protein according to item 16, the host cell according to item 17 or the pharmaceutical composition according to item 18 for use in the treatment of cancer

[0048] In addition the invention relates to the following items:

[0049] 1. An (in vitro) method for treating a patient or determining a treatment option / modality comprising: (a) biopsying a sample or group of samples of clinically relevant tissue from the patient,

[0050] (b) isolating and analyzing RNA, proteins, and / or DNA in the sample or the group of samples of the clinically relevant tissue,

[0051] (c) isolating and analyzing MHC-presented peptides in the sample or group of samples used in step (b),

[0052] (d) matching the data obtained in steps (b) and (c) against quantitative human reference data, wherein the quantitative human reference data is generated as described in steps (b) and (c) based on at least 25 reference donors; and

[0053] (e) identifying the differential MHC-presented peptide on the basis of the matched data of step (d);

[0054] (f) treating the patient or determining the treatment option / modality based on the identified differential MHC-presented peptide of step (e). The method according to item 1, wherein the quantitative human reference data comprises data on at least one MHC-allotype and wherein each of the at least one MHC-allotype in the quantitative human reference data is based on the analysis of at least 20 samples of this MHC-allotype from the at least 25 reference donors. The method according to item 1, wherein the analyzing MHC-presented peptides in step (c) comprises mass spectrometry analysis using data-independent acquisition (DIA), wherein analysis of DIA data comprises use of at least one peptide spectral library based on predicted MHC-presented peptides specific for a single MHC allotype. The method according to item 1, wherein the isolating MHC-presented peptides in step (c) comprises: affinity purification with one or more binders specific for a class of MHC- molecules or specific MHC-allotypes; or acidic elution of MHC peptides. The method according to item 1, wherein the quantitative human reference data is univariate or multivariate reference data, optionally wherein the human reference data comprises data on MHC-presented peptide expression, nucleic acid expression, protein expression, and / or DNA variation.

[0055] 6. The method according to item 1, wherein the quantitative human reference data is combined with further datasets obtained with other methods, optionally wherein the further datasets comprise gene expression data.

[0056] 7. The method according to item 1, wherein the analyzing of step (b) is performed by methods that allow analysis of a complete set of analytes in the sample of clinically relevant tissue, optionally by RNA-sequencing, DNA-sequencing and LC-MS / MS.

[0057] 8. The method according to item 1, wherein the clinically relevant tissue is tissue affected by cancer, inflammation, and / or auto-immune disease.

[0058] 9. The method according to item 3, wherein the at least one peptide spectral library is obtained by a method comprising:

[0059] (I) obtaining a protein database representing the proteome of the biological sample to be analyzed;

[0060] (II) applying a prediction algorithm to the protein database of step (I) for generating a database of potential MHC-presented peptides,

[0061] (III) generating the at least one peptide spectral library based on the database of potential MHC-presented peptides.

[0062] 10. The method according to item 9, wherein the generating the at least one peptide spectral library further comprises empirical optimization of the spectral library.

[0063] 11. The method according to item 3, wherein the at least one spectral library comprises target and decoy spectra and associated retention times.

[0064] 12. The method according to item 9, wherein the at least one peptide spectral library is generated by a method comprising the steps: (a) based on the database of potential MHC-presented peptides of step (II) a database of negative controls is generated;

[0065] (b) a chromatogram for each MHC-presented peptide and negative control is generated and elution peaks are identified therefrom and scored;

[0066] (c) the elution peaks with the highest scores are selected;

[0067] (d) a deep neural network algorithm is applied to the selected elution peaks to distinguish between the MHC-presented peptides and negative controls resulting in the at least one peptide spectral library.

[0068] 13. The method according to item 9, wherein the at least one spectral library is generated using DIA-NN, optionally in library free mode.

[0069] List of Figures

[0070] In the following, the content of the figures comprised in this specification is described. In this context please also refer to the detailed description of the invention above and / or below.

[0071] Figure 1: refers to the improved sensitivity achieved by using DIA (see Example 1). (A) and (B) shows sensitivity to detect peptide CYP2W 1-001 increased by lowering the limit of detection (LOD) from 77.96 fmol (DDA, (A)) to 8.3 frnol (DIA, (B)). LOD was estimated using the probit method. The average reduction in LOD is about 10 fold (log2(DDA / DIA) = - 3.5) based on 45 peptides with at least 60 data points for LOD estimation. Depicted as a scatter plot (C) as well as a box plot (D).

[0072] Figure 2: refers to the generation of spectral libraries (see Example 2). Spectral libraries generated by MHCFlurry predictions and empirical optimization on 100 HLA-A*02:01 donors using DIA-NN-MBR shows 40% improvement (1.4 fold) for donors with more than 1000 peptides and more than 140% improvement (2.4 fold) for donors with less than 1000 peptides. The pronounced improvement in sensitivity at the low end is of particular relevance for comprehensive characterization of healthy tissues, since these frequently express lower levels ofMHC.

[0073] Figure 3: refers to the relevance of paired data. The top panel depicts peptide presentation data of the peptide YSGQLKVLI (SEQ ID NO: 001) that was identified as derived from the gene ANKRD30A, suggesting a relevance in gastric cancer. The middle panel shows the gene expression profile for ANKRD30A, also known as breast cancer antigen NY-BR-1, based on RNAseq data underlining the relevance of the gene in the context of breast cancer. The bottom panel depicts the peptide presentation profile for the peptide KILDTVHSC (SEQ ID NO: 002) derived from the gene ANKRD30A displaying the same preferential breast cancer pattern as the transcriptomic data. The contrast of the peptide presentation plots indicates when compared to the equivalent transcriptomic data that the peptide shown in the top panel is potentially a misidentification.

[0074] Figure 4: shows the statistical power in terms of detection probability in relation to the number of donors for DDA by the example of HLA-A*25:01 peptide YVDGTQFVRF (SEQ ID NO: 003) derived from SNP rs9266183 (D4G, MAF 3.3%) in HLA-B. For this rare peptide n donors are selected at random (sub sampling). This selection was repeated 1000 times and the probability of successful detection was determined.

[0075] Figure 5: shows the number of donors against the statistical power of the method in terms of immunopeptidome coverage. This reflects the ability to detect the indicated percentage of peptides using a spectral library of the present invention with empirical optimization for DIA (A) and DDA (B) data analysis.

[0076] Figure 6: shows the required number of donors against the statistical power to enable correlation analysis between peptide data and surrogate data (protein, mRNA, DNA) to allow biomarker discovery and investigation of mechanism of action for a immunopeptide target.

[0077] Figure 7: refers to comparison of proteomics, immunopeptidomics, and transcriptomic data to define a potential therapeutic modality, i.e. CAR-T vs TCR-T based on the therapeutic window. MHC peptide presentation profile of peptide KVFAGIPTV (SEQ ID NO: 004) from DDA data (top panel) and DIA (second panel) normalized to the input tissue weight, the third panel shows the proteomics data for PTPRZ1 being the source protein of the peptide KVFAGIPTV (SEQ ID NO: 004) as relative iBAQ intensity depicted as ppm, the bottom panel depicts the RNAseq expression data for KVFAGIPTV (SEQ ID NO: 004) on exon level. Overall, the plot is based on data obtained for 807 donors.

[0078] Figure 8: refers to presence of a confirmed off-target peptide ILGTHNITV (SEQ ID NO: 005) derived from the gene INTS7. The top panel depicts the MHC peptide presentation profile of peptide ILGTHNITV (SEQ ID NO: 005) from DIA data normalized to the input tissue weight, the bottom panel depicts the RNAseq expression data for ILGTHNITV (SEQ ID NO: 005) on exon level.

[0079] Figure 9: refers to an overview of the disclosed examples of the present invention and the corresponding figures.

[0080] Figure 10: refers to the detection of the peptide VLLGVKLSGV (SEQ ID NO: 008) from Example 9 using the DIA method of the present invention. Depicted is the intensity of the MSI signal for the peptide VLLGVKLSGV (SEQ ID NO: 008) (left chart) and the respective run-average MSI signal (right chart).

[0081] Figure 11: refers to a comparison of the acquired MS2 spectrum for peptide VLLGVKLSGV (SEQ ID NO: 008) from Example 9 using the DIA method of the present invention (top panel) compared to a spectrum derived from a fragmentation predictor (lower panel).

[0082] Figure 12: refers to the detection of the peptide VLLGVKLSGV (SEQ ID NO: 008) from Example 9 using the DIA method of the present invention. Depicted is the intensity of the MS2 signal for the peptide VLLGVKLSGV (SEQ ID NO: 008) (upper panel) and RNAseq expression data (lower panel) of the corresponding gene (COL18A1, ENSG00000182871) in different normal (left side, light gray) and tumor tissues (right side, dark gray).

[0083] Detailed Descriptions of the Invention

[0084] Before the present invention is described in detail below, it is to be understood that this invention is not limited to the particular methodology, protocols and reagents described herein as these may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the present invention which will be limited only by the appended claims. Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art.

[0085] Throughout this specification and the claims which follow, unless the context requires otherwise, the word "comprise", and variations such as "comprises" and "comprising", will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps. In the following passages, different aspects of the invention are defined in more detail. Each aspect so defined may be combined with any other aspect or aspects unless clearly indicated to the contrary. In particular, any feature indicated as being optional, preferred or advantageous may be combined with any other feature or features indicated as being optional, preferred or advantageous.

[0086] Several documents are cited throughout the text of this specification. Each of the documents cited herein (including all patents, patent applications, scientific publications, manufacturer's specifications, instructions etc.), whether supra or infra, is hereby incorporated by reference in its entirety. Nothing herein is to be construed as an admission that the invention is not entitled to antedate such disclosure by virtue of prior invention. Some of the documents cited herein are characterized as being “incorporated by reference ” . In the event of a conflict between the definitions or teachings of such incorporated references and definitions or teachings recited in the present specification, the text of the present specification takes precedence.

[0087] In the following, the elements of the present invention will be described. These elements are listed with specific embodiments; however, it should be understood that they may be combined in any manner and in any number to create additional embodiments. The variously described examples and preferred embodiments should not be construed to limit the present invention to only the explicitly described embodiments. This description should be understood to support and encompass embodiments which combine the explicitly described embodiments with any number of the disclosed and / or preferred elements. Furthermore, any permutations and combinations of all described elements in this application should be considered disclosed by the description of the present application unless the context indicates otherwise.

[0088] Definitions

[0089] In the following, some definitions of terms frequently used in this specification are provided. These terms will, in each instance of its use, in the remainder of the specification have the respectively defined meaning and preferred meanings.

[0090] As used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents, unless the content clearly dictates otherwise.

[0091] The term “about” when used in connection with a numerical value is meant to encompass numerical values within a range having a lower limit that is 5% smaller than the indicated numerical value and having an upper limit that is 5% larger than the indicated numerical value.

[0092] The “major histocompatibility complex” (MHC) in the context of the present invention is a set of cell surface proteins essential for the acquired immune system to recognize foreign molecules in vertebrates, which in turn determines histocompatibility. The main function of MHC molecules is to bind to antigens derived from pathogens and display them on the cell surface for recognition by the appropriate T cells. The human MHC is also called the HLA (human leukocyte antigen) complex (often just the HLA). The MHC gene family is divided into three subgroups: class I, class II, and class III. Complexes of peptide and MHC class I are recognized by CD8-positive T cells bearing the appropriate T cell receptor (TCR), whereas complexes of peptide and MHC class II molecules are recognized by CD4- positive-helper-T cells bearing the appropriate TCR. Since both types of response, CD8 and CD4 dependent, contribute jointly and synergistically, the identification and characterization ofMHC-presented peptides and corresponding T cell receptors is important in the development of immunotherapies such as vaccines and cell therapies. The HLA-A gene is located on the short arm of chromosome 6 and encodes the larger, a-chain, constituent of HLA-A. Variation of HLA-A a-chain is key to HLA function. This variation promotes genetic diversity in the population. Since each HLA has a different affinity for peptides of certain structures, greater variety of HLAs means greater variety of antigens to be 'presented' on the cell surface. Each individual can express up to two types of HLA-A, one from each of their parents. Some individuals will inherit the same HLA-A from both parents, decreasing their individual HLA diversity; however, the majority of individuals will receive two different copies of HLA-A. This same pattern follows for all HLA groups. In other words, every single person can only express either one or two of the over 7000 known HLA-A alleles.

[0093] A preferred MHC class I HLA protein in the context of the present invention may be an HLA-A, HLA-B or HLA-C protein, more preferably HLA-A protein, even more preferably the allotype group HLA-A*02, most preferably the specific allotype HLA-A*02:01. Other preferred HLA-A allotypes include HLA-A*24:02, HLA-A*01:01, HLA-A*03:01, HLA- B*07:02, HLA-B*08:01, HLA-B*44:02a and HLA-B*44:03.

[0094] “HLA-A*02:01” signifies a specific HLA allele or allotype, wherein the letter A signifies the gene and the suffix “02” indicates the allele or allotype group and “01” a specific HLA allele or allotype. The rules for nomenclature of HLA proteins is well known in the art and can for example be found at https: / / hla.alleles.org / nomenclature / naming.html.

[0095] In the MHC class I dependent immune reaction, peptides not only have to be able to bind to certain MHC class I molecules expressed by tumor cells, they subsequently also have to be recognized by T cells bearing specific T cell receptors (TCR).

[0096] A “MHC-presented peptide” in the context of the present invention is thus a peptide that is presented by a MHC molecule and that can be bound by a binding moiety (in particular a TCR or an antibody). Preferably a MHC class I presented peptide has a length of 8 to 11 amino acids, preferably 9 to 10, most preferably 9 amino acids. Preferably a MHC class II presented peptide has a length of 13 to 25 amino acids.

[0097] A “differential MHC-presented peptide” in the context of the present invention refers to a MHC-presented peptide that differs in abundance between two samples or group of samples. Preferably, the difference in abundance between the samples is a statistically significant difference. Preferably, the difference is sufficient to be of therapeutic relevance. An example of a differential MHC-presented peptide is a MHC-presented peptide that is of different abundance in the diseased and healthy state of an individual. Preferably, the compared samples or group of samples represent the healthy and diseased state of an individual, organ or tissue.

[0098] The term “immunopeptidome” as used herein refers to the set of peptides presented by major histocompatibility complex (MHC) molecules on the surface of cells to enable T-cell immunosurveillance. The immunopeptidome is thus a collection of all MHC-presented peptides in a particular sample or organism. The immunopeptidome can be further subdivided into sets of peptides presented by specific MHC molecules (or HLA-molecules in the human context). For example, the HLA-A*02 immunopeptidome would only encompass the peptides presented by HLA-A*02 molecules. Likewise the HLA-A*02:01 immunopeptidome would only encompass the peptides presented by HLA-A*02:01 molecules. In a preferred embodiment the immunopeptidome refers to all MHC (or HLA) presented peptides. In another preferred embodiment the immunopeptidome is restricted to peptides presented by the HLA molecule selected from: HLA-A*02:01, HLA-A*24:02, HLA-A*01:01, HLA-A*03:01, HLA-B*07:02, HLA-B*08:01, HLA-B*44:02a and HLA-B*44:03.

[0099] The term "peptide spectral library" or “spectral library” as used herein, refers to a collection of annotated peptide mass spectra that serves as a reference database for identifying peptides in a sample. Thus a spectral library may be used to match a mass spectrum measured in mass spectrometry to a particular peptide, which mass spectrum is comprised within the spectral library. A spectral library can be an experimental spectral library containing experimentally obtained ion spectra. In a preferred embodiment the peptide spectral library is constructed from sequence data. In a particular preferred embodiment the peptide spectral library is constructed from the amino acid sequences of predicted MHC-presented peptides.

[0100] In a preferred embodiment the spectral library is a HLA-allotype specific spectra library comprising mass spectra of peptides predicted or demonstrated to be presented by a particular HLA-allotype. Preferred alleles for a HLA-allotype specific library are HLA-A*02:01, HLA- A*24:02, HLA-A*01:01, HLA-A*03:01, HLA-B*07:02, HLA-B*08:01, HLA-B*44:02 and HLA-B*44:03.

[0101] In another preferred embodiment the spectral library is a spectral library comprising peptides presented by all HLA-allotypes.

[0102] In a preferred embodiment the spectral library is a sample specific spectral library. Such a library contains mass spectra of MHC-presented peptides present in a particular type of sample. Examples of preferred sample types are particular tissues and / or samples of clinically relevant tissues and their benign / healthy counterparts. In a preferred embodiment a method for constructing a peptide specific library is provided. This method comprises the provision of a protein database containing the peptides relevant for the peptide spectral library. In a next step the MHC-presented peptides are predicted based on the sequence provided in the protein database.

[0103] The terms ‘univariate’ and ‘multivariate’ as used herein, refer to a data set or an analysis where a single parameter (i.e. univariate) or multiple parameters (i.e. multivariate) are determined for the same experimental unit. A non-limiting example of a univariate dataset would be the measurement of a single parameter (e.g. peptide expression) in a certain type of experiment (e.g. measurement of a single parameter in the samples comprised in a dataset). Accordingly, a non-limiting example of a multivariate dataset would be the measurement of multiple parameters (e.g. peptide expression and RNA expression) in a certain type of experiment (e.g. measurement of the multiple parameters in the samples comprised in a dataset).

[0104] The term “paired data” as used herein refers to data acquired for a certain entity on different levels and therefore the data being linked. For example, measuring a transcript level as well as the level of the encoded peptide would result in paired data, having a measurement on the transcript level and on the peptide level for the same target.

[0105] The term “omics” as used herein refers to the collective characterization of pools of biological molecules. Non-limiting examples of “omics” are genomics, transcriptomics, proteomics, peptidomics and immunopeptidomics. Omics data therefore typically present large scale data sets that aim at characterizing all entities of a certain pool of biological molecules, e.g. transcriptomics aims at characterizing all transcripts of a pool of molecules (e.g. a sample).

[0106] Embodiments

[0107] In the following different aspects of the invention are defined in more detail. Each aspect so defined may be combined with any other aspect or aspects unless clearly indicated to the contrary. In particular, any feature indicated as being preferred or advantageous may be combined with any other feature or features indicated as being preferred or advantageous.

[0108] Method for identifying one or more differential MHC-presented peptides

[0109] In a first aspect, the present invention provides an in vitro method for identifying a differential MHC-presented peptide comprising the following steps:

[0110] (a) providing a sample or group of samples of clinically relevant tissue, (b) isolation and analysis of RNA, proteins, and / or DNA in the sample or group of samples of clinically relevant tissue,

[0111] (c) isolation and analysis of MHC-presented peptides in the sample or group of samples used in step (b),

[0112] (d) matching the paired data obtained in steps b) and c) against quantitative human reference data generated as described in steps b) and c), wherein the quantitative human reference data is based on at least 25 donors; and

[0113] (e) identification of the differential MHC-presented peptide on the basis of the matched data of step (d).

[0114] The method of the first aspect of the invention relies on the combination of measuring MHC-presented peptides in a sample with other parameters in these samples to generate paired data. Examples of such other parameters are the analysis of RNA, proteins or DNA. Thus the MHC-presented peptides and the other parameters are measured in the same sample (or in parts of the same sample). The use of paired data allows to corroborate the differential effect on additional biological levels (e.g. proteomics, transcriptomics, genomics) and thus establishing consistent evidence and confirming clinical relevance of a differential MHC-presented peptide. The data used in the method relating to the MHC-presented peptide or the other parameters (e.g. RNA, protein, DNA) can be either measured for individual species or on a large-scale using omics data.

[0115] In a preferred embodiment the analysis of the parameters in (b) and (c) are large-scale studies (preferably on an ‘omics’ level). This is particularly preferred for the generation of the quantitative human reference data. For large-scale studies of RNA preferably methods such as RNA-seq or microarray studies are used. For large scale studies of proteins, preferably mass spectrometry is used. For large scale studies of DNA preferably DNA-sequencing or microarrays are used. In a preferred embodiment, the analysis of the parameters in (b) and (c) is on a small scale, in particular for measurements not used for the generation of the quantitative human reference data.

[0116] In a preferred embodiment, the analysis of step (b) comprises the analysis of at least 1, at least 5, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 10,000, at least 20,000 or at least 50,000 entities of each of RNA, protein and DNA, independently of each other.

[0117] In a preferred embodiment, the analysis of step (c) comprises the analysis of at least 1, at least 5, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 10,000, at least 20,000 or at least 50,000 MHC-presented peptides. In a preferred embodiment the MHC-presented peptides of step (c) are analyzed by mass spectrometry.

[0118] In a preferred embodiment of the first aspect of the present invention, the clinically relevant tissue is affected by cancer, inflammation and / or auto-immune disease. In a preferred embodiment of the first aspect of the present invention, the clinically relevant tissue is affected by cancer. This is particularly preferred for the sample or group of samples in which the differential MHC-presented peptide(s) are identified.

[0119] In a preferred embodiment of the first aspect of the present invention, the clinically relevant tissue is a benign tissue. In a preferred embodiment of the first aspect of the present invention, the clinically relevant tissue is a normal tissue (i.e. benign or healthy tissue). In another preferred embodiment the clinically relevant tissue contains benign tissue and tissue affected by cancer, inflammation and / or auto-immune disease. This is particularly preferred for the sample or group of samples used to obtain the quantitative human reference data. In a preferred embodiment of the first aspect of the present invention, the clinically relevant tissue used to generate the quantitative human reference data is a benign tissue. In another preferred embodiment the sample or group of samples of clinically relevant tissue are affected by disease (e.g. by cancer, inflammation and / or auto-immune disease). In another preferred embodiment the sample or group of samples of clinically relevant tissue not used to generate the quantitative human reference data are affected by disease (e.g. by cancer, inflammation and / or auto-immune disease). Cancer is used herein in the broadest sense and includes metastatic cancer, advanced cancer, unresectable cancer, recurrent cancer and refractory cancer. Although evident for the skilled person it is pointed out that also solid tumors are encompassed.

[0120] In a preferred embodiment of the first aspect of the present invention, a group of samples is provided in step (a) and used in steps (b) and (c).

[0121] In a preferred embodiment of the first aspect of the present invention, the group of samples comprises at least 2 samples, at least 3 samples, at least 4 samples, at least 5 samples, at least 10 samples, at least 15 samples, at least 20 samples, at least 25 samples, at least 30 samples, at least 35 samples, at least 40 samples, at least 45 samples or at least 50 samples. Preferably the group of samples comprises at least 25 samples. More preferably the group of samples comprises 20-30 samples.

[0122] In a preferred embodiment of the first aspect of the present invention, the group of samples used to obtain the quantitative human reference data comprises samples of at least 20 donors, at least 25 donors, at least 26 donors, at least 27 donors, at least 28 donors, at least 29 donors, at least 30 donors, at least 35 donors, at least 40 donors, at least 45 donors, at least 50 donors, at least 75 donors, at least 100 donors, at least 150 donors, at least 200 donors, at least 250 donors, at least 300 donors, at least 400 donors, at least 500 donors, at least 750 donors, at least 1000 donors, at least 1500 donors, at least 2000 donors. Preferably, the group of samples used to obtain the quantitative human reference data comprises samples of at least 25 donors. Preferably, the group of samples used to obtain the quantitative human reference data comprises samples of 25 to 100, 30 to 75, or 35 to 50 donors.

[0123] In a preferred embodiment of the first aspect of the invention the quantitative human reference data is based on at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 samples per donor.

[0124] In a preferred embodiment of the first aspect of the invention the quantitative human reference data is based on at least 30 samples (at least 35, at least 40, at least 45, at least 50 samples) obtained from at least 20 donors (at least 25, at least 30 donors).

[0125] In a preferred embodiment of the first aspect of the invention the quantitative human reference data is based on at least 30 samples obtained from at least 20 donors.

[0126] In a preferred embodiment of the first aspect of the present invention, RNA is isolated from the sample of clinically relevant tissue. The isolated RNA may comprise all types of RNA, including but not being limited to mRNA, pre-mRNA, small nuclear RNA, microRNA, ribosomal RNA and circular RNA.

[0127] In a preferred embodiment of the first aspect of the present invention, protein is isolated from the sample of clinically relevant tissue. The isolated protein may comprise full length proteins, polypeptides, and / or peptides.

[0128] In a preferred embodiment of the first aspect of the present invention, DNA is isolated from the sample of clinically relevant tissue.

[0129] In a preferred embodiment of the first aspect of the present invention, protein and RNA is isolated from the sample of clinically relevant tissue.

[0130] In a preferred embodiment of the first aspect of the present invention, protein and DNA is isolated from the sample of clinically relevant tissue.

[0131] In a preferred embodiment of the first aspect of the present invention, DNA and RNA is isolated from the sample of clinically relevant tissue.

[0132] In a preferred embodiment of the first aspect of the present invention, protein, DNA and RNA is isolated from the sample of clinically relevant tissue.

[0133] In a preferred embodiment of the first aspect of the present invention, the analysis of protein is performed by mass spectrometry, preferable tandem mass spectrometry, more preferably tandem mass spectrometry combined with liquid chromatography. In a preferred embodiment of the first aspect of the present invention, the analysis of protein is performed by microarray based methods.

[0134] In a preferred embodiment of the first aspect of the present invention, the analysis of protein is performed by immunochemistry.

[0135] In a preferred embodiment of the first aspect of the present invention, the analysis of protein is a (semi) quantitative analysis.

[0136] In a preferred embodiment of the first aspect of the present invention, the analysis of protein includes the whole proteome comprised in the sample of clinically relevant tissue.

[0137] In a preferred embodiment of the first aspect of the present invention, the analysis of RNA is performed by methods based on sequencing of the RNA, preferably by RNAseq.

[0138] In a preferred embodiment of the first aspect of the present invention, the analysis of RNA is performed by microarray based methods.

[0139] In a preferred embodiment of the first aspect of the present invention, the analysis of RNA is a (semi) quantitative analysis.

[0140] In a preferred embodiment of the first aspect of the present invention, the analysis of RNA includes the whole transcriptome comprised in the sample of clinically relevant tissue.

[0141] In a preferred embodiment of the first aspect of the present invention, the analysis of DNA is performed by methods based on sequencing of the RNA

[0142] In a preferred embodiment of the first aspect of the present invention, the analysis of DNA is performed by microarray based methods.

[0143] In a preferred embodiment of the first aspect of the present invention, the analysis of DNA is a (semi) quantitative analysis.

[0144] In a preferred embodiment of the first aspect of the present invention, the analysis of DNA is a qualitative analysis.

[0145] In a preferred embodiment of the first aspect of the present invention, the analysis of DNA includes the whole genome comprised in the sample of clinically relevant tissue.

[0146] In a preferred embodiment of the first aspect of the present invention, the sample of clinically relevant tissue is fresh or frozen tissue.

[0147] In a preferred embodiment of the first aspect of the present invention, the isolation of the MHC-presented peptides from the sample is performed by affinity purification, more preferably by immunoprecipitation using MHC-specific antibodies. Preferably the MHC- specific antibodies can specifically bind to a class of MHC-molecules or specific MHC- allotypes. In a preferred embodiment affinity purification is performed by using a multitude of different affinity binders (preferably antibodies or antigen-binding fragments thereof). Suitable examples of affinity binders can be found inter alia in Apps et al (Immunology; 2009 May; 127(l):26-39).

[0148] In a preferred embodiment the isolation of the MHC-presented peptides from the sample is performed by acidic elution of MHC peptides from the surface of intact cells. An example of a suitable method is disclosed in Sturm et al. (J. Proteome Res. 2021, 20, 1, 289-304).

[0149] In a preferred embodiment of the first aspect of the present invention, the analysis of the MHC-presented peptides is performed by mass spectrometry, preferable tandem mass spectrometry, more preferably tandem mass spectrometry combined with liquid chromatography.

[0150] In a preferred embodiment of the first aspect of the present invention, the differential MHC-presented peptides are enriched or exclusive for the sample of clinically relevant tissue.

[0151] In a preferred embodiment of the first aspect of the present invention, the differential MHC-presented peptides are enriched in the sample of clinically relevant tissue.

[0152] In a preferred embodiment of the first aspect of the present invention, the differential MHC-presented peptides are exclusive for the sample of clinically relevant tissue.

[0153] In a preferred embodiment, enrichment is expressed as a factor that indicates x-fold- enrichment of the MHC-presented peptide between the sample of clinically relevant tissue and the median (or mean) abundance in the quantitative human reference data.

[0154] In a preferred embodiment the peptide is enriched by a factor of at least 2, 3, 4, 5, 10, 20, 50, 100, 200 in the sample of clinically relevant tissue as compared to corresponding healthy tissue or the quantitative human reference data. In a more preferred embodiment the peptide is enriched by a factor of at least 2, 3, 4, 5 or 10 in the sample of clinically relevant tissue as compared to corresponding healthy tissue or the quantitative human reference data.

[0155] In a preferred embodiment, the differential MHC-presented peptide has a prevalence of at least 20% in the clinically relevant tissue.

[0156] In a preferred embodiment of the first aspect of the invention the quantitative human reference data is based on data obtained from at least 20 donors per MHC-allotype (preferably per MHC-allotype comprised in the quantitative human reference data). Preferably at least 20 donors, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100 donors per MHC-allotype.

[0157] Preferably in case of multivariate quantitative human reference data each of the multiple parameters was determined in at least 20 donors per MHC-allotype.

[0158] Preferably the MHC-allotype of the donors is defined to the level of the allele group, i.e. for example wherein the MHC-allotype is HLA-A*02 signifying not only the HLA-gene A but also the 02 allele group. More preferably the single MHC-allotype is defined to the level of the specific HLA protein (e.g. HLA-A*02:05).

[0159] Sensitivity is a key factor in target discovery as well as target validation since false negative identification in the reference data set will result in false positive target identification. Sufficient sensitivity is preferably ensured by use of data-independent acquisition (DIA) mass spectrometry. To enable discovery- mode DIA for immunopeptidomics it is necessary to predict spectra as well as peptides of the spectral library used for DIA. Yet, such libraries would be too large for an efficient search as well as not sensitive enough due to FDR estimation. Therefore, the methods of the invention according to all aspects preferably employ an MHC-specific empirically optimized library.

[0160] In a preferred embodiment, the quantitative human reference data is obtained exclusively with the identical methodology as the paired data obtained in steps b) and c).

[0161] In a preferred embodiment of the first aspect of the present invention, the analysis of MHC-presented peptides in step (c) comprises mass spectrometry analysis using data- independent acquisition (DIA), wherein analysis of DIA data comprises the use of at least one spectral library based on predicted MHC-presented peptides specific for a single MHC- allotype. Preferably the single MHC-allotype is defined at least to the level of the allele group (e.g. HLA-A*02). More preferably the single MHC-allotype is defined to the level of the specific HLA protein (e.g. HLA-A*02:01).

[0162] Data-independent acquisition (DIA) is a method used in mass spectrometry, wherein precursor ions are isolated in pre-defined isolation windows across a mass range of interest in a systematic and repeated manner and are fragmented simultaneously. Fragment ions of each isolation window are recorded by a mass analyzer, preferably by Orbitrap or time-of-fhght technology, typically resulting in complex multiplexed tandem mass spectra. Peptides are identified by matching the ion peaks in a mass spectrum to a spectral library that contains information of the peptide fragment ions’ pattern and optionally its chromatography elution time. DIA analysis is characterized by systematic acquisition of all fragment ions across a predefined mass range resulting in a permanent map of all detectable analytes present in a sample thus enabling deep coverage, high reproducibility, and high accuracy. The data resulting from DIA is more complex than that of other mass spectrometry methods such as for example data- dependent acquisition (DDA). In DIA data there is substantial overlap of ion spectra. Matching peaks with particular peptides is thus more difficult. In one particular preferred embodiment the DIA data obtained is analyzed by using a peptide spectral library, preferably the spectral library represents only (predicted) MHC-presented peptides. Preferably the spectral library represents only (predicted) MHC-presented peptides presented by a specific MHC allotype. Preferably the specific MHC-allotype is defined at least to the level of the allele group, i.e. for example wherein the MHC-allotype is HLA-A*02 signifying not only the HLA-gene A but also the 02 allele group. More preferably the single MHC-allotype is defined to the level of the specific HLA protein (e.g. HLA-A*02:01).

[0163] In a preferred embodiment the spectral library is a HLA-allotype specific spectra library comprising mass spectra of peptides predicted or demonstrated to be presented by a particular HLA-allotype. Preferred alleles for a HLA-allotype specific library are HLA-A*02:01, HLA- A*24:02, HLA-A*01:01, HLA-A*03:01, HLA-B*07:02, HLA-B*08:01, HLA-B*44:02a and HLA-B*44:03. Preferably the allele for a HLA-allotype specific library is HLA-A*02:01.

[0164] In another preferred embodiment the spectral library is a spectral library comprising peptides presented by all HLA-allotypes.

[0165] In a preferred embodiment the DIA data is analyzed by at least one (i.e. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10) peptide spectral library, preferably each peptide spectral library used in the analysis is limited to ion spectra representing peptides presented by a single MHC-allotype.

[0166] In a preferred embodiment the DIA data is analyzed by 1 to 10, preferably 1 to 8, 1 to 6, 1 to 4 or 1 to 3, spectral libraries, preferably each peptide spectral library used in the analysis is limited to ion spectra representing peptides presented by a single MHC-allotype.

[0167] In a preferred embodiment the DIA data is analyzed by 2 to 10, preferably 2 to 8, 2 to 6, 2 to 4 or 2 to 3, spectral libraries, preferably each peptide spectral library used in the analysis is limited to ion spectra representing peptides presented by a single MHC-allotype.

[0168] In a preferred embodiment of the first aspect of the present invention, the method of the second aspect of the invention may be used to analyze DIA data of the first aspect of the invention. Preferably the analysis of MHC-presented peptides in step (c) can involve the collection of DIA data, which can be analyzed with the method of the second aspect of the invention.

[0169] In a preferred embodiment of the first aspect of the present invention, the quantitative human reference data is univariate or multivariate reference data. In a preferred embodiment the quantitative human reference data is univariate reference data. For univariate quantitative human reference data it is preferred that the reference data is based on MHC-presented peptide expression or nucleic acid expression, preferably MHC-presented peptide expression. The univariate quantitative human reference data may in another embodiment be based on nucleic acid expression or DNA variation. In a more preferred embodiment the quantitative human reference data is multivariate reference data. Preferably in multivariate reference data MHC-presented peptide expression is one of the multiple parameters. Other suitable parameters that may be included are nucleic acid expression and DNA variation. Nucleic acid expression preferably refers to expression of mRNA, pre-mRNA, small nuclear RNA, microRNA, ribosomal RNA and circular RNA. DNA variations typically refer to variations in the genomic DNA. In one preferred embodiment DNA variations are selected from: single nucleotide variants (SVs), insertion and / or deletion of nucleic acids and copy number variations (CNV). Protein expression preferably refers to the expression of full length proteins. In a preferred embodiment, the quantitative human reference data is paired data.

[0170] In a preferred embodiment the quantitative human reference data is multivariate reference data based on MHC-presented peptide expression and nucleic acid expression (preferably RNA expression). In another preferred embodiment the quantitative human reference data is multivariate reference data based on MHC-presented peptide expression and DNA variations. In another preferred embodiment the quantitative human reference data is multivariate reference data based on MHC-presented peptide expression and protein expression.

[0171] In another preferred embodiment the quantitative human reference data is multivariate reference data based on MHC-presented peptide expression, nucleic acid expression (preferably RNA expression) and DNA variations.

[0172] In another preferred embodiment the quantitative human reference data is multivariate reference data based on MHC-presented peptide expression, nucleic acid expression (preferably RNA expression) and protein expression.

[0173] In another preferred embodiment the quantitative human reference data is multivariate reference data based on MHC-presented peptide expression, nucleic acid expression (preferably RNA expression), protein expression and DNA variations.

[0174] In a preferred embodiment of the first aspect of the invention the quantitative human reference data is based on samples of clinically relevant tissue and / or benign / healthy tissue.

[0175] In a preferred embodiment of the first aspect of the invention the quantitative human reference data is based on samples of clinically relevant tissue.

[0176] In a preferred embodiment of the first aspect of the invention the quantitative human reference data is based on samples of benign / healthy tissue.

[0177] In a preferred embodiment of the first aspect of the invention the quantitative human reference data is based on samples of clinically relevant tissue and benign / healthy tissue. In a preferred embodiment of the first aspect of the invention the quantitative human reference data is based on samples of clinically relevant tissue or benign / healthy tissue.

[0178] In a preferred embodiment of the first aspect of the present invention, the analysis in step (b) (i.e. analysis of RNA, proteins and / or DNA) is performed by methods that allow the analysis of the complete set of analytes in the sample of clinically relevant tissue. Preferably RNA is analyzed by RNA- sequencing and protein is analyzed by mass spectrometry, more preferably LC-MS / MS.

[0179] In a preferred embodiment of the first aspect of the present invention, the clinically relevant issue is tissue affected by cancer, inflammation and / or auto-immune disease.

[0180] In a preferred embodiment of the first aspect of the present invention, the clinically relevant issue is tissues affected by cancer.

[0181] In a preferred embodiment of the first aspect of the present invention, the clinically relevant issue is tissues affected by inflammation.

[0182] In a preferred embodiment of the first aspect of the present invention, the clinically relevant issue is tissues affected by auto-immune disease.

[0183] Analysis of data independent acquisition (DIA) data

[0184] In a second aspect, the present invention provides a method for the analysis of a data- independent acquisition (DIA) dataset obtained by mass spectrometry analysis of MHC- presented peptides in a biological sample comprising the steps of:

[0185] (i) providing at least one peptide spectral library based on predicted MHC- presented peptides specific for a single MHC allotype;

[0186] (ii) analyzing the DIA-dataset based on the at least one peptide spectral library of step (i).

[0187] The second aspect of the present invention, regards the analysis of data-independent acquisition (DIA) datasets resulting from mass spectrometry of MHC-presented peptides. The analysis of DIA data comprises the use of at least one peptide spectral library based on predicted MHC-presented peptides specific for a single MHC allotype. Preferably the single MHC- allotype is defined at least to the level of the allele group (e.g. HLA-A*02). More preferably the single MHC-allotype is defined to the level of the specific HLA protein (e.g. HLA- A*02:05). Any method known in the art can be used to predict the MHC-presented peptides. In a preferred embodiment the MHC-presented peptides are not predicted by experimentally determined spectral libraries. In the context of the present invention DIA data obtained is preferably analyzed by using a peptide spectral library, wherein the spectral library represents only (predicted) MHC- presented peptides. Preferably the spectral library represents only (predicted) MHC-presented peptides presented by a specific MHC allotype. Preferably the specific MHC-allotype is defined at least to the level of the allele group, i.e. for example wherein the MHC-allotype is HLA- A*02 signifying not only the HLA-gene A but also the 02 allele group. More preferably the single MHC-allotype is defined to the level of the specific HLA protein (e.g. HLA-A*02:05).

[0188] In a preferred embodiment the spectral library is a HLA-allotype specific spectra library comprising mass spectra of peptides predicted or demonstrated to be presented by a particular HLA-allotype. Preferred alleles for a HLA-allotype specific library are HLA-A*02:01, HLA- A*24:02, HLA-A*01:01, HLA-A*03:01, HLA-B*07:02, HLA-B*08:01, HLA-B*44:02a and HLA-B*44:03. Preferably the allele for a HLA-allotype specific library is HLA-A*02:01.

[0189] In another preferred embodiment the spectral library is a spectral library comprising peptides presented by all HLA-allotypes.

[0190] In a preferred embodiment the DIA data is analyzed by at least one (i.e. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 etc.) peptide spectral library, preferably each peptide spectral library used in the analysis is limited to ion spectra representing peptides presented by a single MHC-allotype.

[0191] In a preferred embodiment the DIA data is analyzed by 1 to 10, preferably 1 to 8, 1 to 6, 1 to 4 or 1 to 3, spectral libraries, preferably each peptide spectral library used in the analysis is limited to ion spectra representing peptides presented by a single MHC-allotype.

[0192] In a preferred embodiment the DIA data is analyzed by 2 to 10, preferably 2 to 8, 2 to 6, 2 to 4 or 2 to 3, spectral libraries, preferably each peptide spectral library used in the analysis is limited to ion spectra representing peptides presented by a single MHC-allotype.

[0193] In a preferred embodiment of the second aspect of the invention, the at least one peptide spectral library comprises predicted ion spectra based on (preferably predicted) peptide sequences for MHC-presented peptides and optionally empirical optimization.

[0194] In a preferred embodiment of the second aspect of the invention, the peptide spectral library is constructed by a method comprising the following steps:

[0195] (I) obtaining a protein database representing the proteome of the biological sample to be analyzed;

[0196] (II) applying a prediction algorithm to the protein database of step (I) for generating a database of potential MHC-presented peptides,

[0197] (III) generating a spectral library based on the database of potential MHC-presented peptides. In a preferred embodiment of the second aspect of the invention, the prediction algorithm used in the method of the invention is selected from: MHCflurry (see O’Donnell et al, Cell Systems, vol 11(1), p.42-48, 2020), NetMHC (see Jurtz et al, J Immunol. 2017 Nov l;199(9):3360-336; EP application 23191424.3).

[0198] A preferred example of a suitable method for generating the spectral library uses a database of (preferably predicted) MHC-presented peptides (as in step (II)). Based on this sequence database a spectral library consisting of target and decoy spectra and associated retention times is generated by prediction. Chromatograms for each target or decoy peptide in the database are extracted from the DIA data, and elution peaks for each peptide and negative control are identified and scored. For each peptide the elution peak with the highest score is selected, optionally taking the interference between multiple precursors matched to the same retention time into account.

[0199] An algorithm (preferably a deep neural network algorithm) is trained with the selected elution peaks to distinguish between MHC-presented peptides and decoy controls. The trained networks can then assign to experimental data a statistical significance score to enable comprehensive coverage of the database of potential MHC-presented peptides at strict FDR (false discovery rate) thresholds.

[0200] In a preferred embodiment of the second aspect of the invention, the peptide spectral library is generated by a method comprising the steps:

[0201] (a) based on the database of potential MHC-presented peptides of step (II) a database of negative controls is generated;

[0202] (b) a chromatogram for each MHC-presented peptide and negative control is generated and elution peaks are identified therefrom and scored;

[0203] (c) the elution peaks with the highest scores are selected;

[0204] (d) a deep neural network algorithm is applied to the selected elution peaks to distinguish between the MHC-presented peptides and negative controls resulting in a peptide spectral library.

[0205] A preferred method for generating the spectral library is the use of DIA-neural network (DIA-NN; Demichev et al., Nat Methods. 2020 January; 17(1): 41-44.), preferably DIA-NN in library free mode. More preferably using DIA-NN with MBR (‘match-between-runs’).

[0206] Preferably library generation in step (III) further includes empirical optimization of the spectral library. A preferred example of a suitable empirical optimization method creates a peptide spectral library from predicted MHC-presented peptides and processes the entire DIA dataset using match-between-runs to restrict the predicted sequence database to identifiable HLA peptides and optimize / recalibrate the corresponding predicted spectra and retention times.

[0207] A preferred example of an empirical optimization method is ‘match-between-runs’ as described for example in https: / / github.eom / vdemichev / diann#match-between-runs. Preferably the empirical optimization method uses a two-step process in which a dataset is searched with an initial spectral library (preferably predicted MHC-presented peptides specific for a single MHC allotype) and then the same dataset is reprocessed with an optimized spectral library that results from incorporating only hits from the initial search in order to reduce the search space.

[0208] In a preferred embodiment the protein database is restricted to cancer (preferably tumor) associated proteins.

[0209] Although evident for the skilled person it is pointed out that the subject-matter of the described aspects can be combined.

[0210] It is envisaged that the differential MHC-presented peptides or pharmaceutically acceptable salts thereof identified by the herein described methods are used as a medicament e.g. as an anti-cancer vaccine. Accordingly, the invention relates to a differential MHC- presented peptide identified by the herein described methods or a pharmaceutically acceptable salt thereof for use in the (manufacture of a medicament for the) treatment of cancer or a tumorous disease and / or disorder, such as a solid tumor.

[0211] It is further envisaged that the information obtained by the herein described methods is used to develop cancer therapies. Accordingly, when a differential MHC-presented peptide is identified, an antigen-binding protein may be provided targeting said peptide. Said antigenbinding protein may be used for the treatment of a disease or clinical condition. An antigen binding protein may be an antibody, a TCR, or combination thereof. The antigen binding protein may also be a multi-specific (e.g. a bispecific) protein, for example comprising variable domains of an TCR and an antibody, such as an TCER® .

[0212] For example, when the differential analysis (using the herein described methods) determines peptides that show (pronounced) differences in e.g. presentation by MHC between a sample type of clinical interest (e.g. various tumor types or autoimmune diseases) and unaffected / benign samples and / or a reference dataset said peptides may be used as targets for anti-cancer-immunotherapy. In other words, when the differential analysis (using the herein described methods) determines a peptide that shows (pronounced) differences in e.g. presentation by MHC between a sample type of clinical interest (e.g. various tumor types or autoimmune diseases) and unaffected / benign samples and / or a reference dataset an antigenbinding protein may be developed targeting said peptide. Especially when by the herein described methods a peptide is identified that is only presented on cancer cells and not on healthy tissue or the presentation on cancer cells is significantly increased compared to healthy tissue said peptide or an antigen-binding protein targeting said peptide may be used for the treatment of the corresponding cancer.

[0213] Accordingly, the invention relates to an (in vitro) method for identifying an antigenbinding protein binding to a peptide (preferably bound to a MHC protein) identified by the herein described methods comprising: a) contacting a plurality of antigen-binding proteins with a peptide (preferably bound to a MHC protein) identified by the herein described methods, b) identifying an antigen-binding protein that binds to said peptide (preferably bound to a MHC protein), and c) selecting the antigen-binding protein identified in b).

[0214] Furthermore, the invention relates to an (in vitro) method for providing an antigenbinding protein binding to a peptide (preferably bound to a MHC protein) identified by the herein described methods comprising: a) providing an antigen-binding protein, b) contacting the antigen-binding protein with a peptide (preferably bound to a MHC protein) identified by the herein described methods, b) identifying an antigen-binding protein that binds to said peptide (preferably bound to a MHC protein), and c) selecting the antigen-binding protein identified in b).

[0215] The invention also relates to an antigen-binding protein binding to a peptide (preferably bound to a MHC protein) identified by the herein described methods for use as a medicament, preferably for use in the treatment of cancer.

[0216] The antigen-binding protein identified as described above may be expressed in a host cell. In particular when the antigen-binding protein is a TCR said TCR may be expressed in a host cell and presented by said host cell. Said host cell may be used as a medicament, preferably for use in the treatment of cancer.

[0217] It is evident for the skilled person that the described peptide, antigen-binding protein or host cell may be formulated in a pharmaceutical composition. Said pharmaceutical composition may be used as a medicament, preferably for use in the treatment of cancer.

[0218] It is also envisaged that by the herein described methods a therapeutic modality is determined. Example 7 shows that by the herein described methods a therapeutic window for a TCR- T approach for the treatment of a given cancer and a given peptide can be determined. Accordingly, the invention relates to a method for identifying a therapeutic modality, wherein a peptide is identified by the herein described methods.

[0219] Examples

[0220] The present invention relates to a method for identifying differential MHC-presented peptides through population-scale immunopeptidomics paired with orthogonal omics or single species measurements.

[0221] General methods:

[0222] For immunopeptidomics analysis in the following examples data-independent acquisition (DIA) mass spectrometry was used to ensure sensitivity and robustness (see Example 1 and Figure 1). Each scan cycle consisted of 37 Orbitrap MS2 scans covering the m / z range of 280-720m / z with dynamic isolation widths of 8 m / z, 16.2 m / z and 17.5 m / z, depending on precursor density across the m / z range. Precursor fragmentation was performed by collision- induced dissociation (CID) at 35% normalized collision energy (NCE) and higher collisional dissociation (HCD) at 27% NCE. To allow discovery- mode DIA data processing, DIA-NN (vl.8.1) was used with an empirically optimized spectral library consisting of predicted HLA peptides as described below for Example 2. Novel peptide findings based on this approach are depicted in Example 9 (Figures 10-12).

[0223] Immunopeptidomics data was paired with proteomics, transcriptomics and genomics measurements by measuring homogenous aliquots of the same sample used for immunopeptidomics.

[0224] For transcriptomics, total RNA was isolated and purified. RNA samples with a concentration of > 25 ng / pl and a RNA Integrity Number (RIN) or RNA Quality Number (RQN) of > 6.0 were considered acceptable for downstream library preparation and sequencing. Library preparation, including mRNA selection, RNA fragmentation, cDNA conversion and addition of sequencing adaptors, as well as the sequencing process itself were performed by GENEWIZ Germany GmbH (Leipzig, Germany). Sequencing libraries were generated using the NEBNext® UltraTM II Directional RNA Library Prep Kit for Illumina (New England Biolabs, Ipswich, MA, USA) following the manufacturer’s instructions. For sequencing, libraries were multiplexed and loaded onto the Illumina NovaSeq6000 sequencer (Illumina, San Diego, CA, USA) according to the manufacturer’s instructions, targeting a minimum of 80 million paired-end reads per sample with a read length of 150 bp.

[0225] The obtained reads were trimmed using BBDuk (BBTools v38.81, March 24, 2020) to remove adapter sequences and low-quality bases. Subsequently, the trimmed reads were then mapped to the Genome Reference Consortium Human Build 38 patch release 13 (GRCh38.pl3) using STAR (v.2.7.3a, Dobin et al 2013, Bioinformatics 29: 15-21) and the exon expression was quantified using featureCounts (Liao et al 2014, Bioinformatics 30:923-930) For quantifying transcript expression Kallisto (Bray et al 2016, Nature Biotechnology 34:525-527) was utilized. For the quantifications the Ensembl 99 (Yates et al 2020, Nucleic Acids Research 48:D682-D688) reference annotations were used and the results were normalized using DEseq (Anders et al 2010, Genome Biol 11 :R106) to a predefined reference sample.

[0226] For proteomics, total protein was isolated from respective samples in CHAPS buffer. After cell lysis, protein concentration was spectrophotometrically determined using BCA assay. To acquire untargeted proteomic data, samples were processed using the SP3 protocol (Hughes et al 2019, Nat Protoc 14:68-85). First, an isotopically labeled protein surrogate for monitoring proteolytic efficiency was added to the protein lysate. Subsequently, disulfide bridges were disrupted using TCEP along with CAA for irreversible blocking of free thiol groups. Proteolytic breakdown was induced by the addition of trypsin / LysC mix followed by incubation overnight at 37°C. Samples were analyzed on an Evosep One (Evosep, Odense, Denmark) LC system online coupled to an Orbitrap Eclipse Tribrid MS (Thermo Fisher Scientific, Waltham, MA, USA) instrument. A standardized 44 min gradient (Evosep 30 SPD) in combination with 15cm x 150pm EVI 106 columns was used for separation. Untargeted proteomics data were acquired for each sample using DIA mode, acquired with HCD-FT / FT at 27 NCE. Full MS spectra (350- 1650 m / z) were acquired in OT at a resolution of 120,000 at m / z 200 and AGC target value of 300% with a maximum IT of 100 ms. For fragment spectra (120-1800 m / z), OT resolution was set to 30,000 at m / z 200, the AGC target value to 1000% with an IT of 54 ms. Each MS cycle consisted of 44 DIA MS / MS scans with different isolation windows, along the entire mass range. DIA data were searched against the human Ensembl 99 database using DIA-NN with the following settings: enzyme set to trypsin with maximum of 2 missed cleavages, peptide length range between 7 and 35, Cys carbamidomethylation was set as fixed and Met oxidation defined as variable modification. Protein quantity was assessed using the relative iB AQ method as described elsewhere (Krey et al. 2018, Sci Data 5: 180128). Through the paired architecture of the approach, the target effect can be corroborated on additional biological levels and thus confirming consistent evidence and clinical relevance of a peptide (see Figure 3).

[0227] The paired data architecture of the approach was applied to the samples of interest as well as to a population-scale reference set of donors. This reference empowers the differential analysis and determination of peptides that show pronounced differences in terms of fold change between a sample type of clinical interest (e.g. various tumor types or autoimmune diseases) and the reference dataset. Since immunotherapies address certain HLA-peptide targets it is essential to understand the differential abundance on the peptide level to judge their therapeutic window and potential off-tumor toxicity.

[0228] Reference size is relevant due to the donor and HLA variation that needs to be covered. Considering the three examples below, the inventors determined 25 donors for a given HLA allele to be sufficient to achieve a median average power of 80%. The following table depicts the power in all three examples for varying donor numbers. The numbers indicated in brackets for Example 5 refer to the use of DDA instead of DIA.

[0229] Table ESummary of the power of Example 4, 5 and 6.

[0230] The reference data set should be sized in a way that the following objectives can be achieved.

[0231] • Enable detection of rare peptides to ensure safety with regard to detection of on- target and off-target toxicity (see Example 4 and Figure 4). • Enable empirical optimization of spectral library for DIA data analysis (see Example 5 and Figure 5A).

[0232] • Enable correlation analysis between peptide data and surrogate data (protein, mRNA, DNA) to allow biomarker discovery and investigation of mechanism of action for a immunopeptide target (see Example 6 and Figure 6).

[0233] The last objective is in particular relevant for patient stratification in clinical settings because measurement by small scale methods such as qPCR for RNA or protein by IHC is preferred. To be able to utilize such measurements for screening for peptide-positive patients, it is necessary to establish a quantitative relationship between the immunopeptidome and the surrogate protein, RNA or DNA measurement. This again ideally requires a certain number of data points (see Example 6 and Figure 6).

[0234] Beyond target and biomarker discovery, the use of multi-omics reference data provides the means to select the appropriate therapeutic modality by deciding on the most promising biological entity to address a given target (Example 7 and Figure 7). Over-expression is analyzed for peptide, protein and transcriptome level enabling to select the appropriate therapeutic modality (e.g. CAR-T in case of over-expression on protein level vs. TCR-T in case of over-expression on peptide level).

[0235] Furthermore, in yet another use this reference data set allows to discover off-target peptides without the need for additional laborious experimentation (Example 8).

[0236] In summary, the presented invention provides an approach to target discovery and validation of MHC-presented peptides that extends existing approaches in the following way:

[0237] • Sensitive Differential Immunopeptidomics using DIA for immunopeptidomics (Example 1 and Figure 1).

[0238] • Discovery-mode DIA for immunopeptidomics by empirical optimization of MHC-predicted spectral library DIA (Examples 2 and 9 and Figures 2, 10, 11 and 12).

[0239] • Consistency assurance of differential effect by incorporation of paired omics (or single entity level) measurements (Example 3 and Figure 3).

[0240] • Ability to select appropriate biomarker by enabling correlation between the different omics measurements, e.g. mRNA biomarker allowing prediction of MHC peptide presentation (Example 6 and Figure 6). Ability to select appropriate therapeutic modality, e.g. CAR-T in case of overexpression on protein level vs. TCR-T in case of over-expression on peptide level (Example 7 and Figure 7).

[0241] In the following examples the different uses and their advantages are exemplified.

[0242] Example 1: DIA for improved sensitivity.

[0243] To illustrate the sensitivity of the method of the present invention, the inventors determined the limit of detection (LOD) for the peptide GLIDEVMVL (SEQ ID NO: 006) from cytochrome P450 2W1 (CYP2W1) using untargeted DDA and DIA as well as absolute quantitation targeted MS data (AbsQuant®, US10545154B2) acquired from matched samples.

[0244] The LOD was estimated using the probit method as recommended in CLSI EP17-A2 (see Evaluation of Detection Capability for Clinical Laboratory Measurement Procedures, 2nd Edition; 2012; ISBN 1-56238-796-0).

[0245] Using DIA, the inventors observed an increased sensitivity lowering the LOD from 77.96 fmol (DDA; Figure 1 A) to 8.3 frnol (DIA; Figure IB).

[0246] To investigate this reduction in LOD more in general, the inventors inspected 45 peptides that had at least 60 matched DDA, DIA and AbsQuant® data points for LOD estimation. The average reduction in LOD was about 10 fold (log2(DDA / DIA) = -3.5). This reduction is depicted as a scatter plot (Figure 1C, one point is one peptide) as well as a box plot (Figure ID).

[0247] The use of the DIA method disclosed herein thus allows to significantly increase the sensitivity of peptide detection by mass spectrometry as used in the method of the invention.

[0248] In the presented example, untargeted DDA and DIA LC-MS data were acquired on a nanoAcquity UPLC system (Waters, Milford, USA) coupled to an Orbitrap Tribrid mass spectrometer (Thermo Fisher, Waltham, USA). A trapping setup using Waters 25 cm x 75 ,m BEH C18 analytical columns was used employing a stepped gradient ranging from 1 to 34.5 acetonitrile. MSI scan range was set to 200-1500 m / z with an AGC target of le5 and 120k MSI scan resolution.

[0249] For DDA analysis, data-dependent MS2 scans were acquired using a “top speed” method with a maximum cycle time of 3 s. Precursor isolation width was set to 2 m / z for the topN precursors within a m / z range of 280-720 m / z and charge states of 2+ and 3+ at 30k resolution and an automatic gain control (AGC) target of 5e4 in the Orbitrap. Precursor fragmentation was performed by collision-induced dissociation (CID) at 35% normalized collision energy (NCE) and higher collisional dissociation (HCD) at 27% NCE. For ion trap (IT) runs, MS2 spectra were acquired with identical precursor selection and isolation settings and CID at 35% NCE was performed with le4 AGC target in the IT with the scan speed set to normal.

[0250] For determination of absolute peptide amount in a given sample AbsQuant® targeted MS analysis was performed as described in US10545154B2. In brief, the total peptide amount in a sample is determined by nanoLC-MS / MS (scheduled parallel reaction monitoring in the ion trap) by calculating the ratio of the peptide variant of interest and a fixed amount of an isotope-labelled version of the peptide, the so-called internal standard, which is added after elution and prior to MS analysis to the sample. The efficiency of peptide isolation - during immunoprecipitation and further sample processing prior to LC / MS analysis - was determined by spiking of refolded peptide-HLA complexes of the investigated peptide into the sample lysate at the earliest possible point of time during the peptide isolation procedure for all assays.

[0251] DDA data processing was performed using the search engines X! Tandem, MSGF+ and Comet against the human Ensembl 77 database without enzymatic restriction. Precursor mass tolerance was set to lOppm. For XITandem and MSGF+ fragment mass tolerance was set to 15 ppm and 600 ppm for Orbitrap and Ion Trap MS2 spectra, respectively. For Comet, fragment mass bin width was set to 1.0005 Da with bin offset of 0.4 Da for Ion Trap MS2 spectra and to 0.02 Da with bin offset of 0.0 Da for Orbitrap MS2 spectra, respectively. Peptide length was restricted to 6-15 amino acids length and methionine oxidation was set as variable modification. The search results were integrated and false discovery rates limited to 1% using iProphet (Shteynberg, D. et al. 2011 / For peptide quantitation, MSI features were extracted using SuperHirn (Mueller, L. N. et al. 2007).

[0252] Example 2: Predicted peptides for Discovery-mode DIA.

[0253] To illustrate the benefits of using large scale reference data to refine peptide discovery from DIA data, the inventors generated a spectral library of predicted HLA peptides and empirically optimized it for DIA data analysis by incorporating 100 MS runs (comprising of 84 unique HLA-A*02:01 donors) using the match-between-runs (MBR) option of DIA-NN. The performance of this library was compared against a current state-of-the-art approach for generating predicted HLA peptide spectral libraries, AlphaP eptDeep.

[0254] For creation of the initial predicted HLA peptide spectral library, DIA-NN was used to predict peptide fragmentation patterns from all 8-11 mer peptides found in Ensembl 99 with a MHCFlurry %rank presentation of <5.0. The precursor and fragment m / z ranges were set to 280-720 and 150-1250, respectively. Oxidation of methionine was set as a variable modification with up to 2 modifications per peptide, and DIA-NN’s “smart profiling” mode was enabled.

[0255] This large library was then used in an MBR search by DIA-NN of 100 HLA-A*02:01 MS runs (84 unique HLA-A*02:01 donors) to optimize the library for the A*02:01 search space. MBR is a two-step process in which multiple runs are first searched with a large library (the initial library described above), and the results from each run are compiled into a smaller, refined library. In the MBR search, DIA-NN was set to use individual mass accuracies and retention time windows for each run, cross-run normalization was disabled, and protein inference was switched off.

[0256] For comparison, a spectral library of predicted HLA peptides was created using the state-of-the-art approach described in (Zeng et al 2022, Nat Commun 13 :7238) as follows. DIA- Umpire (Tsou et al 2015, Nat Methods 12:258-264) was used to extract MS2 spectra from the MS runs and the spectra were searched using standard open-source database search engines. The peptides identified at 1% FDR from this library-free search were used to train a run-specific ligand predictor using AlphaP eptDeep. This run-specific ligand-predictor was then used to predict HLA peptides for each run. The peptides with probability greater than 0.7 were then used to build a spectral library using AlphaP eptDeep ’s built in predictors.

[0257] Compared against the libraries created with the AlphaPeptDeep procedure, the use of the library created using the method of the present invention resulted in a 40% improvement (1.4 fold) for donors with more than 1000 peptides and more than 140% improvement (2.4 fold) for donors with less than 1000 peptides (see Figure 2). The pronounced improvement in sensitivity at the lower end is of particular relevance for comprehensive characterization of healthy tissues, since these frequently express lower levels of HLA.

[0258] Example 3: Paired multi-omics for orthogonal evidence

[0259] The relevance of paired data (in particular omics data) may be best depicted with the identification of the peptide YSGQLKVLI (SEQ ID NO: 001) derived from ankyrin repeat domain 30A (ANKRD30A).

[0260] This gene was described in the prior art as breast cancer antigen (NY-BR-1); see e.g. Jager et al, 2001, Cancer Res 61 :2055-2061. This observation is consistent with gene expression data obtained by the inventors (see Figure 3 middle panel). Nevertheless, the peptide presentation data (see Figure 3 top panel) indicates a gastric cancer specific target and does not match with the observed profile on the mRNA level. In contrast, an alternative peptide KILDTVHSC (SEQ ID NO: 002) derived from the same gene displayed a pattern in the peptide presentation profile (see Figure 3 bottom panel) as seen with the transcriptomic data. Finally, the experimental follow-up of the peptide sequence revealed that the sequence was wrongly identified which is in line with the observations deduced from the paired data set. The use of paired data thus allows to confirm the correctness of identified peptides.

[0261] Example 4: Population-scale for rare peptide detection

[0262] Peptides with low prevalence are quite frequent in context of target selection and particularly critical in context of off-target effects. There are various reasons for low prevalence, one of which is that the peptide might be derived from a SNP resulting in potentially low minor allele frequency (MAF). The HLA-A*25:01 peptide YVDGTQFVRF (SEQ ID NO: 003) derived from rs9266183 (D4G, MAF 3.3%) in HLA-B was used as a representative example in Figure 4 showing the statistical power (probability to detect this peptide) depending on the number of donors included in the reference data. In a simulation donors were selected at random (sub sampling). This selection was repeated 1000 times and the statistical power in terms of probability of successful detection was determined. At 22 donors a power of more than 80% was achieved (see Table 1).

[0263] Example 5: Population-scale for high peptide coverage

[0264] Mass spectrometry data analysis is another aspect which requires a certain data size. Peptide-centric analysis for data-independent mass spectrometry can be optimized by using existing raw data measurements to calibrate machine learning algorithms to enhance identification rate in DIA and DDA data analysis. For this to work it is necessary to cover a relevant portion of the immunopeptidome and address the biological variation.

[0265] To simulate the effect of data size the inventors ran another simulation which looked at the coverage of peptides achieved when using only a fraction of the available data. Using repeated subsampling of 1 to 1708 donors measured by mass spectrometry, the inventors counted the number of unique total peptides covered. For DIA data, at 25 donors a statistical power of more than 80% was achieved meaning that the approach detected more than 80% of the possible immunopeptidome (see Table 1 and Figure 5A). To achieve a similar power of more than 80% using DDA data, at least 86 donors were required (see Figure 5B). Example 6: Population-scale for biomarker discovery

[0266] To further illustrate the relevance of paired data and the connection to the required data size, the inventors ran a simulation study. In this study the inventors further demonstrated the identification of a biomarker using the methods of the invention.

[0267] The HLA-A*02:01 peptide KLSEIDVAL (SEQ ID NO: 007) derived from EF-hand domain-containing protein DI (EFHD1) shows a high correlation (R=0.64) between peptide and mRNA in 421 samples. In a biomarker discovery setting, an association is not known in advance. This might be simulated by using the known connection between peptide and coding gene and trying to recover this by correlating all protein-coding genes against the peptide abundance and ranking them according to p-value. Success can be defined if the coding gene (EFHD1) is amongst the top 100 genes providing focus on follow-up experiments for biomarker screening. This success depends on the data size used in this biomarker discovery analysis. Repeated testing of subsampled datasets of 5 to 500 donors for which the peptide was identified (bootstrapping analysis) allows to estimate the power in terms of success rate in dependence on data size (see Figure 6). Thus aiming at a power of at least 80% requires to include 32 donors in the analysis. The ability to determine appropriate biomarkers (e.g. mRNA biomarkers) predictive for peptide presentation is therefore another use of the methods of the present invention.

[0268] Example 7: Selection of therapeutic modality

[0269] Ability to select appropriate therapeutic modality, e.g. CAR-T in case of overexpression on protein level vs. TCR-T in case of over-expression on peptide level. The presentation profiles for the peptide KVFAGIPTV (SEQ ID NO: 004) obtained with DIA (see Figure 7 second panel) show a clear therapeutic window for a TCR-T approach for the treatment of glioblastoma (GB), the fold change computed for GB versus healthy brain or healthy tissues in general is about 11.95 Similarly, the transcriptomic expression pattern (see Figure 7 bottom panel) points in the same direction (7.47). In contrast, when looking at proteomics data for the corresponding source gene PTPRZ1 (see Figure 7 third panel, relative iBAQ intensity depicted as ppm), no therapeutic window becomes obvious with a fold change GBM vs healthy brain of only 0.88. The analysis is based on data obtained for 807 donors.

[0270] Example 8: Identification of relevant off-target peptides

[0271] A more complex use case underscoring the relevance of a comprehensive reference dataset is the identification and characterization of potential off-target peptides, which may be cross-recognized by HLA binding moieties such as TCRs intended for application in immunotherapy. In the present example, peptides cross-recognized by a soluble TCR-bispecific molecule were experimentally determined in a binder capture assay that led to the identification of 25 potential off-targets, i.e. 9-mer peptides that are potentially bound by the lead molecule. Further characterization of these off-targets in binding kinetics measurements showed 14 / 25 peptides to be bound by the lead molecule with potentially relevant affinity, leading to the characterization of their presentation profiles and associated therapeutic windows using the reference dataset. Of note, investigation of these off-targets in a smaller external dataset (HLA Ligand Atlas) only identified two of these 14 off-targets. Figure 8 displays the HLA peptide presentation and source exon RNA expression of the cross-recognized peptide ILGTHNITV (SEQ ID NO: 005) from INTS7, which shows broad and prevalent peptide presentation (top panel) and RNA expression (bottom panel) across a range of relevant normal tissues. The fact that this broadly presented off-target is missing from the smaller external dataset illustrates the relevance of the reference dataset described herein for the discovery and characterization of relevant off-targets.

[0272] The experimental identification and validation of off-targets described above was performed as disclosed in WO 2021 / 028503 AL In brief, the TCR-bispecific molecule was coupled to a solid Sepharose® matrix and used for immunoaffinity purification of HLA peptides from T98G cells. The resultant peptide extracts were analyzed by LC-MS / MS as described above.

[0273] Example 9: Novel peptide findings by disclosed DIA method

[0274] A limitation of data-dependent acquisition (DDA) is its dependence on a relatively clear MSI signal to trigger the acquisition of an MS2 fragment ion spectrum in tandem mass spectrometry. If a peptide does not provide a MSI signal or the signal is extremely low, for example due to the peptide being low in abundance, then the peptide will not be detected using DDA. This has been disclosed in above Example 1.

[0275] This has an impact on the construction of spectral libraries for data-independent acquisition (DIA) of immunopeptidomics samples, because the standard approach for immunopeptidomics in the art is to use DDA identifications to build a DIA library. However, a peptide that is not detectable in DDA then will not appear in such an empirically derived DIA library.

[0276] The predicted library approach described in Example 2 above overcomes this limitation. To illustrate the benefits of the method in Example 2, the inventors selected a peptide which was detected in DIA using a predicted library analogous to that described in Example 2 above. This peptide could not be detected in DDA, and which has -to the best of the inventor’s knowledge- never been described in the art to be detected by mass spectrometry.

[0277] The selected peptide has the amino acid sequence VLLGVKLSGV (SEQ ID NO: 008) and is publicly disclosed in IEDB, the standard public database of immune epitopes (see https: / / iedb.org / epitope / 859569). It has been studied for immune reactivity, tested in T cell assays and MHC binding ligand assay. There is -to the best of the inventor’s knowledge- no public disclosure of it ever having been detected by mass spectrometry. However, using a predicted library and DIA (the predicted library being analogous to that described in Example 1), the inventors have identified it in 129 tested samples.

[0278] The peptide of SEQ ID NO: 008 has very low MSI intensity, meaning it would be very difficult to detect in DDA. In fact, % of the DIA detections using the method of the present invention are MS2-only, meaning there was no MSI signal at all for the peptide in the respective run. Compared to the respective run-average MSI signal the signal of the peptide is very low (see Figure 10).

[0279] The MS2 fragmentation spectrum of the peptide of SEQ ID NO: 008 from a DIA run is shown in Figure 11. The acquired spectrum, and the prediction from Prosit, a state-of-the-art fragmentation predictor, agree very well. Comparison of the quantitative profile of the peptide with the RNAseq-based gene expression profile of the corresponding gene COL18A1 (Figure 12) reveals a high correspondence between both data sets highlighting the correct identification of the peptide. The peptide as well as the gene show highest expression in normal liver as well as hepatocellular carcinoma (HCC).

Claims

Claims1. An in vitro method for identifying a differential MHC-presented peptide comprising the following steps:(a) providing a sample or group of samples of clinically relevant tissue,(b) isolation and analysis of RNA, proteins, and / or DNA in the sample or group of samples of clinically relevant tissue,(c) isolation and analysis of MHC-presented peptides in the sample or group of samples used in step (b),(d) matching the paired data obtained in steps b) and c) against quantitative human reference data generated as described in steps b) and c), wherein the quantitative human reference data is based on at least 25 donors; and(e) identification of the differential MHC-presented peptide on the basis of the matched data of step (d).

2. The method according to claim 1, wherein the quantitative human reference data comprises data on at least one MHC-allotype and wherein each MHC-allotype in the quantitative human reference data is based on the analysis of at least 20 samples of this MHC-allotype.

3. The method according to any one of the preceding claims, wherein the analysis of MHC- presented peptides in step (c) comprises mass spectrometry analysis using data- independent acquisition (DIA), wherein analysis of DIA data comprises the use of at least one peptide spectral library based on predicted MHC-presented peptides specific for a single MHC allotype.

4. The method according to any one of the preceding claims, wherein the isolation of MHC-presented peptides of the sample or group of samples of clinically relevant tissue in step (c) comprises: affinity purification with one or more binders specific for a class of MHC- molecules or a specific MHC-allotypes; or acidic elution of MHC peptides.

5. The method according to any one of the preceding claims, wherein the quantitative human reference data is univariate or multivariate reference data, preferably whereinthe human reference data comprises data on MHC-presented peptide expression, nucleic acid expression, protein expression, and / or DNA variation.

6. The method according to any one of the preceding claims, wherein the quantitative human reference data is combined with further datasets obtained with other methods, preferably wherein the further datasets comprise gene expression data.

7. The method according to any one of the preceding claims, wherein the analysis in step (b) is performed by methods that allow the analysis of the complete set of analytes in the sample of clinically relevant tissue, preferably by RN A- sequencing, DNA- sequencing and LC-MS / MS.

8. The method according to any one of the preceding claims, wherein the clinically relevant issue is tissues affected by cancer, inflammation, auto-immune disease.

9. The method according to any one of claims 3 to8, wherein the peptide spectral library is obtained by a method comprising the following steps:(I) obtaining a protein database representing the proteome of the biological sample to be analyzed;(II) applying a prediction algorithm to the protein database of step (I) for generating a database of potential MHC-presented peptides,(III) generating a peptide spectral library based on the database of potential MHC- presented peptides.

10. The method according to claim 9, wherein the generation of the peptide spectral library further includes empirical optimization of the spectral library.

11. The method according to any one of claims 3 to 10, wherein the spectral library comprises of target and decoy spectra and associated retention times.

12. The method according to any one of claims 9 to 11, wherein the peptide spectral library is generated by a method comprising the steps:(a) based on the database of potential MHC-presented peptides of step (II) a database of negative controls is generated;(b) a chromatogram for each MHC-presented peptide and negative control is generated and elution peaks are identified therefrom and scored;(c) the elution peaks with the highest scores are selected;(d) a deep neural network algorithm is applied to the selected elution peaks to distinguish between the MHC-presented peptides and negative controls resulting in a peptide spectral library.

13. The method according to any one of claims 9 to 12, wherein the spectral library is generated using DIA-NN, preferably in library free mode.

14. A peptide as identified by the method according to any one of claims 1 to 13.

15. An (in vitro) method for identifying an antigen-binding protein binding to a peptide (preferably bound to a MHC protein) identified by the method according to any one of claims 1 to 13 comprising: a) contacting a plurality of antigen-binding proteins with said peptide (preferably bound to a MHC protein), b) identifying an antigen-binding protein that binds to said peptide (preferably bound to a MHC protein), and c) selecting the antigen-binding protein identified in b).

16. An antigen-binding protein as obtained by the method according to claim 15.

17. A host cell comprising the antigen-binding protein according to claim 16.

18. A pharmaceutical composition comprising the peptide according to claim 14, the antigen-binding protein according to claim 16 or the host cell according to claim 17.

19. The peptide according to claim 14, the antigen-binding protein according to claim 16, the host cell according to claim 17 or the pharmaceutical composition according to claim 18 for use as a medicament.

20. The peptide according to claim 14, the antigen-binding protein according to claim 16, the host cell according to claim 17 or the pharmaceutical composition according to claim 18 for use in the treatment of cancer.

Citation Information

Patent Citations

  • Method for the absolute quantification of naturally processed HLA-restricted cancer peptides

    US10545154B2

  • Method for the characterization of peptide:MHC binding polypeptides

    WO2021028503A1

  • Method for differentially quantifying naturally processed HLA-restricted peptides for cancer, autoimmune and infectious diseases immunotherapy development

    US9791444B2

  • EP23191424A