Absolute copy number prediction of MHC-presented peptides by relative quantitative measurement
By employing statistical modeling of relative quantitative measurements of surrogate parameters, the method addresses the challenge of determining absolute copy numbers of MHC-presented peptides with reduced effort and material, enhancing the efficiency of immunotherapy and personalized medicine.
Patent Information
- Application Number
- PCT/EP2024/086147
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-13
- Filing Date
- 2024-12-13
- Publication Date
- 2025-06-19
AI Technical Summary
Current methods for determining the absolute copy numbers of MHC-presented peptides are labor-intensive and require significant biological material, whereas relative quantitative measurements are less demanding but do not provide absolute quantities.
A method that uses relative quantitative measurements of surrogate parameters to predict the absolute copy numbers of MHC-presented peptides through statistical modeling, allowing for the establishment of statistical models that can predict absolute quantities in new biological samples.
Enables the prediction of absolute copy numbers of MHC-presented peptides with reduced experimental effort and sample material, facilitating decision-making in immunotherapy and personalized medicine.
Smart Images

Figure IMGF000038_0001 
Figure IMGF000040_0001 
Figure IMGF000050_0001
Abstract
Description
[0001] Absolute copy number prediction of MHC-presented peptides by relative quantitative measurement
[0002] The present invention relates to a method for the absolute quantification of MHC- presented peptides by measuring relative quantities in a biological sample.
[0003] Background of the Invention
[0004] The major histocompatibility complex (MHC) plays a critical role in the immune system of vertebrates, including humans. The MHC genes encode proteins that are present on the surface of cells and are responsible for the presentation of antigens to the immune system. Proteins are constantly synthesized and proteosomally degraded within cells and short peptides derived from these degraded proteins are presented by the MHC molecules on the cell surface. These MHC-presented peptides can be detected by T lymphocytes with their T cell receptor (TCR) and non-self antigens may initiate immune reactions.
[0005] The entirety of MHC-presented peptides is highly complex and dynamic, varying with the type of cell, its developmental stage, and the presence of diseases or infections. The total number of unique MHC-presented peptides of an individual is estimated to be in the millions.
[0006] MHC-presented peptides are of great interest as targets for immunotherapy to initiate an immune response against specific cells characterized by certain MHC-presented peptides. In this regard it is of importance to quantify specific MHC-presented peptides in a biological sample to determine if, for example, a particular tissue in a patient might benefit from a particular immunotherapy. In fact, such knowledge is highly relevant for decision making in immunotherapy and, in particular, in personalized immunotherapy.
[0007] However, experimental determination of absolute copy numbers of MHC-presented peptides is labor-intensive and requires sufficient biological starting material. On the other hand, acquisition of relative quantitative data requires less effort and sample material. Thus, it would be beneficial if such data can be used to determine the absolute quantity of MHC- presented peptides.
[0008] In the prior art efforts were made to utilize e.g., mRNA expression levels for patient and target selection (Fritsche et al., 2018, Proteomics, 18: 12; 1700284). Such methods do however not provide absolute copy numbers but do rather predict the presence or absence of an MHC- presented peptide in a given sample. The present invention provides a novel method that allows the prediction of absolute copy numbers of one or more MHC-presented peptides in a biological sample based on the relative quantification of a surrogate parameter.
[0009] Summary of the Invention
[0010] In a first aspect, the present invention provides a method for absolute quantification of at least one MHC-presented peptide in a further biological sample comprising the following steps:
[0011] (a) determining the absolute quantity of the at least one MHC-presented peptide in at least two, preferably at least four, biological samples;
[0012] (b) determining in the at least two, preferably at least four, biological samples of step (a):
[0013] (i) the relative quantity of the at least one MHC-presented peptide; and / or
[0014] (ii) the relative or absolute quantity of at least one surrogate for each of the at least one MHC-presented peptide;
[0015] (c) applying statistical modelling to the at least one absolute quantity of step (a) and the at least one relative or absolute quantity of step (b) to establish at least one statistical model;
[0016] (d) determining in the further biological sample the at least one relative or absolute quantity as in step (b);
[0017] (e) predicting the absolute quantity of the at least one MHC-presented peptide in the further biological sample using the at least one statistical model.
[0018] In particular the invention relates to the following items:
[0019] 1. A method for absolute quantification of at least one MHC-presented peptide in a further biological sample comprising the following steps:
[0020] (a) determining the absolute quantity of the at least one MHC-presented peptide in at least two, preferably at least four, biological samples;
[0021] (b) determining in the at least two, preferably at least four, biological samples of step (a):
[0022] (i) the relative quantity of the at least one MHC-presented peptide; and / or (ii) the relative or absolute quantity of at least one surrogate for each of the at least one MHC-presented peptides;
[0023] (c) applying statistical modelling to the at least one absolute quantity of step (a) and the at least one relative or absolute quantity of step (b) to establish at least one statistical model;
[0024] (d) determine in the further biological sample the at least one relative or absolute quantity as in step (b);
[0025] (e) predict the absolute quantity of the at least one MHC-presented peptide in the further biological sample using the at least one statistical model. A method for generating a statistical model for absolute quantification of at least one MHC-presented peptide in a further biological sample comprising the following steps:
[0026] (a) determining the absolute quantity of the at least one MHC-presented peptide in at least two, preferably at least four, biological samples;
[0027] (b) determining in the at least two, preferably at least four, biological samples of step (a):
[0028] (i) the relative quantity of the at least one MHC-presented peptide; and / or
[0029] (ii) the relative or absolute quantity of at least one surrogate for each of the at least one MHC-presented peptides;
[0030] (c) applying statistical modelling to the at least one absolute quantity of step (a) and the at least one relative or absolute quantity of step (b) to establish at least one statistical model. The method of any of the preceding items, wherein the surrogate is a primary surrogate or a secondary surrogate. The method of item 3, wherein the primary surrogate is derived from:
[0031] (I) the nucleic acid sequence encoding the at least one MHC-presented peptide, preferably selected from gene expression, genomic variation, transcript variation, alternative splicing; or
[0032] (II) the protein of which the at least one MHC-presented peptide is derived from, preferably selected from protein expression. The method of item 3 or 4, wherein the secondary surrogate is another parameter relevant for the at least one MHC-presented peptide, preferably selected from HLA, proteasome, immunoproteasome, beta-2 microglobulin, TAP and ERAP. The method of any of the preceding items, wherein the secondary surrogate is the HLA that is involved in presenting the at least one MHC-presented peptide. The method of any of the preceding items, wherein in step (c): only one quantity of step (b) is used for statistical modelling (i.e. univariate analysis); or at least two quantities of step (b) are used for statistical modelling (i.e. multivariate analysis). The method of any of the preceding items, wherein, in case of determining more than one quantity in step (b) (i.e. multivariate analysis), this includes the quantity of at least one primary surrogate or the relative quantity of the at least one MHC-presented peptide, optionally further including at least one secondary surrogate. The method of any of the preceding items, wherein the more than one quantity determined in step (b) (i.e. multivariate analysis) includes: relative quantity of the at least one MHC-presented peptide and relative mRNA expression; or relative mRNA expression and genomic variation. The method of any of the preceding items, wherein the more than one quantity determined in step (b) (i.e. multivariate analysis) includes the relative quantity of
[0033] - the at least one MHC-presented peptide and
[0034] - the specific HLA- Allele that is presenting the at least on MHC-presented peptide. The method of any of the preceding items, wherein the statistical modelling of step (c) is a regression analysis, preferably linear regression. 12. The method of any of the preceding items, wherein the statistical modelling of step (c), preferably regression analysis, comprises the determination of peptide-specific off-set parameter, preferably the peptide-specific off-set parameter is determined based on: physicochemical and biochemical properties of amino acids, preferably selected from amino acid index or Hydrophobicity score; peptide retention time in liquid chromatography; sequence specific peptide bias, or binding affinity.
[0035] 13. The method according to any one of the preceding items, wherein the absolute quantity of more than one MHC-presented peptide is predicted by establishing a universal statistical model in step (c).
[0036] 14. An antigen-binding protein binding to the at least one MHC-presented peptide, at least one MHC-presented peptide or the pharmaceutical salt thereof for use in the treatment of a subject suffering from cancer, wherein the absolute quantity of at least one MHC-presented peptide in a further biological sample of the subject is determined according to any one of the preceding items, wherein the absolute quantity of at least one MHC-presented peptide is indicative of the suitability to treat the subject with an antigen-binding protein binding to the at least one MHC-presented peptide, the at least one MHC-presented peptide or the pharmaceutical salt thereof, wherein an antigen binding protein binding to the at least one MHC-presented peptide, the at least one MHC-presented peptide or the pharmaceutical salt thereof is to be administered to the subject.
[0037] In addition the invention relates to the following items:
[0038] 1. A method for treating a patient or determining a treatment option / modality comprising:
[0039] (a) determining an absolute quantity of at least one MHC-presented peptide in at least two, optionally at least four, biological samples of the patient;
[0040] (b) determining in the at least two biological samples of step (a):
[0041] (i) a relative quantity of the at least one MHC-presented peptide; and / or (ii) a relative or absolute quantity of at least one surrogate for each of the at least one MHC-presented peptides;
[0042] (c) training at least one statistical model with the absolute quantity of the at least one MHC-presented peptide of step (a) and the relative quantity of (i) and / or the relative or absolute quantity of (ii) to establish at least one trained statistical model;
[0043] (d) determining in a further biological sample of the patient, the relative quantity of (i) and / or the relative or absolute quantity of (ii) as in step (b);
[0044] (e) predict the absolute quantity of the at least one MHC-presented peptide in the further biological sample using the at least one trained statistical model; and
[0045] (f) treating the patient or determining a treatment option / modality based on the predicted absolute quantity of the at least one MHC-presented peptide of step (e). A method for generating a statistical model for absolute quantification of at least one MHC-presented peptide in a further biological sample comprising the following steps:
[0046] (a) determining the absolute quantity of the at least one MHC-presented peptide in at least two, preferably at least four, biological samples;
[0047] (b) determining in the at least two, preferably at least four, biological samples of step (a):
[0048] (i) the relative quantity of the at least one MHC-presented peptide; and / or
[0049] (ii) the relative or absolute quantity of at least one surrogate for each of the at least one MHC-presented peptides;
[0050] (c) applying statistical modelling to the at least one absolute quantity of step (a) and the at least one relative or absolute quantity of step (b) to establish at least one statistical model. The method of item 1, wherein the relative or absolute quantity of (ii) is determined, and wherein the surrogate is a primary surrogate or a secondary surrogate. The method of item 3, wherein the surrogate is the primary surrogate, and wherein the primary surrogate is derived from: (I) a nucleic acid sequence encoding the at least one MHC-presented peptide, optionally selected from gene expression, genomic variation, transcript variation, alternative splicing; or
[0051] (II) a protein of which the at least one MHC-presented peptide is derived from, optionally selected from protein expression. The method of item 3, wherein the surrogate is the secondary surrogate, and wherein the secondary surrogate is another parameter relevant for the at least one MHC-presented peptide, optionally selected from HLA, proteasome, immunoproteasome, beta-2 microglobulin, TAP and ERAP. The method of item 5, wherein the secondary surrogate is the HLA that is involved in presenting the at least one MHC-presented peptide. The method of item 1, wherein in step (c): only one quantity of step (b) is used for training (i.e. univariate analysis); or at least two quantities of step (b) are used for training (i.e. multivariate analysis). The method of item 1, wherein, in case of determining more than one quantity in step (b) (i.e. multivariate analysis), the determining of step (b) includes determining a quantity of at least one primary surrogate or a relative quantity of the at least one MHC- presented peptide, optionally further including determining at least one secondary surrogate. The method of item 1, wherein, in case of determining more than one quantity in step (b) (i.e. multivariate analysis), the determining of step (b) includes: relative quantity of the at least one MHC-presented peptide and relative mRNA expression; or relative mRNA expression and genomic variation. The method of item 1, wherein, in case of determining more than one quantity in step (b) (i.e. multivariate analysis), the determining of step (b) includes a relative quantity of
[0052] - the at least one MHC-presented peptide and
[0053] - a specific HLA- Allele that is presenting the at least on MHC-presented peptide. The method of item 1, wherein the training of step (c) comprises a regression analysis, optionally linear regression. The method of any of item 1, wherein the training of step (c), optionally comprising regression analysis, further comprises a determination of peptide-specific off-set parameter, optionally the peptide-specific off-set parameter is determined based on: physicochemical and biochemical properties of amino acids, optionally selected from amino acid index or hydrophobicity score; peptide retention time in liquid chromatography; sequence specific peptide bias, or binding affinity. The method according to item 11 , wherein the ab solute quantity of more than one MHC- presented peptide is predicted by establishing a universal statistical model in step (c). An antigen-binding protein binding to the at least one MHC-presented peptide, at least one MHC-presented peptide or the pharmaceutical salt thereof for use in the treatment of a subject suffering from cancer, wherein the absolute quantity of at least one MHC-presented peptide in a further biological sample of the subject is determined according to any one of the preceding items, wherein the absolute quantity of at least one MHC-presented peptide is indicative of the suitability to treat the subject with an antigen-binding protein binding to the at least one MHC-presented peptide, the at least one MHC-presented peptide or the pharmaceutical salt thereof, wherein an antigen binding protein binding to the at least one MHC-presented peptide, the at least one MHC-presented peptide or the pharmaceutical salt thereof is to be administered to the subject. List of Figures
[0054] In the following, the content of the figures comprised in this specification is described. In this context please also refer to the detailed description of the invention below.
[0055] Figure 1: refers to peptide-specific regression models using surrogates based on relative peptide quantity (example 1). Univariate regression model for prediction of copies per cell (CpC) for Peptide 1 based on relative peptide quantity.
[0056] Figure 2: refers to universal regression models using surrogates based on relative peptide quantity (example 2). Universal regression model for prediction of copies per cell (CpC) for Peptide 1 based on relative peptide quantity.
[0057] Figure 3: refers to peptide-specific regression models using surrogates based on relative protein quantity (example 3). Univariate regression models for prediction of copies per cell (CpC) for Peptide 4 based on relative protein quantity.
[0058] Figure 4: refers to peptide-specific regression models using surrogates based on gene expression (example 4). Univariate regression models for prediction of copies per cell (CpC) for Peptide 3 based on gene expression.
[0059] Figure 5: refers to Peptide-specific regression models using surrogates based on transcript expression (example 5). Univariate regression models for prediction of copies of per cell (CpC) for Peptide 3 based on transcript expression.
[0060] Figure 6: refers to peptide-specific regression models using surrogates based on exon expression (example 6). Univariate regression models for prediction of copies per cell (CpC) for Peptide 2 based on exon expression.
[0061] Figure 7: refers to peptide-specific regression models using surrogates derived from peptide-encoding nucleotide sequences (PNS) (example 7). Univariate regression models for prediction of copies per cell (CpC) of Peptide 2 based on PNS.
[0062] Figure 8: refers to universal regression models using surrogates based on transcript expression (example 8). Multivariate regression models for prediction of copies per cell (CpC) for Peptide 2 based on transcript expression.
[0063] Figure 9: refers to peptide-specific regression models using surrogates based on qPCR (example 9). Univariate regression models for prediction of copies per cell (CpC) for Peptide 005 based on gene quantity derived from qPCR. Detailed Descriptions of the Invention
[0064] Before the present invention is described in detail, it is to be understood that this invention is not limited to the particular methodology, protocols and reagents described herein as these may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the present invention which will be limited only by the appended claims. Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art.
[0065] Throughout this specification and the appended claims, unless the context requires otherwise, the word "comprise", and variations such as "comprises" and "comprising", will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps. In the following passages, different aspects of the invention are defined in more detail. Each aspect so defined may be combined with any other aspect or aspects unless clearly indicated to the contrary. In particular, any feature indicated as being optional, preferred or advantageous may be combined with any other feature or features indicated as being optional, preferred or advantageous.
[0066] Several documents are cited throughout the text of this specification. Each of the documents cited herein (including all patents, patent applications, scientific publications, manufacturer's specifications, instructions etc.), whether supra or infra, is hereby incorporated by reference in its entirety. Nothing herein is to be construed as an admission that the invention is not entitled to antedate such disclosure by virtue of prior invention. Some of the documents cited herein are characterized as being “incorporated by reference ” . In the event of a conflict between the definitions or teachings of such incorporated references and definitions or teachings recited in the present specification, the text of the present specification takes precedence.
[0067] In the following, the elements of the present invention will be described. These elements are listed with specific embodiments; however, it should be understood that they may be combined in any manner and in any number to create additional embodiments. The variously described examples and preferred embodiments should not be construed to limit the present invention to only the explicitly described embodiments. This description should be understood to support and encompass embodiments which combine the explicitly described embodiments with any number of the disclosed and / or preferred elements. Furthermore, any permutations and combinations of all described elements in this application should be considered disclosed by the description of the present application unless the context indicates otherwise. Definitions
[0068] In the following, some definitions of terms frequently used in this specification are provided. These terms will, in each instance of their use, in the remainder of the specification have the respectively defined meaning and preferred meanings.
[0069] As used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents, unless the content clearly dictates otherwise.
[0070] The term “about” when used in connection with a numerical value is meant to encompass numerical values within a range having a lower limit that is 5% smaller than the indicated numerical value and having an upper limit that is 5% larger than the indicated numerical value.
[0071] The “major histocompatibility complex” (MHC) in the context of the present invention is a set of cell surface proteins essential for the acquired immune system to recognize foreign molecules in vertebrates, which in turn determines histocompatibility. The main function of MHC molecules is to bind to antigens derived from pathogens and display them on the cell surface for recognition by the appropriate T cells. The human MHC is also called the HLA (human leukocyte antigen) complex (often just the HLA). The MHC gene family is divided into three subgroups: class I, class II, and class III. Complexes of peptide and MHC class I are recognized by CD8-positive T cells bearing the appropriate T cell receptor (TCR), whereas complexes of peptide and MHC class II molecules are recognized by CD4- positive-helper-T cells bearing the appropriate TCR. Since both types of response, CD8 and CD4 dependent, contribute jointly and synergistically, the identification and characterization of MHC-presented peptides and corresponding T cell receptors is important in the development of immunotherapies such as vaccines and cell therapies. The HLA-A gene is located on the short arm of chromosome 6 and encodes the larger, a-chain, constituent of HLA-A. Variation of HLA-A a-chain is key to HLA function. This variation promotes genetic diversity in the population. Since each HLA has a different affinity for peptides of certain structures, greater variety of HLAs means greater variety of antigens to be 'presented' on the cell surface. Each individual can express up to two types of HLA-A, one from each of their parents. Some individuals will inherit the same HLA-A from both parents, decreasing their individual HLA diversity; however, the majority of individuals will receive two different copies of HLA-A. This same pattern follows for all HLA groups.
[0072] “HLA-A*02:01” signifies a specific HLA allele or allotype, wherein the letter A signifies the gene and the suffix “02” indicates the allele or allotype group and “01” a specific HLA allele or allotype. The rules for nomenclature of HLA proteins is well known in the art and can for example be found at htps: / / hla.alleles.0rg / n0menclature / naming.html.A “MHC- presented peptide” in the context of the present invention is a peptide that is presented by an MHC molecule and that can be bound by a binding moiety (in particular a TCR or an antibody). Preferably an MHC class I presented peptide has a length of 8 to 11 amino acids, preferably 9 to 10, most preferably 9 amino acids. Preferably an MHC class II presented peptide has a length of 13 to 25 amino acids.
[0073] The term “immunoproteasome” refers to a specialized version of the proteasome, a cellular complex responsible for degrading proteins that are no longer needed or are damaged. By generating peptides that can be presented by major histocompatibility complex (MHC) molecules, the immunoproteasome plays a critical role in antigen processing and presentation to T cells.
[0074] The terms “absolute quantity” and “relative quantity” refer to the amount of a particular substance (such as a specific MHC-presented peptide) present in a sample. Absolute quantity may refer to a concentration in terms of a known absolute unit (such as copy number or copy number per cell).
[0075] The terms “absolute quantification” and “relative quantification” are used to describe methods of measuring / determining the amount of a particular substance (such as a specific MHC-presented peptide) present in a sample. Absolute quantification may refer to a method of measuring the amount of a substance in a sample by determining its concentration in terms of a known absolute unit (such as copy number or copy number per cell). Relative quantification may refer to a method of comparing the relative abundance of a particular substance between different samples or conditions. Relative quantification typically involves the use of a reference sample or condition to which the other samples are compared or an internal reference / control. The relative abundance of the substance in each sample is then determined by comparing its signal to the signal of the reference sample or the internal reference / control. Thus, absolute quantification measures the amount of a substance in a sample in absolute terms, while relative quantification compares the relative abundance of a substance between different samples or conditions. In certain embodiments the relative quantity is expressed in relation to a reference (reference sample or internal reference / control) that is known to account for the experimental variation. Relative quantification and absolute quantification may be performed as described in the appended examples. In other words, the relative quantity and absolute quantity may be determined by the methods described in the appended examples. For relative quantification for example, the signal intensity of a specific MHC-presented peptide detected in mass spectrometry may be normalized to the signal intensity of an internal control or a reference sample. The internal control may be a synthetic spike-in peptide(s) added in fixed amounts. Accordingly, the relative quantity of an MHC-presented peptide in context of the invention may be the normalized signal intensity of a mass spectrometry experiment. The unit of the signal intensity may be a relative proportion or arbitrary units.
[0076] The term “surrogate” generally refers to a substitute or alternative that is used to represent a particular variable. In the context of the present invention, it preferably refers to a substitute or proxy for the quantity of an MHC-presented peptide. A surrogate has a relation to the parameter it substitutes (such as for example the quantity of an MHC-presented peptide), which allows to make certain predictions about the actual parameter based on the value of the surrogate. A surrogate may be a quantitative measure (e.g. an expression value) or a qualitative measure (e.g. variations in the genome or transcript). The relationship between one or more surrogates (and additional variables) and the actual parameter can be established by statistical modelling.
[0077] The term “statistical model” in the context of the present invention refers to a mathematical model that includes a set of statistical assumptions concerning the prediction of sample data. A statistical model thus specifies a mathematical relationship between one or more variables and therefore allows for the prediction of variables in case only some of the variables are known due to their relationship. The term “statistical modelling” thus refers to the process of establishing a statistical model. Non-limiting examples for methods of statistical modelling include regression analysis.
[0078] The term “regression analysis” refers to a statistical method used to estimate the relationship between one or more independent variables (also called predictor variables) and a dependent variable (also called the response variable). The goal of regression analysis is to create a statistical model that can accurately predict the value of the dependent variable based on the values of the independent variables. In a preferred embodiment of the present invention the dependent variable is the absolute quantity of the at least one MHC-presented peptide in a biological sample. In a preferred embodiment of the present invention the independent variables are preferably those determined in steps (a) and (b).
[0079] There are many types of regression analysis known in the art. The most preferred regression analysis in the context of the present invention is linear regression, which assumes that the relationship between the variables is linear. In general, linear regression estimates the slope and intercept of the line that best fits the data and uses these estimates to predict the value of the dependent variable for a given set of values of the independent variable. The term “biological sample” may refer to any type of sample derived from an organism including but not limited to tissue samples, blood, isolated cells, cultured cells, cell lines and xenografts.
[0080] Disclosed embodiments may involve use of one or more computers. A computer as used in this disclosure may include a general purpose computer, a personal computer, a workstation, a mainframe computer, a notebook, a global positioning device, a laptop computer, a smart phone, a personal digital assistant, a network server, and any other electronic device that may interact with a user to develop programming code.
[0081] In some embodiments, any of the one or more computers may include at least one processor, a memory, and other components including those components that facilitate electronic communication. Other components may include user interface devices such as an input and output devices. Any of the one or more computers may include computer hardware components such as a combination of Central Processing Units (CPUs) or processors, buses, memory devices, storage units, data processors, input devices, output devices, network interface devices, and other types of components that will become apparent to those skilled in the art. Any of the one or more computers may further include application programs that may include software modules, sequences of instructions, routines, data structures, display interfaces, and other types of structures that execute operations of the present disclosure.
[0082] The at least one processor may be implemented in hardware, firmware, or a combination of hardware and software. The at least one processor may be one or more of a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and another type of processing component. The at least one processor may include one or more processors capable of being programmed to perform a function.
[0083] The memory may include a random access memory (RAM), a read only memory (ROM), and / or another type of dynamic or static storage device (e.g., a flash memory, a magnetic memory, and / or an optical memory) that stores information and / or instructions for use by the at least one processor. The memory may also store information and / or software related to the operation and use of the one or more computers. For example, the memory may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optic disk, and / or a solid state disk), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cartridge, a magnetic tape, and / or another type of non-transitory computer-readable medium, along with a corresponding drive. The one or more computers may perform one or more processes described herein. The one or more computers may perform operations based on the at least one processor executing software instructions stored by a non-transitory computer-readable medium, such as the memory. A computer-readable medium is defined herein as a non-transitory memory device. A memory device includes memory space within a single physical storage device or memory space spread across multiple physical storage devices.
[0084] Software instructions may be read into the memory from another computer-readable medium or from another device via, for example, one or more transceivers. When executed, software instructions stored in the memory may cause the at least one processor to perform one or more processes described herein.
[0085] Additionally, or alternatively, hardwired circuitry may be used in place of or in combination with software instructions to perform one or more processes described herein. Thus, embodiments described herein are not limited to any specific combination of hardware circuitry and software.
[0086] Disclosed embodiments may involve receiving through a communication network. A communication network as used in this disclosure may include a set of computers (such at least some of the one or more computers) sharing resources located on or provided by network nodes. This set of computers may use common communication protocols over digital interconnections to communicate with each other. These interconnections may be made up of telecommunication network technologies, based on physically wired, optical, and wireless radio-frequency methods that may be arranged in a variety of network topologies. For example, these interconnections may take place through databases, servers, RF (radio frequency) signals, cellular technology, Ethernet, telephone, “TCP / IP” (transmission control protocol / intemet protocol), and any other electronic communication format. For example, the network 130 may include a cellular network (e.g., a fifth generation (5G) network, a long-term evolution (LTE) network, a third generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., the Public Switched Telephone Network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber optic-based network, or the like, and / or a combination of these or other types of networks.
[0087] The number and arrangement of computers and networks may be adjustable.
[0088] In some embodiments, the communications network may be set up as a neural network. A neural network may be based on a collection of connected units or nodes called artificial neurons, which loosely model the neurons in a biological brain. Each connection, like the synapses in a biological brain, may transmit a signal to other neurons. An artificial neuron may receive signals to process and may then signal other neurons connected to it. These signals at a connection may be real numbers, and the output of each neuron may be computed by some nonlinear function of the sum of its inputs. These connections may be edges. Neurons and edges may have a weight that adjusts as learning proceeds. The weight may increase or decrease the strength of the signal at a connection. Neurons may have a threshold such that a signal may be sent only if the aggregate signal crosses that threshold. Neurons may be aggregated into layers. Different layers may perform different transformations on their inputs. Signals may travel from a first layer (e.g., an input layer), to a last layer (e.g., an output layer), through potential intermediate layers and may do so multiple times.
[0089] Embodiments
[0090] In the following different aspects of the invention are defined in more detail. Each aspect so defined may be combined with any other aspect or aspects unless clearly indicated to the contrary. In particular, any feature indicated as being preferred or advantageous may be combined with any other feature or features indicated as being preferred or advantageous.
[0091] Method for absolute quantification of at least one MHC-presented peptide
[0092] In a first aspect, the present invention provides a method for absolute quantification of at least one MHC-presented peptide in a further biological sample comprising the following steps:
[0093] (a) determining the absolute quantity of the at least one MHC-presented peptide in at least two, preferably at least four, biological samples;
[0094] (b) determining in the at least two, preferably at least four, biological samples of step (a):
[0095] (i) the relative quantity of the at least one MHC-presented peptide; and / or
[0096] (ii) the relative or absolute quantity of at least one surrogate for each of the at least one MHC-presented peptide;
[0097] (c) applying statistical modelling to the at least one absolute quantity of step (a) and the at least one relative or absolute quantity of step (b) to establish at least one statistical model;
[0098] (d) determining in the further biological sample the at least one relative or absolute quantity as in step (b); (e) predicting the absolute quantity of the at least one MHC-presented peptide in the further biological sample using the at least one statistical model.
[0099] In other words, the first aspect of the invention includes two phases: model generation (i.e., steps (a) to (c)) and prediction phase (i.e., steps (d) and (e)). The biological samples used in the model generation phase are different from the further biological sample used in the prediction phase.
[0100] In some embodiments, the method of the first aspect of the invention is a computer- implemented method.
[0101] In a second aspect of the invention a method for generating a statistical model for absolute quantification of at least one MHC-presented peptide in a biological sample is provided that includes steps (a), (b) and (c) (i.e. the model generation phase) of the first aspect of the invention.
[0102] In some embodiments, the method of the second aspect of the invention is a computer- implemented method.
[0103] For the model generation the absolute quantity (preferably peptide copy number) is determined for a particular MHC-presented peptide using a suitable method. This absolute measurement is paired with the relative quantitative measurement of the particular MHC- presented peptide and / or with relative or absolute quantitative measurements of a surrogate that can be used for predicting absolute quantity such as copy numbers. The surrogate data may be derived from various sources including RNA, DNA, proteins, peptides, or combinations thereof.
[0104] Next a statistical model is established for associating absolute data with relative and / or surrogate data. A specialized statistical model allows prediction of copy numbers for a specific MHC-presented peptide in new samples based on the relative quantitative measurement of the particular MHC-presented peptide and / or the relative or absolute quantitative measurements of a surrogate.
[0105] Alternatively, a universal statistical model may be established that allows prediction of absolute quantities (e.g., peptide copy numbers) for any peptide in a new sample (i.e., more than one MHC-presented peptide). This is accomplished by including sequence-derived properties in the model (e.g., peptide-specific off-set parameter). These properties may be derived from existing models or trained based on experimental measurements. Existing models for predicting peptide detectability from proteomics data have shown to be not transferable to immunopeptidomics (see, Serrano et al. 2019; Bioinformatics, 36(4), 2020, 1279-1280). The statistical model (for both the specialized statistical model and the universal statistical model) is preferably generated by using regression analysis.
[0106] In a preferred embodiment, the statistical model established in step (c) is a specialized model restricted to the prediction of the absolute quantity of one specific MHC-presented peptide; or a universal model allowing the prediction of the absolute quantity of more than one MHC-presented peptide.
[0107] In a preferred embodiment, the statistical model established in step (c) is a specialized model restricted to the prediction of the absolute quantity of one specific MHC-presented peptide.
[0108] In a preferred embodiment, the statistical model established in step (c) is a universal model allowing the prediction of the absolute quantity of more than one MHC-presented peptide.
[0109] In a preferred embodiment, at least one, at least two, at least five or at least 10 statistical models are established in step (c).
[0110] In a preferred embodiment, one statistical model is established in step (c).
[0111] In a preferred embodiment of the first and second aspect of the invention, the absolute quantity of the at least one MHC-presented peptide determined in step (a) involves experimental measurement in the biological sample. In other words, the absolute quantity of the at least one MHC-presented peptide determined in step (a) may be determined by experimental measurement / analysis of the biological sample such as mass spectrometry. The absolute quantity of the at least one MHC-presented peptide determined in step (a) may be determined using the AbsQuant® technology platform [as described in US10,545,154 B2, which is herein incorporated by reference].
[0112] In a preferred embodiment the absolute quantities of step (a) are determined in at least 2, at least 4, at least 6, at least 8, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 500, at least 1000, at least 2000, at least 3000 or at least 5000 samples. In a preferred embodiment the absolute quantities of step (a) are determined in at least 4 samples. In a preferred embodiment the absolute quantities of step (a) are determined in at least 10 samples.
[0113] In a preferred embodiment the absolute quantities of step (a) are determined in 4 to 100, 10 to 80, 20 to 60, or 30 to 50 samples. In a preferred embodiment the absolute quantities of step (a) are determined in 10 to 20 samples.
[0114] In a preferred embodiment the relative quantity of the at least one surrogate in step (b)(ii) of the first and second aspect is determined in relation to a reference. In a preferred embodiment the relative quantity of the at least one MHC-presented peptide in step (b)(i) of the first and second aspect is determined in relation to an internal control, preferably a spike in control.
[0115] In a preferred embodiment the relative quantity of the at least one MHC-presented peptide in step (b)(i) of the first and second aspect is determined in relation to a reference sample, preferably the reference sample is a standardized sample, preferably the reference sample is measured immediately before or after the acquisition of the at least one MHC- presented peptide.
[0116] In a preferred embodiment of the first and second aspect the relative quantity of the at least one surrogate in step (b)(ii) is determined in relation to a reference that is known to have stable expression levels in the biological samples used, preferably the reference has a difference of not more than 5%, 10%, 15% or 20% between the biological samples used.
[0117] In a preferred embodiment the relative and / or absolute quantities of step (b) are determined in at least 2, at least 4, at least 6, at least 8, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90 or at least 100 samples.
[0118] In a preferred embodiment the relative and / or absolute quantities of step (b) are determined in 4 to 100, 10 to 80, 20 to 60, or 30 to 50 samples. In a preferred embodiment the relative and / or absolute quantities of step (b) are determined in 1 to 20 samples.
[0119] In the second phase, i.e. the prediction phase, the statistical model (either specialized or universal) is applied to new samples to predict absolute peptide quantities. The new sample may be referred to as the further biological sample.
[0120] This approach provides fast and easy prediction for copy number compared to the experimental determination of copies. The method of the present invention therefore allows inter alia investigation of samples for research purposes as well as patient and target selection in clinical studies. For example, patient selection can be based on inclusion of patients with a minimal predicted peptide copy number. As another example, target selection can be based on selecting the target with the highest predicted copies out of those targets for which therapeutics are available enabling personalized immunotherapies.
[0121] Surrogates used in the methods of the invention
[0122] In a preferred embodiment of the first aspect of the present invention, the surrogate is a primary surrogate or a secondary surrogate. The surrogate is considered a primary surrogate if it is derived from the MHC presented peptide (e.g., the corresponding source protein, transcript or gene). A secondary surrogate may be any parameter of relevance for the particular MHC- peptide.
[0123] In a preferred embodiment of the first aspect of the present invention, the surrogate is a primary surrogate.
[0124] In a preferred embodiment of the first aspect of the present invention, the surrogate is a secondary surrogate.
[0125] In a preferred embodiment, the surrogate is selected from the nucleic acid sequence encoding the at least one MHC-presented peptide and the protein of which the at least one MHC-presented peptide is derived from.
[0126] In a preferred embodiment of the first aspect of the present invention, the primary surrogate is derived from the nucleic acid sequence encoding the at least one MHC-presented peptide, preferably mRNA expression. In a preferred embodiment, the primary surrogate is derived from the nucleic acid sequence encoding the at least one MHC-presented peptide and is selected from exon transcript or gene, preferably determined by RNAseq. In a preferred embodiment, the primary surrogate is derived from the genomic sequence encoding the at least one MHC-presented peptide, preferably selected from genomic variation, transcript variation or alternative splicing.
[0127] In a preferred embodiment of the first aspect of the present invention, the primary surrogate is derived from the source protein that the at least one MHC-presented peptide is derived from. In a preferred embodiment, the primary surrogate is the (absolute or relative) quantification of protein expression.
[0128] In a preferred embodiment of the first aspect of the present invention, the secondary surrogate is a parameter relevant for the at least one MHC-presented peptide, preferably the secondary surrogate relates to genes relevant for antigen processing and presentation. A parameter is considered relevant if it improves the prediction in step (e). The performance of a model is evaluated using the adjusted R2and the associated p-value. A parameter is deemed relevant if it leads to an increase in the adjusted R2or enhances the statistical significance of the model. Examples thereof can be found in Tables 2 to 9. Preferably the secondary surrogate is selected from HLA, TAP (transporter associated with antigen processing), Beta-2M (Beta-2- microglobulin), ERAP (ER-associated aminopeptidase), proteasome and immunoproteasome (such as the proteasome endopeptidase subunits alpha type and proteasome subunits beta type). Of particular relevance are the secondary surrogates that are associated with the at least one MHC-presented peptide, such as for example the particular HLA molecule that presents the at least one MHC-presented peptide. In a preferred embodiment, more than one (e.g. 2, 3, 4 or 5) secondary surrogate is used.
[0129] In a preferred embodiment, the secondary surrogate is HLA. In such an embodiment, preferably the specific HLA-allele that is presenting the at least one MHC-presented peptide (e.g., HLA-A*02:01) is quantified. In another preferred embodiment the HLA is determined on the level of the HLA-gene (e.g. HLA-A, HLA-B, HLA-C ).
[0130] In another preferred embodiment, the secondary surrogate is the proteasome. In such an embodiment, preferably the measurement would be on the level of the proteasome complex as a whole or its subunits. Relevant subunits include components of the constitutive proteasome such as the proteasome subunit alpha type-1 to type -8 and the proteasome subunit beta type-1 to type-7. In a preferred embodiment at least one (at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15) subunits are selected from subunit alpha type- 1 to type-8 and subunit beta type- 1 to type- 7. The proteasome can be determined on the protein level as well as on the RNA level. Suitable methods to measure the proteasome and in particular its subunits on the protein level are all proteomic methods known in the art to determine specific proteins and protein complexes. Particularly preferred methods are mass spectrometry, immunohistochemistry, immunocytochemistry and ELISA. Suitable methods to measure the proteasome and in particular its subunits on the RNA level are sequencing based methods (preferably RNA-seq), microarrays and PCR-based methods (preferably qPCR).
[0131] In another preferred embodiment, the secondary surrogate is the immunoproteasome. In such an embodiment, preferably the measurement would regard the quantification of the immunoproteasome complex as a whole or its subunits. Preferably subunits specific for the immunoproteasome (i.e. not present in the proteasome) are determined. Specific subunits include the proteasome subunit beta type-8, beta type-9 and beta-type 10 In a preferred embodiment, at least one subunit (at least 2 or 3) specific for the immunoproteasome is selected from proteasome subunit beta type-8, type-9 and type- 10. The immunoproteasome can be determined on the protein level as well as on the RNA level. Suitable methods to measure the immunoproteasome and in particular its subunits on the protein level are all proteomic methods known in the art to determine specific proteins and protein complexes. Particularly preferred methods are mass spectrometry, immunohistochemistry, immunochemistry and ELISA. Suitable methods to measure the immunoproteasome and in particular its subunits on the RNA level are sequencing based methods (preferably RNA-seq), microarrays and PCR-based methods (preferably qPCR).
[0132] In another preferred embodiment, the secondary surrogate is related to processing and / or transport of MHC-presented peptides. In a preferred embodiment the method of the invention uses a combination of primary and secondary surrogates. Particularly preferred combinations are a primary surrogate derived from the nucleic acid sequence encoding the at least one MHC-presented peptide, preferably determined by RNAseq, and the preferred secondary surrogate is selected from HLA. Beta-2M, TAP, Immunoproteasone, Proteasome or ERAP.
[0133] Statistical modelling in step (c)
[0134] In a preferred embodiment of the first aspect of the present invention, in step (c) only one quantity of step (b) is used for statistical modelling (i.e. univariate analysis). In other words, in the model generation phase (i.e. steps (a), (b) and (c)) the absolute quantity determined in step (a) and one quantity of step (b) is used for statistical modelling in step (c). Preferably the one quantity of step (b) is the relative quantity of the at least one MHC-presented peptide or the relative or absolute quantity of at least one surrogate for each of the at least one MHC-presented peptides. In a preferred embodiment the one quantity of step (b) is the relative quantity of the at least one MHC-presented peptide. In another preferred embodiment the one quantity of step
[0135] (b) is the relative or absolute quantity of at least one surrogate for each of the at least one MHC- presented peptide. In another preferred embodiment the one quantity of step (b) is the relative quantity of at least one surrogate for each of the at least one MHC-presented peptides. In another preferred embodiment the one quantity of step (b) is the absolute quantity of at least one surrogate for each of the at least one MHC-presented peptide. In a preferred embodiment, the one quantity of step (b) is the relative quantity of the mRNA encoding the at least one MHC- presented peptide. In a preferred embodiment, the one quantity of step (b) is the relative quantity of the protein that the at least one MHC-presented peptide is derived from.
[0136] In a preferred embodiment of the first aspect of the present invention, wherein in step
[0137] (c) at least two quantities of step (b) are used for statistical modelling (i.e. multivariate analysis). In other words, in the model generation phase (i.e. steps (a), (b) and (c)) the absolute quantity determined in step (a) and at least two quantities of step (b) are used for statistical modelling in step (c). In a preferred embodiment at least three, at least four, at least five or at least six quantities of step (b) are used for statistical modelling in step (c). In another preferred embodiment two, three, four, five, or six quantities of step (b) are used for statistical modelling in step (c). In another preferred embodiment two quantities of step (b) are used for statistical modelling in step (c). In another preferred embodiment three quantities of step (b) are used for statistical modelling in step (c). In another preferred embodiment, four quantities of step (b) are used for statistical modelling in step (c). In another preferred embodiment, five quantities of step (b) are used for statistical modelling in step (c). In another preferred embodiment, six quantities of step (b) are used for statistical modelling in step (c).
[0138] In a preferred embodiment, wherein, in case of determining more than one quantity in step (b) (i.e. multivariate analysis), this includes the relative or absolute quantity of at least one primary surrogate or the relative quantity of the at least one MHC-presented peptide optionally further including at least one secondary surrogate. For a multivariate analysis these quantities may be combined with any other primary or secondary surrogate disclosed herein.
[0139] In a preferred embodiment, multivariate analysis comprises determining in step (a) the absolute quantity of the at least one MHC-presented peptide and in step (b) the relative quantity of the at least one MHC-presented peptide (preferably determined by mass spectrometry or liquid chromatography-mass spectrometry).
[0140] In another preferred embodiment, multivariate analysis comprises determining in step
[0141] (a) the absolute quantity of the at least one MHC-presented peptide and in step (b) the relative quantity of mRNA encoding the MHC-presented peptide (preferably by RNAseq or qPCR).
[0142] In a preferred embodiment, multivariate analysis comprises determining in step (a) the absolute quantity of the at least one MHC-presented peptide and in step (b) the relative quantity of the protein that the at least one MHC-presented peptide is derived from (preferably determined by mass spectrometry or liquid chromatography-mass spectrometry).
[0143] In a preferred embodiment, multivariate analysis comprises determining the relative quantity of the at least one MHC-presented peptide and as a secondary surrogate the relative quantity of the specific HLA- Allele that is presenting the at least on MHC-presented peptide. Preferably the HLA-allele is determined on mRNA level. Preferably the HLA-allele is determined on peptide level. In a preferred embodiment the HLA-allele is determined that is presenting the at least one MHC-presented peptide.
[0144] In a preferred embodiment, multivariate analysis comprises determining the relative or absolute quantity of a primary surrogate (preferably selected from gene expression, genomic variation, transcript variation and alternative splicing, more preferably gene expression) and as a secondary surrogate the relative quantity of the specific HLA- Allele that is presenting the at least on MHC-presented peptide. Preferably the HLA-allele is determined on mRNA level. Preferably the HLA-allele is determined on peptide level. In a preferred embodiment the HLA- allele is determined that is presenting the at least one MHC-presented peptide.
[0145] In a preferred embodiment multivariate analysis may further include determining in step
[0146] (b) mRNA expression, genomic variation, transcript variation and / or alternative splicing. In a preferred embodiment the statistical modelling of step (c) is regression analysis. In a preferred embodiment the regression analysis is linear regression. Linear regression might also be used by transforming the data (e.g. log or log-log transformation) used prior to analysis so to enable the use of linear regression.
[0147] In a particular embodiment the invention relates to a method for absolute quantification of at least one MHC-presented peptide in a further biological sample comprising the following steps:
[0148] (a) determining the absolute quantity of the at least one MHC-presented peptide in at least two, preferably at least four, biological samples;
[0149] (b) determining in the at least two, preferably at least four, biological samples of step (a) the relative quantity of expression of the source gene of the at least one MHC-presented peptide;
[0150] (c) applying statistical modelling to the at least one absolute quantity of step (a) and the relative quantity of the source gene of step (b) to establish at least one statistical model;
[0151] (d) determine in the further biological sample the relative quantity of expression of the source gene of the at least one MHC-presented peptide;
[0152] (e) predict the absolute quantity of the at least one MHC-presented peptide in the further biological sample using the at least one statistical model.
[0153] In a particular embodiment the invention relates to a method for absolute quantification of at least one MHC-presented peptide in a further biological sample comprising the following steps:
[0154] (a) determining the absolute quantity of the at least one MHC-presented peptide in at least two, preferably at least four, biological samples;
[0155] (b) determining in the at least two, preferably at least four, biological samples of step (a) the relative or absolute quantity of one primary surrogate and of at least one secondary surrogate for each of the at least one MHC-presented peptides;
[0156] (c) applying statistical modelling to the at least one absolute quantity of step (a) and the relative or absolute quantities of step (b) to establish at least one statistical model;
[0157] (d) determine in the further biological sample the relative or absolute quantity of one primary surrogate and of at least one secondary surrogate for each of the at least one MHC-presented peptides; (e) predict the absolute quantity of the at least one MHC-presented peptide in the further biological sample using the at least one statistical model; wherein the primary surrogate is the relative quantity of expression of the source gene of the at least one MHC-presented peptide; and wherein the secondary surrogate is relative or absolute quantity of expression of the gene of a HLA, the gene of TAP (transporter associated with antigen processing), the gene of Beta-2M (Beta-2-microglobulin), the gene of ERAP (ER-associated aminopeptidase), the gene of a subunit of the proteasome and / or the gene of a subunit of the immunoproteasome.
[0158] In a particular embodiment the invention relates to a method for absolute quantification of at least one MHC-presented peptide in a further biological sample comprising the following steps:
[0159] (a) determining the absolute quantity of the at least one MHC-presented peptide in at least two, preferably at least four, biological samples;
[0160] (b) determining in the at least two, preferably at least four, biological samples of step (a) the relative or absolute quantity of one primary surrogate and of at least one secondary surrogate for each of the at least one MHC-presented peptides;
[0161] (c) applying statistical modelling to the at least one absolute quantity of step (a) and the relative or absolute quantities of step (b) to establish at least one statistical model;
[0162] (d) determine in the further biological sample the relative or absolute quantity of one primary surrogate and of at least one secondary surrogate for each of the at least one MHC-presented peptides;
[0163] (e) predict the absolute quantity of the at least one MHC-presented peptide in the further biological sample using the at least one statistical model; wherein the primary surrogate is the relative quantity of expression of the source gene of the at least one MHC-presented peptide; and wherein the secondary surrogate is relative or absolute quantity of protein level of a HLA, of TAP (transporter associated with antigen processing), of Beta-2M (Beta-2-microglobulin), of ERAP (ER-associated aminopeptidase), of a subunit of the proteasome and / or of a subunit of the immunoproteasome.
[0164] In a particular embodiment the invention relates to a method for absolute quantification of at least one MHC-presented peptide in a further biological sample comprising the following steps: (a) determining the absolute quantity of the at least one MHC-presented peptide in at least two, preferably at least four, biological samples;
[0165] (b) determining in the at least two, preferably at least four, biological samples of step (a) the relative or absolute quantity of one primary surrogate and of at least one secondary surrogate for each of the at least one MHC-presented peptides;
[0166] (c) applying statistical modelling to the at least one absolute quantity of step (a) and the relative or absolute quantities of step (b) to establish at least one statistical model;
[0167] (d) determine in the further biological sample the relative or absolute quantity of one primary surrogate and of at least one secondary surrogate for each of the at least one MHC-presented peptides;
[0168] (e) predict the absolute quantity of the at least one MHC-presented peptide in the further biological sample using the at least one statistical model; wherein the primary surrogate is the relative or absolute quantity of the protein containing the at least one MHC-presented peptide; and wherein the secondary surrogate is relative or absolute quantity of expression of the gene of a HLA, the gene of TAP (transporter associated with antigen processing), the gene of Beta-2M (Beta-2-microglobulin), the gene of ERAP (ER-associated aminopeptidase), the gene of a subunit of the proteasome and / or the gene of a subunit of the immunoproteasome.
[0169] In a particular embodiment the invention relates to a method for absolute quantification of at least one MHC-presented peptide in a further biological sample comprising the following steps:
[0170] (a) determining the absolute quantity of the at least one MHC-presented peptide in at least two, preferably at least four, biological samples;
[0171] (b) determining in the at least two, preferably at least four, biological samples of step (a) the relative or absolute quantity of one primary surrogate and of at least one secondary surrogate for each of the at least one MHC-presented peptides;
[0172] (c) applying statistical modelling to the at least one absolute quantity of step (a) and the relative or absolute quantities of step (b) to establish at least one statistical model;
[0173] (d) determine in the further biological sample the relative or absolute quantity of one primary surrogate and of at least one secondary surrogate for each of the at least one MHC-presented peptides; (e) predict the absolute quantity of the at least one MHC-presented peptide in the further biological sample using the at least one statistical model; wherein the primary surrogate is the relative or absolute quantity of the protein containing the at least one MHC-presented peptide; and wherein the secondary surrogate is relative or absolute quantity of protein level of a HLA, of TAP (transporter associated with antigen processing), of Beta-2M (Beta-2-microglobulin), of ERAP (ER-associated aminopeptidase), of a subunit of the proteasome and / or of a subunit of the immunoproteasome.
[0174] Universal statistical model
[0175] In a preferred embodiment the statistical modelling of step (c) (preferably regression analysis) comprises the determination of at least one peptide-specific off-set parameter (synonymously referred to herein as peptide-specific parameter). Preferably the at least one peptide-specific off-set parameter is determined for establishing a universal statistical model allowing prediction of absolute quantities for any peptide in a new sample.
[0176] Suitable peptide-specific off-set parameters in the context of the present invention include physicochemical and biochemical properties of amino acids. Preferred examples are Amino Acid index (AAindex; see Kawashima et al, Nucleic Acids Res. 2000 Jan 1 ;28(1):374), hydrophobicity score (see Kyte et al, J Mol Biol. 1982 May 5; 157(1): 105-32), peptide retention time in liquid chromatography (see Ma et al, Anal. Chem. 2018, 90, 18, 10881-10888), sequence specific peptide bias (see Dincer et al, J. Proteome Res. 2022, 21, 7, 1771-1782) or binding affinity. In a preferred embodiment the peptide specific off-set parameter is sequence specific and is determined on the basis of the physicochemical and biochemical properties of the amino acids comprised in the MHC-presented peptide to be quantified in the method of the present invention. Obtaining the peptide-specific off-set parameter may include the use of deep learning algorithms.
[0177] In a preferred embodiment the peptide-specific off-set parameter is a physicochemical property. In a preferred embodiment the peptide-specific off-set parameter is a biochemical property. In a preferred embodiment the peptide-specific off-set parameter is AAindex. In a preferred embodiment the peptide-specific off-set parameter is hydrophobicity score.
[0178] In a preferred embodiment the peptide-specific off-set parameter is peptide retention time. In a more preferred embodiment the peptide-specific off-set parameter is peptide retention time in liquid chromatography, preferably standardized retention time (uRT). Examples of how retention time and standardized retention time can be determined / predicted can be found in section 2.71 below. One preferred example of how retention time is predicted is the use of the Prosit model.
[0179] In a preferred embodiment the peptide-specific off-set parameter is sequence specific peptide bias.
[0180] In a preferred embodiment the peptide-specific off-set parameter is binding affinity, preferably HLA binding affinity. Preferably the binding affinity is predicted using suitable methods. Examples of HLA binding affinity can be determined / predicted can be found in section 2.7.2 below. One preferred example of how the HLA binding affinity can be predicted is NetMHCpan.
[0181] The peptide-specific off-set parameter may be used to amend the statistical model established in step (c), preferably a universal statistical model.
[0182] A simple linear regression model might be expressed using the formula Y = a X X + b, wherein Y represents the predicted value (i.e. absolute quantity of the MHC-presented peptide), X is the value of the predictor variable (i.e. the quantity determined in step (d)) and a is the change in the predicted value for a one unit increase in X and b is the value of the predicted value when X = 0. The variable a is sometimes also referred to as slope, variable b is sometimes also referred to as intercept. In a universal model the formula is extended to Y = a x X + . The predicted value in a universal model is determined by an universal slope a and an universal intercept as well as a peptide-specific off-set factor. The peptide-specific off-set parameter is preferably sequence specific. The combination of both form the peptide-specific intercept bi. The peptide-specific off-set facilitates the prediction process by modeling peptide-specific variation and is estimated preferably based on the known amino acid sequence.
[0183] In a preferred embodiment the linear regression model used for a univariate model is Y = pO + pi * XI + e wherein XI is the primary surrogate (substitute for direct absolute peptide quantity) and represents relative peptide quantities or protein (source protein) or mRNA (source gene, transcript, exon or peptide-encoding nucleotide sequences) that the peptides are derived from.
[0184] In a preferred embodiment the linear regression model used for a multivariate model is
[0185] Y = P0 + pi * XI + p2 * X2 + e wherein XI is the primary surrogate (substitute for direct absolute peptide quantity) and represents relative peptide quantities or protein (source protein) or mRNA (source gene, transcript, exon or peptide-encoding nucleotide sequences) that the peptides are derived from. X2 is the secondary surrogate and represents mRNA or protein quantities that are relevant for antigen processing and presentation. In a particular embodiment the invention relates to a method for absolute quantification of at least one MHC-presented peptide in a further biological sample comprising the following steps:
[0186] (a) determining the absolute quantity of the at least one MHC-presented peptide in at least two, preferably at least four, biological samples;
[0187] (b) determining in the at least two, preferably at least four, biological samples of step (a) the relative or absolute quantity of one primary surrogate and of at least one peptidespecific off-set parameter for each of the at least one MHC-presented peptides;
[0188] (c) applying statistical modelling to the at least one absolute quantity of step (a) and the relative or absolute quantities and off-set parameter(s) of step (b) to establish at least one universal statistical model;
[0189] (d) determine in the further biological sample the relative or absolute quantity of the primary surrogate and of the at least one peptide-specific off-set parameter for each of the at least one MHC-presented peptides;
[0190] (e) predict the absolute quantity of the at least on MHC-presented peptide in the further biological sample using the at least one universal statistical model; wherein the primary surrogate is the relative quantity of the at least one MHC-presented peptide; and wherein the off-set parameter is the estimated peptide retention time in liquid chromatography. In the embodiment described directly above it is preferred that at least 2, at least 10, at least 50 or at least 100, preferably at least 100 MHC-presented peptides are considered in steps (a), (b) and (c).
[0191] In a particular embodiment the invention relates to a method for absolute quantification of at least one MHC-presented peptide in a further biological sample comprising the following steps:
[0192] (a) determining the absolute quantity of the at least one MHC-presented peptide in at least two, preferably at least four, biological samples;
[0193] (b) determining in the at least two, preferably at least four, biological samples of step (a) the relative or absolute quantity of one primary surrogate, the relative or absolute quantity of at least one secondary surrogate, and at least one peptide-specific off-set parameter for each of the at least one MHC-presented peptides;
[0194] (c) applying statistical modelling to the at least one absolute quantity of step (a) and the relative or absolute quantities and off-set parameter(s) of step (b) to establish at least one universal statistical model; (d) determine in the further biological sample the relative or absolute quantity of the primary surrogate, the relative or absolute quantity of the at least one secondary surrogate, the at least one peptide-specific off-set parameter for each of the at least one MHC-presented peptides;
[0195] (e) predict the absolute quantity of the at least on MHC-presented peptide in the further biological sample using the at least one universal statistical model; wherein the primary surrogate is the relative quantity of expression of the source gene of the at least one MHC-presented peptide; andwherein the secondary surrogate is the relative or absolute quantity of the source gene or protein level of a HLA, of TAP (transporter associated with antigen processing), of Beta-2M (Beta-2-microglobulin), of ERAP (ER-associated aminopeptidase), of a subunit of the proteasome and / or of a subunit of the immunoproteasome; and wherein the peptide-specific off-set parameter is determined from the estimated peptide retention time in liquid chromatography and the estimated binding affinity of the peptide to the MHC.
[0196] In the embodiment described directly above it is preferred that at least 2, at least 10, at least 50 or at least 100, preferably at least 100 MHC-presented peptides are considered in steps (a), (b) and (c).
[0197] Determining treatment options
[0198] It is envisaged that the peptides or pharmaceutically acceptable salts thereof quantified by the herein described methods or pharmaceutically acceptable salts thereof are used as a medicament e.g. as an anti-cancer vaccine. Accordingly, the invention relates to a peptide quantified by the herein described methods or a pharmaceutically acceptable salt thereof for use in the (manufacture of a medicament for the) treatment of cancer or a tumorous disease and / or disorder, such as a solid tumor.
[0199] It is further envisaged that the information obtained by the herein described methods is used to develop cancer therapies. Accordingly, when a (MHC-presented) peptide is quantified, an antigen-binding protein may be developed / provided targeting said peptide. Said antigenbinding protein may be used for the treatment of a disease or clinical condition. An antigen binding protein may be an antibody, a TCR, or combination thereof. The antigen binding protein may also be a multi-specific (e.g. a bispecific) protein, for example comprising variable domains of an TCR and an antibody, such as an TCER®
[0200] For example, when the quantification (using the herein described methods) determines peptides that show (pronounced) differences in e.g. presentation by MHC between a sample type of clinical interest (e.g. various tumor types or autoimmune diseases) and unaffected / benign samples and / or a reference dataset said peptides may be used as targets for anti-cancer-immunotherapy. In other words, when the quantification (using the herein described methods) determines a peptide that shows (pronounced) differences in e.g. presentation by MHC between a sample type of clinical interest (e.g. various tumor types or autoimmune diseases) and unaffected / benign samples and / or a reference dataset an antigenbinding protein may be developed targeting said peptide. Especially when by the herein described methods a peptide is identified that is only presented on cancer cells and not on healthy tissue or the presentation on cancer cells is significantly increased compared to healthy tissue said peptide or an antigen-binding protein targeting said peptide may be used for the treatment of the corresponding cancer.
[0201] Accordingly, the invention relates to an (in vitro) method for identifying an antigenbinding protein binding to a peptide (preferably bound to a MHC protein) quantified by the herein described methods comprising: a) contacting a plurality of antigen-binding proteins with a peptide (preferably bound to a MHC protein) quantified by the herein described methods, b) identifying an antigen-binding protein that binds to said peptide (preferably bound to a MHC protein), and c) selecting the antigen-binding protein identified in b).
[0202] Furthermore, the invention relates to an (in vitro) method for providing an antigenbinding protein binding to a peptide (preferably bound to a MHC protein) quantified by the herein described methods comprising: a) providing an antigen-binding protein, b) contacting the antigen-binding protein with a peptide (preferably bound to a MHC protein) quantified by the herein described methods, c) identifying an antigen-binding protein that binds to said peptide (preferably bound to a MHC protein), and d) selecting the antigen-binding protein identified in b).
[0203] The invention also relates to an antigen-binding protein binding to a peptide (preferably bound to a MHC protein) quantified by the herein described methods for use as a medicament, preferably for use in the treatment of cancer.
[0204] The antigen-binding protein identified as described above may be expressed in a host cell. In particular when the antigen-binding protein is a TCR said TCR may be expressed in a host cell and presented by said host cell. Said host cell may be used as a medicament, preferably for use in the treatment of cancer.
[0205] It is evident for the skilled person that the described peptide, antigen-binding protein or host cell may be formulated in a pharmaceutical composition. Said pharmaceutical composition may be used as a medicament, preferably for use in the treatment of cancer.
[0206] It is envisaged that the information obtained by the herein described methods is used to determine treatment options for e.g. cancer patients.
[0207] The peptides quantified by the herein described methods may relate to peptides targeted in immunotherapy.
[0208] Accordingly, from the e.g. copies per cell of a given peptide determined by the herein described methods the skilled person may infer whether a cancer may be treated by said peptide, a pharmaceutical salt thereof or an antigen-binding protein binding said peptide.
[0209] Thus, the invention relates to an antigen-binding protein binding to the at least one MHC-presented peptide, at least one MHC-presented peptide or the pharmaceutical salt thereof for use in the treatment of a subject suffering from cancer, wherein the absolute quantity of at least one MHC-presented peptide in a further biological sample of the subject is determined according to the described methods, wherein the absolute quantity of at least one MHC-presented peptide is indicative of the suitability to treat the subject with an antigen-binding protein binding to the at least one MHC- presented peptide, the at least one MHC-presented peptide or the pharmaceutical salt thereof, wherein an antigen binding protein binding to the at least one MHC-presented peptide, the at least one MHC-presented peptide or the pharmaceutical salt thereof is to be administered to the subject.
[0210] The invention also relates to an antigen-binding protein binding to the at least one MHC- presented peptide, at least one MHC-presented peptide or the pharmaceutical salt thereof for use in the treatment of a subject suffering from cancer, wherein the absolute quantity of at least one MHC-presented peptide in a further biological sample of the subject is determined according to the described methods, wherein the absolute quantity of at least one MHC-presented peptide is indicative of the need of the subject for treatment, wherein an antigen binding protein binding to the at least one MHC-presented peptide, the at least one MHC-presented peptide or the pharmaceutical salt thereof is to be administered to the subject. The invention further relates to an antigen-binding protein that specifically binds to a MHC-presented peptide for use in a method of treating cancer in a subject wherein the method comprises: a) obtaining a sample from the subject, b) determining the absolute quantity of said MHC-presented peptide in said sample according to the herein described methods, c) comparing the absolute quantity of said MHC-presented peptide to a reference value or threshold indicative of treatment suitability, and d) administering the antigen-binding protein to the subject, wherein the administration is based on the determination that the absolute quantity of said MHC-presented peptide meets or exceeds the reference value.
[0211] The invention also relates to an (in vitro) method to identify a subject suffering from cancer that is suitable to be treated against the cancer (by e.g. a given peptide, pharmaceutical salt thereof or an antigen-binding protein).
[0212] Although evident for the skilled person it is pointed out that the subject-matter of the described aspects and embodiments can be combined.
[0213] Examples
[0214] 1. Sample collection
[0215] Immatics Biotechnologies GmbH procured all samples and specimens with written informed consent and appropriate committee approvals. Observations from 1132 samples were utilized during the generation and evaluation of the described models. 781 primary samples were taken surgically or post-mortem from 154 normal donors and 476 cancer patients. In addition, measurements from 208 cultivated cell line and 143 xenograft samples were included.
[0216] 2. Data Generation
[0217] 2.1 HLA peptide isolation for LC- MS Analysis
[0218] HLA peptides were obtained by immune precipitation according to a slightly modified protocol (Falk et al., 1991; Nature, 351(6324), 290-296; Seeger et al., 1999, Immunogenetics, 49, 571-576) using the HLA-A*02-specific antibody BB7.2 or the HLA- A, -B, C-specific antibody W6 / 32 immobilized onto a stationary phase of CNBr-activated Sepharose. Peptides were eluted by acid treatment and further purified using low-mass cut-off filter units, followed by lyophilization and reconstitution in 5 % formic acid (FA).
[0219] 2.2 Absolute Quantification of HLA-presented peptides
[0220] Absolute quantity of HLA-presented peptides was determined using the AbsQuant® technology platform [as described in US10,545,154 B2, which is herein incorporated by reference].
[0221] In general, absolute quantity of HLA-presented peptides (expressed as ‘copies per cell’) by AbsQuant® is calculated using three different parameters: the number of cells within the investigated sample specimen, the total peptide amount in the sample, and the peptide isolation efficiency upon sample processing.
[0222] Each of the three parameters was determined experimentally as outlined below.
[0223] The total peptide amount in a sample is determined by nanoLC-MS / MS (scheduled parallel reaction monitoring in the ion trap) by calculating the ratio of the peptide variant of interest and a fixed amount of an isotope-labelled version of the peptide, the so-called internal standard, which is added after elution and prior to MS analysis to the sample.
[0224] For samples endogenously expressing and presenting the peptide of interest (e.g. tumor tissues or cell lines), the unlabeled peptide variant will be monitored as the compound of interest via nanoLC-MS / MS.
[0225] The efficiency of peptide isolation - during immunoprecipitation and further sample processing prior to LC / MS analysis - was determined by spiking of refolded peptide-HLA complexes of the investigated peptide into the sample lysate at the earliest possible point of time during the peptide isolation procedure for all assays.
[0226] The total cell count was inferred from measurements of the total DNA content of the analyzed tissue sample or, if applicable, by manually counting the cells.
[0227] The number of peptide copies per cell is calculated according to the following equation: total peptide 100% peptide copies per cell = - - - x — - - - — - cell count % isolation efficiency
[0228] 2.3 Relative Quantification of HLA peptides
[0229] LC-MS analysis of HLA peptide extracts was performed using (Waters, Milford, USA) coupled to an Orbitrap Tribrid mass spectrometer (Thermo Fisher, Waltham, USA). A trapping setup using Waters 25 cm * 75 pm BEH Cl 8 analytical columns was used employing a stepped gradient ranging from 1 to 34.5 acetonitrile over 70 min. MS acquisition was performed in data- dependent mode (DDA) using a “top speed” method with a maximum cycle time of 3s. MSI scan range was set to 200-1500 m / z with an AGC target of le5 and 120k MSI scan resolution. Data dependent MS2 scans were acquired with an isolation width of 2 m / z for the topN precursors within a m / z range of 280-720 m / z and charge states of 2+ and 3+ at 30k resolution and an automatic gain control (AGC) target of 5e4 in the Orbitrap. Precursor fragmentation was performed by collision-induced dissociation (CID) at 35% normalized collision energy (NCE) or higher collisional dissociation (HCD) at 27% NCE. For ion trap (IT) runs, MS2 spectra were acquired with identical precursor selection and isolation settings and CID at 35% NCE was performed with le4 AGC target in the IT with the scan speed set to normal.
[0230] 2.3.1 Data processing and relative quantification
[0231] For peptide identification, untargeted HL A peptidomics data were processed using the search engines X! Tandem, MSGF+ and Comet against the human Ensembl database without enzymatic restriction. Precursor mass tolerance was set to lOppm and fragment mass tolerance to 15 ppm and 600 ppm for Orbitrap and Ion Trap MS2 spectra, respectively. Peptide length was restricted to 6-15 amino acids length and methionine oxidation was set as variable modification. The search results were integrated and false discovery rates limited to 5% using iProphet (Shteynberg, D. et al. 2011, Mol Cell Proteomics 10, Ml 11.007690 (201 1 ) / For peptide quantitation, MSI features were extracted using SuperHirn (Mueller, L. N. et al. 2007, Proteomics 7, 3470-80 (2007)).
[0232] 2.3.2 Data Normalization
[0233] A set of synthetic spike-in peptides were added in fixed amounts as reference peptides. These reference spike-ins were used to normalize peptide MS areas across runs. Spike-in raw areas mirror fluctuations in MS performance, and thus technical variability. For each spike-in peptide, the geometric mean across MS runs is the expected spike-in area. For each run, the median ratio between expected and measured spike-in areas represents a multiplicative scaling factor for peptide MS areas. Runs are scaled down if the measured areas are higher than the expected areas and scaled up otherwise. 2.4 Proteomics
[0234] 2.4.1 Proteolytic Digestion for LC / MS Acquisition of Proteomics Data
[0235] Total protein was isolated from respective cultured cell lines, xenograft in-vivo mouse models and tissue samples in CHAPS buffer. After cell lysis, protein concentration was spectrophotometrically determined using BCA assay. To acquire proteomic data, samples were processed using the SP3 protocol as described in Hughes et al. 2019 (Nature Protocols 14, 68- 85).
[0236] First, an isotopically labeled protein surrogate for monitoring proteolytic efficiency was added to the protein lysate. Subsequently, disulfide bridges were disrupted using TCEP along with CAA for irreversible blocking of free thiol groups at 70°C.
[0237] After bead-based protein clean-up and prior to proteolytic digestion, a fixed amount of isotopically labeled internal standard containing relevant HLA peptide isoforms was added. Proteolytic breakdown was induced by the addition of trypsin / LysC mix followed by incubation overnight at 37°C. The digestion was irreversibly inactivated via addition of FA to a final concentration of 5%.
[0238] 2.4.2 LC-MS Data Acquisition
[0239] Untargeted proteomics data are acquired using DIA mode, acquired with HCD-OT / OT at 27 NCE following an MS method as described in Bruderer et al., 2017 (Mol Cell Proteomics 16, 2296-230) with minor adjustments.
[0240] Full MS spectra (350-1650 m / z) are acquired in OT at a resolution of 120,000 at m / z 200 and AGC target value of 300% with a maximum IT of 100 ms. For fragment spectra (120- 1800m / z), OT resolution is set to 30,000 at m / z 200, the AGC target value to 1000% with a maximum injection time of 54 ms. Each MS cycle takes 44 DIA MS scans with different isolation windows, along the entire mass range. The tryptic digest is loaded onto the column at a calculated default amount of 0.5 pg per injection. A standardized 44 min gradient (Evosep 30SPD) with increasing amount of organic solvent is used for reversed-phase separation. Total method length is 48 min per run.
[0241] Targeted proteomics data for absolute HLA quantification are acquired using the same 44 min chromatographic separation gradient as described above. Scheduled parallel reaction monitoring (sPRM) MS / MS data of a predefined precursor list containing tryptic b2m and HLA peptides of interest are acquired in OT at a resolution of 30,000 at m / z 200 and a maximum injection time of 100 ms. 2.4.3 LC-MS Data Analysis
[0242] Untargeted proteomics data are searched using DIA-NN (Demichev et al. 2020, Nat. Methods., 17(1) :41 -44) against an in-silico-predicted spectral library from the human Ensembl database. Identified peak groups are filtered by 0.01 precursor q-value threshold and identified proteins are filtered by 0.01 protein group q-value.
[0243] Enzyme is set to “trypsin” (C -terminal cleavage after lysine and arginine residues) with a maximum of 2 missed cleavages. Carbamidomethylation of cysteine (+57 Da) is set as a fixed and oxidation of methionine (+ 16 Da) is set as a variable modification. N-terminal methionine excision is enabled. Fragment ion m / z range is set to 120-1,800 m / z and full MS m / z range to 350-1,650. Relative quantification of resulting protein groups is performed using the intensitybased absolute quantification (IBAQ) algorithm on MSI areas as described in Schwanhausser et al., 2011 (Nature 473, 337-342). Relative IBAQ values were calculated by normalizing the protein IBAQ values per MS run to the sum of IBAQ values over all proteins detected in the given run: 106
[0244] The resulting run-level Relative IBAQ values were then aggregated onto the samplelevel by selecting the maximum run-level RellBAQ value:
[0245] RelIBAQpsample) = maxrun( elIBAQp( m )
[0246] Finally, the sample-level RellBAQ values were aggregated by averaging over all samples that are homogeneous to each other.
[0247] Targeted proteomics data are analyzed in Skyline (MacCoss Lab, University of Washington). Further data processing and calculation of absolute peptide abundance is carried out as described in patent application US20220283176A1.
[0248] 2.5 RNASeq
[0249] 2.5.1 Data Acquisition
[0250] The extraction of RNA was performed via TRIzol® (Invitrogen, Karlsruhe, Germany).
[0251] The total RNA was purified using the RNeasy mini kit (QIAGEN, Hilden, Germany) following the providers instructions. RNA samples had to match the criteria of 25ng / pl RNA concentration and a RNA Quality Number (RQN) or Integrity Number (RIN) of 6.0 to be eligible for subsequent sequencing.
[0252] Further processing and sequencing of the samples were performed by GENEWIZ Germany GmbH (Leipzig, Germany). This includes the library preparation with mRNA selection, RNA fragmentation, cDNA conversion and the addition of sequencing adaptors. The NEBNext® Ultra™ II Directional RNA Library Prep Kit for Illumina (New England Biolabs, Ipswich, MA, USA) was used to generate sequencing libraries according to the manufacturer’s protocol. After multiplexing, the sample libraries were loaded on the Illumina NovaSeq 6000 sequencer (Illumina, San Diego, CA, USA) aiming at 80 million paired-end reads with a length of 150bp for each individual sample.
[0253] The Genome Reference Consortium Human Build 38 patch release 13 (GRCh38.pl3) was used as the reference for both exon and transcript quantification. 671,462 exons, 196,165 transcripts and 60,107 genes are included in the annotation.
[0254] The raw sequencing reads are obtained from the provider and prepared for either exon or transcript level expression quantification.
[0255] 2.5.2 Processing and Quantification
[0256] For exon quantification the raw sequencing reads are processed via BBDuk (BBTools v38.91, March 24, 2020) first for adapter trimming and quality pruning. Subsequently, the reads are mapped to the reference using the STAR aligner (v.2.7.3a, Dobin et al., 2013, Bioinformatics, 29(1), 15-21). After further processing via Samtools vl.7 (Danecek et al., 2021, Gigascience, 10(2)) the BAM files, containing only properly aligned read pairs, are generated.
[0257] The read counts on exon level are determined via featureCounts (v2.0.0, Liao et al., 2014, Bioinformatics, 30(7), 923-930). The estimated counts are normalized with regard to length and total number of fragments and converted to transcript per million (TPM) values.
[0258] For transcript quantification, Kallisto (vO.46.1, Bray et al, 2016, Nature biotechnology, 34(5), 525-527) is used to run pseudo-alignment and quantification steps to estimate transcript level expression from the raw sequencing reads and stored as estimated read counts and TPM values.
[0259] The individual exon and transcript expression values are determined for each sequenced sample, as described above and subsequently stored in the Discovery data warehouse. 2.5.3 Data Quality Control and Normalization
[0260] The processing includes the calculation of summary statistics, quality control via the IMAQC pipeline and normalization to allow for comparison across samples. The normalization approach is based on the between-samples normalization scheme DESeq by Anders and Huber (2010, Nature Precedings, 1-1). Sample-specific normalization factors are calculated to in- or decrease expression values to achieve the same expression levels as determined in a selected subset of reference samples. The normalization factors are calculated separately on exon and transcript level.
[0261] 2.5.4 Estimation of HLA abundance via arcasHLA
[0262] Reads mapping to the chromosome 6 based on the genome reference were taken from RNAseq BAM files and used to determine HLA-mapping reads. For quantification, a customized reference containing all previously annotated HLA alleles of a given sample, was used. Based on HLA-mapping reads and this customized reference, read counts and TPMs for all annotated alleles were calculated using arcasHLA (version 0.4.0, default parameters, Reference Orenbuch, R et al, 2020, Bioinformatics, 36(1), 33-40).
[0263] 2.5.5 Quantification of Peptide-encoding nucleotide sequences
[0264] Peptide amino acid (AA) sequences were converted to all possible peptide-encoding nucleotide sequences (PNS) using canonical AA codons and including forward and reverse- complemented sequences. For a given sample, the RNAseq BAM file was processed to count the total number of reads containing any of the PNS. Read counts were normalized by the total number of reads recorded in the BAM file as PNS reads per million (RPM).
[0265] 2.6 Quantification of mRNA via RT-qPCR mRNA levels were determined via quantitative real-time RT-PCR (qPCR) based assays. Following isolation, the mRNA was transcribed to cDNA via High-Capacity cDNA Reverse Transcription kit (Life Technologies, Carlsbad, USA) according to manufacturer’s instructions. For each target gene, sample triplicates were measured in addition to a negative reverse transcription control (-RT) and negative control (NTC). Measurements were performed using TaqMan Universal Master Mix II no Uracil-N Glycosidase (Life Technologies, Cat No. 4440047) via the Life Technologies 7500 Real-Time PCR System (Life Technologies, Carlsbad, CA, USA). For relative expression analysis three reference genes are measured in parallel (RPLPO, RPL37A and OAZ1).
[0266] Cycling conditions were selected according to the standard method and run up to 40 cycles. A cycle included 2min of 50°C and lOmin at 95°C for holding, as well as 15sec of 95°C and Imin at 60°C as cycling states.
[0267] Ct values were reported once the threshold of 0.2 was exceeded. The detected Ct values were normalized with regard to the measurements of a set of reference genes. Delta Ct (DCt) values were calculated by subtracting the mean Ct of the refence genes (see Table 1). Measurements below the limit of detection and outliers (exceeding the span of mean DCt ± 2.5 standard deviations (SD) or a SD above 0.5 for all three samples) were not considered for further analysis.
[0268] To determine potential contamination with genomic DNA (gDNA), the percentage of gDNA was estimated based on the -RT DCt. Samples with gDNA contamination above the threshold of 10% were not considered for further analysis.
[0269]
[0270] Table 1: Target Assays including reference genes 0AZ1, RPLPO andRPL37A
[0271] 2.7 Peptide-specific parameters
[0272] 2. 7.1 Retention Time
[0273] During liquid chromatography, peptides are separated based on their hydrophobicity. The retention time of a peptide is defined as the elapsed time from the start of the chromatography process and the peak of the detection signal of the target peptide. The retention time of a peptide can be estimated based on the physico- and biochemical properties of its amino acid components. Prosit, a machine learning based model (Gessulat, S. et al, 2019, Nature methods, 16(6), 509-518), was used to generate the standardized retention time (uRT). uRT estimations are based on the amino acid sequence of the peptides and can hence be generated even if the peptide has not been previously detected. The uRT, among other peptide-specific properties, can be a representative variable for peptide-specific quantification variability.
[0274] 2. 7.2 HLA Binding affinity
[0275] The binding affinity of a peptide to the surface receptor can be estimated using NetMHCpan, a machine learning based model (Reynisson B et al, 2020, Nucleic acids research, 48(W1), W449-W454). The predictions are based on the amino acid sequence of the peptides as well as the HLA typing of the target sample. The interaction occurring between peptide and surface receptor has an impact on the detection of peptide quantities, i. e. peptides with high affinity are expected to bind to receptors at lower concentrations which facilitates their detection, and vice versa for low affinity peptides. The binding affinity was utilized as a representative variable for such peptide- and sample specific effects. 3 Predictive Models
[0276] Peptide (absolute and relative), protein and RNAseq quantities were log-transformed prior to model fitting. Predicted peptide femtomoles were converted to copies per cell using sample-specific cell count and peptide-specific isolation efficiency, as described in 2.2.
[0277] 3.1 Peptide-specific models for the prediction of the absolute quantity for one MHC- presented peptide
[0278] To predict the absolute peptide quantity for specific peptides, robust linear regression models with one or more predictor variables were fitted.
[0279] Univariate model:
[0280] Y = / 30 + / 31 * XI + €
[0281] Multivariate model:
[0282] Y = [30 + pi * XI + (32 * X2 + e
[0283] The predictor variables allow the substitute of the measurement of the absolute peptide quantity by surrogate measurements. In both model formulae, XI is the primary surrogate (substitute for direct absolute peptide quantity) and represents relative peptide quantities or protein (source protein) or mRNA (source gene, transcript, exon or peptide-encoding nucleotide sequences) that the peptide is derived from. X2 is the secondary surrogate and represents mRNA or protein quantities that are relevant for antigen processing and presentation.
[0284] Only the univariate model was fitted in case the primary surrogate represented a peptide quantity directly. For all other primary surrogates (not representing direct peptide measurements), twelve multivariate models were generated in addition to the univariate model. Of those, six models comprise secondary surrogates on RNA level and six comprise secondary surrogates on protein level. In both cases, the six models correspond to six different secondary surrogates representing the HLA presenting the peptide (e.g. HLA-A*02:01), Beta-2M (Beta- 2-microglobulin), TAP (transporter associated with antigen processing), immunoproteasome, proteasome and ERAP (ER-associated aminopeptidase). The utilized secondary surrogates were based on gene-level quantifications for RNA-based and on source protein-level (RellBAQ) quantifications for protein-based primary level surrogates. As a negative control, multivariate models with the house-keeping enzyme GAPDH (Glyceraldehyde-3 -Phosphate Dehydrogenase) on gene or protein level as the secondary surrogate were fitted.
[0285] While Ordinary Least Squares (OLS) regression, as described in 3.2, attempts to minimize the sum of squared residuals, robust regression minimizes another function of the residuals which is solved using Iteratively Reweighted Least Squares. Robust regression assigns samples with small residuals a weight of 1 which corresponds to OLS regression where all samples have equal weights. Samples with larger residuals get assigned smaller weights. This alleviates the effect of influential samples that strongly affect the model fit.
[0286] Robust linear regression models were fitted using the R function rim in the R package MASS (Venables W. N. and Ripley B. D., 2002, Modern Applied Statistics with S. Springer. R package version 7.3-59) with the parameters wt.method = "inv.var", maxit = 10000, method = "M", init = "Is", scale. est = "Huber", and psi = "psi.huber". A weighted coefficient of determination (R2) that incorporates robust regression sample weights was computed as a model performance measure. To test for model significance, a robust F-test was applied (function frobftest in the R package sfsmisc; Maechler M., 2023, Utilities from 'Seminar fuer Statistik' ETH Zurich. R package version 1.1-15.). The inferred linear relationship was used to predict total peptide amount in copies per cell or femtomole.
[0287] 3.1.1 Example 1: Peptide-specific regression models using surrogates based on relative peptide quantity
[0288] Evaluable and QC-checked HLA-A*02 positive samples with matched relative and absolute peptide data were used in model fitting. For absolute peptide data, all values with detections were included. Relative quantities were normalized as described in 2.3.2 Replicate values were aggregated using the mean per sample for absolute quantities and the median for relative quantities. Relative peptide quantities were used as primary surrogate. Model fitting and evaluation were performed as described in 3.1.
[0289] 3.1.2 CpC prediction based on RNAseq quantities
[0290] Evaluable and QC-checked HLA-A*02 positive samples with matched RNAseq and absolute peptide data were used in model fitting. For absolute peptide data, all samples with detections were included. Model fitting and evaluation were performed as described in 3.1. 3.1.2.1 Example 4: Peptide-specific regression models using surrogates based on gene expression
[0291] The dataset described in 3.1.2 was used. To generate the primary surrogate, gene expression was computed by summing up the normalized TPM of all transcripts that belong to the source gene. Model fitting and evaluation were performed as described in 3.1
[0292] Table 2: Multivariate regression models for prediction of copies per cell (CpC) for Peptide 003 based on gene expression (Pri) and secondary surrogate (Sec) measured either on RNA (column 2) or protein level (column 3)
[0293] 3.1.2.2 Example 5: Peptide-specific regression models using surrogates based on transcript expression
[0294] The dataset described in 3.1.2 was used. To generate the primary surrogate, transcript expression was defined as the maximum normalized TPM of all source transcripts. Model fitting and evaluation were performed as described in 3.1
[0295] Table 3: Multivariate regression models for prediction of copies per cell (CpC) for Peptide 003 based on transcript expression (Pri) and secondary surrogate (Sec) measured either on RNA (column 2) or protein level (column 3) 3.1.2.3 Example 6: Peptide-specific regression models using surrogates based on exon expression
[0296] The dataset described in 3.1.2 was used. To generate the primary surrogate, exon expression was defined as the maximum normalized TPM of all source exons. Model fitting and evaluation were performed as described in 3.1.
[0297] Table 4: Multivariate regression models for prediction of copies per cell (CpC) for Peptide 002 based on exon expression (Pri) and secondary surrogate (Sec) measured either on RNA (column 2) or protein level (column 3)
[0298] 3.1.2.4 Example 7: Peptide-specific regression models using surrogates derived from peptide- encoding nucleotide sequences
[0299] The dataset described in 3.1.2 was used. The primary surrogate was generated based on the customized aggregation of reads from the corresponding peptide-encoding nucleotide sequences (PNS), as described in 2.5.5. Model fitting and evaluation were performed as described in 3.1.
[0300] Table 5: Multivariate regression models for prediction of copies per cell (CpC) for Peptide 002 based on PNS levels and Secondary surrogate (Sec) measured either on RNA (column 2) or protein level (column 3) 3.1.2.5 Average performance of regression for models for prediction of copies per cell based on RNASeq-based primary surrogates
[0301] Table 6: Average performance (R2) of significant (p-value < 0.05) univariate and multivariate regression models for prediction of copies per cell (CpC) across all peptides based on RNASeq surrogates (Pri) and secondary surrogate (Sec) measured either on RNA (column 3) or protein level (column 4)
[0302] 3.1.3 Example 3: Peptide-specific regression models using surrogates based on relative protein quantity
[0303] The dataset described in 3.1.2 was used. The relative protein quantity (RellBAQ values) of the protein containing the target peptide was utilized as a primary surrogate. The customized processing is described in 2.4.3. Model fitting and evaluation were performed as described in 3.1.
[0304] Table 7: Multivariate regression models for prediction of copies per cell (CpC) of Peptide 004 based on the relative protein quantity (Pri) and secondary surrogate (Sec) measured either on RNA (column 2) or protein level (column 3)
[0305] 3.1.4 Example 9: Peptide-specific regression models using surrogates based on qPCR
[0306] Samples with matched qPCR and absolute peptide data were selected. For absolute peptide data, all samples with detections were included. The primary surrogates were qPCR DCt values without log-transformation. Replicate values for absolute peptide quantities, primary and secondary surrogate data were aggregated using the mean per sample. Model fitting and evaluation were performed as described in 3.1.
[0307] Table 8: Multivariate regression models for prediction of copies per cell (CpC) of Peptide 005 based on gene quantity derived from qPCR (Pri) and Secondary surrogate (Sec) measured either on RNA (column 2) or protein level (column 3)
[0308] 3.2 Universal models for the prediction of the absolute quantity for more than one MHC- presented peptide
[0309] Linear regression was utilized to predict the absolute quantity of MHC-presented peptides based on the corresponding relative quantities either as determined by MS, such as described in 3.1.1 or as determined by RNASeq, such as described in 3.1.2.
[0310] The equation of a standard linear regression model can be described as: Y = (30 + (31 * Xl + e.
[0311] The absolute quantity or dependent variable is represented by Y. It is determined based on the independent variable XI, the relative quantity of the corresponding peptide in the corresponding sample. (31 describes the slope or regression coefficient of X and (30 the intercept of the model. An error term e is included to represent the discrepancy between the predicted values and the actual observed quantities. The parameters intercept, slope and additional regression coefficients, if multiple variables are included in the model, are estimated using the ordinary least squares (OLS) approach. During OLS, the parameters are determined by minimizing the sum of squared errors (Montgomery, D. C., Peck, E. A., & Vining, G. G., 2021; Introduction to linear regression analysis. John Wiley & Sons).
[0312] The performance of the model was determined based on the adjusted R-squared (R2). The metric is an extension of the regular R2value which is a statistical measure reporting the goodness of fit of a model based on the proportion of variance in the dependent variable explained by the independent predictor variables. The adjusted R2accounts for the inclusion of additional variables. In addition, the p-value of the F-test for the generated model is reported.
[0313] 3.2.1 Example 2: Universal regression models using surrogates based on relative peptide quantity
[0314] A generalized model was developed to facilitate the prediction of absolute peptides quantities of multiple MHC-presented peptides. To this end, the univariate model, described in
[0315] 3.1.1 was extended with an additional variable to account for peptide-specific variabilities.
[0316] Evaluable and QC-checked HLA-A*02 positive samples with matched relative and absolute peptide data measured via MS, as described in 2.2 and 2.3, were included in the training dataset. All peptides which fulfilled the QC criteria were included. The predicted uRT was used to address differences between the peptides. For each peptide the uRT was estimated via Prosit, as described in 2.7.
[0317] The generalized model is formally described as
[0318] Y = (30 + (31 * XI + (32 * uRT + e
[0319] The dependent variable Y represents the observed absolute peptide quantities, while XI represents the primary surrogate in this case the relative peptide quantities. The universal slope (31 and intercept (30 apply to all peptides. The regression coefficient (32 measures the effect of the peptide-specific parameter, such as the uRT. Predicted absolute peptide quantities in femtomole were transformed to copies per cell as described in 2.2. 3.2.2 Example 8: Universal regression models using surrogates based on transcript expression
[0320] Absolute quantities for individual peptides can be predicted based on the expression of the corresponding transcripts in the same sample as described in chapter 3.1.2. A generalization of the model extends the applicability from singular peptides to multiple peptides.
[0321] Evaluable and QC-checked HLA-A*02 positive samples with matched absolute peptide data measured via MS, as described in 2.3, as well as relative RNASeq expression of the corresponding transcript, as described in 2.5, were included in the training dataset.
[0322] The generalized model is formally described as:
[0323] K = ?0 + ?1 * XI + ?2 * uURT + 3 * MHC Affinity + 4 *
[0324] The linear regression model, as described in 3.2, was extended by additional variables.
[0325] The dependent variable Y represents the observed absolute peptide quantities, while XI represents the primary surrogate, in this case the expression of the corresponding transcripts as quantified by RNA Sequencing. The model includes a universal slope / 31 and a universal intercept / 30 which is identical across all peptides. Further variables were added to consider peptide- and sample-specific variability and facilitate the generalization. The uRT was added to account for peptide-specific variations. The abundance of MHC was determined based on RNASeq data via arcasHLA, as described in 2.5.4, and added to account for sample-specific variation. Lastly, the binding affinity of peptides to the MHC was determined via NetMHCpan and added to account for the interaction between both the peptide and the surface receptor. Estimation of binding affinity is described in 2.7.
[0326] Finally, the model was extended in a stepwise approach, by adding variables accounting for peptide processing and presentation. These sample-specific variables were included during model generation, corresponding regression coefficients were estimated, and the model was evaluated. Performance metrics for judging the benefit of these additional factors include the adjusted R2and p-value of the model, as described in 3.1.2 and a variance inflation estimate. Only factors of significance, which increased the predictive performance without adding collinearity were included in the finalized model. 4 Overview peptide sequences with Alias
[0327] Table 9: Amino acid sequences of peptides used in the examples
Claims
Claims1. A method for absolute quantification of at least one MHC-presented peptide in a further biological sample comprising the following steps:(a) determining the absolute quantity of the at least one MHC-presented peptide in at least two, preferably at least four, biological samples;(b) determining in the at least two, preferably at least four, biological samples of step (a):(i) the relative quantity of the at least one MHC-presented peptide; and / or(ii) the relative or absolute quantity of at least one surrogate for each of the at least one MHC-presented peptides;(c) applying statistical modelling to the at least one absolute quantity of step (a) and the at least one relative or absolute quantity of step (b) to establish at least one statistical model;(d) determine in the further biological sample the at least one relative or absolute quantity as in step (b);(e) predict the absolute quantity of the at least one MHC-presented peptide in the further biological sample using the at least one statistical model.
2. A method for generating a statistical model for absolute quantification of at least one MHC-presented peptide in a further biological sample comprising the following steps:(a) determining the absolute quantity of the at least one MHC-presented peptide in at least two, preferably at least four, biological samples;(b) determining in the at least two, preferably at least four, biological samples of step (a):(i) the relative quantity of the at least one MHC-presented peptide; and / or(ii) the relative or absolute quantity of at least one surrogate for each of the at least one MHC-presented peptides;(c) applying statistical modelling to the at least one absolute quantity of step (a) and the at least one relative or absolute quantity of step (b) to establish at least one statistical model.
3. The method of any of the preceding claims, wherein the surrogate is a primary surrogate or a secondary surrogate.
4. The method of claim 3, wherein the primary surrogate is derived from:(I) the nucleic acid sequence encoding the at least one MHC-presented peptide, preferably selected from gene expression, genomic variation, transcript variation, alternative splicing; or(II) the protein of which the at least one MHC-presented peptide is derived from, preferably selected from protein expression.
5. The method of claim 3 or 4, wherein the secondary surrogate is another parameter relevant for the at least one MHC-presented peptide, preferably selected from HLA, proteasome, immunoproteasome, beta-2 microglobulin, TAP and ERAP.
6. The method of any of the preceding claims, wherein the secondary surrogate is the HLA that is involved in presenting the at least one MHC-presented peptide.
7. The method of any of the preceding claims, wherein in step (c): only one quantity of step (b) is used for statistical modelling (i.e. univariate analysis); or at least two quantities of step (b) are used for statistical modelling (i.e. multivariate analysis).
8. The method of any of the preceding claims, wherein, in case of determining more than one quantity in step (b) (i.e. multivariate analysis), this includes the quantity of at least one primary surrogate or the relative quantity of the at least one MHC-presented peptide, optionally further including at least one secondary surrogate.
9. The method of any of the preceding claims, wherein the more than one quantity determined in step (b) (i.e. multivariate analysis) includes: relative quantity of the at least one MHC-presented peptide and relative mRNA expression; or relative mRNA expression and genomic variation.
10. The method of any of the preceding claims, wherein the more than one quantity determined in step (b) (i.e. multivariate analysis) includes the relative quantity of- the at least one MHC-presented peptide and- the specific HLA- Allele that is presenting the at least on MHC-presented peptide.
11. The method of any of the preceding claims, wherein the statistical modelling of step (c) is a regression analysis, preferably linear regression.
12. The method of any of the preceding claims, wherein the statistical modelling of step (c), preferably regression analysis, comprises the determination of peptide-specific off-set parameter, preferably the peptide-specific off-set parameter is determined based on: physicochemical and biochemical properties of amino acids, preferably selected from amino acid index or Hydrophobicity score; peptide retention time in liquid chromatography; sequence specific peptide bias, or binding affinity.
13. The method according to any one of the preceding claims, wherein the absolute quantity of more than one MHC-presented peptide is predicted by establishing a universal statistical model in step (c).
14. An antigen-binding protein binding to the at least one MHC-presented peptide, at least one MHC-presented peptide or the pharmaceutical salt thereof for use in the treatment of a subject suffering from cancer, wherein the absolute quantity of at least one MHC-presented peptide in a further biological sample of the subject is determined according to any one of the preceding claims, wherein the absolute quantity of at least one MHC-presented peptide is indicative of the suitability to treat the subject with an antigen-binding protein binding to the at least one MHC-presented peptide, the at least one MHC-presented peptide or the pharmaceutical salt thereof, wherein an antigen binding protein binding to the at least one MHC-presented peptide, the at least one MHC-presented peptide or the pharmaceutical salt thereof is to be administered to the subject.
Citation Information
Patent Citations
Method for the absolute quantification of naturally processed HLA-restricted cancer peptides
US10545154B2
Method for the absolute quantification of MHC molecules
US20220283176A1
Machine-learning techniques for predicting surface-presenting peptides
US20230115039A1
Methods and systems for prediction of HLA epitopes
WO2024036308A1