Method for identifying splicing peptides

The method simplifies and accelerates the identification of splicing peptides by using mass spectrometry and a secondary library of likely variants, addressing efficiency and personalization challenges in cancer vaccine therapy.

JP7727735B2Active Publication Date: 2025-08-21PROVIDENCE HEALTH & SEVICES OREGON +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023538484
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-07-28
Filing Date
2022-07-22
Publication Date
2025-08-21
Estimated Expiration
2042-07-22

AI Technical Summary

Technical Problem

Existing methods for identifying antigenic peptides, including splicing peptides, face challenges such as low throughput, complex data processing, and difficulty in efficiently identifying splicing peptides due to isobaric amino acids, making it difficult to select appropriate peptides for cancer vaccine therapy based on patient-specific HLA types.

Method used

A method involving mass spectrometry, a primary library search for parent peptides, creation of a secondary library of candidate spliced peptides, and a search within this library to identify spliced peptides, focusing on splicing reactions and likely variants.

Benefits of technology

Enables efficient and simple identification of unknown splicing peptides, reducing processing time and cost, and facilitating personalized cancer vaccine development by identifying splicing peptides likely to bind strongly to patient-specific HLA types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007727735000005
    Figure 0007727735000005
  • Figure 0007727735000006
    Figure 0007727735000006
  • Figure 0007727735000007
    Figure 0007727735000007
Patent Text Reader

Abstract

Provided is a method for identifying an unknown splicing peptide derived from a known parent peptide included in a biological sample, wherein the method includes: (1) a first step for obtaining mass spectrometry data pertaining to peptides included in the sample by implementing mass spectrometry of the sample; (2) a second step for searching a primary library, which is a database of known amino acid sequences, for an amino acid sequence that matches the mass spectrometry data, and identifying the parent peptide included in the sample; (3) a third step for creating a secondary library including candidates of splicing peptides that can be produced from the identified parent peptide; and (4) a fourth step for searching the secondary library for an amino acid sequence that matches the mass spectrometry data, and identifying the splicing peptide included in the sample.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for identifying splicing peptides. [Background technology]

[0002] A complex process is required for the immune system to recognize self and non-self and defend itself. The immune system is activated by two mechanisms: innate immunity and adaptive immunity. Innate immunity primarily functions to quickly recognize and eliminate foreign substances. A complex adaptive immune system exists to complement innate immunity.

[0003] In the adaptive immune system, when a foreign substance enters the body or when foreign substances such as cancer cells are generated in the body, the foreign substance is first digested (broken down) by macrophages and other cells, and the decomposed foreign substance is recognized and taken up by dendritic cells. These cells are called antigen-presenting cells.

[0004] Foreign substances (proteins) taken up by dendritic cells are transported into the cell and degraded into peptides by lysosomes, proteasomes, etc. in the cytoplasm. These peptides are presented to lymphocytes (T cells and B cells) bound to the major histocompatibility complex (MHC) present within the cell, and are recognized by the lymphocytes as antigens. In humans, the MHC is the human leukocyte antigen (HLA).

[0005] The functions of lymphocytes (T cells and B cells) consist of cytotoxic immunity (T cells) and humoral immunity (B cells). Mature T cells have the ability to target recognized antigens and launch a concentrated attack on foreign substances. B cells produce antibodies that bind to recognized antigens and have the ability to neutralize or inactivate foreign substances (antigens).

[0006] In recent years, cancer immunotherapy has been attracting attention as a way to enhance the effectiveness of cancer treatments (surgery, chemotherapy, radiation therapy, etc.). One example of this cancer immunotherapy is a method known as peptide vaccine therapy, in which cancer-specific peptides (peptides that are found only in cancer cells and not in normal cells) are administered as a vaccine to enhance the effectiveness of cancer treatment through the function of adaptive immunity.

[0007] HLA has a genotype (haplotype) determined by a pair of gene loci inherited from each parent, and it is said that there are tens of thousands of different haplotypes. HLA has a variety of phenotypes (HLA types) corresponding to the various haplotypes. It is also known that the peptides that bind to HLA as antigens presented by antigen-presenting cells (HLA-binding peptides) differ depending on the HLA type.

[0008] Therefore, whether or not a peptide vaccine binds to HLA and is presented as an antigen to lymphocytes after vaccination, thereby triggering adaptive immunity, may depend on the patient's HLA type. Therefore, it is desirable to select peptides used as vaccines in cancer vaccine therapy appropriately according to the patient's HLA type.

[0009] Furthermore, protein splicing was previously thought to occur during the editing of mRNA transcribed from template DNA. However, in recent years, it has been discovered that a reaction (peptide splicing reaction) occurs during the degradation of foreign proteins in the proteasome, resulting in substitutions or other changes in peptide sequences (see, for example, Patent Document 1 (European Patent Application Publication No. 2362225)). This peptide splicing reaction results in even greater diversity in antigen peptides, which may further increase the diversity of immune responses between individuals.

[0010] Considering the existence of such diverse splicing peptides and the diversity of HLA-binding peptides, approaches are being attempted to accurately identify as many antigen peptides as possible. [Prior art documents] [Patent documents]

[0011] [Patent Document 1] European Patent Application Publication No. 2362225 Summary of the Invention [Problem to be solved by the invention]

[0012] For example, a method has been investigated for identifying antigenic peptides, including splicing peptides, by synthesizing a protein containing a candidate antigenic peptide region and then actually subjecting the protein to proteasome degradation, etc. This method can reliably obtain antigen information and confirm splicing, but has the drawback of low throughput and making efficient identification difficult.

[0013] Another method is to perform de novo sequencing to identify unknown antigenic peptides, including spliced ​​peptides, using the large number of unidentified MS / MS spectra generated in mass spectrometry data. The amino acid sequences of the identified peptides are then identified as antigenic peptides by setting a cutoff value based on a predetermined identification score (ALC: average local confidence, by Peaks Studio software). However, this method requires complex data processing and is difficult to perform efficient identification. Furthermore, the inability to separate isobaric amino acids is a fundamental limitation.

[0014] Identification of splicing peptide sequences that are candidates for neoantigens cannot be achieved by cancer-specific mutation information obtained by genetic analysis or by predicting or matching sequences based on genetic templates. However, antigenic peptides identified by mass spectrometry, peptides identified by de novo sequence analysis, and information on the identification of their splicing sites and affinity with HLA are stored in resource databases (amino acid sequence databases), and are used as prediction engines for predicting affinity between HLA and antigenic peptides, structural identification, etc.

[0015] The present invention has been made to solve the problems of the conventional identification methods described above, and aims to provide a method that can simply and efficiently identify unknown splicing peptides. [Means for solving the problem]

[0016] The present invention provides a method for identifying unknown spliced ​​peptides derived from known parent peptides contained in a biological sample, comprising the steps of: (1) a first step of performing mass spectrometry on the sample to obtain mass spectrometry data of peptides contained in the sample; (2) a second step of searching a primary library, which is a database of known amino acid sequences, for amino acid sequences that match the mass spectrometry data to identify the parent peptide contained in the sample; (3) a third step of generating a secondary library containing candidate spliced ​​peptides that can be generated from the identified parent peptides; (4) a fourth step of searching the secondary library for an amino acid sequence that matches the mass spectrometry data to identify the spliced ​​peptide contained in the sample. [Effects of the Invention]

[0017] According to the present invention, unknown splicing peptides can be identified simply and efficiently by using a group of candidate splicing peptides that can arise from a specific parent peptide identified by a primary library as a secondary library. [Brief explanation of the drawings]

[0018] [Figure 1] FIG. 1 is a flow chart illustrating a method for identifying a splicing peptide according to an embodiment. [Figure 2] FIG. 1 is a schematic diagram showing the reaction mechanism of peptide splicing via an acyl intermediate. [Figure 3] FIG. 1 is a schematic diagram illustrating an in vitro test for confirming peptide splicing in Reference Example 1. [Figure 4] FIG. 1 shows the types of amino acid residues at the N-terminus of splicing peptides generated from S-220. [Figure 5] FIG. 1 shows the types of amino acid residues at the N-terminus of splicing peptides generated from S-230. [Figure 6] FIG. 1 shows the types of amino acid residues at the N-terminus of splicing peptides generated from S-260. [Figure 7] FIG. 1 shows the types of amino acid residues at the N-terminus of splicing peptides generated from S-280. [Figure 8] FIG. 1 shows the types of amino acid residues at the N-terminus of splicing peptides generated from S-300. [Figure 9] FIG. 1 shows the types of amino acid residues at the N-terminus of splicing peptides generated from S-310. [Figure 10] FIG. 1 shows the types of amino acid residues at the N-terminus of splicing peptides generated from S-320. [Figure 11] FIG. 1 shows the types of amino acid residues at the N-terminus of splicing peptides generated from S-330. DETAILED DESCRIPTION OF THE INVENTION

[0019] As shown in FIG. 1, the method for identifying a splicing peptide according to this embodiment includes the following steps: 1. A method for identifying unknown spliced ​​peptides derived from known parent peptides contained in a biological sample, comprising: (1) A first step (S1) of performing mass spectrometry on a sample to obtain mass spectrometry data of peptides contained in the sample; (2) A second step (S2) of searching a primary library, which is a database of known amino acid sequences, for amino acid sequences that match the mass spectrometry data to identify parent peptides contained in the sample; (3) A third step (S3) of creating a secondary library containing candidate spliced ​​peptides that can be generated from the identified parent peptides; (4) A fourth step (S4) of searching the secondary library for amino acid sequences that match the mass spectrometry data to identify spliced ​​peptides contained in the sample.

[0020] The identification method of this embodiment is a method for identifying unknown spliced ​​peptides derived from known parent peptides contained in a biological sample. Biological samples are, for example, samples obtained from a living organism, such as fluids containing cancer cells, culture samples of cancer cell lines, excised cancer tissues, xenograft tissues, blood, and exosome fractions.

[0021] Each step of the identification method will be described in detail below.

[0022] (First step: S1) In this embodiment, in the first step (S1), Mass spectrometry is performed on the sample to obtain mass spectrometry data for peptides contained in the sample.

[0023] Examples of mass spectrometry include ionizing proteins using ionization methods such as matrix-assisted laser desorption / ionization (MALDI) and electrospray ionization (ESI), and then quantitatively analyzing the signal intensity of each peak resolved according to the mass / charge ratio (m / z) of the protein ion. Mass spectrometry may also be performed by liquid chromatography-mass spectrometry (LC-MS) or liquid chromatography-tandem mass spectrometry (LC-MS / MS).

[0024] As the mass spectrometer, in addition to a general single-type mass spectrometer, a tandem mass spectrometer such as a triple quadrupole (QqQ) mass spectrometer, a quadrupole time-of-flight (Q-TOF) mass spectrometer, a tandem time-of-flight (TOF-TOF) mass spectrometer, a quadrupole ion trap (QIT) mass spectrometer, or a quadrupole ion trap time-of-flight (QIT / TOF) mass spectrometer can be suitably used.

[0025] In LC-MS / MS, multiple reaction monitoring (MRM) may be performed using a triple quadrupole mass spectrometer. The MRM method is a more sensitive analytical method (LC-MS / MS) than the SIM method (LC / MS).

[0026] The mass spectrometry data may be a mass spectrometry spectrum (MS spectrum, MS / MS spectrum, etc.).

[0027] (Second step: S2) In this embodiment, in the second step (S2), The parent peptides contained in the sample are identified by searching a primary library, which is a database (dataset) of known amino acid sequences, for amino acid sequences that match the mass spectrometry data.

[0028] The parent peptide may be a known HLA-binding peptide (such as an HLA Class I-binding peptide or an HLA Class II-binding peptide). The parent peptides may be known peptides derived from cancer cells (cancer antigen peptides, a set of predicted sequences of peptides containing cancer mutated regions). The parent peptide may be a known HLA-binding peptide derived from a cancer cell.

[0029] HLA types are mainly classified into two classes (HLA Class I and HLA Class II). The length of peptides that can bind to HLA is approximately 8 to 13 residues for HLA Class I and approximately 10 to 25 residues for HLA Class II. HLA Class I is present in almost all cells, binds to peptides generated by proteasomal degradation of intracellular foreign substances (proteins), and presents the peptides as antigens. HLA Class II binds to peptides generated by lysosomal degradation of exogenous foreign substances, and presents the peptides as antigens.

[0030] In addition, peptides (antigens) that bind to HLA Class I are directly involved in the activation of T cells, and therefore, peptides that bind to HLA Class I are extremely important in cancer treatment.

[0031] Many resource databases (amino acid sequence databases) have been developed to identify cancer-specific peptides, etc. Representative examples include web tools such as NetMHCpan (http: / / www.cbs.dtu.dk / services / NetMHCpan / ), IEDB (https: / / www.iedb.org / ), HLA Ligand Atlas (https: / / hla-ligand-atlas.org / ), SysteMHC (https: / / systemhcatlas.org / ), WebLogo (https: / / weblogo.berkeley.edu / logo.cgi), and HLAthena (http: / / hlathena.tools / ).

[0032] An "amino acid sequence that matches mass spectrometry data" is, for example, when the mass spectrometry data is a mass spectrometry spectrum (MS spectrum, MS / MS spectrum, etc.), an amino acid sequence that corresponds to the mass spectrometry spectrum falls within an acceptable range such that it is determined to be highly correlated with the mass spectrometry spectrum by multivariate analysis (principal component analysis, etc.).

[0033] In other words, an "amino acid sequence that matches mass spectrometry data" means that the mass spectrometry data of the amino acid sequence has a correlation with the mass spectrometry data of the sample within a range that is acceptable in the art when identifying an amino acid sequence based on mass spectrometry data, and does not necessarily have to have a perfect correlation. For example, a certain degree of difference in the mass spectrometry spectra between the two due to small amounts of impurities remaining (that could not be completely removed) in the sample ultimately subjected to mass spectrometry spectrum measurement is acceptable.

[0034] An amino acid sequence that matches mass spectrometry data may be an amino acid sequence that corresponds to a mass spectrometry spectrum having a peak at the same m / z value (mass-to-charge ratio) as at least one peak in the mass spectrometry spectrum of the sample, or at an m / z value whose deviation is within a predetermined range.

[0035] Here, the threshold used when comparing m / z values ​​to determine an amino acid sequence that matches the mass spectrometry data may be set to an absolute deviation of 0.5 or less, or 0.2 or less, between two corresponding m / z values. The m / z value is calculated based on the molecular weight m and charge state z (1, 2, 3, or 4) of the peptide in the mass spectrometry data.

[0036] Various known identification engines can be used to identify the amino acid sequences of peptides, including Mascot and Peaks Studio. For example, the Mascot system is protein identification software that searches protein and genome sequence databases for amino acid sequences that match peptide mass spectrometry data obtained from a mass spectrometer, and identifies proteins or peptides contained in a measurement sample. It employs a probabilistic scoring algorithm, which allows statistically significant proteins or peptides to be clearly distinguished and visualized by score.

[0037] (Third step: S3) In this embodiment, in the third step (S3), A secondary library is created containing candidate spliced ​​peptides that can arise from the identified parent peptide.

[0038] Preferably, the candidate group consists of spliced ​​peptides that are likely to arise from the parent peptide.

[0039] In this case, the computer processing time for identification can be reduced to a practical level compared to when all possible combinations of variant peptides that can be generated by splicing from the parent peptide are added as candidates, thereby reducing the cost of identification.

[0040] The candidate group preferably consists of variant peptides generated by at least one of the following: deletion of a predetermined number of amino acid residues from at least one of the C-terminus and the N-terminus of the parent peptide; and addition of a predetermined number of amino acid residues to the C-terminus and the N-terminus. In this case, variant peptides generated by "substitution" where the number of added and deleted amino acid residues is the same are also included in the candidate group.

[0041] In this case, the candidate group will consist of spliced ​​peptides that are likely to arise from the parent peptide.

[0042] The candidate group may be all or some of the combinations of variant peptides that can be generated by at least one of deleting a predetermined number of amino acid residues or less from the C-terminus or N-terminus of the parent peptide, and / or adding a predetermined number of amino acid residues or less to the C-terminus or N-terminus.

[0043] The predetermined number is preferably 2. That is, the candidate group is preferably generated by at least one of deleting 2 or less amino acid residues from at least either the C-terminus or the N-terminus of the parent peptide and adding 2 or less amino acid residues to the C-terminus or the N-terminus (see Table 1 below).

[0044] In this case, the candidate group will consist of spliced ​​peptides that are more likely to arise from the parent peptide, which further reduces the computer processing time required for identification and reduces the cost of identification.

[0045] Previous methods for identifying neoantigens, such as unknown splicing peptides, relied on mass spectrometry data, such as MS (mass spectrometry) spectra and MS / MS (tandem mass spectrometry) spectra, to obtain mass spectrometry data for many antigenic peptides and identify their amino acid sequences. Amino acid sequence identification was primarily achieved by sequence database matching or de novo sequence analysis.

[0046] On the other hand, the present inventors focused on the reaction mechanism of peptide splicing. As shown in Figure 2, the peptide splicing reaction is thought to be a reaction in which a part of the peptide is substituted by a nucleophilic substitution reaction with an acyl intermediate at the enzyme active center during the pathway leading to hydrolysis (see the peptide splicing pathway indicated by the black arrow in Figure 2).

[0047] Thus, if peptide splicing occurs via an acyl intermediate, the probability of the splicing reaction occurring is overwhelmingly favorable if the nucleophilic substituent attacks the acyl intermediate, which has a short half-life. Based on this principle, it is thought that the nucleophilic substituent can be generated more quickly and randomly with smaller molecules.

[0048] In addition, there are many candidate molecular species present in biomolecules that can attack the acyl intermediate in the peptide splicing reaction. However, considering that trypsin activity, chymotrypsin activity, peptidylglutamyl aminopeptidase activity, etc. are locally maintained within the cavity of the proteasome (enzyme complex) and have the property of actively incorporating ubiquitinated proteins, the nucleophilic substituent is thought to be limited to peptides or amino acids.

[0049] Therefore, it is believed that the molecules that are favorable as nucleophilic substituents that attack the acyl intermediate are amino acids and peptides with a relatively small number of residues (dipeptides, etc.), with amino acids being the most favorable.

[0050] However, because previous methods for identifying splicing peptides have relied on MS / MS spectrum matching, amino acid sequence homology searches, and de novo sequence analysis techniques, the peptides that could be identified were essentially limited to those for which local alignment could be recognized. In other words, if there was no amino acid sequence of three or more residues, it was not possible to identify the peptide sequence.

[0051] Therefore, when an amino acid or dipeptide becomes a nucleophilic substituent in the above-mentioned peptide splicing reaction (when the number of amino acids constituting the nucleophilic substituent is one or two), it is difficult to identify the spliced ​​peptide using the conventional identification approach based on local alignment.

[0052] Therefore, the identification method of this embodiment is particularly useful when identifying splicing peptides (variant peptides) that are likely to be generated by substitution of one or two amino acid residues, etc.

[0053] Furthermore, no method for analyzing spliced ​​peptides that focuses on the basic mechanism of the splicing reaction has been known to date.

[0054] The number of residues in the amino acid sequences constituting the above candidate group may be, for example, 5 to 15 residues.

[0055] For example, when the parent peptide is composed of 9 amino acid residues and the predetermined number is 2, a preferred example of a candidate group of splicing peptides to be added to the primary library to obtain a secondary library is shown in Table 1. Note that the parent peptide is not included in the candidate group. In Table 1, N1, N2, C1 and C2 each represent an amino acid selected from 20 types of amino acids, and the peptide sequence shown in each row has the number of combinations shown in the right column.

[0056] [Table 1]

[0057] (Fourth step: S4) In this embodiment, in the fourth step (S4), The secondary library is searched for amino acid sequences that match the mass spectrometry data to identify spliced ​​peptides contained in the sample.

[0058] Specifically, steps 3 and 4 involve randomly generating new sequences as FASTA files by substituting one or two amino acid residues at either end of a peptide list (parent peptide list) identified by database search. The FASTA files (secondary libraries) are then reconstructed on the Mascot server (Matrix Science) and searched again to identify spliced ​​peptides.

[0059] The above describes a method for randomizing the sequence of an identified peptide and then performing a re-search. However, with the development of software in the future, it may be possible to develop analytical sequences that virtually construct predicted sequences and perform re-searches, as in error-tolerant searches.

[0060] (Selection of peptide vaccines in cancer vaccine therapy) The ultimate goal of the search for cancer neoantigens is, for example, the development and evaluation of effective cancer vaccines and cell therapies. Considering the nature of cancer, attenuated vaccines such as viruses cannot be used as cancer vaccine candidates. Therefore, trials have been conducted using classical cancer antigens (molecules that are overexpressed in cancer) and oncogenes (genes that are prone to accumulating cancer-specific mutations) as cancer vaccine candidates, but unfortunately, they have not been successful.

[0061] Currently conceivable new cancer vaccines focus on fragments containing cancer gene mutations, incompletely translated proteins, splicing peptides, etc., and it is desirable to implement personalized medicine in which peptide vaccines are selectively administered by comprehensively utilizing information on the HLA-binding peptides presented by these, HLA typing for each individual (using peptide vaccines with high affinity for HLA alleles expressed in each HLA type), and assay results such as microsatellite instability. Here, the identification method of this embodiment can be applied to the efficient search for splicing peptide antigens of cancer.

[0062] Numerous machine learning-based predictors have been developed to identify immunogenic T cell epitopes based on binding affinity to major MHC (HLA) classes I and II. Various known tools can be used to predict HLA affinity, including NetMHCpan (DTU Health Tech, Center for Biological Sequence Analysis, Bioinformatic unit, Technical University of Denmark), Mascot proteome server (Matrix Sciences), and PEAKS X (Bioinformatics Solutions).

[0063] In the identification method of this embodiment, a public database of protein amino acid sequences can be used as the database for identification, but this can easily be replaced with individual cancer clinical sequence information. This not only makes it possible to search for HLA-binding peptides directly linked to genetic information specific to cancer patients, but it is also believed that the identification method of this embodiment can be fully applied to new medical treatments corresponding to personalized medicine, such as evaluation and monitoring of new cancer vaccines, viral vaccine development, drug efficacy and safety evaluation, and immune monitoring.

[0064] Furthermore, by using the identification method of this embodiment, it is possible to efficiently detect splicing fragments of random HLA-binding peptides whose local alignment cannot be predicted, which is thought to enable the search (identification) of novel cancer antigens (cancer neoantigens). [Example]

[0065] The present invention will be described in more detail below with reference to examples, but the present invention is not limited to these. In the following, amino acids (residues) may be represented by three-letter abbreviations.

[0066] <Reference example 1> In vitro assay to identify spliced ​​peptides generated during the proteasome reaction

[0067] (i) Proteasome reaction (peptide splicing reaction) 20S immunoproteasome (15 nM, R&D Systems), PA28α activator (150 nM, R&D Systems), a mixture of 20 amino acids (each at 25 μM, Cambridge Isotope and Sigma-Aldrich), and proteasome substrates (1 μM of an equal mixture of the following eight substrates) were reacted in 25 mM Tris-HCl buffer at 37°C for 5 hours or overnight. The reaction was terminated by adding FA (final concentration: 1% by mass) and ACN (final concentration: 10% by mass).

[0068] As proteasome substrates (peptides degraded by the proteasome reaction), an equal mixture of the following eight substrates (S-220, S-230, S-260, S-280, S-300, S-310, S-320, and S-330: R&D Systems) was used. S-220: Z-LLL-AMC S-230: Z-LLE-AMC S-260: Suc-LY-AMC S-280: Suc-LLVY-AMC S-300: Boc-LRR-AMC S-310: Ac-PAL-AMC S-320: Ac-ANW-AMC S-330: Ac-WLA-AMC In the notation shown on the right side above, AMC (7-Amino-4-methylcoumarin) on the C-terminal side is a fluorescent chromophore of aminocoumarin, which is released by hydrolysis of the substrate and can be monitored at a fluorescence wavelength (Em) of 345 nm / excitation wavelength (Ex) of 445 nm. Z on the N-terminal side is a benzoyl group, Suc is a succinyl group, Boc is a tert-butoxycarbonyl group, and Ac is an acetyl group. Z, Suc, Boc, and Ac are protecting groups, which may be abbreviated as PG below. Others are one-letter amino acid abbreviations (eg, L for leucine, E for glutamic acid).

[0069] (ii) Mass spectrometry of each peptide contained in the proteasome reaction mixture For each of the above substrates (peptides), the C-terminal AMC was hydrolyzed to release the amino acid, and the addition of an amino acid to the C-terminus of the peptide was monitored by MRM (multiple reaction monitoring). Optimal MRM conditions were set for each substrate.

[0070] Mass spectrometry of each peptide contained in the reaction mixture after the proteasome reaction was performed using a liquid chromatograph mass spectrometer (LC-MS) (LCMS-8050, Shimadzu Corporation) using the MRM method described above. The transitions detected by LC-MS were the substrate PG-AAs-AMC (PG: protecting group, AAs: amino acids), the free chromophore Free AMC, the substrate hydrolysis product PG-AAs-COOH, and the splicing product PG-AAs-X-COOH (X represents a random amino acid residue).

[0071] The analytical conditions using the LCMS-8050 are as follows: Solvent A: A solvent containing 0.1% by mass of FA, the remainder being water Solvent B: a solvent containing 0.1% by weight of FA and 100% by weight of ACN Separation column: Shimpack GISS, 2μm, 2×50mm Flow rate: 0.4mL / min Interface: ESI interface Transition time: 10msec Collision gas pressure: 270kPa Interface temperature: 300℃ DL temperature: 250℃ Heat block temperature: 400℃ Nebulizer gas: 3L / min Heating gas: 10L / min Dry gas: 10L / min Interface voltage: 4kV

[0072] In this way, we confirmed the sequences of the peptides actually produced by splicing for each of the eight proteasome substrates (S-220, S-230, S-260, S-280, S-300, S-310, S-320, and S-330). The analytical results for each substrate are shown in Figures 4 to 11, respectively.

[0073] 4 to 11, the intensity on the vertical axis is the measured intensity (unit: cps) value in the MRM transition corresponding to each amino acid (single-letter abbreviation) shown on the horizontal axis for X (the N-terminal amino acid residue) of the splicing product PG-AAs-X-COOH (X represents a random amino acid residue). In the figures, black bars show the value when the proteasome reaction time was 5 hours (5 h), and white bars show the value when the proteasome reaction time was overnight (O / N).

[0074] The results shown in Figures 4 to 11 confirm the production of splicing peptides by the proteasome.

[0075] Example 1 In order to carry out the identification method of this example, the following preparations were first made.

[0076] (Preparation of W6 / 32 antibody) (i) Mouse hybridoma W6 / 32 cells (ATCC: American Type Culture Collection) were cultured in RPMI1640 synthetic medium (Sigma-Aldrich) containing 10% by weight of FBS (Fetal Bovine Serum, Gibco). The mouse hybridoma W6 / 32 cells produce the W6 / 32 antibody (IgG2a) that binds to HLA-A, HLA-B, and HLA-C. (ii) When 70% confluence is reached, the cells are washed to remove the FBS and replaced with serum-free hybridoma medium (Hybrigro, Corning). (iii) Continue culturing for 72 hours. (iv) The medium is collected, filtered and clarified. (v) The W6 / 32 antibody contained in the medium is purified by adsorption onto a Protein A Sepharose (GE) column. After washing the column, the W6 / 32 antibody is quickly eluted with 100 mM glycine-HCl buffer (pH 2.7). (vi) The solution from which the W6 / 32 antibody has been eluted is immediately adjusted to neutral pH with 1 M Tris-HCl buffer (pH 9.0) and then replaced with 200 mM Na phosphate buffer (pH 7.0). (vii) The protein concentration in the resulting solution is measured by bicinchoninic acid assay, and the solution containing the W6 / 32 antibody is stored at 4°C.

[0077] (Biological sample preparation) (i) 2×10 8 Human epidermoid carcinoma-derived A431 cells (ATCC) were suspended in 100 mM Tris-HCl buffer (pH 8.5) containing 2% by weight of OTG (octyl-D-1-thioglucopyranoside) and a protease inhibitor cocktail (Sigma-Aldrich), and the suspension was incubated on ice for 30 minutes to lyse the cells. (ii) The cell lysate is centrifuged (20,000 g, 30 min) and the supernatant is collected. This supernatant is believed to contain HLA (HLA-A, HLA-B, and HLA-C) bound to peptides specific to A431 cells. (iii) 200 μg of W6 / 32 antibody is added to the obtained supernatant and incubated at 4°C for 16 hours to form an immune complex (antigen peptide-HLA-W6 / 32 complex) between HLA (HLA-A, HLA-B, and HLA-C) bound to a peptide specific to A431 cells and the W6 / 32 antibody. (iv) The antigen peptide-HLA-W6 / 32 complex is recovered using "TOYOPEARL AF-rProteinA" (TOSOH), a resin onto which Protein A, which adsorbs IgG, is immobilized. (v) The Protein A-immobilized resin is washed five times with 1 mL of PBS (phosphate-buffered saline) and five times with 1 mL of Tris buffer. (vi) After washing, 500 μL of 10% acetic acid is added to the Protein A-immobilized resin and incubated at room temperature for 10 minutes. The supernatant containing the antigen peptide-HLA-W6 / 32 complex is then collected and completely dried in a centrifugal dryer to obtain a solid containing the antigen peptide-HLA-W6 / 32 complex. (vii) The resulting solid is redissolved in 0.1% formic acid in water. The solution containing the antigen peptide-HLA-W6 / 32 complex thus obtained is used as a sample for mass spectrometry (biological sample).

[0078] [Identification of splicing peptides] Next, the identification method of this example, that is, a method for identifying unknown splicing peptides (unknown antigen peptides: neoantigens) derived from known parent peptides contained in the above biological samples, will be described.

[0079] (1) Mass spectrometry (first step in the above embodiment) As mass spectrometry data for the peptides contained in the above samples, MS / MS spectra were measured by LC-MS / MS using a liquid chromatograph mass spectrometer system (Nexera-Mikros: Shimadzu Corporation) and a Q-TOF mass spectrometer (LCMS-9030: Shimadzu Corporation).

[0080] The LC-MS / MS analysis conditions are as follows: Solvent A: A solvent containing 0.1% by mass of FA (formic acid) and 5% by mass of ACN (acetonitrile), with the remainder being water. Solvent B: A solvent containing 0.1% by mass of FA (formic acid) and 80% by mass of ACN (acetonitrile), with the remainder being water. Trap column: "L-column2 ODS" (Chemicals Evaluation and Research Institute, Japan), 5 μm, 0.3 × 5 mm Separation column: "L-column2 ODS", 2 μm, 0.3 × 150 mm Flow rate: 5μL / min Interface: ESI micro interface Number of events: 8-14 Event duration: 100msec Collision voltage: 25±10V Collision gas pressure: 230kPa Interface voltage: 3kV Interface temperature: 100℃ DL temperature: 200℃ Heat block temperature: 250℃ Scan range: 400-700Da Undetermined ions: Excluded Parent ion resolution: 20 ppm Exclusion time: 5 seconds Nebulizer gas: 1L / min Heating gas: 3L / min Dry Gas: 0 Data conversion: mzML format

[0081] (2) Identification of parent peptides contained in the sample (second step in the above embodiment) Using the mzML-formatted MS / MS spectra, we performed standard peptidome analysis to identify HLA Class I-binding peptides (parent peptides), known as cancer antigen peptides, from among the peptides contained in the sample. For peptidome analysis, we used "Mascot Proteome Server version 2.6.2" (Matrix Science) and "Peaks Studio software version 10.0" (Bioinformatics Solutions Inc.) as the identification engine, and "SwissProt Human protein sequence version 2019.9" as the public protein database (amino acid sequence database). A search was performed using the SwissProt database (primary library), and peptides with significant peptide scores (P<0.05) were identified as parent peptides. The SwissProt database is a public protein sequence database, not a peptide-ligand database.

[0082] (3) Creating a secondary library (the third step in the above embodiment) The list of parent peptides identified as described above was exported. Randomized sequences (candidates for splicing peptides) were generated for the parent peptides (631 types), and a secondary library consisting of these candidates was created.

[0083] (4) Identification of spliced ​​peptides (fourth step in the above embodiment) Using the created secondary library, the above mass spectrometry data (MS / MS spectra) was reanalyzed (to identify splicing peptides contained in the sample).

[0084] The peptide identification results in this step are shown in Table 2. As shown in the left column of Table 2, the candidate splicing peptides constituting the secondary library are divided into those that can be generated by addition, deletion, deletion and addition, or substitution of the parent peptide. In the table, the "Identified peptides" column indicates the total number of identified peptides, and the "Identified spliced ​​peptides" column indicates the number of identified peptides that were identified as spliced ​​peptides (amino acid sequences that were not in the primary library and were added to the secondary library) (the same applies to Table 3 described below).

[0085] [Table 2]

[0086] <Reference example 2> Using the same mass spectrometry data as in Example 1, a feasibility test was performed on the method for identifying splicing peptides. First, we analyzed the splicing frequency using a primary library consisting of the sequences of ligands (1,231,612 entries) accumulated in the IEDB (Immune Epitope Database) and the sequences of ligands (64,201 entries) accumulated in HLAtlas, which are known amino acid sequence databases. Specifically, for all of these sequences, randomized sequences (candidate splicing peptides) were created using each of the substitutions or additions shown in the second and subsequent lines of Table 3, and these randomized sequences were added to the original FASTA file to create a reorganized amino acid sequence database (secondary library). Using this secondary library, amino acid sequences matching the mass spectrometry data of each peptide contained in the proteasome reaction mixture were identified from the secondary library. The results are shown in Table 3. For feasibility studies, the results of "SwissProt protein seq" are also shown.

[0087] [Table 3]

[0088] The results shown in Table 3 indicate that splicing peptides can be identified with an average frequency of about 5%. However, when randomizing sequences to create the secondary library, a group of candidate splicing peptides that could arise from all sequences in the primary library was created, which required a huge amount of work (approximately 10 to 30 days) and data volume (approximately 3 to 10 TB). Therefore, the above method using the IEDB / HLAtlas database is considered difficult to put into practical use with ordinary machine power.

[0089] In contrast, in Example 1, the primary library was used to first identify known peptides (parent peptides) contained in the sample, and a list (candidate group) of splicing peptides that could be generated from only the identified parent peptides was added to create a secondary library. As a result, in Example 1, the database size of the secondary library could be compressed to approximately 300 MB, and the work time could be significantly reduced (by approximately 30 minutes). Furthermore, since the work of creating the secondary library is thought to be automatable, further reductions in processing time are expected. At this level, it is believed that practical application is feasible.

[0090] Example 2: HLA affinity prediction analysis Although the identification of spliced ​​peptides in Example 1 indicated that the mass spectrometry data may have matched the structure of hypothetical spliced ​​peptides derived from HLA-binding peptides, the affinity of the identified spliced ​​peptides for HLA remains unclear. Therefore, in this example, the list of spliced ​​peptides identified in Example 1 was subjected to the NetMHCpan algorithm to calculate predicted affinities for HLA alleles (HLA-A03:01, HLA-B07:02, HLA-C07:02) expressed by A431 cells. From the calculation results, the ranking of each splicing peptide identified in Example 1 was scored, and when the obtained "%Rank score" was less than 0.5, it was evaluated as having strong binding ("S" in the table), and when the "%Rank score" was 0.5 or more but less than 2, it was evaluated as having weak binding ("W" in the table). Table 4 shows a list of the evaluation results (total number of peptides evaluated as S or W).

[0091] [Table 4]

[0092] From the S (strong binding) evaluation results shown in Table 4, approximately 140 splicing peptides were identified as peptides that may bind strongly to HLA (HLA-A03:01, HLA-B07:02, HLA-C07:02) expressed by A431 cells.

[0093] In this way, by obtaining analytical data on affinity with HLA, it is thought that newly identified splicing peptides (neoantigens) can be applied as peptide vaccines for cancer vaccine therapy. Furthermore, by obtaining the affinity data for each HLA allele as described above, it is believed that it is possible to select splicing peptides (neoantigens) with high affinity for the expressed HLA alleles for each HLA type of each individual for the splicing peptides identified by the identification method described in the above embodiment. As a result, for example, when selecting new peptide vaccine candidates to be used in cancer vaccine therapy from splicing peptides, it is believed that it is possible to select the optimal peptide vaccine for each patient based on information on splicing peptides identified from cancer cells, etc., and the patient's HLA type.

[0094] [Aspect] It will be appreciated by those skilled in the art that the exemplary embodiments and examples described above are examples of the following aspects.

[0095] (Section 1) A method for identifying a splicing peptide according to one embodiment is a method for identifying an unknown splicing peptide derived from a known parent peptide contained in a biological sample, comprising the steps of: (1) a first step of performing mass spectrometry on the sample to obtain mass spectrometry data of peptides contained in the sample; (2) a second step of searching a primary library, which is a database of known amino acid sequences, for amino acid sequences that match the mass spectrometry data to identify the parent peptide contained in the sample; (3) a third step of generating a secondary library containing candidate spliced ​​peptides that can be generated from the identified parent peptides; (4) a fourth step of searching the secondary library for an amino acid sequence that matches the mass spectrometry data to identify the spliced ​​peptide contained in the sample.

[0096] According to the identification method described in paragraph 1, by using a secondary library containing candidate splicing peptides that can be generated from a parent peptide, unknown splicing peptides can be identified simply and efficiently without performing proteasome degradation experiments or analysis of the large number of unidentified MS / MS spectra that are generated.

[0097] (Section 2) 2. The method according to claim 1, wherein the candidate group consists of spliced ​​peptides that are likely to arise from the parent peptide.

[0098] According to the identification method described in paragraph 2, the time required for computer processing and other steps related to identification can be reduced, and identification can be performed more efficiently, compared to when all combinations of variant peptides that can be generated by splicing from a parent peptide are added as candidates.

[0099] (Section 3) 3. The method according to claim 1 or 2, wherein the candidate group consists of variant peptides generated by at least one of deleting a predetermined number of amino acid residues or less from the C-terminus and / or the N-terminus of the parent peptide and adding a predetermined number of amino acid residues or less to the C-terminus and / or the N-terminus.

[0100] According to the identification method described in paragraph 3, the candidate group consists of splicing peptides that are likely to be generated from the parent peptide. Therefore, as in paragraph 2, the time required for computer processing and other tasks related to identification can be shortened, and identification can be performed more efficiently, compared to adding all combinations of variant peptides that can be generated from the parent peptide by splicing as a candidate group.

[0101] (Section 4) 4. The method of claim 3, wherein the predetermined number is two residues.

[0102] According to the identification method described in Section 4, the candidate group will consist of splicing peptides that are more likely to arise from the parent peptide, which further reduces the time required for computer processing and other tasks related to identification, allowing for efficient identification.

[0103] (Section 5) the mass spectrometry data is a mass spectrometry spectrum, 5. The method according to any one of items 1 to 4, wherein the amino acid sequence that matches the mass spectrometry data is an amino acid sequence that corresponds to a mass spectrometry spectrum having a peak at the same m / z value or at an m / z value whose deviation is within a predetermined range with respect to the m / z value of at least one peak in the mass spectrometry spectrum.

[0104] (Section 6) 6. The method according to any one of items 1 to 5, wherein the parent peptide is a known HLA-binding peptide.

[0105] (Section 7) 7. The method according to any one of items 1 to 6, wherein the parent peptide is a known HLA-binding peptide.

[0106] (Section 8) 8. The method of claim 7, wherein the parent peptide is a known HLA-binding peptide derived from a cancer cell.

[0107] A program for executing the above-mentioned identification method, and a medium (non-transitory computer-readable medium) storing the program.

[0108] The embodiments and examples disclosed herein should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined by the claims, not by the above description, and is intended to include all modifications within the meaning and scope of the claims.

Claims

1. 1. A method for identifying unknown spliced ​​peptides derived from known parent peptides contained in a biological sample, comprising: (1) a first step of performing mass spectrometry on the sample to obtain mass spectrometry data of peptides contained in the sample; (2) a second step of searching a primary library, which is a database of known amino acid sequences, for amino acid sequences that match the mass spectrometry data to identify the parent peptide contained in the sample; (3) a third step of generating a secondary library containing candidate splicing peptides that can be generated from the identified parent peptides; (4) a fourth step of searching the secondary library for an amino acid sequence that matches the mass spectrometry data to identify the spliced ​​peptide contained in the sample; the parent peptide is a known HLA-binding peptide; The candidate group consists of variant peptides generated by at least one of deleting two or less amino acid residues from at least either the C-terminus or the N-terminus of the parent peptide and adding two or less amino acid residues to the C-terminus or the N-terminus.

2. the mass spectrometry data is a mass spectrometry spectrum, The method of claim 1, wherein the amino acid sequence that matches the mass spectrometry data is an amino acid sequence that corresponds to a mass spectrometry spectrum having a peak at the same m / z value or at an m / z value within a predetermined range of deviation from the m / z value of at least one peak in the mass spectrometry spectrum.

3. The method of claim 1 , wherein the parent peptide is a known HLA-binding peptide derived from a cancer cell.

Citation Information

Patent Citations

  • Method for indentification of proteasome generated spliced peptides

    EP2362225A1