Method for identifying or inferring glycopeptides
By using liquid chromatography-mass spectrometry and a reference glycopeptide list for retention time and mass correction, the method addresses the limitations of conventional glycopeptide detection, enabling sensitive and reproducible identification of glycopeptides.
Patent Information
- Application Number
- PCT/JP2025/024159
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-12
- Filing Date
- 2025-07-04
- Publication Date
- 2026-01-15
AI Technical Summary
Conventional methods for identifying glycopeptides are limited by their inability to detect low-abundance peptides and glycosylation effectively, leading to difficulties in comprehensive glycopeptide analysis.
A method involving liquid chromatography-mass spectrometry analysis to obtain retention time and mass information for each peak, comparing it with a reference glycopeptide list, and correcting the data to identify or predict glycopeptides based on matching retention time and mass information.
This approach allows for the simple and sensitive identification of glycopeptides, particularly in samples derived from multiple subjects, enhancing detection sensitivity and reproducibility.
Smart Images

Figure JP2025024159_15012026_PF_FP_ABST
Abstract
Description
Method for identifying or predicting glycopeptides
[0001] The present invention relates to a method for identifying or predicting glycopeptides, and the like.
[0002] Genomes (nucleic acids), proteins, and glycans are known as the three chains of life, but because glycan structures are diverse and complex, their analysis is extremely difficult, and the amount of information accumulated so far about glycans is overwhelmingly less than that about genomes and proteins. The Human Glycome Project (HGA) will build a comprehensive database of all glycans present in humans and catalogue glycans that change during processes such as aging and disease.
[0003] In addition to glycoprotein glycans, which are addressed in glycoproteomics, there are various other categories of glycans, such as glycolipids, proteoglycans, and even free glycans. However, glycoproteomics is the most important glycan target that directly links genomic information to the extended central dogma. Glycoproteomics utilizes the latest precision techniques (such as liquid chromatography-mass spectrometry (LC-MS)) to perform large-scale, comprehensive analysis of what glycan structure is attached to what protein and at what position.
[0004] Conventional methods (e.g., Non-Patent Document 1) use MS / MS (tandem MS) to separate glycopeptides by LC. The ions generated in the first ionization step are used as precursor ions, and then collision-induced dissociation is performed in the second step to separate small ions (product ions) as fragment ions, thereby identifying the amino acid sequence and glycosylation of the peptide. However, conventional methods detect peptides derived from proteins with higher abundance first, which poses problems in the detection and reproducibility of low-expression proteins. Conventional methods are extremely difficult to detect peptides with relatively low abundance or low ionization efficiency, especially glycosylation, and are therefore not suitable for the purpose of comprehensive glycopeptide analysis.
[0005] Kawahara, R. et al. Community evaluation of glycoproteomics informatics solutions reveals high-performance search strategies for serum glycopeptide analysis. Nat. Methods 18, 1304-1316 (2021).
[0006] An objective of the present invention is to provide a technique for identifying glycopeptides in a glycopeptide sample more simply and with higher sensitivity.
[0007] In light of the above-mentioned problems, the present inventors have conducted extensive research and have found that the above-mentioned problems can be solved by a method for identifying or predicting glycopeptides, comprising the steps of: (1) subjecting a glycopeptide sample prepared from a biological specimen collected from a subject to liquid chromatography-mass spectrometry analysis to obtain retention time information Xa and mass information Xb for each peak; (2) preparing a reference glycopeptide list including retention time information Ya and mass information Yb obtained by liquid chromatography-mass spectrometry analysis of a plurality of previously identified glycopeptides; and (3) for each peak, comparing the retention time information Xa with the retention time information Ya after correcting at least one of them, and comparing the mass information Xb with the mass information Yb, and identifying or predicting the peak as a glycopeptide if both are consistent. Based on this finding, the present inventors have conducted further research and have completed the present invention. Specifically, the present invention encompasses the following aspects.
[0008] Item 1. A method for identifying or predicting a glycopeptide, comprising: (1) subjecting a glycopeptide sample prepared from a biological sample collected from a subject to liquid chromatography-mass spectrometry to obtain retention time information Xa and mass information Xb for each peak; (2) preparing a reference glycopeptide list including retention time information Ya and mass information Yb in liquid chromatography-mass spectrometry of a plurality of glycopeptides previously identified; and (3) comparing, for each peak, the retention time information Xa and the retention time information Ya after correcting at least one of them, and comparing the mass information Xb and the mass information Yb, and identifying or predicting the peak as a glycopeptide if both of them match.
[0009] Item 2. The identification method according to Item 1, wherein the reference glycopeptide list contains 500 or more glycopeptides.
[0010] Item 3. The identification or estimation method according to Item 1 or 2, wherein the reference glycopeptide list includes a glycopeptide list obtained by liquid chromatography-mass spectrometry of a glycopeptide sample prepared from the same biological sample as the biological sample in step (1).
[0011] Item 4. The identification or estimation method according to any one of Items 1 to 3, wherein step (1) is performed for each of a plurality of the test subjects.
[0012] Item 5. The identification or estimation method according to Item 4, wherein the number of subjects is 10 or more.
[0013] Item 6. The identification or estimation method according to any one of Items 1 to 5, wherein the correction in step (3) is performed based on the retention time of a reference internal standard.
[0014] Item 7. A reagent for use as a reference internal standard in the identification method according to Item 6, comprising a plurality of types of reference internal standard substances having different retention times in liquid chromatography-mass spectrometry.
[0015] Item 8. The reagent according to Item 7, wherein the reference internal standard is a glycopeptide comprising a base glycopeptide and a hydrocarbon group added to the base glycopeptide.
[0016] Item 9. The reagent according to Item 8, wherein the molecular weights of the hydrocarbon groups differ among the multiple types of glycopeptides.
[0017] According to the present invention, a technique for identifying glycopeptides in a glycopeptide sample more simply and with higher sensitivity can be provided, which is particularly useful when identifying or predicting glycopeptides in each of glycopeptide samples derived from multiple subjects.
[0018] Schematic diagram of retention time correction. The experimental time of the RT marker is plotted on the x-axis and the RT index value to be corrected on the y-axis, and the slope between each point is calculated. As indicated by the arrows, signals between marker points (e.g., S1a and S1b, 14 minutes and 31 minutes) are corrected by the slope of the two markers surrounding that point (both 24 minutes). This process converts the retention times of two experiments with different LC analysis times into RT index values. A. The retention times of the RT markers in the refined map (x-axis) and the RT markers in the measured data (y-axis) are plotted. The refined map data has a longer gradient in LC, and therefore the correlation between the two has a larger slope. The correlation value between the two is 0.99. The slope values between each point were 0.326, 0.338, 0.343, 0.346, 0.369, 0.384, 0.416, 0.494, and 0.483, with slight differences in slope between the first and second halves. B. The retention times of all glycopeptide signals listed in the refined map were plotted before (x-axis) and after (y-axis) correction. The correction process consolidated the retention times, which had been distributed over a wide range (40-100 minutes), into a range of 15-40 minutes. A schematic diagram of the alkylamidation reaction of sialylglycopeptides is shown. The results of confirming the retention time of alkylamidated SGP by reversed-phase nano-LC / MS are shown. Chromatograms of plasma-derived glycopeptides and alkylamidated SGP by nano-LC / MS are shown. (A) Total ion chromatogram (TIC, m / z 500-1800) obtained by measuring glycopeptides prepared from albumin- and immunoglobulin G-depleted plasma. (B) Total ion chromatogram (TIC, m / z 500-1800) obtained by measuring a sample in which alkylamidated SGP was added to the glycopeptide in (A). (C) Extracted ion chromatogram (EIC) of the alkylamidated SGP signal in a sample in which alkylamidated SGP was added to plasma glycopeptides. The index value and retention time of alkylamidated SGP are shown. (A) 40-minute gradient conditions. (B) 120-minute gradient conditions. The index value of alkylamidated SGP was determined based on the carbon number of the alkylamine used for alkylamidation.
[0019] In this specification, the expressions "contain" and "comprise" include the concepts of "contain," "comprise," "consist essentially of," and "consist only of."
[0020] 1. Glycopeptide Identification Method In one aspect, the present invention relates to a method for identifying or predicting a glycopeptide (sometimes referred to herein as the "identification method of the present invention"), comprising the steps of: (1) subjecting a glycopeptide sample prepared from a biological sample collected from a subject to liquid chromatography-mass spectrometry analysis to obtain retention time information Xa and mass information Xb for each peak; (2) preparing a reference glycopeptide list including retention time information Ya and mass information Yb obtained by liquid chromatography-mass spectrometry analysis of a plurality of previously identified glycopeptides; and (3) for each peak, comparing the retention time information Xa with the retention time information Ya after correcting at least one of them, and comparing the mass information Xb with the mass information Yb, and identifying or predicting the peak as a glycopeptide if both are consistent. This method is described below.
[0021] The biological species of the subject and the biological species from which the protein is derived include various mammals such as humans, monkeys, mice, rats, dogs, cats, and rabbits, with humans being preferred.
[0022] The biological sample is not particularly limited as long as it is a sample collected from a living organism and contains proteins, and examples thereof include body fluids, skin, mucous membranes, and internal tissues. Among these, from the viewpoints of ease of collection and minimal invasiveness, body fluids, skin, and mucous membranes are preferred, and body fluids are more preferred. Examples of body fluids include blood, follicular fluid, menstrual blood, saliva, cerebrospinal fluid, synovial fluid, urine, interstitial fluid, sweat, and tears. Examples of mucous membranes include oral mucosa and nasal mucosa. Biological samples are generally subjected to appropriate pretreatment depending on the sample, and treatments to remove some components and to concentrate and / or purify proteins are preferred.
[0023] In one aspect of the present invention, proteins contained in a blood sample or a sample obtained by removing some components contained in blood (e.g., albumin, antibodies such as IgG, etc.) can be used as the biological sample.
[0024] Glycopeptide samples can be obtained by treating proteins in biological samples with proteases. The proteases are not particularly limited as long as they are capable of fragmenting proteins, and examples include endoproteases such as trypsin, Asp-N, Lys-C, and Glu-C.
[0025] From the viewpoints of protease hydrolysis efficiency, suitability for mass spectrometry, etc., it is preferable that proteins in a biological sample have been subjected to denaturation, reduction, alkylation, or other treatments. Denaturation is a treatment that unfolds a protein and is not particularly limited thereto, and is preferably carried out using a protein denaturant such as a surfactant. Reduction is a treatment that cleaves disulfide bonds within a protein and / or between proteins and is not particularly limited thereto, and is preferably carried out using a reducing agent. Alkylation is a treatment that irreversibly alkylates cysteine residues in a protein and is not particularly limited thereto, and is preferably carried out using an alkylating agent.
[0026] The protease treatment is usually carried out in a liquid, but the protease treatment can also be carried out by supporting the protein on a carrier.
[0027] In step (1), a glycopeptide sample is subjected to liquid chromatography-mass spectrometry.
[0028] The liquid chromatography conditions are not particularly limited as long as they are applicable to liquid chromatography-mass spectrometry of glycopeptide samples. Liquid chromatography can be performed using a reversed-phase column such as a C18 column and an appropriate organic solvent (e.g., acetonitrile) as the mobile phase.
[0029] The ionization method in mass spectrometry is not particularly limited, but electrospray ionization (ESI) is often used as an ionization method for peptides.
[0030] The method for measuring mass information (e.g., m / z) in mass spectrometry is not particularly limited, and for example, time-of-flight, magnetic deflection, quadrupole, ion trap, Fourier transform ion cyclotron resonance, etc. can be used.
[0031] By liquid chromatography-mass spectrometry, retention time information (retention time information Xa) and mass information Xb can be obtained for each peak in liquid chromatography.
[0032] In this specification, the retention time information of a peak is not particularly limited as long as it is information that directly or indirectly represents the retention time of the peak. The retention time information of a peak can be, for example, the retention time of the peak top, the retention time of the peak start, the retention time of the peak end, the retention time of any point in the peak, the retention time width between any two points in the peak, an arithmetic processed value of the retention times of multiple points in the peak (for example, an average value, etc.), etc.
[0033] In this specification, the mass information of a peak is not particularly limited as long as it is information that directly or indirectly represents the mass of the peak. The mass information of a peak can be, for example, molecular weight, m / z (and valence), etc.
[0034] A preferred embodiment of step (1) is to perform step (1) for each of a plurality of analytes. According to the identification method of the present invention, glycopeptides can be identified easily and with high sensitivity by simply obtaining retention time information and mass information for each peak of a glycopeptide sample derived from the analyte. Therefore, the greater the number of analytes, the greater the benefits of the identification method of the present invention. From this perspective, the number of analytes is, for example, 10 or more, preferably 100 or more, and more preferably 1,000 or more. The upper limit of the number is not particularly limited, and is, for example, 10,000,000 or less, 1,000,000 or less, 100,000 or less, or 10,000 or less.
[0035] In step (2), a reference glycopeptide list is prepared, which includes retention time information Ya and mass information Yb of a plurality of previously identified glycopeptides in liquid chromatography-mass spectrometry.
[0036] The reference glycopeptide list can be obtained by subjecting a glycopeptide sample to liquid chromatography-mass spectrometry.
[0037] The reference glycopeptide list preferably includes a glycopeptide list obtained by liquid chromatography-mass spectrometry of a glycopeptide sample prepared from the same biological sample as the biological sample in step (1). For example, when a blood sample is used in step (1), the glycopeptide sample used to generate the reference glycopeptide list is preferably also prepared from the blood sample. This can further increase the sensitivity of the identification method of the present invention (e.g., further increase the number of glycopeptides that can be identified).
[0038] To prepare a reference glycopeptide list, various methods can be used to identify glycopeptides, including determining the amino acid and glycan sequences of glycopeptides from both precursor and fragment ion information obtained by tandem MS, or comprehensively obtaining glycopeptide core peptides in advance and then using this information to identify groups (clusters) of glycopeptides that share a common core peptide (Glyco-RIDGE: Glycan heterogeneity-based Relational IDentification of Glycopeptide signals on Elution profile). To further enhance the depth of the reference glycopeptide list, various methods can be used, such as excluding substances present in large amounts in the glycopeptide sample (e.g., proteins present in large amounts such as albumin) beforehand during glycopeptide sample preparation, or fractionating the glycopeptide sample itself or the sample used to prepare the glycopeptide sample by chromatography (e.g., hydrophilic interaction chromatography (HILIC)) and then performing liquid chromatography-mass spectrometry on each of the obtained fractions.
[0039] Increasing the depth of the reference glycopeptide list can further enhance the sensitivity of the identification method of the present invention, which is characterized by using the reference glycopeptide list for identification. Once the reference glycopeptide list is created, it can be used repeatedly, thereby improving the ease of use when applying the identification method of the present invention to a large number of subjects. From these perspectives, the number of glycopeptides included in the reference glycopeptide list is preferably as large as possible, for example, 100 or more, preferably 200 or more, more preferably 500 or more, even more preferably 1000 or more, even more preferably 2000 or more, and particularly preferably 3000 or more. The upper limit of this number is not particularly limited, and may be, for example, 100,000, 50,000, 20,000, or 10,000.
[0040] In step (3), the retention time information Xa and mass information Xb of each peak in step (1) are compared with the retention time information Ya and mass information Yb of the reference glycopeptide list, and if the retention time information Xa of a peak matches the retention time information Ya of a glycopeptide and the mass information Xb of the peak matches the mass information Yb of the glycopeptide, the peak can be identified as the glycopeptide.
[0041] In the identification method of the present invention, at least one of the retention time information Xa and the retention time information Ya is corrected before comparison. This allows for a more sensitive comparison of retention time information that may vary depending on various conditions. Specifically, for example, the retention time information Xa may be corrected to approximate the retention time information Ya, the retention time information Ya may be corrected to approximate the retention time information Xa, or both the retention time information Xa and the retention time information Ya may be corrected to approximate another retention time information Z (information obtained by another liquid chromatography-mass spectrometry).
[0042] The correction in step (3) is preferably performed based on a reference internal standard. Specifically, for example, a calibration curve can be drawn by plotting the retention times of the reference internal standard, with retention time information Xa of the reference internal standard on the vertical axis and retention time information Ya of the reference internal standard on the horizontal axis, and then correcting the retention time information Xa to approximate the retention time information Ya based on the calibration curve, or correcting the retention time information Ya to approximate the retention time information Xa. The type of reference internal standard is not particularly limited, and examples include glycopeptides, peptides, and other low-molecular-weight compounds that are stably observed.
[0043] The calibration curve is not necessarily linear, and therefore, from the viewpoint of increasing the accuracy of the correction, it is preferable to use a larger number of reference internal standards for correction. From this viewpoint, the number of reference internal standards is preferably 3 or more, more preferably 5 or more, and even more preferably 7 or more. From the viewpoint of efficiency, this number is preferably 100 or less, more preferably 50 or less, even more preferably 20 or less, and even more preferably 15 or less. It is preferable that the retention times of the reference internal standards are separated by a certain degree. From this viewpoint, the retention time interval between the reference internal standards is, for example, about 1 to 15 minutes.
[0044] The "match" in step (3) does not require a strict match. Certain criteria can be set, and if the data falls within that range, it can be determined that there is a "match." Furthermore, the information being compared does not have to be the same. For example, if retention time information Xa is the retention time range between any two points on a peak, and retention time information Ya is the retention time of the peak top, then the two can be determined to "match" if the retention time information Ya falls within the range of retention time information Xa.
[0045] In step (3), if it is determined that the retention time information Xa and mass information Xb of a peak match only the retention time information Ya and mass information Yb of one glycopeptide, the peak can be "identified" as that glycopeptide. On the other hand, if it is determined in step (3) that the retention time information Xa and mass information Xb of a peak match the retention time information Ya and mass information Yb of multiple glycopeptides, the peak can be "presumed" to be one of those glycopeptides. However, in either case, the peak can still be identified as some kind of glycopeptide.
[0046] 2. Reagent In one aspect, the present invention relates to a reagent for use as a reference internal standard in the identification method of the present invention, which comprises multiple types of reference internal standard substances having different retention times in liquid chromatography-mass spectrometry.
[0047] The type of reference internal standard, the number of reference internal standards, and the retention time intervals between the reference internal standards are as described above.
[0048] A preferred reference internal standard substance used as a reagent is a glycopeptide containing a base glycopeptide and a hydrocarbon group added to the base glycopeptide. By varying the molecular weight of the hydrocarbon group, multiple types of internal standard glycopeptides with different retention times can be obtained.
[0049] The base glycopeptide is not particularly limited as long as it has functional groups to which a hydrocarbon group can be added (preferably a plurality of functional groups, for example, 2 to 5, 2 to 4, or 3 to 4).
[0050] The functional group is not particularly limited, and examples thereof include a carboxy group, a hydroxy group, an amino group, an acetyl group, and a sulfate group.
[0051] As the base glycopeptide, a sialylglycopeptide can be preferably used.
[0052] The hydrocarbon group is not particularly limited, and examples thereof include alkyl groups, alkenyl groups, alkynyl groups, aryl groups, linking groups thereof, etc. The number of carbon atoms in the hydrocarbon group is, for example, 1 to 50, 1 to 40, 1 to 30, or 1 to 20.
[0053] The addition of a hydrocarbon group to a base glycopeptide can be carried out according to or in accordance with a known method. The desired internal standard peptide can be obtained by reacting a compound in which an appropriate functional group is linked to a hydrocarbon group with the base glycopeptide.
[0054] The above-mentioned reagents can be in the form of a kit, which can include other reagents, instruments, and documentation explaining how to use the kit, which can be used in the identification method of the present invention.
[0055] The present invention will be described in detail below based on examples, but the present invention is not limited to these examples.
[0056] Test Example 1. Preparation of a Reference Glycopeptide List A reference glycopeptide list (precise glycopeptide map, precision map) was prepared. The outline of the preparation method is as follows.
[0057] Preparation of glycopeptides derived from human blood glycoproteins for reference mapping Commercially available human serum and plasma were applied to a major protein removal column (Agilent MARS column, Human 7 and Human 14) or a Proteominer protein enrichment spin column (Bio-Rad ProteoMiner). The flow-through fraction for the MARS column and the bound fraction for the Proteominer were collected according to the manufacturer's recommendations. The following experiments were performed on four protein fractions, including the unfractionated fraction (whole sample).
[0058] Blood-derived proteins were subjected to reduction of disulfide bonds (dithiothreitol) and alkylation of the resulting thiol groups (iodoacetamide) in a buffer containing phase transfer surfactant (PTS), then desalted onto SP3 beads (Single-Pot Solid-Phase-enhanced Sample Preparation: Cytiva Sera-Mag magnetic beads (carboxylic acid modified)) and subsequently digested on the beads with Lys-C protease and trypsin.
[0059] A portion of the digest was fractionated by hydrophilic interaction chromatography using an Amide 80 column (TOSOH), and glycopeptides were collected in 1 to 20 fractions.
[0060] <Identification of glycopeptides by LC / MS and construction of a reference map> Glycopeptides were analyzed using an LC / MS system consisting of a nanoflow HPLC (Thermo Scientific, UltiMate3000 RSLCnano) and a tribrid mass spectrometer (Thermo Scientific, Orbitrap Eclipse). A 0.075 mm inner diameter, 12-25 cm long C18 column was used, and glycopeptides were separated using an acetonitrile gradient in 0.1% formic acid. Glycopeptides ionized by electrospray were introduced into the analyzer and data-dependently selected and fragmented by collision-induced dissociation, respectively, to obtain fragment ion (MS2) spectra. Based on these spectra, glycopeptides were identified using the database search software Byonic (Proteinmetrics).
[0061] Glycopeptide fractions prepared from each of the four protein fractions were analyzed using the above method, and the results were compiled to create a reference map. The map is a list of identified glycopeptides, and for each glycopeptide, its structural information (glycoprotein accession number, protein name, amino acid sequence of the peptide portion, and glycan composition) is linked to its retention time (or a numerical index normalized using a standard substance; see other sections) and accurate mass (accurate to four or more decimal places) calculated from the chemical structure.
[0062] The resulting reference glycopeptide list contains 648 glycopeptides, and for each glycopeptide, the list also includes the peak retention time in liquid chromatography-mass spectrometry, molecular weight information, protein of origin, peptide sequence, and glycan composition.
[0063] Test Example 2. Identification of glycopeptides in glycopeptide samples A glycopeptide sample prepared from a blood sample collected from a subject was subjected to liquid chromatography-mass spectrometry, and the retention time and mass information of each peak was obtained. The retention time and mass information were compared with those of the reference glycopeptide list, and if both sets of information matched, the peak was identified as a glycopeptide. Specifically, this was done as follows.
[0064] Test Example 2-1. Preparation of Glycopeptide Samples Test Example 2-1-1. Removal of Abundant Proteins from Blood Samples 10 μL of plasma was loaded onto a column packed with High Select Top14 Abundant Protein Depletion Resin (Thermo) to remove 14 abundant proteins. The column was tightly capped and placed on a rotating rotator while mixing by end-over-end mixing for 10 minutes. The column was then placed in a 1.5-mL low-protein binding polypropylene tube (ProteinLobind, Eppendorf), the cap loosened, and spun down in a centrifuge at 1000 x g for 2 minutes. The recovered blood sample-derived proteins were then centrifuged for drying.
[0065] Test Example 2-1-2. Denaturation, reduction, and alkylation of proteins. 100 μL of a solution (2% (w / v) sodium dodecyl sulfate (SDS), 10 mM tris(2-carboxyethyl)phosphine hydrochloride (TCEP-HCl), 40 mM chloroacetamide (CAA)) for denaturation, disulfide bond reduction, and alkylation of the sample protein was added. The sample was redissolved by shaking at 1,000 rpm at 25°C for 1 minute in a thermomixer, and then the denaturation, reduction, and alkylation reactions were carried out by heating (95°C, 10 minutes).
[0066] Test Example 2-1-3. Washing and Digestion of Proteins on Magnetic Beads After cooling the sample solution to approximately room temperature, 30 μL of a 55 μg / μL magnetic particle suspension, ReliaPrep Resin (equivalent to 1.5 mg of magnetic particles; manufactured by Promega), was added. The solution was mixed by shaking in a thermomixer at 1,000 rpm and 25°C for 1 min. Subsequently, to aggregate the protein onto the beads, 520 μL of ethanol (final concentration 80% v / v) was added and the solution was shaken in a thermomixer at 1,000 rpm and 25°C for 10 min. The tube containing the sample solution was placed on a magnetic stand and allowed to stand for 2 min. After the magnetic beads were accumulated on the tube wall by the magnet, the supernatant was removed with a pipette. To wash the solution, 800 μL of 80% (v / v) ethanol was added and allowed to stand for 2 min, after which the supernatant was removed. This washing procedure was repeated twice. Next, 50 μL of a 2x concentrated detergent buffer (4.8 mM SDC, 4.8 mM SLS, 200 mM Tris-HCl pH 8.5) was added to the sample for reconstitution and digestion. The sample was resuspended by shaking in a thermomixer at 2,000 rpm at 25°C for 5 min. To this was added 10 μL of 20 mM CaCl2, 38 μL of distilled water, and 2 μL of enzyme solution for protein digestion (0.5 μg / μL Trypsin / Lys-C mix, 2 mM acetic acid (acetic acid comes from the enzyme reconstitution solution)). The digestion reaction was carried out in a thermomixer at 1,000 rpm at 37°C for 16 h.
[0067] Test Example 2-1-4. Purification of glycopeptide First, droplets were dropped down by flash centrifugation (tabletop compact centrifuge).
[0068] The digestion reaction was stopped by first adding 200 μL of glycopeptide loading solution to retain glycopeptides on the beads. Acetonitrile (ACN) containing 1% (v / v) trifluoroacetic acid (TFA) was used as the glycopeptide loading solution. 5 μL of a standard glycopeptide solution (4 μM sialylglycopeptide; Tokyo Chemical Industry Co., Ltd.) was added, followed by the remaining 350 μL of glycopeptide loading solution (1% TFA in ACN). The glycopeptides were captured by shaking in a thermomixer at 1,500 rpm at 25°C for 30 min. The droplets were removed by flash centrifugation, and the tube containing the sample solution was placed on a magnetic stand and left to stand for 2 min before the supernatant was removed.
[0069] To wash the beads, 800 μL of 82.5% acetonitrile (ACN) containing 1% (v / v) trifluoroacetic acid (TFA) was added as a wash solution. The tube was shaken at 1,000 rpm at 25°C for 5 min. The tube was then placed on a magnetic stand and allowed to stand for 2 min before the supernatant was removed. This procedure was repeated twice. To elute the glycopeptides from the beads, 300 μL of Neutral Elution solution (30% (v / v) acetonitrile containing 20 mM ammonium formate) was added and the tube was shaken at 1,500 rpm at 25°C for 10 min in a thermomixer. After 5 min on the magnetic stand, the supernatant was collected, transferred to a new 1.5-mL low-binding polypropylene tube, and dried in a centrifugal dryer.
[0070] Test Example 2-1-5. Reconstitution of Glycopeptides The dried glycopeptide sample was reconstituted in 50 μL of 0.5% (v / v) aqueous acetic acid, transferred to a filter (Millex LH 45 μm, 4 mm, Merck Millipore) attached to a low-adsorption 0.5-mL PP tube, and filtered by centrifugation at 2,000 × g, 25°C, for 2 min. The sample solution was mixed with a vortex to prepare a 1 / 4-diluted plasma glycopeptide solution (0.1 μL plasma / μL).
[0071] <Test Example 2-2. Measurement and analysis by LC-MS / MS> <Test Example 2-2-1. Measurement by LC-MS / MS> For the measurement, a 1 / 4 diluted plasma glycopeptide solution was transferred to a low-adsorption TPX resin HPLC vial (0.3-mL Proteosave-coated, AMR) and set in the autosampler of the high-performance liquid chromatography (HPLC).
[0072] Glycopeptide samples were analyzed using a liquid chromatography-mass spectrometry-based technique using a nanoflow high-performance liquid chromatography (HPLC) system (Vanquish Neo, Thermo Fisher Scientific) and a quadrupole-Orbitrap mass spectrometer (Orbitrap Exploris 240, Thermo Fisher Scientific).
[0073] The mobile phases used for HPLC were ultrapure water containing 0.1% (v / v) formic acid (A) and acetonitrile containing 0.1% (v / v) formic acid (B). A nano-HPLC capillary column (C18, 0.075 × 150 mm, 3 μm; Nikkyo Technos) with an integrated emitter was used. The flow rate during sample separation was 0.30 μL / min, and the column temperature was set at 40°C. The proportion of mobile phase B was fixed at 0% from 0 to 2 min, increased linearly to 1% from 2 to 2.5%, 40% from 2.5 to 35 min, and 95% from 35 to 40 min.
[0074] The mass spectrometer was set with the ion source at a spray voltage of 1800 V and an ion transfer tube temperature of 275°C. 2 (ddMS2) was performed, and the sample injection volume was 5 μL (equivalent to 0.1 μL of plasma).
[0075] Full MS scan was performed with Orbitrap Resolution: 120,000, Scan Range: 500-1800 m / z, RF Lens: 70%, Normalized AGC Target: 300%, Maximum Injection Time: 50 ms, Microscans: 1, and Positive Polarity. Data-dependent MS 2 The filter settings were: MIPS: Peptide, Intensity Threshold: 5.0e3, Charge state: 2-8, Dynamic Exclusion duration: 15 s, Dynamic Exclusion Mass Tolerance: High-10, Low-10 ppm. The data-dependent MS2 (ddMS2) acquisition settings were: Data Dependent mode: Cycle time (1.5 s), Isolation window: 1.6 m / z, Collision Energy Type: Normalized, HCD collision energy: 20, 30, 40 (%), Orbitrap resolution: 15,000, Scan range: 150-2000 m / z, Normalized AGC target: 200%, Maximum injection time: 50 ms.
[0076] <Test Example 2-2-2. Rapid Glycoproteomics Data Output> The raw LC / MS data files obtained were analyzed using the Byonic search engine version 2.6.46 (Protein Metrics), which was integrated as a node into Proteome Discoverer version 2.5 (Thermo Scientific). The search conditions were as follows: precursor ion mass tolerance: 6 ppm; fragment ion mass tolerance: 20 ppm; static modification: carbamidomethylation (C); dynamic modification: oxidation (M) and deamidation (Q); and the maximum number of missed cleavages per peptide: 2. After analysis was complete, a PSM file was output and used for the following analysis.
[0077] <Test Example 2-3. Retention time correction for comparing analytical data with precision map data> <Test Example 2-3-1. Selection of glycopeptide markers> In this test example, commonly occurring endogenous glycopeptides were selected as markers. The raw files measured by the Date Dependent Acquisition (DDA) method using LC-MS / MS in Test Example 2-2 were analyzed using Proteome discoverer-Byonic, and a list of the molecular weights, retention times, peptide sequences, and glycosylation modifications of the glycopeptides contained in the blood samples was created. This list was compared with the glycopeptides contained in the precision map obtained in Test Example 1, and glycopeptides were selected as markers from the commonly detected glycopeptides so that they were scattered throughout the entire measurement time.
[0078] A list of the selected glycopeptides is shown in Table 1. The molecular weight (Da), protein name, peptide sequence, glycan modification, and retention time of the glycopeptide used for retention time correction are shown. The retention time indicated is the time recorded on the refined map. The measured retention time was determined by analyzing the raw file of the measured sample using Proteome discoverer-Byonic, and the ApexRT (min) of the peptide information identified based on the MS / MS data was used.
[0079]
[0080] <Test Example 2-3-2. Correction of Retention Time> The retention time in the analytical data of the blood sample was used as the reference (index) for the corrected value. The slope between two adjacent points was calculated, and based on this slope, the retention time in the refined map was adjusted to the actual retention time measured by LC-MS / MS to calculate the corrected retention time. The correction was performed using the following formula (Figure 1). The retention times of the marker glycopeptides in the refined map and the retention times of the measured marker glycopeptide signals were plotted on a graph (Figure 2A).
[0081]
[0082] The actual analytical values were obtained by analyzing the raw files measured by LC-MS / MS using 2D-ICAL (MKI, Japan) and extracting peak information (m / z, charge, and retention time information). In this analysis, the retention time of the analytical data was used directly as the RT index.
[0083] Among the information in the precise map, the retention time was calculated using Excel according to the above formula, and the corrected retention time was calculated.
[0084] The results of correcting the retention times in the refined map are shown in Figure 2B. The retention times of all glycopeptide signals are plotted as before (x-axis) / after (y-axis). The correction process consolidated the retention times, which were distributed over a wide range (40-100 min), into a range of 15-40 min.
[0085] <Test Example 2-4. Identification of glycopeptides> First, MS signals were picked up using MS analysis software 2D-ICAL (MKI, Japan). The raw file obtained by analyzing the blood sample was loaded into 2D-ICAL, and the molecular weight (m / z), retention time, and charge were listed. Analysis using 2D-ICAL was performed according to the default settings.
[0086] Next, the measurement data was compared with the precision map created by the reference identification software Glyco2D. The precision map data was rearranged so that it could be read by Glyco2D (MKI, Japan). Glyco2D is software that reads information such as molecular weight, retention time, peptide sequence, protein information, and glycosylation listed in an Excel file, and compares it with the molecular weight (m / z), retention time, and charge list created by 2D-ICAL.
[0087] The Excel file conforms to the output format of PSM data (including modified glycan information) created using the Byonic node in Proteome-discoverer. In this example, for comparison with the precise map information, the peptide sequence (annotated sequence), modification information (modification), protein accession number (master protein accessions), and retention time (RT [min]) were added to the designated rows and saved as an Excel file.
[0088] The 2D-ICAL analysis file and the Excel file containing the precise map information were loaded into Glyco2D, and reference identification was performed.
[0089] A summary of the mutual information used for reference identification is shown in Table 2. The peptide sequences in Table 2 are shown as SEQ ID NOs: 7 to 9.
[0090]
[0091] First, there is the mutual molecular weight. In the precision map data, the molecular weight is shown as a molecular weight, while in the 2D-ICAL measurement data, it is shown in the form of m / z and charge. In the precision map data, the molecular weight is expressed as a monoisotopic value, while in 2D-ICAL, it is shown as the value of the peak with the maximum signal intensity. Therefore, the values in the table shown do not show a perfect match, but the values are converted during the Glyco2D processing process and compared with each other. Molecular weight error was set to allow for an error of up to 10 ppm. Second, there is the retention time. If the corrected retention time on the precision map falls between the minimum and maximum retention times in the measurement data, it is considered to be a "match." If the above two pieces of data "match," they are considered to be the same, and identification is performed by reference identification.
[0092] The file after reference identification was loaded into 2D-ICAL, the data was exported, and then compared with the results of MS / MS data analysis in Excel to list the newly identified glycopeptides (Tables 3 and 4). In this example, 97 newly identified glycopeptides that had not been identified by conventional MS / MS were observed.
[0093]
[0094]
[0095] Test Example 3: Elution Time Indexing Using Alkylamidated Sialylglycopeptides <Summary> ・Retention time indexing is not possible if the sample does not contain an elution position reference peptide (glycopeptide). ・It is necessary to add a reference peptide that will not be confused with the peptide in the sample. ・A glycopeptide derivative that can be used as a peptide (glycopeptide) elution time reference in a reversed-phase liquid chromatography-mass spectrometry (LC-MS) measurement system is prepared. ・Carboxyl groups on the glycopeptide are alkylamide-derivatized by mixing and reacting a glycopeptide, alkylamine, and condensing agent (Figure 3). ・Sialylglycopeptide (SGP; purified from hen's egg yolk), an inexpensive, commercially available glycopeptide, is used as the starting material. ・SGP has three carboxyl groups (COOH) per molecule, which are alkylamidation sites. ・Alkylamidation of glycans and glycopeptides is a technique being studied for stabilizing sialic acid in mass spectrometry and for linkage-specific alkylamidation of sialic acid. This is a simple derivatization method.・Retention time in reversed-phase HPLC can be adjusted by the hydrocarbon chain length (level of hydrophobicity) of the alkylamine reacted with glycopeptides. ・Multiple types of alkylamidated SGPs are combined, and their retention times in reversed-phase HPLC are used as a retention time index. ・The introduction of artificial modifications reduces the likelihood of confusion with natural peptides (glycopeptides) in the sample. ・In this study, alkylamines consisting of linear alkyl groups with an amino group at one end were used. ・Reverse-phase LC / MS retention time indexing using alkylamidated SGPs can be used for both external calibration, in which they are injected and analyzed independently of the peptide sample, and internal calibration, in which they are mixed with the peptide sample and then injected and analyzed.
[0096] <Materials and Methods> <Materials> 1-(3-Dimethylaminopropyl)-3-ethylcarbodiimide hydrochloride (EDC-HCl) was purchased from Peptide Institute. Sialylglycopeptide (SGP), 1-hydroxybenzotriazole monohydrate (HOBt-H2O), ethyl cyano(hydroxyimino)acetate (Oxyma), methylamine hydrochloride, ethylamine hydrochloride, propylamine hydrochloride, butylamine hydrochloride, pentylamine hydrochloride, hexylamine, heptylamine, octylamine hydrochloride, nonylamine, decylamine, and undecylamine were purchased from Tokyo Chemical Industry Co., Ltd. Hydrochloric acid, dimethyl sulfoxide (DMSO), acetonitrile, 2-propanol, acetic acid, and trifluoroacetic acid were purchased from Fujifilm Wako Pure Chemical Industries, Ltd. Milli-Q water (Merck-Millipore) was used as ultrapure water.
[0097] <Method> SGP was dissolved in ultrapure water and dispensed in 10 nmol equivalent amounts into 1.5-mL tubes with low protein binding capacity. The solution was then dried in a centrifugal dryer. Alkylamine hydrochlorides were dissolved in DMSO, while unsalted alkylamines were converted to hydrochlorides by adding concentrated hydrochloric acid to the same concentration and then dissolved in DMSO. Decylamine and undecylamine hydrochlorides had low solubility in DMSO, so they were dissolved in 2-propanol instead. The condensation agent EDC-HCl and the condensation promoters HOBt-HO and Oxyma were each dissolved in DMSO. For the reaction, the alkylamine solution and HOBt or Oxyma solution were added to the dried SGP tube and mixed using a vortex mixer, followed by the addition of EDC-HCl. The solution was reacted in a thermomixer at 37°C and 1,000 rpm for 2 to 16–24 hours. For short hydrocarbon chains (C5 or less), HOBt-HO was used as a condensation promoter and the reaction time was 2 hours. For long chains (C6 or more), a longer reaction time of 16-24 hours was used to improve yield, and Oxyma was used as a condensation promoter. After the reaction, the reaction reagents were removed by glycopeptide purification. Acetonitrile containing 1% trifluoroacetic acid was added to the reaction solution (final acetonitrile concentration 85 vol% or more), and the reaction reagents were removed by solid-phase extraction using hydrophilic interaction liquid chromatography (HILIC).
[0098] Each alkylamidated SGP prepared was centrifuged and dried, then dissolved in a glycopeptide redissolving solution (e.g., 0.1 vol% aqueous acetic acid solution) and mixed. The resulting alkylamidated SGP mixed solution was used for LC / MS measurement, or the mixed solution was added to a glycopeptide sample solution.
[0099] <Results> <Retention Behavior of Alkylamidated Sialylglycopeptides> The prepared alkylamidated SGPs were mixed and subjected to reversed-phase nano-LC / MS (40-minute gradient conditions) to generate extracted ion chromatograms (EICs) using the m / z values shown in Table 5. The retention time behavior of each alkylamidated SGP increased with the carbon number of the alkylamine used for alkylamidation (Fig. 4). Furthermore, reversed-phase nano-LC / MS analysis of glycopeptide solutions prepared from normal human plasma was performed to measure plasma glycopeptides alone and plasma glycopeptides plus alkylamidated SGP (Fig. 5A and B). Extracted ion chromatograms of the alkylamidated SGP signal were also generated for the plasma glycopeptide plus alkylamidated SGP sample (Fig. 5C).
[0100] <Elution time indexing of alkylamidated sialylglycopeptides> To index the elution time, a unique reference value was assigned to each alkylamidated SGP. In this study, the alkylamidated SGP was assigned a value corresponding to the number of carbon atoms in the alkylamine used for amidation (e.g., methylamidated 1, ethylamidated 2, decylamidated 10). Figure 6 shows the plot of the index values and retention times for each alkylamidated SGP under different gradient conditions of 40 minutes and 120 minutes.
[0101] In addition, the retention times of representative glycopeptides contained in the plasma glycopeptide sample were obtained under different gradient conditions of 40 minutes and 120 minutes, respectively, and converted into index values using alkylamidated SGP (Table 6). Information based on the Byonic (v4.3.4) search results for the glycopeptides shown in Table 6 is shown in Table 7. The amino acid sequences of the peptides shown in Table 7 are shown in SEQ ID NOS: 59-76. As a result, by converting the retention times, which differ greatly depending on the chromatographic gradient conditions, into index values, it became possible to make them comparable.
[0102]
[0103] Table 5 shows the m / z values used to generate the extracted ion chromatogram of alkylamidated SGP. All ion species are triply protonated ([M+3H] 3+ ) is equivalent to
[0104]
[0105] Table 6 shows the m / z values for creating extracted ion chromatograms for representative glycopeptides contained in glycopeptide samples prepared from plasma (treated with Top14 Abundant Protein Depletion), retention times under 40-minute and 120-minute gradient conditions, and retention times indexed using the retention times of alkylamidated SGP.
[0106]
[0107] Table 7 shows the peptide sequences, glycan compositions, and protein information for the glycopeptides shown in Table 6. This information is based on the results obtained by searching and processing the LC / MS / MS measurement data of plasma glycopeptides using Byonic.
Claims
1. A method for identifying or predicting a glycopeptide, comprising: (1) subjecting a glycopeptide sample prepared from a biological sample collected from a subject to liquid chromatography-mass spectrometry to obtain retention time information Xa and mass information Xb for each peak; (2) preparing a reference glycopeptide list including retention time information Ya and mass information Yb in liquid chromatography-mass spectrometry of a plurality of glycopeptides previously identified; and (3) comparing, for each peak, the retention time information Xa and the retention time information Ya after correcting at least one of them, and comparing the mass information Xb and the mass information Yb, and identifying or predicting the peak as a glycopeptide if both are consistent.
2. The identification method according to claim 1, wherein the reference glycopeptide list contains 500 or more glycopeptides.
3. The identification or estimation method described in claim 1, wherein the reference glycopeptide list includes a glycopeptide list obtained by liquid chromatography-mass spectrometry of a glycopeptide sample prepared from the same biological sample as the biological sample in step (1).
4. The identification or estimation method according to claim 1, wherein step (1) is performed for each of a plurality of subjects.
5. The identification or estimation method according to claim 4, wherein the number of subjects is 10 or more.
6. The identification or estimation method according to claim 1, wherein the correction in step (3) is performed based on the retention time of a reference internal standard substance.
7. A reagent for use as a reference internal standard in the identification method according to claim 6, comprising a plurality of types of reference internal standard substances having different retention times in liquid chromatography-mass spectrometry.
8. The reagent of claim 7, wherein the reference internal standard is a glycopeptide comprising a base glycopeptide and a carbohydrate group added to the base glycopeptide.
9. The reagent according to claim 8, wherein the molecular weight of the hydrocarbon group differs among the multiple types of glycopeptides.
Citation Information
Patent Citations
Method of analyzing sugar peptide tandem mass data
JP2008232650A
Hydrophilic interaction chromatography materials, preparation method therefor, and use thereof for analysis of glycoprotein and glycopeptide
JP2015121539A
Identification method for bonding area of sugar chain of glycoprotein and kit
JP2019052995A
Identification and Use of Glycopeptides as Biomarkers for Diagnostics and Therapeutic Monitoring
JP2020532732A
Glycopeptide analyzer
JP2021081365A