Method for preparing trace protein-containing sample for proteome analysis by using lectin

By employing a lectin-based method to selectively recover trace proteins from serum and plasma, the limitations of current proteome analysis techniques are overcome, allowing for a more comprehensive detection of biomarker proteins.

WO2025126730A1PCT designated stage expired Publication Date: 2025-06-19KAZUSA DNA RES INST
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/039643
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-15
Filing Date
2024-11-07
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Current methods for proteome analysis of serum and plasma are limited by the large dynamic range of protein concentrations, which prevents the detection of trace proteins that could serve as biomarkers.

Method used

A method involving the use of a column immobilized with N-acetylglucosamine oligomer-binding lectins or fucose-binding lectins to selectively recover trace proteins while minimizing contamination from high-abundance proteins.

Benefits of technology

This method efficiently removes high-abundance proteins and preferentially concentrates trace proteins, enhancing the depth of proteome analysis and increasing the number of detectable proteins by LC-MS/MS.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024039643_19062025_PF_FP_ABST
    Figure JP2024039643_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a method which is for preparing a sample for proteome analysis and by which trace proteins in a biological sample can be efficiently recovered while removing as many highly abundant proteins, hindering more in-depth analysis, as possible in proteome analysis. This method for preparing a sample for proteome analysis involves: bringing a biological sample into contact with a carrier in which any lectin selected from the group consisting of N-acetylglucosamine oligomer-binding lectin and fucose-binding lectin, or a binding fragment thereof is immobilized; and thereby recovering a glycoprotein having binding properties to the lectin in the biological sample. In one aspect, a biological sample is brought into contact with a carrier in which a combination containing potato lectin or a binding fragment thereof and tomato lectin or a binding fragment thereof is immobilized.
Need to check novelty before this filing date? Find Prior Art

Description

Method for preparing trace protein-containing samples for proteome analysis using lectins

[0001] The present invention relates to a method for preparing a sample containing trace amounts of protein for proteome analysis from a biological sample using a lectin.

[0002] Serum and plasma can be collected minimally invasively and are the most common specimens used in clinical testing. In fact, most biomarker-based tests use serum and plasma, so it is expected that new biomarkers will be discovered from serum and plasma. Another advantage of biomarker discovery from serum and plasma is that since serum and plasma are collected in most biobanks, using these samples can save time and effort in specimen collection.

[0003] Although biomarker discovery has been actively conducted through proteome analysis of serum and plasma, there have been few successful examples of discovering biomarkers that can be used in actual clinical practice. This is presumably because the dynamic range of protein concentrations in serum and plasma is extremely large, so only proteins present in large amounts in serum or plasma can be analyzed, which does not ensure sufficient breadth in biomarker discovery. Of course, as the depth of proteomics analysis improves and the breadth of biomarker discovery (the number of detectable proteins) expands, the probability of finding promising biomarkers will certainly increase. Therefore, in order to search for biomarker proteins using mass spectrometry, it is necessary to narrow the dynamic range and perform in-depth proteome analysis.

[0004] The most common method for reducing the dynamic range of protein concentrations in serum and plasma is to use antibody columns to remove highly abundant proteins such as albumin, IgG, and transferrin (Non-Patent Documents 1-3). While depletion of highly abundant proteins is an effective method, no antibody column has yet been developed that can completely deplete these proteins, making it insufficient for detecting trace proteins in serum. One approach to overcome this problem and observe trace proteins in serum and plasma without being affected by highly abundant proteins is to collect extracellular vesicles and analyze them by LC-MS (Non-Patent Document 4). Recently, methods using chemical nanoparticle cocktails to collect trace proteins in serum and plasma have attracted attention (Non-Patent Documents 5 and 6). Meanwhile, target-based analysis using antibodies and aptamers has also developed, with companies such as Olink and SomaLogic offering services to simultaneously measure thousands of proteins in serum and plasma (Non-Patent Documents 7 and 8). Thus, various technologies are currently being developed to broaden the scope of biomarker discovery.

[0005] It is estimated that over 50% of the human proteome is glycosylated (Non-Patent Document 9). Glycosylation is a highly complex post-translational modification that plays an important role in various biological phenomena, such as cell differentiation, intercellular signaling, and immunity (Non-Patent Documents 10 and 11). Furthermore, many of the proteins reported as biomarkers for cancer and other diseases are known to have glycans attached (Non-Patent Documents 12-14). However, there have been few studies using LC-MS to identify biomarkers focusing on glycoproteins.

[0006] Lectin is a general term for proteins that reversibly bind to monosaccharides or sugar chains. A vast variety of lectins have been discovered in animals, plants, bacteria, viruses, and other organisms. Among these, plant lectins are classified into seven families of structurally and evolutionarily closely related proteins (Non-Patent Document 15). These families include (i) the amaranthin lectin family, (ii) chitin-binding lectins, (iii) Cucurbitaceae nodal lectins, (iv) jacalin-related lectins, (v) legume lectins, (vi) monocotyledonous mannose-binding lectins, and (vii) type 2 RIPs. Among these, the chitin-binding lectin family consists of plant proteins containing at least one so-called hevein domain and are frequently found in higher plants. Chitin-binding lectins include those from Solanaceae plants (potato, tomato, Datura, etc.), wheat germ agglutinin (WGA), nettle lectin, mistletoe lectin, celandine lectin, and pokeweed lectin (Non-Patent Document 16). N-acetylglucosamine oligomer ((GlcNAc)n)-binding lectins belong to the families of chitin-binding lectins with hevein domains, Cucurbitaceae nodal lectins, and Legume lectins.

[0007] Lectins are useful for isolating and purifying glycoproteins and sugar chains, and various lectin columns are commercially available.

[0008] Methods Mol Biol, 1002, 1-11 (2013)ACS Omega, 7, 7012-7023 (2022)Rheumatology (Oxford) 62, 3501-3506 (2023)Mass Spectrom Rev, 42, 779-795 (2023)Pediatr Hematol Oncol, 39, 658-671 (2022)Nat Commun, 11, 3662 (2020)Nat Commun 12, 2493 (2021)Nat Commun 12, 6822 (2021)Biochim Biophys Acta, 1473, 4-8 (1999)Cell 126, 855-867 (2006)Nat Rev cancer, 15, 540-555 (2015)Electrophoresis, 33, 1746-1754 (2012)Glycoconj J, 29, 249-258 (2012)Technol Cancer Res Treat 22, 15330338221148811(2023)Crit Rev Pit Sci 17, 575-692 (1998)Trends in Glycoscience and Glycotechnology, 12, No.64, pp.83-101 (2000)The Plant Journal, 37, 34-45 (2003)

[0009] The present invention aims to provide a method for preparing a sample for proteome analysis, which is capable of efficiently recovering trace proteins in a biological sample, such as serum or plasma, while removing as many abundant proteins as possible that hinder further analysis in proteome analysis of the biological sample.

[0010] The present inventors have conducted extensive research to solve the above-mentioned problems and have succeeded in recovering trace proteins with high efficiency while minimizing contamination by highly abundant proteins by applying serum samples to a column immobilized with N-acetylglucosamine oligomer-binding lectins such as potato lectin (Solanum Tuberosum Lectin: STL), tomato lectin (Lycopersicon Esculentum Lectin: LEL), Datura Stramonium Lectin (DSL), or wheat germ agglutinin (WGA), or fucose-binding lectins such as Ulex Europaeus Agglutinin-1 (Ulex Europaeus Agglutinin-1), Aleuria Aurantia lectin (AAL), or Aspergillus oryzae lectin (AOL). LC-MS / MS analysis of samples prepared with N-acetylglucosamine oligomers or fucose-binding lectins yielded a greater number of identified proteins than samples prepared by antibody column depletion of 14 highly abundant proteins (TOP14D). Although highly abundant proteins in serum were also glycosylated, their binding to N-acetylglucosamine oligomers or fucose-binding lectins was limited. Meanwhile, trace proteins, such as membrane proteins, released in serum were preferentially enriched. Combining N-acetylglucosamine oligomer-binding lectins with other lectins with different sugar specificities was expected to recover glycoproteins with diverse glycan structures and increase the number of protein identifications. However, surprisingly, this effect was largely absent. Instead, combining multiple N-acetylglucosamine oligomer-binding lectins improved the number of protein identifications. In particular, STL and LEL both belong to the Solanaceae lectin family and have similar sugar specificities, structures, and physical properties (Non-Patent Document 17). However, the combination of STL and LEL significantly increased the number of protein identifications compared to the use of either STL or LEL alone.When proteins bound to a lectin column are eluted with SDS, the SP3 method requires the recovery of the eluted proteins to remove the SDS, which inhibits enzymatic reactions. However, by eluting the proteins from the lectin column with an acidic aqueous solution and neutralizing the eluate with a buffer to return the pH to the optimal pH for enzymatic digestion, enzymatic digestion and subsequent proteome analysis can be performed without the SP3 step. Furthermore, adding LMNG to the acidic aqueous solution improved the yield of proteins eluted from the lectin column and increased the number of protein identifications. Based on these findings, the inventors conducted further research and completed the present invention.

[0011] That is, the present invention relates to the following: [1] A method for preparing a protein- or peptide-containing sample for identifying proteins contained in a biological sample by mass spectrometry, comprising contacting the biological sample with a support to which a lectin or a binding fragment thereof selected from the group consisting of N-acetylglucosamine oligomer-binding lectins and fucose-binding lectins has been immobilized, thereby recovering glycoproteins from the biological sample that have binding affinity to the lectin. [2] The preparation method of [1], wherein the lectin or binding fragment thereof is an N-acetylglucosamine oligomer-binding lectin or binding fragment thereof selected from the group consisting of potato lectin, tomato lectin, Datura stramonium lectin, and wheat germ agglutinin. [3] The preparation method of [1], wherein the lectin or binding fragment thereof is a fucose-binding lectin or binding fragment thereof selected from the group consisting of Ustilago sativa agglutinin-1, Pichia japonica lectin, and Aspergillus oryzae lectin. [4] The preparation method of [1], wherein the lectin or binding fragment thereof is a combination comprising 1) potato lectin or a binding fragment thereof, and 2) any lectin or binding fragment thereof selected from the group consisting of tomato lectin, Datura stramonium lectin, wheat germ agglutinin, and P. sieboldii lectin. [5] The preparation method of [4], wherein the lectin or binding fragment thereof is a combination comprising 1) potato lectin or a binding fragment thereof, and 2) tomato lectin or a binding fragment thereof. [6] The preparation method of any of [1] to [5], further comprising subjecting glycoproteins in the collected biological sample to a reductive alkylation treatment. [7] The preparation method of any of [1] to [6], further comprising digesting glycoproteins in the collected biological sample with a protease to obtain a peptide mixture as a digestion product. [8] The preparation method of any of [1] to [7], further comprising subjecting the obtained peptide mixture to a desalting column to obtain a purified peptide mixture. [9] The preparation method according to any one of [1] to [8], wherein the biological sample is any one selected from the group consisting of serum, plasma, whole blood, cerebrospinal fluid, ascites, synovial fluid, and lymphatic fluid.

[10] The preparation method of any of [1] to [9], which comprises eluting proteins bound to a lectin or a binding fragment thereof with an acidic aqueous solution to recover glycoproteins in a biological sample that have binding ability to the lectin.

[11] The preparation method of

[10] , wherein the acidic aqueous solution contains a sugar-based nonionic surfactant.

[12] The preparation method of

[11] , wherein the sugar-based nonionic surfactant is LMNG.

[13] The preparation method of any of

[10] to

[12] , further comprising neutralizing the eluate and treating it with a protease to digest glycoproteins in the recovered biological sample in the eluate with the protease, thereby obtaining a peptide mixture as a digestion product.

[14] A method for identifying proteins contained in a biological sample by mass spectrometry, comprising the following steps: 1) contacting the biological sample with a support having immobilized thereon a lectin or a binding fragment thereof selected from the group consisting of N-acetylglucosamine oligomer-binding lectins and fucose-binding lectins, thereby recovering glycoproteins in the biological sample that have binding properties to the lectin; 2) digesting the glycoproteins in the recovered biological sample with a protease to obtain a peptide mixture as a digestion product; 3) subjecting the obtained peptide mixture to a desalting treatment to obtain a purified peptide mixture; 4) subjecting the purified peptide mixture to a mass spectrometer, and identifying peptides contained in the peptide mixture based on the obtained mass spectrometry results; and 5) identifying the proteins in the biological sample from which the identified peptides are derived based on information about the identified peptides.

[15] The method of

[14] , wherein the lectin or binding fragment thereof is any N-acetylglucosamine oligomer-binding lectin or binding fragment thereof selected from the group consisting of potato lectin, tomato lectin, Datura stramonium lectin, and wheat germ agglutinin.

[16] The method of

[14] , wherein the lectin or binding fragment thereof is any fucose-binding lectin or binding fragment thereof selected from the group consisting of Ustilago sativa agglutinin-1, Pichia japonica lectin, and Aspergillus oryzae lectin.

[17] The method of

[14] , wherein the lectin or binding fragment thereof is a combination comprising 1) potato lectin or a binding fragment thereof, and 2) any lectin or binding fragment thereof selected from the group consisting of tomato lectin, Datura stramonium lectin, wheat germ agglutinin, and Pichia stramonium lectin.

[18] The method of

[17] , wherein the lectin or binding fragment thereof is a combination comprising 1) potato lectin or a binding fragment thereof, and 2) tomato lectin or a binding fragment thereof.

[19] Any of the methods of

[14] to

[18] , further comprising subjecting glycoproteins in the biological sample collected in step 1 to a reductive alkylation treatment.

[20] Any of the methods of

[14] to

[19] , wherein the biological sample is any one selected from the group consisting of serum, plasma, whole blood, cerebrospinal fluid, ascites, synovial fluid, and lymph.

[21] The method of any of

[14] to

[20] , wherein in step 1, proteins bound to the lectin or a binding fragment thereof are eluted with an acidic aqueous solution to recover glycoproteins in a biological sample that have binding properties to the lectin.

[22] The method of

[21] , wherein the acidic aqueous solution contains a sugar-based nonionic surfactant.

[23] The method of

[22] , wherein the sugar-based nonionic surfactant is LMNG.

[24] The method of any of

[21] to

[23] , wherein the eluate is neutralized and subjected to digestion with a protease in step 2.

[0012] According to the present invention, it is possible to efficiently remove high-abundance proteins that hinder the in-depth of LC-MS analysis from biological samples such as serum and plasma, and to preferentially concentrate trace proteins that are useful as biomarkers. Samples prepared by the methods of the present invention contain highly concentrated trace proteins, are free of contaminating high-abundance proteins, and have a narrow dynamic range of protein concentrations. Therefore, when used in LC-MS analysis, these samples can improve the depth of proteome analysis and enable the detection of a large number of proteins.

[0013] Figure 1 shows a schematic diagram of the automated enrichment method for trace serum proteins using biotinylated lectins. The Maelstrom 9610 can accommodate eight 96-deep-well plates. One of these plates requires a dedicated tip for this instrument. Each plate can be rotated into a predetermined position under a 96-pin magnetic head that is lowered into the 96-well plate to bind, release, or mix the magnetic beads in solution. The Maelstrom 9610 was used to automatically process the reaction of biotinylated lectins with streptavidin beads, the reaction of lectin-bound beads with serum, and the washes between each step. After reduction and alkylation, the lectin-enriched serum was automatically processed with SP3 using the Maelstrom 9610. Figure 2 shows the number of proteins detected in serum enriched with 37 lectins. As a representative of a typical method, TOP14 high-abundance protein depletion (TOP14D)-treated serum was also analyzed. Each treated sample was analyzed by LC-MS / MS after trypsin digestion using 200 ng of peptides with a 60-minute activity gradient. The horizontal, diagonal, and gray bars represent lectins with specificity for N-acetylglucosamine, fucose, and other sugars, respectively. Figure 3 shows the number of proteins detected in serum enriched with two or three lectin combinations. Serum treated with STL alone was also analyzed as a control. Each treated sample was analyzed by LC-MS / MS after trypsin digestion using 200 ng of peptides with a 60-minute activity gradient. Figure 4 shows the number of proteins detected in serum from various species enriched with STL / LEL. 200 ng of tryptic peptides were analyzed by LC-MS / MS using a 60-minute activity gradient. Figure 5A shows a Venn diagram comparing the number of proteins observed in crude serum, TOP14D-treated serum, and STL / LEL-enriched serum using a 120-minute activity gradient LC-MS / MS. 500 ng of tryptic peptides were measured. The numbers in the figure indicate the overlap of proteins identified in each treatment. Figure 5B shows the cellular components of the GO enrichment analysis of proteins observed only in STL / LEL-enriched serum.Figures 6A–D show the results of protein analysis in the serum of SLE model mice and control mice using the STL / LEL and TOP2D methods. Figure 6A shows a Pearson correlation coefficient heat map of protein intensities for each sample. Figure 6B shows a volcano plot of protein intensities obtained from the serum of SLE model mice and control mice. Red dots (points in the upper right area of ​​each graph) represent proteins with increased expression in the serum of R-SLE model mice, and blue dots (points in the upper left area of ​​each graph) represent proteins with decreased expression in the serum of SLE model mice. Figure 6C shows a Venn diagram showing the overlap of differentially expressed proteins in the serum of SLE model mice detected by the STL / LEL and TOP2D methods. Figure 6D shows a graph showing disease ontology enrichment analysis of proteins elevated in the serum of SLE model mice. Disease terms related to SLE and its complications are shown. Figure 7 shows a schematic diagram of the automated enrichment method for trace serum proteins using an improved biotinylated lectin. Figure 8 is a graph comparing the number of identified proteins when lectin column-bound proteins were eluted with the SDS-SP3 method, 0.2% TFA, and 0.2% TFA / 0.05% LMNG aqueous solution, with the SDS-SP3 method. Figure 9 is a graph comparing the number of identified proteins when lectin column-bound proteins were eluted with various types of acidic aqueous solutions.

[0014] The present invention provides a method for preparing a protein- or peptide-containing sample for identifying proteins contained in the biological sample by mass spectrometry, the method comprising contacting the biological sample with a support on which a lectin selected from the group consisting of N-acetylglucosamine oligomer-binding lectins and fucose-binding lectins, or a binding fragment thereof, has been immobilized, and thereby recovering glycoproteins from the biological sample that have binding affinity to the lectin.

[0015] A biological sample is treated with a carrier on which a lectin selected from the group consisting of N-acetylglucosamine oligomer-binding lectins and fucose-binding lectins, or a binding fragment thereof, is immobilized; glycoproteins in the biological sample that have binding properties to these lectins (which may be N-acetylglucosamine oligomer-containing glycoproteins or fucose-containing glycoproteins) are adsorbed onto the carrier; and the carrier is washed to remove non-adsorbed components. This allows for efficient removal of highly abundant proteins that hinder in-depth analysis in proteome analysis, and allows preferential enrichment of trace proteins that are useful as biomarkers.

[0016] Examples of biological samples used in the present invention include, but are not limited to, body fluids such as blood (serum, plasma, whole blood), cerebrospinal fluid, ascites, synovial fluid, lymph, saliva, urine, sweat, sputum, tears, nasal discharge, semen, pleural effusion, pericardial fluid, and tissue fluid, cells, tissues or parts thereof, cell or tissue homogenates, cell or tissue extracts, biopsy samples, swab samples, feces, cell cultures, bacteria, viruses, and fungi. The preparation method of the present invention can efficiently remove highly abundant proteins such as albumin, immunoglobulins, and transferrin from a biological sample and preferentially concentrate trace proteins. Therefore, the preparation method of the present invention is advantageous for preparing trace protein-containing samples for proteome analysis from biological samples rich in highly abundant proteins (e.g., serum, plasma, whole blood, cerebrospinal fluid, ascites, synovial fluid, lymph, and the like), particularly serum and plasma.

[0017] The biological sample may be a sample obtained by extracting or concentrating a specific biological material fraction, or a sample diluted with a buffer solution, etc. Removal of highly abundant proteins (e.g., albumin, IgG, antitrypsin, IgA, transferrin, haptoglobin, fibrinogen, α2-macroglobulin, α1-acid glycoprotein, IgM, apolipoprotein AI, apolipoprotein AII, complement protein C3, transthyretin 9, etc.) from a biological sample using an affinity removal system with an antibody column is widely used. Therefore, a biological sample after removal of highly abundant proteins using such an antibody column may be used as a starting material in the present invention. However, the preparation method of the present invention does not require such an antibody column, and allows the preparation of a sample for proteome analysis that is rich in trace proteins and suppresses contamination with highly abundant proteins. Therefore, in one embodiment, the preparation method of the present invention uses a biological sample from which at least one highly abundant protein selected from the group consisting of albumin, IgG, antitrypsin, IgA, transferrin, haptoglobin, fibrinogen, α2-macroglobulin, α1-acid glycoprotein, IgM, apolipoprotein AI, apolipoprotein AII, complement protein C3, and transthyretin 9 has not been removed by antibody column treatment.

[0018] The biological species from which the biological sample is derived is not particularly limited, and examples thereof include animals, plants, bacteria, fungi, viruses, etc. The type of animal is also not particularly limited, and examples thereof include mammals, birds, reptiles, amphibians, fish, etc., with mammals or birds being preferred. Examples of mammals include rodents such as mice, rats, hamsters, and guinea pigs; lagomorphs such as rabbits; ungulates such as pigs, cows, goats, horses, and sheep; carnivores such as dogs and cats; and primates such as humans, monkeys, rhesus monkeys, cynomolgus monkeys, marmosets, orangutans, and chimpanzees. The mammal is preferably a primate (such as a human) or a rodent (such as a mouse).

[0019] Examples of N-acetylglucosamine oligomer-binding lectins include, but are not limited to, lectins from Solanaceae plants (potato, tomato, Datura stramonium, etc.), wheat germ agglutinin (WGA), nettle lectin, mistletoe lectin, celandine lectin, and pokeweed lectin (PWM) (Non-Patent Document 18) (Trends in Glycoscience and Glycotechnology, Vol. 12, No. 64 (March 2000), pp. 83-101). The N-acetylglucosamine oligomer-binding lectin is preferably potato lectin (STL), tomato lectin (LEL), Datura stramonium lectin (DSL), or wheat germ agglutinin (WGA), more preferably potato lectin (STL) or tomato lectin (LEL). A representative amino acid sequence of potato lectin is shown in SEQ ID NO: 1, and a representative amino acid sequence of tomato lectin is shown in SEQ ID NO: 2. The amino acids 22 to 323 in the amino acid sequence shown in SEQ ID NO: 1 may correspond to mature potato lectin. The amino acids 24 to 365 in the amino acid sequence shown in SEQ ID NO: 2 may correspond to mature tomato lectin (The Plant Journal (2004) 37, 34-45; Biosci. Biotechnol. Biochem., 72, 80310-1-11, 2008).

[0020] Examples of fucose-binding lectins include, but are not limited to, Urea agglutinin-1 (UEA-1), Atractylodes macrocarpa lectin (AAL), and Aspergillus oryzae lectin (AOL).

[0021] As used herein, "N-acetylglucosamine oligomer" refers to a multimer of N-acetylglucosamine in which N-acetylglucosamine units are linked together via β1-4 bonds. The length of the multimer is not particularly limited, but includes dimers, trimers, tetramers, and the like. As used herein, when a lectin has the ability to bind to a sugar chain with a specific structure, the dissociation constant (Kd value) of the binding affinity of the lectin to the sugar chain is generally 1 × 10 -2M or less (e.g., 1×10 -3 M or less, 1×10 -4 M or less, 1×10 -5 M or less). The binding affinity of an "N-acetylglucosamine oligomer-binding lectin" can be evaluated by assessing its binding affinity to N-acetylglucosamine oligomers of a specific length (e.g., tetramers), allowing for relative comparison of binding affinity between different types of lectins. Binding affinity can be determined, for example, using surface plasmon resonance (BIAcore™) analysis.

[0022] As used herein, a "binding fragment" of a lectin refers to a protein or peptide containing a portion (partial fragment) of a lectin, which maintains the glycan-binding ability of the parent lectin. In one aspect, the "binding fragment" of a lectin comprises a continuous partial sequence that is 30% or more (e.g., 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, or 90% or more) in length of the full-length amino acid sequence of the mature form of the parent lectin.

[0023] In the present invention, any lectin or binding fragment thereof selected from the group consisting of N-acetylglucosamine oligomer-binding lectins and fucose-binding lectins may be used alone, but the use of two or more lectins in combination is expected to improve the number of protein identifications in proteome analysis. In one embodiment, 1) potato lectin (STL) or a binding fragment thereof and 2) any lectin or binding fragment thereof selected from the group consisting of tomato lectin (LEL), Datura stramonium lectin (DSL), wheat germ agglutinin (WGA), and Acanthus arvensis lectin (AAL) are used in combination.

[0024] In particular, potato lectin (STL) and tomato lectin (LEL), both of which belong to the Solanaceae family of lectins, are similar in terms of sugar specificity, structure, physical properties, etc. (Non-Patent Document 19) (The Plant Journal, 37, 34-45 (2003)). However, as shown in the Examples below, unexpectedly, when STL and LEL were combined, the number of protein identifications was significantly increased compared to when STL or LEL was used alone. Therefore, in a preferred embodiment, 1) potato lectin (STL) or a binding fragment thereof and 2) tomato lectin (LEL) or a binding fragment thereof are used in combination.

[0025] When multiple types of lectins or binding fragments thereof are used in combination, the blending ratio is not particularly limited. Preferably, the blending amount (molar amount) of the lectin or binding fragment thereof with the highest content is 10 times or less (preferably, 5 times or less, more preferably, 2 times or less) that of the lectin or binding fragment thereof with the lowest content.

[0026] As used herein, the term "carrier" refers to a substance that serves as a base for adsorbing or immobilizing other substances. Examples of carriers include, but are not limited to, beads, filters, cartridges, columns, porous material plates (e.g., multi-well plates), etc. Examples of beads include, but are not limited to, silica particles, agarose particles, polymer particles such as polystyrene, polyacrylamide, polydivinylbenzene, and silicon, magnetic particles, and glass particles. Various functional groups may be bound to the surface of the carrier.

[0027] Immobilization of a lectin or its binding fragment on a carrier can be carried out by protein immobilization methods well known to those skilled in the art. For example, by biotinylating a lectin or its binding fragment and mixing it with a carrier to which streptavidin has been bound, the biotin on the lectin or its binding fragment binds to the streptavidin on the carrier, thereby immobilizing the lectin or its binding fragment on the carrier. Alternatively, a crosslinker such as N-hydroxysuccinimide (NHS) can be used to immobilize the lectin or its binding fragment on the carrier.

[0028] When multiple types of lectins or binding fragments thereof are used in combination, a mixture of multiple types of lectins or binding fragments thereof may be immobilized on a carrier, or each lectin or binding fragment thereof may be immobilized individually on a carrier, and then multiple types of carriers on which the resulting lectins or binding fragments thereof are immobilized may be mixed.

[0029] To suppress non-specific binding of proteins to the carrier surface, the carrier on which the lectin or its binding fragment is immobilized may be blocked with BSA, FBS, skim milk, Block Ace (manufactured by Dainippon Sumitomo Pharma Co., Ltd.), or the like.

[0030] The carrier on which the resulting lectin or its binding fragment is immobilized is mixed with a biological sample, thereby contacting the biological sample with a carrier on which a lectin or its binding fragment selected from the group consisting of N-acetylglucosamine oligomer-binding lectins and fucose-binding lectins is immobilized. As a result, glycoproteins in the biological sample that have binding affinity to the lectin or its binding fragment on the carrier are adsorbed to the lectin or its binding fragment on the carrier and recovered. To prevent nonspecific adsorption and contamination with highly abundant proteins, the carrier on which the glycoproteins are adsorbed may be washed with an appropriate buffer. Examples of buffers that can be used include, but are not limited to, Tris-HCl buffer, phosphate buffer, citrate buffer, acetate buffer, HEPES buffer, borate buffer, and tartrate buffer. The pH of the buffer is approximately 6 to 9. A surfactant may be added to the buffer for effective washing. Examples of surfactants include, but are not limited to, Tween 20, Triton X-100, and NP-40. An inorganic salt may also be added to the buffer. Examples of inorganic salts include, but are not limited to, sodium chloride, potassium chloride, calcium chloride, magnesium sulfate, ammonium sulfate, and the like.

[0031] The carrier, on which glycoproteins are adsorbed to immobilized lectins or binding fragments thereof, may be treated with an elution solution to release and elute the adsorbed glycoproteins into the elution solution. In one embodiment, the elution solution is an aqueous solution of a denaturant. Denaturants that can be used include, but are not limited to, sodium lauryl sulfate (SDS), urea, guanidine hydrochloride, and the like. In another embodiment, the elution solution is an aqueous solution of a hapten sugar corresponding to the lectin or binding fragment thereof immobilized on the carrier. For example, when an N-acetylglucosamine oligomer-binding lectin or binding fragment thereof is used, the carrier is treated with an aqueous solution of N-acetylglucosamine oligomers (e.g., N-acetylglucosamine dimers, trimers, tetramers, pentamers, hexamers, etc.), and when a fucose-binding lectin or binding fragment thereof is used, the carrier is treated with an aqueous fucose solution. The elution solution may contain buffers such as Tris-HCl, phosphate, citrate, acetate, HEPES, borate, or tartrate, or inorganic salts such as sodium chloride, potassium chloride, calcium chloride, magnesium sulfate, or ammonium sulfate. To avoid loss of trace proteins in the sample, the next treatment may be carried out while the N-acetylglucosamine oligomer-containing glycoprotein remains adsorbed on the carrier, without the elution step.

[0032] In a further aspect, the elution solution is an acidic aqueous solution. The pH of the acidic aqueous solution is usually 5 or less, preferably 4 or less, and more preferably 3 or less. The type of acidic aqueous solution is not particularly limited as long as it can release glycoproteins adsorbed to lectins or their binding fragments. Examples include aqueous solutions of organic acids such as trifluoroacetic acid, acetic acid, and formic acid, and aqueous solutions of inorganic acids such as hydrochloric acid, sulfuric acid, and nitric acid. The acid concentration in the acidic aqueous solution is not particularly limited as long as it can release glycoproteins adsorbed to lectins or their binding fragments. The acid concentration is usually 0.01 to 1.0% (w / w). The eluate containing the released glycoproteins obtained by elution with the acidic aqueous solution may be neutralized. The pH of the eluate after neutralization is preferably adjusted to a pH at which the protease described below is active (preferably the optimal pH). This varies depending on the type of protease, but is usually within the range of 7.0 to 10.0. Neutralization can be performed using a buffer (e.g., Tris-HCl) or an alkaline aqueous solution (e.g., NaOH aqueous solution) well known to those skilled in the art.

[0033] A surfactant may be added to the elution solution to suppress loss of the eluted glycoprotein and improve the yield of the recovered glycoprotein. The type of surfactant is not particularly limited, but from the viewpoint of suppressing adsorption of glycoproteins or peptides to the container, etc. while avoiding contamination of the mass spectrometer, a sugar-based nonionic surfactant is preferred. A sugar-based nonionic surfactant is a nonionic surfactant having a sugar group as a hydrophilic group unit. Examples of sugar-based nonionic surfactants include, but are not limited to, 2,2-didecylpropane-1,3-bis-β-D-maltopyranoside (lauryl maltose neopentyl glycol) (LMNG), 2,2-dioctylpropane-1,3-bis-β-D-maltopyranoside (decyl maltose neopentyl glycol) (DMNG), 2,2-dihexylpropane-1,3-bis-β-D-glucopyranoside (octyl glucose neopentyl glycol), n-dodecyl-β-D-maltoside (DDM), n-undecyl-β-D-maltoside, 3-oxatridecyl-α-D-mannoside, and α-D-glucopyranosyl-α-D-glucopyranoside monododecanoate (trehalose C12). Among these, LMNG has a strong affinity for reversed-phase columns and is retained in the column even under elution conditions that would elute most peptides from the column. Therefore, LMNG can be removed from peptides while avoiding peptide loss during the desalting step using the reversed-phase column, thereby minimizing the possibility of contamination of the LC-MS instrument ( WO 2024 / 127796 A1 ). The concentration of the sugar-based nonionic surfactant can be adjusted appropriately depending on the type of sugar-based nonionic surfactant and the concentration of glycoproteins or peptides in the eluate, but is, for example, 0.00125% (w / w) or more, 0.0025% (w / w) or more, 0.005% (w / w) or more, 0.01% (w / w) or more, or 0.05% (w / w) or more. Sugar-based nonionic surfactants are typically used at a concentration of 0.5% (w / w) or less, e.g., 0.1% (w / w) or less.

[0034] To cleave intermolecular and intramolecular disulfide bonds between cysteines in the protein and prevent disulfide formation, the recovered glycoprotein may be subjected to reductive alkylation. Glycoprotein reduction can be achieved by heating the glycoprotein adsorbed to or eluted from the carrier with a reducing agent such as 1,4-dithiothreitol (DTT), 2-mercaptoethanol, or tris(2-carboxyethyl)phosphine. Glycoprotein alkylation can also be achieved by reacting the reduced glycoprotein with an alkylating agent such as iodoacetamide, iodoacetic acid, acrylamide, or chloroacetamide. To promote reductive alkylation, a chaotropic agent may be added during the reduction and / or alkylation. Examples of chaotropic agents include, but are not limited to, urea, thiourea, guanidine hydrochloride, thiocyanate, and sarcosine. After the reductive alkylation reaction, unreacted reducing agent and alkylating agent may be removed. For example, the carrier onto which the reduced and alkylated glycoprotein is adsorbed is washed with an appropriate buffer. The buffer may be any of those listed as buffers used for washing carriers onto which glycoproteins are adsorbed. The reduced and alkylated treatment is an optional step, and may not be performed when speed and labor saving are important.

[0035] To facilitate subsequent processing of the eluted glycoproteins and minimize sample loss during buffer exchange, the glycoproteins may be adsorbed onto a new carrier. For example, the single-pot solid-phase-enhanced sample preparation (SP3) method, in which glycoproteins are bound to the hydrophilic surface of carboxylic acid-coated beads, is known (Mol. Syst. Biol. 2014, 10, 757). Adsorption onto a new carrier can be performed either before or after the reductive alkylation treatment.

[0036] When an acidic aqueous solution is used as the elution solution, the eluate can be neutralized to a pH (preferably an optimal pH) at which the protease described below is active, and then treated with the protease. This allows the recovered proteins in the extract to be digested with the protease while still in the liquid phase, yielding a peptide mixture as a digestion product. In this case, there is no need to adsorb the eluted proteins onto a new carrier, which eliminates the need for buffer exchange or the removal of denaturants that inhibit the activity of the protease contained in the elution solution. This avoids the risk of sample loss associated with adsorption onto the carrier, and can increase the number of protein identifications. Therefore, in one embodiment, the eluted glycoproteins are not adsorbed onto a new carrier.

[0037] Next, the recovered glycoproteins (including those that have been reduced and alkylated) are digested with a protease. Specifically, the carrier with adsorbed glycoproteins is incubated in an aqueous solvent containing the protease, or the protease is added to the glycoprotein-containing eluate and then incubated. Examples of protease include, but are not limited to, endoproteases such as trypsin, Glu-C, Lys-N, Lys-C, Asp-N, and chymotrypsin. Multiple types of proteases may also be used in combination. It is preferable to add an appropriate buffer to the aqueous solvent used for the digestion reaction to maintain pH. Examples of buffers that can be used include, but are not limited to, Tris-HCl, phosphate, citrate, acetate, HEPES, borate, and tartrate. Inorganic salts may also be added to the aqueous solvent. Examples of inorganic salts include, but are not limited to, sodium chloride, potassium chloride, calcium chloride, magnesium sulfate, and ammonium sulfate. Suitable aqueous solvents include, but are not limited to, aqueous buffer solutions, such as Tris-HCl buffer (pH 8.0). By treatment with proteolytic enzymes, glycoproteins are digested and peptides, which are digestion products, are released into an aqueous solvent.

[0038] When a carrier to which glycoproteins are adsorbed is incubated in an aqueous solvent containing a protease, the carrier is removed from the digestion reaction mixture to recover a peptide mixture as a digestion product, which is used as a peptide-containing sample for identifying proteins contained in a biological sample by mass spectrometry. Specifically, for example, a digestion reaction mixture containing the aqueous solvent containing the peptide mixture as a digestion product and the carrier is centrifuged to precipitate the carrier, and the aqueous solvent containing the peptide mixture as a digestion product is recovered as a supernatant.

[0039] To analyze the resulting peptide mixture using a mass spectrometer, it is typically subjected to a desalting process to remove low-molecular-weight substances such as salts, buffers, and chaotropic agents, yielding a purified peptide mixture. Desalting can be performed by applying the aqueous solvent containing the digested peptide mixture to a reversed-phase column. Examples of reversed-phase columns include, but are not limited to, columns in which hydrophobic groups such as octadecyl (C18), octyl (C8), butyl (C3), phenyl, and cyanopropyl groups are bonded to a solid support (e.g., silica gel), and styrene-divinylbenzene (SDB) copolymer columns. The reversed-phase column is preferably a column in which octadecyl (C18) groups are bonded to a solid support (e.g., silica gel) (referred to as a C18 column) or an SDB copolymer column. Commercially available reversed-phase columns can be used, including commercially available kits such as the EVOSEP ONE and SDB-STAGE chips.

[0040] When an aqueous solvent containing a peptide mixture is applied to a reverse-phase column, the peptide mixture is adsorbed onto the column, while low-molecular-weight substances such as salts, buffers, and chaotropic agents are eluted and removed without binding to the column. The peptide mixture adsorbed on the reverse-phase column can then be eluted to obtain a purified peptide mixture. The solvent used to elute the peptides is typically a mixture of water and an organic solvent such as acetonitrile, methanol, tetrahydrofuran, isopropanol, or acetone. To promote protonation, it is preferable to add an acid such as trifluoroacetic acid, formic acid, acetic acid, or hydrochloric acid to the water.

[0041] The eluted peptide mixture is dried to remove the solvent, which can be done using, for example, a centrifugal evaporator.

[0042] The resulting dry peptide mixture is redissolved in water. To promote protonation, an acid such as formic acid, trifluoroacetic acid, acetic acid, or hydrochloric acid may be added to the water. Alternatively, an organic solvent such as acetonitrile, methanol, tetrahydrofuran, isopropanol, or acetone may be added to the water. To promote dissolution of the dry peptide in water and to prevent adsorption of the dissolved peptide to the tube, a surfactant such as LMNG, DMNG, or DDM may be added to the water. For example, the dry peptide mixture is dissolved in water containing 0.1% (v / v) TFA and 0.01% (w / v) DMNG to obtain an aqueous solution of the peptide mixture.

[0043] Next, the resulting aqueous solution of the peptide mixture is subjected to mass spectrometry. Based on the mass spectrometry results, the peptides contained in the peptide mixture are identified, and the proteins in the biological sample from which they originate are identified based on the information on the identified peptides. Mass spectrometry of the peptide mixture is typically performed using a liquid chromatography-mass spectrometer (LC-MS), preferably a tandem mass spectrometer online coupled to liquid chromatography (LC-MS / MS). The measurement mode can be appropriately selected depending on the purpose of the measurement. For comprehensive identification of proteins contained in a biological sample using non-targeted proteomics, measurement modes such as data-dependent acquisition (DDA) or data-independent acquisition (DIA) are selected. For identification and quantification of specific proteins using targeted proteomics, measurement modes such as S / MRM (selected or multiple reaction monitoring) can be selected.

[0044] DDA selects multiple precursor ions with the strongest signal intensity and automatically acquires product ion spectra. From the obtained data, a file (peak list) is created in which the product ion spectra are linked to peaks in the MS spectrum, and a database search is then performed. In the database search, the peptide sequences obtained when all protein sequences registered in a given sequence database are enzymatically predicted (for example, trypsin cleaves at the C-terminus of K / R), and candidate peptide sequences with masses matching those of the precursor ions are narrowed down. Next, the m / z values ​​of the N-terminal ion (b-ion) and C-terminal ion (y-ion) produced when a single peptide bond in the candidate peptide sequence is cleaved are calculated for each peptide bond. The theoretical product ion spectrum obtained is then compared with the measured product ion spectrum to identify the specific peptide sequence.

[0045] In DIA, without selecting a precursor ion, mixed product ion spectra of all precursor ions that fit within the Q1 window of the set m / z width are acquired repeatedly while shifting the window. In DIA, chromatograms of fragment ions derived from the peptide sequence of interest are extracted using an MS / MS information library created from separately acquired DDA data, and the chromatographic peak of the target peptide is identified from the coelution pattern of the fragment ions.

[0046] In S / MRM, the combination of the m / z of the precursor ion of the target peptide identified by DDA and the m / z of the fragment generated by CID is specified in the measurement method, and the ions of the target peptide in the sample (fragment ion signals) are selectively monitored to obtain a chromatogram.

[0047] In DDA, when multiple analytes elute simultaneously and their abundances vary significantly, there is a relatively high risk that low-abundance analytes will not be detected in the initial MS spectrum, or that the mass spectrometer speed will be too slow for the complexity of the sample to acquire MS / MS spectra for all peaks detected in MS mode. In contrast, DIA collects all MS / MS spectra for all detectable analytes passing through each Q1 window, and the entire mass range is analyzed within the LC analysis time. This allows for the acquisition of complete MS and MS / MS spectra for virtually all detectable peaks in the sample, which is advantageous for the detection and identification of low-abundance analytes. In the preparation method of the present invention, affinity purification using a support immobilized with N-acetylglucosamine oligomers or fucose-binding lectin or its binding fragments effectively removes highly abundant proteins, such as albumin, from a biological sample while preserving trace proteins. This allows the preparation of a peptide-containing sample with a high content of peptides derived from trace proteins. Therefore, by subjecting a peptide-containing sample prepared by the preparation method of the present invention to LC-MS / MS analysis in the measurement mode of DIA, it becomes possible to detect and identify trace proteins contained in biological samples and peptides derived from them with higher sensitivity, thereby achieving more sensitive and in-depth proteome analysis.

[0048] Therefore, the preparation method of the present invention can also be considered as a method for identifying proteins contained in a biological sample by mass spectrometry, comprising the following steps: 1) contacting the biological sample with a support on which a lectin selected from the group consisting of N-acetylglucosamine oligomer-binding lectins and fucose-binding lectins, or a binding fragment thereof, has been immobilized, thereby recovering glycoproteins in the biological sample that have binding properties to the lectin; 2) digesting the glycoproteins in the recovered biological sample with a protease to obtain a peptide mixture as a digestion product; 3) subjecting the obtained peptide mixture to a desalting treatment to obtain a purified peptide mixture; 4) subjecting the purified peptide mixture to a mass spectrometer and identifying peptides contained in the peptide mixture based on the obtained mass spectrometry results; and 5) identifying the proteins in the biological sample from which the identified peptides are derived based on information about the identified peptides.

[0049] In one embodiment, in step 1, proteins bound to the lectin or its binding fragment are eluted with an acidic aqueous solution to recover glycoproteins capable of binding to the lectin in the biological sample. The eluate may be neutralized and subjected to digestion with a protease in step 2.

[0050] As described in detail in the Examples, by subjecting a peptide-containing sample prepared from human serum by the preparation method of the present invention to mass spectrometry, it was possible to identify the proteins listed in Tables 2-1 to 2-12, which could not be detected by conventional depletion of highly abundant proteins. Thus, in one embodiment, by subjecting a peptide-containing sample prepared from a human-derived biological sample (preferably human serum, plasma, or whole blood) by the preparation method of the present invention to mass spectrometry, at least one, for example, at least 5, 10, 20, 50, 100, 200, 300, 400, 500, or 1,000 proteins selected from the proteins listed in Tables 2-1 to 2-12 are identified in the biological sample.

[0051] All references cited herein, including publications, patent documents, and the like, are incorporated herein by reference to the same extent as if each was individually and specifically incorporated by reference and the contents thereof were specifically set forth in their entirety.

[0052] The present invention will be explained in more detail below with reference to examples, but the present invention is not limited thereto.

[0053] Example 1 1. Experimental Procedure Serum Three different lots of human serum pools were purchased from Kohjin Bio (Saitama, Japan). These three types of serum were pooled and used in the experiments. Normal mouse, rat, cow, rabbit, goat, sheep, and chicken sera were purchased from TK Craft (Gunma, Japan). Systemic lupus erythematosus (SLE) model mice were purchased from female (NZB x NZW) F1 hybrid mice that spontaneously developed a polygenic autoimmune disease, and presymptomatic control mice were purchased from Hooke Laboratories (Lawrence, MA, USA).

[0054] For the lectin enrichment method, glycoproteins in serum were automatically enriched with lectins using a Maelstrom 9610 (Taiwan Advanced Nanotech, Taoyuan, Taiwan). First, 950 μL of blocking buffer (TBST, 1% BSA, 0.2 mM CaCl2) was added to 50 μL of streptavidin (SA) bead suspension (CAT# 21152104010350; Cytiva, Marlborough, MA, USA), followed by the addition of 40 μL of 0.5 μg / μL biotinylated lectin (Table 1).

[0055]

[0056]

[0057] 20 μL of each mixture was added to the two-lectin mixture, and 13.3 μL was added to the three-lectin mixture. The mixture was mixed for 30 minutes, and the beads were washed with wash buffer (TBST and 0.2 mM CaCl2). Next, 50 μL of serum diluted in 150 μL of wash buffer was mixed with the beads for 60 minutes. The beads were then washed three times with 1000 μL of wash buffer and removed by mixing with 200 μL of elution buffer (100 mM Tris-HCl pH 8.0, 4% SDS, 20 mM NaCl) for 10 minutes.

[0058] For depletion of high-abundance proteins, human serum was treated with Top14 Abundant Protein Depletion Mini Spin Columns (Thermo Fisher Scientific, Waltham, MA, USA) according to the manufacturer's instructions, and mouse serum was treated with Proteome Purify 2 Mouse Serum Protein Immunodepletion Resin (R&D System, Minneapolis, MN, USA) according to the manufacturer's instructions.

[0059] Protein Digestion. Treated serum was treated with 20 mM tris(2-carboxyethyl)phosphine at 80°C for 10 minutes, alkylated with 35 mM iodoacetamide at room temperature for 30 minutes in the dark, and then subjected to SP3 cleanup and digestion using a Maelstrom 9610. Briefly, two types of Sera-Mag SpeedBead carboxylate-modified magnetic particles (hydrophilic particles, CAT# 45152105050250; hydrophobic particles, CAT# 65152105050250; Cytiva) were used. These beads were combined in a 1:1 (v / v) ratio, washed twice with distilled water, and reconstituted in distilled water to a concentration of 8 μg solids / μL. 20 μL of the reconstituted beads (SP3 beads) were added to the alkylated protein sample, followed by the addition of 99.5% ethyl alcohol to a final concentration of 75% (v / v) and vortexing for 15 minutes. The supernatant was discarded, and the pellet was washed twice with 80% ethyl alcohol. The beads were suspended in 80 μL of 50 mM Tris-HCl (pH 8.0), 10 mM CaCl2 (containing 0.02% LMNG) (20). Then, 500 ng of trypsin / Lys-C Mix (CAT# V5072, Promega, Madison, WI, USA) was added and gently mixed overnight at 37°C to digest the proteins. The digested sample was acidified with 20 μL of 5% trifluoroacetic acid (TFA) and then sonicated at high power for 5 min at room temperature using a Bioruptor II (Cosmo Bio, Tokyo, Japan). The sample was desalted using an SDB-STAGE chip (GL Science, Tokyo, Japan). The SDB-STAGE chip was washed with 25 μL of 80% ACN in 0.1% TFA and equilibrated with 50 μL of 3% ACN in 0.1% TFA. The sample was then loaded onto a chip, washed with 80 μL of 3% ACN in 0.1% TFA, and eluted with 50 μL of 36% ACN in 0.1% TFA. The eluate was dried using a miVac Duo concentrator. The dried sample was redissolved in 0.1% TFA containing 0.01% DMNG.The reconstituted samples were analyzed for peptide concentration using a Lunatic instrument (Unchained Labs, Pleasanton, CA, USA) and transferred to LC vials (Thermo Fisher Scientific).

[0060] LC-MS / MS and DIA. Reconstituted peptides were directly injected onto a 75 μm × 12 cm nanoLC column (Nikkyo Technos Co., Ltd., Tokyo, Japan) at 50 °C and separated using an UltiMate 3000 RSLC nano LC system with a 60-min gradient (A = 0.1% FA in water, B = 0.1% FA in 80% ACN) at a flow rate of 200 nl / min. The gradient consisted of 6% B at 0 min, 36% B at 50 min, 70% B at 57 min, and 70% B at 60 min. Peptides eluted from the column were analyzed using a Q-Exactive HF-X with an InSpIon system (J Proteome Res 22, 1564–1569, 2023). The precursor range for DIA was based on previous parameters (J Proteome Res 21, 1418–1427, 2022). MS1 spectra were collected in the range of 495–745 m / z at a resolution of 15,000 and an autogain control target of 3 × 10 6 MS2 spectra were collected above 200 m / z at a resolution of 30,000, with an autogain control target of 3 × 10 6 The MS2 isolation width was set to 4 m / z, and the window pattern from 500 to 740 m / z was used with the optimized window arrangement in Scaffold DIA (Proteome Software, Inc., Portland, OR, USA).

[0061] For deep proteomic analysis, reconstituted peptides were directly injected onto a 75 μm × 30 cm nanoLC column (CoAnn Technologies, Richland, WA, USA) at 60 °C and separated using an UltiMate 3000 RSLC nano LC system with a 120-min gradient (A = 0.1% FA in water, B = 0.1% FA in 80% ACN) at a flow rate of 200 nl / min, consisting of 3% B at 0 min, 33% B at 108 min, 65% B at 114 min, and 65% B at 120 min. Peptides eluted from the column were analyzed on an Orbitrap Exploris 480 with an InSpIon system. MS1 spectra were collected in the m / z range of 495–745 at a resolution of 15,000, with an autogain control target of 3 × 10. 6 The maximum injection time was set to "auto." MS2 spectra were collected in the m / z range of 200–1,800 at a resolution of 45,000, with the autogain control target set to 3 × 10. 6 The maximum injection time was set to "auto," and the stepwise normalized collision energies were set to 22, 26, and 30%. The MS2 isolation width was set to 4 m / z, and the window pattern from 500 to 740 m / z used the window placement optimized by Scaffold DIA.

[0062] Data Analysis: DIA-MS files were searched against in silico human or mouse spectral libraries using DIA-NN (version 1.8.1, https: / / github.com / vdemichev / DiaNN) (Nat Methods 17, 41-44, 2020). First, spectral libraries were generated using DIA-NN from the UniProt database of human or mouse protein sequences.

[0063] The parameters for spectral library creation were as follows: digestion enzyme, trypsin; miss-cleavage, 1; peptide length range, 7-35; precursor charge range, 2-4; precursor m / z range, 495-745; fragment ion m / z range, 200-1800; and "FASTA digest for library-free search / library generation," "deep learning-based spectrum, RT and IM prediction," "N-terminal M-cleavage," and "C-carbamidomethylation" were enabled.

[0064] DIA-NN search parameters were as follows: mass accuracy, 10 ppm; MS1 accuracy, 10 ppm; protein prediction, gene; neural network classification, single-pass mode; quantitation strategy, robust LC (high accuracy); cross-run normalization, off; and "unrelated runs," "use isotopes," "heuristic protein prediction," and "no shared spectra" were enabled.

[0065] MBR was turned off. The threshold for protein identification was set to ≤1% for both precursor and protein FDR. Protein quantification values ​​were calculated by summing the quantification values ​​of unique peptides calculated by DIA-NN. Protein quantification data were transformed to log2 (protein intensity) and filtered so that for each protein, at least one group contained at least 70% valid values. Remaining missing values ​​were imputed with random numbers drawn from a normal distribution (width, 0.3; downshift, 1.8) in Perseus v1.6.15.0 (24). The threshold for altered proteins was a >2-fold change between two groups and a difference of p < 0.05 (Welch's test). GO enrichment analysis was performed using DAVID (https: / / david.ncifcrf.gov / tools.jsp). Enrichment analysis of disease ontologies (based on the HumanPSD database) was performed using the geneXplain platform (GeneXplain GmbH, Wolfenbüttel, Germany).

[0066] 2. Results and Discussion: Enrichment of Trace Proteins from Serum with Lectins We first established an automated method for the enrichment of trace proteins in serum using lectins (Figure 1). Using a Maelstrom 9610 automated magnetic bead system, biotinylated lectins were reacted with streptavidin magnetic beads, and then the beads were reacted with serum to enrich lectin-binding proteins. Furthermore, the SP3 method was also automated using the Maelstrom 9610. This system allowed for the automated processing of 96 samples simultaneously, enabling the examination of various conditions. The method established with this system is also suitable for the processing of multiple samples, such as clinical specimens. Due to the limited number of commercially available lectin-immobilized beads, we did not use commercially available lectin-immobilized beads in this study. Instead, we conjugated widely available biotinylated lectins to streptavidin magnetic beads. Although the enrichment in this study focused on lectins, the same system can also be used for automated affinity purification of other biotinylated substances.

[0067] We used this automated method to attempt to recover trace proteins in serum using 37 lectins (Table 1). In this study, commercially available pooled serum was used. Proteins recovered using the lectins were trypsin-digested and subjected to LC-MS / MS (Figure 2). Serum depleted of TOP14 proteins (TOP14D) was used as a control. The lectins that identified more proteins than TOP14D were STL, LEL, WGA, UEA-I, DSL, AAL, AOL, sWGA, and MAL-I. We found that these lectins could enrich trace proteins beyond those typically used with TOP14D. These lectins recognized N-acetylglucosamine or fucose. Among these, STL and LEL, which had the highest number of protein identifications (IDs), strongly recognized N-acetylglucosamine oligomers, and targeting them was found to be important for recovering trace proteins in serum. Therefore, we investigated whether combining STL, which had the highest number of identified proteins, with other lectins would further improve the number of observed proteins (Figure 3). The combination of STL and LEL resulted in a higher number of protein identifications than STL alone, making it the most effective combination. We also examined the effect of adding another lectin to STL and LEL. As a result, STL / LEL / DEL showed the highest number of protein IDs, but the number of protein identifications was almost the same as STL / LEL, suggesting that there was little benefit to combining three lectins. Based on these results, we concluded that STL / LEL is the optimal combination for the enrichment of trace proteins in serum using lectins. Next, we investigated whether this STL / LEL enrichment method could be applied to animal species other than humans (Figure 4). More than 1,500 proteins were identified in all animal species tested, confirming that it can enrich trace proteins regardless of the animal species. Because few antibody-based depletion columns for animal species other than humans and mice are commercially available, the adaptability of this method to various animal species is a major advantage.

[0068] Deep proteomics using the STL / LEL enrichment method. The 60-minute activity gradient used in the study to establish the STL / LEL method was extended to 120 minutes, modifying the setup to detect trace serum proteins. Figure 5A shows a comparison of crude serum, TOP14D-treated serum, and STL / LEL-enriched serum using this 120-minute gradient. 1,339, 2,049, and 3,122 proteins were identified in crude serum, TOP14D-treated serum, and STL / LEL-enriched serum, respectively. Proteins detected only in STL / LEL-enriched serum but not in crude serum or TOP14D-treated serum are listed in Tables 2-1 through 2-12. GO analysis of proteins detected only in STL / LEL-enriched serum indicated that membrane proteins were recovered. Because membrane proteins are known to be frequently glycosylated (Electrophoresis 37, 1407-1419, 2016), it was speculated that STL / LEL enriched membrane proteins released into serum. Interleukin receptors, including IL1R1, IL2RA, IL2RB, IL2RG, IL3RA, IL10RB, IL15RA, IL17RA, IL17RC, and IL27RA, were also found to be enriched in STL / LEL.

[0069] Next, we compared serum samples from SLE model mice and control mice using the STL / LEL method and a 120-minute activity gradient LC-MS / MS method. To evaluate the STL / LEL method, the same samples were also processed using the top 2 high-abundance protein depletion (TOP2D) method. Comparing the protein profiles between each group, the difference between the STL / LEL and TOP2D methods was greater than the difference between the mouse groups, confirming that the STL / LEL method is distinct from conventional methods (Figure 6A). The STL / LEL method identified 1,332 proteins that were altered in serum samples from SLE model mice, while the TOP2D method identified 1,005. The STL / LEL method identified more altered proteins than the TOP2D method (Figure 6B). Furthermore, numerous non-overlapping altered proteins were identified by each method (Figure 6C), demonstrating its potential for biomarker discovery from a different angle than the TOP2D method. Furthermore, disease enrichment analysis was performed on proteins elevated in SLE (Supplementary Table 3). Figure 6D shows the results for diseases associated with SLE and its complications. For all diseases shown, the STL / LEL method classified more proteins and was more likely to observe the true nature of the disease than the TOP2D method. These results demonstrate that the STL / LEL enrichment method is useful for biomarker discovery.

[0070] 3. Conclusions To broaden the scope of biomarker discovery in serum, we established an automated method for enriching trace proteins in serum by combining STL and LEL. The STL / LEL method was able to observe approximately 1.5 times the number of proteins in serum compared with the TOP14D method. Furthermore, when comparing the serum of SLE model mice and control mice using the STL / LEL enrichment method, over 1,300 differential proteins were identified. By collecting N-acetylglucosamine oligomer-attached proteins using STL and LEL, we were able to observe variations different from those observed with conventional depletion methods, demonstrating that this method is an effective method for biomarker discovery.

[0071]

[0072]

[0073]

[0074]

[0075]

[0076]

[0077]

[0078]

[0079]

[0080]

[0081]

[0082]

[0083] Example 2 In Example 1, lectin was denatured with SDS as shown in Figure 1, and bound proteins were eluted from the lectin column. However, because SDS inhibits enzymatic digestion, SDS was removed from the eluted proteins using the SP3 method after elution. In this example, an acidic aqueous solution (0.2% TFA) was used as the elution solution. After acid-denaturing the lectin column to elute bound proteins, the resulting eluate was neutralized with a buffer to return the pH to an optimal pH for enzymatic digestion, enabling proteome analysis without the SP3 step. Figure 7 shows a schematic diagram of this improved method for concentrating trace serum proteins using a biotinylated lectin column. To prevent the eluted proteins from being contaminated with surfactants that are not suitable for pretreatment in proteome analysis, the lectin column must be washed with an appropriate buffer before the elution solution is applied. In this example, TBST containing 0.2 mM CaCl2 was used as the wash buffer, and a 0.2% TFA aqueous solution was used as the acidic aqueous solution for elution. This acidic aqueous solution for elution contained 0.05% LMNG depending on the test conditions. LMNG is a surfactant that suppresses nonspecific adsorption of proteins to columns and tubes without inhibiting enzymatic digestion. The eluted protein solution was neutralized by adding the same volume of 500 mM Tris-HCl aqueous solution (pH 8.0) and then directly digested with trypsin / Lys-C without SP3 treatment. Other test conditions were the same as in Example 1, and the number of proteins identified by the improved method was compared with that of Example 1.

[0084] The results are shown in Figure 8. When the elution solution was 0.2% TFA aqueous solution without SP3 treatment, the number of identified proteins was only approximately 8% lower than in Example 1, where SDS elution was followed by SP3 treatment (SDS-SP3). Furthermore, when LMNG was added to the 0.2% TFA aqueous solution, more proteins were detected than with the SDS-SP3 method. LMNG can be easily removed using a reverse-phase spin column, eliminating the need for additional pretreatment. The 0.2% TFA / 0.05% LMNG elution condition reduced the number of steps compared to the SDS-SP3 method, and further increased the number of identified proteins by recovering and digesting the eluate without the loss associated with SP3 treatment.

[0085] Next, we investigated the difference in the number of identified proteins eluted depending on the type of acid. The results are shown in Figure 9. Although there was a tendency for many proteins to be detected with TFA and hydrochloric acid, which are highly acidic, it was confirmed that protein elution from the lectin column was enhanced with any acid.

[0086] These results suggest that elution of lectin-binding proteins with an acidic aqueous solution containing LMNG enables simple and in-depth proteome analysis.

[0087] According to the present invention, it is possible to efficiently remove high-abundance proteins that hinder the in-depth of LC-MS analysis from biological samples such as serum and plasma, and to preferentially concentrate trace proteins that are useful as biomarkers. Samples prepared by the methods of the present invention contain highly concentrated trace proteins, are free of contaminating high-abundance proteins, and have a narrow dynamic range of protein concentrations. Therefore, when used in LC-MS analysis, these samples can improve the depth of proteome analysis and enable the detection of a large number of proteins.

[0088] This application is based on patent application No. 2023-211819 filed in Japan (filing date: December 15, 2023), the contents of which are incorporated in full herein.

Claims

1. A method for preparing a protein- or peptide-containing sample for identifying proteins contained in a biological sample by mass spectrometry, the method comprising contacting the biological sample with a carrier on which a lectin selected from the group consisting of N-acetylglucosamine oligomer-binding lectins and fucose-binding lectins, or a binding fragment thereof, is immobilized, and recovering glycoproteins from the biological sample that have a binding affinity to the lectin.

2. The preparation method according to claim 1, wherein the lectin or binding fragment thereof is any N-acetylglucosamine oligomer-binding lectin or binding fragment thereof selected from the group consisting of potato lectin, tomato lectin, Datura stramonium lectin, and wheat germ agglutinin.

3. The preparation method according to claim 1, wherein the lectin or its binding fragment is any fucose-binding lectin or its binding fragment selected from the group consisting of Ustilago sativa agglutinin-1, Acanthus obtusifolia lectin, and Aspergillus oryzae lectin.

4. The preparation method according to claim 1, wherein the lectin or binding fragment thereof is a combination comprising: 1) potato lectin or a binding fragment thereof, and 2) any lectin or binding fragment thereof selected from the group consisting of tomato lectin, Datura stramonium lectin, wheat germ agglutinin, and Anemone lectin.

5. The method of claim 4, wherein the lectin or binding fragment thereof is a combination comprising: 1) potato lectin or binding fragment thereof, and 2) tomato lectin or binding fragment thereof.

6. The preparation method according to claim 1, further comprising subjecting the glycoprotein in the collected biological sample to a reductive alkylation treatment.

7. The preparation method according to claim 1, further comprising digesting glycoproteins in the collected biological sample with a proteolytic enzyme to obtain a peptide mixture as a digestion product.

8. The preparation method according to claim 7, further comprising subjecting the obtained peptide mixture to a desalting column to obtain a purified peptide mixture.

9. The preparation method according to any one of claims 1 to 8, wherein the biological sample is any one selected from the group consisting of serum, plasma, whole blood, cerebrospinal fluid, ascites, synovial fluid, and lymphatic fluid.

10. A preparation method according to any one of claims 1 to 9, comprising recovering glycoproteins capable of binding to the lectin in a biological sample by eluting the proteins bound to the lectin or its binding fragment with an acidic aqueous solution.

11. The preparation method according to claim 10, wherein the acidic aqueous solution contains a sugar-based nonionic surfactant.

12. The preparation method according to claim 11, wherein the sugar-based nonionic surfactant is LMNG.

13. The preparation method according to any one of claims 10 to 12, further comprising neutralizing the eluate and treating it with a proteolytic enzyme to digest glycoproteins in the biological sample recovered in the eluate with the proteolytic enzyme, thereby obtaining a peptide mixture as a digestion product.

14. A method for identifying proteins contained in a biological sample by mass spectrometry, comprising the steps of: 1) contacting the biological sample with a support having a lectin or a binding fragment thereof selected from the group consisting of N-acetylglucosamine oligomer-binding lectins and fucose-binding lectins immobilized thereon, thereby recovering glycoproteins in the biological sample that have a binding affinity to the lectin; 2) digesting the glycoproteins in the recovered biological sample with a protease to obtain a peptide mixture as a digestion product; 3) subjecting the obtained peptide mixture to a desalting treatment to obtain a purified peptide mixture; 4) subjecting the purified peptide mixture to a mass spectrometer, and identifying peptides contained in the peptide mixture based on the obtained mass spectrometry results; and 5) identifying proteins in the biological sample from which the identified peptides are derived based on information on the identified peptides.

15. The method according to claim 14, wherein the lectin or binding fragment thereof is any N-acetylglucosamine oligomer-binding lectin or binding fragment thereof selected from the group consisting of potato lectin, tomato lectin, Datura stramonium lectin, and wheat germ agglutinin.

16. The method according to claim 14, wherein the lectin or binding fragment thereof is any fucose-binding lectin or binding fragment thereof selected from the group consisting of Ustilago sativa agglutinin-1, Acanthus obtusifolia lectin, and Aspergillus oryzae lectin.

17. The method of claim 14, wherein the lectin or binding fragment thereof is a combination comprising: 1) potato lectin or a binding fragment thereof, and 2) any lectin or binding fragment thereof selected from the group consisting of tomato lectin, Datura stramonium lectin, wheat germ agglutinin, and Anemone lectin.

18. The method of claim 17, wherein the lectin or binding fragment thereof is a combination comprising: 1) potato lectin or binding fragment thereof, and 2) tomato lectin or binding fragment thereof.

19. The method of claim 14, further comprising subjecting the glycoproteins in the biological sample collected in step 1 to a reductive alkylation treatment.

20. The method of any one of claims 14 to 19, wherein the biological sample is any one selected from the group consisting of serum, plasma, whole blood, cerebrospinal fluid, peritoneal fluid, synovial fluid, and lymphatic fluid.

21. A preparation method according to any one of claims 14 to 20, wherein in step 1, a glycoprotein capable of binding to the lectin in the biological sample is recovered by eluting the protein bound to the lectin or its binding fragment with an acidic aqueous solution.

22. The preparation method according to claim 21, wherein the acidic aqueous solution contains a sugar-based nonionic surfactant.

23. The preparation method according to claim 22, wherein the sugar-based nonionic surfactant is LMNG.

24. The preparation method according to any one of claims 21 to 23, wherein the eluate is neutralized and subjected to proteolytic enzyme digestion in step 2.

Citation Information

Patent Citations

  • Method for preparing peptide-containing sample

    WO2024127796A1

  • High-affinity receptor for Helicobacter pylori and its use

    JP2006511497A

  • Method for detecting advance degree of diabetic nephropathy, kit for diagnosing advance degree of diabetic nephropathy, substance becoming index of advance degree of diabetic nephropathy, and method for sorting the substance

    JP2010256132A

  • Analytical methods for glycoproteins

    JP2021519923A

  • Removal of serum albumin from human serum or plasma

    US20050272643A1