A method for identifying and verifying a protein sialic acid
Through the mass spectrometry method of targeted screening at the molecular level of monosaccharides, monosaccharide sequences and glycoproteomes, the accuracy and efficiency of sialic acid identification are solved, and efficient and accurate sialic acid identification and verification are achieved.
Patent Information
- Application Number
- CN202310097993.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-10
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2043-02-10
AI Technical Summary
The prior art is difficult to efficiently and accurately identify and verify the presence of sialic acid in proteins in high resolution mass spectrometry, especially at the incomplete N-glycopeptide level.
Through targeted screening at the three molecular levels of monosaccharides, monosaccharide sequences and glycoproteomes, combined with mass spectrometry technology, the identification and verification of sialic acid is achieved, including precursor ion isotope profile matching, fragment ion isotope profile alignment and random matching probability control, complete N-glycopeptides containing sialic acid were screened out.
It greatly improves the accuracy and efficiency of sialic acid in complete N-glycopeptide analysis, shortens the search time, and is suitable for high accuracy qualitative and quantitative analysis of large cohort samples.
Smart Images

Figure CN115980168B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of accurate analysis of protein structures, and particularly relates to a method for identifying and verifying protein sialic acid. Background Art
[0002] Glycosylation is one of the most abundant post-translational modifications of proteins in the human body. Accurate analysis of the glycoprotein structure is a prerequisite for understanding its biochemical properties and physiological functions. As a monosaccharide widely present at the end of the branched chain of glycoprotein modification, sialic acid plays a key role in various biological recognition processes in the human body.
[0003] The widespread use of high-resolution mass spectrometry enables the analysis of sialic acid-containing N-linked glycoproteins to obtain isotope resolution at the level of intact N-glycopeptides and accurate measurement of the isotope profiles of precursor ions in the first-order mass spectrum and fragment ions in the second-order mass spectrum. Based on this progress and characteristics, the present invention provides a method for analyzing intact N-glycopeptides containing sialic acid. Summary of the Invention
[0004] To solve the above technical problems, the purpose of the present invention is to provide a method for identifying and verifying protein sialic acid. This method realizes the selective search for intact N-glycopeptides containing sialic acid based on the targeted screening of the structural fingerprints of sialic acid at three molecular levels: monosaccharide, monosaccharide sequence, and glycoproteome, saves the search time, and greatly improves the accuracy and efficiency of the analysis of intact N-glycopeptides containing sialic acid.
[0005] To achieve the above technical purposes and achieve the above technical effects, the present invention is realized through the following technical solutions:
[0006] A method for identifying and verifying protein sialic acid, comprising the following steps:
[0007] S1. Prepare an intact N-glycopeptide sample from the biological sample to be analyzed;
[0008] S2. After the intact N-glycopeptide sample is subjected to electrospray ionization, precursor ions are obtained. The precursor ions are sent into a mass spectrometer to obtain an experimental first-order mass spectrum containing the isotope profile fingerprint of the precursor ions;
[0009] S3. Match the isotope profile fingerprint of the precursor ions with the corresponding targeted positive theoretical database, and screen out the primary candidate intact N-glycopeptide IDs that meet the matching conditions and have the same molecular composition;
[0010] S4. Send the precursor ions into an ion trap for gas-phase dissociation to obtain fragment ions. The fragment ions are sent into a mass spectrometer to obtain a second-order mass spectrum containing the isotope profile fingerprint of the fragment ions;
[0011] S5. Perform pairwise isotope profile fingerprint comparison and matching on the molecular composition fingerprints of the experimental and theoretical fragment ions of the primary candidate complete N-glycopeptide IDs;
[0012] S6. Targetedly screen for sialic acid characteristic oxonium ions and experimental structure diagnostic fragment ions of the monosaccharide sequence from the primary candidate complete N-glycopeptide IDs with the same molecular composition fingerprints, and classify the screened primary candidate complete N-glycopeptide IDs as candidate complete N-glycopeptide IDs;
[0013] S7. Perform random match probability scoring on the candidate complete N-glycopeptide IDs, and finally classify the IDs that meet the pre-set complete N-glycopeptide spectrum matching conditions as targeted GPSMs;
[0014] S8. Obtain decoy GPSMs in the decoy library according to the steps of S3 - S7;
[0015] S9. Combine the targeted GPSMs and decoy GPSMs, sort them in ascending order of P score, select a threshold Pscore such that the false positive rate is not greater than 1%, and remove duplicates from the targeted GPSMs with P score not greater than the threshold Pscore to obtain complete N-glycopeptide IDs containing sialic acid.
[0016] Furthermore, the biological sample to be analyzed is a sample containing sialylated glycoprotein.
[0017] Furthermore, in step S2, before performing electrospray ionization and mass spectrometry analysis, perform liquid phase separation on the complete N-glycopeptide sample first.
[0018] Furthermore, the experimental molecular composition fingerprint of the precursor ion is measured in the first-order mass spectrum, and the experimental molecular composition fingerprint of the fragment ion is measured in the second-order mass spectrum.
[0019] Furthermore, the theoretical molecular composition fingerprint is generated through the following steps:
[0020] Calculate the molecular formula of each molecule according to the theoretical molecular library of the system under study;
[0021] Refer to the standard element list and calculate the corresponding molecular composition fingerprint according to the types and quantities of elements in the molecular formula.
[0022] Furthermore, the matching in steps S3 and S5 refers to the one-to-one comparison of the m / z value and relative peak intensity value of each isotope peak in the experimental molecular composition fingerprint with the corresponding theoretical values.
[0023] Furthermore, the matching criteria in steps S3 and S5 are controlled by the isotope peak intensity cut-off value, the maximum allowable error of the isotope peak mass-to-charge ratio, and the maximum allowable error of the isotope peak intensity.
[0024] Further, the duplicate removal in step S9 is carried out according to the standards of the amino acid sequence of the polypeptide backbone, modification, glycosylation sites, and the structure of N-linked sugar sequences.
[0025] The beneficial effects of the present invention are as follows:
[0026] The identification and verification method of the present invention identifies and confirms sialic acid based on the molecular composition and structural fingerprint of mass spectrometry at three molecular levels of intact N-glycopeptides, monosaccharide sequences, and monosaccharides, including: 1) at the level of intact N-glycopeptides, identifying the monosaccharide composition of intact N-glycopeptides containing sialic acid in the first-order mass spectrometry based on the isotope profile of precursor ions; 2) at the molecular level of monosaccharide sequences, performing false positive control and identification at the spectral level on the sequence structure based on fragment ions in the second-order mass spectrometry and targeted-decoy library search, and confirming the sequence structure of diagnostic fragment ions based on the sequence structure; 3) at the molecular level of monosaccharides, confirming the characteristic oxonium ions of sialic acid in the second-order mass spectrometry. Through the identification and verification at the above three molecular levels, the accurate analysis of protein sialic acid based on mass spectrometry is maximally realized.
[0027] The method of the present invention realizes the selective search of effective tandem mass spectrometry based on the targeted screening of molecular composition and structural fingerprint fragment ions. By searching and filtering the molecular composition and structural fingerprint of sialic acid before the comprehensive analysis and false positive control of tandem mass spectrometry, it saves the search for sialic acid-containing candidate molecules that do not observe the corresponding molecular composition and structural fingerprint in the experiment, greatly shortens the search time, and greatly improves the bioinformatics identification throughput. It is applicable to high-accuracy qualitative identification and quantitative analysis based on molecular composition and structural fingerprint for large cohort sample mass spectrometry analysis.
[0028] The method of the present invention realizes the selective search of intact N-glycopeptides containing sialic acid based on the targeted screening of the structural fingerprint of sialic acid at three molecular levels of monosaccharides, monosaccharide sequences, and glycoproteomes, thereby saving the search time used for invalid monosaccharide compositions in the existing "search first - screen later" process, and greatly improving the accuracy and efficiency of the analysis of intact N-glycopeptides containing sialic acid. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0030] Figure 1 It is a schematic flow chart of the method of the present invention.
[0031] Figure 2 It is the first-order mass spectrum containing the precursor ion m / z 1182.32373 (z = 4).
[0032] Figure 3 It is the comparison graph of the experimental and theoretical isotope profile fingerprints of the precursor ion m / z 1182.32373 (z = 4).
[0033] Figure 4 It is the annotated second-order mass spectrum containing the complete N-glycopeptide VDKDLQSLEDILHQVENK_01Y41Y41M(31M41Y – 41L)61M61Y41L32S with matching fragment ions of the polypeptide backbone and N-linked glycan moiety.
[0034] Figure 5 It is the graphical dissociation map of the polypeptide backbone with annotated matching fragment ions.
[0035] Figure 6 It is the graphical dissociation map of the N-linked glycan moiety with annotated matching fragment ions.
[0036] Figure 7 It is the identification and verification graph of the oxonium ion m / z 274 after the neutral loss of one molecule of water from the sialic acid monosaccharide based on the isotope profile fingerprint comparison.
[0037] Figure 8 It is the identification and verification graph of the oxonium ion m / z 292 of the sialic acid monosaccharide based on the isotope profile fingerprint comparison. Detailed implementation manners
[0038] The technical solutions in the present invention will be clearly and completely described below with reference to specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0039] As Figure 1 shown, specifically, a method for the identification and verification of protein sialic acid includes the following steps:
[0040] S1. Prepare a complete N-glycopeptide sample from the biological sample to be analyzed according to the existing method; wherein, the biological sample to be analyzed is a sample containing sialylated glycoprotein.
[0041] S2. First, separate the complete N-glycopeptide sample by high performance liquid chromatography, then obtain precursor ions after electrospray ionization of the complete N-glycopeptide sample, and send the precursor ions into a mass spectrometer to obtain an experimental first-order mass spectrum containing the isotope profile of the precursor ions.
[0042] S3. Match the isotope profile of the precursor ion with the corresponding targeted forward theoretical database, and screen out multiple primary candidate intact N-glycopeptide IDs with the same molecular composition that meet the matching conditions.
[0043] S4. Send the precursor ion into the ion trap for gas-phase dissociation to obtain fragment ions, and send the fragment ions into the mass spectrometer to obtain a second-order mass spectrum containing the isotope profile of the fragment ions.
[0044] S5. Perform one-by-one isotope profile fingerprint comparison and matching on the molecular composition fingerprints of the experimental and theoretical fragment ions of the primary candidate intact N-glycopeptide IDs.
[0045] The experimental molecular composition fingerprint of the precursor ion is measured in the first-order mass spectrum, and the experimental molecular composition fingerprint of the fragment ion is measured in the second-order mass spectrum.
[0046] The theoretical molecular composition fingerprint is generated through the following steps:
[0047] Calculate the molecular formula of each molecule according to the theoretical molecular library of the system under study;
[0048] Refer to the standard element list and calculate the corresponding molecular composition fingerprint according to the types and quantities of elements in the molecular formula.
[0049] The matching in steps S3 and S5 refers to the one-by-one comparison of the m / z value and relative peak intensity value of each isotope peak in the experimental molecular composition fingerprint with the corresponding theoretical values.
[0050] The matching criteria in steps S3 and S5 are controlled by the isotope peak intensity cut-off value, the maximum allowable error of the isotope peak mass-to-charge ratio, and the maximum allowable error of the isotope peak intensity.
[0051] S6. Targetedly screen out the sialic acid characteristic oxonium ions and monosaccharide sequence experimental structure diagnostic fragment ions from the primary candidate intact N-glycopeptide IDs with the same molecular composition, and classify the screened primary candidate intact N-glycopeptide IDs as candidate intact N-glycopeptide IDs.
[0052] S7. Perform a random matching probability score (i.e., calculate the P score) on the candidate intact N-glycopeptide IDs, and finally classify the candidate intact N-glycopeptide IDs that meet the pre-set intact N-glycopeptide spectrum matches (GPSMs) conditions (such as the number of fragment ions matching the polypeptide backbone is not less than 5, and the number of fragment ions matching the N-linked sugar part is not less than 1) as targeted GPSMs.
[0053] S8. Obtain decoy GPSMs in the decoy library (reverse library or random library) according to the steps of S3 - S7.
[0054] S9. Combine the targeted GPSMs and decoy GPSMs, sort them in ascending order of P score, and select a threshold Pscore such that the false positive rate is no greater than 1% (the calculation method is 2 times the number of decoy GPSMs with P score not greater than this threshold divided by the total number of targeted and decoy GPSMs). Dedup the targeted GPSMs with P score not greater than the threshold P score according to the standards of the polypeptide backbone amino acid sequence, modification, glycosylation site, and N-linked sugar sequence structure to obtain intact sialic acid-containing N-glycopeptide IDs.
[0055] The following describes the method for identifying and verifying protein sialic acid of the present invention with reference to an embodiment.
[0056] S1. Prepare an intact N-glycopeptide sample from human liver cancer tissue by the existing method;
[0057] S2. After the intact N-glycopeptide sample is separated by high performance liquid chromatography and electrospray ionization, positively charged precursor ions are obtained. The precursor ions are sent into a mass spectrometer to obtain an experimental first-order mass spectrum containing the isotope profile fingerprint of the precursor ions, as Figure 2 shown;
[0058] S3. Match the isotope profile fingerprint of the precursor ions with the corresponding targeted positive theoretical database, and screen out the primary candidate intact N-glycopeptide IDs with the same molecular composition that meet the matching conditions: VDKDLQSLEDILHQVENK_N4H5F0S1, as Figure 3 shown, where N represents N-acetylglucosamine, H represents hexose (including mannose M, glucose G, and galactose L), F represents fucose, and S represents sialic acid; the monosaccharide composition N4H5F0S1 corresponds to 4 monosaccharide sequence structures, as shown in Table 1;
[0059] Table 1
[0060] 01Y41Y41M(31M)61M(21Y41L32S)61Y41L 01Y41Y41M(31M41Y41L)61M61Y41L32S 01Y41Y41M(31M(21Y41L32S)41Y)61M61M 01Y41Y41M(31M41Y41L32S)(41Y)61M61M
[0061] In Table 1, Y represents N-acetylglucosamine, M represents mannose, L represents galactose, and S represents sialic acid.
[0062] S4. Send the precursor ions into an ion trap for gas-phase dissociation to obtain fragment ions, and send the fragment ions into a mass spectrometer to obtain a second-order mass spectrum containing the isotope profile fingerprint of the fragment ions;
[0063] S5. Perform pairwise isotope profile fingerprint comparison on the molecular composition fingerprints of the experimental and theoretical fragment ions of the primary candidate intact N-glycopeptide IDs, as Figure 4 , Figure 5 and Figure 6as shown;
[0064] S6. Targetedly screen experimental structure diagnostic fragment ions containing sialic acid monosaccharide sequence ( Figure 4 the ions in the thick-line rectangular box in, including 0,3 AI5-4+, 0,3 AI5-2+, CII3-1+, 3,5 AI5-1+) and characteristic oxonium ions ( Figure 7 , Figure 8 ), and classify the screened primary candidate intact N-glycopeptide IDs as candidate intact N-glycopeptide IDs;
[0065] S7. Perform a random matching probability score (i.e., calculate the P score) on the candidate intact N-glycopeptide IDs, and finally classify those that meet the pre-set intact N-glycopeptide spectrum matches (GPSMs) conditions (such as the number of fragment ions matching the polypeptide backbone is not less than 5, and the number of fragment ions matching the N-linked sugar part is not less than 1) as targeted GPSMs;
[0066] S8. Follow the steps of S3-S7 in the decoy library (reverse library or random library) to obtain decoy GPSMs;
[0067] S9. Combine the targeted and decoy GPSMs, sort them in ascending order of the P score, select a threshold P score such that the false positive rate is not greater than 1% (the calculation method is 2 times the number of decoy GPSMs with a P score not greater than this threshold divided by the total number of targeted and decoy GPSMs), and remove duplicates from the targeted GPSMs with a P score not greater than the threshold P score according to the standards of the polypeptide backbone amino acid sequence and modification, glycosylation site, and N-linked sugar sequence structure to obtain the final intact N-glycopeptide IDs.
[0068] In summary, the method of the present invention realizes the selective search for intact N-glycopeptides containing sialic acid based on the targeted screening of the molecular composition and structural fingerprints of sialic acid at the three molecular levels of monosaccharide, monosaccharide sequence, and glycoproteome, thereby saving the search time for invalid monosaccharide compositions in the existing "search first - screen later" process, and greatly improving the accuracy and efficiency of the analysis of intact N-glycopeptides containing sialic acid.
[0069] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
[0070] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for identifying and verifying protein sialic acid, characterized in that, It includes the following steps: S1. Prepare a complete N-glycopeptide sample from the biological sample to be analyzed; S2. After the complete N-glycopeptide sample is subjected to electrospray ionization, precursor ions are obtained. The precursor ions are sent into a mass spectrometer to obtain an experimental first-order mass spectrum containing the isotope profile fingerprint of the precursor ions; S3. Match the isotope profile fingerprint of the precursor ions with the corresponding targeted forward theoretical database, and screen out the primary candidate complete N-glycopeptide IDs with the same molecular composition that meet the matching conditions; S4. Send the precursor ions into an ion trap for gas-phase dissociation to obtain fragment ions. The fragment ions are sent into a mass spectrometer to obtain a second-order mass spectrum containing the isotope profile fingerprint of the fragment ions; S5. Perform one-by-one isotope profile fingerprint comparison and matching on the molecular composition fingerprints of the experimental and theoretical fragment ions of the primary candidate complete N-glycopeptide IDs; S6. Targetedly screen out the sialic acid-containing characteristic oxonium ions and monosaccharide sequence experimental structure diagnostic fragment ions from the primary candidate complete N-glycopeptide IDs with the same molecular composition, and classify the screened primary candidate complete N-glycopeptide IDs as candidate complete N-glycopeptide IDs; S7. Perform a random matching probability scoring on the candidate complete N-glycopeptide IDs, and finally classify those that meet the preset complete N-glycopeptide spectrum matching conditions as targeted GPSMs; S8. According to the steps of S3-S7 in the decoy library, obtain decoy GPSMs; S9. Combine the targeted GPSMs and decoy GPSMs, sort them in ascending order of P score, select a threshold Pscore such that the false positive rate is not greater than 1%, and remove duplicates of the targeted GPSMs with P score not greater than the threshold P score to obtain sialic acid-containing complete N-glycopeptide IDs.
2. The method for identifying and verifying a protein sialic acid according to claim 1, wherein The biological sample to be analyzed is a sample containing sialylated glycoprotein.
3. The method for identifying and verifying a protein sialic acid according to claim 1, wherein In step S2, before electrospray ionization and mass spectrometry analysis, the complete N-glycopeptide sample is first subjected to liquid phase separation.
4. The method for identifying and verifying a protein sialic acid according to claim 1, characterized in that, The experimental molecular composition fingerprint of the precursor ions is measured in the first-order mass spectrum, and the experimental molecular composition fingerprint of the fragment ions is measured in the second-order mass spectrum.
5. The method for identifying and verifying a protein sialic acid according to claim 1, wherein The theoretical molecular composition fingerprint is generated through the following steps: Calculate the molecular formula of each molecule according to the theoretical molecular library of the system under study; Referring to the standard element list, calculate the corresponding molecular composition fingerprint according to the types and quantities of elements in the molecular formula.
6. The identification and verification method of a protein sialic acid according to claim 1, characterized in that The matching in steps S3 and S5 refers to the one-by-one comparison of the m / z value and relative peak intensity value of each isotope peak in the experimental molecular composition fingerprint with the corresponding theoretical values.
7. The method for identifying and verifying a protein sialic acid according to claim 1, wherein The matching criteria in steps S3 and S5 are controlled by the isotope peak intensity cut-off value, the maximum allowable error of the isotope peak mass-to-charge ratio, and the maximum allowable error of the isotope peak intensity.
8. The method for identifying and verifying a protein sialic acid according to claim 1, characterized in that, The duplicate removal in step S9 is carried out according to the standards of the polypeptide backbone amino acid sequence and modification, glycosylation site, and N-linked sugar sequence structure.
Citation Information
Patent Citations
Glycoprotein sialic acid link specific analysis method
CN112461987A
Identification and Quantification of Intact Glycopeptides in Complex Samples
US20150160233A1