A method for discriminating between n-linked glycan sialic acid alpha-2,3 and alpha-2,6 linkages using a signature ion model
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-14
- Publication Date
- 2026-08-11
AI Technical Summary
虽然化学衍生化方法(如特定的还原胺化或点击化学)可以引入质量标签来区分这两种异构体,但这增加了实验操作的复杂性和成本,且不适用于回顾性分析已有的非特异性标记数据
Smart Images

Figure CN122551875A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of identification technology of sialic acid linkage isomers, and specifically relates to a method for identifying the linkage isomers of N-linked sugar sialic acid α-2,3 and α-2,6 using a characteristic ion model. Background Technology
[0002] Glycoproteomics, as an important branch of proteomics research, is a fundamental research field driving the advancement of precision medicine. It plays a crucial role in cutting-edge life science research, including the discovery of disease biomarkers, antibody drug development, the elucidation of viral invasion mechanisms, and cell signaling. Glycosylation modification of proteins is not only one of the most common and complex post-translational modifications, but it also profoundly affects protein structural stability, solubility, immunogenicity, and biological activity. N-linked glycosylation, in particular, refers to the process by which glycan chains are covalently linked to asparagine residues in proteins. Its structure exhibits high microheterogeneity and macroheterogeneity, posing a significant challenge to the fine structural analysis of glycoproteins.
[0003] Mass spectrometry (MS) has become the mainstream method for glycoproteomics analysis due to its advantages of high sensitivity, high throughput, and high accuracy. In particular, liquid chromatography-mass spectrometry (LC-MS / MS), combined with data-dependent acquisition (DDA) or data-independent acquisition (DIA) modes, enables large-scale identification of intact glycopeptides in complex biological samples. Currently, mature commercial or open-source analytical tools such as PEAKS GlycanFinder, StrucGP, and Byonic are available for peptide sequence identification, glycosylation site confirmation, and inference of glycan composition (e.g., the number of monosaccharide units such as Hex, HexNAc, and Neu5Ac). These tools primarily rely on constructing theoretical spectral databases and inferring the molecular formula and basic structure of glycopeptides by matching fragment ions (such as Y, B, and oxonium ions) in experimental spectra.
[0004] However, as research into the relationship between glycoprotein molecular structure and function deepens, the scientific community's need for elucidation of the fine structure of glycans is becoming increasingly urgent. The biological function of glycans depends not only on their monosaccharide composition but also on the linkage and isomer configurations between monosaccharides. For example, sialic acid, a common terminal modification group in glycans, is directly related to physiological processes such as protein half-life, cell recognition, and immune escape in its N-glycan form. Existing mass spectrometry tools, when dealing with sialic acid modifications, can typically only identify the presence and quantity of sialic acid but struggle to precisely distinguish the specific types of its linked isomers.
[0005] Specifically, the terminal sialic acid linkages are mainly divided into two isomers: α-2,3 linkage and α-2,6 linkage. These two isomers are extremely similar in chemical properties, exhibiting only minor differences in signal intensity or specific fragment ion ratios in mass spectra. Furthermore, they lack specific "diagnostic ions" for direct identification in conventional collision-induced dissociation (CID) or high-energy collisional dissociation (HCD) spectra. While chemical derivatization methods (such as specific reductive amination or click chemistry) can introduce mass tags to distinguish these two isomers, this increases the complexity and cost of experimental procedures and is not suitable for retrospective analysis of existing non-specific labeled data. Therefore, how to utilize existing spectral information to mine deeper features in high-throughput, non-targeted LC-MS / MS glycoproteomics data to achieve accurate and automated identification of N-linked glycosyl sialic acid α-2,3 / 2,6 linkage isomers is a pressing technical challenge in glycoproteomics data analysis and the core pain point that this invention aims to address.
[0006] The information disclosed in this background section is intended only to enhance the understanding of the overall background of the invention and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention
[0007] The purpose of this invention is to provide a method for identifying the α-2,3 and α-2,6 linkage isomers of N-linked sugar sialic acid using a characteristic ion model, thereby overcoming the defects in the prior art.
[0008] To achieve the above objectives, the present invention provides a method for identifying the α-2,3 and α-2,6 linkage isomers of N-linked sugar sialic acid using a characteristic ion model, comprising the following steps: S1: Construct a terminal sialic acid diagnostic model, train a support vector machine (SVM) model using the background dataset, and establish a predictive relationship between feature ion signals and terminal sialic acid glycoforms; S2: Calculate the weights of the heterogeneous identification scoring function, and using the background dataset that has undergone connection-specific derivatization, optimize the weight coefficients of the scoring function based on the terminal sialic acid diagnostic model to obtain the optimal weight combination; S3: Determine the identification threshold for connection heterogeneity. Based on the optimal weight combination, statistically analyze the score distribution of α-2,3 connection and α-2,6 connection isomers in the background dataset, and calculate the identification threshold according to the confidence principle. S4: Target sample identification. Extract the LC-MS / MS spectral feature signals of the sample to be tested, use the terminal sialic acid diagnostic model for screening, and combine the optimal weight combination and identification threshold to determine the α-2,3 or α-2,6 linkage type of sialic acid.
[0009] Preferably, the construction of the terminal sialic acid diagnostic model in step S1 includes the following steps: S11: Feature signal extraction and distribution fitting. Extract the relative intensities of sialic acid-related oxonium ions Sia274, Sia292 and NHS ions from the LC-MS / MS spectra in the background dataset, and statistically analyze and fit the intensity distribution patterns of each feature signal in sialic acid-terminated and non-sialic acid-terminated glycoforms. S12: Determine the false positive threshold, use the background dataset to identify labels, and calculate the 95% confidence false positive assessment threshold corresponding to each feature signal. S13: Model training and optimization. Construct a scoring function containing learnable weights, use an SVM machine learning model to learn the weight coefficients of the scoring function on the background dataset, and select the optimal weights based on the AUC value of the ROC working curve to obtain the terminal sialic acid prediction model.
[0010] Preferably, in step S11, Sia274 is a Neu5Ac-H2O ion, Sia292 is a Neu5Ac ion, and NHS is a B ion with a sugar composition of HexNAc+Hex+Neu5Ac.
[0011] Preferably, the mathematical form of the scoring function in step S13 is: In the formula P sia274 P sia292 and P NHS The true positive confidence level is calculated by fitting a statistical distribution model based on the relative intensity of the corresponding ion detected in the target candidate glycopeptide in the experiment. w1 and w2 are the weighting coefficients, and w1+w2=100.
[0012] Preferably, in step S13, the SVM machine learning model is implemented using a radial basis function kernel rbf, and is set as an SVC object that performs probability output.
[0013] Preferably, the calculation of the heterogeneity discrimination scoring function weights in step S2 includes the following steps: S21: Calculate the confidence level of the feature channel by using the intensity distribution pattern described in claim 2 to calculate the true positive probability of each feature ion in the specific derivatized background dataset. S22: Weight optimization, input the true positive probability into the scoring function, calculate the score under different weight combinations, use the derived label as the true label, evaluate and screen the weight combination with the largest AUC value through the ROC working curve.
[0014] Preferably, determining the connectivity heterogeneity identification threshold in step S3 includes the following steps: S31: Distribution of scores for α-2,3 linkage isomers and α-2,6 linkage isomers under optimal weights in the statistical background dataset; S32: Based on the 95% confidence level principle, calculate the scoring threshold value of the α-2,3 linkage isomer with a confidence level greater than 95% as the lower threshold value, and calculate the scoring threshold value of the α-2,6 linkage isomer with a confidence level greater than 95% as the upper threshold value.
[0015] Preferably, the target sample identification in step S4 includes the following steps: S41: Extract the Sia274, Sia292 and NHS feature signals from the target LC-MS / MS spectrum and input them into the terminal sialic acid diagnostic model described in claim 2 for screening; S42: Calculate the score for the selected terminal sialic acid glycopeptides; S43: If the score is less than the lower threshold, it is determined to be an α-2,3 connection; if the score is greater than the upper threshold, it is determined to be an α-2,6 connection.
[0016] Compared with the prior art, one aspect of the present invention has the following beneficial effects: (1) This invention utilizes non-linkage-specific universal signals (Sia274, Sia292, NHS) and mines their relative intensity pattern features in LC-MS / MS spectra to achieve high-precision diagnosis of sialic acid linker isomers for the first time using non-specific signals. This technology fills the gap in high-throughput glycoproteomics where there is a lack of effective isomer identification methods. (2) This invention constructs a dual-model architecture of "terminal sialic acid diagnostic model" and "α-2,3 / 2,6 connection heterogeneity identification model"; by combining the SVM (Support Vector Machine) algorithm with ROC curve optimization, the optimal weight combination is selected so that the AUC value of terminal sialic acid identification reaches the optimal level, which significantly reduces the false positive rate; and a scoring threshold calculation method based on 95% confidence level is established. By setting clear "upper threshold" and "lower threshold", the ambiguous spectral signal is transformed into a clear connection type judgment result (score value < lower threshold is judged as α-2,3; score value > upper threshold is judged as α-2,6), realizing the standardization and quantification of the identification process; (3) This invention can perform preliminary screening and identification without the need for complex specific chemical derivatization of samples, which greatly reduces experimental costs and operational complexity. It is particularly suitable for retrospective analysis of massive amounts of existing non-specific label mass spectrometry data and can be widely applied to glycoproteomics research of complex biological samples such as human tissues and cell lines. It has extremely high throughput processing capacity. (4) The feature ion scoring function and machine learning training strategy proposed in this invention are not limited to the identification of sialic acid isomers. Their core idea can be extended to the analysis of other glycan isomers that are difficult to distinguish by traditional fragment ions, providing a general algorithm framework for solving a wider range of glycomic structural isomer problems. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the overall process of the present invention for identifying terminal sialic acid glycoforms and α-2,3 / 2,6 linkage isomers by means of characteristic signals; Figure 2 This diagram illustrates the statistical distribution fitting of the feature signals extracted from the background dataset of the training terminal sialic acid diagnostic model in this invention, along with the calculation of the 95% confidence boundary value. Figure 3 ROC working curves and AUC evaluation values of the terminal sialic acid SVM prediction model under different weight combination scoring functions; Figure 4 Annotations for non-sialic acid α-2,3 / 2,6 linkage-specific characteristic ions and signal annotations containing specific derivatization characteristic ions for determining linkage isomerism in an example LC-MS / MS spectrum of a sialic acid α-2,3 / 2,6 linkage-specific derivatization background dataset; Figure 5 The expected probability of characteristic channels and the clustering relationship of connection isomers were calculated for three non-sialic acid α-2,3 / 2,6-linked characteristic ions. Figure 6 ROC curves and AUC values of the sialic acid α-2,3 / 2,6 linkage prediction function under different weights; Figure 7The statistical distribution of predicted values for the two linkage isomers of sialic acid α-2,3 / 2,6 under the weights of the optimal prediction function and the corresponding isomer judgment thresholds are presented. Detailed Implementation
[0018] The specific embodiments of the present invention will be described in detail below, but it should be understood that the scope of protection of the present invention is not limited to the specific embodiments.
[0019] Unless otherwise expressly stated, throughout the specification and claims, the term "comprising" or its variations such as "including" or "comprises" shall be understood to include the stated elements or components without excluding other elements or other components.
[0020] like Figure 1 As shown, a method for identifying the α-2,3 and α-2,6 linkage isomers of N-linked sugar sialic acid using a characteristic ion model includes the following steps: S1: Construct a terminal sialic acid diagnostic model, train a support vector machine (SVM) model using the background dataset, and establish a predictive relationship between feature ion signals and terminal sialic acid glycoforms; S2: Calculate the weights of the heterogeneous identification scoring function, and using the background dataset that has undergone connection-specific derivatization, optimize the weight coefficients of the scoring function based on the terminal sialic acid diagnostic model to obtain the optimal weight combination; S3: Determine the identification threshold for connection heterogeneity. Based on the optimal weight combination, statistically analyze the score distribution of α-2,3 connection and α-2,6 connection isomers in the background dataset, and calculate the identification threshold according to the confidence principle. S4: Target sample identification. Extract the LC-MS / MS spectral feature signals of the sample to be tested, use the terminal sialic acid diagnostic model for screening, and combine the optimal weight combination and identification threshold to determine the α-2,3 or α-2,6 linkage type of sialic acid.
[0021] The specific method for constructing the terminal sialic acid diagnostic model in S1 is as follows: S11: Feature signal extraction and distribution fitting. Extract the relative intensities of sialic acid-related oxonium ions Sia274, Sia292 and NHS ions from the LC-MS / MS spectra in the background dataset, and statistically analyze and fit the intensity distribution patterns of each feature signal in sialic acid-terminated and non-sialic acid-terminated glycoforms. S12: Determine the false positive threshold, use the background dataset to identify labels, and calculate the 95% confidence false positive assessment threshold corresponding to each feature signal. S13: Model training and optimization. Construct a scoring function containing learnable weights, use an SVM machine learning model to learn the weight coefficients of the scoring function on the background dataset, and select the optimal weights based on the AUC value of the ROC working curve to obtain the terminal sialic acid prediction model.
[0022] The specific method for calculating the weights of the heterogeneity discrimination scoring function in S2 is as follows: S21: Calculate the confidence level of the feature channel by using the intensity distribution pattern described in claim 2 to calculate the true positive probability of each feature ion in the specific derivatized background dataset. S22: Weight optimization, input the true positive probability into the scoring function, calculate the score under different weight combinations, use the derived label as the true label, evaluate and screen the weight combination with the largest AUC value through the ROC working curve.
[0023] The specific method for determining the connection heterogeneity discrimination threshold in S3 is as follows: S31: Distribution of scores for α-2,3 linkage isomers and α-2,6 linkage isomers under optimal weights in the statistical background dataset; S32: Based on the 95% confidence level principle, calculate the scoring threshold value of the α-2,3 linkage isomer with a confidence level greater than 95% as the lower threshold value, and calculate the scoring threshold value of the α-2,6 linkage isomer with a confidence level greater than 95% as the upper threshold value.
[0024] The specific method for identifying S4 target samples is as follows: S41: Extract the Sia274, Sia292 and NHS feature signals from the target LC-MS / MS spectrum and input them into the terminal sialic acid diagnostic model described in claim 2 for screening; S42: Calculate the score for the selected terminal sialic acid glycopeptides; S43: If the score is less than the lower threshold, it is determined to be an α-2,3 connection; if the score is greater than the upper threshold, it is determined to be an α-2,6 connection.
[0025] The background dataset required for S11 is taken from complex, intact N-glycopeptide samples labeled with human TMT, acquired via HCD positive ion mode DDA, and the spectra are processed by gPeptide. TM The database search yielded complete glycoform identification results for N-glycopeptides; By extracting the relative intensities of the characteristic ions Sia274, Sia292, and NHS from all LC-MS / MS spectra annotated with glycopeptides, statistical results were obtained as follows: Figure 2 The density distribution map shown uses the identification results of the search software to distinguish the terminal sialic acid tag. Then, the characteristic signal density distribution of sialic acid and non-sialic acid glycotypes is fitted to obtain the fitted density distribution function. The relative intensity value is used as the target boundary value with a 95% confidence level when the error rate of judging the terminal sialic acid glycotype is less than 5% within the interval of relative intensity from 0 to the target boundary value. The error rate is calculated by the percentage of the overlapping area between the interval integral of the non-sialic acid glycotype distribution function (false positive) and the interval integral of the terminal sialic acid glycotype distribution function (true positive). Using the relative intensity density distribution function and a 95% confidence threshold, the expected probability of identifying a sialic acid characteristic signal as a sialic acid terminus can be calculated for any glycopeptide-spectrum identification result. The calculation method involves using the relative intensity true positive and false positive density distribution functions of this feature, along with the aforementioned integration method, to calculate the probability. Then, the three probability values corresponding to the three characteristic signals are input in the following format: In the scoring function, the corresponding score value is calculated under different weights w1 and w2; An SVM model is constructed using an SVC object with a kernel function of rbf and a gamma parameter of 0.05 as the SVM model architecture. The background dataset is randomly and uniformly divided into training and test sets. The SVM model is trained on the training set with different weight values, and its output prediction probability is used as the prediction score. The true scores under the corresponding weights are then used as the true labels to validate the model. The ROC curve and corresponding AUC value of the model are calculated, and the model performance under different weights is statistically obtained as follows: Figure 3 As shown, the optimal weight combination selected by AUC value is w1=70 and w2=30.
[0026] The background dataset required for the target sample was taken from human lung cancer tissue, which underwent specific derivatization with sialic acid α-2,3 / 2,6 linkage, and the complete N-glycopeptide was identified using gPeptide. TM The database search revealed that the spectra used in model training consisted of glycopeptide spectra identification results determined by the aforementioned method of using sialic acid-related oxonium ions to diagnose N-linked glycoforms as terminal sialic acid glycoforms. The true labels for sialic acid α-2,3 / 2,6 linkages in the background dataset were determined by the identification results of the derived labels, such as... Figure 4 As shown.
[0027] Using the same method as described above, the expected probabilities of the characteristic channels were used to calculate three characteristic ions (non-sialic acid α-2,3 / 2,6 linkage-specific characteristic ions). Clustering can visually assess their specificity in distinguishing true labels from sialic acid α-2,3 / 2,6 linkages. Figure 5 The expected probabilities corresponding to the values shown are input into a scoring function of the same form with unoptimized weights to calculate the scores corresponding to the screening glycopeptide-spectrum matching results under different weight combinations.
[0028] Using the scoring function as a classifier and the scores as predicted values, the ROC curves and AUC values of the scoring function under different weights were calculated using the true labels connected by sialic acid α-2,3 / 2,6. Figure 6 As shown in the figure, after evaluation, the optimal weight combination was selected as w1=20 and w2=80, and the scoring function under this weight was used as the final classification function.
[0029] Based on the true labels of sialic acid α-2,3 / 2,6 linkages, the score distribution of the two isomers under their optimal weights was statistically analyzed. Observation of the distribution revealed that the median value of α-2,3 linkages was in the low-value range, while α-2,6 linkages were in the high-value range. Therefore, a scoring threshold of α-2,3 with a confidence level greater than 95 was used as the lower threshold for judging α-2,3 linkage, and a scoring threshold of α-2,6 with a confidence level greater than 95 was used as the upper threshold for judging α-2,6 linkage. For experimental spectral data from non-sialic acid α-2,3 / 2,6 linkage-specific derivatization, the score corresponding to the glycopeptide-spectral matching result was determined to be either α-2,3 or α-2,6 linkage if it was below the lower threshold or above the upper threshold (e.g., ...). Figure 7 (As shown).
[0030] The foregoing description of specific exemplary embodiments of the invention is for illustrative and explanatory purposes. These descriptions are not intended to limit the invention to the precise forms disclosed, and it will be apparent that many changes and variations can be made in accordance with the foregoing teachings. The exemplary embodiments were chosen and described in order to explain the specific principles of the invention and its practical application, thereby enabling those skilled in the art to implement and utilize various different exemplary embodiments of the invention, as well as various different choices and variations. The scope of the invention is intended to be defined by the claims and their equivalents.
Claims
1. A method for identifying the α-2,3 and α-2,6 linkage isomers of N-linked sugar sialic acid using a characteristic ion model, characterized in that, Includes the following steps: S1: Construct a terminal sialic acid diagnostic model, train a support vector machine (SVM) model using the background dataset, and establish a predictive relationship between feature ion signals and terminal sialic acid glycoforms; S2: Calculate the weights of the heterogeneous identification scoring function, and using the background dataset that has undergone connection-specific derivatization, optimize the weight coefficients of the scoring function based on the terminal sialic acid diagnostic model to obtain the optimal weight combination; S3: Determine the identification threshold for connection heterogeneity. Based on the optimal weight combination, statistically analyze the score distribution of α-2,3 connection and α-2,6 connection isomers in the background dataset, and calculate the identification threshold according to the confidence principle. S4: Target sample identification. Extract the LC-MS / MS spectral feature signals of the sample to be tested, use the terminal sialic acid diagnostic model for screening, and combine the optimal weight combination and identification threshold to determine the α-2,3 or α-2,6 linkage type of sialic acid.
2. The method for identifying N-linked sugar sialic acid α-2,3 and α-2,6 linkage isomers using a characteristic ion model according to claim 1, characterized in that, The step S1 of constructing the terminal sialic acid diagnostic model includes the following steps: S11: Feature signal extraction and distribution fitting. Extract the relative intensities of sialic acid-related oxonium ions Sia274, Sia292 and NHS ions from the LC-MS / MS spectra in the background dataset, and statistically analyze and fit the intensity distribution patterns of each feature signal in sialic acid-terminated and non-sialic acid-terminated glycoforms. S12: Determine the false positive threshold, use the background dataset to identify labels, and calculate the 95% confidence false positive assessment threshold corresponding to each feature signal. S13: Model training and optimization. Construct a scoring function containing learnable weights, use an SVM machine learning model to learn the weight coefficients of the scoring function on the background dataset, and select the optimal weights based on the AUC value of the ROC working curve to obtain the terminal sialic acid prediction model.
3. The method for identifying N-linked sugar sialic acid α-2,3 and α-2,6 linkage isomers using a characteristic ion model according to claim 2, characterized in that, In step S11, Sia274 is a Neu5Ac-H2O ion, Sia292 is a Neu5Ac ion, and NHS is a B ion with a sugar composition of HexNAc+Hex+Neu5Ac.
4. The method for identifying N-linked glycosialic acid α-2,3 and α-2,6 linkage isomers using a characteristic ion model according to claim 2, characterized in that, The mathematical form of the scoring function in step S13 is: where P sia274 , P sia292 and P NHS are the relative intensities of the corresponding ions experimentally detected for the target candidate glycopeptide, the true positive confidence calculated by the fitted statistical distribution model, and w1, w2 are the weight coefficients, respectively, with w1 + w2 = 100.
5. The method for identifying the α-2,3 and α-2,6 linkage isomers of N-linked glycosyl sialic acid using a characteristic ion model according to claim 2, characterized in that, In step S13, the SVM machine learning model is implemented using a radial basis function kernel rbf, and is set as an SVC object that outputs probabilities.
6. The method for identifying N-linked glycosialic acid α-2,3 and α-2,6 linkage isomers using a characteristic ion model according to claim 2, characterized in that, The calculation of the heterogeneity discrimination scoring function weights in step S2 includes the following steps: S21: Calculate the confidence level of the feature channel by using the intensity distribution pattern described in claim 2 to calculate the true positive probability of each feature ion in the specific derivatized background dataset. S22: Weight optimization, input the true positive probability into the scoring function, calculate the score under different weight combinations, use the derived label as the true label, evaluate and screen the weight combination with the largest AUC value through the ROC working curve.
7. The method for identifying N-linked sugar sialic acid α-2,3 and α-2,6 linkage isomers using a characteristic ion model according to claim 6, characterized in that, Determining the connection heterogeneity discrimination threshold in step S3 includes the following steps: S31: Distribution of scores for α-2,3 linkage isomers and α-2,6 linkage isomers under optimal weights in the statistical background dataset; S32: Based on the 95% confidence level principle, calculate the scoring threshold value of the α-2,3 linkage isomer with a confidence level greater than 95% as the lower threshold value, and calculate the scoring threshold value of the α-2,6 linkage isomer with a confidence level greater than 95% as the upper threshold value.
8. The method for identifying N-linked glycosialic acid α-2,3 and α-2,6 linkage isomers using a characteristic ion model according to claim 7, characterized in that, The target sample identification in step S4 includes the following steps: S41: Extract the Sia274, Sia292 and NHS feature signals from the target LC-MS / MS spectrum and input them into the terminal sialic acid diagnostic model described in claim 2 for screening; S42: Calculate the score for the selected terminal sialic acid glycopeptides; S43: If the score is less than the lower threshold, it is determined to be an α-2,3 connection; if the score is greater than the upper threshold, it is determined to be an α-2,6 connection.