A combination of aging-related glycosylation markers and applications thereof
By constructing a combination of site-specific glycopeptides and characteristic glycan structures as glycosylation biomarkers, the shortcomings of existing technologies in analyzing the patterns of glycosylation modification changes have been addressed, enabling precise assessment and dynamic tracking of the aging process and improving the accuracy and reliability of aging biomarkers.
Patent Information
- Application Number
- CN202611100434.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-23
- Publication Date
- 2026-08-25
AI Technical Summary
Existing technologies struggle to analyze age-related glycosylation modification patterns at specific glycosylation sites of specific glycoproteins with precision, and lack longitudinal dynamic tracking across multiple age stages throughout the life cycle, resulting in insufficient accuracy and robustness of aging-related biomarker screening strategies.
This invention provides a combination of aging-related glycosylation biomarkers, including site-specific glycopeptides and characteristic glycan structures. Through differential expression analysis, time series analysis, biological function enrichment analysis, and co-expression network analysis, dynamic change features significantly related to physiological age are screened out, and a multi-level biomarker combination is constructed for the preparation of reagent kits that distinguish different physiological age stages.
This study enabled precise correlation analysis between glycoproteins, glycosylation sites, and glycan structures, improving the resolution of glycosylation information, enhancing the accuracy and robustness of aging assessment, and providing crucial data support for elucidating the molecular mechanisms of glycosylation reprogramming in the aging process.
Smart Images

Figure CN122631804A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical technology, specifically to a combination of aging-related glycosylation biomarkers and their applications. Background Technology
[0002] Aging is a complex physiological process that occurs in organisms with increasing age, accompanied by the progressive decline in the function of tissues and organs and an increased risk of various age-related diseases. Since the aging process is influenced by multiple factors, including genetic background, environmental exposure, and physiological state, developing aging-related biomarkers that can objectively reflect the dynamic changes at the molecular level as the body ages is of great significance for a deeper understanding of the biological mechanisms of aging, assessing biological age, and providing early warning of age-related diseases.
[0003] Blood, as a systemic circulatory medium connecting multiple tissues and organs, contains a wealth of information about proteins and their post-translational modifications that reflect the body's physiological and pathological states. Glycosylation, as one of the most important post-translational modifications, is widely involved in key biological processes such as protein stability regulation, cell recognition, signal transduction, and immune responses. Furthermore, the level of glycosylation modification can dynamically remodel with age. Therefore, age-related glycosylation characteristics in the blood glycoproteome have significant potential value as biomarkers of aging.
[0004] However, existing research largely focuses on age-related analyses of overall glycoform profile changes or single glycoprotein levels, making it difficult to analyze age-related glycosylation modification patterns at the precision of specific glycosylation sites for particular glycoproteins. Since the same glycan structure can modify different glycoproteins or different glycosylation sites of the same glycoprotein, its biological function depends not only on the glycan structure itself but also on the specific combination of "glycoprotein-glycosylation site-glycan structure." Glycosylation analysis detached from site-specificity cannot accurately reveal the precise molecular regulatory mechanisms of glycosylation modification during aging. Furthermore, existing biomarker screening strategies are mostly based on cross-sectional studies or single-time-point data, lacking longitudinal dynamic tracking across multiple age stages throughout the lifespan and failing to effectively integrate overall abundance changes in the blood glycoproteome with fine structural changes at specific sites. Summary of the Invention
[0005] In view of the shortcomings of the prior art, the purpose of this invention is to provide a combination of aging-related glycosylation biomarkers and their applications, aiming to solve at least one of the problems in the above-mentioned background art.
[0006] This invention provides a combination of aging-related glycosylation biomarkers, the combination of biomarkers comprising one or more of the following glycosylation features: Site-specific glycopeptides are composed of glycoproteins, age-related glycosylation sites located on the glycoproteins, and corresponding glycan structures. The glycoproteins include one or more of SPA3K, MUG1, and HPT. The glycosylation sites include positions 39, 185, and 270 of SPA3K, positions 294 and 313 of MUG1, and positions 148, 182, and 256 of HPT. The corresponding glycan structures are selected from one or more of H5N4, H6N4F1G1, H6N4F1, H4N4, H9N2, H5N4G1, H6N5, and H6N5G1. Characteristic glycan structures, selected from one or more of H4N4F1, H3N4F1, H4N3F1, H3N3F1, H7N7F1, H5N4F2, H5N3F1 and H6N4, are used to characterize changes in glycan abundance in the overall blood glycoprotein community. The glycosylation features exhibit dynamic changes in isolated blood samples that are significantly correlated with physiological age, and were obtained through differential expression analysis, time series analysis, biological function enrichment analysis, and co-expression network analysis.
[0007] According to one aspect of the above technical solution, the dynamic change feature is selected from the following modes: With increasing physiological age, it shows one or more of the following trends: a continuous upward trend, a continuous downward trend, or a phased change trend.
[0008] According to one aspect of the above technical solution, the combination of markers satisfies at least one of the following dependencies: Includes at least one of the site-specific glycopeptides; It includes at least two of the aforementioned characteristic glycan structures, and the characteristic glycan structures exhibit statistically significant and unidirectional dynamic changes during physiological aging.
[0009] According to one aspect of the above technical solution, the ex vivo blood sample is derived from a naturally aging model mouse or an induced aging model mouse; Physiological age stages include adolescence, middle age, and old age; The period of adolescence is 2 to 6 months of age, the period of middle age is 10 to 15 months of age, and the period of old age is 18 to 30 months of age.
[0010] According to one aspect of the above technical solution, the differential expression analysis includes missing value filtering, missing value imputation, data standardization and logarithmic transformation to base 2, and uses one-way ANOVA for significance testing. The screening threshold is that the absolute value of the logarithmic multiple change to base 2 is ≥1.0 and the corrected P-value is ≤0.05.
[0011] According to one aspect of the above technical solution, the biological function enrichment analysis is obtained through gene ontology function enrichment analysis and / or Kyoto Encyclopedia of Genes and Genomes pathway enrichment analysis.
[0012] According to one aspect of the above technical solution, the co-expression network analysis is obtained through weighted gene co-expression network analysis, and the specific steps include: Based on the quantitative expression profile of glycans, the correlation coefficients between glycans are calculated, and a glycan co-expression network is constructed; the co-regulatory modules of glycans are identified by the dynamic pruning tree algorithm. Calculate the correlation coefficient between the module feature vector of each glycan co-regulatory module and physiological age, and screen glycan co-regulatory modules that are significantly related to physiological age.
[0013] Another aspect of the present invention provides an application of the combination of aging-related glycosylation biomarkers described above for the preparation of a kit to distinguish different physiological age stages.
[0014] Furthermore, the kit includes: The sample processing unit is used to separate plasma or serum from isolated blood samples, extract total protein, and then reduce and alkylate it before hydrolyzing it with proteases to obtain a polypeptide mixture. The glycopeptide enrichment unit is used to solid-phase enrich the polypeptide mixture using glycopeptide enrichment materials, remove non-glycopeptide components, and collect the enriched glycopeptide eluate. The data detection unit is used to detect the enriched glycopeptide eluent using liquid chromatography-tandem mass spectrometry to obtain the abundance data of the combination of aging-related glycosylation biomarkers. The data analysis unit is used to output evaluation results based on the abundance data.
[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention overcomes the limitations of existing glycomics technologies that only focus on overall glycoform or single protein levels, by treating glycoproteins, glycosylation sites, and glycan structures as a unified whole for correlation analysis. By precisely targeting specific sites on proteins such as SPA3K, MUG1, and HPT, it achieves a leap from macroscopic overall abundance to microscopic site specificity, enabling the analysis of differential glycosylation remodeling events at different sites of the same protein, and greatly improving the resolution of glycosylation information.
[0016] 2. The biomarker combination provided by this invention offers both microscopic and macroscopic perspectives: it includes site-specific glycopeptides reflecting functional changes at specific sites of specific proteins, as well as characteristic glycan structures reflecting the overall glycoproteome of blood. This multi-level biomarker combination overcomes the randomness and one-sidedness of single-type biomarkers, and significantly enhances the accuracy and robustness of aging assessment through the mutual corroboration of microscopic and macroscopic information.
[0017] 3. Based on longitudinal analysis across multiple physiological age stages, this invention clarifies the dynamic changes in various glycosylation characteristics as physiological age increases, such as continuous upregulation, downregulation, or phased changes. This provides crucial data support for elucidating the molecular mechanisms of glycosylation reprogramming in the aging process and lays a solid foundation for developing novel biomarkers that can quantify physiological age or assess the effects of aging interventions. Attached Figure Description
[0018] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is a schematic flowchart of the method for constructing a combination of aging-related glycosylation biomarkers in this invention; Figure 2 This is a statistical chart showing the glycoproteomics identification results of samples at different physiological age stages in this invention; Figure 3 This is a principal component analysis (PCA) diagram of site-specific glycopeptide expression profiles of samples from different physiological age stages in this invention. Figure 4 The image shows the UpSet intersection analysis of the site-specific glycopeptide identification results of samples from different physiological age stages in this invention. Figure 5 This is a graph showing the changes in the characteristic glycan structure types of samples at different physiological age stages in this invention. Figure 6 This is a graph showing the results of time-series clustering analysis of blood glycoproteomics data in this invention. Figure 7 This is a heatmap showing the results of weighted gene co-expression network analysis (WGCNA) in this invention and its correlation with physiological age; Figure 8 This is a graph showing the time-series clustering analysis results of site-specific glycopeptides in this invention; Figure 9 This is a graph showing the time-series clustering analysis results of the characteristic glycan structures in this invention; Figure 10 This is a graph showing the Spearman rank correlation (Spearman grade correlation) results between the characteristic glycan structure and physiological age in this invention. Figure 11This is a graph showing the principal component analysis (PCA) results of the expression profiles of all characteristic glycan structure combinations in this invention. Figure 12 This is a graph showing the Spearman rank correlation analysis results between site-specific glycopeptides and physiological age in this invention. Figure 13 This is a graph showing the principal component analysis (PCA) results of the expression profiles of all site-specific glycopeptide combinations in this invention. In the figure: the statistical differences between samples of different physiological age stages were tested using one-way ANOVA. The significance level was defined as: not significant (ns), indicating P≥0.05; * indicates P<0.05; ** indicates P<0.01; *** indicates P<0.001; **** indicates P<0.0001. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described below in conjunction with specific embodiments. It should be understood that the following embodiments are for illustrative purposes only and are not intended to limit the scope of protection of this invention. Unless otherwise defined, the technical terms used in this specification have the meanings commonly understood by those skilled in the art.
[0020] This invention provides a combination of aging-related glycosylation biomarkers, the combination of biomarkers comprising one or more of the following glycosylation features: Site-specific glycopeptides are composed of glycoproteins, age-related glycosylation sites located on the glycoproteins, and corresponding glycan structures. The glycoproteins include one or more of SPA3K, MUG1, and HPT. The glycosylation sites include positions 39, 185, and 270 of SPA3K, positions 294 and 313 of MUG1, and positions 148, 182, and 256 of HPT. The corresponding glycan structures are selected from one or more of H5N4, H6N4F1G1, H6N4F1, H4N4, H9N2, H5N4G1, H6N5, and H6N5G1. Characteristic glycan structures, selected from one or more of H4N4F1, H3N4F1, H4N3F1, H3N3F1, H7N7F1, H5N4F2, H5N3F1 and H6N4, are used to characterize changes in glycan abundance in the overall blood glycoprotein community. The glycosylation features exhibit dynamic changes in isolated blood samples that are significantly correlated with physiological age, and were obtained through differential expression analysis, time series analysis, biological function enrichment analysis, and co-expression network analysis.
[0021] It should be noted that the biomarker combination of the present invention breaks through the limitation of existing technologies that rely solely on the overall glycan abundance. Its core lies in simultaneously considering the specific correlation between glycoproteins, glycosylation sites, and corresponding glycan structures. Among them, site-specific glycopeptides, as microscopic biomarkers, are covalently defined by a specific glycoprotein, a specific glycosylation site on that glycoprotein, and the corresponding glycan structure connected to that glycosylation site, forming an indivisible molecular entity. The characteristic glycan structure, as a macroscopic biomarker, characterizes the changes in glycan abundance of the overall blood glycoprotein group. This characteristic does not depend on a specific parent protein or a specific site, but reflects the overall glycosylation profile of glycoproteins. The organic combination of the above microscopic and macroscopic characteristics constitutes the multi-level biomarker combination described in the present invention.
[0022] Furthermore, the dynamic change characteristics are selected from the following patterns: With increasing physiological age, it shows one or more of the following trends: a continuous upward trend, a continuous downward trend, or a phased change trend.
[0023] Accordingly, screening methods for combinations of aging-related glycosylation biomarkers include: Step S1: Obtain in vitro blood samples from at least three different physiological age stages, separate plasma or serum, extract total protein, and then reduce and alkylate it before hydrolyzing it with protease to obtain a polypeptide mixture. Specifically, mice of different physiological ages were selected as research subjects, including young, middle-aged and old groups.
[0024] The age groups were defined as follows: juvenile (2-6 months), middle-aged (10-15 months), and old-aged (18-30 months). Peripheral blood samples were collected from mice at each age, and serum samples were obtained by centrifugation after natural clotting and stored at -80℃ for later use.
[0025] Serum samples from each group of mice were collected and denatured and dissolved in a protein denaturing solution (such as ammonium bicarbonate solution containing urea). The protein concentration was then determined using the Bicinchoninic Acid Assay (BCA).
[0026] Dithiothreitol (DTT) was added to the resulting protein solution for protein reduction, followed by iodoacetamide (IAA) for protein alkylation. The reducing agent used for reduction was an ammonium bicarbonate solution containing dithiothreitol (1 mM-10 mM), and the ammonium bicarbonate solution concentration was 40 mM-60 mM. The reduction reaction conditions were: temperature 80℃-100℃, time 5 min-30 min. The alkylating agent used for alkylation was an ammonium bicarbonate solution containing iodoacetamide (5 mM-20 mM), and the ammonium bicarbonate solution concentration was 40 mM-60 mM. The alkylation reaction conditions were: temperature 20℃-37℃, reaction in the dark, 15 min-60 min.
[0027] Finally, trypsin is added to the protein sample for enzymatic hydrolysis. After hydrolysis, formic acid is added to terminate the reaction, yielding a plasma enzymatically hydrolyzed polypeptide sample, i.e., a polypeptide mixture. The mass ratio of trypsin to protein sample is 1:20 to 1:100, and the hydrolysis reaction conditions are: temperature 35℃-40℃, time 4h-24h.
[0028] Before glycopeptide enrichment, the polypeptide mixture after proteolysis and desalting was quantitatively analyzed to determine the actual mass of polypeptides entering the glycopeptide enrichment process. Specifically, the polypeptide mixture after proteolysis was terminated was desalted using a C18 solid-phase extraction column. The resulting polypeptide solution was concentrated by vacuum centrifugation, and the concentration was measured using a nanodrop ultraviolet spectrophotometer. The actual mass of polypeptides entering the glycopeptide enrichment process was calculated based on the measured polypeptide concentration and the volume of the polypeptide solution.
[0029] Step S2: The polypeptide mixture is enriched in a solid phase using a glycopeptide enrichment material to remove non-glycopeptide components and collect the enriched glycopeptide eluent. Specifically, the glycopeptide enrichment material includes one or more of the following: ZIC-HILIC (Zwitterionic Hydrophilic Interaction Liquid Chromatography), amino silica gel, maleimide-glycosyl affinity material (Click-Mal), and cysteine-glycosyl affinity material (Click-Cys).
[0030] The mass ratio of the glycopeptide enrichment material to the polypeptide mixture is 10:1 to 100:1. During the glycopeptide enrichment process, selective retention of glycopeptides is achieved through the affinity between the glycans and the enrichment material. In a specific embodiment, 1 μL of peripheral blood sample yields a polypeptide mixture with a mass of 50 μg-60 μg, followed by the addition of 1.5 mg of glycopeptide enrichment material.
[0031] Glycopeptide enrichment can be performed using solid-phase extraction or dispersion solid-phase extraction. The specific steps are as follows: The peptide mixture was evaporated to dryness, desalted using a C18 solid phase extraction column (C18 SPE), and then redissolved in an equilibration solution. The equilibration solution was an organic phase solution containing organic acids, with a volume concentration of 0.1%-5% and a volume concentration of 60%-95%. The organic acids were formic acid, acetic acid, or trifluoroacetic acid, and the organic phase was acetonitrile, methanol, or ethanol. The glycopeptide enrichment material was packed into a pipette tip or gel loading tip to prepare a solid phase extraction column.
[0032] After equilibrating the peptide mixture with equilibration buffer, the equilibrated peptide mixture is loaded into the solid-phase extraction column. The solid-phase extraction column is washed with 5-200 column volumes of equilibration buffer to remove phosphorylated peptides and unmodified peptides. The loading solution and eluent during the enrichment process are collected and combined. Glycopeptides are eluted with 5-100 column volumes of glycopeptide eluent, wherein the glycopeptide eluent is an organic phase solution containing organic acid, wherein the volume concentration of organic acid is 0.1%-5%, the volume concentration of organic phase is 20%-55%, the organic acid is formic acid, acetic acid, or trifluoroacetic acid, and the organic phase is acetonitrile, methanol, or ethanol.
[0033] When enriching glycopeptides in dispersion solid-phase extraction mode, the polypeptide mixture is directly mixed with the glycopeptide enrichment material, incubated for 0.5 min to 60 min, centrifuged, and then the material is rinsed with equilibration buffer. After centrifugation, the sample solution and the eluent are collected and combined. Then, the glycopeptides are eluted with glycopeptide elution buffer.
[0034] The glycopeptide eluent obtained by solid-phase extraction or dispersion solid-phase extraction is the enriched glycopeptide eluent after being evaporated to dryness.
[0035] Step S3: The enriched glycopeptide eluent is detected by liquid chromatography-tandem mass spectrometry to obtain blood glycoproteomics data; Specifically, liquid chromatography-tandem mass spectrometry uses any one of the following platforms: Orbitrap mass spectrometry platform, Trapped Ion Mobility Spectrometry-Time of Flight (TIMS-TOF) platform, or Quadrupole-Time of Flight (Q-TOF) platform.
[0036] Furthermore, the liquid chromatography-tandem mass spectrometry separation and analysis process includes the use of any one of the Thermo Fisher Orbitrap series mass spectrometers, the Bruker TIMS-TOF series mass spectrometers, or other liquid chromatography-mass spectrometry systems with high-resolution detection capabilities, and is equipped with a nano-level liquid chromatography system.
[0037] The liquid chromatography analysis column and pre-column are reversed-phase C18 capillary columns with an outer diameter of 100 μm, 360 μm, or 1 mm and an inner diameter of 10 μm, 50 μm, or 100 μm. The column packing length is 10 cm-50 cm, preferably 15 cm-30 cm, and the packing particle size is 1.7 μm-5 μm, preferably 1.9 μm. The column temperature is controlled at 30℃-60℃, preferably 40℃. The injection volume is 100 ng-5 μg, preferably 200 ng-1 μg, and the flow rate is controlled at 100 nL / min-500 nL / min, preferably 300 nL / min.
[0038] The reversed-phase C18 capillary column was used for separation in gradient elution mode. Mobile phase A consisted of deionized water containing 0.1% (v / v) formic acid or acetic acid, and mobile phase B consisted of 80%-100% acetonitrile solution containing 0.1% (v / v) formic acid or acetic acid. The gradient elution program included gradually increasing the concentration of mobile phase B from 2%-10% to 20%-35%, then to 35%-80%, and finally to 80%-95% to complete column washing, followed by column equilibration at the initial concentration. The total gradient time was 30-180 min, preferably 60-120 min.
[0039] During quantitative analysis, stable isotope-labeled peptides, stable isotope-labeled glycopeptides, indexed retention time (iRT) internal standards, or other mass spectrometry internal standards can be added to the sample as needed for retention time correction, instrument performance monitoring, and quantitative correction. Alternatively, label-free quantification strategies can be used for glycoproteomics analysis.
[0040] Mass spectrometry detection employed electrospray ionization (ESI) in positive ion mode for data acquisition. The spray voltage was set to 1.8 kV–2.5 kV, and the capillary temperature to 250 °C–320 °C. The first-stage mass spectrometry (MS1) scanned within the range of 350 m / z–2000 m / z, with a resolution of 60,000–120,000. The second-stage mass spectrometry (MS / MS) used high-energy collision-induced dissociation (HCD) or collision-induced dissociation (CID) for fragmentation, with a normalized collision energy (NCE) set to 20–40, preferably 28–32; the MS / MS resolution was set to 15,000–60,000.
[0041] Data acquisition modes include Data Dependent Acquisition (DDA) and Data Independent Acquisition (DIA). In DDA mode, the first 10-20 precursor ions are selected for fragmentation analysis; in DIA mode, a continuous isolation window of 5-40 Da covers the MS1 scan range, achieving high-coverage quantitative analysis of glycopeptides.
[0042] Step S4: Match the blood glycoproteomics data with the glycosylation-specific database retrieval algorithm to identify the identity of glycoproteins, the location of glycosylation sites and the corresponding glycan structure, and quantitatively normalize the abundance of site-specific glycopeptides among different samples to construct a quantitative blood glycoproteomics dataset. Specifically, the glycoproteomics data were retrieved using one or more software programs selected from Byonic, pGlyco, MSFragger-Glyco, GlycoDecipher, or MetaMorpheus, with the UniProt Mouse reference proteome database being used. In the search parameters, cysteine residue carbamidomethylation was set as a fixed modification, methionine oxidation as a variable modification, and N-glycosylation as a variable modification for searching. The precursor ion mass error was preferably set to 5ppm-20ppm, and the fragment ion mass error was preferably set to 10ppm-50ppm. The false discovery rate (FDR) for both precursor ion and peptide identification was controlled to be within 1%.
[0043] After the database search was completed, site-specific glycopeptides were quantitatively analyzed using either the label-free quantitative algorithm built into the search software or a peak area integration algorithm based on extracted ion current (XIC). Glycopeptide abundance was determined using the precursor ion intensity or the extracted ion current peak area as the quantitative basis.
[0044] To eliminate systematic errors among different samples, site-specific glycopeptide quantification data are normalized. The normalization method includes one of the following: Total Ion Current (TIC), Median Normalization, Quantile Normalization, or Loose Normalization, with Median Normalization being preferred. After normalization, a log2 transformation is further performed to reduce data skewness and stabilize variance.
[0045] By using quantitative normalization, systematic biases introduced by experimental factors such as sample loading amount, enzymatic digestion efficiency, and mass spectrometry response are eliminated between different samples, making the glycopeptide abundance between different samples comparable.
[0046] Step S5: Perform differential expression analysis and time series analysis on the blood glycoproteomics quantitative dataset to screen candidate glycosylation features whose abundance changes at different physiological age stages reach statistical significance and show a stable dynamic change pattern. Specifically, differential expression analysis includes missing value filtering, missing value imputation, standardization, and logarithmic transformation to base 2. One-way ANOVA (Analysis of Variance) is used for significance testing; the differential expression analysis includes the following steps: Missing values were filtered from glycoproteomics quantitative data, and glycoproteins, glycosylation sites or glycan structures with valid quantitative values of no less than 50% of the sample size in at least one physiological age stage were retained.
[0047] Missing values are filled using one of the following methods: random minimum imputation based on normal distribution, K-nearest neighbor imputation, local minimum imputation, or local random minimum imputation.
[0048] The imputed data is then standardized using one of the following methods: median normalization, quantile normalization, or Z-score standardization. The standardized quantitative data is further transformed using log2 to reduce data skewness and stabilize variance.
[0049] One-way ANOVA was used for significance testing, with multiple comparisons adjusted as needed. False discovery rate (FDR) control was performed using the Benjamini-Hochberg (BH) method. Screening criteria included: adjusted p-value ≤ 0.05; and a logarithmic change to base 2 |log2FoldChange| ≥ 1.0.
[0050] Furthermore, to analyze the dynamic changes in glycosylation characteristics with age, a time series clustering algorithm was used to classify the trends of glycosylation characteristics at different age stages.
[0051] The time series clustering algorithm used is the Mfuzz soft clustering algorithm. Based on the clustering results, the glycosylation features related to physiological age are classified into: types that show a continuous upward trend with increasing physiological age; types that show a continuous downward trend with increasing physiological age; and types that show a phased change trend with increasing physiological age.
[0052] Step S6: Perform biological function enrichment analysis and co-expression network analysis on the candidate glycosylation features, and construct the aging-related glycosylation biomarker combination based on the results of the biological function enrichment analysis and co-expression network analysis.
[0053] Specifically, the biological function enrichment analysis is obtained through Gene Ontology (GO) functional enrichment analysis and / or Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis, and is used to screen candidate glycosylation features enriched in aging-related biological function pathways and / or signaling pathways.
[0054] The co-expression network analysis was obtained through weighted gene co-expression network analysis, and included the following steps: First, using quantitative matrices of glycoproteins, glycosylation sites, or glycan structures as input data, the Pearson correlation coefficient matrix is calculated on the data after missing value processing and standardization.
[0055] Subsequently, the pickSoftThreshold function from the WGCNA package was used to filter the soft-thresholding power (β value) and then the scale-free topology fit index (R0) was used. 2The optimal soft threshold β was determined based on the principle of achieving a value of 0.85 or higher while maintaining a high average connectivity. The β value was 4-12, preferably 6-10, and β=8 was used in Example 1. A weighted adjacency matrix was constructed using the determined β value, and then converted into a topological overlap matrix (TOM). The TOM dissimilarity (1-TOM) was used as the distance matrix for hierarchical clustering analysis.
[0056] A dynamic pruning tree algorithm is used to identify co-expressed modules. The minimum module size (minModuleSize) is set to 20-50, preferably 30; the module pruning sensitivity parameter (deepSplit) is set to 2; the correlation between module feature vectors is used for module merging. Module merging is performed when the distance between module feature vectors is less than 0.25, that is, the module merging threshold (mergeCutHeight) is set to 0.25.
[0057] After constructing the feature vector (Module Eigengene, ME) for each module, physiological age was used as a continuous phenotypic variable. Pearson correlation analysis was employed to calculate the correlation coefficient (r) and significance level (P-value) between each module and age. Preferably, |r| ≥ 0.50 and P < 0.05 were used as the screening criteria for age-related modules, and modules with an absolute correlation coefficient ≥ 0.60 were preferred as significantly related modules.
[0058] For the age-related modules obtained through screening, key glycoproteins, glycosylation sites, and corresponding glycan structures are screened by combining module membership (MM) and gene significance (GS). Among them, MM ≥ 0.80 and GS ≥ 0.50 are preferred as the screening criteria for core molecules of the module.
[0059] Finally, the selected candidate glycosylation features were subjected to biological function enrichment analysis and co-expression network analysis to screen glycosylation features significantly associated with the aging process, and a ensemble of aging-related glycosylation biomarkers was constructed. This ensemble of aging-related glycosylation biomarkers reveals information at both the microscopic and macroscopic levels, namely site-specific glycopeptides and characteristic glycans.
[0060] Furthermore, this application also provides the application of the aforementioned combination of aging-related glycosylation biomarkers for the preparation of kits that differentiate between different physiological age stages. The kit includes: The sample processing unit is used to separate plasma or serum from isolated blood samples, extract total protein, and then reduce and alkylate it before hydrolyzing it with proteases to obtain a polypeptide mixture. The glycopeptide enrichment unit is used to solid-phase enrich the polypeptide mixture using glycopeptide enrichment materials, remove non-glycopeptide components, and collect the enriched glycopeptide eluate. The data detection unit is used to detect the enriched glycopeptide eluent using liquid chromatography-tandem mass spectrometry to obtain the abundance data of the combination of aging-related glycosylation biomarkers. The data analysis unit is used to output evaluation results based on the abundance data.
[0061] As an example rather than a limitation, abundance data can be compared with reference abundance ranges for different physiological age stages. Based on the reference abundance range into which the abundance data falls, the corresponding physiological age stage can be determined, and the evaluation result can be output.
[0062] The present invention is further illustrated below with specific embodiments: Example 1 like Figure 1 As shown, serum samples were collected from mice aged 2 months (juvenile group), 12 months (middle-aged group), and 24 months (old group). Low-abundance glycoproteins were captured and enriched from the serum samples. The enriched glycoproteins underwent sample processing (including protein extraction, reduction, alkylation, and enzymatic digestion). Liquid chromatography-tandem mass spectrometry (LC-MS / MS) was used for detection and analysis. Proteomics data analysis software (PROTEIN METRICS, Spectronaut) was employed. TM Data parsing is performed.
[0063] Specifically as follows: Experimental materials C57BL / 6J mice were used as an aging model. Three age groups were set up: young group (2 months old, n=6), middle-aged group (12 months old, n=6), and old group (24 months old, n=6).
[0064] Sample preparation and enzymatic digestion Blood was collected from the orbital venous plexus of mice, allowed to stand at room temperature for 30 minutes, and then centrifuged at 4℃ and 3000 rpm for 10 minutes to separate the supernatant serum, which was then stored at -80℃ for later use.
[0065] Take 50 μL of serum and add 200 μL of lysis buffer (50 mM ammonium bicarbonate solution containing 8 M urea) for protein denaturation. Determine protein concentration using the BCA method.
[0066] Take a solution containing 100 μg of protein, add a 50 mM ammonium bicarbonate solution containing 5 mM dithiothreitol, and reduce at 95 °C for 10 min; after cooling to room temperature, add an ammonium bicarbonate solution containing 10 mM iodoacetamide, and alkylate in the dark for 30 min.
[0067] Add trypsin at a ratio of trypsin:protein sample = 1:50 (w / w) and incubate at 37°C for 16 hours. After incubation, add 5% formic acid to terminate the reaction and evaporate to dryness under vacuum.
[0068] Glycopeptide enrichment The evaporated peptides were resuspended in 80% acetonitrile / 1% trifluoroacetic acid solution. 1.5 mg of Click-Mal material was packed into a pipette tip to prepare a solid-phase extraction (SPE) column. The peptide sample was loaded into the SPE column and washed four times with 50 μL of 80% acetonitrile / 1% trifluoroacetic acid solution. The eluent and wash were collected and combined. The column was then eluted twice with 30 μL of 30% acetonitrile / 1% formic acid solution. The resulting glycopeptide eluent was evaporated to dryness to obtain the glycopeptide sample.
[0069] Liquid chromatography-tandem detection Glycopeptide samples were analyzed using an EASY-nLC 1200 nano-level liquid chromatography system (Thermo Scientific) coupled with Orbitrap Exploris. TM Liquid chromatography-tandem mass spectrometry (LC-MS / MS) analysis was performed using a Thermo Scientific 480 high-resolution mass spectrometer. Liquid chromatography separation was performed using an Acclaim instrument. TM PepMap TM A C18 reversed-phase analytical column (50 μm × 150 mm, 2 μm, ThermoScientific). Mobile phase A was 0.1% (v / v) formic acid aqueous solution, and mobile phase B was 80% acetonitrile / 0.1% (v / v) formic acid solution. Separation was performed using reversed-phase gradient elution mode, with mobile phase B gradually increased from a low proportion to a high proportion, followed by column elution under a high proportion of organic phase, and then column equilibration was performed by restoring the initial proportion; the total gradient time was 60 min, and the flow rate was 400 nL / min.
[0070] Mass spectrometry data acquisition was performed using an electrospray ionization source in positive ion mode, employing a data-dependent acquisition mode. The primary mass spectrometry scan mass range was set to 350 m / z–1500 m / z, with a resolution of 60,000. Secondary mass spectrometry used high-energy collision-induced dissociation mode for fragmentation analysis, with a resolution of 30,000, and normalized collision energies were achieved using a multi-energy collision mode. A dynamic exclusion strategy was employed. The obtained raw mass spectrometry data were used for subsequent database searching and quantitative analysis.
[0071] Database retrieval and quantitative analysis Raw glycoproteomics mass spectrometry data were retrieved using Byonic software, with the UniProt mouse reference protein database employed. In the search parameters, cysteine residue aminomethylation was set as a fixed modification, while methionine oxidation and N-glycosylation were set as variable modifications. Trypsin was selected as a specific digestion method, allowing a maximum of two missed cleavage sites. The precursor ion mass error was set to ±10 ppm, and the fragment ion mass error was set to 20 ppm. A target-decoy database strategy was used to control the false positive rate, with both peptide and protein false positive rates kept below 1%.
[0072] Glycopeptide quantification employed a label-free quantification strategy based on the peak area of precursor ion extraction ions. Site-specific glycopeptides were used as independent quantification units to calculate the abundance of glycopeptides in different samples. Different glycosylation sites of the same glycoprotein and different corresponding glycan structures at the same glycosylation site were quantified separately without being combined.
[0073] To reduce the impact of systematic errors among different samples, the quantitative data were preprocessed as follows: missing values were imputed using the random minimum imputation method based on normal distribution; median normalization was used for standardization; and log2 transformation was further performed to reduce the skewness of the data distribution and stabilize the variance.
[0074] This embodiment employs a label-free quantification strategy without adding a stable isotope internal standard; quality control samples are set up during the detection of each batch of samples to monitor liquid chromatography retention time, mass spectrometry response stability, and instrument repeatability.
[0075] Data Analysis and Biomarker Screening Intergroup difference analysis and time series analysis were performed on the preprocessed quantitative data. Intergroup difference analysis used one-way ANOVA, with the selection criteria being |log2Fold Change| ≥ 1.0 and corrected P ≤ 0.05. Time series analysis used the Mfuzz soft clustering algorithm to obtain candidate glycosylation features.
[0076] GO functional enrichment analysis and KEGG pathway enrichment analysis were performed on candidate glycosylation features. A glycan co-expression module was constructed using weighted gene co-expression network analysis, and co-expression modules significantly associated with physiological age were screened. The results of differential analysis, time-series analysis, functional enrichment analysis, and co-expression network analysis were integrated to construct a ensemble of aging-related glycosylation biomarkers, namely site-specific glycopeptides and characteristic glycan structures.
[0077] like Figure 2 As shown, a total of 419 glycoproteins, 17,120 site-specific glycopeptides, and 1,256 N-glycosylation sites were identified. This indicates that the deep glycoproteomics analysis workflow established in this invention can achieve high-coverage identification of serum glycoproteomes.
[0078] like Figure 3 As shown, principal component analysis was performed on the site-specific glycopeptide expression profiles of samples from different physiological age stages. The results showed that the samples from youth, middle age and old age exhibited a clear separation trend in the principal component space, indicating that the serum glycoproteome underwent significant remodeling with increasing physiological age.
[0079] like Figure 4 As shown, Upset intersection analysis of site-specific glycopeptides identified at different physiological age stages revealed 2602, 3680, and 1085 unique site-specific glycopeptides in youth, middle age, and old age, respectively, suggesting that there are physiological age-specific glycosylation modification changes during aging.
[0080] like Figure 5 As shown, a statistical analysis of characteristic glycan structures at different physiological age stages revealed significant changes in characteristic glycan structures such as high-mannose glycans, fucosylated glycans, sialylated glycans, and complex glycans during aging. Among these, heterozygous glycans exhibited a significant trend of first increasing and then decreasing with age, indicating a characteristic glycan remodeling process accompanying aging.
[0081] like Figure 6 As shown, time-series cluster analysis revealed that glycoproteins, glycosylation sites, and corresponding glycan structures could be clustered into six different expression patterns. Cluster 3 showed a continuous increase with age; Cluster 2 and Cluster 5 showed an initial upregulation followed by a downregulation with age; and Cluster 1 showed a peak expression in middle age followed by a decline. These results indicate that changes in aging-related glycoproteins in serum exhibit significant dynamic temporal characteristics.
[0082] like Figure 7 As shown, glycan co-expression modules were constructed using weighted gene co-expression network analysis (WGCNA), yielding five site-specific glycan regulatory modules. Modules 2 and 3 showed significant correlations with age. Further analysis revealed that modules significantly associated with physiological age were primarily enriched in three glycoproteins: SPA3K (serum amyloid A3 kinase), MUG1 (mouse urinary globulin), and HPT (hopoiesis-binding globulin). This suggests that site-specific glycosylation modifications on these glycoproteins are closely related to the aging process and could serve as key candidate molecules for a combination of aging-related glycosylation biomarkers.
[0083] like Figure 8As shown, this is the time-series clustering analysis result of some representative site-specific glycopeptides related to physiological age. Based on the Mfuzz soft clustering algorithm, the expression levels of the selected site-specific glycopeptides were analyzed, revealing the dynamic evolution of site-specific glycopeptides on different glycoproteins (such as SPA3K, MUG1, and HPT) with physiological age. Site-specific glycopeptides are jointly defined by specific glycoproteins, specific glycosylation sites, and the corresponding glycan structures linked to those glycosylation sites. The results show that site-specific glycopeptides from different sources exhibit a significantly heterogeneous dynamic change pattern: such as... Figure 8 As shown by the clustering trends and the data in Table 1, some site-specific glycopeptides exhibit a continuous upregulation trend with age; others show a continuous downregulation trend; and still others exhibit a phased change characteristic of first upregulation and then downregulation. These results confirm that glycosylation modifications at specific sites have a clear physiological age-dependent remodeling characteristic, providing direct evidence for understanding site-specific glycosylation reprogramming in the aging process.
[0084] Table 1: Site-specific glycopeptides related to physiological age
[0085] like Figure 9 As shown, time-series cluster analysis of characteristic glycan structures revealed the following: H4N4F1 (tetrahexose tetraN-acetyxosamine monofucose), H3N4F1 (trihexose tetraN-acetyxosamine monofucose), H4N3F1 (tetrahexose triN-acetyxosamine monofucose), H3N3F1 (trihexose triN-acetyxosamine monofucose), H7N7F1 (heptahexose heptaN-acetyxosamine monofucose), H5N4F2 (pentahexose tetraN-acetyxosamine difucose), and H5N3F1 (pentahexose triN-acetyxosamine monofucose). Eight N-glycans, including fucose (monofucose) and H6N4 (hexahexose tetra-N-acetylhexosamine, without fucose), are characterized by H representing hexose (Hexo), N representing N-acetylhexosamine (HexNAc), F representing fucose, G representing sialic acid (NeuGc), and A representing sialic acid (NeuAc). The numbers represent the number of corresponding monosaccharide residues. These characteristic glycan structures characterize the macroscopic abundance changes of the overall blood glycoprotein group, which show a significant trend evolution with physiological age.
[0086] like Figure 10As shown, to further verify the association between the characteristic glycan structures screened in this invention and physiological age, Spearman rank correlation analysis was performed on the screened characteristic glycan structures. The results showed that many characteristic glycan structures were significantly correlated with physiological age. Among them, some characteristic glycan structures showed a positive correlation with age, while others showed a negative correlation, suggesting that different characteristic glycan structures exhibit differentiated dynamic remodeling characteristics during aging. The above results indicate that the characteristic glycan structures screened in this invention can reflect changes in glycosylation composition with physiological age at the overall glycoproteome level, providing a molecular-level basis for constructing a combination of aging-related glycosylation biomarkers.
[0087] like Figure 11 As shown, to verify the characterization ability of multiple characteristic glycan structure combinations on physiological age, principal component analysis was performed based on the expression profiles of all selected characteristic glycan structure combinations. The first principal component (PC1) was extracted as a comprehensive characteristic index of all characteristic glycan structure combinations. Statistical analysis showed that the samples from different physiological age stages (youth, middle age, and old age) exhibited a significant dispersion trend, and the PC1 scores showed significant differences among the three groups, showing a consistent trend with the increase of physiological age. The above results confirm that the joint analysis of all characteristic glycan structure combinations can effectively capture the systematic remodeling characteristics of the overall blood glycoprotein community during the aging process, indicating that the characteristic glycan structures of the present invention have the potential to quantify physiological age.
[0088] like Figure 12 As shown, to clarify the association between site-specific glycopeptides and physiological age, Spearman rank correlation analysis was performed on the abundance of the screened site-specific glycopeptides and physiological age. The results showed that multiple site-specific glycopeptides were significantly correlated with physiological age: some site-specific glycopeptides showed a significant positive correlation (i.e., their abundance increased with age), while others showed a significant negative correlation (i.e., their abundance decreased with age). These results indicate that changes in site-specific glycopeptides have a clear physiological age dependence and can reflect the glycosylation remodeling characteristics that occur during aging at specific glycoprotein and site levels, further validating the physiological age correlation of the site-specific glycopeptides screened in this invention.
[0089] like Figure 13As shown, to evaluate the characterization ability of combinations of multiple site-specific glycopeptides for physiological age, principal component analysis (PCA) was performed based on the expression profiles of all selected site-specific glycopeptide combinations. The results showed that samples from different physiological age stages exhibited a certain separation trend in the principal component space, indicating that with increasing physiological age, all site-specific glycopeptide combinations undergo overall remodeling. Comparative results showed significant differences in PC1 scores among samples from youth, middle age, and old age. In conclusion, biomarkers composed of combinations of all site-specific glycopeptides can comprehensively reflect specific site glycosylation changes induced by physiological aging and can be used for molecular-level assessment of aging.
[0090] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0091] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this patent should be determined by the appended claims.
Claims
1. A combination of aging-related glycosylation biomarkers, characterized in that, The combination of biomarkers includes one or more of the following glycosylation features: Site-specific glycopeptides are composed of glycoproteins, age-related glycosylation sites located on the glycoproteins, and corresponding glycan structures. The glycoproteins include one or more of SPA3K, MUG1, and HPT. The glycosylation sites include positions 39, 185, and 270 of SPA3K, positions 294 and 313 of MUG1, and positions 148, 182, and 256 of HPT. The corresponding glycan structures are selected from one or more of H5N4, H6N4F1G1, H6N4F1, H4N4, H9N2, H5N4G1, H6N5, and H6N5G1. Characteristic glycan structures, selected from one or more of H4N4F1, H3N4F1, H4N3F1, H3N3F1, H7N7F1, H5N4F2, H5N3F1 and H6N4, are used to characterize changes in glycan abundance in the overall blood glycoprotein community. The glycosylation features exhibit dynamic changes in isolated blood samples that are significantly correlated with physiological age, and were obtained through differential expression analysis, time series analysis, biological function enrichment analysis, and co-expression network analysis.
2. The combination of aging-related glycosylation biomarkers according to claim 1, characterized in that, The dynamic change characteristics are selected from the following patterns: With increasing physiological age, it exhibits one or more of the following trends: a continuous upward trend, a continuous downward trend, or a phased change trend.
3. The combination of aging-related glycosylation biomarkers according to claim 1, characterized in that, The combination of markers satisfies at least one of the following dependencies: Includes at least one of the site-specific glycopeptides; It includes at least two of the aforementioned characteristic glycan structures, and the characteristic glycan structures exhibit statistically significant and unidirectional dynamic changes during physiological aging.
4. The combination of aging-related glycosylation biomarkers according to claim 1, characterized in that, The ex vivo blood samples were derived from naturally aging model mice or induced aging model mice; The different physiological age stages include youth, middle age, and old age; The period of adolescence is 2 to 6 months of age, the period of middle age is 10 to 15 months of age, and the period of old age is 18 to 30 months of age.
5. The combination of aging-related glycosylation biomarkers according to claim 1, characterized in that, The differential expression analysis includes missing value filtering, missing value imputation, data standardization, and logarithmic transformation to base 2. One-way ANOVA is used for significance testing, and the screening threshold is the absolute value of the logarithmic change to base 2 ≥ 1.0 and the corrected P value ≤ 0.
05.
6. The combination of aging-related glycosylation biomarkers according to claim 1, characterized in that, The biological function enrichment analysis was obtained through gene ontology function enrichment analysis and / or Kyoto Encyclopedia of Genes and Genomes pathway enrichment analysis.
7. The combination of aging-related glycosylation biomarkers according to claim 1, characterized in that, The co-expression network analysis was obtained through weighted gene co-expression network analysis, and the specific steps included: Based on the quantitative expression profile of glycans, the correlation coefficients between glycans are calculated, and a glycan co-expression network is constructed; the co-regulatory modules of glycans are identified by the dynamic pruning tree algorithm. Calculate the correlation coefficient between the module feature vector of each glycan co-regulatory module and physiological age, and screen glycan co-regulatory modules that are significantly related to physiological age.
8. The application of a combination of aging-related glycosylation biomarkers as described in any one of claims 1-7, characterized in that, Used to prepare reagent kits that differentiate between different physiological age stages.
9. The application of the combination of aging-related glycosylation biomarkers according to claim 8, characterized in that, The kit includes: The sample processing unit is used to separate plasma or serum from isolated blood samples, extract total protein, and then reduce and alkylate it before hydrolyzing it with proteases to obtain a polypeptide mixture. The glycopeptide enrichment unit is used to solid-phase enrich the polypeptide mixture using glycopeptide enrichment materials, remove non-glycopeptide components, and collect the enriched glycopeptide eluent. The data detection unit is used to detect the enriched glycopeptide eluent using liquid chromatography-tandem mass spectrometry to obtain the abundance data of the combination of aging-related glycosylation biomarkers. The data analysis unit is used to output evaluation results based on the abundance data.