Kidney biological age calculation method based on glycosylation marker and integrated machine learning

By utilizing N-glycosylation profile data of plasma immunoglobulin G, combined with adaptive group elasticity network and ensemble machine learning, a kidney biological age calculation model was constructed, which solves the accuracy and cost problems of kidney physiological aging assessment in existing technologies and realizes individualized assessment of kidney physiological status.

CN122050801APending Publication Date: 2026-05-15BEIJING BAIYANG TANGKE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING BAIYANG TANGKE TECHNOLOGY CO LTD
Filing Date
2026-01-12
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies are insufficient to accurately reflect the physiological aging state of the kidney organ, and biological age calculation methods suffer from technical complexity, high cost, or lack of causal verification.

Method used

Using N-glycosylation profile data of plasma immunoglobulin G, combined with adaptive group elastic network and ensemble machine learning algorithm, feature dimensionality reduction and causal inference were performed to construct a kidney biological age calculation model.

Benefits of technology

This approach enables specific and accurate assessment of the physiological aging state of the kidneys, improves the model's biological interpretability and predictive robustness, reduces costs, and facilitates widespread application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122050801A_ABST
    Figure CN122050801A_ABST
Patent Text Reader

Abstract

The invention discloses a kidney biological age calculation method based on glycosylation markers and integrated machine learning, and belongs to the technical field of bioinformatics and health monitoring. The method comprises the following steps: firstly, acquiring N-glycosylation spectral pattern data of plasma immunoglobulin G of an individual and preprocessing; secondly, carrying out feature dimension reduction on the glycosyl data by adopting a self-adaptive group elastic network, and screening candidate glycosyl structures; then, further screening out a target glycosyl structure associated with kidney functions by using an integrated machine learning algorithm, and performing causal inference through Mendel randomization to determine a final glycosyl structure; and finally, based on the final glycosyl structure and the individual age, constructing a multi-factor Logistic regression model, and calculating to obtain the kidney glycosyl biological age of the individual. According to the method, quantitative evaluation of the kidney specific biological age is realized by integrating glycomics data and a machine learning algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of renal glycosylation biological age calculation technology, and particularly relates to a method for calculating renal biological age based on glycosylation markers and integrated machine learning. Background Technology

[0002] Biological age aims to quantify an individual's biological aging level. Compared to chronological age, it provides a more accurate and individualized assessment of physiological status, which is of great significance in the fields of health monitoring and risk assessment. However, human aging is a non-uniform process; the rate of functional decline varies significantly among different organs and systems, a phenomenon known as "asynchronous aging." As a key metabolic organ, the kidney's functional state is not synchronized with the overall aging process, and significant individual differences exist. Therefore, developing an assessment method that specifically reflects the aging state of the kidney is of outstanding value for a deeper understanding of organ aging heterogeneity and for targeted health management.

[0003] In methods for calculating biological age, omics-based analytical techniques have shown great potential. Among them, glycomics, by analyzing the dynamic evolution of protein glycosylation modifications, has become an important tool for constructing aging-related assessment models. The N-glycosylation profile of plasma immunoglobulin G (IgG) has been established as a core biomarker reflecting immune and inflammatory states due to its high quantifiability and age-related regular changes.

[0004] Existing technologies include methods for calculating the overall or specific system (such as the cardiovascular system) biological age using the N-glycosylation profile of plasma IgG. For example, machine learning is used to integrate specific glycan features to distinguish chronological age from cardiovascular glycomic biological age and construct related indices. However, such methods are mainly targeted at specific systems and fail to provide an assessment specifically for the aging status of the kidneys.

[0005] Furthermore, existing biological age calculation techniques, such as epigenetic clocks, telomere length analysis, and transcriptomic or proteomic methods, generally suffer from limitations such as technical complexity, high cost, or difficulty in reflecting the aging state of specific organs. In particular, many methods rely heavily on statistical correlation when constructing models for biomarkers, lacking in-depth validation of the causal relationship between biomarkers and physiological states, which may affect the robustness of the models and the reliability of biological interpretations.

[0006] Therefore, there is a need in this field for a new technical solution that can construct an assessment model that specifically and accurately reflects the physiological aging state of the kidneys based on stable and quantifiable biological markers, combined with efficient machine learning and causal inference methods, in order to make up for the shortcomings of existing technologies. Summary of the Invention

[0007] This invention proposes a method for calculating the biological age of the kidney based on glycosylation markers and integrated machine learning, in order to solve the problems existing in the prior art.

[0008] To achieve the above objectives, this invention provides a method for calculating kidney biological age based on glycosylation biomarkers and integrated machine learning, comprising the following steps:

[0009] Obtain N-glycosylation profile data of individual plasma immunoglobulin G;

[0010] Based on the N-glycosylation profile data, processing is performed to obtain preprocessed glycosylation data;

[0011] Based on the glycosylation data and combined with preset kidney function reference parameters, an adaptive group elastic network is used for feature dimensionality reduction to screen out candidate glycosylation structures related to kidney function.

[0012] Based on the candidate glycosylation structures, an integrated machine learning algorithm was used to screen and determine the target glycosylation structures associated with abnormal kidney function.

[0013] Based on the target glycosyl structure, Mendelian randomization method is used to perform causal inference to determine the final glycosyl structure;

[0014] Based on the final glycosylation structure and individual age information, a multivariate logistic regression model is constructed to calculate the individual's renal glycosylation biological age.

[0015] Optionally, obtaining the N-glycosylation profile data of an individual's plasma immunoglobulin G includes:

[0016] The N-glycosylation profile of the immunoglobulin G was obtained by detecting its N-glycosylation.

[0017] Based on the N-glycosylation pattern, the proportion of the chromatographic peak area of ​​each glycosyl group to the total chromatographic peak area is determined as the glycosyl data.

[0018] Optionally, the processing to obtain preprocessed glycosyl data includes:

[0019] The glycosyl data were corrected using a batch effect correction method.

[0020] The corrected glycosylation data were subjected to normalization transformation to obtain preprocessed glycosylation data that conforms to a normal distribution.

[0021] Optionally, the feature dimensionality reduction using an adaptive grouped elastic network includes:

[0022] Use the preprocessed glycosyl data as input;

[0023] By using an adaptive group elastic network model to remove irrelevant variables, a sparse feature set is obtained as the candidate glycosylation structure.

[0024] Optionally, the screening using an integrated machine learning algorithm includes:

[0025] Construct an ensemble machine learning model that includes support vector machines, random forests, XGBoost, LightGBM, and deep neural networks;

[0026] The ensemble machine learning model is trained based on the candidate glycosylation structures.

[0027] Based on the model performance evaluation results, the target glycosyl structure is determined from the candidate glycosyl structures.

[0028] Optionally, the use of Mendelian randomization for causal inference includes:

[0029] Obtain individual genotype data and perform quality control;

[0030] Genetic variations associated with the target glycosyl structure were screened as instrumental variables based on genome-wide association analysis;

[0031] The association between the target glycosylation structure and the preset kidney function reference parameters was evaluated using Mendelian randomization analysis.

[0032] The target glycosylation structure with causal relationship is determined as the final glycosylation structure.

[0033] Optionally, the construction of a multivariate logistic regression model to calculate an individual's renal glycobiological age includes:

[0034] Based on the final glycosylation structure and individual calendar age, a multivariate logistic regression model is constructed to obtain the corresponding regression coefficients;

[0035] Based on the regression coefficients, a formula for calculating the renal glycosylation biological age is established, with the final glycosylation structure measurement and individual calendar age as input variables.

[0036] The individual's final glycosylation structure data and calendar age are substituted into the calculation formula to obtain the renal glycosylation biological age.

[0037] Optionally, the formula for calculating the renal glycosyl biological age is:

[0038] GlyKage=(β1*(X1-m1)+β2*(X2-m2)+β3*(X3-m3)+β4*(X4-m4)) / β0+age;

[0039] Where X1, X2, X3, and X4 are the relative content ratios of the four final glycosyl structures, m1, m2, m3, and m4 are the average values ​​of the four final glycosyl structures in the age-matched population, β0, β1, β2, β3, and β4 are the regression coefficients of the multivariate logistic regression model, and age is the calendar age.

[0040] Compared with the prior art, the present invention has the following advantages and technical effects:

[0041] This invention provides a method for calculating renal glycosylation biological age based on plasma IgG N-glycosylation profiles, constructing the first specific biological age assessment model for the kidney organ, which can accurately reflect the physiological changes in kidney status. This method employs an adaptive group elastic network and ensemble machine learning algorithms to robustly screen key features from high-dimensional glycosylation data and uses Mendelian randomization to verify causal associations, significantly improving the model's biological interpretability and predictive robustness. Based on stable and quantifiable glycomic biomarkers, this invention avoids the limitations of high cost and complexity, facilitating widespread application. The calculated renal glycosylation biological age and its difference from calendrical age can quantitatively assess the rate and trend of changes in renal physiological status, providing a new tool for personalized and dynamic organ age assessment. Attached Figure Description

[0042] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0043] Figure 1 This is a chromatogram showing the results of ultra-high performance liquid chromatography (UHPLC) analysis of plasma IgG N-glycosyl groups according to an embodiment of the present invention.

[0044] Figure 2 This is a diagram showing the results of adaptive elastic mesh screening of plasma IgG N-glycosyl structures according to an embodiment of the present invention.

[0045] Figure 3 This is a diagram showing the cross-validation results of the adaptive group elastic network according to an embodiment of the present invention;

[0046] Figure 4 This is a diagram showing the age distribution of renal glycosyl organisms according to an embodiment of the present invention;

[0047] Figure 5 This is a flowchart illustrating the calculation of renal glycosyl biological age according to an embodiment of the present invention. Detailed Implementation

[0048] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0049] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0050] Example 1

[0051] like Figure 5 As shown, this embodiment provides a method for calculating kidney biological age based on glycosylation biomarkers and integrated machine learning, including the following steps:

[0052] Obtain N-glycosylation profile data of individual plasma immunoglobulin G;

[0053] Based on the N-glycosylation profile data, the data was processed to obtain preprocessed glycosylation data;

[0054] Based on glycosylation data and combined with preset kidney function reference parameters, an adaptive group elastic network was used for feature dimensionality reduction to screen out candidate glycosylation structures related to kidney function.

[0055] Based on candidate glycosylation structures, an integrated machine learning algorithm was used to screen and identify target glycosylation structures associated with abnormal kidney function.

[0056] Based on the target glycosyl structure, Mendelian randomization method is used to perform causal inference to determine the final glycosyl structure with causal relationship;

[0057] Based on the final glycosylation structure and individual age information, a multivariate logistic regression model was constructed to calculate the individual's renal glycosylation biological age.

[0058] Further, obtain individual plasma immunoglobulin G N-glycosylation profile data, including:

[0059] The N-glycosylation profile of the immunoglobulin G was obtained by detecting its N-glycosylation.

[0060] Based on the N-glycosylation pattern, the proportion of the chromatographic peak area of ​​each glycosyl group to the total chromatographic peak area is determined as the glycosyl data.

[0061] Specifically, a hydrophilic-interactive high-performance liquid chromatography-ultra-high-performance liquid chromatography (HPLC-UHPLC) method was used for the quantitative detection of plasma IgG glycosyl modifications. The method mainly consists of four steps: separation and purification of IgG using a 96-well Protein G extraction plate; extraction and purification of glycans using PNGase F enzyme; 2-AB fluorescent labeling of the glycans; and quantitative detection using an UPLC fluorescence detector. The specific experimental procedure is as follows:

[0062] ①IgG isolation;

[0063] IgG was separated using a 96-well Protein G extraction plate.

[0064] ② IgG concentration detection;

[0065] The purity of the isolated IgG molecules was determined by sodium dodecyl sulfate polyacrylamide gel electrophoresis (SDS-PAGE), using Precision Plus Protein All Blue Standards as molecular weight markers, stained with GelCode Blue, and imaged using a VersaDoc imaging system. Finally, the IgG concentration was determined using a BioSpec-nano scanner.

[0066] ③ IgG deglycosylation;

[0067] The N-glycosyl group is removed from IgG using glycosidases (N-Glycosidase F, PNGase F).

[0068] ④ Glycosyl labeling and purification;

[0069] The glycosyl groups after enzyme digestion were further labeled with 2-AB (2-Aminobenzamide) fluorescent labeling.

[0070] ⑤UPLC glycosyl detection and structural identification;

[0071] Waters ACQUITY Ultra Performance Liquid Chromatography (UPLC) instrument was used to separate 2-AB fluorescently labeled N-glycosyl groups based on the principle of hydrophilic interaction. Following a conventional integrated algorithm, the detection data was automatically processed, and peaks for each glycosyl group were plotted according to the peak area and retention time of the detected sugar chains. The final chromatogram was divided into 24 peaks (GP1-GP24) using the same method (e.g., Figure 1 ).

[0072] The N-glycosyl structure of IgG is represented by the following letters and numbers commonly used in glycosylation research: F represents a core fucose linked to the core N-acetylglucosamine; Mx represents “x” mannose linked to the core N-acetylglucosamine; Ax represents “x” branches (N-acetylglucosamine) linked to the trimannose core; A2 represents N-acetylglucosamine linked to both branches of the core mannose; B represents bisecting N-acetylglucosamine linked to the core mannose; Gx represents “x” galactose linked to the branches; [3]G1 and [6]G1 represent galactose linked to α1-3 and α1-6 mannose, respectively; Sx represents “x” sialic acid linked to galactose. Table 1 shows the meaning of the 24 glycosyl peaks measured in the experiment and the structural meaning of the glycosyl peaks. See Table 2 for the structural diagrams.

[0073] Table 1

[0074] Sugar peak code Glycosyl peak structure Structural meaning GP1 FA1 Core fucoidanized dual-antenna glycan GP2 A2 Two-antenna-type sugar chains GP3 A2B Two-antenna type with branched sugar chains GP4 FA2 Core fucoidanized two-elemental glycan chain GP5 M5 Pentamyl sugar chain GP6 FA2B The core fucoidan is a two-sided, branched glycan. GP7 A2G1 Two antenna-shaped sugar chains, each containing a galactose molecule. GP8 FA2[6]G1 The core fucoidan has a two-sided antenna-shaped glycan with a galactose glycoside at the 6-position branch. GP9 FA2[3]G1 The core fucoidan has a two-sided antenna-shaped glycan with a galactose molecule at the 3-position branch. GP10 FA2[6]BG1 The core fucoidan has a two-sided antenna-shaped glycan with a galactose and a branch at position 6. GP11 FA2[3]BG1 The core fucoidan has a two-sided antenna-shaped glycan with a galactose and a branch at position 3. GP12 A2G2 Two antenna-shaped sugar chains, each containing two galactoses GP13 A2BG2 Two-antenna-shaped sugar chains, with two galactoses and branches. GP14 FA2G2 The core fucoidan is a two-pronged galactose chain with two galactoses. GP15 FA2BG2 The core fucoidan is a two-sided antennae-shaped glycan with two galactoses and branches. GP16 FA2G1S1 The core fucoidan is a two-sided antennae-shaped glycan with a galactose and a sialic acid. GP17 A2G2S1 Two antenna-shaped sugar chains, containing two galactoses and one sialic acid. GP18 FA2G2S1 The core fucoidan is a two-antenna-shaped glycan with two galactoses and one sialic acid. GP19 FA2BG2S1 The core fucoidan is a two-sided antennae-shaped glycan with two galactoses, a branch, and a sialic acid. GP20 FA2FG2S1 The core fucoidan is a two-sided antennae-shaped glycan containing fucose, two galactoses, and one sialic acid. GP21 A2G2S2 Two antenna-shaped sugar chains, containing two galactoses and two sialic acids. GP22 A2BG2S2 Two antenna-shaped sugar chains, with two galactoses, branches, and two sialic acids. GP23 FA2G2S2 The core fucoidan is a two-antenna-shaped glycan with two galactoses and two sialic acids. GP24 FA2BG2S2 The core fucoidan is a two-antenna-shaped glycan with two galactoses, branches, and two sialic acids.

[0075] Table 2

[0076] Sugar peak Sub-peak structure Structural diagram GP1 FA1 GP2 A2 GP3 A2B GP4 FA2 GP5 M5 GP6 FA2B GP7 A2G1 GP8 FA2[6]G1 GP9 FA2[3]G1 GP10 FA2[6]BG1 GP11 FA2[3]BG1 GP12 A2G2 GP13 A2BG2 GP14 FA2G2 GP15 FA2BG2 GP16 FA2G1S1 GP17 A2G2S1 GP18 FA2G2S1 GP19 FA2BG2S1 GP20 FA2FG2S1 GP21 A2G2S2 GP22 A2BG2S2 GP23 FA2G2S2 GP24 FA2BG2S2

[0077] Different oligosaccharide residues are represented using graphic symbols, and these symbols are arranged according to the relationship between the oligosaccharide residues and their bonds to represent different glycosyl structures. A square represents N-acetylglucosamine, a circle represents mannose, a triangle represents fucose, a circle represents galactose, and a rhombus represents sialic acid.

[0078] The model for the biological age of kidney glycosylation is constructed as follows:

[0079] Feature dimensionality reduction using an adaptive group elastic network includes: taking the preprocessed glycosylation data as input; and using the adaptive group elastic network model to remove irrelevant variables to obtain a sparse feature set as the candidate glycosylation structure.

[0080] Specifically, the level of each glycosyl group is calculated by dividing the peak area under the total peak area by the chromatographic area under the total peak area. In other words, the content of each glycosyl group is expressed as a percentage. After obtaining the glycosyl data, batch effects were corrected for using the ComBat method in R software. The GlycanR program was used for quantitative calculations of glycosyl-derived structures. Z-score variation, logarithmic variation, and polynomial variation were used to preprocess the glycosyl data to ensure it conformed to a normal distribution.

[0081] Reference parameters for kidney function are set based on past experience data. The past experience data is obtained by calculating the individual's glomerular filtration rate (GFR) using serum creatinine, thereby assessing the individual's kidney function health status. If the estimated GFR is less than 60, kidney function is considered abnormal.

[0082] Furthermore, feature dimensionality reduction is performed using an adaptive group elastic network, including: taking the preprocessed glycosylation data as input; and using the adaptive group elastic network model to remove irrelevant variables to obtain a sparse feature set as the candidate glycosylation structure.

[0083] Specifically, the Adaptive Group Elastic-Net (AGENET) introduces a combination of adaptive L1 and L2 penalty terms into the mean squared error loss function. This allows for both sparse feature selection and group selection of relevant features, enabling it to handle highly correlated biomarkers in high-dimensional glycomics data, reducing the impact of multicollinearity, and improving model stability and interpretability. The Adaptive Elastic-Net model obtains a non-zero feature set by compressing the coefficients of irrelevant biomarkers to zero and eliminating irrelevant variables. For a given set of data (… , Biomarkers ( =1, …, The ANENT coefficients are shown in formula (1):

[0084]

[0085] biomarkers The adaptive weights are shown in formula (2):

[0086]

[0087] Furthermore, an ensemble machine learning algorithm is used for screening, including: constructing an ensemble machine learning model that includes support vector machines, random forests, XGBoost, LightGBM, and deep neural networks; training the ensemble machine learning model based on candidate glycosylation structures; and determining the target glycosylation structure from the candidate glycosylation structures based on the model performance evaluation results.

[0088] Specifically, the ensemble machine learning algorithm is constructed as follows:

[0089] This paper employs an ensemble machine learning algorithm, combining five methods—Support Vector Machine, Random Forest, XGBoost, LightGBM, and DNN—to construct a model. The model's performance is evaluated using its AUC (Average Acceptance Value). Ensemble machine learning is a machine learning paradigm that accomplishes learning tasks by building and combining multiple machine learning algorithms (called "base learners" or "individual learners"), rather than relying on a single model. The final decision is generated by the combined voting or averaging of these models, aiming to achieve superior generalization performance, stronger robustness, and higher prediction accuracy than any single model.

[0090] ① Support Vector Machines (SVMs) are powerful supervised learning models primarily used for classification tasks. Their core idea is to find a hyperplane that optimally separates data points from different classes. This "optimal" aspect is reflected in SVM's requirement not only for correct classification but also for maximizing the classification margin—that is, maximizing the sum of the distances from the boundary points of the two classes (called "support vectors") to the hyperplane. This maximizes the model's generalization ability and robustness. For linearly inseparable data, SVM uses kernel tricks to map the data to a high-dimensional feature space, making it linearly separable in the new space. Furthermore, by introducing the concept of soft margin, SVM allows a small number of samples to exist within the margin or even be misclassified, balancing model complexity and fault tolerance, and avoiding overfitting.

[0091] To handle noise and linearly inseparable cases, slack variables are usually introduced. The formula is shown in formula (3):

[0092]

[0093]

[0094] Among them, the penalty parameter It controls the tolerance for misclassification. The larger the value, the more the model tends to use a hard margin, making it prone to overfitting. The smaller the value, the more error is allowed, and the stronger the model's generalization ability may be.

[0095] ② The Random Forest model is a supervised learning algorithm that uses ensemble learning to construct multiple decision trees and integrates their predictions to obtain a more accurate and stable prediction result. It has the advantages of high accuracy, resistance to overfitting, strong interpretability, and applicability to large-scale datasets. The Random Forest model uses Bootstrap sampling and the CART algorithm to construct multiple decision trees and improves the model's generalization ability through ensemble prediction. In the Random Forest, each decision tree uses a randomly sampled sample of the same size as the original dataset as the training set. During the construction of the decision tree, the feature selection of each node is based on randomly selected features rather than selecting the optimal features, which enhances the stability of the model. Finally, the Random Forest combines the prediction results of multiple decision trees through averaging (for regression problems) or voting (for classification problems) to obtain the final prediction. The function formula for the averaging method is shown in formula (4), and the function formula for the voting method is shown in formula (5):

[0096]

[0097]

[0098] ③ XGBoost (eXtreme Gradient Boosting) is an ensemble learning algorithm based on gradient boosting and decision trees. It constructs multiple decision trees to form a model, with each decision tree correcting the prediction errors of the previous decision trees. During training, XGBoost improves model performance by optimizing the objective function. The objective function consists of two parts: a loss function L(θ) and a regularization term Ω(θ). The loss function measures the difference between the predicted values ​​of the constructed model and the true values; the regularization term measures the complexity of the model and prevents overfitting. XGBoost does not support the handling of missing data, so missing data needs to be deleted or imputed before model construction. This study uses multiple imputation to impute missing values. The objective function of XGBoost is shown in formulas (6) and (7):

[0099]

[0100]

[0101] ④ LightGBM (Light Gradient Boosting Machine) is a high-performance gradient boosting decision tree framework. Its core advantages are extremely fast training speed and low memory consumption, making it particularly suitable for large-scale datasets. Its success is attributed to two key technologies: a gradient-based one-sided sampling strategy, which efficiently approximates the full data gain by retaining data samples with large gradients, reducing computational cost; and a mutually exclusive feature bundling technique, which merges sparse mutually exclusive features into a single feature, significantly reducing the number of features. These innovations enable LightGBM to achieve training speedups of several times to tens of times while maintaining the same or even higher accuracy as XGBoost.

[0102] LightGBM's core objective aligns with the gradient boosting framework: adding a tree at each step. The objective function is minimized by including the loss function and the regularization term. The overall objective function is shown in Equation (8):

[0103]

[0104] The first gradient of the loss function. The second gradient of the loss function. Let be the predicted value brought by the t-th tree for sample i. is the regularization term for the t-th tree, used to control model complexity and prevent overfitting.

[0105] ⑤ Deep Neural Networks (DNNs) are machine learning models inspired by the neural networks of the human brain and are the core architecture of deep learning. Their "depth" is reflected in their inclusion of an input layer, an output layer, and multiple hidden layers in between. The power of DNNs stems from the hierarchical feature learning enabled by this layered structure: shallow hidden layers learn local features at lower levels, while deep hidden layers combine these features into more complex, abstract global patterns. By using backpropagation and gradient descent to minimize prediction errors, DNNs can automatically learn complex mapping relationships from massive amounts of data, achieving groundbreaking performance in fields such as image recognition and natural language processing.

[0106] The core of DNN consists of two processes: forward propagation and backward propagation.

[0107] The forward propagation formula describes the process of data flowing from the input layer to the output layer, calculating and obtaining the prediction result layer by layer. For the l-th layer, the calculation is shown in formulas (9) and (10):

[0108]

[0109]

[0110] This is the output of the (l-1)th layer (i.e., the input of this layer). It means inputting data X; and The weights and bias parameters of the l-th layer are the core parameters that the model needs to learn. This represents the linear calculation result for the l-th layer; A nonlinearity is introduced into the activation function of the l-th layer, enabling the network to learn complex patterns; This is the final output of the l-th layer.

[0111] Backpropagation and parameter update formulas are the process of adjusting network parameters in reverse based on prediction errors. The core is to use the chain rule to calculate the loss function J for each parameter (…). and The gradient of ).

[0112] The gradient is calculated as shown in equations (11) and (12):

[0113]

[0114]

[0115] It is the error term that propagates back from the output layer to the l-th layer through the chain rule, and m is the number of training samples.

[0116] The updated parameters are shown in formulas (13) and (14):

[0117]

[0118]

[0119] The learning rate controls the step size for parameter updates.

[0120] The dimensionality reduction process and cross-validation results of adaptive grouped elastic networks are as follows: Figure 2 and Figure 3 As shown, GP3, GP11, GP13, GP22, and GP24 were selected. The glycosyl structures selected after adaptive group elastic network dimensionality reduction and ensemble machine learning are GP3, GP11, GP13, and GP24, respectively.

[0121] Furthermore, causal inference is performed using Mendelian randomization, including: obtaining individual genotype data and performing quality control; screening genetic variations related to the target glycosyl structure as instrumental variables based on genome-wide association analysis; determining the causal association between the target glycosyl structure and abnormal kidney function through Mendelian randomization analysis; and identifying the target glycosyl structure with causal association as the final glycosyl structure.

[0122] Specifically, independent sample Mendelian randomization was used to make causal inferences between the dimensionality-reduced glycosyl structure and CKD.

[0123] Instrumental variable screening first uses PLINK version 1.90 for quality control. Samples and SNPs that do not meet the quality control criteria will be deleted. The specific quality control process for genotype data is as follows:

[0124] ① SNP Call Rate: This is the percentage of a particular SNP in all samples. It requires that 98% of the samples have these SNPs. If the distribution of a particular SNP in the sample is less than 98%, the SNP should be deleted.

[0125] ② Sample Call Rate: This refers to the percentage of SNPs in each sample out of all SNPs. Generally, the sample call rate is 90%, meaning that the number of SNPs in each sample is no less than 98% of the total number of SNPs. If the number of SNPs in a sample does not reach 98%, that sample must be deleted.

[0126] ③ Sex check: This involves verifying the sex determination of a sample based on genotype data against the sample's actual sex data. If the two are inconsistent, the sample must be deleted.

[0127] ④ Sibling pairs identification: This mainly identifies whether there is a familial relationship between samples. It is calculated using the IBS matrix correlation. If two samples are related, one sample needs to be randomly deleted.

[0128] ⑤ Minimum allele frequency (MAF): This refers to the frequency of an uncommon allele in a given population. In association studies, a smaller MAF will reduce statistical power, leading to false negative results. Generally, a minimum allele frequency greater than 5% is required. If the MAF is less than 5%, all such SNPs should be deleted.

[0129] ⑥ Hardy-Weinberg equilibrium test (HWE): Hardy-Weinberg equilibrium means that in an infinitely large interbreeding population where mutations, migration, and selection do not occur, gene frequencies and genotype frequencies will remain constant from generation to generation. GWAS requires an HWE parameter of 1 × 10⁻⁶. -4 All SNPs that do not meet this requirement will be deleted.

[0130] ⑦ Heterozygosity Quality Control: In general, in natural populations, excessively high or low heterozygosity in individuals is abnormal and requires filtering based on heterozygosity. Deviation may indicate sample contamination or inbreeding. Individuals deviating from ±3 SD of the average heterozygosity rate in the sample should be removed.

[0131] After quality control of SNPs, significant loci were screened based on genome-wide association studies (GWAS). Population-level statistical analysis was then performed on genotypes and observable traits, and the results were determined based on a significance p-value (P < 5 × 10⁻⁶). -8 The genetic variations most likely to affect the trait are screened out, and genes associated with the trait variations and their estimated effects are identified.

[0132] Independent-sample Mendelian randomization uses a two-stage least squares regression model to quantitatively estimate the magnitude of the association effect between exposure factor X and outcome Y:

[0133] Establish a G-X regression model to obtain the predicted value (P) of the exposure factors.

[0134] Establish a P-Y regression model to obtain the regression equation between the predicted values ​​of exposure factors P and the outcome Y.

[0135] Independent-sample Mendelian randomization avoids the Beavis effect and confounding factors introduced by population stratification that may occur in two-sample Mendelian randomization, thus obtaining a more realistic estimate of the causal association between glycosylation biomarkers and renal function abnormalities. The glycosyl structures with causal associations are GP3, GP11, GP13, and GP24.

[0136] Furthermore, a multivariate logistic regression model is constructed to calculate the individual's renal glycosyl biological age, including: constructing a multivariate logistic regression model based on the final glycosylation structure and the individual's calendar age to obtain the corresponding regression coefficients; establishing a formula for calculating the renal glycosyl biological age based on the regression coefficients, wherein the formula uses the measured value of the final glycosylation structure and the individual's calendar age as input variables; and substituting the individual's final glycosylation structure data and calendar age into the calculation formula to obtain the renal glycosyl biological age.

[0137] The renal glycosyl biological age model was constructed based on N-glycosylation structure (GP3, GP11, GP13, GP24) and age, and a multivariate logistic regression model was built.

[0138] Furthermore, the β values ​​of each N-glycosyl group and age were obtained through multivariate logistic regression and used as coefficients to further construct the formula for calculating the biological age of kidney glycosyl groups, as shown in formula (3):

[0139]

[0140] β0 is the coefficient of Age obtained from multivariate logistic regression; β1-β4 are the coefficients of GP3, GP11, GP13, and GP24 obtained from multivariate logistic regression; X1, X2, X3, and X4 are the percentages of the chromatographic peak area of ​​the four glycosyl structures GP3, GP11, GP13, and GP24 to the total chromatographic peak area; m1, m 2、 m 3、 m4 represents the mean values ​​of GP3, GP11, GP13, and GP24 for all samples within the sample age range of ±3 years; Age is the calendar age. The renal glycosyl biological age distribution obtained based on multivariate logistic regression is as follows: Figure 4 As shown.

[0141] The continuous variable predicted by the model is the individualized biological age (BioAge). The method for calculating the biological age acceleration value (AgeAccel) is: AgeAccel = BioAge – ChronAge, where ChronAge is the actual calendar age. A positive AgeAccel indicates accelerated renal biological aging, while a negative value indicates delayed renal aging.

[0142] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for calculating kidney biological age based on glycosylation biomarkers and ensemble machine learning, characterized in that, Includes the following steps: Obtain N-glycosylation profile data of individual plasma immunoglobulin G; Based on the N-glycosylation profile data, processing is performed to obtain preprocessed glycosylation data; Based on the glycosylation data and combined with preset kidney function reference parameters, an adaptive group elastic network is used for feature dimensionality reduction to screen out candidate glycosylation structures related to kidney function. Based on the candidate glycosylation structures, an integrated machine learning algorithm was used to screen and determine the target glycosylation structures associated with abnormal kidney function. Based on the target glycosyl structure, Mendelian randomization method is used to perform causal inference to determine the final glycosyl structure; Based on the final glycosylation structure and individual age information, a multivariate logistic regression model is constructed to calculate the individual's renal glycosylation biological age.

2. The method according to claim 1, characterized in that, The acquisition of N-glycosylation profile data of plasma immunoglobulin G from individuals includes: The N-glycosylation profile of the immunoglobulin G was obtained by detecting its N-glycosylation. Based on the N-glycosylation pattern, the proportion of the chromatographic peak area of ​​each glycosyl group to the total chromatographic peak area is determined as the glycosyl data.

3. The method according to claim 1, characterized in that, The process to obtain the preprocessed glycosyl data includes: The glycosyl data were corrected using a batch effect correction method. The corrected glycosylation data were subjected to normalization transformation to obtain preprocessed glycosylation data that conforms to a normal distribution.

4. The method according to claim 1, characterized in that, The feature dimensionality reduction using adaptive elastic mesh includes: Use the preprocessed glycosyl data as input; By using an adaptive group elastic network model to remove irrelevant variables, a sparse feature set is obtained as the candidate glycosylation structure.

5. The method according to claim 1, characterized in that, The screening using ensemble machine learning algorithms includes: Construct an ensemble machine learning model that includes support vector machines, random forests, XGBoost, LightGBM, and deep neural networks; The ensemble machine learning model is trained based on the candidate glycosylation structures. Based on the model performance evaluation results, the target glycosyl structure is determined from the candidate glycosyl structures.

6. The method according to claim 1, characterized in that, The use of Mendelian randomization for causal inference includes: Obtain individual genotype data and perform quality control; Genetic variations associated with the target glycosyl structure were screened as instrumental variables based on genome-wide association analysis; The association between the target glycosylation structure and the preset kidney function reference parameters was evaluated using Mendelian randomization analysis. The target glycosylation structure with causal relationship is determined as the final glycosylation structure.

7. The method according to claim 1, characterized in that, The multivariate logistic regression model is constructed to calculate the individual's renal glycosyl age, including: A multivariate logistic regression model was constructed with renal dysfunction as the dependent variable and the final glycosylation structure and individual calendar age as independent variables. Based on the regression coefficients of the multivariate logistic regression model, a formula for calculating the glycosyl biological age of the kidney is established. The individual's final glycosylation structure data and calendar age are substituted into the calculation formula to obtain the renal glycosylation biological age.

8. The method according to claim 7, characterized in that, The formula for calculating the renal glycosyl biological age is as follows: GlyKage=(β1*(X1-m1)+β2*(X2-m2)+β3*(X3-m3)+β4*(X4-m4)) / β0+age; Where X1, X2, X3, and X4 are the relative content ratios of the four final glycosyl structures, m1, m2, m3, and m4 are the average values ​​of the four final glycosyl structures in the age-matched population, β0, β1, β2, β3, and β4 are the regression coefficients of the multivariate logistic regression model, and age is the calendar age.