Biomarker combinations for the discrimination of diabetic nephropathy and type ii diabetes and uses thereof
Patent Information
- Application Number
- CN202310528590.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-10
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2043-05-10
AI Technical Summary
不同生物标志物各具优缺点,但是现有的生物标志物检测常缺乏足够的敏感性和特异性,临床应用价值受限
[0028]与现有技术相比,本发明第一方面提供的用于糖尿病肾病和II型糖尿病判别的生物标志物组合中,不论是代谢物组还是多肽组作为单组学标志物,经过验证,均对二型糖尿病和糖尿病肾病判别以及糖尿病肾病病程的判别具有足够的敏感性和特异性;如果将代谢物组和多肽组作为双组学标志物,则可以进一步提高准确性。基于生物标志物组合构建的不同机器学习算法模型,对于糖尿病肾病风险预测均具有较高的准确率。
Smart Images

Figure CN117147672B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biotechnology, and more specifically, to a combination of biomarkers for the differentiation of diabetic nephropathy and type 2 diabetes, and their application in the preparation of diagnostic kits for diabetic nephropathy and type 2 diabetes, and in the construction of predictive models for diabetic nephropathy and type 2 diabetes. Background Technology
[0002] Diabetic nephropathy is one of the most serious complications of diabetes and a leading cause of death in diabetic patients. It is a complex metabolic disease clinically characterized by proteinuria and a progressive decrease in glomerular filtration rate (GFR). Currently, the urine albumin-to-creatinine ratio (UACR) is the recommended indicator for diagnosing early kidney damage in the Chinese guidelines for the prevention and treatment of type 2 diabetes. However, a significant increase in UACR levels was not observed in 20% of type 2 diabetic patients with decreased glomerular filtration rate. Glomerular filtration rate directly reflects the extent of glomerular damage and overall renal function impairment, and is also a predictive indicator for some complications. However, its application is limited due to the difficulty in obtaining baseline GFR information and the fact that GFR changes usually only appear in the late stages of renal insufficiency.
[0003] More than 30 biomarkers related to kidney injury have been explored and discovered in recent years, mainly including: (1) biomarkers reflecting glomerular dysfunction, such as cystatin C; and (2) biomarkers reflecting renal tubular damage, such as β2-microglobulin and α1-microglobulin. Different biomarkers have their own advantages and disadvantages, but existing biomarker detection often lacks sufficient sensitivity and specificity, thus limiting their clinical application value.
[0004] Therefore, discovering new biomarkers for the analysis of diabetic nephropathy and obtaining intermediate results information is of great significance for guiding the early treatment of diabetic nephropathy and improving prognosis. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide a combination of biomarkers and their applications, which have sufficient sensitivity and specificity for the differentiation of diabetic nephropathy and type II diabetes, providing intermediate result information or reference for clinical applications. The specific scheme is as follows:
[0006] In a first aspect, the present invention provides a combination of biomarkers for the differentiation of diabetic nephropathy and type II diabetes, comprising a metabolite group and / or a polypeptide group; the metabolite group comprising: adenine, arginine, cytosine, gentic acid, histidine, imidazole lactate, aspartic acid, proline, pyroglutamic acid, threonine, malonic acid, tryptophan, uracil, 3-hydroxy-an-aminobenzoic acid, and proline betaine; the polypeptide group comprising: polypeptides with amino acid sequences as shown in SEQ ID NO: 1~10.
[0007] In a second aspect, the present invention also provides a biomarker combination for the differentiation of diabetic nephropathy and type II diabetes, comprising a first metabolite group and a first polypeptide group, wherein the first metabolite group comprises: adenine, arginine, cytosine, gentic acid, histidine, imidazole lactate, aspartic acid, proline, pyroglutamic acid, threonine, malonic acid, tryptophan and uracil; and the first polypeptide group comprises: polypeptides with amino acid sequences as shown in SEQ ID NO: 1-8.
[0008] Preferably, it further includes a second metabolite group and a second polypeptide group, wherein the second metabolite group includes: 3-hydroxy-o-aminobenzoic acid, histidine, malonic acid, proline, proline betaine and uracil; and the second polypeptide group includes: polypeptides with amino acid sequences as shown in SEQ ID NO: 6~10.
[0009] Thirdly, the present invention also provides the application of the above-mentioned combination of biomarkers for the differentiation of diabetic nephropathy and type II diabetes in the preparation of a kit for the differentiation of diabetic nephropathy and type II diabetes.
[0010] The kit may consist of: a list of biomarkers, silicon micro / nano particles, a metabolic chip, a sample incubation solution, a buffer solution, and a matrix solution.
[0011] Fourthly, the present invention also provides the application of the combination of biomarkers described in the first aspect for the differentiation of diabetic nephropathy and type II diabetes in the construction of predictive models for diabetic nephropathy and type II diabetes.
[0012] Preferably, the method for constructing the predictive model for diabetic nephropathy and type II diabetes includes:
[0013] Urine samples were collected from patients with type 2 diabetes, early-stage diabetic nephropathy, and late-stage diabetic nephropathy.
[0014] Metabolites and peptides in urine samples from different types of patients are detected, and differentially expressed metabolites and peptides are screened as biomarkers, wherein the biomarkers include the combination of biomarkers for diabetic nephropathy risk assessment provided in the first aspect of the present invention.
[0015] Urine samples from patients with type 2 diabetes, early diabetic nephropathy, and late diabetic nephropathy were randomly selected as training sets. Different machine learning algorithms were used to construct corresponding prediction models, and the critical value of the prediction model was obtained based on the ROC curve.
[0016] The predictive performance of the prediction models is evaluated, and the prediction model with the best predictive performance is selected as the risk prediction model for diabetic nephropathy.
[0017] Preferably, the predictive model for diabetic nephropathy and type II diabetes is a model constructed based on the LASSO regression algorithm.
[0018] Fourthly, the present invention also provides the application of the combination of biomarkers described in the second aspect for the differentiation of diabetic nephropathy and type II diabetes in the construction of predictive models for diabetic nephropathy and type II diabetes.
[0019] Preferably, the predictive model for diabetic nephropathy and type II diabetes includes:
[0020] The first prediction step includes a first risk prediction function, which takes the Inten of the markers in the first metabolite group and the first polypeptide group in the test sample as input values to the first risk prediction function and calculates a first risk prediction value; if the first risk prediction value of the test sample is greater than or equal to a first threshold value, then there is a risk of diabetic nephropathy.
[0021] The second prediction step includes a second risk prediction function, which uses the Inten of the markers in the second metabolite group and the second polypeptide group in the test sample with the risk of diabetic nephropathy as the input value of the second risk prediction function to calculate a second risk prediction value; if the second risk prediction value of the test sample is greater than or equal to a second threshold value, then the sample has the risk of advanced diabetic nephropathy.
[0022] Where Inten is the normalized value of the mass spectrometry signal peak of the marker.
[0023] In some specific embodiments of the present invention, the first risk prediction function is:
[0024] First risk prediction value = 0.61 – 1.22 × 10 -5 Inten (adenine) + 9.18×10 -8 Inten (arginine) + 7.20 × 10 -7 Inten (cytosine) – 6.16 × 10 -7 Inten (gentic acid) – 3.35 × 10 -6 Inten (histidine) – 2.11 × 10 -6 Inten (imidazolium lactate) – 3.67 × 10 -5 Inten (aspartic acid) + 9.06×10 -5 Inten (proline) +1.40×10 -5 Inten (pyroglutamic acid) – 1.99 × 10 -5 Inten (threonine) – 5.13 × 10 -5 Inten (malonic acid) +1.18×10 -6Inten (tryptophan) – 1.12 × 10 -7 Inten (uracil) – 9.66 × 10 -5 Inten(SEQ ID NO: 1) +1.70×10 -4 Inten(SEQ ID NO: 2) - 1.85 × 10 -5 Inten(SEQ ID NO: 3) - 7.68 × 10 -5 Inten(SEQID NO:4)–4.27×10 -6 Inten(SEQ ID NO:7)–1.26×10 -5 Inten(SEQ ID NO: 5) + 6.74×10 - 6 Inten(SEQ ID NO: 8) + 4.38×10 -5 Inten (SEQ ID NO: 6);
[0025] The second risk prediction function is:
[0026] Second risk prediction value = –0.088 + 2.76 × 10 -5 Inten(3-hydroxy-o-aminobenzoic acid) – 1.22 × 10 - 5 Inten (histidine) – 3.80 × 10 -5 Inten (malonic acid) + 4.89×10 -5 Inten (proline) + 3.98×10 - 5 Inten (proline betaine) – 2.65 × 10 -6 Inten (uracil) + 1.00 × 10 -5 Inten(SEQ ID NO: 9) - 8.94 × 10 -7 Inten(SEQ ID NO: 7) + 2.64 × 10 -4 Inten(SEQ ID NO: 10)+1.02×10 -5 Inten(SEQ ID NO: 8) + 1.25 × 10 -4 Inten (SEQ ID NO: 6).
[0027] Preferably, the diabetic nephropathy and type II diabetes prediction model is used to predict the risk of atypical diabetic nephropathy.
[0028] Compared with existing technologies, the biomarker combination for distinguishing diabetic nephropathy and type II diabetes provided in the first aspect of this invention, whether using metabolomics or peptide groups as single-omics biomarkers, has been validated to have sufficient sensitivity and specificity for distinguishing type II diabetes and diabetic nephropathy, as well as for determining the course of diabetic nephropathy. If metabolomics and peptide groups are used as dual-omics biomarkers, the accuracy can be further improved. Different machine learning algorithm models constructed based on the biomarker combination all show high accuracy in predicting the risk of diabetic nephropathy.
[0029] The predictive model for diabetic nephropathy and type 2 diabetes, constructed based on a combination of biomarkers for differentiating diabetic nephropathy and type 2 diabetes provided in the second aspect of this invention, can accurately predict the risk and duration of diabetic nephropathy. Compared to existing eGFR index methods, the predictive model for diabetic nephropathy and type 2 diabetes provided by this invention is faster and more accurate, showing great potential for future diabetes health management. Furthermore, the predictive model for diabetic nephropathy and type 2 diabetes provided by this invention is particularly effective in predicting the risk of atypical diabetic nephropathy, solving the current challenge of identifying atypical diabetic nephropathy.
[0030] The biomarkers for differentiating diabetic nephropathy and type II diabetes provided by this invention, and their application in constructing predictive models for diabetic nephropathy and type II diabetes, are all methods for obtaining information as intermediate results, providing a reference for health management and clinical applications. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0032] Figure 1 This is a schematic diagram of the data acquisition process for urinary metabolites and polypeptides in Embodiment 1 of the present invention, wherein: (A) urinary polypeptides, (B) urinary metabolites;
[0033] Figure 2 The urinary metabolism and polypeptide fingerprint profiles of patients with type 2 diabetes and diabetic nephropathy in Example 1 of the present invention, wherein: (A) urinary metabolites, (B) urinary polypeptides;
[0034] Figure 3The intra-batch and inter-batch stability of urine metabolites and peptide data collected in Example 1 of the present invention, wherein: (A) intra-batch stability of peptides, (B) inter-batch stability of peptides, (C) intra-batch stability of metabolites, and (D) inter-batch stability of metabolites;
[0035] Figure 4 This is a performance comparison of different machine learning algorithms in binary classification prediction of the model set in Embodiment 3 of the present invention, wherein: (A) type 2 diabetes and diabetic nephropathy, (B) early diabetic nephropathy and late diabetic nephropathy, and (C) type 2 diabetes and late diabetic nephropathy.
[0036] Figure 5 This is a performance comparison of different machine learning algorithms in binary classification prediction on a validation set in Embodiment 3 of the present invention, wherein: (A) type 2 diabetes and diabetic nephropathy, (B) early diabetic nephropathy and late diabetic nephropathy, and (C) type 2 diabetes and late diabetic nephropathy.
[0037] Figure 6 The average accuracy of the risk prediction models constructed using metabolomics, peptideomics, and dual-omics data in Example 3 of this invention for binary classification prediction is given, where: (A) is the modeling set, and (B) is the validation set.
[0038] Figure 7 The accuracy of the risk prediction models constructed using metabolomics, peptidomics, and dual-omics data in Example 3 of this invention for predicting each group of type 2 diabetes and early diabetic nephropathy is shown in (A) Metabolomics, (B) Peptidomics, and (C) Dual-omics.
[0039] Figure 8 The accuracy of the risk prediction models constructed using metabolomics, peptidomics, and dual-omics data in Example 3 of this invention for predicting early diabetic nephropathy and late diabetic nephropathy groups, wherein: (A) metabolomics, (B) peptidomics, and (C) dual-omics;
[0040] Figure 9 The following are the prediction scores and confusion matrices of the modeling set and validation set samples in the first prediction step in Embodiment 4 of the present invention, wherein: (A) the prediction scores of the modeling set and validation set samples in the first prediction step, (B) the prediction confusion matrix of the modeling set, and (C) the prediction confusion matrix of the validation set.
[0041] Figure 10 The following are the prediction scores and confusion matrices of the modeling set and validation set samples in the second prediction step in Embodiment 4 of the present invention, wherein: (A) the prediction scores of the modeling set and validation set samples in the second prediction step, (B) the prediction confusion matrix of the modeling set, and (C) the prediction confusion matrix of the validation set.
[0042] Figure 11The overall prediction accuracy of each group of samples in the modeling set and validation set in the two-step prediction model in Embodiment 4 of the present invention is given by: (A) modeling set, (B) validation set;
[0043] Figure 12 In Example 4 of this invention, the results of a two-step prediction model for the modeling set and validation set samples are compared with the clinical eGFR index, wherein: (A) is the modeling set, and (B) is the validation set;
[0044] Figure 13 The prediction score of the risk prediction model trained in Embodiment 5 of the present invention for patients with atypical diabetic nephropathy;
[0045] Figure 14 The prediction results of patients with atypical diabetic nephropathy based on the models of metabolomics, peptidomics, and dual omics in Example 5 of the present invention are as follows: (A) metabolomics, (B) peptidomics, and (C) dual omics.
[0046] Figure 15 This is a simulation diagram based on Simulink in Embodiment 6 of the present invention, showing the time estimation for data preprocessing, including the time estimation for file conversion, peak reading, and data normalization.
[0047] Figure 16 This refers to the prediction module based on the LASSO regression algorithm pre-built in Simulink in Embodiment 6 of the present invention;
[0048] Figure 17 This is a simulated real-time prediction analysis process for 96 samples in the first and second prediction steps of Embodiment 6 of the present invention. Detailed Implementation
[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0050] Example 1: Data Acquisition of Urinary Metabolites and Peptides
[0051] The data collection process for urinary metabolites and peptides is as follows: Figure 1 As shown.
[0052] A total of 349 urine samples were collected, including 149 from patients with type 2 diabetes, 106 from patients with early-stage diabetic nephropathy, 55 from patients with late-stage diabetic nephropathy, and 39 from patients with atypical diabetic nephropathy who had normal urine albumin-to-creatinine ratios but abnormal microalbumin concentrations. All urine samples were midstream morning urine collected on an empty stomach. Patients avoided eating, drinking alcohol, or taking medications for 8 hours prior to sample collection.
[0053] (1) Collection of urine metabolite data
[0054] The method disclosed in patent CN202110804335.9 is used to extract urinary metabolites, and the specific steps include:
[0055] Top-contact extraction of urinary metabolites was performed using a metabolic chip. The metabolic chip was inverted and placed on the surface of a urine sample, with the tips of the nanowires contacting the sample to adsorb and extract the metabolites. After standing for 20 minutes and complete extraction, the remaining droplets on the surface were slowly dried with N2. The metabolic chip with the extracted and adsorbed metabolic molecules was then attached to a target plate using carbon conductive adhesive and stored in a dry vacuum environment for mass spectrometry analysis.
[0056] Urinary metabolic profiles were acquired using a matrix-assisted laser desorption / ionization time-of-flight mass spectrometer (MALDI-TOF MS) from Bruker Daltonics. The analysis and processing of the metabolic data were performed using Clinprotools software provided by Bruker Daltonics. The data acquisition process is as follows:
[0057] The aforementioned metabolic chip was attached to an aluminum plate and inserted into an ultrafleXtreme MALDI-MS / TOF instrument equipped with a 355 nm Nd:YAG laser beam. Data acquisition was performed in reflected negative ion mode, with the molecular weight range set to 20–350 Da, and the relative laser pulse energy for mass spectrometry acquisition set to 57% of the total energy. The lens voltage was set to 8.50 kV, and the voltages of ion source 1, ion source 2, reflector 1, and reflector 2 were set to 20.00, 17.75, 21.10, and 10.70 kV, respectively. The ion extraction time was 120 ns, the laser parameters were set to 4_large, and the total number of mass spectrometry shots for each sample was 2000.
[0058] (2) Urine polypeptide data collection
[0059] The method disclosed in patent CN201210252926.0 for enriching and capturing urinary peptides includes the following steps:
[0060] Centrifuged urine samples were added to wells containing silicon micro / nanoparticles pre-treated with water. After incubation with shaking for 15 min, the samples were centrifuged and the supernatant was discarded. The retained silicon micro / nanoparticles were washed with ultrapure water, centrifuged, resuspended in matrix dissolution solution, and then spotted. After the samples dried, 1 μL of CHCA matrix solution was added to each well. The spotted target plate was stored in a dry vacuum environment for peptide profiling.
[0061] Urinary peptide profiles were acquired using a ClinMS-Plat I MALDI-MS / TOF mass spectrometer developed by Hangzhou Huijian Technology Co., Ltd., and the peptide data analysis and processing were performed using ClinMS Analyzer software developed by Hangzhou Huijian Technology Co., Ltd. The data acquisition process is as follows:
[0062] Urinary peptide mass spectrometry was performed using a ClinMS-Plat I MALDI-MS / TOF mass spectrometer equipped with a 337 nm nitrogen laser. (Mass spectrometry data were acquired in linear positive ion mode, with a molecular weight range of 600–20000 Da. The voltages of the detector, repulsion electrode, extractor, and focusing electrode were set to 2.77, 20.00, 1.95, and 7.00 kV, respectively. The frequency was 60 Hz, and the delay time was 250 ns. The laser pulse energy was 41.8 μJ, and the total number of shots for each mass spectrum was 800.)
[0063] Urinary metabolism and polypeptide fingerprinting in patients with type 2 diabetes and diabetic nephropathy, such as Figure 2 As shown.
[0064] Data acquisition quality control: For each acquired raw spectrum, the number of peaks with an S / N ≥ 3 was set as the standard for judging spectrum quality; only spectra with more than 100 peaks were retained, and spectra with less than 100 peaks were discarded. For the entire experimental procedure, the relative standard deviation (relative standard deviation = standard deviation / mean) of the pooled quality control urine samples (all included urine samples were mixed in equal volumes) was used to ensure experimental consistency. Figure 3 As shown, the intra-batch relative standard deviation of the metabolic profile of the quality control urine samples in this embodiment was 10.17%, and the inter-batch relative standard deviation was 18.74%; the intra-batch relative standard deviation of the peptide profile of the quality control urine samples was 7.55%, and the inter-batch relative standard deviation was 9.37%, which met the consistency range (<30%), indicating that the experimental consistency was good.
[0065] Example 2: Analysis of urinary metabolites and peptide data
[0066] (1) Data preprocessing
[0067] Raw urine metabolomics and peptide profiles were processed using FlexAnalysis (Bruker Daltonics) and ClinMS Analyzer (Hangzhou Huijian Technology Co., Ltd.), respectively. Peaks with an S / N ≥ 3 were selected for subsequent statistical analysis. Normalization was performed using the cubic spline method under the affy algorithm package in R 3.5.2 software.
[0068] (2) Selection of characteristic metabolites and peptides in diabetic nephropathy
[0069] Differentially expressed urinary metabolites and peptides were screened using t-tests in MATLAB software, and p-values were corrected using the p.adjust function in R 3.5.2 software. Metabolites and peptides with p < 0.05 after correction using the Benjamini-Hochberg method are potential biomarkers for diabetic nephropathy.
[0070] (3) Identification of characteristic metabolites and polypeptides of diabetic nephropathy
[0071] Identification of Characteristic Metabolites: The structures of differentially identified metabolites were identified by matching secondary mass spectrometry data with the human metabolome database and metabolic standards. First, UPLC-MS / MS analysis provided the accurate molecular weights and secondary fragment peaks of metabolites in urine samples, which were used to identify metabolites through database searches. The relative error of the accurate molecular weights was controlled within 30 ppm to obtain a preliminary list of differentially identified metabolites. Subsequently, the preliminarily identified metabolites were validated using MALDI-TOF / TOF tandem mass spectrometry, matching metabolite standards with metabolites in urine samples, including matching the accurate molecular weights and secondary mass spectrometry fragment peaks of the metabolites. The identification results of characteristic metabolites are shown in Table 1.
[0072] Table 1. Identification results of characteristic metabolites
[0073] 103.0065 MH malonic acid 110.0363 MH Cytosine 111.0203 MH Uracil 114.0566 MH proline 118.0498 MH threonine 128.0343 MH Pyroglutamic acid 132.0305 MH Aspartic acid 134.0473 MH adenine 142.0875 MH Proline betaine 152.0356 MH 3-Hydroxy-o-aminobenzoic acid 153.0188 MH Gentian acid 154.0626 MH Histidine 155.0463 MH Imidazole lactic acid 173.1052 MH Arginine 203.0818 MH Tryptophan
[0074] Identification of Characteristic Peptides: Urinary peptides captured in porous silica microparticles were identified by nano-LC-ESI MS / MS analysis. Before injection, 10K ultrafiltration centrifuge tubes were used to remove macromolecules larger than 10K from the urine. The open pFind3 search engine was used to search the Swiss-Prot database for the exact molecular weight and secondary spectra of matching peptide molecules, with no enzymes selected in the digestion phase. The maximum deletion cleavage number was set to 3, and the error range for exact mass and fragment peaks was set to 20 ppm. Automated filtering was performed, and a list of peptides with a q value ≤ 0.01 was retained. The identification results of characteristic peptides are shown in Table 2.
[0075] Table 2. Identification results of characteristic peptides
[0076] 1047.361 DFSFLPQPP 1 1149.796 LVQEVTDFAK 2 1615.784 NELRVAPEEHPVLL 3 1790.735 SYELPDGQVITIGNER 4 1950.251 IIVDTYGGWGAHGGGAFSGK 5 2520.430 LMIEQNTKSPLFMGKVVNPTQK 6 1912.077 SGSVIDQSRVLNLGPITR 7 2386.211 MTVSTLVLGEGATEAEISMTSTR 8 1103.988 LPGVPLARPAL 9 2176.330 WNTDNTLGTEITVEDQLAR 10
[0077] The characteristic metabolites and peptides identified in this embodiment can be used as biomarkers to prepare kits for predicting the risk of diabetic nephropathy.
[0078] Example 3: Construction of a model for predicting the risk of diabetic nephropathy
[0079] (1) Constructing training and validation sets
[0080] From 349 urine samples, 75 samples from type 2 diabetes, 53 samples from early diabetic nephropathy, and 28 samples from late diabetic nephropathy were randomly selected as the training set (modeling set) for model building.
[0081] Another 154 cases from the remaining samples were selected as the validation set for blind selection testing. Among them, 74 cases were known to be from patients with type 2 diabetes, 53 cases were from patients with early diabetic nephropathy, and the remaining 27 cases were from people with late diabetic nephropathy.
[0082] The remaining 39 urine samples were from atypical diabetic nephropathy patients with normal urine albumin-to-creatinine ratios but abnormally high levels of microalbumin in their urine. These samples were used to test the model's ability to predict samples missed by clinical indicators.
[0083] (2) Evaluation of the predictive performance of different machine learning algorithms
[0084] Accuracy, F-metric, Kappa coefficient, and precision were used as metrics to evaluate model performance. Based on the feature metabolites and peptides screened in Example 2, the prediction performance of dual-omics models constructed using different machine learning algorithms, including Support Vector Machine (SVM), Decision Tree (DT), Naive Bayes (NB), Logistic Regression (logi), Linear Regression (LDA), Nearest Neighbor (KNN), and LASSO regression, was compared on the modeling and validation sets. The binary classifications included: risk prediction of type 2 diabetes and early diabetic nephropathy, risk prediction of type 2 diabetes and late diabetic nephropathy, and risk prediction of early diabetic nephropathy and late diabetic nephropathy.
[0085] from Figure 4 and Figure 5 As can be seen, the models built based on the LASSO regression algorithm all exhibit the best performance in binary classification prediction.
[0086] (3) Comparison of prediction results between single-omics and dual-omics
[0087] The predictive models constructed using single-omics (characteristic metabolites or characteristic peptides from Example 2 as biomarkers) or dual-omics (characteristic metabolites and characteristic peptides from Example 2 as biomarkers) data were evaluated for their predictive performance in binary classification on the validation set.
[0088] Algorithms including Support Vector Machine (SVM), Decision Tree (DT), Naive Bayes classification (NB), Logistic Regression (logi), Linear Regression (LDA), Nearest Neighbor Classification (KNN), and LASSO Regression were used for modeling, and the average accuracy was calculated.
[0089] like Figure 6 As shown, both the prediction models built from single-omics and dual-omics data have high prediction accuracy. However, compared with the prediction model built from single-omics data, the prediction model built from dual-omics data can significantly improve the prediction accuracy on both the training and validation sets.
[0090] The accuracy rates of models constructed using metabolomics, peptidomics, and dual-omics data in predicting early diabetic nephropathy and detecting the progression of diabetic nephropathy in each group are shown in the following results: Figure 7 and Figure 8 As shown, the prediction accuracy of single-omics models all exceeded 0.6, while the prediction accuracy of dual-omics models all exceeded 0.85, indicating that dual-omics models significantly improved prediction accuracy.
[0091] Example 4: Predicting the risk of diabetic nephropathy based on the LASSO regression algorithm
[0092] Based on the evaluation results of the performance of different machine learning algorithms in Example 3, this example establishes a two-step urine prediction model based on the LASSO regression algorithm to achieve risk prediction of diabetic nephropathy.
[0093] In the first prediction step, a biomarker group including 21 metabolites and peptides is used to predict the risk of diabetic nephropathy, specifically for type 2 diabetes and diabetic nephropathy. The biomarker group specifically includes: adenine, arginine, cytosine, gentic acid, histidine, imidazole lactate, aspartic acid, proline, pyroglutamic acid, threonine, malonic acid, tryptophan, uracil, and SEQ ID NOs: 1-8. The specific first risk prediction function is as follows:
[0094] First risk prediction value = 0.61 – 1.22 × 10 -5 Inten (adenine) + 9.18×10 -8 Inten (arginine) + 7.20 × 10 -7 Inten (cytosine) – 6.16 × 10 -7 Inten (gentic acid) – 3.35 × 10-6 Inten (histidine) – 2.11 × 10 -6 Inten (imidazolium lactate) – 3.67 × 10 -5 Inten (aspartic acid) + 9.06×10 -5 Inten (proline) +1.40×10 -5 Inten (pyroglutamic acid) – 1.99 × 10 -5 Inten (threonine) – 5.13 × 10 -5 Inten (malonic acid) +1.18×10 -6 Inten (tryptophan) – 1.12 × 10 -7 Inten (uracil) – 9.66 × 10 -5 Inten(SEQ ID NO: 1) +1.70×10 -4 Inten(SEQ ID NO: 2) - 1.85 × 10 -5 Inten(SEQ ID NO: 3) - 7.68 × 10 -5 Inten(SEQID NO:4)–4.27×10 -6 Inten(SEQ ID NO:7)–1.26×10 -5 Inten(SEQ ID NO: 5) + 6.74×10 - 6 Inten(SEQ ID NO: 8) + 4.38×10 -5 Inten (SEQ ID NO: 6).
[0095] If the first risk prediction value is greater than or equal to the first threshold value, the urine sample carries a risk of diabetic nephropathy. The first threshold value is preferably 0.482.
[0096] Figure 9 The figure shows the prediction scores and confusion matrix for the modeling and validation sets in the first prediction step, with a first critical value of 0.482. As can be seen from the figure, the prediction results for the training samples are: 64 out of 75 cases of type 2 diabetes were predicted accurately, with a specificity of 85.33%; and 72 out of 81 cases of diabetic nephropathy were predicted accurately, with a sensitivity of 88.89%. The prediction results for the validation samples are: 65 out of 74 cases of type 2 diabetes were predicted accurately, with a specificity of 87.84%; and 74 out of 80 cases of diabetic nephropathy were predicted accurately, with a sensitivity of 92.50%.
[0097] In the first prediction step, urine samples identified as having a risk of diabetic nephropathy are then subjected to a second prediction step. In this second step, a biomarker set including 11 metabolites and peptides is used to assess the disease progression of diabetic nephropathy and predict the risk of early and late-stage diabetic nephropathy. The biomarker set specifically includes: 3-hydroxy-an-aminobenzoic acid, histidine, malonic acid, proline, proline betaine, uracil, and SEQ ID NOs: 6-10. The specific second risk prediction function is as follows:
[0098] Second risk prediction value = –0.088 + 2.76 × 10 -5 Inten(3-hydroxy-o-aminobenzoic acid) – 1.22 × 10 - 5 Inten (histidine) – 3.80 × 10 -5 Inten (malonic acid) + 4.89×10 -5 Inten (proline) + 3.98×10 - 5 Inten (proline betaine) – 2.65 × 10 -6 Inten (uracil) + 1.00 × 10 -5 Inten(SEQ ID NO: 9) - 8.94 × 10 -7 Inten(SEQ ID NO: 7) + 2.64 × 10 -4 Inten(SEQ ID NO: 10)+1.02×10 -5 Inten(SEQ ID NO: 8) + 1.25 × 10 -4 Inten (SEQ ID NO: 6).
[0099] If the second risk predictive value is greater than or equal to the second threshold, the urine sample carries a risk of advanced diabetic nephropathy. The second threshold is preferably 0.369.
[0100] Figure 10 The prediction scores and confusion matrix for the modeling and validation sets in the second prediction step are shown, with a second critical value of 0.369. As can be seen from the figure, among the 72 training samples, 38 out of 44 cases of early-stage diabetic nephropathy were accurately predicted, and 24 out of 28 cases of late-stage diabetic nephropathy were accurately predicted, for an overall accuracy of 86.11%. Among the 74 validation samples, 39 out of 47 cases of early-stage diabetic nephropathy were accurately predicted, and all 27 cases of late-stage diabetic nephropathy were accurately predicted, for an overall accuracy of 89.19%.
[0101] Figure 11The figure shows the overall prediction accuracy of each group of samples in the modeling and validation sets using the two-step prediction model. As can be seen from the figure, for the 156 urine samples in the modeling set, after prediction analysis using the two-step prediction model, 85.3% of the samples in the type 2 diabetes group were accurately predicted, 71.7% of the samples in the early diabetic nephropathy group were accurately predicted, and 85.7% of the samples in the late diabetic nephropathy group were accurately predicted. For the 154 urine samples in the validation set, after prediction analysis using the two-step prediction model, 87.8% of the samples in the type 2 diabetes group were accurately predicted, 73.6% of the samples in the early diabetic nephropathy group were accurately predicted, and all samples in the late diabetic nephropathy group were accurately predicted.
[0102] Figure 12 The figure shows a comparison between the risk prediction results of the two-step prediction model and the discrimination results of the clinical indicator eGFR. As can be seen from the figure, the two-step prediction model can significantly distinguish between type 2 diabetes and diabetic nephropathy that could not be distinguished by eGFR in the modeling set and validation set.
[0103] Example 5: Risk Prediction of Atypical Diabetic Nephropathy
[0104] In patients with atypical diabetic nephropathy, the urine albumin-to-creatinine ratio was normal, but the concentration of trace amounts of urinary albumin was abnormal.
[0105] Using the two-step prediction model trained in Example 4, predictive analysis was performed on urine samples from 39 patients with atypical diabetic nephropathy. The results are as follows: Figure 13 As shown, the model's prediction accuracy reached 79.5%, confirming its potential in predicting the risk of atypical diabetic nephropathy.
[0106] Based on the LASSO regression algorithm, risk prediction models constructed using metabolomics, peptidomics, and dual-omics data were tested for their accuracy in predicting the risk of atypical diabetic nephropathy. The results are as follows: Figure 14 As shown, the dual-omics model has a higher prediction accuracy than the single-omics model, demonstrating a significant advantage.
[0107] Example 6: Simulation of a two-step prediction model for predicting the risk of diabetic nephropathy
[0108] MATLAB was used to simulate the risk prediction of diabetic nephropathy using a two-step prediction model. Simulink is a software package for modeling, simulating, and analyzing dynamic systems. Simulink has a built-in prediction module based on the LASSO regression algorithm. After normalizing the mass spectrometry detection data, it was imported into the Sample module. Clicking the Run button and then Scope outputs the predicted value.
[0109] A simulation of the entire workflow from data collection to final prediction was conducted using 96 urine samples. Figure 15 The time estimates for file conversion, peak reads, and data normalization are shown, all of which can be completed in seconds. Subsequently, near real-time predictions are made using a pre-deployed stepwise LASSO regression model in Simulink, such as... Figure 16 As shown.
[0110] Figure 17 The simulation demonstrates the real-time predictive analysis process of the first and second steps of prediction for 96 urine samples, requiring an average of only 0.2 seconds per sample. The simulation shows that the two-step prediction model can be applied to predict the risk of early diabetic nephropathy and the progression of diabetic nephropathy.
[0111] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A combination of biomarkers for differentiating diabetic nephropathy and type II diabetes, characterized in that, Including the first metabolite group and the first polypeptide group, The first metabolite group includes: adenine, arginine, cytosine, gentic acid, histidine, imidazole lactate, aspartic acid, proline, pyroglutamic acid, threonine, malonic acid, tryptophan, and uracil. The first polypeptide group includes polypeptides with amino acid sequences as shown in SEQ ID NO: 1~8.
2. The biomarker combination as described in claim 1, characterized in that, It also includes the second metabolite group and the second polypeptide group. The second metabolite group includes: 3-hydroxy-an-aminobenzoic acid, histidine, malonic acid, proline, proline betaine, and uracil; The second polypeptide group includes polypeptides with amino acid sequences as shown in SEQ ID NO: 6~10.
3. The use of the combination of biomarkers for the differentiation of diabetic nephropathy and type II diabetes as described in claim 1 or 2 in the preparation of a kit for the differentiation of diabetic nephropathy and type II diabetes.
4. The application of the combination of biomarkers for the differentiation of diabetic nephropathy and type II diabetes as described in claim 1 or 2 in the construction of predictive models for diabetic nephropathy and type II diabetes.
5. The application as described in claim 4, characterized in that, The predictive models for diabetic nephropathy and type 2 diabetes include: The first prediction step includes a first risk prediction function, which takes the Inten of the markers in the first metabolite group and the first polypeptide group as described in claim 3 in the test sample as the input value of the first risk prediction function, and calculates a first risk prediction value; if the first risk prediction value of the test sample is greater than or equal to a first threshold value, then there is a risk of diabetic nephropathy. The second prediction step includes a second risk prediction function, which uses the Inten of the markers in the second metabolite group and the second polypeptide group as described in claim 3 in the test sample with the risk of diabetic nephropathy as the input value of the second risk prediction function to calculate a second risk prediction value; if the second risk prediction value of the test sample is greater than or equal to a second threshold value, then it has the risk of advanced diabetic nephropathy. Where Inten is the normalized value of the mass spectrum signal peak of the marker; The first risk prediction function is: First risk prediction value = 0.61 – 1.22 × 10 -5 Inten (adenine) + 9.18×10 -8 Inten (arginine) +7.20×10 -7 Inten (cytosine) – 6.16 × 10 -7 Inten (gentic acid) – 3.35 × 10 -6 Inten (histidine) – 2.11 × 10 -6 Inten (imidazolium lactate) – 3.67 × 10 -5 Inten (aspartic acid) + 9.06×10 -5 Inten (proline) + 1.40×10 -5 Inten (pyroglutamic acid) – 1.99 × 10 -5 Inten (threonine) – 5.13 × 10 -5 Inten (malonic acid) + 1.18 × 10 -6 Inten (tryptophan) – 1.12 × 10 -7 Inten (uracil) – 9.66 × 10 -5 Inten(SEQ ID NO: 1) + 1.70×10 -4 Inten(SEQ ID NO: 2) - 1.85 × 10 -5 Inten(SEQ ID NO: 3) - 7.68 × 10 -5 Inten(SEQ ID NO: 4)–4.27×10 -6 Inten(SEQ ID NO:7)–1.26×10 -5 Inten(SEQ ID NO: 5) + 6.74×10 -6 Inten(SEQ ID NO: 8) + 4.38×10 -5 Inten (SEQ ID NO: 6); The second risk prediction function is: Second risk prediction value = –0.088 + 2.76 × 10 -5 Inten(3-hydroxy-o-aminobenzoic acid) – 1.22 × 10 -5 Inten (histidine) – 3.80 × 10 -5 Inten (malonic acid) + 4.89×10 -5 Inten (proline) + 3.98×10 -5 Inten (proline betaine) – 2.65 × 10 -6 Inten (uracil) + 1.00 × 10 -5 Inten(SEQ ID NO: 9) - 8.94 × 10 - 7 Inten(SEQ ID NO: 7) + 2.64 × 10 -4 Inten(SEQ ID NO: 10)+1.02×10 -5 Inten(SEQ ID NO: 8) + 1.25 × 10 -4 Inten (SEQ ID NO: 6).
Citation Information
Patent Citations
Kit for detecting low-abundance and low-molecular-weight protein spectrum
CN102788833A
Laser desorption ionization mass spectrometry kit for rapidly detecting small molecular substances in liquid sample and use method of laser desorption ionization mass spectrometry kit
CN113533492A