A specific plasma metabolic marker combination for early-onset type 2 diabetes and application thereof
Patent Information
- Application Number
- CN202610606814.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-06
- Publication Date
- 2026-08-18
AI Technical Summary
[0008]针对现有技术中关于血浆游离氨基酸或脂肪酸多集中于单一代谢物分析、缺乏联合整合分析手段,且现有模型多基于一般人群、未能体现早发2型糖尿病特异性代谢特征,同时传统统计方法难以刻画多种代谢物之间复杂交互关系等问题,本发明的目的在于提供一种早发2型糖尿病的特异性血浆代谢标志物组合及其应用
[0034] This invention significantly improves the identification ability of early-onset type 2 diabetes by screening plasma free amino acid and fatty acid biomarkers and combining them with a stratified analysis strategy based on age of onset. Results show that the 14 candidate metabolic biomarkers screened demonstrated significantly better diagnostic efficacy in early-onset type 2 diabetes than in late-onset type 2 diabetes. The area under the receiver operating characteristic (AUC) curve for a single metabolite in the early-onset population ranged from 0.645 to 0.816, significantly higher than the 0.505–0.687 in the late-onset population. Specifically, γ-aminobutyric acid (GABA) exhibited the highest diagnostic efficacy, with an AUC of 0.816 (95% CI: 0.761–0.871), followed by leucine, with an AUC of 0.781 (95% CI: 0.722–0.841). These results indicate that the metabolic biomarkers screened in this invention have higher sensitivity and specificity for early-onset type 2 diabetes, effectively overcoming the shortcomings of existing technologies in identifying early-onset individuals.
Smart Images

Figure CN122591829A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a specific combination of plasma metabolic markers for early-onset type 2 diabetes and their applications, and belongs to the field of pharmaceutical technology. Background Technology
[0002] With changing lifestyles and an increasing burden of metabolic diseases, the incidence of early-onset type 2 diabetes mellitus (T2DM) is rising annually, becoming a significant global public health issue. Compared to late-onset T2DM, early-onset T2DM has an earlier age of onset, a longer disease duration, and is usually accompanied by more pronounced metabolic abnormalities, such as a higher incidence of obesity, more prominent lipid metabolism disorders, and more severe insulin resistance, leading to a higher risk of complications and a greater impact on individual health and the social healthcare burden. However, despite the increasingly evident trend of early-onset T2DM, its identification and diagnosis remain significantly inadequate.
[0003] Currently, the diagnosis of early-onset type 2 diabetes mellitus (T2DM) mainly relies on blood glucose-related indicators such as fasting blood glucose, oral glucose tolerance test (OGTT), and glycated hemoglobin (HbA1c). These indicators are largely based on late-onset T2DM populations and are not entirely applicable to children and young adults. In practice, early-onset T2DM often has a long subclinical phase, with patients lacking typical symptoms. Furthermore, young adults often lack awareness of disease risk and have low screening adherence, leading to the disease being easily overlooked. In addition, HbA1c has limited sensitivity to blood glucose abnormalities, potentially underestimating prediabetes and the prevalence of diabetes; fasting blood glucose testing requires multiple repeated measurements, and a single test can easily result in missed diagnoses. Therefore, the existing diagnostic system has insufficient applicability and sensitivity in early-onset T2DM populations, lacking effective identification methods.
[0004] Recent studies have shown that the development of type 2 diabetes mellitus (T2DM) is not only related to abnormal blood glucose levels but also closely associated with amino acid and fatty acid metabolism disorders. In particular, plasma free amino acids and fatty acids, as important components of energy metabolism and insulin signaling regulation, play crucial roles in insulin resistance, β-cell dysfunction, and chronic inflammation. For example, elevated levels of branched-chain amino acids (such as leucine, isoleucine, and valine) and aromatic amino acids (such as phenylalanine and tyrosine) are considered closely related to insulin resistance, while elevated saturated fatty acid levels and an imbalance in the proportion of unsaturated fatty acids can participate in the development of T2DM by inducing lipotoxicity and oxidative stress. Further research indicates that these metabolic abnormalities often appear before a significant increase in blood glucose levels, suggesting their potential value as biomarkers. Compared to late-onset T2DM, early-onset type 2 diabetes patients typically exhibit more significant metabolic disorders, with more prominent amino acid and fatty acid metabolism abnormalities. Therefore, based on the combined characteristics of plasma free amino acids and fatty acids, a new technical approach may be provided for the identification of early-onset T2DM.
[0005] Existing technologies for detecting free amino acids and fatty acids in plasma typically include sample collection, pretreatment, and instrumental analysis. Pretreatment includes protein precipitation, derivatization, and organic solvent extraction. Detection methods primarily employ liquid chromatography-mass spectrometry (LC-MS) or gas chromatography-mass spectrometry (GC-MS). The detection process involves process conditions such as column type, mobile phase composition, gradient elution program, ion source parameters, and scan range. Regarding data analysis, most studies use univariate analysis or multiple regression models to assess the relationship between single or a few amino acids or fatty acids and the risk of T2DM. While some studies have constructed predictive models, they are often limited to single metabolite categories and lack a systematic integrated analysis of the combined characteristics of amino acids and fatty acids.
[0006] However, the aforementioned existing technologies still have significant shortcomings: First, most studies focus only on single metabolic pathways of amino acids or fatty acids, failing to integrate information from both key metabolites and thus making it difficult to comprehensively reflect the overall characteristics of insulin resistance and energy metabolism disorders; second, existing analytical methods mostly employ traditional statistical models, which are difficult to effectively characterize the nonlinear relationships and interactions between multiple metabolites, limiting the improvement of model performance; third, existing studies are mostly based on the general T2DM population, lacking stratified analysis strategies for early-onset T2DM and failing to reveal its specific metabolic patterns; in addition, different studies vary greatly in sample processing procedures and detection parameters, resulting in insufficient reproducibility and comparability of results.
[0007] Therefore, existing technologies have not yet provided a systematic diagnostic method based on the combined characteristics of plasma free amino acids and fatty acids, combined with specific analysis of early-onset populations, which makes it difficult to meet the needs of screening and accurate identification of early-onset type 2 diabetes. Summary of the Invention
[0008] In view of the problems that existing technologies for analyzing plasma free amino acids or fatty acids are mostly focused on single metabolite analysis, lacking integrated analysis methods, and that existing models are mostly based on the general population and fail to reflect the specific metabolic characteristics of early-onset type 2 diabetes, while traditional statistical methods are difficult to characterize the complex interactions between multiple metabolites, the purpose of this invention is to provide a specific plasma metabolic biomarker combination for early-onset type 2 diabetes and its application.
[0009] To achieve the above objectives, the present invention employs the following technical means:
[0010] This invention integrates protein metabolism and lipid metabolism information and introduces a stratified comparison strategy based on the age of onset, dividing research subjects into an early-onset type 2 diabetes group, an early-onset control group, a late-onset type 2 diabetes group, and a late-onset control group. This allows for the identification of age-specific metabolic abnormalities, improving the model's specificity and diagnostic accuracy. First, plasma samples are collected from the subjects and pre-processed. For amino acid detection, protein precipitation and derivatization are preferred; for fatty acid detection, organic solvent extraction and methyl esterification are preferred. Subsequently, liquid chromatography-mass spectrometry (LC-MS) or gas chromatography-mass spectrometry (GC-MS) is used to quantitatively detect free amino acids (including but not limited to branched-chain amino acids and aromatic amino acids) and fatty acids (including but not limited to saturated fatty acids, unsaturated fatty acids, and polyunsaturated fatty acids) in the plasma. The detection process involves chromatographic separation conditions and mass spectrometry parameters, including column type, mobile phase composition, gradient elution program, ion source type, scan range, and collision energy. After obtaining metabolite concentration data, considering that they are usually non-normally distributed, univariate statistical analysis was first used to preliminarily screen each metabolite, with nonparametric tests preferred for inter-group comparisons. Subsequently, orthogonal partial least squares discriminant analysis (OPLS-DA) was used to perform multivariate pattern recognition analysis on samples from different groups to reveal the overall metabolic differences between early-onset type 2 diabetes and the control group, and to evaluate the model's classification ability and variable contributions. Based on this, various machine learning methods were further used to construct a diagnostic model, including but not limited to random forest, support vector machine, and other supervised learning algorithms. A multivariate classification model was established by integrating the key metabolite features obtained from the screening, and the model performance was evaluated through methods such as cross-validation. The model with the best performance was selected as the final diagnostic model, thereby realizing the assessment and auxiliary diagnosis of individual risk of early-onset type 2 diabetes.
[0011] This invention discloses a specific plasma metabolic marker combination for early-onset type 2 diabetes mellitus, comprising 14 peripheral blood amino acid and fatty acid markers. The amino acid markers include leucine, valine, isoleucine, γ-aminobutyric acid, dimethylglycine, glutamic acid, and GABA; the fatty acid markers include C18:1, C18:2, C18:3a, C18:4, C18:3r, C22:5, and C22:6.
[0012] The symbols mentioned above (e.g., C18:1) are abbreviations for fatty acid names, in the format "C number: number", where: the first number represents the total number of carbon atoms in the fatty acid molecule; the number after the colon represents the number of carbon-carbon double bonds in the molecule (i.e., the degree of unsaturation); the suffix letters (e.g., a, r) represent the spatial configuration or positional difference of the double bond; a represents the cis configuration, where the hydrogen atoms on both sides of the double bond are on the same side; r represents the trans configuration, where the hydrogen atoms on both sides of the double bond are on opposite sides.
[0013] C18:1 Official name: Oleic acid;
[0014] C18:2 Official name: Linoleic acid;
[0015] C18:3a Official name: α-Linolenic acid;
[0016] C18:3r official name: γ-Linolenic acid;
[0017] C18:4 Official name: Stearidonic acid;
[0018] C22:5 Official name: docosapentaenoic acid (DPA);
[0019] C22:6's official name is docosahexaenoic acid (DHA).
[0020] Preferably, the plasma metabolic markers are derived from plasma, serum, or whole blood.
[0021] Furthermore, the present invention also proposes the application of the aforementioned combination of specific plasma metabolic markers in the preparation of a diagnostic kit for early-onset type 2 diabetes.
[0022] And the use of reagents for detecting the aforementioned specific plasma metabolic marker combination in the preparation of a diagnostic kit for early-onset type 2 diabetes.
[0023] Furthermore, the present invention also proposes a diagnostic kit for early-onset type 2 diabetes, comprising the aforementioned specific plasma metabolic marker combination and / or reagents for detecting the aforementioned specific plasma metabolic marker combination.
[0024] Preferably, the reagent is used to quantitatively detect free amino acids and fatty acids in peripheral blood using liquid chromatography-mass spectrometry, gas chromatography-mass spectrometry, or nuclear magnetic resonance spectroscopy.
[0025] Furthermore, this invention also proposes a system for predicting or diagnosing early-onset type 2 diabetes, comprising:
[0026] (1) Data acquisition module: used to acquire detection data of biological samples to be tested, wherein the detection data is used to characterize the content of specific plasma metabolic markers of early-onset type 2 diabetes in the samples to be tested;
[0027] (2) Prediction module: used to output the prediction or diagnosis results of early-onset type 2 diabetes based on the detection data of the biological sample to be tested;
[0028] The specific plasma metabolic markers include 14 peripheral blood amino acid and fatty acid markers. The amino acid markers include leucine, valine, isoleucine, γ-aminobutyric acid, dimethylglycine, glutamic acid, and aminobutyric acid. The fatty acid markers include C18:1, C18:2, C18:3a, C18:4, C18:3r, C22:5, and C22:6.
[0029] Preferably, peripheral blood samples are collected from the subjects and pre-processed; then, liquid chromatography-mass spectrometry, gas chromatography-mass spectrometry, or nuclear magnetic resonance spectroscopy are used to quantitatively detect free amino acids and fatty acids in the plasma to obtain detection data of the biological sample to be tested.
[0030] Preferably, in the pretreatment steps, when used for amino acid detection, the sample is subjected to protein precipitation and derivatization; when used for fatty acid detection, the sample is subjected to organic solvent extraction and methyl esterification.
[0031] Preferably, if the levels of the 14 peripheral blood amino acid and fatty acid markers detected are significantly higher than those in biological samples without early-onset type 2 diabetes, then the sample is identified as having early-onset type 2 diabetes.
[0032] It should be noted that, without departing from the inventive concept, various alternative solutions can be used to achieve the same inventive objective. For example, in terms of data analysis methods, in addition to univariate analysis, OPLS-DA, and the aforementioned machine learning methods, other multivariate statistical analysis methods or deep learning models can be used; in terms of stratification strategies, adjustments can be made based on different age thresholds or other clinical characteristics; in terms of metabolite selection, the specific types and combinations of amino acids and fatty acids can be optimized according to the characteristics of the research subjects. None of these alternative solutions affect the core technical idea of this invention—integrating amino acid and fatty acid metabolic information and combining it with a multi-step modeling strategy to construct a diagnostic model—and all should be considered within the scope of protection of this invention.
[0033] Compared with the prior art, the beneficial effects of the present invention are:
[0034] This invention significantly improves the identification ability of early-onset type 2 diabetes by screening plasma free amino acid and fatty acid biomarkers and combining them with a stratified analysis strategy based on age of onset. Results show that the 14 candidate metabolic biomarkers screened demonstrated significantly better diagnostic efficacy in early-onset type 2 diabetes than in late-onset type 2 diabetes. The area under the receiver operating characteristic (AUC) curve for a single metabolite in the early-onset population ranged from 0.645 to 0.816, significantly higher than the 0.505–0.687 in the late-onset population. Specifically, γ-aminobutyric acid (GABA) exhibited the highest diagnostic efficacy, with an AUC of 0.816 (95% CI: 0.761–0.871), followed by leucine, with an AUC of 0.781 (95% CI: 0.722–0.841). These results indicate that the metabolic biomarkers screened in this invention have higher sensitivity and specificity for early-onset type 2 diabetes, effectively overcoming the shortcomings of existing technologies in identifying early-onset individuals.
[0035] Furthermore, this invention significantly improves diagnostic performance by integrating the aforementioned 14 metabolic biomarkers to construct a multivariate diagnostic model. A classification model was established based on various machine learning methods, including logistic regression, random forest, support vector machine, and XGBoost. Model parameters were optimized using 5-fold, 5-times repeated cross-validation, and model stability was evaluated using 100 Monte Carlo cross-validations. Results show that all models exhibited good predictive performance (average AUC greater than 0.75). The logistic regression model performed best, with an average AUC of 0.859 (95% CI: 0.852–0.865), and its sensitivity and specificity were 0.811 and 0.807, respectively. The XGBoost model had an average AUC of 0.845 (95% CI: 0.838–0.853), the support vector machine model had an average AUC of 0.836 (95% CI: 0.828–0.844), and the random forest model had an average AUC of 0.833 (95% CI: 0.825–0.840). The above results demonstrate that the present invention, through multi-marker joint modeling and multi-model validation, can significantly improve the diagnostic accuracy and model stability of early-onset type 2 diabetes.
[0036] Furthermore, this invention also found that the diagnostic efficacy of the aforementioned metabolic biomarkers in late-onset type 2 diabetes is significantly reduced. Even when using multiple machine learning methods to construct models, their predictive performance remains limited, further validating the uniqueness of early-onset type 2 diabetes in terms of metabolic characteristics. Therefore, this invention, by introducing a hierarchical analysis strategy and combining amino acid and fatty acid co-modeling, not only improves model performance but also clarifies its specific application value in early-onset populations, demonstrating promising clinical application prospects for screening, risk assessment, and auxiliary diagnosis of early-onset type 2 diabetes. Attached Figure Description
[0037] Figure 1 Principal component analysis plots of free amino acids and free fatty acids for 300 samples and QC samples;
[0038] Where A represents amino acids and B represents fatty acids;
[0039] Figure 2 Differences in the levels of free amino acids and fatty acids in the plasma of patients with type 2 diabetes mellitus (T2DM) and control groups;
[0040] Where A represents amino acids and B represents fatty acids;
[0041] Figure 3 shows the OPLS-DA analysis and 200 permutation test of diabetic patients and control groups based on free amino acid and free fatty acid data;
[0042] Wherein, A represents the separation trend of early-onset T2DM patients and their control group in terms of metabolic profile; B represents the results of 200 Y-transformation tests for early-onset T2DM patients; C represents the separation trend of late-onset T2DM patients and their control group in terms of metabolic profile; and D represents the results of 200 Y-transformation tests for late-onset T2DM patients.
[0043] Figure 4 This study aims to analyze the amino acid and fatty acid VIP values and pathway enrichment in early-onset type 2 diabetes mellitus (T2DM) based on the OPLS-DA model.
[0044] Where A represents the VIP values of metabolites in the OPLS-DA model of early-onset T2DM and the control group; B represents the results of pathway enrichment analysis of these differentially expressed metabolites using MetaboAnalyst. Detailed Implementation
[0045] The present invention will be further described below with reference to specific embodiments, but the present invention is not limited to the following embodiments. Those skilled in the art should understand that modifications or substitutions can be made to the details and form of the technical solutions of the present invention without departing from the spirit and scope of the present invention, but all such modifications and substitutions fall within the protection scope of the present invention.
[0046] Example 1: Screening and application of specific plasma metabolic markers for early-onset type 2 diabetes.
[0047] I. Materials and Methods
[0048] 1. Main instruments and equipment
[0049] Table 1
[0050]
[0051] 2. Main reagents
[0052] Table 2
[0053]
[0054] 3. Sample collection
[0055] The study included 120 patients with early-onset type 2 diabetes mellitus (T2DM) and late-onset T2DM recruited by the Department of Endocrinology, Second Affiliated Hospital of Harbin Medical University between October 1, 2023 and October 1, 2024. Diabetes was diagnosed according to the "Guidelines for the Prevention and Treatment of Type 2 Diabetes in China (2020 Edition)": patients needed to have typical diabetic symptoms and meet the criteria of fasting blood glucose ≥ 7.0 mmol / L or postprandial blood glucose ≥ 11.1 mmol / L. Exclusion criteria included: type 1 diabetes (excluded by immune antibody and gene testing), other special types of diabetes (such as drug-induced, gene mutation, or infection), gestational diabetes, cancer, hereditary diseases, hyperthyroidism, mental illness, heart, liver, and kidney failure, and limb disability (inability to walk normally). The control group was selected from the orthopedic department during the same period and matched with the case group by sex and age ± 3 years. Patients with prediabetes and any type of diabetes were excluded from the control group through blood glucose, glycated hemoglobin testing, and medical history inquiry. This study has been approved by the Ethics Committee of Harbin Medical University, and all participants have signed informed consent forms.
[0056] All participants were required to fast overnight for 8–12 hours before blood collection, during which time they were prohibited from consuming any food, beverages (including sugary drinks), and alcohol. On the day of blood collection, participants were required to avoid taking any hypoglycemic medications (including oral hypoglycemic agents and insulin), eating breakfast, and engaging in strenuous physical activity to minimize the interference of short-term behavioral factors on metabolic indicators. Fasting venous blood was collected by uniformly trained professional nurses before 9:00 AM using 10 mL blood collection tubes containing EDTA anticoagulant. Blood samples were processed within 2 hours of collection, including high-speed centrifugation at 4°C to separate plasma. The plasma was then aliquoted and immediately placed in a deep cryogenic freezer for long-term storage.
[0057] 4. Plasma amino acid detection
[0058] 4.1 Sample Pretreatment
[0059] Plasma samples were thawed in stages before analysis: first transferred from a −80 °C freezer to −20 °C overnight, then slowly thawed at 4 °C, and finally processed on ice to minimize the impact of repeated freeze-thaw cycles on metabolite stability. 30 μL of plasma was taken and 80 μL of sample processing solution (methanol / acetonitrile / formic acid = 24.9:74.9:0.2) was added, mixed thoroughly, and incubated at −20 °C for 2 h. Subsequently, it was centrifuged at 15,000 r / min at 4 °C for 15 min, and the supernatant was collected into a new 1.5 mL EP tube and incubated overnight at −20 °C. The next day, it was centrifuged again under the same conditions for 15 min, and the supernatant was transferred to a sample vial for LC-MS analysis. Fifteen participants were randomly selected from each study group, for a total of 60 plasma samples. Take 50 μL of plasma from each sample and combine them into a new 5 mL centrifuge tube. Mix thoroughly to prepare a mixed QC sample.
[0060] 4.2 Liquid Chromatography-Mass Spectrometry Conditions
[0061] Quantitative analysis of amino acids and biogenic amines was performed using a Waters ACQUITY UPLC system coupled with a Waters Xevo TQD triple quadrupole mass spectrometer (Waters Corporation, Manchester, UK). Chromatographic separation was performed on a HILIC column (100 mm × 2.1 mm × 1.7 μm, Waters Corporation, Milford, CT, USA). Mobile phase A consisted of 0.2% formic acid, 5% acetonitrile, and 5 mmol / L ammonium formate aqueous solution, while mobile phase B consisted of 0.2% formic acid, 5% water, and 5 mmol / L ammonium formate acetonitrile solution. Gradient elution was used to achieve effective separation of target compounds, with an injection volume of 2 μL. Mass spectrometry detection was performed by electrospray ionization (ESI) in positive ion mode, and data acquisition was performed in multiple reaction monitoring (MRM) mode to ensure sensitivity and selectivity. The mass spectrometry conditions and retention times for each amino acid are shown in Table 3.
[0062] Table 3 Mass spectrometry conditions and retention times for amino acids
[0063]
[0064]
[0065] 4.3 Quantitative analysis of amino acids
[0066] The quantification of amino acids in plasma was performed using the external standard method, with a series of standards prepared using Sigma brand amino acid standard reference materials. Standard curves were established by determining the linear relationship between the peak area of the standards and their known concentrations to calculate the absolute content of each amino acid in the plasma samples. The standard curves for all 26 amino acids showed good linearity (R² > 0.99), and the linear range and detection limit for each amino acid are detailed in Table 4.
[0067] Table 4 Standard curves and detection ranges for amino acids
[0068]
[0069]
[0070] Since the plasma sample was diluted 4 times before testing, the actual plasma concentration is converted using the following formula:
[0071]
[0072] 5. Plasma fatty acid detection
[0073] 5.1 Fatty acid sample pretreatment
[0074] Plasma samples were thawed and thoroughly mixed at 4 °C before analysis. 150 μL of sample was placed in a 15 mL centrifuge tube, and an equal volume (150 μL) of internal standard solution (C17:0, 200 μg / mL) and 1.5 mL of 10% H2SO4-methanol were added. After vortexing for 1 minute, the mixture was incubated in a 65 °C water bath for 2 hours. After cooling, an appropriate amount of anhydrous sodium sulfate and 1.5 mL of n-hexane were added, vortexed for 1 minute, and centrifuged (3500 r / min, 5 min). The supernatant of n-hexane was transferred to a new 1.5 mL microcentrifuge tube and dried under nitrogen. The solution was reconstituted with 75 μL of n-hexane, vortexed, centrifuged, and the supernatant was analyzed by GC-MS. To monitor instrument stability, equal volumes of plasma from 60 subjects were mixed to prepare QC samples, processed using the same pretreatment procedure, and thoroughly mixed before drying under nitrogen. The samples were then aliquoted and stored. During the analysis, a QC sample is inserted every 5–10 samples to monitor instrument drift and repeatability, ensuring data reliability.
[0075] 5.2 Detection Method
[0076] Analysis was performed using a TRACE 1310 gas chromatograph in conjunction with a TSQ 8000 Evo triple quadrupole mass spectrometer. Chromatographic separation was performed on a DB-WAX capillary column (30 m × 0.25 mm, 0.25 μm) with helium carrier gas at a flow rate of 1.0 mL / min, an injection volume of 1 μL, a split ratio of 1:10, and a solvent delay of 5 minutes. The column temperature program was: 50 °C for 1 minute, ramping to 200 °C (10 °C / min) for 1–16 minutes, maintaining at 200 °C for 16–26 minutes, ramping to 220 °C (5 °C / min) for 26–30 minutes, and maintaining at 220 °C for 30–41 minutes. Mass spectrometry was performed in EI positive ion mode with SRM scanning, covering a mass range of 30–450 m / z and an electron energy of 70 eV. Two standard injections were administered daily to assess system stability, and one QC sample was inserted every 10 sample injections. The mass spectrometry conditions and retention times for each fatty acid are shown in Table 5.
[0077] Table 5. Detection conditions and retention times for plasma fatty acids
[0078]
[0079] To achieve accurate quantification of fatty acids in plasma samples, this study established a quantitative model based on the external standard method. First, high-purity fatty acid standards were selected and serially diluted with methanol to prepare a series of standard solutions of different concentrations, covering the possible concentration range of the target compounds in the samples. All standard solutions were subjected to the same derivatization, injection, and chromatographic-mass spectrometry detection conditions as the experimental samples to ensure comparability of response values. Subsequently, by measuring the peak areas of the standard solutions at different concentrations, a standard curve was plotted using linear regression. The linear relationship between concentration and signal was obtained by plotting the response value (peak area) on the ordinate and concentration on the abscissa. The correlation coefficients of the standard curves for each fatty acid were all higher than 0.99, indicating good linearity and meeting the quantitative requirements (see Table 6). Finally, by combining the measured peak areas of the samples with the standard curve equation, the actual concentrations of each fatty acid in plasma were calculated.
[0080] Table 6 Standard curves and detection ranges for fatty acids
[0081]
[0082] 6 Data Analysis
[0083] In univariate analysis, we first performed natural logarithmic transformation on the raw metabolite data to mitigate the impact of data skewness on statistical analysis, and used independent samples t-tests to assess differences between groups. For multiple comparisons, all p-values were corrected using the Benjamini-Hochberg method to control for false positives. In multivariate analysis, we used SIMCA software for partial least squares discriminant analysis (PLS-DA) and orthogonal PLS-DA (OPLS-DA). Principal component analysis (PCA) was used to assess the clustering of QC samples and overall data quality to detect potential batch effects and technical variability; the OPLS-DA model was used to observe metabolic differences between groups and screen for potential biomarkers.
[0084] To further evaluate the predictive power of multi-marker integration for disease classification, we constructed four machine learning models—Logistic Regression (GLMNET), Support Vector Machine (SVM), Random Forest (RF), and XGBoost—using the R language. During model construction, we first used 5-fold, 5-repeated cross-validation to tune the hyperparameters of each model on the training set, selecting the parameter group with the highest average AUC in the cross-validation as the optimal hyperparameters. Subsequently, we performed 100 Monte Carlo cross-validations based on these optimal hyperparameters to internally validate the models and evaluate their predictive performance and robustness under different random partitions.
[0085] II. Results
[0086] 1 Univariate analysis
[0087] We performed fatty acid and amino acid omics analysis on plasma samples from 120 patients with early-onset type 2 diabetes mellitus (T2DM) and their age- and sex-matched controls, as well as 30 patients with late-onset T2DM and their controls. PCA score plots showed that amino acids (… Figure 1 A) and fatty acids ( Figure 1 The QC samples in B) were well clustered, and no abnormal sample points were found by DModX diagnosis, indicating that the data quality was reliable.
[0088] Cluster heatmaps showed the differential distribution of amino acids and fatty acids among the four groups of samples. Univariate statistical analysis showed that, compared with the age-matched control group, patients with early-onset T2DM had significantly elevated levels of multiple amino acids and fatty acids in their plasma, with 18 metabolites reaching statistical significance, including nine amino acids (glutamate, tryptophan, β-aminobutyric acid, γ-aminobutyric acid, valine, dimethylglycine, leucine, isoleucine, and alanine). Figure 2A). Among them, glutamate, as a System Xc⁻ substrate, was significantly elevated, which led to restricted glutathione synthesis in β-cells and increased ferroptosis sensitivity; among fatty acids, early-onset T2DM, monounsaturated fatty acids, polyunsaturated fatty acids, n-3 series fatty acids (C18:3r, C18:4, C22:5, C22:6), and n-6 series fatty acids (C18:2, C18:3a, C20:4, C22:4) were significantly elevated. Figure 2 B). Among them, polyunsaturated fatty acids are rich in double bonds and are highly susceptible to lipid peroxidation. Their significant increase can provide sufficient easily oxidized substrates for ferroptosis. In contrast, although late-onset T2DM patients showed an upward trend in some amino acids and fatty acids, the increase did not reach statistical significance after multiple comparison correction (Table 7).
[0089] Table 7. Differences in 40 metabolites among different groups
[0090]
[0091]
[0092] 2. Results of multivariate statistical analysis
[0093] In the targeted fatty acid and amino acid analysis, to further identify the metabolic features that contributed most to the differences between groups, this study used OPLS-DA for multivariate analysis. The results showed a clear separation trend in the metabolic profiles of early-onset T2DM patients and their control group. Figure 3 A) indicates significant differences in amino acid and fatty acid metabolic characteristics (R) 2 Y = 0.599, Q 2 =0.555). The model was further validated using 200 Y-permutation tests, and the results showed that the R-value of the original model was 0.555. 2 and Q 2 The values are all significantly higher than those of the random permutation model ( Figure 3 B) indicates that the model has good stability and no obvious overfitting. In contrast, although late-onset T2DM patients and their control group showed a certain separation trend in OPLS-DA analysis (R... 2 Y = 0.652, Q 2 =0.137), but the Y permutation test results suggest that the model has a risk of overfitting ( Figure 3 C, D).
[0094] In the OPLS-DA model of early-onset T2DM and the control group, a total of 14 metabolites had VIP values greater than 1 ( Figure 4A). Based on the screening criteria of VIP > 1 and FDR < 0.05, 14 metabolites were finally identified, including glutamate, γ-aminobutyric acid (GABA), valine, dimethylglycine, leucine, isoleucine, aminobutyric acid, and C18:1, C18:2, C18:3a, C18:3r, C18:4, C22:5, and C22:6. Further pathway enrichment analysis of these differentially expressed metabolites using MetaboAnalyst revealed that they were mainly enriched in the Valine, leucine, and isoleucine biosynthesis and the Biosynthesis of unsaturated fatty acids pathway. Figure 4 B).
[0095] 3. Constructing a specific diagnostic model for early-onset T2DM based on machine learning and internal validation.
[0096] We evaluated the diagnostic efficacy of 14 single candidate biomarkers for early-onset and late-onset T2DM. The results are shown in Table 8. The AUCs of these metabolites in early-onset T2DM ranged from 0.645 to 0.816, with GABA exhibiting the highest diagnostic efficacy (AUC 0.816, 95% CI: 0.761–0.871), followed by leucine (AUC 0.781, 95% CI: 0.722–0.841). In late-onset T2DM, the AUCs ranged from 0.505 to 0.687, with valine showing the highest AUC (0.687, 95% CI: 0.540–0.833). Notably, all 14 metabolites demonstrated significantly higher diagnostic efficacy in early-onset T2DM than in late-onset T2DM.
[0097] Table 8. Diagnostic efficacy of individual amino acids and fatty acids for early-onset and late-onset T2DM
[0098]
[0099] To further explore whether multi-marker integration can improve diagnostic efficacy, we constructed four classification models based on the aforementioned 14 peripheral blood amino acid and fatty acid markers: logistic regression (GLMNET), random forest (RF), support vector machine (SVM), and XGBoost. To optimize model performance and assess robustness, we first used 5-fold, 5-times repeated cross-validation to adjust the hyperparameters of each model, and selected the parameter group with the highest average AUC in repeated cross-validation as the optimal hyperparameters (see Table 9).
[0100] Table 9 Key parameter settings for four machine learning models
[0101]
[0102] Subsequently, 100 Monte Carlo cross-validations were performed based on these optimal hyperparameters to evaluate the model's predictive performance. The results showed that all four models had good predictive power (mean AUC > 0.75, Table 10). Among them, logistic regression performed best, with an average AUC of 0.859 (95% CI: 0.852–0.865), sensitivity and specificity of 0.811 and 0.807, respectively; XGBoost model had an average AUC of 0.845 (95% CI: 0.838–0.853), sensitivity and specificity of 0.828 and 0.768, respectively; SVM model had an average AUC of 0.836 (95% CI: 0.828–0.844), sensitivity and specificity of 0.787 and 0.813, respectively; random forest model had an average AUC of 0.833 (95% CI: 0.825–0.840), Youden J had an AUC of 0.574, sensitivity and specificity of 0.779 and 0.794, respectively.
[0103] Table 10 Diagnostic efficacy of the combined amino acid and fatty acid diagnostic model in early-onset T2DM
[0104]
[0105] Furthermore, we used these algorithms to evaluate the diagnostic efficacy of these metabolites for late-onset T2DM, and the results showed that these metabolites had limited diagnostic efficacy for late-onset T2DM (Table 11).
[0106] Table 11 Diagnostic efficacy of the combined amino acid and fatty acid diagnostic model in late-onset type 2 diabetes mellitus (T2DM)
[0107]
Claims
1. A specific plasma metabolic marker combination for early-onset type 2 diabetes, characterized in that, The specific plasma metabolic marker combination contains 14 peripheral blood amino acid and fatty acid markers, wherein the amino acid markers include leucine, valine, isoleucine, γ-aminobutyric acid, dimethylglycine, glutamic acid, and aminobutyric acid; and the fatty acid markers include C18:1, C18:2, C18:3a, C18:4, C18:3r, C22:5, and C22:
6.
2. The specific plasma metabolic marker combination as described in claim 1, characterized in that, The plasma metabolic markers mentioned are derived from plasma, serum, and whole blood.
3. The use of the specific plasma metabolic marker combination according to claim 1 in the preparation of a diagnostic kit for early-onset type 2 diabetes.
4. The use of the reagent for detecting the specific plasma metabolic marker combination of claim 1 in the preparation of a diagnostic kit for early-onset type 2 diabetes.
5. A diagnostic kit for early-onset type 2 diabetes, characterized in that, Includes the specific plasma metabolic marker combination as described in claim 1 and / or reagents for detecting the specific plasma metabolic marker combination as described in claim 1.
6. The diagnostic kit as described in claim 5, characterized in that, The reagents described herein are used to quantitatively detect free amino acids and fatty acids in peripheral blood using liquid chromatography-mass spectrometry, gas chromatography-mass spectrometry, or nuclear magnetic resonance spectroscopy.
7. A system for predicting or diagnosing early-onset type 2 diabetes, characterized in that, include: (1) Data acquisition module: used to acquire detection data of biological samples to be tested, wherein the detection data is used to characterize the content of specific plasma metabolic markers of early-onset type 2 diabetes in the samples to be tested; (2) Prediction module: used to output the prediction or diagnosis results of early-onset type 2 diabetes based on the detection data of the biological sample to be tested; The specific plasma metabolic markers include 14 peripheral blood amino acid and fatty acid markers. The amino acid markers include leucine, valine, isoleucine, γ-aminobutyric acid, dimethylglycine, glutamic acid, and aminobutyric acid. The fatty acid markers include C18:1, C18:2, C18:3a, C18:4, C18:3r, C22:5, and C22:
6.
8. The prediction or diagnostic system for early-onset type 2 diabetes as described in claim 7, characterized in that, Peripheral blood samples were collected from the subjects and pretreated. Subsequently, liquid chromatography-mass spectrometry, gas chromatography-mass spectrometry, or nuclear magnetic resonance spectroscopy were used to quantitatively detect free amino acids and fatty acids in the plasma to obtain detection data of the biological samples to be tested.
9. The prediction or diagnostic system for early-onset type 2 diabetes as described in claim 8, characterized in that, In the aforementioned pretreatment steps, for amino acid detection, the sample undergoes protein precipitation and derivatization; for fatty acid detection, the sample undergoes organic solvent extraction and methyl esterification.
10. The prediction or diagnostic system for early-onset type 2 diabetes as described in claim 7, characterized in that, If the levels of the 14 peripheral blood amino acid and fatty acid markers detected are significantly higher than those in biological samples without early-onset type 2 diabetes, then the patient is diagnosed with early-onset type 2 diabetes.