Diabetes early warning method based on metabonomics

Through metabolomics methods, combined with principal component analysis and partial least squares discriminant analysis, key metabolites were screened to construct an early warning model for diabetes, which solved the problem that the existing technology was difficult to accurately identify high-risk people in diabetes in the early stage, and achieved efficient and accurate prediction of diabetes risk.

CN120032880AInactive Publication Date: 2025-05-23李艳
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510109514.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify people at high risk of diabetes in the early stage, and traditional detection methods are difficult to capture subtle and specific metabolic changes signals.

Method used

Using metabolomics-based methods, blood, urine and saliva samples were collected, metabolites were detected using detection platforms such as liquid chromatography-mass spectrometry, principal component analysis and partial least squares discriminant analysis were performed, key metabolites were screened, and early warning model for diabetes was constructed.

Benefits of technology

It realizes early accurate prediction of diabetes risk, improves the accuracy and specificity of detection, simplifies the data processing process, supports high-throughput detection platform, and improves detection efficiency and feasibility of clinical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032880A_ABST
    Figure CN120032880A_ABST
Patent Text Reader

Abstract

The invention discloses a metabonomics-based diabetes early warning method. The metabonomics-based diabetes early warning method comprises the following steps: collecting samples of a total population; detecting metabolites in the sample, and constructing an initial data matrix; preprocessing the initial data matrix; preprocessing the data matrix through singular value decomposition to obtain a difference principal component direction, and excluding outliers; judging the preprocessed data matrix to obtain a maximum projection vector; screening metabolites meeting conditions in the projection vectors to obtain a standard data matrix; and screening samples meeting conditions in the standard data matrix and the to-be-tested sample to obtain a risk level. Therefore, the characteristics of the metabolites in the sample can be systematically analyzed, early accurate prediction of the diabetes risk is realized, the accuracy and the specificity of detection are improved, outliers are effectively eliminated, key metabolites are screened, the reliability and the stability of the model are enhanced, the data processing flow is simplified, and the detection efficiency and the feasibility of clinical application are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical technology, and in particular to an early warning method for diabetes based on metabolomics. Background Art

[0002] Diabetes is a metabolic disease caused by multiple factors and characterized by chronic hyperglycemia. Long-term hyperglycemia can lead to serious complications such as cardiovascular and cerebrovascular diseases, diabetic nephropathy, and diabetic retinopathy, which have a serious impact on the patient's quality of life and increase the social medical burden. Therefore, early diagnosis and intervention are of vital importance to preventing or delaying the occurrence of diabetes and its complications.

[0003] The current diagnosis of diabetes mainly relies on indicators such as fasting plasma glucose (FPG), oral glucose tolerance test (OGTT) and glycated hemoglobin A1c (HbA1c). However, these detection methods are difficult to detect high-risk groups of diabetes with normal blood sugar in a timely manner, so we urgently need to explore and use other biomarkers to identify high-risk groups of diabetes.

[0004] In recent years, with the development of bioinformatics and analytical chemistry, metabolomics has gradually attracted attention as an emerging research paradigm. Metabolomics can comprehensively reflect the changes in the physiological state of the body by systematically analyzing small molecule metabolites in organisms, providing new ideas for early detection of diseases. Diabetes involves changes in complex metabolic networks. Metabolomics can capture more subtle and specific change signals, which helps to identify risk individuals that have not been screened by traditional detection methods.

[0005] Despite this, current metabolomics-based diabetes research still faces many challenges. On the one hand, there are many types of metabolites and the data dimension is high. How to screen out markers that are truly clinically valuable is a difficult problem. On the other hand, factors such as sample processing, selection of detection platforms, and data analysis methods will also have an important impact on the results. Therefore, it is particularly urgent to develop a scientific, reasonable, efficient and accurate metabolomics-based diabetes early warning system. Summary of the invention

[0006] The present invention aims to solve the technical problems in the above-mentioned technology at least to some extent.

[0007] To this end, the present invention discloses a metabolomics-based diabetes early warning method, comprising the following steps:

[0008] S1. Collect samples from a total population n, wherein the total population n is divided into a prediabetic population N1 、Normal blood sugar diabetes high risk group 2 and healthy people N 3 , n=N 1 +N 2 +N 3 ;

[0009] S2. Use the detection platform to detect the m metabolites in the sample to construct the initial data matrix X of the metabolites n×m , where the initial data matrix X n×m Every element x in ij represents the original signal intensity of the j-th metabolite in the i-th sample;

[0010] S3. The initial data matrix X n×m Perform preprocessing to obtain the preprocessed data matrix X' n×m ;

[0011] S4. Using principal component analysis to preprocess the data matrix X' n×m Performing singular value decomposition to obtain the direction of the principal component that characterizes the difference in the sample and to eliminate outliers in the sample;

[0012] S5. Using partial least squares discriminant analysis method to pre-process the data matrix X' n×m Perform discrimination to obtain the projection vector W with the maximum discrimination of the sample in the low-dimensional projection space max ;

[0013] S6. Filter the projection vector W max The k metabolites that satisfy the variable importance projection value VIP greater than the first threshold and the univariate statistical test P less than the second threshold are selected to obtain the standard data matrix X' n×k ;

[0014] S7. For a sample x to be tested test , filter the standard data matrix X' n×k The satisfaction and the test sample x test The distance between the samples is less than the third threshold, so as to predict the sample to be tested x test diabetes risk level.

[0015] The metabolomics-based early warning method for diabetes disclosed in the present invention can achieve early and accurate prediction of diabetes risk by systematically analyzing the metabolite characteristics in samples, thereby improving the accuracy and specificity of detection, effectively eliminating outliers using principal component analysis, and combining partial least squares discriminant analysis to screen key metabolites, thereby enhancing the reliability and stability of the model, simplifying the data processing process, supporting high-throughput detection platforms, and improving detection efficiency and the feasibility of clinical applications.

[0016] In addition, the metabolomics-based diabetes early warning method disclosed in the present invention may also have the following additional technical features:

[0017] In one embodiment of the present invention, in step S1, the sample includes: blood V Bld 、Urine V Ur and saliva V Sal .

[0018] In one embodiment of the present invention, in step S2, the detection platform includes: liquid chromatography-mass spectrometry, gas chromatography-mass spectrometry, nuclear magnetic resonance spectroscopy and capillary electrophoresis.

[0019] In one embodiment of the present invention, in step S4, the mathematical form of the principal component analysis method is Among them, U n×n is the left singular vector matrix, ∑ n×m is a diagonal matrix of singular values, is the right singular vector matrix, by observing that U n×n The score distribution of the samples in , excluding outliers in the samples.

[0020] In one embodiment of the present invention, in step S5, the mathematical form of the partial least squares discriminant analysis method is: Among them, T n×r is the score matrix, is the load matrix, E n×m is the residual matrix, by maximizing the score matrix T n×r The covariance between the column vectors in the is used to obtain the projection vector W with the maximum discrimination of the sample in the low-dimensional projection space. max .

[0021] In one embodiment of the present invention, in step S6, the first threshold value ranges from 1.0 to 1.5, and the second threshold value ranges from 0.01 to 0.05.

[0022] In addition, the present invention also discloses a computer-readable storage medium on which instructions are stored. When the instructions are executed by a computer, the computer executes the metabolomics-based diabetes early warning method.

[0023] In addition, the present invention also discloses an electronic device, including: a processor and a memory, wherein:

[0024] The memory stores instructions, and when the instructions are executed by the processor, the processor executes the metabolomics-based diabetes early warning method.

[0025] Additional contents and advantages of the present invention will be given in the following description or may be understood through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The technical solutions and beneficial effects of the present invention will become apparent and easily understood from the following contents in conjunction with the accompanying drawings, wherein:

[0027] Figure 1 The present invention is a flow chart of the metabolomics-based diabetes early warning method. DETAILED DESCRIPTION

[0028] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0029] The metabolomics-based early warning method for diabetes disclosed in the present invention will be described below with reference to the accompanying drawings.

[0030] like Figure 1 As shown, a metabolomics-based diabetes early warning method comprises the following steps:

[0031] S1. Collect samples from the total population n, where the total population n is divided into prediabetic population N 1 、Normal blood sugar diabetes high risk group 2 and healthy people N 3 , n=N 1 +N 2 +N 3 , samples include: blood V Bld 、Urine V Ur and saliva V Sal ;

[0032] Specifically, we should first recruit a total population of n = 500 people, and then divide them into the following groups based on their health records and family medical history: 1 =150 people, high-risk group of normal blood sugar diabetes N 2 =200 people, healthy population N 3 =150 people.

[0033] Prediabetic population 1The diagnostic criteria are as follows:

[0034] (1) Fasting venous blood glucose ≥6.1mmol / L, <7.0mmol / L;

[0035] (2) OGTT two-hour blood glucose ≥7.8mmol / L, <11.1mmol / L;

[0036] If any one of (1) and (2) is met, the person can be identified as having prediabetes;

[0037] Normal blood sugar diabetes high risk group 2 The diagnostic criteria are as follows:

[0038] (1) Age ≥40 years;

[0039] (2) a history of prediabetes;

[0040] (3) overweight or obese (BMI ≥ 24 kg / m2), waist circumference ≥ 90 cm for men and ≥ 85 cm for women;

[0041] (4) a sedentary lifestyle;

[0042] (5) family history of type 2 diabetes in first-degree relatives;

[0043] (6) Women with a history of giving birth to a macrosomia (birth weight ≥ 4 kg), overt gestational diabetes or gestational diabetes;

[0044] (7) Hypertension [systolic blood pressure ≥140 mmHg (1 mmHg = 0.133 kPa) and / or diastolic blood pressure ≥90 mmHg] or receiving antihypertensive treatment;

[0045] (8) dyslipidemia (high-density lipoprotein cholesterol ≤ 0.91 mmol / L and triglycerides ≥ 2.22 mmol / L, or receiving lipid-lowering treatment);

[0046] (9) Patients with atherosclerotic cerebrovascular disease;

[0047] (10) Those with a history of transient steroid diabetes;

[0048] (11) Patients with polycystic ovary syndrome;

[0049] (12) Patients with severe mental illness and / or long-term treatment with antidepressant drugs;

[0050] For adults (>18 years old), those who meet any of (1) to (12) can be identified as a high-risk group for diabetes:

[0051] Blood collected from each person V Bld5ml of fasting venous blood and 5ml of urine V Ur 10ml of mid-morning urine and saliva V Sal It is 2ml of non-irritating saliva, which is properly stored in a specific container and quickly transported to the laboratory at low temperature for testing, so as to fully obtain the individual metabolite information source and ensure that the sample can accurately reflect the metabolic characteristics of different populations.

[0052] It should be noted that for blood V Bld Sample collection: After collection, blood samples are centrifuged at 4℃ / 3000rpm for 15 minutes. To prepare plasma, samples are collected using blood collection tubes containing anticoagulants. After centrifugation, the upper plasma is aspirated and transferred to a sterile EP tube and marked. It is placed in a -80℃ refrigerator for long-term storage. Thaw at room temperature before use to avoid repeated freezing and thawing more than 3 times.

[0053] If necessary, treat the plasma sample with a protein precipitant (such as methanol, acetonitrile and other organic solvents). Add the precipitant in an amount 3 to 4 times the sample volume, vortex mix for 30 seconds, let stand at -20°C for 30 minutes, centrifuge at 4°C / 12000rpm for 10 minutes, take the supernatant, evaporate the organic solvent, and then re-dissolve it for metabolite detection.

[0054] For urine V Ur Sample collection, urine V Ur Centrifuge at 4℃ / 3000rpm for 10 minutes, transfer the supernatant to a new centrifuge tube, discard the precipitate, and repeat the centrifugation if the supernatant is turbid.

[0055] According to the test requirements, use a vacuum concentrator or nitrogen blower to concentrate the supernatant at low temperature, with a nitrogen flow rate of 10-15ml / min and a temperature of 30-40℃ to concentrate to the original volume The concentrate was stored at -80°C and thawed and mixed at room temperature before testing.

[0056] For saliva V Sal Sample collection, saliva V Sal Centrifuge at 4°C / 10,000 rpm for 10 minutes to separate impurities such as cells, microorganisms, and food residues, and transfer the supernatant to a new tube.

[0057] If protein metabolites are detected, add an appropriate amount of protease inhibitor (phenylmethanesulfonyl fluoride, PMSF), determine the concentration and dosage according to the inhibitor instructions, vortex mix, ice bath for 30 minutes and then store at -80℃. Thaw and vortex to re-dissolve before detection to ensure sample uniformity and stability and improve detection accuracy.

[0058] S2. Use the detection platform to detect m metabolites in the sample to construct the initial data matrix X of the metabolites n×m , where the initial data matrix Xn×m Every element x in ij represents the raw signal intensity of the jth metabolite in the ith sample. The detection platforms include: liquid chromatography-mass spectrometry, gas chromatography-mass spectrometry, nuclear magnetic resonance spectroscopy, and capillary electrophoresis;

[0059] Specifically, taking liquid chromatography-mass spectrometry detection as an example, 200 metabolites (i.e., m = 200) were detected, and a reversed-phase C18 column (2.1×100 mm, 1.8 μm particle size) was selected. Mobile phase A was water containing 0.1% formic acid, and mobile phase B was acetonitrile containing 0.1% formic acid. The flow rate was stabilized at 0.3 mL / min, and the column temperature was maintained at 35°C.

[0060] The mass spectrometer was operated in negative ion mode with an electrospray ion source, a spray voltage of -4.5 kV, an ion transfer tube temperature of 320 °C, a drying gas (nitrogen) flow rate of 10 L / min, and a mass scan range of 50-1500 m / z. 500 samples (n = 500) were tested one by one, and the retention time and mass-to-charge ratio of each metabolite were recorded. For example, the signal intensity of metabolite 1 in sample 1 was 1256.3 (x 11 ), the signal intensity of metabolite 2 is 892.1 (x 12 ) etc., and thus construct the initial data matrix X of 500×200 n×m .

[0061] S3. For the initial data matrix X n×m Perform preprocessing to obtain the preprocessed data matrix X' n×m ;

[0062] Specifically, for the initial data matrix X n×m For missing value processing, take metabolite 3 as an example. If there are missing values ​​in 10 samples, it is determined to be randomly missing after analyzing the sample distribution and metabolite characteristics. The mean interpolation method is used to calculate the mean signal intensity of metabolite 3 in the remaining 490 samples to be 765.4, which is used to fill the missing values. For noise data, take the signal data of metabolite 15 as an example, wavelet transform is used to remove noise, db4 wavelet basis function is selected, 3 layers of decomposition are performed, and threshold processing is performed according to the characteristics of high-frequency coefficients after decomposition (such as hard threshold is set to 30) to remove the noise component and reconstruct the signal;

[0063] For data normalization, linear normalization was used to process the data of metabolite 50. The minimum value of the original signal intensity was 120.5 and the maximum value was 1890.2. The original value of metabolite 50 in sample 2 was 850.6. The normalized value of 0.412 was calculated, and all metabolite data were normalized accordingly;

[0064] S4. Use principal component analysis to preprocess the data matrix X'n×m Perform singular value decomposition to obtain the direction of the principal components that characterize the differences in the sample and eliminate outliers in the sample;

[0065] The mathematical form of principal component analysis is Among them, U n×n is the left singular vector matrix, ∑ n×m is a diagonal matrix of singular values, is the right singular vector matrix, by observing that U n×n The score distribution of the sample is extracted, excluding outliers in the sample;

[0066] Specifically, for the preprocessed 500×200 data matrix X' n×m Perform principal component analysis singular value decomposition After calculating the covariance matrix, the singular values ​​are arranged in descending order. The first principal component has a singular value of 125.6, accounting for 35% of the total variance. The second principal component has a singular value of 82.3, accounting for 23% of the total variance. The cumulative variance contribution rate of the first five principal components is 78%. The main difference direction is determined and U is plotted. n×n The two-dimensional scatter plot of sample scores (first principal component and second principal component) is based on the cluster center and 3 times the standard deviation distance is set as the outlier judgment limit. For example, if sample 45 is more than 3 times the standard deviation (standard deviation is 25.2) from the cluster center in the direction of the first principal component, it is excluded as an outlier based on the comprehensive two-dimensional position judgment. After processing, 480 samples are retained to construct a new matrix. The principal components of the new matrix accurately describe the core dimensions of metabolic differences among people with different risk factors for diabetes.

[0067] S5. Use partial least squares discriminant analysis to preprocess the data matrix X' n×m Perform discrimination to obtain the projection vector W with the highest discrimination of the sample in the low-dimensional projection space max ;

[0068] The mathematical form of partial least squares discriminant analysis is Among them, T n×r is the score matrix, is the load matrix, E n×m is the residual matrix, by maximizing the score matrix T n×r The covariance between the column vectors in the sample is the maximum projection vector W in the low-dimensional projection space. max ;

[0069] Specifically, in partial least squares discriminant analysis In the process, the iterative optimization starts from r = 1. The first calculation is based on X' n×m The correlation between the column vector and the diabetes category dummy variable determines the first latent variable and constructs T n×1 and and residual matrix iteration;

[0070] Gradually increase r, and calculate the score matrix T each time n×r Column vector covariance evaluation model, for example, when r = 2, the covariance value is 45.2, when r = 3, the covariance is 62.8 and the cross-validation error rate is the lowest, thus determining the optimal r = 3 and the corresponding projection vector W max ,This vector accurately distinguishes high risk of diabetes, potential diabetes and healthy diabetes in the three-dimensional projection space, and defines the category boundaries based on the projection distribution;

[0071] S6. Filter the projection vector W max The k metabolites that satisfy the variable importance projection value VIP greater than the first threshold and the univariate statistical test P less than the second threshold are selected to obtain the standard data matrix X' n×k ;

[0072] The first threshold value ranges from 1.0 to 1.5, and the second threshold value ranges from 0.01 to 0.05;

[0073] Specifically, the variable importance projection value VIP and the univariate statistical test P value were calculated to screen key metabolites. For example, the VIP value of metabolite 25 was 1.28 (>1.0), and the concentration of metabolite 25 was significantly different in different diabetes risk groups after t-test (data normality), with a P value of 0.02 (>0.01) ∩ (<0.05), and was selected into the key metabolite set;

[0074] Traverse the projected vector metabolites, screen out 30 species, i.e. k = 30, and construct a 500×30 standard data matrix X' n×k .

[0075] S7. For a sample x to be tested test , screening standard data matrix X' n×k The satisfaction and test sample x in test The distance between the samples is less than the third threshold to predict the sample x to be tested. test diabetes risk level;

[0076] Specifically, calculate the sample x to be tested test and the standard data matrix X' n×k Sample distance, select Euclidean distance, assuming the third threshold is 1.2, such as the sample to be tested x test and the standard data matrix X' n×k The Euclidean distance is calculated for the samples in the middle, and the square root of the corresponding difference values ​​of metabolites 1 to 30 is summed and then the distance value is obtained. If the distance with 280 samples is less than 1.2, the prediabetes population N 1 There are 180 samples, accounting for 64.3% of the total, so the sample to be tested is predicted to be at an extremely high risk level; if the high-risk population of normal blood sugar diabetes N 2 The sample accounts for 40% to 60%, which is considered high risk; if the healthy population N3 If the sample exceeds 60%, it is considered low risk.

[0077] A computer-readable storage medium stores instructions, which, when executed by a computer, enable the computer to execute a metabolomics-based diabetes early warning method.

[0078] An electronic device, comprising: a processor and a memory, wherein:

[0079] The memory stores instructions, which, when executed by the processor, enable the processor to execute a diabetes early warning method based on metabolomics.

[0080] In summary, the metabolomics-based early warning method for diabetes disclosed in the present invention can achieve early and accurate prediction of diabetes risk by systematically analyzing the metabolite characteristics in samples, improve the accuracy and specificity of detection, effectively exclude outliers by principal component analysis, and combine partial least squares discriminant analysis to screen key metabolites, enhance the reliability and stability of the model, simplify the data processing process, support high-throughput detection platforms, and improve detection efficiency and feasibility of clinical applications.

[0081] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.

Claims

1. A metabolomics-based diabetes early warning method, characterized in that: The following steps are involved: S1. Collect samples from a total population n, wherein the total population n is divided into a prediabetic population N1, a high-risk population with normal blood sugar diabetes N2, and a healthy population N3, n = N1 + N2 + N3; S2. Use the detection platform to detect the m metabolites in the sample to construct the initial data matrix X of the metabolites n×m , where the initial data matrix X n×m Every element x in ij represents the original signal intensity of the j-th metabolite in the i-th sample; S3. The initial data matrix X n×m Perform preprocessing to obtain the preprocessed data matrix X′ n×m ; S4. Use principal component analysis to analyze the preprocessed data matrix X' n×m Performing singular value decomposition to obtain the direction of the principal component that characterizes the difference in the sample and to eliminate outliers in the sample; S5. Use partial least squares discriminant analysis to preprocess the data matrix X' n×m Perform discrimination to obtain the projection vector W with the maximum discrimination of the sample in the low-dimensional projection space max ; S6. Filter the projection vector W max The k metabolites that satisfy the variable importance projection value VIP greater than the first threshold and the univariate statistical test P less than the second threshold are selected to obtain the standard data matrix X′ n×k ; S7. For a sample x to be tested test , filter the standard data matrix X′ n×k The satisfaction and the test sample x test The distance between the samples is less than the third threshold, so as to predict the sample to be tested x test diabetes risk level.

2. The metabolomics-based diabetes early warning method according to claim 1, characterized in that: In step S1, the sample includes: blood V Bld 、Urine V Ur and saliva V Sal .

3. The metabolomics-based diabetes early warning method according to claim 1, characterized in that: In step S2, the detection platform includes: liquid chromatography-mass spectrometry, gas chromatography-mass spectrometry, nuclear magnetic resonance spectroscopy and capillary electrophoresis.

4. The metabolomics-based diabetes early warning method according to claim 1, characterized in that: In step S4, the mathematical form of the principal component analysis method is Among them, U n×n is the left singular vector matrix, ∑ n×m is a diagonal matrix of singular values, is the right singular vector matrix, by observing that U n×n The score distribution of the samples in , excluding outliers in the samples.

5. The metabolomics-based diabetes early warning method according to claim 1, characterized in that: In step S5, the mathematical form of the partial least squares discriminant analysis method is Among them, T n×r is the score matrix, is the load matrix, E n×m is the residual matrix, by maximizing the score matrix T n×r The covariance between the column vectors in the is used to obtain the projection vector W with the maximum discrimination of the sample in the low-dimensional projection space. max .

6. The metabolomics-based diabetes early warning method according to claim 1, characterized in that: In step S6, the first threshold value ranges from 1.0 to 1.5, and the second threshold value ranges from 0.01 to 0.

05.

7. A computer-readable storage medium, characterized in that: Instructions are stored thereon, and when the instructions are executed by a computer, the computer is caused to execute the metabolomics-based diabetes early warning method as claimed in any one of claims 1 to 6.

8. An electronic device, characterized in that: include: processor and memory, wherein The memory stores instructions, and when the instructions are executed by the processor, the processor executes the metabolomics-based diabetes early warning method according to any one of claims 1 to 6.