Metabolic aging recognition model construction system, storage medium and kit
By constructing a metabolic aging recognition model and using machine learning of plasma small molecule metabolic biomarkers, the problem of inconsistent assessment of individual metabolic aging status is solved, and efficient and low-cost metabolic aging risk prediction and drug screening are achieved.
Patent Information
- Application Number
- CN202510764292.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The prior art is difficult to accurately evaluate individual metabolic aging status, resulting in inconsistent assessment of metabolic aging risk in individuals of the same age, and lack of reliable metabolic biomarker-based identification models.
A metabolic aging recognition model is constructed, and small molecule metabolic biomarkers in plasma samples are collected and analyzed, characteristic markers are screened using machine learning methods, and identification models are trained to predict metabolic age based on actual age to evaluate the risk of aging.
Accurate assessment of metabolic aging is achieved, the predictive sensitivity and specificity of all-cause death risk is improved, the detection cost is reduced, and it is convenient for large-scale promotion, and can be used for anti-aging drug screening.
Smart Images

Figure CN120280160A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical prediction, and particularly to a system for constructing an identification model of metabolic aging, a storage medium, and a kit. Background Art
[0002] As a key link in the aging process, metabolic aging refers to a biological process in which the metabolic network in an organism gradually malfunctions during the aging degenerative process, leading to the decline of energy balance, oxidative stress, and cell repair functions. Its core feature is the progressive loss of metabolic homeostasis, manifested as a decline in mitochondrial function, abnormal nutrient sensing pathways, accumulation of harmful metabolites, and changes in epigenetic modifications. This process not only accelerates the degradation of organ functions (such as muscle atrophy and decreased liver metabolic capacity), but is also closely related to various age-related diseases such as diabetes, cardiovascular diseases, and neurodegenerative diseases. Recent studies have revealed that by intervening in metabolic pathways, the aging process can be significantly delayed, making metabolic aging a key target in the anti-aging field.
[0003] Traditionally, the assessment of metabolic aging from the perspective of the beginning of life (birth) depends on age. As age increases, metabolic homeostasis is disrupted, and the process of metabolic aging deepens. However, individuals of the same age often exhibit different metabolic aging states, resulting in different individual risk profiles. For example, an elderly person has already developed various metabolic abnormalities such as diabetes and obesity, while another elderly person of the same age is well maintained, has relatively healthy metabolism, and has a longer life expectancy than individuals of the same age, with a relatively mild degree of true metabolic aging. In addition, age is calculated based on the date of birth and increases constantly, making it difficult to judge the increase or decrease in a person's metabolic rate. The assessment of metabolic aging from the perspective of the end of life (death) depends on the length of the remaining life. However, the length of the remaining life requires long-term follow-up records of all-cause death events for individuals, and there is still a lack of a reliable identification model of metabolic aging based on a combination of metabolite biomarkers. Summary of the Invention
[0004] The purpose of the present invention is to provide a system for constructing an identification model of metabolic aging, a storage medium, a kit, and their applications.
[0005] To solve the above problems, the present invention provides a system for constructing an identification model of metabolic aging, including: A collection module, configured to collect plasma samples of each subject who has experienced an all-cause death event and those who have not experienced an all-cause death event during a 5-year follow-up period, and obtain plasma small molecule metabolic biomarkers and their relative concentration values of the subjects who have experienced an all-cause death event and those who have not experienced an all-cause death event during the 5-year follow-up period from each plasma sample; A feature selection module, which is used to perform machine learning based on the plasma small molecule metabolite biomarkers and their relative concentration values of the subjects who do not experience all-cause death events during the 5-year follow-up period and the subjects who experience all-cause death events during the 5-year follow-up period, so as to screen out the plasma small molecule metabolite biomarkers that can identify the subjects who do not experience all-cause death events during the 5-year follow-up period and the subjects who experience all-cause death events during the 5-year follow-up period, and use them as the plasma small molecule metabolite biomarkers after feature selection; A training module, which trains a metabolic aging recognition model for identifying metabolic age based on the plasma small molecule metabolite biomarkers after feature selection, their relative concentration values, and the actual ages of each subject who does not experience all-cause death events during the 5-year follow-up period and the subjects who experience all-cause death events during the 5-year follow-up period; An identification module, which obtains the age of the person to be tested, the plasma small molecule metabolite biomarkers after feature selection, and their relative concentration values, inputs the age of the person to be tested, the plasma small molecule metabolite biomarkers after feature selection, and their relative concentration values into the metabolic aging recognition model to obtain the metabolic age of the person to be tested, and based on the metabolic age, the all-cause death risk of the person to be tested can be predicted.
[0006] Further, in the above system, a collection module is used to collect plasma samples of the subjects who do not experience all-cause death events during the 5-year follow-up period and the subjects who experience all-cause death events during the 5-year follow-up period, and obtain the plasma small molecule metabolite biomarkers and their relative concentration values of the subjects who do not experience all-cause death events during the 5-year follow-up period and the subjects who experience all-cause death events during the 5-year follow-up period from each plasma sample, including: The collection module is used to obtain non-overlapping first and second subject groups, where both the first and second subject groups include each subject who experiences and does not experience all-cause death events during the 5-year follow-up period without overlap; The collection module is used to collect the first plasma samples of the first subject group, and obtain the plasma small molecule metabolite biomarkers and their relative concentration values from each of the first plasma samples as the first set; The collection module is used to collect the second plasma samples of the second subject group, and obtain the plasma small molecule metabolite biomarkers and their relative concentration values from each of the second plasma samples as the test set.
[0007] Further, in the above system, the feature selection module performs machine learning based on the plasma small molecule metabolite biomarkers and their relative concentration values of the subjects who do not experience all-cause death events during the 5-year follow-up period and the subjects who experience all-cause death events during the 5-year follow-up period, so as to screen out the plasma small molecule metabolite biomarkers that can identify the subjects who do not experience all-cause death events during the 5-year follow-up period and the subjects who experience all-cause death events during the 5-year follow-up period, and use them as the plasma small molecule metabolite biomarkers after feature selection, including: The feature selection module is used to divide the first set into 5 non-overlapping data portions. Each time, one of the 5 non-overlapping data portions that has not been selected is used as the internal validation set, and the remaining 4 data portions are used as the training set each time; among them, both the internal validation set and the training set include each subject who has experienced an all-cause death event during the 5-year follow-up period and who has not experienced an all-cause death event during the 5-year follow-up period without overlap. The feature selection module is used to, based on each training set, use the Lasso feature selection algorithm on the R (4.3.1) software, and train the feature selection model in a loop for 5 rounds respectively to obtain the corresponding 5 rounds of feature selection models. For each λn value in the hyperparameter sequence, a corresponding feature selection model is generated Mk ( λn ), n = 1 - 100; k = 1 - 5, k represents the number of rounds; n represents the serial number of the λ value; there are a total of 100; Here, it is assumed that the value sequence of λ is { λ 1, λ 2,…, λ 100}, in the 5-fold cross-validation, for each round of cross-validation, taking the k-th round as an example, k = 1 - 5; The feature selection module is used to, for each value λn in the hyperparameter sequence, input the internal validation set corresponding to each round into the feature selection model of the corresponding round in the 5 feature selection models Mk ( λn ), and respectively obtain a misclassification rate corresponding to each λn value on the internal validation set of the corresponding round; from the 5 misclassification rates corresponding to each value, obtain an average misclassification rate corresponding to each value; based on the minimum average misclassification rate among all values, select the λ .min when the corresponding misclassification rate is the lowest, where λ ( lambda) is one of the hyperparameters; based on λ .min when the misclassification rate is the lowest, obtain the corresponding metabolite biomarker, which is used as the plasma small molecule metabolite biomarker after feature selection.
[0008] Here, based on lambda.min being the best hyperparameter lambda, extract the corresponding features to screen out 14 metabolite biomarkers.
[0009] Furthermore, in the above system, the plasma small molecule metabolite biomarker after feature selection at least includes the following metabolite biomarkers in peripheral plasma: D-Glucose, L-Arginine, Taurine, L-Threonine, Creatinine, L-Acetylcarnitine, L-Serine, Uric acid, Guanidoacetic acid, Pipecolic acid, L-Aspartic acid, Oxoglutaric acid, L-Leucine, and L-Phenylalanine.
[0010] Further, in the above system, the training module selects the plasma small molecule metabolic biomarkers, their relative concentration values, and the actual age of each subject who does not experience an all-cause death event during the 5-year follow-up period and who experiences an all-cause death event during the 5-year follow-up period in the first set, and trains a metabolic aging recognition model for identifying metabolic age, including: The training module is used to train a metabolic aging recognition model based on the plasma small molecule metabolic biomarkers, their relative concentration values, and the actual age selected from the features in each plasma sample in the first set.
[0011] Further, in the above system, the training module trains a metabolic aging recognition model based on the plasma small molecule metabolic biomarkers, their relative concentration values, and the actual age selected from the features in each plasma sample, including: The training module is used to train a metabolic aging recognition model for identifying metabolic age. Among them, the metabolic aging recognition model for identifying metabolic age uses the ElasticNet algorithm on R (4.3.1) software and is based on the plasma small molecule metabolic biomarkers, their relative concentration values, and the actual age selected from the features of each subject who does not experience an all-cause death event during the 5-year follow-up period and who experiences an all-cause death event during the 5-year follow-up period in the first set, and automatically outputs a metabolic-based all-cause death risk score P for each subject in the first set. The range of P is from 0 to 1, representing the low metabolic death risk to the high metabolic death risk; the implementation of the ElasticNet algorithm depends on the glmnet package, where the ElasticNet algorithm hyperparameter alpha for controlling the ratio of L1 regularization and L2 regularization is set to 0.5; the s for controlling the penalty strength on the parameters is set to 0.01; The metabolic aging recognition model for identifying metabolic age calculates the mean and standard deviation of the ages of the subjects in the first set, and calculates the mean and standard deviation of the metabolic-based all-cause death risk scores P of all the subjects in the first set; scales the mean and standard deviation of the metabolic-based all-cause death risk scores P to be consistent with the mean and standard deviation of the ages of the subjects in the first set, thereby obtaining the metabolic age, in years, for identifying metabolic aging. The specific formula is as follows:
[0012] Where: Metabolic age i is the metabolic age after conversion of the i th sample in the first set; P i is the metabolic-based all-cause death risk score value of the i th subject; μ p is the mean of the metabolic-based all-cause death risk scores P of all the subjects in the first set; σ p is the sample standard deviation of the metabolic-based all-cause death risk scores P of all the subjects in the first set; μ Age is the mean of the ages of the subjects in the first set; σ Age is the standard deviation of the ages of the subjects in the first set; The training module is used to evaluate the performance of the metabolic age output by the trained metabolic aging model, and take the metabolic aging model whose performance evaluation meets the preset requirements as the final metabolic aging model for identifying metabolic age and accurately predicting the all-cause death risk based on the metabolic age.
[0013] Further, in the above system, the training module uses the trained metabolic aging model to obtain the recognition performance of the trained metabolic aging model on the first set and the recognition performance of the metabolic age on the test set; if the difference in the recognition performance of the metabolic age of the trained metabolic aging model between the first set and the test set is less than the preset difference threshold and both are greater than the actual age of the subjects themselves, then the trained metabolic aging model is taken as the final metabolic aging model.
[0014] According to another aspect of the present invention, there is also provided a computer-readable storage medium, on which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor performs the following steps: Step S1: The acquisition module collects plasma samples of each subject who has experienced an all-cause death event and those who have not experienced an all-cause death event during the 5-year follow-up period, and obtains the plasma small molecule metabolic biomarkers and their relative concentration values of the subjects who have experienced an all-cause death event and those who have not experienced an all-cause death event from each plasma sample; Step S2: The feature selection module performs machine learning based on the plasma small molecule metabolite biomarkers and their relative concentration values of the subjects who do not experience all-cause death events during the 5-year follow-up period and the subjects who experience all-cause death events during the 5-year follow-up period, so as to screen out the plasma small molecule metabolite biomarkers that can identify the subjects who do not experience all-cause death events and the subjects who experience all-cause death events during the 5-year follow-up period, and use them as the plasma small molecule metabolite biomarkers after feature selection; Step S3: The training module trains the metabolic aging recognition model based on the plasma small molecule metabolite biomarkers after feature selection, their relative concentration values and the actual ages of the subjects who do not experience all-cause death events and the subjects who experience all-cause death events during the 5-year follow-up period to obtain the metabolic age; Step S4: The recognition module obtains the age of the person to be tested and the plasma small molecule metabolite biomarkers after feature selection and their relative concentration values, and inputs the age of the person to be tested and the plasma small molecule metabolite biomarkers after feature selection and their relative concentration values into the metabolic aging recognition model to obtain the metabolic age of the person to be tested, which can predict the all-cause death risk of the person to be tested.
[0015] According to another aspect of the present invention, there is also provided an application of a metabolite biomarker in a product for identifying metabolic aging. The metabolite biomarker at least includes the following metabolite biomarkers: D-Glucose, L-Arginine, Taurine, L-Threonine, Creatinine, L-Acetylcarnitine, L-Serine, Uric acid, Guanidoacetic acid, Pipecolic acid, L-Aspartic acid, Oxoglutaric acid, L-Leucine and L-Phenylalanine.
[0016] According to another aspect of the present invention, there is also provided a test kit for detecting metabolic age. The metabolite biomarkers in the test kit at least include the following metabolite biomarkers: D-Glucose, L-Arginine, Taurine, L-Threonine, Creatinine, L-Acetylcarnitine, L-Serine, Uric acid, Guanidoacetic acid, Pipecolic acid, L-Aspartic acid, Oxoglutaric acid, L-Leucine and L-Phenylalanine; The metabolite biomarkers are subjected to feature selection and identification through the following modules: A collection module that collects plasma samples of each subject who experienced an all-cause death event and those who did not during the 5-year follow-up period, and obtains the plasma small molecule metabolite biomarkers and their relative concentration values of the subjects who experienced an all-cause death event and those who did not during the 5-year follow-up period from each plasma sample; A feature selection module that performs machine learning based on the plasma small molecule metabolite biomarkers and their relative concentration values of the subjects who did not experience an all-cause death event and those who did experience an all-cause death event during the 5-year follow-up period, to screen and obtain the plasma small molecule metabolite biomarkers that can identify the subjects who did not experience an all-cause death event and those who did experience an all-cause death event during the 5-year follow-up period, as the plasma small molecule metabolite biomarkers after feature selection; A training module that trains a metabolic aging recognition model for identifying metabolic age based on the plasma small molecule metabolite biomarkers after feature selection, their relative concentration values, and the actual age of each subject who did not experience an all-cause death event and those who did experience an all-cause death event during the 5-year follow-up period; An identification module that obtains the age of the person to be tested and the plasma small molecule metabolite biomarkers after feature selection and their relative concentration values, inputs the age of the person to be tested and the plasma small molecule metabolite biomarkers after feature selection and their relative concentration values into the metabolic aging recognition model to obtain the metabolic age of the person to be tested, and predicts the all-cause death risk of the person to be tested based on the metabolic age.
[0017] Compared with the prior art, through machine learning of age and plasma metabolite biomarkers, the present invention obtains a metabolic aging recognition model. According to the calculated metabolic age, it shows excellent sensitivity, specificity, and accuracy in predicting the risk of all-cause death, exceeding the actual age itself. Experiments have confirmed that it can intuitively reflect the acceleration of metabolic aging by comparing with the actual age. The present invention only requires a small amount of plasma as a detection sample, the evaluation process is simple and convenient, the cost is low, it is easy to be accepted by subjects, and it is suitable for large-scale promotion of the method. Moreover, due to its high sensitivity and specificity, it can also be used in the screening of anti-aging drugs. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 Performance of metabolic age and actual age obtained based on the metabolic aging recognition model in identifying the 5-year all-cause death risk in the training set in an embodiment of the present invention; Figure 2 Performance of metabolic age and actual age obtained based on the metabolic aging recognition model in identifying the 5-year all-cause death risk in the test set in an embodiment of the present invention; Figure 3 5-year cumulative all-cause death risk curves for the top 25% of metabolic age acceleration (metabolic age - actual age > 2 years) and the remaining individuals based on the test set in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The present invention will be further described in detail below with reference to the accompanying drawings.
[0020] In a typical configuration of the present application, the terminal, the device of the service network, and the trusted party all include one or more processors (CPUs), input / output interfaces, network interfaces, and memories.
[0021] The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash RAM. The memory is an example of the computer-readable medium.
[0022] A computer-readable medium includes permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory media such as modulated data signals and carrier waves.
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will elaborate on each embodiment of the present invention in detail in combination with the embodiments. For experimental methods where specific conditions are not indicated in the embodiments, they are generally in accordance with conventional conditions, such as those described in textbooks and experimental guides, or in accordance with the conditions recommended by the manufacturer. Run under the recommended conditions of the supporting software.
[0024] The present invention provides a system for constructing a metabolic aging recognition model, a storage medium, a kit, and their applications. The method includes: Step S1 to Step S4.
[0025] Step S1, a collection module, collects plasma samples of each subject who experienced an all-cause death event and those who did not during a 5-year follow-up period, and obtains the plasma small molecule metabolic biomarkers and their relative concentration values of the subjects who experienced an all-cause death event and those who did not during the 5-year follow-up period from each plasma sample; Here, metabolomics is an important branch of systems biology, mainly studying the qualitative and quantitative analysis of all metabolites in an organism. It comprehensively detects small molecule metabolites in biological samples, including sugars, lipids, organic acids, etc., through high-throughput analysis techniques. These metabolites reflect the systemic metabolic state of an organism at a specific time point and can be used for disease diagnosis, drug research and development, and biomarker discovery.
[0026] All-cause death is a core indicator in medical and epidemiological research, referring to the aggregation of all death events regardless of the specific cause. The higher the all-cause death risk during a 5-year follow-up, the shorter the lifespan, the more significant the metabolic disorders, the greater the metabolic age, and the deeper the metabolic aging process.
[0027] During the 5-year follow-up period, the population with all-cause death events had a short life expectancy and a high risk of aging, showing a different metabolic aging model from those without all-cause death events during the 5-year follow-up period. Therefore, plasma metabolomics can be a useful resource for metabolic aging assessment and will strongly promote healthy aging.
[0028] The relative concentration value refers to the relative concentration value of a certain type of plasma small molecule metabolite biomarker in plasma, and its unit can be %. Preferably, step S1 includes: Step S11, the acquisition module obtains a non-overlapping first subject group and a second subject group, where both the first subject group and the second subject group include individual subjects who had and did not have all-cause death events during the non-overlapping 5-year follow-up period; Here, the first subject group includes individual subjects who had and did not have all-cause death events during the 5-year follow-up period; the second subject group includes individual subjects who had and did not have all-cause death events during the 5-year follow-up period; the first subject group and the second subject group do not overlap; Step S12, the acquisition module collects the first plasma samples of the first subject group, and obtains plasma small molecule metabolite biomarkers and their relative concentration values from each of the first plasma samples as the first set; Step S13, the acquisition module collects the second plasma samples of the second subject group, and obtains plasma small molecule metabolite biomarkers and their relative concentration values from each of the second plasma samples as the test set; Specifically, for the group of subjects, subjects selected from two centers from January 2019 to January 2020 can be recruited.
[0029] For the subjects in the first set used for training and validating the prediction model, they can be 9365 subjects who did not have all-cause death events during the 5-year follow-up period and 351 subjects who had all-cause death events during the 5-year follow-up period from Center 1; subsequently, the training set can be split into non-overlapping internal training set and internal validation set.
[0030] The test set for the prediction model can be 2227 subjects who did not have all-cause death events during the 5-year follow-up period and 93 subjects who had all-cause death events during the 5-year follow-up period from Center 2.
[0031] The all-cause death events of the subjects can be recorded according to the regular home follow-up of the researchers' families and the registration of local medical and health institutions. Subjects who had all-cause death events and those who did not have all-cause death events during the 5-year follow-up period collected plasma at the beginning of the study (Year 0) and detected metabolite biomarkers in a nanoparticle-enhanced laser desorption / ionization mass spectrometer.
[0032] Preferably, the exclusion criteria within the inclusion criteria are subjects with acute and infectious clinical symptoms within three weeks before sampling, including but not limited to fever, headache, cough, sore throat, loss of smell, runny nose, abdominal pain, diarrhea, etc.
[0033] Step S12 or step S13 may include: steps S121 to S128.
[0034] In step S121, the plasma sample of the subject's peripheral venous blood in the early morning fasting state is anticoagulated with EDTA and stored at -80°C. In step S122, preparation of the instrument and reagents: Prepare a nanoparticle enhanced laser desorption / ionization time-of-flight mass spectrometer (MALDI-TOF-MS) and related reagents. Here, the nanoparticle enhanced laser desorption / ionization time-of-flight mass spectrometer may be, for example, a product of Bruker Corporation in Germany. In step S123, dilution treatment of the plasma sample: Dilute the plasma sample with deionized water in a standard ratio to obtain a diluted plasma sample, ensuring that the metabolite concentration is suitable for mass spectrometry analysis. Specifically, 100 nL of the plasma sample can be diluted 10 times with deionized water to obtain a diluted plasma sample. In step S124, prepare an inorganic nanoparticle matrix solution for enhancing the mass spectrometry signal. Specifically, prepare a 1 mg / mL matrix solution of the matrix (inorganic nanoparticle) with deionized water. In step S125, sample preparation on the mass spectrometry target plate: Spot the diluted plasma sample onto the mass spectrometry target plate and dry it at room temperature. Specifically, sample preparation can be carried out on the mass spectrometry target plate. 500 nL of each diluted plasma sample is spotted and dried at room temperature. In step S126, matrix preparation on the mass spectrometry target plate: Spot the nanoparticle matrix solution onto the mass spectrometry target plate to ensure that the nanoparticle matrix solution uniformly covers the plasma sample spot to obtain a sample. Specifically, matrix preparation can be carried out on the mass spectrometry target plate. 500 nL of each matrix solution is spotted on each plasma sample spot and dried at room temperature. In step S127, data acquisition of the mixture in the nanoparticle enhanced laser desorption / ionization time-of-flight mass spectrometer: Use MALDI-TOF-MS to collect the mass spectrometry data of metabolites in the plasma, and obtain the metabolite peak map of the sample through delayed extraction and time-of-flight analysis.
[0035] Specifically, data acquisition is performed in a nanoparticle-enhanced laser desorption / ionization time-of-flight mass spectrometer. The acquired data is extracted in positive ion mode, with delayed extraction, a repetition rate of 1000 hz, an acceleration voltage of 20 kV, a delay time of 250 ns, and 2000 laser pulses for each analysis.
[0036] For step S128, based on the metabolite peak map in the sample under the recommended conditions of the software supporting the nanoparticle-enhanced laser desorption / ionization time-of-flight mass spectrometer, an in-built data preprocessing pipeline can be run, including: data resampling, spectral smoothing, baseline correction, and spectral peak alignment, so as to obtain the relative concentration values of 303 metabolic biomarkers in the full spectrum.
[0037] Specifically, a Savitzky-Golay (S-G) filter can be used for spectral smoothing; then an adaptive iterative algorithm is used to identify baseline points, polynomial fitting is performed on the identified baseline points in the metabolite peak map, the fitted baseline is subtracted from the original spectrum, and the corrected spectrum is checked to ensure that the baseline is flat and there are no negative values; then the peak detection parameters are set: signal-to-noise ratio threshold = 3, mass accuracy tolerance 50 ppm, the continuous wavelet transform (CWT) is used to identify the possible metabolite biomarker peak positions, and further peak shape fitting is performed to obtain the corresponding peak positions, peak heights, and peak areas, and relative quantification of the metabolite biomarker levels is performed according to the peak height; finally, Min-Max normalization is performed to obtain the relative concentration values of 303 biomarkers.
[0038] The relative concentration values obtained here can be the peak intensity values after min-max normalization, with the unit of %, which are used to measure the relative concentration of metabolites in plasma.
[0039] 303 plasma small molecule metabolic biomarkers can be identified and recognized by comparing with the HMDB (Human Metabolome Database).
[0040] For step S2, based on the plasma small molecule metabolic biomarkers and their relative concentration values of the subjects who do not experience all-cause death events during the 5-year follow-up period and the subjects who experience all-cause death events during the 5-year follow-up period, machine learning is performed to screen and obtain the plasma small molecule metabolic biomarkers that identify the subjects who do not experience all-cause death events during the 5-year follow-up period and the subjects who experience all-cause death events during the 5-year follow-up period, as the plasma small molecule metabolic biomarkers after feature selection; More preferably, step S2 includes: Step S20: Divide the first set into 5 non-overlapping data portions. Each time, take 1 non-selected data portion from the 5 non-overlapping data portions as the internal validation set, and take the remaining 4 data portions as the training set each time. Among them, both the internal validation set and the training set include each subject who has experienced an all-cause death event during the non-overlapping 5-year follow-up period and each subject who has not experienced an all-cause death event during the 5-year follow-up period. Step S21: Based on each training set, the feature selection module uses the Lasso feature selection algorithm on the R (4.3.1) software and trains the feature selection model 5 rounds respectively to obtain the corresponding 5 feature selection models. For each λn value in the hyperparameter sequence in each round of the feature selection model, a corresponding feature selection model is generated. Mk ( λn ), n = 1 - 100; k = 1 - 5, where k represents the number of rounds; n represents the serial number of the λ value; there are 100 in total. Here, assume that the value sequence of λ is { λ 1, λ 2, …, λ 100}. In 5-fold cross-validation, for each round of cross-validation, take the k-th round as an example, k = 1 - 5. For each value in the sequence of 100 values corresponding to the hyperparameter sequence lambda, the feature selection module λn inputs the corresponding internal validation set of each round into one feature selection model of the corresponding round in the 5 feature selection models Mk ( λn ), and respectively obtains a misclassification rate corresponding to each λn value on each internal validation set. From the 5 misclassification rates corresponding to each value, an average misclassification rate corresponding to each value is obtained. Based on the minimum average misclassification rate among all the values, select lambda.min corresponding to the lowest misclassification rate. Among them, lambda is one of the hyperparameters. Based on lambda.min ( λ .min), 14 metabolite markers are obtained.
[0041] Here, based on lambda.min being the best hyperparameter lambda, extract the corresponding features to screen out 14 metabolite markers. Specifically, the lambda hyperparameters of the five feature selection models are the same. For each lambda value, once the lambda value is determined, the features used are determined, and the corresponding metabolite markers can be mapped based on these features. 5 feature selection models can be obtained through the training set. By looking at the average performance of the lambda value in the internal validation sets of the five models, determine which lambda value has the best average performance in the internal validation set, that is, the lowest average misclassification rate.
[0042] R software is a language and environment for statistical computing and graphics. It supports multiple platforms such as Windows, Mac, and Linux. It is a GNU project and is similar to the S language and environment developed by John Chambers and his colleagues at Bell Laboratories (formerly AT&T, now Lucent Technologies). R can be considered as a different implementation of S. There are some important differences, but much of the code written for S remains unchanged when running under R. R provides a wide variety of statistical (linear and non-linear modeling, classical statistical tests, time series analysis, classification, clustering, etc.) and graphical techniques and is highly scalable. The S language is usually the preferred tool for statistical method research, and the R language provides an open-source way to participate in this activity.
[0043] Preferably, the plasma small molecule metabolite biomarkers after feature selection at least include the following metabolite biomarkers in peripheral plasma: D-Glucose, L-Arginine, Taurine, L-Threonine, Creatinine, L-Acetylcarnitine, L-Serine, Uric acid, Guanidoacetic acid, Pipecolic acid, L-Aspartic acid, Oxoglutaric acid, L-Leucine, and L-Phenylalanine.
[0044] The specific information is shown in Table 1.
[0045] Table 1: Information related to 14 metabolite biomarkers
[0046] Here, the 14 plasma metabolite biomarkers and age obtained by the feature selection module of the present invention are used for subsequent training of the metabolic aging recognition model, which can exhibit excellent all-cause death risk prediction ability; by performing machine learning on the plasma small molecule metabolite biomarkers after feature selection, efficient evaluation of metabolic aging can be achieved.
[0047] Step S3: The training module trains a metabolic aging recognition model for identifying metabolic age based on the plasma small molecule metabolite biomarkers, their relative concentration values, and the actual ages of the respective subjects in the first set who did not experience all-cause death events during the 5-year follow-up period and those who did experience all-cause death events during the 5-year follow-up period, after feature selection. Preferably, Step S3 includes: Step S311: The training module is used to train a metabolic aging recognition model for identifying metabolic age. Among them, the metabolic aging recognition model for identifying metabolic age uses the ElasticNet algorithm on R (4.3.1) software and is based on the plasma small molecule metabolite biomarkers, their relative concentration values, and the actual ages of the respective subjects in the first set who did not experience all-cause death events during the 5-year follow-up period and those who did experience all-cause death events during the 5-year follow-up period, after feature selection. It automatically outputs a metabolic-based all-cause death risk score P for each subject. The range of P is from 0 to 1, representing the risk of metabolic death from low to high. The implementation of the ElasticNet algorithm depends on the glmnet package. Among them, the ElasticNet algorithm hyperparameter alpha for controlling the ratio of L1 regularization and L2 regularization is set to 0.5; s for controlling the penalty strength on the parameters is set to 0.01. Here, differentiating between subjects who experienced all-cause death events and those who did not during the 5-year follow-up period is to output a metabolic-based all-cause death risk score P.
[0048] Step S312: The metabolic aging recognition model for identifying metabolic age calculates the mean and standard deviation of the ages of the subjects in the first set, and calculates the mean and standard deviation of the metabolic-based all-cause death risk scores P of all subjects in the first set; scales the mean and standard deviation of the metabolic-based all-cause death risk scores P to be consistent with the mean and standard deviation of the ages of the subjects in the first set, thereby obtaining the metabolic age, in years, for identifying metabolic aging. The formula is as follows: (1) Where: Metabolic age i is the converted metabolic age of the i th sample in the first set; P i is the metabolic-based all-cause death risk score value of the i th subject; μ p is the mean of the metabolic-based all-cause death risk scores P of all subjects in the first set; σ p is the standard deviation of the ages of the subjects in the first set; μ Age is the mean of the ages of the subjects in the first set; σThe standard deviation of the ages of the subjects in the first set; Step S32: The training module performs performance evaluation on the metabolic age output by the trained metabolic aging model, and uses the metabolic age model whose performance evaluation meets the preset requirements as the final metabolic aging recognition model for accurately predicting the all-cause death risk and identifying metabolic aging; More preferably, step S32 includes: Step S321: The training module obtains the performance of the trained metabolic aging recognition model in predicting the 5-year all-cause death risk on the first set and the performance of predicting the 5-year all-cause death risk on the test set on the R (4.3.1) software; if the prediction performances of the trained metabolic aging recognition model on the training set and the test set differ by less than the preset difference threshold and are both greater than the actual age of the subjects themselves, it indicates that there is no overfitting risk and the metabolic aging assessment performance is greater than the age itself, then the trained metabolic aging recognition model is used as the final metabolic aging recognition model.
[0049] Here, all-cause death is a core indicator in medical and epidemiological research, referring to the aggregation of all death events regardless of the specific cause. The higher the all-cause death risk during the 5-year follow-up, the shorter the lifespan, the more significant the metabolic disorders, the greater the metabolic age, and the deeper the metabolic aging process.
[0050] Therefore, a metabolic-based all-cause death risk score P (the range of P is from 0 to 1, representing the all-cause death risk from low to high, and the proportion of the actual number of deaths in the training cohort, i.e., 0.04, is used as the prediction cut-off value) can be trained by recording the 5-year all-cause death events of the subjects in combination with the selected plasma small molecule metabolic biomarkers and machine learning, and the metabolic age can be calculated based on P, as well as the difference between the metabolic age and the actual age to intuitively identify the level of metabolic aging and whether the metabolic aging is accelerating compared with peers. Similarly, the all-cause death risk can be intuitively predicted according to the P corresponding to the metabolic age.
[0051] After obtaining the metabolic age from the plasma small molecule metabolic biomarkers and age in the first set and the test set, the area under the receiver operating characteristic curve (AUC), accuracy, specificity, and sensitivity are used to evaluate the prediction performance of the metabolic age for the all-cause death risk.
[0052] The training module uses the trained metabolic aging recognition model on R (4.3.1) software to make predictions on the first set and the test set respectively, and obtains the recognition performances (recognition capabilities) on the first set and the test set. The recognition performances include: the AUC values are 0.728 for the training set and 0.752 for the test set, showing high accuracy and sensitivity. The specific performance indicators are shown in Table 2. If the recognition performance curves of the first set and the test set are close to each other (Delong test P = 0.26 > 0.05), and both are greater than the actual age (such as Figure 1 and 2 shown), it indicates that there is no overfitting risk and the metabolic aging assessment performance is greater than the actual age, indicating that it is a trained metabolic aging recognition model and can be used as the final metabolic aging recognition model.
[0053] The metabolic aging recognition model provided by the present invention has the characteristics of less sample consumption and high reproducibility. By comparing the performances of the training set and the test set, it is judged whether the model for metabolic aging recognition has the risk of overfitting. If the performances are close, it indicates that the model has generality. If they are not close, it means that the generalization ability of the model is relatively poor and a better model needs to be trained. By comparing the performances of the metabolic age and the actual age, it is judged whether the performance of the model for metabolic aging recognition is better than the age. If it is better, it indicates that the model captures the metabolic disorders that cannot be reflected by the simple age and is a superior metabolic aging assessment model.
[0054] Table 2. Performance of machine learning obtained from the metabolic aging recognition model and age
[0055] The model prediction results are represented by a 2x2 confusion matrix, where the rows represent the actual labels and the columns represent the predicted labels. The confusion matrix classifies the samples into four categories, namely: TP (True Positive) True positive: The number of samples that are actually positive and predicted to be positive.
[0056] TN (True Negative) True negative: The number of samples that are actually negative and predicted to be negative.
[0057] FP (False Positive) False positive: The number of samples that are actually negative and predicted to be positive.
[0058] FN (False Negative) False negative: The number of samples that are actually positive and predicted to be negative.
[0059] Accuracy: It is the proportion of samples with correct predictions in the total samples, and the calculation formula is (TP + FP) / (FN + TN + TP + TN).
[0060] F1 value: It is a comprehensive index of the proportion of samples actually being positive cases that are predicted as positive cases and the proportion of correctly predicted amounts among positive case samples. The calculation formula is 2*P*R / (P + R), where P = TP / (TP + FP) and R = TP / (TP + FN).
[0061] Sensitivity: It is the proportion of samples actually being positive cases that are predicted as positive cases. The calculation formula is TP / (TP + FN).
[0062] Specificity: It is the proportion of samples actually being negative cases that are predicted as negative cases. The calculation formula is TN / (TN + FP).
[0063] AUC: The area under the ROC curve. The ROC curve is a curve composed of 1 - specificity and sensitivity at different thresholds. AUC is not affected by the proportion of positive and negative samples. It reflects the overall performance of the model at different thresholds, ranging from 0 to 1. The larger it is, the better the comprehensive performance of the model.
[0064] In step S4, the recognition module obtains the age of the person to be tested and the plasma small - molecule metabolite biomarkers and their relative concentration values after feature selection, and inputs the age of the person to be tested and the plasma small - molecule metabolite biomarkers and their relative concentration values after feature selection into the final metabolic aging recognition model to obtain the metabolic age of the person to be tested, which can accurately predict the all - cause death risk.
[0065] Here, the probabilities of not having an all - cause death event during the 5 - year follow - up period and having an all - cause death event during the 5 - year follow - up period are actually the all - cause death risk score P based on metabolism mentioned above. The range of P is from 0 to 1, representing the all - cause death risk from low to high. The proportion of the actual number of deaths in the first set, which is 0.04, is used as the prediction cut - off value.
[0066] The metabolic age is P, which can be calculated through the aforementioned formula (1).
[0067] Obviously, the greater the metabolic age, the higher the probability of all - cause death risk. In addition, the metabolic age can be used to intuitively reflect the level of metabolic aging and whether the metabolic aging is accelerating compared with peers, resulting in an increased all - cause death risk. For example, in the test set, a person with an actual age of 65 years has a metabolic age of 65 years (the corresponding P is 0.017), and there is no acceleration of metabolic aging. While another person with an actual age of 65 years has a metabolic age of 70 years (the corresponding P is 0.046), with an acceleration of metabolic aging by 5 years, and thus a higher all - cause death risk.
[0068] The acquisition method of the plasma small - molecule metabolite biomarkers and their relative concentration values after feature selection of the person to be tested can refer to the acquisition method in step 1.
[0069] In addition, an experimental verification of the ability of the model to predict the future all-cause death risk can be carried out: Calculate the metabolic age based on the foregoing method in the test set, and obtain the difference between the metabolic age and the age, that is, the acceleration of metabolic aging. Use the Kaplan-Meier curve to visualize the cumulative all-cause death risks of individuals with metabolic aging acceleration in the top 25% (metabolic aging acceleration > 2 years) and the remaining individuals. As Figure 3 shown, it can be seen that the cumulative all-cause death risk of individuals with metabolic aging acceleration in the top 25% is higher during the 5-year follow-up period, with statistical significance (Log-rank P < 0.05). Through analysis, it is shown that the difference between the metabolic age and the actual age obtained by the final prediction model of the present invention reflects the change in the metabolic aging rate and can further reflect the all-cause death risk based on the actual age.
[0070] In summary, through machine learning of age and plasma metabolite biomarkers, the present invention obtains a metabolic aging recognition model. Based on the metabolic age calculated therefrom, it shows excellent sensitivity, specificity, and accuracy in predicting the all-cause death risk, exceeding the actual age itself. Experiments have confirmed that it can intuitively reflect the acceleration of metabolic aging by comparison with the actual age. The present invention only requires a small amount of plasma as a test sample, the evaluation process is simple and convenient, the cost is low, it is easily acceptable to subjects, and it is suitable for large-scale promotion of the method. Moreover, due to its high sensitivity and specificity, it can also be used in the screening of anti-aging drugs.
[0071] According to another aspect of the present invention, there is also provided a computer-readable storage medium having computer-executable instructions stored thereon, wherein when the computer-executable instructions are executed by a processor, the processor is caused to perform the following steps: Step S1: The collection module collects plasma samples of each subject who has experienced an all-cause death event and those who have not experienced an all-cause death event during the 5-year follow-up period, and obtains the plasma small molecule metabolite biomarkers and their relative concentration values of the subjects who have experienced an all-cause death event and those who have not experienced an all-cause death event during the 5-year follow-up period from each plasma sample; Step S2: The feature selection module performs machine learning based on the plasma small molecule metabolite biomarkers and their relative concentration values of the subjects who have not experienced an all-cause death event during the 5-year follow-up period and those who have experienced an all-cause death event during the 5-year follow-up period, so as to screen and obtain the plasma small molecule metabolite biomarkers that can identify the subjects who have not experienced an all-cause death event and those who have experienced an all-cause death event during the 5-year follow-up period as the plasma small molecule metabolite biomarkers after feature selection; Step S3: The training module trains a metabolic aging recognition model to obtain a metabolic age based on the plasma small molecule metabolite biomarkers, their relative concentration values, and the actual age after feature selection of each subject who did not experience an all-cause death event during the 5-year follow-up period and those who did experience an all-cause death event during the 5-year follow-up period. Step S4: The recognition module obtains the age of the person to be tested and the plasma small molecule metabolite biomarkers and their relative concentration values after feature selection, and inputs the age of the person to be tested and the plasma small molecule metabolite biomarkers and their relative concentration values after feature selection into the metabolic aging recognition model to obtain the metabolic age of the person to be tested, which can predict the all-cause death risk of the person to be tested.
[0072] According to another aspect of the present invention, there is also provided an application of metabolite biomarkers in a product for recognizing metabolic aging. The metabolite biomarkers at least include the following metabolite biomarkers: D-Glucose, L-Arginine, Taurine, L-Threonine, Creatinine, L-Acetylcarnitine, L-Serine, Uric acid, Guanidoacetic acid, Pipecolic acid, L-Aspartic acid, Oxoglutaric acid, L-Leucine, and L-Phenylalanine.
[0073] According to another aspect of the present invention, there is also provided a test kit for detecting metabolic age. The metabolite biomarkers in the test kit at least include the following metabolite biomarkers: D-Glucose, L-Arginine, Taurine, L-Threonine, Creatinine, L-Acetylcarnitine, L-Serine, Uric acid, Guanidoacetic acid, Pipecolic acid, L-Aspartic acid, Oxoglutaric acid, L-Leucine, and L-Phenylalanine; The metabolite biomarker is subjected to feature selection and identification through the following modules: A collection module that collects plasma samples of each subject who experienced an all-cause death event and those who did not during a 5-year follow-up period, and obtains the plasma small molecule metabolite biomarkers and their relative concentration values of the subjects who experienced an all-cause death event and those who did not during the 5-year follow-up period from each plasma sample; A feature selection module that performs machine learning based on the plasma small molecule metabolite biomarkers and their relative concentration values of the subjects who did not experience an all-cause death event during the 5-year follow-up period and those who experienced an all-cause death event during the 5-year follow-up period, so as to screen out the plasma small molecule metabolite biomarkers that identify the subjects who did not experience an all-cause death event during the 5-year follow-up period and those who experienced an all-cause death event during the 5-year follow-up period, as the plasma small molecule metabolite biomarkers after feature selection; A training module that trains a metabolic aging recognition model for identifying metabolic age based on the plasma small molecule metabolite biomarkers after feature selection, their relative concentration values, and the actual age of each subject who did not experience an all-cause death event during the 5-year follow-up period and those who experienced an all-cause death event during the 5-year follow-up period; An identification module that obtains the age of the person to be tested and the plasma small molecule metabolite biomarkers after feature selection and their relative concentration values, inputs the age of the person to be tested and the plasma small molecule metabolite biomarkers after feature selection and their relative concentration values into the metabolic aging recognition model to obtain the metabolic age of the person to be tested, and predicts the all-cause death risk of the person to be tested based on the metabolic age.
[0074] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.
[0075] It should be noted that the present invention can be implemented in software and / or a combination of software and hardware. For example, it can be implemented using an application specific integrated circuit (ASIC), a general purpose computer, or any other similar hardware device. In one embodiment, the software program of the present invention can be executed by a processor to implement the above-mentioned steps or functions. Similarly, the software program of the present invention (including related data structures) can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, or a floppy disk and similar devices. Additionally, some steps or functions of the present invention can be implemented using hardware, for example, as a circuit that cooperates with a processor to execute each step or function.
[0076] In addition, a part of the present invention can be applied as a computer program product, such as computer program instructions which, when executed by a computer, can call or provide the methods and / or technical solutions according to the present invention through the operations of the computer. The program instructions for calling the methods of the present invention may be stored in a fixed or removable recording medium, and / or transmitted through a data stream in a broadcast or other signal-bearing medium, and / or stored in the working memory of a computer device operating according to the program instructions. Here, an embodiment according to the present invention includes a device, which includes a memory for storing computer program instructions and a processor for executing the program instructions. When the computer program instructions are executed by the processor, the device is triggered to operate based on the methods and / or technical solutions according to the foregoing multiple embodiments of the present invention.
[0077] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to cover all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention. Any reference signs in the claims should not be construed as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices described in the apparatus claims can also be implemented by one unit or device through software or hardware. First, second, etc. are used to denote names and do not denote any particular order.
Claims
1. A system for constructing an identification model of metabolic aging, characterized in that, Including: A collection module, configured to collect plasma samples of each subject who experienced an all-cause death event and those who did not during a 5-year follow-up period, and obtain plasma small molecule metabolic biomarkers and their relative concentration values of the subjects who experienced an all-cause death event and those who did not during the 5-year follow-up period from each plasma sample; A feature selection module, configured to perform machine learning based on plasma small molecule metabolic biomarkers and their relative concentration values of subjects who did not experience an all-cause death event during the 5-year follow-up period and those who experienced an all-cause death event during the 5-year follow-up period, so as to screen out plasma small molecule metabolic biomarkers for identifying subjects who did not experience an all-cause death event during the 5-year follow-up period and those who experienced an all-cause death event during the 5-year follow-up period, as the plasma small molecule metabolic biomarkers after feature selection; A training module, based on the plasma small molecule metabolic biomarkers after feature selection, their relative concentration values, and the actual age of each subject who did not experience an all-cause death event during the 5-year follow-up period and those who experienced an all-cause death event during the 5-year follow-up period, trains a metabolic aging recognition model for identifying metabolic age; An identification module, obtains the age of the person to be tested and the plasma small molecule metabolic biomarkers after feature selection and their relative concentration values, and inputs the age of the person to be tested and the plasma small molecule metabolic biomarkers after feature selection and their relative concentration values into the metabolic aging recognition model to obtain the metabolic age of the person to be tested.
2. The system for constructing an identification model of metabolic aging according to claim 1, wherein A collection module, configured to obtain a first group of subjects and a second group of subjects that do not overlap with each other, wherein both the first group of subjects and the second group of subjects include each subject who experienced and did not experience an all-cause death event during a non-overlapping 5-year follow-up period; A collection module, configured to collect first plasma samples of the first group of subjects, and obtain plasma small molecule metabolic biomarkers and their relative concentration values from each first plasma sample as a first set; A collection module, configured to collect second plasma samples of the second group of subjects, and obtain plasma small molecule metabolic biomarkers and their relative concentration values from each second plasma sample as a test set.
3. The system for constructing an identification model of metabolic aging according to claim 2, wherein The feature selection module is configured to divide the first set into 5 non-overlapping data portions, each time using 1 data portion that has not been selected from the 5 non-overlapping data portions as an internal validation set, and each time using the remaining 4 data portions as a training set; wherein both the internal validation set and the training set include each subject who experienced an all-cause death event and those who did not during a non-overlapping 5-year follow-up period; A feature selection module, which is used to use the Lasso feature selection algorithm based on each training set, and train the feature selection model in a loop for 5 rounds respectively to obtain the corresponding 5-round feature selection models. For each λn value in the hyperparameter sequence in each round of the feature selection model, a corresponding feature selection model is generated Mk ( λn ), n = 1 - 100; k = 1 - 5, where k represents the number of rounds; n represents the serial number of the λ value λ is one of the hyperparameters A feature selection module, which is used for each value in the hyperparameter sequence λn , inputting the internal validation set corresponding to each round into the feature selection model of the corresponding round in 5 feature selection models Mk ( λn ), and respectively obtaining a misjudgment rate corresponding to each λn value on the internal validation set of the corresponding round; obtaining an average misjudgment rate corresponding to each value from the 5 misjudgment rates corresponding to each value; based on the minimum average misjudgment rate among all values, selecting the λ .min when the corresponding misclassification rate is the lowest, where, based on the λ .min when the misjudgment rate is the lowest, obtaining the corresponding metabolite biomarker as the plasma small molecule metabolite biomarker after feature selection.
4. The system for constructing an identification model of metabolic aging according to claim 1, wherein The plasma small molecule metabolic biomarkers after feature selection at least include the following metabolite biomarkers in peripheral plasma: D-Glucose, L-Arginine, Taurine, L-Threonine, Creatinine, L-Acetylcarnitine, L-Serine, Uric acid, Guanidoacetic acid, Pipecolic acid, L-Aspartic acid, Oxoglutaric acid, L-Leucine, and L-Phenylalanine.
5. The system for constructing an identification model of metabolic aging according to claim 2, wherein A training module for training a metabolic aging recognition model based on the plasma small molecule metabolite biomarkers selected from each plasma sample in the first set, their relative concentration values, and the actual age.
6. The system for constructing an identification model of metabolic aging according to claim 5, wherein, A training module for training a metabolic aging recognition model for identifying metabolic age. In the metabolic aging recognition model for identifying metabolic age, the elastic net algorithm is used and based on the plasma small molecule metabolite biomarkers selected from the characteristics of each subject who does not have an all-cause death event during the 5-year follow-up period and who has an all-cause death event during the 5-year follow-up period in the first set, their relative concentration values, and the actual age, an all-cause death risk score P based on metabolism is automatically output for each subject in the first set. The range of P is from 0 to 1, representing the low all-cause death risk to the high all-cause death risk; the implementation of the elastic net algorithm depends on the glmnet package, where the ElasticNet algorithm hyperparameter alpha for controlling the ratio of L1 regularization and L2 regularization is set to 0.5; s for controlling the penalty strength on the parameters is set to 0.01; the metabolic aging recognition model for identifying metabolic age calculates the mean and standard deviation of the ages of the subjects in the first set, and calculates the mean and standard deviation of the all-cause death risk scores P based on metabolism of all the subjects in the first set; scales the mean and standard deviation of the all-cause death risk scores P based on metabolism to be consistent with the mean and standard deviation of the ages of the subjects in the first set, thereby obtaining the metabolic age in years. A training module for performing performance evaluation on the metabolic age output by the trained metabolic aging model, and using the metabolic aging model whose performance evaluation meets the preset requirements as the final metabolic aging model for identifying metabolic age.
7. The system for constructing an identification model of metabolic aging according to claim 6, characterized in that, The calculation formula for the metabolic age is as follows: ; Among them, the metabolic age i is the metabolic age after conversion of the i th sample in the first set; P i is the metabolic-based all-cause death risk score value of the i th subject; μ p is the mean of the metabolic-based all-cause death risk scores P of all the subjects in the first set; σ p is the sample standard deviation of the metabolic-based all-cause death risk scores P of all the subjects in the first set; μ Age is the mean of the ages of the subjects in the first set; σ Age is the standard deviation of the ages of the subjects in the first set.
8. The system for constructing an identification model of metabolic aging according to claim 2, wherein The training module uses the trained metabolic aging model to obtain the trained metabolic aging recognition model, and the recognition performance on the first set and the recognition performance of the metabolic age on the test set; if the recognition performance of the trained metabolic aging recognition model on the metabolic age of the first set and the test set differs by less than a preset difference threshold and is greater than the actual age of the subject itself, then the trained metabolic aging recognition model is used as the final metabolic aging recognition model.
9. A computer-readable storage medium having computer-executable instructions stored thereon, wherein, When the computer-executable instructions are executed by the processor, the processor is caused to perform the following steps: The collection module collects plasma samples of each subject who experienced an all-cause death event and those who did not during the 5-year follow-up period, and obtains the plasma small molecule metabolic biomarkers and their relative concentration values of the subjects who experienced an all-cause death event and those who did not during the 5-year follow-up period from each plasma sample; Based on the plasma small molecule metabolic biomarkers and their relative concentration values of the subjects who did not experience an all-cause death event during the 5-year follow-up period and those who experienced an all-cause death event during the 5-year follow-up period, the feature selection module performs machine learning to screen out the plasma small molecule metabolic biomarkers that identify the subjects who did not experience an all-cause death event during the 5-year follow-up period and those who experienced an all-cause death event during the 5-year follow-up period, as the plasma small molecule metabolic biomarkers after feature selection; Based on the plasma small molecule metabolic biomarkers after feature selection, their relative concentration values, and the actual age of each subject who did not experience an all-cause death event during the 5-year follow-up period and those who experienced an all-cause death event during the 5-year follow-up period, the training module trains a metabolic aging recognition model for identifying metabolic age; The recognition module obtains the age of the person to be tested and the plasma small molecule metabolic biomarkers after feature selection and their relative concentration values, and inputs the age of the person to be tested and the plasma small molecule metabolic biomarkers after feature selection and their relative concentration values into the metabolic aging recognition model to obtain the metabolic age of the person to be tested.
Citation Information
Patent Citations
Metabolic age prediction model and application thereof in colorectal cancer diagnosis
CN114334170A
Cardiovascular metabolism risk factor spectrum recognition model construction system, storage medium and kit
CN119181506A
Diagnostic marker for metabolic syndrome and preclinical stage thereof and application of diagnostic marker
CN119246661A
Age estimation method based on metabolic data
CN119324055A
Biomarkers related to metabolic age and methods using the same
US20080124752A1
Cited By
Construction system of prediction model for occurrence of clostridium difficile infection and recurrent infection, storage medium and kit
CN120432173A
Prediction model construction system, storage medium and kit for Clostridium difficile infection and recurrent infection
CN120432173B