Metabolic aging identification model construction system, storage medium and kit
By constructing a metabolic aging identification model and utilizing plasma small molecule metabolic biomarkers and machine learning, we have solved the problem of accuracy in assessing individual metabolic aging status, achieved efficient identification of metabolic aging and accurate prediction of all-cause mortality risk.
Patent Information
- Application Number
- CN202510764292.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Existing technologies make it difficult to accurately assess an individual's metabolic aging status, resulting in the inability to effectively identify differences in metabolic aging risks among people of the same age, and a lack of reliable identification models based on metabolite biomarkers.
By collecting and analyzing plasma samples from individuals who did and did not experience all-cause mortality during the 5-year follow-up, 14 plasma small molecule metabolic biomarkers were screened. A metabolic aging identification model was constructed using machine learning methods. The model was trained in combination with actual age to predict metabolic age and then evaluate the metabolic aging status.
It achieves accurate assessment of metabolic aging, improves the predictive sensitivity and specificity of all-cause mortality risk, is low-cost and easy to promote, and is suitable for anti-aging drug screening.
Smart Images

Figure CN120280160B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical prediction, and in particular to a metabolic aging recognition model construction system, storage medium and kit. Background Art
[0002] Metabolic aging, a key component of the aging process, refers to the biological process in which metabolic networks within an organism gradually become dysregulated during the degenerative aging process, leading to a decline in energy balance, oxidative stress, and cellular repair. Its core characteristic is the progressive loss of metabolic homeostasis, manifested by decreased mitochondrial function, abnormal nutrient-sensing pathways, the accumulation of harmful metabolites, and altered epigenetic modifications. This process not only accelerates organ degeneration (such as muscle atrophy and decreased liver metabolic capacity) but is also closely associated with a variety of age-related diseases, such as diabetes, cardiovascular disease, and neurodegenerative diseases. Recent studies have revealed that intervening in metabolic pathways can significantly slow the aging process, making metabolic aging a key target in the anti-aging field.
[0003] Traditionally, metabolic aging has been assessed from the perspective of the beginning of life (birth) and relies on age. With aging, metabolic homeostasis is disrupted, and the metabolic aging process deepens. However, individuals of the same age often exhibit different metabolic aging states, resulting in different individual risk profiles. For example, one elderly individual may already have multiple metabolic abnormalities such as diabetes and obesity, while another elderly individual of the same age may be well-maintained, metabolically healthy, and have a longer life expectancy than their peers, indicating less severe metabolic aging. Furthermore, age is calculated based on date of birth and increases steadily, making it difficult to assess whether an individual's metabolic rate is increasing or decreasing. Assessing metabolic aging from the perspective of the end of life (death) relies on the length of remaining lifespan. However, this length of remaining lifespan requires long-term follow-up of individuals to record all-cause mortality. Currently, reliable models for identifying metabolic aging based on combinations of metabolite biomarkers are still lacking. Summary of the Invention
[0004] The purpose of the present invention is to provide a metabolic aging identification model construction system, storage medium and kit and their applications.
[0005] To solve the above problems, the present invention provides a metabolic aging recognition model construction system, comprising:
[0006] An acquisition module is used to collect plasma samples from subjects who have or have not experienced all-cause mortality during the 5-year follow-up period, and to obtain plasma small molecule metabolic biomarkers and their relative concentration values from each plasma sample for subjects who have or have not experienced all-cause mortality during the 5-year follow-up period;
[0007] a feature selection module for performing machine learning based on the plasma small molecule metabolic biomarkers and relative concentration values of subjects who did not experience all-cause mortality during the 5-year follow-up period and subjects who experienced all-cause mortality during the 5-year follow-up period, to screen and identify the plasma small molecule metabolic biomarkers of subjects who did not experience all-cause mortality during the 5-year follow-up period and subjects who experienced all-cause mortality during the 5-year follow-up period, as the plasma small molecule metabolic biomarkers after feature selection;
[0008] The training module trains a metabolic aging recognition model for identifying metabolic age based on the plasma small molecule metabolic biomarkers, relative concentrations, and actual age of subjects who did not experience all-cause mortality during the 5-year follow-up period and those who did experience all-cause mortality during the 5-year follow-up period.
[0009] The identification module obtains the age of the subject to be tested and the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection, and inputs the age of the subject to be tested and the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection into the metabolic aging identification model to obtain the metabolic age of the subject to be tested. Based on the metabolic age, the all-cause mortality risk of the subject to be tested can be predicted.
[0010] Furthermore, in the above system, the collection module is used to collect plasma samples from subjects who did not suffer from all-cause mortality during the 5-year follow-up period and subjects who suffered from all-cause mortality during the 5-year follow-up period, and obtain plasma small molecule metabolic biomarkers and their relative concentration values of the subjects who did not suffer from all-cause mortality during the 5-year follow-up period and subjects who suffered from all-cause mortality during the 5-year follow-up period from each plasma sample, including:
[0011] an acquisition module, configured to obtain a first subject group and a second subject group that do not overlap with each other, wherein the first subject group and the second subject group both include subjects who do and do not experience all-cause mortality during a non-overlapping 5-year follow-up period;
[0012] a collection module, configured to collect first plasma samples from a first group of subjects, and obtain plasma small molecule metabolic biomarkers and their relative concentration values from each first plasma sample as a first set;
[0013] The acquisition module is used to collect second plasma samples from the second subject group, and obtain plasma small molecule metabolic biomarkers and their relative concentration values from each second plasma sample as a test set.
[0014] Furthermore, in the above system, the feature selection module performs machine learning based on the plasma small molecule metabolic biomarkers and their relative concentration values of the subjects who did not suffer from all-cause mortality during the 5-year follow-up period and the subjects who suffered from all-cause mortality during the 5-year follow-up period, so as to screen and obtain plasma small molecule metabolic biomarkers for identifying the subjects who did not suffer from all-cause mortality during the 5-year follow-up period and the subjects who suffered from all-cause mortality during the 5-year follow-up period, as the plasma small molecule metabolic biomarkers after feature selection, including:
[0015] The feature selection module is configured to divide the first set into five non-overlapping data sets, each time using one of the five non-overlapping data sets that has not been selected as an internal validation set, and each time using the remaining four data sets as training sets; wherein both the internal validation set and the training set include non-overlapping subjects who experienced all-cause mortality during the five-year follow-up period and subjects who did not experience all-cause mortality during the five-year follow-up period;
[0016] The feature selection module is used to use the Lasso feature selection algorithm on the R (4.3.1) software based on each training set, and to cycle through the feature selection model for 5 rounds to obtain the corresponding 5-round feature selection model. In each round of the feature selection model, each of the hyperparameter sequences is selected. λn Value, generate the corresponding feature selection model Mk ( λn ), n=1-100; k=1-5, k represents the number of rounds; n represents the sequence number of λ value; a total of 100;
[0017] Here, it is assumed that the value sequence of λ is { l 1, l 2,…, l 100}, in 5-fold cross validation, each round of cross validation, taking the kth round as an example, k=1-5;
[0018] Feature selection module for each value in the hyperparameter sequence λn , input the internal validation set corresponding to each round into the feature selection model of the corresponding round in the 5 feature selection models Mk ( λn ), and obtain each λn The error rate corresponding to the value on the internal validation set of the corresponding round; the average error rate corresponding to each value is obtained from the 5 error rates corresponding to each value; based on the minimum average error rate among all values, the value with the lowest error rate is selected l .min, where l ( lambda) is one of the hyperparameters; based on the lowest misjudgment rate l.min, and the corresponding metabolite markers were obtained as plasma small molecule metabolic biomarkers after feature selection.
[0019] Here, based on the fact that lambda.min is the best hyperparameter lambda, the corresponding features are extracted to screen out 14 metabolite markers.
[0020] Furthermore, in the above system, the plasma small molecule metabolite biomarkers after feature selection include at least the following metabolite biomarkers in peripheral plasma:
[0021] D-Glucose, L-Arginine, Taurine, L-Threonine, Creatinine, L-Acetylcarnitine, L-Serine, Uric acid, Guanidoacetic acid, Pipecolic acid, L-Aspartic acid, Oxoglutaric acid, L-Leucine, and L-Phenylalanine.
[0022] Furthermore, in the above system, the training module trains a metabolic aging recognition model for identifying metabolic age based on the plasma small molecule metabolic biomarkers, relative concentration values, and actual age of each subject in the first set who did not experience all-cause mortality and who experienced all-cause mortality during the 5-year follow-up period, including:
[0023] The training module is used to train a metabolic aging recognition model based on the plasma small molecule metabolic biomarkers after feature selection, their relative concentration values and actual age in each plasma sample in the first set.
[0024] Furthermore, in the above system, the training module trains a metabolic aging recognition model based on the plasma small molecule metabolic biomarkers, their relative concentration values, and actual age after feature selection in each plasma sample, including:
[0025] A training module for training a metabolic aging recognition model for identifying metabolic age, wherein the metabolic aging recognition model for identifying metabolic age uses the ElasticNet algorithm on R (4.3.1) software and automatically outputs a metabolic-based all-cause mortality risk score P for each subject in the first set based on the plasma small molecule metabolic biomarkers, their relative concentration values, and actual age after feature selection of each subject in the first set who did not experience all-cause mortality during the 5-year follow-up period and who experienced all-cause mortality during the 5-year follow-up period. P ranges from 0 to 1, representing a range from low metabolic mortality risk to high metabolic mortality risk; the implementation of the ElasticNet algorithm relies on the glmnet package, wherein the ElasticNet algorithm hyperparameter alpha for controlling the ratio of L1 regularization to L2 regularization is set to 0.5; and s for controlling the penalty intensity of the parameter is set to 0.01;
[0026] The metabolic aging identification model for identifying metabolic age calculates the mean and standard deviation of the ages of the subjects in the first set, and calculates the mean and standard deviation of the metabolic-based all-cause mortality risk score P for all subjects in the first set; the mean and standard deviation of the metabolic-based all-cause mortality risk score P are scaled to be consistent with the mean and standard deviation of the ages of the subjects in the first set, thereby obtaining the metabolic age in years for identifying metabolic aging. The specific formula is as follows:
[0027]
[0028] Among them: metabolic age i It is the first set i Metabolic age after conversion of samples; P i It is i Metabolic-based all-cause mortality risk score for each subject; m p is the mean of the metabolic-based all-cause mortality risk score P for all subjects in the first set; s p is the sample standard deviation of the metabolic-based all-cause mortality risk score P for all subjects in the first set; m Age is the mean age of the subjects in the first set; s Age is the standard deviation of the age of the subjects in the first set;
[0029] The training module is used to perform performance evaluation on the metabolic age output by the trained metabolic aging model. The metabolic aging model whose performance evaluation meets the preset requirements is used as the final metabolic aging model to identify the metabolic age and accurately predict the risk of all-cause mortality based on the metabolic age.
[0030] Furthermore, in the above system, the training module uses the trained metabolic aging model to obtain the recognition performance of the trained metabolic aging model on the first set and the recognition performance of metabolic age on the test set; if the difference in the recognition performance of the trained metabolic aging model in the first set and the test set of metabolic age is less than a preset difference threshold, and both are greater than the actual age of the subject itself, the trained metabolic aging model will be used as the final metabolic aging model.
[0031] According to another aspect of the present invention, a computer-readable storage medium is provided, on which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor is caused to perform the following steps:
[0032] Step S1: The collection module collects plasma samples from subjects who have experienced all-cause mortality and those who have not experienced all-cause mortality during the 5-year follow-up period, and obtains plasma small molecule metabolic biomarkers and their relative concentration values of the subjects who have experienced all-cause mortality and those who have not experienced all-cause mortality during the 5-year follow-up period from each plasma sample;
[0033] Step S2: The feature selection module performs machine learning based on the plasma small molecule metabolic biomarkers and relative concentration values of the subjects who did not experience all-cause mortality during the 5-year follow-up period and the subjects who experienced all-cause mortality during the 5-year follow-up period, to screen and obtain the plasma small molecule metabolic biomarkers that identify the subjects who did not experience all-cause mortality during the 5-year follow-up period and the subjects who experienced all-cause mortality during the 5-year follow-up period, as the plasma small molecule metabolic biomarkers after feature selection;
[0034] Step S3: The training module trains a metabolic aging recognition model to obtain metabolic age based on the plasma small molecule metabolic biomarkers, relative concentration values, and actual age of each subject who did not experience all-cause mortality during the 5-year follow-up period and who experienced all-cause mortality during the 5-year follow-up period;
[0035] Step S4: The recognition module obtains the age of the subject to be tested and the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection, and inputs the age of the subject to be tested and the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection into the metabolic aging recognition model to obtain the metabolic age of the subject to be tested, which can predict the all-cause mortality risk of the subject to be tested.
[0036] According to another aspect of the present invention, there is also provided a use of a metabolite biomarker in a product for identifying metabolic aging, wherein the metabolite biomarker comprises at least the following metabolite biomarkers:
[0037] D-Glucose, L-Arginine, Taurine, L-Threonine, Creatinine, L-Acetylcarnitine, L-Serine, Uric acid, Guanidoacetic acid, Pipecolic acid, L-Aspartic acid, Oxoglutaric acid, L-Leucine, and L-Phenylalanine.
[0038] According to another aspect of the present invention, a detection kit for metabolic age is provided, wherein the metabolite biomarkers in the detection kit include at least the following metabolite biomarkers:
[0039] D-Glucose, L-Arginine, Taurine, L-Threonine, Creatinine, L-Acetylcarnitine, L-Serine, Uric acid, Guanidoacetic acid, Pipecolic acid, L-Aspartic acid, Oxoglutaric acid, L-Leucine, and L-Phenylalanine;
[0040] The metabolite biomarkers were feature selected and identified using the following modules:
[0041] The acquisition module collects plasma samples from subjects who have or have not experienced all-cause mortality during the 5-year follow-up period, and obtains plasma small molecule metabolic biomarkers and their relative concentration values from each plasma sample for subjects who have or have not experienced all-cause mortality during the 5-year follow-up period;
[0042] The feature selection module performs machine learning based on the plasma small molecule metabolic biomarkers and their relative concentration values of subjects who did not experience all-cause mortality during the 5-year follow-up and subjects who experienced all-cause mortality during the 5-year follow-up, to screen and identify the plasma small molecule metabolic biomarkers that identify subjects who did not experience all-cause mortality during the 5-year follow-up and subjects who experienced all-cause mortality during the 5-year follow-up, and use them as the plasma small molecule metabolic biomarkers after feature selection;
[0043] The training module trains a metabolic aging recognition model for identifying metabolic age based on the plasma small molecule metabolic biomarkers, relative concentrations, and actual age of subjects who did not experience all-cause mortality during the 5-year follow-up period and those who did experience all-cause mortality during the 5-year follow-up period.
[0044] The identification module obtains the age of the subject to be tested and the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection, and inputs the age of the subject to be tested and the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection into the metabolic aging identification model to obtain the metabolic age of the subject to be tested, and predicts the all-cause mortality risk of the subject to be tested based on the metabolic age.
[0045] Compared to existing technologies, the present invention utilizes machine learning of age and plasma metabolite biomarkers to develop a metabolic aging identification model. The metabolic age calculated from this model demonstrates superior sensitivity, specificity, and accuracy in predicting all-cause mortality risk, exceeding chronological age alone. Experiments have confirmed its ability to intuitively reflect accelerated metabolic aging through comparison with chronological age. The present invention requires only a small amount of plasma as a test sample, resulting in a simple and convenient evaluation process, low cost, and easy acceptance by subjects, making it suitable for large-scale dissemination. Furthermore, its high sensitivity and specificity make it suitable for screening anti-aging drugs. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is the performance of identifying the 5-year all-cause mortality risk on the training set using the metabolic age and actual age obtained based on the metabolic aging recognition model in one embodiment of the present invention;
[0047] Figure 2 This is the performance of identifying the 5-year all-cause mortality risk on the test set using the metabolic age and actual age obtained based on the metabolic aging recognition model in one embodiment of the present invention;
[0048] Figure 3 5-year cumulative risk curves for all-cause mortality of the TOP25% of metabolic age acceleration (metabolic age - chronological age > 2 years) and the remaining individuals based on the test set in one embodiment of the present invention. DETAILED DESCRIPTION
[0049] The present invention is further described in detail below with reference to the accompanying drawings.
[0050] In a typical configuration of the present application, the terminal, the device of the service network and the trusted party all include one or more processors (CPUs), input / output interfaces, network interfaces and memories.
[0051] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0052] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.
[0053] To further clarify the objectives, technical solutions, and advantages of the embodiments of the present invention, various embodiments of the present invention are described in detail below with reference to the examples. Experimental methods in the examples where specific conditions are not specified are generally performed under conventional conditions, such as those described in textbooks and experimental manuals, or under conditions recommended by the manufacturer. The methods are also performed under the recommended conditions of the supporting software.
[0054] The present invention provides a metabolic aging identification model construction system, a storage medium, a kit and applications thereof. The method comprises: steps S1 to S4.
[0055] Step S1, an acquisition module, collecting plasma samples from subjects who experienced all-cause mortality and those who did not experience all-cause mortality during the 5-year follow-up period, and obtaining plasma small molecule metabolic biomarkers and their relative concentration values from each plasma sample for subjects who experienced all-cause mortality and those who did not experience all-cause mortality during the 5-year follow-up period;
[0056] Metabolomics, a key branch of systems biology, focuses on the qualitative and quantitative analysis of all metabolites within an organism. Using high-throughput analytical techniques, it comprehensively detects small molecule metabolites, including carbohydrates, lipids, and organic acids, in biological samples. These metabolites reflect the metabolic state of an organism at a specific point in time and can be used for disease diagnosis, drug development, and biomarker discovery.
[0057] All-cause mortality is a core indicator in medical and epidemiological research, referring to all deaths regardless of the specific cause. A higher risk of all-cause mortality during the five-year follow-up is associated with a shorter lifespan, more significant metabolic disorders, an older metabolic age, and a more advanced metabolic aging process.
[0058] People who experienced all-cause mortality during the 5-year follow-up had a shorter life expectancy and a higher risk of aging, showing a metabolic aging model different from that of people who did not experience all-cause mortality during the 5-year follow-up. Therefore, plasma metabolomics can be a useful resource for the assessment of metabolic aging and will strongly promote healthy aging.
[0059] The relative concentration value refers to the relative concentration value of a certain type of plasma small molecule metabolic biomarker in plasma, and its unit can be %;
[0060] Preferably, step S1 includes:
[0061] Step S11: The acquisition module obtains a first subject group and a second subject group that do not overlap with each other, wherein the first subject group and the second subject group both include subjects who did or did not experience all-cause mortality during the 5-year follow-up period that do not overlap with each other;
[0062] Here, the first subject group includes all subjects who have and have not experienced all-cause mortality during the 5-year follow-up period; the second subject group includes all subjects who have and have not experienced all-cause mortality during the 5-year follow-up period; the first subject group and the second subject group do not overlap with each other;
[0063] Step S12: The collection module collects first plasma samples from the first subject group, and obtains plasma small molecule metabolic biomarkers and their relative concentration values from each first plasma sample as a first set;
[0064] Step S13: The collection module collects second plasma samples from the second subject group, and obtains plasma small molecule metabolic biomarkers and their relative concentration values from each second plasma sample as a test set;
[0065] Specifically, the group of subjects can recruit subjects enrolled in the two centers from January 2019 to January 2020.
[0066] The first set of subjects used for training and validation of the prediction model can be 9,365 subjects from center 1 who did not suffer from all-cause death during the 5-year follow-up period and 351 subjects who suffered from all-cause death during the 5-year follow-up period; the training set can subsequently be split into non-overlapping internal training sets and internal validation sets.
[0067] The test set for the prediction model can be obtained from 2227 subjects who did not suffer from all-cause mortality during the 5-year follow-up and 93 subjects who suffered from all-cause mortality during the 5-year follow-up in center 2.
[0068] All-cause mortality events were recorded over a five-year period based on regular home visits by the investigators' families and registries at local medical institutions. Plasma was collected at the beginning of the study (Year 0) for metabolite biomarkers using a nanoparticle-enhanced laser desorption / ionization mass spectrometer for both subjects who experienced all-cause mortality and those who did not.
[0069] Preferably, the exclusion criteria within the inclusion criteria are subjects with acute and infectious clinical symptoms within three weeks before sampling, including but not limited to fever, headache, cough, sore throat, loss of smell, runny nose, abdominal pain and diarrhea.
[0070] Step S12 or step S13 may include: steps S121 to S128.
[0071] Step S121, treating the plasma sample of peripheral venous blood obtained from the subject in the early morning in a fasting state with EDTA anticoagulation and storing it at -80°C;
[0072] Step S122, preparation of instruments and reagents: preparing a nanoparticle enhanced laser desorption / ionization time-of-flight mass spectrometer (MALDI-TOF-MS) and related reagents;
[0073] Here, the rice particle enhanced laser desorption / ionization time-of-flight mass spectrometer can be a product of Bruker, Germany;
[0074] Step S123, diluting the plasma sample: diluting the plasma sample with deionized water at a standard ratio to obtain a diluted plasma sample, ensuring that the metabolite concentration is suitable for mass spectrometry analysis;
[0075] Specifically, 100 nL of plasma sample can be diluted 10 times with deionized water to obtain a diluted plasma sample;
[0076] Step S124, preparing the inorganic nanoparticles into a nanoparticle matrix solution for enhancing the mass spectrometry signal;
[0077] Specifically, the matrix (inorganic nanoparticles) was prepared into a 1 mg / mL matrix solution with deionized water;
[0078] Step S125, performing sample preparation on the mass spectrometry target plate: spotting the diluted plasma sample onto the mass spectrometry target plate and drying at room temperature;
[0079] Specifically, sample preparation can be performed on a mass spectrometry target plate, with 500 nL of each diluted plasma sample spotted and dried at room temperature;
[0080] Step S126, performing matrix preparation on the mass spectrometry target plate: spotting the nanoparticle matrix solution on the mass spectrometry target plate, ensuring that the nanoparticle matrix solution evenly covers the plasma sample spot to obtain a sample;
[0081] Specifically, the matrix can be prepared on a mass spectrometry target plate by spotting 500 nL of each matrix solution on each plasma sample spot and drying at room temperature;
[0082] Step S127 , performing data acquisition on the mixture in a nanoparticle enhanced laser desorption / ionization time-of-flight mass spectrometer: using MALDI-TOF-MS to acquire mass spectrometric data of metabolites in the plasma, and obtaining a metabolite peak map in the sample through delayed extraction and time-of-flight analysis.
[0083] Specifically, data were collected in a nanoparticle-enhanced laser desorption / ionization time-of-flight mass spectrometer. The collected data were extracted in positive ion mode with delayed extraction, a repetition rate of 1000 Hz, an acceleration voltage of 20 kV, a delay time of 250 ns, and 2000 laser shots per analysis.
[0084] In step S128, the built-in data preprocessing pipeline can be run based on the metabolite peak map in the sample under the recommended conditions of the nanoparticle-enhanced laser desorption / ionization time-of-flight mass spectrometer supporting software, including: data resampling, spectral line smoothing, baseline correction and spectral peak matching, to obtain the relative concentration values of the full spectrum of 303 metabolic biomarkers.
[0085] Specifically, a Savitzky-Golay (SG) filter can be used for spectral smoothing; then an adaptive iterative algorithm is used to identify baseline points, and polynomial fitting is performed on the baseline points identified in the metabolite peak graph. The fitted baseline is subtracted from the original spectrum, and the corrected spectrum is checked to ensure that the baseline is flat and has no negative values; then the peak detection parameters are set: signal-to-noise ratio threshold = 3, mass accuracy tolerance 50ppm, continuous wavelet transform (CWT) is used to identify possible metabolite marker peak positions, and further peak fitting is performed to obtain the corresponding peak position, peak height and peak area, and the metabolite marker levels are relatively quantified according to the peak height; finally, Min-Max normalization is performed to obtain the relative concentration values of 303 markers.
[0086] The relative concentration value obtained here can be the peak intensity value after min-max normalization, in %, which is used to measure the relative concentration of metabolites in plasma.
[0087] By comparing the HMDB (Human Metabolome Database) database, 303 plasma small molecule metabolic biomarkers were identified.
[0088] Step S2: The feature selection module performs machine learning based on the plasma small molecule metabolic biomarkers and relative concentration values of the subjects who did not experience all-cause mortality during the 5-year follow-up period and the subjects who experienced all-cause mortality during the 5-year follow-up period, to screen and identify the plasma small molecule metabolic biomarkers that identify the subjects who did not experience all-cause mortality during the 5-year follow-up period and the subjects who experienced all-cause mortality during the 5-year follow-up period, and use them as the plasma small molecule metabolic biomarkers after feature selection;
[0089] Preferably, step S2 includes:
[0090] Step S20: Divide the first data set into five non-overlapping data sets, each time using one of the five non-overlapping data sets that has not been selected as an internal validation set, and each time using the remaining four data sets as training sets; wherein both the internal validation set and the training set include non-overlapping subjects who experienced all-cause mortality during the five-year follow-up and subjects who did not experience all-cause mortality during the five-year follow-up;
[0091] Step S21: The feature selection module uses the Lasso feature selection algorithm on the R (4.3.1) software based on each training set to train the feature selection model for 5 rounds to obtain the corresponding 5 feature selection models. In each round of feature selection model, for each λn value in the hyperparameter sequence, a corresponding feature selection model is generated. Mk ( λn ), n=1-100; k=1-5, k represents the number of rounds; n represents the sequence number of λ value; a total of 100;
[0092] Here, it is assumed that the value sequence of λ is { l 1, l 2,…, l 100}, in 5-fold cross validation, each round of cross validation, taking the kth round as an example, k=1-5;
[0093] Feature selection module, for each value in the sequence of 100 values corresponding to the hyperparameter sequence lambda λn , input the internal validation set corresponding to each round into one of the five feature selection models corresponding to the round Mk ( λn ), and obtain each λn The error rate corresponding to the value on each internal validation set; from the 5 error rates corresponding to each value, an average error rate corresponding to each value is obtained; based on the minimum average error rate among all values, lambda.min with the lowest error rate is selected, where lambda is one of the hyperparameters; based on lambda.min with the lowest error rate ( l .min), and obtained the corresponding 14 metabolite markers.
[0094] Here, based on lambda.min being the best hyperparameter lambda, the corresponding features were extracted to screen out 14 metabolite markers;
[0095] Specifically, the lambda hyperparameter is the same for all five feature selection models. For each lambda value, the features used are determined, and these features can then be mapped to corresponding metabolite markers. Five feature selection models are generated from the training set. The average performance of the lambda values of these five models on the internal validation set is then analyzed to determine which lambda value achieves the best average performance on the internal validation set, i.e., the lowest average false positive rate.
[0096] R software is a language and environment for statistical computing and graphics. It supports cross-platform platforms such as Windows, Mac, and Linux. A GNU project, it is similar to the S language and environment developed by John Chambers and colleagues at Bell Labs (formerly AT&T, now Lucent Technologies). R can be considered a different implementation of S. While there are some important differences, many codes written for S run unchanged under R. R provides a wide variety of statistical (linear and nonlinear modeling, classical statistical tests, time series analysis, classification, clustering, etc.) and graphical techniques and is highly extensible. S is often the tool of choice for research in statistical methods, and R provides an open source path to participate in this activity.
[0097] Preferably, the plasma small molecule metabolic biomarkers after feature selection include at least the following metabolite biomarkers in peripheral plasma:
[0098] D-Glucose, L-Arginine, Taurine, L-Threonine, Creatinine, L-Acetylcarnitine, L-Serine, Uric acid, Guanidoacetic acid, Pipecolic acid, L-Aspartic acid, Oxoglutaric acid, L-Leucine, and L-Phenylalanine.
[0099] The specific information is shown in Table 1.
[0100] Table 1: Information about 14 metabolic biomarkers
[0101]
[0102] Here, the present invention uses 14 plasma metabolic biomarkers and age obtained by the feature selection module for subsequent training of the metabolic aging recognition model, which can demonstrate excellent prediction capabilities for all-cause mortality risk; by performing machine learning on the plasma small molecule metabolic biomarkers after feature selection, efficient assessment of metabolic aging can be achieved.
[0103] Step S3, a training module, training a metabolic aging recognition model for identifying metabolic age based on the selected plasma small molecule metabolic biomarkers, relative concentrations, and actual age of each subject in the first set who did not experience all-cause mortality during the 5-year follow-up period and who experienced all-cause mortality during the 5-year follow-up period;
[0104] Preferably, step S3 includes:
[0105] Step S311, a training module, is used to train a metabolic aging recognition model for identifying metabolic age, wherein the metabolic aging recognition model for identifying metabolic age uses the ElasticNet algorithm on R (4.3.1) software and is based on the plasma small molecule metabolic biomarkers, their relative concentration values, and actual age of each subject in the first set who did not experience all-cause mortality during the 5-year follow-up period and who experienced all-cause mortality during the 5-year follow-up period. Automatically output a metabolic-based all-cause mortality risk score P for each subject, where P ranges from 0 to 1, representing a range from low metabolic mortality risk to high metabolic mortality risk. The implementation of the ElasticNet algorithm relies on the glmnet package, wherein the ElasticNet algorithm hyperparameter alpha, which is used to control the ratio of L1 regularization and L2 regularization, is set to 0.5; s, which is used to control the penalty intensity of the parameter, is set to 0.01;
[0106] Here, the distinction between those who did and did not experience all-cause mortality during the 5-year follow-up was made in order to output a metabolic-based all-cause mortality risk score P.
[0107] In step S312, the metabolic aging recognition model for identifying metabolic age calculates the mean and standard deviation of the ages of the subjects in the first set, and calculates the mean and standard deviation of the metabolic-based all-cause mortality risk score P for all subjects in the first set; the mean and standard deviation of the metabolic-based all-cause mortality risk score P are scaled to be consistent with the mean and standard deviation of the ages of the subjects in the first set, thereby obtaining the metabolic age, in years, for identifying metabolic aging. The specific formula is as follows:
[0108] (1)
[0109] Among them: metabolic age i It is the first set i Metabolic age after conversion of samples; P i It is i Metabolic-based all-cause mortality risk score for each subject; m p is the mean of the metabolic-based all-cause mortality risk score P for all subjects in the first set; s p is the standard deviation of the age of the subjects in the first set; m Age is the mean age of the subjects in the first set; s Age is the standard deviation of the age of the subjects in the first set;
[0110] Step S32: The training module performs a performance evaluation on the metabolic age output by the trained metabolic aging model. The metabolic age model whose performance evaluation meets the preset requirements is used as the final metabolic aging recognition model for accurately predicting all-cause mortality risk and identifying metabolic aging.
[0111] More preferably, step S32 includes:
[0112] In step S321, the training module obtains the trained metabolic aging recognition model in R (4.3.1) software, and performs the predicted 5-year all-cause mortality risk performance on the first set and the test set respectively. If the difference in the predicted performance of the trained metabolic aging recognition model in the training set and the test set is less than a preset difference threshold and both are greater than the actual age of the subject, indicating that there is no overfitting risk and the metabolic aging assessment performance is greater than the age itself, the trained metabolic aging recognition model is used as the final metabolic aging recognition model.
[0113] All-cause mortality is a core indicator in medical and epidemiological research, encompassing all deaths regardless of their specific cause. A higher risk of all-cause mortality during the five-year follow-up is associated with a shorter lifespan, more significant metabolic disturbances, an older metabolic age, and a more advanced metabolic aging process.
[0114] Therefore, by recording all-cause mortality events over a five-year period, combined with screened plasma small molecule metabolic biomarkers and machine learning, a metabolic-based all-cause mortality risk score (P) can be trained (P ranges from 0 to 1, representing a low to high all-cause mortality risk, with 0.04, the proportion of actual deaths in the training cohort, used as the prediction cutoff). Metabolic age can be calculated based on P, as well as the difference between metabolic age and actual age, to intuitively identify the level of metabolic aging and whether metabolic aging is accelerated compared to peers. Similarly, the correlation between metabolic age and P can provide an intuitive prediction of all-cause mortality risk.
[0115] After obtaining metabolic age through plasma small molecule metabolic biomarkers and age in the first set and test set, the predictive performance of metabolic age on the risk of all-cause mortality was evaluated using the area under the receiver operating characteristic curve (AUC), accuracy, specificity, and sensitivity.
[0116] The training module used the trained metabolic aging recognition model on the R (4.3.1) software to make predictions on the first set and the test set, respectively, and obtained the recognition performance (recognition performance) on the first set and the test set. The recognition performance included: AUC values of 0.728 for the training set and 0.752 for the test set, respectively, showing high accuracy and sensitivity. The specific performance indicators are shown in Table 2. If the recognition performance curves of the first set and the test set are close to each other (Delong test P = 0.26>0.05), and both are greater than the actual age (such as Figure 1 and 2 As shown in the figure, there is no overfitting risk and the metabolic aging assessment performance is greater than the actual age, indicating that it is a well-trained metabolic aging recognition model and can be used as the final metabolic aging recognition model.
[0117] The metabolic aging recognition model provided by the present invention has the characteristics of low sample consumption and high reproducibility. By comparing the performance of the training set and the performance of the test set, it is determined whether the model used for metabolic aging recognition is at risk of overfitting. If the performance is close, it indicates that the model is universal. If it is not close, it indicates that the generalization ability of the model is relatively poor and a better model needs to be trained. By comparing the performance of metabolic age and the performance of actual age, it is determined whether the performance of the model used for metabolic aging recognition is better than that of age. If it is better, it means that the model captures metabolic disorders that age alone cannot reflect, and is a superior metabolic aging assessment model.
[0118] Table 2. Performance of metabolic aging recognition model and machine learning obtained by age
[0119]
[0120] The model prediction results are represented by a 2x2 confusion matrix, where rows represent actual labels and columns represent predicted labels. The confusion matrix divides the samples into four categories:
[0121] TP (True Positive) True Positive: The number of samples that are actually positive examples and predicted as positive examples.
[0122] TN (True Negative) True Negative Examples: The number of samples that are actually negative examples but predicted to be negative examples.
[0123] FP (False Positive): False positive examples: The number of samples that are actually negative examples but predicted to be positive examples.
[0124] FN (False Negative) False negative examples: The number of samples that are actually positive examples but predicted as negative examples.
[0125] Accuracy: The proportion of correctly predicted samples in the total samples, calculated as (TP+FP) / (FN+TN+TP+TN).
[0126] F1 value: A comprehensive indicator of the proportion of samples predicted as positive among actual positive samples and the proportion of correctly predicted samples among positive samples. The calculation formula is 2*P*R / (P+R), where P=TP / (TP+FP) and R=TP / (TP+FN).
[0127] Sensitivity: The proportion of samples predicted to be positive among samples that are actually positive. The calculation formula is TP / (TP+FN).
[0128] Specificity: The proportion of samples predicted to be negative among samples that are actually negative. The calculation formula is TN / (TN+FP).
[0129] AUC: Area under the ROC curve. The ROC curve is a curve consisting of 1-specificity and sensitivity at different thresholds. The AUC is not affected by the ratio of positive and negative samples. It reflects the overall performance of the model at different thresholds. It ranges from 0 to 1. The larger the AUC, the better the overall performance of the model.
[0130] In step S4, the recognition module obtains the age of the subject to be tested and the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection, and inputs the age of the subject to be tested and the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection into the final metabolic aging recognition model to obtain the metabolic age of the subject to be tested, which can accurately predict the risk of all-cause mortality.
[0131] Here, the probability of no all-cause death and the probability of all-cause death during the 5-year follow-up are actually the metabolic-based all-cause death risk score P mentioned above. The range of P is 0 to 1, representing a range from low all-cause death risk to high all-cause death risk. The proportion of the actual number of deaths in the first set, i.e., 0.04, is used as the prediction cutoff value.
[0132] Metabolic age is P which can be calculated using the aforementioned formula (1).
[0133] Clearly, the older the metabolic age, the higher the risk of all-cause mortality. Furthermore, metabolic age can be used to intuitively reflect the level of metabolic aging and whether metabolic aging is accelerated compared to peers, increasing the risk of all-cause mortality. For example, one person in the test set had a chronological age of 65 and a metabolic age of 65 (corresponding to a P of 0.017), indicating no accelerated metabolic aging. Meanwhile, another person had a chronological age of 65 and a metabolic age of 70 (corresponding to a P of 0.046), indicating a 5-year accelerated metabolic aging and, therefore, a higher risk of all-cause mortality.
[0134] The method for obtaining the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection of the subject to be tested can refer to the method for obtaining in step 1.
[0135] In addition, we can also conduct an experimental verification to evaluate the ability of the model to reflect the future all-cause mortality risk: in the test set, we calculate the metabolic age based on the above method, and obtain the difference between metabolic age and age, that is, metabolic aging acceleration. We use the Kaplan-Meier curve to visualize the cumulative all-cause mortality risk of individuals in the top 25% of metabolic aging acceleration (metabolic aging acceleration > 2 years) and the rest of the individuals. Figure 3 As shown, individuals in the top 25% of individuals with accelerated metabolic aging had a statistically significant higher cumulative risk of all-cause mortality during the five-year follow-up period (Log-rank P < 0.05). This analysis shows that the difference between metabolic age and chronological age derived from the final prediction model of the present invention reflects changes in the rate of metabolic aging and can further reflect the risk of all-cause mortality based on chronological age.
[0136] In summary, the present invention develops a metabolic aging recognition model through machine learning of age and plasma metabolite biomarkers. The metabolic age calculated from this model demonstrates excellent sensitivity, specificity, and accuracy in predicting all-cause mortality risk, exceeding chronological age alone. Experiments have confirmed its ability to intuitively reflect accelerated metabolic aging through comparison with chronological age. The present invention requires only a small amount of plasma as a test sample, resulting in a simple and convenient evaluation process, low cost, and easy acceptance by subjects, making it suitable for large-scale dissemination. Furthermore, its high sensitivity and specificity make it suitable for screening anti-aging drugs.
[0137] According to another aspect of the present invention, a computer-readable storage medium is provided, on which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor is caused to perform the following steps:
[0138] Step S1: The collection module collects plasma samples from subjects who have experienced all-cause mortality and those who have not experienced all-cause mortality during the 5-year follow-up period, and obtains plasma small molecule metabolic biomarkers and their relative concentration values of the subjects who have experienced all-cause mortality and those who have not experienced all-cause mortality during the 5-year follow-up period from each plasma sample;
[0139] Step S2: The feature selection module performs machine learning based on the plasma small molecule metabolic biomarkers and relative concentration values of the subjects who did not experience all-cause mortality during the 5-year follow-up period and the subjects who experienced all-cause mortality during the 5-year follow-up period, to screen and obtain the plasma small molecule metabolic biomarkers that identify the subjects who did not experience all-cause mortality during the 5-year follow-up period and the subjects who experienced all-cause mortality during the 5-year follow-up period, as the plasma small molecule metabolic biomarkers after feature selection;
[0140] Step S3: The training module trains a metabolic aging recognition model to obtain metabolic age based on the plasma small molecule metabolic biomarkers, relative concentration values, and actual age of each subject who did not experience all-cause mortality during the 5-year follow-up period and who experienced all-cause mortality during the 5-year follow-up period;
[0141] Step S4: The recognition module obtains the age of the subject to be tested and the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection, and inputs the age of the subject to be tested and the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection into the metabolic aging recognition model to obtain the metabolic age of the subject to be tested, which can predict the all-cause mortality risk of the subject to be tested.
[0142] According to another aspect of the present invention, there is also provided a use of a metabolite biomarker in a product for identifying metabolic aging, wherein the metabolite biomarker comprises at least the following metabolite biomarkers:
[0143] D-Glucose, L-Arginine, Taurine, L-Threonine, Creatinine, L-Acetylcarnitine, L-Serine, Uric acid, Guanidoacetic acid, Pipecolic acid, L-Aspartic acid, Oxoglutaric acid, L-Leucine, and L-Phenylalanine.
[0144] According to another aspect of the present invention, a detection kit for metabolic age is provided, wherein the metabolite biomarkers in the detection kit include at least the following metabolite biomarkers:
[0145] D-Glucose, L-Arginine, Taurine, L-Threonine, Creatinine, L-Acetylcarnitine, L-Serine, Uric acid, Guanidoacetic acid, Pipecolic acid, L-Aspartic acid, Oxoglutaric acid, L-Leucine, and L-Phenylalanine;
[0146] The metabolite biomarkers were feature selected and identified using the following modules:
[0147] The acquisition module collects plasma samples from subjects who have or have not experienced all-cause mortality during the 5-year follow-up period, and obtains plasma small molecule metabolic biomarkers and their relative concentration values from each plasma sample for subjects who have or have not experienced all-cause mortality during the 5-year follow-up period;
[0148] The feature selection module performs machine learning based on the plasma small molecule metabolic biomarkers and their relative concentration values of subjects who did not experience all-cause mortality during the 5-year follow-up and subjects who experienced all-cause mortality during the 5-year follow-up, to screen and identify the plasma small molecule metabolic biomarkers that identify subjects who did not experience all-cause mortality during the 5-year follow-up and subjects who experienced all-cause mortality during the 5-year follow-up, and use them as the plasma small molecule metabolic biomarkers after feature selection;
[0149] The training module trains a metabolic aging recognition model for identifying metabolic age based on the plasma small molecule metabolic biomarkers, relative concentrations, and actual age of subjects who did not experience all-cause mortality during the 5-year follow-up period and those who did experience all-cause mortality during the 5-year follow-up period.
[0150] The identification module obtains the age of the subject to be tested and the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection, and inputs the age of the subject to be tested and the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection into the metabolic aging identification model to obtain the metabolic age of the subject to be tested, and predicts the all-cause mortality risk of the subject to be tested based on the metabolic age.
[0151] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
[0152] It should be noted that the present invention may be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of the present invention may be executed by a processor to implement the steps or functions described above. Similarly, the software program of the present invention (including related data structures) may be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, a floppy disk, or the like. In addition, some steps or functions of the present invention may be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.
[0153] In addition, a portion of the present invention may be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. The program instructions for calling the method of the present invention may be stored in a fixed or removable recording medium, and / or transmitted through a data stream in a broadcast or other signal-carrying medium, and / or stored in a working memory of a computer device that operates according to the program instructions. Here, according to one embodiment of the present invention, a device is included, which includes a memory for storing computer program instructions and a processor for executing the program instructions, wherein, when the computer program instructions are executed by the processor, the device is triggered to operate based on the aforementioned methods and / or technical solutions according to multiple embodiments of the present invention.
[0154] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive, and the scope of the invention is defined by the appended claims rather than the above description, and it is intended that all changes that fall within the meaning and scope of the equivalents of the claims be encompassed within the present invention. Any figure marks in the claims should not be regarded as limiting the claims involved. In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the device claim may also be implemented by one unit or device through software or hardware. Words such as first and second are used to indicate names and do not indicate any particular order.
Claims
1. A metabolic aging recognition model construction system, characterized by: include: An acquisition module is used to collect plasma samples from subjects who have or have not experienced all-cause mortality during the 5-year follow-up period, and to obtain plasma small molecule metabolic biomarkers and their relative concentration values from each plasma sample for subjects who have or have not experienced all-cause mortality during the 5-year follow-up period; a feature selection module for performing machine learning based on the plasma small molecule metabolic biomarkers and relative concentration values of subjects who did not experience all-cause mortality during the 5-year follow-up period and subjects who experienced all-cause mortality during the 5-year follow-up period, to screen and identify the plasma small molecule metabolic biomarkers of subjects who did not experience all-cause mortality during the 5-year follow-up period and subjects who experienced all-cause mortality during the 5-year follow-up period, as the plasma small molecule metabolic biomarkers after feature selection; The training module trains a metabolic aging recognition model for identifying metabolic age based on the plasma small molecule metabolic biomarkers, relative concentrations, and actual age of subjects who did not experience all-cause mortality during the 5-year follow-up period and those who did experience all-cause mortality during the 5-year follow-up period. an identification module, which obtains the age of the subject to be tested and the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection, and inputs the age of the subject to be tested and the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection into the metabolic aging identification model to obtain the metabolic age of the subject to be tested; The plasma small molecule metabolic biomarkers after feature selection include at least the following metabolite biomarkers in peripheral plasma: Glucose, arginine, taurine, threonine, creatinine, acetylcarnitine, serine, uric acid, guanidinoacetic acid, pipecolic acid, aspartic acid, alpha-ketoglutarate, leucine, and phenylalanine.
2. The metabolic aging recognition model construction system according to claim 1, characterized in that: an acquisition module, configured to obtain a first subject group and a second subject group that do not overlap with each other, wherein the first subject group and the second subject group both include subjects who do and do not experience all-cause mortality during a non-overlapping 5-year follow-up period; a collection module, configured to collect first plasma samples from a first group of subjects, and obtain plasma small molecule metabolic biomarkers and their relative concentration values from each first plasma sample as a first set; The acquisition module is used to collect second plasma samples from the second subject group, and obtain plasma small molecule metabolic biomarkers and their relative concentration values from each second plasma sample as a test set.
3. The metabolic aging recognition model construction system according to claim 2, characterized in that: The feature selection module is configured to divide the first set into five non-overlapping data sets, each time using one of the five non-overlapping data sets that has not been selected as an internal validation set, and each time using the remaining four data sets as training sets; wherein both the internal validation set and the training set include non-overlapping subjects who experienced all-cause mortality during the five-year follow-up period and subjects who did not experience all-cause mortality during the five-year follow-up period; The feature selection module is used to train the feature selection model for 5 rounds based on each training set using the Lasso feature selection algorithm to obtain the corresponding 5-round feature selection model. In each round of feature selection model, for each λn value in the hyperparameter sequence, a corresponding feature selection model is generated. Mk ( λn ), n=1-100; k=1-5, k represents the number of rounds; n represents the sequence number of λ value, λ is one of the hyperparameters; Feature selection module for each value in the hyperparameter sequence λn , input the internal validation set corresponding to each round into the feature selection model of the corresponding round in the 5 feature selection models Mk ( λn ), and obtain each λn The error rate corresponding to the value on the internal validation set of the corresponding round; the average error rate corresponding to each value is obtained from the 5 error rates corresponding to each value; based on the minimum average error rate among all values, the value with the lowest error rate is selected λ .min, where the error rate is the lowest. λ .min, and the corresponding metabolite markers were obtained as plasma small molecule metabolic biomarkers after feature selection.
4. The metabolic aging recognition model construction system according to claim 2, characterized in that: The training module is used to train a metabolic aging recognition model based on the plasma small molecule metabolic biomarkers after feature selection, their relative concentration values and actual age in each plasma sample in the first set.
5. The metabolic aging recognition model construction system according to claim 4, characterized in that: A training module is used to train a metabolic aging recognition model for identifying metabolic age, wherein the metabolic aging recognition model for identifying metabolic age uses an elastic network algorithm and is based on the plasma small molecule metabolic biomarkers, relative concentration values, and actual age of each subject in the first set who did not experience all-cause mortality during the 5-year follow-up period and who experienced all-cause mortality during the 5-year follow-up period, and automatically outputs a metabolic-based all-cause mortality risk score P for each subject in the first set, where P ranges from 0 to 1, representing a range from low all-cause mortality risk to high all-cause mortality risk; the implementation of the elastic network algorithm relies on glm net package, wherein the ElasticNet algorithm hyperparameter alpha for controlling the ratio of L1 regularization and L2 regularization is set to 0.5; s for controlling the penalty strength of the parameter is set to 0.01; the metabolic aging recognition model for identifying metabolic age calculates the mean and standard deviation of the age of the subjects in the first set, and calculates the mean and standard deviation of the metabolic-based all-cause mortality risk score P for all subjects in the first set; the mean and standard deviation of the metabolic-based all-cause mortality risk score P are scaled to be consistent with the mean and standard deviation of the age of the subjects in the first set, thereby obtaining the metabolic age in years; The training module is used to perform performance evaluation on the metabolic age output by the trained metabolic aging model, and the metabolic aging model whose performance evaluation meets the preset requirements is used as the final metabolic aging model for identifying metabolic age.
6. The metabolic aging recognition model construction system according to claim 5, characterized in that: The calculation formula of the metabolic age is as follows, Among them, metabolic age i It is the first set i Metabolic age after conversion of samples; P i It is i Metabolic-based all-cause mortality risk score for each subject; μ p is the mean of the metabolic-based all-cause mortality risk score P for all subjects in the first set; σ p is the sample standard deviation of the metabolic-based all-cause mortality risk score P for all subjects in the first set; μ Age is the mean age of the subjects in the first set; σ Age The standard deviation of the ages of the subjects in the first set.
7. The metabolic aging recognition model construction system according to claim 2, characterized in that: The training module uses the trained metabolic aging model to obtain the trained metabolic aging recognition model, and its recognition performance on the first set and the recognition performance on the test set; if the difference in the recognition performance of the trained metabolic aging recognition model in the first set and the test set of metabolic age is less than a preset difference threshold, and both are greater than the actual age of the subject, the trained metabolic aging recognition model will be used as the final metabolic aging recognition model.
8. A computer-readable storage medium having computer-executable instructions stored thereon, wherein: When the computer executable instructions are executed by a processor, the processor performs the following steps: The collection module collects plasma samples from subjects who have or have not experienced all-cause mortality during the 5-year follow-up period, and obtains plasma small molecule metabolic biomarkers and their relative concentration values from each plasma sample; The feature selection module performs machine learning based on the plasma small molecule metabolic biomarkers and their relative concentration values of subjects who did not experience all-cause mortality during the 5-year follow-up and subjects who experienced all-cause mortality during the 5-year follow-up, to screen and identify the plasma small molecule metabolic biomarkers that identify subjects who did not experience all-cause mortality during the 5-year follow-up and subjects who experienced all-cause mortality during the 5-year follow-up, and use them as the plasma small molecule metabolic biomarkers after feature selection; The training module trains a metabolic aging recognition model for identifying metabolic age based on the plasma small molecule metabolic biomarkers, their relative concentration values, and actual age of each subject who did not experience all-cause mortality during the 5-year follow-up period and those who did experience all-cause mortality during the 5-year follow-up period. The recognition module obtains the age of the subject to be tested and the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection, and inputs the age of the subject to be tested and the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection into the metabolic aging recognition model to obtain the metabolic age of the subject to be tested; The plasma small molecule metabolic biomarkers after feature selection include at least the following metabolite biomarkers in peripheral plasma: Glucose, arginine, taurine, threonine, creatinine, acetylcarnitine, serine, uric acid, guanidinoacetic acid, pipecolic acid, aspartic acid, alpha-ketoglutarate, leucine, and phenylalanine.
9. A detection kit for metabolic age, wherein the metabolite biomarkers in the detection kit include at least the following metabolite biomarkers: glucose, arginine, taurine, threonine, creatinine, acetylcarnitine, serine, uric acid, guanidinoacetic acid, pipecolic acid, aspartic acid, alpha-ketoglutarate, leucine, and phenylalanine; The metabolite biomarkers were feature selected and identified using the following modules: The acquisition module collects plasma samples from subjects who have or have not experienced all-cause mortality during the 5-year follow-up period, and obtains plasma small molecule metabolic biomarkers and their relative concentration values from each plasma sample for subjects who have or have not experienced all-cause mortality during the 5-year follow-up period; The feature selection module performs machine learning based on the plasma small molecule metabolic biomarkers and their relative concentration values of subjects who did not experience all-cause mortality during the 5-year follow-up and subjects who experienced all-cause mortality during the 5-year follow-up, to screen and identify the plasma small molecule metabolic biomarkers that identify subjects who did not experience all-cause mortality during the 5-year follow-up and subjects who experienced all-cause mortality during the 5-year follow-up, and use them as the plasma small molecule metabolic biomarkers after feature selection; The training module trains a metabolic aging recognition model for identifying metabolic age based on the plasma small molecule metabolic biomarkers, relative concentrations, and actual age of subjects who did not experience all-cause mortality during the 5-year follow-up period and those who did experience all-cause mortality during the 5-year follow-up period. The identification module obtains the age of the person to be tested and the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection, and inputs the age of the person to be tested and the plasma small molecule metabolic biomarkers and their relative concentration values after feature selection into the metabolic aging identification model to obtain the metabolic age of the person to be tested.
Citation Information
Patent Citations
Metabolic age prediction model and application thereof in colorectal cancer diagnosis
CN114334170A
Cardiovascular metabolism risk factor spectrum recognition model construction system, storage medium and kit
CN119181506A