Method and model for predicting rat age by using rat serum metabolite

By using a multidimensional dataset of rat serum metabolites, combined with univariate linear regression and the Lasso model, an age prediction model was developed. This model addresses the shortcomings of existing technologies in assessing organ-specific aging and enables accurate age prediction and health management support.

CN120998344APending Publication Date: 2025-11-21PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511393871.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing age prediction models are unable to accurately characterize the aging state of individual organs, resulting in a lack of organ-specific aging assessment and an inability to comprehensively assess an individual's aging state.

Method used

Using a multidimensional dataset composed of rat serum metabolites, an age prediction model was developed through univariate linear regression and the Lasso model. The age was calculated using a weighted summation formula and predicted in combination with metabolite content data.

Benefits of technology

It has achieved relatively accurate prediction of the actual age of rats, improved the accuracy and reliability of age prediction, has generalization ability, supports personalized health management, and promotes the development of precision medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The invention relates to the technical field of medicine, in particular to a model and method for predicting the age of a rat, and the model comprises a data collection module which is used for obtaining data of metabolites of a subject; the data processing module is used for performing data standardization processing on the data information acquired in the data acquisition module; and the age calculation module is used for calculating the data information processed in the data processing module so as to calculate the age of the subject.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical technology, specifically to a model and method for predicting the age of rats. Background Technology

[0002] With the aging population, there is a growing concern about extending healthy lifespan. Current mainstream age prediction models primarily rely on epigenetic clocks (such as DNA methylation mapping) and transcriptome analysis, assessing an individual's biological age by detecting specific biomarkers. However, these methods struggle to accurately characterize the aging state of individual organs, leading to a lack of organ-specific aging assessment.

[0003] The decline in the function of the immune system, as the central regulatory network for maintaining systemic homeostasis, has become a key factor in systemic aging. With increasing age, immune capacity gradually weakens, leading to impaired protective immune responses. This process involves dynamic changes in multiple immune tissues and organs, such as the spleen, thymus, and lymph nodes, as well as corresponding cytokine networks. However, the potential of immune cell profiles and cytokine characteristics as innovative biomarkers for assessing biological age has not yet been fully explored.

[0004] In the future, with the Health 2025 plan's emphasis on precision medicine, age prediction models will evolve towards multi-organ, multi-factor, and multi-dimensional approaches. To address the shortcomings of existing models, there is an urgent need to establish a model capable of comprehensively assessing the aging status of an individual's various organs. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this application provides a model capable of predicting the age of rats, the specific scheme of which is as follows:

[0006] 1. A metabolome for predicting age, comprising one or more metabolites selected from the group consisting of: icocarbidol, N-nitrosothiazolidine-4-carboxylic acid, mono-L-menthol glutaric acid, myristic acid, (2E)-4-hydroxy-5-methyl-2-propane-3(2H)-furanone, and (3β,5α,6β,22E,24R)-23-methylergoster-7,22-diene-3,5,6-triol. , 10-[5]-Ladder-alkyl-decanoic acid, 2,2'-dithiodipyridine, gemmaconone, monogalactosyl glycerol (18:2(9Z,12) / 0:0), josinoside, hydroxymethylcytosine, 1-palmitoyl-2-linoleoyl-sn-glycerol-3-phosphocholine, 5,6-dihydroxyprostaglandin F1a, 3-(but-3-en-1-yn-1-yl)-6-methyl-1,2-dithioene, cholic acid 7-sulfate.

[0007] 2. A system for predicting age, comprising:

[0008] The data acquisition module is used to acquire data on the metabolites of the subjects;

[0009] A data processing module is used to perform data standardization processing on the data information acquired by the data acquisition module; and

[0010] An age calculation module is used to calculate the age of the subject by performing calculations on the data information processed in the data processing module.

[0011] The subjects were rats.

[0012] 3. The system according to item 2, wherein:

[0013] In the data acquisition module, the collected metabolite data refers to the types and / or amounts of metabolites detected on any day after the subject's birth.

[0014] 4. The system according to claim 2, wherein the metabolite is any one or more metabolites selected from the group consisting of: icocarbidol, N-nitrosothiazolidine-4-carboxylic acid, mono-L-menthol glutaric acid, myristic acid, (2E)-4-hydroxy-5-methyl-2-propane-3(2H)-furanone, (3β,5α,6β,22E,24R)-23-methylergoster-7,22-diene-3,5,6-triol , 10-[5]-Ladder-alkyl-decanoic acid, 2,2'-dithiodipyridine, gemmaconone, monogalactosyl glycerol (18:2(9Z,12) / 0:0), josinoside, hydroxymethylcytosine, 1-palmitoyl-2-linoleoyl-sn-glycerol-3-phosphocholine, 5,6-dihydroxyprostaglandin F1a, 3-(but-3-en-1-yn-1-yl)-6-methyl-1,2-dithioene, cholic acid 7-sulfate.

[0015] 5. The system according to item 2, wherein:

[0016] The age calculation module contains a pre-stored formula for calculating age (Y) based on data fitted from the metabolites of subjects in an existing database.

[0017] 6. The system according to item 5, wherein the formula for calculating age (Y) is a weighted summation formula.

[0018] 7. The system according to item 6, wherein:

[0019] The formula is as follows: Formula 1

[0020]

[0021] Where Y is the calculated age of the subject, βjA is a unitless parameter, b1 is a constant, and X 1A X represents the level of icoracetin in the serum of the subjects. 2A X represents the concentration of N-nitrosothiazolidine-4-carboxylic acid in the serum of the subjects. 3A X represents the serum glutaric acid mono-L-menthol ester content in the subjects. 4A X represents the level of myristic acid in the serum of the subjects. 5A X represents the concentration of (2E)-4-hydroxy-5-methyl-2-propane-3(2H)-furanone in the serum of the subjects. 6A X represents the concentration of (3β,5α,6β,22E,24R)-23-methylergoster-7,22-diene-3,5,6-triol in the serum of the subjects. 7A X represents the content of 10-[5]-ladder-decanoic acid in the serum of the subjects. 8A X represents the concentration of 2,2'-dithiodipyridine in the serum of the subjects. 9A X represents the level of gemcitabine in the serum of the subject. 10A X represents the serum concentration of monogalactosyl glycerol (18:2(9Z,12) / 0:0) in the subjects. 11A X represents the concentration of ginsenosides in the serum of the subjects. 12A X represents the concentration of hydroxymethylcytosine in the serum of the subject. 13A X represents the serum concentration of 1-palmitoyl-2-linoleoyl-sn-glycerol-3-phosphate choline in the subjects. 14A X represents the serum 5,6-dihydroxyprostaglandin F1a level in the subjects. 15A X represents the concentration of 3-(but-3-en-1-yn-1-yl)-6-methyl-1,2-dithioene in the serum of the subjects. 16A The concentration of cholic acid 7-sulfate in the serum of the subjects.

[0022] 8. The system according to item 7, wherein:

[0023] b1 is selected from any value between -6.4 and -6.2, preferably -6.319993; β 1A β is selected from any value between 13.8 and 14.0, preferably 13.967592; 2A Any value selected from 11.4 to 11.6, preferably 11.568449; β 3A β is selected from any value between -6.4 and -6.2, preferably -6.366745; 4A β is selected from any value between 5.5 and 5.72, preferably 5.626721; 5A Any value selected from -4.4 to -4.2, preferably -4.393360; β6A Any value selected from -4.0 to -3.8, preferably -3.974611; β 7A Any value selected from 3.7 to 3.9, preferably 3.851389; β 8A Any value selected from -3.8 to -3.6, preferably -3.785134; β 9A β is selected from any value between -3.7 and -3.5, preferably -3.600627; 10A β is selected from any value between -3.4 and -3.2, preferably -3.363541; 11A Any value selected from -3.2 to -3.0, preferably -3.169782; β 12A Any value selected from 2.7 to 2.9, preferably 2.890089; β 13A Any value selected from -1.8 to -1.6, preferably -1.716091; β 14A β is selected from any value between 1.3 and 1.5, preferably 1.467066; 15A The value is selected from any value between -0.9 and -0.7, preferably -0.842830; β 16A The value is selected from any value between -0.3 and -0.1, preferably -0.265457.

[0024] 9. A method for predicting age, comprising:

[0025] The data acquisition step involves obtaining data on the subject's metabolites;

[0026] The data processing step standardizes the data information acquired in the data acquisition step; and

[0027] The step of calculating age involves calculating the age of the subject by processing the data information processed in the data processing step.

[0028] The subjects were rats.

[0029] 10. The method according to item 9, wherein:

[0030] In the data acquisition step, the collected metabolite data refers to the types and / or amounts of metabolites detected on any day after the subject's birth.

[0031] 11. The method according to claim 9, wherein the metabolite is any one or more metabolites selected from the group consisting of: icocarbidol, N-nitrosothiazolidine-4-carboxylic acid, mono-L-menthol glutaric acid, myristic acid, (2E)-4-hydroxy-5-methyl-2-propane-3(2H)-furanone, (3β,5α,6β,22E,24R)-23-methylergoster-7,22-diene-3,5,6-triol , 10-[5]-Ladder-alkyl-decanoic acid, 2,2'-dithiodipyridine, gemmaconone, monogalactosyl glycerol (18:2(9Z,12) / 0:0), josinoside, hydroxymethylcytosine, 1-palmitoyl-2-linoleoyl-sn-glycerol-3-phosphocholine, 5,6-dihydroxyprostaglandin F1a, 3-(but-3-en-1-yn-1-yl)-6-methyl-1,2-dithioene, cholic acid 7-sulfate.

[0032] 12. The method according to item 9, wherein:

[0033] The age calculation module contains a pre-stored formula for calculating age (Y) based on data fitted from the metabolites of subjects in an existing database.

[0034] 13. The method according to item 12, wherein the formula for calculating age (Y) is a weighted summation formula.

[0035] 14. The method according to item 13, wherein:

[0036] The formula is as follows: Formula 1

[0037]

[0038] Where Y is the calculated age of the subject, βjA is a unitless parameter, b1 is a constant, and X 1A X represents the level of icoracetin in the serum of the subjects. 2A X represents the concentration of N-nitrosothiazolidine-4-carboxylic acid in the serum of the subjects. 3A X represents the serum glutaric acid mono-L-menthol ester content in the subjects. 4A X represents the level of myristic acid in the serum of the subjects. 5A X represents the concentration of (2E)-4-hydroxy-5-methyl-2-propane-3(2H)-furanone in the serum of the subjects. 6A X represents the concentration of (3β,5α,6β,22E,24R)-23-methylergoster-7,22-diene-3,5,6-triol in the serum of the subjects. 7A X represents the content of 10-[5]-ladder-decanoic acid in the serum of the subjects. 8A X represents the concentration of 2,2'-dithiodipyridine in the serum of the subjects.9A X represents the level of gemcitabine in the serum of the subject. 10A X represents the serum concentration of monogalactosyl glycerol (18:2(9Z,12) / 0:0) in the subjects. 11A X represents the concentration of ginsenosides in the serum of the subjects. 12A X represents the concentration of hydroxymethylcytosine in the serum of the subject. 13A X represents the serum concentration of 1-palmitoyl-2-linoleoyl-sn-glycerol-3-phosphate choline in the subjects. 14A X represents the serum 5,6-dihydroxyprostaglandin F1a level in the subjects. 15A X represents the concentration of 3-(but-3-en-1-yn-1-yl)-6-methyl-1,2-dithioene in the serum of the subjects. 16A The concentration of cholic acid 7-sulfate in the serum of the subjects.

[0039] 15. The method according to item 14, wherein: b1 is selected from any value from -6.4 to -6.2, preferably -6.319993; β 1A β is selected from any value between 13.8 and 14.0, preferably 13.967592; 2A Any value selected from 11.4 to 11.6, preferably 11.568449; β 3A β is selected from any value between -6.4 and -6.2, preferably -6.366745; 4A Any value selected from 5.5 to 5.7, preferably 5.626721; β 5A Any value selected from -4.4 to -4.2, preferably -4.393360; β 6A Any value selected from -4.0 to -3.8, preferably -3.974611; β 7A Any value selected from 3.7 to 3.9, preferably 3.851389; β 8A Any value selected from -3.8 to -3.6, preferably -3.785134; β 9A β is selected from any value between -3.7 and -3.5, preferably -3.600627; 10A β is selected from any value between -3.4 and -3.2, preferably -3.363541; 11A Any value selected from -3.2 to -3.0, preferably -3.169782; β 12A Any value selected from 2.7 to 2.9, preferably 2.890089; β 13A Any value selected from -1.8 to -1.6, preferably -1.716091; β 14Aβ is selected from any value between 1.3 and 1.5, preferably 1.467066; 15A The value is selected from any value between -0.9 and -0.7, preferably -0.842830; β 16A The value is selected from any value between -0.3 and -0.1, preferably -0.265457.

[0040] The beneficial effects of this application are as follows:

[0041] This application, based on a multidimensional dataset of rat serum metabolites, identified age-related characteristic curves and developed an age prediction model based on univariate linear regression and the Lasso model. This resulted in a unified aging prediction index, and the most valuable metabolite for age prediction was determined using minimum mean squared error (MSE). The model demonstrated good predictive ability and high stability on the validation set, exhibiting strong generalization ability. This model can accurately predict the actual age of SD rats and can be effectively extended to new, unseen data, demonstrating significant practical application value.

[0042] The model and method provided in this application significantly improve the accuracy and reliability of age prediction, provide strong support for personalized health management, help promote the development of precision medicine, and ultimately aim to extend healthy lifespan. Attached Figure Description

[0043] The accompanying drawings are provided to better understand this application and do not constitute an undue limitation thereof. Wherein:

[0044] Figure 1 A schematic diagram illustrating the construction of a Lasso age prediction model based on a rat serum metabolite dataset.

[0045] Figure 2 A schematic diagram of the residual histogram of the model established for the minimum mean square error (MSE).

[0046] Figure 3 A schematic diagram showing the relationship between the model residuals and predicted values ​​established for the minimum mean square error (MSE).

[0047] Figure 4 This is a schematic diagram illustrating the validation results of the Lasso age prediction model constructed based on serum metabolites from SD rats. Detailed Implementation

[0048] Specific embodiments of the present application will now be described in more detail with reference to the accompanying drawings. While specific embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0049] It should be noted that certain terms are used in the specification and claims to refer to specific components. Those skilled in the art will understand that different terms may be used to refer to the same component. This specification and claims do not distinguish components based on differences in terminology, but rather on differences in function. The terms "comprising" or "including" used throughout the specification and claims are open-ended and should be interpreted as "comprising but not limited to." The following descriptions in the specification are preferred embodiments for carrying out this application; however, these descriptions are for the purpose of understanding the general principles of the specification and are not intended to limit the scope of this application. The scope of protection of this application shall be determined by the appended claims.

[0050] The type of variable is not static; it can be transformed between different types depending on the research objective. For example, hemoglobin level (g / L) is originally a numerical variable. If it is divided into two categories—normal and low—it can be analyzed as binary categorical data. If it is divided into five levels—severe anemia, moderate anemia, mild anemia, normal, and elevated hemoglobin—it can be analyzed as ordinal data. Sometimes, categorical data can also be quantified. For instance, if a patient's nausea is represented by 0, 1, 2, and 3, it can be analyzed as numerical variable data (quantitative data).

[0051] In this application, the term "linear regression" refers to a statistical, linear method used to model the relationship between a dependent variable and one or more independent variables. In linear regression, the relationship is modeled using a linear predictor function estimated from the data using its unknown model parameters. This model is called a linear model. Typically, for the case of a single independent variable, the (x, y) data points are plotted graphically as a scatter plot, where x is the independent variable and y is the dependent variable. Linear regression aims to obtain a "best-fit line" representing the relationship between the dependent and independent variables. Linear regression models are typically fitted using least squares, but they can also be fitted in other ways, such as by minimizing a "cost function" in some other norm (e.g., L1-norm penalty or L2-norm penalty).

[0052] As used in this paper, the term "minimum absolute shrinkage and operator selection regression (often simply called Lasso regression)" is a compression estimation method based on the idea of ​​reducing the variable set (order reduction). It constructs a penalty function to compress the coefficients of variables and make some regression coefficients zero, thereby achieving variable selection. It is an algorithm that uses a penalty function to improve the predictive power of a model. This algorithm, using 1-norm constraints, not only solves problems of high dimensionality and collinearity but also makes the established model "sparse," meaning the algorithm has an automatic wavelength selection effect during modeling.

[0053] This application provides a system for predicting age, comprising:

[0054] The data acquisition module is used to acquire data on the metabolites of the subjects;

[0055] A data processing module is used to perform data standardization processing on the data information acquired by the data acquisition module; and

[0056] An age calculation module is used to calculate the age of the subject by performing calculations on the data information processed in the data processing module.

[0057] The subjects were rats.

[0058] In one embodiment of this application, obtaining data on the subject's metabolites includes extracting the subject's blood and detecting data on metabolites derived from serum from the blood.

[0059] For example, whole blood can be extracted using methods well-known in the art, and serum can be separated using suitable methods. In one embodiment of this application, the whole blood and / or serum can be extracted from any part of the subject (organ, tissue, blood vessel), such as blood collected from the abdominal aorta, the heart, the tail vein, the orbital venous plexus, the femoral artery / femoral vein, the jugular vein / carotid artery, etc. The metabolites can be detected in a manner well-known to those skilled in the art and should not be construed as limiting this application.

[0060] In one embodiment of this application, the metabolite data of the subjects are standardized in the data processing module. Specifically, the subjects are divided into multiple batches according to their age in months, and for the sample data of the same batch, the average value of the data from each batch is taken as the standardized data for that batch.

[0061] In one embodiment of this application, one subject sample is taken from each batch to form an internal validation set, while the other subject samples are used to build the model.

[0062] In one embodiment of this application, the data collected in the data acquisition module refers to the types and / or amounts of serum-derived metabolites detected on any day after the subject's birth.

[0063] In one embodiment of this application, the type and / or content of the serum-derived metabolites are detected for each subject in each of the aforementioned batches.

[0064] In one embodiment of this application, the terms "metabolite" and "serum-derived metabolite" may be used interchangeably and should not be construed as limiting the scope of this application.

[0065] In one embodiment of this application, after the metabolite detection is completed, further screening of the metabolites is required.

[0066] In one embodiment of this application, the serum data collected in the data acquisition module refers to the types and / or amounts of serum-derived metabolites detected on any day after the subject's birth. In another embodiment of this application, the serum-derived metabolites can be either pre-screening or post-screening metabolites. These serum-derived metabolites constitute a metabolite set for a model used to predict age.

[0067] In one aspect of this application, a metabolome for predicting age is also provided, wherein the metabolites in the metabolome include one or more metabolites selected from the group consisting of: icaridin, N-nitrosothiazolidine-4-carboxylic acid, L-monomenthyl glutarate, and myristoleic acid. acid), (2E)-4-hydroxy-5-methyl-2-propane-3(2H)-furanone, (3β,5α,6β,22E,24R)-23-methylergosta-7,22-diene-3,5,6-triol, 10-[5]-ladderane-decanoic acid One or more of the following: acid), 2,2'-dithiodipyridine, Germacrenone, monogalactosyl glycerol (18:2(9Z,12) / 0:0), choline glycoside, hydroxymethylcytosine, 1-palmitoyl-2-linoleoyl-sn-glycerol-3-phosphocholine, 5,6-dihydroxyprostaglandin F1a, 3-(but-3-en-1-yn-1-yl)-6-methyl-1,2-dithioene, and cholic acid 7-sulfate.

[0068] In one embodiment of this application, the metabolite is any one or more metabolites selected from the group consisting of: icocarbidol, N-nitrosothiazolidine-4-carboxylic acid, mono-L-menthol glutaric acid, myristic acid, (2E)-4-hydroxy-5-methyl-2-propane-3(2H)-furanone, (3β,5α,6β,22E,24R)-23-methylergoster-7,22-diene-3,5,6-triol, 1 0-[5]-Ladder-alkyl-decanoic acid, 2,2'-dithiodipyridine, gemmaconone, monogalactosyl glycerol (18:2(9Z,12) / 0:0), josinoside, hydroxymethylcytosine, 1-palmitoyl-2-linoleoyl-sn-glycerol-3-phosphocholine, 5,6-dihydroxyprostaglandin F1a, 3-(but-3-en-1-yn-1-yl)-6-methyl-1,2-dithioene, cholic acid 7-sulfate.

[0069] In one embodiment of this application, the metabolite group comprises icocarbidol, N-nitrosothiazolidine-4-carboxylic acid, mono-L-menthol glutaric acid, myristic acid, (2E)-4-hydroxy-5-methyl-2-propane-3(2H)-furanone, (3β,5α,6β,22E,24R)-23-methylergoster-7,22-diene-3,5,6-triol, 10-[5]-ladderane- It consists of decanoic acid, 2,2'-dithiodipyridine, gemimazone, monogalactosyl glycerol (18:2(9Z,12) / 0:0), jojoba glycoside, hydroxymethylcytosine, 1-palmitoyl-2-linoleoyl-sn-glycerol-3-phosphocholine, 5,6-dihydroxyprostaglandin F1a, 3-(but-3-en-1-yn-1-yl)-6-methyl-1,2-dithioene, and cholic acid 7-sulfate.

[0070] In one embodiment of this application, the age calculation module contains a pre-stored formula for calculating age (Y) fitted based on metabolite data of subjects in an existing database.

[0071] In one embodiment of this application, the existing database refers to a database of subjects that meet the above-mentioned criteria for batch size, age, and gender.

[0072] In one embodiment of this application, the formula for calculating age (Y) is a weighted summation formula.

[0073] In one embodiment of this application, the formula is as follows:

[0074]

[0075] Where Y is the calculated age of the subject, βjA is a unitless parameter, b1 is a constant, and X 1A X represents the level of icoracetin in the serum of the subjects.2A X represents the concentration of N-nitrosothiazolidine-4-carboxylic acid in the serum of the subjects. 3A X represents the serum glutaric acid mono-L-menthol ester content in the subjects. 4A X represents the level of myristic acid in the serum of the subjects. 5A X represents the concentration of (2E)-4-hydroxy-5-methyl-2-propane-3(2H)-furanone in the serum of the subjects. 6A X represents the concentration of (3β,5α,6β,22E,24R)-23-methylergoster-7,22-diene-3,5,6-triol in the serum of the subjects. 7A X represents the content of 10-[5]-ladder-decanoic acid in the serum of the subjects. 8A X represents the concentration of 2,2'-dithiodipyridine in the serum of the subjects. 9A X represents the level of germaxylone in the serum of the subjects. 10A X represents the serum concentration of monogalactosyl glycerol (18:2(9Z,12) / 0:0)(MGMG(18:2(9Z,12) / 0:0)) in the subject. 11A X represents the level of pinocembroside in the serum of the subjects. 12A X represents the level of hydroxymethylcytosine in the serum of the subject. 13A X represents the serum concentration of 1-palmitoyl-2-linoleoyl-sn-glycerol-3-phosphate choline in the subjects. 14A X represents the serum concentration of 5,6-dihydroxyprostaglandin F1a in the subjects. 15A X represents the concentration of 3-(but-3-en-1-yn-1-yl)-6-methyl-1,2-dithiine in the serum of the subjects. 16A The concentration of 7-sulfocholic acid in the serum of the subjects was measured.

[0076] Specifically, in the age calculation module, the subject's metabolite data, β... jA The values ​​of b1 and b1 can be directly calculated by substituting them into formula (1).

[0077] in:

[0078] b1 is selected from any value between -6.4 and -6.2, preferably -6.319993; β 1AAny value selected from 13.8 to 14.02, preferably 13.967592; β 2A Any value selected from 11.4 to 11.6, preferably 11.568449; β 3A β is selected from any value between -6.4 and -6.2, preferably -6.366745; 4A Any value selected from 5.5 to 5.7, preferably 5.626721; β 5A Any value selected from -4.4 to -4.2, preferably -4.393360; β 6A Any value selected from -4.0 to -3.8, preferably -3.974611; β 7A Any value selected from 3.7 to 3.9, preferably 3.851389; β 8A Any value selected from -3.8 to -3.6, preferably -3.785134; β 9A β is selected from any value between -3.7 and -3.5, preferably -3.600627; 10A β is selected from any value between -3.4 and -3.2, preferably -3.363541; 11A Any value selected from -3.2 to -3.0, preferably -3.169782; β 12A Any value selected from 2.7 to 2.9, preferably 2.890089; β 13A Any value selected from -1.8 to -1.6, preferably -1.716091; β 14A β is selected from any value between 1.3 and 1.5, preferably 1.467066; 15A The value is selected from any value between -0.9 and -0.7, preferably -0.842830; β 16A The value is selected from any value between -0.3 and -0.1, preferably -0.265457;

[0079] This application also provides a method for predicting age, comprising: a data acquisition step, which acquires metabolite data of a subject; a data processing step, which standardizes the data acquired in the data acquisition step; and an age calculation step, which calculates the age of the subject by performing calculations on the processed data in the data processing step. The subject is a rat.

[0080] In one embodiment of this application, the data collection step refers to the types and / or amounts of metabolites collected on any day after the subject's birth.

[0081] In one embodiment of this application, the type and / or content of the serum-derived metabolites are detected for each subject in each of the aforementioned batches.

[0082] In one embodiment of this application, after the metabolite detection is completed, further screening of the metabolites is required.

[0083] In one embodiment of this application, the serum data collected in the data acquisition step refers to the types and / or amounts of serum-derived metabolites detected on any day after the subject's birth. In another embodiment of this application, the serum-derived metabolites can be either pre-screening or post-screening metabolites. These serum-derived metabolites constitute a metabolite set for a model used to predict age.

[0084] In one embodiment of this application, the metabolite is any one or more metabolites selected from the group consisting of: icocarbidol, N-nitrosothiazolidine-4-carboxylic acid, mono-L-menthol glutaric acid, myristic acid, (2E)-4-hydroxy-5-methyl-2-propane-3(2H)-furanone, (3β,5α,6β,22E,24R)-23-methylergoster-7,22-diene-3,5,6-triol, 1 0-[5]-Ladder-alkyl-decanoic acid, 2,2'-dithiodipyridine, gemmaconone, monogalactosyl glycerol (18:2(9Z,12) / 0:0), josinoside, hydroxymethylcytosine, 1-palmitoyl-2-linoleoyl-sn-glycerol-3-phosphocholine, 5,6-dihydroxyprostaglandin F1a, 3-(but-3-en-1-yn-1-yl)-6-methyl-1,2-dithioene, cholic acid 7-sulfate.

[0085] In one embodiment of this application, the metabolite group comprises icocarbidol, N-nitrosothiazolidine-4-carboxylic acid, mono-L-menthol glutaric acid, myristic acid, (2E)-4-hydroxy-5-methyl-2-propane-3(2H)-furanone, (3β,5α,6β,22E,24R)-23-methylergoster-7,22-diene-3,5,6-triol, 10-[5]-ladderane- It consists of decanoic acid, 2,2'-dithiodipyridine, gemimazone, monogalactosyl glycerol (18:2(9Z,12) / 0:0), jojoba glycoside, hydroxymethylcytosine, 1-palmitoyl-2-linoleoyl-sn-glycerol-3-phosphocholine, 5,6-dihydroxyprostaglandin F1a, 3-(but-3-en-1-yn-1-yl)-6-methyl-1,2-dithioene, and cholic acid 7-sulfate.

[0086] In one embodiment of this application, during the step of calculating age, a formula for calculating age (Y) is pre-stored, which is fitted based on data of metabolites of subjects in an existing database.

[0087] In one embodiment of this application, the formula for calculating age (Y) is a weighted summation formula.

[0088] In one embodiment of this application, the formula is as follows:

[0089]

[0090] Where Y is the calculated age of the subject, βjA is a unitless parameter, b1 is a constant, and X 1A X represents the level of icoracetin in the serum of the subjects. 2A X represents the concentration of N-nitrosothiazolidine-4-carboxylic acid in the serum of the subjects. 3A X represents the serum glutaric acid mono-L-menthol ester content in the subjects. 4A X represents the level of myristic acid in the serum of the subjects. 5A X represents the concentration of (2E)-4-hydroxy-5-methyl-2-propane-3(2H)-furanone in the serum of the subjects. 6A X represents the concentration of (3β,5α,6β,22E,24R)-23-methylergoster-7,22-diene-3,5,6-triol in the serum of the subjects. 7A X represents the content of 10-[5]-ladder-decanoic acid in the serum of the subjects. 8A X represents the concentration of 2,2'-dithiodipyridine in the serum of the subjects. 9A X represents the level of gemcitabine in the serum of the subject. 10A X represents the serum concentration of monogalactosyl glycerol (18:2(9Z,12) / 0:0) in the subjects. 11A X represents the concentration of ginsenosides in the serum of the subjects. 12A X represents the concentration of hydroxymethylcytosine in the serum of the subject. 13A X represents the serum concentration of 1-palmitoyl-2-linoleoyl-sn-glycerol-3-phosphate choline in the subjects. 14A X represents the serum 5,6-dihydroxyprostaglandin F1a level in the subjects. 15A X represents the concentration of 3-(but-3-en-1-yn-1-yl)-6-methyl-1,2-dithioene in the serum of the subjects. 16A The concentration of cholic acid 7-sulfate in the serum of the subjects.

[0091] A confidence interval is an estimated interval for a population parameter constructed from a sample statistic. In statistics, a confidence interval for a probability sample is an interval estimate of a population parameter for that sample. A confidence interval shows the degree to which the true value of this parameter has a certain probability of falling within the range of the measured result. A confidence interval gives the degree of confidence of the measured value of the measured parameter, that is, the "probability" mentioned earlier. This probability is called the confidence level; for example, a confidence space at a confidence level of 0.95 can also be expressed as a 95% confidence interval.

[0092] Therefore, it should be understood that in one embodiment of this application, those skilled in the art can determine the preferred value of b1 and the preferred values ​​of each β value through the aforementioned weighted summation formula. Furthermore, given the existence of confidence intervals, those skilled in the art can reasonably determine the preferred range of b1 and the preferred range of each β value; that is, the value of b1 and the value of each β can arbitrarily take values ​​within their respective confidence intervals. Therefore, in one embodiment of this application, b1 is selected from any value between -6.4 and -6.2, preferably -6.319993; β... 1A β is selected from any value between 13.8 and 14.0, preferably 13.967592; 2A Any value selected from 11.4 to 11.6, preferably 11.568449; β 3A β is selected from any value between -6.4 and -6.2, preferably -6.366745; 4A Any value selected from 5.5 to 5.7, preferably 5.626721; β 5A Any value selected from -4.4 to -4.2, preferably -4.393360; β 6A Any value selected from -4.0 to -3.8, preferably -3.974611; β 7A Any value selected from 3.7 to 3.9, preferably 3.851389; β 8A Any value selected from -3.8 to -3.6, preferably -3.785134; β 9A β is selected from any value between -3.7 and -3.5, preferably -3.600627; 10A β is selected from any value between -3.4 and -3.2, preferably -3.363541; 11A Any value selected from -3.2 to -3.0, preferably -3.169782; β 12A Any value selected from 2.7 to 2.9, preferably 2.890089; β 13A Any value selected from -1.8 to -1.6, preferably -1.716091; β 14Aβ is selected from any value between 1.3 and 1.5, preferably 1.467066; 15A The value is selected from any value between -0.9 and -0.7, preferably -0.842830; β 16A The value is selected from any value between -0.3 and -0.1, preferably -0.265457.

[0093] As described above, the specific details of the steps performed in the method of this application, as well as the acquisition and processing of the subject's metabolite data, can all refer to the steps performed in each module of the model involved in this application.

[0094] In one specific embodiment of this application, the calculated age of the subject in the age calculation module can be expressed as, for example, in months.

[0095] In one specific embodiment of this application, in the step of calculating age, the calculated age of the subject can be expressed as, for example, age in months.

[0096] Example

[0097] The following description, in conjunction with specific embodiments, illustrates the content of this application, but the scope of this application is not limited thereto. Unless otherwise specified, the reagents and instruments used in the following embodiments are all conventional reagents and instruments in the art and can be obtained commercially. The methods used are all conventional experimental methods, and those skilled in the art can undoubtedly implement the described schemes and obtain corresponding results based on the embodiments.

[0098] Main instruments and reagents

[0099] Animal isoflurane, anesthesia machine (MIDMARKMatrxVMR), blood collection needle (Biotopped), PAXgene RNAtube (Bio-Rad).

[0100] Low-temperature high-speed centrifuge (Heraeus, Germany), disposable vacuum blood collection tubes (Hebei Kangweishi Medical Technology Co., Ltd.)

[0101] animal samples

[0102] In this embodiment, a total of 119 SPF-grade SD rats were used: the first batch consisted of 10 one-month-old rats, 20 three-month-old rats (half male and half female), and 10 nine-month-old females; the second batch consisted of 10 one-month-old rats, 20 six-month-old rats (half male and half female); the third batch consisted of 14 one-month-old rats, 20 twelve-month-old rats (half male and half female); and the fourth batch consisted of 5 one-month-old males and 10 nine-month-old males. All animals were provided by the Department of Animal Science, Peking University School of Medicine. The animals were housed in a barrier environment, provided with SPF-grade rat and mouse maintenance feed, and had free access to food and water. The barrier environment temperature was (23±2)℃, and the relative humidity was (55±15)%, with a light-dark cycle every 12 hours. This animal experiment complied with the relevant regulations of the Peking University Animal Welfare and Management Committee.

[0103] Experimental methods

[0104] The selected animals included: the first batch consisted of 6 one-month-old females, 12 three-month-old females (half male and half female), and 6 nine-month-old females; the second batch consisted of 10 one-month-old females, 12 six-month-old females (half male and half female); the third batch consisted of 10 one-month-old females, 12 twelve-month-old females (half male and half female); and the fourth batch consisted of 6 one-month-old males and 6 nine-month-old males. All animals were allowed to grow naturally to the required age for the experiment. After being anesthetized with isoflurane, whole blood was collected via the abdominal aorta for later use. The animals were then euthanized by cervical dislocation.

[0105] 1. Isolation of serum

[0106] Whole blood was collected from rats and centrifuged twice at 3000 rpm for 10 min. The supernatant serum was collected and stored in a cryovial, and then sent to Shanghai Ouyi Biomedical Technology Co., Ltd. for metabolomics analysis.

[0107] 2. Statistical Analysis

[0108] To eliminate the influence of animal batches, each batch of samples was standardized using the mean of the 1-month-old sample data. One male and one female were randomly selected from each age group as the internal validation set, and the remaining samples were used to construct the model. Metabolites measured in each batch were used as independent variables, and univariate linear regression was performed using the Rv4.4 "tidyverse" package with the dependent variable (age). Statistically significant regression coefficients (P < 0.05) were retained, while those without statistical significance were discarded.

[0109] Specifically, the metabolite data matrices from the three batches were first merged, and common metabolites, totaling 675, were identified. Then, a univariate linear regression method was used to analyze the relationship between each metabolite and age. Metabolites with statistically significant regression coefficients were retained for further analysis, while those without statistical significance were removed. The results showed that 88 of the 675 metabolites were retained.

[0110] Subsequently, Lasso regression was performed on the age-related metabolite dataset using the Rv4.4 "glmnet" package to screen variables, and the results are as follows: Figure 1 As shown.

[0111] First, different values ​​of λ and mean squared error (MSE) are obtained through cross-validation. Then, based on the recommended range of the number of variables to be retained, two models are established by selecting the λ corresponding to the minimum MSE and the minimum number of variables.

[0112] First, we attempted to build a model using λ (0.1270906) corresponding to the minimum MSE, retaining 16 variables and their corresponding regression coefficients. The independent variables and regression coefficients of the constructed model are shown in Table 1. The coefficient of determination (R²) is... 2 The value is 0.932676.

[0113] Table 1 lists the metabolite variables included in the Lasso age prediction model constructed using the minimum mean squared error (MSE).

[0114]

[0115] The residual histogram of the model is as follows: Figure 2 As shown, the residuals are uniformly distributed around 0, indicating good homogeneity of variance. The residuals also show good normality. The relationship between residuals and predicted values ​​is shown in the graph below. Figure 3 As shown, the variance homogeneity is good.

[0116] 3. Validation of the Lasso age prediction model based on serum metabolites from SD rats

[0117] Each batch of samples was standardized using the mean of the 1-month-old sample data. Then, one male and one female were randomly selected from each month of age as the external validation set for model validation. The data processing method for the external validation set was the same as that for the dataset used to build the model.

[0118] The predicted age and actual age obtained from the external validation set Figure 4 As shown, R2: 0.9374075.

[0119] This application utilizes a Lasso model constructed from serum metabolites of SD rats, which can accurately predict the actual age of rats. The fitting equation for this model is as follows, where Y represents the rat's age in months, and the independent variable is the metabolite cid:

[0120] Y=13.967592*(207690)+11.568449*(63414)+(-6.366745)*(256141)+

[0121] 5.626721*(560)+(-4.393360)*(256800)+(-3.974611)*(57845)+3.851389*(2535)+(-3.785134)*(200787)+(-3.600627)*(59708)+(-3.363541)*(18021)+(-3.169782)*(34325)+2.890089*(207638)+(-1.716091)*(27841)+1.467066*(51835)

[0122] +(-0.842830)*(10145)+(-0.265457)*(46701)-6.319993

[0123] This application uses serum metabolites from SD rats as independent variables and age in months as the dependent variable to construct a Lasso model. The model's coefficient of determination R0 is... 2 =0.932676, the regularization parameter λ corresponding to the minimum MSE is 0.1270906, the residuals follow a normal distribution, and the model has strong predictive ability. The model's R² on the outer validation set is 0.9374075, indicating good predictive ability, high stability, and generalization ability. In summary, this model can accurately predict the actual age of SD rats and can be effectively extended to new and unseen data, demonstrating strong practical application value.

[0124] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A metabolome for predicting age, wherein, It includes one or more metabolites selected from the following group: icocarbidol, N-nitrosothiazolidine-4-carboxylic acid, mono-L-menthol glutaric acid, myristic acid, (2E)-4-hydroxy-5-methyl-2-propane-3(2H)-furanone, (3β,5α,6β,22E,24R)-23-methylergoster-7,22-diene-3,5,6-triol, 10-[5]-titanium Alkyl-decanoic acid, 2,2'-dithiodipyridine, gemimazone, monogalactosyl glycerol (18:2(9Z,12) / 0:0), jojoba glycoside, hydroxymethylcytosine, 1-palmitoyl-2-linoleoyl-sn-glycerol-3-phosphocholine, 5,6-dihydroxyprostaglandin F1a, 3-(but-3-en-1-yn-1-yl)-6-methyl-1,2-dithioene, cholic acid 7-sulfate.

2. A system for predicting age, comprising: The data acquisition module is used to acquire data on the metabolites of the subjects; A data processing module is used to perform data standardization processing on the data information acquired by the data acquisition module. as well as An age calculation module is used to calculate the age of the subject by performing calculations on the data information processed in the data processing module. The subjects were rats.

3. The system according to claim 2, wherein: In the data acquisition module, the collected metabolite data refers to the types and / or amounts of metabolites detected on any day after the subject's birth.

4. The system according to claim 2, wherein, The metabolite is selected from any one or more metabolites from the group consisting of: icocarbidol, N-nitrosothiazolidine-4-carboxylic acid, mono-L-menthol glutaric acid, myristic acid, (2E)-4-hydroxy-5-methyl-2-propane-3(2H)-furanone, (3β,5α,6β,22E,24R)-23-methylergoster-7,22-diene-3,5,6-triol, 10-[5]- Ladder-type alkyl-decanoic acid, 2,2'-dithiodipyridine, gemimazone, monogalactosyl glycerol (18:2(9Z,12) / 0:0), jojoba glycoside, hydroxymethylcytosine, 1-palmitoyl-2-linoleoyl-sn-glycerol-3-phosphocholine, 5,6-dihydroxyprostaglandin F1a, 3-(but-3-en-1-yn-1-yl)-6-methyl-1,2-dithioene, cholic acid 7-sulfate.

5. The system according to claim 2, wherein: The age calculation module contains a pre-stored formula for calculating age (Y) based on data fitted from the metabolites of subjects in an existing database.

6. The system according to claim 5, wherein, The formula used to calculate age (Y) is a weighted summation formula.

7. The system according to claim 6, wherein: The formula is as follows: Formula 1 Where Y is the calculated age of the subject, βjA is a unitless parameter, b1 is a constant, and X 1A X represents the level of icoracetin in the serum of the subjects. 2A X represents the concentration of N-nitrosothiazolidine-4-carboxylic acid in the serum of the subjects. 3A X represents the serum glutaric acid mono-L-menthol ester content in the subjects. 4A X represents the level of myristic acid in the serum of the subjects. 5A X represents the concentration of (2E)-4-hydroxy-5-methyl-2-propane-3(2H)-furanone in the serum of the subjects. 6A X represents the concentration of (3β,5α,6β,22E,24R)-23-methylergoster-7,22-diene-3,5,6-triol in the serum of the subjects. 7A X represents the content of 10-[5]-ladder-decanoic acid in the serum of the subjects. 8A X represents the concentration of 2,2'-dithiodipyridine in the serum of the subjects. 9A X represents the level of gemcitabine in the serum of the subject. 10A X represents the serum concentration of monogalactosyl glycerol (18:2(9Z,12) / 0:0) in the subjects. 11A X represents the concentration of ginsenosides in the serum of the subjects. 12A X represents the concentration of hydroxymethylcytosine in the serum of the subject. 13A X represents the serum concentration of 1-palmitoyl-2-linoleoyl-sn-glycerol-3-phosphate choline in the subjects. 14A X represents the serum 5,6-dihydroxyprostaglandin F1a level in the subjects. 15A X represents the concentration of 3-(but-3-en-1-yn-1-yl)-6-methyl-1,2-dithioene in the serum of the subjects. 16A The concentration of cholic acid 7-sulfate in the serum of the subjects.

8. The system according to claim 7, wherein: b1 is selected from any value from -6.4 to -6.2, preferably -6.319993; β 1A The value is selected from any value between 13.8 and 14.0, preferably 13.967592; β 2A Any value selected from 11.4 to 11.6, preferably 11.568449; β 3A Any value selected from -6.4 to -6.2, preferably -6.366745; β 4A Any value selected from 5.5 to 5.72, preferably 5.626721; β 5A Any value selected from -4.4 to -4.2, preferably -4.393360; β 6A Any value selected from -4.0 to -3.8, preferably -3.974611; β 7A Any value selected from 3.7 to 3.9, preferably 3.851389; β 8A Any value selected from -3.8 to -3.6, preferably -3.785134; β 9A Any value selected from -3.7 to -3.5, preferably -3.600627; β 10A Any value selected from -3.4 to -3.2, preferably -3.363541; β 11A Any value selected from -3.2 to -3.0, preferably -3.169782; β 12A Any value selected from 2.7 to 2.9, preferably 2.890089; β 13A Any value selected from -1.8 to -1.6, preferably -1.716091; β 14A Any value selected from 1.3 to 1.5, preferably 1.467066; β 15A The value is selected from any value between -0.9 and -0.7, preferably -0.842830; β 16A The value is selected from any value between -0.3 and -0.1, preferably -0.265457.

9. A method for predicting age, comprising: The data acquisition step involves obtaining data on the subject's metabolites; The data processing step standardizes the data information acquired in the data acquisition step; and The step of calculating age involves calculating the age of the subject by processing the data information processed in the data processing step. in, The subjects were rats.

10. The method according to claim 9, wherein: In the data acquisition step, the collected metabolite data refers to the types and / or amounts of metabolites detected on any day after the subject's birth.