Weak subtype identification method and device based on machine learning and metabonomics characteristics
Through machine learning and metabolomic characteristics, the identification of new subtypes of weak patients has solved the problem of the inability to accurately identify heterogeneity and chronic disease risks in the prior art, and achieved accurate risk stratification and prognosis assessment of weak patients.
Patent Information
- Application Number
- CN202510337072.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
AI Technical Summary
The existing technical means cannot accurately and easily identify the heterogeneity of weak patients and their chronic disease risks, resulting in large differences in prognosis of weak patients.
Using a method based on machine learning and metabolomics characteristics, metabolic characteristics were analyzed by the CatBoost algorithm, the top 11 key metabolites with the highest SHAP value were selected, combined with principal component analysis and K-mean clustering, weak participants were stratified into four subtypes, and their association with chronic disease and all-cause mortality was evaluated using Kaplan–Meier curves and multivariate Cox proportional hazards model.
Identify the novel subtype of weakness and evaluate its application value in chronic disease management and prognosis, provide a new basis for clinical intervention and health management, and improve the accuracy of risk stratification for weak patients.
Smart Images

Figure CN120260697A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of biomedical technologies, and in particular, to a method and device for identifying frailty subtypes based on machine learning and metabolomic features. Background Art
[0002] Frailty affects approximately 24% of the global population aged 50 and above. It is a characteristic of aging and is associated with multiple systems and organs. Frailty is a risk factor for chronic diseases such as cardiovascular diseases, cerebrovascular diseases, chronic obstructive pulmonary disease, kidney diseases, liver cancer, and colorectal cancer. Therefore, many clinical guidelines advocate for the routine monitoring of frailty.
[0003] Currently, there are two main classical models for defining frailty: the Frailty Phenotype (FP) and the Frailty Index (FI). The FP was proposed by Fried et al. in 2001 and its characteristics include unintentional weight loss, decreased grip strength, slowed walking speed, poor endurance, low energy expenditure, and reduced physical activity. In contrast, the FI is based on the cumulative deficit model and was proposed by Rockwood et al., reflecting age-related cumulative health deficits. Multiple studies have shown that the FI can more accurately predict the risk of adverse events such as falls and hospitalizations than the FP. However, there are significant prognostic differences among frail patients defined by the FI, mainly because frailty itself is highly heterogeneous, with significant differences in its etiology, pathophysiological mechanisms, clinical manifestations, and prognosis. Current technical means such as routine physical examinations or single biomarker detections cannot accurately and simply identify the heterogeneity of frail patients and their chronic disease risks. Therefore, there is an urgent need to develop new technologies or methods to perform more precise risk stratification for frail patients to improve the prognosis of high-risk populations.
[0004] Frailty is associated with disorders of amino acid and fatty acid metabolism. Although studies have shown that metabolic disorders in frail patients affect the prognosis of chronic diseases such as type 2 diabetes, cardiovascular diseases, and metabolic dysfunction-related fatty liver disease, there is currently no consensus on the management and application of blood metabolites in this population. Therefore, identifying novel frailty subtypes based on metabolomic features may contribute to the management of frail patients. Summary of the Invention
[0005] In view of this, it is necessary to provide a method for identifying frailty subtypes based on machine learning and metabolomic features, which can identify novel frailty subtypes based on machine learning and metabolomic features, and evaluate their application value in the management and prognosis of chronic diseases, providing a new basis for clinical intervention and health management.
[0006] In a first aspect, an embodiment of the present application provides a method for identifying frailty subtypes based on machine learning and metabolomic features, the method comprising the following steps:
[0007] (1) Data acquisition and preprocessing steps: Obtain participant data from the local database, construct a frailty index FI based on 49 items to evaluate the degree of frailty and exclude specific ineligible participants. At the same time, analyze 251 metabolites and perform standard normalization processing; and divide the data into a training set and a test set.
[0008] (2) Frailty subtype identification steps: Use the CatBoost algorithm to analyze metabolic characteristics and select the top 11 key metabolites with the highest SHAP value rankings. Stratify frail participants into four subtypes through principal component analysis and K-means clustering analysis.
[0009] (3) Association assessment steps for frailty subtypes with chronic diseases and all-cause mortality: Use Kaplan–Meier curves to show the differences in cumulative incidence of the four subtypes and non-frail participants in 13 chronic diseases and all-cause mortality, and use a multivariable Cox proportional hazards model to test the association of the new frailty subtypes with these outcomes, including adjusting for age and gender factors.
[0010] Optionally, in an implementation manner of the first aspect of the present invention, the step of using the CatBoost algorithm to analyze metabolic characteristics and select the top 11 key metabolites with the highest SHAP value rankings includes:
[0011] Establish an initial frailty subtype prediction and feature analysis model integrating CatBoost and SHAP, and the initial frailty subtype prediction and feature analysis model is established based on all variables.
[0012] Train the initial frailty subtype prediction and feature analysis model through the training set, analyze the importance of input variables for feature selection, continuously eliminate the variable with the lowest importance ranking, retain the input variables with predictive ability for frailty subtypes, and obtain test results through the test set, and construct a frailty subtype prediction and feature analysis model based on highly correlated variables.
[0013] Introduce the SHAP model to perform interpretability analysis on the CatBoost prediction model, and perform weighted summation on the SHAP values of each feature. The formula is:
[0014]
[0015] where f(x) represents the prediction result of the SHAP model, and φ0 represents the prediction benchmark value. represents the cumulative of SHAP values from the i-th to the M-th feature.
[0016] The SHAP value calculation formula for the feature X of the CatBoost prediction model is: i
[0017]
[0018] where CatBoost(X i ) represents the prediction result of the CatBoost model, N represents the total number of features, represents the set of features except X i , T is a subset of the set of features except X i , f(T) represents the SHAP model prediction in T, and f(T∪{i}) represents the SHAP model prediction in T∪{i}.
[0019] Optionally, in an implementation manner of the first aspect of the present invention, the metabolite includes clinically verified indicators such as cholesterol, fatty acids, amino acids, and inflammatory markers, as well as emerging biomarkers such as lipoprotein subclasses.
[0020] Optionally, in an implementation manner of the first aspect of the present invention, the key metabolite is based on the top 11 metabolic features ranked by SHAP value, including:
[0021] GlycA: Glycoprotein acetylate;
[0022] LA / FA: Percentage of linoleic acid in total fatty acids;
[0023] MUFA / FA: Percentage of monounsaturated fatty acids in total fatty acids;
[0024] Alb: Albumin;
[0025] Val: Valine;
[0026] DHA / FA: Percentage of docosahexaenoic acid in total fatty acids;
[0027] PUFA / MUFA: Ratio of polyunsaturated fatty acids to monounsaturated fatty acids;
[0028] LA: Linoleic acid;
[0029] XS-VLDL-FC_%: Percentage of free cholesterol in total lipids in very low density lipoproteins;
[0030] L-HDL-PL_%: Percentage of phospholipids in total lipids in large high density lipoproteins;
[0031] XS-VLDL-CE_%: Percentage of cholesterol esters in total lipids in very low density lipoproteins.
[0032] Optionally, in an implementation of the first aspect of the present invention, when the multivariate Cox proportional hazards model examines the association between the new frailty subtypes and the outcomes, the results show that compared with the participants in subtype I, individuals in subtypes III and IV show significantly higher risks in all outcomes; the outcomes include coronary artery disease, heart failure, major adverse cardiovascular events (MACE), myocardial infarction (MI), type 2 diabetes, metabolic dysfunction-associated fatty liver disease (MASLD), chronic obstructive pulmonary disease (COPD), severe liver disease (SLD), peripheral artery disease (PAD), end-stage renal disease (ESRD), kidney cancer, lung cancer, abdominal aortic aneurysm (AAA), and all-cause mortality.
[0033] Optionally, in an implementation of the first aspect of the present invention, in the frailty subtype identification, when the CatBoost algorithm analyzes the metabolic characteristics, the SHAP interpreter is applied to the best training model, and the SHAP values are calculated to rank the top 20 most influential features, and the top 11 metabolites with the highest SHAP values are selected for subsequent clustering.
[0034] Optionally, in an implementation of the first aspect of the present invention, in the frailty subtype identification step, when performing principal component analysis on the selected 11 metabolic features, principal component 1 and principal component 2 explain 44.54% and 16.63% of the variance, respectively.
[0035] In a second aspect, an embodiment of the present application provides a frailty subtype identification system based on machine learning and metabolomic features, which is applied to the frailty subtype identification method based on machine learning and metabolomic features as described in the first aspect, and includes:
[0036] Data acquisition and preprocessing module: Obtain participant data from a local database, construct a frailty index (FI) according to 49 items to evaluate the frailty level and exclude specific ineligible participants, and at the same time analyze 251 metabolites and perform standard normalization processing; and divide the data into a training set and a test set;
[0037] Frailty subtype identification module: Use the CatBoost algorithm to analyze the metabolic characteristics and select the top 11 key metabolites with the highest ranked SHAP values, and stratify the frail participants into four subtypes through principal component analysis and K-means clustering analysis;
[0038] Frailty subtype and chronic disease and all-cause mortality association evaluation module: Use the Kaplan–Meier curve to show the cumulative incidence differences of the four subtypes and non-frail participants in 13 chronic diseases and all-cause mortality, and use the multivariate Cox proportional hazards model to test the association between the new frailty subtypes and these outcomes, including adjusting for age and gender factors.
[0039] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0040] Processor;
[0041] A memory for storing instructions executable by the processor;
[0042] Wherein, when the processor is configured to execute the instructions, it implements the method for identifying frailty subtypes based on machine learning and metabolomics features as described in the first aspect.
[0043] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a program that instructs a device to execute the method for identifying frailty subtypes based on machine learning and metabolomics features as described in the first aspect.
[0044] In the technical solution provided by the present invention, a method and device for identifying frailty subtypes based on machine learning and metabolomics features are provided, including data acquisition and preprocessing steps: obtaining participant data from a local database, constructing a frailty index FI according to 49 items to evaluate the frailty degree and excluding specific ineligible participants, and at the same time analyzing 251 metabolites and performing standard normalization processing; and dividing the data into a training set and a test set; frailty subtype identification step: analyzing metabolic features using the CatBoost algorithm and selecting the top 11 key metabolites with the highest SHAP value ranking, and stratifying frail participants into four subtypes through principal component analysis and K-means clustering analysis; frailty subtype and chronic disease and all-cause mortality association evaluation step: using the Kaplan–Meier curve to show the cumulative incidence differences of the four subtypes and non-frail participants in 13 chronic diseases and all-cause mortality, and using a multivariable Cox proportional hazards model to test the association of the new frailty subtypes with these outcomes, including adjusting for age and gender factors. Based on machine learning and metabolomics features, new frailty subtypes are identified and their application value in chronic disease management and prognosis is evaluated, providing a new basis for clinical intervention and health management. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a schematic flowchart of a method for identifying frailty subtypes based on machine learning and metabolomics features provided by an embodiment of the present application.
[0046] Figure 2 It is a schematic flowchart of the CatBoost algorithm provided by an embodiment of the present application.
[0047] Figure 3 It is the key frailty features identified by using machine learning provided by an embodiment of the present application.
[0048] Figure 4 A - B is for identifying the optimal number of clusters provided by an embodiment of the present application.
[0049] Figure 5The clustering result graph of frail participants provided by an embodiment of the present application.
[0050] Figure 6 a - e are radar graphs of the characteristics of a new frail subtype provided by an embodiment of the present application.
[0051] Figure 7 The Kaplan - Meier curve graph of 13 chronic diseases and all - cause mortality provided by an embodiment of the present application.
[0052] Figure 8 The multivariate Cox proportional hazards model graph classified according to the results of the new frail subtype provided by an embodiment of the present application, adjusted for age, sex, race, current smoking, current alcohol consumption, systolic blood pressure (SBP), and diastolic blood pressure (DBP).
[0053] Figure 9 The schematic diagram of the module of a frail subtype recognition system based on machine learning and metabolomic characteristics provided by an embodiment of the present application.
[0054] Figure 10 The schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0055] Now, various exemplary implementation manners of the present invention will be described in detail. This detailed description should not be considered as a limitation of the present invention, but rather as a more detailed description of certain aspects, characteristics, and implementation schemes of the present invention.
[0056] It should be understood that the terms described in the present invention are only used to describe specific implementation manners and are not used to limit the present invention. Additionally, for the numerical ranges in the present invention, it should be understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. Any intermediate value within any stated value or stated range, as well as each smaller range between any other stated value or intermediate value within the stated range, is also included in the present invention. The upper and lower limits of these smaller ranges can be independently included or excluded from the range.
[0057] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the present invention pertains. Although the present invention only describes preferred methods and materials, any methods and materials similar or equivalent to those described herein can also be used in the implementation or testing of the present invention. All documents mentioned in this specification are incorporated by reference to disclose and describe the methods and / or materials related to the documents. In case of conflict with any incorporated document, the content of this specification shall prevail.
[0058] Without departing from the scope or spirit of the present invention, various modifications and variations can be made to the specific embodiments of the present invention specification, which are obvious to those skilled in the art. Other embodiments obtained from the present invention specification are obvious to those skilled in the art. The present invention specification and examples are merely exemplary.
[0059] Regarding the use of "comprising", "including", "having", "containing", etc. in this text, they are all open-ended terms, meaning including but not limited to.
[0060] The meanings represented by each English abbreviation in the embodiments of the present invention are as follows:
[0061] GlycA: Glycoprotein acetylation;
[0062] LA / FA: Percentage of linoleic acid in total fatty acids;
[0063] MUFA / FA: Percentage of monounsaturated fatty acids in total fatty acids;
[0064] Alb: Albumin;
[0065] Val: Valine;
[0066] DHA / FA: Percentage of docosahexaenoic acid in total fatty acids;
[0067] PUFA / MUFA: Ratio of polyunsaturated fatty acids to monounsaturated fatty acids;
[0068] LA: Linoleic acid;
[0069] XS-VLDL-FC_%: Percentage of free cholesterol in total lipids in very low density lipoproteins;
[0070] L-HDL-PL_%: Percentage of phospholipids in total lipids in large high density lipoproteins;
[0071] XS-VLDL-CE_%: Percentage of cholesterol esters in total lipids in very low density lipoproteins;
[0072] Tyr: Tyrosine;
[0073] L-HDL-FC_%: Percentage of cholesterol in total lipids in large low density lipoproteins;
[0074] FAw6: Omega-6 fatty acids;
[0075] VLDLsize: Average diameter of very low density lipoprotein particles;
[0076] UnSat: Estimated fatty acid unsaturation;
[0077] XXL-VLDL-CE: Cholesterol esters in chylomicrons and very large-sized VLDL;
[0078] L-LDL-TG_%: Percentage of triglycerides in large low-density lipoproteins among total lipids;
[0079] S-LDL-CE_%: Percentage of cholesterol esters in small low-density lipoproteins among total lipids;
[0080] BCAAs: Total concentration of branched-chain amino acids (leucine + isoleucine + valine);
[0081] MASLD: Metabolic dysfunction-associated fatty liver disease
[0082] MI: Myocardial infarction
[0083] ESRD: End-stage renal disease;
[0084] SLD: Severe liver disease;
[0085] MACE: Major adverse cardiovascular events;
[0086] PAD: Peripheral artery disease
[0087] AAA: Abdominal aortic aneurysm
[0088] SBP: Mean systolic blood pressure;
[0089] DBP: Mean diastolic blood pressure;
[0090] HR: Hazard ratio;
[0091] Example 1
[0092] Figure 1 It is a schematic flow chart of a frailty subtype recognition method based on machine learning and metabolomics features provided by an embodiment of this application. As Figure 1 shown, the frailty subtype recognition method based on machine learning and metabolomics features includes the following steps:
[0093] 1. Data acquisition and preprocessing
[0094] As Figure 1 shown, 160,407 participants in the local database are selected, and the participants are evaluated for frailty, and the frailty is evaluated by the frailty index (FI).
[0095] The frailty index (FI) reflects the accumulation of health deficits. This index is based on indicators in multiple physiological and psychological domains, including symptoms, diagnosed diseases, and dysfunctions. From the UK Biobank, 49 items were selected to construct a frailty index. The health deficits of each participant were evaluated, and the total number of health deficits was divided by 49 to calculate the FI, which ranges from 0 to 1. Participants with FI ≤ 0.10 were classified as non-frail, those with FI between 0.10 - 0.21 were classified as pre-frail, and those with FI > 0.21 were classified as frail.
[0096] Among the included participants, 28,196 individuals were classified as frail. Frail participants were predominantly older, more likely to be female, and less likely to be white. They also had higher smoking rates, lower alcohol consumption, and higher levels of systolic blood pressure (SBP), diastolic blood pressure (DBP), body mass index (BMI), alanine aminotransferase (ALT), aspartate aminotransferase (AST), alkaline phosphatase (ALP), triglycerides, and glycated hemoglobin (HbA1c).
[0097] The demographic and clinical characteristics of the participants are shown in Table 1.
[0098] Table 1. Baseline characteristics of the study population
[0099]
[0100]
[0101] 2. Identification of new frailty subtypes
[0102] (1) Analyze 251 metabolic biomarkers from nuclear magnetic resonance (NMR) metabolomics data in the local database. These biomarkers include clinically validated indicators such as cholesterol, fatty acids, amino acids, and inflammatory markers, as well as emerging biomarkers such as lipoprotein subclasses.
[0103] Perform standard Z-score normalization on these 251 metabolic data points. Use the CatBoost algorithm to analyze these features. Apply the SHAP interpreter to the best trained model to calculate SHAP values to rank the top 20 most influential features ( Figure 3 ). GlycA is the most important metabolite, followed by LA / FA.
[0104] Figure 2 Schematic diagram of the CatBoost algorithm process provided for an embodiment of this application. As Figure 2As shown, the CatBoost algorithm is used to analyze metabolic features. Specifically, the CatBoost algorithm is used to analyze metabolic features and select the top 11 key metabolites with the highest SHAP value rankings, including:
[0105] An initial frailty subtype prediction and feature analysis model integrating CatBoost and SHAP is established, and the initial frailty subtype prediction and feature analysis model is established based on all variables;
[0106] The initial frailty subtype prediction and feature analysis model is trained using a training set, the importance of input variables is analyzed for feature selection, variables with the lowest importance ranking are continuously removed, input variables with predictive ability for frailty subtypes are retained, and test results are obtained through testing with a test set, and a frailty subtype prediction and feature analysis model based on highly correlated variables is constructed;
[0107] The SHAP model is introduced to perform interpretability analysis on the CatBoost prediction model, and the SHAP values of each feature are weighted and summed. The formula is:
[0108]
[0109] where f(x) represents the prediction result of the SHAP model, φ0 represents the prediction benchmark value, represents the cumulative sum of SHAP values from the i-th to the M-th feature;
[0110] The feature X of the CatBoost prediction model i The calculation formula for the SHAP value is:
[0111]
[0112] where CatBoost(X i ) represents the prediction result of the CatBoost model, N represents the total number of features, represents the set of features except X i and T is a subset of the set of features except X i and f(T) represents the SHAP model prediction in T, and f(T∪{i}) represents the SHAP model prediction in T∪{i}.
[0113] Furthermore, the top 11 metabolic features ranked by SHAP are selected for principal component analysis, where principal component 1 (PC1) and principal component 2 (PC2) explain 44.54% and 16.63% of the variance, respectively. Based on PC1 and PC2, K-means clustering analysis is performed. The optimal number of clusters is determined to be four through the elbow method and silhouette coefficient ( Figure 4 ) Subsequently, frail participants are further stratified into four new subtypes for subsequent analysis (Figure 5 )。
[0114] (2) Summarize the demographics, anthropometrics, and clinical characteristics of the four identified frailty subtypes. Participants in subtypes III and IV were older, predominantly male, and more likely to be white. These subtypes were also associated with lower systolic blood pressure (SBP) and diastolic blood pressure (DBP), as well as higher body mass index (BMI), alanine aminotransferase (ALT), and aspartate aminotransferase (AST) levels.
[0115] The demographics, anthropometrics, and clinical characteristics of the four frailty subtypes are shown in Table 2
[0116] Table 2. Baseline characteristics of the new frailty subtypes
[0117]
[0118]
[0119]
[0120] (3) Summarize the metabolite characteristics of the four identified frailty subtypes.
[0121] The radar chart of the metabolite characteristics of the four frailty subtypes is shown in Figure 6 。Specifically:
[0122] Subtype I: This subtype of population showed extremely low glycoprotein acetylation and valine levels, the highest percentage of linoleic acid in total fatty acids, and a relatively high percentage of docosahexaenoic acid in total fatty acids. They also had a relatively high ratio of polyunsaturated fatty acids to monounsaturated fatty acids. Subtype I was named low GlycA and Val of frailty (LGVF).
[0123] Subtype II: This subtype was characterized by the highest albumin and linoleic acid concentrations and was thus named high Alb and LA of frailty (HALF).
[0124] Subtype III: This subtype showed the lowest albumin and linoleic acid levels and was named low LA and Alb of frailty (LALF).
[0125] Subtype IV: This subtype shows a higher level of glycoprotein acetylation, an elevated valine concentration, a higher percentage of monounsaturated fatty acids in total fatty acids, and a significantly lower percentage of linoleic acid in total fatty acids. It is named high GlycA and Val of frailty (HGVF).
[0126] 3. Establish Kaplan–Meier survival curves and Cox proportional hazards models for prognostic evaluation
[0127] (1) Prognostic evaluation by Kaplan–Meier curves
[0128] The outcomes included 13 chronic diseases and all-cause mortality. The chronic diseases were: coronary artery disease, heart failure, major adverse cardiovascular events (MACE), myocardial infarction (MI), type II diabetes, metabolic dysfunction-associated fatty liver disease (MASLD), chronic obstructive pulmonary disease (COPD), severe liver disease (SLD), peripheral artery disease (PAD), end-stage renal disease (ESRD), renal cancer, lung cancer, and abdominal aortic aneurysm (AAA).
[0129] At a median follow-up time of 13.8 years, the Kaplan–Meier curves showed the differences in cumulative incidences of 13 chronic diseases and all-cause mortality between the four new frailty subtypes and non-frail participants (see Figure 7 ). The highest cumulative incidences were observed for coronary artery disease, type 2 diabetes, chronic obstructive pulmonary disease (COPD), and all-cause mortality. The cumulative incidence trajectories of subtypes I and II were similar, and those of subtypes III and IV were similar. The cumulative incidence of non-frail participants was the lowest among all outcomes; in contrast, the cumulative risks of chronic diseases for subtypes III and IV were the highest.
[0130] (2) Risk assessment of the outcomes of the new frailty subtypes using a multivariable Cox proportional hazards model A multivariable Cox proportional hazards model was used to examine the association between the new frailty subtypes and the outcomes. After adjusting for age, gender, race, current smoking and alcohol use, systolic blood pressure (SBP), and diastolic blood pressure (DBP), most of the outcomes remained significant (see Figure 8 ). Among all outcomes, the non-frail group always had the lowest risk. Compared with the participants in subtype I, individuals in subtypes III and IV showed significantly higher risks in all outcomes.
[0131] Specifically, as shown in Table 3:
[0132] Table 3. Multivariate analysis of non-frail participants and four new frailty subtypes for chronic diseases and all-cause mortality
[0133]
[0134]
[0135]
[0136] As can be seen from the data in the table, the chronic disease risks of subtypes III and IV are higher than those of subtypes I and II.
[0137] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
[0138] Embodiment 2
[0139] As Figure 9 shown, the present application provides a frailty subtype identification system based on machine learning and metabolomic features, which is applied to the frailty subtype identification method based on machine learning and metabolomic features as described in Embodiment 1, and includes: a data acquisition and preprocessing module 11, a frailty subtype identification module 12, and a frailty subtype and chronic disease and all-cause mortality association evaluation module 13.
[0140] It can be understood that in this embodiment, the data acquisition and preprocessing module 11 is used to obtain participant data from the local database, construct a frailty index FI according to 49 items to evaluate the frailty degree and exclude specific unqualified participants, and at the same time analyze 251 metabolites and perform standard normalization processing; and divide the data into a training set and a test set.
[0141] It can be understood that in this embodiment, the frailty subtype identification module 12 is used to analyze the metabolic characteristics with the CatBoost algorithm and select the top 11 key metabolites with the highest SHAP value ranking, and stratify the frail participants into four subtypes through principal component analysis and K-means clustering analysis.
[0142] It can be understood that in this embodiment, the frailty subtype and chronic disease and all-cause mortality association evaluation module 13 is used to display the cumulative incidence differences of the four subtypes and non-frail participants in 13 chronic diseases and all-cause mortality with the Kaplan–Meier curve, and use the multivariate Cox proportional hazards model to test the association between the new frailty subtypes and these results, including adjusting for age and gender factors.
[0143] Figure 10 is an electronic device provided by an embodiment of the present application. As Figure 9 shown, the electronic device at least includes the following parts: a processor 101, a memory 100, a communication interface 103, and a bus 102.
[0144] In an embodiment of the present application, the memory 100 is used to store executable instructions for the processor 101. When the processor 101 is configured to execute the instructions, it implements the frailty subtype identification method based on machine learning and metabolomics features as shown in Figure 1 the following.
[0145] In an embodiment of the present application, a computer-readable storage medium includes instructions that direct a device to execute the method of the first aspect. For example, the instructions direct the device to execute the frailty subtype identification method based on machine learning and metabolomics features shown in the process steps in Figure 1 the following.
[0146] In the electronic device involved in an embodiment of the present application, the program that operates can be a program that controls a central processing unit (CPU) or the like to implement the functions of the above-described embodiments related to a solution of the present invention (a program that causes a computer to function). Then, the information processed by these devices is temporarily stored in a random access memory (RAM) during its processing, and then stored in various ROMs such as a read-only memory (FlashROM), a hard disk drive (HDD), etc. It is read, corrected, and written by the CPU as needed.
[0147] It should be noted that a part of the electronic device of the above-described embodiment can also be implemented by a computer. In this case, the program for implementing the control function can be recorded on a computer-readable recording medium, and is implemented by reading the program recorded on the recording medium into the computer and executing it.
[0148] It should be noted that the "computer" mentioned here refers to a computer built into an electronic device, which is a computer including hardware such as an OS and peripheral devices. In addition, the "computer-readable recording medium" refers to removable media such as a floppy disk, a magneto-optical disk, a ROM, a CD-ROM, and a storage device such as a hard disk built into a computer.
[0149] Moreover, the "computer-readable recording medium" can include: a medium that dynamically stores a program for a short time, such as a communication line when sending a program via a network such as the Internet or a communication line such as a telephone line; a medium that stores a program for a fixed time, such as a volatile memory inside a computer of a server or a client in this case. In addition, the above program can be a program for implementing a part of the above functions, and can also be a program that can implement the above functions by combining with a program already recorded in a computer.
[0150] In addition, the electronic device in the above-described embodiment can also be implemented as an aggregate (device group) composed of multiple devices. Each device constituting the device group can have some or all of the functions or function blocks of the electronic device in the above-described embodiment. As the device group, it is sufficient to have all the functions or function blocks of the electronic device.
[0151] Those of ordinary skill in the art should recognize that the above embodiments are only used to illustrate the present application and are not intended to limit the present application. As long as appropriate changes and variations made to the above embodiments fall within the scope of the spirit of the present application, they fall within the scope of protection required by the present application.
Claims
1. A method for identifying frailty subtypes based on machine learning and metabolomic features, characterized in that, The method includes the following steps: (1) Data acquisition and preprocessing step: Obtain participant data from a local database, construct a frailty index (FI) based on 49 items to evaluate the degree of frailty and exclude specific ineligible participants. At the same time, analyze 251 metabolites and perform standard normalization processing; and divide the data into a training set and a test set; (2) Frailty subtype identification step: Use the CatBoost algorithm to analyze metabolic characteristics and select the top 11 key metabolites with the highest SHAP value rankings. Stratify frail participants into four subtypes through principal component analysis and K-means clustering analysis; (3) Association evaluation step of frailty subtypes with chronic diseases and all-cause mortality: Use Kaplan–Meier curves to show the differences in cumulative incidence of the four subtypes and non-frail participants in 13 chronic diseases and all-cause mortality, and use a multivariable Cox proportional hazards model to test the association of the new frailty subtypes with these outcomes, including adjusting for age and gender factors.
2. The frailty subtype identification method based on machine learning and metabolomics features according to claim 1, wherein The use of the CatBoost algorithm to analyze metabolic characteristics and select the top 11 key metabolites with the highest SHAP value rankings includes: Establish an initial frailty subtype prediction and feature analysis model integrating CatBoost and SHAP, and the initial frailty subtype prediction and feature analysis model is established based on all variables; Train the initial frailty subtype prediction and feature analysis model through the training set, analyze the importance of input variables for feature selection, continuously eliminate the variable with the lowest importance ranking, retain the input variables with predictive ability for frailty subtypes, and obtain test results through the test set, and construct a frailty subtype prediction and feature analysis model based on highly correlated variables; Introduce the SHAP model to perform interpretability analysis on the CatBoost prediction model, and perform weighted summation on the SHAP values of each feature. The formula is: Among them, f(x) represents the prediction result of the SHAP model, and φ0 represents the prediction benchmark value. represents the cumulative SHAP value from the i-th to the M-th feature; Feature X of the CatBoost prediction model i The calculation formula for the SHAP value is as follows: Among them, CatBoost(X i ) represents the prediction result of the CatBoost model, N represents the total number of features, represents the feature set except X i and T is a subset of the feature set except X i . f(T) represents the SHAP model prediction in T, and f(T∪{i}) represents the SHAP model prediction in T∪{i}.
3. The frailty subtype identification method based on machine learning and metabolomics features according to claim 1, wherein The metabolites include clinically verified indicators such as cholesterol, fatty acids, amino acids, and inflammatory markers, as well as emerging biomarkers such as lipoprotein subclasses.
4. The method for identifying frailty subtypes based on machine learning and metabolomic features according to claim 2, wherein The key metabolites are based on the top 11 metabolic characteristics ranked by SHAP values and include: GlycA: Glycoprotein acetylate; LA / FA: Percentage of linoleic acid in total fatty acids; MUFA / FA: Percentage of monounsaturated fatty acids in total fatty acids; Alb: Albumin; Val: Valine; DHA / FA: Percentage of docosahexaenoic acid in total fatty acids; PUFA / MUFA: Ratio of polyunsaturated fatty acids to monounsaturated fatty acids; LA: Linoleic acid; XS-VLDL-FC_%: Percentage of free cholesterol in very low density lipoproteins in total lipids; L-HDL-PL_%: Percentage of phospholipids in large high density lipoproteins in total lipids; XS-VLDL-CE_%: Percentage of cholesterol esters in very low density lipoproteins in total lipids.
5. The method for identifying frailty subtypes based on machine learning and metabolomic features according to claim 1, wherein, When the multivariate Cox proportional hazards model was used to examine the association between the new frailty subtypes and the outcomes, the results showed that compared with the participants in subtype I, individuals in subtypes III and IV showed significantly higher risks in all outcomes; the outcomes included coronary artery disease, heart failure, major adverse cardiovascular events (MACE), myocardial infarction (MI), type 2 diabetes, metabolic dysfunction-associated fatty liver disease (MASLD), chronic obstructive pulmonary disease (COPD), severe liver disease (SLD), peripheral artery disease (PAD), end-stage renal disease (ESRD), renal cancer, lung cancer, abdominal aortic aneurysm (AAA), and all-cause mortality.
6. The method for identifying frailty subtypes based on machine learning and metabolomic features according to claim 1, wherein In the frailty subtype identification, when the CatBoost algorithm analyzed the metabolic features, the SHAP explainer was applied to the best training model, and SHAP values were calculated to rank the top 20 most influential features, and the top 11 metabolites with the highest SHAP values were selected for subsequent clustering.
7. The method for identifying frailty subtypes based on machine learning and metabolomic features according to claim 1, wherein In the frailty subtype identification step, when principal component analysis was performed on the selected 11 metabolic features, principal component 1 and principal component 2 explained 44.54% and 16.63% of the variance, respectively.
8. A frailty subtype identification system based on machine learning and metabolomics features, applied to the frailty subtype identification method based on machine learning and metabolomics features according to any one of claims 1 to 7, characterized in that, Including: Data acquisition and preprocessing module: Obtain participant data from a local database, construct a frailty index (FI) based on 49 items to evaluate the degree of frailty and exclude specific ineligible participants, and at the same time analyze 251 metabolites and perform standard normalization processing; and divide the data into a training set and a test set; Frailty subtype identification module: Use the CatBoost algorithm to analyze the metabolic features and select the top 11 key metabolites with the highest SHAP value rankings, and stratify the frail participants into four subtypes through principal component analysis and K-means clustering analysis; Frailty subtype and chronic disease and all-cause mortality association evaluation module: Use Kaplan–Meier curves to show the differences in cumulative incidence of the four subtypes and non-frail participants in 13 chronic diseases and all-cause mortality, and use the multivariate Cox proportional hazards model to examine the association between the new frailty subtypes and these outcomes, including adjusting for age and gender factors.
9. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to implement the frailty subtype identification method based on machine learning and metabolomics features as described in any one of claims 1 to 7 when executing the instructions.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that instructs the device to execute the frailty subtype identification method based on machine learning and metabolomics features as described in any one of claims 1 to 7.