Marker group for screening and diagnosing sarcopenia and application
By using a multi-omics combined with machine learning approach to screen differentially expressed proteins and metabolites, a sarcopenia diagnostic model was constructed, which solved the problems of expensive equipment and poor diagnostic performance in existing technologies, and achieved highly sensitive sarcopenia screening and diagnosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WEST CHINA HOSPITAL SICHUAN UNIV
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-12
AI Technical Summary
Existing diagnostic methods for sarcopenia rely on expensive and cumbersome dual-energy X-ray absorptiometry and other tests, making large-scale population screening difficult. Furthermore, existing plasma biomarkers have poor diagnostic performance and are susceptible to interference from comorbidities, exhibiting high feature dimensionality and poor model generalization ability.
Using a multi-omics combined with machine learning approach, and combining the Olink Explore 384 inflammatory proteomics platform and LC-MS non-targeted metabolomics technology, differentially expressed proteins and metabolites were screened, biomarker groups were constructed, and diagnostic models of four core biomarkers (CCL13, FGF2, N-hexadecanoylpyrrolidine, 1-(cyclohexylmethyl)proline) were developed.
It achieves highly sensitive diagnosis of sarcopenia, with an AUC>0.9 in an independent validation cohort, demonstrating high clinical translational value, simplifying the diagnostic process and improving diagnostic performance.
Smart Images

Figure CN122017251A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical technology, specifically relating to a biomarker group for the screening and diagnosis of sarcopenia and its uses. Background Technology
[0002] Sarcopenia is an age-related muscle disease characterized by a progressive decline in skeletal muscle mass, strength, and function, severely impacting the quality of life of older adults and increasing the risk of falls, fractures, and death. Currently, internationally accepted diagnostic criteria (such as AWGSOP2) rely on dual-energy X-ray absorptiometry (DXA) to measure muscle mass, grip strength testing, and gait rate assessment. These methods have limitations, including expensive equipment, cumbersome operation, and limited accessibility, hindering large-scale population screening.
[0003] In recent years, plasma biomarkers have attracted attention due to their advantages of being non-invasive and repeatable. Previous studies have explored potential biomarkers such as the creatinine / cystatin C ratio, AST / ALT ratio, and irisin, but their diagnostic performance is generally poor (AUC < 0.8) and easily affected by comorbidities. Furthermore, proteomics or metabolomics modeling alone suffers from high feature dimensionality, poor model generalization ability, and insufficient validation. Summary of the Invention
[0004] To address the aforementioned shortcomings in the prior art, this invention provides a set of biomarkers for the screening and diagnosis of sarcopenia and their applications.
[0005] To achieve the above objectives, the technical solution adopted by the present invention to solve its technical problem is as follows: The purpose of this invention is to provide a set of biomarkers for screening and diagnosis of sarcopenia, which are selected from at least four of the following proteins and metabolites, including CXCL8, CCL13, IL18R1, RABGAP1L, CCL20, FGF2 and CLEC4A. Metabolites include 1-(Cyclohexylmethyl)proline, 24-epi-brassinolide, 2-(3,4,5-Trimethoxyphenyl)-1H-benzimidazole, 2(1H)-Pyrimidinone, 5-[3-[(1S,2S,4R)-bicyclo[2.2.1]hept-2-yloxy]-4-methoxyphenyl]tetrahydro-, Propentofylline, Adenosine, and (3-Oxo-2-piperazinyl)acetic acid. acid ((3-oxo-2-piperazinyl)acetic acid (or 2-(3-oxopiperazin-1-yl)acetic acid)).
[0006] Furthermore, the biomarker set consists of two proteins and two metabolites.
[0007] Furthermore, the biomarker group consists of CCL13, FGF2, fatty acid amides, and proline derivatives.
[0008] Furthermore, the proline derivative is 1-(cyclohexylmethyl)proline.
[0009] Furthermore, the biomarker group consists of CXCL8, CCL13, IL18R1, RABGAP1L, CCL20, FGF2, CLEC4A, 1-(Cyclohexylmethyl)proline, 24-epi-brassinolide, 2-(3,4,5-Trimethoxyphenyl)-1H-benzimidazole, 2(1H)-Pyrimidinone, 5-[3-[(1S,2S,4R)-bicyclo[2.2.1]hept-2-yloxy]-4-methoxyphenyl]tetrahydro-, Propentofylline, Adenosine, and (3-Oxo-2-piperazinyl)acetic acid.
[0010] Another object of the present invention is to provide a screening and diagnostic kit for sarcopenia, comprising reagents for detecting the above-mentioned biomarker group.
[0011] Another object of the present invention is to provide the use of reagents for detecting the above-mentioned biomarker group in the preparation of formulations for screening and diagnosis of sarcopenia.
[0012] Furthermore, the preparation is a reagent, kit, or test strip.
[0013] Furthermore, the reagents include reagents for the quantitative detection of protein antibodies and reagents for the quantitative detection of metabolites.
[0014] The beneficial effects of this invention are: This invention provides two highly sensitive combined diagnostic models for sarcopenia based on plasma multi-omics and machine learning. Using the Olink Explore 384 inflammatory proteomics platform and LC-MS non-targeted metabolomics technology, plasma samples from sarcopenia patients and healthy controls were systematically analyzed to identify differentially expressed proteins and metabolites. Multiple machine learning algorithms were employed to screen key features, constructing single-omics and dual-omics combined diagnostic models. One simplified diagnostic model, containing only four core biomarkers (CCL13, FGF2, N-hexadecanoylpyrrolidine, and 1-(cyclohexylmethyl)proline), demonstrated excellent diagnostic performance (AUC>0.9) in an independent validation cohort, possessing high clinical translational value. Attached Figure Description
[0015] Figure 1 This is a flowchart of the present invention; Figure 2 Volcano plots of differentially expressed proteins and metabolites; Figure 3 The AUC detection results are for six machine learning algorithms; Figure 4 The ROC curve for combined model 1 in the discovery and validation queues; Figure 5 The ROC curve for combined model 2 in the discovery and validation queues. Detailed Implementation
[0016] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0017] Example 1. Research Design and Participant Recruitment This invention is based on the "West China Health and Aging Trends Cohort" (WCHAT). A discovery cohort (n=80, sarcopenia:non-sarcopenia = 1:1) and an independent validation cohort (n=60, 1:1 matching) were constructed.
[0018] The inclusion criteria are as follows: (1) Age ≥ 60 years; (2) Complete sarcopenia assessment data; (3) Completion of qualified plasma sample collection.
[0019] The exclusion criteria are as follows: (1) having cognitive impairment; (2) having severe infection, malignant tumor or other major systemic disease; (3) hemolysis of plasma sample.
[0020] 2. Data collection and sarcopenia assessment The participants’ general data came from the HIS electronic medical record system, including gender, age, BMI, drinking history and other basic clinical information. In addition, the following clinical data were collected: (1) information on common comorbidities such as hyperlipidemia, coronary heart disease and hypertension; (2) indicators such as waist circumference, hip circumference and arm circumference; (3) core diagnostic parameters of sarcopenia: skeletal muscle mass index (SMI), grip strength and walking speed.
[0021] The diagnosis of sarcopenia is based on the Asian Working Group on Sarcopenia (AWGS 2019) criteria, including (1) low muscle mass (SMI < 7.0 for men / < 5.7 for women kg / m²). 2 (2) Low grip strength (<28kg male / <18kg female); (3) Slow walking speed (<1.0 m / s).
[0022] 3. Sample collection and multiplex protein detection Plasma was collected from the discovery cohort and the independent validation cohort: peripheral venous blood was drawn in the morning on an empty stomach, centrifuged at 3000 rpm for 10 minutes, and the plasma was separated and immediately aliquoted and stored at -80℃.
[0023] Proteomic analysis: 369 inflammation-related proteins were detected using the Olink Explore 384 Inflammation Panel based on the Proximity Extension Assay (PEA). Data were output as normalized protein expression values (NPX, log2 scale) and used for analysis after batch effect correction. Furthermore, plasma proteins (CCL13 and FGF2) with significant high-performance were further validated using enzyme-linked immunosorbent assay (ELISA).
[0024] Metabolomics assay: Small molecule metabolites were extracted from 50 μL of plasma using untargeted liquid chromatography-mass spectrometry (LC-MS), and analyzed after vacuum drying and reconstitution (Waters UPLC + Orbitrap Exploris 120). Positive and negative ion modes were used for scanning (m / z 70–1050), and metabolite identification was performed using the BiotreeDB database.
[0025] 4. Screening and functional analysis of differential markers Proteomics: Based on the results of homogeneity of variance and normality tests, t-tests or Wilcoxon rank-sum tests were used to screen for 65 differentially expressed proteins (61 upregulated and 4 downregulated). Figure 2 A); Metabolomics: The same method identified 268 differentially regulated metabolites (28 upregulated, 240 downregulated). Figure 2 B); Functional enrichment analysis: GO / KEGG analysis showed that differentially expressed proteins were enriched in the chemokine signaling pathway and the NF-κB pathway; metabolites were enriched in pathways such as glycine-serine-threonine metabolism and branched-chain amino acid synthesis, revealing the core role of inflammation and metabolic disorders in sarcopenia.
[0026] 5. Machine Learning Modeling and Feature Selection Feature selection: Five algorithms (elastic mesh, CFS, Boruta, RFE, SVM) were used to score proteins and metabolites respectively, and the intersection was used to determine important features. The results are shown in Table 1 and Table 2.
[0027] Table 1 Protein Weights
[0028] Table 2 Metabolite Weights
[0029] Model Construction: In the discovery cohort, after adjusting for clinical factors (age, sex, BMI), a binary classification model was trained using six machine learning algorithms (random forest, SVM, LASSO, Naive Bayes, gradient boosting, and KNN). Ten-fold cross-validation was employed, and the area under the receiver operating characteristic (ROC) curve (AUC) and its 95% confidence interval (CI) were calculated. Sensitivity, specificity, positive predictive value (PPV), and F1 score were calculated at the optimal threshold. The results are shown in [Figure number missing]. Figure 3 .
[0030] like Figure 3 As shown, Naïve Bayes performed best in both proteomics and metabolomics, constructing a 7-protein model (AUC=0.743) and a 7-metabolite model (AUC=0.828), respectively.
[0031] 6. Construction of multi-omics fusion models Combinatorial Model 1: Using the predicted probabilities of the optimal protein model and the metabolic model as input, a multi-omics ensemble model is constructed through logistic regression fusion. Its discovery cohort AUC reaches 0.951, and the validation cohort AUC is 0.823 (see...). Figure 4 ).
[0032] Figure 4 In the diagram, A is the AUC power test plot of Bayesian protein, metabolism, and combinatorial model 1 in the discovery cohort; B is the AUC power test plot of Bayesian protein, metabolism, and combinatorial model 1 in the validation cohort; and C is the confusion matrix of combinatorial model 1 in the validation cohort.
[0033] Combinatorial Model 2: The top two most discriminative biomarkers from each omics were selected: proteomics: CCL13 (chemokine), FGF2 (fibroblast growth factor); metabolomics: N-hexadecanoylpyrrolidine (fatty acid amide), 1-(cyclohexylmethyl)proline (proline derivative). A minimal model containing only four biomarkers was constructed, and its validation cohort achieved an AUC of 0.911, significantly better than the single-omics model (P<0.05), and showed no statistical difference compared to the complex model (P=0.124). Furthermore, the clinical utility (decision curve analysis, DCA) of each model was evaluated, see [link to relevant documentation]. Figure 5 .
[0034] Figure 5 A shows the AUC power of Bayesian protein, metabolism, and combinatorial model 2 in the discovery cohort; B shows the AUC power of Bayesian protein, metabolism, and combinatorial model 2 in the validation cohort and the confusion matrix of combinatorial model 2 in the validation cohort; C compares the AUC of combinatorial model 2 and combinatorial model 1 in the validation cohort, showing that combinatorial model 2 achieves higher diagnostic power with fewer biomarkers; D shows the Calibration curves (model accuracy) of combinatorial model 2, metabolism, and protein; E shows the Decision Curve Analysis (DCA) of combinatorial model 2, metabolism, and protein.
[0035] Finally, it should be noted that the above specific embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to examples, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications and substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A biomarker set for screening and diagnosis of sarcopenia, characterized in that, The protein is selected from at least four of the following proteins and metabolites, including CXCL8, CCL13, IL18R1, RABGAP1L, CCL20, FGF2, and CLEC4A. The metabolites include 1-(Cyclohexylmethyl)proline, 24-epi-brassinolide, 2-(3,4,5-Trimethoxyphenyl)-1H-benzimidazole, 2(1H)-Pyrimidinone, 5-[3-[(1S,2S,4R)-bicyclo[2.2.1]hept-2-yloxy]-4-methoxyphenyl]tetrahydro-, Propentofylline, Adenosine and (3-Oxo-2-piperazinyl)acetic acid.
2. The marker set according to claim 1, characterized in that, The biomarker group consists of two proteins and two metabolites.
3. The marker set according to claim 1, characterized in that, The biomarker group consists of CCL13, FGF2, fatty acid amides, and proline derivatives.
4. The marker set according to claim 3, characterized in that, The proline derivative is 1-(cyclohexylmethyl)proline.
5. The set of markers according to claim 1, characterized in that, The biomarker group consists of CXCL8, CCL13, IL18R1, RABGAP1L, CCL20, FGF2, CLEC4A, 1-(Cyclohexylmethyl)proline, 24-epi-brassinolide, 2-(3,4,5-Trimethoxyphenyl)-1H-benzimidazole, 2(1H)-Pyrimidinone, 5-[3-[(1S,2S,4R)-bicyclo[2.2.1]hept-2-yloxy]-4-methoxyphenyl]tetrahydro-, Propentofylline, Adenosine, and (3-Oxo-2-piperazinyl)acetic acid.
6. A reagent kit for screening and diagnosing sarcopenia, characterized in that, The kit includes reagents for detecting the biomarker group according to any one of claims 1 to 5.
7. Use of the reagent for detecting the biomarker group according to any one of claims 1 to 5 in the preparation of preparations for screening and diagnosing sarcopenia.
8. The use according to claim 7, characterized in that, The preparations are reagents, kits, or test strips.
9. The use according to claim 7, characterized in that, The reagents include reagents for the quantitative detection of protein antibodies and reagents for the quantitative detection of metabolites.