Method and system for screening sarcopenia biomarkers
The upregulated gene of sarcopenia skeletal muscle was screened through transcriptome data and Mendel randomized analysis, which solved the problem of biomarker screening in the existing technology that is difficult to determine the causal relationship in the prior art, and achieved high accuracy and reliability of sarcopenia marker screening, which is suitable for disease diagnosis and prognosis judgment.
Patent Information
- Application Number
- CN202510461928.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-08-08
Smart Images

Figure CN120452531A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of diagnosis and treatment of skeletal muscle senile diseases, and specifically relates to a method and system for screening sarcopenia biomarkers. Background Art
[0002] Sarcopenia is a syndrome characterized by a progressive, age-related decline in skeletal muscle mass, strength, and function. Its core features include muscle atrophy, decreased strength, and impaired physical function. With the aging of the global population, sarcopenia has become a major public health concern threatening the health of the elderly. Its prevalence increases significantly with age, reaching 50% in people over 80 years old. In addition to direct consequences such as falls and fractures, sarcopenia is also closely associated with an increased risk of heart failure, respiratory disease, cognitive impairment, and mortality.
[0003] The pathogenesis of sarcopenia is complex, insidious, and nonspecific, making its early diagnosis extremely challenging. Currently, there is no unified and authoritative "gold standard" for the diagnosis of sarcopenia. Existing diagnostic methods generally rely on a comprehensive assessment of muscle mass, muscle strength, and physical function. However, these methods have many limitations in practical application, especially the susceptibility to subjective and environmental factors during the measurement process, which affects the accuracy and reproducibility of the results.
[0004] Therefore, there is an urgent need to find more objective, stable and highly sensitive diagnostic methods.
[0005] In light of this, in recent years, biomarkers, particularly gene expression level measurements, have emerged as potential solutions for the early assessment and intervention of sarcopenia, as they can overcome the subjectivity and environmental interference inherent in traditional methods. For example, by analyzing the mRNA expression levels of muscle-specific genes (such as the myogenic regulatory factors MyoD and Myogenin) and inflammatory factors (such as IL-6 and TNF-α), a patient's muscle health status can be dynamically visualized at the molecular level, minimizing the impact of human manipulation on diagnostic results. Furthermore, high-sensitivity sequencing technologies and multi-omics analyses can reveal underlying molecular events and biological mechanisms, providing a strong basis for the diagnosis of sarcopenia and the evaluation of therapeutic efficacy.
[0006] There is no biomarker screening method specifically for sarcopenia in the prior art, but Chinese patent application CN202411946317.4 discloses a genomics-based biomarker screening method, including: determining biomarker expression data related to significant genes based on transcriptome sequencing data of cancer disease biological samples, and the transcriptome sequencing data is obtained based on functional genomics and clinical genomics; based on the biomarker expression data of the significant genes, estimating the posterior distribution of the association between biomarkers and cancer diseases through a logistic regression model; based on the posterior distribution and the gene interaction network, determining the significant interacting genes in the cancer disease biological samples; based on the significant interacting genes, screening to obtain a list of biomarkers related to cancer diseases.
[0007] The above-mentioned existing technology estimates the "association posterior distribution" between biomarkers and cancer through a logistic regression model. In essence, it is a correlation analysis and cannot eliminate the interference of confounding factors (such as the environment and other genetic mutations), making it difficult to determine the causal relationship between biomarkers and diseases. Summary of the Invention
[0008] The purpose of the present invention is to provide a method and system for screening sarcopenia biomarkers, which can partially solve or alleviate the above-mentioned deficiencies in the prior art and can screen biomarkers that determine the causal relationship with sarcopenia outcomes from a large number of genes.
[0009] In order to solve the above-mentioned technical problems, the present invention specifically adopts the following technical solutions: The first aspect of the present invention is to provide a method for screening sarcopenia biomarkers, comprising: By comparing the transcriptomes of sarcopenia patients and healthy controls, we screened out the upregulated genes in sarcopenic skeletal muscle. Identify the sarcopenia-related genes that have been screened for sarcopenia skeletal muscle upregulation to obtain sarcopenia risk genes; The sarcopenia risk genes are externally validated to screen out sarcopenia biomarkers.
[0010] As an improvement, the method of screening for upregulated genes in sarcopenic skeletal muscle includes: The adapter sequences and low-quality bases of the sequencing reads in the skeletal muscle RNA-seq sequencing data of the sarcopenia patient group and the healthy control group were removed, and the reads with N values exceeding the N threshold ratio or with an average value below the average threshold were removed to correct the sequencing depth difference; The significant p-value of the differential expression of genes between the sarcopenia patient group and the healthy control group was calculated, and genes with a significant p-value greater than the p-threshold and a difference fold greater than the fold threshold were screened as sarcopenia skeletal muscle upregulated genes.
[0011] As an improvement, the negative binomial distribution generalized linear model of the DESeq2 software package was used to model the gene expression data, and the significant p-value of gene differential expression was calculated by the Wald test.
[0012] As an improvement, the method for identifying the relevance of sarcopenia to the screened sarcopenia skeletal muscle upregulated genes includes: Screening SNPs associated with the sarcopenia skeletal muscle up-regulated genes as SNP instrumental variables for Mendelian randomization analysis; The selected SNP instrumental variables were used to conduct a two-sample Mendelian randomization analysis on the association between genes and sarcopenia, inferring the causal effect of sarcopenia skeletal muscle upregulated genes on sarcopenia, and screening genes with a significant p-value less than the significance threshold as candidate genes; Sensitivity analysis was performed on the candidate genes, and genes that passed the sensitivity analysis and whose gene expression was positively correlated with sarcopenia were screened out as sarcopenia risk genes.
[0013] As an improvement, when screening SNPs, the screening condition is that the association p value is less than 5e -5 , linkage disequilibrium threshold was r2<0.05 and window size = 50 kb, and F statistic >10.
[0014] As an improvement, in the two-sample Mendelian randomization analysis, the Wald ratio method was used to calculate the effect size of a single SNP, and the inverse variance weighted method was used to calculate the effect size of two or more SNPs.
[0015] As an improvement, the sensitivity analysis includes horizontal pleiotropy analysis, heterogeneity analysis and reverse causality analysis.
[0016] As an improvement, the method for externally validating the sarcopenia risk gene includes: A two-sample Mendelian randomization analysis was performed on the association between genes and muscle mass, strength loss, and mobility decline to infer the causal effect of sarcopenia risk genes on muscle mass, strength loss, and mobility decline, and to screen out sarcopenia risk genes whose gene expression is positively correlated with muscle mass, strength loss, and mobility decline as sarcopenia biomarkers.
[0017] As an improvement, the sarcopenia biomarkers are genes PRRC2A, SIX5, and DMXL2.
[0018] The present invention also provides a sarcopenia biomarker screening system, comprising: The upregulated gene screening module is used to screen the upregulated genes of sarcopenic skeletal muscle by comparing the transcriptomes of sarcopenic patients and healthy controls; Sarcopenia risk gene screening module, used to identify the sarcopenia relevance of the screened sarcopenia skeletal muscle upregulated genes and obtain sarcopenia risk genes; The sarcopenia biomarker screening module is used to externally verify the sarcopenia risk gene, thereby screening out sarcopenia biomarkers.
[0019] Beneficial effects: The present invention first screens out genes upregulated in sarcopenic skeletal muscle through transcriptome data analysis. It can quickly lock in gene groups with abnormally elevated expression in the sarcopenic state from a massive number of genes, exclude genes with no significant expression changes, greatly narrow the scope of subsequent research, make the research more targeted, improve research efficiency, and avoid wasting resources on irrelevant genes.
[0020] Further screening of risk genes from upregulated genes, using methods such as Mendelian randomization analysis, and using genetic variation as an instrumental variable to infer a causal relationship between genes and sarcopenia, rather than simple association analysis, is performed. This approach effectively mitigates confounding factors and ensures a true causal relationship, rather than a chance association, between the identified risk genes and sarcopenia. This enhances the reliability of genes as risk factors and lays a more solid foundation for mechanistic research and clinical application.
[0021] Finally, when screening biomarkers from risk genes, by analyzing the causal effect of genes on sarcopenia-related traits, we ensure that the selected biomarkers not only have a causal relationship with sarcopenia, but are also closely associated with the typical phenotype of sarcopenia. This makes the biomarkers more closely aligned with the clinical characteristics of sarcopenia, more practical, and can be directly used for disease diagnosis, condition assessment, or prognosis, providing more valuable targets for clinical translation.
[0022] The entire screening process proceeds step by step, with each step based on rigorous data analysis and statistical testing to gradually eliminate false positives and confounding factors. This systematic screening process ensures that the final biomarkers identified undergo multiple rounds of validation, ensuring both biological plausibility and statistical significance. The rigorous and reliable results reduce the risk of misdiagnosis and provide high-quality research leads for in-depth research on sarcopenia and precision medicine.
[0023] In addition, the present invention integrates cis-eQTL and GWAS data, adopts two-sample Mendelian randomization analysis, combines Wald ratio to calculate single SNP effects, inverse variance weighting method to integrate multiple SNP effects, and Bonferroni multiple correction to control multiple false positives, to more accurately infer the causal relationship between genes and sarcopenia.
[0024] In addition to sarcopenia outcomes, this study also conducts causal effect analysis on sarcopenia-related traits such as grip strength and gait speed, assessing risk genes by comparing the consistency of effects across multiple outcomes. This multi-dimensional validation strategy makes the screening of risk genes more convincing. Existing technologies lack systematic analysis of related traits, making it difficult to fully verify the association between genes and sarcopenia phenotypes, which is more in line with the complex pathological phenotype of sarcopenia.
[0025] Through multiple rounds of screening including differential expression analysis, Mendelian randomization analysis, sensitivity analysis and external validation, the present invention ultimately identified candidate biomarker genes PRRC2A, SIX5, and DMXL2, which have both biological plausibility and statistical significance, providing a more precise direction for the mechanism research, clinical diagnosis and therapeutic target development of sarcopenia, and having higher application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for the embodiments or the description of the prior art. In all drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the various elements or parts are not necessarily drawn according to the actual scale. Obviously, the drawings described below are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without inventive work.
[0027] Figure 1 Schematic diagram of differentially expressed genes in sarcopenic skeletal muscle, where red dots represent upregulated genes.
[0028] Figure 2 Schematic diagram of the causal effect of upregulated genes in sarcopenic skeletal muscle on the outcome of sarcopenia.
[0029] Figure 3 Schematic diagram of the effect of SNP instrumental variables on the exposure gene DMXL2 and sarcopenia outcomes.
[0030] Figure 4 Schematic diagram of the effect of SNP instrumental variables on the exposure gene PRRC2A and sarcopenia outcomes.
[0031] Figure 5 Schematic diagram of the effect of SNP instrumental variables on the exposure gene SIX5 and sarcopenia outcomes.
[0032] Figure 6 Schematic diagram of the effect of SNP instrumental variables on the exposure gene DMXL2 and grip strength outcomes.
[0033] Figure 7 Schematic diagram of the effect of SNP instrumental variables on the exposure gene PRRC2A gene and grip strength outcome.
[0034] Figure 8 Schematic diagram of the effect of SNP instrumental variables on the exposure gene SIX5 and grip strength outcomes.
[0035] Figure 9 Schematic diagram of the effect of SNP instrumental variables on the exposure gene DMXL2 and walking speed outcomes.
[0036] Figure 10 Schematic diagram of the effect of SNP instrumental variables on the exposure gene PRRC2A and walking speed outcomes.
[0037] Figure 11 This is a logic diagram for screening sarcopenia biomarkers in the present invention. DETAILED DESCRIPTION
[0038] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0039] Herein, suffixes such as "module," "component," or "unit" used to represent elements are only used to facilitate description of the present invention and have no specific meaning. Therefore, "module," "component," or "unit" may be used interchangeably.
[0040] As used herein, terms such as "upper," "lower," "inner," "outer," "front," "back," "one end," and "the other end" indicate positions or locations based on those shown in the accompanying drawings. These terms are intended solely to facilitate and simplify the description of the present invention and are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on the present invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0041] As used herein, unless otherwise expressly specified or limited, the terms "installed," "provided with," and "connected" should be understood broadly. For example, "connected" may refer to a fixed connection, a detachable connection, or an integral connection; it may refer to a mechanical connection, a direct connection, an indirect connection via an intermediate medium, or internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention on a case-by-case basis.
[0042] As used herein, "and / or" includes any and all combinations of one or more of the associated listed items.
[0043] Herein, "plurality" means two or more than two, ie, it includes two, three, four, five, etc.
[0044] Example 1: This embodiment provides a method for screening biomarkers of sarcopenia, which is used to screen biomarkers that have a causal relationship with the outcome of sarcopenia from skeletal muscle genes. Figure 11 As shown, the intersection of the three sets of sarcopenia skeletal muscle highly expressed genes, sarcopenia risk genes, and genes negatively correlated with muscle mass / strength / performance constitutes the sarcopenia gene biomarkers of the present invention. This means that these genes simultaneously possess the following characteristics: 1. They are highly expressed in the skeletal muscle of sarcopenic patients. 2. They are risk genes causally associated with the risk of sarcopenia. 3. Their elevated expression levels are negatively correlated with muscle mass, strength, and performance.
[0045] In this embodiment, the specific steps of screening sarcopenia biomarkers include: S1 screened out upregulated genes in sarcopenic skeletal muscle by comparing the transcriptomes of sarcopenic patients and healthy controls.
[0046] S11 The gene data of the sarcopenia patient group and the healthy control group in this example comes from the skeletal muscle RNA-seq sequencing data of sarcopenia patients (19 subjects, average age 72 years) and healthy subjects (20 subjects, average age 68 years) downloaded from the NCBI GEO database, with the serial number GSE226151.
[0047] After obtaining the genetic data of the sarcopenia patient group and the healthy control group, S12 needs to preprocess the data and quantify the gene expression, including: First, the adapter sequences and low-quality bases of the sequencing reads in the skeletal muscle RNA-seq sequencing data of the sarcopenia patient group and the healthy control group were removed.
[0048] The adapter sequence is an artificial sequence added during the sequencing process and does not belong to the genetic information of the sample itself. Removing the adapter sequence can avoid its interference with subsequent analysis.
[0049] Low-quality bases are those with a Phred quality value less than 20. Phred quality measures the probability of base sequencing errors. A Phred quality value of 20 corresponds to a 1% probability of base sequencing errors. Removing low-quality bases can reduce the impact of sequencing errors on data, improve data accuracy, and ensure that subsequent analysis is based on high-quality sequencing data.
[0050] Secondly, reads with N values exceeding the N threshold ratio or with an average value below the average threshold are removed. For example, if the proportion of N (unknown bases) in a single sequencing read exceeds 10%, or the average quality value of a single sequencing read is lower than 20, the sequencing read is removed.
[0051] Next, the retained sequencing reads are aligned to the human reference genome and the sequencing read counts for each gene are counted. Bioinformatics tools (such as HISAT2 and STAR) are used to align the reads to the human reference genome (such as GRCh38) and count the corresponding sequencing read counts for each gene.
[0052] Aligning sequencing reads to a reference genome determines the location of each read on the genome, providing coordinates for subsequent gene expression statistics. Only sequencing reads that map to the genome can be assigned to specific genes. Sequencing read counts reflect the expression intensity of that gene in the sample. Higher gene expression levels indicate more mRNA generated by transcription, resulting in higher sequencing read counts.
[0053] Finally, sequencing depth differences were corrected using the Trimmed Mean of M-values (TMM) method.
[0054] TMM is a method for normalizing RNA-seq data. This method removes the extreme M values of highly expressed genes and lowly expressed genes, and only averages the M values of moderately expressed genes as a global correction factor.
[0055] The sequencing depth, or total number of sequencing reads, can vary significantly between samples. For example, one sample may have 10 million sequencing reads, while another may have 5 million. Directly comparing sequencing reads can introduce bias due to differences in sequencing depth. The TMM method corrects for sequencing depth, making gene expression comparable across samples. This ensures the accuracy of subsequent differential expression analyses (such as DESeq2) and allows for a more realistic reflection of biological differences.
[0056] S13 calculates the significant p-value of the differential expression of genes between the sarcopenia patient group and the healthy control group, and screens out genes with a significant p-value greater than the p-threshold and a difference fold greater than the fold threshold as sarcopenia skeletal muscle upregulated genes.
[0057] Specifically, the negative binomial distribution generalized linear model of the DESeq2 software package was used to model the gene expression data, and the significant p-value of the gene differential expression was calculated by the Wald test.
[0058] DESeq2 is a widely used tool for differential expression analysis of RNA-seq data. It can effectively process count data such as gene sequencing read counts and take into account factors such as differences between samples and sequencing depth.
[0059] Gene expression data were modeled using a generalized linear model with a negative binomial distribution. The negative binomial distribution is suitable for describing count data and can handle overdispersion, a common phenomenon in RNA-seq data. Generalized linear models can flexibly model the relationship between gene expression and sample grouping, namely, the sarcopenia and healthy groups in this example, thereby accurately assessing differences in gene expression between different groups.
[0060] The Wald test is a commonly used hypothesis testing method used in generalized linear models to test whether model parameters are zero. In this example, the Wald test was used to determine whether there was a significant difference in the expression of each gene between the sarcopenia group and the healthy group.
[0061] In this example, the criteria for up-regulated genes are p < 0.05 (i.e., the difference is statistically significant, the null hypothesis is rejected, and it is considered that there is a difference in gene expression between the two groups) and the difference fold > 0 (indicating that the gene expression level in the sarcopenia group is higher than that in the healthy group, i.e., gene expression is up-regulated).
[0062] A total of 947 genes with upregulated expression were identified in the skeletal muscle of patients with sarcopenia. These genes are believed to play an important role in the development and progression of sarcopenia and are the focus of subsequent in-depth research.
[0063] Figure 1 In the figure, the horizontal axis (log2) represents the fold difference, reflecting the fold change in gene expression between the sarcopenia group and the healthy group, taking the logarithm to base 2. A positive value indicates that the gene expression in the sarcopenia group is higher than that in the healthy group; a negative value indicates that the gene expression in the sarcopenia group is lower than that in the healthy group.
[0064] vertical axis-log 10 The p-value measures the statistical significance of differential gene expression. The negative logarithm is used to magnify small values for easier observation. A larger value indicates a smaller p-value and more significant differential gene expression.
[0065] like Figure 1 As shown in the figure, multiple differentially expressed genes were identified in the skeletal muscle of patients with sarcopenia. Among them, 947 genes were upregulated with significant differences (represented by red dots). These genes may play an important role in the development and progression of sarcopenia and are the focus of subsequent in-depth research on the molecular mechanisms of sarcopenia and the search for biomarkers.
[0066] S2 conducts sarcopenia-related identification on the screened sarcopenia skeletal muscle upregulated genes to obtain sarcopenia risk genes.
[0067] S21 In this step, download the human skeletal muscle cis-eQTL summary data (v8 version) and sarcopenia GWAS summary data from the GTEx database and GWAS catalog database, respectively, with the number GCST90007526.
[0068] The GTEx (Genotype-Tissue Expression) database is a large public database designed to study gene expression patterns in different human tissues and their associations with genetic variation. Download the human skeletal muscle cis-eQTL summary data (version 8) from the database. Cis-eQTLs (cis-acting expression quantitative trait loci) are genetic variants (such as single nucleotide polymorphisms (SNPs) in this example) located near a gene (typically within 1000 kb upstream or downstream) that can affect the expression levels of nearby genes.
[0069] The GWAS catalog database collects results from numerous genome-wide association studies aimed at identifying genetic variants associated with various diseases and traits. Download the Sarcopenia GWAS summary data, numbered GCST90007526, from the database. GWAS, through genotyping and phenotyping of large numbers of samples, can identify genetic variants associated with sarcopenia risk.
[0070] S22 screens SNPs associated with the sarcopenia skeletal muscle up-regulated genes as SNP instrumental variables for Mendelian randomization analysis.
[0071] The purpose of instrumental variable screening is to select genetic variants (SNPs) that can reliably represent the expression levels of upregulated genes in sarcopenic skeletal muscle as instrumental variables for use in subsequent Mendelian randomization analysis and other studies.
[0072] In this embodiment, when screening SNPs, the screening condition is that the association p value is less than 5e -5 , linkage disequilibrium threshold was r2 < 0.05 and window size (sliding window width) = 50 kb, and F statistic > 10.
[0073] The p-value is an indicator used to measure the significance of statistical results. In this embodiment, the p-value of the correlation is required to be less than 5e - 5 , indicating that the association between the screened SNP instrumental variables and sarcopenia skeletal muscle upregulated genes is highly statistically significant.
[0074] Linkage disequilibrium refers to the phenomenon of non-random association between alleles at different loci. r² is a measure of linkage disequilibrium. An r² value of <0.05 indicates low linkage disequilibrium between the selected SNP instrumental variables, meaning weak correlation. This avoids excessive redundant information between the selected instrumental variables and ensures that each instrumental variable independently reflects gene expression. A 50 kb window size assesses linkage disequilibrium within a 50 kb genomic region centered on each SNP. The 1000 Genomes reference panel provides a widely recognized reference standard for genomic variants, ensuring accuracy and consistency in the screening process.
[0075] The F statistic is used to test the strength of an instrumental variable. An F statistic greater than 10 indicates a strong association between the selected SNP instrumental variable and the expression of genes upregulated in sarcopenic skeletal muscle, effectively serving as an instrumental variable to infer a causal relationship between gene expression and sarcopenia. A low F statistic may indicate a weak association between the instrumental variable and gene expression, thus compromising the reliability of subsequent analysis results.
[0076] S23 uses the screened SNP instrumental variables to perform a two-sample Mendelian randomization analysis on the association between genes and sarcopenia, inferring the causal effect of sarcopenia skeletal muscle upregulated genes on sarcopenia, and screening genes with significant p-values less than the significance threshold as candidate genes.
[0077] Mendelian randomization analysis is based on Mendel's laws of inheritance and uses genetic variation, such as SNPs in this example, as instrumental variables to infer causal relationships between gene expression and disease outcomes, such as sarcopenia in this example. In the two-sample Mendelian randomization analysis in this example, gene expression data were obtained from human skeletal muscle cis-eQTL summary data, and outcome data were obtained from sarcopenia GWAS summary data. By comparing the effects of SNPs on genes and sarcopenia, it was determined whether gene expression exacerbated or reduced the risk of sarcopenia, that is, whether the gene was causal for the sarcopenia outcome.
[0078] More specifically, when a gene has only a single SNP instrumental variable, its effect size is calculated using the Wald ratio method. This method estimates the causal effect of gene expression on sarcopenia by comparing the effect of the SNP on sarcopenia with the effect of the SNP on gene expression. Simply put, the effect size of the SNP on the outcome is divided by the effect size of the SNP on the exposure to obtain an estimate of the causal effect of the gene on the sarcopenia outcome.
[0079] When a gene has two or more SNP instruments, the inverse variance weighting method is used to calculate the effect size. The IVW method takes a weighted average of the effect sizes of each SNP based on its variance. The calculation begins by first calculating the weight of each SNP, then multiplying the effect size by its weight, adding these products, and finally dividing by the sum of the weights to obtain the combined effect size. This method integrates information from multiple SNPs, reducing potential bias introduced by individual SNPs and improving the stability and reliability of the results.
[0080] Since the analysis involves testing multiple genes, multiple testing correction is required to control the false positive rate caused by multiple testing. Specifically, the Bonferroni correction method is used in this embodiment to perform multiple testing correction.
[0081] The original significance level is usually set to α = 5. In this example, 374 independent tests were performed, so the significance p threshold after Bonferroni correction is 0.05 / 374≈1.337×10 -4 Only when the p-value of the gene's causal analysis is less than this corrected threshold is the causal relationship between the gene and sarcopenia considered statistically significant. That is, the gene is preliminarily identified as a possible causal gene for sarcopenia and can enter subsequent sensitivity analysis and other steps for further verification.
[0082] S24 performed sensitivity analysis on candidate genes and screened out genes that passed the sensitivity analysis and whose gene expression was positively correlated with sarcopenia as sarcopenia risk genes.
[0083] The purpose of performing sensitivity analysis is to further verify the reliability and stability of the results and to eliminate the interference of other potential factors on the conclusions. In this embodiment, sensitivity analysis includes horizontal pleiotropy analysis, heterogeneity analysis and reverse causality analysis.
[0084] Horizontal pleiotropy analysis (MR-Egger regression) refers to the fact that SNP instrumental variables affect sarcopenia outcomes through other independent pathways besides influencing gene expression. By fitting a regression model, the intercept term can reflect the presence of horizontal pleiotropy. If the instrumental variable does not exhibit horizontal pleiotropy, then the intercept should theoretically be 0; if the intercept deviates significantly from 0, horizontal pleiotropy is indicated. When the p-value of the MR-Egger intercept test is >0.05, it indicates that the intercept does not deviate significantly from 0, indicating that there is insufficient evidence for horizontal pleiotropy. In this case, the results of Mendelian randomization analysis are more reliable.
[0085] Qualitative analysis (Cochran Q test), when using multiple SNPs as instrumental variables, heterogeneity analysis is used to test whether the estimates of causal effects of these instrumental variables are consistent. If heterogeneity exists, it means that different SNPs may affect the outcome in different ways or degrees, which may affect the reliability of the results of Mendelian randomization analysis. The Cochran Q test assesses heterogeneity by calculating the differences between the effect values of each SNP. Specifically, it compares the actual observed effect value differences with the expected random differences. When the p-value of the Q test is >0.05, it indicates that there is no significant difference between the effect values of different instrumental variables, that is, the heterogeneity is not significant. This means that the estimates of the causal effects of each SNP are relatively consistent, and the results are relatively stable and reliable; conversely, if the p-value is ≤0.05, it indicates that there is significant heterogeneity, and further analysis of the cause or adjustment of the analysis method may be required.
[0086] Reverse causation analysis (Steiger filter test) refers to the hypothesis that reverse causation does not cause the outcome of sarcopenia, but rather that the outcome causes changes in the exposure factor. The Steiger filter test is used to determine whether this reverse causal relationship exists. This test determines the direction of causation by comparing the variance of the effect of gene expression on sarcopenia with the variance of the effect of sarcopenia on gene expression. If the variance of the effect of gene expression on sarcopenia is greater than the variance of the effect of sarcopenia on gene expression, this strengthens the causal relationship in which gene expression is the cause and sarcopenia is the effect; conversely, reverse causation is possible. When the p-value of the Steiger filter test is <0.05, it indicates that the variance of the effect of gene expression on sarcopenia is significantly greater than the variance of the effect of sarcopenia on gene expression, indicating no evidence of reverse causation, thus supporting the causal conclusion originally drawn through Mendelian randomization analysis. If the p-value is ≥0.05, the possibility of reverse causation cannot be ruled out, and the reliability of the conclusion is questioned.
[0087] like Figure 2 As shown in this example, a total of 8 genes passed the sensitivity analysis and showed a robust causal effect on sarcopenia. Among them, the effect OR values of 5 genes were greater than 1, indicating that increased expression of these genes was positively correlated with sarcopenia outcomes. Therefore, they were preliminarily identified as risk genes for sarcopenia, including: DLG5 (OR=1.0788, p=3.748e-06), PRRC2A (OR=1.3419, p=5.752e-18), MFHAS1 (OR=1.0350, p=1.387e-05), SIX5 (OR=1.1460, p=4.199e-06), and DMXL2 (OR=1.0744, p=2.592e-06).
[0088] S3 externally validates the sarcopenia risk genes to identify sarcopenia biomarkers. Specifically, a two-sample Mendelian randomization analysis is performed to investigate the association between the genes and muscle mass, strength loss, and mobility decline, inferring the causal effects of sarcopenia risk genes on muscle mass, strength loss, and mobility decline. Furthermore, sarcopenia risk genes whose gene expression is positively correlated with muscle mass, strength loss, and mobility decline are identified as sarcopenia biomarkers.
[0089] The main clinical manifestations of sarcopenia are decreased muscle mass and strength, as well as decreased mobility. The reliability of previously identified sarcopenia risk genes was assessed by analyzing their causal effects on these sarcopenia-related traits and comparing the consistency of effects across multiple outcomes.
[0090] The limb lean mass GWAS data (GCST90000025) were downloaded from the GWAS catalog database, and the grip strength GWAS data (ukb-b-10215) and walking speed GWAS data (ukb-b-4711) were downloaded from the IEUOpenGWAS database.
[0091] The two-sample Mendelian randomization analysis method mentioned in the previous step was used to calculate the causal effects of the five identified risk genes on the three outcomes of limb lean mass, grip strength, and gait speed. The detailed analysis process is not repeated in this step.
[0092] See also Figures 3 to 10 The results showed that increased expression of PRRC2A (beta = -0.0862, p = 6.882e-40), SIX5 (beta = -0.0384, p = 1.986e-10), and DMXL2 (beta = -0.0173, p = 1.655e-08) genes was associated with lower grip strength, and increased expression of PRRC1A (beta = -0.0345, p = 5.236e-10) and DMXL2 (beta = -0.0116, p = 9.263e-06) was associated with slower walking speed. Therefore, a total of three genes, PRRC2A, SIX5, and DMXL2, showed consistent causal effects between sarcopenia and related traits. Based on the above results, these three genes are highly expressed in the skeletal muscle of sarcopenia patients, and the increased expression levels of these genes are associated with an increased risk of multiple sarcopenia phenotypes. They are considered as candidate biomarkers for sarcopenia.
[0093] Example 2: This embodiment also provides a sarcopenia biomarker screening system, comprising: The upregulated gene screening module is used to screen the upregulated genes of sarcopenic skeletal muscle by comparing the transcriptomes of sarcopenic patients and healthy controls; Sarcopenia risk gene screening module, used to identify the sarcopenia relevance of the screened sarcopenia skeletal muscle upregulated genes and obtain sarcopenia risk genes; The sarcopenia biomarker screening module is used to externally verify the sarcopenia risk gene, thereby screening out sarcopenia biomarkers.
[0094] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0095] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a computer terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0096] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.
Claims
1. A method for screening sarcopenia biomarkers, characterized in that include: By comparing the transcriptomes of sarcopenia patients and healthy controls, we screened out the upregulated genes in sarcopenic skeletal muscle. Identify the sarcopenia-related genes that have been screened for sarcopenia skeletal muscle upregulation to obtain sarcopenia risk genes; The sarcopenia risk genes are externally validated to screen out sarcopenia biomarkers.
2. The method for screening sarcopenia biomarkers according to claim 1, characterized in that Methods for screening for upregulated genes in sarcopenic skeletal muscle include: The adapter sequences and low-quality bases of sequencing reads in the skeletal muscle RNA-seq sequencing data of sarcopenia patients and healthy controls were removed, and reads with N values exceeding the N threshold ratio or with an average value below the average threshold were removed; the retained sequencing reads were aligned with the human reference genome and the sequencing read counts of each gene were counted; and differences in sequencing depth were corrected; The significant p-value of the differential expression of genes between the sarcopenia patient group and the healthy control group was calculated, and genes with a significant p-value greater than the p-threshold and a difference fold greater than the fold threshold were screened as sarcopenia skeletal muscle upregulated genes.
3. The method for screening sarcopenia biomarkers according to claim 2, wherein: The gene expression data were modeled using the negative binomial distribution generalized linear model of the DESeq2 software package, and the significant p-values of gene differential expression were calculated using the Wald test.
4. The method for screening sarcopenia biomarkers according to claim 1, characterized in that Methods for identifying the relevance of sarcopenia to the screened sarcopenia skeletal muscle upregulated genes include: Screening SNPs associated with the sarcopenia skeletal muscle up-regulated genes as SNP instrumental variables for Mendelian randomization analysis; The selected SNP instrumental variables were used to conduct a two-sample Mendelian randomization analysis on the association between genes and sarcopenia, inferring the causal effect of sarcopenia skeletal muscle upregulated genes on sarcopenia, and screening genes with a significant p-value less than the significance threshold as candidate genes; Sensitivity analysis was performed on the candidate genes, and genes that passed the sensitivity analysis and whose gene expression was positively correlated with sarcopenia were screened out as sarcopenia risk genes.
5. The method for screening sarcopenia biomarkers according to claim 4, wherein: When screening SNPs, the screening condition is that the association p value is less than 5e -5 , linkage disequilibrium threshold was r2<0.05 and window size = 50kb, F statistic >10.
6. The method for screening sarcopenia biomarkers according to claim 4, wherein: In the two-sample Mendelian randomization analysis, the Wald ratio method was used to calculate the effect size of a single SNP, and the inverse variance weighted method was used to calculate the effect size of two or more SNPs.
7. The method for screening sarcopenia biomarkers according to claim 4, wherein: The sensitivity analysis included horizontal pleiotropy analysis, heterogeneity analysis, and reverse causality analysis.
8. The method for screening sarcopenia biomarkers according to claim 1, characterized in that The method for externally validating the sarcopenia risk gene includes: A two-sample Mendelian randomization analysis was performed on the association between genes and muscle mass, strength loss, and mobility decline to infer the causal effect of sarcopenia risk genes on muscle mass, strength loss, and mobility decline, and to screen out sarcopenia risk genes whose gene expression is positively correlated with muscle mass, strength loss, and mobility decline as sarcopenia biomarkers.
9. The method for screening sarcopenia biomarkers according to claim 1, wherein: The sarcopenia biomarkers are genes PRRC2A, SIX5, and DMXL2.
10. A sarcopenia biomarker screening system, characterized in that include: The upregulated gene screening module is used to screen the upregulated genes of sarcopenic skeletal muscle by comparing the transcriptomes of sarcopenic patients and healthy controls; Sarcopenia risk gene screening module, used to identify the sarcopenia relevance of the screened sarcopenia skeletal muscle upregulated genes and obtain sarcopenia risk genes; The sarcopenia biomarker screening module is used to externally verify the sarcopenia risk gene, thereby screening out sarcopenia biomarkers.
Citation Information
Patent Citations
Biomarker screening method and system based on genomics
CN119360969A