A screening method and system for alzheimer's disease diagnostic markers
By combining Mendelian randomization analysis based on aggregated data with multi-omics information, diagnostic biomarkers and therapeutic targets for Alzheimer's disease were screened, solving the problem of screening difficulties in existing technologies and realizing a deeper understanding of the molecular mechanisms of the disease and personalized treatment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANDONG UNIV
- Filing Date
- 2024-10-31
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies are insufficient for effectively screening diagnostic biomarkers and therapeutic targets for Alzheimer's disease at the multi-omics level, and Mendelian randomization methods are prone to confounding and reverse causality biases in observational studies.
Using Mendelian randomization (SMR) analysis based on aggregated data, combined with gene methylation, gene expression, and protein abundance information, we obtained GWAS data of hereditary trait variant sites (SNPs) associated with Alzheimer's disease. We then used multi-omics genetic information data for analysis to screen for diagnostic biomarkers and therapeutic targets for Alzheimer's disease.
It has accurately identified diagnostic biomarkers and therapeutic targets for Alzheimer's disease, deepened our understanding of the disease's molecular mechanisms, provided innovative strategies for personalized treatment, and improved patients' quality of life.
Smart Images

Figure BDA0005113146900000081 
Figure BDA0005113146900000084 
Figure BDA0005113146900000087
Abstract
Description
Technical Field
[0001] This invention relates to the fields of biomedicine and data analysis, and in particular to a method and system for screening diagnostic biomarkers for Alzheimer's disease. Background Technology
[0002] The information disclosed in this background section is intended only to enhance understanding of the overall background of the invention and is not necessarily to be construed as an admission or in any way implying that such information constitutes prior art known to those skilled in the art.
[0003] The incidence of neurodegenerative diseases is increasing year by year, and they are more common in the elderly. These include Alzheimer's disease (AD) and Parkinson's disease, among others. Patients experience symptoms such as memory loss, language impairment, disorientation in time and space, intellectual disability, slowed thinking, poor calculation and judgment, personality changes, and loss of initiative in daily life and work. In particular, AD has become the fourth leading cause of death after cardiovascular disease, cancer, and stroke. The causes of AD are multifactorial, including genetics, infection, traumatic brain injury, environmental factors, and lifestyle factors. However, the underlying mechanisms of each factor remain unclear, posing significant challenges to the discovery and treatment of the disease. Biomolecular information (such as protein and gene expression levels and methylation) can reflect the disease's progression and provide physiological and pathological information about patients. It can also serve as an effective intervention target, providing insights for drug development and precision medicine. These biomolecular information provide a more comprehensive understanding of the pathogenesis of Alzheimer's disease (AD), and also offer favorable conditions for the diagnosis, prediction, and treatment of AD. In recent years, researchers have been continuously searching for more accurate biomarkers and exploring treatment options for AD. With the deepening of epigenetic research in the life sciences, discovering epigenetic changes with research value in AD has also become a research focus.
[0004] DNA methylation is a crucial component of epigenetics. Researchers have explored the DNA methylation processes of disease-causing genes in Alzheimer's disease (AD), elucidating the roles of gene methylation modifications in influencing Aβ and tau protein expression, and potentially affecting the normal expression and function of various genes. However, AD is a complex, multifactorial disease that may involve the modification and regulation of many key genes during its pathogenesis, and the underlying molecular regulatory network remains to be elucidated. Therefore, further exploration of potential risk genes, modification processes, and their specific roles in Alzheimer's disease is essential.
[0005] Mendelian randomization (MR) is a genetic approach that uses genetic variation as an instrumental variable to study how exposure factors affect outcome events, elucidating the causal relationship between exposure (factor) and outcome (event). Randomized controlled trials are impractical for assessing causality, and observational studies may produce biased association results due to confounding or reverse causality. MR addresses these issues by using genetic variation as an instrumental variable for the exposure under study, ensuring that the alleles of these exposure-related genetic variations are randomly assigned and unaffected by reverse causality.
[0006] Quantitative trait loci (QTLs) can serve as functional intermediates for studying the underlying biological mechanisms of disease inheritance. Expression quantitative trait loci (eQTLs), methylation quantitative trait loci (mQTLs), and protein quantitative trait loci (pQTLs), derived from multi-omics integrated data, reveal changes in gene expression, methylation, and protein expression levels. By selecting QTL-related single nucleotide polymorphisms (SNPs) as instrumental variables (IVs) and performing summary-data-based Mendelian randomization (SMR), direct causal relationships between certain gene methylation, expression, or protein levels and Alzheimer's disease (AD) can be inferred. Currently, some studies have used pQTL to obtain biomarkers for mental illnesses through a two-sample Mendelian randomization method, but how to screen biomarkers and therapeutic targets for Alzheimer's disease at the multi-omics level is still unclear. Summary of the Invention
[0007] To overcome the shortcomings of existing technologies, this invention provides a method and system for screening diagnostic biomarkers for Alzheimer's disease (AD). Based on Mendelian randomization (SMR) analysis of pooled data, combined with multi-omics research methods (risk gene methylation, expression, and protein abundance information), it identifies key genes in AD pathogenesis and elucidates their regulatory mechanisms in AD development. The method and system provided by this invention further refine the diagnostic biomarkers and therapeutic targets for AD, deepen the understanding of the molecular mechanisms in AD, identify new therapeutic targets, provide innovative strategies and tools for personalized treatment, and ultimately improve patients' quality of life.
[0008] To achieve the above technical objectives, the present invention adopts the following technical solution:
[0009] In a first aspect, the present invention provides a method for screening diagnostic biomarkers for Alzheimer's disease, the method comprising the following steps:
[0010] (1) Obtain GWAS data of hereditary trait variant sites (SNPs) associated with Alzheimer's disease; and simultaneously obtain multi-omics genetic information data of healthy individuals, including gene methylation genetic information data, gene expression genetic information data, and protein abundance genetic information data.
[0011] (2) Based on the summary data, Mendelian randomization was used to analyze GWAS data and multi-omics genetic information data to obtain gene methylation biomarkers, gene expression biomarkers and protein abundance biomarkers.
[0012] (3) Perform hierarchical analysis on the omics biomarkers obtained in step (2) to obtain diagnostic biomarkers for Alzheimer's disease.
[0013] In one or more embodiments, a method for obtaining GWAS data of genetic trait variant sites (SNPs) associated with Alzheimer's disease includes: performing genomic variation detection to correlate genetic variations in Alzheimer's disease populations with health data from a healthy subject population to obtain GWAS data of genetic trait variant sites (SNPs) associated with Alzheimer's disease.
[0014] Preferably, genomic information of Alzheimer's patients and healthy subjects is obtained separately, the sample reads are aligned to a reference genome, and then variant sites are detected by SNP calling, and the genotype of each variant site in an individual is identified by genotype calling.
[0015] In one or more embodiments, methods for obtaining gene methylomics genetic information data include:
[0016] Whole blood samples were obtained from the first batch of healthy individuals, and whole blood methylation was measured using an Illumina HumanMethylation450 array to screen for highly correlated Top Cis-mQTLs.
[0017] Preferably, during the Top Cis-mQTL screening process, the screening gene interval is ±1000kb, and the P-value threshold is 5.0×10⁻⁶. -8 And remove data with an allele frequency difference >0.2.
[0018] In one or more embodiments, methods for acquiring gene expression omics genetic information data include:
[0019] Blood and peripheral blood mononuclear cell (PBMC) samples were obtained from a second batch of healthy individuals. Gene expression levels in the samples were analyzed using Illumina, Affyte™ U219, Affyte™ Hu-Ex v1.0 ST expression arrays and RNA-seq to screen for highly correlated Top Cis-eQTLs.
[0020] Preferably, during the Top Cis-eQTL screening process, the screening gene interval is ±1000kb, and the P-value threshold is 5.0×10⁻⁶. -8 And remove data with an allele frequency difference greater than 0.2.
[0021] In one or more embodiments, methods for acquiring proteomics genetic information data include:
[0022] Plasma proteins were obtained from a third batch of healthy individuals. The plasma levels were analyzed using SomaScan and subjected to GWAS to screen for highly correlated Top Cis-pQTLs.
[0023] Preferably, during the Top Cis-pQTL screening process, the screening gene interval is ±1000kb, and the P-value threshold is 5.0×10⁻⁶. -8 And remove data with an allele frequency difference greater than 0.2.
[0024] In one or more embodiments, methods for analyzing GWAS data and gene methylomics genetic information data based on Mendelian randomization of aggregated data to obtain gene methylomics biomarkers include:
[0025] Using gene methylomics genetic information data as the exposure factor and GWAS data as the outcome factor, Mendelian randomization analysis based on the summary data was conducted to obtain the hypothesis test standard value P-value.
[0026] Exposure factors with HEIDI > 0.01 were identified using the HEIDI test.
[0027] The Benjamini-Hochberg method was used to adjust the standard value of hypothesis testing (P-value) to obtain exposure factors with a false discovery rate (FDR) corrected P-value < 0.05.
[0028] Exposure factors with FDR-corrected P-value < 0.05 and HEIDI > 0.01 were selected as biomarkers for gene methylation.
[0029] In one or more embodiments, methods for analyzing GWAS data and gene expression genomics genetic information data based on Mendelian randomization of aggregated data to obtain gene expression genomics biomarkers include:
[0030] Using gene expression omics genetic information data as exposure factors and GWAS data as outcome factors, Mendelian randomization analysis based on the summary data was conducted to obtain the hypothesis test standard value P-value.
[0031] Exposure factors with HEIDI > 0.01 were identified using the HEIDI test.
[0032] The Benjamini-Hochberg method was used to adjust the standard value of hypothesis testing (P-value) to obtain exposure factors with a false discovery rate (FDR) corrected P-value < 0.05.
[0033] Exposure factors with FDR-corrected P-value < 0.05 and HEIDI > 0.01 were selected as biomarkers for gene expression.
[0034] In one or more embodiments, methods for analyzing GWAS data and proteomics genetic information data based on Mendelian randomization of aggregated data to obtain proteomics biomarkers include:
[0035] Using proteomics genetic information data as the exposure factor and GWAS data as the outcome factor, Mendelian randomization analysis based on the summary data was conducted to obtain the hypothesis test standard value (P-value).
[0036] Exposure factors with HEIDI > 0.01 were identified using the HEIDI test.
[0037] The Benjamini-Hochberg method was used to adjust the standard value of hypothesis testing (P-value) to obtain exposure factors with a false discovery rate (FDR) corrected P-value < 0.05.
[0038] Exposure factors with FDR-corrected P-value < 0.05 and HEIDI > 0.01 were selected as proteomics biomarkers.
[0039] In one or more embodiments, the method for obtaining Alzheimer's disease diagnostic biomarkers by performing hierarchical analysis on the various omics biomarkers obtained in step (2) includes:
[0040] The same markers among proteomics markers, gene methylation markers, and gene expression markers are classified as first-tier markers;
[0041] Markers that are the same as those in proteomics and gene methylation, or those that are the same as those in proteomics and gene expression, are classified as second-tier markers.
[0042] A second aspect of the present invention provides a screening system for diagnostic biomarkers of Alzheimer's disease, comprising:
[0043] The data acquisition module is used to acquire GWAS data of hereditary trait variant sites (SNPs) associated with Alzheimer's disease; and simultaneously acquire multi-omics genetic information data of healthy individuals, including gene methylation omics genetic information data, gene expression omics genetic information data, and proteomics genetic information data.
[0044] The data analysis module is used to analyze GWAS data and multi-omics genetic information data based on Mendelian randomization of the aggregated data to obtain gene methylation biomarkers, gene expression biomarkers, and proteomics biomarkers.
[0045] The biomarker acquisition module is used to perform hierarchical analysis on the various omics biomarkers acquired in the data analysis module to obtain diagnostic biomarkers for Alzheimer's disease.
[0046] A third aspect of the present invention provides a diagnostic biomarker for Alzheimer's disease, obtained by screening using the screening method described in the first aspect and / or the screening system described in the second aspect;
[0047] The diagnostic biomarkers for Alzheimer's disease include: TMEM106B, IBSP, FCRLB, SIGLEC9, CD33, LILRB1, PLIRA, ACE, GRN, NSF, SIRPA, LILRA4, UBASH3B, or LY6G6D.
[0048] The beneficial effects of this invention are as follows:
[0049] This invention utilizes genetic information data from three omics domains—gene methylation, gene expression, and protein abundance levels—from healthy individuals, combined with genetic information data from Alzheimer's disease patients, to perform a pooled data-based Mendelian randomization analysis. This analysis identifies key genes and corresponding single nucleotide polymorphisms (SNPs) in Alzheimer's disease, providing important references for predicting disease progression, diagnosis, and intervention treatment of Alzheimer's disease. Attached Figure Description
[0050] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0051] Figure 1 This is a forest plot of gene methylation and SMR results in Alzheimer's disease in Example 1 of the present invention;
[0052] Figure 2 In the diagrams A through F, the results of gene-Alzheimer's disease SMR in Example 1 of this invention are forest plots.
[0053] Figure 3 This is a forest plot of the protein-to-SMR results for Alzheimer's disease in Example 1 of the present invention;
[0054] Figure 4 This is an MR forest plot of the corresponding methylation sites and the ACE and CD33 genes in the primary evidence of Example 1 of this invention;
[0055] Figure 5 This is the result of multi-omics evidence integration in Embodiment 1 of the present invention;
[0056] Figure 6 The results of database verification in Example 2 of this invention are shown; where A is the verification of gene methylation level, B and C are the verification of gene expression level, and D is the verification of plasma protein level.
[0057] Figure 7 In Example 3 of this invention, A and B both describe the differential expression of important risk genes in AD and healthy individuals.
[0058] Figure 8 In both A and B, the results of sensitivity and specificity analysis of important risk genes described in Example 3 of this invention are as follows. Detailed Implementation
[0059] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0060] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0061] To enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention will be described in detail below with reference to specific embodiments.
[0062] Example 1
[0063] This embodiment provides a method for screening diagnostic biomarkers for Alzheimer's disease, the method comprising the following steps:
[0064] (1) Obtain GWAS data of hereditary trait variant sites (SNPs) associated with Alzheimer's disease; and simultaneously obtain multi-omics genetic information data of healthy individuals, including gene methylation genetic information data, gene expression genetic information data, and protein abundance genetic information data.
[0065] (2) Based on the summary data, Mendelian randomization was used to analyze GWAS data and multi-omics genetic information data to obtain gene methylation biomarkers, gene expression biomarkers and protein abundance biomarkers.
[0066] (3) Perform hierarchical analysis on the omics biomarkers obtained in step (2) to obtain diagnostic biomarkers for Alzheimer's disease.
[0067] In step (1), the method for obtaining GWAS data of genetic trait variant sites (SNPs) associated with Alzheimer's disease is as follows:
[0068] By performing genomic variation detection, genetic variations in Alzheimer's patients are correlated with health data from healthy subjects to obtain GWAS data on hereditary trait variant sites (SNPs) associated with Alzheimer's disease.
[0069] Specifically, the genetic variations in Alzheimer's patients and healthy participants were collected from the European Alzheimer & Dementia Biobank (EADB) consortium, comprising 85,934 AD patients and 401,577 healthy participants. Genomic information was obtained from both populations, sample reads were aligned to a reference genome, variant sites were detected using SNP calling, and the genotype of each variant site in each individual was identified using genotype calling.
[0070] In step (1), the methods for obtaining gene methylation omics genetic information data include:
[0071] Whole blood samples were obtained from the first batch of healthy individuals, and whole blood methylation was measured using an Illumina HumanMethylation450 array to screen for highly correlated Top Cis-mQTLs.
[0072] Specifically:
[0073] The first batch of healthy human samples came from the following databases (Brisbane Systems Genetics Study (BSGS) and Lothian Birth Cohorts (LBC)), totaling 1,980 subjects. Whole blood samples were obtained from the subjects, and whole blood methylation was measured using the Illumina HumanMethylation450 array, yielding 52,916 cis-DNA methylation quantitative trait sites (mQTLs).
[0074] Based on the collected mQTL data, the gene selection interval was ±1000kb, and the P-value threshold was 5.0×10⁻⁶. -8 Furthermore, data with an allele frequency difference > 0.2 were removed as highly correlated Top Cis-mQTLs. In this process, any SNPs with an allele frequency difference greater than a specified threshold (0.2) between paired datasets were excluded.
[0075] In step (1), the method for obtaining gene expression omics genetic information data is as follows:
[0076] Blood and peripheral blood mononuclear cell (PBMC) samples were obtained from a second batch of healthy individuals and preprocessed in a standardized manner. Gene expression levels of the samples were analyzed using Illumina, Affyte™ U219, Affyte™ Hu-Ex v1.0 ST expression arrays and RNA-seq to screen for highly correlated Top Cis-eQTLs.
[0077] Specifically:
[0078] The first batch of healthy human samples came from the eQTLGen consortium database, comprising 31,684 subjects. Blood and peripheral blood mononuclear cell (PBMC) samples were obtained and preprocessed using standardized methods. Gene expression levels were analyzed using Illumina, Affyte™ U219, Affyte™ Hu-Ex v1.0 ST expression arrays, and RNA-seq. Genotyping and expression data preprocessing, PGS calculation, and cis-eQTL, trans-eQTL, and eQTL mapping were performed. Standardized preprocessing included genotype inference, data harmonization, SNP screening, and labeling using specific database versions (using a unified genome version; for GRCh37 / hg19 data, the GRCh37 genome version was used; for GRCh38 / hg38 data, the GRCh38 genome version was used) to ensure the accuracy and reliability of the study.
[0079] The gene selection interval was ±1000kb, and the P-value threshold was 5.0×10⁻⁶. -8 Data with a threshold of allele frequency difference >0.2 were removed as highly correlated Top Cis-eQTLs.
[0080] In step (1), the methods for obtaining proteomics genetic information data include:
[0081] Plasma proteins were obtained from a third batch of healthy individuals. The plasma levels were analyzed using SomaScan and subjected to GWAS to screen for highly correlated Top Cis-pQTLs.
[0082] Specifically:
[0083] The third batch of healthy individuals was obtained from the Iceland Cancer Project and the deCODE genetics, Reykjavík, Iceland project databases, totaling 35,559 participants. Plasma proteins were extracted and isolated, and GWAS analysis was performed on plasma levels of 27 million variants and 4,907 aptamers using SomaScan. This identified 28,191 pQTL associations, of which 27% were cis, 73% were trans, and 93% were unreported. pQTLs were associated with the levels of one or more proteins.
[0084] Data with a gene range of ±1000kb, a P-value threshold of 5.0×10-8, and a removal threshold of allele frequency difference >0.2 were considered highly correlated Top Cis-pQTLs.
[0085] In step (2), the GWAS data and gene methylomics genetic information data are analyzed based on Mendelian randomization of the summarized data to obtain gene methylomics biomarkers. The methods include:
[0086] Using gene methylomics genetic information data as the exposure factor and GWAS data as the outcome factor, Mendelian randomization analysis based on the summary data was conducted to obtain the hypothesis test standard value P-value.
[0087] Exposure factors with HEIDI > 0.01 were identified using the HEIDI test.
[0088] The Benjamini-Hochberg method was used to adjust the standard value of hypothesis testing (P-value) to obtain exposure factors with a false discovery rate (FDR) corrected P-value < 0.05.
[0089] Exposure factors with FDR-corrected P-value < 0.05 and HEIDI > 0.01 were selected as biomarkers for gene methylation.
[0090] In step (2), the GWAS data and gene expression genomics genetic information data are analyzed based on Mendelian randomization of the summarized data to obtain gene expression genomics biomarkers. The methods include:
[0091] Using gene expression omics genetic information data as exposure factors and GWAS data as outcome factors, Mendelian randomization analysis based on the summary data was conducted to obtain the hypothesis test standard value P-value.
[0092] Exposure factors with HEIDI > 0.01 were identified using the HEIDI test.
[0093] The Benjamini-Hochberg method was used to adjust the standard value of hypothesis testing (P-value) to obtain exposure factors with a false discovery rate (FDR) corrected P-value < 0.05.
[0094] Exposure factors with FDR-corrected P-value < 0.05 and HEIDI > 0.01 were selected as biomarkers for gene expression.
[0095] In step (2), the GWAS data and proteomics genetic information data are analyzed based on Mendelian randomization of the summarized data to obtain proteomics biomarkers. The methods include:
[0096] Using proteomics genetic information data as the exposure factor and GWAS data as the outcome factor, Mendelian randomization analysis based on the summary data was conducted to obtain the hypothesis test standard value (P-value).
[0097] Exposure factors with HEIDI > 0.01 were identified using the HEIDI test.
[0098] The Benjamini-Hochberg method was used to adjust the standard value of hypothesis testing (P-value) to obtain exposure factors with a false discovery rate (FDR) corrected P-value < 0.05.
[0099] Exposure factors with FDR-corrected P-value < 0.05 and HEIDI > 0.01 were selected as proteomics biomarkers.
[0100] In step (2) above, the formula for obtaining exposure factors with HEIDI > 0.01 using the HEIDI test is as follows:
[0101]
[0102] The subscript "0" indicates causal variation, and the subscript "i" indicates SNP i, r 0i It is a linkage disequilibrium correlation between the causal variant and SNP i, and h = 2p(1-p), where p is the allele frequency. Therefore, a linkage test for pleiotropy (two distinct causal variants, one affecting gene expression and the other affecting the trait) is equivalent to testing b. xx(top) (b estimated using top-related cis-eQTL) xy ) and b xy(i) (b estimated using any other significant SNP in the cis-eQTL region) xy Are there differences between )? If we define d i =b xy(i) —b xy(top) This is further equivalent to testing the d values of all significant cis-eQTLs (excluding the top cis-eQTLs). i Is it equal to 0? If we define... Where m is the number of significant eQTLs (excluding top cis-eQTLs), making Here, V is the (co)variance matrix, where the ij-th element is
[0103]
[0104] and The covariance between them is calculated as follows:
[0105]
[0106] Where r ij This represents the linkage disequilibrium correlation between SNPs i and j. Under the null hypothesis of no heterogeneity, i.e., d = 0, the standard normal vector is represented as z. d ={z d(1) ,z d(2) ,…,z d(m)}, z d ~MVN(0,R), where R is the correlation matrix, can be represented as To test for d=0, a HEIDI test is constructed. Right now Where I is the identity matrix.
[0107] The HEIDI test was used to distinguish between true conjoint effects and non-conjoint effects. Cases with a p-HEIDI < 0.01 were considered non-conjoint effects and therefore excluded from the analysis. This procedure ensured the reliability and accuracy of our analytical results.
[0108] In step (2) above, the method for adjusting the hypothesis test standard value P-value using the Benjamini-Hochberg method is as follows:
[0109] ① Sort all P-values in ascending order;
[0110] ② Assign rankings to P-values: the smallest P-value is ranked 1, the second smallest is ranked 2, and so on;
[0111] ③ Calculate the BH critical value for each p-value using the formula: (i / m)Q; where i is the rank of the p-value, m is the total number of tests, and Q is the false positive rate (FDR, a percentage chosen by the researcher).
[0112] ④ Compare the original P-value with the BH critical value calculated in step 3: find the largest P-value that is "less than its corresponding BH critical value".
[0113] The method for correcting the false discovery rate (FDR) is as follows:
[0114] The results of a total of m tests are sorted in ascending order, where k is the rank corresponding to the P-value of one of the test results.
[0115] Find the largest k value that meets the original threshold α, satisfying P(k)<=α*k / m. Consider all tests from rank 1 to k to have significant differences, and calculate the corresponding q value using the formula q=p*(m / k).
[0116] In step (2) above, before analyzing GWAS data and multi-omics genetic information data based on Mendelian randomization of the summarized data to obtain gene methylation biomarkers, gene expression biomarkers, and proteomics biomarkers, colocalization analysis is performed using the "coloc" R package to determine whether two or more phenotypes (such as diseases or molecular phenotypes) are affected by the same causal variant site in a specific region. Colocalization analysis yields four scenarios: (1) neither phenotype is significantly associated with any SNP site in the genomic region (H0); (2) one phenotype is significantly associated with an SNP site in a genomic region (H1); (3) both phenotypes are significantly associated with an SNP site in a genomic region, but driven by different causal variant sites (H3); (4) both phenotypes are significantly associated with an SNP site in a genomic region, and driven by the same causal variant site (H4). For mQTL-GWAS colocalization, the colocalization region window is ±500kb. The prior probabilities of a SNP being associated only with trait 1 (mQTL), a SNP being associated only with trait 2 (AD), and a SNP being associated with both traits are set to 10, respectively. -4 10 -4 and 10 -5 In our study, a posterior probability H4 (PP.H4) greater than 0.5 indicates successful co-localization, suggesting that the two phenotypes are driven by the same causal variant. For eQTL-GWAS co-localization, the co-localization region window was ±1000 kb. The prior probabilities of SNPs associated only with trait 1 (eQTL), SNPs associated only with trait 2 (AD), and SNPs associated with both traits were set to 10, respectively. -4 10 -4 and 10 -5 In our study, a posterior probability H4 (PP.H4) greater than 0.5 indicates successful co-localization, suggesting that the two phenotypes are driven by the same causal variant. For pQTL-GWAS co-localization, the co-localization region window was ±1000 kb. The prior probabilities of SNPs associated only with trait 1 (pQTL), SNPs associated only with trait 2 (AD), and SNPs associated with both traits were set to 10, respectively. -4 10 -4 and 10 -5 In our study, a posterior probability H4 (PP.H4) greater than 0.5 indicates successful co-localization, suggesting that the two phenotypes are driven by the same causal variants.
[0117] result:
[0118] Visualizing causal gene methylation information using forest plots: The "forestploter" package in R was used to visualize the methylation and related information of genes with a strong causal relationship (P<0.05) with AD. Specific information included methylation sites, genes, SNP sites, forest plots, odds ratios (OR), 95% confidence intervals, hypothesis testing standard values (P-values), and co-localization results (PP.H4). Ultimately, 39 CpG sites (25 genes) were found to be associated with the development of AD (Table 1). Among these, multiple CpG sites within the same gene can predict its association with AD risk. For example, methylation predicted at 24 CpG sites on the TMEM184A gene (cg00411097,cg01726448,cg02879083,cg03564557,cg03664589,cg03723321,cg03958363,cg04501217,cg04633141,cg05868365) can be used. Genes (cg06194808, cg07071809, cg08729755, cg09329079, cg09495282, cg09990600, cg10633906, cg12503394, cg14480560, cg18294109, cg18648645, cg18705808, cg21368161, cg23404435) have been found to either increase or decrease the risk of developing Alzheimer's disease (AD). Based on the odds ratio (OR) values, different gene methylations may increase or decrease the risk of AD. For some genes, methylation predicted by different CpG sites may play diametrically opposed roles in disease development. For example, in the RSPO4 gene, methylation at cg00862894 is predicted to increase the risk of Alzheimer's disease (AD), while tests at cg03569404, cg04199533, and cg26278454 show that RSPO4 methylation reduces the risk. Figure 1 ).
[0119] Table 1: AD biomarkers in gene methylomics
[0120]
[0121]
[0122] Visualizing gene expression information with causal relationships using forest plots: The "forestploter" package in R program is used to visualize genes and related information that have a strong causal relationship with AD (P<0.05). Specific information includes related genes, SNP sites, forest plots, odds ratios (OR), 95% confidence intervals, and hypothesis testing standard values (P-values).
[0123] A total of 125 genes were found to have a causal relationship between their expression levels and the onset of Alzheimer's disease (AD). Elevated expression levels of 61 genes were found to increase the risk of AD (OR>1), while high expression levels of the remaining 64 genes may decrease the risk of AD. Figure 2 ), of which 125 genes and loci are shown in Table 2:
[0124] Table 2: Biomarkers of AD at Genomic Expression Levels
[0125]
[0126]
[0127]
[0128] Visualizing causal protein expression information using forest plots: The abundance and related information of proteins with a strong causal relationship with AD (P<0.05) are visualized using the "forestploter" package in the R program. Specific information includes related proteins, SNP sites, forest plots, odds ratio (OR), 95% confidence intervals, and hypothesis test standard values (P-values).
[0129] Ultimately, 18 risk-related proteins for AD were identified (Table 3). Genetic prediction showed that high expression levels of eight proteins increased the risk of Alzheimer's disease (AD): TMEM106B (OR 1.44 CI 1.24, 1.68), IBSP (OR 1.35 CI 1.13, 1.61), FCRLB (OR 1.11 CI 1.04, 1.17), BPNT2 (OR 1.11 CI 1.05, 1.17), CTSH (OR 1.04 CI 1.03, 1.06), SIGLEC9 (OR 1.03 CI 1.02, 1.04), CD33 (OR 1.03 CI 1.01, 1.05), and SIRPA (OR 1.03 CI 1.02, 1.04). Increased expression levels of ten proteins were associated with a decreased prevalence of AD: LILRB1 (OR 0.95 CI 0.93, 0.97), PILRA (OR 0.94 CI...). 0.93,0.96),ACE(OR 0.92CI 0.89,0.94),LGALS3(OR 0.91CI 0.87,0.96),LILRA4(OR 0.9CI 0.85,0.95),UBASH3B(OR 0.77CI 0.66,0.88),GRN(OR 0.74CI 0.68,0.81),CLN5(OR 0.69CI 0.58,0.83),NSF(OR 0.6CI 0.5,0.72),LY6G6D(OR 0.27CI 0.15,0.48)( Figure 3 ).
[0130] Table 3: Risk-related proteins for 18 AD markers
[0131]
[0132] In step (3), to further improve accuracy, SMR data of different omics and AD are organized to form evidence at different levels. Through the previous SMR studies, the correlation between gene methylation, gene expression, the abundance level of proteins after gene translation and AD pathogenesis has been obtained. That is, from the perspective of three omics, it is explained that there is a causal relationship between changes in the expression of certain genes and the pathological progression of AD. The regulation and expression of genes is a complex and continuous process. To prove that a certain gene itself plays a key role in the pathological progression of the disease, a complete chain of evidence must be formed. Each piece of evidence is obtained by combining statistically significant data results (SMR1: gene methylation and AD; SMR2: gene expression level and AD; SMR3: protein abundance level and AD). The more rigorous the data acquisition process to support the evidence, the more persuasive the evidence.
[0133] Therefore, the obtained causal relationships were organized and summarized into two different tiers of evidence. Proteins are the basic functional units and have the most direct relationship with disease; therefore, the causal relationship between protein abundance levels and disease was taken as the most fundamental common condition.
[0134] The study aims to determine the causal relationship between gene methylation and gene expression, and between gene expression and protein abundance, by using SMR to infer how gene methylation affects gene expression and to supplement the obtained evidence.
[0135] Results: The same markers among proteomics markers, gene methylation markers, and gene expression markers were classified as first-tier markers;
[0136] Markers that are the same as those in proteomics and gene methylation, or those that are the same as those in proteomics and gene expression, are classified as second-tier markers.
[0137] The first tier of evidence: There is a strong causal relationship between protein abundance level and AD (FDR-corrected P-value < 0.05 and P-HEIDI > 0.01), and there is also a strong causal relationship between gene methylation and gene expression level and AD (FDR-corrected P-value < 0.05 and P-HEIDI > 0.01);
[0138] The second tier of evidence shows a strong causal relationship between protein abundance levels and AD (FDR-corrected P-value < 0.05 and P-HEIDI > 0.01) and a strong causal relationship between gene methylation or gene expression levels and AD (FDR-corrected P-value < 0.05 and P-HEIDI > 0.01).
[0139] By integrating and grading multi-omics evidence, the persuasiveness and accuracy of some evidence can be enhanced, and the underlying molecular mechanisms can be preliminarily inferred. Figure 5 The causal relationship between the two key genes ACE and CD33 and AD was classified as Level 1 evidence because these two genes are causally related to AD progression in terms of methylation, expression levels, and abundance of related proteins. This causal relationship meets the multiple correction criteria, excludes in vitro effects in HEIDI verification, and shows good co-localization with AD at the three omics levels. Therefore, it is the most convincing and can be considered a key research target. Figure 4 ).
[0140] The causal relationship between the levels of five gene-related proteins—CTSH, SIRPA, TMEM106B, FCRLB, and UBASH3B—and the progression of Alzheimer's disease can be demonstrated by SMR (FDR-corrected P-value < 0.05 and P-HEIDI > 0.01), which meets the most basic assumption of protein-function influence. However, due to the lack of evidence regarding corresponding methylation or gene expression levels, it is classified as Level 2 evidence. Figure 6 ).
[0141] Elevated CLN5 protein levels may be associated with a reduced probability of Alzheimer's disease (AD) (OR 0.96, CI 0.58, 0.83). However, due to a lack of evidence regarding gene expression or methylation levels, this cannot be classified as Level 1 or Level 2 evidence. Nevertheless, the co-localization results between CLN5 protein levels and AD indicate that the causal relationship is driven by the same variant site, reinforcing the causal link between CLN5 protein levels and AD. Therefore, the crucial role of CLN5 in AD should not be overlooked and warrants further research.
[0142] The discovery of the aforementioned key genes, methylation sites, and corresponding SNPs provides insights into the molecular mechanisms of Alzheimer's disease and is of great research significance.
[0143] Example 2
[0144] To verify the feasibility of the above-obtained markers, this embodiment provides a method for verifying the three aspects of the obtained markers using a database.
[0145] Exposure data: mQTL (methylated sequences obtained from sequencing by McRae et al.), eQTL (obtained from the eQTLGen consortium), pQTL (protein quantitative trait loci study by Ferkingstad et al.); Outcome data: The FinnGenStudy G6_AD_WIDE. There was no overlap between the exposed and outcome populations.
[0146] Specifically,
[0147] (1) Obtain the validation data required in this study:
[0148] AD Data: Validation data for AD was obtained from the Finnish biobank (Kurki, MI, Karjalainen, J., Palta, P. et al. FinnGen provides genetic insights from a well-phenotyped isolated population. Nature 613, 508–518 (2023). https: / / doi.org / 10.1038 / s41586-022-05473-8). The FinnGen study is a large-scale genomics project that has analyzed over 500,000 samples from the Finnish Biobank and correlated genetic variations with health data to understand disease mechanisms and susceptibility. The dataset used in this study was “G6_AD_WIDE,” which contains genetic variation information from 15,617 AD patients and 396,564 healthy subjects.
[0149] mQTL data are derived from methylated sequences that have already been sequenced by McRae et al.
[0150] eQTL data is obtained from the eQTLGen consortium.
[0151] pQTL data comes from a study by Ferkingstad et al. on protein quantitative trait loci (pQTL) in individuals from Iceland.
[0152] (2) Gene methylation biomarkers, gene expression biomarkers and protein abundance biomarkers were obtained using the same processing method as in Example 1.
[0153] result:
[0154] Visualizing molecular information with causal relationships using forest plots: The "forestploter" package in R is used to visualize molecular information with a strong causal relationship (P<0.05) with AD. Specific information includes methylation sites, gene expression, related proteins, SNP sites, forest plots, odds ratio (OR), 95% confidence intervals, and hypothesis testing standard values (P-values).
[0155] At the methylation level, the causal relationship between four CpG sites (three genes) and AD was successfully verified; at the gene expression level, the causal relationship between five genes and AD was successfully verified; and at the protein abundance level, the causal relationship between five plasma proteins and AD was successfully verified.
[0156] At the methylation level, 12 CpG sites (11 genes) were found to be associated with the development of AD. At the gene level, a total of 52 genes were found to have a causal relationship between their expression levels and the onset of AD. We found that elevated expression levels of 19 genes increased the risk of AD (OR>1), while high expression levels of the remaining 33 genes may have a decreased risk of AD. The causal relationship between four CpG sites (3 genes) and AD was successfully verified: cg26167930 (DAPK2), cg00443543 (SERPINF2), cg18278729 (SERPINF2), and cg23422036 (TMEM106B). Figure 6 A). The causal relationship between five genes and AD was successfully verified: GRINA, RITA1, APH1B, ACE, SIGLEC18P, and GRINA. Figure 6 B). At the protein abundance level, 14 risk-related proteins for AD were ultimately identified, among which the causal relationship between 5 proteins and AD was verified: PILRA, GRN, ACE, SIGLEC9, and CD33. Figure 6 C).
[0157] Example 3
[0158] To further determine the diagnostic efficacy of the above biomarkers, ROC analysis was performed.
[0159] Data from post-mortem temporal cortical brain samples of 71 individuals investigated in a study conducted by the Neuro Center, Kuopio University Hospital, included expression matrices and clinical data. This data comprised transcript, protein, and phosphopeptide expression information from temporal lobe tissues of 54 AD patients and 6 healthy individuals. After grouping and annotating the data with probes, the expression of 17 genes—TMEM106B, IBSP, FCRLB, CTSH, SIGLEC9, CD33, LILRB1, PILRA, ACE, LGALS3, GRN, CLN5, NSF, SIRPA, LILRA4, UBASH3B, and LY6G6D—in 60 samples was further analyzed (BPNT2 expression was not present in this dataset and was therefore ignored). Wilcox analysis revealed differences in the expression of some genes between the two populations, which were visualized using the R package "ggplot" and ROC curves and AUC calculations were performed using the R package "pROC".
[0160] The expression of six proteins—CD33, LILRA4, NSF, PILRA, SIRPA, and TMEM106B—did differ significantly between AD and healthy individuals. Figure 7 The study found that the AUC of 17 genes was above 0.5, indicating that these genes all have certain diagnostic capabilities. Among them, 14 genes, including TMEM106B, IBSP, FCRLB, SIGLEC9, CD33, LILRB1, PLIRA, ACE, GRN, NSF, SIRPA, LILRA4, UBASH3B, and LY6G6D, had an AUC above 0.6, demonstrating high accuracy. Figure 8 ).
[0161] Example 4
[0162] A screening system for diagnostic biomarkers of Alzheimer's disease includes:
[0163] The data acquisition module is used to acquire GWAS data of hereditary trait variant sites (SNPs) associated with Alzheimer's disease; and simultaneously acquire multi-omics genetic information data of healthy individuals, including gene methylation omics genetic information data, gene expression omics genetic information data, and proteomics genetic information data.
[0164] The data analysis module is used to analyze GWAS data and multi-omics genetic information data based on Mendelian randomization of the aggregated data to obtain gene methylation biomarkers, gene expression biomarkers, and proteomics biomarkers.
[0165] The biomarker acquisition module is used to perform hierarchical analysis on the various omics biomarkers acquired in the data analysis module to obtain diagnostic biomarkers for Alzheimer's disease.
[0166] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for screening diagnostic biomarkers for Alzheimer's disease, characterized in that, The method includes the following steps: (1) Obtain GWAS data of genetic trait variation sites associated with Alzheimer's disease; and simultaneously obtain multi-omics genetic information data of healthy individuals, including gene methylation genetic information data, gene expression genetic information data, and protein abundance genetic information data; Among them, the samples for gene methylation omics genetic information data were whole blood samples from the first batch of healthy individuals; the samples for gene expression omics genetic information data were blood and peripheral blood mononuclear cell samples from the second batch of healthy individuals; and the samples for protein abundance omics genetic information data were plasma proteins from the third batch of healthy individuals. (2) Based on the Mendelian randomization of the summary data, GWAS data and multi-omics genetic information data were analyzed to obtain gene methylation biomarkers, gene expression biomarkers and protein abundance biomarkers. Methods for obtaining gene methylation biomarkers include: Using gene methylomics genetic information data as the exposure factor and GWAS data as the outcome factor, Mendelian randomization analysis based on the summary data was conducted to obtain the P-value. Exposure factors with HEIDI > 0.01 were identified using the HEIDI test. The Benjamini-Hochberg method was used to adjust the hypothesis test standard value P-value to obtain exposure factors with a false discovery rate corrected FDR (Frequency Detection Rate) of <0.
05. Exposure factors with FDR-corrected P-value < 0.05 and HEIDI > 0.01 were selected as biomarkers for gene methylation. Methods for obtaining gene expression biomarkers include: Using gene expression omics genetic information data as exposure factors and GWAS data as outcome factors, Mendelian randomization analysis based on the summary data was conducted to obtain the hypothesis test standard value P-value. Exposure factors with HEIDI > 0.01 were identified using the HEIDI test. The Benjamini-Hochberg method was used to adjust the hypothesis test standard value P-value to obtain exposure factors with a false discovery rate corrected FDR (Frequency Detection Rate) of <0.
05. Exposure factors with FDR-corrected P-value < 0.05 and HEIDI > 0.01 were screened as biomarkers for gene expression genomics. Methods for obtaining proteomics biomarkers by analyzing GWAS data and proteomics genetic information data using Mendelian randomization based on aggregated data include: Using proteomics genetic information data as exposure factors and GWAS data as outcome factors, Mendelian randomization analysis based on the summary data was conducted to obtain the hypothesis test standard value (P-value). Exposure factors with HEIDI > 0.01 were identified using the HEIDI test. The Benjamini-Hochberg method was used to adjust the hypothesis test standard value P-value to obtain exposure factors with a false discovery rate corrected FDR (Frequency Detection Rate) of <0.
05. Exposure factors with FDR-corrected P-value < 0.05 and HEIDI > 0.01 were screened as proteomics biomarkers; (3) Perform hierarchical analysis on the omics biomarkers obtained in step (2) to obtain diagnostic biomarkers for Alzheimer's disease; The method for obtaining diagnostic biomarkers for Alzheimer's disease includes: The same markers among proteomics markers, gene methylation markers, and gene expression markers are classified as first-tier markers; Markers that are the same as those in proteomics and gene methylation, or those that are the same as those in proteomics and gene expression, are classified as second-tier markers.
2. The screening method as described in claim 1, characterized in that, Methods for obtaining GWAS data on genetic trait variants associated with Alzheimer's disease include: performing genomic variation detection, associating genetic variations in Alzheimer's patients with health data from healthy subjects, and obtaining GWAS data on genetic trait variants associated with Alzheimer's disease.
3. The screening method as described in claim 2, characterized in that, Genomic information was obtained from individuals with Alzheimer's disease and healthy subjects. The sample reads were aligned to a reference genome. Then, variant sites were detected by SNP calling, and the genotype of each variant site in an individual was identified by genotype calling.
4. The screening method as described in claim 1, characterized in that, Methods for obtaining genetic information data from gene methylation omics include: Whole blood methylation was measured using an Illumina HumanMethylation450 array, and highly correlated TopCis-mQTLs were screened.
5. The screening method as described in claim 4, characterized in that, During the Top Cis-mQTL screening process, the gene selection interval was ±1000 kb, and the P-value threshold was 5.0 × 10⁻⁶. -8 And remove data with an allele frequency difference greater than 0.
2.
6. The screening method as described in claim 1, characterized in that, Methods for obtaining genetic information data from gene expression omics include: Gene expression levels in the samples were analyzed using Illumina, Affyte™ U219, Affyte™ Hu-Ex v1.0 ST expression arrays, and RNA-seq to screen for highly correlated Top Cis-eQTLs.
7. The screening method as described in claim 6, characterized in that, During the Top Cis-eQTL screening process, the gene selection interval was ±1000 kb, and the P-value threshold was 5.0 × 10⁻⁶. -8 And remove data with an allele frequency difference greater than 0.
2.
8. The screening method as described in claim 1, characterized in that, Methods for obtaining proteomics genetic information data include: Highly correlated Top Cis-pQTLs were identified by GWAS screening of plasma levels using SomaScan.
9. The screening method as described in claim 8, characterized in that, During the Top Cis-pQTL screening process, the gene selection interval was ±1000 kb, and the P-value threshold was 5.0 × 10⁻⁶. -8 And remove data with an allele frequency difference greater than 0.
2.
10. A screening system for diagnostic biomarkers of Alzheimer's disease, characterized in that, include: The data acquisition module is used to acquire GWAS data of genetic trait variation sites related to Alzheimer's disease; at the same time, it acquires multi-omics genetic information data of healthy individuals, including gene methylation omics genetic information data, gene expression omics genetic information data, and protein abundance omics genetic information data. Among them, the samples for gene methylation omics genetic information data were whole blood samples from the first batch of healthy individuals; the samples for gene expression omics genetic information data were blood and peripheral blood mononuclear cell samples from the second batch of healthy individuals; and the samples for protein abundance omics genetic information data were plasma proteins from the third batch of healthy individuals. The data analysis module is used to analyze GWAS data and multi-omics genetic information data based on Mendelian randomization of the aggregated data to obtain gene methylation biomarkers, gene expression biomarkers, and proteomics biomarkers. Methods for obtaining gene methylation biomarkers include: Using gene methylomics genetic information data as the exposure factor and GWAS data as the outcome factor, Mendelian randomization analysis based on the summary data was conducted to obtain the P-value. Exposure factors with HEIDI > 0.01 were identified using the HEIDI test. The Benjamini-Hochberg method was used to adjust the hypothesis test standard value P-value to obtain exposure factors with a false discovery rate corrected FDR (Frequency Detection Rate) of <0.
05. Exposure factors with FDR-corrected P-value < 0.05 and HEIDI > 0.01 were selected as biomarkers for gene methylation. Methods for obtaining gene expression biomarkers include: Using gene expression omics genetic information data as exposure factors and GWAS data as outcome factors, Mendelian randomization analysis based on the summary data was conducted to obtain the hypothesis test standard value P-value. Exposure factors with HEIDI > 0.01 were identified using the HEIDI test. The Benjamini-Hochberg method was used to adjust the hypothesis test standard value P-value to obtain exposure factors with a false discovery rate corrected FDR (Frequency Detection Rate) of <0.
05. Exposure factors with FDR-corrected P-value < 0.05 and HEIDI > 0.01 were screened as biomarkers for gene expression genomics. Methods for obtaining proteomics biomarkers by analyzing GWAS data and proteomics genetic information data using Mendelian randomization based on aggregated data include: Using proteomics genetic information data as exposure factors and GWAS data as outcome factors, Mendelian randomization analysis based on the summary data was conducted to obtain the hypothesis test standard value (P-value). Exposure factors with HEIDI > 0.01 were identified using the HEIDI test. The Benjamini-Hochberg method was used to adjust the hypothesis test standard value P-value to obtain exposure factors with a false discovery rate corrected FDR (Frequency Detection Rate) of <0.
05. Exposure factors with FDR-corrected P-value < 0.05 and HEIDI > 0.01 were screened as proteomics biomarkers; The biomarker acquisition module is used to perform hierarchical analysis on the various omics biomarkers acquired in the data analysis module to obtain diagnostic biomarkers for Alzheimer's disease. The method for obtaining diagnostic biomarkers for Alzheimer's disease includes: The same markers among proteomics markers, gene methylation markers, and gene expression markers are classified as first-tier markers; Markers that are the same as those in proteomics and gene methylation, or those that are the same as those in proteomics and gene expression, are classified as second-tier markers.