A method and system for screening pathogenic genes and biomarkers of retinopathy
Through multi-dimensional data analysis, combined with GWAS and multi-tissue-specific gene expression weight data, the pathogenic genes and biomarkers of retinopathy were screened out, which solved the problems of complex interpretation of GWAS and limited TWAS data in the existing technology, achieved scientific and reliable screening of pathogenic genes and biomarkers, and promoted the progress of retinopathy research.
Patent Information
- Application Number
- CN202411222355.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-02
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-09-02
AI Technical Summary
When screening pathogenic genes and biomarkers of diabetic retinopathy, the prior art has problems such as complex interpretation of GWAS results, limited sources of TWAS data, and inability to identify the causal relationship of the disease in PWAS, and lacks effective screening methods.
Multidimensional data analysis method was used, combined with GWAS summary data and multi-tissue-specific gene expression weight data, transcriptome association analysis and Mendel randomization analysis were performed, and combined with cis-eQTL or pQTL data, and MR-Phewas screening, pathogenic genes and biomarkers related to retinopathy were obtained.
Scientifically and reliably excavate significantly related pathogenic genes and biomarkers, promote the progress of retinopathy research, and provide important scientific data support for clinical practice.
Smart Images

Figure CN119207575B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information processing and analysis, and particularly to a method and system for screening pathogenic genes and biomarkers of retinopathy. Background Art
[0002] In recent years, the incidence of diabetes has been increasing year by year, and it is estimated that the number of global diabetes patients will increase to about 800 million in 2045. Diabetes can lead to many serious eye complications, including diabetic retinopathy (DR), neovascular glaucoma, and cataracts, etc. Among them, DR is the most important microvascular complication and an important cause of blindness for patients. By 2045, it is estimated that the number of global DM patients complicated with DR will increase to about 160 million. Factors leading to the progression of DR include hypertension, obesity, hyperlipidemia, and anemia, etc. DR is a highly specific neurovascular complication, which can be divided into non-proliferative retinopathy (NPDR) and proliferative retinopathy (PDR) characterized by vision loss according to the pathological progression.
[0003] Currently, the main methods for evaluating drug targets and biomolecular markers from the perspective of heritability are as follows: ① Genome Wide Association Study (GWAS), ② Transcriptome Wide Association Study (TWAS), ③ Proteome Wide Association Study (PWAS), etc.
[0004] However, the applicant's research found that for diabetic retinopathy, the interpretation of GWAS results usually requires further functional annotation and bioinformatics analysis to understand the specific effects of these loci on traits and the biological mechanisms behind them. Although the TWAS results are presented in the form of genes, overcoming the main drawbacks of GWAS, the data source of mRNA is the most restrictive part. Although PWAS uses protein data, it cannot identify their causal relationship with the disease.
[0005] Based on this, there is an urgent need for a method for screening pathogenic genes and biomarkers of retinopathy. Summary of the Invention
[0006] In order to overcome the above technical defects, the present invention provides a method and system for screening pathogenic genes and biomarkers of retinopathy.
[0007] In order to solve the above problems, the present invention is implemented according to the following technical solutions:
[0008] In a first aspect, the present invention provides a method for screening pathogenic genes and biomarkers for retinopathy, comprising the following steps:
[0009] S100: Obtain GWAS summary data related to retinopathy and multi-tissue specific gene expression weight data;
[0010] S200: Combine the GWAS summary data with the multi-tissue specific gene expression weight data for transcriptome-wide association analysis and output TWAS positive results;
[0011] S300: Combine the GWAS summary data with the multi-tissue specific gene expression weight data for Mendelian randomization analysis based on summary genes to obtain tissue-specific positive associated genes;
[0012] S400: Based on the TWAS positive results and the positive associated genes, incorporate transcriptome cis-eQTL data or proteome pQTL data for Mendelian randomization analysis based on drug targets to obtain significant drug target genes or target proteins;
[0013] S510: Screen and associate the first SNPs of the TWAS positive results, the positive associated genes, and the significant drug target genes in MR-Phewas for multi-disease association to obtain associated causal molecules; or,
[0014] S520: Screen and associate the first SNPs of the TWAS positive results, the positive associated genes, and the target proteins in MR-Phewas for multi-disease association to obtain associated causal molecules;
[0015] Wherein, the causal molecules are pathogenic genes or biomarkers related to retinopathy.
[0016] In combination with the first aspect, the present invention also provides a first specific embodiment of the first aspect. Specifically, the screening method further includes:
[0017] S600: Obtain GWAS summary data of biomarker levels, and the GWAS summary data of biomarker levels includes metabolomics data and immunomics data;
[0018] S700: Perform large-scale two-sample Mendelian randomization analysis based on the associated causal molecules, GWAS summary data of biomarker levels, and outcome indicators to obtain significantly associated causal molecules; the outcome indicator is diabetic retinopathy.
[0019] In combination with the first aspect, the present invention also provides a second specific implementation manner of the first aspect. Specifically, in step S400, transcriptome cis-eQTL data or proteome pQTL data is incorporated for Mendelian randomization analysis based on drug targets, which specifically includes:
[0020] S401: Use Steiger filtering analysis and co-localization analysis to verify the obtained significant drug target genes or target proteins, and output the verified significant drug target genes or target proteins.
[0021] In combination with the first aspect, the present invention also provides a third specific implementation manner of the first aspect. Specifically, in step S200, transcriptome association analysis is performed on the GWAS summary data, which specifically includes:
[0022] S201: Combine multiple verification algorithms to verify the TWAS positive results, and output the verified TWAS positive results;
[0023] Among them, the verification algorithms include co-localization analysis, conditional analysis, permutation test, and fine-mapping analysis.
[0024] In combination with the first aspect, the present invention also provides a fourth specific implementation manner of the first aspect. Specifically, the GWAS summary data related to retinopathy is diabetes retinopathy data and proliferative diabetic retinopathy data screened from the FinnGen R9 and R6 datasets.
[0025] In combination with the first aspect, the present invention also provides a fifth specific implementation manner of the first aspect. Specifically, the multi-tissue specific gene expression weight data includes pancreatic, kidney, whole blood, and sCCA cross-histology gene expression weight data aggregated based on the GTEx-v8 database.
[0026] In the second aspect, the present invention provides a screening system for pathogenic genes and biomarkers of retinopathy, including:
[0027] A first acquisition module, which is used to acquire GWAS summary data related to retinopathy and multi-tissue specific gene expression weight data;
[0028] A first processing module, which is used to perform transcriptome association analysis on the GWAS summary data combined with multi-tissue specific gene expression weight data, and output TWAS positive results;
[0029] A second processing module, which is used to perform Mendelian randomization analysis based on the summary genes on the GWAS summary data combined with multi-tissue specific gene expression weight data to obtain tissue-specific positive associated genes;
[0030] A third processing module, which is used to incorporate transcriptome cis-eQTL data or proteome pQTL data based on the TWAS positive results and the positive associated genes to perform Mendelian randomization analysis based on drug targets, so as to obtain significant drug target genes or target proteins;
[0031] A fourth processing module, which is used to screen and perform multi-disease association of the first SNP of the TWAS positive result, the positive associated gene positive, and the significant drug target gene in MR-Phewas to obtain associated causal molecules; alternatively, screen and perform multi-disease association of the first SNP of the TWAS positive result, the positive associated gene positive, and the target protein in MR-Phewas to obtain associated causal molecules;
[0032] Wherein, the causal molecule is a pathogenic gene or biomarker related to retinopathy.
[0033] Combined with the second aspect, the present invention also provides the first specific implementation manner of the second aspect. Specifically, the screening system further includes:
[0034] A second acquisition module, which is used to acquire GWAS summary data of biomarker levels, and the GWAS summary data of biomarker levels includes metabolomics data and immunomics data;
[0035] A sixth processing module, which is used to perform large-scale two-sample Mendelian randomization analysis based on the associated causal molecules, GWAS summary data of biomarker levels, and outcome indicators to obtain significantly associated causal molecules; the outcome indicator is diabetic retinopathy.
[0036] Combined with the second aspect, the present invention also provides the second specific implementation manner of the second aspect. Specifically, the third processing module incorporates transcriptome cis-eQTL data or proteome pQTL data to perform Mendelian randomization analysis based on drug targets, and specifically executes the following steps:
[0037] Use Steiger filtering analysis and co-localization analysis to verify the obtained significant drug target genes or target proteins, and output the verified significant drug target genes or target proteins.
[0038] Combined with the second aspect, the present invention also provides the third specific implementation manner of the second aspect. Specifically, the second processing module performs transcriptome association analysis on the GWAS summary data, and specifically executes the following steps:
[0039] Combine multiple verification algorithms to verify the TWAS positive results, and output the verified TWAS positive results;
[0040] Among them, the verification algorithm includes co-localization analysis, conditional analysis, permutation test, and fine-mapping analysis.
[0041] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0042] The present invention provides a method for screening pathogenic genes and biomarkers of retinopathy, comprising the following steps: S100: Obtain GWAS summary data related to retinopathy and multi-tissue specific gene expression weight data; S200: Combine the GWAS summary data with the multi-tissue specific gene expression weight data to perform transcriptome-wide association analysis and output TWAS positive results; S300: Combine the GWAS summary data with the multi-tissue specific gene expression weight data to perform Mendelian randomization analysis based on summary genes to obtain tissue-specific positive associated genes; S400: Based on the TWAS positive results and the positive associated genes, incorporate transcriptome cis-eQTL data or proteome pQTL data to perform Mendelian randomization analysis based on drug targets to obtain significant drug target genes or target proteins; S510: Screen the first SNPs of the TWAS positive results, the positive associated genes, and the significant drug target genes in MR-Phewas and perform association analysis with multiple diseases to obtain associated causal molecules; alternatively, S520: Screen the first SNPs of the TWAS positive results, the positive associated genes, and the target proteins in MR-Phewas and perform association analysis with multiple diseases to obtain associated causal molecules; wherein, the causal molecules are pathogenic genes or biomarkers related to retinopathy.
[0043] This screening method reflects the application of data science in the field of biomedicine. The core technical concept mainly conducts scientific and reliable in-depth analysis and excavation of significantly associated pathogenic genes and biomarkers of retinopathy through multi-dimensional data analysis. The screening method of the present invention has profound significance and role at the technical level, promotes the progress of retinopathy research, and also provides important scientific data support for clinical practice. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The following further elaborates on the specific embodiments of the present invention in conjunction with the drawings, where:
[0045] Figure 1 is a flow schematic diagram of a method for screening pathogenic genes and biomarkers of retinopathy of the present invention;
[0046] Figure 2 is a schematic diagram of the technical concept of the screening method of the present invention Figure 1 ;
[0047] Figure 3 Schematic illustration of the technical concept of the screening method of the present invention Figure 2 ;
[0048] Figure 4 are the significantly high-confidence genes of TWAS of the present invention;
[0049] Figure 5 are the high-confidence genes of drug-targeted MR in the eQTLgen consortium of the present invention;
[0050] Figure 6 is that the MR of the present invention reveals the causal relationship between metabolomics and DR. Specific embodiments
[0051] The following specifically illustrates the embodiments of the present invention in conjunction with the accompanying drawings. The examples are given only for illustrative purposes and should not be construed as limiting the present invention. The accompanying drawings are for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.
[0052] Reference Figure 1 , the embodiments of the present invention provide a schematic flowchart of a method for screening pathogenic genes and biomarkers of retinopathy. The core technical concept of the present invention mainly conducts scientific and reliable in-depth analysis and excavation of significantly associated pathogenic genes and biomarkers of retinopathy through multi-dimensional data analysis. The screening method of the present invention has profound significance and role at the technical level, promotes the progress of retinopathy research, and also provides important scientific data support for clinical practice.
[0053] A method for screening pathogenic genes and biomarkers of retinopathy according to the present invention, which can be executed by a screening system for pathogenic genes and biomarkers of retinopathy. The screening system can be implemented in the form of hardware and / or software, and the screening system can be configured in an electronic device. As Figure 1 shown, the screening method includes:
[0054] S100: Obtain GWAS summary data related to retinopathy and multi-tissue-specific gene expression weight data;
[0055] S200: Combine the GWAS summary data with the multi-tissue-specific gene expression weight data to perform transcriptome-wide association analysis and output TWAS positive results;
[0056] S300: Combine the GWAS summary data with the multi-tissue-specific gene expression weight data to perform Mendelian randomization analysis based on summary genes to obtain tissue-specific positive associated genes;
[0057] S400: Based on the TWAS positive results and the positive associated genes, incorporate transcriptome cis-eQTL data or proteome pQTL data for Mendelian randomization analysis based on drug targets to obtain significant drug target genes or target proteins;
[0058] S510: Screen and associate the first SNPs of the TWAS positive results, the positive associated genes, and the significant drug target genes in MR-Phewas for multi-disease association to obtain associated causal molecules; or,
[0059] S520: Screen and associate the first SNPs of the TWAS positive results, the positive associated genes, and the target proteins in MR-Phewas for multi-disease association to obtain associated causal molecules;
[0060] Among them, the causal molecules are pathogenic genes or biomarkers related to retinopathy.
[0061] Specifically, the method for screening pathogenic genes and biomarkers of retinopathy in the present invention is mainly a method for screening pathogenic genes and biomarkers provided for diabetic retinopathy (DR). Specifically, in combination with diabetic retinopathy, each step of this screening method will be described in detail.
[0062] S100: Obtain GWAS summary data related to retinopathy and multi-tissue specific gene expression weight data.
[0063] In the present invention, GWAS summary data generally refers to the summary statistical data of Genome-Wide Association Study. These data contain the analysis results of searching for genetic variations related to specific phenotypes or diseases within the entire genome. GWAS summary data is statistical genetics data, which involves detecting genetic variation (marker) polymorphisms of multiple individuals within the entire genome, usually for single nucleotide polymorphisms (SNPs), to obtain the genotypes of each individual. These data are further subjected to population-level statistical analysis with phenotypes, and the genetic variations most likely to affect specific traits are screened out through statistics or significance (p-value), so as to discover genes related to trait variations. The data types of GWAS mainly include genotype data and phenotype data. Genotype data usually exists in the vcf file format, while phenotype data exists in the txt file format containing sample names and traits. It can be obtained from public databases, such as GWAS Catalog and OpenGWAS, which collect published GWAS data.
[0064] Specifically, GWAS summary data mainly includes genetic variation and genotype data, trait and disease-related data, and meta-analysis data. In addition, GWAS summary data also includes genetic variant identifiers (SNP ids), allele information, effect sizes, statistical significance (p-values), sample information, etc. These information are of great significance for understanding the relationship between gene variations and phenotypes. (1) Genetic variation and genotype data: These data are usually presented in the form of single nucleotide polymorphisms (SNPs) and are associated with various traits, diseases, and characteristics. The GWAS Catalog provides information on these SNPs, including genotype frequencies, statistical association coefficients, p-values, etc. (2) Trait and disease-related data: The database contains data related to various traits and diseases, such as cardiovascular diseases, cancers, autoimmune diseases, metabolic diseases, etc. These data are crucial for studying the relationship between genes and diseases. (3) Meta-analysis data: The GWAS Catalog also provides meta-analysis data from multiple independent studies. These data combine the results of multiple studies and provide a more comprehensive perspective for analyzing the relationship between genes and diseases.
[0065] In the present invention, the GWAS summary data related to retinopathy are the diabetic retinopathy data and proliferative diabetic retinopathy data screened from the FinnGen R9 and R6 datasets.
[0066] In a specific implementation, S100 obtained SNPs related to diabetic retinopathy (DR) from the FinnGen R9 and R6 datasets. Among them, the SNPs related to diabetic retinopathy (DR) include 10,413 cases of DR and 10,860 cases of proliferative diabetic retinopathy (PDR).
[0067] In the present invention, multi-tissue specific gene expression weight data generally refers to data on gene expression levels in different tissues or cell types, which can reveal the activity patterns of genes in specific tissues. The multi-tissue specific gene expression weight data contains the weights of tissue-specific genes and the results of gene set tests. The multi-tissue specific gene expression weight data is calculated based on the tissue-specific gene activity information in the Human Protein Atlas, and these data reflect the specific weights of gene expression in different tissues. Using these data, tissue-specific gene set weights can be generated, which reflect the expression specificity and importance of genes in different tissues.
[0068] In the present invention, the multi-tissue specific gene expression weight data includes pancreatic, kidney, whole blood, and sCCA cross-histology gene expression weight data aggregated based on the GTEx-v8 database.
[0069] In the present invention, the cross-tissue gene expression weight data, obtained through sCCA analysis, refers to the weights of each gene in the gene expression data across multiple tissues or samples when forming the maximum correlation linear combination. These weights reflect the importance of different genes in the association pattern, that is, which genes' expression changes contribute the most to the association between different tissues or samples.
[0070] Specifically, the sCCA cross-tissue gene expression weight data includes:
[0071] (1) Gene weights: The weights of each gene in the linear combination obtained through sCCA analysis, indicating the importance of the gene in the association pattern.
[0072] (2) Association between tissues or samples: Through sCCA analysis, the association patterns of gene expression between different tissues or samples can be revealed, and these association patterns may be related to specific biological processes or disease states.
[0073] (3) Sparsity: A key feature of sCCA is sparsity, that is, the resulting linear combination usually contains fewer non-zero weights, which helps to identify the genes that play a key role in the association pattern.
[0074] S200: Combine the GWAS summary data with the multi-tissue specific gene expression weight data for transcriptome-wide association analysis and output the positive TWAS results.
[0075] In the present invention, transcriptome-wide association studies (TWAS) is a method that uses genotype data and gene expression data to identify genetic variations that affect specific traits. TWAS usually uses the summary statistics from genome-wide association studies (GWAS) to infer the association between gene expression and specific traits.
[0076] In a specific implementation, for each tissue analysis, the positive TWAS P-value < 0.05 / number of traits (pancreas: 5734, whole blood: 7938, renal cortex: 1203) and the P-value of sCCA < 1.32e-06.
[0077] In a preferred implementation, in step S200, the combination of the GWAS summary data with the multi-tissue specific gene expression weight data for transcriptome-wide association analysis specifically includes:
[0078] S201: Verify the positive TWAS results by combining multiple verification algorithms and output the verified positive TWAS results;
[0079] Among them, the verification algorithm includes colocalization analysis, conditional analysis, permutation test, and fine-mapping analysis.
[0080] In a specific implementation, regarding the post-TWAS verification analysis:
[0081] (1) Colocalization analysis: The present invention evaluates the probability that the TWAS association reflects the following two situations: (1) PP.H3: Linkage between different causal single nucleotide polymorphisms (SNPs). (2) PP.H4: A single causal SNP. By comparing PP.H3 and PP.H4, the present invention discerns whether there is a true causal relationship between the trait and diabetic retinopathy (DR). Further, the present invention defines positive genes as genes with PP.H4 exceeding 0.9.
[0082] (2) Permutation test: To verify the analysis results of the present invention, a permutation test is adopted, and a P-value < 0.05 is regarded as a positive indicator.
[0083] (3) Conditional analysis: Conditional / marginal traits are only significantly associated with DR in the unadjusted model, and their association completely depends on the expression of other neighboring traits. After correction, independent / jointly significant traits are still associated with the phenotype, at a significant level (p < 0.05 is defined as passing the conditional analysis test).
[0084] (4) Fine-mapping analysis: FOCUS estimates the posterior inclusion probability (PIP) of causal associations for each trait within the associated region. The PIP of individual elements is also of interest, and a value > 0.5 indicates that this locus is more likely to be causally related to DR than any other locus in this region. FOCUS naturally allows for multiple causal SNPs and genes, while integrating gene effect sizes by using conjugate priors. In this example, we define genes with PIP > 0.8, significant TWAS P-values, and significant conditional analysis P-values as high-confidence genes.
[0085] In a specific implementation, the analysis can be performed by using plink and FOCUS software on Linux.
[0086] As Figure 4 shown, based on the above analysis process, the present invention screened 10 positive loci, namely the TWAS positive results.
[0087] S300: Combine the GWAS summary data with multi-tissue-specific gene expression weight data to perform Mendelian randomization analysis based on summary genes, and obtain tissue-specific positive associated genes.
[0088] In the present invention, the summary databased Mendelian randomization method (SMR) is a tool for inferring the causal relationship of genetic variants closely related to an exposure on an outcome. Another advantage of SMR is that it can integrate data from two independent samples.
[0089] In a specific implementation, SMR can establish a connection between expression quantitative trait loci (eQTL) and phenotypes. Specifically, the present invention uses SMR for analysis. To evaluate linkage disequilibrium (LD) and estimate potential colocalization, we performed the HEIDI test using an external reference. For this purpose, in this step of the present invention, the criteria for defining the final significant genes are as follows:
[0090] (1) SMR - FDR: Genes with an SMR - FDR - P value < 0.05 are considered significant.
[0091] (2) Genome - wide significance: We require that there are significant differences (P < 1×10 -5 ) between the eQTL and genome - wide association study (GWAS) results at the genome - wide level.
[0092] (3) HEDI test result: If the HEDI test shows a result greater than 0.05, the gene is retained. These strict criteria enable us to identify genes related to diabetic retinopathy.
[0093] In a specific implementation, for the two diseases DR and PDR, we found that 92 key DR regulatory genes showed 446 significant positives in 49 tissues, and 55 key PDR regulatory genes showed 327 significant positives in 49 tissues.
[0094] S400: Based on the TWAS positive results and the positive associated genes, transcriptome cis - eQTL data or proteome pQTL data are incorporated for Mendelian randomization analysis based on drug targets to obtain significant drug target genes or target proteins.
[0095] In the present invention, transcriptome cis-eQTL data or proteome pQTL data is incorporated. During the drug development process, paying attention to the use of cis-expression quantitative trait loci (cis-eQTL) can better bind to key target genes. The present invention obtained filtered significant cis-eQTL data from the eQTLgen consortium. The eQTLgen consortium collected and measured peripheral blood samples of 31,684 subjects. These data provide valuable insights into the genetic regulation of gene expression. By extracting multi-dimensional genomics data from the PsychENCODE project, which focuses on the developing and adult brain, including healthy and diseased states, the positive genes extracted from the eQTLgen consortium were verified.
[0096] In a specific implementation, in order to identify gene-related instrumental variables, the present invention selected SNPs within a range of 100 kb. The P value of these SNPs and the target gene is < 5e-8. And set r 2 = 0.001, a total of 2,499 genes were screened from eQTLgen, and 1,367 genes were screened from PsychENCODE. The screening method of the present invention utilizes these eQTL data sets to enhance the understanding of gene regulation and potential drug targets.
[0097] In another implementation, pQTL data can also be incorporated, and the data processing method is similar to the above. The eQTL data is used in this example.
[0098] In a specific implementation, given a sufficient number of high-confidence SNPs, the present invention avoids using proxies in this step. To ensure the robustness of the instrument, the present invention sets the F-value statistic to > 10 in this step, excluding SNPs with insufficient statistical power. For traits or genes with a single SNP as an instrumental variable, the present invention uses the Wald ratio to analyze the causal relationship in this step. When dealing with two SNPs, the present invention uses inverse-variance weighting (IVW) analysis in this step.
[0099] For more than two SNPs, the present invention centers on IVW in this step, supplemented by MR Egger, weighted median, simple mode, and weighted mode analysis. To enhance robustness and minimize confounding, the present invention corrects the p-value using FDR in this step. The MR Egger intercept test enables us to observe horizontal pleiotropy. Specifically, the Cochran Q heterogeneity test was used to evaluate the heterogeneity of the causal relationship between each instrument and diabetic retinopathy.
[0100] In a specific implementation, the MR analysis can be performed using the R software package TwoSampleMR v0.5.11.
[0101] In a specific implementation, in step S400, transcriptome cis-eQTL data or proteome pQTL data is incorporated for Mendelian randomization analysis of drug targets, which specifically includes:
[0102] S401: Use Steiger filtering analysis and colocalization analysis to verify the obtained significant drug target genes or target proteins, and output the verified significant drug target genes or target proteins.
[0103] Specifically, the present invention conducts an MR Steiger test for direction verification and explores reverse causation. All results are tested by colocalization analysis, and a total of 6 positive genes, namely significant drug target genes, are found, as shown in Figure 5 shown.
[0104] S510: Screen and associate the TWAS positive results, the positive associated genes, and the first SNP of the significant drug target genes in MR-Phewas for multiple diseases to obtain associated causal molecules; or,
[0105] S520: Screen and associate the TWAS positive results, the positive associated genes, and the first SNP of the target proteins in MR-Phewas for multiple diseases to obtain associated causal molecules.
[0106] Among them, the causal molecules are pathogenic genes or biomarkers related to retinopathy.
[0107] In a specific implementation of the present invention, the functional significance of the drug target sites significantly related to DR in steps S510 / S520 (significance threshold P < 0.05) can be explored by using MR PheWAS. By selecting the top snp of each gene as input, their associations with GWAS summary data of different traits are explored on a large scale. In a specific implementation, the analysis is all performed using the ieugwasr v0.2.2 R software package.
[0108] In a preferred implementation of the present invention, the screening method further includes further analysis and screening steps, specifically including:
[0109] S600: Obtain GWAS summary data of biomarker levels, and the GWAS summary data of biomarker levels includes metabolomics data and immunomics data.
[0110] In the present invention, the GWAS summary data of biomarker levels includes metabolomic exposure data of 8,229 individuals from the Canadian Longitudinal Study on Aging (CLSA) cohort, which is used to select 1,091 metabolites and 309 metabolite ratios. As exposure data of 3,757 individuals from Sardinia, 731 immune cell (118 absolute cell (AC) counts, 389 median fluorescence intensities (MFI) reflecting surface antigen levels, 32 morphological parameters, and 192 relative cell (RC) counts) - related features were selected. The DR - related features were used as outcomes. The present invention screened SNPs (P = 5E - 08) related to 1,400 metabolites and SNPs (P = 1E - 05) related to 731 immune markers. r^2 was set to 0.001, and clumping was within 100 kb. These biomarkers were selected to study whether there is an accidental association between them and DR, which will help us understand the deeper mechanism of DR progression.
[0111] S700: Conduct a large - scale two - sample Mendelian randomization analysis based on association - based causal molecules, biomarker - level GWAS summary data, and outcome indicators to obtain significantly associated causal molecules; the outcome indicator is diabetic retinopathy.
[0112] In a specific implementation, to ensure that there are sufficient SNPs for analysis, the present invention proxy the 1000 Genomes Project European reference sample. The screening threshold is R^2 = 0.8. SNPs in the MHC region are not removed. SNPs with NSNP <= 3 will be deleted.
[0113] TwoSampleMR v0.5.11 was used for MR analysis of metabolomic and immunomic biomarkers for DR. If not deliberately mentioned, other steps of this method are the same as those of drug target MR.
[0114] In a specific implementation, in addition to the Egger intercept test, the present invention also uses MR - PRESSO as a test for horizontal pleiotropy. By leveraging the power of multiple genetic variants to account for linkage disequilibrium (LD) between variants. To understand whether the association between gene variants and outcomes weakens after adjusting for exposure, a "leave - one - out" analysis was then performed, gradually deleting each SNP, calculating the meta - effect of the remaining SNPs, and observing whether the outcome changed after deleting each SNP. The results were mainly positive for metabolomics, and the immunomics results showed negative after correction. Specifically, the results can be seen in Figure 6 , to obtain significantly associated causal molecules.
[0115] The present invention also provides a screening system for pathogenic genes and biomarkers of retinopathy, including:
[0116] The first acquisition module is used to acquire GWAS summary data related to retinopathy and multi-tissue specific gene expression weight data;
[0117] The first processing module is used to perform transcriptome-wide association analysis by combining the GWAS summary data with multi-tissue specific gene expression weight data, and output TWAS positive results;
[0118] The second processing module is used to perform Mendelian randomization analysis based on summary genes by combining the GWAS summary data with multi-tissue specific gene expression weight data to obtain tissue-specific positive associated genes;
[0119] The third processing module is used to perform drug target Mendelian randomization analysis by incorporating transcriptome cis-eQTL data or proteome pQTL data based on the TWAS positive results and the positive associated genes to obtain significant drug target genes or target proteins;
[0120] The fourth processing module is used to screen and associate with multiple diseases in MR-Phewas for the first SNP of the TWAS positive results, the positive associated genes, and the significant drug target genes to obtain associated causal molecules; alternatively, screen and associate with multiple diseases in MR-Phewas for the first SNP of the TWAS positive results, the positive associated genes, and the target proteins to obtain associated causal molecules;
[0121] Wherein, the causal molecules are pathogenic genes or biomarkers related to retinopathy.
[0122] Preferably, the screening system further includes:
[0123] The second acquisition module is used to acquire GWAS summary data of biomarker levels, and the GWAS summary data of biomarker levels includes metabolomics data and immunomics data;
[0124] The sixth processing module is used to perform large-scale two-sample Mendelian randomization analysis based on the associated causal molecules, GWAS summary data of biomarker levels, and outcome indicators to obtain significantly associated causal molecules; the outcome indicator is diabetic retinopathy.
[0125] Preferably, the third processing module incorporates transcriptome cis-eQTL data or proteome pQTL data to perform drug target Mendelian randomization analysis, and specifically performs the following steps:
[0126] Verify the obtained significant drug target genes or target proteins using Steiger filtering analysis and co-localization analysis, and output the verified significant drug target genes or target proteins.
[0127] Preferably, the second processing module performs transcriptome association analysis on the GWAS summary data, and specifically executes the following steps:
[0128] Verify the TWAS positive results by combining multiple verification algorithms, and output the verified TWAS positive results;
[0129] Among them, the verification algorithms include co-localization analysis, conditional analysis, permutation test, and fine-mapping analysis.
[0130] For other structures of the method and system for screening pathogenic genes and biomarkers of retinopathy in this embodiment, refer to the prior art.
[0131] The above are only preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Therefore, any modifications, equivalent changes, and decorations made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A method for screening pathogenic genes and biomarkers of retinopathy, characterized in that It includes the following steps: S100: Obtain GWAS summary data related to retinopathy and multi-tissue specific gene expression weight data; S200: Combine the GWAS summary data with multi-tissue specific gene expression weight data for transcriptome-wide association analysis and output TWAS positive results; S3 S300: Combine the GWAS summary data with multi-tissue specific gene expression weight data for Mendelian randomization analysis based on summary genes to obtain tissue-specific positive associated genes; S400: Based on the TWAS positive results and the positive associated genes, incorporate transcriptome cis-eQTL data or proteome pQTL data for Mendelian randomization analysis based on drug targets to obtain significant drug target genes or target proteins; S510: Screen and associate the first SNPs of the TWAS positive results, the positive associated genes, and the significant drug target genes in MR-Phewas for association with multiple diseases to obtain associated causal molecules; or S520: Screen and associate the first SNPs of the TWAS positive results, the positive associated genes, and the genes corresponding to the target proteins in MR-Phewas for association with multiple diseases to obtain associated causal molecules; wherein the causal molecules are pathogenic genes or biomarkers related to retinopathy; S600: Obtain GWAS summary data of biomarker levels, and the GWAS summary data of biomarker levels includes metabolomics data and immunomics data; S700: Conduct large-scale two-sample Mendelian randomization analysis based on the associated causal molecules, GWAS summary data of biomarker levels, and outcome indicators to obtain significantly associated causal molecules; the outcome indicator is diabetic retinopathy; wherein in step S400, incorporating transcriptome cis-eQTL data or proteome pQTL data for Mendelian randomization analysis based on drug targets specifically includes: S401: Use Steiger filtering analysis and colocalization analysis to verify the obtained significant drug target genes or target proteins and output the verified significant drug target genes or target proteins.
2. The screening method of pathogenic genes and biomarkers for retinopathy according to claim 1, characterized in that In step S200, for transcriptome-wide association analysis of the GWAS summary data, it specifically includes: S201: Combine multiple validation algorithms to verify the TWAS positive results and output the verified TWAS positive results; wherein the validation algorithms include colocalization analysis, conditional analysis, permutation test, and fine-mapping analysis.
3. The method for screening pathogenic genes and biomarkers of retinopathy according to claim 1, wherein: The GWAS summary data related to retinopathy is diabetic retinopathy data and proliferative diabetic retinopathy data screened from FinnGen R9 and R6 datasets.
4. The method for screening pathogenic genes and biomarkers of retinopathy according to claim 1, wherein: The multi-tissue specific gene expression weight data includes pancreatic, kidney, whole blood, and sCCA cross-histology gene expression weight data aggregated based on the GTEx-v8 database.
5. A screening system for pathogenic genes and biomarkers of retinopathy, characterized in that, Including: A first acquisition module for acquiring GWAS summary data related to retinopathy and multi-tissue specific gene expression weight data; A first processing module for performing transcriptome association analysis by combining the GWAS summary data with multi-tissue specific gene expression weight data and outputting TWAS positive results; A second processing module for performing Mendelian randomization analysis based on summary genes by combining the GWAS summary data with multi-tissue specific gene expression weight data to obtain tissue-specific positive associated genes; A third processing module for performing Mendelian randomization analysis based on drug targets by incorporating transcriptome cis-eQTL data or proteome pQTL data based on the TWAS positive results and the positive associated genes to obtain significant drug target genes or target proteins; A fourth processing module for screening and associating with multiple diseases in MR-Phewas for the first SNPs of the TWAS positive results, the positive associated genes, and the significant drug target genes to obtain associated causal molecules; alternatively, screening and associating with multiple diseases in MR-Phewas for the first SNPs of the TWAS positive results, the positive associated genes, and the genes corresponding to the target proteins to obtain associated causal molecules; Wherein, the causal molecules are pathogenic genes or biomarkers related to retinopathy; A second acquisition module for acquiring GWAS summary data of biomarker levels, and the GWAS summary data of biomarker levels includes metabolomics data and immunomics data; A sixth processing module for performing large-scale two-sample Mendelian randomization analysis based on the associated causal molecules, GWAS summary data of biomarker levels, and outcome indicators to obtain significantly associated causal molecules; the outcome indicator is diabetic retinopathy; Wherein, the third processing module incorporates transcriptome cis-eQTL data or proteome pQTL data to perform Mendelian randomization analysis based on drug targets, and specifically executes the following steps: Using Steiger filtering analysis and colocalization analysis to verify the obtained significant drug target genes or target proteins, and outputting the verified significant drug target genes or target proteins.
6. The screening system for pathogenic genes and biomarkers of retinopathy according to claim 5, wherein The second processing module performs transcriptome association analysis on the GWAS summary data, and specifically executes the following steps: Combining multiple verification algorithms to verify the TWAS positive results and outputting the verified TWAS positive results; Wherein, the verification algorithms include colocalization analysis, conditional analysis, permutation test, and fine-mapping analysis.
Citation Information
Patent Citations
Mental disease biomarker prediction method and system and storage medium
CN117476096A
Multi-part chronic pain potential treatment target point determination system
CN118298911A