Neurodegenerative disease comprehensive analysis platform based on multi-source data fusion and application

By establishing a comprehensive analysis platform for neurodegenerative diseases that integrates multi-source data, the problems of incomplete data integration and limited analytical functions in neurodegenerative disease research have been solved. This has enabled efficient and accurate data analysis and the discovery of treatment strategies, thus promoting disease research and clinical applications.

CN120824035APending Publication Date: 2025-10-21HENAN UNIV OF CHINESE MEDICINE
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510975250.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing research on neurodegenerative diseases suffers from incomplete data integration, limited analytical functions, and a lack of interactive visualization, failing to meet the needs of comprehensive research. Furthermore, current treatments cannot fundamentally stop or reverse the disease progression.

Method used

Establish a comprehensive analysis platform for neurodegenerative diseases based on multi-source data fusion, including a data integration module, a gene name conversion module, a one-click generation analysis tool module, a network analysis module, a drug screening module, and an omics visualization module. It supports multiple analysis functions, integrates multiple data sources, provides interactive visualization and customized analysis, and dynamically updates data.

Benefits of technology

It has enabled the systematic integration and efficient analysis of data, improved research efficiency, discovered new disease targets and potential treatment strategies, provided accurate and reliable analytical results, supported scientific research collaboration and data reuse, and promoted the study of the pathogenesis and clinical treatment of neurodegenerative diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120824035A_ABST
    Figure CN120824035A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of bioinformatics and medical data analysis, and discloses a neurodegenerative disease comprehensive analysis platform based on multi-source data fusion, and the platform comprises a data integration module; a gene name conversion module; a core analysis module; a network analysis module; a drug screening module; the omics visualization module systematically catalogs data of genes related to at least 10 main neurodegenerative diseases, 486 drug-derived 3, 957 bioactive components and 18 lifestyle factors through the data integration module, and the data source is wide; according to the method, artificially sorted literature evidence, reanalyzed batch transcriptome data, single-cell RNA sequencing data and standardized data from a public database are covered, new disease targets and potential treatment strategies can be found easily, and the neurodegenerative disease research efficiency and comprehensiveness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of bioinformatics and medical data analysis technology, specifically to a comprehensive analysis platform and application of neurodegenerative diseases based on multi-source data fusion. Background Art

[0002] Neurodegenerative diseases are a major category of illnesses that pose a serious threat to human health and quality of life. They encompass at least 10 major conditions, including Alzheimer's disease, Parkinson's disease, and Huntington's disease. These diseases are typically characterized by the gradual loss of neuronal structure and function. As the disease progresses, patients develop symptoms such as cognitive impairment and movement disorders, ultimately leading to the inability to care for themselves, placing a heavy burden on patients' families and society.

[0003] Currently, the pathogenesis of neurodegenerative diseases remains largely unknown, posing a significant challenge to the development of effective treatments. Despite significant research efforts, existing treatments can only alleviate symptoms and are unable to fundamentally halt or reverse the disease's progression. Therefore, further investigation of the pathogenesis of neurodegenerative diseases, discovery of new therapeutic targets, and screening for potential therapeutic drugs and active ingredients have become urgent tasks in current medical research.

[0004] To fully utilize these multi-source data and advance research on neurodegenerative diseases, it is necessary to establish a comprehensive analysis platform that can systematically integrate data from different sources, provide a unified data processing and analysis process, and support a variety of analytical functions such as gene name conversion, network analysis, drug screening, and omics visualization. Currently, although some relevant databases and analysis tools already exist, most of them suffer from problems such as incomplete data integration, limited analysis functions, and lack of interactive visualization, which cannot meet the needs of comprehensive neurodegenerative disease research.

[0005] For example, some genetic databases only provide basic information about genes and lack association analysis with diseases, drugs, and lifestyle factors; drug screening tools are often based on a single data source or simple screening strategies, making it difficult to accurately identify candidate drugs and active ingredients with therapeutic potential; and omics data analysis platforms focus on data storage and display, lacking in-depth data mining and interactive exploration capabilities. Summary of the Invention

[0006] The purpose of the present invention is to provide a comprehensive analysis platform for neurodegenerative diseases based on multi-source data fusion and its application to solve the problems raised in the above background technology.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a comprehensive analysis platform for neurodegenerative diseases based on multi-source data fusion, the comprehensive analysis platform for neurodegenerative diseases based on multi-source data fusion comprising: A data integration module to systematically catalog data on genes, drugs, active ingredients, and lifestyle interventions associated with neurodegenerative diseases. The data covers at least 10 major neurodegenerative diseases, 3,957 bioactive ingredients derived from 486 drugs, and 18 lifestyle factors. The data are integrated from manually curated literature evidence, reanalyzed bulk transcriptome data, single-cell RNA sequencing data, and standardized data from public databases. Gene name conversion module, used to convert non-standard gene identifiers from literature and external resources into official gene symbols to ensure consistency in gene naming; One-click generative analysis tool modules, including drug query tools, brain tissue eQTL visualization analysis tools, bioinformatics analysis tools, pathway enrichment analysis tools, protein interaction analysis tools, and gene regulatory network analysis tools; The network analysis module supports one-click identification of gene sets that intersect disease targets with targets associated with drugs, ingredients, or lifestyle factors, and provides enrichment analysis and interactive visualization of protein-protein interaction networks; The drug screening module maps the therapeutic targets of drugs and ingredients to specific neurodegenerative diseases through mechanistic network analysis based on multi-omics evidence, identifying and prioritizing candidate drugs and active ingredients; The Omics Visualization module provides interactive exploration of transcriptomic landscapes across tissues and cell types, at both bulk and single-cell resolution.

[0008] Preferably, the data integration module further comprises: Genes associated with neurodegenerative diseases were compiled from seven authoritative databases, including DisGeNET, GeneCards, GWAS Catalog, HPO, MalaCards, OMIM, and Open Targets. Gene symbols were standardized and duplicate entries were removed to generate an integrated non-redundant candidate target gene set. Systematically organize batch RNA sequencing data related to human neurodegenerative diseases from the GEO database, focusing on tissue types closely related to the disease, and perform data integration and normalization, differential expression analysis, and weighted gene co-expression network analysis; We systematically organize single-cell RNA sequencing data from human disease-related tissues from GEO, Zenodo, and the Human MTG 10x SEA-AD project, and perform strict quality control, dimensionality reduction, and cell map construction.

[0009] Preferably, the gene name conversion module converts non-standard gene identifiers to official gene symbols through the alias2Symbol function in the R language package limma.

[0010] Preferably, the drug screening module further performs the following steps: S1. Calculate the intersection genes between the complete drug target library and the disease-specific gene set; S2. Reverse mapping the intersection genes to their corresponding active ingredients and drugs; S3. Prioritize candidate drugs / ingredients based on the frequency of intersection gene hits; S4. Construct mechanistic networks for the top-ranked drugs to demonstrate how their active ingredients regulate disease-related targets and infer their potential pathological impact pathways.

[0011] Preferably, the omics visualization module provides an interactive exploration interface for batch transcriptome data and single-cell transcriptome data, supporting users to freely access and download all data sets and analysis results.

[0012] Preferably, the network analysis module further includes: The enrichment analysis of the intersection gene set supports KEGG pathway enrichment and Gene Ontology enrichment statistics at three levels: biological process, cellular component, and molecular function. When constructing the protein-protein interaction network, the STRING database was used as the main source of interaction relationships, and the igraph package was combined for network topology analysis and visualization; Provides network modularization analysis function to identify key functional modules and hub nodes in the interaction network.

[0013] Preferably, the data integration module implements cross-database cross-validation when integrating public database data and applies filtering criteria to specific data sources, wherein the filtering criteria include: For the GeneCards database, only genes with a combined score higher than the median were retained; For GWAS data, genetic variants were annotated using ANNOVAR, and strict significance thresholds were used to screen genes highly associated with the phenotype; For transcriptome data, the ComBat-seq method was used to correct for batch effects to ensure comparability between data from different batches.

[0014] Preferably, the neurodegenerative disease comprehensive analysis platform based on multi-source data fusion further includes a user-defined analysis module, which allows the user to: Customize analysis workflows based on research needs, including selecting specific diseases, tissue types, datasets, and analysis parameters; Save and share customized analysis plans and results to promote scientific research collaboration and data reuse; Connect to third-party analysis tools or databases to expand the platform's analytical capabilities and data sources.

[0015] Preferably, the neurodegenerative disease comprehensive analysis platform based on multi-source data fusion further includes a dynamic data update module, which is capable of: Regularly review and update the integrated public database and literature evidence to ensure the timeliness and accuracy of the data; Automatically identify and integrate newly released data on genes, drugs, and lifestyle factors associated with neurodegenerative diseases; Provides data version control function, allowing users to trace back to historical versions of data and analysis results.

[0016] Application of a comprehensive neurodegenerative disease analysis platform based on multi-source data fusion, including: Utilizing the comprehensive neurodegenerative disease analysis platform based on multi-source data fusion to perform target discovery and hypothesis generation for neurodegenerative diseases; Identifying and prioritizing candidate drugs and active ingredients for treating neurodegenerative diseases through the drug screening module of the neurodegenerative disease comprehensive analysis platform based on multi-source data fusion; Utilize the omics visualization module of the multi-source data fusion-based neurodegenerative disease comprehensive analysis platform to explore the transcriptome landscape of neurodegenerative diseases and assist in mechanistic research; Combining multi-omics data, we conduct research on the pathogenesis of neurodegenerative diseases and provide new strategies and targets for clinical treatment.

[0017] The beneficial effects of the present invention are as follows: 1. In this invention, a data integration module is used to systematically catalog data on genes associated with at least 10 major neurodegenerative diseases, 3,957 bioactive ingredients derived from 486 drugs, and 18 lifestyle factors. The data comes from a wide range of sources, including manually curated literature evidence, reanalyzed batch transcriptome data, single-cell RNA sequencing data, and standardized data from public databases. This comprehensive and systematic data integration approach breaks the fragmented and isolated situation of previous research data and provides researchers with a "one-stop" data acquisition channel. Researchers no longer need to tediously search and organize data in multiple databases and literature, which greatly saves time and energy and allows them to focus more on studying and analyzing disease mechanisms, thereby significantly improving the efficiency of neurodegenerative disease research. At the same time, the rich data resources also provide a more comprehensive perspective for research, helping to discover new disease targets and potential treatment strategies, and improving the efficiency and comprehensiveness of neurodegenerative disease research.

[0018] 2. In the present invention, during the data processing and analysis process, this platform has taken a series of effective measures to ensure the accuracy and reliability of the results. The gene name conversion module uses the alias2Symbol function in the R language package limma to uniformly convert non-standard gene identifiers from literature and external resources into official gene symbols, ensuring the consistency of gene naming and avoiding analysis errors caused by inconsistent gene names. When integrating public database data, the data integration module implements cross-database cross-validation and applies strict filtering criteria to specific data sources. For example, for the GeneCards database, only genes with a comprehensive score higher than the median are retained. For GWAS data, a strict significance threshold is used to screen genes that are highly correlated with the phenotype. For transcriptome data, the ComBat-seq method is used to correct batch effects. These measures effectively remove noise and redundant information in the data, improve data quality, make subsequent analysis results more accurate and reliable, and provide a solid data foundation for the study of neurodegenerative diseases.

[0019] 3. In the present invention, the platform not only has powerful core analysis functions, such as the network analysis module that supports one-click identification of the intersection gene set between disease targets and drug, ingredient or lifestyle factor-related targets, and provides enrichment analysis and interactive visualization of protein-protein interaction networks; the drug screening module can identify and prioritize candidate drugs and active ingredients based on multi-omics evidence; the omics visualization module provides interactive exploration of transcriptome landscapes across tissues and cell types, etc. It also has a user-defined analysis module and a dynamic data update module. The user-defined analysis module allows users to customize the analysis process according to research needs, save and share customized analysis plans and results, and promote scientific research collaboration and data reuse; the function of accessing third-party analysis tools or databases further expands the platform's analysis capabilities and data sources. The dynamic data update module can regularly check and update data and provide data version control functions to ensure that users can always obtain the latest and most accurate data and analysis results. These features make the platform highly practical and scalable, able to meet the research needs of different researchers, accelerate the study of the pathogenesis of neurodegenerative diseases, provide new strategies and targets for clinical treatment, and promote the transformation of scientific research results into clinical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 Schematic diagram of the composition of the comprehensive analysis platform of the present invention; Figure 2 This is a flow chart of transcriptome analysis in the present invention; Figure 3 This is a flowchart of single-cell transcriptome analysis in the present invention; Figure 4 This is a workflow diagram of the network analysis module in the present invention; Figure 5 This is a workflow diagram of the drug screening module in the present invention; Figure 6 Schematic diagram of the data integration module in the present invention. DETAILED DESCRIPTION

[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0022] like Figures 1 to 6 As shown, an embodiment of the present invention provides a comprehensive analysis platform for neurodegenerative diseases based on multi-source data fusion. The comprehensive analysis platform for neurodegenerative diseases based on multi-source data fusion includes: A data integration module is used to systematically catalog data on genes, drugs, active ingredients, and lifestyle interventions associated with neurodegenerative diseases. The data covers at least 10 major neurodegenerative diseases, 3,957 bioactive ingredients derived from 486 drugs, and 18 lifestyle factors. The data are integrated from manually curated literature evidence, reanalyzed bulk transcriptome data, single-cell RNA sequencing data, and standardized data from public databases. Specifically, phenotype-gene association data come from a wide range of sources, including the following authoritative databases: DisGeNET, GeneCards, GWAS Catalog, Human Phenotype Ontology (HPO), MalaCards, OMIM, Open Targets, Korean Genome and Epidemiology Study (KoGES), UK Biobank (UKBB), GEO, Human MTG 10xSEA-AD, and ZENODO; Gene name conversion module, used to convert non-standard gene identifiers from literature and external resources into official gene symbols to ensure consistency in gene naming; One-click generative analysis tool modules, including the drug query tool (HerbDex), the brain tissue eQTL visualization analysis tool (IGV), the biological function enrichment analysis tool (GO Enrichment), the pathway enrichment analysis tool (KEGGEnrichment), the protein interaction analysis tool (PPI), and the gene regulatory network analysis tool (GRN); The network analysis module supports one-click identification of gene sets that intersect disease targets with targets associated with drugs, ingredients, or lifestyle factors, and provides enrichment analysis and interactive visualization of protein-protein interaction networks; The drug screening module maps the therapeutic targets of drugs and ingredients to specific neurodegenerative diseases through mechanistic network analysis based on multi-omics evidence, identifying and prioritizing candidate drugs and active ingredients; The Omics Visualization module provides interactive exploration of transcriptomic landscapes across tissues and cell types, at both bulk and single-cell resolution.

[0023] The data integration module further includes: Genes associated with neurodegenerative diseases were compiled from seven authoritative databases, including DisGeNET, GeneCards, GWAS Catalog, HPO, MalaCards, OMIM, and Open Targets. Gene symbols were standardized and duplicate entries were removed to generate an integrated non-redundant candidate target gene set. Systematically organize batch RNA sequencing data related to human neurodegenerative diseases from the GEO database, focusing on tissue types closely related to the disease, and perform data integration and normalization, differential expression analysis, and weighted gene co-expression network analysis; We systematically organize single-cell RNA sequencing data from human disease-related tissues from GEO, Zenodo, and the Human MTG 10x SEA-AD project, and perform strict quality control, dimensionality reduction, and cell map construction.

[0024] To remove redundant and low-confidence entries, cross-database cross-validation was performed. Furthermore, to ensure data quality, filtering criteria were applied to specific data sources: (1) for GeneCards, only genes with a composite score above the median were retained; (2) for GWAS data (UKBB / KoGES), genetic variants were annotated using ANNOVAR, and a strict significance threshold (p < 1E-6) was used to screen genes highly associated with the phenotype.

[0025] For human neurodegenerative diseases (NDDs) target integration: 1. Organizing Disease-Related Genes The comprehensive analysis platform initially aggregated genes associated with neurodegenerative diseases from seven authoritative databases, including DisGeNET, GeneCards, GWAS Catalog, HPO, MalaCards, OMIM, and Open Targets. Gene symbols were then standardized and duplicate entries removed to generate an integrated, non-redundant candidate target gene set (Unigenes). 2. Batch transcriptome feature analysis The comprehensive analysis platform systematically organizes bulk RNA sequencing (bulk RNA-seq) data related to human neurodegenerative diseases (NDDs) from the GEO database, focusing on tissue types closely related to the disease. The analysis process includes the following three steps: Data integration and standardization: Integrate multiple batches of data and use the ComBat-seq method to correct for batch effects; Differential expression analysis: Screening of disease-related differentially expressed genes based on statistical significance (p < 0.05) and expression change magnitude (|log2FC| > 1); Weighted gene co-expression network analysis (WGCNA): constructing a scale-free network topology by selecting the optimal soft threshold, and then identifying key regulatory modules based on hierarchical clustering of principal components and phenotypic correlation analysis; During the data collation process, the comprehensive analysis platform integrated a total of 5,132 bulk RNA-seq samples from the GEO database, covering 7 neurodegenerative diseases and 28 tissue types. All samples were processed through strict case-control matching (total including 3,240 disease samples and 1,892 healthy samples). The specific distribution is as follows: Alzheimer's disease (cases = 2,317, controls = 988); Parkinson's disease (cases = 524, controls = 442); Huntington's disease (cases = 157, controls = 157); Amyotrophic Lateral Sclerosis (cases = 145, controls = 17); Multiple Sclerosis (cases = 40, controls = 10); Prion Disease (cases = 4, controls = 10); Progressive Supranuclear Palsy (cases = 53, controls = 268).

[0026] 3. Single-cell transcriptome analysis The integrated analysis platform systematically organizes single-cell RNA sequencing data from human disease-related tissues from GEO, Zenodo, and the Human MTG 10x SEA-AD project. The analysis workflow includes the following steps: Strict quality control: retained cells must meet the following criteria: nFeature_RNA > 250, log10(GenesPerUMI) > 0.80, mitochondrial gene ratio < 0.10; Dimensionality reduction process: Data normalization (LogNormalize, scale factor = 10,000) → Selection of highly variable genes (using the VST method to screen 2,000 genes) → Harmony algorithm for batch effect correction → UMAP projection (using the first 30 principal components, resolution = 0.3); Cell atlas construction: Use the FindAllMarkers function to annotate clustering results → Interactively visualize cell type characteristics through UMAP embedding and word cloud diagrams based on marker genes.

[0027] The comprehensive analysis platform integrates 13 single-cell RNA sequencing datasets from GEO, Zenodo and the Human MTG 10x SEA-AD project, covering 244 samples (including 131 disease samples and 113 control samples), 14 disease-related tissues, and involving three major neurodegenerative diseases.

[0028] Among them, the gene name conversion module uses the alias2Symbol function in the R language package limma to convert non-standard gene identifiers to official gene symbols.

[0029] The drug screening module further performs the following steps: S1. Calculate the intersection genes between the complete drug target library and the disease-specific gene set; S2. Reverse mapping the intersection genes to their corresponding active ingredients and drugs; S3. Prioritize candidate drugs / ingredients based on the frequency of intersection gene hits; S4. Construct mechanistic networks for the top-ranked drugs to demonstrate how their active ingredients regulate disease-related targets and infer their potential pathological impact pathways.

[0030] Among them, the omics visualization module provides an interactive exploration interface for batch transcriptome data and single-cell transcriptome data, supporting users to freely access and download all data sets and analysis results.

[0031] The network analysis module further includes: The enrichment analysis of the intersection gene set supports KEGG pathway enrichment and Gene Ontology enrichment statistics at three levels: biological process, cellular component, and molecular function. When constructing the protein-protein interaction network, the STRING database was used as the main source of interaction relationships, and the igraph package was combined for network topology analysis and visualization; Provides network modularization analysis function to identify key functional modules and hub nodes in the interaction network.

[0032] The data integration module implements cross-database cross-validation when integrating data from public databases and applies filtering criteria to specific data sources. The filtering criteria include: For the GeneCards database, only genes with a combined score higher than the median were retained; For transcriptome data, the ComBat-seq method was used to correct for batch effects to ensure comparability between data from different batches.

[0033] The comprehensive neurodegenerative disease analysis platform based on multi-source data fusion also includes a user-defined analysis module, which allows users to: Customize analysis workflows based on research needs, including selecting specific diseases, tissue types, datasets, and analysis parameters; Save and share customized analysis plans and results to promote scientific research collaboration and data reuse; Connect to third-party analysis tools or databases to expand the platform's analytical capabilities and data sources.

[0034] The comprehensive neurodegenerative disease analysis platform based on multi-source data fusion also includes a dynamic data update module, which can: Regularly review and update the integrated public database and literature evidence to ensure the timeliness and accuracy of the data; Automatically identify and integrate newly released data on genes, drugs, and lifestyle factors associated with neurodegenerative diseases; Provides data version control function, allowing users to trace back to historical versions of data and analysis results.

[0035] Applications of a comprehensive neurodegenerative disease analysis platform based on multi-source data fusion include: Utilize a comprehensive neurodegenerative disease analysis platform based on multi-source data fusion to discover targets and generate hypotheses for neurodegenerative diseases; Identify and prioritize candidate drugs and active ingredients for the treatment of neurodegenerative diseases through the drug screening module of the neurodegenerative disease comprehensive analysis platform based on multi-source data fusion; Exploring the transcriptomic landscape of neurodegenerative diseases using the omics visualization module of the comprehensive neurodegenerative disease analysis platform based on multi-source data fusion to assist in mechanistic research; Combining multi-omics data, we conduct research on the pathogenesis of neurodegenerative diseases and provide new strategies and targets for clinical treatment.

[0036] Target integration for neurodegenerative diseases: 1. Summary of Drug-Gene Interactions The comprehensive analysis platform is based on the Traditional Chinese Medicine Systems Pharmacology (TCMSP) database. It systematically organizes drug-ingredient-target interaction information and implements strict pharmacokinetic screening criteria for the screening of active ingredients: 1) Oral bioavailability (OB) ≥ 30%; 2) Drug-likeness (DL) ≥ 0.18.

[0037] After screening, a total of 66,819 high-confidence interaction data were obtained, covering 486 traditional Chinese medicines, 3,957 active ingredients, and 397 unique drug targets. Relevant details can be interactively browsed in the "Drug" module of the database; 2. Organizing lifestyle-related genes The comprehensive analysis platform systematically integrated 18 genes associated with lifestyle interventions, categorized by intervention type into two groups: 12 dietary interventions and 6 non-dietary interventions. This compilation integrated evidence from six authoritative databases: GeneCards, GWAS Catalog, OMIM, Open Targets, KoGES, and UK Biobank (UKBB). After cross-database integration and removal of redundant entries, a non-redundant set of candidate target genes (Unigenes) was generated.

[0038] The database usage methods include: 1. Home page: The homepage of the comprehensive analysis platform provides a comprehensive overview of the database content and is the core entrance for exploring various resources. The platform integrates data involving 10 neurodegenerative diseases, covers 486 drugs and 3,957 active ingredients, and incorporates information on 18 types of lifestyle factors. The comprehensive analysis platform supports interactive exploration of multi-omics data, including batch transcriptome data (covering 23 tissues, 5,132 samples) and single-cell transcriptome data (covering 14 tissues, 244 samples, 1,534,598 cells). To facilitate data analysis and hypothesis generation, the comprehensive analysis platform provides two core analysis modules (network analysis and drug screening) and six dedicated tools (HerbDex, IGV, GO enrichment, KEGG enrichment, PPI, GRN); 2. Browse diseases: The "Disease" page of the comprehensive analysis platform integrates disease-related genes from seven authoritative public databases. Users can access specific disease content through three intuitive ways: the homepage search bar, the "Neurodegenerative Disease" icon on the homepage, and the "Disease" module in the top menu. Once on a disease page, users can explore candidate genes through dynamic visualizations and interactive tables: (1) Candidate gene details: Click on the gene count in any database, or the corresponding area in the circular chart to view the detailed annotation information of the candidate target; (2) Integrated non-redundant targets (Unigenes): Click the "Unigenes" area in the circular diagram to view the unique, non-redundant disease-related gene sets integrated from seven major databases; (3) Cross-database comparison: Click on any disease name to compare the differences in the gene sets of the disease in multiple databases.

[0039] 3. Browse drugs: The "Drugs" page integrates 66,819 drug-gene interaction data, covering 486 drugs, 3,957 unique active ingredients (MOLs), and 397 independent gene symbols. Users can access specific drug information in the following ways: After entering the "Drug" module in the top menu of the search bar, the page will display key information in the database in a visual chart: 1) Drug frequency chart: A bar chart showing the top 20 drugs, ranked by the number of associated active ingredients and target genes; 2) Ingredient frequency chart: displays the top 20 ingredients, sorted by the number of their targets; 3) Target frequency map: Displays the top 20 gene symbols that frequently appear in drug-ingredient interactions.

[0040] Users can search by drug name or ingredient name, and the system will dynamically display the corresponding top 20 active ingredients and targets, or reversely display the top 20 drugs related to a certain ingredient and their targets, thereby achieving rapid insight into the compound mechanism and multi-target characteristics.

[0041] 4. Browse Lifestyle: The "Lifestyle" page aggregates genes associated with 18 lifestyle factors, categorized by intervention type: dietary and non-dietary. Users can access the page by selecting the "Lifestyle" module from the top menu in the comprehensive search bar. Similar to the disease module, the page's interactive interface allows users to explore in-depth relationships between lifestyle factors and genes through charts and tables.

[0042] 5. Browse transcriptome data: The "Transcriptome" page integrates batch RNA sequencing data from human tissues related to neurodegenerative diseases compiled from the GEO database. Users can access it by selecting the target disease from the "Transcriptome" search bar drop-down menu, then selecting a tissue or dataset. This module features: if a tissue is represented by only a single GEO dataset, its differential expression and WGCNA analysis results are directly displayed; if a tissue corresponds to multiple datasets, the following are provided separately: 1) Independent analysis results of each sub-dataset; 2) Unified analysis results after integrating multiple data sets from the same organization.

[0043] Click on any tissue or dataset name to further view the corresponding differential expression profile and co-expression network module information.

[0044] 6. Browse single-cell data: The "Single Cell" page integrates scRNA-seq data related to neurodegenerative diseases, covering 14 tissues, 244 samples, and 1,534,598 cells. Users can select a disease from the "Single Cell" drop-down menu and click on the corresponding tissue or dataset to access the corresponding tissue atlas.

[0045] The processing flow is as follows: 1) If a tissue corresponds to only one dataset, its single-cell atlas is displayed directly; 2) If a tissue contains multiple datasets, batch effect correction and integrated analysis are applied to generate a unified tissue-level map.

[0046] Click on any tissue or dataset to explore its single-cell expression landscape in depth, including cell type composition, marker gene expression, and UMAP visualization.

[0047] 7. Browsing tools: The "Tools" page integrates eight analytical tools: HerbDex, IGV, GO enrichment, KEGG enrichment, PPI network, GRN regulatory network, drug screening tools, and network analysis tools. Users can access the corresponding tool module through the top menu bar, supporting interactive hypothesis generation and mechanism exploration.

[0048] HerbDex: This database contains 240 traditional Chinese medicines, providing information on their efficacy, properties, flavor, meridians, origin, growing environment, diseases they are suitable for, and chemical composition. Advanced filtering allows users to quickly locate herbs that meet their research needs.

[0049] IGV: An integrated genome browser for visualizing GTEx brain eQTL data. It supports querying expression regulatory sites by gene name or chromosome interval.

[0050] GO enrichment analysis: Based on the enrichR and ggplot2 packages, GO functional annotation is performed on the input gene set to show the biological processes, cellular components and molecular functions in which it participates.

[0051] KEGG enrichment analysis: We also used enrichR and ggplot2 to perform KEGG pathway analysis on the input gene set to reveal NDD-related dysregulated signaling pathways.

[0052] PPI network: Call the STRING database and igraph package to construct and visualize protein interaction networks between genes, supporting topological analysis and functional module discovery.

[0053] GRN regulatory network: Based on JASPAR, ENCODE and enrichR data, infer the transcription factor-gene regulatory relationship and explore potential core regulatory factors and co-expression modules.

[0054] Drug screening tool: Same as the drug screening module, except that in this tool, users can upload the gene set of interest. The platform will then automatically generate drug and active ingredient screening results based on this gene set, as well as which active ingredients of high-frequency drugs affect the expression of which target genes, and thus participate in the disease process involved in the gene set.

[0055] Network analysis tool: Same as the network analysis module, except that in this tool, users can upload the gene set of interest. Then, after the user selects a drug, active ingredient, or lifestyle, the platform will generate enrichment analysis results and protein interaction network results of the intersection genes of the two with one click.

[0056] 8. Data download: The "Download" page supports downloading the following data modules: disease data, lifestyle data, transcriptome data, and single-cell data. Click on the corresponding module to access the detailed page for data source information and data download links.

[0057] 9. Network Analysis Module: Users can access the Network Analysis Module details page by selecting "Yin Yang" from the homepage menu > "Network Analysis Module". In this module, users can customize the following parameters based on their research needs: 1) Select the target disease; Specify differentially expressed gene sets at the tissue or dataset level; At the same time, disease-related gene sets compiled from public databases are included; 2) Enter the target drug name, active ingredient, or lifestyle factor to be analyzed.

[0058] After the user completes the above selections, the system will automatically generate the following analysis content: 1) Gene set intersection analysis: identifying common targets between diseases and interventions (drugs or lifestyle); 2) Protein interaction network analysis: constructing a PPI network of intersection genes based on the STRING database; 3) GO enrichment analysis: identifying biological processes, cellular components, and molecular functions that are significantly enriched in the intersection genes; 4) KEGG pathway analysis: revealing the key signaling pathways and potential disease mechanisms involved in the intersection genes.

[0059] All analysis results support interactive visualization and result download, facilitating mechanism interpretation and scientific research report writing.

[0060] 10. Drug Screening Module Users can access the Drug Screening Module from the homepage menu > Yin-Yang > Drug Screening Module. This module supports the following complete analysis workflow: 1) Disease selection The user can specify one or more target diseases.

[0061] 2) Data configuration Differentially expressed gene sets can be selected at tissue or dataset level; Disease-related gene sets from multiple public databases can be selected in parallel.

[0062] 3) Automated analysis process The system will automatically complete the following analysis tasks: Gene set intersection analysis: identifying overlapping regions between drug targets and disease genes Identification of drugs and their active ingredients: Locating all drugs and ingredients related to the intersection genes; Top drug / ingredient statistics: output the top 20 drugs and active ingredients with the highest hit frequency; Target mechanism deconstruction: Displays the active ingredients and targets of the most frequently identified drugs, assisting in analyzing the potential molecular mechanisms of drug intervention.

[0063] All analysis modules support interactive chart display and data export, suitable for scientific research reporting, project design and drug mechanism exploration.

[0064] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0065] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A comprehensive neurodegenerative disease analysis platform based on multi-source data fusion, characterized by: The comprehensive analysis platform for neurodegenerative diseases based on multi-source data fusion includes: A data integration module to systematically catalog data on genes, drugs, active ingredients, and lifestyle interventions associated with neurodegenerative diseases. The data covers at least 10 major neurodegenerative diseases, 3,957 bioactive ingredients derived from 486 drugs, and 18 lifestyle factors. The data are integrated from manually curated literature evidence, reanalyzed bulk transcriptome data, single-cell RNA sequencing data, and standardized data from public databases. Gene name conversion module, used to convert non-standard gene identifiers from literature and external resources into official gene symbols to ensure consistency in gene naming; One-click generative analysis tool modules, including drug query tools, brain tissue eQTL visualization analysis tools, bioinformatics analysis tools, pathway enrichment analysis tools, protein interaction analysis tools, and gene regulatory network analysis tools; The network analysis module supports one-click identification of gene sets that intersect disease targets with targets associated with drugs, ingredients, or lifestyle factors, and provides enrichment analysis and interactive visualization of protein-protein interaction networks; The drug screening module maps the therapeutic targets of drugs and ingredients to specific neurodegenerative diseases through mechanistic network analysis based on multi-omics evidence, identifying and prioritizing candidate drugs and active ingredients; The Omics Visualization module provides interactive exploration of transcriptomic landscapes across tissues and cell types, at both bulk and single-cell resolution.

2. The neurodegenerative disease comprehensive analysis platform based on multi-source data fusion according to claim 1, characterized in that: The data integration module further includes: Genes associated with neurodegenerative diseases were compiled from seven authoritative databases, including DisGeNET, GeneCards, GWAS Catalog, HPO, MalaCards, OMIM, and Open Targets. Gene symbols were standardized and duplicate entries were removed to generate an integrated non-redundant candidate target gene set. Systematically organize batch RNA sequencing data related to human neurodegenerative diseases from the GEO database, focusing on tissue types closely related to the disease, and perform data integration and normalization, differential expression analysis, and weighted gene co-expression network analysis; We systematically organize single-cell RNA sequencing data from human disease-related tissues from GEO, Zenodo, and the Human MTG 10x SEA-AD project, and perform strict quality control, dimensionality reduction, and cell map construction.

3. The neurodegenerative disease comprehensive analysis platform based on multi-source data fusion according to claim 1, characterized in that: The gene name conversion module converts non-standard gene identifiers to official gene symbols through the alias2Symbol function in the R language package limma.

4. The neurodegenerative disease comprehensive analysis platform based on multi-source data fusion according to claim 1, characterized in that: The drug screening module further performs the following steps: S1. Calculate the intersection genes between the complete drug target library and the disease-specific gene set; S2. Reverse mapping the intersection genes to their corresponding active ingredients and drugs; S3. Prioritize candidate drugs / ingredients based on the frequency of intersection gene hits; S4. Construct mechanistic networks for the top-ranked drugs to demonstrate how their active ingredients regulate disease-related targets and infer their potential pathological impact pathways.

5. The neurodegenerative disease comprehensive analysis platform based on multi-source data fusion according to claim 1, characterized in that: The omics visualization module provides an interactive exploration interface for batch transcriptome data and single-cell transcriptome data, allowing users to freely access and download all data sets and analysis results.

6. The neurodegenerative disease comprehensive analysis platform based on multi-source data fusion according to claim 1, characterized in that: The network analysis module further includes: The enrichment analysis of the intersection gene set supports KEGG pathway enrichment and Gene Ontology enrichment statistics at three levels: biological process, cellular component, and molecular function. When constructing the protein-protein interaction network, the STRING database was used as the main source of interaction relationships, and the igraph package was combined for network topology analysis and visualization; Provides network modularization analysis function to identify key functional modules and hub nodes in the interaction network.

7. The neurodegenerative disease comprehensive analysis platform based on multi-source data fusion according to claim 1, characterized in that: When integrating data from public databases, the data integration module implements cross-database cross-validation and applies filtering criteria to specific data sources. The filtering criteria include: For the GeneCards database, only genes with a combined score higher than the median were retained; For GWAS data, genetic variants were annotated using ANNOVAR, and strict significance thresholds were used to screen genes highly associated with the phenotype; For transcriptome data, the ComBat-seq method was used to correct for batch effects to ensure comparability between data from different batches.

8. The neurodegenerative disease comprehensive analysis platform based on multi-source data fusion according to claim 1, characterized in that: The neurodegenerative disease comprehensive analysis platform based on multi-source data fusion also includes a user-defined analysis module, which allows the user to: Customize analysis workflows based on research needs, including selecting specific diseases, tissue types, datasets, and analysis parameters; Save and share customized analysis plans and results to promote scientific research collaboration and data reuse; Connect to third-party analysis tools or databases to expand the platform's analytical capabilities and data sources.

9. The neurodegenerative disease comprehensive analysis platform based on multi-source data fusion according to claim 1, characterized in that: The neurodegenerative disease comprehensive analysis platform based on multi-source data fusion also includes a dynamic data update module, which can: Regularly review and update the integrated public database and literature evidence to ensure the timeliness and accuracy of the data; Automatically identify and integrate newly released data on genes, drugs, and lifestyle factors associated with neurodegenerative diseases; Provides data version control function, allowing users to trace back to historical versions of data and analysis results.

10. Application of the neurodegenerative disease comprehensive analysis platform based on multi-source data fusion according to any one of claims 1 to 9, characterized in that: The applications include: Utilizing the comprehensive neurodegenerative disease analysis platform based on multi-source data fusion to perform target discovery and hypothesis generation for neurodegenerative diseases; Identifying and prioritizing candidate drugs and active ingredients for treating neurodegenerative diseases through the drug screening module of the neurodegenerative disease comprehensive analysis platform based on multi-source data fusion; Utilize the omics visualization module of the multi-source data fusion-based neurodegenerative disease comprehensive analysis platform to explore the transcriptome landscape of neurodegenerative diseases and assist in mechanistic research; Combining multi-omics data, we conduct research on the pathogenesis of neurodegenerative diseases and provide new strategies and targets for clinical treatment.

Citation Information

Cited By

  • Virtual cell analysis platform based on cell disturbance data

    CN122117066A