System for analyzing relation between trialkanoic acid and cognition by using GEO and NHANES databases

By designing a system including data acquisition, processing, analysis and integration modules, the problems of data quality decline and information loss after data integration of NHANES and GEO databases are solved, and a comprehensive and accurate analysis of the relationship between trialanoic acid and cognitive is achieved.

CN120048364APending Publication Date: 2025-05-27SHANGHAI CITY PUDONG NEW AREA GONGLI HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411810176.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The data format, structure and source of NHANES and GEO databases are different, resulting in a decline in data quality and loss of information after data integration, affecting the accuracy of subsequent analysis.

Method used

Design a system, including data acquisition module, data processing module, GEO data analysis module, NHANES data analysis module, data integration module and data verification module, through these modules, the data is preprocessed, analyzed and integrated to ensure the accuracy and consistency of the data.

Benefits of technology

Through the use of the system, the data from GEO and NHANES databases can be effectively integrated to form a comprehensive understanding of the relationship between trialanoic acid and cognitive, avoiding the problems of data quality and information loss, and improving the reliability and accuracy of the analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048364A_ABST
    Figure CN120048364A_ABST
Patent Text Reader

Abstract

The invention discloses a system for analyzing a relationship between trialkanoic acid and cognition by using GEO and NHANES databases. The system comprises a GEO database, an NHANES database, a data acquisition module, a data processing module, a GEO data analysis module, an NHANES data analysis module, a data integration module and a data verification module. The signal end of the GEO database and the signal end of the NHANES database are connected with the signal end of the data acquisition module. An analysis result of GEO data and an analysis result of NHANES data are integrated through the data integration module, comprehensive understanding of the trialkanoic acid and cognitive relationship is formed, in the integration process, the module adopts an integration algorithm to ensure the accuracy and consistency of the data, and meanwhile, the data integration module has the advantages of being simple in structure and convenient to use. And the integrated data is compared with a known standard data set or a reliable data source through the data verification module, so that the reliability of the integrated data is further verified, and the problems of data quality reduction and information loss cannot occur after the data is integrated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a system, specifically a system for analyzing the relationship between triacontanoic acid and cognition by using GEO and NHANES databases, belonging to the technical field of data analysis systems. Background Technique

[0002] NHANES is conducted by the Centers for Disease Control and Prevention of the United States, which estimates the health and nutritional conditions and potential risk factors in the local civilian population. The survey adopts a complex probability sampling design, combining extensive household interviews, clinical examinations, and laboratory evaluations to collect information. The GEO database, namely the Gene Expression Omnibus, collects gene expression data from research institutions around the world, covering multiple fields such as tumors, non-tumors, microarrays, NGS, differential analysis, and molecular validation. Researchers can use the GEO database to find relevant gene expression data for a certain disease and then analyze these data to find potential biomarkers or therapeutic targets for the disease.

[0003] Triacontanoic acid is a saturated fatty acid containing 23 carbon atoms and plays an important physiological role in the human body. Triacontanoic acid has a potential impact on cognitive dysfunction in the human body. By using GEO and NHANES databases, the relationship between triacontanoic acid and cognitive function can be analyzed. However, the data formats, structures, and sources of NHANES and GEO databases are different. After integrating these data, problems such as decreased data quality and information loss often occur, affecting the accuracy of subsequent analysis. Therefore, a system for analyzing the relationship between triacontanoic acid and cognition by using GEO and NHANES databases is proposed. Summary of the Invention

[0004] The purpose of the present invention is to provide a system for analyzing the relationship between triacontanoic acid and cognition by using GEO and NHANES databases to solve one of the problems mentioned in the above background technique.

[0005] The present invention is implemented by the following technical solutions: A system for analyzing the relationship between triacontanoic acid and cognition by using GEO and NHANES databases includes a GEO database, an NHANES database, a data acquisition module, a data processing module, a GEO data analysis module, an NHANES data analysis module, a data integration module, and a data verification module;

[0006] The signal terminals of the GEO database and the NHANES database are both connected to the signal terminal of the data acquisition module. The signal terminal of the data processing module is respectively connected to the signal terminals of the GEO data analysis module and the NHANES data analysis module. The signal terminal of the data integration module is respectively connected to the signal terminals of the data verification module, the GEO data analysis module, and the NHANES data analysis module;

[0007] The data acquisition module is used to separately acquire gene expression data and health and nutrition data associated with triacids and cognitive function from the GEO database and the NHANES database;

[0008] The data processing module is used to preprocess the acquired data, including data cleaning, format conversion, missing value filling, and standardization, screen out datasets related to triacid metabolism and cognitive function from the preprocessed data, and further filter the screened data to remove irrelevant or noisy data;

[0009] The GEO data analysis module is used to perform statistical analysis on the screened gene expression data, compare the differences in gene expression between different sample groups, and identify genes related to triacids and cognitive function;

[0010] The NHANES data analysis module is used to statistically analyze the health and nutrition data in the NHANES database and analyze the relationship between triacid intake and the results of cognitive function assessment;

[0011] The data integration module uses an integration algorithm to integrate the analysis results of the GEO data analysis module and the NHANES data analysis module to form a comprehensive understanding of the relationship between triacids and cognition.

[0012] As a further preference of this technical solution: The integration algorithm includes the following steps:

[0013] Data preparation: Obtain the preprocessed gene expression data and health and nutrition data related to triacids and cognitive function;

[0014] Data matching: Match the gene expression data in the GEO database with the health and nutrition data in the NHANES database according to the common characteristics of the data;

[0015] Data integration: Merge the matched and aligned data into a complete dataset, check the merged dataset, remove duplicate data records, and perform standardization processing on the merged data to eliminate the dimensional differences and distribution differences between different datasets;

[0016] Data verification: Check whether the integrated data is logically consistent, including whether there is a reasonable association between the gene expression data and the health and nutrition data, and verify whether the integrated data is complete.

[0017] As a further preference of this technical solution: The signal end of the data verification module is connected to the signal end of the data integration module;

[0018] The data verification module is used to check whether the number of records in the integrated dataset meets the expectation, and compare the integrated data with a known standard dataset or a reliable data source to verify the accuracy of the data.

[0019] As a further preference of this technical solution: The signal terminals of the data integration module are respectively connected to a bioinformatics analysis module, an epidemiological information analysis module, and a comprehensive analysis module;

[0020] The bioinformatics analysis module is used to perform bioinformatics analysis on the screened key genes in combination with gene expression data, and identify gene networks, metabolic pathways, and biomarkers related to triacylglycerol metabolism and cognitive function.

[0021] As a further preference of this technical solution: The epidemiological information analysis module is used to utilize the health and nutrition data in the NHANES database to conduct epidemiological analysis and analyze the relationship between triacylglycerol intake and cognitive function.

[0022] As a further preference of this technical solution: The comprehensive analysis module is used to comprehensively analyze the relationship between triacylglycerol and cognition by combining gene expression data, health and nutrition data, bioinformatics analysis results, and epidemiological analysis results.

[0023] As a further preference of this technical solution: The signal terminal of the comprehensive analysis module is connected to a result visualization module;

[0024] The result visualization module is used to visually display the analysis results in the form of charts and images, and intuitively present the association and trend between the relationship of triacylglycerol and cognition.

[0025] As a further preference of this technical solution: It further includes a data security module, which is used to perform security guardianship on the system to protect personal privacy and information security.

[0026] As a further preference of this technical solution: The NHANES database is used to store demographic data, dietary data, examination data, laboratory data, questionnaire data, and restricted access data.

[0027] As a further preference of this technical solution: The GEO database is used to store, share, and analyze gene expression data, and the gene expression data in the GEO database is organized by gene.

[0028] Advantages of the present invention:

[0029] 1. The present invention respectively performs statistical analysis on gene expression data, health and nutrition data through a GEO data analysis module and an NHANES data analysis module. By comparing the differences in gene expression between different sample groups and analyzing the relationship between trienoic acid intake and cognitive function assessment results, key analysis results can be provided for data integration. The data integration module integrates the analysis results of GEO data and NHANES data to form a comprehensive understanding of the relationship between trienoic acid and cognition. During the integration process, the data integration module adopts an integration algorithm to ensure the accuracy and consistency of the data;

[0030] 2. The present invention compares the integrated data with known standard data sets or reliable data sources through a data verification module, further verifying the reliability of the integrated data, and thus ensuring that there will be no problems such as data quality degradation and information loss after data integration. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0032] Figure 1 It is a schematic structural diagram of a system for analyzing the relationship between trienoic acid and cognition using GEO and NHANES databases according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0034] Embodiment

[0035] In the prior art, the relationship between trienoic acid and cognitive function can be analyzed by using GEO and NHANES databases. However, the data formats, structures and sources of NHANES and GEO databases are different. After integrating these data, problems such as data quality degradation and information loss often occur, affecting the accuracy of subsequent analysis;

[0036] Please refer to Figure 1, the present invention provides a technical solution: a system for analyzing the relationship between triacids and cognition using GEO and NHANES databases, including a GEO database, an NHANES database, a data collection module, a data processing module, a GEO data analysis module, an NHANES data analysis module, a data integration module, and a data verification module;

[0037] The signal terminals of the GEO database and the NHANES database are both connected to the signal terminal of the data collection module. The signal terminal of the data processing module is respectively connected to the signal terminals of the GEO data analysis module and the NHANES data analysis module. The signal terminal of the data integration module is respectively connected to the signal terminals of the data verification module, the GEO data analysis module, and the NHANES data analysis module;

[0038] The data collection module is used to collect gene expression data and health and nutrition data associated with triacids and cognitive function from the GEO database and the NHANES database respectively, providing a reliable data basis for subsequent analysis;

[0039] The data processing module is used to preprocess the collected data, including data cleaning, format conversion, missing value filling, and standardization, to improve data quality and consistency, enhance the efficiency and accuracy of data analysis, and reduce analysis biases caused by data quality problems;

[0040] Screen out datasets related to triacid metabolism and cognitive function from the preprocessed data, further filter the screened data to remove irrelevant or noisy data, improve the pertinence and relevance of the data, ensure the accuracy and reliability of the analysis results, and reduce the interference of irrelevant data;

[0041] The GEO data analysis module is used to perform statistical analysis on the screened gene expression data, compare the differences in gene expression between different sample groups, identify genes related to triacids and cognitive function. Through the GEO data analysis module, the relationship between gene expression and triacids and cognitive function can be revealed, providing key gene information for subsequent bioinformatics analysis and mechanism research;

[0042] The NHANES data analysis module is used to statistically analyze the health and nutrition data in the NHANES database, analyze the relationship between triacid intake and cognitive function assessment results. Through the NHANES data analysis module, the association between triacid intake and cognitive function can be revealed, providing data support for subsequent epidemiological research and clinical applications;

[0043] The data integration module uses an integration algorithm to integrate the analysis results of the GEO data analysis module and the NHANES data analysis module, forming a comprehensive understanding of the relationship between triacids and cognition. Furthermore, it can solve the problem of difficult data integration, improve the reliability and accuracy of the analysis results, and provide comprehensive data support for subsequent comprehensive analysis and discussion.

[0044] The signal terminal of the data verification module is connected to the signal terminal of the data integration module.

[0045] The data verification module is used to check whether the number of records in the integrated dataset meets the expectations, compare the integrated data with known standard datasets or reliable data sources, and verify the accuracy of the data.

[0046] The method of the data verification module includes the following steps:

[0047] Record count verification: Verify whether the number of records in the integrated dataset matches the expectations. By comparing the record counts before and after integration, it is possible to initially determine whether the data is complete.

[0048] Field integrity: Check whether each field in the dataset has a value. The key fields include sample ID, gene name, triacid level, and cognitive function assessment results.

[0049] Internal consistency: Verify whether the logical relationships between different fields in the dataset are reasonable: there should be some correlation or association between gene expression data and triacid levels or cognitive function assessment results.

[0050] External consistency: Compare the integrated data with known standard datasets or reliable data sources.

[0051] Range verification: Check whether the data values are within a reasonable range. Gene expression values, triacid levels, and cognitive function assessment results usually have specific value ranges, and values outside these ranges may be incorrect and require further verification and correction.

[0052] Precision verification: For numerical data, verify whether its precision meets the requirements. Gene expression data often needs to be accurate to a certain number of decimal places, while cognitive function assessment results may need to be rounded to an integer or a specific number of decimal places.

[0053] For the data verified by the data verification module, it can be directly used for subsequent analysis or applications. For the data that fails to pass the verification of the data verification module, further investigation and analysis are required to determine the cause of the problem and the solution, including data correction, data filling, and data deletion.

[0054] To solve the problems existing in the prior art, an embodiment of the present invention provides a system for analyzing the relationship between trialkyl acid and cognition by using GEO and NHANES databases, and the problems are solved through the above technical solutions:

[0055] The GEO data analysis module and the NHANES data analysis module respectively perform statistical analysis on gene expression data, health and nutrition data. By comparing the differences in gene expression between different sample groups and analyzing the relationship between trialkyl acid intake and cognitive function evaluation results, key analysis results are provided for data integration. The data integration module integrates the analysis results of GEO data and NHANES data to form a comprehensive understanding of the relationship between trialkyl acid and cognition. During the integration process, this module adopts an integration algorithm to ensure the accuracy and consistency of the data. At the same time, the integrated data is compared with known standard data sets or reliable data sources through the data verification module to further verify the reliability of the integrated data.

[0056] In this embodiment, specifically: the integration algorithm includes the following steps:

[0057] Data preparation: Obtain preprocessed gene expression data and health and nutrition data related to trialkyl acid and cognitive function;

[0058] Data matching: According to the common characteristics of the data, match the gene expression data in the GEO database with the health and nutrition data in the NHANES database. For the successfully matched data, perform alignment operations to ensure that the data is consistent in dimension and format for subsequent analysis;

[0059] Data integration: Merge the matched and aligned data into a complete data set, check the merged data set, remove duplicate data records, and perform standardization processing on the merged data to eliminate the dimensional differences and distribution differences between different data sets;

[0060] Data verification: Check whether the integrated data is logically consistent, including whether there is a reasonable association between gene expression data and health and nutrition data, verify whether the integrated data is complete, and verify the accuracy of the integrated data by comparing it with other reliable data sources;

[0061] Through the above steps, the data integration module can effectively integrate the analysis results of GEO and NHANES databases, providing comprehensive and accurate data support for subsequent comprehensive analysis.

[0062] In this embodiment, specifically: the signal terminals of the data integration module are respectively connected to a bioinformatics analysis module, an epidemiological information analysis module, and a comprehensive analysis module;

[0063] The bioinformatics analysis module is used to perform bioinformatics analysis on the selected key genes by combining gene expression data, identify gene networks, metabolic pathways, and biomarkers related to tricarboxylic acid metabolism and cognitive function. Through the bioinformatics analysis module, the biological mechanism between tricarboxylic acid and cognitive function can be deeply revealed, providing potential targets for subsequent drug development and clinical applications.

[0064] In this embodiment, specifically: the epidemiological information analysis module is used to utilize the health and nutrition data in the NHANES database to conduct epidemiological analysis and analyze the relationship between tricarboxylic acid intake and cognitive function. Through the epidemiological information analysis module, the epidemiological association between tricarboxylic acid intake and cognitive function can be revealed, providing a scientific basis for subsequent public health policies and intervention measures.

[0065] In this embodiment, specifically: the comprehensive analysis module is used to comprehensively analyze the relationship between tricarboxylic acid and cognition by combining gene expression data, health and nutrition data, bioinformatics analysis results, and epidemiological analysis results. Through the comprehensive analysis module, a comprehensive and in-depth understanding of the relationship between tricarboxylic acid and cognition can be formed, providing strong support for subsequent research and clinical applications.

[0066] In this embodiment, specifically: the signal end of the comprehensive analysis module is connected to a result visualization module;

[0067] The result visualization module is used to visually display the analysis results in the form of charts and images, intuitively presenting the association and trend between the relationship of tricarboxylic acid and cognition. Through the result visualization module, the readability and understandability of the analysis results can be improved, providing an intuitive display for researchers and clinicians.

[0068] In this embodiment, specifically: it further includes a data security module, which is used to monitor the security of the system, protect personal privacy and information security, improve the security and credibility of the data, and provide a strong guarantee for the stable operation and legal use of the system;

[0069] The data security module ensures the privacy and security of the data in the following ways;

[0070] Formulate strict data security policies: clarify the security requirements for all aspects such as data classification, storage, use, transmission, and destruction, and ensure that the entire life cycle of the data is properly managed;

[0071] Establish data access permission control: set different data access permissions according to the roles and responsibilities of users to ensure that only authorized users can access sensitive data;

[0072] Data Encryption: Sensitive data is encrypted for storage to ensure that even if the data is illegally obtained, its content cannot be directly read. The encryption method uses asymmetric encryption;

[0073] Firewall and Intrusion Detection System: A firewall is deployed to prevent unauthorized access. At the same time, an intrusion detection system is used to monitor network traffic to detect and respond to potential security threats in a timely manner;

[0074] Fine-grained Access Control: Based on the roles and permissions of users, fine-grained control is exerted over data access to ensure that users can only access the data within their authorized scope.

[0075] In this embodiment, specifically: The NHANES database is used to store demographic data, dietary data, examination data, laboratory data, questionnaire data, and limited access data;

[0076] The GEO database is used to store, share, and analyze gene expression data. The gene expression data in the GEO database is organized by gene.

[0077] Working principle or structural principle: In use, gene expression data and health and nutrition data associated with triacids and cognitive function are respectively collected from the GEO database and the NHANES database through a data collection module. The collected data is preprocessed by a data processing module, and datasets related to triacid metabolism and cognitive function are screened out from the preprocessed data. The gene expression data and health and nutrition data are respectively statistically analyzed by a GEO data analysis module and an NHANES data analysis module to compare the differences in gene expression between different sample groups and analyze the relationship between triacid intake and cognitive function assessment results. Then, a data integration module uses an integration algorithm to integrate the analysis results of the GEO data analysis module and the NHANES data analysis module. Bioinformatics analysis and epidemiological analysis are respectively carried out by a bioinformatics analysis module and an epidemiological information analysis module. Then, through a comprehensive analysis module, combining gene expression data, health and nutrition data, bioinformatics analysis results, and epidemiological analysis results, a comprehensive analysis of the relationship between triacids and cognition is performed, and the analysis results are presented in the form of charts and images.

[0078] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A system for analyzing the relationship between trialkanoic acid and cognition using GEO and NHANES databases, characterized in that: It includes GEO database, NHANES database, data acquisition module, data processing module, GEO data analysis module, NHANES data analysis module, data integration module and data verification module; The signal end of the GEO database and the signal end of the NHANES database are both connected to the signal end of the data acquisition module, the signal end of the data processing module is respectively connected to the signal end of the GEO data analysis module and the signal end of the NHANES data analysis module, and the signal end of the data integration module is respectively connected to the signal end of the data verification module, the signal end of the GEO data analysis module, and the signal end of the NHANES data analysis module; The data collection module is used to collect gene expression data and health nutrition data associated with trialkanoic acid and cognitive function from the GEO database and the NHANES database respectively; The data processing module is used to preprocess the collected data, including data cleaning, format conversion, missing value filling and standardization, screen out data sets related to trialkanoic acid metabolism and cognitive function from the preprocessed data, and further filter the screened data to remove irrelevant or noise data; The GEO data analysis module is used to perform statistical analysis on the screened gene expression data, compare the differences in gene expression between different sample groups, and identify genes related to trialkanoic acid and cognitive function; The NHANES data analysis module is used to collect statistics on health and nutrition data in the NHANES database and analyze the relationship between trialkanoic acid intake and cognitive function assessment results; The data integration module uses an integration algorithm to integrate the analysis results of the GEO data analysis module and the NHANES data analysis module to form a comprehensive understanding of the relationship between trialkanoic acids and cognition.

2. The system for analyzing the relationship between trialkanoic acid and cognition using GEO and NHANES databases according to claim 1, characterized in that: The integration algorithm comprises the following steps: Data preparation: Obtain pre-processed gene expression data and health nutrition data related to trialkanoic acid and cognitive function; Data matching: Gene expression data in the GEO database were matched with health and nutrition data in the NHANES database based on common features of the data; Data integration: merge the matched and aligned data into a complete data set, check the merged data set, remove duplicate data records, standardize the merged data, and eliminate the dimensional and distribution differences between different data sets; Data verification: Check whether the integrated data is logically consistent, including whether there is a reasonable association between gene expression data and health nutrition data, and verify whether the integrated data is complete.

3. The system for analyzing the relationship between trialkanoic acid and cognition using GEO and NHANES databases according to claim 1, characterized in that: The signal end of the data verification module is connected to the signal end of the data integration module; The data verification module is used to check whether the number of records in the integrated data set meets expectations, compare the integrated data with a known standard data set or a reliable data source, and verify the accuracy of the data.

4. The system for analyzing the relationship between trialkanoic acid and cognition using GEO and NHANES databases according to claim 1, characterized in that: The signal end of the data integration module is respectively connected to the biological information analysis module, the epidemic information analysis module and the comprehensive analysis module; The bioinformatics analysis module is used to perform bioinformatics analysis on the selected key genes in combination with gene expression data, and identify gene networks, metabolic pathways and biomarkers related to trialkanoic acid metabolism and cognitive function.

5. The system for analyzing the relationship between trialkanoic acid and cognition using GEO and NHANES databases according to claim 4, characterized in that: The epidemiological information analysis module is used to use the health and nutrition data in the NHANES database to conduct epidemiological analysis and analyze the relationship between trialkanoic acid intake and cognitive function.

6. The system for analyzing the relationship between trialkanoic acid and cognition using GEO and NHANES databases according to claim 4, characterized in that: The comprehensive analysis module is used to combine gene expression data, health and nutrition data, bioinformatics analysis results and epidemiological analysis results to conduct a comprehensive analysis of the relationship between trialkanoic acid and cognition.

7. The system for analyzing the relationship between trialkanoic acid and cognition using GEO and NHANES databases according to claim 4, characterized in that: The signal end of the comprehensive analysis module is connected to a result visualization module; The result visualization module is used to visualize the analysis results in the form of charts and images, and intuitively present the association and trend between trialkanoic acid and cognitive relationship.

8. The system for analyzing the relationship between trialkanoic acid and cognition using GEO and NHANES databases according to claim 1, characterized in that: It also includes a data security module, which is used to monitor the security of the system and protect personal privacy and information security.

9. The system for analyzing the relationship between trialkanoic acid and cognition using GEO and NHANES databases according to claim 1, characterized in that: The NHANES database is used to store demographic data, dietary data, examination data, laboratory data, questionnaire data, and limited access data.

10. The system for analyzing the relationship between trialkanoic acid and cognition using GEO and NHANES databases according to claim 1, characterized in that: The GEO database is used to store, share and analyze gene expression data. The gene expression data in the GEO database is organized on a gene basis.