Microorganism-disease relation mining research and system construction method based on Neo4j

By constructing a microbial-disease relation database and data query system in the Neo4j graph database, and using graph machine learning algorithms to mine the potential relationship between microbial-disease, the problem of difficulty in exploring the relationship between microbial and brain diseases in the existing technology is solved, and more accurate predictions and better scientific research value are achieved.

CN120183728APending Publication Date: 2025-06-20GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510329378.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing technologies are difficult to effectively explore and predict the potential relationship between microbial and brain diseases, limiting the development of new strategies for disease prevention and treatment.

Method used

The graph machine learning algorithm based on Neo4j is used to build a microbial-disease relational database and data query system, and the relationship knowledge graph is displayed through the Neo4j graph database, and the graph algorithm is used to mine the potential relationship between microbial-disease.

Benefits of technology

It realizes dynamic modeling and visual display of complex relationships between microorganisms and diseases, improves the prediction accuracy of microorganism-disease relationships, and has good scientific research value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183728A_ABST
    Figure CN120183728A_ABST
Patent Text Reader

Abstract

The invention mainly relates to an algorithm, a relational database and a non-relational database used for predicting the relationship between microorganisms and diseases. The method comprises the following steps: S1, constructing a microorganism-disease relational database; s2, establishing a microorganism-disease data query system; s3, connecting a Neo4j graph database, and displaying a relational knowledge graph; s4, mining a microorganism-disease potential relationship by adopting a graph algorithm; according to the method, a system, a graph database, an embedding algorithm and the like are combined to excavate the potential relationship between the microorganisms and the brain diseases, new insights are provided for researchers in the field, guidance of prevention and diagnosis of the diseases is facilitated, and valuable reference is provided for further research of the potential relationship between the microorganisms and the diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to algorithms, relational and non-relational databases used for predicting the relationships between microorganisms and diseases. Background Art

[0002] Microorganisms: are microscopic living organisms, covering various types. These living organisms participate in ecological operations through material cycles in the human body, especially forming unique microecosystems in the host body. The intestinal microbiota, as a typical representative, establishes a symbiotic and mutually beneficial relationship with the host through bacteria, and plays key functions in nutrient metabolism, neurotransmitter synthesis, and immune balance. This microbial population is mainly composed of genera such as Clostridium, Lactobacillus, Bifidobacterium, Salmonella, Enterobacter, etc. The stable state of these genera will directly affect physical health. Therefore, a reasonable proportion of microbial genera can not only maintain intestinal health, but also have potential intervention value for disease prevention and control.

[0003] Brain diseases: refer to neurological disorders caused by abnormal brain function or structure, including neurodegenerative diseases (such as Alzheimer's disease, Parkinson's disease), cerebrovascular diseases (such as stroke, cerebral hemorrhage), and mental and psychological diseases (such as anxiety disorder, bipolar disorder), etc. The pathogenesis of these diseases is complex, and their symptoms often manifest as cognitive function decline, mood swings, etc. The gut microbiota forms a two-way interaction with the brain through the "gut-brain axis", affecting the development and function of the nervous system. Therefore, intervening in the gut microbiota by adjusting the diet structure or using specific probiotics may become a new strategy for regulating the function of the nervous system and preventing or alleviating brain diseases. This mechanism provides a new research direction and treatment potential for the prevention and treatment of brain diseases.

[0004] Machine learning: is the core field of artificial intelligence. Its core idea is to extract patterns from data and use these patterns for prediction, etc. Machine learning is mainly divided into supervised learning (classification or regression tasks), unsupervised learning, semi-supervised learning, and reinforcement learning. Graph machine learning is an important branch of machine learning. Its graph data consists of nodes, edges, and global attributes, and exists in scenarios such as molecular structures, knowledge graphs, etc. The core tasks include node classification, link prediction, graph classification, etc., and also include key technologies such as graph embedding and graph neural networks (GNN). Therefore, the present invention uses a graph machine learning algorithm embedded in Neo4j to predict the potential relationships between microorganisms and diseases, providing new technical methods and guidance for the development of this field.

[0005] A proportionally coordinated gut microbiota not only improves the digestive function of the gut in the human body but also has a profound impact on the occurrence of diseases. Among them, brain diseases are directly or indirectly maintained by the microbiota. Exploring the potential relationship between the microbiota and diseases has a profound impact on the field of bioinformatics. Therefore, using technical fields such as machine learning algorithms and databases to mine the relationship is a new technical method, which can improve the value of the relationship between the microbiota and diseases and discover more potential relationships, providing value reference and efficiency for this field. Summary of the Invention

[0006] The purpose of the present invention is to provide a method for mining the relationship between microbiota and diseases based on Neo4j and constructing a system, which is convenient for researchers to conduct the latest research and provide new insights.

[0007] To achieve the above purpose, the present invention provides a method for mining the relationship between microbiota and diseases based on Neo4j and constructing a system, including:

[0008] S1: Design a suitable database for the structured data, and after screening the data, use MySQL to construct a microbiota-disease relationship database;

[0009] S2: Based on the convenient query of the data, a microbiota-disease data query system needs to be established, where the content of the query includes all its field attributes;

[0010] S3: Use Python syntax tools to connect to the Neo4j graph database, set the login user and password, and construct a relationship knowledge graph in the form of "microbiota → disease";

[0011] S4: Mine the potential relationship between microbiota and diseases based on the graph algorithms in the graph database Neo4j;

[0012] Among them, the construction of the microbiota-disease relationship database includes:

[0013] S11: Design a suitable and overall database architecture according to the specific types of data such as the structured form;

[0014] S12: Establish the fields of the database table according to the name, length, data type, name ID, and introduction, etc.;

[0015] S13: Select the structured MySQL database as the management system according to the microbiota-disease data in the present invention;

[0016] Among them, the establishment of the microbiota-disease data query system includes:

[0017] S21: Adopt a search box and classification labels, and design the query system style using keywords and filtering functions;

[0018] S22: Integrate microorganism-disease data and establish a GUI query system that supports querying in multiple ways such as by name and ID;

[0019] Among them, connecting to the Neo4j graph database and displaying the relational knowledge graph includes:

[0020] S31: Configure corresponding information according to Python language tools to connect to the Neo4j graph database and ensure secure access to the system;

[0021] S32: Use microorganisms and diseases as nodes respectively and include various attributes, and the microorganism-disease relationship will be used as an edge to construct the microorganism-disease graph together with the nodes;

[0022] S33: Display the relational knowledge graph based on the Neo4j visualization interface and support user interaction operations;

[0023] Among them, the use of graph algorithms to mine potential microorganism-disease relationships includes:

[0024] S41: Use the Path Finding algorithm embedded in Neo4j to define graph traversal and explore potential associations between microorganisms and diseases;

[0025] S42: Calculate the similarity between microorganisms and diseases using graph embedding algorithms based on the node data of microorganisms and diseases;

[0026] A method for mining research and system construction of microorganism-disease relationships based on Neo4j of the present invention includes: constructing a microorganism-disease relationship database; establishing a microorganism-disease data query system; connecting to the Neo4j graph database and displaying the relational knowledge graph; using graph algorithms to mine potential microorganism-disease relationships. The beneficial effects of the present invention are: by constructing a microorganism-disease relationship database and a system to clearly query structured relational data, using the Neo4j graph database to improve storage and query efficiency, realizing dynamic modeling and visual display of the complex relationships between microorganisms and diseases, and mining potential microorganism-disease relationships through graph algorithms to improve the prediction accuracy, which has good scientific research value. Description of the Drawings

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0028] Figure 1 is a flowchart of a method for mining research and system construction of microorganism-disease relationships based on Neo4j.

[0029] Figure 2 is the flowchart of constructing the microorganism-disease relationship database.

[0030] Figure 3 is the flowchart of establishing the microorganism-disease data query system.

[0031] Figure 4 is the flowchart of connecting to the Neo4j graph database to display the relationship knowledge graph.

[0032] Figure 5 is the flowchart of mining potential relationships between microorganisms and diseases using graph algorithms. Detailed implementation manners

[0033] In order to enable researchers in this field to more clearly understand the technical solution of the present invention and be able to implement it, the present invention will be described in detail below in combination with specific embodiments and the accompanying drawings, where the same symbols or labels represent the same content. The technical solutions described below and their specific embodiments are all exemplary, only used to represent the technical solutions of the present publication, and the embodiments all belong to the scope of the present invention application, and should not be construed as a limitation to the present invention.

[0034] Before introducing the technical solution of the present invention, the noun terms involved in the embodiments will be explained first.

[0035] Neo4j: It is a management system of a graph database, which stores and processes data in nodes and edges in the graph, rather than a common structured database. Its core advantage is that it can efficiently process complex relationship networks, making it widely used in fields such as bioinformatics. For example, by modeling the relationships between microorganisms, diseases, genes, and metabolites as a graph, Neo4j can more intuitively reveal the potential relationships between these entities, thereby increasing the efficiency of drug research and development and disease prevention. The Neo4j graph database is not limited to data storage and query, but can also mine potential biological meanings through graph algorithms (such as path analysis, centrality calculation, etc.), providing new insights for scientific research.

[0036] Please refer to Figures 1 to 4 , the present invention is a method for extracting the relationship between microorganisms and brain diseases based on deep learning, including the following steps of the embodiment:

[0037] S1 Construct a microorganism-disease relationship database;

[0038] S11: Design the overall database architecture;

[0039] Specifically, first design a suitable database according to the specific type of data, define the overall database architecture, and design different data tables to store different data types, including brain disease data, microbiome data, and relationship data, etc.

[0040] S12: Establish the fields of the data table;

[0041] Specifically, to establish a data table, it is necessary to determine the fields of each data table, including name, length, data type, name ID, and introduction, etc., to comprehensively display the structured microorganism-disease data table, and three data tables of microorganisms, diseases, and relationships have been determined.

[0042] S13: Select a database management system;

[0043] Specifically, select a suitable data management system according to the data type. For a huge amount of data, a distributed relational database can be selected, which is suitable for scenarios such as strong consistency and complex business; for the key-value pair data type, a NoSQL database can be selected, which has characteristics such as high concurrent read and write, caching, etc. and supports horizontal expansion. For the present invention, the MySQL database is selected, which has advantages such as flexibility and reliability, is suitable for medium-sized data volumes such as microorganism-disease, can run on multiple platforms, and ensures data consistency and integrity.

[0044] Constructing a database is a high-quality processing of data, which helps with applied research and scientific innovation, and enables researchers to clearly see the changes and associations of data.

[0045] S2 Establish a microorganism-disease data query system;

[0046] S21: Design the style of the query system;

[0047] Specifically, first design the interface layout of the system, define the functional areas including the navigation bar, query input area, result display area, and data export, to ensure that users can use it smoothly in this scenario and meet the simple retrieval style of microorganism-disease relationship data;

[0048] S22: Establish a microorganism-disease data query system;

[0049] Specifically, build a query system on the basis of the system architecture design, use the Python Flask framework to build a RESTful API service, including functions such as query, download, and display, and connect to the Neo4j graph database to help researchers better process data and provide reliable technical support for the research of the microbiome-disease.

[0050] S3: Connect to the Neo4j graph database and display the relationship knowledge graph;

[0051] S31: Connect to the Neo4j graph database;

[0052] Specifically, before connecting to the graph database, the server needs to be configured, including parameter configuration, port configuration, etc., and user permissions are set to achieve fine-grained access control, and then a test connection is made;

[0053] S32: Construct microbe-disease relationship nodes and edges;

[0054] Specifically, first define the attributes of microbes and diseases to construct nodes M (microbe nodes) and D (disease nodes), including information attributes such as name, ID, intro, etc., determine the relationship edges through the relationship type, such as correlation, inhibition, etc., and establish a unified data standard to ensure the correctness and uniqueness of nodes and edges.

[0055] S33: Display the relationship knowledge graph;

[0056] Specifically, use the native visualization component technology of the graph database to render the front end, and through the Cypher statement:

[0057] "MATCH path=(m:Microbe)-[r]->(d:Disease) RETURN path"

[0058] Query to obtain all microbe-disease graph data, and click on the data in the graph to obtain the node attributes and the node and edge information of the affiliated relationship graph.

[0059] S4 Graph algorithm to mine potential microbe-disease relationships;

[0060] S41: Use the PathFinding algorithm to discover potential associations;

[0061] Specifically, use the Dijkstra algorithm in the shortest path to calculate the shortest path from microbes to diseases. For a given microbe M i and disease D j , the formula that can be represented is:

[0062]

[0063] where P is the path from M i to D j , which can be expressed as the weight of each edge in the path P, and d(M i , D j ) is the shortest path between microbe M i and D j . Its Neo4j syntax is:

[0064] "MATCH (m:Microbe{name:'Microbial Name'}),(d:Disease{name:'Disease Name'}) MATCH p = shortestPath((m)-[*]-(d)) RETURN p"

[0065] Discover the relationship path indirectly or directly through this method to improve the efficiency of path discovery.

[0066] S42: Calculate the similarity between microbes and diseases;

[0067] Specifically, to calculate the similarity, first use the Node2Vec node embedding algorithm to generate paths through random walk and update the microbial embedding vector v and the disease embedding vector v'. Its Neo4j syntax is:

[0068] "CALL gds.node2Vec.stream('microbeDiseaseGraph',{embeddingDimension:128,walkLength:10,walksPerNode:10,randomSeed:42}) YIELD nodeId,embedding WITH gds.util.asNode(nodeId) AS node,embedding RETURN node.name AS nodeName,embedding"

[0069] Specifically, use cosine to calculate the microbe-disease similarity. The formula can be expressed as:

[0070]

[0071] The embedding similarity can further strengthen the potential of the relationship.

[0072] Specifically, according to the calculated shortest path and similarity, comprehensively measure the strength of the potential relationship between microbes and diseases by combining the two:

[0073]

[0074] By combining the shortest path and the node embedding graph algorithm, comprehensively consider the distance information and semantic information between nodes, and directly or indirectly mine the potential relationship between microbes and diseases.

Claims

1. A method for studying and building a system for mining microorganism-disease relationships based on Neo4j, characterized in that: include: S1: Design a suitable database for structured data, and use MySQL to build a microorganism-disease relationship database after screening the data; S2: Based on the convenient query of data, a microorganism-disease data query system needs to be established, in which the query content includes all its field attributes; S3: Use Python syntax tools to connect to the Neo4j graph database, set the login user and password, and build a relational knowledge graph in the form of "microorganism → disease"; S4: Mining potential relationships between microorganisms and diseases based on graph algorithms in the graph database Neo4j.

2. A method for studying and building a system for mining microorganism-disease relationships based on Neo4j as described in claim 1, characterized in that: The construction of a microorganism-disease relationship database comprises: S11: Design an appropriate and overall database architecture based on the specific data type such as structured form; S12: Establishing fields of a database table according to the name, length, data type, name ID, and description; S13: According to the microorganism-disease data of the present invention, a structured MySQL database is selected as a management system.

3. A method for studying and building a system for mining microorganism-disease relationships based on Neo4j as described in claim 2, characterized in that: The establishment of a microorganism-disease data query system comprises: S21: Use search boxes and category labels, and use keywords and filtering functions to design the query system style; S22: Integrate microbial-disease data and support multiple query methods such as name and ID to establish a GUI query system.

4. A method for studying and building a system for mining microorganism-disease relationships based on Neo4j as described in claim 3, characterized in that: The connection to the Neo4j graph database to display the relational knowledge graph includes: S31: Configure corresponding information according to the Python language tool to connect to the Neo4j graph database and ensure secure access to the system; S32: Microorganisms and diseases are taken as nodes respectively and contain various attributes. The microorganism-disease relationship is taken as an edge to combine the nodes together to construct a microorganism-disease graph; S33: Display the relational knowledge graph based on the Neo4j visualization interface and support user interactive operations.

5. A method for studying and building a system for mining microorganism-disease relationships based on Neo4j as described in claim 4, characterized in that: The method of using graph algorithms to mine potential relationships between microorganisms and diseases includes: S41: Use the Path Finding algorithm embedded in Neo4j to define graph traversal and explore potential microbial-disease associations; S42: Calculate the similarity between microorganisms and diseases using a graph embedding algorithm based on the node data of microorganisms and diseases.