A Recommendation Method for Microbial Synthesis of Nanomaterials Based on Knowledge Graph

By constructing a recommendation method for microbial synthesis of nanomaterials based on knowledge maps, the BERT-RGCN model is used to extract the characteristic vectors of microorganisms and nanomaterials, which solves the problems of low efficiency and high cost in traditional methods, and achieves efficient and accurate analysis of the relationship between microorganisms and nanomaterials.

CN119598006BActive Publication Date: 2025-07-22COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411534489.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-30
Publication Date
2025-07-22
Estimated Expiration
2044-10-30

AI Technical Summary

Technical Problem

The existing technology is difficult to efficiently and systematically explore and analyze the complex relationship between microorganisms and nanomaterials. The traditional methods are inefficient, costly and have great uncertainty in results, making it difficult to adapt to diversified research needs.

Method used

A recommendation method for microbial synthetic nanomaterials based on knowledge graph is constructed. By obtaining microbial and nanomaterial information, a knowledge graph is constructed, structural and semantic feature vectors are extracted using the BERT-RGCN model, and node scoring is performed to judge the relationship.

Benefits of technology

It realizes the hidden relationship between microorganisms and potential nanomaterials from large-scale knowledge graphs, provides comprehensive and systematic analysis support, reduces research costs and time, and improves the accuracy and consistency of results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119598006B_ABST
    Figure CN119598006B_ABST
Patent Text Reader

Abstract

The present invention relates to the fields of bioinformatics and artificial intelligence technologies, and particularly relates to a method for recommending microbial synthesized nanomaterials based on a knowledge graph. Microbial information and nanomaterial information are obtained; a knowledge graph is constructed through the microbial information and the nanomaterial information; wherein, the nodes of the knowledge graph are composed of microorganisms, nanomaterials, synthesis methods, and elements; a structural feature vector and a semantic feature vector are obtained based on the knowledge graph; the structural feature vector and the semantic feature vector are spliced to obtain a representation vector of each node; each node is scored based on the representation vector, and the relationship between the microorganism and the nanomaterial is judged according to the scoring result. The present invention can mine the implicit association between microorganisms and potential nanomaterials from a large-scale knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of bioinformatics and artificial intelligence, and particularly relates to a method for recommending microbial synthesized nanomaterials based on a knowledge graph. Background Art

[0002] The rapid development of synthetic biology has promoted the wide application of microorganisms in industrial production. Microorganisms are widely used in the synthesis of nanomaterials and have gradually become an important tool in multiple fields such as chemical engineering, medicine, energy, and materials due to their high production efficiency, mild reaction conditions, environmental friendliness, etc. For example, microorganisms such as lactic acid bacteria, yeast, and streptomyces have been successfully applied to the synthesis of various biological materials. However, with the expansion of the application fields, the complex relationship between microorganisms and nanomaterials has become increasingly prominent, and researchers urgently need an efficient tool to systematically explore and analyze these relationships.

[0003] In traditional research, the synthesis of microbial nanomaterials mainly relies on the knowledge accumulation of experts and experimental verification. These research methods usually include steps such as literature retrieval, experimental verification, and data analysis. Researchers screen out microorganisms that may synthesize nanomaterials by consulting a large number of scientific research literatures and then verify them through experiments. The main advantage of this method is the relatively high reliability of its research results, and the experimental data directly supports the correlation between microorganisms and nanomaterials. However, the disadvantages of this method are also very obvious.

[0004] For example, the efficiency of manual literature retrieval and screening is low. With the growth of the number of scientific research literatures, the amount of information faced by researchers is increasing, making it difficult to comprehensively and systematically consult and analyze all literatures related to the research. This not only increases the workload of researchers but also may lead to the omission of some potential correlation relationships between microorganisms and nanomaterials. Secondly, the experimental verification has a long cycle and high cost. The correlation between each microorganism and nanomaterial needs to be verified through experiments, which usually involves complex biological culture, chemical analysis, etc. These experiments are not only time-consuming but also require a large amount of experimental materials and equipment. In addition, due to the complexity of experimental conditions, the results between different laboratories may vary, further increasing the uncertainty of the research. Therefore, traditional research methods are difficult to meet diverse research needs. Researchers not only focus on the direct relationship between microorganisms and nanomaterials but also hope to understand their potential correlations. These complex correlation relationships are difficult to fully reveal through traditional methods.

[0005] In recent years, with the development of big data technology, knowledge graphs have gradually become an important tool for studying complex systems and association mining. A knowledge graph is a structured knowledge representation method that visually expresses different entities and their relationships in the form of nodes and edges. A knowledge graph can not only integrate multi-source heterogeneous data but also reveal potential associations in complex systems through graph algorithms.

[0006] In the field of bioinformatics, the application of knowledge graphs has gradually deepened. Researchers have used knowledge graphs to integrate a large amount of bioinformatics data, such as genomic data, proteomic data, metabolic network data, etc. These data are integrated and analyzed through knowledge graphs, which can provide researchers with a new perspective and reveal some potential associations that are difficult to discover by traditional methods. For example, researchers have used knowledge graphs to discover potential associations between certain genes and diseases, promoting the development of personalized medicine.

[0007] In the research of microbial synthesis of nanomaterials, knowledge graphs also have broad application prospects. By constructing a knowledge graph of the association between microorganisms and nanomaterials, researchers can systematically integrate multi-source data such as literature and experimental data, and structurally express microorganisms, nanomaterials, and their properties and relationships. A knowledge graph can not only intuitively display the direct relationship between microorganisms and nanomaterials but also mine the potential association between them through graph algorithms.

[0008] Although knowledge graphs have shown great potential in the association mining of microorganisms and nanomaterials, their applications still face many challenges. For example, the acquisition and processing of data are still one of the main difficulties in constructing knowledge graphs. In the research of microorganisms and nanomaterials, data sources are extensive, including scientific research literature, experimental data, databases, etc. These data often have heterogeneity, with different data formats, structures, and contents. How to effectively collect, clean, and standardize these data to ensure their accuracy and consistency is the primary task in constructing knowledge graphs. In addition, how to maintain the simplicity of its structure and the accuracy of information while ensuring the scale of the graph is the key to constructing a high-quality knowledge graph. With the continuous progress of scientific research and the emergence of new data and new knowledge, knowledge graphs need to be continuously updated and maintained to ensure their timeliness and practicality. The association mining of knowledge graphs relies on advanced graph algorithms and machine learning technologies. Traditional graph algorithms, such as path analysis and clustering analysis, can reveal some simple association relationships in the graph, but for complex biological systems, the capabilities of traditional algorithms are limited. Summary of the Invention

[0009] A method for recommending microbial synthesis of nanomaterials based on a knowledge graph, comprising:

[0010] Obtaining microbial information and nanomaterial information;

[0011] Construct a knowledge graph based on the microbial information and the nanomaterial information; wherein, the nodes of the knowledge graph are composed of microorganisms, nanomaterials, synthesis methods, and elements;

[0012] Obtain structural feature vectors and semantic feature vectors based on the knowledge graph;

[0013] Concatenate the structural feature vectors and the semantic feature vectors to obtain the representation vectors of the respective nodes; score the respective nodes based on the representation vectors, and judge the relationship between microorganisms and nanomaterials according to the scoring results.

[0014] Preferably, in the ontology of the knowledge graph, the entity labels include: organisms, habitats, mechanisms, functions, metal resistance, nanomaterials, synthesis methods, precursors, pH values, temperatures, rotation speeds, positions, sizes, shapes, applications, antibacterial agents, electrons, and energy.

[0015] Preferably, the knowledge graph uses the BERT-RGCN model to obtain the structural feature vectors and the semantic feature vectors, wherein the BERT-RGCN model includes a graph structure processing branch and a text data processing branch.

[0016] Preferably, extract the features of the graph structure through the graph structure processing branch.

[0017] Preferably, capture the semantic information of the text data through the text data processing branch.

[0018] Preferably, the extraction steps include:

[0019] Input the multi-relational graph obtained based on the knowledge graph;

[0020] Update the node representations in the multi-relational graph based on the relationship types;

[0021] Output the structural feature vectors of the nodes based on the updated node representations.

[0022] Preferably, the capture steps include:

[0023] Input the text sequence obtained based on the knowledge graph;

[0024] Convert the text sequence into word vectors;

[0025] The word vectors output the semantic feature vectors of the nodes through an encoder.

[0026] Preferably, use the mean reciprocal rank as the evaluation index of the method.

[0027] The beneficial effects of the present invention are as follows:

[0028] By applying emerging technologies such as deep learning and graph neural networks to the association mining of knowledge graphs, researchers can mine more complex potential associations from large-scale knowledge graphs, such as the implicit associations between microorganisms and potential nanomaterials. Researchers can use knowledge graphs to comprehensively and systematically analyze the complex relationships between microorganisms and nanomaterials, thereby providing theoretical support for the development and optimization of nanomaterials. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 This is a flowchart of an embodiment of the present invention.

[0030] Figure 2 This is a flowchart of the model training stage in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The following will Figure 1 describe the embodiments of the present invention in detail. These embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.

[0032] I. Obtaining Data

[0033] Obtain data, and the sources of the data include at least literature data and NCBI data.

[0034] The literature data includes at least titles, microorganisms, synthesis methods, and nanomaterials, and the above information is manually annotated.

[0035] The NCBI data includes at least microorganism information and microorganism mechanism information.

[0036] The above data provides basic data support for the construction of the knowledge graph. For the missing data, it is obtained by querying the LLM.

[0037] II. Data Preprocessing

[0038] Since the sources of the above data are different, the data formats and structures vary greatly.

[0039] Preprocess the obtained data, including: data cleaning, data standardization, and data integration operations.

[0040] Data cleaning includes: removing duplicate data, handling missing values, and correcting errors, aiming to remove or correct the noise, redundancy, and incorrect data in the data.

[0041] In the steps of missing value processing, different handling methods are adopted for missing data according to different situations. If the microorganism or nanomaterial is missing in the data annotation, the annotated data will be deleted; if the synthesis method is missing, the synthesis method will be directly replaced with a number during processing; if the mechanism of the microorganism is missing, the name node of the microorganism will be searched in the NCBI database for the corresponding microorganism mechanism information. If there is no corresponding microorganism mechanism in the NCBI database, the LLM will be used to generate the mechanism corresponding to the microorganism.

[0042] Data standardization aims to make data from different sources have a unified format and structure for subsequent graph construction. For example, the annotated data "ag", "Ag", "Agg", "silver" all represent the "silver" material, and they are unified as "silver".

[0043] Data integration aims to integrate data from different sources to form a unified whole logically. The process of data integration includes data mapping, entity alignment, and relationship integration operations.

[0044] III. Construction of Knowledge Graph

[0045] The construction process of the knowledge graph includes steps of defining nodes, defining edges, entity selection, relationship selection, graph construction, and graph storage.

[0046] Define nodes. Nodes represent entities in the knowledge graph. In the present invention, the node classification is as follows:

[0047] Microorganism nodes: Represent the types of microorganisms involved in the production of nanomaterials.

[0048] Material nodes: Represent the nanomaterials synthesized by microorganisms.

[0049] Synthesis method nodes: Represent how microorganisms synthesize nanomaterials in a piece of literature.

[0050] Element nodes: Represent which elements are included in each nanomaterial.

[0051] Define edges. Edges represent the relationships between entities. In the present invention, the edge classification is as follows:

[0052] Synthesis relationship: Represents that a certain microorganism can synthesize a certain specific nanomaterial.

[0053] Containment relationship: Represents which elements are included in each nanomaterial.

[0054] Select microorganisms, nanomaterials, synthesis methods, and elements from the above data as entity nodes in the knowledge graph to complete entity selection; connect the microorganism and material nodes in the same document through a synthesis relationship, and connect the material with its constituent elements to complete relationship selection.

[0055] After both entities and relationships are selected, combine them to construct a knowledge graph. The constructed graph is a multi-attribute directed graph, where nodes are connected by edges. Each node and edge contain multiple attribute information. For example, the material node contains shape, size, and location attributes. The knowledge graph can not only display the entities and relationships in the data but also provide support for subsequent algorithms.

[0056] To construct a domain knowledge dataset, the present invention designs an ontology. In the ontology, the present invention provides entity labels, including data such as microorganisms, habitats, mechanisms, functions, metal resistance, nanomaterials, synthesis methods, precursors, pH value, temperature, rotation speed, position, size, shape, applications, antibacterial agents, electrons, and energy.

[0057] Graph storage. The constructed knowledge graph needs to be stored for subsequent querying and analysis. In the present invention, the Neo4j graph database is used for storage.

[0058] IV. Model Training

[0059] To process graph structure information and text information, the present invention designs a BERT-RGCN model, as Figure 2 shown. It includes two branches: the graph structure processing branch and the text processing branch. These two branches are responsible for processing graph data and text data respectively, and fusing their features in subsequent layers.

[0060] 1. Graph Structure Processing Branch

[0061] Use RGCN as the basic model to extract the features of the graph structure. Pooling layer: Perform a pooling operation on the feature map output by the convolutional layer to reduce the feature dimension and improve the robustness of the model. Fully connected layer: Connect the output of the pooling layer to one or more fully connected layers to learn higher-level image feature representations.

[0062] The RGCN model is a neural network model specifically designed to process multi-relational graph data. It can capture the complex relationships between nodes and their neighboring nodes in the graph through convolutional operations. The RGCN model takes nodes and their connection relationships as input and gradually updates the representation vectors of nodes through multiple layers of convolutional operations. Each layer of convolutional operation can further extract and integrate the local neighborhood information of nodes, so that the finally output node representation can fully reflect its position and role in the entire graph structure. Through training, the RGCN model can generate the structural feature vectors of each node, and these vectors represent the local and global information of nodes in the graph structure. Since the RGCN model captures the dependency relationships between nodes through relational convolutions, it is particularly suitable for processing knowledge graph data with complex relationships.

[0063] The node representation vectors generated by the RGCN model are further subjected to non-linear transformation (ReLU activation function) and linear transformation. Through these steps, the structural features of nodes are mapped into a new feature space for integration with subsequent semantic features. The processing at this stage helps the model to more effectively fuse feature information from different sources in subsequent steps.

[0064] RGCN is an extension of the graph neural network (GNN) for processing heterogeneous graphs, that is, graphs with multiple different types of nodes and edges. A multi-relational graph is a complex graph structure in which different types of nodes are connected by multiple types of relationships. In the standard graph convolutional network (GCN), the edge type information is usually ignored, while RGCN retains the edge type information by introducing relational matrices and corresponding transformations, so as to better capture the rich semantic relationships in the graph. The basic structure of RGCN is similar to that of the standard GCN, but its core difference lies in the introduction of transformation matrices for different relationship types. The input of the RGCN model is a multi-relational graph, represented as:

[0065] G=(V, E, R)

[0066] where V is the set of nodes, E is the set of edges, and R is the set of relationships. Each layer of RGCN updates the representation of nodes, and this update process depends on the neighboring nodes of the nodes and their connection relationships. For each relationship type, RGCN defines an independent linear transformation for it, and then combines the information from different relationship types to update the node representation. The key of RGCN lies in how to effectively update the node representations in the multi-relational graph. Assume that the set of nodes in the graph is v, the set of edges is ε, and the set of relationships is R. For the representation of node v At layer l, RGCN updates by aggregating information from neighboring nodes. The specific update formula is:

[0067]

[0068] where represents the set of neighbor nodes connected to node v through relation r; is the transformation matrix of relation r at the l-th layer; the eigenvector of node u in the l-th layer; is the transformation matrix of the node's own information; c v,r is the normalization constant, used to adjust the contribution of different relation types to the node representation update; σ(·) is the non-linear activation function, such as ReLU.

[0069] In RGCN, the contributions of different types of relations to the node representation update may vary. By applying different weights to each relation type, RGCN can flexibly aggregate information from different relations, and these weights are passed through the transformation matrix W r , achieving:

[0070]

[0071] This formula shows how to perform weighted summation on information from different relation types and perform non-linear transformation on the node representation through the activation function.

[0072] Since the number of parameters in the RGCN model increases with the increase in the number of relation types, it may lead to overfitting, especially when the training data is scarce. To solve this problem, RGCN introduces edge type-based regularization by decomposing the transformation matrix W r into the product of low-rank matrices to reduce the model parameters:

[0073] W r = W base + A r B r

[0074] where W base is the shared basic transformation matrix, A r and B r are matrices specific to relation r.

[0075] 2. Text data processing branch

[0076] The BERT model receives the text description of the node, processes the text through its multi-layer Transformer encoder, generates the semantic eigenvector corresponding to the node, and supplements the structural features generated by the RGCN model.

[0077] BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model that significantly improves the performance of natural language processing (NLP) tasks through bidirectional encoders and deep learning techniques. At the core of BERT is the Transformer architecture, which is pre-trained on a large amount of unsupervised text data to generate context-sensitive word vector representations, thus performing well in downstream tasks.

[0078] The architecture of BERT is based on the encoder part of the Transformer, which is a deep neural network structure entirely based on the attention mechanism. The input to the BERT model is a tokenized text sequence, which is converted into word vectors and added to the position encoding to retain the position information in the sequence. The basic units of the Transformer are the multi-head self-attention mechanism and the feed-forward neural network. BERT constructs a deep model by stacking multiple such basic units (called "layers"). In the standard version of BERT, 12 encoder layers are used, with each layer containing 768 hidden units and the total number of parameters exceeding 100 million.

[0079] The core mathematical formulas in BERT mainly include the multi-head self-attention mechanism, the feed-forward neural network, and the position encoding. The self-attention mechanism is used to calculate the dependencies between each word in the input sequence and other words. Given the word vector representation of the input sequence:

[0080] X = [x1, x2, x3,..., x n

[0081] The calculation of self-attention is as follows:

[0082] Query, Key, Value are represented as:

[0083] Q = XW Q

[0084] K = XW K

[0085] V = XW V

[0086] where W Q 、W K 、W V are trainable weight matrices. The calculation formula for the attention scores is as follows:

[0087]

[0088] ​where dk is the dimension of K. The multi-head attention calculates the attention through multiple different heads and concatenates the results and then maps them back to the original dimension:

[0089] MultiHead(Q, K, V) = Concat(head1, …, head h )W O

[0090] After the self-attention mechanism, the input will be processed by a feed-forward neural network. The form of the feed-forward network for each layer is:

[0091] FFN(x) = ReLU(xW1 + b1)W2 + b2

[0092] where W1, W2, b1, and b2 are trainable parameters.

[0093] 3. Vector concatenation

[0094] Perform a concatenation (concat) operation on the structural feature vectors of the RGCN model and the semantic feature vectors from the BERT model. In this way, the model can retain both the structural information and semantic information of the nodes, thereby generating a comprehensive node representation vector. It can not only reflect the structural position of the nodes in the knowledge graph but also embody the content and semantics of the nodes. For example, if the vector of the microbial node is m1 and the vector of the nanomaterial node is m2, then the scoring formula for this microbial and nanomaterial node is as follows:

[0095]

[0096] For a given nanomaterial, use the above scoring formula to score each microorganism and then sort the scores. During the training process, the present invention uses MRR (Mean Reciprocal Rank) as the evaluation metric. MRR is used to measure the effectiveness of an information retrieval system or model, especially focusing on the quality of the first relevant result returned. The value of MRR ranges between 0 and 1, and the larger the value, the better the model ranks the relevant results in the front. Specifically, MRR measures the overall performance of the system for all queries by calculating the reciprocal of the rank of the first relevant result for each query and then taking the average of the results for all queries. Its calculation formula is as follows:

[0097]

[0098] where |Q| represents the total number of query sets. rank i represents the ranking position of the first relevant result for the i-th query.

[0099] V. Examples of achievements

[0100] Examples are given using publicly available partial invention results. For a given nanomaterial, the present invention predicts candidate microbial entities and obtains the conclusion that "the microorganism Shewanella oneidensis MR-1 can synthesize the nanomaterial AuPdPt", and the conclusion has been verified by experts in the field.

[0101] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Those of ordinary skill in the art can modify or equivalently replace the technical solutions of the present invention without departing from the spirit and scope of the present invention. The protection scope of the present invention shall be subject to what is described in the claims.

Claims

1. A method for recommending microbial synthesis of nanomaterials based on a knowledge graph, comprising: Obtaining microbial information and nanomaterial information; Constructing a knowledge graph through the microbial information and the nanomaterial information; wherein, the nodes of the knowledge graph are composed of microorganisms, nanomaterials, synthesis methods, and elements; The knowledge graph obtains structural feature vectors and semantic feature vectors through a BERT-RGCN model, where the BERT-RGCN model includes a graph structure processing branch and a text data processing branch; Extracting the features of the graph structure through the graph structure processing branch, including: Inputting a multi-relational graph obtained based on the knowledge graph; different types of nodes in the multi-relational graph are connected by multiple relationship types; Updating the node representations in the multi-relational graph based on the relationship types; the update includes performing multi-layer convolutional operations with the nodes and their connection relationships as inputs; Outputting the structural feature vectors of the nodes based on the updated node representations; Capturing the semantic information of the text data through the text data processing branch, including: Inputting a text sequence obtained based on the knowledge graph; Converting the text sequence into word vectors; The word vectors output the semantic feature vectors of the nodes through an encoder; Concatenating the structural feature vectors and the semantic feature vectors to obtain the representation vectors of the respective nodes; Scoring the respective nodes based on the representation vectors, and judging the relationship between the microorganisms and the nanomaterials according to the scoring results.

2. The method according to claim 1, characterized in that, In the ontology of the knowledge graph, the entity labels include: organisms, habitats, mechanisms, functions, metal resistance, nanomaterials, synthesis methods, precursors, pH values, temperatures, rotation speeds, positions, sizes, shapes, applications, antibacterial agents, electrons, and energy.

3. The method according to claim 1, wherein Using the mean reciprocal rank as the evaluation metric for the method.

4. An electronic device, comprising a memory and a processor, the memory storing a computer program, the computer program being configured to be executed by the processor, the computer program including instructions for executing the method according to any one of claims 1 to 3.

5. A storage medium storing a computer program, which when executed by a computer, implements the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Static knowledge reasoning method and system fusing semantic and structural information

    CN116306939A

  • Knowledge graph processing

    WO2023071845A1