Methods, devices, equipment and products for constructing multidimensional knowledge graphs

By acquiring multiple biomedical databases, standardizing and reconstructing entity relationships, and constructing a multidimensional knowledge graph, the problem of the single structure of knowledge graphs in existing technologies is solved, and a comprehensive reconstruction and in-depth analysis of biomedical knowledge is achieved.

CN116992037BActive Publication Date: 2025-12-02TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211216862.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-30
Publication Date
2025-12-02
Estimated Expiration
2042-09-30

AI Technical Summary

Technical Problem

Most existing biomedical knowledge graphs are centered on diseases or drugs, have a simple structure, and mainly target a certain disease or a few biomedical entities, making it difficult to comprehensively integrate and display complex biomedical knowledge.

Method used

By acquiring multiple biomedical databases, standardizing similar entities, and reconstructing entity relationships, a multidimensional knowledge graph with gene-related entities as the main component is constructed, containing entity relationships between gene-related and non-gene-related entities, and using artificial intelligence models to predict unknown relationships.

Benefits of technology

It has achieved a comprehensive reconstruction of biomedical knowledge at the molecular and cellular level, enhanced the ability to analyze the correlation between genes and diseases, drugs, etc., and provided a deeper and broader display of biomedical knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116992037B_ABST
    Figure CN116992037B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and product for constructing a multidimensional knowledge graph, applicable to the field of data processing technology. The method includes: acquiring at least two biomedical databases, wherein the biomedical databases store different entities and entity relationships connecting the different entities; standardizing similar entities in the at least two biomedical databases to obtain at least two standardized entities; reconstructing entity relationships between different standardized entities based on the entity relationships between different entities in the at least two biomedical databases; and constructing the multidimensional knowledge graph based on the at least two standardized entities and the entity relationships between different standardized entities. This method can construct a multidimensional knowledge graph, primarily based on gene-related entities, by integrating databases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, device and product for constructing a multidimensional knowledge graph. Background Technology

[0002] In the biomedical field, knowledge is vast and fragmented, mostly scattered across unstructured academic literature, textbooks, non-standardized data repositories, and clinical guidelines. Knowledge graphs can integrate and summarize this existing knowledge and structure.

[0003] Based on expertise in relevant topics within the biomedical field, the related technologies have developed knowledge graphs for multiple specific domains.

[0004] Most of the aforementioned knowledge graphs are constructed around diseases or drugs, have relatively simple structures, and are mostly targeted at a specific disease or several biomedical entities. Summary of the Invention

[0005] This application provides a method, apparatus, device, and product for constructing a multidimensional knowledge graph, capable of generating a multidimensional knowledge graph primarily composed of gene-related entities. The technical solution is as follows:

[0006] According to one aspect of this application, a method for constructing a multidimensional knowledge graph is provided, the method comprising:

[0007] Obtain at least two biomedical databases, wherein the biomedical databases store different entities and entity relationships connecting the different entities;

[0008] Standardize the similar entities in the at least two biomedical databases to obtain at least two standardized entities;

[0009] Based on the entity relationships between different entities in the at least two biomedical databases, reconstruct the entity relationships between the different standardized entities;

[0010] Based on the at least two standardized entities and the entity relationships between different standardized entities, the multidimensional knowledge graph is constructed, and the multidimensional knowledge graph includes entity relationships between gene-type entities and non-gene-type entities;

[0011] The genetic entities include genetic entities and genetic variation entities, and the non-genetic entities include at least one of disease entities, drug entities, pathway entities, and anatomical entities.

[0012] According to another aspect of this application, a search method for a multidimensional knowledge graph is provided, the method comprising:

[0013] Receive entity search terms;

[0014] Based on the multidimensional knowledge graph query, at least one result triple is associated with the entity search term. The multidimensional knowledge graph includes entity relationships between gene-type entities and non-gene-type entities. The gene-type entities include gene entities and gene variation entities. The non-gene-type entities include at least one of disease entities, drug entities, pathway entities, and anatomical entities.

[0015] The search results page displays at least one result triple associated with the entity keyword.

[0016] According to another aspect of this application, an apparatus for constructing a multidimensional knowledge graph is provided, the apparatus comprising:

[0017] The acquisition module is used to acquire at least two biomedical databases, wherein the biomedical databases store different entities and entity relationships connecting the different entities;

[0018] A standardization module is used to standardize similar entities in the at least two biomedical databases to obtain at least two standardized entities.

[0019] The reconstruction module is used to reconstruct the entity relationships between different standardized entities based on the entity relationships between different entities in the at least two biomedical databases.

[0020] A construction module is used to construct the multidimensional knowledge graph based on the at least two standardized entities and the entity relationships between different standardized entities. The multidimensional knowledge graph includes entity relationships between gene-type entities and non-gene-type entities.

[0021] The genetic entities include genetic entities and genetic variation entities, and the non-genetic entities include at least one of disease entities, drug entities, pathway entities, and anatomical entities.

[0022] According to another aspect of this application, a search apparatus for a multidimensional knowledge graph is provided, the apparatus comprising:

[0023] The receiving module is used to receive entity search terms;

[0024] The query module is used to query at least one result triple associated with the entity search term based on the multidimensional knowledge graph. The multidimensional knowledge graph includes entity relationships between gene-type entities and non-gene-type entities. The gene-type entities include gene entities and gene variation entities, and the non-gene-type entities include at least one of disease entities, drug entities, pathway entities, and anatomical entities.

[0025] The display module is used to display at least one result triple associated with the entity keyword in the search results interface.

[0026] According to another aspect of this application, a computer device is provided, the computer device comprising: a processor and a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to implement the method for constructing a multidimensional knowledge graph as described above and the method for searching a multidimensional knowledge graph as described above.

[0027] According to another aspect of this application, a computer-readable storage medium is provided, wherein at least one computer program is stored therein, the at least one computer program being loaded and executed by the processor to implement the method for constructing a multidimensional knowledge graph as described above and the method for searching a multidimensional knowledge graph as described above.

[0028] According to another aspect of this application, a computer program product is provided, the computer program product comprising a computer program stored in a computer-readable storage medium; the computer program is read from the computer-readable storage medium and executed by a processor of a computer device to implement the method for constructing a multidimensional knowledge graph as described above and the method for searching a multidimensional knowledge graph as described above.

[0029] The beneficial effects of the technical solutions provided in this application include at least the following:

[0030] By acquiring at least two biomedical databases and standardizing similar entities in these databases, standardized entities are obtained. Simultaneously, based on the entity relationships between different entities in the at least two biomedical databases, the entity relationships between standardized entities are reconstructed. Based on the different standardized entities and the relationships between the reconstructed standardized entities, a multidimensional knowledge graph is constructed. This multidimensional knowledge graph primarily reconstructs knowledge entities in the biomedical field from the perspective of genes and biological cells and molecules, focusing on gene variations and their impact on diseases. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This is a structural block diagram of a computer system provided in an exemplary embodiment of this application;

[0033] Figure 2 This is a flowchart of a multidimensional knowledge graph construction method provided in an exemplary embodiment of this application;

[0034] Figure 3 This is a flowchart of a multidimensional knowledge graph construction method provided in an exemplary embodiment of this application;

[0035] Figure 4 This is a block diagram of a multidimensional knowledge graph construction method provided in an exemplary embodiment of this application;

[0036] Figure 5 This is a flowchart of a method for constructing a multidimensional knowledge graph provided in an exemplary embodiment of this application;

[0037] Figure 6 This is a structural block diagram of a multidimensional knowledge graph provided in an exemplary embodiment of this application;

[0038] Figure 7 This is a flowchart of a multidimensional knowledge graph search method provided in an exemplary embodiment of this application;

[0039] Figure 8 This is a schematic diagram of the search results interface of a multidimensional knowledge graph provided in an exemplary embodiment of this application;

[0040] Figure 9 This is a flowchart of the triggering operation of a multidimensional knowledge graph provided in another exemplary embodiment of this application;

[0041] Figure 10 This is a schematic diagram of the search results interface of a multidimensional knowledge graph provided in an exemplary embodiment of this application;

[0042] Figure 11 This is a block diagram of a multidimensional knowledge graph construction apparatus provided in an exemplary embodiment of this application;

[0043] Figure 12 This is a block diagram of a multidimensional knowledge graph search device provided in an exemplary embodiment of this application;

[0044] Figure 13 This is a structural block diagram of a computer device provided in an exemplary embodiment of this application. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0046] First, let's introduce the terms used in the embodiments of this application:

[0047] A database can be viewed as an electronic file storage cabinet, where multiple files or data are stored together electronically. Users can perform operations such as adding, extracting, updating, and deleting files or data.

[0048] A biomedical database is a database primarily used to extract literature and data based on relationships within the biomedical field. The data sources, types, and other attributes within this database all belong to the biomedical field. This biomedical database can be a bibliographic database or a web database. The following example illustrates this using a web database as an example.

[0049] Knowledge graph: A knowledge graph is a data graph used to accumulate and communicate knowledge about the real world. This data graph consists of multiple nodes and edges connecting each node; nodes represent entities, and edges represent the relationships between these entities.

[0050] Figure 1 This is a structural block diagram of a computer system 100 provided in an exemplary embodiment of this application. The computer system 100 includes: a terminal 120 and a server 140.

[0051] Terminal 120 has an application (or client) installed and running that supports multidimensional knowledge graph search. Terminal 120 is the terminal used by the user to perform keyword searches.

[0052] Typically, terminal 120 is connected to server 140 via a wireless or wired network.

[0053] Server 140 includes at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. For example, server 140 includes a processor 144 and a memory 142. Memory 142 further includes a receiving module 1421, a control module 1422, and a sending module 1423. The receiving module 1421 receives requests sent by clients, such as receiving keyword information to be queried. The control module 1422 controls the display of the network database. The sending module 1423 sends responses to clients, such as sending information about the search keywords found. Server 140 provides background services for applications supporting multidimensional knowledge graph search. Optionally, server 140 performs the primary computational work, and terminal 120 performs secondary computational work; or, server 140 performs secondary computational work, and terminal 120 performs primary computational work; or, server 140 and terminal 120 collaborate using a distributed computing architecture.

[0054] The device type of terminal 120 includes at least one of the following: smartphone, smartwatch, smart TV, tablet computer, e-book reader, MP3 player, MP4 player, laptop computer, and desktop computer. The following embodiments use a smartphone as an example.

[0055] Those skilled in the art will understand that the number of terminals described above can be more or less. For example, there may be only one terminal, or there may be dozens or hundreds of terminals, or even more. This application does not limit the number of terminals or the type of device.

[0056] Figure 2 This is a flowchart illustrating a method for constructing a multidimensional knowledge graph according to an exemplary embodiment of this application. This method can be implemented by, for example... Figure 1 The method is executed by the terminal or an application on the terminal, or by a client on the terminal. The method includes:

[0057] Step 210: Obtain at least two biomedical databases, which store different entities and entity relationships connecting the different entities;

[0058] For example, the at least two biomedical databases differ based on the different attributes of the information data within them. Different biomedical databases store different entities, and these entities have different attributes or categories. At least one of the at least two biomedical databases contains entities that are genetic entities.

[0059] For example, the at least two biomedical databases may include several types, including gene / protein, gene variant, drug, disease, pathway, anatomy, gene / gene variant-drug, gene / gene variant-disease, and gene-anatomy. Specifically, gene / protein databases include Ensembl and HGNC; gene variant databases include dbSNP; drug databases include PubChem; disease databases include MONDO, Experimental Factor Ontology (EFO), PATO, and Human Phenotype Ontology (HPO); pathway databases include R-HAS; anatomy databases include BRENDA Tissue Ontology (BTO); gene / gene variant-drug databases include Pharmacogenetics and Pharmacogenomics Knowledge Base (PharmGKB), SIGNOR, and Drug-Gene Interaction database (DGIdb); gene / gene variant-disease databases include GWAS Catalog, ClinVar, ClinGen, and DISEASE; and gene-anatomy databases include TISSUES.

[0060] For example, a biomedical database stores information from different biomedical fields. This information from different biomedical fields constructs different types of knowledge graphs, which define different entities, concepts, and relationships connecting these entities.

[0061] For example, the information quality varies among different biomedical databases. To ensure the authority and accuracy of the selected and used biomedical databases, eligibility criteria for data screening have been established. These criteria include: 1. The selected biomedical database should be open and downloadable; 2. The selected biomedical database should be of good quality, assessed by the authority of the database maintainer and the popularity of the literature; 3. The selected biomedical database should be regularly updated, with transparent data management and processing methods, and should be continuously adjusted as individual data resources develop and new data emerges.

[0062] Step 220: Standardize the same type of entities in at least two biomedical databases to obtain at least two standardized entities;

[0063] For example, the entity types or concepts in at least two biomedical databases may be the same or different. Entities of the same type in a database are entities of the same or similar type obtained from at least two biomedical databases. Entities of the same type may have different names, definitions, or categories in different biomedical databases to meet different functional requirements, but their attributes should be the same.

[0064] For example, at least two biomedical databases are used to standardize similar entities, unifying those with different names, definitions, or categories. This involves unifying the names, definitions, and categories of similar entities, defining each entity according to a unified standard, and selecting a common ontology for each entity type, resulting in at least two standardized entities.

[0065] For example, at least two standardized entities include gene-type entities and non-gene-type entities. The distinction between gene-type and non-gene-type entities can be mainly determined by the attributes of the entities; gene-type entities include entities with gene attributes, and non-gene-type entities include entities other than gene-type entities.

[0066] For example, gene entities and protein entities are standardized as gene entities; variant entities are standardized as gene variant entities; disease entities, phenotype entities, trait entities, variable entities, and rare disease entities are standardized as disease entities; drug entities and small chemical molecule entities are standardized as drug entities; biological process entities and molecular pathway entities are standardized as pathway entities; and anatomical tissue entities and cell type entities are standardized as anatomical entities.

[0067] For example, gene-related entities include at least one of gene entities and gene variation entities, while non-gene-related entities include at least one of disease entities, drug entities, pathway entities, and anatomical entities.

[0068] For example, different databases can use different ontology for each entity. For instance, database A uses ontology X to label diseases, while database B uses a cross-ontology mapping of ontology Y to label diseases. The standardization process involves linking different ontology of the same entity in different databases to align the ontology, thereby producing a unified dataset in a standardized format.

[0069] Step 230: Based on the entity relationships between different entities in at least two biomedical databases, reconstruct the entity relationships between different standardized entities;

[0070] Based on the at least two standardized entities obtained in step 220, it is possible to further obtain gene-type entities and entity relationships between gene-type entities, and / or entity relationships between gene-type entities and non-gene-type entities.

[0071] For example, based on the entity relationships between different entities in at least two biomedical databases, these relationships are mapped to different standardized entities. These standardized entities are then reconstructed to obtain new one-to-one entity relationships. For instance, the entity relationship in the biomedical database that corresponds to the gene entity and the protein entity corresponds to the entity relationship between gene entities in the standardized entity database.

[0072] For example, reconstructing the entity relationships between different standardized entities requires mapping the standardized entities to their respective ontologies, filtering out missing entries after ontology mapping, and removing duplicates.

[0073] Step 240: Construct a multidimensional knowledge graph based on at least two standardized entities and the entity relationships between different standardized entities.

[0074] Based on the standardized entity relationships between at least two standardized entities obtained in step 220 and different standardized entities obtained in step 230, at least one set of triples can be generated.

[0075] For example, at least one set of triples includes at least two normalized entities and a normalized entity relationship between the two normalized entities. A set of triples is generated based on the two normalized entities and the normalized entity relationship.

[0076] For example, a multidimensional knowledge graph is constructed based on at least one set of triples.

[0077] For example, at least one set of triplets includes at least one of the following types: gene variant-disease, gene variant-drug, gene variant-gene, gene-disease, gene-drug, gene-gene, gene-pathway, and gene-tissue.

[0078] For example, this multidimensional knowledge graph is a biomedical knowledge graph at the molecular level, which mainly includes gene entities, gene mutation entities, disease entities, drug entities, etc. This multidimensional knowledge graph is constructed at the molecular and cellular level, connecting genes with genes and gene mutations, and to a certain extent highlighting the correlation between genes and diseases and drugs. It can analyze and display biomedical pathogenesis and treatment mechanisms through this multidimensional knowledge graph.

[0079] In summary, the method provided in this embodiment acquires at least two biomedical databases, standardizes similar entities within them to obtain at least two standardized entities, and reconstructs the relationships between the standardized entities based on the entity relationships of different entities in the acquired at least two biomedical databases. Based on the at least two standardized entities and their relationships, a multidimensional knowledge graph is constructed. This multidimensional knowledge graph integrates multiple biomedical databases, primarily focusing on gene-related entities, and includes entity relationships between gene-related and non-gene-related entities.

[0080] Figure 3 This is a partial flowchart of a method for constructing a multidimensional knowledge graph provided in an exemplary embodiment of this application. According to the above... Figure 2 Step 230, based on the entity relationships between different entities in at least two biomedical databases, reconstructs the entity relationships between different standardized entities, which may optionally include the following sub-steps 232 to 238:

[0081] Step 232: Based on any triple from at least two biomedical databases, determine the first normalized entity corresponding to the first entity in the triple and the second normalized entity corresponding to the second entity in the triple;

[0082] For example, at least two selected biomedical databases contain multiple distinct entities and entity relationships between them, which can generate multiple triples. Each triple contains a first entity and a second entity.

[0083] Based on the standardization of different entities, the first entity and the second entity have corresponding first standardized entities and second standardized entities in this multidimensional knowledge graph.

[0084] Step 234: Based on the entity relationship between the first entity and the second entity, determine the standardized entity relationship between the first standardized entity and the second standardized entity;

[0085] For example, for any triple from at least two selected biomedical databases, there is a corresponding entity relationship between the first entity and the second entity. Here, the first entity corresponds to the first normalized entity in the multidimensional knowledge graph, and the second entity corresponds to the second normalized entity in the multidimensional knowledge graph. Therefore, based on the entity relationship between the first entity and the second entity, the normalized entity relationship between the first normalized entity and the second normalized entity can be determined.

[0086] For example, gene-type entities and non-gene-type entities have a certain recursive relationship, such as gene variation > gene > anatomy > pathway > drug > disease. That is, if entity a > entity b, then entity a is the first standardized entity and entity b is the second standardized entity.

[0087] For example, a third entity relationship can be predicted based on the first entity relationship and the second entity relationship.

[0088] In some embodiments, if the first entity relationship is an entity relationship in the gene mutation-disease triplet and the second entity relationship is an entity relationship in the gene mutation-drug triplet, then the entity relationship between the disease entity and the drug entity can be predicted as the third entity relationship.

[0089] In some embodiments, if the first entity relationship is an entity relationship in a gene variation-pathway triplet and the second entity relationship is an entity relationship in a gene variation-disease triplet, then the entity relationship between the pathway entity and the disease entity can be predicted as the third entity relationship.

[0090] For example, the first entity relationship and the second entity relationship are input into the first artificial intelligence (AI) model to obtain the third entity relationship predicted by the first AI model.

[0091] For example, the first AI model is trained using a sample training set, which includes at least one pair of sample sets. Each pair of sample sets includes a first entity relationship and a second entity relationship. The first entity relationship and the second entity relationship are input into the first AI model, and the first AI model predicts a third entity relationship. The predicted third entity relationship is compared with the third entity relationship used as a label, and the error loss is calculated. The model parameters in the first AI model are trained based on the error loss using an error backpropagation algorithm.

[0092] Feature extraction is performed on the first entity relationship of the sample to obtain the corresponding first feature representation; feature extraction is performed on the second entity relationship of the sample to obtain the corresponding second feature representation. The first feature representation and the second feature representation are input into the first AI model.

[0093] After multiple training sessions, such as training with 10,000 samples, or when the model converges, a first AI model is obtained that can predict the third entity relationship based on the first entity relationship and the second entity relationship.

[0094] In some embodiments, the gene mutation-disease entity relationship is selected as the first entity relationship of the sample, and the gene mutation-drug entity relationship is selected as the second entity relationship of the sample.

[0095] Feature extraction is performed on DNA sequences, protein structures, chromosomes, alleles, etc., in gene variant entities to obtain gene variant feature representations.

[0096] Feature extraction is performed on disease symptoms, phenotypes, variables, and rare pathological structures in disease entities to obtain disease feature representations.

[0097] The pharmacological structure and small chemical molecules in the drug entity are extracted to obtain the drug feature representation.

[0098] The first entity relationship in the sample includes features related to diseases caused by gene mutations, such as a change in a DNA sequence leading to a disease symptom. The second entity relationship includes features related to the association between gene mutations and drug molecular structures, such as a change in a DNA sequence being related to the molecular structure of a drug. The first AI model, after receiving the sample input, predicts the third entity relationship, which includes the association between disease symptoms and drug molecular structures predicted based on the first and second entity relationships. The predicted third entity relationship is compared with the labeled third entity relationship, and the error loss is calculated. The model parameters are then trained using an error backpropagation algorithm based on this error loss, ultimately resulting in a first AI model capable of predicting the third entity relationship based on the first and second entity relationships.

[0099] For example, the fourth and fifth entity relationships are input into the second AI model to obtain the sixth entity relationship predicted by the second AI model.

[0100] For example, the second AI model is trained using a sample training set, which includes at least one pair of sample sets. Each pair of sample sets includes a fourth entity relation and a fifth entity relation. The fourth and fifth entity relations are input into the second AI model, which predicts the sixth entity relation. The predicted sixth entity relation is compared with the sixth entity relation used as a label, and the error loss is calculated. The model parameters in the second AI model are trained based on the error loss using an error backpropagation algorithm.

[0101] Feature extraction is performed on the fourth entity relation of the sample to obtain the corresponding fourth feature representation; feature extraction is performed on the fifth entity relation of the sample to obtain the corresponding fifth feature representation. The fourth feature representation and the fifth feature representation are then input into the second AI model.

[0102] After multiple training sessions, such as training with 10,000 samples, or when the model converges, a second AI model is obtained that can predict the sixth entity relationship based on the fourth and fifth entity relationships.

[0103] In some embodiments, the gene mutation-anatomy entity relationship is selected as the fourth entity relationship of the sample, and the gene mutation-drug entity relationship is selected as the fifth entity relationship of the sample.

[0104] Feature extraction is performed on DNA sequences, protein structures, chromosomes, alleles, etc., in gene variant entities to obtain gene variant feature representations.

[0105] Features such as anatomical tissues and cell types in anatomical entities are extracted to obtain anatomical feature representations.

[0106] The pharmacological structure and small chemical molecules in the drug entity are extracted to obtain the drug feature representation.

[0107] The fourth entity relationship in the sample includes features of anatomical tissue changes caused by gene mutations, such as a change in a DNA sequence leading to a change in a cell type. The fifth entity relationship in the sample includes features of the correlation between gene mutations and drug molecular structures, such as a change in a DNA sequence enabling treatment through the molecular structure of a drug. The second AI model, after receiving the sample input, predicts the sixth entity relationship, which includes the correlation between cell types and drug molecular structures predicted based on the fourth and fifth entity relationships. The predicted sixth entity relationship is compared with the sixth entity relationship used as a label, and the error loss is calculated. The model parameters in the second AI model are then trained using an error backpropagation algorithm based on the error loss, ultimately resulting in a second AI model capable of predicting the sixth entity relationship based on the fourth and fifth entity relationships.

[0108] For example, when the number of training samples for the AI ​​model is large enough or the model converges, inputting a gene-related entity can predict a non-gene-related entity, indicating a correlation between the gene-related and non-gene-related entities. For instance, inputting a gene-related entity or a gene variant entity can predict a disease-related entity, drug-related entity, or pathway-related entity.

[0109] For example, when the number of training samples for the AI ​​model is large enough or the model converges, inputting a non-gene entity can predict a gene entity, and there is a correlation between gene and non-gene entities. For instance, inputting a disease entity, drug entity, or pathway entity can predict related genes or gene variant entities.

[0110] While the AI ​​model predictions mentioned above have some errors, they can, to a certain extent, help researchers in the biomedical field to conduct further studies on gene molecular structures.

[0111] Step 236: Based on the entity type, determine the start entity and the end entity from the first and second standardized entities;

[0112] For example, such as Figure 4As shown, for any triple in at least two selected biomedical databases, if the corresponding entity relationship between the first entity 410 and the second entity 420 starts from the first entity 410 and ends at the second entity 420, then the corresponding first normalized entity 430 and the second normalized entity 440 also start from the first normalized entity 430 and end at the second normalized entity 440; or, if the corresponding entity relationship between the first entity 410 and the second entity 420 starts from the second entity 420 and ends at the first entity 410, then the corresponding first normalized entity 430 and the second normalized entity 440 also start from the second normalized entity 440 and end at the first normalized entity 430.

[0113] For example, if both the first standardized entity and the second standardized entity are gene-related entities, then the first standardized entity serves as the starting entity and the second standardized entity as the ending entity; or, the second standardized entity serves as the starting entity and the first standardized entity as the ending entity. If the first standardized entity is a gene-related entity and the second standardized entity is a non-gene-related entity, then the first standardized entity serves as the starting entity and the second standardized entity as the ending entity. If the first standardized entity is a non-gene-related entity and the second standardized entity is a gene-related entity, then the second standardized entity serves as the starting entity and the first standardized entity as the ending entity.

[0114] For example, this multidimensional knowledge graph is mainly composed of gene-related entities, so gene-related entities can be used as start entities or end entities.

[0115] Step 238: Based on the starting entity, ending entity, and standardized entity relationship, generate entity relationships between standardized entities that have an association relationship.

[0116] For example, based on the start entity and end entity determined in step 236 above, the directional relationship between standardized entities is determined. Generally speaking, the start entity and end entity are related, or they are causally related, with the start entity being the cause and the end entity being the result. For instance, the standardized entity gene entity and the standardized entity disease entity are determined because pathological changes in the human system are usually caused by changes in the genome, that is, changes in genes lead to the occurrence of diseases. Therefore, the gene entity is usually used as the start entity and the disease entity as the end entity.

[0117] In summary, the method provided in this embodiment can reconstruct entity relationships between different standardized entities based on entity relationships between different entities in at least two biomedical databases. Using any triple from at least two biomedical databases, the method determines the first standardized entity corresponding to the first entity in the triple and the second standardized entity corresponding to the second entity in the triple; then, based on the entity relationship between the first and second entities, it determines the standardized entity relationship between the first and second standardized entities.

[0118] Figure 5 This is a flowchart of a multidimensional knowledge graph construction method provided in an exemplary embodiment of this application. According to the above... Figure 2 Step 240 in the method involves constructing a multidimensional knowledge graph based on at least two standardized entities and the entity relationships between different standardized entities. This method includes:

[0119] Step 242: Generate at least one set of triples based on at least two normalized entities and the entity relationships between different normalized entities;

[0120] For example, generating a set of triples includes at least two normalized entities and the normalized entity relationship between those two normalized entities, such as... Figure 6 As shown, the multiple triplets with gene entities as the main component include: a triplet starting from gene entity 601 and ending with drug entity 604; a triplet starting from gene entity 601 and ending with pathway entity 605; a triplet starting from gene entity 601 and ending with anatomical entity 606; a triplet starting from gene entity 601 and ending with gene entity 601; a triplet starting from gene entity 601 and ending with disease entity 603; a triplet starting from gene mutation entity 602 and ending with drug entity 604; a triplet starting from gene mutation entity 602 and ending with gene entity 601; and a triplet starting from gene mutation entity 602 and ending with disease entity 603.

[0121] Each triple category contains a corresponding standardized entity relation, and each relation text box records the corresponding entity relation within the triple category. The relation text box is... Figure 6 The text box is drawn on the arrow line.

[0122] For example, the entity relationships in the gene-drug triplet mainly include data source, cell, and tissue; the entity relationships in the gene-pathway triplet mainly include data source; the entity relationships in the gene-anatomy triplet mainly include at least one of data source and confidence level; the entity relationships in the gene-gene triplet mainly include at least one of mechanism, cell, tissue, and data source; the entity relationships in the gene-disease triplet mainly include at least one of inheritance, confidence level, number of case reports, and data source; the entity relationships in the gene-variation-drug triplet mainly include at least one of number of case reports and data source; the entity relationships in the gene-variation-gene triplet mainly include at least one of variation, amino acid, and data source; and the entity relationships in the gene-variation-disease triplet mainly include at least one of clinical, inheritance pattern, number of case reports, and data source.

[0123] Step 244: Construct a multidimensional knowledge graph based on at least one set of triples;

[0124] For example, at least one set of triples can be selected to construct a multidimensional knowledge graph. Any set of triples includes at least one of the following: gene-drug, gene-pathway, gene-anatomy, gene-gene, gene-disease, gene mutation-drug, gene mutation-gene, and gene mutation-disease. For example, assuming the gene-drug triple is selected, this triple includes at least a gene entity, a drug entity, and the entity relationship between the gene entity and the drug entity. Based on the gene entity, the drug entity, and the entity relationship between the gene entity and the drug entity, a multidimensional knowledge graph can be constructed.

[0125] For example, this multidimensional knowledge graph is constructed based on at least one set of triples, which includes at least one of gene entity or gene mutation entity. This multidimensional knowledge graph focuses on gene-gene and gene-gene mutation, describes the association between genes and diseases, genes and drugs, etc., and analyzes and displays the mechanisms of biomedical pathogenesis and treatment at the molecular and cellular level. It belongs to the molecular level of biomedical knowledge graph.

[0126] Step 246: Generate node type files corresponding to the entity types of the normalized entities in the triples;

[0127] For example, the standardized entities generated in the above steps can be regarded as entity categories. For instance, gene entities can be regarded as a category, which includes a series of gene entities, protein entities, etc. related to genes; gene variation entities are a category, which includes a series of variation entities related to variation; drug entities are a category, which includes a series of drug entities, small chemical molecule entities, etc. related to drugs; anatomy entities are a category, which includes some anatomical tissue entities, cell type entities, etc. related to anatomy; pathway entities are a category, which includes a series of biological process entities, analytical pathway entities, etc. related to pathways; and disease entities are a category, which includes a series of disease entities, phenotype and trait entities, experimental variable entities, rare disease entities, etc. related to diseases.

[0128] For example, each standardized entity represents an entity type and can be regarded as a node. A node type file corresponding to the entity type of the standardized entity is generated. Each node type file corresponds to a standardized entity category. Each node type file includes at least one of the following: a unique identifier for the corresponding standardized entity, the name of the standardized entity, and the type of the standardized entity.

[0129] Step 248: Generate edge annotation files corresponding to the types in the triples.

[0130] For example, each triple includes a start entity, an end entity, and a one-to-one normalized entity relationship between the start and end entities. For instance, in a triple starting with a gene entity and ending with a disease entity, there is a corresponding entity relationship between the gene and disease entities. Different triples can be considered as different types based on their different start entities, end entities, and normalized entity relationships, generating edge annotation files corresponding to the types within the triples. Each triple corresponds to one edge annotation file. The edge annotation file includes the source of the normalized entities, the relationships between the normalized entities in the triple, and the relational subtypes of the triple's type.

[0131] For example, the relation subtypes include at least one of pharmacokinetics, pharmacokinetics, inhibitors, allosteric agents, and agonists. Relation subtypes may also include efficacy, dose, toxicity, association, pathogenesis, activation, regulation, antagonism, blockade, positive regulation, inverse activation, partial activation, vaccine, negative regulation, induction, allosteric regulation, inhibitory allosteric regulation, upregulation, downregulation, downregulation by destabilization, upregulation of activity, downregulation of activity, upregulation of quantity by expression, downregulation of quantity by inhibition, low expression - longer survival, high expression - shorter survival, high expression - longer survival, low expression - shorter survival, etc.

[0132] In summary, the method provided in this embodiment constructs a multidimensional knowledge graph based on at least two standardized entities and the entity relationships between different standardized entities. The construction of the multidimensional knowledge graph is achieved through at least one set of triples, each triple including at least one start entity, one end entity, and the relationship between the start and end entities. Based on the entity relationships between at least two standardized entities and different standardized entities, at least one triple is generated. This triple includes triples connecting gene-type entities and non-gene-type entities, thus reconsidering relationships within the biomedical field from a gene perspective.

[0133] For example, triples, node type files, and edge annotation files are all planar files that can be applied to downstream graph machine learning to predict edge types, node types, etc. For instance, to predict the side effects / toxicity of multiple drugs, based on triples protein-protein, protein-drug, and drug-drug, 645 drugs, 19085 proteins, and 5.39M relationships were found in multiple databases. A graph convolutional network (GCN) can then be used to simulate a tensor decomposition model of drug side effects / toxicity.

[0134] For example, multi-relationship link prediction methods are also used in multidimensional knowledge graphs to predict interactions between entities. For instance, it can predict whether atorvastatin (a lipid-lowering drug) and amlodipine (a blood pressure-lowering drug) may cause muscle inflammation based on the interaction between the two drugs, combined with recent clinical reports showing that muscle tissue may be damaged due to the drugs.

[0135] In summary, the method provided in this embodiment constructs a multidimensional knowledge graph based on standardized entities and the relationships between them. This not only integrates and describes the relationships between biomedical entities at a deeper level, but also uses multidimensional patterns to link diseases with drugs, genes / proteins, gene variations, pathways, cellular components / tissues, etc., thereby enhancing the ability of the gene graph to characterize diseases to a certain extent.

[0136] The method provided in this embodiment is expected to be further expanded for computing node types, novel edge prediction, etc., with the implementation of machine learning technology. For example, it can identify new drug targets or high-risk groups through edge prediction between genes / gene mutations and disease nodes, and discover biomarkers or predict side effects through edge prediction between genes / gene mutations and drugs, which can help more research and development in the biomedical field to a certain extent.

[0137] Figure 7 This is a flowchart illustrating a multidimensional knowledge graph search method provided in an exemplary embodiment of this application. The method can be derived from, for example... Figure 1The method is executed on a terminal or a client on a terminal in the system shown. The method includes:

[0138] Step 710: Receive entity search terms;

[0139] The multidimensional knowledge graph constructed through the above steps can ultimately be displayed as a visual product. This visual product can be displayed by an application in electronic products such as smartphones, tablets, or smartwatches, or by a web page, or other display products that can be equipped with search functionality and display the knowledge graph. This embodiment uses a web page to display the multidimensional knowledge graph as an example.

[0140] For example, a multidimensional knowledge graph, as an integrated large dataset, has a search function. First, users need to create suitable search keywords according to their needs. The multidimensional knowledge graph needs to receive the search keywords provided by the user and perform data processing and analysis in the background program.

[0141] Optionally, search keywords can be the name of a standardized entity, the category of a standardized entity, the unique identifier of an entity within a standardized entity, or other biomedical information keywords.

[0142] Step 720: At least one result triple associated with the entity search term based on the multidimensional knowledge graph query;

[0143] When the backend database of the multidimensional knowledge graph receives the entity search terms that the user needs to query, the multidimensional knowledge graph query constructed based on the above steps will query at least one result triple that is related to the entity keywords that the user needs to query. The result triple is under the category of triple generated by the standardized entity in the above steps, that is, the result triple is included under the category of standardized entity triple.

[0144] For example, such as Figure 8 As shown, it can be phospholipid metabolism, metabolism, etc. under the pathway entity category, or rs3757970, etc. under the gene mutation entity category, or a specific knowledge entity under other standardized entity categories.

[0145] Step 730: Display at least one result triple associated with the entity keyword in the search results interface.

[0146] After the multidimensional knowledge graph has undergone data reading, data retrieval, and data analysis in the background, the search results interface will display at least one result triple associated with the entity keyword, such as... Figure 8As shown, the entity keyword that the user needs to query is DGAT1, a type of glycerol acyltransferase (DGAT). The results found in the multidimensional knowledge graph are shown in the figure, which include multiple result triplets associated with the DGAT1 gene.

[0147] In some embodiments, the multidimensional knowledge graph displays a customized graph about entity search terms.

[0148] For example, when searching for the keyword "DGAT1", the interface of this multidimensional knowledge graph displays a keyword search area and a results display area. In the keyword search area, the user can enter the keyword they want to search for, for example, they can enter "DGAT1". In the results display area, different entities associated with the keyword can be displayed. For example, the graph displays standardized entity categories such as gene mutation entity, pathway entity, anatomical entity, and disease entity. Under each standardized entity category, multiple specific entities are displayed.

[0149] For example, the results display area of ​​the multidimensional knowledge graph is divided into different colors according to different results: search keywords are displayed in dark red, gene mutation entities are displayed in blue, anatomical entities are displayed in green, pathway entities are displayed in brown, and disease entities are displayed in pink.

[0150] In some embodiments, gene entities and drug entities may also be displayed, and distinguished by other colors.

[0151] Figure 9 This is a flowchart of the triggering operation of a multidimensional knowledge graph provided in another exemplary embodiment of this application;

[0152] Step 910: In response to a triggering operation on the first entity in at least one result triple, display the attribute details of the first entity.

[0153] For example, in response to a triggering operation on any knowledge entity in the multidimensional knowledge graph search results interface, the results interface will display the attribute details of the triggered knowledge entity. These attribute details may include cell type, chromosome type, phenotypic characteristics, metabolic type, identification unit, and evaluation criteria. Specifically, based on the different attributes of different knowledge entities, and according to biological principles, combined with genetic factors, the attributes of different knowledge entities are summarized and displayed.

[0154] For example, clicking on different entities associated with the keyword in the results display area will bring up the corresponding entity's attribute details. For instance... Figure 10As shown, for example, clicking on the gene mutation entity rs863225093 will display the corresponding entity's attribute details in a pop-up window of the search results interface. These details include the mutation site, mutation frequency, mutation location, mutation harmfulness assessment, and the mutation's unique identification unit. The mutation location is represented by chromosome and sequence locus. Mutation harmfulness assessment includes two methods: Polymorphism Phenotyping (PolyPhen), a conservation-based prediction method, and Scale Invariant Feature Transform (SIFT), a rule-based prediction method. The accuracy of these two assessment methods is determined by their numerical values, which range from 0 to 1. For the PolyPhen method, a larger value indicates higher accuracy; for example, 0.999 indicates a harmful mutation. For the SIFT method, a smaller value indicates higher accuracy; for example, 0 indicates a harmful mutation.

[0155] For example, Figure 10 The multidimensional knowledge graph displayed is based on a gene search term and analyzes and displays multiple entity categories, including disease entities, gene mutation entities, anatomical entities, and drug entities. Each entity category also displays multiple specific entity tissues. It can be seen intuitively that this multidimensional knowledge graph studies the mechanisms of biomedical pathogenesis and treatment at the molecular and cellular level, and belongs to the molecular level of biomedical knowledge graph.

[0156] Figure 11 This is a block diagram of a multidimensional knowledge graph construction apparatus provided in an exemplary embodiment of this application.

[0157] The device includes:

[0158] The acquisition module 1110 is used to acquire at least two biomedical databases, which store different entities and entity relationships connecting the different entities;

[0159] Standardization module 1120 is used to standardize similar entities in at least two biomedical databases to obtain at least two standardized entities.

[0160] For example, gene entities and protein entities are standardized as gene entities; variant entities are standardized as gene variant entities; disease entities, phenotype entities, trait entities, variable entities, and rare disease entities are standardized as disease entities; drug entities and small chemical molecule entities are standardized as drug entities; biological process entities and molecular pathway entities are standardized as pathway entities; and anatomical tissue entities and cell type entities are standardized as anatomical entities.

[0161] Reconstruction module 1130 is used to reconstruct entity relationships between different standardized entities based on entity relationships between different entities in at least two biomedical databases;

[0162] Refactoring module 1130 also includes:

[0163] The determination submodule 11301 is used to determine, for any triple in the at least two biomedical databases, a first normalized entity corresponding to the first entity in the triple and a second normalized entity corresponding to the second entity in the triple.

[0164] The determination submodule 11301 is further configured to determine the standardized entity relationship between the first standardized entity and the second standardized entity based on the entity relationship between the first entity and the second entity;

[0165] The determination submodule 11301 is also used to determine the start entity and the end entity in the first standardized entity and the second standardized entity based on the entity type;

[0166] The first generation submodule 11302 is used to generate entity relationships between the standardized entities that have an association relationship based on the starting entity, the ending entity, and the standardized entity relationship.

[0167] Module 1140 is used to construct a multidimensional knowledge graph based on at least two standardized entities and entity relationships between different standardized entities. The multidimensional knowledge graph includes entity relationships between gene-type entities and non-gene-type entities.

[0168] Module 1140 also includes:

[0169] The second generation submodule 11401 is used to generate at least one set of triples based on at least two standardized entities and entity relationships between different standardized entities. The at least one set of triples includes a triple connecting the gene-type entity and the non-gene-type entity.

[0170] Construct submodule 11402, which is used to construct a multidimensional knowledge graph based on at least one set of triples;

[0171] The generation module 1150 is used to generate a node type file corresponding to the entity type of the normalized entity in the triplet. The node type file includes at least one of the unique identifier of the normalized entity, the name of the normalized entity, and the type of the normalized entity.

[0172] The generation module 1150 is further configured to generate an edge annotation file corresponding to the type in the triplet, wherein the edge annotation file includes at least one of the source of the normalized entity, the relationship of the normalized entity quality inspection in the triplet, and the relationship subtype of the type of the triplet;

[0173] For example, the relation subtype includes at least one of pharmacokinetics, pharmacokinetics, inhibitors, allosteric agents, and agonists. Relation subtypes may also include efficacy, dose, toxicity, association, pathogenesis, activation, regulation, antagonism, blockade, positive regulation, inverse activation, partial activation, vaccine, negative regulation, induction, allosteric regulation, inhibitory allosteric regulation, upregulation, downregulation, downregulation by destabilization, upregulation of activity, downregulation of activity, upregulation of quantity by expression, downregulation of quantity by inhibition, low expression - longer survival, high expression - shorter survival, high expression - longer survival, low expression - shorter survival, etc.

[0174] The device also includes:

[0175] The prediction module 1160 is used to input the first entity relationship and the second entity relationship into the first AI model to obtain the third entity relationship predicted by the first AI model.

[0176] For example, the first AI model is trained using a sample training set, which includes at least one pair of sample sets. Each pair of sample sets includes a first entity relationship and a second entity relationship. The first entity relationship and the second entity relationship are input into the first AI model, and the first AI model predicts a third entity relationship. The predicted third entity relationship is compared with the third entity relationship used as a label, and the error loss is calculated. The model parameters in the first AI model are trained based on the error loss using an error backpropagation algorithm.

[0177] Feature extraction is performed on the first entity relationship of the sample to obtain the corresponding first feature representation; feature extraction is performed on the second entity relationship of the sample to obtain the corresponding second feature representation. The first feature representation and the second feature representation are input into the first AI model.

[0178] After multiple training sessions, such as training with 10,000 samples, or when the model converges, a first AI model is obtained that can predict the third entity relationship based on the first entity relationship and the second entity relationship.

[0179] In some embodiments, the gene mutation-disease entity relationship is selected as the first entity relationship of the sample, and the gene mutation-drug entity relationship is selected as the second entity relationship of the sample.

[0180] Feature extraction is performed on DNA sequences, protein structures, chromosomes, alleles, etc., in gene variant entities to obtain gene variant feature representations.

[0181] Feature extraction is performed on disease symptoms, phenotypes, variables, and rare pathological structures in disease entities to obtain disease feature representations.

[0182] The pharmacological structure and small chemical molecules in the drug entity are extracted to obtain the drug feature representation.

[0183] The first entity relationship in the sample includes features related to diseases caused by gene mutations, such as a change in a DNA sequence leading to a disease symptom. The second entity relationship includes features related to the association between gene mutations and drug molecular structures, such as a change in a DNA sequence being related to the molecular structure of a drug. The first AI model, after receiving the sample input, predicts the third entity relationship, which includes the association between disease symptoms and drug molecular structures predicted based on the first and second entity relationships. The predicted third entity relationship is compared with the labeled third entity relationship, and the error loss is calculated. The model parameters are then trained using an error backpropagation algorithm based on this error loss, ultimately resulting in a first AI model capable of predicting the third entity relationship based on the first and second entity relationships.

[0184] The prediction module 1160 is also used to input the fourth entity relationship and the fifth entity relationship into the second AI model to obtain the sixth entity relationship predicted by the second AI model.

[0185] For example, the second AI model is trained using a sample training set, which includes at least one pair of sample sets. Each pair of sample sets includes a fourth entity relation and a fifth entity relation. The fourth and fifth entity relations are input into the second AI model, which predicts the sixth entity relation. The predicted sixth entity relation is compared with the sixth entity relation used as a label, and the error loss is calculated. The model parameters in the second AI model are trained based on the error loss using an error backpropagation algorithm.

[0186] Feature extraction is performed on the fourth entity relation of the sample to obtain the corresponding fourth feature representation; feature extraction is performed on the fifth entity relation of the sample to obtain the corresponding fifth feature representation. The fourth feature representation and the fifth feature representation are then input into the second AI model.

[0187] After multiple training sessions, such as training with 10,000 samples, or when the model converges, a second AI model is obtained that can predict the sixth entity relationship based on the fourth and fifth entity relationships.

[0188] In some embodiments, the gene mutation-anatomy entity relationship is selected as the fourth entity relationship of the sample, and the gene mutation-drug entity relationship is selected as the fifth entity relationship of the sample.

[0189] Feature extraction is performed on DNA sequences, protein structures, chromosomes, alleles, etc., in gene variant entities to obtain gene variant feature representations.

[0190] Features such as anatomical tissues and cell types in anatomical entities are extracted to obtain anatomical feature representations.

[0191] The pharmacological structure and small chemical molecules in the drug entity are extracted to obtain the drug feature representation.

[0192] The fourth entity relationship in the sample includes features of anatomical tissue changes caused by gene mutations, such as a change in a DNA sequence leading to a change in a cell type. The fifth entity relationship in the sample includes features of the correlation between gene mutations and drug molecular structures, such as a change in a DNA sequence enabling treatment through the molecular structure of a drug. The second AI model, after receiving the sample input, predicts the sixth entity relationship, which includes the correlation between cell types and drug molecular structures predicted based on the fourth and fifth entity relationships. The predicted sixth entity relationship is compared with the sixth entity relationship used as a label, and the error loss is calculated. The model parameters in the second AI model are then trained using an error backpropagation algorithm based on the error loss, ultimately resulting in a second AI model capable of predicting the sixth entity relationship based on the fourth and fifth entity relationships.

[0193] For example, when the number of training samples for the AI ​​model is large enough or the model converges, inputting a gene-related entity can predict a non-gene-related entity, indicating a correlation between the gene-related and non-gene-related entities. For instance, inputting a gene-related entity or a gene variant entity can predict a disease-related entity, drug-related entity, or pathway-related entity.

[0194] For example, when the number of training samples for the AI ​​model is large enough or the model converges, inputting a non-gene entity can predict a gene entity, and there is a correlation between gene and non-gene entities. For instance, inputting a disease entity, drug entity, or pathway entity can predict related genes or gene variant entities.

[0195] While the AI ​​model predictions mentioned above have some errors, they can, to a certain extent, help researchers in the biomedical field to conduct further studies on gene molecular structures.

[0196] Figure 12This is a block diagram of a multidimensional knowledge graph search apparatus provided in an exemplary embodiment of this application. The apparatus includes:

[0197] Receiver module 1210 is used to receive entity search terms;

[0198] The query module 1220 is used to query at least one result triple associated with an entity search term based on a multidimensional knowledge graph, which includes entity relationships between gene-type entities and non-gene-type entities.

[0199] Gene-related entities include gene entities and gene variation entities, while non-gene-related entities include at least one of disease entities, drug entities, pathway entities, and anatomical entities.

[0200] For example, at least one of the following types of the resulting triples:

[0201] At least one of the following categories: gene mutation-disease, gene mutation-drug, gene mutation-gene, gene-disease, gene-drug, gene-gene, gene-pathway, and gene-tissue.

[0202] Display module 1230 is used to display at least one result triple associated with entity keywords in the search results interface.

[0203] The response module 1240 is also configured to respond to a trigger operation on the first entity in at least one result triple and display the attribute details of the first entity;

[0204] The attribute details include at least one of the following: the first entity's unique identifier, name, tag, mutation site, mutation frequency, and mutation harmfulness assessment.

[0205] Figure 13 This is a structural block diagram of a computer device 1300 provided in an exemplary embodiment of this application. The computer device 1300 may be a terminal device, such as a smartphone, tablet computer, smart bracelet, etc.

[0206] Typically, computer device 1300 includes a processor 1301 and a memory 1302.

[0207] Processor 1301 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1301 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1301 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1301 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1301 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0208] The memory 1302 may include one or more computer-readable storage media, which may be tangible and non-transitory. The memory 1302 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1302 are used to store at least one instruction, which is executed by the processor 1301 to implement at least one of the multidimensional knowledge graph construction method or multidimensional knowledge graph search method provided in the embodiments of this application.

[0209] This application also provides a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement at least one of the multidimensional knowledge graph construction method or multidimensional knowledge graph search method provided in the above-described method embodiments.

[0210] This application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium; the computer program is read from and executed by a processor of a computer device from the computer-readable storage medium, causing the computer device to perform at least one of the multidimensional knowledge graph construction method or multidimensional knowledge graph search method provided in the above-described method embodiments.

[0211] It should be understood that "multiple" as used in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0212] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0213] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for constructing a multidimensional knowledge graph, characterized in that, The method includes: Acquire at least two biomedical databases, wherein the biomedical databases store different entities and entity relationships connecting the different entities; the at least two biomedical databases include at least two of the following types: gene / protein database, gene variation database, drug database, disease database, pathway database, anatomy database, gene / gene variation-drug database, gene / gene variation-disease database, and gene-anatomy database; Standardize the similar entities in the at least two biomedical databases to obtain at least two standardized entities; the at least two standardized entities include gene entities and non-gene entities, and the gene entities and the non-gene entities have a recursive relationship, which includes the following: from gene variation entity to gene entity to anatomical entity to pathway entity to drug entity to disease entity. Based on any triple from the at least two biomedical databases, determine the first normalized entity corresponding to the first entity in the triple and the second normalized entity corresponding to the second entity in the triple; Based on the entity relationship between the first entity and the second entity, the standardized entity relationship between the first standardized entity and the second standardized entity is determined; wherein, based on the entity relationship in the gene mutation-disease triplet and the entity relationship in the gene mutation-drug triplet, the entity relationship between the disease entity and the drug entity is predicted; and, based on the entity relationship in the gene mutation-pathway triplet and the entity relationship in the gene mutation-disease triplet, the entity relationship between the pathway entity and the disease entity is predicted. Based on entity type, determine start entity and end entity in the first standardized entity and the second standardized entity; based on the relationship between the start entity, the end entity and the standardized entity, generate entity relationship between the standardized entities with association relationship; Based on the at least two standardized entities and the entity relationships between different standardized entities, a multidimensional knowledge graph is constructed, wherein the multidimensional knowledge graph includes the entity relationships between gene-type entities and non-gene-type entities; The gene-type entities include at least one of the gene entities and the gene mutation entities, and the non-gene-type entities include at least one of the disease entities, the drug entities, the pathway entities, and the anatomical entities.

2. The method according to claim 1, characterized in that, The construction of the multidimensional knowledge graph based on the at least two standardized entities and the entity relationships between different standardized entities includes: Based on the at least two standardized entities and the entity relationships between different standardized entities, at least one set of triples is generated, wherein the at least one set of triples includes triples connecting the genetic entity and the non-gene entity; The multidimensional knowledge graph is constructed based on at least one set of triples.

3. The method according to claim 2, characterized in that, The generation of at least one set of triples based on the at least two standardized entities and the entity relationships between different standardized entities includes: Based on the entity relationships between the gene-type entities and / or the entity relationships between the gene-type entities and the non-gene-type entities, at least one set of triples is generated; The at least one set of triplets includes at least one of the following types: gene mutation-disease, gene mutation-drug, gene mutation-gene, gene-disease, gene-drug, gene-gene, gene-pathway, and gene-tissue.

4. The method according to claim 3, characterized in that, The method further includes: Generate a node type file corresponding to the entity type of the normalized entity in the triplet, wherein the node type file includes at least one of the unique identifier of the normalized entity, the standard name of the normalized entity, and the type of the normalized entity; Generate an edge annotation file corresponding to the types in the triples, wherein the edge annotation file includes at least one of the source of the normalized entity, the relationship between the normalized entities in the triples, and the relational subtype of the type of the triples; The relationship subtypes include at least one of the following: drug metabolism dynamics, drug effect dynamics, inhibitors, allosteric agents, and agonists.

5. The method according to any one of claims 1 to 4, characterized in that, The standardization of similar entities in the at least two biomedical databases to obtain at least two standardized entities includes: Perform at least one of the following steps in the at least two biomedical databases: Gene entities and protein entities are standardized to the gene entity; The variant entity is standardized to the aforementioned gene variant entity; The disease entity, phenotype entity, trait entity, variable entity, and rare disease entity are standardized into the disease entity. The pharmaceutical entity and the small chemical molecule entity are standardized as the pharmaceutical entity; Biological process entities and molecular pathway entities are standardized as the pathway entities; Anatomical tissue entities and cell type entities are standardized as the anatomical entity.

6. The method according to claim 1, characterized in that, The step of determining the start entity and end entity from the first standardized entity and the second standardized entity based on entity type includes: The gene-type entities in the first standardized entity and the second standardized entity are determined as the starting entity, and the gene-type entities and non-gene-type entities in the first standardized entity and the second standardized entity are determined as the ending entity.

7. A search method for a multidimensional knowledge graph, characterized in that, The method includes: Receive entity search terms; Based on the multidimensional knowledge graph query, at least one result triple is associated with the entity search term. The multidimensional knowledge graph includes entity relationships between gene-type entities and non-gene-type entities. The gene-type entities include gene entities and gene variation entities. The non-gene-type entities include at least one of disease entities, drug entities, pathway entities, and anatomical entities. Display at least one result triple associated with the search term for the entity in the search results interface; The multidimensional knowledge graph is constructed as follows: At least two biomedical databases are acquired, each storing different entities and entity relationships connecting them. These databases include at least two of the following types: gene / protein database, gene variation database, drug database, disease database, pathway database, anatomy database, gene / gene variation-drug database, gene / gene variation-disease database, and gene-anatomy database. Entities of the same type in the at least two biomedical databases are standardized to obtain at least two standardized entities. These standardized entities include gene-related entities and non-gene-related entities, with a recursive relationship between them, from gene variation entity to gene entity to anatomy entity to pathway entity to drug entity to disease entity. Based on any triple in the at least two biomedical databases, a first standard is determined corresponding to the first entity in the triple. The system identifies entities and their corresponding second standardized entities within the triplet. Based on the entity relationships between the first and second entities, it determines the standardized entity relationships between the first and second standardized entities. Specifically, it predicts the entity relationships between the disease entity and the drug entity based on the entity relationships in the gene mutation-disease triplet and the gene mutation-drug triplet. It also predicts the entity relationships between the pathway entity and the disease entity based on the entity relationships in the gene mutation-pathway triplet and the gene mutation-disease triplet. Based on entity types, it identifies the start entity and the end entity within the first and second standardized entities. Based on the start entity, the end entity, and the standardized entity relationships, it generates entity relationships between the standardized entities with associated relationships. Finally, it constructs the multidimensional knowledge graph based on the at least two standardized entities and the entity relationships between different standardized entities.

8. The method according to claim 7, characterized in that, The method further includes: In response to a triggering operation on the first entity in the at least one result triple, the attribute details of the first entity are displayed; The attribute details include at least one of the following: the first entity's unique identifier, name, tag, mutation site, mutation frequency, and mutation harmfulness assessment.

9. A device for constructing a multidimensional knowledge graph, characterized in that, The device includes: The acquisition module is used to acquire at least two biomedical databases, wherein the biomedical databases store different entities and entity relationships connecting the different entities; the at least two biomedical databases include at least two of the following types: gene / protein database, gene variation database, drug database, disease database, pathway database, anatomy database, gene / gene variation-drug database, gene / gene variation-disease database, and gene-anatomy database; A standardization module is used to standardize similar entities in the at least two biomedical databases to obtain at least two standardized entities; the at least two standardized entities include gene entities and non-gene entities, and the gene entities and non-gene entities have a recursive relationship, which includes the following progression from gene variation entities to gene entities to anatomical entities to pathway entities to drug entities to disease entities. The reconstruction module is configured to: determine a first standardized entity corresponding to a first entity in any triplet from the at least two biomedical databases, and a second standardized entity corresponding to a second entity in the triplet; determine a standardized entity relationship between the first and second standardized entities based on the entity relationship between the first and second entities; predict the entity relationship between the disease entity and the drug entity based on the entity relationship in the gene mutation-disease triplet and the gene mutation-drug triplet; predict the entity relationship between the pathway entity and the disease entity based on the entity relationship in the gene mutation-pathway triplet and the gene mutation-disease triplet; determine a start entity and an end entity in the first and second standardized entities based on entity type; and generate an entity relationship between the standardized entities with an association relationship based on the start entity, the end entity, and the standardized entity relationship. A construction module is used to construct the multidimensional knowledge graph based on the at least two standardized entities and the entity relationships between different standardized entities, wherein the multidimensional knowledge graph includes the entity relationships between the gene-type entities and the non-gene-type entities; The gene-related entities include the gene entity and the gene mutation entity, and the non-gene-related entities include at least one of the disease entity, the drug entity, the pathway entity, and the anatomical entity.

10. A search device for a multidimensional knowledge graph, characterized in that, The device includes: The receiving module is used to receive entity search terms; The query module is used to query at least one result triple associated with the entity search term based on the multidimensional knowledge graph. The multidimensional knowledge graph includes entity relationships between gene-type entities and non-gene-type entities. The gene-type entities include gene entities and gene variation entities, and the non-gene-type entities include at least one of disease entities, drug entities, pathway entities, and anatomical entities. The display module is used to display at least one result triple associated with the entity search term in the search results interface; The multidimensional knowledge graph is constructed as follows: At least two biomedical databases are acquired, each storing different entities and entity relationships connecting them. These databases include at least two of the following types: gene / protein database, gene variation database, drug database, disease database, pathway database, anatomy database, gene / gene variation-drug database, gene / gene variation-disease database, and gene-anatomy database. Entities of the same type in the at least two biomedical databases are standardized to obtain at least two standardized entities. These standardized entities include gene-related entities and non-gene-related entities, with a recursive relationship between them, from gene variation entity to gene entity to anatomy entity to pathway entity to drug entity to disease entity. Based on any triple in the at least two biomedical databases, a first standard is determined corresponding to the first entity in the triple. The system identifies entities and their corresponding second standardized entities within the triplet. Based on the entity relationships between the first and second entities, it determines the standardized entity relationships between the first and second standardized entities. Specifically, it predicts the entity relationships between the disease entity and the drug entity based on the entity relationships in the gene mutation-disease triplet and the gene mutation-drug triplet. It also predicts the entity relationships between the pathway entity and the disease entity based on the entity relationships in the gene mutation-pathway triplet and the gene mutation-disease triplet. Based on entity types, it identifies the start entity and the end entity within the first and second standardized entities. Based on the start entity, the end entity, and the standardized entity relationships, it generates entity relationships between the standardized entities with associated relationships. Finally, it constructs the multidimensional knowledge graph based on the at least two standardized entities and the entity relationships between different standardized entities.

11. A computer device, characterized in that, The computer device includes a processor and a memory, wherein the memory stores at least one computer program, the at least one computer program being loaded and executed by the processor to implement the method for constructing a multidimensional knowledge graph as described in any one of claims 1 to 6, or the method for searching a multidimensional knowledge graph as described in any one of claims 7 to 8.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the method for constructing a multidimensional knowledge graph as described in any one of claims 1 to 6, or the method for searching a multidimensional knowledge graph as described in any one of claims 7 to 8.

13. A computer program product, characterized in that, The computer program product includes a computer program stored in a computer-readable storage medium; the computer program is read from and executed by a processor of a computer device from the computer-readable storage medium, causing the computer device to perform the method for constructing a multidimensional knowledge graph as described in any one of claims 1 to 6, or the method for searching a multidimensional knowledge graph as described in any one of claims 7 to 8.

Citation Information

Patent Citations

  • Pathological knowledge graph construction method and device

    CN113742493A

  • Systems And Methods For Prioritizing The Selection Of Targeted Genes Associated With Diseases For Drug Discovery Based On Human Data

    US20210174906A1