Information retrieval method, medium and equipment
By using a knowledge graph of material relationships for multi-dimensional keyword retrieval and association expansion, the problem of multi-dimensional association query in drug information retrieval in existing technologies has been solved. This enables comprehensive retrieval and visualization of multi-dimensional information, improving the accuracy of search results and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YUYANG ZHISHU (BEIJING) TECHNOLOGY CO LTD
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-05
AI Technical Summary
Existing drug information retrieval methods cannot meet the needs of multi-dimensional related queries, lack visualization of multiple elements such as drugs, genes, pathways, diseases, and treatment plans, and do not fully integrate clinical diagnosis and treatment guidelines, resulting in a lack of accuracy and reliability in search results.
It employs a material relationship knowledge graph for multi-dimensional keyword retrieval, integrating preset material nodes, target nodes, gene nodes, pathway nodes, disease nodes, and treatment plan nodes. Through association expansion, it generates associated sub-graphs, combines structured material information tables and visual output, integrates authoritative clinical evidence, and provides multi-dimensional association queries and detailed information.
It enables comprehensive retrieval of multi-dimensional information, improves the accuracy and reliability of retrieval results, enhances information acquisition efficiency and user experience, and meets the needs of multiple scenarios in clinical diagnosis and treatment and scientific research.
Smart Images

Figure CN121979929A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information retrieval technology, and in particular to an information retrieval method, medium, and device. Background Technology
[0002] In drug development, clinical diagnosis and treatment, and pharmaceutical research, drug information retrieval is a core and fundamental step, requiring users to quickly obtain accurate drug-related data. Current drug information retrieval methods primarily rely on keyword matching or simple association queries within a single database, which have several significant drawbacks: First, most methods only support single-dimensional searches based on drug name or disease name, making it difficult to meet the cross-dimensional association query needs of "gene-target-drug-disease" in precision medicine scenarios, and failing to achieve integrated retrieval of multi-dimensional information. Second, search results are mostly presented as plain text lists or isolated data entries, lacking a visual representation of the logical relationships between multiple elements such as drugs, genes, pathways, diseases, and treatment plans. Users need to manually integrate multi-source data to sort out the core logical chain, which is extremely inefficient. In addition, existing methods do not fully integrate authoritative evidence such as clinical practice guidelines, and the association between drugs and diseases and treatment plans lacks clinical rationale support. Furthermore, the lack of quantitative differentiation of association strength makes it easy for false associations or irrelevant information to interfere. Moreover, existing technologies often directly return "no related drugs" results for queries on non-target genes, failing to explore indirect drug associations at the gene pathway level, and thus failing to meet the needs of researchers and clinicians to explore potential drug targets.
[0003] Therefore, how to support multi-dimensional keyword retrieval and integrate authoritative clinical evidence to improve the accuracy and reliability of information retrieval has become an urgent problem to be solved. Summary of the Invention
[0004] To address the aforementioned technical problems, the present invention provides an information retrieval method, which includes the following steps: S1. Based on the search request initiated by the user using at least one type of keyword, such as substance name, gene name, or disease name, determine at least one target substance node associated with the keyword from the pre-constructed substance relationship knowledge graph. The substance relationship knowledge graph contains at least preset substance nodes, target nodes, gene nodes, pathway nodes, disease nodes, and treatment plan nodes.
[0005] S2, based on the target material node, retrieves the corresponding marketed material data and / or material data under development from the preset material information database, and generates a structured material information table.
[0006] S3 outputs the first search results page, which presents the material information table, to the user.
[0007] S4, in response to the user's selection of the target substance on the first search results page, retrieves detailed information about the target substance from the substance information database, wherein the detailed information includes at least the basic information of the substance, related substance information and related target information.
[0008] S5. Centered on the material node corresponding to the target material, perform association expansion in the material relationship knowledge graph to extract the associated sub-graph containing the associated treatment plan node.
[0009] S6 outputs a second search results page to the user, presenting the associated sub-graph and detailed information.
[0010] The present invention also provides a non-transitory computer-readable storage medium storing at least one instruction or at least one program, wherein the at least one instruction or at least one program is loaded and executed by a processor to implement the above-described information retrieval method.
[0011] The present invention also provides an electronic device, including a processor and the aforementioned non-transitory computer-readable storage medium.
[0012] This invention offers at least the following advantages: By supporting cross-dimensional retrieval of keywords including substance name, gene name, and disease name, and integrating multiple types of nodes in the substance relationship knowledge graph, it can meet users' needs for related queries in various scenarios, improving the comprehensiveness of retrieval dimensions and allowing complete information to be obtained without multiple cross-platform searches; By pre-constructing a substance relationship knowledge graph containing multi-dimensional nodes and relationships, combined with the structured substance information table and the visual output of the related subgraphs, it intuitively and clearly displays the logical chain between drugs and various related elements, enabling users to quickly understand the core relationships and significantly improving information acquisition efficiency; By implicitly integrating authoritative evidence such as clinical treatment guidelines during the construction of the substance relationship knowledge graph, the association between drugs, treatment plans, and diseases possesses clinical rationality and authority, while extracting precise related subgraphs through association expansion effectively filters out false associations and irrelevant information, improving the accuracy and reliability of search results; By outputting the first and second search result pages in stages, it not only meets users' needs for quickly filtering target substances but also provides in-depth detailed information, achieving a balance between search efficiency and information depth, and improving the user experience. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a flowchart of an information retrieval method provided in Embodiment 1 of the present invention. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It is understood that, where appropriate, the terms used to distinguish similar objects can be interchanged so that the invention can also be implemented in other embodiments besides the illustrated or described embodiments. Furthermore, the terms "including," "having," and any variations are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices.
[0017] Example 1 This embodiment provides an information retrieval method, such as Figure 1 As shown, this information retrieval method includes the following steps: S1. Based on the search request initiated by the user using at least one type of keyword, such as substance name, gene name, or disease name, determine at least one target substance node associated with the keyword from the pre-constructed substance relationship knowledge graph. The substance relationship knowledge graph contains at least preset substance nodes, target nodes, gene nodes, pathway nodes, disease nodes, and treatment plan nodes.
[0018] Among them, the substance name specifically refers to the generic name, trade name, or English name of the drug, serving as the core keyword directly associated with drug nodes to meet users' needs for quickly searching for specific drugs. The gene name is a gene identifier with clearly defined naming conventions (such as "EGFR", "ALG10I", "RRM2"), serving as a core hub linking targets, pathways, and drugs, supporting cross-dimensional searches of "gene-target-drug". The disease name is a standardized clinical term for a disease (such as "non-small cell lung cancer", "pancreatic cancer", "breast cancer"), used to locate corresponding treatment drugs and regimens, meeting the search needs of clinical diagnosis and research scenarios.
[0019] The Material Relationship Knowledge Graph is a structured knowledge network that uses nodes to represent drug-related entities and edges to represent the relationships between entities. It integrates multi-dimensional drug-related information to provide logical support for keyword matching and target substance deduction. Specifically, preset material nodes represent core drug nodes, containing basic drug attributes (such as name and type); target nodes represent biological targets of drug action, such as gene-encoded proteins like "EGFR" and "TYMS," used to establish bridges between drugs and genes / diseases, supporting precise matching between targets and drugs; gene nodes represent gene entities, containing attributes such as gene name, function, and pathway affiliation, used to achieve association retrieval of genes, targets, and drugs, especially supporting pathway association deduction for non-target genes; and pathway nodes represent biological metabolic or signal transduction pathways (such as "Pyrimidine"). The "metabolism" node is used to support the indirect association between non-target genes and drugs, filling the technical blind spots in non-target gene retrieval; the disease node represents the disease entity, containing attributes such as disease name, indication, and subtype, used to locate the corresponding treatment drugs and regimens for the disease, realizing direct association retrieval of diseases and drugs; the treatment regimen node represents the clinical treatment regimen (such as "GEM regimen", "GP regimen + sintilimab"), used to integrate treatment logic and establish the association chain of diseases, regimens, and drugs; the target substance node is the drug node associated with the user's keywords, which is the core object for subsequent data retrieval and display, used to clarify the core scope of the search results, and ensure the relevance and accuracy of the subsequent output data.
[0020] This embodiment utilizes the preset association relationships of each node in the material relationship knowledge graph to transform single or multiple keywords input by the user into specific nodes in the material relationship knowledge graph. Then, the target material node is located through the association path between nodes, realizing the rapid conversion of keywords, graph nodes, and target materials, laying the foundation for subsequent data retrieval and visualization.
[0021] Specifically, it receives query commands from users using one or more of the following keywords: substance name (i.e., drug name / trade name), gene name, and disease name, and supports multiple input formats such as single keyword and combined keyword.
[0022] The keywords input by the user are standardized, including Chinese-English matching, synonym expansion, and error correction, to ensure that the keywords are consistent with the naming conventions of the graph nodes. Those skilled in the art will recognize that the standardization methods such as Chinese-English matching, synonym expansion, and error correction in the prior art fall within the protection scope of this invention, and will not be elaborated upon here.
[0023] The standardized keywords are compared one by one with nodes in the material relationship knowledge graph to determine the initial node corresponding to each keyword. For example, if the keyword is "non-small cell lung cancer," it matches the "disease node," and if the keyword is "EGFR," it matches the "gene node." Based on the association between the initial node and the material node, the target material node is derived. If the keyword is a drug name / brand name and the initial node is a "substance node", the initial node is directly identified as the target substance node; if the keyword is a disease name and the initial node is a "disease node", the recommended drug node corresponding to the disease node is extracted as the target substance node through the association path of "disease node-treatment plan node-substance node".
[0024] As described above, by supporting multi-dimensional searches using three types of keywords—substance name, gene name, and disease name—the search entry point becomes more comprehensive, covering users' search needs in different scenarios such as clinical diagnosis and treatment, drug development, and scientific research. Multi-dimensional queries can be achieved without switching between platforms. By relying on a substance relationship knowledge graph containing multi-dimensional nodes for keyword matching and node derivation, the association logic between keywords and target substances becomes clearer, avoiding the information fragmentation problem of searching a single database. By identifying the target substance node at the initial stage of the search, a clear basis is provided for subsequent structured data extraction and association sub-graph generation, making the entire search process logically coherent, more efficient, and shortening the time for users to obtain core information.
[0025] In one specific embodiment, S1 includes the following steps: When the keyword is a gene name, the corresponding gene node is queried in the material relationship knowledge graph.
[0026] If a gene node has target functional attributes, then based on the association between the target node and the material node, the material node directly connected to the gene node is identified as the target material node. The target functional attributes are determined by a preset target gene database.
[0027] If a gene node does not have target functional attributes, then based on the association between gene nodes and pathway nodes, several target nodes located in the same pathway as the gene node will be considered as candidate nodes.
[0028] Based on the relationship between target nodes and material nodes, the material nodes that are directly connected to each candidate node are identified as target material nodes.
[0029] If the keyword is a gene name and the initial node is "gene node", a preset target gene database is invoked. The located gene node is compared with the target gene list in the database to determine whether the gene node has target functional attributes. This preset target gene database integrates authoritative drug target resources, including target information for drugs approved by the FDA / EMA / National Medical Products Administration, target genes recommended by clinical practice guidelines such as CSCO / ASCO, and experimentally validated drug targets from databases such as NCBI and KEGG, ensuring the authority and accuracy of target functional attribute determination.
[0030] If it is determined that the gene node has target function attributes (such as the "EGFR" gene), then based on the preset "target node-substance node" association relationship in the substance relationship knowledge graph, all substance nodes directly connected to the gene node are directly screened. The drugs corresponding to these substance nodes are all drugs that directly target the gene. All of them are identified as target substance nodes, realizing the direct association retrieval of "gene target-drug".
[0031] If, after assessment, the gene node is determined not to possess target functional attributes (e.g., the "ALG10I" gene), then based on the "gene node-pathway node" association in the material relationship knowledge graph, one or more biological pathway nodes (e.g., metabolic pathways, signal transduction pathways) to which the gene node belongs are first located. These pathway nodes are then traversed, and all nodes within the pathway that have been proven to be drug targets are extracted and designated as candidate target nodes. Based on the "target node-material node" association in the material relationship knowledge graph, material nodes directly connected to each candidate target node are screened one by one. The drugs corresponding to these material nodes are all drugs that act on the core targets of the pathway to which the gene belongs, and all of them are identified as target material nodes, achieving indirect association retrieval of "non-target gene-pathway-target-drug".
[0032] As described above, by searching for gene name keywords, we have achieved direct and accurate matching between target genes and drugs, and filled the drug search blind spots for non-target genes through pathway association. This ensures that the search results comprehensively cover both directly acting drugs and potentially related drugs, meeting the needs of multiple scenarios such as clinical diagnosis and treatment, drug development and scientific research.
[0033] In one specific implementation, the material relationship knowledge graph is constructed through the following steps: S10: Perform entity recognition and relation extraction on the raw data from several data sources to obtain a set of entity nodes and a set of relationships between entity nodes. The set of entity nodes includes at least material nodes, target nodes, gene nodes, pathway nodes, and disease nodes.
[0034] S20: For each relationship in the set of relationships, determine the confidence weight of the current relationship based on the reliability level of the data source corresponding to the current relationship.
[0035] S30: Based on the treatment recommendation information extracted from standardized treatment guidelines, construct corresponding treatment plan nodes and the relationship between each treatment plan node and the disease node.
[0036] S40 sets the patient applicability conditions included in the treatment recommendation information as a constraint attribute on the association relationship.
[0037] S50 generates a material relationship knowledge graph based on the set of entity nodes, treatment plan nodes, confidence weights and constraint attributes corresponding to each relationship.
[0038] The data sources include substance databases, genomics databases, disease databases, and standardized treatment guidelines. Substance databases provide basic information such as the generic name, English name, type, target, dosage form, and specifications of drugs; genomics databases provide data such as gene names, protein names, gene functions, and variant types; disease databases provide information such as disease names, precise subtypes, and indications; and standardized treatment guidelines (such as the 2024 CSCO Non-Small Cell Lung Cancer Treatment Guidelines) provide core clinical information such as treatment recommendations, patient eligibility criteria, and levels of evidence.
[0039] Furthermore, clinical trial databases (such as ClinicalTrials.gov), pathway databases (such as KEGG), and real-world medical data can be supplemented to ensure that the data covers the entire process of drug development, clinical diagnosis and treatment, and basic research.
[0040] The integrated raw data undergoes cleaning and standardization, including: removing duplicate data, logically contradictory data (such as drug data unrelated to indications and diseases), and data lacking key attributes; standardizing entity naming conventions (e.g., using HGNC for gene names, Chinese-English glossary codes for drug categories, and ICD-10 codes for disease names); and converting unstructured data (such as guideline text) into a semi-structured format to lay the foundation for subsequent entity recognition. Those skilled in the art will recognize that any data cleaning and standardization method in the prior art falls within the scope of this invention, and will not be elaborated further here.
[0041] Natural language processing (NLP) techniques are employed to perform deep analysis of standardized data, enabling entity recognition and relation extraction. Specifically, pre-trained models such as Bidirectional Encoder Representations from Transformers (BERT) are used to label entity types, including substances, targets, genes, pathways, and diseases, forming a set of entity nodes. Each entity node contains unique core attributes; for example, substance nodes include "common name (Chinese name), common name (English name), and type," while gene nodes include "gene name and functional annotation." Relationship extraction algorithms based on semantic similarity (such as sequence labeling models fusing BERT and Conditional Random Fields) are used to mine the relationships between entities, forming a set of associations, including association types such as "substance-target," "gene-pathway," "pathway-disease," and "substance-disease," to clarify the logical connections between entities. Those skilled in the art will recognize that the BERT model and Conditional Random Fields in the prior art fall within the scope of this invention, and will not be elaborated upon further here.
[0042] In one specific embodiment, S20 includes the following steps: S201, obtain the reliability level and basic weight coefficient corresponding to each data source.
[0043] S202, For each relationship, determine at least one data source from which the current relationship originates.
[0044] S203. Calculate the confidence weight corresponding to the current relationship based on each data source, reliability level, and basic weight coefficient corresponding to the current relationship.
[0045] Based on the authority, data validation level, and clinical / research acceptance of the data sources, all data sources are divided into three levels to ensure the objectivity and rationality of the classification: Level 1 data sources include standardized treatment guidelines, substance databases, drug instructions, and approval documents from regulatory agencies such as the FDA / EMA / National Medical Products Administration. These data sources have undergone rigorous review and practical verification, and their information accuracy is the highest. Level 2 data sources include Phase III and above clinical trial results, meta-analysis literature, genomics databases, and disease databases. These data sources are based on large-sample studies or multi-study integration analyses, and their conclusions are highly reliable, but their applicability may have certain limitations. Level 3 data sources include basic research literature, KEGG and other pathway database annotations, and preliminary real-world data collection. These data sources provide theoretical or preliminary practical support for correlations, but have not yet undergone large-scale clinical validation.
[0046] The base weight coefficient directly reflects the credibility of the data source itself, and the specific value can be set by the implementer according to the actual situation. For example, in this embodiment, the base weight coefficient corresponding to the first-level data source is set to 0.8, the base weight coefficient corresponding to the second-level data source is set to 0.6, and the base weight coefficient corresponding to the third-level data source is set to 0.4.
[0047] For each relationship in the set of relationships, at least one data source is identified as its origin through data tracing technology. Specifically, if a relationship originates from only a single data source, that data source is directly labeled; if a relationship originates from multiple data sources, all relevant data sources are fully labeled to ensure that subsequent weight calculations cover all sources of evidence. Those skilled in the art will recognize that any existing data tracing technology falls within the scope of this invention, such as metadata embedding technology, data fingerprint matching technology, semantic association tracing technology, and block tracing technology, which will not be elaborated upon here.
[0048] If all data sources corresponding to a certain relationship are consistent and have no logical conflicts, then the confidence weight is calculated by weighting and summing the reliability levels of all data sources corresponding to the relationship based on the basic weight coefficient of each data source. If different data sources corresponding to a certain relationship contradict each other (e.g., data source A supports the relationship, data source B denies it), then the product of the reliability level and the basic weight coefficient of the data source with the higher reliability level is used as the confidence weight. If the conflicting data sources have the same reliability level, then the average of the products of the reliability levels and the basic weight coefficients of all conflicting data sources is used as the confidence weight.
[0049] Furthermore, this embodiment uses the basic weight coefficient as the core and adjusts it based on the data publication / update time in the data source to obtain the update weight coefficient corresponding to each data source. Specifically, a time threshold is set to divide the data publication / update time into three time intervals and determine their corresponding update weight coefficients: for data sources published / updated within the last year (inclusive), the update weight coefficient is increased by 10% based on the corresponding basic weight coefficient; for data sources published / updated within 1-3 years (inclusive), the update weight coefficient remains unchanged based on the corresponding data source's basic weight coefficient; for data sources published / updated more than 3 years ago, the update weight coefficient is decreased by 10% based on the corresponding data source's basic weight coefficient.
[0050] It should be noted that the specific values for the time range, the upward adjustment value, and the downward adjustment value can be set by the implementer according to the actual situation.
[0051] The above-mentioned standardized confidence weight calculation process assigns precise confidence weights to each relationship, ensuring the authority and reliability of the relationships in the material relationship knowledge graph and providing a quantitative basis for the accurate screening of subsequent search results.
[0052] In one specific embodiment, S30 includes the following steps: S301, extract treatment recommendation information from standardized treatment guidelines data, wherein the treatment recommendation information includes at least a description of the treatment plan, an identifier of the applicable disease, and an identifier of one or more substances included in the treatment plan.
[0053] S302, Based on the treatment plan description, create or update the corresponding treatment plan node in the entity node set.
[0054] S303, Establish the first association between the treatment plan node and the disease node determined by the applicable disease identifier.
[0055] S304, Establish a second association between the treatment plan node and the substance nodes identified by each substance identifier.
[0056] This process involves selecting chapters related to drug therapy from standardized treatment guidelines. The guidelines are then segmented into sentences, words, and stop words, retaining core medical terminology. A rule-based and machine learning-based hybrid extraction algorithm is used to extract treatment plan descriptions (the specific treatment plan name and core features recommended by the guidelines), applicable disease identifiers (the disease name and precise subtype corresponding to the treatment plan), and substance identifiers (the generic, English, or brand names of the drugs included in the treatment plan). These extracted core elements are stored in a standardized data table according to the correspondence between "treatment plan description - applicable disease identifier - substance identifier," ensuring traceable relationships between elements.
[0057] The applicable disease identifier is the name and precise subtype of the disease corresponding to the treatment plan. It is used to clarify the scope of application of the treatment plan and provide a basis for establishing the association between the treatment plan and the disease node. The substance identifier is the drug identifier (generic name, English name, brand name) included in the treatment plan. It serves as a key link connecting the treatment plan and the drug node, ensuring the accurate mapping of the plan's components.
[0058] The core attributes of a treatment plan node include the plan name, plan composition, administration method, administration cycle, guideline evidence level, and applicable treatment lines. If there is no node corresponding to the treatment plan in the entity node set, a new node is created and the above attributes are filled in. If the treatment plan node already exists in the entity node set, the newly extracted attribute information is compared, and missing attributes are supplemented and changed attributes are updated.
[0059] Based on the applicable disease identifier, the corresponding disease node is accurately matched in the entity node set (e.g., if the applicable disease identifier is "EGFR mutation positive lung adenocarcinoma", then the "EGFR mutation positive" subtype node under the "lung adenocarcinoma" node is located). Through the node association function of the graph database, the treatment plan node is connected with the matched disease node to form the first association relationship of "disease node-treatment plan node". The semantics of the first association relationship is "recommended treatment plan for a certain disease / subtype". A reliability level attribute is attached to the association relationship to clarify the applicable disease of the treatment plan, so that the corresponding treatment plan can be quickly located by disease during retrieval.
[0060] Based on the substance identifier, the corresponding substance node is matched in the entity node set (e.g., if the substance identifier is "gemcitabine", then the substance node with the generic name "gemcitabine" is located), supporting cross-dimensional matching of Chinese and English, trade names and generic names; The treatment plan node is connected one by one with all the corresponding substance nodes to form a second association relationship of "treatment plan node - substance node". The semantics of the second association relationship is "the drugs included in the treatment plan", and the role attributes of the drugs in the plan (such as "main drug" or "auxiliary drug") are added to clarify the drug composition of the treatment plan. This allows the corresponding drugs to be located through the plan or related plans to be inferred from the drugs during retrieval.
[0061] The above-mentioned methods, through structured analysis of standardized treatment guidelines, extract core treatment recommendation elements, transforming unstructured guideline texts into standardized data, and providing an operational data source for integrating the atlas into clinical logic; by creating or updating treatment plan nodes and aggregating core plan attributes, treatment plan information becomes more centralized and standardized, avoiding the confusion caused by scattered plan information and improving the logic of the atlas; by establishing a two-level association relationship of "disease node - treatment plan node - substance node", the logical chain of "disease-plan-drug" in clinical diagnosis and treatment is visualized in the atlas, enabling search results to reflect the recommendation logic of authoritative guidelines and ensuring the timeliness and authority of the atlas's clinical guidance.
[0062] In one specific embodiment, the treatment recommendation information also includes patient eligibility criteria, which include at least one of the following: disease stage, performance status score, specific biomarker status, gene mutation type, and previous treatment history. S40 includes the following steps: S401 transforms the patient eligibility criteria into a structured logical expression that can be parsed by the graph query engine.
[0063] S402 stores the structured logic expression as the constraint attribute of the first association.
[0064] Among these criteria, patient eligibility refers to the patient selection standards specified in the guidelines for a particular treatment regimen, used to clarify the applicable population for that regimen. Disease staging is a standardized classification of disease progression (e.g., stages I-IV of non-small cell lung cancer), used to differentiate treatment options at different disease stages. Performance status score is a quantitative indicator assessing a patient's physical function, ranging from 0 to 2 points; a lower score indicates better physical condition and reflects the patient's tolerance to the treatment regimen. Specific biomarker status refers to the results of biomarker testing related to disease diagnosis and treatment (e.g., EGFR mutation positive / negative), used to support the selection of precision targeted therapy regimens and realize the logic of "personalized treatment." Gene mutation type refers to abnormal changes at the gene level (e.g., EGFR p.L858R mutation, ALK fusion), used to identify the applicable gene mutation population for targeted drug therapy regimens. Previous treatment history refers to the treatment methods previously received by the patient (e.g., "no previous chemotherapy history," "previously received platinum-based chemotherapy"), used to differentiate between first-line, second-line, and other treatment lines, conforming to clinical treatment pathways.
[0065] The graph query engine is a core tool for retrieving nodes, relationships, and attributes in a knowledge graph of material relationships. It can parse structured logical expressions and filter out "disease-treatment plan" relationships that meet the criteria.
[0066] Therefore, the applicable conditions for patients are standardized by using unified attribute naming and attribute value formats. Based on the combination relationships of conditions in the guidelines, explicit logic between conditions is achieved using "AND" and "OR" operators. Structured logical expressions are formed by combining logical operators in the structure of "attribute name + operator + attribute value," transforming natural language conditions into a format recognizable by the graph, thus providing a basis for filtering association relationships. For example, disease staging is standardized to the "I-IV" format, physical performance status scores to the "0-2" format, and biomarker status to the "positive / negative / none" format.
[0067] By binding the structured logical expression with the first association relationship of "disease node - treatment plan node", the association relationship is upgraded from "unconditional association" to "conditional association". This allows the system to automatically match the treatment plan node that meets the patient's conditions during retrieval, ensuring the clinical accuracy of the retrieval results.
[0068] As described above, by transforming the patient application conditions described in natural language into structured logical expressions, non-standardized clinical screening conditions become quantifiable rules that can be parsed by the graph, thus solving the problem of association confusion caused by ambiguity in natural language. By binding structured logical expressions as constraint attributes of the first association, the "disease-treatment plan" association is no longer an unconditional match, but has a precise patient condition filtering function, providing data support for the accuracy of subsequent search results.
[0069] In one specific embodiment, S50 includes the following steps: S501 maps each entity in the entity node set to a node in the material relationship knowledge graph.
[0070] S502 maps each association in the association set to a directed or undirected edge connecting two corresponding nodes.
[0071] S503 uses confidence weights and constraint attributes as attributes of corresponding directed or undirected edges to construct a knowledge graph of material relationships.
[0072] Each entity in the entity node set is assigned a unique ID to ensure that the node can be accurately located in the graph, avoiding duplication or confusion. Each node is also labeled with a type (such as "drug", "gene", "disease") to facilitate the graph query engine to quickly identify the node category and improve retrieval efficiency.
[0073] Simultaneously, the core attributes of each entity are bound to the graph nodes. The attributes of different types of nodes strictly correspond to their definitions. For example, substance nodes are filled with attributes such as Chinese generic name, English generic name, molecular structure, molecular type, chemical formula, PubChem ID, target, category, and ATC code. Gene nodes are filled with attributes such as gene name, HGNC ID, functional annotation, and pathway affiliation. Disease nodes are filled with attributes such as disease name, ICD-10 code, precise subtype, and indication. Treatment regimen nodes are filled with attributes such as regimen name, regimen composition, administration method, administration cycle, and reliability level. Other nodes (target, pathway, patient type, etc.) are filled with corresponding information according to preset attribute templates (e.g., the patient type node is filled with "stage: stage IV" and "performance status score: 0-1").
[0074] Directed edges are nodes that connect with a clear logical direction, indicating the direction of the association. They are used to represent unidirectional logical associations (such as drug → target), making the graph logic clearer and more traceable. Undirected edges are nodes that connect without a clear logical direction, only indicating that there is an association between entities. They are used to represent bidirectional associations (such as gene-pathway), adapting to semantic scenarios without a clear direction.
[0075] Based on the semantics of the association, define the edge type. For example, the edge type corresponding to the "substance-target" association is "acts on", the edge type corresponding to the "disease-treatment plan" association is "recommendation", the edge type corresponding to the "treatment plan-substance" association is "included", and the edge type corresponding to the "gene-pathway" association is "participate". Based on the entity pair information in the association set, connect the corresponding nodes through the defined edge type and direction.
[0076] By embedding the confidence weights that quantify the reliability of the relationships and the constraint attributes that limit the scope of application of the relationships into the attribute fields of the corresponding edges, the edges are endowed with quantitative features and conditional filtering functions, ultimately forming a complete and accurate knowledge graph of material relationships.
[0077] The above describes how, using a graph database as a carrier, entities are mapped to graph nodes, relationships are mapped to connecting edges between nodes, and quantitative attributes are embedded in the attribute fields of the edges. This process transforms data elements into a graph structure, ultimately forming a knowledge graph with logical coherence, authoritative associations, and clinical adaptability, providing core data support for drug information retrieval.
[0078] S2, based on the target material node, retrieves the corresponding marketed material data and / or material data under development from the preset material information database, and generates a structured material information table.
[0079] Specifically, the core identifiers of the target material nodes (such as common name, English name, and unique ID) are used as search keywords to perform precise queries in the material information database, ensuring a unique association between the spectral nodes and the database data and avoiding data matching bias. Specifically, if the database is in tabular format, precise matching is performed using the "common name" and "English name" fields; fuzzy matching adaptation: if there are scenarios where product names and common names are used interchangeably, fuzzy matching is performed using a thesaurus to ensure association with the target data entry.
[0080] Marketed substance data refers to drug-related data that has been approved by regulatory agencies (such as the FDA / National Medical Products Administration) and is clinically applicable. It is used to meet the information needs of clinical diagnosis and treatment, drug procurement and other scenarios, and is one of the core search output data. Specifically, it includes information such as generic name (Chinese name), generic name (English name), type, category, Chinese category, target name and so on.
[0081] Investigational drug data refers to drug-related data that is in the clinical trial stage and has not yet been approved for marketing. It is used to meet the information needs of drug development, scientific research and other scenarios, and to expand the search coverage. Specifically, it includes information such as drug name, trial phase, applicant institution, dosage form, specifications, indications, and application region.
[0082] The extracted standardized data is categorized into marketed drugs and investigational drugs and organized into tables. With clear field column names and pagination display logic, users can quickly locate core information and improve data acquisition efficiency.
[0083] Additionally, data volume statistics and pagination functions can be added to the information table, allowing users to jump to page numbers and avoid display chaos caused by excessive data volume; "Generic Name (Chinese Name)" and "Drug Name" can also be set as clickable interactive fields, reserving jump interfaces for subsequent detailed information queries and related sub-map displays.
[0084] As described above, by using target material nodes as indexes to associate material information databases, data retrieval becomes more targeted, avoiding indiscriminate data traversal and improving retrieval efficiency and data accuracy. By extracting differentiated core fields according to the classification of marketed drugs / investigation drugs, the field design of the material information table is made more suitable for different scenarios. Clinical users can quickly obtain target and category information of marketed drugs, while research users can focus on data such as the trial phase and application institution of investigational drugs, thereby improving the reliability and usability of search results.
[0085] S3 outputs the first search results page, which presents the material information table, to the user.
[0086] Guided by user search needs, it integrates functions such as pagination display, category switching, pagination control, and interactive navigation to transform structured data into a visual page. This ensures both the clarity and standardization of information display while reserving entry points for subsequent detailed information queries, thus balancing information coverage with user experience.
[0087] Specifically, the keyword search display module can show the user-entered search keywords and search type (such as "keyword: non-small cell lung cancer (disease name)" or "keyword: ALG10I (gene name)") at the top of the page, allowing users to clearly understand the current search context and avoid confusion about the search target during multiple rounds of searching. It also includes total search result statistics (such as "63 marketed drugs" or "1042 drugs under development"), enabling users to quickly understand the data scale.
[0088] The category switching module allows users to set two tabs: "Marketed Drugs" and "Drugs in Development." It defaults to displaying all categories or prioritizes categories based on user needs, such as showing marketed drugs to clinical users and drugs in development to research users. The switching tabs are highlighted, and clicking them refreshes the information table below in real time, eliminating the need for page reloads and improving operational smoothness.
[0089] The pagination control module can be configured with pagination functionality for scenarios with large amounts of data, making it easy for users to accurately locate target data.
[0090] The interactive redirect module can set the "Generic Name (Chinese Name)" (for marketed drugs) and "Drug Name" (for drugs under development) in the information table as clickable links, with each link linked to a unique identifier of the target drug (such as a node ID). When a user clicks the link, the page redirects to the second search results page, providing trigger conditions for displaying detailed information and rendering associated sub-maps.
[0091] Based on the principle of structured data visualization, the standardized material information table is transformed into a search results page that users can operate intuitively. Users can quickly grasp key information without having to interpret complex formats, which reduces the cost of information understanding and improves the information retrieval experience.
[0092] S4, in response to the user's selection of the target substance on the first search results page, retrieves detailed information about the target substance from the substance information database.
[0093] Specifically, the system captures user selection behavior through page interaction components, extracts the unique identifier of the selected target substance, and establishes a precise association with the substance information database to ensure the accuracy of data retrieval. The extracted unique identifier is compared with the index field of the substance information database to confirm the existence of the target substance. If it exists, subsequent data retrieval is performed; otherwise, a "No relevant detailed information available" message is returned.
[0094] Detailed information includes at least material basis information, related material information, and related target information.
[0095] Specifically, the core attribute information of the target substance is extracted to form standardized material basis information, which intuitively displays the core characteristics of the substance and provides users with key basic knowledge. This includes basic information for marketed drugs such as Chinese name, English name, molecular structure, molecular type, chemical formula, PubChem ID, target, category, and ATC code; and basic information for investigational drugs such as drug name, molecular type, target, indication, applicant institution, trial phase, dosage form, and specifications. Standardized field representation is used to ensure the professionalism and readability of the information.
[0096] The system mines related drug data for the target substance, identifying drugs with the same target, category, or indication. This data is organized into structured tables to provide users with information on related substances, offering references for drug comparison and alternative selection. Specifically, "same target association" refers to screening other drugs that act on the same target; "same category association" refers to screening drugs belonging to the same therapeutic category; and "same indication association" refers to screening drugs with the same indication as the target substance. Furthermore, core fields of related substances can be extracted, including drug name, dosage form, manufacturer, English manufacturer's name, brand name, specifications, indication, market status, and market region. This information is displayed in paginated format and supports filtering by fields such as dosage form, manufacturer, and market region.
[0097] This system extracts comprehensive information on the target substance's action target and related drugs, constructing a logical chain of "drug-target-related drug" to integrate relevant target information and meet the needs of scientific research and clinical exploration of target mechanisms. Specifically, it retrieves target information corresponding to the target substance from a substance information database, including gene name, protein name, pathway, and pharmacological mechanism of action; based on a substance relationship knowledge graph, it filters other related drugs corresponding to the target and adds them to the "related drugs" field; and it organizes the relevant target information into a table according to the field order of "gene name-protein name-pathway-pharmacological mechanism of action-related drug," supporting filtering by gene name, pathway, and other fields, facilitating users to quickly locate core target mechanism information.
[0098] As described above, the user's selection operation on the first search results page serves as the trigger signal. By linking the target substance's unique identifier to the substance information database, the database is accessed to retrieve the full-dimensional information corresponding to the target substance. This provides complete data support for the visualization display on the second search results page, satisfying the user's search needs from initial screening to in-depth understanding.
[0099] S5. Centered on the material node corresponding to the target material, perform association expansion in the material relationship knowledge graph to extract the associated sub-graph containing the associated treatment plan node.
[0100] In one specific embodiment, S5 includes the following steps: Starting from the material node corresponding to the target substance, traverse along the relationship edges in the material relationship knowledge graph, and identify the target node, gene node, disease node that are directly connected to the material node corresponding to the target substance, as well as the treatment plan node that is associated through the disease node and whose constraint attributes on the association relationship match the current context.
[0101] In this process, based on the unique identifier of the target substance in the material relationship knowledge graph, the corresponding substance node is identified as the starting point for traversal, ensuring the accuracy of the traversal direction. The traversal depth is limited to 2-3 layers. The first layer traverses the nodes directly connected to the target substance node (target, gene, disease), and the second layer traverses the treatment plan nodes associated with the disease node, avoiding over-expansion that could lead to subgraph chaos. The traversal direction adopts "bidirectional traversal + targeted filtering." Bidirectional traversal (e.g., substance → target, target → gene) is performed on target, gene, and pathway nodes, while targeted traversal ("disease → treatment plan") is performed only on treatment plan nodes, ensuring that the association logic aligns with clinical scenarios. Pre-set filtering rules can also be used to remove nodes unrelated to the core logic (e.g., pathway nodes that only participate in basic research and have no clinical relevance), while retaining associated nodes with diagnostic or research value.
[0102] Specifically, target node extraction involves traversing the "acting on" relationship edges of target substance nodes to extract directly related target nodes, which correspond to the direct targets of the drug; gene node extraction involves extracting corresponding gene nodes through the "encoding" relationship edges of target nodes, while also supplementing the extraction of core gene nodes that are indirectly related to the target substance (such as gene nodes involved in drug metabolism); and disease node extraction involves traversing the "indication" relationship edges of target substance nodes to extract directly related disease nodes and precise subtypes, thus clarifying the clinically applicable disease range of the drug.
[0103] The current context is the search background information that includes the user's search keywords, search type (such as disease / gene search), and implicit needs (such as stage IV disease treatment). It is used to provide reference standards for matching constraint attributes and ensure the accuracy of treatment plan selection.
[0104] Furthermore, using disease nodes as intermediaries, and combining the constraint attributes of the association with the current search context, treatment plan nodes that are suitable for the target substance are selected to ensure that the sub-graph contains authoritative clinical diagnosis and treatment logic. Specifically, for each extracted disease node, its "recommendation" relationship edges are traversed to obtain all associated treatment plan nodes. The constraint attributes of the "disease node-treatment plan node" association are extracted and compared with the current search context (e.g., the user's search keyword is "non-small cell lung cancer," and the context includes "stage IV" and PS score "0-1"). If the constraint attributes completely match the context, the treatment plan node is retained; if there is a partial match or no match, the nodes are sorted by confidence weight, and high-authority treatment plan nodes with weights greater than a preset threshold are retained to ensure the clinical suitability and authority of the screening results. The specific value of the preset threshold can be set by the implementer according to the actual situation; for example, it is set to 0.6 in this embodiment.
[0105] The extracted core association nodes (targets, genes, diseases, treatment plans) and target substance nodes are recombined according to the original association relationship to form a structurally complete and logically clear association sub-graph, providing structured data for visualization.
[0106] As described above, by traversing radially around the target substance node, the associated sub-graph focuses on the core drug, allowing users to quickly locate key nodes related to the target drug and improving information acquisition efficiency. By combining constraint attributes with the current context to filter treatment plan nodes, the treatment plans included in the associated sub-graph have clinical suitability and authority, enhancing the clinical reference value of the search results.
[0107] S6 outputs a second search results page to the user, presenting the associated sub-graph and detailed information.
[0108] With user in-depth query needs as the core, the system is divided into clearly defined display modules to achieve a two-way presentation of detailed data and related logic. This allows users to obtain accurate, full-dimensional information about drugs and to intuitively understand the clinical logic chain of drugs, targets, diseases, and treatment plans through graphs, thus meeting the dual needs of clinical diagnosis and treatment as well as scientific research.
[0109] In this embodiment, the second search results page adopts a three-column layout of "top navigation + left detailed information area + right-side graph visualization area," balancing information density and ease of operation. Specifically, the top navigation and basic identification module can display the core identifiers of the target substance, including the common Chinese name, common English name, and source of search keywords, to clarify the current query context; it can also be equipped with function buttons such as "Homepage Return," "Data Download," "Full Screen Graph," and "Help Instructions," supporting users to quickly return to the search entry, download detailed information tables, and view related sub-graphs in full screen, improving operational flexibility.
[0110] The detailed information display module on the left can be categorized into drug information and related information at two levels, using a collapsible panel structure that allows users to expand / collapse as needed, avoiding information redundancy. Specifically, the drug information panel, when expanded, displays basic material information, with fields arranged in the order of "Chinese name, English name, molecular structure, molecular type, chemical formula, PubChem ID, target, category, ATC code." The molecular structure is displayed intuitively as an image, and core fields (such as target and category) are highlighted in bold. The related information panel contains two sub-panels: "Related Drugs" and "Related Targets," both displayed in structured table format. The "Related Drugs" table has columns named as follows: "Drug Name, Dosage Form, Manufacturer, English Manufacturer, Trade Name, Specification, Indication, Market Status, Market Region." It supports filtering by fields such as dosage form, manufacturer, and market region, and supports pagination. The "Related Targets" table has columns named as follows: "Gene Name, Protein Name, Pathway, Pharmacological Mechanism of Action, Related Drugs." It supports filtering by fields such as gene name and pathway. Related drugs are displayed separated by semicolons for easy and quick searching.
[0111] The right-hand sub-graph visualization module, centered on the sub-graph, provides interactive visualization capabilities. Specifically, The map layout adopts a "central radial" layout, with the target substance node located at the center of the map, and target, gene, and disease nodes distributed around it. Treatment plan nodes are displayed in association with disease nodes to ensure a clear logical chain. Different types of nodes use differentiated styles (e.g., drug nodes are blue circles, disease nodes are red squares, and treatment plan nodes are green diamonds). Node size is positively correlated with the strength of the association. When the mouse hovers over a node, its core attributes are displayed (e.g., a disease node displays "Non-small cell lung cancer (Stage IV)"). The association type and confidence weight are also labeled. The "disease-treatment plan" edge is additionally labeled with constraint attributes. Edges with different confidence weights use different thicknesses: weights ≥0.8 are thick, 0.6-0.7 are medium, and ≤0.5 are thin. Interactive functions such as node dragging, graph zooming, and link highlighting are also supported. Clicking on a node highlights its core association path with the target substance, and the detailed information area automatically locates the relevant information for that node.
[0112] The auxiliary explanation module can be located below the associated subgraph, annotating the style meaning of different types of nodes and edges, as well as the confidence weight explanation, to help users quickly understand the graph logic; annotating detailed information and the data source of the associated subgraph to enhance the authority of the information; and displaying the last update time of the data to remind users of the timeliness of the information.
[0113] Based on the principles of deep information integration and logical visualization, the detailed information of the target substance and the associated sub-maps are structurally integrated and transformed into a second search results page that users can intuitively operate and browse in depth. This allows users to obtain accurate, full-dimensional information about the drug and to intuitively understand the clinical logical chain of "drug-target-disease-treatment plan" through the map, meeting the dual needs of clinical diagnosis and scientific research and improving the depth and efficiency of information exploration.
[0114] The above-mentioned features support cross-dimensional retrieval of three types of keywords: substance name, gene name, and disease name. By integrating multiple types of nodes into the substance relationship knowledge graph, it can meet users' needs for related queries in various scenarios, improving the comprehensiveness of retrieval dimensions and allowing users to obtain complete information without multiple cross-platform searches. By pre-constructing a substance relationship knowledge graph containing multi-dimensional nodes and relationships, combined with a structured substance information table and the visual output of related subgraphs, it intuitively and clearly displays the logical chain between drugs and various related elements, enabling users to quickly understand core relationships and significantly improving information acquisition efficiency. By implicitly integrating authoritative evidence such as clinical treatment guidelines during the construction of the substance relationship knowledge graph, the association between drugs, treatment plans, and diseases possesses clinical rationality and authority. Simultaneously, by extracting precise related subgraphs through association expansion, it effectively filters out false associations and irrelevant information, improving the accuracy and reliability of search results. By outputting the first and second search result pages in stages, it not only meets users' needs for quickly filtering target substances but also provides in-depth detailed information, achieving a balance between search efficiency and information depth, thus improving the user experience.
[0115] Example 2 Embodiment 2 of the present invention provides a non-transitory computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a method in the method embodiment. The at least one instruction or at least one program is loaded and executed by the processor to implement the information retrieval method provided in the above embodiment.
[0116] Example 3 Embodiment 3 of the present invention provides an electronic device, which includes a processor and the non-transitory computer-readable storage medium of Embodiment 2 of the present invention.
[0117] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. An information retrieval method, characterized in that, The information retrieval method includes the following steps: S1. Based on the search request initiated by the user using at least one type of keyword from substance name, gene name, or disease name, determine at least one target substance node associated with the keyword from the pre-constructed substance relationship knowledge graph, wherein the substance relationship knowledge graph includes at least preset substance nodes, target nodes, gene nodes, pathway nodes, disease nodes, and treatment plan nodes. S2, based on the target material node, retrieve the corresponding marketed material data and / or research material data from the preset material information database, and generate a structured material information table; S3, output the first search results page displaying the material information table to the user; S4, in response to the user's selection of a target substance on the first search results page, obtain detailed information about the target substance from the substance information database, wherein the detailed information includes at least basic substance information, related substance information and related target information; S5, taking the material node corresponding to the target substance as the center, perform association expansion in the material relationship knowledge graph to extract the association sub-graph containing the associated treatment plan node; S6, output a second search results page to the user, presenting the associated sub-map and the detailed information.
2. The information retrieval method according to claim 1, characterized in that, The material relationship knowledge graph is constructed through the following steps: S10, perform entity recognition and relationship extraction on the raw data from several data sources to obtain a set of entity nodes and a set of relationships between entity nodes. The data sources include material databases, genomics databases, disease databases and standardized diagnosis and treatment guidelines data. The set of entity nodes includes at least material nodes, target nodes, gene nodes, pathway nodes and disease nodes. S20, for each relationship in the set of relationships, determine the confidence weight of the current relationship based on the reliability level of the data source corresponding to the current relationship; S30, Based on the treatment recommendation information extracted from the standardized diagnosis and treatment guidelines data, construct corresponding treatment plan nodes and the association between each treatment plan node and the disease node; S40, set the patient applicability conditions contained in the treatment recommendation information as a constraint attribute on the association; S50, the material relationship knowledge graph is generated based on the entity node set, the treatment plan node, the confidence weight and constraint attributes corresponding to each relationship.
3. The information retrieval method according to claim 2, characterized in that, S20 includes the following steps: S201, obtain the reliability level and basic weight coefficient corresponding to each data source; S202, For each relationship, determine at least one data source from which the current relationship originates; S203, calculate the confidence weight corresponding to the current association based on each data source corresponding to the current association, the reliability level, and the basic weight coefficient.
4. The information retrieval method according to claim 2, characterized in that, S30 includes the following steps: S301, Treatment recommendation information is parsed from the standardized diagnosis and treatment guideline data, wherein the treatment recommendation information includes at least a description of the treatment plan, an applicable disease identifier, and an identifier of one or more substances included in the treatment plan; S302, according to the treatment plan description, create or update the corresponding treatment plan node in the entity node set; S303, establish a first association relationship between the treatment plan node and the disease node determined by the applicable disease identifier; S304, establish a second association between the treatment plan node and the substance node determined by each substance identifier.
5. The information retrieval method according to claim 4, characterized in that, The treatment recommendation information also includes patient eligibility criteria, which include at least one of the following: disease stage, performance status score, specific biomarker status, gene mutation type, and previous treatment history. S40 includes the following steps: S401, the patient's applicable conditions are transformed into a structured logical expression that can be parsed by the graph query engine; S402, the structured logical expression is stored as a constraint attribute of the first association relationship.
6. The information retrieval method according to claim 2, characterized in that, S50 includes the following steps: S501, map each entity in the entity node set to a node in the material relationship knowledge graph; S502, map each association in the association set to a directed edge or an undirected edge connecting two corresponding nodes; S503, the confidence weight and the constraint attribute are used as attributes of the corresponding directed or undirected edges to construct the material relationship knowledge graph.
7. The information retrieval method according to claim 2, characterized in that, S1 includes the following steps: When the keyword is a gene name, the gene node corresponding to the gene name is queried in the material relationship knowledge graph; If the gene node has target functional attributes, then based on the association between the target node and the material node, the material node directly connected to the gene node is determined as the target material node, wherein the target functional attributes are determined by a preset target gene database. If the gene node does not have target functional attributes, then based on the association between the gene node and the pathway node, several target nodes located on the same pathway as the gene node will be considered as candidate nodes. Based on the association between target nodes and material nodes, the material nodes that are directly connected to each candidate node are identified as the target material nodes.
8. The information retrieval method according to claim 2, characterized in that, S5 includes the following steps: Starting from the material node corresponding to the target substance, traverse along the relationship edges in the material relationship knowledge graph, and identify the target node, gene node, disease node that are directly connected to the material node corresponding to the target substance, as well as the treatment plan node that is associated through the disease node and whose constraint attributes on the association relationship match the current context, as associated nodes.
9. A non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores at least one instruction or at least one program segment, characterized in that, The at least one instruction or the at least one program segment is loaded and executed by the processor to implement the information retrieval method as described in any one of claims 1-8.
10. An electronic device, characterized in that, Includes a processor and the non-transitory computer-readable storage medium as described in claim 9.
Citation Information
Patent Citations
Knowledge graph construction method and device and knowledge graph retrieval method and device
CN117150029A
Drug knowledge retrieval method and device
CN117370596A
Tumor knowledge base construction system and method based on large model and cue word optimization
CN119168040A
Knowledge graph-based content generation and optimization method and device, equipment and medium
CN120579627A
Mental disorder electronic medical record structured information extraction system based on knowledge graph and natural language
CN121281724A