Medical knowledge graph determination method, apparatus, device, medium, and product
Patent Information
- Application Number
- CN202610825019.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本申请提供一种医疗知识图谱确定方法、装置、设备、介质及产品,解决现有的医疗知识图谱中的知识来源与版本绑定薄弱,导致撤销与审计困难的缺陷
[0014]本申请还提供一种计算机程序产品,包括计算机程序,所述计算机程序被处理器执行时实现如上述任一种所述医疗知识图谱确定方法。
Smart Images

Figure CN122819418A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of medical knowledge engineering and artificial intelligence, and in particular to a method, apparatus, equipment, medium and product for determining medical knowledge graphs. Background Technology
[0002] Medical knowledge graph construction technology is used to structure static authoritative knowledge such as pharmacopoeias and clinical guidelines, define core entities such as diseases and drugs and their relationships, and achieve this through methods such as manually written rules, extraction by a single AI model, pipeline-style multi-AI collaboration, and enhancement of conventional retrieval.
[0003] However, the existing medical knowledge graphs have weak binding between knowledge sources and versions, making revocation and auditing difficult. Summary of the Invention
[0004] This application provides a method, apparatus, equipment, medium, and product for determining medical knowledge graphs, which addresses the shortcomings of existing medical knowledge graphs where the binding between knowledge sources and versions is weak, leading to difficulties in revocation and auditing.
[0005] This application provides a method for determining a medical knowledge graph, including: The medical instruction manual and / or medical guide are parsed and knowledge is extracted to determine seed triples of multiple extracted knowledge and source metadata bound to the seed triples, wherein the source metadata includes the fingerprints of the medical instruction manual and / or the medical guide respectively. A medical knowledge graph is determined based on each of the seed triples and the source metadata.
[0006] As one embodiment, determining the medical knowledge graph based on each of the seed triples and the source metadata includes: The seed triples and the source metadata are merged to obtain the medical knowledge graph, or... Missing content analysis is performed on each of the seed triples. If missing content exists in a seed triple, the missing content is supplemented to obtain an enhanced triple. Each of the seed triples / enhanced triples and the source metadata are merged to obtain the medical knowledge graph.
[0007] As an example, the missing content analysis of each seed triplet, and the supplementation of missing content in the seed triplet to obtain an enhanced triplet, includes: Based on the main intelligent agent, the missing content analysis is performed on each of the seed triplets. The missing content includes one or more of the following: missing or vague pharmacological mechanism of action, missing or vague dosage description, missing or vague description of applicable population, single or outdated source of evidence, and knowledge conflict. If it is determined that the seed triple has missing content, the professional sub-agent corresponding to the missing content is invoked to supplement the missing content of the seed triple in parallel, and the supplemented seed triple is used as the enhanced triple.
[0008] As an example, the step of any of the specialized sub-agents supplementing the missing content of the seed triple includes: The specialized sub-agent submits a structured retrieval request to the retrieval enhancement generation engine based on the missing content of the seed triple, obtains the evidence fragment content and the source of the evidence fragment returned by the retrieval enhancement generation engine, and supplements the missing content of the seed triple based on the evidence fragment content and the source of the evidence fragment.
[0009] As one embodiment, it also includes: Based on the prior confidence of the enhanced triple, determine the Bayesian prior confidence of the enhanced triple; Based on the source of the evidence fragment, the Bayesian prior confidence is incrementally updated to determine the mean and variance of the Bayesian posterior confidence. Based on the mean and variance of the Bayesian posterior confidence, it is determined whether the enhanced triplet can be written into the medical knowledge graph.
[0010] As one embodiment, merging the seed triples and the source metadata to obtain the medical knowledge graph includes: Perform on each of the seed triplets Figure 1 Consistency verification and / or counterfactual testing are performed. The seed triples that pass the verification and / or testing, along with the corresponding source metadata, are merged to obtain the medical knowledge graph, or... The process of merging the seed triples / enhanced triples and the source metadata to obtain the medical knowledge graph includes: Perform on each of the seed triples / the enhancement triples Figure 1 Consistency verification and / or counterfactual testing are performed, and the seed triples / enhanced triples that pass the verification and / or testing, along with the corresponding source metadata, are merged to obtain the medical knowledge graph.
[0011] This application also provides a medical knowledge graph determination device, comprising: The first determining module is used to parse and extract knowledge from medical instructions and / or medical guidelines, and determine seed triples of multiple extracted knowledge and source metadata bound to the seed triples, wherein the source metadata includes the fingerprints of the medical instructions and / or the medical guidelines respectively. The second determining module is used to determine the medical knowledge graph based on each of the seed triples and the source metadata.
[0012] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the medical knowledge graph determination method as described above.
[0013] This application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the medical knowledge graph determination method as described above.
[0014] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the medical knowledge graph determination method as described above.
[0015] This application provides a method, apparatus, equipment, medium, and product for determining a medical knowledge graph. The method includes: parsing and extracting knowledge from medical instructions and / or medical guidelines; determining seed triples of multiple extracted knowledge and source metadata bound to the seed triples, wherein the source metadata includes the fingerprints of each medical instruction and / or medical guide; and determining a medical knowledge graph based on each seed triple and the source metadata. This application binds source metadata to the seed triples simultaneously with their determination. When new medical instructions and / or new medical guidelines are published, this facilitates the rapid retrieval of seed triples corresponding to existing knowledge, achieving reproducibility, auditability, and revocability of knowledge sources. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating the medical knowledge graph determination method provided in this application.
[0018] Figure 2 This is a schematic diagram of the medical knowledge graph determination device provided in this application.
[0019] Figure 3 This is a schematic diagram of the medical knowledge graph determination system provided in this application.
[0020] Figure 4 This is a flowchart illustrating the medical knowledge graph determination method implemented by the medical knowledge graph determination system provided in this application.
[0021] Figure 5 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0023] Figure 1 This is a flowchart illustrating the medical knowledge graph determination method provided in this application, such as... Figure 1 As shown, this application provides a method for determining a medical knowledge graph, which is used to construct and update an accurate medical knowledge graph. The medical knowledge graph includes knowledge such as drug interactions, contraindications, mechanisms, dosage and usage, and population applicability. The medical knowledge graph is applicable to rational drug use systems, drug use assistance systems, and drug use decision-making systems. The method includes steps S110-S120.
[0024] Step S110: Parse and extract knowledge from the medical instruction manual and / or medical guide, determine the seed triples of the extracted knowledge and the source metadata bound to the seed triples, wherein the source metadata includes the fingerprints of the medical instruction manual and / or the medical guide respectively.
[0025] Optionally, medical instructions and / or medical guidelines, including drug instructions and / or drug guidelines such as pharmacopoeias and clinical guidelines published by the Chinese Medical Association, contain authoritative medical knowledge. Furthermore, authoritative electronic medical records and other dynamic diagnostic and treatment data can be introduced, recording information such as medication dosage and treatment adjustments through time attributes. Correspondingly, one or more of the medical instructions, medical guidelines, and dynamic diagnostic and treatment data can be analyzed and knowledge extracted to expand the medical knowledge graph.
[0026] Optionally, this application converts multiple extracted knowledge sources into the same format to verify the validity of units, dimensions, or time across medical instructions and / or guidelines. Furthermore, this application uses various versions of medical instructions and / or guidelines as anchor points. When medical instructions and / or guidelines are updated, a medical knowledge graph corresponding to the updated medical instructions and / or guidelines is constructed. This allows for the rapid retrieval of seed triples corresponding to older knowledge, achieving reproducibility, auditability, and revocability of knowledge sources.
[0027] Furthermore, the unified format of knowledge includes required fields, recommended fields, source / version fields, calculated key fields, and optional evidence fingerprint fields. Required fields include: subject, subject_type, relation, object, object_type, confidence, and source. Recommended fields include: indication, mechanism, clinical recommendation, effect, severity, and monitoring. Source / version fields include: source_id, filename, source_ver, knowledge version, valid_from, and valid_to. Calculated key fields include: triple_id = MD5(subject|relation|object|source_ver), represented as triple_id = MD5(subject|relation|object|source_ver). Optional evidence fingerprint fields include: section, page, paragraph_hash, and locator. It should be noted that the unified format of knowledge is compatible with the comma-separated values (CSV) format.
[0028] Optionally, the binding relationship between the source metadata and the seed triple persists throughout the entire lifecycle of the seed triple. In a preferred embodiment, the seed triple is also bound to an evidence fragment throughout its lifecycle; the evidence fragment is used to demonstrate the credibility of the seed triple.
[0029] Optionally, mapping relationships between different languages can be constructed for disease entities in medical instructions and / or medical guidelines to facilitate subsequent multilingual semantic retrieval and reasoning.
[0030] Optionally, embodiments of this application may employ a medical pre-trained model to parse and extract knowledge from medical instructions and / or medical guidelines. The medical pre-trained model refers to a large-scale model pre-trained using massive amounts of professional data such as medical literature, electronic medical records, and image reports. Preferably, the medical pre-trained model can be adjusted based on the enterprise's private medical data to improve the F1 score (harmonic mean of precision and recall) of entity relation extraction. Preferably, the medical pre-trained model can be combined with Retrieval-Augmented Generation (RAG) technology to improve the accuracy of seed triples extracted by the medical pre-trained model from medical instructions and / or medical guidelines.
[0031] Step S120: Determine the medical knowledge graph based on each of the seed triples and the source metadata.
[0032] Optionally, the seed triples and their corresponding source metadata can be directly combined to obtain a medical knowledge graph. Alternatively, the seed triples can be optimized and verified before being combined with their corresponding source metadata to obtain a medical knowledge graph.
[0033] Optionally, the medical knowledge graph integrates more than 40,000 drug instructions, clinical guidelines, and off-label drug use lists, containing more than 2 million structured early warning rules, which can be reviewed through multi-dimensional combinations of drug attributes.
[0034] Understandably, this application binds source metadata to the seed triples while determining them. When new medical instructions and / or new medical guidelines are published, this facilitates the rapid retrieval of seed triples corresponding to old knowledge, thereby enabling the knowledge source to be reproducible, auditable, and revocable.
[0035] As an example, the source metadata includes one or more of the document identifier, chapter, page number fingerprint, and paragraph fingerprint of each of the medical instruction manual and / or medical guide.
[0036] Optionally, in this embodiment of the application, the medical instruction manual and / or the medical guide are used as trust anchors. Based on the document identifier, chapter, page number, and paragraph of each of the medical instruction manual and / or the medical guide, a stable fingerprint is generated, and one or more of the stable fingerprints are used as source metadata.
[0037] Optionally, the entire lifecycle of a seed triple includes three stages: initial extraction, extended verification, and final fusion. The source metadata corresponding to the seed triple at any stage is used to form a source chain. Any update to the seed triple requires the addition of new source metadata at the end of the source chain to ensure that the knowledge source is traceable, the iteration process is trackable, and the version evolution is auditable.
[0038] Optionally, in this embodiment of the application, the generation and decision trajectory of the seed triple are recorded through audit logs. The decision trajectory includes trajectory points such as retrieval query, evidence weight, calibration parameters and threshold routing, and the seed triple can be replayed based on the trajectory points.
[0039] It is understood that this application uses one or more of the document identifier, chapter, page number fingerprint, and paragraph fingerprint of each of the medical instruction manual and / or medical guide as source metadata of the seed triple. When a new medical instruction manual and / or a new medical guide is published, it is beneficial to quickly find the seed triple corresponding to the old knowledge, so as to realize the reproducibility, auditability and revocability of the knowledge source.
[0040] As one embodiment, determining the medical knowledge graph based on each of the seed triples and the source metadata includes: The seed triples and the source metadata are merged to obtain the medical knowledge graph, or... Missing content analysis is performed on each of the seed triples. If missing content exists in a seed triple, the missing content is supplemented to obtain an enhanced triple. Each of the seed triples / enhanced triples and the source metadata are merged to obtain the medical knowledge graph.
[0041] Optionally, the seed triples and the source metadata are merged to obtain the medical knowledge graph, including: entity alignment of the seed triples, mapping the subjects and objects in different seed triples to nodes, mapping the relationship between subjects and objects to directed edges connecting nodes, merging nodes pointing to the same entity into the same node, and obtaining a network graph, i.e., the medical knowledge graph, while retaining the source metadata during the node merging process.
[0042] Optionally, the missing content of the seed triples may include mechanisms, dosages, populations, and single-source evidence. To improve efficiency and accuracy, multiple agents can be configured to supplement the missing content of the seed triples in parallel. Furthermore, in the construction and completion process of the medical knowledge graph, external authoritative knowledge alignment and domain mechanism rules are prioritized to generate high-precision candidate triples to ensure the accuracy and interpretability of the knowledge; while the link prediction model is only used as an auxiliary means, responsible for large-scale candidate recall and clue discovery at a coarse-grained level, and does not directly participate in the final high-confidence knowledge finalization.
[0043] It is understood that this application merges each of the seed triples and the source metadata to obtain the medical knowledge graph, or supplements the missing content of the seed triples to obtain enhanced triples, and merges each of the seed triples / enhanced triples and the source metadata to obtain the medical knowledge graph. The medical knowledge graph can be flexibly constructed, and the accuracy of the triples can be improved in the scheme of supplementing the missing content of the seed triples.
[0044] As an example, the missing content analysis of each seed triplet, and the supplementation of missing content in the seed triplet to obtain an enhanced triplet, includes: Based on the main intelligent agent, the missing content analysis is performed on each of the seed triplets. The missing content includes one or more of the following: missing or vague pharmacological mechanism of action, missing or vague dosage description, missing or vague description of applicable population, single or outdated source of evidence, and knowledge conflict. If it is determined that the seed triple has missing content, the professional sub-agent corresponding to the missing content is invoked to supplement the missing content of the seed triple in parallel, and the supplemented seed triple is used as the enhanced triple.
[0045] Optionally, the main agent and specialized sub-agents can be deployed on the same device or different devices. Furthermore, the main agent, specialized sub-agents, and medical pre-trained models can be deployed on the same device or different devices.
[0046] Optionally, the main agent is the pharmacist agent, which deeply integrates pharmaceutical knowledge graphs, drug instruction databases, and clinical guidelines to simulate the thinking mode of licensed pharmacists. It can provide full-process auxiliary services such as prescription review, medication guidance, and adverse reaction monitoring. Therefore, the agent can perform missing content analysis on seed triplets to determine whether the seed triplets have one or more problems such as missing or ambiguous pharmacological mechanism of action, missing or ambiguous dosage description, missing or ambiguous description of applicable population, single or outdated evidence source, or knowledge conflict.
[0047] Optionally, specialized sub-agents include alignment agents, mechanism-rule agents, drug-interaction agents, contraindication / population agents, pharmacology agents, and dosage agents. The alignment agent outputs alignment results, a thesaurus, and external identifiers, and generates searchable relational clues. The mechanism-rule agent generates candidates based on known mechanisms / pathways and rule closures, outputting inference rule identifiers (inference_rule_id) and path evidence; candidates are then verified and calibrated using the RAG. The drug-interaction agent extracts and completes the mechanisms of drug / food interactions. The contraindication / population agent standardizes disease / population contraindications and terminology. The pharmacology agent strengthens evidence related to metabolic pathways, targets, and mechanisms. The dosage agent refines dosage / administration regimens for children, those with hepatic or renal impairment, and the elderly.
[0048] Optionally, during the process of the main agent calling specialized sub-agents, task allocation can be driven by the objective function CoverageGain × ΔConfidence × ConflictImpact / Cost. Here, CoverageGain represents the maximum extent to which the specialized sub-agent can supplement missing content, ΔConfidence represents the knowledge confidence that the specialized sub-agent can improve, ConflictImpact represents the task execution efficiency of the specialized sub-agent in resolving knowledge conflicts, and Cost represents the overhead required for the specialized sub-agent to execute the task. The main agent also employs UpperConfidence Bound (UCB) / Thompson sampling, early stopping / retrying, caching, and fingerprint deduplication to reduce latency and cost, prioritizing computational resources for tasks that can fill in critical relationships and reduce uncertainty. The main intelligent agent also has an early stop and retry mechanism. The early stop and retry mechanism means that the current process is terminated immediately when the task execution effect does not meet expectations or the confidence level is sufficient to avoid wasting resources. If it is determined that there is a temporary anomaly or insufficient information, it will automatically re-initiate the attempt to avoid blind search and reduce the average call and end-to-end latency.
[0049] Understandably, this application achieves efficiency and accuracy in supplementing missing content in seed triples through an architecture that combines a main intelligent agent with specialized sub-intelligent agents.
[0050] As an example, the step of any of the specialized sub-agents supplementing the missing content of the seed triple includes: The specialized sub-agent submits a structured retrieval request to the retrieval enhancement generation engine based on the missing content of the seed triple, obtains the evidence fragment content and the source of the evidence fragment returned by the retrieval enhancement generation engine, and supplements the missing content of the seed triple based on the evidence fragment content and the source of the evidence fragment.
[0051] Optionally, the specialized sub-agent generates a structured retrieval request based on the missing content of the seed triple, and sends the structured retrieval request to the retrieval enhancement generation engine for submission.
[0052] Understandably, this application supplements the missing content of the seed triple based on a specialized sub-agent and a retrieval-enhanced generation engine, which can avoid the specialized sub-agent from generating illusions and ensure the accuracy and timeliness of the supplemented content.
[0053] As an example, the medical knowledge graph determination method provided in this application further includes the following steps: Based on the prior confidence of the enhanced triple, determine the Bayesian prior confidence of the enhanced triple; Based on the source of the evidence fragment, the Bayesian prior confidence is incrementally updated to determine the mean and variance of the Bayesian posterior confidence. Based on the mean and variance of the Bayesian posterior confidence, it is determined whether the enhanced triplet can be written into the medical knowledge graph.
[0054] Optionally, the prior confidence of the augmented triples is the prior confidence of the seed triples, which is output by the medical pre-trained model. The prior confidence of the augmented triples can be expressed as: The Bayesian prior confidence of the enhanced triple can be expressed as Beta( , ), , ,in, This can be understood as an estimate of the probability of the triple being true, with a value ranging from 0 to 1. The weights represent the outputs of the medical pre-trained model, used to measure the credibility of the outputs within the overall evidence system.
[0055] Optionally, based on the source of the evidence fragment, the Bayesian prior confidence is incrementally updated to determine the mean and variance of the Bayesian posterior confidence, including: determining the stance of each source of evidence fragment, i.e., determining whether stance∈{support,neutral,refute} is true, where stance is the stance of the source of evidence fragment, support indicates support for the establishment of the enhancement triple, neutral indicates neither support nor opposition to the establishment of the enhancement triple, and refute indicates opposition to the establishment of the enhancement triple.
[0056] The Bayesian prior confidence is incrementally updated based on the confidence level and overall weight of the evidence fragments to determine the mean and variance of the Bayesian posterior confidence.
[0057] The rule for incremental updates is: if stage = support: , β Unchanged; if stance = refute: , Unchanged; if stance = neutral: , ,in, To support the cumulative strength of evidence for establishing a stronger triple, To increase the cumulative value of evidence against establishing a triplet, Indicates the first i The combined weight of each piece of evidence, where ε is the minimum value.
[0058] The formula for calculating the mean of the Bayesian posterior confidence score is: The formula for calculating variance is: .
[0059] Preferably, the formula for calculating the overall weight of the evidence fragments is as follows: ,in, Determined according to the source level of the evidence fragments, the source level of the evidence fragments is related to... Examples of the correspondence between values are as follows: National Pharmacopoeia / Authoritative Guidelines: = 3.0; RCT / System Overview: = 2.5; Expert consensus: = 1.5; Observational study: = 1.0; Case report: = 0.5. This represents the time decay weight, used to reduce the influence of evidence fragments from older dates. λ is the attenuation coefficient. This indicates the number of days since the evidence was released.
[0060] Preferably, after determining the mean and variance of the Bayesian posterior confidence, a decision threshold routing method is used to determine whether the enhanced triples should be written into the medical knowledge graph. In the event of automatic data entry, In certain cases, manual verification will be conducted; otherwise, a supplementary certificate will be required. This represents the confidence threshold for automatic writing into the knowledge graph, for example, 0.80. This indicates the confidence threshold at which further verification is needed, for example, 0.65. Furthermore, a variance threshold can also be set. Used to determine whether the evidence is sufficient. Even if it meets the requirements Manual verification is also required.
[0061] Understandably, this application determines the Bayesian prior confidence of the enhanced triple based on its prior confidence; incrementally updates the Bayesian prior confidence based on the source of the evidence fragment to determine the mean and variance of the Bayesian posterior confidence; and determines whether the enhanced triple is allowed to be written into the medical knowledge graph based on the mean and variance of the Bayesian posterior confidence. This solves the problems that single AI model extraction is prone to producing illusions, the model's confidence cannot be used as a thresholdable probability, and the cross-source conflict resolution and evidence fusion capabilities are insufficient.
[0062] As one embodiment, merging the seed triples and the source metadata to obtain the medical knowledge graph includes: Perform on each of the seed triplets Figure 1 Consistency verification and / or counterfactual testing are performed. The seed triples that pass the verification and / or testing, along with the corresponding source metadata, are merged to obtain the medical knowledge graph, or... The process of merging the seed triples / enhanced triples and the source metadata to obtain the medical knowledge graph includes: Perform on each of the seed triples / the enhancement triples Figure 1 Consistency verification and / or counterfactual testing are performed, and the seed triples / enhanced triples that pass the verification and / or testing, along with the corresponding source metadata, are merged to obtain the medical knowledge graph.
[0063] Optional, Figure 1 Consistency verification includes domain / range / unit / time constraint verification based on the shape constraint language SHACL / web ontology language OWL, and satisfiability checks on dose, threshold, and mutual exclusion conditions based on satisfiability module theory SMT / constraint programming-satisfiability solver CP-SAT. If these conditions are not met, writing is blocked and targeted retrieval correction is triggered. Specifically, this application uses SHACL / OWL to perform strong constraint verification on the domain / range of relations, ensuring the semantic and structural legality of triples, and merging synonyms and removing loops from circular dependencies. Secondly, at the unit and dimension level, all numerical data such as dose, frequency, and threshold must be uniformly converted to standard units before being stored; any inconsistencies result in direct rejection of writing. Finally, at the numerical satisfiability level, solvers such as SMT or CP-SAT are introduced to determine the logical satisfiability of dose ranges, threshold conditions, and mutual exclusion constraints, ensuring that the knowledge is not self-contradictory in numerical logic, thereby constructing a high-quality medical knowledge graph with clear semantics, standardized structure, and numerical satisfiability.
[0064] Optionally, counterfactual testing refers to automatically generating high-impact scenarios such as drug use / population / dosage, triggering targeted retrieval and re-fusion of high-impact and low-confidence conclusions, and feeding the results back into confidence and queue routing.
[0065] Optionally, before writing the seed triples / enhanced triples into the medical knowledge graph, it is necessary to perform population and terminology standardization operations. Population and terminology standardization operations refer to mapping fuzzy populations to standard ontology identifiers and binding quantification thresholds, writing the standardized structure and participating in consistency and decision-making logic.
[0066] Optional, in Figure 1 Following consistency verification and / or counterfactual testing, the seed triples / enhanced triples that have passed the verification and / or testing are written to the medical knowledge graph using idempotent key writing. Specifically, idempotent key writing refers to write / update / revert APIs using triple_id=md5(subject|relation|object|source_ver) as the idempotent key.
[0067] Optionally, after constructing the medical knowledge graph, versioning and chained undo operations can be performed on the medical knowledge graph. Specifically, the triples carry fields such as version, source_ver, valid_from, and valid_to. When a guideline / drug alert is triggered, the affected triples are chained undo and replacement, while retaining derived_from / replaced_by to achieve historical traceability.
[0068] The following is a detailed description of a specific embodiment of the medical knowledge graph determination method provided in this application.
[0069] The aspirin instruction manual is input into a medical pre-trained model, yielding seed triples. This seed triples are then fed into the main agent. The main agent first aligns the entities in the seed triples using synonyms / external IDs, then generates a search request based on the alignment results. Based on the search request, five specialized sub-agents are invoked. It's important to note that these specialized sub-agents act as high-precision candidates; their output is not directly written to the database but serves only as search clues and candidates.
[0070] The main agent generates sub-prompt words for each sub-agent by "pruning the relationship whitelist" and "object type whitelist" from the total prompt words and injecting search clues. The remaining rules are consistent with the output format.
[0071] For example, taking aspirin as an example, the sub-prompt word template is as follows: 1) Align the agents (candidates, no library required) - Relationship whitelist: None (alignment and clues only) - Objective: To provide external IDs, synonyms, category information, and searchable clues. ```text Task: Align the term "Aspirin (ASA)" with an external knowledge base and compile a list of synonyms, and generate a searchable list of leads.
[0072] Output JSON: { "alignments": { "DrugBank": "DB00820", "RxNorm": "1191", "KEGG": "D00109", "ATC": ["B01AC06","N02BA01"] }, "synonyms": ["Aspirin","Acetylsalicylic acid","ASA","Acetylsalicylic acid"], "class_hints": ["Antibacterial drugs","Salicylates","NSAIDs (Special: Irreversible COX-1)"], "interaction_hints": [ "warfarin INR bleeding", "Cardiovascular protective effects of ibuprofen", "clopidogrel dual antiplatelet therapy", "methotrexate protein binding / renal excretion" ] } ``` 2) Mechanism rule-based intelligent agent (candidate, not written into the library) - Relationship whitelist: Only generates candidates, not final relationships. - Objective: Generate candidates using the rule closure of "mechanism → effect → relationship" and label them with `inference_rule_id`. ```text Drug to be tested: Aspirin Known mechanism: Irreversible acetylation of COX-1 → TXA2 decrease → platelet aggregation decrease; platelet function recovery takes 7–10 days. Output JSON.candidates (candidates only): [ { "subject":"Aspirin","relation":"interacts_with","object":"ibuprofen", "effect": "Cardiovascular protective effect ↓", "mechanism": "Reversible NSAIDs compete for COX-1 space, weakening the irreversible inhibition of platelets by aspirin", "recommendation": "If both must be used, take aspirin first and then stagger the ibuprofen dose." "relation_source_type":"inferred","inference_rule_id":"RULE_COX1_IRREV_COMPETE_NSAID" }, { "subject":"Aspirin","relation":"not_recommended_in","object":"Population-high-risk perioperative bleeding population","object_type":"Population", "mechanism": "Irreversible platelet inhibition; functional recovery requires 7–10 days." "recommendation": "Discontinue use 5–7 days before elective surgery". "relation_source_type":"inferred","inference_rule_id":"RULE_PLATELET_RECOVERY_7D" } ] ``` 3) Drug interaction intelligent agent (candidate for writing to the library, requires retrieval → judgment → fusion → threshold) - Relationship whitelist: interacts_with / contraindicated_with / enhances / antagonizes / interacts_with_class - Search query examples: ["Aspirin warfarin INR bleeding guideline", "Aspirinibuprofen cardioprotection", "Aspirin methotrexate protein binding renalexcretion"] - Sub-tips (using the same output format, omitting general rules, and only restricting relation types): ```text Extract only the following relations: interacts_with, contraindicated_with, enhances, antagonizes, interacts_with_class Object type restriction: Drug or DrugClass Main medication: Aspirin (ASA) Returns an array of triples (containing mechanism / severity / recommendation / monitoring). ``` - Expected output snippet: json [ { "subject":"Aspirin","subject_type":"Drug", "relation":"interacts_with","object":"Warfarin","object_type":"Drug", "mechanism": "Cooperative anticoagulation and competition for protein binding sites → INR↑ / Bleeding risk↑", "effect":"INR↑ / Bleeding risk↑","severity":"high", "recommendation":"When used in combination, close monitoring of INR and signs of bleeding is required.", "monitoring":"INR" "confidence": 0.86 }, { "subject":"Aspirin","subject_type":"Drug", "relation":"interacts_with","object":"ibuprofen","object_type":"Drug", "mechanism": "Ibuprofen occupies COX-1, leading to a decrease in the antiplatelet effect of aspirin", "effect":"Cardiovascular protective effect ↓","severity":"medium", "recommendation": "If both must be used, take aspirin first and then stagger the ibuprofen dose." "confidence": 0.80 }, { "subject":"Aspirin","subject_type":"Drug", "relation":"interacts_with","object":"Methotrexate","object_type":"Drug", "mechanism": "Competition between protein binding and renal tubular secretion → Methotrexate concentration ↑", "effect":"toxicity ↑","severity":"high", "recommendation":"Monitor blood drug concentration and adjust dosage as necessary", "monitoring":"MTX concentration / liver and kidney function", "confidence": 0.78 } ] ``` 4) Contraindications / Crowd-based intelligent agents (candidates for writing libraries) - Relationship whitelist: contraindicated_in_population, contraindicated_in_disease, not_recommended_in - Search query examples: ["Aspirin active peptic ulcer contraindication", "Aspirin asthma sensitivity", "Aspirin pregnancy trimester bleeding"] json [ { "subject":"Aspirin","subject_type":"Drug", "relation":"contraindicated_in_disease","object":"active peptic ulcer","object_type":"Disease", "mechanism":"Mucosal damage and bleeding risk ↑","severity":"high", "recommendation":"Disable","confidence":0.92 }, { "subject":"Aspirin","subject_type":"Drug", "relation":"contraindicated_in_population","object":"Aspirin-induced asthma patients","object_type":"Population", "mechanism":"COX-1 inhibition → leukotriene pathway upregulation → bronchospasm","severity":"high", "recommendation":"Discontinue or switch to a non-inducible alternative medication","confidence":0.83 } ] ``` 5) Pharmacological intelligent agents (candidate for library writing) - Relationship whitelist: treats / has_mechanism / belongs_to_class / induces - Search query examples: ["Aspirin low-dose antiplatelet guideline", "Aspirinmechanism COX-1 TXA2", "Aspirin adverse bleeding risk factors"] json [ { "subject":"Aspirin","subject_type":"Drug", "relation":"has_mechanism","object":"Inhibits COX-1-mediated platelet aggregation","object_type":"Mechanism", "effect":"TXA2↓→platelet aggregation↓","confidence":0.90 }, { "subject":"Aspirin","subject_type":"Drug", "relation":"treats","object":"Secondary prevention of atherosclerotic cardiovascular events","object_type":"Disease", "indication":"Low-dose antiplatelet therapy","mechanism":"Irreversible COX-1 inhibition" "recommendation":"75–100 mg / day for long-term use","confidence":0.88 }, { "subject":"Aspirin","subject_type":"Drug", "relation":"induces","object":"bleed","object_type":"SideEffect", "mechanism":"mucosal damage / platelet inhibition","severity":"medium","monitoring":"fecal occult blood / hemoglobin", "confidence": 0.78 } ] The method for determining the medical knowledge graph provided in this application includes steps S1-S8.
[0073] Step S1: Construct seed triples.
[0074] Authoritative sources: National Pharmacopoeia / Authoritative Guidelines / Drug Instructions (Chapters: Indications, Interactions, Contraindications, Dosage and Administration) Initial triples (example): (Aspirin has_mechanism, inhibits COX-1 mediated platelet aggregation) (Aspirin, interacts with class, anticoagulant, effect = increased bleeding risk; recommendation = use with caution and monitor INR) ) (Aspirin, treats, mild to moderate pain) (Aspirin interacts with warfarin); effect = increased risk of bleeding; recommendation = use with caution and monitor INR. (Aspirin, contraindicated_for, active peptic ulcer); severity=high The following is a CSV example of a seed triple: csv subject,subject_type,relation,object,object_type,confidence,indication,mechanism,recommendation,filename,source,source_id,effect,severity,monitoring,source_ver,valid_from,valid_to Aspirin, Drug, has_mechanism, Inhibits COX-1-mediated platelet aggregation, Mechanism, 0.9, Inhibits COX-1 to reduce TXA2 synthesis, triples_aspirin_0001.json, MANUAL, ASP-001, v1, 2024-01-01T00:00:00Z, Aspirin, Drug, Treats, Mild to Moderate Pain, Symptom, 0.88, Analgesic and Antipyretic, triples_aspirin_0002.json, MANUAL, ASP-002, v1, 2024-01-01T00:00:00Z, Aspirin, Drug, interactions_with, Warfarin, Drug, 0.85, synergistic anticoagulation and competition for protein binding sites, avoid or closely monitor INR, triples_aspirin_0003.json, GUIDELINE, GL-2024-CARD, effect=INR↑ / bleeding risk↑, high, monitor INR and signs of bleeding, v1, 2024-01-01T00:00:00Z, Aspirin, Drug, contraindicated_for, active peptic ulcer, Condition, 0.93, disabled, triples_aspirin_0004.json, MANUAL, ASP-004, high, v1, 2024-01-01T00:00:00Z, ``` Step S2: Expand and standardize the seed triplet, including mechanism reinforcement, population standardization, and interaction reinforcement.
[0075] Mechanism enhancement: Add the pathway “irreversible acetylation of COX-1 → platelet function recovery requires 7–10 days” for counterfactual scenarios.
[0076] Population standardization: "Aspirin-induced asthma" is mapped to `SNOMED: Asthma with sensitivity to aspirin` Threshold structure for "renal insufficiency": {metric:eGFR, op:<, value:30, unit:mL / min / 1.73m 2} Interaction reinforcement: (Aspirin, interacts with, clopidogrel); effect = increased bleeding risk; recommendation = assess bleeding risk and weigh it against cardiovascular benefits. (ibuprofen, interacts with aspirin); effect = cardiovascular protective effect ↓; recommendation = stagger or avoid using ibuprofen first.
[0077] Step S3: RAG evidence fusion and calibration.
[0078] a priori: =0.88, =1 → =1 + 0.88 = 1.88, =1 + 0.12 = 1.12 evidence: Authoritative guidance (support): w=3, p=0.95 → Δα=2.85, Δβ=0.15 RCT / Systematic Overview (support): w=2.5, p=0.92 → Δα=2.30, Δβ=0.20 Case report (refute, rare abnormal INR not rising): w=0.5, p=0.20, after conflict penalty γ=0.7 and taking p'=1-p=0.8 → equivalent w=0.35 → Δα=0.28, Δβ=0.07 Posterior: α=7.31, β=1.54 → =7.31 / (7.31+1.54)=0.826; Variance≈0.0146 Routing: Threshold =0.80 → Write to the database; if 0.65≤ If the score is less than 0.80, it will be subject to manual review; if it is less than 0.65, it will be subject to supplementary documentation.
[0079] The key points for writing back fields are as follows: - confidence=0.826 (After calibration, the labeling method can be: temperature / sequence preservation) - sources=[{guidelineID, page / section fingerprint}, {PMID list}, {caseID}] - Set version, source_ver, and valid_from to the guide release date or extraction time.
[0080] Step S4: Figure 1 Consistency and numerical satisfiability verification.
[0081] Semantic constraints: Both ends of `interacts_with` must be `Drug`; the `contraindicated_for` object must be `Condition`.
[0082] Units / dimensions: Low-dose antiplatelet therapy “75–100 mg / day” is normalized to `mg / day` and decoupled from “analgesia 325–650 mg q4–6h” to the `has_dosage` relationship.
[0083] SMT check: If the same patient has both "active bleeding" and "aspirin treatment" → cannot be satisfied, block writing or force downgrade to `not_recommended` and trigger targeted search.
[0084] Step S5: Counterfactual stress test.
[0085] Combined medications: Aspirin + Warfarin + Ibuprofen Trigger test: Increased bleeding risk coexists with decreased cardiovascular protection. Results: The conclusion of "high risk with aspirin × warfarin" is retained; a staggered intake recommendation is added for "aspirin × ibuprofen"; evidence retrieval for "PPI protection" recommendation is added if necessary. Perioperative period: Generate rule candidates for "discontinue aspirin ≥5–7 days before surgery" in the `valid_from` window. If the evidence is insufficient, proceed to the stage of needing to supplement evidence.
[0086] Step S6: Versioning and Chained Undo.
[0087] Initial database entry: valid_from=2024-01-01T00:00:00Z; valid_to=null (Active) Event: An updated guideline was released on June 1, 2025, clarifying the "Bleeding Management Protocol When Used Concurrently with Direct Oral Anticoagulants (DOACs)" Operation: Set valid_to = 2025-06-01T00:00:00Z (Deprecated) for the old triple; write a new triple (containing the new recommendation), valid_from = 2025-06-01T00:00:00Z, and replace_by the record chain. Historical search: as_of=2025-05-01 Returns to the old version suggestion; as_of=2025-07-01 Returns to the new version suggestion Step S7: Perform the write operation in JSON format.
[0088] json { "triples": [ { "subject": "aspirin", "subject_type": "Drug", "relation": "has_mechanism", "object": "Inhibits COX-1-mediated platelet aggregation", "object_type": "Mechanism", "confidence": 0.90, "mechanism": "Irreversible acetylation of COX-1, TXA2 decrease, platelet aggregation decrease", "source": "MANUAL", "source_id": "ASP-001", "filename": "triples_aspirin_0001.json", "source_ver": "v1", "valid_from": "2024-01-01T00:00:00Z", "valid_to": null }, { "subject": "aspirin", "subject_type": "Drug", "relation": "interacts_with", "object": "warfarin", "object_type": "Drug", "confidence": 0.826, "effect": "INR↑ / Bleeding risk↑", Recommendation: Use as needed, closely monitor INR and signs of bleeding; assess benefit / risk. "severity": "high", "monitoring": "INR and signs of bleeding", "source": "GUIDELINE", "source_id": "GL-2024-CARD", "filename": "triples_aspirin_0003.json", "source_ver": "v1", "valid_from": "2024-01-01T00:00:00Z", "valid_to": null }, { "subject": "aspirin", "subject_type": "Drug", "relation": "contraindicated_for", "object": "active peptic ulcer", "object_type": "Condition", "confidence": 0.93, "severity": "high", "recommendation": "Disabled", "source": "MANUAL", "source_id": "ASP-004", "filename": "triples_aspirin_0004.json", "source_ver": "v1", "valid_from": "2024-01-01T00:00:00Z", "valid_to": null }, { "subject": "ibuprofen", "subject_type": "Drug", "relation": "interacts_with", "object": "Aspirin", "object_type": "Drug", "confidence": 0.78, "Effect": "Cardiovascular protective effect ↓", Recommendation: If both medications are needed, prioritize aspirin and stagger ibuprofen use, or switch to a non-interfering NSAID. "source": "GUIDELINE", "source_id": "GL-NSAID-2024", "filename": "triples_ibuprofen_aspirin.json", "source_ver": "v1", "valid_from": "2024-01-01T00:00:00Z", "valid_to": null } ] } ``` Step S8: Auditing and Observability.
[0089] For each aspirin-related triple, retain {doc_hash, section, page, paragraph_hash, queries, weights, calibration}.
[0090] Indicators: Conflict rate, ECE / ACE, withdrawal delay, coverage; when the conflict rate of "aspirin × anticoagulant" increases or the ECE deteriorates, an alarm and targeted supplementary certification are triggered.
[0091] The medical knowledge graph determination device provided in this application is described below. The medical knowledge graph determination device described below and the medical knowledge graph determination method described above can be referred to in correspondence.
[0092] Figure 2 This is a schematic diagram of the medical knowledge graph determination device provided in this application, such as... Figure 2 As shown, this application also provides a medical knowledge graph determination device, comprising: The first determining module 210 is used to parse and extract knowledge from medical instructions and / or medical guidelines, and determine seed triples of multiple extracted knowledge and source metadata bound to the seed triples, wherein the source metadata includes the fingerprints of the medical instructions and / or the medical guidelines respectively. The second determining module 220 is used to determine a medical knowledge graph based on each of the seed triples and the source metadata.
[0093] As one embodiment, the second determining module 220 is configured to: The seed triples and the source metadata are merged to obtain the medical knowledge graph, or... Missing content analysis is performed on each of the seed triples. If missing content exists in a seed triple, the missing content is supplemented to obtain an enhanced triple. Each of the seed triples / enhanced triples and the source metadata are merged to obtain the medical knowledge graph.
[0094] As one embodiment, the second determining module 220 is configured to: Based on the main intelligent agent, the missing content analysis is performed on each of the seed triplets. The missing content includes one or more of the following: missing or vague pharmacological mechanism of action, missing or vague dosage description, missing or vague description of applicable population, single or outdated source of evidence, and knowledge conflict. If it is determined that the seed triple has missing content, the professional sub-agent corresponding to the missing content is invoked to supplement the missing content of the seed triple in parallel, and the supplemented seed triple is used as the enhanced triple.
[0095] As one embodiment, the second determining module 220 is configured to: The specialized sub-agent submits a structured retrieval request to the retrieval enhancement generation engine based on the missing content of the seed triple, obtains the evidence fragment content and the source of the evidence fragment returned by the retrieval enhancement generation engine, and supplements the missing content of the seed triple based on the evidence fragment content and the source of the evidence fragment.
[0096] As one embodiment, the second determining module 220 is configured to: Based on the prior confidence of the enhanced triple, determine the Bayesian prior confidence of the enhanced triple; Based on the source of the evidence fragment, the Bayesian prior confidence is incrementally updated to determine the mean and variance of the Bayesian posterior confidence. Based on the mean and variance of the Bayesian posterior confidence, it is determined whether the enhanced triplet can be written into the medical knowledge graph.
[0097] As one embodiment, the second determining module 220 is configured to: Perform on each of the seed triplets Figure 1 Consistency verification and / or counterfactual testing are performed. The seed triples that pass the verification and / or testing, along with the corresponding source metadata, are merged to obtain the medical knowledge graph, or... The process of merging the seed triples / enhanced triples and the source metadata to obtain the medical knowledge graph includes: Perform on each of the seed triples / the enhancement triples Figure 1 Consistency verification and / or counterfactual testing are performed, and the seed triples / enhanced triples that pass the verification and / or testing, along with the corresponding source metadata, are merged to obtain the medical knowledge graph.
[0098] Figure 3 This is a schematic diagram of the medical knowledge graph determination system provided in this application. Figure 4 This is a flowchart illustrating the medical knowledge graph determination method implemented using the medical knowledge graph determination system provided in this application, as shown below. Figure 3 and Figure 4 As shown, the medical knowledge graph determination system provided in this application includes a medical pre-trained model, a main intelligent agent, a group of professional sub-intelligent agents, a RAG engine, and a graph verifier.
[0099] A medical pre-trained model is used to parse and extract knowledge from medical instructions and / or medical guidelines, identify seed triples of multiple extracted knowledge and source metadata bound to the seed triples, the source metadata including the fingerprints of the medical instructions and / or medical guidelines respectively.
[0100] The main agent is used to analyze the missing content of each seed triple. If there is missing content in the seed triple, it calls a group of specialized sub-agents to supplement the missing content of the seed triple, thus obtaining an enhanced triple.
[0101] Any specialized sub-agent in the specialized sub-agent group is used to submit a structured retrieval request to the RAG engine based on the missing content of the seed triple, obtain the evidence fragment content and evidence fragment source returned by the RAG engine, and supplement the missing content of the seed triple based on the evidence fragment content and evidence fragment source.
[0102] The graph validator is used to perform a graph check on each of the seed triplets. Figure 1 Consistency verification and / or counterfactual testing are performed. The main agent merges the seed triples that pass the verification and / or testing with the corresponding source metadata to obtain the medical knowledge graph.
[0103] It should be noted that the medical knowledge graph determination device and system provided in this application have the same technical effects as the medical knowledge graph determination method, which will not be elaborated further.
[0104] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a medical knowledge graph determination method, which includes: The medical instruction manual and / or medical guide are parsed and knowledge is extracted to determine seed triples of multiple extracted knowledge and source metadata bound to the seed triples, wherein the source metadata includes the fingerprints of the medical instruction manual and / or the medical guide respectively. A medical knowledge graph is determined based on each of the seed triples and the source metadata.
[0105] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0106] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the medical knowledge graph determination method provided by the above methods, the method including: The medical instruction manual and / or medical guide are parsed and knowledge is extracted to determine seed triples of multiple extracted knowledge and source metadata bound to the seed triples, wherein the source metadata includes the fingerprints of the medical instruction manual and / or the medical guide respectively. A medical knowledge graph is determined based on each of the seed triples and the source metadata.
[0107] Furthermore, this application also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is implemented to perform the medical knowledge graph determination method provided by the methods described above, the method comprising: The medical instruction manual and / or medical guide are parsed and knowledge is extracted to determine seed triples of multiple extracted knowledge and source metadata bound to the seed triples, wherein the source metadata includes the fingerprints of the medical instruction manual and / or the medical guide respectively. A medical knowledge graph is determined based on each of the seed triples and the source metadata.
[0108] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0109] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for determining a medical knowledge graph, characterized in that, include: The medical instruction manual and / or medical guide are parsed and knowledge is extracted to determine seed triples of multiple extracted knowledge and source metadata bound to the seed triples, wherein the source metadata includes the fingerprints of the medical instruction manual and / or the medical guide respectively. A medical knowledge graph is determined based on each of the seed triples and the source metadata.
2. The method for determining a medical knowledge graph according to claim 1, characterized in that, The process of determining a medical knowledge graph based on each of the seed triples and the source metadata includes: The seed triples and the source metadata are merged to obtain the medical knowledge graph, or... Missing content analysis is performed on each of the seed triples. If missing content exists in a seed triple, the missing content is supplemented to obtain an enhanced triple. Each of the seed triples / enhanced triples and the source metadata are merged to obtain the medical knowledge graph.
3. The method for determining a medical knowledge graph according to claim 2, characterized in that, The step involves performing missing content analysis on each of the seed triples. If any seed triple has missing content, the missing content is supplemented to obtain an enhanced triple, including: Based on the main intelligent agent, the missing content analysis is performed on each of the seed triplets. The missing content includes one or more of the following: missing or vague pharmacological mechanism of action, missing or vague dosage description, missing or vague description of applicable population, single or outdated source of evidence, and knowledge conflict. If it is determined that the seed triple has missing content, the professional sub-agent corresponding to the missing content is invoked to supplement the missing content of the seed triple in parallel, and the supplemented seed triple is used as the enhanced triple.
4. The method for determining a medical knowledge graph according to claim 3, characterized in that, The step of any of the aforementioned specialized sub-agents supplementing the missing content of the seed triple includes: The specialized sub-agent submits a structured retrieval request to the retrieval enhancement generation engine based on the missing content of the seed triple, obtains the evidence fragment content and the source of the evidence fragment returned by the retrieval enhancement generation engine, and supplements the missing content of the seed triple based on the evidence fragment content and the source of the evidence fragment.
5. The method for determining a medical knowledge graph according to claim 4, characterized in that, Also includes: Based on the prior confidence of the enhanced triple, determine the Bayesian prior confidence of the enhanced triple; Based on the source of the evidence fragment, the Bayesian prior confidence is incrementally updated to determine the mean and variance of the Bayesian posterior confidence. Based on the mean and variance of the Bayesian posterior confidence, it is determined whether the enhanced triplet can be written into the medical knowledge graph.
6. The method for determining a medical knowledge graph according to claim 2, characterized in that, The process of merging the seed triples and the source metadata to obtain the medical knowledge graph includes: Each seed triple is subjected to graph consistency verification and / or counterfactual testing. The seed triples that pass the verification and / or testing, along with their corresponding source metadata, are merged to obtain the medical knowledge graph, or... The process of merging the seed triples / enhanced triples and the source metadata to obtain the medical knowledge graph includes: Each seed triple / enhanced triple is subjected to graph consistency verification and / or counterfactual testing. The seed triples / enhanced triples that pass the verification and / or testing, along with the corresponding source metadata, are merged to obtain the medical knowledge graph.
7. A medical knowledge graph determination device, characterized in that, include: The first determining module is used to parse and extract knowledge from medical instructions and / or medical guidelines, and determine seed triples of multiple extracted knowledge and source metadata bound to the seed triples, wherein the source metadata includes the fingerprints of the medical instructions and / or the medical guidelines respectively. The second determining module is used to determine the medical knowledge graph based on each of the seed triples and the source metadata.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the medical knowledge graph determination method as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the medical knowledge graph determination method as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the medical knowledge graph determination method as described in any one of claims 1 to 6.