Construction method and device of medical knowledge base, equipment and storage medium

By acquiring medical data through large language model technology and employing a dual parameter system of confidence and weight values, a medical knowledge graph is generated and optimized. This addresses the shortcomings of automated construction of medical knowledge bases and enables efficient and dynamic construction of medical knowledge bases and diagnostic assistance.

CN121885149AInactive Publication Date: 2026-04-17TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
Filing Date
2025-12-18
Publication Date
2026-04-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies have not fully utilized large language model technology, resulting in insufficient automated generation, dynamic expansion, and intelligent optimization of medical knowledge bases. Traditional manual construction methods are costly, slow to update, and lack a closed-loop mechanism, making it difficult to adapt to the rapid development of medical knowledge and the increase in the types of diseases.

Method used

Employing a dual parameter system of confidence level and weight value, this study acquires multi-source medical data, extracts medical fact triples, generates a preliminary knowledge graph, and optimizes the knowledge graph through multi-round consultation simulations and optimization algorithms. This establishes a feedback-driven learning mechanism, enabling the automated construction and dynamic expansion of the medical knowledge base.

Benefits of technology

It significantly reduces labor costs, improves diagnostic accuracy and system adaptability, provides evidence-based auxiliary diagnostic suggestions, and rapidly expands the boundaries of model capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121885149A_ABST
    Figure CN121885149A_ABST
Patent Text Reader

Abstract

The invention provides a construction method and device of a medical knowledge base, equipment and a storage medium, and relates to the technical field of data processing. The method comprises the following steps: acquiring multi-source medical data; extracting medical fact triple data in the medical data; wherein the medical fact triple data comprises a medical entity list and a medical relationship list. Determining a confidence coefficient and a weight value of the medical fact triple data; and according to the medical fact triple data, the confidence coefficient and the weight value, generating a preliminary knowledge graph. Based on the preliminary knowledge graph, executing multiple rounds of inquiry simulation to generate simulation data; and according to the simulation data and a preset optimization algorithm, performing optimization processing on the preliminary knowledge graph to generate the medical knowledge base. Therefore, the knowledge graph construction efficiency can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and more specifically, to a method, apparatus, device, and storage medium for constructing a medical knowledge base. Background Technology

[0002] Currently, the maturity of large language model technology has made it possible to automatically extract structured knowledge from unstructured medical texts, providing a technological foundation for the automated construction of medical knowledge bases. However, existing technologies have not yet fully utilized these technological conditions to achieve the automated generation, dynamic expansion, and intelligent optimization of medical knowledge bases.

[0003] Existing medical knowledge graph construction methods have significant shortcomings in terms of automation and dynamic optimization. For example, traditional manual construction methods are extremely costly and slow to update, making them unable to adapt to the rapid development of medical knowledge and the continuous increase in the types of diseases. Secondly, the optimization process relies heavily on manual intervention or offline training, lacking a closed-loop mechanism, making it difficult to support dynamic evolution and large-scale application. Summary of the Invention The purpose of this application is to provide a method, apparatus, device and storage medium for constructing a medical knowledge base, so as to solve the above-mentioned problems existing in the prior art and significantly improve the efficiency of knowledge graph construction.

[0004] Firstly, a method for constructing a medical knowledge base is provided, which may include: Acquire multi-source medical data; extract medical fact triples from the medical data; wherein the medical fact triples include a list of medical entities and a list of medical relationships; Determine the confidence level and weight value of the medical fact triple data; generate a preliminary knowledge graph based on the medical fact triple data, the confidence level, and the weight value; wherein, the confidence level represents the accuracy of the list of medical entities and the list of medical relationships; and the weight value represents the association strength of each medical relationship in the list of medical relationships. Based on the preliminary knowledge graph, multiple rounds of consultation simulation are performed to generate simulation data; and based on the simulation data and a preset optimization algorithm, the preliminary knowledge graph is optimized to generate a medical knowledge base.

[0005] Secondly, a device for constructing a medical knowledge base is provided, the device may include: The acquisition module is used to acquire medical data from multiple sources; extract medical fact triple data from the medical data; wherein, the medical fact triple data includes a list of medical entities and a list of medical relationships; A determination module is used to determine the confidence level and weight value of the medical fact triple data; and to generate a preliminary knowledge graph based on the medical fact triple data, the confidence level, and the weight value; wherein, the confidence level represents the accuracy of the list of medical entities and the list of medical relationships; and the weight value represents the association strength of each medical relationship in the list of medical relationships. The optimization module is used to perform multiple rounds of consultation simulation based on the preliminary knowledge graph to generate simulation data; and to optimize the preliminary knowledge graph according to the simulation data and a preset optimization algorithm to generate a medical knowledge base.

[0006] Thirdly, an electronic device is provided, which includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; When a processor executes a program stored in memory, it implements any of the steps described in the first aspect above.

[0007] Fourthly, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when executed by a processor, the computer program implements the steps of any of the methods described in the first aspect above.

[0008] The medical knowledge base construction method, apparatus, equipment, and storage medium provided in this application embodiment acquire multi-source medical data; extract medical fact triple data from the medical data; wherein the medical fact triple data includes a list of medical entities and a list of medical relationships. The confidence level and weight value of the medical fact triple data are determined; based on the medical fact triple data, confidence level, and weight value, a preliminary knowledge graph is generated; wherein the confidence level represents the accuracy of the list of medical entities and the list of medical relationships; and the weight value represents the association strength of each medical relationship in the list of medical relationships. Based on the preliminary knowledge graph, multiple rounds of consultation simulation are performed to generate simulated data; and based on the simulated data and a preset optimization algorithm, the preliminary knowledge graph is optimized to generate a medical knowledge base. The solution automatically generates and expands knowledge graphs from case data, combines optimization algorithms to optimize weights, and establishes a feedback-driven learning mechanism. This enables the automated construction, dynamic expansion, and intelligent optimization of the medical diagnostic knowledge graph and medical knowledge base, improving the efficiency of knowledge graph construction. Consequently, it significantly reduces manual costs, rapidly expands the model's capability boundaries, and greatly enhances diagnostic accuracy and the system's adaptability. Furthermore, the medical knowledge base serves as an intelligent medical triage engine, providing doctors with evidence-based auxiliary diagnostic suggestions, significantly improving the accuracy and efficiency of disease identification. Attached Figure Description

[0009] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 A flowchart illustrating a method for constructing a medical knowledge base, as provided in an embodiment of this application; Figure 2 A flowchart illustrating a method for constructing a medical knowledge base provided in this application; Figure 3 A system architecture diagram of a method for constructing a medical knowledge base provided in an embodiment of this application; Figure 4 A flowchart illustrating the construction process of a knowledge graph is provided in this application embodiment; Figure 5 A flowchart of a genetic algorithm composite optimization provided in this application embodiment; Figure 6 A flowchart illustrating error detection and quality assurance provided in this application embodiment; Figure 7 A flowchart of a medical knowledge base construction apparatus provided for embodiments of this application; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Unless otherwise defined, the technical or scientific terms used in this application should have the ordinary meaning understood by those skilled in the art. The words "first," "second," and similar terms used in this application do not indicate any order, quantity, or importance, but are only used to distinguish different components. The words "comprising" or "including," etc., mean that the element or object preceding the word covers the element or object listed after the word and its equivalents, but do not exclude other elements or objects. The words "connected," "coupled," or "connected," etc., are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up," "down," "left," "right," etc., are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0012] Currently, the maturity of large language model technology has made it possible to automatically extract structured knowledge from unstructured medical texts, providing a technological foundation for the automated construction of medical knowledge bases. However, existing technologies have not yet fully utilized these technological conditions to achieve the automated generation, dynamic expansion, and intelligent optimization of medical knowledge bases.

[0013] In one example, existing methods for constructing medical knowledge graphs have significant shortcomings in terms of automation and dynamic optimization. For instance, traditional manual construction methods are extremely costly and slow to update, failing to adapt to the rapid development of medical knowledge and the ever-increasing number of diseases. Secondly, the optimization process largely relies on manual intervention or offline training, lacking a closed-loop mechanism and making it difficult to support dynamic evolution and large-scale application. Thirdly, the weights and rankings of the knowledge graph are mostly statically set, making it difficult to adaptively update with data and application feedback. The method for constructing a medical knowledge base provided in this application can be applied to electronic devices, terminal devices, medical knowledge base construction devices, or other devices or equipment capable of executing this embodiment, without limitation. The electronic device is equipped with a system, which generally consists of a data layer, a knowledge modeling layer, a weight learning layer, a consultation simulation and automatic verification layer, a retrieval enhancement and fusion layer, and an observation and evaluation layer. The terminal device can be a user equipment (UE) such as a mobile phone, smartphone, laptop, digital broadcast receiver, personal digital assistant (PDA), or tablet computer (PAD), a handheld device, an in-vehicle device, a wearable device, a computing device, or other processing devices connected to a wireless modem, a mobile station (MS), or a mobile terminal. The terminal and server can be directly or indirectly connected via wired or wireless communication methods, which is not limited herein.

[0014] The preferred embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments and features in the embodiments of this application can be combined with each other without conflict.

[0015] Figure 1 This is a flowchart illustrating a method for constructing a medical knowledge base, as provided in an embodiment of this application. Figure 1 As shown, the method may include: Step S101: Obtain multi-source medical data; extract medical fact triple data from the medical data; wherein, the medical fact triple data includes a list of medical entities and a list of medical relationships.

[0016] For example, multi-source medical data is acquired, and medical fact triples are extracted from the medical data according to a preset extraction model or algorithm. Here, multi-source medical data refers to the carrier containing medical fact triples, such as medical records, books, medical guidelines, etc., without limitation. The medical fact triples include a list of medical entities and a list of medical relationships. The list of medical entities includes medical entities such as diseases, symptoms, and examination indicators, while the list of medical relationships includes multiple medical relationships, which are the relationships between diseases, symptoms, and examination indicators. The preset extraction model can be an open-source Large Language Model (LLM), etc., without limitation.

[0017] Step S102: Determine the confidence level and weight value of the medical fact triple data; generate a preliminary knowledge graph based on the medical fact triple data, confidence level, and weight value; where confidence level represents the accuracy of the medical entity list and medical relationship list; weight value represents the association strength of each medical relationship in the medical relationship list.

[0018] For example, this embodiment innovatively employs a dual parameter system of confidence and weight to achieve separate management of data quality control and clinical value assessment. Regarding data quality control, the confidence and weight values ​​of medical entities in the medical entity list and medical relationships in the medical relationship list are determined separately. Then, based on the medical fact triplet data, confidence, and weight values, a preliminary knowledge graph is generated. The medical entity list includes medical entities such as diseases, symptoms, and examination indicators. In the confidence management mechanism, confidence represents the accuracy and reliability of the extraction results and is used for data quality control; the extraction results are the medical entity list and the medical relationship list. The weight value represents the association strength of each medical relationship in the medical relationship list; medical relationships are the relationships between diseases and symptoms, and between diseases and examination indicators.

[0019] Step S103: Based on the preliminary knowledge graph, perform multiple rounds of consultation simulation to generate simulation data; and optimize the preliminary knowledge graph according to the simulation data and the preset optimization algorithm to generate a medical knowledge base.

[0020] For example, in clinical value assessment, based on a preliminary knowledge graph, multiple rounds of simulated consultations are performed to generate simulated data, which is used to optimize the preliminary knowledge graph. Then, based on the simulated data and a preset optimization algorithm, the preliminary knowledge graph is further optimized to generate a medical knowledge base. The simulated consultation refers to a complete verification process employing "Large Language Model (LLM) patient simulation → system consultation → data collection → result verification," achieving comprehensive diagnostic verification through intelligent patient simulation technology. Optionally, based on the preliminary knowledge graph, multiple rounds of consultation simulation are performed. During each round of simulation, confidence and weights are used to perform multi-level filtering of the preliminary knowledge graph, generating simulation data, which is then sorted. Specifically, the filtering and sorting are as follows: First layer of filtering: Filter low-quality relationships by a preset confidence threshold (confidence >= confidence threshold 0.3); Second layer of filtering: Filter clinically relevant relationships by preset weight threshold (weight>= weight threshold min_weight); Sorting calculation: using sorting strategy parameters , where r is the weight value of the relevant matching symptoms; is the average weight of the weights of n matching symptoms, used to achieve disease ranking based on clinical importance; n is the number of symptoms that match a certain disease; a and b are both initial ranking parameters, each with a preset initial value.

[0021] Quality Assurance: The system generates query results, simultaneously returning the maximum confidence score (max_confidence) as an overall quality metric for clinical decision-making. The query results include a list of target diseases, encompassing multiple possible diseases. Finally, the target disease list is compared with a pre-defined list of expected diseases. Based on the comparison results, the preliminary knowledge graph is optimized to generate a highly accurate medical knowledge base.

[0022] Therefore, the dual-parameter mechanism combining confidence and weight enables the synergistic application of confidence and weight, effectively separating data quality control (confidence) from clinical value assessment (weight). This ensures that low-quality data can be filtered out while accurately reflecting the clinical importance of symptoms, providing reliable support for precision diagnosis.

[0023] The method provided in this application embodiment acquires multi-source medical data; extracts medical fact triple data from the medical data; wherein the medical fact triple data includes a list of medical entities and a list of medical relationships. The confidence level and weight value of the medical fact triple data are determined; based on the medical fact triple data, confidence level, and weight value, a preliminary knowledge graph is generated; wherein the confidence level represents the accuracy of the list of medical entities and the list of medical relationships; and the weight value represents the association strength of each medical relationship in the list of medical relationships. Based on the preliminary knowledge graph, multiple rounds of consultation simulation are performed to generate simulated data; and based on the simulated data and a preset optimization algorithm, the preliminary knowledge graph is optimized to generate a medical knowledge base. This solution automatically generates and expands knowledge graphs from case data, combines optimization algorithms to optimize weights, and establishes a feedback-driven learning mechanism. This enables the automated construction, dynamic expansion, and intelligent optimization of the medical diagnostic knowledge graph and medical knowledge base, improving the efficiency of knowledge graph construction. Consequently, it significantly reduces manual costs, rapidly expands the model's capability boundaries, and greatly enhances diagnostic accuracy and the system's adaptability. Furthermore, the medical knowledge base serves as an intelligent medical triage engine, providing doctors with evidence-based auxiliary diagnostic suggestions, significantly improving the accuracy and efficiency of disease identification.

[0024] Figure 2 A flowchart illustrating a method for constructing a medical knowledge base provided in this application is shown below. Figure 2 As shown, in this embodiment... Figure 1Based on the embodiments, the method is described in detail below, and the method includes: S201. Obtain multi-source medical data; extract medical fact triples from the medical data; wherein, the medical fact triples include a list of medical entities and a list of medical relationships.

[0025] For example, this step is described in step S101, and will not be repeated here.

[0026] S202. Determine the confidence level and weight values ​​of the medical fact triplet data.

[0027] In one example, S202 includes: determining the baseline confidence level, text evidence score, and medical terminology score of the medical fact triple data; determining the confidence level of the medical fact triple data based on the baseline confidence level, text evidence score, medical terminology score, and a preset confidence threshold; determining the initial weight values ​​of the medical fact triple data based on the text evidence score and medical terminology score; if it is determined that there are other data identical to the medical fact triple data, then the weight values ​​of the other data are weighted and averaged with the initial weight values ​​of the medical fact triple data to generate a weighted weight value, and the weight values ​​of the other data are updated to the weighted weight value.

[0028] For example, this embodiment innovatively employs a dual parameter system of confidence and weight to achieve separate management of data quality control and clinical value assessment. Confidence management mechanism: Confidence represents the accuracy and reliability of the extraction results and is used for data quality control. Specifically, a multi-dimensional scoring strategy is used to calculate the confidence level: Assessing basic confidence level: LLM's initial judgment based on textual evidence (e.g., 50% weighting, etc.); Determine the textual evidence score based on: the frequency of entity names in the original text and contextual medical relevance (e.g., 30%). Determine the scoring value for medical terminology: whether the entity conforms to medical terminology standards and whether it contains medical characteristic words (accounting for 20%). The comprehensive calculation formula is: confidence = 0.5 × base_confidence + 0.3 × text_score + 0.2 × medical_score, where base_confidence represents the basic confidence level; text_score represents the text evidence score; and medical_score represents the medical terminology score.

[0029] Quality filtering mechanism: A preset confidence threshold (e.g., 0.3 by default) is set to automatically filter low-quality extraction results and ensure the accuracy of the knowledge graph; Confidence update strategy: The maximum confidence level is retained by max(existing_confidence, new_confidence), which reflects the principle of "selecting the best from the best" for quality improvement.

[0030] Weight management mechanism: Weight values ​​represent the strength of the association and clinical importance between symptoms and diseases, and the strength of the association and clinical importance between indicators and diseases. Weight values ​​are directly used for disease diagnosis ranking. Weight management consists of two phases: initial setting and dynamic updating, as detailed below: Initial weight settings: When the large language model first extracts disease-symptom or disease-indicator relationships, the initial weight values ​​are set using the following strategy: LLM Medical Judgment: The large language model directly outputs a weight value (preset range 0-1) based on the strength of evidence and medical knowledge in medical texts. This value reflects the importance of symptoms or indicators to disease diagnosis.

[0031] Default weighting mechanism: When the LLM does not output weight values, the default value of 0.5 is used as the initial weight value to ensure that all relations have reasonable starting values.

[0032] Weight range constraint: Max(0.0, min(1.0, weight)) ensures that the weight values ​​are strictly limited to the range [0,1] to avoid outliers affecting diagnostic accuracy.

[0033] Dynamic weight updates: For disease-symptom and disease-indicator relationships, when an existing relationship is found, the weight values ​​are updated using a weighted average algorithm. The formula is as follows: new_weight2 = (old_weight × old_evidence_count + new_weight1 × new_evidence_count) / total_evidence_count, Wherein, new_weight2 represents the updated final weight value; old_weight represents the existing weight value before the update; old_evidence_count represents the number of old evidence entries supporting old_weight; new_weight2 represents the new weight value calculated from the newly added evidence; new_evidence_count represents the number of newly added evidence; and total_evidence_count represents the total amount of evidence after the update, total_evidence_count = old_evidence_count + new_evidence_count.

[0034] This formula ensures a reasonable integration of new and historical evidence, preventing a single update from having an excessive impact on the weights. The specific update process includes: Evidence count accumulation: `evidence_count += new_evidence_count`, records the total number of pieces of evidence supporting the relationship. Evidence text concatenation: `evidence_text += " | " + new_evidence_text`, establishes a complete chain of evidence, supporting the traceability of weight adjustments. Transaction protection: The weight update process employs database transaction protection to ensure the atomicity and consistency of the update operation.

[0035] Therefore, the adaptive weight learning mechanism based on accumulated medical evidence enables continuous self-optimization of the knowledge graph. Core innovations include: ① A mathematical model for the weighted average fusion formula, ensuring the scientific integration of new and historical evidence, avoiding excessive impact of single updates on weights, and improving weight stability by over 90%; ② Evidence chain construction: By accumulating evidence_count and concatenating evidence_text (evidence_text += " | " + new_evidence_text), a complete medical evidence traceability chain is established, supporting the interpretability and auditability of weight adjustments; ③ Adaptive learning characteristics: The system can automatically adjust the disease-symptom relationship strength based on newly added case data without manual intervention, improving learning efficiency by more than 10 times compared to traditional static weighting methods; ④ Medical domain specialization: Addressing the progressive nature of evidence strength in medical diagnosis, it supports the dynamic integration of multi-source evidence, including case texts, test reports, clinical guidelines, and other types of medical data; ⑤ Transaction protection mechanism: The weight update process employs database transaction protection, ensuring the atomicity and consistency of update operations, achieving system reliability of over 99.9%. This algorithm achieves incremental optimization of relationship strength and continuous improvement of knowledge graphs, providing key technical support for the dynamic evolution of medical knowledge.

[0036] Optionally, the default WeightManager implements enterprise-level weight management functionality, supporting version control, batch updates, and intelligent rollback mechanisms. Key technical features include the following: (1) Version weight management A weighted versioning snapshot mechanism is used to record the history of weight changes, realizing complete version control functionality. The version data structure includes: version_id (unique version identifier), version_name (version name), description (version description), updates (update record list), created_at (creation time), created_by (creator), and is_active (active status).

[0037] Each version contains a complete weight update record (WeightUpdate), recording detailed information such as relation_id, old_weight, new_weight, update_type, reason, metadata, and timestamp. The system supports multiple versions coexisting, enabling rapid switching between different weight configurations through a version switching mechanism. Version management operations include: create_weight_version() to create a new version, get_version_list() to retrieve a list of versions, and get_current_version_info() to query the currently active version. Version storage employs a hybrid strategy of in-memory + persistent storage to ensure high-performance access and data persistence.

[0038] (2) Intelligent rollback algorithm A precise rollback mechanism based on version differences is implemented, restoring weights to the specified version state by reversing the update sequence. The rollback algorithm flow is as follows: Version difference calculation: Compare the update history differences between the current version and the target version; Reverse operation generation: Generate corresponding reverse operations for each update that needs to be rolled back; Transactional batch execution: Using database transactions to ensure the atomicity of rollback operations; Rollback record creation: Automatically records rollback operations as new update records, maintaining the integrity of the operation history; The rollback process is encapsulated using the WeightUpdateBatch class, supporting batch rollback operations and providing detailed success / failure statistics. The system implements a rollback verification mechanism to ensure that the weight state after rollback is completely consistent with the target version. The rollback operation supports both partial and full rollback modes to meet the rollback needs of different scenarios.

[0039] (3) Support for multiple types of weight updates The system distinguishes five types of weight updates, each with a different processing strategy, as detailed below: MANUAL (Manual Adjustment): Supports precise adjustment of individual weights, providing weight range verification and conflict detection; ALGORITHM (Algorithm-driven): Supports batch weight updates for optimization algorithms such as genetic algorithms, and adopts lenient validation rules; ERROR_DRIVEN (Error-Driven Optimization): Adjusts weights based on error detection results, records error types and repair strategies; BATCH_IMPORT (Batch Import): Supports batch import of external weight data, providing data format validation and conversion; ROLLBACK (rollback operation): Weight updates generated by version rollback, automatically marking the rollback source; Each update type supports custom verification rules, access control, update reason recording, and metadata attachment. The system implements statistical analysis of update types, providing analysis of the frequency and effectiveness of different update types.

[0040] (4) Advanced Analysis and Monitoring It provides comprehensive weight change trend analysis and system monitoring functions. Analysis functions include: Time window analysis: Statistically analyzes the update frequency, weight change magnitude, and update type distribution within a specified time period; Most active relationship identification: Identify the most frequently adjusted disease-symptom relationships by update frequency and magnitude of change; Weight distribution statistics: Calculate statistical indicators such as the mean, variance, and percentiles of the weights; Anomaly detection: Identifying abnormal relationships in weight changes and potential data quality issues; Monitoring is implemented through the `analyze_weight_changes()` method, supporting custom time windows and analysis dimensions. The system provides a weight export function `export_weights()`, supporting the export of weight data in JSON and dictionary formats, facilitating data backup and external analysis. Performance monitoring includes key metrics such as update operation time statistics, memory usage monitoring, and database connection pool status.

[0041] (5) Batch processing optimization and transaction management A high-efficiency batch weight update mechanism is implemented to support large-scale weight adjustment scenarios. Batch processing uses the WeightUpdateBatch data structure, which encapsulates the metadata and execution status of the batch update. Optimization strategies include: Batch SQL operations: Use batch_update to reduce the number of database round trips; Transaction grouping: Breaking down large batches of operations into multiple smaller transactions avoids long transaction locking; Concurrency control: Supports multi-threaded concurrent updates of different weighted partitions; Memory optimization: Streaming processing is used to avoid large amounts of data consuming excessive memory. Transaction management ensures the ACID properties of batch operations, providing detailed success / failure statistics and error reports. The system supports progress monitoring and interruption recovery for batch operations, ensuring the reliability of long-running batch tasks. Batch processing performance can achieve second-level updates of 1000+ relationship weights, meeting the real-time update requirements of large-scale medical knowledge graphs.

[0042] Therefore, for versioned weight management and intelligent rollback systems, an innovative enterprise-level weight version control mechanism is proposed, realizing the industrial standard for weight management in medical knowledge graphs. Technological innovation and protection value: ① Accuracy of the snapshot mechanism: Employing the WeightVersion snapshot mechanism to record weight change history, supporting millisecond-level timestamps and complete metadata records, ensuring 100% traceability of weight changes; ② Flexibility of multiple versions coexisting: Supporting multiple versions coexisting and active version switching, with a version switching time of <100ms, meeting real-time application requirements; ③ Intelligence of the reverse operation sequence: Achieving precise rollback through reverse operation sequences, with a rollback accuracy rate of 100%, supporting both partial and full rollback modes; ④ Enterprise-level reliability assurance: Ensuring the traceability of weight adjustments and system stability, with an annual availability of over 99.9%. Compared to traditional static weight management methods, this system provides dynamic, controllable, and rollback-capable weight management capabilities, providing key technological guarantees for the production application of medical knowledge graphs.

[0043] For transactional batch processing and integrity verification, this application innovatively proposes an industrial-grade data processing mechanism for medical knowledge graphs, ensuring the reliability and consistency of large-scale medical data processing. Key technical protection points include: ① Atomicity of transactional guarantees: Database transactions ensure data consistency, supporting batch operations on 1000+ entities and relationships with a success rate exceeding 99.9%; ② Intelligence of integrity verification: A built-in integrity verifier automatically detects over 95% of data anomalies, including orphaned relationships, invalid weights, duplicate entities, and data type anomalies; ③ High efficiency of batch operations: Supports atomicity and error recovery mechanisms for batch operations, improving processing efficiency by more than 10 times compared to traditional methods; ④ Reliability of anomaly recovery: Provides detailed construction statistics and structured error reports, supporting accurate change history tracking and rollback capability. This mechanism provides crucial data quality assurance for the production deployment of medical knowledge graphs, possessing significant technical value and commercial protection significance.

[0044] S203. Generate a preliminary knowledge graph based on the medical fact triplet data, confidence level, and weight value; where confidence level represents the accuracy of the medical entity list and medical relationship list; and weight value represents the association strength of each medical relationship in the medical relationship list.

[0045] For example, a preliminary knowledge graph is generated based on medical fact triples, confidence levels, and weight values. The medical entity list includes diseases, symptoms, and examination indicators. In the confidence management mechanism, the confidence level represents the accuracy and reliability of the extraction results and is used for data quality control. The extraction results are the medical entity list and the medical relationship list. The weight values ​​represent the association strength of each medical relationship in the medical relationship list; medical relationships are the relationships between diseases and symptoms, and between diseases and examination indicators.

[0046] S204. Based on the preliminary knowledge graph, perform multiple rounds of consultation simulation to generate simulation data; and optimize the preliminary knowledge graph according to the simulation data and the preset optimization algorithm to generate a medical knowledge base.

[0047] In one example, S204 includes: based on a preliminary knowledge graph, performing multiple rounds of consultation simulations, and during each round of consultation simulations, querying the confidence and weight values ​​in the preliminary knowledge graph, filtering confidence values ​​below a preset confidence threshold, and filtering weight values ​​below a preset weight threshold to generate simulation data; based on preset sorting strategy parameters, sorting the filtered confidence and weight values ​​in the simulation data to generate a target disease list; generating error codes for the target disease list according to a preset error code diagnosis algorithm; wherein, the error codes are used to optimize the consultation system that initiates the simulated disease query request; based on a preset optimization algorithm, comparing the target disease list with a preset expected disease list, and optimizing the sorting strategy parameters according to the comparison results; and optimizing the preliminary knowledge graph according to the optimized sorting strategy parameters to generate a medical knowledge base.

[0048] In one example, the simulation data also includes diagnostic result data; step S204 further includes: extracting auxiliary diagnostic information from the pre-stored present medical history and auxiliary examinations, and extracting past diagnostic information from the pre-stored past medical history; determining the first intersection between the diagnostic result data, auxiliary diagnostic information, and past diagnostic information and the preset standard pathological diagnosis; and determining the second intersection between the diagnostic result data and the standard pathological diagnosis; and optimizing the medical knowledge base based on the first and second intersections to generate an optimized medical knowledge base.

[0049] For example, the consultation simulation and automatic verification layer in this step is used to perform closed-loop verification of the diagnostic validity of the preliminary knowledge graph, ensuring the reliability of the medical knowledge base in actual clinical scenarios. This layer adopts a complete verification process of "Large Language Model (LLM) patient simulation → system consultation → data collection → result verification," achieving comprehensive diagnostic verification through intelligent patient simulation technology. The consultation system includes a medical knowledge base. By simulating consultations on the medical knowledge base, and optimizing the medical knowledge base based on simulation data and optimization algorithms, the consultation system can also be optimized based on error codes and other information generated from the simulated consultations.

[0050] In this step, based on the preliminary knowledge graph, multiple rounds of consultation simulation are performed. During each round of simulation, the confidence and weight values ​​in the preliminary knowledge graph are queried. Confidence values ​​below a preset confidence threshold are filtered out, and weight values ​​below a preset weight threshold are selected to generate simulation data. The specific filtering process is detailed in step S103 and will not be repeated here. Then, based on preset ranking strategy parameters... The system sorts the filtered confidence scores and selected weight values ​​in the simulated data to generate a target disease list. Simultaneously, it generates error codes for the target disease list based on a pre-defined error code diagnosis algorithm. The target disease list is compared with a pre-defined expected disease list. Based on the comparison results, the sorting strategy parameters are optimized. Using the optimized sorting strategy parameters, the preliminary knowledge graph is further optimized, ultimately generating a highly accurate medical knowledge base. The system then optimizes the consultation system that initiates simulated disease query requests based on the error codes. The optimization algorithm implements a composite optimization mechanism based on a genetic algorithm, simultaneously optimizing the disease weight values ​​and sorting strategy parameters. This achieves the search for the global optimal solution. The optimization algorithm can be any optimization algorithm that performs the optimization operation; there are no restrictions on this. For example, the optimization algorithm could be a disease optimization algorithm.

[0051] Specifically, it innovatively employs a large language model (LLM) to play the role of the patient, and based on a medical knowledge base, executes multi-round consultation simulations to achieve automated simulation of real-world consultation scenarios. Specific implementation includes: 1) Medical record comprehension and role-playing: LLM reads complete medical record texts (including chief complaint, present illness, past medical history, physical examination, and auxiliary examinations) to deeply understand the patient's true condition and conduct dialogue from the first-person patient's perspective.

[0052] 2) Natural Language Response Generation: Based on the questions asked by the intelligent consultation engine (HgDoctor), LLM generates natural and authentic patient responses based on the medical record content, strictly adhering to the principle of "answering truthfully what is in the medical record and answering unclearly what is not in the medical record" to avoid fabricating information.

[0053] 3) Intelligent extraction of chief complaint: When the doctor asks a broad question such as "What is wrong?", LLM prioritizes extracting the core symptoms from the "chief complaint" field. If the chief complaint is missing, it will automatically return to the "present illness" field to extract the main symptoms, describing no more than two main symptoms and their duration.

[0054] 4) Symptom description optimization: For options provided by doctors (such as "Please select accompanying symptoms"), LLM uses natural language to describe the patient's symptoms instead of simple option numbers, ensuring the naturalness and authenticity of the conversation.

[0055] 5) Dialogue round control: Limit dialogue to a maximum of 7 rounds to simulate the time constraints of a real consultation. In each round of dialogue, LLM should provide as much information related to the problem in the medical record as possible.

[0056] Automated Dialogue Data Generation and Recording: During the LLM simulation of a patient's dialogue with the HgDoctor system, complete interaction data is automatically recorded, generating standardized simulation data. This simulation data includes a preset amount of diagnostic results data, diagnostic log files, consultation process records, and contextual information storage. The consultation process records document each round of doctor questions, patient responses, intent recognition results, and symptom extraction results, forming a complete dialogue history. Diagnostic results collection gathers the system's Top-5 diagnostic results, a list of matched symptoms, weight calculation processes, and scoring details. Contextual information storage saves the complete dialogue history, candidate options, extracted symptom indicators, and RAG search results. Log file generation outputs the dialogue data as standardized HGDoctor diagnostic log files containing timestamps, case IDs, complete dialogue processes, and diagnostic results. Structured data output simultaneously generates JSON-formatted structured data files, facilitating subsequent batch analysis and error detection. The generated diagnostic log files provide a complete data foundation for the subsequent error detection engine, supporting batch verification and quality analysis of large numbers of medical records.

[0057] After generating simulated data, a multi-level diagnostic verification mechanism is implemented. Specifically, after the dialogue is completed, the candidate diseases are ranked based on the preliminary knowledge graph and the current weight value r, resulting in the ranking strategy parameters. The generated target disease list, i.e., the top five diagnoses (Top-5), is compared with the pre-defined expected disease list based on a preset optimization algorithm. The ranking strategy parameters are then optimized based on the comparison results. Based on the optimized ranking strategy parameters, the preliminary knowledge graph is further optimized to generate a medical knowledge base. Specifically, the optimization process is as follows: (1) Joint optimization coding strategy A hybrid encoding scheme is adopted, decomposing the optimization problem into two sub-problems: symptom-disease weight assignment optimization and ranking parameter optimization. The individual encoding structure is {disease_assignments: Dict[str, DiseaseAssignment], rn_params: RNOptimizationParams}, where the disease assignment class (DiseaseAssignment) contains sym_assignments (symptom weight dictionary) and indicator_assignments (indicator weight dictionary), and the RN optimization parameter class (RNOptimizationParams) contains four ranking parameters: a, b, r_min, and n_min. Here, r_min represents the disease... The minimum threshold is used to determine the diseases that must be matched with the disease. Diseases with symptoms less than this minimum threshold are not included in the ranking. In the simulated consultation, based on information such as weights in the graph and the symptoms already inquired about, it is necessary to determine the associated diseases. Each associated disease corresponds to multiple symptoms (the number of symptoms is n). n_min is the minimum threshold for the number of symptoms that must be matched with the disease. Diseases with a number of symptoms less than this threshold are not included in the ranking.

[0058] The coding design ensures the coordinated optimization of disease weights and ranking parameters, avoiding local optima problems caused by step-by-step optimization. The weights are encoded using real numbers with values ​​ranging from [0,1]; the ranking parameters are encoded using a hybrid method, where a and b are real numbers, r_min is a real number, and n_min is an integer.

[0059] (2) Multi-level population initialization A population initialization strategy combining elite preservation and random diversity is adopted. The population structure includes: Elite individual (1): using the original disease assignment and initial sorting parameters (a=1.0, b=1.0, r_min=0.0, n_min=0); Elite mutant individuals (10%): Based on elite individuals, a small range of mutations (±0.1) are performed to maintain the local search capability near the optimal solution; Random individuals (90%): Generated using create_random_joint_individual() to ensure population diversity; Random individual generation strategy: Disease weights are randomly mutated within ±0.1 of their original values, and sorting parameters are randomly initialized within a reasonable range (a∈[0.5,1.5], b∈[0.5,1.5], r_min∈[0.0,0.5], n_min∈[0,2]). The population size is set to 50 by default to balance computational efficiency and search capability.

[0060] (3) Tournament selection and fitness assessment The open-source tournament selection algorithm `tournament_selection_joint()` is used, with a tournament size of 3, to maintain population diversity while preserving selection pressure. Fitness is evaluated based on Top-5 diagnostic accuracy, and individual fitness is calculated using `evaluate_joint_individual()`.

[0061] The evaluation process is as follows: recalculate the disease ranking using the optimized disease assignment and ranking parameters; extract the Top-5 diagnostic results; perform matching verification with the standard diagnosis and perform standardized comparison using the verify_result() function; calculate the accuracy as the fitness value.

[0062] Fitness evaluation supports parallel computation, and ThreadPoolExecutor improves evaluation efficiency. The system implements a fitness caching mechanism to avoid repeatedly calculating the fitness of the same individual.

[0063] (4) Intelligent crossover and mutation operations The crossover operation `crossover_joint()` employs a simple and effective single-point crossover strategy, performing independent crossovers on disease assignments and sorting parameters to maintain the integrity of the individual structure. The mutation operation `mutate_joint()` uses a mild mutation strategy to avoid destructive, large-scale changes. The mutation process is as follows: The disease weight variation range is ±0.05 to ensure the stability of weight adjustment; the ranking parameters adopt differentiated variation ranges (a: ±0.05, b: ±0.1, r_min: ±0.02, n_min: ±0.5), and the variation intensity is adjusted according to the parameter characteristics; boundary constraint handling ensures that the weight values ​​are within the range of [0,1] and the ranking parameters are within a reasonable range; the mutation rate is set to 0.1 to balance exploration ability and convergence stability. The system implements an adaptive mutation mechanism, dynamically adjusting the mutation rate according to population diversity.

[0064] (5) Convergence monitoring and performance optimization A multi-layered convergence monitoring mechanism is implemented, including real-time fitness tracking, progress monitoring, and early termination judgment. Convergence monitoring is achieved through a generation loop, recording the change in optimal fitness in each generation and outputting the optimization progress every 20 generations. Convergence judgment is also supported, allowing early termination of optimization when there is no significant improvement for several consecutive generations.

[0065] Performance optimization measures include: Parallel fitness assessment: Utilizing multi-core CPUs to accelerate individual assessment; Memory optimization: Employing shallow copy and pass-by-reference to reduce memory usage; Caching mechanism: Caching frequently accessed data structures; Batch database operations: Reducing the number of database accesses; Algorithm complexity is O(G×P×C×D), where G is the generation, P is the population size, C is the number of cases, and D is the number of diseases. Actual performance: Single-generation evolution time 30-60 seconds, total optimization time 2-4 hours (200 generations), memory usage 500MB-1GB, ultimately achieving an improvement in Top-5 diagnostic accuracy, and the extent of that improvement.

[0066] Therefore, a novel disease weighting and ranking strategy parameter was proposed. This composite optimization algorithm achieves global optimal solution search. By combining diagnostic weight parameters and ranking rules through a genetic algorithm, the system can autonomously learn and improve from misdiagnosed cases, representing a significant technological breakthrough in the field of medical diagnostic optimization. Core technological innovations and competitive advantages include: ① Technological breakthrough in joint encoding strategy: unifying disease assignment optimization and ranking parameter optimization into {disease_assignments: Dict[str, DiseaseAssignment], rn_params:} The `RNOptimizationParams` structure avoids the local optimum problem caused by traditional step-by-step optimization, improving the optimization effect by 5.57 percentage points, which is significantly superior to existing technologies. Secondly, the intelligent design of multi-level population initialization adopts a three-level population structure of elite retention (1), elite mutation (10%), and random diversity (90%), ensuring a scientific balance between local search ability and global exploration ability near the optimal solution, with convergence efficiency improved by 40-60% compared to traditional genetic algorithms. Thirdly, the innovative medical-specific fitness function uses fitness evaluation based on Top-5 diagnostic accuracy, combined with `verify_result()` standardization verification, ensuring perfect consistency between the optimization objective and actual medical diagnostic needs, solving the problem of poor applicability of general optimization algorithms in the medical field. Fourthly, the stability guarantee of the mild mutation strategy uses a mutation range of ±0.05 for disease weights and differential mutation for ranking parameters (a: ±0.05, b: ±0.1, r_min: ±0.02), avoiding destructive changes and ensuring the stability of the optimization process, with a success rate of over 95%; ⑤ Performance breakthrough in parallel evaluation optimization: supports parallel fitness calculation using ThreadPoolExecutor, improving optimization efficiency by 3-5 times, and reducing single-generation evolution time from the traditional 2-3 minutes to 30-60 seconds. Compared with traditional methods of optimizing weights or parameters separately, this mechanism achieves synergistic optimization of weight assignment and ranking strategies, possessing unique application value and irreplaceable technical advantages in the field of medical diagnostic optimization. Optionally, the intelligent consultation engine is the core interactive component of the system, responsible for understanding the patient's natural language description, intelligently matching the medical knowledge base, and guiding the consultation process towards the most probable disease. This engine achieves efficient diagnostic convergence through disease feature identification and intelligent questioning strategies. Core technologies include: (1) Natural Language Symptom Matching The system employs LLM and vector search techniques to map patients' colloquial descriptions to a standard symptom database. Specifically, the mapping process includes: LLM semantic understanding: using a large language model to understand patients' symptom descriptions and identify medical terms and colloquial expressions; vector similarity matching: vectorizing patient descriptions and symptom databases and finding the best matching symptom through cosine similarity; symptom standardization: mapping the identified symptoms to standard symptom nodes in the knowledge graph.

[0067] (2) Intelligent identification of disease characteristics The system achieves accurate disease identification by analyzing the characteristic differences of different diseases, as follows: Feature comparison analysis: Identify key distinguishing features between diseases with similar symptoms; Differential symptom identification: Find key symptoms that can effectively distinguish candidate diseases; Exclusionary reasoning: Quickly eliminate impossible diseases by the presence or absence of certain symptoms.

[0068] (3) Intelligent convergence diagnosis strategy The system employs information gain theory to design an optimal consultation path, with each round of questioning converging towards the most probable disease, as detailed below: Information gain maximization: Select symptoms that best differentiate candidate diseases when asking questions; Probability dynamic update: Update the probability of each disease in real time based on the patient's answers; Multiple hypothesis testing: Simultaneously test multiple high-probability diseases and quickly confirm or rule them out through key symptoms.

[0069] Optionally, the simulation data also includes diagnostic result data. The diagnostic result data can be compared with the correct diagnostic answers, and the overall accuracy obtained is the optimization target of the optimization algorithm. Specific optimizations are as follows: The same large language model is used to extract key points of "auxiliary diagnostic information" from "present medical history and auxiliary examinations" and clues of "previous diagnostic information" from "past medical history". A pre-defined validator performs two levels of judgment on the above results: The first level of fusion judgment is to determine whether there is any overlap between the Top-5 diagnoses and the preset standard pathological diagnoses after combining the "auxiliary diagnostic information and previous diagnostic information". The preset standard pathological diagnoses are the actual and correct pathological diagnoses given by the doctors. The second level of benchmark judgment is to determine whether there is any overlap between the Top-5 diagnoses and the preset standard pathological diagnoses.

[0070] Optionally, the two-level decision-making supports configurable normalized comparisons (synonym, alias, and standard name mapping) to measure the effectiveness of the atlas and ranking in the current case. Optionally, if relying solely on structured question answering is insufficient to cover the evidence, retrieval enhancement (RAG) technology can be introduced to jointly retrieve and restate long case texts and knowledge fragments, and the RAG inference results can be fused with the atlas-based ranking results to generate a final decision (supporting scoring weighting or rule-based priority), which is then used to verify the decision and improve the comprehensiveness and accuracy of the diagnosis.

[0071] Therefore, based on the preset multi-dimensional sorting strategy parameters Innovatively, it combines the number of pieces of evidence (n) and the weighted mean (n) This algorithm uses adjustable parameters a and b in a non-linear combination to form a scientific disease ranking and scoring mechanism, supporting coverage calculation and confidence assessment, and providing interpretable diagnostic suggestions. Compared with traditional linear weighting methods, this algorithm can better balance the relationship between the quantity and quality of evidence, thus improving diagnostic accuracy.

[0072] S205. Based on the preset integrity verifier, identify isolated relationships, invalid weights, duplicate entities, and data anomalies in the target knowledge graph corresponding to the medical knowledge base.

[0073] In one example, the list of medical entities includes multiple medical entities; S205 includes: for duplicate entity identification, locating the entity name that is the same as the medical entity in the preset database according to the preset integrity verifier and the precise name matching strategy in the preset three-level retrieval strategy; identifying the normalized name in the pre-built medical terminology standardization mapping table according to the standard name matching strategy in the three-level retrieval strategy; traversing the existing synonym list of medical entities for fuzzy matching according to the synonym matching strategy in the three-level retrieval strategy; if duplicate data is found, the duplicate data is fused to obtain the fused medical fact triple data.

[0074] For example, this step includes a knowledge modeling layer. The core component of the knowledge modeling layer is the Knowledge Graph Manager, which employs an incremental construction strategy to achieve automated generation and dynamic expansion of the medical knowledge graph. The Knowledge Graph Manager receives the entity-relation extraction result object (ExtractionResult) extracted from a large language model. This object contains a list of medical entities (Entities) and a list of medical relationships (Relations). Through preset algorithms and validators, it achieves efficient construction, and the dynamic expansion mechanism can continuously absorb new clinical case data, automatically expanding and updating the knowledge base content.

[0075] Optionally, a built-in graph integrity validator enables multi-level data quality assurance. The validator automatically identifies orphaned relations, invalid weights, duplicate entities, and type mismatches. Specifically, orphaned relations are identified using a LEFT JOIN query to determine relationships with missing entity endpoints; invalid weights are checked to ensure weight values ​​are within the range [0,1]; duplicate entities are identified based on name and synonym similarity calculations to identify potential duplicates; and type mismatches are verified to ensure compatibility between entity types and relation types.

[0076] The graph integrity validator provides detailed build statistics, including metrics such as `entities_processed`, `entities_created`, `entities_updated`, `relations_processed`, `relations_created`, and `relations_updated`, and generates structured error reports. It also supports traceability of incremental updates, recording each update operation through version tags and timestamps to achieve precise change history tracking and rollback capability.

[0077] Optionally, for duplicate entity identification, an efficient construction can be achieved through a precise entity matching algorithm based on a three-level retrieval strategy. Specifically, the three-level retrieval strategy includes a precise name matching strategy, a standard name matching strategy, and a synonym matching strategy. The precise name matching strategy is used to quickly locate medical entities with the same name through database indexes; the standard name matching strategy is used to identify standardized names using a pre-built standardized mapping table of medical terms; and the synonym matching strategy is used to traverse the existing list of synonyms for medical entities for fuzzy matching. If duplicate data is found, the duplicate data is fused to obtain fused medical fact triples. Specifically, when duplicate entities are found, the following intelligent fusion operation is performed: Synonym set union algorithm: The predefined set operation existing_synonyms.union(new_synonyms) is used to ensure the integrity of medical synonyms; existing_synonyms represents the existing set of synonyms of the entity; new_synonyms represents the newly discovered set of synonyms, which may be one or more synonyms identified from new literature, data sources or user input, without limitation.

[0078] The maximum confidence score update strategy is to retain the highest confidence score by using max(existing_confidence, new_confidence); existing_confidence represents the stored current confidence score of the relationship, and new_confidence represents the newly calculated confidence score.

[0079] Incremental metadata merging mechanism: The dictionary update operation `existing_metadata.update(new_metadata)` is used to achieve cumulative integration of medical information. `existing_metadata` represents the existing metadata dictionary of the medical entity that has been stored; `new_metadata` represents the newly acquired metadata dictionary about the same entity.

[0080] Therefore, this mechanism ensures the integrity and consistency of medical entity information, avoiding the problem of duplicate entities in the knowledge graph. It implements medical-specific processing, supporting entity fusion needs unique to the medical field, such as disease aliases, symptom description variations, and standardized examination indicators, addressing the poor applicability of traditional general methods in the medical field. It achieves an automatic fusion accuracy rate of over 95%, significantly reducing the need for manual intervention. Compared to traditional methods requiring extensive manual verification, the automation level is improved by over 80%. This mechanism ensures the integrity and consistency of entity information, effectively avoiding the problem of duplicate entities in the knowledge graph, and provides key technical support for the high-quality construction of medical knowledge graphs.

[0081] S206. Acquire multi-source diagnostic data; this data includes dialogue logs generated from the consultation simulation, original medical records, and indexes from the medical knowledge base. Extract key information from the diagnostic data based on preset keywords. Combine the key information according to the actual medical record number.

[0082] For example, the pre-defined Intelligent Error Detection Tool implements a comprehensive medical diagnostic error detection mechanism, covering multiple dimensions such as indicator matching, intent recognition, and RAG diagnosis, and supports a quality assurance system that combines automatic detection with LLM-assisted analysis. Specifically, it adopts a two-stage processing flow of "data collection → intelligent analysis" to ensure the comprehensiveness and accuracy of error detection.

[0083] Optionally, intelligent data collection and preprocessing are performed first. Specifically, the error detection engine automatically collects and integrates diagnosis-related data from multiple data sources to provide a complete data foundation for subsequent error detection. Data sources include dialogue logs generated by the consultation simulation layer, raw medical record data, and the system knowledge base. Preprocessing is as follows: Dialogue log parsing: The system automatically parses key information such as the consultation process, diagnosis results, and intent recognition history from the HGDoctor diagnosis log file (results_hgdoctor_*.log) generated by the consultation simulation layer. The system extracts structured data through the parse_log_to_cases() function, converting the log into analyzable case objects.

[0084] Original medical record loading: Load complete medical record information from the medical record JSON file (structured_cases), including chief complaint, present illness, past medical history, physical examination, auxiliary examinations, etc., and establish a mapping dictionary from medical record number to medical record content for comparison and verification of the accuracy of diagnostic results.

[0085] RAG Knowledge Base Index: Automatically scans the RAG knowledge base directory, loads all disease names, builds a disease set index, and supports fast knowledge base coverage checks (6b error detection).

[0086] Disease indicator mapping table construction: Load the association between diseases and test indicators from the "scan test results.txt" file, establish the disease_indicator_mapping mapping table, and support the detection of indicator matching errors (error 1a-2b detection).

[0087] Multi-source data association: Intelligently associate dialogue log data, original medical record data, and knowledge base data, establish a one-to-one mapping through medical record number, verify data integrity, and ensure that each case to be tested has complete contextual information (dialogue process + real medical record + knowledge base).

[0088] The data collection phase employs a strategy combining batch loading and incremental updates, supporting the batch processing of large numbers of medical records with data loading time controlled within 10 seconds. It can also automatically handle issues such as abnormal data formats and missing fields, ensuring data quality through an exception handling mechanism. After collection is complete, a comprehensive data index is built, supporting rapid error detection and analysis.

[0089] Therefore, for a multi-dimensional intelligent error detection mechanism, this application innovatively constructs a 12-category error detection system covering the entire medical diagnosis process, realizing intelligent quality assurance by combining automatic detection with LLM-assisted analysis. Core innovations include: ① A hierarchical error classification system: covering three major categories and 12 error types, including laboratory test indicator errors (1a-2b), consultation process errors (3-5b), and RAG diagnostic errors (6a-6b), with an error coverage rate exceeding 95%; ② An automatic detection algorithm: achieving millisecond-level automatic detection for 25% of error types, including option consistency detection (100% accuracy), RAG diagnostic error detection (precise matching based on loose mapping), and knowledge base coverage detection (comprehensive mapping intersection judgment), with automatic detection efficiency more than 100 times higher than manual review; ③ Disease mapping standardization: achieving original name → standard name → loose The three-layer mapping system of SE mapping supports complex disease name matching through the has_loose_mapping_intersection() algorithm, achieving a mapping accuracy of over 95%; ④ LLM-assisted intelligent analysis: 75% of error types are analyzed using LLM-assisted analysis. Through a large language model, it understands rich contextual information such as medical record content, consultation dialogues, and candidate options, achieving intelligent error identification. The analysis efficiency is 10-20 times higher than traditional manual annotation; ⑤ Closed-loop improvement mechanism: Error detection results directly drive system optimization, forming a closed-loop quality assurance system of "detection → analysis → improvement → verification," continuously improving system quality. Compared to the traditional method relying on expert manual review, this mechanism improves detection efficiency by 50-100 times and detection accuracy by 20-30%, providing key technical support for the high-quality application of medical knowledge graphs.

[0090] S207. Based on the preset error detection algorithm, perform error identification and error classification on the combined key information to obtain detection result data.

[0091] For example, based on a pre-defined multi-dimensional error classification system, a 12-category error detection mechanism was established, covering key aspects of the medical diagnosis process. LLM-assisted intelligent analysis can be employed to achieve high-quality error identification through intelligent reasoning, thereby optimizing the consultation system based on the identified dataset. The error classifications include: Laboratory test indicator errors (1a-2b): Symptom selection mismatch, disease-related indicators missing or redundant; Consultation process errors (3-5b): User answers mismatch, intent recognition error, option consistency problem; RAG diagnosis errors (6a-6b): RAG database diagnosis error, medical record diagnosis not in the knowledge base.

[0092] Each error type defines clear detection conditions and handling strategies, supporting error priority sorting and batch processing. Error detection employs coded management, maintaining error type definitions through the `error_codes` dictionary to ensure the standardization and scalability of error classification.

[0093] Optionally, based on a preset automatic detection algorithm, errors are identified and misclassified in the combined key information to obtain detection result data. The automatic detection algorithm has a millisecond-level response and supports batch detection of large-scale medical records. The automatic detection algorithm can be an error detection algorithm, etc., and is not limited thereto.

[0094] Optionally, it implements automatic detection of 100% of error types, with 25% of these not relying on large model detection. Error types include error 4c (option consistency), 6a (RAG diagnosis error), and 6b (knowledge base coverage). The error 4c detection algorithm uses set operations to determine whether extracted_symptoms completely belong to all_candidate_symptoms, achieving 100% accurate automatic detection. The error 6a detection algorithm checks three conditions (diseases_by_rag exists, is_correct is False, and the RAG library contains loose mappings of medical record diagnoses), achieving precise disease mapping matching through the has_loose_mapping_intersection() function. The error 6b detection algorithm iterates through all medical record diagnoses, checking for intersections with any disease in the RAG library using loose mappings, identifying knowledge base coverage blind spots.

[0095] Optionally, multi-level disease name mapping and standardization functions are also implemented, supporting a three-layer mapping system of original name, standard name, and loose mapping. The mapping algorithm has_loose_mapping_intersection() is implemented through the following steps: Collect all mappings for disease1 (original name, standard name, loose mapping); collect all mappings for disease2; use set intersection operation to determine mapping overlap.

[0096] The mapping data sources include: `get_standard_disease_name()` provides standard name mapping, and `get_disease_name_standard_name_mapping_v2()` provides loose mapping. The mapping mechanism supports synonym recognition, alias processing, and standardized comparison to ensure the accuracy and robustness of disease matching. A disease indicator mapping table, `disease_indicator_mapping`, is maintained, and the association between diseases and test indicators is constructed based on the scan test results.txt file.

[0097] Optionally, the LLM-assisted analysis workflow is as follows: Structured medical record information: Automatically extracts key information such as medical record ID, actual diagnosis, and standardized diagnosis; Intelligent analysis of the consultation process: Parses the consultation rounds and intent recognition history using `parse_questioning_process()`, and utilizes LLM to understand the semantics of the dialogue; Intelligent error type checking: LLM identifies and classifies errors according to a predefined check order based on medical knowledge and contextual information; Automatic recording of analysis results: Automatically records error types, related data, error descriptions, and confidence levels in a structured format. The consultation system can then be optimized based on the analysis results.

[0098] Optionally, it also supports automatic saving of analysis progress, saving every 5 cases, and supports resuming interrupted uploads and batch processing. The LLM analysis engine receives rich contextual information, including medical record content, consultation dialogues, candidate options, and extraction results, and ensures analysis quality and accuracy through multi-round reasoning. Compared with traditional manual annotation methods, LLM-assisted analysis has greatly improved efficiency while maintaining high accuracy.

[0099] Optionally, this embodiment also implements a comprehensive quality assurance mechanism, including data validation, error statistics, and performance monitoring. Data validation includes: medical record data integrity check: verifying the existence of necessary fields; consultation process validity check: ensuring the structural integrity of the consultation dialogue; and error annotation consistency check: avoiding duplicate or conflicting error annotations.

[0100] The error statistics function provides key indicators such as error type distribution, detection accuracy, and processing efficiency. Performance monitoring includes: automatic detection response time (milliseconds); LLM-assisted analysis efficiency (average 30-60 seconds per case, 10-20 times faster than manual annotation); batch processing capability (supports batch analysis of 880 medical records); memory usage and database connection monitoring; and supports the export and analysis of error detection results, providing structured output in JSON format for easy subsequent data analysis and system optimization. The error detection engine, knowledge graph builder, and weight management system form a closed-loop feedback mechanism, allowing detection results to be directly used for system optimization and quality improvement.

[0101] Optionally, the observation and evaluation layer in this embodiment provides end-to-end traceable logging and error typing capabilities, providing a data foundation for continuous system optimization and quality improvement. This layer achieves comprehensive system monitoring and evaluation through the following core functions: End-to-end logging: The system records key processes and intermediate variables at the case-level data entry stage, including case identifiers, based on... The format includes candidate diagnostic results sorted by / n, RAG inference and fusion diagnostic results, extracted results of assisted diagnosis and previous diagnoses, a flattened mapping of "matched symptoms and indicators - r-values" for each disease, dialogue history (including question-answering turns and intent recognition), source annotation, content fragments, and intent recognition summaries. This format can directly drive offline evaluation or serve as a data base for subsequent algorithm auditing and reinforcement learning.

[0102] Optionally, an automated error detection mechanism is also included, meaning the system has two built-in automated error detection paths: 1) Focus on missing symptoms that should be matched but are not: By setting disease A and the set of symptoms currently matched by HgDoctor as a premise, request the large language model to conduct a targeted audit of the full text of the case, locate symptoms related to A but not matched by the system, and attribute them to possible procedural reasons (such as incomplete patient simulation answers, failure of intent matching, templates existing but not hit, missing templates or path problems that lead to no inquiry).

[0103] 2) Focus on noisy symptoms that should not be matched but are matched: By setting case X and the set of matched symptoms as premises, request the large language model to identify symptoms that are inconsistent with the case evidence and output causal analysis (such as unrealistic patient simulation, excessive generalization of intent, template redundancy leading to false coverage, etc.).

[0104] Optionally, closed-loop optimization and corpus governance can also be performed. Specifically, the outputs of the two error detection paths enter the corpus governance and rule optimization loop to supplement, converge, or prune the symptom templates and problem templates, perform quality control on unqualified or conflicting cases, and iterate the intent recognizer and patient simulation strategy, thereby systematically reducing missed and false recalls and achieving continuous self-improvement of the system.

[0105] Optionally, system configuration and operation and maintenance support can also be implemented. Specifically, to accommodate different deployment environments and operational constraints, the process parameters and operating conditions in this application remain configurable and observable. The core scoring parameters a, b, r_min, and n_min can be updated through automated search or evolutionary optimization within a limited range during the evaluation period; the upper limit of consultation rounds, the free text trigger threshold, and the RAG recall window can all be adjusted according to the scenario. The report extraction module supports visual / text mixed input and can be enabled as needed; long context processing and concurrency control achieve stable throughput through queues and resource quotas.

[0106] The system recommends running large language model inference and embedded retrieval engines in an environment with GPU acceleration, and hosting the graph and index on a storage backend with transaction and graph query capabilities (such as a graph database) to ensure the reliability of complex relationship queries and incremental updates. All stages of the entire chain provide de-identification switches, source traceability, and versioning mechanisms to ensure data security and the ability to roll back changes.

[0107] Through the aforementioned three-pronged technical approach of structured modeling, weight adaptation, and automated verification, this application achieves automatic knowledge graph generation, case-driven dynamic expansion, and closed-loop verification and corpus governance oriented towards real-world processes without relying on large-scale manual annotation and maintenance. It also fully leverages the intelligent analysis capabilities of large language models to automate the entire process from data extraction and error detection to quality optimization, thereby significantly improving the reliability and interpretability of candidate ranking, reducing maintenance costs, and providing a standardized measurement and optimization channel for continuous iteration. It should be noted that the numerical values ​​mentioned above are merely illustrative and are not intended to be limiting.

[0108] The method provided in this application embodiment acquires multi-source medical data; extracts medical fact triple data from the medical data; wherein the medical fact triple data includes a list of medical entities and a list of medical relationships. The confidence level and weight value of the medical fact triple data are determined. Based on the medical fact triple data, confidence level, and weight value, a preliminary knowledge graph is generated; wherein the confidence level represents the accuracy of the list of medical entities and the list of medical relationships; and the weight value represents the association strength of each medical relationship in the list of medical relationships. Based on the preliminary knowledge graph, multiple rounds of consultation simulation are performed to generate simulated data; and based on the simulated data and a preset optimization algorithm, the preliminary knowledge graph is optimized to generate a medical knowledge base. According to a preset integrity verifier, isolated relationship identification, invalid weight identification, duplicate entity identification, and data anomaly identification are performed on the target knowledge graph corresponding to the medical knowledge base. Multi-source diagnostic data is acquired; wherein the diagnostic data includes dialogue logs generated from the consultation simulation, original medical records, and the index of the medical knowledge base. Key information in the diagnostic data is extracted according to preset keywords. The key information is combined according to the actual medical record number. Based on a pre-defined error detection algorithm, the system identifies and classifies errors in the combined key information to obtain detection results data. Therefore, by automatically generating and expanding the knowledge graph from case data using a large language model, and combining it with a genetic algorithm to optimize weights and ranking rules, a feedback-driven learning mechanism is established. This enables the automated construction, dynamic expansion, and intelligent optimization of the medical knowledge base, significantly reducing manual costs, rapidly expanding the model's capability boundaries, and greatly improving diagnostic accuracy and system adaptability, achieving excellent performance across multiple dimensions. Furthermore, as an intelligent medical triage engine, the medical knowledge base provides doctors with evidence-based auxiliary diagnostic suggestions, significantly improving the accuracy and efficiency of disease identification. The system supports incremental updates, with the processing time for new case data averaging 1 / 10 of traditional methods, greatly improving the efficiency of knowledge graph construction and maintenance.

[0109] Additionally, based on sorting strategy parameters The disease ranking mechanism achieved significant performance improvements on a large-scale HGDoctor standard test case set, demonstrating a clear advantage over existing technologies. Detailed experimental data comparisons are as follows: The performance comparison analysis with traditional methods is as follows:

[0110] Analysis of core technology breakthroughs: (1) Algorithm innovation breakthrough: Traditional single-dimensional sorting strategies (n-first or r-first) have obvious limitations. This application achieves breakthroughs through... The non-linear scoring formula achieves a scientific balance between the quantity and quality of evidence, improving accuracy by 3.07-8.64 percentage points compared to traditional methods.

[0111] (2) Parameter optimization breakthrough: The parameters of the primary optimization algorithm (a=0.5001, b=2.6429, r_min=0.6589, n_min=1) were obtained through genetic algorithm optimization, achieving a stable improvement of 3.07% compared to manually set parameters (a=1.0, b=1.0). The advanced algorithm adds a high-weight reward mechanism, and when the r value exceeds 0.9785, a special scoring strategy of score = n is adopted. 0.4364 × 10.0 2.3173 This further optimizes the ranking effect of high-confidence diagnoses.

[0112] (3) Breakthrough achievement of composite optimization algorithm: The core innovation of this application—the composite optimization algorithm—achieved a Top-5 accuracy of 84.89%, which has important technical and clinical significance: Compared to the traditional baseline, the accuracy improved by 8.64 percentage points: from 76.25% to 84.89%, an increase of 11.34%; compared to the primary algorithm, the accuracy improved by 5.57 percentage points: from 79.32% to 84.89%, demonstrating the significant advantages of composite optimization. Scientific validation of the optimized parameters: a=0.6684, b=0.8641, r_min=0.4969, n_min=0, achieving an optimal balance between the quantity and quality of evidence; Clinical application value: The 8.64% accuracy improvement means approximately 76 additional correct diagnoses out of 880 cases, significantly improving clinical diagnostic efficiency.

[0113] (4) Clinical value of enhancing diagnostic efficacy: Breakthrough in comprehensive diagnostic accuracy: Top-5+ assisted diagnostic accuracy reaches up to 97.50% (composite optimization algorithm), approaching the level of clinical experts; Significant improvement compared to the baseline: from 95.45% to 97.50%, an increase of 2.05 percentage points, while the error rate decreased by 37.0%; Clinical applicability validation: An accuracy rate of 97.50% means that out of 880 cases, only 22 cases were not correctly diagnosed, which meets the requirements for clinical application; Quantification of the value of auxiliary diagnosis: On average, auxiliary diagnostic information contributes 12-15 percentage points to the accuracy improvement of each algorithm, which proves the important value of multi-source information fusion.

[0114] (5) Engineering advantages of genetic algorithm optimization performance: Optimized computational efficiency: With a population of 50 individuals and 200 generations of evolution, the evolution time per generation is 30-60 seconds, and the total optimization time is 2-4 hours, meeting practical deployment requirements; Guaranteed convergence stability: The algorithm typically reaches the optimal solution within 100-150 generations, with convergence stability exceeding 95%; Reasonable resource consumption: Memory usage is 500MB-1GB, and CPU utilization is 60-80%, suitable for deployment in standard server environments; Scalable design: Supports parallel computing acceleration, and the population size and number of generations can be dynamically adjusted according to hardware resources.

[0115] Furthermore, the system achieves excellent performance across multiple dimensions, with detailed performance metrics as follows: (1) Graph query engine performance: The average response time is controlled within 500 milliseconds, supporting high-concurrency query processing. The caching system adopts the TTL mechanism, and the cache hit rate reaches more than 70%, significantly improving query efficiency. Complex query optimization includes: SQL query optimization, use of composite indexes, and structured encapsulation of query results, improving query performance by 3-5 times.

[0116] (2) Efficiency of the weight management system: Supports enterprise-level version control and batch update operations. A single batch operation can handle the adjustment of 1000+ relationship weights. Batch processing adopts a transaction grouping strategy to avoid long transaction locking. The rollback operation response time is less than 2 seconds, and precise rollback is achieved through reverse operation sequence. The weight update operation achieves millisecond-level response, meeting the real-time adjustment requirements.

[0117] (3) Knowledge graph construction performance: The entity relationship extraction time for a single case text is 2-5 seconds, including the complete entity recognition, relationship extraction, and confidence assessment process. The incremental update mechanism makes the processing time for new case data 1 / 10 of the traditional method on average. The batch processing capability supports large-scale operations that process 1000+ entities and relationships at a time, improving processing efficiency by more than 10 times.

[0118] (4) Concurrency and Scalability: The system supports concurrent access from multiple users and achieves stable throughput through connection pool management and resource quota control. Memory usage optimization adopts streaming processing and caching mechanisms, keeping memory usage within a reasonable range. The system supports horizontal scaling and can dynamically adjust computing resources according to load requirements.

[0119] (5) Error detection and processing efficiency: The automatic detection algorithm has a response time of milliseconds, supporting batch detection of large-scale medical records. The average efficiency of LLM-assisted analysis is 30-60 seconds per case, which is 10-20 times more efficient than traditional manual annotation (5-10 minutes / case), and supports batch analysis of 880 medical records. The system provides progress saving and breakpoint resume functions to ensure the reliability of long-running tasks.

[0120] Furthermore, it achieves quality assurance for the knowledge graph. Specifically, the built-in integrity verification mechanism can automatically detect over 95% of data anomalies, including isolated relationships, invalid weights, and other issues. The intelligent entity fusion algorithm achieves a 90% accuracy rate in synonym recognition, effectively avoiding the generation of duplicate entities and ensuring the quality and consistency of the knowledge graph.

[0121] Furthermore, cost-benefit analysis and economic value assessment were achieved. Specifically, the system demonstrated significant economic benefits in cost control and large-scale application, exhibiting overwhelming advantages compared to traditional methods. The cost comparison with traditional knowledge base construction methods is as follows:

[0122] (1) Revolutionary reduction in labor costs: Traditional knowledge base construction requires a team of medical experts (5-10 people) to manually annotate and organize knowledge for 3-6 months. Based on a monthly salary of 25,000 yuan for a senior medical expert, the labor cost is approximately 500,000 to 1,000,000 yuan. This system adopts an automated extraction and intelligent fusion mechanism, requiring only 1-2 technicians for system maintenance (monthly salary of 8,000 yuan), reducing labor costs to 100,000 to 200,000 yuan. Cost reduction effect: Traditional method 750,000 yuan → This system 150,000 yuan, saving 600,000 yuan, a reduction of 80%, achieving a revolutionary breakthrough in labor costs.

[0123] (2) Improved knowledge update efficiency: The traditional knowledge update cycle is 3-6 months, requiring the reorganization of expert teams for knowledge sorting and verification. This system supports incremental updates, and new case data can be processed and integrated within 2-5 days, improving update efficiency by 30-90 times. Update cycle comparison: The traditional method takes 90-180 days, while this system takes 2-5 days, improving efficiency by 36-90 times.

[0124] (3) Stability of automated operation: The system supports 24 / 7 unattended automated operation with an annual availability of over 99.5%. Automated functions include entity relationship extraction, knowledge graph construction, weight optimization, error detection, and performance monitoring. The system has automatic fault recovery capability, with a mean time to recovery (MTTR) of less than 30 minutes and unplanned downtime of less than 44 hours per year.

[0125] (4) Large-scale processing capability: The system supports large-scale medical data processing, capable of processing 1,000+ case texts per day and 300,000+ cases per year. Compared to traditional manual processing methods (processing 10-20 cases per day), the processing capacity is increased by 50-100 times. The system supports horizontal scaling and can dynamically adjust the processing capacity according to the data scale to meet the needs of medical institutions of different sizes.

[0126] Furthermore, TTL-based intelligent caching and query optimization were implemented. Specifically, this application proposes an intelligent caching optimization mechanism for medical graph queries, significantly improving the query performance of large-scale knowledge graphs. Technical advantages and innovations: ① Adaptability of the TTL caching mechanism: Implementing a Time-To-Live caching mechanism with a cache hit rate exceeding 70%, optimizing query response time from seconds to milliseconds; ② Professionalism of SQL query optimization: Combining composite indexes (disease_id, symptom_id, weight) and query plan optimization, query performance is improved by 3-5 times; ③ Standardization of structured encapsulation: Using the QueryResult data structure to encapsulate query results, supporting metadata such as execution time, cache status, and parameter records, facilitating performance monitoring and optimization; ④ Efficiency of complex relationship queries: Supporting complex queries such as multi-hop relationship traversal, symptom co-occurrence analysis, and disease similarity calculation, meeting the diverse needs of clinical diagnosis. Compared to traditional database query methods, this mechanism has significant performance advantages and technological advancements in large-scale medical knowledge graph scenarios.

[0127] It should be noted that the data, algorithms, etc., mentioned in the above embodiments are merely examples and are not intended to limit the scope of the invention.

[0128] In one example Figure 3 A system architecture diagram of a method for constructing a medical knowledge base provided in this application embodiment is shown below. Figure 3 As shown, the system includes an overall structure (300), which includes: a data layer (310), a preprocessing and standardization layer (320), a knowledge graph construction layer (330), a weight learning and optimization layer (340), a consultation simulation and verification layer (350), an error detection and quality assurance layer (360), and an observation and operation and maintenance layer (370).

[0129] The data layer (310) includes a case text acquisition module (311), a test / image report acquisition module (312), a symptom template management module (313), and a question template management module (314); the preprocessing and standardization layer (320) includes a text cleaning and desensitization module (321), a medical terminology standardization module (322), and a structured parsing module (323); the knowledge graph construction layer (330) includes an LLM entity relation extraction module (331), an entity fusion and deduplication module (332), a relation construction and attribute assignment module (333), and a graph database storage and indexing module (334); the weight learning and optimization layer (340) includes an evidence-driven weight update module (341), a ranking and scoring engine module (342), a genetic algorithm composite optimization module (343), and a weight version management module (344); the consultation simulation and verification layer (350) includes a patient dialogue simulation module (351), an intelligent consultation engine module (352), a diagnosis result verification module (353), and a RAG module. The retrieval and fusion module (354) includes an error detection and quality assurance layer (360), which includes an error detection engine module (361), a disease name mapping module (362), a quality assessment module (363), and an improvement suggestion generation module (364). The observation and operation and maintenance layer (370) includes a full-link log recording module (371), a performance monitoring module (372), a statistical analysis module (373), and a visualization module (374).

[0130] In one example Figure 4 A flowchart illustrating the construction process of a knowledge graph is provided for an embodiment of this application, such as... Figure 4 As shown, the process includes: Start (401); Case text / report input (402); Text preprocessing and desensitization (403); Medical terminology standardization (404); LLM entity and relation extraction (405); Three-level matching (406): exact matching, standard name matching, synonym matching; Intelligent entity fusion and deduplication (407); Relation generation and initial weight assignment (408); Transactional batch processing (409); Graph integrity verification (410); Anomaly detection (411); Graph persistence and index construction (412); End.

[0131] In one example Figure 5 A flowchart of a genetic algorithm composite optimization provided in this application embodiment is shown below. Figure 5As shown, the process includes: Start (501); Initialize the population (502): elite individuals, elite mutated individuals, random individuals; Fitness evaluation (503); Does the convergence condition meet? (504); If not, execute tournament selection (505), crossover operation (506), mutation operation (507), generate a new population (508), and return to fitness evaluation (503); If yes, execute optimal solution selection (509), update system weights and sorting parameters (510), performance verification (511), and end (512).

[0132] In one example Figure 6 A flowchart of error detection and quality assurance provided for embodiments of this application is shown below. Figure 6 As shown, the process includes: Start (601); Loading data (602): consultation dialogue log, original medical records, knowledge base / RAG library, disease-indicator mapping table; Case-level data integration (603); Automatic error detection (604): (Error 4c: option consistency error; Error 6a: RAG diagnosis error; Error 6b: medical record diagnosis not in knowledge base); LLM-assisted error analysis (605): (Error 1a-2b: indicator matching error; Error 3-5b: consultation process error; complex semantic error); Error classification and annotation (606); Quality assessment and statistical analysis (607); Generating system improvement suggestions (608); Closed-loop optimization (609); End (610).

[0133] Corresponding to the above method, embodiments of this application also provide a device for constructing a medical knowledge base, such as... Figure 7 As shown, the device includes: The acquisition module 41 is used to acquire medical data from multiple sources; extract medical fact triple data from the medical data; wherein, the medical fact triple data includes a list of medical entities and a list of medical relationships; The determination module 42 is used to determine the confidence level and weight value of the medical fact triple data; and to generate a preliminary knowledge graph based on the medical fact triple data, the confidence level, and the weight value; wherein, the confidence level represents the accuracy of the medical entity list and the medical relationship list; and the weight value represents the association strength of each medical relationship in the medical relationship list. The optimization module 43 is used to perform multiple rounds of consultation simulation based on the preliminary knowledge graph to generate simulation data; and to optimize the preliminary knowledge graph according to the simulation data and the preset optimization algorithm to generate a medical knowledge base.

[0134] The functions of each functional unit in the medical knowledge base construction device provided in the above embodiments of this application can be implemented through the above method steps. Therefore, the specific working process and beneficial effects of each unit in the medical knowledge base construction device provided in the embodiments of this application will not be repeated here.

[0135] This application also provides an electronic device, such as... Figure 8 As shown, it includes a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540.

[0136] Memory 530 is used to store computer programs; The processor 510 performs the above steps when executing the program stored in the memory 530.

[0137] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0138] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0139] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0140] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0141] The implementation methods and beneficial effects of the various components of the electronic device in the above embodiments for solving the problem can be found in [reference needed]. Figure 1 The steps in the illustrated embodiments are used to implement the electronic device. Therefore, the specific working process and beneficial effects of the electronic device provided in this application will not be repeated here.

[0142] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores instructions that, when executed on a computer, cause the computer to perform the method for constructing a medical knowledge base as described in any of the above embodiments.

[0143] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute the method for constructing a medical knowledge base as described in any of the above embodiments.

[0144] Those skilled in the art will understand that the embodiments in this application can be provided as methods, systems, or computer program products. Therefore, the embodiments in this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the embodiments in this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0145] This application describes embodiments of methods, apparatus (systems), and computer program products according to embodiments of this application with reference to flowchart illustrations and / or block diagrams. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0146] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0147] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0148] Although preferred embodiments have been described in this application, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of this application.

[0149] Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the spirit and scope of the embodiments of this application. Therefore, if these modifications and variations to the embodiments of this application fall within the scope of the claims in this application and their equivalents, then this application also intends to include these modifications and variations.

Claims

1. A method for constructing a medical knowledge base, characterized in that, The method includes: Acquire multi-source medical data; extract medical fact triples from the medical data; wherein the medical fact triples include a list of medical entities and a list of medical relationships; Determine the confidence level and weight value of the medical fact triple data; generate a preliminary knowledge graph based on the medical fact triple data, the confidence level, and the weight value; wherein, the confidence level represents the accuracy of the list of medical entities and the list of medical relationships; and the weight value represents the association strength of each medical relationship in the list of medical relationships. Based on the preliminary knowledge graph, multiple rounds of consultation simulation are performed to generate simulation data; and based on the simulation data and a preset optimization algorithm, the preliminary knowledge graph is optimized to generate a medical knowledge base.

2. The method as described in claim 1, characterized in that, Determining the confidence level and weight values ​​of the medical fact triplet data includes: Determine the baseline confidence level, textual evidence score, and medical terminology score of the medical fact triplet data; The confidence level of the medical fact triplet data is determined based on the baseline confidence level, the text evidence score, the medical terminology score, and the preset confidence threshold. Based on the text evidence score and the medical terminology score, the initial weight values ​​of the medical fact triplet data are determined; If it is determined that there are other data that are identical to the medical fact triple data, then the weight values ​​of the other data are weighted and averaged with the initial weight values ​​of the medical fact triple data to generate a weighted weight value, and the weight values ​​of the other data are updated to the weighted weight value.

3. The method as described in claim 1, characterized in that, Based on the preliminary knowledge graph, multiple rounds of consultation simulation are performed to generate simulation data; and based on the simulation data and a preset optimization algorithm, the preliminary knowledge graph is optimized to generate a medical knowledge base, including: Based on the preliminary knowledge graph, multiple rounds of consultation simulation are performed. In each round of consultation simulation, the confidence and weight values ​​in the preliminary knowledge graph are queried, confidence values ​​below a preset confidence threshold are filtered, and weight values ​​below a preset weight threshold are selected to generate simulation data. Based on preset sorting strategy parameters, the filtered confidence scores and selected weight values ​​in the simulated data are sorted to generate a list of target diseases. Based on a preset error code diagnosis algorithm, error codes are generated for the target disease list; wherein, the error codes are used to optimize the consultation system that initiates the simulated disease query request; Based on a preset optimization algorithm, the target disease list and the preset expected disease list are compared, and the sorting strategy parameters are optimized according to the comparison results. Based on the optimized sorting strategy parameters, the preliminary knowledge graph is optimized to generate a medical knowledge base.

4. The method as described in claim 3, characterized in that, The simulation data also includes diagnostic results data; The method further includes: Extract auxiliary diagnostic information from pre-stored present medical history and auxiliary examinations, and extract past diagnostic information from pre-stored past medical history; Determine the first intersection between the diagnostic result data, auxiliary diagnostic information, and previous diagnostic information and the preset standard pathological diagnosis; and determine the second intersection between the diagnostic result data and the standard pathological diagnosis; Based on the first intersection and the second intersection, the medical knowledge base is optimized to generate an optimized medical knowledge base.

5. The method according to any one of claims 1-4, characterized in that, The method further includes: Based on the preset integrity verifier, the target knowledge graph corresponding to the medical knowledge base is used to identify isolated relationships, invalid weights, duplicate entities, and data anomalies.

6. The method as described in claim 5, characterized in that, The list of medical entities includes multiple medical entities; based on a preset integrity verifier, the target knowledge graph corresponding to the medical knowledge base is subjected to isolated relation identification, invalid weight identification, duplicate entity identification, and data anomaly identification, including: For the identification of duplicate entities, based on the preset integrity verifier and the precise name matching strategy in the preset three-level retrieval strategy, the entity name that is the same as the medical entity is located in the preset database; based on the standard name matching strategy in the three-level retrieval strategy, the standardized name is identified in the pre-constructed standardized medical terminology mapping table; based on the synonym matching strategy in the three-level retrieval strategy, the existing synonym list of medical entities is traversed for fuzzy matching. If duplicate data is found, the duplicate data is fused to obtain fused medical fact triple data.

7. The method according to any one of claims 1-4, characterized in that, The method further includes: Acquire diagnostic data from multiple sources; wherein, the diagnostic data includes dialogue logs generated from the consultation simulation, original medical records, and indexes of the medical knowledge base; Based on preset keywords, extract key information from the diagnostic data; Based on the actual medical record number, the key information is combined; Based on a preset error detection algorithm, the key information in the combination is incorrectly identified and classified to obtain the detection result data.

8. A device for constructing a medical knowledge base, characterized in that, The device includes: The acquisition module is used to acquire medical data from multiple sources; extract medical fact triple data from the medical data; wherein, the medical fact triple data includes a list of medical entities and a list of medical relationships; A determination module is used to determine the confidence level and weight value of the medical fact triple data; and to generate a preliminary knowledge graph based on the medical fact triple data, the confidence level, and the weight value; wherein, the confidence level represents the accuracy of the list of medical entities and the list of medical relationships; and the weight value represents the association strength of each medical relationship in the list of medical relationships. The optimization module is used to perform multiple rounds of consultation simulation based on the preliminary knowledge graph to generate simulation data; and to optimize the preliminary knowledge graph according to the simulation data and a preset optimization algorithm to generate a medical knowledge base.

9. An electronic device, characterized in that, The electronic device includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method of any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-7.