A full medical record quality control system based on a knowledge graph and LLM hybrid architecture

By adopting a hybrid architecture based on knowledge graph and LLM, the medical record quality control system achieves equivalent identification and result reuse in the scenario of repeated storage, which solves the problems of inconsistent analysis results and resource waste in the existing technology and ensures the stability and efficiency of the quality control system.

CN122135867APending Publication Date: 2026-06-02SUZHOU FUXIN MEDICALXIN ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU FUXIN MEDICALXIN ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
Filing Date
2026-03-02
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing medical record quality control systems are prone to inconsistent analysis results and unstable interpretations when triggered multiple times. They lack a mechanism for determining the equivalence of quality control inputs and reusing inference results, leading to waste of system resources and unreproducible results.

Method used

A full medical record quality control system based on a knowledge graph and LLM hybrid architecture is adopted. The equivalence judgment module determines whether the saved events are triggered equivalently, decides whether to start the rule detection path and evidence extraction path, and performs pattern rule mismatch analysis and evidence anchoring consistency analysis to achieve deduplication and traceable archiving of results.

Benefits of technology

It realizes equivalent identification and result reuse of quality control tasks in repeated storage scenarios, ensures consistent output of structured rule conclusions and semantic evidence, solves the problems of inconsistent results and resource waste in existing technologies, and improves the stability and efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135867A_ABST
    Figure CN122135867A_ABST
Patent Text Reader

Abstract

This invention discloses a full medical record quality control system based on a hybrid architecture of knowledge graph and LLM, relating to the field of semantic processing technology. This full medical record quality control system based on a hybrid architecture of knowledge graph and LLM includes: a data acquisition and preprocessing module for acquiring and preprocessing data; an equivalence determination module for determining whether an equivalence trigger exists, whether to initiate a pathway, and whether to reuse quality control results; a consistency determination module for analyzing consistency status, executing rule detection pathways, and handling entry restrictions; an evidence anchoring execution module for performing secondary reasoning and verification; and a result reuse, deduplication, and audit archiving module for deduplication, merging, and archiving. This solves the problem that existing technologies use a save-and-trigger quality control method for document duplication detection. Without quality control input equivalence determination and reasoning result reuse, multiple triggers can easily lead to unstable output of the large language model, resulting in inconsistencies between quality control conclusions and interpretations, and difficulty in deduplicating abnormal results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of semantic processing technology, specifically to a full medical record quality control system based on a hybrid architecture of knowledge graph and LLM. Background Technology

[0002] With the continuous development of medical informatization and electronic medical record systems, medical record quality control has gradually shifted from a post-event review model primarily based on manual spot checks to an automated, end-to-end quality control model relying on information systems. Existing medical record quality control systems typically focus on verifying the completeness, standardization, consistency, and logical rationality of medical record content during the generation, editing, storage, and archiving stages of electronic medical records to meet the needs of medical management, quality assessment, and compliance auditing. As the coverage of electronic medical records continues to expand, medical record content has evolved from a single document to a complex collection of texts containing multiple saves, multiple versions, and cross-stage records. Quality control systems need to simultaneously process structured field information and a large amount of unstructured text content. At the technical implementation level, rule engines are widely used to detect medical record structure and fields based on established standards and specifications, while natural language processing technology is used to assist in analyzing text semantics and contextual relationships. In recent years, with the development of artificial intelligence technology, large-scale language models have been gradually introduced into the field of medical record analysis to improve the ability to understand and interpret complex text content, complementing structure rule-based analysis methods. Under the existing technical architecture, medical record quality control systems typically need to operate in scenarios where documents are frequently saved and edited in real time, and record, display and manage the quality control results triggered multiple times, while taking into account system response efficiency, result consistency and audit traceability requirements, so as to support the continuous operation of clinical business and the implementation of medical quality management.

[0003] For example, the invention patent with announcement number CN117057350B discloses a method and system for named entity recognition in Chinese electronic medical records. This method acquires and preprocesses a Chinese electronic medical record dataset, introduces a Chinese electronic medical record knowledge graph, and embeds knowledge graph triples corresponding to the medical record text into the original data to generate a medical record dataset fused with the knowledge graph. At the model level, a generative adversarial network model, a bidirectional gated recurrent unit model, and a bidirectional long short-term memory network model are trained respectively to extract character features, word features, and character feature sequence vectors fused with the knowledge graph. Furthermore, a graph attention network model is used to model contextual features, and finally, a conditional random field model is combined to complete named entity recognition. This system achieves automatic recognition of named entities in Chinese electronic medical record text through multi-model collaboration and knowledge graph embedding, improving the accuracy of entity recognition and the ability to model contextual semantic relationships.

[0004] For example, the invention patent with publication number CN119380973A discloses an auxiliary diagnosis and treatment method, system, device, and medium based on a large language model. This method obtains the user's self-reported medical condition information, uses a large language model to perform semantic extraction and construct medical record information, and matches examination suggestions based on this. At the same time, it combines the symptom extraction results of the large language model with a convolutional neural network image model to comprehensively analyze the examination data, and obtains further diagnosis and treatment related information through a follow-up questioning mechanism. Finally, the large language model makes a comprehensive judgment and outputs a diagnosis and treatment opinion. This system realizes semantic analysis of medical condition information, auxiliary judgment, and generation of diagnosis and treatment suggestions by combining a large language model with multiple model selection and reasoning processes, and is used to support intelligent decision-making in auxiliary diagnosis and treatment scenarios.

[0005] While existing technologies have applied knowledge graphs and large language models to electronic medical record analysis and intelligent decision support to improve the understanding of unstructured text, most solutions still focus on single-inference or independent task processing, failing to fully consider the characteristics of electronic medical records in practical applications, such as frequent saving, continuous content evolution, and the coexistence of multiple versions. In this context, the lack of mechanisms for determining input equivalence and reusing results in repeatedly triggered scenarios easily leads to inconsistent analysis results or interpretations when the same document is processed multiple times. Furthermore, the collaboration between knowledge graphs and large language models in existing systems is often loosely coupled, lacking stable binding and consistency constraints between structured conclusions and semantic evidence, affecting the reproducibility and audit consistency of results, and failing to meet the needs of complex medical record quality control scenarios.

[0006] Therefore, in order to address the above problems, there is an urgent need for a full medical record quality control system based on a hybrid architecture of knowledge graph and LLM. Summary of the Invention

[0007] Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a full medical record quality control system based on a hybrid architecture of knowledge graph and LLM. This system solves the problem that existing technologies use a save-and-trigger approach to perform duplicate checks on documents. In the absence of quality control input equivalence judgment and reasoning result reuse, multiple triggers can easily lead to unstable output of the large language model, resulting in inconsistencies between quality control conclusions and interpretations, and difficulty in deduplicating abnormal results.

[0008] Technical solution To achieve the above objectives, the present invention provides the following technical solution: a full medical record quality control system based on a hybrid architecture of knowledge graph and LLM, comprising: a data acquisition and preprocessing module, used to acquire and save trigger document data and auxiliary evidence data, obtain quality control rule configuration data and knowledge graph schema definition data, and perform preprocessing; an equivalence determination module, used to perform quality control trigger equivalence determination on the saved trigger document data and quality control rule configuration data, determine whether the saved event constitutes an equivalent trigger, and decide whether to start the rule detection path and evidence extraction path and reuse existing quality control results; a consistency discrimination module, used to perform schema rule mismatch analysis based on the quality control rule configuration data and knowledge graph schema definition data, analyze the consistency status of rule references and schema definitions based on the schema rule mismatch analysis results, and implement the execution and entry restriction processing strategy of the rule detection path; an evidence anchoring execution module, used to perform span drift evidence anchoring consistency analysis on the output results of the evidence extraction path, and perform secondary reasoning and review processes based on the span drift evidence anchoring consistency analysis results; and a result reuse deduplication and audit archiving module, used to perform deduplication and archiving of cross-version results and traceability.

[0009] Furthermore, the specific process for collecting and saving trigger document data and auxiliary evidence data, and obtaining quality control rule configuration data and knowledge graph schema definition data is as follows: Collecting and saving trigger document data includes: a unique identifier for the in-patient document, patient visit serial number data, document saving timestamp, document submission status data, and the full text content of the document; obtaining quality control rule configuration data includes: rule code data, quality control focus node name data, and quality control focus node field identifiers; obtaining knowledge graph schema definition data includes: a unique schema definition identifier, schema definition version number, schema definition node type data, schema definition node attribute label data, node attribute prompt information data, and schema definition relationship constraint set; collecting auxiliary evidence data includes: examination result text data and test result values.

[0010] Further, the specific preprocessing steps are as follows: The full-text document data is segmented into nodes to obtain document node name sequence data and document node content sequence data; uniform text normalization is performed on the document node content sequence data, generating corresponding document version fingerprint data; disambiguation and alignment are performed on duplicate content and positional offsets in the document node content sequence data to obtain node content span sequence data; consistency correction and missing data completion are performed on document save timestamps and document submission status data to obtain save event sequence data; a distributional normalization algorithm is used to standardize and map save trigger document data, quality control rule configuration data, knowledge graph pattern definition data, and auxiliary evidence data; and a min-max linear normalization algorithm is used to normalize the test result values ​​and document save timestamps.

[0011] Furthermore, the specific process for determining the equivalence of quality control triggers for saved trigger document data and quality control rule configuration data is as follows: The number of saved triggers is obtained by using a sliding count algorithm on the saved event sequence data; a character counting algorithm is performed on the full-text content data of the document to obtain the length of the current and previous versions of the full-text content, and the difference is calculated to obtain the document length difference. The previous version refers to the version immediately preceding the current saved event after sorting by document save timestamp under the unique identifier of the document within the same medical visit; the number of entries in the content span sequence data within a node is counted, and the difference between adjacent versions is calculated to obtain the content span quantity difference; the set intersection algorithm is performed on the quality control focus node name data with the current document node name sequence data and the previous version document node name sequence data, and the difference between adjacent versions is calculated to obtain the quality control node quantity difference; the difference between adjacent timestamps is calculated after sorting by the unique identifier of the document within the medical visit. Interval; the median of historical intervals is obtained by statistically analyzing the time intervals in seconds between the unique identifiers of the medical records in the saved event sequence data; the natural logarithm of the number of saved triggers plus one is calculated to obtain the intensity logarithm; the absolute value of the difference in document length is calculated plus one and divided by the length of the full text of the previous version of the document plus one to obtain the change ratio; the product of the intensity logarithm and the change ratio is calculated and its opposite is taken to obtain the equivalent input quantity; the equivalent input quantity is exponentially operated on to obtain the equivalent index; the reciprocal of the absolute value of the difference in the number of content spans is calculated and one is taken to obtain the content suppression; the reciprocal of the absolute value of the difference in the number of quality control nodes is calculated and one is taken to obtain the node suppression; the ratio of the adjacent time interval to the median of historical intervals is calculated, one is taken, and the reciprocal is taken to obtain the time suppression; the equivalent index, content suppression, node suppression, and time suppression are multiplied together to obtain the quality control trigger equivalence judgment value.

[0012] Furthermore, the specific process for determining whether a saved event constitutes an equivalent trigger, and deciding whether to activate the rule detection path, evidence extraction path, and reuse existing quality control results, is as follows: Real-time comparison of the quality control trigger equivalence judgment value and the quality control trigger equivalence judgment threshold: When the quality control trigger equivalence judgment value is less than the quality control trigger equivalence judgment threshold, it is determined to be a non-equivalence trigger; static consistency verification is performed on the quality control attention node name data and the quality control attention node field identifier; the rule detection path and evidence extraction path are activated. When the quality control trigger equivalence judgment value is greater than or equal to the quality control trigger equivalence judgment threshold, it is determined to be an equivalent trigger; the activation of the rule detection path and the large language model evidence extraction path is stopped, and the reuse hit is performed. The reuse of rule code data, defect level data, location node data, and evidence anchoring quadruples is conditional upon the corresponding pattern definition version number being consistent with the rule configuration. Specifically, in the rule detection pathway, structured rule detection is performed on the document node content sequence data based on the knowledge graph pattern definition data and quality control focus node name data, outputting the matched rule code data, defect level data, and location node data, and generating corresponding quality control conclusions. In the evidence extraction pathway, auxiliary evidence data is filtered based on medical visit serial number data, the auxiliary evidence data is deterministically sorted, and a large language model is invoked for evidence extraction and interpretation generation, outputting evidence anchoring quadruples.

[0013] Furthermore, the specific process of performing pattern rule mismatch analysis based on quality control rule configuration data and knowledge graph pattern definition data is as follows: For the quality control attention node name data and quality control attention node field identifiers, respectively, an existence match counting algorithm is performed on the pattern definition node type data and pattern definition node attribute label data to obtain the number of missing entries. These numbers are then summed to form the rule reference missing count. Finally, the number of entries with successful existence matches is counted to obtain the rule reference hit count. For the missing quality control attention node name data and quality control attention node field identifiers, a minimum edit distance normalization algorithm is performed on them respectively. The similarity corresponding to the minimum distance is taken as the renaming suspicion. The degree of renaming suspicion is calculated as follows: the existence of the field identifiers of the quality control concern nodes in the attribute label data of the pattern definition nodes is matched by the identifiers being completely identical. For fields that do not match, the renaming suspicion is calculated by the minimum edit distance normalization algorithm; the ratio of the rule reference missing count to the rule reference hit count plus one is calculated to obtain the missing ratio term; the inverse of the missing ratio term is taken and exponentially to obtain the missing decay term; the renaming suspicion is calculated by adding one to the renaming suppression term to obtain the renaming suppression term; the missing decay term and the renaming suppression term are multiplied to obtain the match preservation term; and the pattern rule mismatch risk value is obtained by subtracting the match preservation term from one.

[0014] Furthermore, based on the pattern rule mismatch analysis results, the consistency status of rule references and pattern definitions is analyzed. The specific process for executing the rule detection path and entering the restricted processing strategy is as follows: Real-time comparison of the pattern rule mismatch risk value and the pattern rule mismatch risk threshold: When the pattern rule mismatch risk value is greater than or equal to the pattern rule mismatch risk threshold, the execution of the rule detection corresponding to this rule code is suspended, the knowledge graph pattern definition cache refresh and rule configuration rollback check are triggered, and the document task is marked as having insufficient structural constraints, entering the evidence anchoring execution module; When the pattern rule mismatch risk value is less than the pattern rule mismatch risk threshold, the rule detection path is executed according to the rule code data, and the consistency verification result is created and archived to the pattern consistency database, entering the evidence anchoring execution module.

[0015] Furthermore, the specific process of performing span-drift evidence anchoring consistency analysis on the output results of the evidence extraction pathway is as follows: The evidence anchoring quadruple is aggregated and counted according to the quality control focus node name data and the pattern definition node attribute label data to obtain the distribution probability, which is then normalized using the total number of evidence anchoring quadruples in this round. The evidence anchoring distribution entropy is then obtained using the Shannon entropy calculation algorithm. The evidence anchoring quadruple and the fixed evidence candidate set, composed of evidence fragment sets formed by rule-driven text fragment extraction and node attribute mapping from inspection result text data, test result values, and document node content sequence data, are subjected to a set intersection-union ratio calculation algorithm to obtain the evidence overlap. The longest common substring determination is performed on the explanatory text output by the large language model in two adjacent rounds within the document node content sequence data. The algorithm calculates and normalizes the difference between the start and end positions of characters to obtain the explanatory text span drift intensity. The two adjacent rounds refer to the two inference outputs corresponding to the same quality control task under two consecutive save triggers. The algorithm calculates the reciprocal of the evidence anchoring distribution entropy plus one to obtain the distribution entropy suppression term. The algorithm calculates the difference between the evidence overlap degree of the current round and the evidence overlap degree of the previous round and takes the absolute value to obtain the change amplitude term. The algorithm takes the opposite number of the change amplitude term and performs an exponential operation to obtain the evidence overlap index term. The algorithm calculates the ratio of the explanatory text span drift intensity to the explanatory text span drift intensity plus one to obtain the drift ratio term. The algorithm takes the opposite number of the drift ratio term and performs an exponential operation to obtain the drift index term. The algorithm multiplies the distribution entropy suppression term, the evidence overlap index term, and the drift index term to obtain the evidence anchoring consistency value.

[0016] Furthermore, the specific process of secondary reasoning and review based on the results of the span drift evidence anchoring consistency analysis is as follows: Real-time comparison of evidence anchoring consistency value and evidence anchoring consistency threshold: When the evidence anchoring consistency value is less than the evidence anchoring consistency threshold, a constrained secondary reasoning execution is triggered for the current quality control task. If the evidence anchoring consistency value is still less than the evidence anchoring consistency threshold after secondary reasoning, an evidence anchoring insufficiency flag is output. When the evidence anchoring consistency value is greater than or equal to the evidence anchoring consistency threshold, the explanatory text and evidence anchoring quadruple are written back along with the quality control conclusion, and an abnormal record deduplication flag is generated. The abnormal record deduplication flag is hash digested according to the concatenation order and written into the violation information entity.

[0017] Furthermore, the specific process of deduplicating and archiving cross-version results is as follows: the results of deduplication of quality control conclusions and abnormal records are written and merged: when the deduplication marker of abnormal records exists, only the document storage timestamp and reuse hit record are added; when the deduplication marker of abnormal records does not exist, the violation information entity is written and the rule detection result is bound as the structured trace of the knowledge graph path, the evidence anchoring quadruple and the explanation text are bound as the semantic trace of the large language model path, and the pattern definition version number and rule code data are archived to the pattern consistency database.

[0018] Beneficial effects The present invention has the following beneficial effects: (1) By introducing a quality control trigger equivalence judgment mechanism based on the characteristics of saving events and changes in document content, this invention achieves the effect of equivalent identification and result reuse of quality control tasks in repeated saving scenarios, effectively solving the problem of inconsistent results and waste of system resources caused by repeated quality control triggering due to multiple saving in the prior art.

[0019] (2) This invention constructs a hybrid architecture in which the rule detection pathway and the evidence extraction pathway operate in tandem, and establishes a stable anchoring and binding relationship between the structured rule conclusions and semantic evidence, thereby achieving the effect of consistent output of structural conclusions and semantic interpretations from the same source, effectively solving the problem of the disconnect between knowledge graph results and large language model interpretations in the prior art.

[0020] (3) This invention achieves the effect of verifiable structural constraints before rule execution by conducting consistency judgment and risk assessment on the reference relationship between the quality control rule configuration and the knowledge graph pattern definition, thereby effectively solving the problem of silent failure or mis-execution of rules due to pattern evolution or rule mismatch in the prior art.

[0021] (4) This invention performs stability and consistency analysis on the evidence anchoring results of the large language model, and triggers constrained secondary reasoning or enters the review process when the consistency conditions are not met, thereby achieving the effect of stable and controllable evidence location and interpretation output, effectively solving the problem of large fluctuations and difficulty in reproducing semantic reasoning results in the prior art.

[0022] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0023] Figure 1 This is a structural diagram of a full medical record quality control system based on a hybrid architecture of knowledge graph and LLM according to the present invention; Figure 2 This is a schematic diagram illustrating how the save trigger equivalence determination value changes with the save event in this invention; Figure 3 This is a closed-loop flowchart of the hybrid architecture quality control system of the present invention. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] Please see Figures 1-3 This invention provides a technical solution: a full medical record quality control system based on a hybrid architecture of knowledge graph and LLM, comprising: a data acquisition and preprocessing module, used to acquire and save trigger document data and auxiliary evidence data, obtain quality control rule configuration data and knowledge graph pattern definition data, and perform preprocessing; an equivalence determination module, used to perform quality control trigger equivalence determination on the saved trigger document data and quality control rule configuration data, determine whether the saved event constitutes an equivalent trigger, and decide whether to start the rule detection path and evidence extraction path and reuse existing quality control results; a consistency discrimination module, used to perform pattern rule mismatch analysis based on the quality control rule configuration data and knowledge graph pattern definition data, analyze the consistency status of rule references and pattern definitions according to the pattern rule mismatch analysis results, and implement the execution of the rule detection path and entry restriction processing strategy; an evidence anchoring execution module, used to perform span drift evidence anchoring consistency analysis on the output results of the evidence extraction path, and perform secondary reasoning and review processes according to the span drift evidence anchoring consistency analysis results; and a result reuse deduplication and audit archiving module, used to perform cross-version result deduplication and traceable archiving.

[0026] Specifically, the process of collecting and saving trigger document data and auxiliary evidence data, and obtaining quality control rule configuration data and knowledge graph pattern definition data is as follows: Collect and save trigger document data, which includes: unique identifier of the document in the medical visit, medical visit serial number data, document saving timestamp, document submission status data and full text content data of the document. This serves as a unified input for knowledge graph structured modeling and large language model semantic reasoning, enabling the two pathways of the hybrid architecture to run under the same document version and the same trigger event.

[0027] Obtain quality control rule configuration data, which includes rule code data, quality control attention node name data, and quality control attention node field identifiers. This data serves as the structural constraint basis for the knowledge graph rule detection pathway and as the source of prompt constraints for evidence extraction and interpretation generation in the large language model.

[0028] Acquire knowledge graph pattern definition data, which includes: unique pattern definition identifier, pattern definition version number, pattern definition node type data, pattern definition node attribute label data, node attribute hint information data, and pattern definition relation constraint set. This data is used to map the full text content of a document into a structured expression of nodes and attributes, and serves as the node attribute boundaries and relation constraints for the evidence anchoring quadruple output by the large language model, thereby forming structural semantic collaboration in the hybrid architecture. The pattern definition relation constraint set describes the relationship types and directional constraints between node types allowed at the pattern definition level, and is used to limit the legal scope of relation generation in knowledge graph structured modeling, rule detection path execution, and evidence anchoring quadruple.

[0029] Collect auxiliary evidence data, including textual data of inspection results and numerical data of test results, as input to the evidence candidate set of the large language model pathway, and as a supplementary source of external evidence nodes or attribute values ​​of the knowledge graph.

[0030] This implementation plan constructs a comprehensive input foundation covering document content, rule constraints, structural semantic boundaries, and external evidence by uniformly collecting and organizing trigger document data, quality control rule configuration data, knowledge graph pattern definition data, and auxiliary evidence data. This ensures that the knowledge graph rule detection pathway and the large language model semantic reasoning pathway obtain consistent data support under the same document version and the same trigger event context. This eliminates the risk of reasoning bias caused by asynchronous multi-source data, version inconsistency, or contextual fragmentation at the system entry level. It also provides a stable, traceable, and reproducible input environment for subsequent trigger equivalence determination, structured rule detection, evidence anchoring execution, and result reuse and deduplication, ensuring that the hybrid architecture has a unified cognitive baseline and reliable operating premise in the whole medical record quality control scenario.

[0031] Specifically, the preprocessing process is as follows: The full-text content data of the document is segmented into document node name sequence data and document node content sequence data using document node template matching and hierarchical parsing algorithms. By mapping the full-text text into node-level payloads with clear semantic boundaries and hierarchical relationships, the impact of different document structures on the consistency of subsequent rule detection and evidence localization is eliminated, providing a stable structured input foundation for the hybrid architecture. Uniform text normalization processing is performed on the document node content sequence data, generating corresponding document version fingerprint data. By generating a unique version-level fingerprint within a unified normalized corpus space, document versions with only surface format differences or no semantic changes possess discriminability and traceability in subsequent equivalence judgment and result reuse processes. Disambiguation and alignment processing is performed on duplicate content and positional offsets in the document node content sequence data to obtain content span sequence data within nodes. By constraining the unique span expression of the same semantic content within a node, structural noise introduced by copying, pasting, or paragraph rearrangement is suppressed, preventing positional drift of evidence anchoring during multiple saving processes. Consistency correction and missing data completion are performed on document retention timestamps and document submission status data to obtain retention event sequence data. By constructing a continuous, monotonic, and state-discriminable retention event time series, a reliable time benchmark is provided for subsequent trigger equivalence analysis, task idempotency control, and result write-back validity determination. A distribution standardization algorithm is used to standardize and map retention trigger document data, quality control rule configuration data, knowledge graph pattern definition data, and auxiliary evidence data; this ensures comparability of indicators with different dimensions and numerical distribution characteristics within the same statistical space, reducing the interference of extreme values ​​or distribution skewness on evidence ranking and consistency assessment. A min-maximum linear normalization algorithm is used to normalize the test result values ​​and document retention timestamps; this maps various numerical features to a fixed interval, ensuring a stable and controllable numerical contribution in multi-factor joint and judgment processes.

[0032] This implementation plan constructs a stable and consistent data foundation by conducting unified structured, standardized, and numerically constrained analysis on the content, temporal behavior, and supporting evidence of the entire medical record. This foundation ensures that different versions of documents, different storage schedules, and evidence from different sources can be uniformly understood and reliably aligned within the hybrid architecture. This effectively eliminates uncertainties introduced by differences in document structure, format noise, repeated editing, and temporal discontinuity. It provides a reproducible, comparable, and auditably credible foundation for subsequent trigger equivalence determination, rule detection consistency, evidence anchoring stability, and cross-version result reuse.

[0033] Specifically, the process for determining the equivalence of quality control triggers for saved trigger document data and quality control rule configuration data is as follows: The number of save triggers is obtained by using a sliding count algorithm on the saved event sequence data. Based on the time precision resolution of the document save timestamp and the minimum distinguishable time interval of the system save event, combined with the save event trigger mechanism, the minimum time window that is greater than the minimum distinguishable interval of the system timestamp and can cover one continuous editing operation cycle is taken as the state window length, with a value range between 10 and 600 seconds. A character counting algorithm is then performed on the full-text content data of the document to obtain the length of the current full-text content and the length of the previous version of the document. The previous version of the document's full-text content data refers to the document's full-text content data immediately preceding the current save event, sorted by document save timestamp under the unique identifier of the document within the same medical visit. The difference is used to obtain the document length difference. The character counting algorithm is then used to obtain the length of the previous version of the document's full-text content data. The algorithm calculates the length of the full text of the previous version of the document; counts the number of spans in the content span sequence data within each node to obtain the current span count and the previous version span count, and calculates the difference to obtain the content span count difference; performs a set intersection algorithm on the quality control focus node name data and the current document node name sequence data to obtain the current quality control focus node set and counts; performs a set intersection algorithm on the quality control focus node name data and the document node name sequence data obtained by parsing the full text content data of the previous version of the document selected by the unique identifier of the document within the same medical visit and sorted by document save timestamp to obtain the previous version quality control focus node set and counts; calculates the difference between the current count and the previous version count to obtain the quality control node count difference; sorts the document save timestamps by the unique identifier of the document within the medical visit and calculates the difference to obtain the adjacent time interval; and uses a median statistical algorithm to obtain the historical interval median of the time interval sequence of the unique identifier of the document within the medical visit in the saved event sequence data.

[0034] The natural logarithm of the save trigger count plus one is calculated to obtain the intensity logarithm term. By performing a logarithmic mapping on the save trigger count, high-frequency repeated save behavior is compressed into a numerical range where its impact on the judgment value gradually diminishes. The absolute value of the document length difference plus one and divided by the full text length of the previous version plus one yields the change ratio term. By introducing a relative ratio, the content change amplitude of documents of different sizes is scaled to ensure that length changes are comparable and consistent in judgment semantics across different document lengths. The product of the intensity logarithm term and the change ratio term is multiplied and its negative is taken to obtain the equivalent input quantity. The save trigger intensity and content change amplitude are coupled and modeled to form a synergistic decay relationship between high-frequency saves and large content changes in the numerical space. An exponential operation is performed on the equivalent input quantity to obtain the equivalent exponential term. The coupled decay quantity is mapped to a continuous interval between zero and one through an exponential function, making the trigger equivalence exhibit smooth, monotonic, and threshold-separable numerical characteristics as content and trigger intensity change. The absolute value of the difference in the number of content spans is calculated. The reciprocal of the value plus one yields the content suppression term, which models the suppression of span changes introduced by internal document structure rearrangement, paragraph copying, or cross-node adjustments using the reciprocal form. The reciprocal of the absolute value of the difference in the number of quality control nodes plus one yields the node suppression term, characterizing whether the scope of quality control focus has shifted or expanded before and after the save event. This constrains the equivalence judgment result from the perspective of rule focus, avoiding misjudgments of equivalence triggers due to changes in quality control focus nodes. The reciprocal of the ratio of adjacent time intervals to the median of historical intervals plus one yields the time suppression term. By comparing the current save rhythm with the statistical characteristics of historical save behavior, an adaptive constraint in the time dimension is introduced to suppress the interference of abnormal save rhythms on the stability of equivalence judgment. Multiplying the equivalence index term, content suppression term, node suppression term, and time suppression term yields the quality control trigger equivalence judgment value. The combined constraint judgment of save trigger behavior under multiple dimensions such as content changes, structural changes, rule focus changes, and time rhythm is achieved through the product of multiple suppression factors. The specific calculation formula is as follows: ; In the formula, This represents the quality control trigger equivalence judgment value, which characterizes the degree of equivalence between the currently saved trigger event and the previous saved event at the quality control semantic level; This indicates the number of times the event is saved, reflecting the intensity of the continuous save operation within the status window. It represents the difference in document length, used to describe the absolute magnitude of text content changes between adjacent saved events; This indicates the length of the full text of the previous version of the document, serving as a benchmark for the change ratio to eliminate the impact of differences in document size. It represents the difference in the number of content spans, reflecting the rearrangement or repetition of internal structural fragments in adjacent save events; This represents the difference in the number of quality control nodes, indicating the degree of change in the scope of concern of the quality control rules before and after the event is saved; It represents the time interval between adjacent events, depicting the temporal rhythm characteristics of the current saved event relative to the previous saved event; The median of the historical interval is used as a robust statistical benchmark for historical preservation behavior to suppress the impact of occasional time anomalies on the judgment results.

[0035] In this embodiment, Table 1 is an example data table for saving trigger equivalence determination. It records in detail the number of saving triggers, the change value of the full text length of the document, the change value of the number of content spans, the change value of the number of quality control attention nodes, the time interval between adjacent saving, and the finally calculated quality control trigger equivalence determination value for the same document during continuous saving. It is used to quantify the degree of equivalence of document saving trigger behavior under different saving events. Specifically: the first save had 0 save triggers, a change of 5 in the total document length, a change of 0 in the number of content spans, a change of 0 in the number of quality control focus nodes, a save interval of 25 seconds, and a quality control trigger equivalence judgment value of 0.545455; the second save had 1 save trigger, a change of 20 in the total document length, a change of 1 in the number of content spans, a change of 0 in the number of quality control focus nodes, a save interval of 20 seconds, and a quality control trigger equivalence judgment value of 0.291409; the third save had 3 save triggers, a change of 120 in the total document length, a change of 4 in the number of content spans, a change of 1 in the number of quality control focus nodes, a save interval of 10 seconds, and a quality control trigger equivalence judgment value of 0.05435. 4; The fourth save corresponds to 5 save triggers, the change in the total document length is 0, the change in the number of content spans is 0, the change in the number of quality control focus nodes is 0, the time interval between adjacent saves is 5 seconds, and the quality control trigger equivalence judgment value is 0.854200; The fifth save corresponds to 2 save triggers, the change in the total document length is 300, the change in the number of content spans is 8, the change in the number of quality control focus nodes is 3, the time interval between adjacent saves is 40 seconds, and the quality control trigger equivalence judgment value is 0.006311; The sixth save corresponds to 1 save trigger, the change in the total document length is 15, the change in the number of content spans is 0, the change in the number of quality control focus nodes is 0, the time interval between adjacent saves is 90 seconds, and the quality control trigger equivalence judgment value is 0.246646.

[0036] Table 1 contains example data for saving trigger equivalence determination. like Figure 2The diagram illustrates the changes in the equivalence judgment value of the save trigger as a function of save events. Table 1 shows that, in the continuous save scenario, there are significant differences in the equivalence judgment values ​​of the quality control trigger corresponding to different save events. The fourth save has the highest equivalence judgment value at 0.854200, indicating that despite numerous save triggers in this event, the full text content and structural features of the document did not undergo substantial changes, making the save triggering behavior highly equivalent. The fifth save has the lowest equivalence judgment value at only 0.006311, reflecting significant changes in the document's content length, structural span, and quality control focus points under this save event, representing a typical non-equivalence triggering situation. The equivalence judgment value for the third save is 0.054354, also at a low level, indicating that equivalence decreases rapidly with increasing save triggers and cumulative changes in document structure. The equivalence judgment values ​​for the first, second, and sixth saves are in the middle range, reflecting certain differences in the magnitude of content changes and the save rhythm of their save triggering behaviors. Overall, the diagram showing the change of the save trigger equivalence judgment value with the save event intuitively reflects the distinction between equivalent and non-equivalent triggers during continuous saving, and can serve as an important basis for determining whether to reuse historical results in subsequent quality control tasks.

[0037] In this implementation plan, by jointly quantifying and modeling multi-dimensional information such as event behavior, changes in document content, structural adjustment characteristics, changes in the scope of quality control concern, and the timing of saving, a unified quality control trigger equivalence judgment value is formed. This enables the system to stably determine whether the saving trigger introduces substantial changes in quality control input in scenarios of high-frequency editing and repeated saving of the entire medical record. Thus, at the entry point of the quality control process, it distinguishes between non-equivalent triggers that require re-execution of hybrid inference and equivalent triggers that can directly reuse existing results. This effectively suppresses the problems of large language model output fluctuations, inconsistent quality control conclusions, and difficulty in deduplicating abnormal results caused by repeated triggers. At the same time, it significantly reduces the number of invalid quality control tasks and the overall computational cost of the system, providing a reliable scheduling basis for the orderly execution of subsequent rule detection and evidence extraction pathways.

[0038] Specifically, the process of determining whether a saved event constitutes an equivalent trigger, and deciding whether to activate the rule detection pathway, the evidence extraction pathway, and reuse existing quality control results, is as follows: Real-time comparison of the quality control trigger equivalence judgment value and the quality control trigger equivalence judgment threshold. When the quality control trigger equivalence judgment value is less than the quality control trigger equivalence judgment threshold, it is judged as a non-equivalence trigger. A quality control task identifier is generated and entered into the quality control task execution queue. This is used to indicate that the currently saved event has undergone substantial changes in content, structure, or time rhythm, and the complete quality control process needs to be re-executed to ensure the accuracy and timeliness of the judgment result. Static consistency verification is performed on the quality control attention node name data and quality control attention node field identifier, combined with the pattern definition node type data and pattern definition node attribute label data. By completing the structural constraint verification of the rule and pattern definition before the rule is executed, incomplete or unreproducible rule detection results due to pattern evolution or configuration abnormalities are avoided. Under the premise of passing the verification, the rule detection path and evidence extraction path are started, and the path output is bound to the unique identifier of the document within the same medical visit and the document's save timestamp to create and archive it to the quality control result archive record data. If there are incomplete tasks for the same document unique identifier, the write-back qualification is canceled and only the write-back qualification of the latest task is retained. The write-back qualification constraint ensures that the structured conclusions and semantic interpretations under the hybrid architecture always correspond to the latest document version, avoiding result overwriting or version mismatch due to asynchronous execution.

[0039] When the quality control trigger equivalence judgment value is greater than or equal to the quality control trigger equivalence judgment threshold, it is determined to be an equivalence trigger. The startup of new rule detection pathways and large language model evidence extraction pathways is stopped. By directly blocking repeated inference requests, the number of large language model calls and the overall system computational overhead in high-frequency saving scenarios are significantly reduced. Based on the unique identifier of the same medical record and the document saving timestamp corresponding to the previous saving event, the archived hit rule code data, defect level data, location node data and evidence anchoring quadruples are read and reused from the quality control result archive record data. The reuse is conditional on the corresponding pattern definition version number being consistent with the rule configuration, so as to achieve the same source reuse of rule-level conclusions and semantic-level interpretations, and ensure the consistency and stability of output results under repeated triggering conditions. At the same time, an equivalence reuse event log is generated based on the current saving event, and the reuse results are written back to the front-end page of the inpatient electronic medical record system for display, ensuring that the knowledge graph conclusions and large language model interpretations are of the same source under repeated triggering scenarios.

[0040] like Figure 3The diagram shows the closed-loop flowchart of the hybrid architecture quality control system. In the rule detection pathway, structured rule detection is performed on the document node content sequence data based on knowledge graph pattern definition data and quality control focus node name data. This outputs the hit rule code data, defect level data, and location node data, generating corresponding quality control conclusions. The explicit structural constraints of the knowledge graph ensure the determinism and interpretability of the rule judgment results. In the evidence extraction pathway, auxiliary evidence data is filtered based on medical visit serial number data. This auxiliary evidence data is deterministically sorted according to time sequence and rule citation order, constructing ranked evidence candidates. The sequence reduces the randomness in the reasoning process of the large language model by fixing the order of evidence input. It calls the large language model to extract and interpret evidence and outputs evidence anchoring quadruples. The output of the large language model is then attached back to the attribute tags of the knowledge graph nodes through the evidence anchoring quadruples. This constrains the freely generated semantic interpretations to the scope of clear structural nodes and attributes, achieving the homologous binding of structured conclusions and semantic evidence. This forms a closed-loop hybrid architecture of knowledge graph and large language model, thus simultaneously ensuring reasoning stability, result consistency and system operating efficiency in real business scenarios with multiple saves and repeated triggers.

[0041] In this implementation plan, a diversion control mechanism based on the equivalent judgment value of quality control triggers is introduced to realize the pre-judgment and process scheduling of the saved trigger behavior in the hybrid architecture quality control system. This enables the system to stably distinguish between effective quality control triggers that require re-execution of rule detection and evidence extraction pathways and equivalent triggers that can directly reuse existing quality control conclusions and semantic interpretations in the operating environment of high-frequency editing and repeated saving of the entire medical record. This significantly reduces the result fluctuations and computing power consumption caused by repeated reasoning while ensuring that the knowledge graph structured judgment results and the semantic interpretation of the large language model are of the same origin. Furthermore, through a unified result archiving, reuse, and write-back mechanism, the consistency, traceability, and overall operating efficiency of quality control conclusions in multiple saves, multiple versions of documents, and real business scenarios are ensured.

[0042] Specifically, the process of performing pattern rule mismatch analysis based on quality control rule configuration data and knowledge graph pattern definition data is as follows: For the quality control attention node name data and quality control attention node field identifiers, an existence matching counting algorithm is executed in the pattern definition node type data and pattern definition node attribute label data respectively to obtain the number of missing entries, and the sum is used to obtain the rule reference missing count. By performing existence verification at both the node and attribute levels, the system systematically characterizes the coverage completeness of the rule configuration in the structural dimension, thereby avoiding the underestimation of mismatch risk caused by single-dimensional matching. An existence matching counting algorithm is then executed in the pattern definition node type data and pattern definition node attribute label data respectively to count the number of successfully matched entries, obtaining the rule reference hit count, which is used to quantify the current rule configuration and the knowledge graph pattern definition. The established effective reference relationships between definitions provide a stable cardinal constraint for subsequent risk ratio calculations. For missing quality control attention node name data and quality control attention node field identifiers, the minimum edit distance normalization algorithm is applied to the pattern definition node type data and pattern definition node attribute label data, respectively. The similarity of the minimum distance is taken as the renaming suspicion. By modeling the naming differences using edit distance, a fault-tolerant characterization of field aliases, naming evolution, and semantic similarity is introduced, so that the rule reference analysis has the necessary semantic robustness while maintaining structural rigor. Among them, the existence matching of quality control attention node field identifiers in pattern definition node attribute label data is performed in a way that the identifiers are completely consistent. Fields that fail to match are then renaming suspicion calculated using the minimum edit distance normalization algorithm.

[0043] The ratio of the rule reference missing count to the rule reference hit count plus one is used to obtain the missing ratio term. By proportionalizing the missing and hit counts, the rule reference integrity is made comparable under different rule sizes and configuration complexities. The inverse of the missing ratio term is then exponentially calculated to obtain the missing decay term. This exponential mapping is used to non-linearly suppress the growth of the rule missing ratio, making the impact of a small number of missing rules on the risk value more gradual. The exponential decay term ranges from zero to one and decreases monotonically with the rule missing ratio. Finally, the renaming suspicion level is calculated by adding one to obtain the renaming suppression term. This reciprocal form is used to suppress potential rule renamings. Constraint modeling is used for cases involving field aliases to provide a reasonable buffer for semantically similar but inconsistently named situations in risk calculation. The renaming suppression term ranges from zero to one and decreases monotonically with the renaming suspicion level. The missing term is multiplied by the renaming suppression term to obtain the match preservation term. By jointly analyzing rule integrity and naming consistency, the comprehensive degree to which the current rule configuration maintains effective matching at the structural and semantic levels is characterized. The match preservation term is subtracted from the value of one to obtain the pattern rule mismatch risk value, which increases monotonically as the rule matching degree decreases, thus forming an intuitive and thresholdable mismatch risk metric. The specific calculation formula is as follows: ; In the formula, The indicator rule mismatch risk value represents the overall degree of mismatch between the current rule configuration and the knowledge graph pattern definition at the structural and naming levels. This indicates the rule reference missing count, reflecting the number of references in the quality control rule that could not be found in the current schema definition; It represents the rule reference hit count, characterizing the scale of valid reference relationships established between rule configuration and pattern definition; This indicates the degree of suspicion of renaming, quantifying the extent of potential interpretable bias caused by naming differences, field aliases, or version evolution when rule references are not hit.

[0044] This implementation scheme systematically models and quantifies the risk of node and attribute reference relationships between quality control rule configuration and knowledge graph schema definition, constructing a schema rule mismatch risk value. This enables the system to accurately identify potential mismatches in rule configuration at the structural coverage and naming consistency levels before rule detection and execution. This avoids problems such as incomplete, unstable, or unreproducible rule detection results due to schema evolution, field renaming, or configuration drift. It also provides thresholdable decision-making basis for whether to suspend rule execution, trigger schema refresh, or enter a conservative quality control path, ensuring that the hybrid architecture of knowledge graph and large language model has a stable structural constraint foundation and sustainable operation capability in the whole medical record quality control scenario.

[0045] Specifically, the process of analyzing the consistency status between rule references and pattern definitions based on the pattern rule mismatch analysis results, and then executing the rule detection path and implementing the restricted processing strategy, involves real-time comparison of the pattern rule mismatch risk value and the pattern rule mismatch risk threshold. When the pattern rule mismatch risk value is greater than or equal to the pattern rule mismatch risk threshold, a mismatch risk is determined, and the execution of the rule detection corresponding to this rule code is suspended. This is used to prevent the continued execution of rule detection under insufficient structural constraints, thereby preventing incomplete or unreproducible structured quality control conclusions. This triggers a knowledge graph pattern definition cache refresh and rule configuration rollback check. The cache refresh only applies to the currently used cache instance corresponding to the unique identifier and version number of the current pattern definition. The rule configuration rollback check verifies whether the current rule configuration deviates from the most recently confirmed stable and effective rule version state. This is achieved through proactively refreshing the pattern definition and rolling back the rule configuration. By tracing the rule configuration status, the impact of mode version drift on the stability of the hybrid architecture is reduced, and a list of mismatched items and a list of suspicious renaming items are generated to explicitly list missing rule references and potential naming differences, providing a basis for subsequent manual review or configuration correction. Document tasks are marked as having insufficient structural constraints, causing the task to enter a restricted execution mode at the system level. The input evidence anchoring execution module executes only the conservative strategy of the evidence extraction path or generates a conflict item log, avoiding the triggering of unstable large language model inference under structural uncertainty conditions while ensuring the traceability of semantic evidence, and avoiding inconsistencies in audit traces caused by silent rule skipping.

[0046] When the pattern rule mismatch risk value is less than the pattern rule mismatch risk threshold, the pattern and rule reference are determined to be consistent. The rule detection path is executed according to the rule code data. Under the premise that the structure and rule reference relationship have been confirmed to be consistent, the rule detection result is allowed to participate in the subsequent mixed quality control process as a deterministic judgment basis. The consistency verification result is created and archived in the pattern consistency database to record the stable correspondence between the pattern definition and the rule configuration at the current point in time, so as to support the reuse of subsequent results and audit reproducibility. Then, the evidence anchoring execution module is entered, and the collaborative reasoning closed loop of knowledge graph and large language model is completed under the condition that the structural constraints are clear.

[0047] In this implementation plan, by introducing a pre-judgment and diversion control mechanism based on the risk value of pattern rule mismatch, the hybrid architecture quality control system can identify and isolate rule configurations with insufficient structural constraints before the rule detection and execution stage. This avoids unstable structured judgment and semantic reasoning results caused by pattern evolution or rule reference drift during the whole medical record quality control process. When structural consistency is established, the rule detection path and evidence anchoring execution module are released in an orderly manner, ensuring that rule-level conclusions and large language model interpretations are generated collaboratively under stable structural constraints. This provides reliable protection for the reuse, audit traceability and long-term consistency of quality control results, while reducing the risk of repeated reasoning and operation caused by the failure of implicit rules.

[0048] Specifically, the process of conducting span-drift evidence anchoring consistency analysis on the output results of the evidence extraction pathway is as follows: The evidence anchoring quadruple is aggregated and counted according to the data of quality control focus node names and pattern definition node attribute labels to obtain the distribution probability. This probability is then normalized using the total number of evidence anchoring quadruples in this round. The evidence anchoring distribution entropy is then obtained using the Shannon entropy calculation algorithm. By modeling the distribution state of evidence at different nodes and attribute labels using information entropy, the concentration and uncertainty level of the evidence anchoring results are characterized, thus providing a quantitative basis for the evidence dispersion risk in subsequent consistency assessments. The process involves analyzing the evidence anchoring quadruple and the fixed evidence candidate set, which consists of evidence fragments formed by rule-driven text fragment extraction and node attribute mapping based on inspection result text data, test result values, and document node content sequence data. The intersection-union-comparison (IUCN) algorithm calculates the evidence overlap. By comparing the evidence selected in the current reasoning with the available evidence set across the entire medical record, it measures whether the evidence selection stably falls within the high-confidence candidate space, thereby reducing the probability of accidental evidence being adopted during the reasoning process. The longest common substring localization algorithm is executed on the explanatory text output by the large language model in two adjacent rounds in the document node content sequence data. The difference between the start and end positions of characters is calculated and normalized to obtain the explanatory text span drift intensity. The two adjacent rounds refer to the two reasoning outputs corresponding to the same quality control task under two consecutive save triggers. By quantifying the change in the location of the explanatory text in the document structure, it characterizes whether the semantic landing point of the large language model undergoes structural shift in multi-round reasoning, thereby providing a calculable basis for identifying explanation instability or context drift.

[0049] The reciprocal of the evidence anchoring distribution entropy plus one is calculated to obtain the distribution entropy suppression term. By suppressing the dispersion of evidence anchoring results across node and attribute dimensions, anchoring results with higher evidence concentration receive higher weight in consistency assessment. The difference between the evidence overlap degree of the current round and the evidence overlap degree of the previous round is calculated, the absolute value is taken, the inverse is taken, and then an exponential operation is performed to obtain the evidence overlap index term. By comparing the consistency degree of the evidence candidate set in adjacent reasoning rounds, a cross-round consistency constraint is introduced to ensure the continuity of evidence selection in multiple reasoning processes. The interpretation is calculated. The drift index is obtained by taking the ratio of the text span drift intensity to the explanatory text span drift intensity plus one, taking the opposite value, and then performing an exponential operation. This exponential index term is used to nonlinearly suppress the change in the positioning span of the explanatory text within the document node content, ensuring that the explanatory text maintains a relatively stable anchoring position in multiple rounds of generation. The evidence anchoring consistency value is obtained by multiplying the distribution entropy suppression term, the evidence overlap index term, and the drift index term. By jointly modeling the evidence distribution stability, cross-round evidence consistency, and explanatory positioning stability, a comprehensive quantitative index of the reliability of evidence anchoring is formed. The specific calculation formula is as follows: ; In the formula, The evidence anchoring consistency value measures the degree of consistency between the current round of large language model inference results and historical inference results in terms of evidence selection and positioning. It represents the evidence anchoring distribution entropy, reflecting the degree of dispersion and central tendency of the evidence anchoring results at different nodes and attributes; It indicates the degree of evidence overlap, characterizing the matching degree between the selected set of evidence and the fixed set of candidate evidence in the current round of reasoning; It indicates the degree of overlap of evidence in the previous round, serving as a reference benchmark for cross-round consistency comparison, and is used to measure the stability of evidence selection as the reasoning rounds progress; It indicates the intensity of the explanatory text span drift, quantifying the change in the location of the explanatory text within the document node content between adjacent reasoning rounds.

[0050] This implementation plan constructs an evidence anchoring consistency value by jointly quantifying the evidence anchoring results across multiple dimensions, including distribution concentration, candidate space consistency, and cross-round interpretation positioning stability. This enables the system to objectively determine whether the evidence selection and semantic landing point remain stable and convergent during the reasoning process of the large language model in a full medical record quality control scenario. This effectively identifies the risk of interpretation inconsistency caused by random generation, context drift, or evidence dispersion, and provides clear numerical basis for whether to trigger constrained re-inference, restricted interpretation write-back, or execution result reuse. This ensures that the structured conclusions of the knowledge graph and the semantic interpretation of the large language model are consistent, reproducible, and auditable under conditions of multi-round reasoning, multiple saving, and cross-document joint quality control.

[0051] Specifically, the process of secondary reasoning and verification based on the results of the span drift evidence anchoring consistency analysis is as follows: real-time comparison of evidence anchoring consistency values ​​and evidence anchoring consistency thresholds: When the evidence anchoring consistency value is less than the evidence anchoring consistency threshold, a constrained secondary inference execution is triggered for the current quality control task. This secondary inference execution includes: truncating the context content of the documents involved in the inference process according to length rules within the evidence extraction pathway to fix the context range; truncating the context content forward and backward, not exceeding a certain character length range, centered on the document node corresponding to the location node data output by the current rule detection pathway; and limiting the inference context by surrounding the key document nodes determined by the rule detection pathway to reduce interference from irrelevant medical record fragments on the inference path of the large language model. The character length range is determined based on the statistical distribution of the text span corresponding to the evidence anchoring quadruples in historical quality control tasks. The auxiliary evidence data, consisting of 300 to 1000 characters, is deterministically sorted according to chronological order and rule citation order. A stable evidence input order is established for inspection, testing, and treatment records on the timeline to avoid introducing uncertainty when multiple sources of medical record data participate in inference simultaneously. When calling the large language model, the same generation parameter configuration as the first inference is used to reduce the impact of randomness on the evidence anchoring result. If the evidence anchoring consistency value is still less than the evidence anchoring consistency threshold after the second inference, an evidence anchoring insufficiency marker is output, the task is transferred to the review queue, and only the structured conclusions and positioning nodes of the rule detection path are written back, without writing back the free interpretation text, to avoid the difficulty of deduplication of abnormal results and inconsistency in audit traces caused by inconsistent interpretation.

[0052] When the evidence anchoring consistency value is greater than or equal to the evidence anchoring consistency threshold, the explanatory text and evidence anchoring quadruple are allowed to be written back along with the quality control conclusion. Under the premise of stable evidence anchoring, the semantic interpretation across document nodes is output as part of the whole medical record quality control result and an abnormal record deduplication identifier is generated. The abnormal record deduplication identifier is generated by hash digest in the following order: first concatenating the unique identifier of the inpatient document, then concatenating the document version fingerprint data, then concatenating the hit rule code data, and finally concatenating the evidence anchoring quadruple. By introducing a unified abnormal uniqueness identification mechanism at the whole medical record scale, the same abnormality is avoided from being recorded repeatedly in multiple saves or multi-document associated quality control. It is used to uniquely identify and deduplicate the same abnormal result and write it into the violation information entity.

[0053] This implementation plan introduces a judgment and triage control mechanism centered on the consistency value of evidence anchoring. This enables the whole medical record quality control system to dynamically constrain the stability of evidence and the reliability of interpretation at the critical stages of reasoning involving the large language model. When evidence anchoring is insufficient, controlled re-inference and interpretation suppression are used to prevent the spread of unstable semantic results. When evidence anchoring is sufficient, the unified writing back and archiving of structured conclusions and semantic interpretations are allowed. Ultimately, this achieves unique identification, deduplication storage, and traceable auditing of abnormal results at the whole medical record level. This ensures that the knowledge graph structured judgment and the large language model semantic interpretation maintain consistency, reproducibility, and system-level credibility in multi-round reasoning, multiple saving, and cross-document joint quality control scenarios.

[0054] Specifically, the process of deduplicating and archiving cross-version results is as follows: The deduplication results of quality control conclusions and abnormal records are written and merged. Through unified archiving management of structured conclusions and semantic interpretations across the entire medical record, a converged view of quality control results across documents and versions is formed. When a deduplication marker for an abnormal record already exists, only the document's save timestamp and reuse hit record are appended to characterize the recurrence of the same abnormality in different save events. When a deduplication marker for an abnormal record does not exist, a violation information entity is written and bound to the rule detection result as a structured trace of the knowledge graph pathway. Simultaneously, evidence anchoring quadruples and explanatory text are bound as semantic traces of the large language model pathway. The pattern definition version number and rule code data are archived to the pattern consistency database, providing a reproducible versioned basis for long-term auditing, review, and retrospective analysis of the entire medical record's quality control results, achieving reproducible auditing and closed-loop review under a hybrid architecture.

[0055] In this implementation plan, by uniformly writing and merging the deduplication markers of quality control conclusions and abnormal records, the whole medical record quality control system achieves centralized aggregation and consistent evidence storage of structured judgments and semantic interpretations at the result output level. This avoids the same abnormality being repeatedly recorded or presented in a scattered manner during multiple storage or cross-document quality control processes. It also establishes a clear distinction mechanism between the continuous recurrence of abnormalities and the first appearance of new abnormalities. This ensures that the results generated by the knowledge graph pathway and the large language model pathway have a unified audit basis that is traceable, alignable, and verifiable throughout the whole medical record. Ultimately, it forms a stable quality control result management system that supports long-term operation, historical retrospection, and closed-loop review.

[0056] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0057] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A full medical record quality control system based on a hybrid architecture of knowledge graph and LLM, characterized in that, include: The data acquisition and preprocessing module is used to collect and save trigger document data and auxiliary evidence data, obtain quality control rule configuration data and knowledge graph pattern definition data, and perform preprocessing. The equivalence determination module is used to perform quality control trigger equivalence determination on the saved trigger document data and quality control rule configuration data, determine whether the saved event constitutes an equivalent trigger, and decide whether to start the rule detection path and evidence extraction path and reuse the existing quality control results. The consistency discrimination module is used to perform pattern rule mismatch analysis based on quality control rule configuration data and knowledge graph pattern definition data. Based on the pattern rule mismatch analysis results, it analyzes the consistency status between rule references and pattern definitions, and implements the execution of rule detection paths and entry restriction handling strategies. The evidence anchoring execution module is used to perform span drift evidence anchoring consistency analysis on the output results of the evidence extraction path, and to perform secondary reasoning and review processes based on the span drift evidence anchoring consistency analysis results; The result reuse, deduplication, and audit archiving module is used to perform deduplication and merging of results across versions for traceable archiving.

2. The full medical record quality control system based on a hybrid architecture of knowledge graph and LLM as described in claim 1, characterized in that: The specific process for collecting and saving trigger document data and auxiliary evidence data, and obtaining quality control rule configuration data and knowledge graph pattern definition data is as follows: Collect and save trigger document data, which includes: unique identifier of the document in the medical visit, medical visit serial number data, document saving timestamp, document submission status data, and full text content of the document; Obtain the quality control rule configuration data, which includes: rule code data, quality control attention node name data, and quality control attention node field identifier; Obtain knowledge graph pattern definition data, which includes: unique identifier of pattern definition, pattern definition version number, pattern definition node type data, pattern definition node attribute label data, node attribute prompt information data, and pattern definition relationship constraint set; Collect supporting evidence data, which includes: textual data of inspection results and numerical data of test results.

3. The full medical record quality control system based on a hybrid architecture of knowledge graph and LLM as described in claim 1, characterized in that: The specific process of preprocessing is as follows: The full text content of the document is segmented into nodes to obtain document node name sequence data and document node content sequence data; uniform text normalization processing is performed on the document node content sequence data, and corresponding document version fingerprint data is generated; Disambiguation and alignment are performed on duplicate content and positional offsets in the document node content sequence data to obtain the content span sequence data within the node; Consistency correction and missing completion processing are performed on document saving timestamps and document submission status data to obtain saved event sequence data; a distribution standardization algorithm is used to standardize and map saved trigger document data, quality control rule configuration data, knowledge graph pattern definition data, and auxiliary evidence data; and a min-max linear normalization algorithm is used to normalize the test result values ​​and document saving timestamps.

4. The full medical record quality control system based on a hybrid architecture of knowledge graph and LLM as described in claim 1, characterized in that: The specific process for determining the equivalence of quality control triggers between the saved trigger document data and the quality control rule configuration data is as follows: The number of save triggers is obtained by using a sliding count algorithm on the saved event sequence data; the character count algorithm is used on the full text content data of the document to obtain the length of the current and previous versions of the document, and the difference is calculated to obtain the document length difference. The previous version refers to the version immediately preceding the current save event after sorting by document save timestamp under the unique identifier of the document within the same medical visit; the number of entries in the content span sequence data within a node is counted, and the difference between adjacent versions is calculated to obtain the content span quantity difference; the set intersection algorithm is performed on the quality control focus node name data with the current document node name sequence data and the previous version document node name sequence data, and the difference between adjacent versions is calculated to obtain the quality control node quantity difference; the adjacent time intervals are obtained by sorting the document save timestamps by the unique identifier of the document within the medical visit; the median of the historical intervals is obtained by statistically analyzing the time interval (in seconds) sequence of the unique identifier of the document within the medical visit in the saved event sequence data. Calculate the natural logarithm of the number of saved triggers plus one to obtain the intensity logarithm term; calculate the absolute value of the document length difference plus one and divide by the full text length of the previous version of the document plus one to obtain the change ratio term; calculate the product of the intensity logarithm term and the change ratio term and take the opposite to obtain the equivalent input quantity; perform an exponential operation on the equivalent input quantity to obtain the equivalence index term; calculate the reciprocal of the absolute value of the content span quantity difference plus one to obtain the content suppression term; calculate the reciprocal of the absolute value of the quality control node quantity difference plus one to obtain the node suppression term; calculate the ratio of the adjacent time interval to the median of the historical interval, add one and take the reciprocal to obtain the time suppression term; multiply the equivalence index term, content suppression term, node suppression term, and time suppression term to obtain the quality control trigger equivalence judgment value.

5. The full medical record quality control system based on a hybrid architecture of knowledge graph and LLM as described in claim 1, characterized in that: The specific process for determining whether a saved event constitutes an equivalent trigger, and deciding whether to activate the rule detection pathway, the evidence extraction pathway, and reuse existing quality control results, is as follows: Real-time comparison of quality control trigger equivalence judgment value and quality control trigger equivalence judgment threshold: When the quality control trigger equivalence judgment value is less than the quality control trigger equivalence judgment threshold, it is judged as non-equivalence trigger. Static consistency verification is performed on the quality control attention node name data and the quality control attention node field identifier, and the rule detection path and evidence extraction path are started. When the quality control trigger equivalence judgment value is greater than or equal to the quality control trigger equivalence judgment threshold, it is determined to be an equivalence trigger. The rule detection path and the large language model evidence extraction path are stopped. The hit rule code data, defect level data, location node data and evidence anchoring quadruple are reused. The reuse is subject to the condition that the corresponding mode definition version number is consistent with the rule configuration. In the rule detection pathway, structured rule detection is performed on the document node content sequence data based on the knowledge graph pattern definition data and the quality control focus node name data, outputting the hit rule code data, defect level data, and location node data, and generating the corresponding quality control conclusions; in the evidence extraction pathway, auxiliary evidence data is screened based on the medical visit serial number data, the auxiliary evidence data is deterministically sorted, and a large language model is called to extract and interpret the evidence, outputting the evidence anchoring quadruple.

6. The full medical record quality control system based on a hybrid architecture of knowledge graph and LLM as described in claim 1, characterized in that: The specific process of performing pattern rule mismatch analysis based on quality control rule configuration data and knowledge graph pattern definition data is as follows: For the quality control attention node name data and the quality control attention node field identifier, an existence matching counting algorithm is performed on the pattern definition node type data and the pattern definition node attribute label data, respectively, to obtain the number of missing entries. The sum of these numbers forms the rule reference missing count. Then, the number of entries with successful existence matching is counted to obtain the rule reference hit count. For the missing quality control attention node name data and the quality control attention node field identifier, a minimum edit distance normalization algorithm is performed. The similarity corresponding to the minimum distance is taken as the renaming suspicion. The existence matching of the quality control attention node field identifier in the pattern definition node attribute label data is performed in a way that the identifiers are completely identical. For fields that do not match successfully, the renaming suspicion is calculated again using the minimum edit distance normalization algorithm. The missing rule reference count is calculated as the ratio of the rule reference hit count plus one to obtain the missing ratio term. The missing ratio term is then inversely calculated and exponentially calculated to obtain the missing decay term. The renaming suspicion value is calculated as the reciprocal of the renaming suppression term. The missing decay term and the renaming suppression term are multiplied to obtain the match preservation term. The match preservation term is then subtracted from the missing ratio to obtain the pattern rule mismatch risk value.

7. The full medical record quality control system based on a hybrid architecture of knowledge graph and LLM as described in claim 1, characterized in that: The specific process of analyzing the consistency status between rule references and pattern definitions based on the pattern rule mismatch analysis results, and executing the rule detection path and entering the restricted processing strategy is as follows: Real-time comparison of pattern rule mismatch risk value and pattern rule mismatch risk threshold: When the pattern rule mismatch risk value is greater than or equal to the pattern rule mismatch risk threshold, the execution of the rule detection corresponding to the rule code is suspended, the knowledge graph pattern definition cache refresh and rule configuration rollback check are triggered, and the document task is marked as having insufficient structural constraints, and then enters the evidence anchoring execution module. When the pattern rule mismatch risk value is less than the pattern rule mismatch risk threshold, the rule detection path will be executed according to the rule code data, and the consistency verification result will be created and archived to the pattern consistency database, and then enter the evidence anchoring execution module.

8. The full medical record quality control system based on a hybrid architecture of knowledge graph and LLM as described in claim 1, characterized in that: The specific process of performing span drift evidence anchoring consistency analysis on the output results of the evidence extraction pathway is as follows: The evidence anchoring quadruple is aggregated and counted according to the data of quality control focus node name and pattern definition node attribute label data to obtain the distribution probability. It is then normalized with the total number of evidence anchoring quadruples in this round, and the evidence anchoring distribution entropy is obtained by Shannon entropy calculation algorithm. The evidence overlap is calculated by performing a set intersection and union ratio algorithm on the evidence anchoring quadruple and the fixed evidence candidate set, which is composed of the evidence fragment set formed by the text data of the inspection results, the numerical data of the inspection results, and the content sequence data of the document nodes through rule-driven text fragment extraction and node attribute mapping. The longest common substring positioning algorithm is performed on the explanatory text output by the large language model in two adjacent rounds in the content sequence data of the document nodes to calculate the difference between the start and end positions of the characters and normalize it to obtain the span drift intensity of the explanatory text. The two adjacent rounds refer to the two inference outputs corresponding to the same quality control task under two consecutive save triggers. Calculate the reciprocal of the evidence anchoring distribution entropy plus one to obtain the distribution entropy suppression term; calculate the difference between the current round of evidence overlap and the previous round of evidence overlap and take the absolute value to obtain the change amplitude term; take the opposite of the change amplitude term and perform an exponential operation to obtain the evidence overlap index term; calculate the ratio of the explanatory text span drift intensity to the explanatory text span drift intensity plus one to obtain the drift ratio term; take the opposite of the drift ratio term and perform an exponential operation to obtain the drift index term; multiply the distribution entropy suppression term, the evidence overlap index term, and the drift index term to obtain the evidence anchoring consistency value.

9. A full medical record quality control system based on a hybrid architecture of knowledge graph and LLM as described in claim 1, characterized in that: The specific process of conducting secondary reasoning and verification based on the anchoring consistency analysis results of the span drift evidence is as follows: Real-time comparison of evidence anchoring consistency value and evidence anchoring consistency threshold: When the evidence anchoring consistency value is less than the evidence anchoring consistency threshold, a constrained secondary inference is triggered for the current quality control task. If the evidence anchoring consistency value is still less than the evidence anchoring consistency threshold after the secondary inference, an evidence anchoring insufficiency flag is output. When the evidence anchoring consistency value is greater than or equal to the evidence anchoring consistency threshold, the explanatory text and the evidence anchoring quadruple are written back along with the quality control conclusion, and an abnormal record deduplication identifier is generated. The abnormal record deduplication identifier is generated by hash digest according to the concatenation order and written into the violation information entity.

10. A full medical record quality control system based on a hybrid architecture of knowledge graph and LLM as described in claim 1, characterized in that: The specific process of deduplicating cross-version results and archiving them in a traceable manner is as follows: The process of writing and merging deduplication results for quality control conclusions and deduplication markers of abnormal records is as follows: When a deduplication marker for an abnormal record exists, only the document storage timestamp and reuse hit record are appended; when a deduplication marker for an abnormal record does not exist, the violation information entity is written and the rule detection result is bound as a structured trace of the knowledge graph pathway, and the evidence anchoring quadruple and explanation text are bound as semantic traces of the large language model pathway. The pattern definition version number and rule code data are archived to the pattern consistency database.