Intelligent medical data processing method and system based on knowledge graph and model reasoning
Patent Information
- Application Number
- CN202610889727.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-06-18
AI Technical Summary
然而由于医学知识体系自身具有庞杂交叉且高度非线性的固有属性,在构建包含海量实体关系的庞大医疗知识图谱后实体之间往往交织成错综复杂的网状结构,甚至存在大量隐性的循环依赖闭环
本发明提供了一种基于知识图谱与模型推理的智慧医疗数据处理方法,对医学文档解析提取三元组构建带有初始置信度属性的有向基础图谱,并从目标患者病例中精准锚定目标医学实体,基于最小开销优先原则计算最优连通路径,并对偏离路径的非最优关联关系边执行逻辑阻塞断开处理,将存在循环依赖闭环的图谱降维重构为以目标医学实体为源点的有向无环推理树,有效排除了冗余交叉路径的干扰,从根本上避免了逻辑死循环导致的推理偏差,提高了推理方向的准确性。同时,本发明在推理树中分别提取正向支持路径与反向排斥路径,并针对逻辑层级跳数配置衰减补偿系数执行均衡补偿计算,量化并弥补了多跳长路径推理带来的逻辑信息衰减,使得生成的最终置信度更加精准。此外,当最终置信度处于模糊阈值域时,通过构建并比对标准特征布尔向量与实际特征布尔向量,定位缺失的医疗实体维度并生成补充建议信息。该设计赋予了系统在信息不充分条件下主动引导补充验证的能力,大幅提升了整体诊断建议的可靠性。
Smart Images

Figure CN122417265B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical data processing technology, specifically relating to a smart medical data processing method and system based on knowledge graphs and model reasoning. Background Technology
[0002] To extract valuable information from massive amounts of medical literature and electronic medical records, constructing medical knowledge graphs and performing logical reasoning based on them is a core development direction for intelligent healthcare systems. Knowledge graphs can structurally represent complex medical concepts and their relationships, theoretically providing doctors with intuitive knowledge references. However, due to the inherently complex, intersecting, and highly nonlinear nature of medical knowledge systems, after constructing a vast medical knowledge graph containing massive entity relationships, entities often intertwine into intricate network structures, even containing numerous implicit circular dependencies. When the system performs path searching and multi-step logical reasoning based on the patient's physical characteristics within the graph, the algorithm is prone to getting stuck in redundant calculations and path loops, causing the reasoning path to deviate from the correct diagnostic direction. Ultimately, when faced with complex and ever-changing clinical case data, existing intelligent healthcare data processing methods generally suffer from low accuracy and reliability of reasoning results. Summary of the Invention
[0003] This invention provides a smart healthcare data processing method and system based on knowledge graphs and model reasoning to solve the above-mentioned technical problems.
[0004] In a first aspect, the present invention provides a smart healthcare data processing method based on knowledge graphs and model reasoning, the method comprising the following steps: The obtained complete medical documents are parsed and logically cleaned to extract medical triples containing medical entities and the relationships between them. A directed basic graph is then constructed based on the medical triples, where the relationships have an initial confidence attribute. Receive the case data of the target patient, extract medical feature entities from the case data through a natural language model, and anchor the target medical entity that is most closely related to the pathological state of the target patient in the directed basic graph based on the medical feature entities. Global path retrieval is performed in the directed basic graph based on the target medical entity. The topological path cost is determined according to the initial confidence of the association relationship. The optimal connectivity path from each associated entity node in the directed basic graph to the target medical entity is calculated based on the minimum cost priority principle. Logical blocking and disconnection processing is performed on non-optimal association edges that deviate from the optimal connected path in the directed basic graph. The dimensionality of the directed basic graph with cyclic dependency loops is reduced and reconstructed into a directed acyclic reasoning tree with the target medical entity as the source point. In the directed acyclic inference tree, forward support paths and reverse rejection paths pointing to the terminal medical entity are extracted along the topological hierarchy. Attenuation compensation coefficients are configured based on the logical hop count of the forward support paths and reverse rejection paths. According to the attenuation compensation coefficients, the initial confidence of the multi-hop paths in the forward support paths and reverse rejection paths is calculated. The final confidence of the terminal medical entity is generated by combining the forward and reverse compensation results. When the final confidence level is within the preset fuzzy threshold range, the supplementary suggestion set corresponding to the end medical entity in the directed basic graph is extracted. A standard feature Boolean vector is generated based on the supplementary suggestion set. The standard feature Boolean vector is compared with the actual feature Boolean vector composed of medical feature entities. The missing medical entity dimension is located based on the difference between the standard feature Boolean vector and the actual feature Boolean vector. Supplementary suggestion information for doctors to refer to is generated based on the missing medical entity dimension.
[0005] Optionally, the step of parsing and logically cleaning the obtained complete medical document, extracting medical triples containing medical entities and the relationships between them, and constructing a directed basic graph based on the medical triples includes the following steps: The layout analysis algorithm is used to preprocess the input complete medical document to obtain a clean text stream; The pure text stream is truncated into multiple text stream segments using a preset semantic sliding window, and the text stream segments are then input into a pre-trained medical big language model. Using a medical big language model, implicit medical subjects in text stream paragraphs are explicitly completed based on the medical context, spelling errors are corrected, and enhanced text blocks are generated. A parallel extraction thread is started to identify medical entities and relationships in the enhanced text block, and a set of candidate medical triples consisting of subject, predicate and object is output. Calculate the initial extraction confidence of each candidate medical triple in the candidate medical triple set, filter out candidate medical triples whose initial extraction confidence is lower than a preset quality threshold, and obtain medical triples; By mapping medical triples to medical entity nodes and related edges in a graph database, a directed basic graph is constructed.
[0006] Optionally, the step of calculating the initial extraction confidence of each candidate medical triplet in the candidate medical triplet set, and filtering candidate medical triplets with initial extraction confidence below a preset quality threshold to obtain medical triplet sets includes the following steps: Scan all candidate medical triples and use a collision detection algorithm to identify conflicting triple pairs that have the same medical entity but whose relationship is mutually exclusive; For the identified conflicting triple pairs, extract the corresponding preceding and following time adverbs and condition adverbs in the complete medical document; Logical disambiguation is performed on conflicting triple pairs based on pre- and post-temporal adverbs and conditional adverbs, and the initial extraction confidence of the time dimension is redistributed for conflicting triple pairs based on the disambiguation results. Candidate medical triples that are still below the preset quality threshold after updating the initial extraction confidence level will be automatically discarded, and candidate medical triples that are above the preset quality threshold after updating the initial extraction confidence level will be converted into medical triples.
[0007] Optionally, receiving the target patient's case data, extracting medical feature entities from the case data using a natural language model, and anchoring the target medical entity most closely related to the target patient's pathological state in the directed basic graph based on the medical feature entities includes the following steps: Receive case data of target patients and perform word segmentation and stop word filtering on the case data to extract medical feature words containing symptoms, signs and past medical history; Medical feature terms are input into a natural language model, which then maps these terms into standardized medical feature entities. Extract vital sign parameters from case data and calculate the global importance score for each medical feature entity using a term frequency-inverse document frequency algorithm; The global importance score is weighted and fused with vital sign parameters to generate the core weight value of each medical feature entity; Select the target medical entity with the largest core weight value, and match the corresponding medical entity node in the directed basic graph with a unique identifier as the target medical entity.
[0008] Optionally, the step of performing global path retrieval in the directed basic graph based on the target medical entity, determining the topological path cost based on the initial confidence of the association, and calculating the optimal connectivity path from each associated entity node in the directed basic graph to the target medical entity based on the minimum cost priority principle includes the following steps: Broadcast path exploration probes from the target medical entity as the origin to adjacent nodes in the directed basic graph; Record each relationship edge traversed by the path exploration probe and extract the initial confidence level of each relationship edge; The topological path cost corresponding to each association edge is calculated based on the initial confidence level. For any associated entity node in the directed basic graph that is associated with the target medical entity, the topological path cost of all candidate paths from the associated entity node to the target medical entity through different intermediate nodes is accumulated to obtain the accumulated topological path cost. Compare the cumulative topological path costs of all candidate paths for the same associated entity node, and mark the candidate path with the smallest cumulative topological path cost as the optimal connectivity path from the associated entity node to the target medical entity.
[0009] Optionally, the step of calculating the topological path cost corresponding to each association edge based on the initial confidence level includes the following steps: Extract the starting entity node corresponding to each relationship edge in the directed basic graph, and count the total number of out-degree connections from the starting entity node to other different entity nodes; The logical divergence entropy of the starting entity node is calculated based on the total number of out-degree connections, and the logical divergence entropy is used to quantify the characteristic specificity of the associated edges in the medical reasoning process. Construct an overhead evaluation function that includes a specific penalty term, input the logic divergence entropy as the specific penalty term into the dynamic overhead evaluation function, and input the initial confidence level as the basic variable into the overhead evaluation function; The topological path cost of each associated edge, after feature-specific calibration, is calculated using the cost evaluation function.
[0010] Optionally, the steps of extracting forward support paths and reverse rejection paths pointing to the final medical entity along the topological hierarchy in the directed acyclic inference tree, configuring attenuation compensation coefficients based on the logical hop count of the forward support paths and reverse rejection paths, performing balanced compensation calculations on the initial confidence of multi-hop paths in the forward support paths and reverse rejection paths according to the attenuation compensation coefficients, and generating the final confidence of the final medical entity by combining the forward and reverse compensation results include the following steps: In a directed acyclic reasoning tree, identify all unidirectional connected branches that point to the terminal medical entity from all leaf nodes, and classify the unidirectional connected branches into positive support paths and negative repulsion paths according to medical logical attributes. Count the number of edges traversed from the leaf node to the terminal medical entity for each positive support path and negative rejection path, and define the number of edges as the logical level hop count of the path; Extract the basic network connectivity before the directed acyclic reasoning tree is generated as a global parameter, and generate a decay compensation coefficient that is positively correlated with the number of logical level hops based on the basic network connectivity. The initial confidence level after the cumulative attenuation of each positive support path and negative repulsion path is calculated with the attenuation compensation coefficient to obtain the single-path balanced confidence level. The total support is obtained by summing the single-path balanced confidence scores of all positive support paths in the forward direction, and the total rejection score is obtained by summing the single-path balanced confidence scores of all negative rejection paths in the same dimension. The final confidence level of the end medical entity is obtained by subtracting the total rejection from the total support and then normalizing the result.
[0011] Optionally, the formula for calculating the single-path equilibrium confidence level is as follows: In the formula, For single-path equilibrium confidence, The initial confidence level is the product of the cumulative decay of each positive support path and the negative repulsion path. is the attenuation compensation coefficient, n is the number of logical level hops, e is the natural constant, and tanh is the hyperbolic tangent function.
[0012] In a second aspect, the present invention also provides a smart medical data processing system based on knowledge graphs and model reasoning, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the smart medical data processing method based on knowledge graphs and model reasoning as described in any one of the first aspects.
[0013] Thirdly, the present invention also provides a computer-readable storage medium storing instructions that, when executed by a processor, configure the processor to perform the intelligent medical data processing method based on knowledge graph and model reasoning according to any one of the first aspects.
[0014] The beneficial effects of this invention are: This invention provides a smart healthcare data processing method based on knowledge graphs and model reasoning. It extracts triples from medical documents to construct a directed basic graph with initial confidence attributes. It then accurately anchors target medical entities from patient cases, calculates the optimal connectivity path based on the minimum cost priority principle, and performs logical blocking to disconnect non-optimal association edges that deviate from the path. This dimensionality reduction and reconstruction of the graph with cyclic dependencies into a directed acyclic reasoning tree with the target medical entity as the source point effectively eliminates interference from redundant cross paths, fundamentally avoiding reasoning bias caused by logical dead loops and improving the accuracy of the reasoning direction. Simultaneously, this invention extracts positive support paths and negative rejection paths from the reasoning tree and performs balanced compensation calculations based on the number of hops at logical levels, quantifying and compensating for the logical information attenuation caused by multi-hop long path reasoning, resulting in a more accurate final confidence score. Furthermore, when the final confidence score is in the fuzzy threshold domain, it constructs and compares standard feature Boolean vectors with actual feature Boolean vectors to locate missing medical entity dimensions and generate supplementary suggestion information. This design gives the system the ability to proactively guide supplementary verification when information is insufficient, which greatly improves the reliability of the overall diagnostic recommendations. Attached Figure Description
[0015] Figure 1This is a flowchart illustrating a smart healthcare data processing method based on knowledge graphs and model reasoning in one embodiment of this application.
[0016] Figure 2 This is a flowchart illustrating the construction process of a directed basic graph in one embodiment of this application. Detailed Implementation
[0017] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0018] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0019] Figure 1 This is a flowchart illustrating a smart healthcare data processing method based on knowledge graphs and model reasoning in one embodiment. It should be understood that, although... Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps. For example Figure 1 As shown, the intelligent medical data processing method based on knowledge graphs and model reasoning disclosed in this invention specifically includes the following steps: S101. Parse and logically clean the obtained complete medical documents, extract medical triples containing medical entities and relationships between them, and construct a directed basic graph based on the medical triples, wherein the relationships have an initial confidence attribute.
[0020] The process involves preprocessing the input complete medical document using a layout analysis algorithm. This algorithm can identify different layout elements in the document, such as tables, images, and text blocks, and mark or remove non-text content to obtain a clean text stream. A clean text stream refers to a continuous text sequence after removing formatting marks, special symbols, and redundant whitespace characters. The clean text stream is then truncated into multiple text stream segments using a pre-defined semantic sliding window. The window size of the semantic sliding window is configured according to the input length limit of the medical language model, and overlapping areas exist between windows to avoid semantic truncation. After the text stream segments are input into the pre-trained medical language model, the model explicitly completes the implicit medical subjects in the text stream segments based on the medical context. For example, it completes symptom descriptions with omitted subjects into complete subject-verb-object structures, while also correcting spelling errors and non-standard terminology, generating enhanced text blocks. A parallel extraction thread is initiated to identify medical entities and their relationships in the enhanced text blocks. Medical entities include categories such as disease names, drug names, examination items, and symptoms / signs. Relationships include types such as causal relationships, accompanying relationships, and treatment relationships. The identification results are output as a set of candidate medical triples consisting of a subject, a predicate, and an object. The initial extraction confidence of each candidate medical triple in the set is calculated. This confidence is derived from a combination of the probability distribution of the model output and the accuracy of entity boundary recognition. Candidate medical triples with initial extraction confidence below a preset quality threshold are filtered out. The quality-filtered medical triples are mapped to medical entity nodes and relationship edges in a graph database. Medical entity nodes store the standardized name, type label, and attribute information of the entity, while relationship edges store the relationship type, direction, and initial confidence attributes, thus constructing a directed basic graph.
[0021] S102. Receive the case data of the target patient, extract medical feature entities from the case data through a natural language model, and anchor the target medical entity most closely related to the pathological state of the target patient in the directed basic graph based on the medical feature entities.
[0022] After receiving the case data of the target patients, the data is first processed through word segmentation and stop word filtering. Word segmentation divides the continuous case text into independent lexical units, while stop word filtering removes function words and conjunctions that contribute little to medical semantics, extracting medical feature words containing symptoms, signs, and medical history. These medical feature words are then input into a natural language model, which, based on a medical ontology and standard terminology system, maps them to standardized medical feature entities. For example, it converts colloquial symptom descriptions into standardized medical terminology. Vital signs, including numerical indicators such as body temperature, blood pressure, heart rate, and respiratory rate, are extracted from the case data. These parameters quantify the severity of the patient's physiological state. A global importance score for each medical feature entity is calculated using a term frequency-inverse document frequency (IF-IVF) algorithm. IF reflects the frequency of the feature in the current case, while IVF reflects the rarity of the feature in the overall medical document set. The two are multiplied to obtain the global importance score. The global importance score is weighted and fused with vital sign parameters to generate core weight values for each medical feature entity. The weighting coefficients are differentiated according to the clinical importance of different medical feature entity types; for example, the weight coefficient for chief complaint symptoms is higher than that for general signs. The medical feature entity with the largest core weight value is selected as the target medical entity, and the corresponding medical entity node is matched in the directed basic graph using a unique identifier. The unique identifier typically adopts a standard medical coding system such as ICD coding or SNOMED coding to ensure the accuracy and uniqueness of entity matching.
[0023] S103. Perform global path retrieval in the directed basic graph based on the target medical entity, determine the topological path cost based on the initial confidence of the association, and calculate the optimal connectivity path from each associated entity node in the directed basic graph to the target medical entity based on the minimum cost priority principle.
[0024] The path exploration probe is a virtual traversal marker used to record path information during graph traversal. It records each relational edge traversed by the probe and extracts the initial confidence score for each edge, derived from attribute values assigned during the knowledge graph construction phase. Based on the initial confidence score, the topological path cost corresponding to each relational edge is calculated. This calculation first extracts the starting entity node corresponding to each relational edge in the directed basic graph and counts the total number of out-degree connections from the starting entity node to other different entity nodes. The total number of out-degree connections reflects the connection complexity of that node in the graph. Based on the total number of out-degree connections, the logical divergence entropy of the starting entity node is calculated. This logical divergence entropy is quantified using the information entropy formula; a larger value indicates a more dispersed relational relationship. The logical divergence entropy is used to quantify the characteristic specificity of relational edges in the medical reasoning process. A cost evaluation function incorporating a specificity penalty term is constructed. Logical divergence entropy is input as the specificity penalty term into the cost evaluation function, along with the initial confidence level as a basic variable. The cost evaluation function calculates the topological path cost of each association edge after feature-specific calibration. For any associated entity node in the directed basic graph that is linked to the target medical entity, the topological path costs of all candidate paths from the associated entity node to the target medical entity through different intermediate nodes are accumulated to obtain the accumulated topological path cost. The accumulated topological path costs of all candidate paths for the same associated entity node are compared, and the candidate path with the smallest accumulated topological path cost is marked as the optimal connected path from the associated entity node to the target medical entity.
[0025] S104. Perform logical blocking and disconnection processing on the non-optimal association edges that deviate from the optimal connected path in the directed basic graph, reduce the dimensionality of the directed basic graph with cyclic dependency loops, and reconstruct it into a directed acyclic reasoning tree with the target medical entity as the source.
[0026] The specific implementation of the logic blocking disconnection process involves traversing all the associated edges in the directed basic graph and determining whether each edge belongs to the optimal connected path from any associated entity node to the target medical entity. If it does not, the edge is marked as a non-optimal edge and removed from the active topology of the graph. By removing non-optimal edges, the directed basic graph, which originally contained cyclic dependency loops, is reduced in dimensionality. Cyclic dependency loops refer to circular paths in the graph that can cause the reasoning process to fall into an infinite loop. The reduced-dimensional graph structure no longer contains any path that starts from a node, goes through several edges, and returns to that node, thus satisfying the topological properties of a directed acyclic graph. The reconstruction process uses the target medical entity as the source node, retains all nodes and edges in the optimal connected paths, and forms a tree structure with the target medical entity as the root node. In this tree structure, each node has at most one edge pointing to its parent node, and all edges point to the root node, forming a directed acyclic reasoning tree. The hierarchical structure of the directed acyclic reasoning tree reflects the logical reasoning distance between medical entities. Nodes closer to the target medical entity have stronger reasoning relevance. The leaf nodes of the tree represent the starting entities of the reasoning chain, and the root node is the target medical entity. The reconstructed directed acyclic reasoning tree retains the most critical reasoning paths in the original directed basic graph while eliminating redundancy and circular dependencies.
[0027] S105. Extract the forward support path and reverse rejection path pointing to the terminal medical entity along the topological hierarchy in the directed acyclic inference tree. Configure the attenuation compensation coefficient based on the logical hop count of the forward support path and the reverse rejection path. Perform equalization compensation calculation on the initial confidence of the multi-hop path in the forward support path and the reverse rejection path according to the attenuation compensation coefficient. Generate the final confidence of the terminal medical entity by combining the forward and reverse compensation results.
[0028] In this process, positive support paths and negative rejection paths pointing to the final medical entity are extracted along the topological hierarchy in the directed acyclic inference tree. A positive support path refers to a path in which the associations support the establishment of the final medical entity, while a negative rejection path refers to a path in which the associations exclude or negate the establishment of the final medical entity. All unidirectional connected branches pointing from leaf nodes to the final medical entity are identified and classified into positive support paths and negative rejection paths according to medical logical attributes. The classification is based on the semantic type of the association edges; for example, relationships such as "cause" and "trigger" belong to positive support, while relationships such as "exclude" and "prohibit" belong to negative rejection. The number of edges traversed from the leaf node to the final medical entity for each positive support path and negative rejection path is counted, and this number is defined as the logical level hop count of the path, reflecting the length of the inference chain. The basic network connectivity before the generation of the directed acyclic inference tree is extracted as a global parameter. The basic network connectivity is obtained by calculating the average node degree or graph density of the directed basic graph, and this parameter reflects the overall complexity of the knowledge graph. A decay compensation coefficient, positively correlated with the number of hops in the logical hierarchy, is generated based on the basic network connectivity. This coefficient balances the confidence contributions of long and short paths, preventing short paths from dominating the final result. The initial confidence after cumulative decay of each forward support path and reverse rejection path is balanced with the decay compensation coefficient to obtain the single-path balanced confidence. The formula is the hyperbolic tangent function applied to the product of the cumulative confidence and the exponential compensation term. The total support is obtained by summing the single-path balanced confidence of all forward support paths in the forward direction, and the total rejection is obtained by summing the single-path balanced confidence of all reverse rejection paths in the same dimension. The final confidence of the terminal medical entity is obtained by subtracting the total rejection from the total support and then normalizing the result.
[0029] S106. When the final confidence level is within the preset fuzzy threshold range, extract the supplementary suggestion set corresponding to the end medical entity in the directed basic graph, generate a standard feature Boolean vector based on the supplementary suggestion set, compare the standard feature Boolean vector with the actual feature Boolean vector composed of medical feature entities, locate the missing medical entity dimension based on the difference between the standard feature Boolean vector and the actual feature Boolean vector, and generate supplementary suggestion information for doctors to refer to based on the missing medical entity dimension.
[0030] When the final confidence level falls within a preset fuzzy threshold range, it indicates that the current evidence is insufficient to make a definitive judgment, and more information is needed. The fuzzy threshold range is typically set to a confidence level between 0.4 and 0.6, where the confidence level is neither sufficient to confirm nor exclude the terminal medical entity. A supplementary suggestion set corresponding to the terminal medical entity is extracted from the directed basic knowledge graph. This supplementary suggestion set contains all medical entities related to the terminal medical entity that may affect the judgment; these entities have direct or indirect connections with the terminal medical entity in the knowledge graph. A standard feature Boolean vector is generated based on the supplementary suggestion set. This standard feature Boolean vector is a Boolean array, where each element corresponds to a medical entity in the supplementary suggestion set. A true value indicates that the entity should exist, and a false value indicates that the entity does not affect the judgment. The standard feature Boolean vector is compared with an actual feature Boolean vector composed of medical feature entities. The actual feature Boolean vector is generated based on medical feature entities extracted from the case data; entities present in the case are considered true, and those not present are considered false. Boolean discrepancy sites are extracted from the standard feature Boolean vector where the state value is true and the corresponding state value in the actual feature Boolean vector is false. The medical entities corresponding to these sites are the missing entities. These Boolean discrepancy sites are mapped to a directed basic graph to generate a candidate missing entity set. For each candidate missing entity in the candidate missing entity set, a bi-state inference branch is constructed in the directed acyclic inference tree, presenting the candidate missing entity as support and rejection states. Based on the bi-state inference branch, the expected confidence of the final medical entity under the bi-state is simulated and calculated, and the information entropy gain of the corresponding candidate missing entity is calculated according to the preset state prior probability in the directed basic graph. The candidate missing entity set is sorted in descending order according to the magnitude of the information entropy gain. The highest gain item is accumulated one by one using a greedy algorithm until the accumulated expected result makes the final confidence jump out of the fuzzy threshold domain. The optimal missing entity combination corresponding to the final confidence jumping out of the fuzzy threshold domain is extracted, and the dimension of the optimal missing entity combination is taken as the missing medical entity dimension. The missing medical entity dimension is then converted into supplementary suggestion information for doctors' reference.
[0031] In one implementation, the process of parsing and logically cleaning the obtained complete medical document to extract medical triples containing medical entities and the relationships between them, and constructing a directed basic graph based on the medical triples, includes the following steps: The layout analysis algorithm is used to preprocess the input complete medical document to obtain a clean text stream; The pure text stream is truncated into multiple text stream segments using a preset semantic sliding window, and the text stream segments are then input into a pre-trained medical big language model. Using a medical big language model, implicit medical subjects in text stream paragraphs are explicitly completed based on the medical context, spelling errors are corrected, and enhanced text blocks are generated. A parallel extraction thread is started to identify medical entities and relationships in the enhanced text block, and a set of candidate medical triples consisting of subject, predicate and object is output. Calculate the initial extraction confidence of each candidate medical triple in the candidate medical triple set, filter out candidate medical triples whose initial extraction confidence is lower than a preset quality threshold, and obtain medical triples; By mapping medical triples to medical entity nodes and related edges in a graph database, a directed basic graph is constructed.
[0032] In this embodiment, refer to Figure 2 The layout analysis algorithm first scans the medical document, identifying different layout element types, including text blocks, table areas, image areas, headers and footers, annotation boxes, and other structured components. The algorithm uses a deep learning-based object detection model to divide the page into regions, extracts layout features through a convolutional neural network, and uses bounding box regression technology to accurately locate the spatial position of each layout element. For table areas, the algorithm identifies the row and column structure of the table and extracts cell content, converting the structured table data into a linear text sequence with logical delimiters. For image areas, the algorithm marks the image content and inserts placeholders into the text stream, retaining the image's reference position information but not including the image pixel data in the text stream. The algorithm strips the document of formatting markup information, including font styles, color attributes, paragraph indentation, and other typesetting control symbols, while removing special symbols such as tabs, page breaks, and redundant whitespace characters. After layout analysis processing, a clean text stream is obtained. A clean text stream refers to a continuous sequence of characters that retains only semantic content. The text stream maintains the original document's logical order and paragraph boundaries, but removes all formatting information irrelevant to content understanding.
[0033] The size of the semantic sliding window is configured based on the maximum input length of the pre-trained medical language model, typically ranging from 512 to 2048 characters, with the specific value depending on the model's positional encoding capabilities and computational resource limitations. As the sliding window moves across the clean text stream, it extracts a fixed-length character sequence as a text stream segment at a time, with the window movement step being smaller than the window size, resulting in overlapping regions between adjacent text stream segments. The length of the overlapping region is usually set to 10% to 30% of the window size. This overlap design aims to prevent semantic truncation by window boundaries, ensuring that medical entities and relationships crossing window boundaries appear completely in at least one text stream segment. When performing window truncation, the algorithm prioritizes segmentation at sentence or paragraph boundaries. By recognizing sentence terminators such as periods, question marks, and exclamation marks, as well as paragraph separators such as line breaks, the window boundaries are adjusted to the nearest natural segmentation point, avoiding truncation in the middle of words or within medical terms. The process of explicitly completing implicit medical subjects in text stream segments using the medical language model based on the medical context leverages the model's semantic understanding and generation capabilities. Medical documents often employ expressions that omit the subject. For example, when describing a series of symptoms, the patient is mentioned only once, and subsequent sentences directly use verbs or adjectives to describe the symptoms. Medical large language models analyze the context of text stream segments to identify sentence structures missing subjects and infer the missing subject based on previously mentioned subject entities.
[0034] The completion operation is performed at the model's internal representation level. The model reconstructs sentences with implicit subjects into complete subject-verb-object structures, enabling subsequent entity recognition and relation extraction to accurately locate sentence components. The spelling error correction function addresses potential input errors, OCR recognition errors, and non-standard terminology spelling issues in medical documents. Based on a medical vocabulary and language model probability distribution, the model identifies erroneous words that have an edit distance from standard medical terminology and replaces them with the most probable correct form. The edit distance calculation uses the Levenstein distance algorithm, quantifying the similarity between words by statistically analyzing the minimum number of insertion, deletion, and replacement operations. The generated enhanced text blocks supplement missing grammatical components and correct spelling errors while maintaining the original text's semantic content.
[0035] The process of identifying medical entities and their relationships in enhanced text blocks employs a multi-threaded concurrent processing architecture to improve processing efficiency. The number of parallel extraction threads is dynamically configured based on available computing resources and the number of enhanced text blocks. Each thread independently processes one or more enhanced text blocks, with no data dependencies between threads, allowing for fully parallel execution. Medical entity recognition utilizes a sequence labeling-based named entity recognition model. The model labels each lexical unit in the enhanced text block as an entity type tag, using the BIO tagging format. B tags indicate the start position of an entity, I tags indicate its internal position, and O tags indicate non-entity positions. Entity types include core medical concept categories such as diseases, symptoms, drugs, examinations, treatments, and anatomical locations, with each category corresponding to a set of B and I tags. Relationship recognition builds upon entity recognition, analyzing the syntactic dependencies and semantic roles between entity pairs to determine whether specific medical relationships exist between them. The types of associations include causal relationships, accompanying relationships, treatment relationships, and examination relationships. The relationship recognition model uses a graph neural network-based relationship classifier, representing sentences as a graph structure with entities as nodes and dependency arcs as edges. Node features are aggregated and edge relationship types are predicted through graph convolution operations. The output set of candidate medical triples consists of three elements: subject, predicate, and object. The subject and object are the identified medical entities, and the predicate is the association type between entity pairs. Each triple is accompanied by an initial extraction confidence score, which is jointly calculated from the entity recognition confidence score and the relationship recognition confidence score.
[0036] Calculating the initial extraction confidence for each candidate medical triplet in the candidate medical triplet set requires considering the influence of multiple factors. The formula for calculating the initial extraction confidence is as follows: In the formula, The initial extraction confidence level for candidate medical triples, The confidence level for identifying the main entity. The confidence level for identifying object entities, To determine the confidence level for identifying associations, Scoring the accuracy of entity boundaries, , , The weighting coefficients are and satisfy the following conditions: Entity recognition confidence is derived from the softmax probability distribution output by the named entity recognition model, and the geometric mean of the probabilities of all lexical labels for the entity is taken as the entity-level confidence. Relationship recognition confidence is derived from the relationship type probability output by the relationship classifier, and the probability value corresponding to the predicted relationship type is used directly. Entity boundary accuracy is calculated by comparing the consistency between the entity boundary position and the syntactic component boundary; entities with perfectly aligned boundaries score 1, while the score for entities with deviated boundaries decreases linearly according to the degree of deviation.
[0037] The graph database uses an attribute graph model, supporting nodes and edges carrying any number of attribute key-value pairs. The creation process for medical entity nodes first checks if an entity node with the same identifier already exists in the graph database. The entity identifier is a unique identifier generated by combining the entity's standardized name and type. If the entity node already exists, its attribute information is updated, including increasing the entity's occurrence count in the document and updating the entity's contextual description. If the entity node does not exist, a new node is created and its attributes are set, including metadata such as entity name, entity type, standard code, and first occurrence position. The creation process for relationship edges establishes a directed edge between the subject node and the object node, with the edge pointing from the subject to the object, reflecting the directionality of the relationship. Relationship edges carry a relationship type attribute and an initial confidence attribute; the value of the initial confidence attribute is the initial extraction confidence of the triple. When multiple relationship edges of the same type exist between the same pair of entities, a confidence-weighted fusion strategy is used to merge them into a single edge. The formula for calculating the fused confidence is: In the formula, Here, m represents the initial confidence level after fusion, and m is the number of association edges to be fused. Let be the initial extraction confidence score for the j-th association edge. The constructed directed basic graph is stored in the form of a graph database.
[0038] In one implementation, calculating the initial extraction confidence of each candidate medical triplet in the candidate medical triplet set, and filtering candidate medical triplets with initial extraction confidence below a preset quality threshold to obtain medical triplets includes the following steps: Scan all candidate medical triples and use a collision detection algorithm to identify conflicting triple pairs that have the same medical entity but whose relationship is mutually exclusive; For the identified conflicting triple pairs, extract the corresponding preceding and following time adverbs and condition adverbs in the complete medical document; Logical disambiguation is performed on conflicting triple pairs based on pre- and post-temporal adverbs and conditional adverbs, and the initial extraction confidence of the time dimension is redistributed for conflicting triple pairs based on the disambiguation results. Candidate medical triples that are still below the preset quality threshold after updating the initial extraction confidence level will be automatically discarded, and candidate medical triples that are above the preset quality threshold after updating the initial extraction confidence level will be converted into medical triples.
[0039] In this implementation, the collision detection algorithm first constructs an index for the candidate medical triplet set, using the combination of subject and object entities as the index key to cluster triples with the same entity pair into the same index bucket. For each triplet set within an index bucket, the algorithm extracts the association type of each triplet and queries a pre-built mutual exclusion rule library to determine whether there is logical mutual exclusion between the relationship types. The mutual exclusion rule library defines relationship pairs that cannot be simultaneously valid in the medical field, such as mutual exclusion between facilitating and inhibiting relationships, mutual exclusion between increasing and decreasing relationships, and mutual exclusion between indication and contraindication relationships. The algorithm traverses all triplet pairs within the index bucket. For each triplet pair, it extracts the association type between the two and searches the mutual exclusion rule library for a matching mutual exclusion rule. If a matching mutual exclusion rule is found, the triplet pair is marked as a conflicting triplet pair and the conflict type is recorded. The conflict detection process also considers the temporal and conditional attributes of entities. If two triplet entities have the same name but different time ranges or conditional constraints, they are not considered to be in conflict.
[0040] For identified conflicting triplet pairs, the process of extracting the corresponding pre- and post-temporal adverbs and conditional adverbs in the complete medical document employs syntactic analysis and semantic role labeling techniques. Temporal adverbs are phrases or words expressing time concepts, including absolute time expressions such as specific dates, relative time expressions such as before and after treatment, and time period expressions such as lasting three months. Conditional adverbs are phrases or words expressing conditions or preconditions, usually introduced by conjunctions such as "if," "when," or "under certain circumstances." The extraction process first locates the sentence position corresponding to the conflicting triplet pair in the complete medical document, using document offset information recorded during triplet generation. Dependency parsing is performed on the located sentences to construct a dependency tree structure. Nodes in the dependency tree represent words, and edges represent syntactic dependency relationships between words. Dependency arcs labeled as temporal or conditional adverbs are searched in the dependency tree, and the lexical subtrees connected by these arcs are extracted as adverbial components. The identification of time adverbs is based on part-of-speech features such as time nouns, time adverbs, and time prepositional phrases, as well as syntactic features indicating a time modification relationship with the predicate in the dependency tree. The identification of condition adverbs is based on grammatical features such as conditional conjunctions, hypothetical mood markers, and conditional clause structures. The extracted pre- and post-time and condition adverbs are normalized into structured representations; time adverbs are converted into standard formats for time intervals or time points, and condition adverbs are converted into logical conditional expressions. For adverbial components spanning multiple sentences, the algorithm expands the search scope to adjacent sentences, determining the scope of the adverb through reference resolution and discourse coherence analysis.
[0041] The logical disambiguation algorithm first compares the time adverbs of conflicting triplet pairs to determine whether the two triples describe medical facts in different time periods. If the time adverbs of the two triples represent non-overlapping time intervals, the two triples are considered to coexist in the time dimension and do not constitute a true logical conflict; the disambiguation result is time separation. If the time adverbs overlap or are missing, the algorithm further compares the condition adverbs to determine whether the two triples are true under different conditions. The comparison of condition adverbs is performed through the satisfiability analysis of logical expressions. If the two condition expressions are mutually exclusive, the triples are considered to coexist in the condition dimension; the disambiguation result is condition separation. For conflicting triplet pairs that have neither time separation nor condition separation, the algorithm adopts an evidence strength comparison strategy, calculating the quantity and quality of supporting evidence for each triple in the document. Supporting evidence includes factors such as repetition frequency, contextual consistency, and source authority. Based on the disambiguation result, the initial extraction confidence of the time dimension is redistributed for the conflicting triplet pairs. The redistribution formula is: In the formula, The initial extraction confidence level after reallocation, The conflict penalty coefficient ranges from 0.2 to 0.5. The number of triplets that conflict with the current triplet. This represents the total number of triples in the index bucket. Triples separated by time or condition are not penalized with confidence level and retain their original confidence level. constant.
[0042] The quality threshold filtering operation iterates through all candidate medical triples that have undergone confidence reassignment, comparing the initial extraction confidence of each triple after reassignment. The relationship between the value and the preset quality threshold. The preset quality threshold is configured based on the application scenario and quality requirements of the knowledge graph. In applications requiring high-precision knowledge, the threshold is set to 0.8, and in applications requiring high-coverage knowledge, the threshold is set to 0.7. For Candidate medical triples below a preset quality threshold are marked as low-quality triples and removed from the candidate set. This removal does not affect the processing of the original document or other triples. For candidate medical triples that are at or above a preset quality threshold, the algorithm transforms their state from candidate to confirmed, generating official medical triples. The medical triples have the same data structure as the candidate medical triples, containing three core elements: subject, predicate, and object. However, they also include a quality verification tag and a timestamp attribute. The quality verification tag indicates that the triple has passed the quality threshold test, and the timestamp attribute records the triple's generation and last update times. The transformed set of medical triples constitutes the input data for knowledge graph construction.
[0043] In one implementation, receiving case data of a target patient, extracting medical feature entities from the case data using a natural language model, and anchoring the target medical entity most closely related to the target patient's pathological state in a directed basic graph based on the medical feature entities includes the following steps: Receive case data of target patients and perform word segmentation and stop word filtering on the case data to extract medical feature words containing symptoms, signs and past medical history; Medical feature terms are input into a natural language model, which then maps these terms into standardized medical feature entities. Extract vital sign parameters from case data and calculate the global importance score for each medical feature entity using a term frequency-inverse document frequency algorithm; The global importance score is weighted and fused with vital sign parameters to generate the core weight value of each medical feature entity; Select the target medical entity with the largest core weight value, and match the corresponding medical entity node in the directed basic graph with a unique identifier as the target medical entity.
[0044] In this implementation, case data typically exists in free text format, containing descriptive content across multiple sections, including chief complaint, present illness, past medical history, physical examination, and auxiliary examinations. Word segmentation divides the continuous case text into independent lexical units, employing either a maximum matching algorithm based on a medical dictionary or a segmentation algorithm based on a statistical model to ensure that medical terminology is not incorrectly segmented. Medical terms, such as multi-character disease names and drug names, need to be retained as complete units during the segmentation process. Pre-loading a medical domain dictionary enables priority identification and boundary protection of these terms. Stop word filtering removes function words that contribute little to medical semantics, including auxiliary words, conjunctions, and modal particles. The stop word list is customized based on the linguistic characteristics of medical texts, retaining special function words with medical significance, such as negation words and degree adverbs. When extracting medical feature words containing symptoms, signs and medical history, the keywords describing the patient's condition are located by using part-of-speech tagging and named entity recognition technology. Symptom words describe the patient's subjective feelings such as pain and fatigue, sign words describe objective examination findings such as swelling and tenderness, and medical history words describe the patient's historical medical information such as surgical history and allergy history.
[0045] The natural language model is pre-trained on a large-scale medical text corpus, learning the semantic representations and synonym relationships of medical terms, and can map medical concepts in different forms to unified standard entities. The mapping process first converts medical feature words into the model's input representation. A word embedding layer converts the words into high-dimensional vectors, where semantically similar words are closer together in the vector space. The encoder layer learns context-aware representations of the input vectors, capturing the specific meaning of words in the case text and eliminating ambiguity for polysemous words. The decoder layer maps the encoded semantic representations to entity identifiers in a standard medical ontology library. This ontology library uses internationally recognized medical coding systems such as ICD coding and SNOMEDCT coding to ensure the uniqueness and interoperability of entity identifiers. The mapping output is a standardized set of medical feature entities, each containing attribute information such as a standard name, coded identifier, and entity type.
[0046] Vital signs parameters are extracted from case data, including basic physiological indicators such as body temperature, pulse, respiration, and blood pressure, as well as laboratory test values such as blood routine and biochemical indicators. These parameters are extracted from structured fields or semi-structured tables in the case data. Parameter extraction employs regular expression matching and template parsing techniques to identify numerical data and their corresponding units and reference ranges. The extracted parameter values are compared with normal reference ranges, and the degree of deviation is calculated as an indicator of abnormality severity. The term frequency-inverse document frequency (IF-IVF) algorithm calculates the global importance score of medical feature entities. Term frequency reflects the number of times an entity appears in the current case, while IVF reflects the rarity of the entity in the overall case database. The formula for calculating the global importance score is: In the formula, For the first Global importance score of each medical feature entity. The frequency of this entity's occurrence in the current cases, The total number of words in the current case. The total number of cases in the case database. The formula represents the number of cases containing the entity. It avoids bias from long cases through word frequency normalization and highlights rare but important medical features through inverse document frequency. Next, vital sign parameters are normalized, mapping parameters of different dimensions and numerical ranges to a standard interval of 0 to 1. The normalization method uses min-max normalization or deviation normalization based on a reference range. For each medical feature entity, the vital sign parameters associated with that entity are identified. These associations are established using an entity-parameter mapping table in a medical knowledge base; for example, fever is associated with body temperature parameters, and anemia is associated with hemoglobin parameters. The degree of normalization abnormality of the associated parameters is used as a weight for physiological severity, along with the global importance score. Weighted fusion is performed, and the fusion formula is as follows: In the formula, For the first The core weight value of each medical feature entity, The normalized abnormality level of the vital sign parameters associated with this entity. and The weighted coefficients are and satisfy the following conditions: , usually set and To balance the contributions of textual features and physiological indicators. For entities with no associated parameters, Setting it to the default value of 0.5 indicates medium importance. The core weight value comprehensively reflects the significance and clinical importance of the medical feature entity in the current case.
[0047] Select the medical feature entity with the largest core weight value, and then iterate through all medical feature entities to select the core weight value. The system compares and identifies the entity with the highest numerical value, which is considered the core feature most representative of the target patient's current pathological state. When multiple entities have the same maximum weight value, a priority rule based on entity type is used for selection, with the priority order being chief complaint symptoms over physical examination findings over past medical history. The selected medical feature entities are matched against nodes in the directed basic graph using standard coded identifiers. The matching operation performs an exact query in the graph database, retrieving the corresponding medical entity node based on the entity's unique identifier. The unique identifiers use an international standard medical coding system to ensure the accuracy and uniqueness of the matching, avoiding erroneous matching due to terminological ambiguity. After a successful match, the retrieved medical entity node is marked as the target medical entity, serving as the root node for subsequent global path retrieval and inference tree construction. The position of the target medical entity in the directed basic graph determines the direction and scope of the inference analysis. Path exploration starting from the target medical entity can discover medical knowledge and inference chains related to the patient's state, providing knowledge support for decision support.
[0048] In one implementation, global path retrieval is performed in the directed basic graph based on the target medical entity. The topological path cost is determined according to the initial confidence of the association relationship, and the optimal connectivity path from each associated entity node in the directed basic graph to the target medical entity is calculated based on the minimum cost priority principle, including the following steps: Broadcast path exploration probes from the target medical entity as the origin to adjacent nodes in the directed basic graph; Record each relationship edge traversed by the path exploration probe and extract the initial confidence level of each relationship edge; The topological path cost corresponding to each association edge is calculated based on the initial confidence level. For any associated entity node in the directed basic graph that is associated with the target medical entity, the topological path cost of all candidate paths from the associated entity node to the target medical entity through different intermediate nodes is accumulated to obtain the accumulated topological path cost. Compare the cumulative topological path costs of all candidate paths for the same associated entity node, and mark the candidate path with the smallest cumulative topological path cost as the optimal connectivity path from the associated entity node to the target medical entity.
[0049] In this embodiment, the path exploration probe is a virtual traversal marker carrying attributes such as source node information, traversed path information, and cumulative cost information. The broadcast operation begins with the target medical entity node. First, the target medical entity node is added to the queue to be visited and marked as visited. The target medical entity node is then removed from the queue, and all outgoing edges of that node in the directed basic graph are queried. The nodes pointed to by these outgoing edges are the adjacent nodes. For each adjacent node, a path exploration probe instance is created. The probe records the path information from the target medical entity to that adjacent node, including the sequence of nodes traversed and the sequence of edges. The adjacent node is added to the queue to be visited, and the same broadcast operation is performed on the next node removed from the queue. The broadcast process proceeds hierarchically: the first layer consists of the direct adjacent nodes of the target medical entity, the second layer consists of the adjacent nodes of the nodes in the first layer, and so on, until all reachable nodes in the directed basic graph have been traversed. To avoid getting stuck in an infinite loop in the circular structure of the graph, the broadcast probe is not repeated for nodes that have already been visited.
[0050] When a path exploration probe moves from one node to an adjacent node, it needs to traverse the edge connecting the two nodes. This edge stores the relationship type and initial confidence attribute in the directed basic graph. The recording operation adds the edge identifier traversed by the probe to the probe's path sequence. The edge identifier uniquely identifies each edge in the graph and is typically generated by combining the starting node identifier, ending node identifier, and relationship type. When extracting the initial confidence, the attribute information of that edge is queried from the graph database based on the edge identifier, and the initial confidence attribute value is retrieved. The initial confidence originates from the quality assessment of medical triples during the knowledge graph construction phase, reflecting the reliability of the relationship. The value ranges from 0 to 1, with values closer to 1 indicating higher confidence. The probe associates and stores the extracted initial confidence with the corresponding edge identifier, forming an edge-confidence mapping table. For multiple edges traversed by the probe, the mapping table records the confidence information of each edge on the complete path. When the path contains multi-hop edges, the overall reliability of the path is affected by the confidence of all edges on the path; edges with lower confidence will reduce the overall reliability of the path.
[0051] Topological path cost is a core concept in path optimization algorithms; a smaller cost value indicates better path quality. The calculation process first extracts the starting entity node corresponding to the associated edges in the directed basic graph. The total number of out-degree connections from the starting entity node to other different entity nodes is then counted. This total number of out-degree connections reflects the node's connection complexity and information divergence. Based on the total number of out-degree connections, the logical divergence entropy of the starting entity node is calculated. Logical divergence entropy is calculated using the information entropy method, with the formula being the negative logarithm-weighted sum of the probabilities of each outgoing edge. The outgoing edge probabilities are obtained by normalizing the initial confidence of the outgoing edges. A larger logical divergence entropy value indicates a more dispersed association relationship between the node, and a lower specificity of the reasoning path originating from that node. A cost evaluation function containing a specificity penalty term is constructed, using logical divergence entropy as the specificity penalty term and the initial confidence as the basic variable. The form of the cost evaluation function is: In the formula, The topological path cost of the associated edges. The divergence penalty coefficient ranges from 0.1 to 0.3. Let be the logical divergence entropy of the starting entity node. This formula achieves the conversion of high confidence to low cost through the reciprocal of the confidence level, and reduces the path priority of outgoing edges from high divergence nodes through the divergence entropy penalty term.
[0052] Associative entity nodes are nodes visited by the path exploration probe during graph traversal, and these nodes are connected to the target medical entity through one or more paths. For each associated entity node, the algorithm enumerates all possible paths from that node to the target medical entity, and the path information is extracted from the path exploration probe's records. Each candidate path consists of a series of continuous edges, and the topological path cost of the path is equal to the sum of the topological path costs of all edges on the path. The accumulation operation calculates the total path cost for all candidate paths of the same associated entity node, and sums the edge costs of each path. The cumulative topological path cost is obtained by adding them sequentially. For paths containing... The formula for calculating the cumulative topological path cost for a path with edges is: In the formula, The cumulative topology path cost for candidate paths, For the first on the path The topological path cost of each edge. After the accumulation operation, each associated entity node corresponds to a set of candidate paths. Each path in the set has a corresponding accumulated topological path cost value. These values are used for subsequent path comparison and optimal path selection. The accumulated topological path cost comprehensively reflects the path length, the reliability of the associations on the path, and the specificity of the nodes traversed by the path.
[0053] The comparison operation is performed independently for each associated entity node, extracting all candidate paths corresponding to that node and their accumulated topological path costs. Find out by numerical comparison The minimum cost path represents the most reliable and specific reasoning chain from associated entity nodes to the target medical entity, exhibiting optimal connectivity within the knowledge graph's topology. The labeling operation records the selected optimal path in a path index table, using the associated entity node identifier as the key and the edge sequence and accumulated topological path cost of the optimal path as values. For cases where multiple paths share the same minimum cost, path length is used as a secondary sorting criterion, prioritizing shorter paths with fewer edges, as these have more direct logical connections during reasoning. The set of optimal connected paths forms the backbone of the directed basic graph; non-optimal paths are logically blocked, simplifying graph complexity and eliminating redundancy and circular dependencies in the reasoning process. The process of determining the optimal connected path ensures that reasoning analysis focuses on the most reliable knowledge connections, improving the accuracy and interpretability of the reasoning results.
[0054] In one implementation, calculating the topological path cost corresponding to each association edge based on the initial confidence level includes the following steps: Extract the starting entity node corresponding to each relationship edge in the directed basic graph, and count the total number of out-degree connections from the starting entity node to other different entity nodes; The logical divergence entropy of the starting entity node is calculated based on the total number of out-degree connections, and the logical divergence entropy is used to quantify the characteristic specificity of the associated edges in the medical reasoning process. Construct an overhead evaluation function that includes a specific penalty term, input the logic divergence entropy as the specific penalty term into the dynamic overhead evaluation function, and input the initial confidence level as the basic variable into the overhead evaluation function; The topological path cost of each associated edge, after feature-specific calibration, is calculated using the cost evaluation function.
[0055] In this implementation, the edges of the relationships in the directed graph have a clear directionality, pointing from the starting entity node to the ending entity node. The starting entity node is the source node of the edge. The extraction operation traverses the path exploration probe to record all the edges of the relationships. For each edge, the identifier of the starting node is read, and the corresponding entity node is located in the graph database using the identifier. When counting the total number of out-degree connections, all outgoing edges from the starting entity node are queried. Outgoing edges refer to the relationships that point from that node to other nodes. The statistical process calculates the number of outgoing edges, i.e., how many different target nodes the node connects to. The total number of outgoing degrees reflects the breadth of connections of the node in the knowledge graph. Nodes with high outgoing degrees are usually general concepts or high-frequency entities in the medical field, such as common symptoms or basic examination items. These nodes have relationships with a large number of other entities. Nodes with low outgoing degrees are usually specific medical concepts or rare entities. The relationships of these nodes are more specific and particular. The statistical results are stored in the form of a node-outgoing degree mapping table. The mapping table records the identifier of each starting entity node and its corresponding total number of outgoing degrees. This mapping table is repeatedly queried and used in subsequent logical divergence entropy calculations. The statistics of the total number of out-degree connections provide a direct numerical basis for quantifying the degree of information divergence of nodes and are a key indicator for evaluating the specificity of association relationships.
[0056] Logical divergence entropy borrows from the concept of information entropy to measure the uncertainty and dispersion of the distribution of node relationships. The calculation process first obtains the total number of out-degree connections of the starting entity node, assuming the node has... There are several outgoing edges, each connecting to a different target node. Extract the initial confidence score for each outgoing edge. The initial confidence scores of all outgoing edges are normalized to obtain the probability distribution of each outgoing edge. The normalization formula is the outgoing edge confidence score divided by the sum of the confidence scores of all outgoing edges. Based on the normalized probability distribution, the logical divergence entropy is calculated using the following formula: In the formula, For the first Normalized probability of outgoing edges. Logical divergent entropy. The entropy value depends on the total number of outgoing edges. It is highest when all outgoing edges have equal probabilities and lowest when the probability of one outgoing edge is close to 1 while the probabilities of others are close to 0. High entropy indicates highly dispersed node relationships, and the reasoning path originating from that node lacks specificity. Low entropy indicates concentrated node relationships, and the reasoning path has a clear direction. When using logical divergence entropy to quantify feature specificity, the entropy value is used as a penalty factor. Additional overhead is applied to the outgoing edges of highly divergent nodes in the path cost calculation, reducing the priority of these edges in optimal path selection, thereby guiding path search towards more specific relationships.
[0057] The cost evaluation function is designed as a multi-factor weighted function, with the initial confidence scores of the association edges as the basic variables. The initial confidence level reflects the reliability of the association; a higher confidence level indicates a more trustworthy relationship. The specificity penalty term is the logical divergence entropy of the starting entity node. The penalty term is used to adjust path cost to reflect the specificity of the association. The cost evaluation function uses a reciprocal transformation to convert confidence into a cost metric, because high confidence should correspond to low cost. It also introduces a linear penalty term for divergent entropy, and the function has the following form: In the formula, This is the divergence penalty coefficient, used to control the intensity of the specificity penalty. Its value ranges from 0.1 to 0.3, with larger values preferred. The value significantly increases the outgoing edge cost of highly divergent nodes. The multiplicative structure in the function ensures the combined effect of confidence and specificity; when the confidence is low or the divergence entropy is high, the topological path cost... As the value increases, the edge's priority in path selection decreases. The design of the cost evaluation function balances the reliability of knowledge with the specificity of reasoning, avoiding path search from falling into general but non-specific associations and improving the clinical relevance of the reasoning results.
[0058] The topological path cost of each association edge, after feature-specific calibration, is calculated using a cost evaluation function. The calculation operation is performed independently for each association edge traversed by the path exploration probe. First, the initial confidence of the edge is extracted. This value is already assigned to the edge's attributes during the knowledge graph construction phase. Then, the logical divergence entropy of the entity node at the edge's starting point is extracted. ,Will and Substitute the cost evaluation function, combined with the preset divergence penalty coefficient. Calculate the topology path cost The calculation results are stored in an edge-cost mapping table, with edge identifiers as keys and topological path costs as values, for subsequent path accumulation and comparison operations. The topological path cost, calibrated for feature specificity, comprehensively reflects both the reliability and specificity of the association. Compared to a simple inverse transformation based solely on confidence, the calibrated cost can more accurately assess the actual value of a path in medical reasoning. For edges connecting highly divergent nodes, even with high initial confidence, the topological path cost increases due to the divergence entropy penalty, causing the path search algorithm to tend to select more targeted reasoning paths.
[0059] In one implementation, forward support paths and reverse rejection paths pointing to the final medical entity are extracted along the topological hierarchy in the directed acyclic inference tree. Attenuation compensation coefficients are configured based on the logical hop count of the forward support paths and reverse rejection paths. Balanced compensation calculations are performed on the initial confidence of multi-hop paths in the forward support paths and reverse rejection paths according to the attenuation compensation coefficients. The final confidence of the final medical entity is generated by combining the forward and reverse compensation results, including the following steps: In a directed acyclic reasoning tree, identify all unidirectional connected branches that point to the terminal medical entity from all leaf nodes, and classify the unidirectional connected branches into positive support paths and negative repulsion paths according to medical logical attributes. Count the number of edges traversed from the leaf node to the terminal medical entity for each positive support path and negative rejection path, and define the number of edges as the logical level hop count of the path; Extract the basic network connectivity before the directed acyclic reasoning tree is generated as a global parameter, and generate a decay compensation coefficient that is positively correlated with the number of logical level hops based on the basic network connectivity. The initial confidence level after the cumulative attenuation of each positive support path and negative repulsion path is calculated with the attenuation compensation coefficient to obtain the single-path balanced confidence level. The total support is obtained by summing the single-path balanced confidence scores of all positive support paths in the forward direction, and the total rejection score is obtained by summing the single-path balanced confidence scores of all negative rejection paths in the same dimension. The final confidence level of the end medical entity is obtained by subtracting the total rejection from the total support and then normalizing the result.
[0060] In this implementation, after graph dimensionality reduction and reconstruction, the directed acyclic reasoning tree forms a tree structure with the target medical entity as the root node. Leaf nodes represent the initial evidence entities in the reasoning chain, and terminal medical entities represent the reasoning conclusion entities whose confidence level needs to be evaluated. A unidirectional connected component refers to a complete path starting from a leaf node, sequentially passing through intermediate nodes along the tree's edges, and finally reaching the terminal medical entity. This path is unique and unidirectional in the directed acyclic tree. The identification operation traverses all leaf nodes in the tree. For each leaf node, a depth-first search or breadth-first search is performed, tracing upwards along the edges pointing to the parent node until the terminal medical entity node is reached. The tracing process records all nodes and edges traversed, forming a complete unidirectional connected component. The classification operation judges based on the medical logical attributes of the relational edges in the unidirectional connected component. Medical logical attributes include causal relationships, accompanying relationships, exclusionary relationships, taboo relationships, etc. A positive support path refers to a path where the relational relationship supports the establishment of the terminal medical entity. A negative exclusion path refers to a path where the relational relationship excludes or negates the establishment of the terminal medical entity.
[0061] The logical hop count reflects the reasoning distance from the evidence entity to the conclusion entity. Fewer hops indicate more direct reasoning, while more hops indicate more intermediate steps involved. The statistical operation is performed independently for each unidirectional connected component, traversing the edge sequence recorded in the path and calculating the number of edges, which is the logical hop count for that path. For components containing... The path of an edge, the logical level hop count is defined as follows: In a directed acyclic inference tree (DAI), due to the hierarchical structure of the tree, the path length from a leaf node to the root node is equal to the depth of the tree containing the leaf node. The statistical result of the logical level hop count is appended to the attribute information of each path. The hop count information plays a crucial role in subsequent attenuation compensation calculations because the confidence level of multi-hop paths accumulates and decays due to multiple inference passes, requiring correction through a compensation mechanism. Basic network connectivity refers to the graph connectivity measure of a directed basic graph before dimensionality reduction and reconstruction into a directed acyclic inference tree, reflecting the overall density and complexity of the knowledge graph. Connectivity calculation methods include average node degree, graph density, and the number of connected components. The average node degree is obtained by dividing the sum of the in-degree and out-degree of all nodes by the total number of nodes, and the graph density is obtained by dividing the actual number of edges by the maximum possible number of edges.
[0062] The extraction of basic network connectivity involves retrieving pre-calculated values from the metadata of the graph database or performing connectivity analysis on the original graph after the inference tree is constructed. Basic network connectivity, as a global parameter, reflects the richness of relationships between medical entities in the knowledge graph. Higher connectivity indicates a complex knowledge network and relatively high reliability of multi-hop paths, while lower connectivity indicates a sparse knowledge network and potentially greater uncertainty in multi-hop paths. The attenuation compensation coefficient is generated based on the joint relationship between basic network connectivity and the number of hops in the logical hierarchy. The compensation coefficient is designed as an exponential function positively correlated with the number of hops; the larger the number of hops, the larger the compensation coefficient, used to offset the confidence attenuation effect of multi-hop paths. The specific value of the compensation coefficient is adjusted through the basic network connectivity; higher connectivity results in greater compensation strength, and lower connectivity results in less compensation strength, achieving adaptive adjustment of the compensation mechanism to the characteristics of the knowledge graph.
[0063] The initial confidence level of each forward support path and reverse repulsion path after cumulative attenuation is calculated by combining it with the attenuation compensation coefficient to obtain the single-path balanced confidence level. Specifically, the initial confidence level after cumulative attenuation refers to the initial confidence level of all edges on the path. The results obtained by multiplying sequentially reflect the cumulative attenuation effect of confidence during multi-hop propagation. The overall reliability of the path is affected by the confidence of each edge traversed. For paths containing... The initial confidence score of a path with edges, after cumulative decay, is calculated as the sum of all edges on the path. The product of consecutive products is denoted as The attenuation compensation coefficient is denoted as... Based on the number of hops in the logical hierarchy The design of the compensation coefficient, along with the generation of basic network connectivity, allows multi-hop paths to gain additional confidence, partially offsetting the effect of cumulative attenuation. The formula for calculating the single-path equilibrium confidence is as follows: In the formula, For single-path equilibrium confidence, The initial confidence level is the product of the cumulative decay of each positive support path and the negative repulsion path. is the attenuation compensation coefficient, n is the number of logical hops, e is the natural constant, and tanh is the hyperbolic tangent function used to limit the result to the range of 0 to 1 to avoid numerical overflow. Single-path equilibrium confidence level. It comprehensively reflects the original reliability and structural complexity of the path, providing a fair confidence assessment benchmark for paths of different lengths.
[0064] The forward summation operation iterates through all paths in the set of forward support paths and extracts the single-path equilibrium confidence for each path. , all The values are summed arithmetically, and the sum is the total support, denoted as . Total support reflects the sum of the strengths of all evidence supporting the establishment of the terminal medical entity; a higher value indicates stronger supporting evidence. The same summation operation is performed on the set of reverse exclusion paths to obtain the total exclusion score, denoted as... The total rejection score reflects the sum of the strength of evidence against the medical entity at the end; a higher value indicates stronger rejection evidence. The summation operation uses a linear overlay method, assuming that the contributions of evidence from different paths are independent and can be directly accumulated. The process of subtracting the total rejection score from the total support score and then normalizing the result to obtain the final confidence score of the medical entity integrates the adversarial effect of both positive and negative evidence. The subtraction operation is calculated... The net support score is obtained. A positive net support score indicates that supporting evidence is dominant, while a negative net support score indicates that refutating evidence is dominant. The absolute value reflects the strength of the evidence's tendency. Normalization maps the net support score to a standard confidence interval of 0 to 1. Normalization is achieved using the sigmoid function or linear transformation. A commonly used normalization formula is: In the formula, This represents the final confidence level for the end-stage medical entity. The formula uses a sigmoid function to convert the net support of any real value into a probability value between 0 and 1. Much larger hour A value close to 1 indicates a high degree of confidence. much smaller hour A value close to 0 indicates a high degree of negation. and When close A score close to 0.5 indicates insufficient evidence. The final confidence score, as the core output of the reasoning analysis, quantifies the credibility of the end medical entity under the current evidentiary conditions.
[0065] In one implementation, when the final confidence level is within a preset fuzzy threshold range, a supplementary suggestion set corresponding to the terminal medical entity in the directed basic graph is extracted. A standard feature Boolean vector is generated based on the supplementary suggestion set. The standard feature Boolean vector is compared with the actual feature Boolean vector composed of medical feature entities. The missing medical entity dimension is located based on the difference between the standard feature Boolean vector and the actual feature Boolean vector. Supplementary suggestion information for doctors' reference is generated based on the missing medical entity dimension, including the following steps: When the final confidence level is within the preset fuzzy threshold range, the supplementary suggestion set corresponding to the terminal medical entity in the directed basic graph is extracted, and a standard feature Boolean vector is generated based on the supplementary suggestion set. The standard feature Boolean vector is compared with the actual feature Boolean vector composed of medical feature entities. Boolean difference sites with true state values in the standard feature Boolean vector and false state values in the actual feature Boolean vector are extracted. The Boolean difference sites are then mapped to the directed basic graph to generate a candidate missing entity set. For each candidate missing entity in the candidate missing entity set, construct a bi-state reasoning branch in the directed acyclic reasoning tree, where the state of the candidate missing entity is presented as support and rejection. Based on the bi-state reasoning branch, the expected confidence of the terminal medical entity under the bi-state is simulated and calculated respectively, and the information entropy gain of the corresponding candidate missing entity is calculated according to the preset state prior probability in the directed basic graph. The candidate missing entity set is sorted in descending order according to the magnitude of information entropy gain, and the highest gain item is accumulated one by one using a greedy algorithm until the expected result is accumulated and the final confidence jumps out of the fuzzy threshold domain. The best missing entity combination corresponding to the final confidence jumps out of the fuzzy threshold domain is then extracted. The dimension containing the optimal combination of missing entities is taken as the dimension of the missing medical entity, and the dimension of the missing medical entity is converted into supplementary suggestion information for doctors' reference.
[0066] In this embodiment, the fuzzy threshold domain refers to the final confidence level. The value falls within the middle range, where it cannot be confirmed or ruled out; this range is typically set between 0.4 and 0.6. This range indicates that the current evidence is insufficient to make a definitive judgment. The judgment operation compares the calculated values. With the upper and lower boundaries of the fuzzy threshold domain, if If the entity is located within the specified interval, the supplementary suggestion generation process is triggered. When extracting the supplementary suggestion set, all medical entities directly or indirectly related to the final medical entity are queried in the directed basic graph. These entities may be key factors influencing the judgment of the final medical entity, including related symptoms, signs, examination items, and past medical history. The extraction range of the supplementary suggestion set is controlled by setting the maximum number of hops in the graph traversal, typically limited to 2 to 3 hops to ensure the relevance of the suggestions. The supplementary suggestion set is organized as an entity list, with each entity accompanied by attribute information such as entity type, standard code, and clinical significance. When generating the standard feature Boolean vector, a Boolean array is created, with a length equal to the number of entities in the supplementary suggestion set. Each element of the array corresponds to a supplementary suggestion entity. The element value is assigned according to the standard requirements for judging the final medical entity in medical knowledge. If the entity is a necessary condition or important reference for judging the final medical entity, the corresponding element is assigned a value of true; otherwise, it is assigned a value of false.
[0067] The actual feature Boolean vector is generated based on medical feature entities extracted from the target patient's case data. The generation method involves traversing each entity in the supplementary suggestion set, checking if the entity exists in the extracted medical feature entity set, assigning a value of true to the corresponding position if it exists, and false if it does not. The actual feature Boolean vector reflects the existing information state in the current case data. The comparison operation performs an element-wise comparison between the standard feature Boolean vector and the actual feature Boolean vector, using logical AND and NOT operations to extract the difference sites, specifically the positions where the standard vector element is true and the corresponding element in the actual vector is false. The entities corresponding to these Boolean difference sites represent the missing but important medical information in the current case. The mapping operation converts the indices of the difference sites into corresponding entity identifiers in the supplementary suggestion set. The corresponding medical entity nodes are retrieved in the directed basic graph using these entity identifiers, and the retrieved node set is organized into a candidate missing entity set. Each entity in the candidate missing entity set represents a potential direction for information supplementation. Obtaining information from these entities may change the final confidence of the final medical entity, transforming the judgment from an ambiguous state to a definite state.
[0068] Bi-state inference branching refers to assuming two states for candidate missing entities: supporting the final medical entity and rejecting the final medical entity. For each state, a corresponding inference path is constructed. The construction operation first queries the directed base graph to find the association between the candidate missing entity and the final medical entity, identifying the path connecting them and the type of relationship. For the supporting branch, assuming the candidate missing entity is positive or exists, the entity is added as a new leaf node to the directed acyclic inference tree. Based on the positive association between the candidate missing entity and the final medical entity, an inference path from the candidate missing entity to the final medical entity is established in the tree. For the rejecting branch, assuming the candidate missing entity is negative or does not exist, a corresponding rejecting path is established in the tree based on the negative association between the candidate missing entity and the final medical entity. The construction process of bi-state inference branching generates two virtual inference scenarios for each candidate missing entity. Each scenario corresponds to a new inference tree structure, with the new tree adding inference paths related to the candidate missing entity to the original inference tree. The construction of bi-state branches allows the algorithm to predict the impact of obtaining different missing information on the final confidence level, providing a basis for prioritizing information acquisition.
[0069] Based on the bi-state inference branches, the expected confidence of the terminal medical entity in the bi-state is simulated and calculated. The simulation operation performs the same confidence calculation process as described above for each bi-state inference branch, including path identification, single-path equilibrium confidence calculation, summation of total support and total rejection, and final confidence normalization, to obtain the expected confidence in the support state. Expected confidence level under exclusionary state The prior probability of a state refers to the probability distribution of whether a candidate missing entity presents a supportive or exclusionary state in medical statistics. This probability is extracted from the statistical metadata of the directed basic graph or set according to medical epidemiological data. The prior probability of the supportive state is denoted as... The prior probability of the repulsive state is denoted as and satisfy The information entropy gain is calculated as the difference between the confidence entropy of the current fuzzy state and the expected confidence entropy after obtaining the missing entity information. The entropy value is calculated using the Shannon entropy formula. The formula for calculating the information entropy gain is as follows: , In the formula, This represents the information entropy gain of the candidate missing entity; a larger value indicates a greater contribution of obtaining information about that entity to eliminating uncertainty.
[0070] The candidate missing entity set is sorted in descending order according to the magnitude of information entropy gain. The descending sorting operation applies the information entropy gain to all entities in the candidate missing entity set. Sort the list from largest to smallest, placing the entities with the highest information value at the top. The greedy algorithm starts with the first entity in the sorted list and adds entities sequentially to the selected entity set. After each addition, a new inference tree containing all selected entities is constructed, and a new expected confidence score is calculated. This accumulation operation continues until the expected confidence score is reached. Jumping out of the fuzzy threshold region, i.e. A value greater than the upper boundary or less than the lower boundary indicates that the evidence is sufficiently clear. The set of selected entities that causes the confidence level to jump above the threshold is extracted from the operation records. This set is the optimal missing entity combination, representing the optimal subset of entities required to achieve a clear judgment with minimal information acquisition cost. The greedy algorithm achieves global approximation optima through local optimal selection, providing a practical decision-making solution when information acquisition costs are limited. Then, the dimension containing the optimal missing entity combination is taken as the missing medical entity dimension. The missing medical entity dimension refers to the type and attributes of each entity in the optimal missing entity combination; for example, some entities belong to the laboratory examination dimension, some to the imaging examination dimension, and some to the medical history inquiry dimension. The dimension identification operation extracts the type label and clinical classification of each missing entity, merging entities of the same type into the same dimension. The transformation operation converts the abstract entity identifiers and dimension classifications into supplementary suggestion information described in natural language. The suggestion information includes specific recommended examination items, recommended medical history content, and recommended vital signs indicators. The transformation process uses a predefined template library, selecting appropriate expression templates based on entity type and clinical context, and filling the templates with the standard names of the entities to generate complete suggestion statements. The generated supplementary suggestions are output in the form of structured lists or natural language paragraphs for doctors to refer to in their clinical work. This helps doctors identify important information that may have been missed in the current case, guides subsequent information collection and examination arrangements, and improves the completeness of medical data processing and the scientific nature of decision-making.
[0071] The present invention also discloses a smart medical data processing system based on knowledge graph and model reasoning, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the smart medical data processing method based on knowledge graph and model reasoning as described above.
[0072] The processor can be a central processing unit (CPU). Of course, depending on the actual use, it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc., and this application does not limit it.
[0073] The memory can be an internal storage unit of a computer device, such as a hard disk or RAM, or an external storage device, such as a plug-in hard disk, smart memory card (SMC), secure digital card (SD), or flash memory card (FC) provided on the computer device. Furthermore, the memory can be a combination of internal storage units and external storage devices of a computer device. The memory is used to store computer programs and other programs and data required by the computer device. The memory can also be used to temporarily store data that has been output or will be output. This application does not limit this.
[0074] The present invention also discloses a computer-readable storage medium storing instructions that, when executed by a processor, configure the processor to perform the intelligent medical data processing method based on knowledge graph and model reasoning as described in any of the above embodiments.
[0075] The computer program can be stored in a machine-readable medium. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or certain middleware. The machine-readable medium includes any entity or device capable of carrying computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the machine-readable medium includes, but is not limited to, the above-mentioned components.
[0076] The intelligent medical data processing method based on knowledge graph and model reasoning in the above embodiments is stored in the computer-readable storage medium and loaded and executed on the processor to facilitate the storage and application of the above method.
[0077] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of protection of this application is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of one or more embodiments of this application as described above, which are not provided in detail for the sake of brevity.
[0078] One or more embodiments in this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments in this application should be included within the protection scope of this application.
Claims
1. A smart medical data processing method based on a knowledge graph and model reasoning, characterized in that, Includes the following steps: The obtained complete medical documents are parsed and logically cleaned to extract medical triples containing medical entities and the relationships between them. A directed basic graph is then constructed based on the medical triples, where the relationships have an initial confidence attribute. Receive the case data of the target patient, extract medical feature entities from the case data through a natural language model, and anchor the target medical entity that is most closely related to the pathological state of the target patient in the directed basic graph based on the medical feature entities. Global path retrieval is performed in the directed basic graph based on the target medical entity. The topological path cost is determined according to the initial confidence of the association relationship. The optimal connectivity path from each associated entity node in the directed basic graph to the target medical entity is calculated based on the minimum cost priority principle. Remove non-optimal association edges that deviate from the optimal connected path from the graph topology in the directed basic graph, reduce the dimension of the directed basic graph with cyclic dependency closed loops, and reconstruct it into a directed acyclic reasoning tree with the target medical entity as the root node. In the directed acyclic inference tree, forward support paths and reverse rejection paths pointing to the terminal medical entity are extracted along the topological hierarchy. Attenuation compensation coefficients are configured based on the logical hop count of the forward support paths and reverse rejection paths. According to the attenuation compensation coefficients, the initial confidence of the multi-hop paths in the forward support paths and reverse rejection paths is calculated. The final confidence of the terminal medical entity is generated by combining the forward and reverse compensation results. When the final confidence level is within the preset fuzzy threshold range, the supplementary suggestion set corresponding to the end medical entity in the directed basic graph is extracted. A standard feature Boolean vector is generated based on the supplementary suggestion set. The standard feature Boolean vector is compared with the actual feature Boolean vector composed of medical feature entities. The missing medical entity dimension is located based on the difference between the standard feature Boolean vector and the actual feature Boolean vector. Supplementary suggestion information for doctors to refer to is generated based on the missing medical entity dimension.
2. The intelligent medical data processing method based on knowledge graphs and model reasoning according to claim 1, characterized in that, The process of parsing and logically cleaning the obtained complete medical documents, extracting medical triples containing medical entities and the relationships between them, and constructing a directed basic graph based on the medical triples includes the following steps: The layout analysis algorithm is used to preprocess the input complete medical document to obtain a clean text stream; The pure text stream is truncated into multiple text stream segments using a preset semantic sliding window, and the text stream segments are then input into a pre-trained medical big language model. Using a medical big language model, implicit medical subjects in text stream paragraphs are explicitly completed based on the medical context, spelling errors are corrected, and enhanced text blocks are generated. A parallel extraction thread is started to identify medical entities and relationships in the enhanced text block, and a set of candidate medical triples consisting of subject, predicate and object is output. Calculate the initial extraction confidence of each candidate medical triple in the candidate medical triple set, filter out candidate medical triples whose initial extraction confidence is lower than a preset quality threshold, and obtain medical triples; By mapping medical triples to medical entity nodes and related edges in a graph database, a directed basic graph is constructed.
3. The intelligent medical data processing method based on knowledge graphs and model reasoning according to claim 2, characterized in that, The process of calculating the initial extraction confidence of each candidate medical triplet in the candidate medical triplet set, and filtering out candidate medical triplets with initial extraction confidence below a preset quality threshold to obtain medical triplet sets includes the following steps: Scan all candidate medical triples and use a collision detection algorithm to identify conflicting triple pairs that have the same medical entity but whose relationship is mutually exclusive; For the identified conflicting triple pairs, extract the corresponding preceding and following time adverbs and condition adverbs in the complete medical document; Logical disambiguation is performed on conflicting triple pairs based on pre- and post-temporal adverbs and conditional adverbs, and the initial extraction confidence of the time dimension is redistributed for conflicting triple pairs based on the disambiguation results. Candidate medical triples that are still below the preset quality threshold after updating the initial extraction confidence level will be automatically discarded, and candidate medical triples that are above the preset quality threshold after updating the initial extraction confidence level will be converted into medical triples.
4. The intelligent medical data processing method based on knowledge graphs and model reasoning according to claim 1, characterized in that, The process of receiving the target patient's case data, extracting medical feature entities from the case data using a natural language model, and anchoring the target medical entity most closely related to the target patient's pathological state in the directed basic graph based on the medical feature entities includes the following steps: Receive case data of target patients and perform word segmentation and stop word filtering on the case data to extract medical feature words containing symptoms, signs and past medical history; Medical feature terms are input into a natural language model, which then maps these terms into standardized medical feature entities. Extract vital sign parameters from case data and calculate the global importance score for each medical feature entity using a term frequency-inverse document frequency algorithm; The global importance score is weighted and fused with vital sign parameters to generate the core weight value of each medical feature entity; Select the target medical entity with the largest core weight value, and match the corresponding medical entity node in the directed basic graph with a unique identifier as the target medical entity.
5. The intelligent medical data processing method based on knowledge graphs and model reasoning according to claim 1, characterized in that, The process of performing global path retrieval in the directed basic graph based on the target medical entity, determining the topological path cost based on the initial confidence of the association, and calculating the optimal connectivity path from each associated entity node in the directed basic graph to the target medical entity based on the minimum cost priority principle includes the following steps: Broadcast path exploration probes from the target medical entity as the origin to adjacent nodes in the directed basic graph; Record each relationship edge traversed by the path exploration probe and extract the initial confidence level of each relationship edge; The topological path cost corresponding to each association edge is calculated based on the initial confidence level. For any associated entity node in the directed basic graph that is associated with the target medical entity, the topological path cost of all candidate paths from the associated entity node to the target medical entity through different intermediate nodes is accumulated to obtain the accumulated topological path cost. Compare the cumulative topological path costs of all candidate paths for the same associated entity node, and mark the candidate path with the smallest cumulative topological path cost as the optimal connectivity path from the associated entity node to the target medical entity.
6. The intelligent medical data processing method based on knowledge graphs and model reasoning according to claim 5, characterized in that, The process of calculating the topological path cost corresponding to each association edge based on the initial confidence level includes the following steps: Extract the starting entity node corresponding to each relationship edge in the directed basic graph, and count the total number of out-degree connections from the starting entity node to other different entity nodes; The logical divergence entropy of the starting entity node is calculated based on the total number of out-degree connections, and the logical divergence entropy is used to quantify the characteristic specificity of the associated edges in the medical reasoning process. Construct an overhead evaluation function that includes a specific penalty term, input the logic divergence entropy as the specific penalty term into the dynamic overhead evaluation function, and input the initial confidence level as the basic variable into the overhead evaluation function; The topological path cost of each associated edge, after feature-specific calibration, is calculated using the cost evaluation function.
7. The intelligent medical data processing method based on knowledge graphs and model reasoning according to claim 1, characterized in that, The steps of extracting forward support paths and reverse rejection paths pointing to the final medical entity along the topological hierarchy in the directed acyclic inference tree, configuring attenuation compensation coefficients based on the logical hop count of the forward support paths and reverse rejection paths, performing balanced compensation calculations on the initial confidence of multi-hop paths in the forward support paths and reverse rejection paths according to the attenuation compensation coefficients, and generating the final confidence of the final medical entity by combining the forward and reverse compensation results include the following steps: In a directed acyclic reasoning tree, identify all unidirectional connected branches that point to the terminal medical entity from all leaf nodes, and classify the unidirectional connected branches into positive support paths and negative repulsion paths according to medical logical attributes. Count the number of edges traversed from the leaf node to the terminal medical entity for each positive support path and negative rejection path, and define the number of edges as the logical level hop count of the path; Extract the basic network connectivity before the directed acyclic reasoning tree is generated as a global parameter, and generate a decay compensation coefficient that is positively correlated with the number of logical level hops based on the basic network connectivity. The initial confidence level after the cumulative attenuation of each positive support path and negative repulsion path is calculated with the attenuation compensation coefficient to obtain the single-path balanced confidence level. The total support is obtained by summing the single-path balanced confidence scores of all positive support paths in the forward direction, and the total rejection score is obtained by summing the single-path balanced confidence scores of all negative rejection paths in the same dimension. The final confidence level of the end medical entity is obtained by subtracting the total rejection from the total support and then normalizing the result.
8. The intelligent medical data processing method based on knowledge graphs and model reasoning according to claim 7, characterized in that, The formula for calculating the single-path equilibrium confidence level is as follows: ; In the formula, For single-path equilibrium confidence, The initial confidence level is the product of the cumulative decay of each positive support path and the negative repulsion path. is the attenuation compensation coefficient, n is the number of logic level hops, e is the natural constant, and tanh is the hyperbolic tangent function.
9. A smart medical data processing system based on knowledge graphs and model reasoning, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the intelligent medical data processing method based on knowledge graph and model reasoning as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing instructions thereon, characterized in that, When executed by a processor, the instruction causes the processor to be configured to perform the intelligent medical data processing method based on knowledge graph and model reasoning according to any one of claims 1 to 8.
Citation Information
Patent Citations
Medical record content real-time intelligent auditing and error correcting method based on knowledge reasoning
CN121390040A
Large model knowledge graph completion method and system based on causal guidance
CN121787526A