Blood disease transplantation diagnosis and treatment auxiliary system based on knowledge graph
By optimizing the structure and applying path constraints to the hematological disease knowledge graph, an enhanced structure for the hematological disease transplantation knowledge graph is generated. This solves the problems of incomplete information and inaccurate reasoning in traditional systems, and enables efficient and personalized decision support for the hematological disease transplantation diagnosis and treatment auxiliary system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- LIAONING TAIYANG PHARMA TECH DEV
- Filing Date
- 2026-03-05
- Publication Date
- 2026-05-05
AI Technical Summary
Traditional knowledge graph-based hematology transplantation diagnostic and treatment support systems lack dynamic updates and intelligent reasoning capabilities, making it difficult to handle constantly changing and growing medical information. This results in insufficient practicality and accuracy of recommended solutions, and an inability to effectively support personalized optimization decisions.
By optimizing the structure of the knowledge graph in the field of hematology, identifying and correcting the feature density and sparse connection regions between entity nodes, an enhanced structure for the hematology transplantation knowledge graph is generated. Medical rule boundary conditions are applied to screen legal paths, a candidate strategy set is generated, and strategy consistency is verified. Finally, the knowledge graph structure is optimized through graph repair.
It improves the adaptability and information density of knowledge graphs, optimizes the accuracy and comprehensiveness of reasoning paths, ensures the practicality and accuracy of recommended solutions, strengthens the feasibility and rationality of diagnosis and treatment strategies, and solves the problems of incomplete information and inaccurate reasoning in traditional systems.
Smart Images

Figure CN121789962B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical informatics technology, and in particular to a knowledge graph-based auxiliary system for the diagnosis and treatment of hematological transplantation. Background Technology
[0002] Medical informatics is a cross-disciplinary field that integrates information technology and medicine. It aims to improve the efficiency and quality of healthcare services using computer science and information technology. This field includes medical data management, electronic health record systems, medical image processing, and intelligent decision support systems. The focus is on supporting medical research, healthcare services, and management decisions through information systems, knowledge management, and data mining. Core aspects include the automated acquisition and processing of disease information, the analysis and mining of medical data, the graphical representation and intelligent application of medical knowledge, and the application of intelligent systems in the medical process. The overall goal of medical informatics is to improve the efficiency of healthcare services, reduce medical errors, and promote the development of precision medicine.
[0003] Traditional knowledge graph-based hematological transplantation support systems involve constructing a hematological disease-related knowledge graph and combining it with the patient's clinical data to optimize the transplantation process. These systems rely on the integration of large-scale medical literature, case data, and clinical expert knowledge. By establishing a hematological disease knowledge graph covering etiology, pathogenesis, treatment methods, and postoperative management, the system first parses and integrates the patient's medical record data, then compares and infers with relevant information in the knowledge graph, ultimately outputting a suitable transplantation plan for the patient. Traditionally, such systems rely on rule engines or expert systems for reasoning and decision-making, without involving the application of machine learning or deep learning models.
[0004] Traditional knowledge graph-based hematology transplantation diagnostic and treatment support systems rely on rule engines or expert systems for reasoning and decision-making. While these systems can integrate vast amounts of medical literature, case data, and expert knowledge, their processing capabilities are limited by the size of the knowledge base and the complexity of the rules, making it difficult to handle constantly changing and growing medical information. This approach lacks dynamic updates and intelligent reasoning capabilities, resulting in poor adaptability to emerging medical research or complex cases. Furthermore, due to the lack of effective path constraints and optimization, the reasoning results cannot fully meet the actual needs of the transplantation stage, leading to insufficient practicality and accuracy of recommended protocols. This makes it difficult to cope with changing and complex clinical situations. Existing systems have significant deficiencies in the accuracy and comprehensiveness of transplantation protocols and cannot effectively support personalized optimization decisions. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing a knowledge graph-based auxiliary system for the diagnosis and treatment of hematological diseases via transplantation.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: a knowledge graph-based hematological transplantation diagnosis and treatment assistance system includes:
[0007] The knowledge graph structure optimization module obtains the initial knowledge graph in the field of hematology, identifies the feature density and sparse connection regions of entity nodes in the graph, calls the entity attribute consistency rules, hierarchical transmission rules and clinical semantic association rules in the medical ontology logical rule library, and performs topological reconstruction of the relationship between nodes by filling in logically missing edges in sparse connection regions or inserting intermediate nodes to generate an enhanced structure of the hematology transplantation knowledge graph.
[0008] The reasoning path constraint module, based on the enhanced structure of the hematological transplantation knowledge graph, extracts candidate reasoning paths from clinical phenotype nodes to transplantation strategy nodes. Using the preset hematopoietic stem cell transplantation time sequence logic and treatment access conditions as medical rule boundary conditions, it filters out path branches that do not conform to the order of transplantation stages, selects a set of paths that satisfy the logical order of transplantation stages, and establishes a transplantation knowledge reasoning path constraint model.
[0009] The candidate solution generation module, based on the transplantation knowledge reasoning path constraint model, traverses entities and relational chains in the graph that meet the constraint conditions, aggregates paths of the same strategy type and marks their support weights, and generates a set of candidate strategies for hematological transplantation.
[0010] The strategy consistency verification module, based on the blood disease transplant candidate strategy set, performs backtracking verification on the medical conditions on which each strategy depends, compares the logical closure between the strategy node and its upstream etiology and disease course nodes, records the strategy number and breakpoint position that do not meet the closure conditions, and generates the source tracing result of logical closure defects in transplantation diagnosis and treatment strategies.
[0011] As a further aspect of the present invention, the enhanced structure of the hematological transplantation knowledge graph includes an entity feature density distribution matrix, a sparse connection region identifier set, and a relation topology reconstruction mapping table; the transplantation knowledge reasoning path constraint model includes a state transition matrix composed of transplantation stage encodings, a path feasibility determination function, and a list of effective path indices; the hematological transplantation candidate strategy set includes strategy node identifiers, the number of supporting paths, and support weight values; and the transplantation diagnosis and treatment strategy logical closed-loop defect tracing results include abnormal strategy numbers, logical breakpoint locations, and closure verification state markers.
[0012] As a further aspect of the present invention, the knowledge graph structure optimization module includes:
[0013] The feature density calculation submodule obtains all entity nodes in the initial knowledge graph, counts the co-occurrence frequency of each node within a text co-occurrence window of a preset size, calculates the feature density value of the node by combining the number of adjacent edges of the node, and generates the entity feature density distribution result.
[0014] The sparse region identification submodule sets a density threshold based on the entity feature density distribution results, clusters nodes whose feature density values are lower than the density threshold, detects relationship loss patterns inside and at the boundaries of the clusters, and outputs a set of sparse connection region identifiers.
[0015] The topology reconstruction submodule, based on the sparse connection region identifier set, uses the medical ontology logic rule base to identify missing "belongs to", "cause", or "concurrent" logical relationships, inserts virtual intermediary nodes or supplements missing edges in the sparse region, adjusts the adjacency matrix structure of the original graph, and generates an enhanced structure for the hematological transplantation knowledge graph.
[0016] As a further aspect of the present invention, the inference path constraint module includes:
[0017] The stage rule loading submodule is based on the enhanced structure of the hematologic transplantation knowledge graph and loads the medical stage division criteria of the hematologic transplantation process. It encodes the five stages of pre-transplantation assessment, donor matching, pre-treatment plan, implantation operation and post-operative monitoring into an ordered state sequence and constructs a stage transition legal matrix.
[0018] The path filtering submodule traverses the path from the cause or symptom node representing the patient's initial state to the strategy node in the knowledge graph. Based on the stage transition legality matrix, it verifies the transition legality of the stage to which adjacent nodes belong in the path, eliminates paths that violate time sequence or treatment dependency, and obtains a set of legal paths.
[0019] The constraint model construction submodule assigns a unique index to each path based on the set of legal paths, and records its starting point, ending point, and intermediate stage node sequence. It integrates the path index and stage sequence information to generate a path constraint model for transplanted knowledge reasoning.
[0020] As a further aspect of the present invention, the candidate solution generation module includes:
[0021] The path traversal submodule, based on the transplanted knowledge reasoning path constraint model, sequentially visits the end node of each legal path, extracts the strategy type label and path length information of the end node, and generates the original set of strategy nodes.
[0022] The strategy aggregation submodule merges nodes with the same strategy type label based on the original set of strategy nodes, counts the number of their corresponding valid paths, and then sums them using a weighted average based on the inverse of the path length, using the formula:
[0023] ;
[0024] Calculate the support weight values to obtain the strategy support evaluation results;
[0025] in, Represents the support weight value. The number of nodes representing strategy types. Representing the The number of valid paths for each policy node. Representing the The path length of each policy node. Representing the The weight values of each strategy node;
[0026] The candidate set output submodule organizes node information by strategy type and associates support path indexes and weight values based on the strategy support evaluation results to generate a candidate strategy set for hematological transplantation.
[0027] As a further aspect of the present invention, the policy consistency verification module includes:
[0028] The backtracking path extraction submodule traces back along the knowledge graph to the root cause node for each strategy node in the blood disease transplantation candidate strategy set, extracts the condition nodes on the complete backtracking path, and constructs a strategy dependency condition chain.
[0029] The closure determination submodule verifies whether there is a key condition coverage relationship between the strategy node and the condition node constrained by a predefined set of medical rules, based on the strategy dependency condition chain. If there are uncovered key conditions, the strategy is marked as a logical anomaly. The strategy nodes marked as logical anomalies are summarized, and their numbers, first breakpoint positions, and closure verification failure types are recorded to generate the source tracing results of logical closure defects in transplantation diagnosis and treatment strategies.
[0030] As a further aspect of the present invention, the step of tracing back along the knowledge graph to the root cause node refers to taking the strategy node as the starting node, traversing layer by layer according to the preset reverse edge type, limiting the maximum tracing level to no more than six levels, and determining the root cause node when encountering a node without an upstream causal edge, thereby obtaining a complete backtracking path.
[0031] The key condition coverage relationship constrained by the predefined set of medical rules refers to the fact that in the policy dependency condition chain, adjacent condition nodes must satisfy the keyness and sufficiency mapping relationship registered in the rule base, and each policy node corresponds to at least two condition nodes marked as key, resulting in a verifiable logical mapping set.
[0032] The uncovered critical condition refers to a condition node in the policy dependency condition chain that is marked as critical by the rule base but does not appear in the complete backtracking path. When the number of missing nodes is greater than or equal to one, the marking is triggered, and a missing condition identifier is obtained.
[0033] As a further aspect of the present invention, the system also includes a map repair guidance module:
[0034] Based on the source tracing results of the logical closed-loop defects of the transplantation diagnosis and treatment strategy, the knowledge graph repair guidance module locates the missing condition nodes corresponding to the abnormal strategy, analyzes the candidate equivalent entities of the missing nodes in the knowledge graph, and generates a set of knowledge graph structure repair suggestion instructions.
[0035] The knowledge graph structure repair suggestion instruction set includes semantic descriptions of missing nodes, a list of candidate equivalent entities, and coordinates of relation insertion positions.
[0036] As a further aspect of the present invention, the map restoration guidance module includes:
[0037] The missing location submodule extracts the logical breakpoint position of each abnormal strategy based on the source tracing results of the logical closed loop defects of the transplantation diagnosis and treatment strategy, analyzes the missing condition constraint type at the breakpoint, and generates a feature vector of the missing node.
[0038] The equivalent entity retrieval submodule performs feature similarity matching in the same-level entity pool of the knowledge graph based on the feature vector of the missing node, uses the cosine similarity algorithm to filter candidate entities with similarity higher than a preset threshold, and outputs a list of candidate equivalent entities.
[0039] The repair instruction generation submodule constructs a relation insertion operation instruction based on the candidate equivalent entity list and the topological coordinates of the breakpoint in the graph, specifies the source node, target node and relation type, and generates a set of knowledge graph structure repair suggestion instructions.
[0040] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0041] This invention enhances the hematology transplantation knowledge graph by optimizing its structure, identifying and correcting the feature density and sparse connection regions between entity nodes. This improves the graph's adaptability and information density, and optimizes the accuracy and comprehensiveness of its reasoning paths. By applying medical rule boundary conditions to feasible reasoning paths, a set of paths conforming to the logical order of the transplantation stage is selected, further strengthening the medical rationality and reasoning accuracy of the decision-making process and effectively ensuring the practicality and precision of the recommended solutions. Furthermore, through candidate strategy generation and consistency verification, the logical closure between different strategies is strengthened, ensuring the feasibility and rationality of transplantation treatment strategies. Simultaneously, graph repair guidance optimizes the structure and integrity of the knowledge graph, effectively solving the technical problems of incomplete information and inaccurate reasoning in traditional systems. Attached Figure Description
[0042] Figure 1 This is a system flowchart of the present invention;
[0043] Figure 2 This is a flowchart of the knowledge graph structure optimization module in this invention;
[0044] Figure 3 This is a flowchart of the reasoning path constraint module in this invention;
[0045] Figure 4 This is a flowchart of the candidate solution generation module in this invention;
[0046] Figure 5 This is a flowchart of the strategy consistency verification module in this invention;
[0047] Figure 6 This is a flowchart of the map repair guidance module in this invention. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0049] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0050] Please see Figure 1 Knowledge graph-based diagnostic and treatment support systems for hematological transplantation include:
[0051] The knowledge graph structure optimization module obtains the initial knowledge graph in the field of hematology, identifies the feature density and sparse connection regions of entity nodes in the graph, calls the entity attribute consistency rules, hierarchical transmission rules and clinical semantic association rules in the medical ontology logical rule library, and performs topological reconstruction of the relationship between nodes by filling in logically missing edges in sparse connection regions or inserting intermediate nodes to generate an enhanced structure of the hematology transplantation knowledge graph.
[0052] The reasoning path constraint module is based on the enhanced structure of the hematological transplantation knowledge graph. It extracts candidate reasoning paths from clinical phenotype nodes to transplantation strategy nodes. Using the preset hematopoietic stem cell transplantation time sequence logic and treatment access conditions as medical rule boundary conditions, it filters out path branches that do not conform to the order of transplantation stages, selects a set of paths that meet the logical order of transplantation stages, and establishes a transplantation knowledge reasoning path constraint model.
[0053] The candidate strategy generation module traverses entities and relational chains that meet the constraints in the graph based on the transplantation knowledge reasoning path constraint model, extracts the intervention strategy node at the end of each chain, aggregates paths of the same strategy type and marks their support weights, and generates a set of candidate strategies for hematological transplantation.
[0054] The strategy consistency verification module is based on the candidate strategy set for blood disease transplantation. It backtracks and verifies the medical conditions on which each strategy depends, compares the logical closure between the strategy node and its upstream etiology and disease course nodes, records the strategy number and breakpoint position that do not meet the closure conditions, and generates the source tracing results of logical closure defects in transplantation diagnosis and treatment strategies.
[0055] Logical closure refers to the ability of the transplantation decision conditions corresponding to the strategy nodes in the hematological transplantation knowledge graph to form a complete, uninterrupted, and medically causal and temporal relationship path that conforms to the constraints of medical causality and time sequence among its upstream etiology nodes, disease stage nodes, and treatment precondition nodes.
[0056] Based on the source tracing results of logical loop defects in transplantation and treatment strategies, the knowledge graph repair guidance module locates the missing condition nodes corresponding to abnormal strategies, analyzes the candidate equivalent entities of the missing nodes in the knowledge graph, and generates a set of knowledge graph structure repair suggestion instructions.
[0057] The enhanced structure of the hematologic transplantation knowledge graph includes an entity feature density distribution matrix, a sparse connection region identifier set, and a relation topology reconstruction mapping table. The transplantation knowledge reasoning path constraint model includes a state transition matrix composed of transplantation stage encodings, a path feasibility judgment function, and a list of valid path indices. The hematologic transplantation candidate strategy set includes strategy node identifiers, the number of supporting paths, and support weight values. The transplantation diagnosis and treatment strategy logical closed-loop defect tracing results include abnormal strategy numbers, logical breakpoint locations, and closure verification state markers. The knowledge graph structure repair suggestion instruction set includes semantic descriptions of missing nodes, a list of candidate equivalent entities, and coordinates of relation insertion positions.
[0058] Please see Figure 2 The knowledge graph structure optimization module includes:
[0059] The feature density calculation submodule obtains all entity nodes in the initial knowledge graph, counts the co-occurrence frequency of each node within a preset-sized text co-occurrence window, and uses the formula:
[0060] ;
[0061] The feature density value of a node is calculated by combining the number of adjacent edges of the node, and the entity feature density distribution result is generated.
[0062] in, The feature density value representing the node, Representative node and nodes Edge weights between them Representative node and nodes Co-occurrence frequency between them Representative node The number of adjacent edges, Represents the total number of nodes;
[0063] Retrieve all entity nodes from the initial knowledge graph, and select entity nodes under the category of "leukemia chemotherapy drugs". Using "cyclophosphamide" as the target (e.g.), the size of the medical context window is set to include the target node and the text spanning 50 words before and after it. The medical literature corpus is traversed to search for entries within this window that correspond to the target node. Simultaneously appearing adjacent nodes (For example, "hemorrhagic cystitis", "mesna", "hematopoietic stem cells"), statistical nodes Co-occurrence frequency within a text co-occurrence window of a preset size Call the graph database query interface to obtain nodes Current number of adjacent edges Extracting nodes from the edge attributes of a pre-trained medical knowledge graph With each adjacent node Edge weights between This weight value Based on the product of Pearson correlation coefficient and medical entity confidence, with values strictly limited to between 0.1 and 1.0, a dataset containing a list of adjacent nodes, corresponding weights, and co-occurrence frequencies is constructed using the formula: Here, in the formula This indicates that for all nodes... There are direct connections Perform a weighted summation operation on each of the adjacent nodes. Representative node and nodes The edge weights between them are used to measure the strength of the medical association between the two. Representative node and nodes The co-occurrence frequency between them reflects the actual co-existence activity of the two in the corpus, and the product is... Quantify the adjacency pairs of nodes The contribution of feature density, Representative node The number of adjacent edges, with the denominator set to . The aim is to smooth and normalize the node degree to avoid excessive inflation of the density values of high-frequency connected nodes, thereby truly reflecting the average characteristic strength under unit connectivity. A practical example is presented using the data in Table 1, setting the nodes... The number of adjacent edges of "cyclophosphamide" is... It is 3 (i.e.) The adjacent nodes are respectively (Hemorrhagic cystitis) (Mestna) (Hematopoietic stem cells), specific parameters are shown in Table 1;
[0064] Substitute into the formula to calculate:
[0065] First, calculate the molecular part. ;
[0066] Next, calculate the denominator. ;
[0067] Finally, it was concluded
[0068] ;
[0069] The calculation result of 63.8 indicates that "cyclophosphamide" has a high feature information carrying density in this local map structure, and the value is significantly higher than the preset low density benchmark value of 20.0. The feature density value of the node is calculated by combining the number of adjacent edges of the node, and the entity feature density distribution result is generated.
[0070] Table 1: Cyclophosphamide Node Adjacency Data Table
[0071] ;
[0072] Table 1 shows the basic parameters used to calculate the characteristic density of cyclophosphamide and the specific co-occurrence frequency data.
[0073] The sparse region identification submodule clusters nodes with feature density values below the density threshold based on the entity feature density distribution results, detects relationship loss patterns inside and at the boundaries of clusters, and outputs a set of sparse connection region identifiers.
[0074] Based on the entity feature density distribution results, the feature density values of all nodes in the entire graph are statistically analyzed, and their arithmetic mean (45.5) and standard deviation (12.2) are calculated. A density threshold is set as the mean minus 1.5 times the standard deviation, i.e., 45.5 - 1.5 × 12.2 = 27.2. The graph nodes are traversed, and all nodes with feature density values below 27.2 (below the density threshold), such as auxiliary treatment nodes like "dietary care" and "psychological counseling," are selected. The K-Means clustering algorithm is then applied to map these nodes to a high-dimensional vector space. The number of clusters is set to 5, and nodes with feature density values below the density threshold are divided into different clusters based on Euclidean distance. For each cluster, the average path length between nodes within the cluster is calculated. If the length is greater than 3.0, the cluster is considered sparsely connected. The number of connections between the cluster boundary nodes and high-density region nodes is detected. If the number of boundary connections is less than 2, the boundary is considered broken. The "postoperative rehabilitation care" region is identified as a sparsely connected region. Although the nodes in this region are semantically related, they lack direct medical logical connections. The relationship missing patterns inside and at the boundaries of the cluster are detected, and a set of sparsely connected region identifiers is output. This process accurately locates the weak links in the knowledge graph with insufficient information coverage by quantifying the density distribution and topological connection characteristics. It provides clear spatial coordinates and logical ranges for subsequent targeted completion, ensuring the directionality and effectiveness of graph optimization, avoiding the waste of computing resources caused by global blind search, and realizing the intelligent identification and extraction of low-density, weakly connected regions.
[0075] The topology reconstruction submodule is based on the sparse connection region identifier set. It uses the medical ontology logical rule base to identify the missing "belongs to", "cause", or "concurrent" logical relationships, inserts virtual mediator nodes or supplements missing edges in the sparse region, adjusts the adjacency matrix structure of the original graph, and generates an enhanced structure for the hematological transplantation knowledge graph.
[0076] Based on the sparsely connected region identifier set, the list of isolated nodes within the sparse region of "Postoperative Rehabilitation Nursing" is read, including "High-Protein Diet" and "Infection Prevention". The external authoritative medical ontology library SNOMEDCT is loaded, and the hidden logic between nodes is retrieved. It is found that both "High-Protein Diet" and "Infection Prevention" belong to the concept of "Supportive Treatment" in the ontology library, and there is a hidden equivalent description of a "Promoting" relationship. A virtual mediator node "Supportive Treatment Strategy" is inserted into the sparse region. According to the ontology hierarchy, "Belongs" edges pointing from "High-Protein Diet" to "Supportive Treatment Strategy" and "Cooperates" edges pointing from "Supportive Treatment Strategy" to "Infection Prevention" are established. For missing edges, the ontology is used to... The reasoning rules directly supplement the "positive correlation effect" edge between "psychological counseling" and "immunity recovery," update the graph database, reallocate node IDs and refresh the adjacency list, adjust the adjacency matrix structure of the original graph, and fill the structural gaps in specific sub-domains of the original graph by introducing external standardized ontology knowledge. This enhances the semantic connectivity of the graph in the dimensions of auxiliary treatment and nursing care, generating an enhanced structure for the hematology transplantation knowledge graph. This structure not only repairs broken semantic links but also enriches the hierarchical structure of knowledge through newly added mediator nodes, enabling previously isolated nursing knowledge points to be linked through standard medical concepts. This improves the graph's ability to describe and support reasoning in complex diagnosis and treatment scenarios.
[0077] Please see Figure 3 The inference path constraint module includes:
[0078] The stage rule loading submodule is based on the enhanced structure of the hematologic transplantation knowledge graph. It loads the medical stage division standards of the hematologic transplantation process and encodes the five stages of pre-transplantation assessment, donor matching, pretreatment plan, implantation operation and post-operative monitoring into an ordered state sequence, and constructs a stage transition legal matrix.
[0079] Based on the knowledge graph enhancement structure of hematological transplantation, the "Management Standards for Hematopoietic Stem Cell Transplantation Technology" is read, and the medical stage division standards of the hematological transplantation process are loaded. Five standard stages are defined: Stage 1 "Pre-transplantation Assessment", Stage 2 "Donor Matching", Stage 3 "Pre-treatment Protocol", Stage 4 "Implantation Procedure", and Stage 5 "Post-operative Monitoring". These five stages are sequentially encoded as a state set S={1, 2, 3, 4, 5}, and a state transition function is defined. Construct a 5×5 stage transition legal matrix M, initialized to all zeros. Based on time irreversibility and treatment guidelines, only transitions to the current state or the next state are allowed. This is achieved by setting the elements on the main diagonal and the first superdiagonal element above the main diagonal to 1 (i.e., when...). or hour, The remaining positions are kept at 0. For example, an element of 1 in the first row and second column indicates that the transition from the evaluation stage to the matching stage is allowed, while an element of 0 in the third row and first column indicates that the transition from the preprocessing stage back to the evaluation stage is prohibited. This matrix strictly constrains the temporal logic of the diagnosis and treatment events, encoding the five stages of pre-transplantation evaluation, donor matching, preprocessing plan, implantation operation, and post-operative monitoring into an ordered state sequence, and constructing a stage transition legality matrix. This matrix serves as the core logical filter for subsequent path verification, ensuring that all retrieved diagnosis and treatment paths conform to the time flow of actual clinical operations, eliminating interference from invalid paths that are logically reversed or jump across stages, and laying a solid rule foundation for building a highly reliable diagnosis and treatment knowledge reasoning model.
[0080] The path filtering submodule traverses the path in the knowledge graph from the cause or symptom node representing the patient's initial state to the strategy node. It verifies the legality of the transition between adjacent nodes in the path based on the stage transition legality matrix, and removes paths that violate the time sequence or treatment dependency relationship to obtain a set of legal paths.
[0081] In practice, a depth-first search (DFS) algorithm is used to extract all potential path sequences from the etiology node "acute myeloid leukemia" to the strategy node "allogeneic hematopoietic stem cell transplantation". The system obtains the stage code of each node through a preset entity-stage mapping table. For example, node A belongs to stage 1, node B belongs to stage 2, node C belongs to stage 1, and node D belongs to stage 3. It then generates the transition sequence 1→2, 2→1, 1→3, based on the stage transition legal matrix. Verify the validity of transitions in the stages to which adjacent nodes belong in the verification path, i.e., verify... To check if it equals 1, examine the sequence 1→2. The result is deemed valid; sequences 2 to 1 are checked. The system determines illegal paths. If any illegal transition exists in a path, the path is immediately marked as invalid and pruned, removed from the candidate list. Only paths that conform to unidirectional progressive logic, such as sequences 1→2→3→4→5 or 1→1→2→3, are retained. Paths that violate time sequence or treatment dependence are eliminated, resulting in a set of legal paths. This process effectively eliminates logically paradoxical paths in the graph caused by automatic extraction or noisy data, ensuring that each path ultimately retained represents a diagnostic and treatment evolution route that is executable in terms of medical ethics and operational norms. This greatly improves the clinical reference value and safety of the knowledge reasoning results.
[0082] The constraint model construction submodule assigns a unique index to each path based on the set of legal paths, and records its starting point, ending point and intermediate stage node sequence. It integrates the path index and stage sequence information to generate a path constraint model for transplanted knowledge reasoning.
[0083] The model construction process is as follows: A path retrieval structure based on a multi-dimensional vector space is established. Each legal path in the set (e.g., Path_001 to Path_150) is transformed into a triple vector consisting of "node features - stage weights - transition probabilities" and stored in an inverted index database. For Path_001, its starting point is identified as "CMV virus infection" and its ending point as "antiviral treatment strategy". The intermediate stage node sequence includes "immunosuppression assessment" and "drug sensitivity test". These stage nodes are mapped to attribute labels with timestamp weights. A fast index model from stage state to path ID is established through hash mapping. This model, through a structured indexing mechanism and logical gating units, transforms discrete graph paths into a probabilistic inference network with temporal constraints. It not only records the start and end points of the path but also retains the temporal position information of each node in the path, providing an efficient data access interface for subsequent policy support calculation and solving the policy offset problem caused by the lack of logical constraints in graph inference.
[0084] Please see Figure 4 The candidate solution generation module includes:
[0085] The path traversal submodule is based on the transplanted knowledge reasoning path constraint model. It visits the end node of each legal path in turn, extracts the policy type label and path length information of the end node, and generates the original set of policy nodes.
[0086] Based on the transplanted knowledge reasoning path constraint model, the traversal pointer is initialized, and the terminal node "Ganciclovir Treatment Plan" of Path_001 is locked. The attribute table of this node is queried to extract the strategy type label "Antiviral Drug" and the hop length of this path, for example, 4 hops. The terminal node "Fosfocarboxylate Treatment Plan" of another path Path_002 is locked, and the label "Antiviral Drug" and length 5 are extracted. The extracted metadata is temporarily stored in an in-memory hash table, with the key being the strategy ID and the value containing the type label and path length list. The strategy type label and path length information of the terminal node are extracted to generate the original set of strategy nodes. This set gathers all potential treatment strategies and their corresponding source path features. Through the organization of the in-memory hash table, the rapid classification and caching of large-scale path data is realized, providing standardized data input for subsequent aggregation calculations and ensuring that the strategy evaluation process can fully cover all diagnosis and treatment clues mined from the graph.
[0087] The strategy aggregation submodule merges nodes with the same strategy type label based on the original set of strategy nodes, counts the number of their corresponding valid paths, and then sums them weighted by the inverse of the path length using the formula:
[0088] ;
[0089] Calculate the support weight values to obtain the strategy support evaluation results;
[0090] in, Represents the support weight value. The number of nodes representing strategy types. Representing the The number of valid paths for each policy node. Representing the The path length of each policy node. Representing the The weight values of each strategy node;
[0091] Based on the original set of policy nodes, identify multiple policy nodes belonging to the same policy type label "preprocessing scheme" (e.g., policy A "BuCy scheme", policy B "TBI+Cy scheme"). For policy A "BuCy scheme", count the number of legal paths pointing to this policy. Calculate the length of each path And obtain the baseline weight value of the strategy node preset in the expert system. The formula used is: Here, This represents the support weight value; a higher value indicates a higher recommendation priority for this strategy. The number of policy type nodes involved in the aggregation calculation (here, for a single policy calculation, then...). This represents the diversity dimension of path sources, or can be understood as a weighted sum of different path sources for the same strategy. In this example... (Refers to the total number of paths that support this strategy). Representing the The number of times a path is referenced repeatedly or the number of paths (here, the strength of a single path is taken and set to 1; if aggregated by group, it is the number). Representing the The length of the path (number of hops). Representing the The confidence weights of the policy nodes corresponding to each path are shown in the formula. This item reflects the logic that "the shorter and more direct the path, the stronger the evidence." The term introduces a length penalty again and combines it with weights, giving non-linear high-score support to strategies with short paths and high weights. Now, based on the data in Table 2, the support of strategy A "BuCy scheme" is calculated, assuming that there are 2 paths supporting this strategy (i.e., ), path 1 length Path weight Number of paths (Single); Path 2 length Path weight Number of paths ;
[0092] The specific parameters are shown in Table 2. Substitute them into the formula to calculate:
[0093] First item ;
[0094] Second item ;
[0095] Total support ;
[0096] The result of 0.09025 will be used as the final ranking score for the "BuCy scheme". The support weight value will be calculated to obtain the strategy support evaluation result.
[0097] Table 2: Supported Path Parameters for the BuCy Solution
[0098] ;
[0099] Table 2 details the number of paths, their lengths, and the corresponding node weight parameters used in calculating the support of the BuCy scheme.
[0100] The candidate set output submodule organizes node information by strategy type based on the strategy support evaluation results, associates support path indexes with weight values, and generates a candidate strategy set for hematological transplantation.
[0101] Based on the strategy support evaluation results, the S-values of all strategy nodes are obtained. Classification buckets are established based on strategy type (e.g., "anti-infection," "anti-rejection," "pretreatment"), and strategy nodes are archived by type. Within each category, strategies are quickly sorted from highest to lowest S-value. For example, in the "pretreatment" category, strategy A has a score of 0.09025, ranking ahead of strategy B with a score of 0.075. A structured object containing strategy name, score, and a list of supporting path IDs is constructed, and the supporting path index and weight value are associated to generate a candidate strategy set for hematological transplantation. This set is categorized according to medical logic and sorted by confidence level, providing clinicians with a clearly structured and focused decision reference list. This allows doctors to prioritize advantageous strategies with shorter evidence chains, higher weights, and more supporting paths, significantly improving the efficiency and accuracy of treatment plan formulation and effectively transforming knowledge reasoning results into clinical application value.
[0102] Please see Figure 5 The policy consistency verification module includes:
[0103] The backtracking path extraction submodule targets each strategy node in the blood disease transplant candidate strategy set, traces back along the knowledge graph to the root cause node, extracts the condition nodes on the complete backtracking path, and constructs the strategy dependency condition chain.
[0104] Tracing back along the knowledge graph to the root cause node means taking the strategy node as the starting node, traversing layer by layer according to the preset reverse edge type, limiting the maximum tracing level to no more than six levels, and determining the root cause node when encountering a node without an upstream causal edge, thus obtaining the complete backtracking path;
[0105] For each strategy node in the candidate strategy set for hematologic transplantation, such as selecting the strategy node "cyclosporine A injection", its associated graph edge data is read, and the causal edge type pointing to the node is identified as "treated_by" or "mitigated_by". The knowledge graph is traced backward to the root cause node. The first hop backtracks to "acute graft-versus-host disease (GVHD)", the second hop backtracks to "allogeneic immune response", the third hop backtracks to "HLA mismatch", and the fourth hop backtracks to the root node "acute lymphoblastic leukemia". At this point, the node has no in-degree causal edge and is determined to be the root cause. The complete node sequence is recorded as [cyclosporine A, GVHD, immune response, HLA mismatch, leukemia]. The condition nodes on the complete backtracking path are extracted, and a strategy dependency condition chain is constructed. This chain completely reproduces the pathological evolution logic from the root cause of the disease to the specific treatment method, providing detailed contextual basis for subsequent logical closure verification. This ensures that the review of the rationality of the strategy can be based on the complete causal evidence chain, rather than judging the validity of a single node in isolation.
[0106] The closure determination submodule verifies whether there is a key condition coverage relationship between the strategy node and the condition node constrained by a predefined set of medical rules based on the strategy dependency condition chain. If there are uncovered key conditions, the strategy is marked as a logical anomaly. The strategy nodes marked as logical anomalies are summarized, and their numbers, first breakpoint positions and closure verification failure types are recorded to generate the source tracing results of logical closure defects in transplantation diagnosis and treatment strategies.
[0107] The key condition coverage relationship constrained by a predefined set of medical rules refers to the fact that in the policy dependency condition chain, adjacent condition nodes must satisfy the keyness and sufficiency mapping relationship registered in the rule base, and each policy node corresponds to at least two condition nodes marked as key, resulting in a verifiable logical mapping set.
[0108] Uncovered critical conditions refer to the condition nodes in the policy dependency condition chain that are marked as critical by the rule base but do not appear in the complete backtracking path. The marking is triggered when the number of missing nodes is greater than or equal to one, and a missing condition identifier is obtained.
[0109] Based on the policy dependency chain, a logical closed-loop template from the medical rule base is loaded. For the policy "Cyclosporine A," the rules require the existence of the precondition "liver and kidney function assessment" with a "tolerable" status, and the existence of a "blood drug concentration monitoring" plan. Scanning the extracted policy dependency chain reveals that it contains "GVHD" and "immune response," but lacks the crucial precondition "liver and kidney function assessment." Verification is performed to determine if there is a key condition coverage relationship between the policy node and the condition node, constrained by a predefined set of medical rules. The current path is deemed not to meet the "sufficiency" constraint, and the policy is flagged. There are slight logical anomalies of the "missing preconditions" type. The breakpoint is recorded between "cyclosporine A" and "GVHD". If a "fracture history" node unrelated to treatment is found in the chain, it is marked as a "redundant condition". The strategy nodes marked as logical anomalies are summarized, and their numbers, first breakpoint positions and closure verification failure types are recorded. The result of tracing the logical closed loop defects of the transplantation diagnosis and treatment strategy is generated. The result accurately points out the specific defects in the logical integrity of the recommended strategy, prevents medical risks caused by missing conditions or logical breaks, and reflects the rigor of knowledge reasoning in terms of security review.
[0110] Please see Figure 6 The map restoration guidance module includes:
[0111] The missing node localization submodule extracts the logical breakpoint location of each abnormal strategy based on the source tracing results of the logical closed-loop defects in the transplantation diagnosis and treatment strategy, analyzes the missing condition constraint type at the breakpoint, and generates a feature vector of the missing node.
[0112] Based on the source tracing results of the logical closed-loop defects in the transplantation diagnosis and treatment strategy, the logical breakpoint of the abnormal strategy "Cyclosporine A" was located. A query of the rule base revealed that the missing condition constraint type was "Pre-check_Condition". The vector dimension was set to 128 dimensions. Based on the semantic association between "Cyclosporine A" (drug-related) and "liver and kidney function" (physiological indicator-related) in the context, the vector values were initialized, with a focus on strengthening the feature values of the "examination", "assessment", and "organ function" dimensions. The missing condition constraint type at the breakpoint was analyzed, generating a feature vector for the missing node. This vector mathematically represents the semantic features of the missing information, transforming the abstract medical logical defect into high-dimensional spatial coordinates that can be processed by computers. This provides quantified retrieval keys for subsequent accurate retrieval and matching of potential repair entities in the knowledge graph, building a technical bridge from logical diagnosis to entity repair.
[0113] The equivalent entity retrieval submodule performs feature similarity matching in the entity pool of the same level of the knowledge graph based on the feature vector of the missing node, uses the cosine similarity algorithm to filter candidate entities with similarity higher than a preset threshold, and outputs a list of candidate equivalent entities.
[0114] Based on the feature vectors of missing nodes, the system connects to the entity embedding vector database of the knowledge graph. The search scope is limited to the entity pool at the same level as "clinical examination." Feature similarity matching is performed, calculating the similarity between the missing node vector and the entity vectors in the pool. This calculation is achieved by measuring the cosine of the angle between the two vectors in multidimensional space, i.e., calculating the dot product of the two vectors divided by the product of their moduli. The similarity of the entity "liver and kidney function combined detection" is calculated to be 0.92, and the similarity of the entity "blood routine" is 0.65. Candidate entities with similarity higher than the preset threshold of 0.85 are filtered, and "liver and kidney function combined detection" is selected as the best match. A list of candidate equivalent entities is output. This process utilizes entity embedding representations generated by deep learning, overcoming the limitations of traditional keyword matching. It can accurately find functionally equivalent or logically complementary medical entities based on semantic connotation, achieving precise recall even with differences in naming, ensuring the accuracy and medical rationality of the knowledge graph repair suggestions.
[0115] The repair instruction generation submodule constructs relation insertion operation instructions based on the candidate equivalent entity list and the topological coordinates of the breakpoints in the graph, specifies the source node, target node and relation type, and generates a set of knowledge graph structure repair suggestion instructions.
[0116] Based on the candidate equivalent entity list, the node to be inserted is determined to be "Liver and Kidney Function Combined Detection". Combining the topological coordinates, i.e., located after the "GVHD" node and before the "Cyclosporine A" node, a relation insertion operation instruction is constructed, clearly specifying the source node, target node, and relation type. The generated instruction contains two specific operations: first, inserting the "Need to be checked" relation between "GVHD" and "Liver and Kidney Function Combined Detection", and then inserting the "As a prerequisite" relation between "Liver and Kidney Function Combined Detection" and "Cyclosporine A". This generates a knowledge graph structure repair suggestion instruction set. This instruction set encapsulates specific graph editing actions in a standardized data format, which can be directly parsed and executed by the graph database execution engine. It realizes a closed-loop process from logical defect discovery to automatic structural repair, providing an executable technical foothold for the continuous self-evolution and quality improvement of the knowledge graph.
[0117] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A knowledge graph-based diagnostic and treatment support system for hematological transplantation, characterized in that, The system includes: The knowledge graph structure optimization module obtains the initial knowledge graph in the field of hematology, identifies the feature density and sparse connection regions of entity nodes in the graph, calls the entity attribute consistency rules, hierarchical transmission rules and clinical semantic association rules in the medical ontology logical rule library, and performs topological reconstruction of the relationship between nodes by filling in logically missing edges in sparse connection regions or inserting intermediate nodes to generate an enhanced structure of the hematology transplantation knowledge graph. The reasoning path constraint module, based on the enhanced structure of the hematological transplantation knowledge graph, extracts candidate reasoning paths from clinical phenotype nodes to transplantation strategy nodes. Using the preset hematopoietic stem cell transplantation time sequence logic and treatment access conditions as medical rule boundary conditions, it filters out path branches that do not conform to the order of transplantation stages, selects a set of paths that satisfy the logical order of transplantation stages, and establishes a transplantation knowledge reasoning path constraint model. The candidate solution generation module, based on the transplantation knowledge reasoning path constraint model, traverses entities and relational chains in the graph that meet the constraint conditions, aggregates paths of the same strategy type and marks their support weights, and generates a set of candidate strategies for hematological transplantation. The strategy consistency verification module, based on the blood disease transplant candidate strategy set, performs backtracking verification on the medical conditions on which each strategy depends, compares the logical closure between the strategy node and its upstream etiology and disease course nodes, records the strategy number and breakpoint position that do not meet the closure conditions, and generates the source tracing result of logical closure defects in transplantation diagnosis and treatment strategies.
2. The knowledge graph-based blood disease transplantation diagnostic and treatment assistance system according to claim 1, characterized in that, The enhanced structure of the hematological transplantation knowledge graph includes an entity feature density distribution matrix, a sparse connection region identifier set, and a relation topology reconstruction mapping table. The transplantation knowledge reasoning path constraint model includes a state transition matrix composed of transplantation stage encodings, a path feasibility judgment function, and a list of effective path indices. The hematological transplantation candidate strategy set includes strategy node identifiers, the number of supporting paths, and support weight values. The source tracing results of logical closed-loop defects in transplantation diagnosis and treatment strategies include abnormal strategy numbers, logical breakpoint locations, and closure verification state markers.
3. The knowledge graph-based hematology transplantation diagnostic and treatment support system according to claim 1, characterized in that, The knowledge graph structure optimization module includes: The feature density calculation submodule obtains all entity nodes in the initial knowledge graph, counts the co-occurrence frequency of each node within a text co-occurrence window of a preset size, calculates the feature density value of the node by combining the number of adjacent edges of the node, and generates the entity feature density distribution result. The sparse region identification submodule sets a density threshold based on the entity feature density distribution results, clusters nodes whose feature density values are lower than the density threshold, detects relationship loss patterns inside and at the boundaries of the clusters, and outputs a set of sparse connection region identifiers. The topology reconstruction submodule, based on the sparse connection region identifier set, uses the medical ontology logic rule base to identify missing "belongs to", "cause", or "concurrent" logical relationships, inserts virtual mediator nodes or supplements missing edges in the sparse region, adjusts the adjacency matrix structure of the original graph, and generates an enhanced structure for the hematological transplantation knowledge graph.
4. The knowledge graph-based hematology transplantation diagnostic and treatment support system according to claim 3, characterized in that, The inference path constraint module includes: The stage rule loading submodule is based on the enhanced structure of the hematologic transplantation knowledge graph and loads the medical stage division criteria of the hematologic transplantation process. It encodes the five stages of pre-transplantation assessment, donor matching, pre-treatment plan, implantation operation and post-operative monitoring into an ordered state sequence and constructs a stage transition legal matrix. The path filtering submodule traverses the path from the cause or symptom node representing the patient's initial state to the strategy node in the knowledge graph. Based on the stage transition legality matrix, it verifies the transition legality of the stage to which adjacent nodes belong in the path, eliminates paths that violate time sequence or treatment dependency, and obtains a set of legal paths. The constraint model construction submodule assigns a unique index to each path based on the set of legal paths, and records its starting point, ending point, and intermediate stage node sequence. It integrates the path index and stage sequence information to generate a path constraint model for transplanted knowledge reasoning.
5. The knowledge graph-based hematology transplantation diagnostic and treatment support system according to claim 4, characterized in that, The candidate solution generation module includes: The path traversal submodule, based on the transplanted knowledge reasoning path constraint model, sequentially visits the end node of each legal path, extracts the strategy type label and path length information of the end node, and generates the original set of strategy nodes. The strategy aggregation submodule merges nodes with the same strategy type label based on the original set of strategy nodes, counts the number of their corresponding valid paths, and then sums them using a weighted average based on the inverse of the path length, using the formula: ; Calculate the support weight values to obtain the strategy support evaluation results; in, Represents the support weight value. The number of nodes representing strategy types. Representing the The number of valid paths for each policy node. Representing the The path length of each policy node. Representing the The weight values of each strategy node; The candidate set output submodule organizes node information by strategy type and associates support path indexes and weight values based on the strategy support evaluation results to generate a candidate strategy set for hematological transplantation.
6. The knowledge graph-based blood disease transplantation diagnostic and treatment assistance system according to claim 5, characterized in that, The policy consistency verification module includes: The backtracking path extraction submodule traces back along the knowledge graph to the root cause node for each strategy node in the blood disease transplantation candidate strategy set, extracts the condition nodes on the complete backtracking path, and constructs a strategy dependency condition chain. The closure determination submodule verifies whether there is a key condition coverage relationship between the strategy node and the condition node constrained by a predefined set of medical rules, based on the strategy dependency condition chain. If there are uncovered key conditions, the strategy is marked as a logical anomaly. The strategy nodes marked as logical anomalies are summarized, and their numbers, first breakpoint positions, and closure verification failure types are recorded to generate the source tracing results of logical closure defects in transplantation diagnosis and treatment strategies.
7. The knowledge graph-based hematological transplantation diagnostic and treatment support system according to claim 6, characterized in that, The process of tracing back along the knowledge graph to the root cause node refers to using the strategy node as the starting node, traversing layer by layer according to the preset reverse edge type, limiting the maximum tracing level to no more than six levels, and determining the root cause node when encountering a node without an upstream causal edge, thus obtaining the complete backtracking path. The key condition coverage relationship constrained by the predefined set of medical rules refers to the fact that in the policy dependency condition chain, adjacent condition nodes must satisfy the keyness and sufficiency mapping relationship registered in the rule base, and each policy node corresponds to at least two condition nodes marked as key, resulting in a verifiable logical mapping set. The uncovered critical condition refers to a condition node in the policy dependency condition chain that is marked as critical by the rule base but does not appear in the complete backtracking path. When the number of missing nodes is greater than or equal to one, the marking is triggered, and a missing condition identifier is obtained.
8. The knowledge graph-based blood disease transplantation diagnostic and treatment assistance system according to claim 1, characterized in that, The system also includes a map restoration guidance module: Based on the source tracing results of the logical closed-loop defects of the transplantation diagnosis and treatment strategy, the knowledge graph repair guidance module locates the missing condition nodes corresponding to the abnormal strategy, analyzes the candidate equivalent entities of the missing nodes in the knowledge graph, and generates a set of knowledge graph structure repair suggestion instructions. The knowledge graph structure repair suggestion instruction set includes semantic descriptions of missing nodes, a list of candidate equivalent entities, and coordinates of relation insertion positions.
9. The knowledge graph-based blood disease transplantation diagnostic and treatment assistance system according to claim 8, characterized in that, The atlas restoration guidance module includes: The missing location submodule extracts the logical breakpoint position of each abnormal strategy based on the source tracing results of the logical closed loop defects of the transplantation diagnosis and treatment strategy, analyzes the missing condition constraint type at the breakpoint, and generates a feature vector of the missing node. The equivalent entity retrieval submodule performs feature similarity matching in the same-level entity pool of the knowledge graph based on the feature vector of the missing node, uses the cosine similarity algorithm to filter candidate entities with similarity higher than a preset threshold, and outputs a list of candidate equivalent entities. The repair instruction generation submodule constructs a relation insertion operation instruction based on the candidate equivalent entity list and the topological coordinates of the breakpoint in the graph, specifies the source node, target node and relation type, and generates a set of knowledge graph structure repair suggestion instructions.
Citation Information
Patent Citations
Node set of a value-based systematized and fully-typed frequency calibration data map and method for determining topology structure thereof
CN109376217A
Medical clinical decision support method and system based on knowledge graph
CN121393835A