Man-machine collaborative intelligent writing method and system for medical documents
By using cross-modal semantic alignment and multi-level TCM knowledge graph reasoning extension, a structured dialectical framework is generated, which solves the problems of multimodal data processing and dynamic optimization in the TCM medical document generation system, realizes the professional and standardized expression of TCM medical documents, and improves the clinical adaptability of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NEWLINK TECH INC
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-24
AI Technical Summary
Existing intelligent medical document generation systems are unable to meet the specific needs of TCM syndrome differentiation and treatment, cannot effectively process multimodal clinical data, and lack dynamic optimization mechanisms. As a result, there is a significant gap between the generated content and actual clinical needs, affecting the accuracy of syndrome differentiation and the rationality of treatment plans.
By acquiring patients' multimodal clinical data for cross-modal semantic alignment, utilizing a multi-level TCM knowledge graph for reasoning expansion and path constraint propagation, a structured diagnostic framework is generated. Semantic conflicts and contradictions between syndrome differentiation and treatment are detected and resolved. The knowledge graph is then optimized in conjunction with actual clinical efficacy data to achieve dynamic optimization of the medical record generation model.
It achieves professional and standardized expression in TCM medical documents, reduces subjective bias and omissions, and enhances the system's clinical adaptability and practical value.
Smart Images

Figure CN121390080B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to human-computer collaboration technology, and more particularly to a human-computer collaborative intelligent writing method and system for medical documents. Background Technology
[0002] With the rapid development of medical informatization, electronic medical records have become an important part of clinical medical work. In traditional Chinese medicine (TCM) clinical practice, medical documentation is characterized by high professionalism, complex diagnostic reasoning, and standardized expression, placing high demands on doctors' professional competence and writing skills. Traditional TCM electronic medical record systems mainly rely on structured templates and manual data entry, which is insufficient to meet the personalized needs of TCM clinical diagnosis and treatment. In recent years, with the development of artificial intelligence technology, intelligent text generation technology based on deep learning and natural language processing has provided new possibilities for the automated writing of medical documents.
[0003] Current intelligent methods for writing medical documents suffer from the following shortcomings: First, most existing intelligent medical document generation systems are designed for Western medicine clinical scenarios, failing to fully consider the unique characteristics of traditional Chinese medicine (TCM) diagnosis and treatment. They struggle to accurately express the thought processes and professional knowledge system of TCM, resulting in a significant gap between the generated content and actual clinical needs. Second, existing technologies lack the ability to deeply integrate multimodal clinical data, making it difficult to effectively process heterogeneous data such as images, text, and audio from the four diagnostic methods of TCM. This hinders the comprehensive capture of patients' syndrome characteristics, affecting the accuracy of diagnosis and the rationality of treatment plans. Finally, existing medical document generation systems generally lack dynamic optimization mechanisms, failing to adaptively adjust their knowledge structure and generation strategies based on clinical practice feedback. This makes them ill-equipped to handle individual differences and complex syndrome variations in TCM clinical practice, limiting the system's practicality and sustainable development capabilities.
[0004] In view of the above problems, there is an urgent need to develop a human-computer collaborative intelligent medical record writing method that can integrate the theoretical system of traditional Chinese medicine, support multimodal data processing, and have dynamic optimization capabilities, so as to improve the writing efficiency and professional quality of TCM electronic medical records. Summary of the Invention
[0005] This invention provides a human-computer collaborative intelligent writing method and system for medical documents, which can solve the problems in the prior art.
[0006] A first aspect of this invention provides a human-computer collaborative intelligent writing method for medical documents, comprising:
[0007] Acquire multimodal clinical data of patients, perform cross-modal semantic alignment on the multimodal clinical data, and obtain a fused semantic vector by extracting TCM domain semantic features from each modality of data and mapping them to a unified semantic representation space.
[0008] Based on a pre-constructed multi-level TCM knowledge graph, the fused semantic vector is extended through reasoning, and a path constraint propagation mechanism is used to select combinations of diagnostic and treatment knowledge nodes that conform to the dialectical logic, thereby generating a structured dialectical framework.
[0009] Based on the semantic association strength and clinical applicability constraints of each knowledge node in the structured dialectical framework, a set of candidate medical record fragments is generated, and a logical consistency verification graph of the set of candidate medical record fragments is constructed. By detecting semantic conflicts and syndrome-treatment contradictions between medical record fragments in the logical consistency verification graph, and using the dialectical rules in the multi-level TCM knowledge graph to replace or merge conflicting fragments, a logically self-consistent complete medical record text is generated.
[0010] The complete medical record text is structurally broken down according to symptom characteristics, syndrome differentiation elements, and treatment plans, and a corresponding relationship is established with actual clinical efficacy data. Based on the association rule mining algorithm, the diagnosis and treatment paths and syndrome differentiation rules in the multi-level TCM knowledge graph are optimized to achieve dynamic optimization of the medical record generation model.
[0011] Based on a pre-constructed multi-level TCM knowledge graph, the fused semantic vector is extended through reasoning, and a path constraint propagation mechanism is used to select combinations of diagnostic and treatment knowledge nodes that conform to the dialectical logic, generating a structured dialectical framework including:
[0012] In a multi-level TCM knowledge graph, an initial set of syndrome nodes with a correlation degree exceeding a preset correlation threshold is located using a dynamic semantic matching strategy. Starting from the initial set of syndrome nodes, a multi-hop expansion is performed along the inter-level reasoning edges in the multi-level TCM knowledge graph. During the expansion process, the expansion path is pruned in real time according to the rules of dialectical logic constraints to generate a set of candidate reasoning paths.
[0013] For each reasoning path in the candidate reasoning path set, the semantic consistency score between adjacent nodes and the evidence-theory correspondence score between cross-level nodes are calculated. The semantic consistency score and the evidence-theory correspondence score are bidirectionally propagated and cumulatively updated in the path through the path constraint propagation mechanism to obtain the comprehensive credibility score of each reasoning path.
[0014] The candidate reasoning path set is hierarchically screened based on the comprehensive credibility score. The syndrome nodes, pathogenesis nodes, and treatment principle nodes in the effective reasoning paths that meet the screening conditions are extracted. The syndrome nodes, pathogenesis nodes, and treatment principle nodes are then organized in a structured manner according to the hierarchical order of TCM syndrome differentiation logic. Causal relationships and semantic dependencies between nodes are established to generate a structured dialectical framework.
[0015] Starting from the initial set of syndrome nodes, a multi-hop expansion is performed along the inter-level reasoning edges in the multi-level TCM knowledge graph. During the expansion process, the expansion path is pruned in real time according to the rules of dialectical logic constraints, generating a set of candidate reasoning paths, including:
[0016] The syndrome type identifier of each node is parsed from the initial syndrome node set, and the lower-level pathogenesis node is located in the multi-level TCM knowledge graph based on the syndrome type identifier.
[0017] For the lower-level pathogenesis nodes, an extended state vector containing the trajectory of syndrome evolution is constructed. The extended state vector is input into the syndrome differentiation logic constraint rules to judge the rationality of syndrome mechanism transformation. When the judgment shows that there is a causal contradiction between syndrome and pathogenesis or the transformation of disease nature violates the laws of traditional Chinese medicine theory, the extended path of the corresponding pathogenesis node is intelligently pruned.
[0018] For the pathogenesis nodes identified through the rationality analysis of the mechanism of disease progression, a progressive analysis is carried out along the inter-level reasoning edge towards the treatment principle level and the prescription level. At each cross-level extension, the extended state vector is enriched to integrate the treatment method adaptation characteristics and prescription meridian tropism elements of the newly traversed nodes. The updated extended state vector is evaluated for treatment principle fit and prescription compatibility is verified through the dialectical logic constraint rules. When it is found that the treatment method does not match the pathogenesis or that the prescription compatibility is contradictory or antagonistic, the path extension is stopped.
[0019] During the multi-hop expansion process, a path dynamic score is calculated for each retained path, and the expansion priority sequence is adjusted based on the path dynamic score. Paths with scores exceeding the retention threshold are included in the candidate inference path set.
[0020] Based on the semantic association strength and clinical applicability constraints of each knowledge node in the structured dialectical framework, a set of candidate medical record fragments is generated, and a logical consistency verification graph of the candidate medical record fragment set is constructed, including:
[0021] Syndrome nodes, pathogenesis nodes, and treatment principle nodes are extracted from the structured dialectical framework, and the semantic association and hierarchical inheritance relationship between any two knowledge nodes are determined through multi-dimensional semantic feature analysis.
[0022] Based on the core semantic features of each knowledge node, candidate expression templates with semantic expression fit higher than a preset fit threshold are identified in the pre-stored medical record text template library. Clinical applicability assessment criteria are constructed by combining the patient's age characteristics, physical constitution classification and disease evolution stage. The candidate expression templates are then screened and optimized and recombined at multiple levels to generate candidate medical record fragments corresponding to each knowledge node.
[0023] The knowledge node pairs with direct reasoning relationships and their corresponding candidate medical record fragments in the structured dialectical framework are organized into a fragment pair set. The semantic consistency representation and the compatibility characteristics of the diagnosis and treatment logic between the two medical record fragments in each fragment pair are determined through semantic content association analysis and diagnosis and treatment logic deduction.
[0024] Using each medical record fragment in the candidate medical record fragment set as a base node, and taking the semantic consistency representation and the diagnostic and treatment logic compatibility feature as the link, a logical consistency verification graph with a hierarchical reasoning structure is constructed.
[0025] By detecting semantic conflicts and contradictions in diagnosis and treatment between medical record fragments in the logical consistency verification graph, and using the diagnostic rules in the multi-level TCM knowledge graph to replace or merge conflicting fragments, a logically consistent complete medical record text is generated, including:
[0026] The edge weight features are analyzed in the logical consistency verification graph to construct a multi-level scoring matrix that integrates semantic similarity and syndrome-treatment correlation. Based on the multi-level scoring matrix, semantic conflicts and syndrome-treatment contradictions between medical record fragments are identified, and the inheritance of the fragments with semantic conflicts and syndrome-treatment contradictions in the multi-level TCM knowledge graph is traced.
[0027] Based on the syndrome differentiation rule chain in the multi-level TCM knowledge graph, the segments with semantic conflicts and syndrome-treatment contradictions are determined. The expression fit and syndrome-treatment consistency of the segments with the nodes in the syndrome differentiation rule chain are quantified by the multi-level scoring matrix. When the expression fit is lower than the fit threshold, it is classified as a semantic deepening segment. When the syndrome-treatment consistency deviates from the benchmark, it is classified as a rule reshaping segment.
[0028] For the semantically deepened segments, segments that highly resonate with the dialectical rule chain are selected from the candidate medical record segment set for semantic transformation, and the scoring elements are updated in the multi-level scoring matrix; for the pattern reshaping segments, the segment level is adjusted based on the core pathogenesis nodes in the dialectical rule chain, and the adjusted segments are re-placed into the multi-level scoring matrix for verification.
[0029] The validated medical record fragments are integrated according to the hierarchical sequence of the structured dialectical framework and their semantic coherence is evaluated to generate a logically consistent complete medical record text.
[0030] Based on the multi-level scoring matrix, semantic conflicts and contradictions in diagnosis and treatment between medical record fragments are identified. Tracing the inheritance of these semantically conflicting and contradictory fragments within a multi-level TCM knowledge graph includes:
[0031] Scoring elements for medical record fragment pairs are extracted from a multi-level scoring matrix. A two-level evaluation criterion is constructed based on semantic association threshold and syndrome-treatment correspondence threshold. By jointly judging the semantic expression similarity and syndrome-treatment rule fit of medical record fragment pairs in the multi-level scoring matrix, medical record fragments with semantic conflicts and syndrome-treatment contradictions are identified.
[0032] For the medical record fragments with semantic conflicts and contradictory treatments, a transmission tracing network is constructed in a multi-level TCM knowledge graph. The transmission tracing network maps the syndrome hierarchy relationship and treatment principle transformation process based on the chain of diagnostic rules between nodes, and determines the transmission path of the semantic conflicts and contradictory treatments at different levels.
[0033] The complete medical record text is structurally broken down according to symptom characteristics, syndrome differentiation elements, and treatment plans, and a correspondence is established with actual clinical efficacy data. The diagnosis and treatment paths and syndrome differentiation rules in the multi-level TCM knowledge graph are optimized based on association rule mining algorithms, including:
[0034] The complete medical record text is structured and broken down according to symptom characteristics, syndrome differentiation elements and treatment plan, and a symptom-syndrome differentiation-treatment correlation matrix is constructed. The actual clinical efficacy data is written into the corresponding position of the symptom-syndrome differentiation-treatment correlation matrix.
[0035] The symptom-diagnosis-treatment association matrix is analyzed using an association rule mining algorithm. The support and confidence of each node combination are calculated. Node combinations whose support and confidence exceed the optimization threshold are marked as items to be optimized. The items to be optimized include effectiveness score and reliability score.
[0036] Based on the analysis results of the item to be optimized by the association rule mining algorithm, the diagnosis and treatment path and syndrome differentiation and treatment rules corresponding to the item to be optimized are located in the multi-level TCM knowledge graph. The effectiveness score and the reliability score are written into the diagnosis and treatment path and the syndrome differentiation and treatment rules respectively, thus completing the clinical verification and optimization update of the knowledge graph.
[0037] A second aspect of the present invention provides a human-computer collaborative intelligent writing system for medical documents, comprising:
[0038] Unit 1 is used to acquire patients' multimodal clinical data, perform cross-modal semantic alignment on the multimodal clinical data, and obtain a fused semantic vector by extracting the semantic features of the TCM domain from each modality of data and mapping them to a unified semantic representation space.
[0039] Unit 2 is used to perform reasoning and expansion on the fused semantic vector based on a pre-constructed multi-level TCM knowledge graph, and to filter combinations of diagnostic and treatment knowledge nodes that conform to the dialectical logic through a path constraint propagation mechanism to generate a structured dialectical framework.
[0040] Unit 3 is used to generate a set of candidate medical record fragments based on the semantic association strength and clinical applicability constraints of each knowledge node in the structured dialectical framework, and to construct a logical consistency verification graph of the set of candidate medical record fragments; by detecting semantic conflicts and syndrome-treatment contradictions between medical record fragments in the logical consistency verification graph, and using the syndrome differentiation rules in the multi-level TCM knowledge graph to replace or merge conflicting fragments, a logically self-consistent complete medical record text is generated.
[0041] Unit 4 is used to structurally split the complete medical record text according to symptom characteristics, syndrome differentiation elements and treatment plans, and establish a corresponding relationship with actual clinical efficacy data. Based on the association rule mining algorithm, it optimizes the diagnosis and treatment path and syndrome differentiation rules in the multi-level TCM knowledge graph to achieve dynamic optimization of the medical record generation model.
[0042] A third aspect of the present invention provides an electronic device, comprising:
[0043] processor;
[0044] Memory used to store processor-executable instructions;
[0045] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0046] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0047] The beneficial effects of this application are as follows:
[0048] By performing cross-modal semantic alignment on multimodal clinical data of patients, the semantic features of traditional Chinese medicine in each modality of data are mapped to a unified semantic representation space, realizing the deep integration of different types of medical data, overcoming the limitations of traditional methods in handling heterogeneous data, and improving information utilization efficiency and representation accuracy.
[0049] By using a pre-constructed multi-level TCM knowledge graph for reasoning extension and employing a path constraint propagation mechanism to select combinations of diagnostic and treatment knowledge nodes that conform to the logic of syndrome differentiation, the computerized expression of TCM syndrome differentiation thinking is realized, ensuring the professionalism and standardization of the generated content and effectively reducing subjective bias and omissions in medical document writing.
[0050] By constructing a logical consistency verification graph of the candidate medical record fragment set, semantic conflicts and contradictions between evidence and treatment are detected and resolved. Combined with actual clinical efficacy data, the diagnosis and treatment paths and evidence and treatment rules in the knowledge graph are dynamically optimized, realizing the self-improvement and iterative optimization of the medical document generation process, and enhancing the clinical adaptability and long-term practical value of the system. Attached Figure Description
[0051] Figure 1 This is a flowchart illustrating the human-computer collaborative intelligent writing method for medical documents according to an embodiment of the present invention;
[0052] Figure 2 This is a flowchart illustrating the process of detecting and optimizing conflicts in medical record segments in an embodiment of the present invention. Detailed Implementation
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0055] Figure 1 This is a flowchart illustrating the human-computer collaborative intelligent writing method for medical documents according to an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes:
[0056] Acquire multimodal clinical data of patients, perform cross-modal semantic alignment on the multimodal clinical data, and obtain a fused semantic vector by extracting TCM domain semantic features from each modality of data and mapping them to a unified semantic representation space.
[0057] Based on a pre-constructed multi-level TCM knowledge graph, the fused semantic vector is extended through reasoning, and a path constraint propagation mechanism is used to select combinations of diagnostic and treatment knowledge nodes that conform to the dialectical logic, thereby generating a structured dialectical framework.
[0058] Based on the semantic association strength and clinical applicability constraints of each knowledge node in the structured dialectical framework, a set of candidate medical record fragments is generated, and a logical consistency verification graph of the set of candidate medical record fragments is constructed. By detecting semantic conflicts and syndrome-treatment contradictions between medical record fragments in the logical consistency verification graph, and using the dialectical rules in the multi-level TCM knowledge graph to replace or merge conflicting fragments, a logically self-consistent complete medical record text is generated.
[0059] The complete medical record text is structurally broken down according to symptom characteristics, syndrome differentiation elements, and treatment plans, and a corresponding relationship is established with actual clinical efficacy data. Based on the association rule mining algorithm, the diagnosis and treatment paths and syndrome differentiation rules in the multi-level TCM knowledge graph are optimized to achieve dynamic optimization of the medical record generation model.
[0060] In one optional implementation, based on a pre-constructed multi-level TCM knowledge graph, the fused semantic vector is extended through reasoning, and a path constraint propagation mechanism is used to filter combinations of diagnostic and treatment knowledge nodes that conform to the dialectical logic, generating a structured dialectical framework, including:
[0061] In a multi-level TCM knowledge graph, an initial set of syndrome nodes with a correlation degree exceeding a preset correlation threshold is located using a dynamic semantic matching strategy. Starting from the initial set of syndrome nodes, a multi-hop expansion is performed along the inter-level reasoning edges in the multi-level TCM knowledge graph. During the expansion process, the expansion path is pruned in real time according to the rules of dialectical logic constraints to generate a set of candidate reasoning paths.
[0062] For each reasoning path in the candidate reasoning path set, the semantic consistency score between adjacent nodes and the evidence-theory correspondence score between cross-level nodes are calculated. The semantic consistency score and the evidence-theory correspondence score are bidirectionally propagated and cumulatively updated in the path through the path constraint propagation mechanism to obtain the comprehensive credibility score of each reasoning path.
[0063] The candidate reasoning path set is hierarchically screened based on the comprehensive credibility score. The syndrome nodes, pathogenesis nodes, and treatment principle nodes in the effective reasoning paths that meet the screening conditions are extracted. The syndrome nodes, pathogenesis nodes, and treatment principle nodes are then organized in a structured manner according to the hierarchical order of TCM syndrome differentiation logic. Causal relationships and semantic dependencies between nodes are established to generate a structured dialectical framework.
[0064] A dynamic semantic matching strategy is used to accurately locate initial syndrome nodes in a multi-level TCM knowledge graph. The knowledge graph adopts a hierarchical architecture, with the syndrome layer containing 2847 standard syndrome nodes. Each node is encoded as a 768-dimensional semantic vector using a domain-optimized BERT model. The fused semantic vector represents a unified representation of the patient's multimodal clinical data and is also a 768-dimensional dense vector. The matching process first performs L2 norm normalization on the fused semantic vector and all syndrome node vectors, and then calculates the dot product similarity. To improve computational efficiency, an index structure based on locality-sensitive hashing is constructed, projecting the 768-dimensional space onto a 128-dimensional subspace and establishing 16 hash buckets for coarse screening. After coarse screening, precise similarity calculation is performed within the candidate set using a parallel computing framework, with the time taken for a single match controlled within 80 milliseconds. The preset association threshold is set to 0.72, determined based on statistical analysis of 5000 historical cases, which corresponds to an 85% accurate matching rate. Nodes exceeding the threshold are sorted by similarity, and the top 25 are selected to form the initial set of syndrome nodes. At the same time, the similarity score and ranking information of each node are recorded as prior knowledge for subsequent reasoning.
[0065] The multi-hop expansion mechanism performs directed traversal along the reasoning edges of the knowledge graph to discover complete dialectical reasoning links. These reasoning edges define the logical relationships between nodes, including "pathological deduction" edges from syndrome to pathogenesis, "treatment establishment" edges from pathogenesis to treatment principles, and "prescription selection" edges from treatment principles to prescriptions. Each edge is assigned a confidence weight, calculated using expert ratings and clinical validation data, ranging from 0.1 to 1.0. The expansion algorithm employs breadth-first search, traversing downwards layer by layer from the initial syndrome node, with a maximum expansion depth of 5 layers. During traversal, a path stack and access markers are maintained to prevent loops. The reachability and validity of the target node are checked during each expansion; invalid nodes include deleted nodes, low-frequency nodes, and experimental nodes. The expansion strategy prioritizes high-weight edges; when multiple candidate nodes exist at the same layer, expansion is performed in descending order of edge weight. To control the search space, an upper limit is set on the number of expansion branches, with each node expanding a maximum of 8 successor nodes. During the expansion process, the number and depth distribution of paths are counted in real time, and the expansion is terminated early when the number of candidate paths exceeds 300.
[0066] The dialectical logic constraint rules enable real-time pruning and quality control of the expanded paths. These rules are divided into hard and soft constraints. Hard constraints must be strictly satisfied, while soft constraints allow for moderate violations but will lower the path score. Hard constraints include mutual exclusion constraints and hierarchical constraints. Mutual exclusion constraints prohibit the simultaneous occurrence of contradictory syndromes, such as cold-heat syndromes, deficiency-excess syndromes, and exterior-interior syndromes, defining a total of 327 mutually exclusive relationships. Hierarchical constraints ensure that the reasoning path follows the order of TCM theory, prohibiting direct connections and reverse reasoning across levels. Soft constraints include co-occurrence constraints and frequency constraints. Co-occurrence constraints require that specific syndromes must be accompanied by corresponding treatments; a violation deducts 0.2 points. Frequency constraints require that core nodes in the path appear at a clinical frequency no less than a threshold; low-frequency nodes deduct 0.1 points. The pruning strategy performs constraint checks immediately after each node expansion. Path branches that violate hard constraints are discarded directly, while paths that violate soft constraints record the degree of violation and continue expansion. Constraint checks use efficient hash table lookups, with a single check taking less than 1 millisecond. The pruning process statistically analyzes the types and frequencies of constraint violations, providing data support for the optimization of constraint rules.
[0067] The semantic consistency score and syndrome-treatment correspondence score provide quantitative indicators for path quality assessment. The semantic consistency score evaluates the semantic coherence of adjacent nodes in the path, calculated using weighted cosine similarity. Weighting factors are determined based on the importance of nodes in the knowledge graph, and a modified PageRank algorithm is used to calculate node centrality, converging after 200 iterations. Core syndrome nodes have a weight of 1.0, important pathogenesis nodes have a weight of 0.9, general treatment principle nodes have a weight of 0.8, and auxiliary prescription nodes have a weight of 0.7. The semantic similarity of adjacent node pairs is calculated using vector dot products, with results ranging from -1 to 1; negative values indicate semantic conflict. The consistency score is a weighted average of the similarities of all adjacent node pairs in the path, with weights representing edge confidence. The syndrome-treatment correspondence score reflects the degree of conformity between TCM theories across nodes at different levels, calculated based on a correspondence matrix constructed by experts. The matrix is stored in a COO sparse format and contains 450,000 valid correspondences. Each correspondence includes a base score and conditional adjustments. The base score is assessed by domain experts, while the adjustments consider individual factors such as patient age, gender, disease duration, and physical condition. The correspondence score is calculated by querying relevant correspondences and taking a weighted average as the final score.
[0068] The path constraint propagation mechanism enables bidirectional flow and cumulative updating of score information within the reasoning path. Forward propagation begins at the syndrome node at the path's starting point, transmitting semantic consistency information along the reasoning direction. The propagation formula combines the current node's score, outgoing edge weights, and a propagation attenuation factor, with the attenuation factor set to 0.95 to simulate the natural loss of information transmission. When each node receives propagation information from multiple predecessor nodes, a maximum value selection strategy is adopted to avoid information redundancy. Backward propagation begins at the prescription node at the path's ending point, transmitting syndrome-treatment correspondence information in reverse. Backward propagation has a larger weight, with an attenuation factor set to 0.98, reflecting the strong constraint effect of treatment effect on the syndrome differentiation process. Bidirectional propagation information is weighted and fused when converging at intermediate nodes, with a forward weight of 0.4 and a backward weight of 0.6, reflecting the characteristic of TCM's "efficacy verifying pathogenesis." The propagation process employs an asynchronous update mechanism, where node score updates do not depend on the synchronization status of other nodes. Convergence is determined based on the global score change; convergence is considered achieved when the change is less than 0.008 for five consecutive iterations, with a maximum iteration count limited to 30 rounds. The propagation algorithm supports incremental updates, eliminating the need to recalculate all scores when adding new inference paths. The overall credibility score is calculated using the harmonic mean, effectively suppressing the negative impact of abnormally low-scoring nodes and ensuring the reliability of the overall path quality.
[0069] The hierarchical screening of candidate reasoning paths employs a multi-stage filtering strategy to progressively improve path quality. The first stage, basic filtering, sets a comprehensive credibility threshold of 0.6 to eliminate obviously unreasonable reasoning paths, retaining approximately 60% of the candidate paths. The second stage, relative filtering, ranks paths based on their scores, retaining the top 70% of high-scoring paths to avoid data loss caused by absolute thresholds. The third stage, diversity filtering, prevents excessive path similarity by merging similar paths using a clustering algorithm. Similarity calculation is based on a linear combination of the Jaccard coefficient of the path node set and semantic similarity, with a weighting ratio of 6:4. Paths with a Jaccard coefficient greater than 0.75 are considered highly similar, and the highest-scoring representative path is retained from each similarity cluster. The fourth stage, integrity verification, requires that retained paths cover the three core levels of syndrome, pathogenesis, and treatment principles, with a path length between 3 and 6 nodes. Detailed statistical information is recorded throughout the screening process, including filtering ratios, score distributions, and length distributions for each stage. Screening criteria also include clinical feasibility checks, requiring that the treatment methods and prescriptions in the paths have precedents in actual clinical practice; rare or experimental treatment methods are marked as low priority.
[0070] The extraction of syndrome nodes, pathogenesis nodes, and treatment principle nodes in the effective reasoning path is categorized and organized according to the TCM theoretical system. Syndrome nodes are extracted by distinguishing between primary and secondary symptoms; the primary symptom reflects the fundamental pathogenesis of the disease, while secondary symptoms describe accompanying manifestations. Pathogenesis nodes are divided into superficial and deep pathogenesis based on pathological levels; superficial pathogenesis describes dysfunction of the internal organs, while deep pathogenesis reveals imbalances in Qi, blood, Yin, and Yang. Treatment principle nodes are divided into treatment method nodes and prescription / medication nodes; treatment method nodes define the basic principles of treatment, while prescription / medication nodes specify the specific medication regimen. During the extraction process, the distribution frequency and co-occurrence patterns of various nodes are statistically analyzed to provide a data foundation for structured organization. Node deduplication uses a semantic similarity threshold method; nodes of the same type with a similarity greater than 0.9 are merged into a single representative node. The extraction results are sorted according to the importance of TCM theory, with core nodes ranked first and secondary nodes ranked last.
[0071] The structured dialectical framework organizes extracted nodes according to the logical order of TCM syndrome differentiation and treatment. The framework adopts a hierarchical directed graph structure, where nodes represent TCM conceptual entities, and directed edges represent the logical relationships between concepts. The syndrome layer is located at the top level, containing the patient's main syndrome manifestations and symptom characteristics. The pathogenesis layer is located in the middle layer, describing the internal mechanisms of disease occurrence and development. The treatment principle layer is located at the bottom level, specifying the corresponding treatment principles and specific measures. Inter-layer connections are achieved through causal association edges; the direction of the edge indicates the direction of reasoning, and the edge weight reflects the strength of the association. Causal associations are divided into direct causality and indirect causality. Direct causality indicates reachability in one step of reasoning, while indirect causality indicates reachability in multiple steps of reasoning. Semantic dependencies characterize the intrinsic connections between concepts and are quantified through a semantic similarity matrix. Similarity greater than 0.8 establishes a strong dependency, 0.5 to 0.8 establishes a weak dependency, and less than 0.5 establishes no dependency. The framework construction process implements strict consistency checks, including logical consistency, semantic consistency, and structural integrity. Logical consistency requires the framework to not violate TCM theoretical rules; semantic consistency requires related concepts to be semantically compatible; and structural integrity requires the framework to have a complete reasoning chain. The checking algorithm uses depth-first traversal, starting from the symptom node to check the reachability to the treatment node, and at the same time verifying the rationality of the path.
[0072] In one optional implementation, starting from the initial set of syndrome nodes, multi-hop expansion is performed along the inter-level reasoning edges in the multi-level TCM knowledge graph. During the expansion process, the expansion path is pruned in real time according to the rules of dialectical logic constraints, generating a set of candidate reasoning paths, including:
[0073] The syndrome type identifier of each node is parsed from the initial syndrome node set, and the lower-level pathogenesis node is located in the multi-level TCM knowledge graph based on the syndrome type identifier.
[0074] For the lower-level pathogenesis nodes, an extended state vector containing the trajectory of syndrome evolution is constructed. The extended state vector is input into the syndrome differentiation logic constraint rules to judge the rationality of syndrome mechanism transformation. When the judgment shows that there is a causal contradiction between syndrome and pathogenesis or the transformation of disease nature violates the laws of traditional Chinese medicine theory, the extended path of the corresponding pathogenesis node is intelligently pruned.
[0075] For the pathogenesis nodes identified through the rationality analysis of the mechanism of disease progression, a progressive analysis is carried out along the inter-level reasoning edge towards the treatment principle level and the prescription level. At each cross-level extension, the extended state vector is enriched to integrate the treatment method adaptation characteristics and prescription meridian tropism elements of the newly traversed nodes. The updated extended state vector is evaluated for treatment principle fit and prescription compatibility is verified through the dialectical logic constraint rules. When it is found that the treatment method does not match the pathogenesis or that the prescription compatibility is contradictory or antagonistic, the path extension is stopped.
[0076] During the multi-hop expansion process, a path dynamic score is calculated for each retained path, and the expansion priority sequence is adjusted based on the path dynamic score. Paths with scores exceeding the retention threshold are included in the candidate inference path set.
[0077] The syndrome type identifier parsing module extracts the standardized syndrome code and attribute label for each node from the initial syndrome node set. The syndrome nodes employ a five-layer nested coding structure, including organ affiliation coding, disease classification coding, deficiency / excess attribute coding, cold / heat attribute coding, and specific syndrome coding. Each layer has a fixed coding length of three digits. The organ affiliation coding ranges from 001 to 012, corresponding to the twelve organ systems. The disease classification coding ranges from 101 to 199, covering pathological products such as qi, blood, body fluids, phlegm, and blood stasis. The deficiency / excess attribute coding uses binary representation, with 0 indicating deficiency and 1 indicating excess. The cold / heat attribute coding also uses binary representation, with 0 indicating cold and 1 indicating heat. The syndrome type identifier parsing module receives a 768-dimensional initial syndrome node vector and extracts the coding information through a pre-trained syndrome classifier. The classifier adopts a multi-task learning architecture, containing five parallel fully connected branches, each corresponding to a layer of coding classification tasks. The classifier input layer accepts normalized node vectors, the hidden layers use a two-layer network of 512 dimensions and 256 dimensions, and the output layer generates the probability distribution of each coding layer through a softmax activation function. When extracting codes, a confidence threshold of 0.85 is set. Codes below the threshold are marked as uncertain and trigger a manual review process.
[0078] The localization of lower-level pathogenesis nodes is performed through directed graph traversal in a multi-level TCM knowledge graph based on syndrome type identifiers. The knowledge graph uses an adjacency list storage structure, and each syndrome node maintains a list of outgoing edges pointing to lower-level pathogenesis nodes. Each outgoing edge contains three fields: target node identifier, edge weight, and association type. The pathogenesis node hierarchy contains 2156 standard pathogenesis entities, classified in two dimensions according to the location and nature of the disease. The location dimension includes 15 categories (five Zang-Fu organs and six Fu organs), and the nature dimension includes eight categories (deficiency / excess, cold / heat, dampness / dryness, wind / fire). The localization algorithm traverses all outgoing edges of the initial syndrome node, filtering lower-level pathogenesis nodes with edge weights greater than 0.6 and an association type of "syndrome-mechanism association." The association strength calculation integrates expert ratings and clinical data statistics. Expert ratings were collected from 15 TCM experts using the Delphi method, and the statistical results were based on co-occurrence frequency analysis of 30,000 medical records. The pathogenesis node selection process also considers node activation and stability indicators. Activation reflects the node's activity level in the inference network, while stability reflects the consistency of the node's associations with other nodes. Pathogenic nodes with activation levels below 0.3 or stability levels below 0.5 are marked as low-quality nodes and ranked lower in the localization results.
[0079] The extended state vector construction integrates syndrome node information and syndrome evolution trajectory features into a unified vector representation. The state vector adopts a 1024-dimensional fixed-length structure, with the first 768 dimensions retaining the semantic features of the original syndrome nodes, and the last 256 dimensions encoding the syndrome evolution trajectory information. The trajectory features include four components: disease stage identifiers, syndrome transformation direction, evolution intensity index, and time-dependent factors. Disease stage identifiers use 64-dimensional one-hot encoding, covering major stages such as acute phase, chronic phase, recovery phase, and remission phase. Syndrome transformation direction is represented by a 96-dimensional dense vector, with each dimension corresponding to a syndrome transformation pattern, such as from superficial to deep, from excess to deficiency, and from heat to cold. Evolution intensity index uses a 32-dimensional numerical vector to record the quantitative value of the degree of various pathological changes, ranging from 0 to 1, where 0 represents no change and 1 represents drastic change. Time-dependent factors use 64-dimensional temporal encoding to reflect the temporal characteristics and periodic patterns of syndrome evolution. The vector construction process uses a feature fusion network, which contains two parallel branches to process static syndrome features and dynamic trajectory features respectively. The fusion layer adaptively adjusts the weight ratio of the two types of features through an attention mechanism.
[0080] The rationality assessment of syndrome progression is performed by multi-dimensionally verifying the extended state vector using dialectical logic constraint rules. The constraint rule base contains 867 expert-defined progression rules, represented by condition-action-result triplets. The condition describes the matching pattern of syndrome features, the action defines the logical operation of the progression, and the result gives the expected pathogenesis state. Causal contradiction detection is achieved by calculating the cosine distance between semantic vectors. A causal contradiction is identified when the cosine distance between the syndrome vector and the corresponding pathogenesis vector is greater than 0.4. Pathogenesis transformation verification is based on a state transition matrix established according to the Five Elements theory of Traditional Chinese Medicine. The matrix is 64x64, with rows and columns corresponding to various pathogenesis states. Matrix elements represent the rationality score of state transformations. The score ranges from -1 to 1, with positive values indicating reasonable transformations and negative values indicating violations of theory. The absolute value indicates the strength of credibility. The assessment process uses a parallel computing framework, with a single verification time controlled within 15 milliseconds. Transition paths that violate theoretical rules are recorded in a pruning log, including the type, severity, and specific cause of the violation, providing feedback data for rule optimization.
[0081] The progressive analysis module performs a directed traversal along the inter-level reasoning edges towards the treatment principle level and the prescription / medicine level. The treatment principle level contains 498 standard treatment method nodes, categorized according to the three principles of tonification, purgation, and harmonization. Each treatment method node is associated with 3 to 8 specific treatment methods. The prescription / medicine level contains 1247 classic prescription nodes and 2834 single-herb nodes. Prescription nodes record composition, usage, dosage, and indications, while herb nodes record properties, channels, efficacy, indications, and contraindications. The inter-level reasoning edges adopt a directed weighted graph structure, with edge weights reflecting the tightness of the association and the reliability of clinical validation. The weights of the reasoning edges from the pathogenesis level to the treatment principle level are calculated based on the targeting of the treatment method to the pathogenesis, with weights ranging from 0.1 to 1.0. Edges with weights greater than 0.7 are considered to have strong associations. The weights of the reasoning edges from the treatment principle level to the prescription / medicine level are calculated based on the degree to which the prescription / medicine implements the treatment method, also using a weight range of 0.1 to 1.0. The progressive analysis employs a dynamic programming algorithm to maintain the optimal arrival path and cumulative weight for each node, avoiding redundant calculations and improving traversal efficiency. The traversal depth is limited to 5 hops; paths exceeding this depth are automatically terminated.
[0082] The extended state vector update integrates the treatment method adaptation features and the meridian tropism elements of the new traversal nodes. The treatment method adaptation features are represented by 128-dimensional vectors, encoding information such as the mechanism of action, applicable pathogenesis, and contraindications of the treatment method. These features are generated through a pre-trained treatment method encoder, trained on a large-scale TCM literature database containing 30,000 TCM papers and 150,000 medical case records. The meridian tropism elements are represented by 192-dimensional vectors, containing attributes such as the four natures and five flavors of the drug, its meridian tropism, toxicity level, and compatibility / contraindications. The meridian tropism element vectors are constructed using a traditional Chinese medicine database, which includes complete information on over 5,000 kinds of Chinese medicinal materials. The state vector update employs a residual connection mechanism. New features are mapped to the same dimensional space as the original vector through a fully connected layer, and then weighted and added to the original vector. The weight coefficients are dynamically adjusted according to the importance of the new node in the inference path: the weight of the core treatment method node is set to 0.8, the weight of the auxiliary treatment method node is set to 0.5, the weight of the main drug node is set to 0.7, and the weight of the adjuvant drug node is set to 0.3.
[0083] The assessment of treatment principle fit is achieved through a dual mechanism of semantic matching and theoretical verification. Semantic matching calculates the vector similarity between treatment principle nodes and pathogenesis nodes; a similarity greater than 0.75 is considered a good fit. Theoretical verification is based on rule matching using a TCM theoretical knowledge base, which contains 236 treatment principle theories and 1540 specific application norms. The fit score is calculated by combining the semantic matching score and the theoretical verification results using a weighted average, with semantic matching having a weight of 0.6 and theoretical verification having a weight of 0.4. Treatment principles with a score below 0.65 are judged as non-fit, and the corresponding expansion path is marked as low confidence. The assessment process supports multi-threaded parallel computation, utilizing vectorization operations to accelerate similarity calculation, with the time taken for a single assessment controlled within 8 milliseconds.
[0084] The formula compatibility verification examines the interactions and contraindications between drugs within a formula. The compatibility rules are established based on the traditional theories of "Eighteen Incompatibilities" and "Nineteen Antagonisms," as well as modern pharmacological research. The rule base contains 327 pairs of incompatible combinations and 198 pairs of antagonistic combinations. Incompatible combinations refer to drugs that, when used together, will produce toxic side effects or reduce efficacy; antagonistic combinations refer to one drug that can reduce the toxicity of another. The verification algorithm traverses all drug pairs in the formula, querying the compatibility rule base to check for contraindications. When incompatible or antagonistic combinations are found, the specific drug combination and contraindication type are recorded, the path is marked as inappropriate, and further exploration is stopped. The verification process also checks the rationality of drug dosages, verifying whether the dosage of a single herb is within the safe range based on a traditional Chinese medicine dosage database. The database contains information on the commonly used dosage, maximum dosage, and toxic dosage of 4200 kinds of traditional Chinese medicines; formulas with dosages exceeding the safe range are judged as unreasonable.
[0085] The path dynamic scoring system calculates the quality indicators of each retained path in real time during multi-hop expansion. The scoring system includes four dimensions: semantic coherence score, theoretical compliance score, clinical applicability score, and path integrity score. Semantic coherence score assesses the semantic association strength between adjacent nodes in the path, calculated using the cosine similarity of node vectors; the average similarity of all adjacent node pairs in the path is used as the coherence score. Theoretical compliance score assesses whether the path follows the basic principles of TCM syndrome differentiation and treatment, using a matching verification based on a TCM theoretical rule base. Node pairs that conform to the rules receive positive scores, while those that violate the rules receive negative scores. Clinical applicability score reflects the frequency and effectiveness of the treatment methods and prescriptions in clinical practice, calculated statistically based on an electronic medical record database; treatment methods and prescriptions with high application frequency and good efficacy receive higher scores. Path integrity score assesses the logical completeness and hierarchical coverage of the reasoning chain. A complete path should include four levels: syndrome, pathogenesis, treatment principle, and prescription; paths lacking key levels receive lower scores. The dynamic scoring adopts a real-time update mechanism. Whenever a new node is added to the path, the scores of each dimension are recalculated immediately. The total score is a weighted average of the scores of the four dimensions, with weights of 0.3, 0.3, 0.2, and 0.2, respectively.
[0086] The priority sequence adjustment is based on a dynamic path rating system to implement an intelligent search strategy. The priority sequence is maintained using a max-heap data structure, where the top element is the path with the highest current rating. New paths are inserted into appropriate positions based on their ratings. The priority sequence adjustment frequency is set to a global reordering every 100 node expansions, balancing computational efficiency and search quality. The reordering process uses the quicksort algorithm, with a time complexity of O(nlogn), where n is the number of current paths. To avoid getting trapped in local optima, a random perturbation mechanism is introduced, randomly adjusting the priority of paths with similar ratings (defined as a rating difference of less than 0.05). During the expansion process, high-rated paths are prioritized for extension while retaining a certain proportion of low-rated paths to maintain search diversity. The probability of selecting a high-rated path is 0.7, and the probability of selecting a low-rated path is 0.3.
[0087] The retention threshold setting and candidate inference path set construction ensure the quality of the final output paths. The retention threshold is determined based on historical data statistics, set according to the path score distribution of 5000 labeled cases. The threshold selection ensures that the accuracy of retained paths reaches over 85%. A dynamic threshold adjustment mechanism adaptively corrects the threshold based on the quality distribution of the current expanded results. When the proportion of high-quality paths is too low, the threshold is lowered to increase the number of candidates; when there are too many candidate paths, the threshold is raised to control the scale. The candidate inference path set is stored using a linked list structure. Each path node records the complete inference chain, cumulative score, confidence level, and other information. The set size is limited to 200 paths; if the limit is exceeded, the path with the lowest score is automatically discarded. Path deduplication is implemented based on the hash value of the node sequence; only the path with the highest score is retained for paths with the same hash value.
[0088] In one optional implementation, a set of candidate medical record fragments is generated based on the semantic association strength and clinical applicability constraints of each knowledge node in the structured dialectical framework, and the logical consistency verification graph of the set of candidate medical record fragments is constructed, including:
[0089] Syndrome nodes, pathogenesis nodes, and treatment principle nodes are extracted from the structured dialectical framework, and the semantic association and hierarchical inheritance relationship between any two knowledge nodes are determined through multi-dimensional semantic feature analysis.
[0090] Based on the core semantic features of each knowledge node, candidate expression templates with semantic expression fit higher than a preset fit threshold are identified in the pre-stored medical record text template library. Clinical applicability assessment criteria are constructed by combining the patient's age characteristics, physical constitution classification and disease evolution stage. The candidate expression templates are then screened and optimized and recombined at multiple levels to generate candidate medical record fragments corresponding to each knowledge node.
[0091] The knowledge node pairs with direct reasoning relationships and their corresponding candidate medical record fragments in the structured dialectical framework are organized into a fragment pair set. The semantic consistency representation and the compatibility characteristics of the diagnosis and treatment logic between the two medical record fragments in each fragment pair are determined through semantic content association analysis and diagnosis and treatment logic deduction.
[0092] Using each medical record fragment in the candidate medical record fragment set as a base node, and taking the semantic consistency representation and the diagnostic and treatment logic compatibility feature as the link, a logical consistency verification graph with a hierarchical reasoning structure is constructed.
[0093] Knowledge node extraction parses three types of core nodes and their attribute information from the structured dialectical framework. Syndrome node extraction employs a depth-first traversal algorithm, recursively visiting all syndrome type nodes starting from the top-level syndrome layer of the framework. Each syndrome node contains five basic fields: node identifier, syndrome name, affiliated organ, deficiency / excess attribute, and cold / heat attribute. The node identifier uses an eight-digit code: the first two digits indicate the affiliated organ, the middle two digits indicate the syndrome type, and the last four digits are the specific syndrome number. Pathogenesis node extraction traverses the pathogenesis entities in the middle layer of the framework. The pathogenesis node structure contains five core fields: node identifier, pathogenesis description, disease location information, disease characteristics, and pathogenesis mechanism. Disease location information uses hexadecimal encoding to represent the involved organs and meridians, and disease characteristics are marked with attributes such as deficiency / excess, cold / heat, dampness, and dryness through bitwise operations. Treatment principle node extraction accesses the treatment principles and specific prescriptions at the bottom layer of the framework. Treatment principle nodes contain five key fields: node identifier, treatment name, applicable pathogenesis, contraindicated syndromes, and prescription composition. The prescription composition field stores a list of drugs in JSON format, with each drug entry containing detailed information such as drug name, dosage, and processing method. The node extraction process employs multi-threaded concurrent processing; the extraction tasks for syndrome nodes, pathogenesis nodes, and treatment principle nodes are assigned to three independent threads, exchanging intermediate results through a shared memory area. After extraction, nodes are deduplicated and their validity is checked, removing duplicate nodes and invalid nodes with incomplete information.
[0094] Multi-dimensional semantic feature analysis calculates the degree of association and hierarchical relationship between any two knowledge nodes. The semantic feature vectors are generated using a pre-trained BERT model in the field of Traditional Chinese Medicine (TCM). The model is trained on 500,000 ancient TCM texts and 300,000 modern medical records, with a vector dimension of 768. The degree of association between node pairs is calculated using vector cosine similarity, ranging from -1 to +1. Positive values indicate semantic similarity, negative values indicate semantic opposites, and the absolute value represents the strength of the association. Hierarchical inheritance relationships are determined through directed graph analysis. The path length from a higher-level node to a lower-level node reflects the inheritance distance, and the cumulative path weight reflects the inheritance strength. Inheritance strength is calculated using the geometric mean of path weights to avoid weight dilution caused by long paths. Semantic association analysis also considers the centrality of nodes in the knowledge graph, using betweenness centrality and close centrality to measure the importance of nodes. Betweenness centrality reflects the bridging role of a node in the entire graph, while close centrality reflects the average distance between a node and other nodes. Nodes with high centrality receive greater weight in association calculations, reflecting their core position in the dialectical system. The correlation degree calculation adopts a weighted fusion strategy, with semantic similarity weight of 0.6, inheritance strength weight of 0.3, and centrality index weight of 0.1.
[0095] The medical record text template library stores standardized expression templates corresponding to different types of knowledge nodes. The library employs a hierarchical index structure. The first-level index is divided into three main categories based on knowledge node type: syndrome templates, pathogenesis templates, and treatment principle templates. The second-level index is further subdivided according to organ affiliation, and the third-level index is organized according to specific syndrome, pathogenesis, and treatment type. Each template contains three forms: core expression, synonym expression, and negation expression. The core expression is the most commonly used standard description, the synonym expression provides diverse expressions, and the negation expression is used to exclude opposite cases. The template content is generated based on 50,000 high-quality medical record texts, and quality is ensured through a combination of frequency statistics and expert review. Semantic expression fit is measured by the similarity between the template vector and the node vector. The similarity calculation uses an improved cosine distance, and length normalization is added to eliminate the influence of text length on similarity. The fit threshold is set at 0.72, determined based on statistical analysis of 1000 labeled samples. This threshold corresponds to a 78% exact match rate and an 85% recall rate. Template retrieval is accelerated using an inverted index, mapping node keywords to a candidate template set and quickly narrowing the search scope through Boolean queries.
[0096] The clinical applicability assessment criteria integrate patient individual characteristics and disease status information for template selection. Age characteristics are divided into four stages: childhood, adolescence, middle age, and old age, each with different physiological characteristics and medication characteristics. In the childhood stage, the internal organs are delicate, and the syndrome description emphasizes external pathogens and food stagnation, with treatment tending towards mild and warming tonification. In the old age stage, the internal organs are weak, and the syndrome description emphasizes deficiency and phlegm stagnation, with treatment emphasizing nourishment and blood circulation. The constitution classification uses nine constitution types: balanced, qi deficiency, yang deficiency, yin deficiency, phlegm-dampness, damp-heat, blood stasis, qi stagnation, and special constitution. Each constitution corresponds to different susceptible diseases and suitable treatments. Constitution information is obtained through a constitution scale assessment, which contains 60 standardized questions, each corresponding to a specific constitution assessment dimension. The disease evolution stages are divided into four phases: acute, subacute, chronic, and recovery. The syndrome characteristics and treatment focuses differ significantly between these stages. The acute phase focuses on eliminating pathogens, the chronic phase emphasizes strengthening the body's resistance, and the recovery phase stresses recuperation. The applicability assessment uses a multilayer perceptron model. The input layer receives a fused vector of node features, patient features, and disease features. The hidden layers consist of two network layers: a 256-dimensional layer and a 128-dimensional layer. The output layer generates an applicability score. The score ranges from 0 to 1, with higher scores indicating a more suitable template for the current patient's condition.
[0097] The candidate description templates are screened and optimized to generate personalized medical record fragment descriptions. The screening process employs a multi-level filtering mechanism. The first level is based on semantic fit, retaining candidate templates with a fit higher than a threshold. The second level is based on clinical applicability assessment, excluding templates unsuitable for the current patient. The third level is based on description diversity, avoiding the generation of overly similar repetitive descriptions. Diversity is measured by calculating the edit distance between templates; templates with an edit distance less than 5 are considered highly similar, and only the one with the highest score is retained. The optimization and restructuring uses template fusion technology to merge the core information of multiple related templates to generate new descriptions. The fusion algorithm is based on an attention mechanism, automatically identifying key information fragments in each template and combining them according to their importance weight. The restructuring process also considers the language habits and logical order of traditional Chinese medicine descriptions to ensure that the generated fragments conform to the narrative norms of traditional medicine. Each knowledge node generates 3 to 5 candidate medical record fragments, with fragment lengths controlled between 20 and 100 characters, ensuring both information integrity and ease of subsequent processing.
[0098] The fragment pair set construction identifies knowledge node pairs and their corresponding fragments with direct inference associations. These direct inference associations are determined through edge relationships within a structured dialectical framework. Edge types include "symptom-mechanism association" (syndrome-pathogenesis association), "mechanism-treatment association" (pathogenesis-treatment principle association), and "synergistic association" within treatment principles. Association strength is quantified by edge weights; edges with weights greater than 0.5 are considered significantly related. The fragment pair set is stored using an adjacency list, with each node maintaining a list of directly associated nodes and their corresponding fragment combinations. Fragment combinations are generated using a Cartesian product approach, pairing all candidate fragments from the source node with all candidate fragments from the target node. To control the combination size, each node pair retains a maximum of 25 fragment combinations, selected based on semantic relevance. Fragment pairs also contain contextual information, recording the positional relationship and semantic roles of the two nodes in the inference path, providing a reference for subsequent consistency analysis. The set construction process employs parallel computing, utilizing multi-core processors to process different node pairs simultaneously, improving processing efficiency through task sharding and result merging.
[0099] Semantic consistency is used to characterize the content association between two medical record fragments within a fragment pair. A dual-encoder architecture is employed, with two encoders processing the source and target fragments respectively, generating 512-dimensional fragment vector representations. The encoders utilize a Transformer-based pre-trained model, trained on large-scale TCM text data for domain adaptation. Semantic consistency is calculated using a combination of the dot product and cosine similarity of the fragment vectors. The dot product reflects the absolute correlation between the vectors, while the cosine similarity reflects directional consistency. A weighted average is used for consistency scores, with a dot product weight of 0.4 and a cosine similarity weight of 0.6. Scores range from 0 to 1, with high scores indicating high semantic consistency and low scores indicating semantic conflict. A consistency threshold of 0.68 is set based on expert-annotated fragment pair samples, corresponding to 82% accuracy and 79% completeness. The characterization analysis also considers semantic role matching, identifying key components in the fragments through named entity recognition and semantic role annotation, and checking whether the components of the two fragments are semantically compatible.
[0100] The logical compatibility feature assessment evaluates whether the fragment pairs conform to the logical relationship of syndrome differentiation and treatment in Traditional Chinese Medicine (TCM). The feature assessment is based on logical reasoning using a TCM theoretical rule base, which contains 1247 expert-defined syndrome-treatment association rules and 934 contraindication / conflict rules. Each rule is represented in an if-then format, with the condition describing the characteristic pattern of the syndrome or pathogenesis and the conclusion providing the corresponding treatment principles or contraindications. Logical compatibility is assessed through rule matching and conflict detection. Successfully matched rules increase the compatibility score, while conflicting rules decrease it. The score calculation employs an evidence accumulation mechanism: the weights of multiple supporting rules are accumulated as positive scores, and the weights of conflicting rules are accumulated as negative scores. The final score is the normalized difference between positive and negative scores. Rule matching uses a pattern matching algorithm, supporting both exact and fuzzy matching modes. Fuzzy matching allows for partial differences in the condition parts, with the degree of difference measured by edit distance. The compatibility feature also includes temporal logic verification, checking whether the syndrome-treatment relationship described by the fragment pairs conforms to the chronological order of disease development and treatment.
[0101] The logical consistency verification graph construction uses candidate medical record fragments as nodes and consistency representations and compatibility features as edges. The graph employs a directed weighted graph structure, where node weights reflect the quality and importance of fragments, and edge weights reflect the strength of association and logical consistency between fragments. Node weights comprehensively consider three dimensions: semantic richness, clinical applicability, and standardization of expression, and are calculated using linear weighting. Edge weights integrate semantic consistency representations and diagnostic-treatment logical compatibility features, with a semantic consistency weight of 0.6 and a logical compatibility weight of 0.4. Graph construction uses an incremental method, adding fragment nodes and associated edges one by one, detecting and eliminating loops and conflicts in the graph in real time during the addition process. The hierarchical reasoning structure is implemented through topological sorting, arranging nodes according to reasoning levels to ensure that the reasoning results of upper-level nodes can be passed to lower-level nodes. The graph also supports dynamic updates, automatically adjusting relevant nodes and edges when fragments are added or modified to maintain the integrity and consistency of the graph.
[0102] In one optional implementation, by detecting semantic conflicts and diagnostic-treatment contradictions between medical record fragments in the logical consistency verification graph, and using the diagnostic rules in the multi-level TCM knowledge graph to replace or merge conflicting fragments, a logically consistent complete medical record text is generated, including:
[0103] The edge weight features are analyzed in the logical consistency verification graph to construct a multi-level scoring matrix that integrates semantic similarity and syndrome-treatment correlation. Based on the multi-level scoring matrix, semantic conflicts and syndrome-treatment contradictions between medical record fragments are identified, and the inheritance of the fragments with semantic conflicts and syndrome-treatment contradictions in the multi-level TCM knowledge graph is traced.
[0104] Based on the syndrome differentiation rule chain in the multi-level TCM knowledge graph, the segments with semantic conflicts and syndrome-treatment contradictions are determined. The expression fit and syndrome-treatment consistency of the segments with the nodes in the syndrome differentiation rule chain are quantified by the multi-level scoring matrix. When the expression fit is lower than the fit threshold, it is classified as a semantic deepening segment. When the syndrome-treatment consistency deviates from the benchmark, it is classified as a rule reshaping segment.
[0105] For the semantically deepened segments, segments that highly resonate with the dialectical rule chain are selected from the candidate medical record segment set for semantic transformation, and the scoring elements are updated in the multi-level scoring matrix; for the pattern reshaping segments, the segment level is adjusted based on the core pathogenesis nodes in the dialectical rule chain, and the adjusted segments are re-placed into the multi-level scoring matrix for verification.
[0106] The validated medical record fragments are integrated according to the hierarchical sequence of the structured dialectical framework and their semantic coherence is evaluated to generate a logically consistent complete medical record text.
[0107] like Figure 2 As shown, the method includes:
[0108] Edge weight feature parsing extracts the correlation information between segments from the logical consistency verification graph, constructing the data foundation for conflict detection. Edge weights consist of two core components: semantic similarity and causal correlation. The semantic similarity component uses 64-bit double-precision floating-point numbers, ranging from 0 to 1, with precision maintained to six decimal places. The causal correlation component also uses 64-bit double-precision floating-point numbers, ranging from -1 to +1. Positive values indicate logical support between the two sides, while negative values indicate logical conflict. The absolute value represents the correlation strength. The edge weight parsing module uses an adjacency matrix data structure. The matrix dimension equals the total number of nodes in the graph, and the matrix elements are composite weight objects, containing semantic similarity, causal correlation, confidence, and timestamp fields. The confidence field records the reliability of the weight calculation, determined based on the number of samples involved in the calculation and the number of expert verifications, ranging from 0 to 1. The timestamp field records the last update time of the weights, used for cache invalidation and incremental update control. The parsing process employs a parallel reading strategy, which divides the graph data into pieces and distributes them to multiple worker threads for processing. Each thread is responsible for processing a specific set of edges, and intermediate results are exchanged through a shared memory pool between threads.
[0109] A multi-level scoring matrix is constructed to create a comprehensive evaluation system integrating semantic similarity and the correlation between evidence and governance. The matrix adopts a three-dimensional tensor structure. The first dimension represents the source fragment index, the second dimension represents the target fragment index, and the third dimension contains four channels: comprehensive score, semantic score, evidence and governance score, and conflict identifier. The comprehensive score is calculated using a weighted average, with a semantic similarity weight of 0.45 and an evidence and governance correlation weight of 0.55. The weighting is determined based on regression analysis of 3000 expert-annotated samples. Score normalization uses the Min-Max standardization method to map the original scores to the interval between 0 and 1, avoiding the influence of different scales on the comprehensive score. The conflict identifier uses binary encoding: the first bit represents semantic conflict, the second bit represents evidence and governance contradiction, the third bit represents logical inconsistency, and the fourth bit represents temporal conflict. The matrix is stored in a sparse matrix format, storing only non-zero elements and their coordinates, compressing storage space and improving access efficiency. Matrix updates support both incremental and full modes. Incremental mode only updates changed matrix elements, while full mode recalculates all elements. The update strategy is automatically selected based on the range of changes and time constraints.
[0110] Semantic conflict and treatment-symptom contradiction identification are based on abnormal pattern detection in a multi-level scoring matrix. Semantic conflict identification employs a combination of threshold comparison and cluster analysis. Segment pairs with semantic scores below 0.3 are identified as potential conflicts. Cluster analysis, using the K-means algorithm, categorizes conflicting segments into three types: strong conflict, moderate conflict, and mild conflict. Strong conflict corresponds to segment pairs with semantic scores below 0.15, indicating complete semantic incompatibility; moderate conflict corresponds to segment pairs with scores between 0.15 and 0.25; and mild conflict corresponds to segment pairs with scores between 0.25 and 0.3. Treatment-symptom contradiction identification uses the sign and absolute value analysis of the correlation between treatment and symptoms. Segment pairs with negative correlation values and absolute values greater than 0.4 are identified as treatment-symptom contradictions. Negative values indicate conflicting treatment principles, and absolute values reflect the severity of the conflict. Contradiction types are categorized into three types: drug property conflict, treatment method conflict, and pathogenesis conflict. Drug property conflict refers to conflicting properties and effects of drugs involved in the segment; treatment method conflict refers to conflicting treatment methods; and pathogenesis conflict refers to inconsistent understanding of pathological mechanisms. Conflict detection also considers the topological position of fragments in the graph, with conflicts between adjacent nodes having a higher priority than conflicts between distant nodes.
[0111] Tracing the lineage of traditional Chinese medicine (TCM) knowledge graphs identifies the theoretical origins and evolutionary paths of conflicting fragments. The tracing algorithm employs a reverse breadth-first search, starting from the knowledge node corresponding to the conflicting fragment and backtracking upwards along the edges of the knowledge graph until the theoretical root is found or the maximum search depth is reached. The search depth is limited to eight levels, determined based on the hierarchical structure of TCM theory; tracing paths exceeding eight levels typically lack direct guiding significance. The tracing process records the weight and association strength of each node along the path. Weight reflects the node's importance within the theoretical system, while association strength reflects the logical coherence between nodes. Path scoring is calculated using the geometric mean of weight and association strength; the path with the highest score is considered the lineage of traditional Chinese medicine. The tracing results include four key pieces of information: the lineage path, confidence level, branch points, and convergence points. The lineage path records the complete node sequence from the root to the conflicting node; confidence level reflects the reliability of the tracing results; branch points identify the locations of theoretical divergences; and convergence points indicate the intersections of different theoretical paths.
[0112] The dialectical rule chain assessment evaluates the degree of conformity between conflicting fragments and theoretical rules in the knowledge graph. The assessment module maintains a rule base containing 2847 dialectical rules. Each rule uses a production-style representation and includes four fields: condition, conclusion, confidence level, and scope of application. The condition describes the combination pattern of symptoms, signs, and pathogenesis characteristics; the conclusion provides the corresponding diagnosis or treatment suggestion; the confidence level reflects the rule's reliability; and the scope of application defines the rule's application scenario. Rule matching uses a pattern matching engine, supporting three modes: exact matching, fuzzy matching, and partial matching. Exact matching requires the fragment content to be completely consistent with the rule conditions; fuzzy matching allows synonym substitution and word order adjustment; and partial matching accepts subset matching of conditions. The matching process generates a matching score: exact matching scores 1.0, fuzzy matching scores 0.7 to 0.9, and partial matching scores 0.3 to 0.6. Rule chain construction is achieved through forward reasoning, starting from initial symptoms and gradually applying applicable rules to deduce pathogenesis diagnosis and treatment plans, forming a complete reasoning chain. The validity of the chain is verified through loop detection and contradiction detection to ensure the logical consistency of the reasoning process.
[0113] The expression fit and treatment consistency metrics quantify the matching degree between the fragment and the nodes in the syndrome differentiation rule chain. Expression fit is calculated through a weighted fusion of text similarity and semantic similarity. Text similarity uses the normalized result of edit distance, and semantic similarity uses the cosine similarity of word vectors. The text similarity weight is 0.3, and the semantic similarity weight is 0.7, taking into account the richness of synonymous expressions in TCM texts. The fit calculation also considers keyword matching, and TCM terminology matching is given additional weight. The terminology weight is determined using the TF-IDF algorithm. The fit threshold is set at 0.65, derived from regression analysis of 500 sets of sample data evaluated by experts. Fragments below this threshold are classified as semantically deepened fragments. The syndrome differentiation and treatment consistency metrics assess the consistency between the treatment plan described in the fragment and the recommended plan in the rule chain. The consistency calculation considers three dimensions: consistency of treatment method, drug similarity, and dosage rationality. Consistency of treatment method is assessed through semantic matching of treatment principles, drug similarity is calculated through the overlap of drug efficacy and meridian tropism, and dosage rationality is judged by the degree of conformity of the dosage range. The baseline fit was set at 0.72, and segments that deviated from the baseline by more than 0.15 were classified as regular remodeling segments.
[0114] Semantic enhancement fragment processing selects highly resonant replacement fragments from a set of candidate medical record fragments. The selection process employs a multi-round screening strategy: the first round performs a coarse screening based on semantic similarity, retaining candidate fragments with a semantic similarity higher than 0.75 to the dialectical rule chain; the second round performs a fine screening based on contextual adaptability, evaluating the semantic coherence and logical consistency of candidate fragments with surrounding fragments; and the third round performs a final screening based on professional standardization, ensuring that the selected fragments conform to TCM expression habits and clinical norms. Resonance quantification calculates the comprehensive matching degree between candidate fragments and multiple nodes in the rule chain using a weighted summation method, with weights allocated according to the importance of the nodes in the rule chain. Semantic transformation employs a template-based replacement method, mapping key information from the original fragment to the corresponding position in the new fragment, maintaining the core semantic content while improving expression fit. The transformation process maintains semantic consistency checks to ensure that the transformed fragments do not generate new semantic conflicts. The multi-level scoring matrix update employs a local recalculation strategy, recalculating only the matrix rows and columns involved in the replacement fragments, avoiding the performance overhead of full recalculation.
[0115] The pattern reshaping fragment processing relies on core pathogenesis nodes in the dialectical rule chain to adjust the fragment hierarchy. Core pathogenesis nodes are identified through a comprehensive evaluation of three dimensions: node degree, centrality index, and expert weight. Node degree reflects the breadth of node connections, centrality index reflects the node's important position in the graph, and expert weight reflects the subjective assessment of the node's importance by clinical experts. The hierarchy adjustment algorithm reorganizes the logical hierarchy of fragments according to the guidance of core pathogenesis nodes, moving fragments that do not conform to the pattern to appropriate hierarchical positions. The adjustment process considers the dependencies between fragments to ensure that the adjusted hierarchical structure maintains logical integrity. Adjustment operations include three types: moving up, moving down, and reorganizing. Moving up elevates a fragment to a higher logical level, moving down demotes a fragment to a more specific level, and reorganizing rearranges the order of fragments within the same level. After adjustment, the fragments are re-placed into a multi-level scoring matrix for verification. The verification process uses incremental calculation, only calculating the matrix portion affected by the adjustment. Fragments that pass verification retain their adjusted positions, while fragments that fail verification are rolled back to their original positions or other adjustment schemes are attempted.
[0116] The medical record fragment integration process organizes the verified fragments according to a hierarchical sequence within a structured diagnostic framework. The integration process follows the logical order of traditional Chinese medicine's syndrome differentiation and treatment, starting with symptoms and signs, proceeding through syndrome diagnosis and pathogenesis analysis, and culminating in a complete chain of treatment plans. The hierarchical sequence is represented using a directed acyclic graph (DAG), with the topological sorting within the graph determining the integration order. The integration algorithm employs a greedy strategy, prioritizing the combination of fragments with the highest scores and no conflicts. When multiple equally divided candidates exist, the combination with stronger clinical applicability is selected. Fragment connection utilizes natural language generation technology, inserting appropriate conjunctions and transitional sentences between adjacent fragments to ensure the fluency and readability of the integrated text. Conjunction selection is based on the logical relationship between fragments: causal relationships use words like "due to" and "therefore," progressive relationships use words like "moreover" and "furthermore," and adversative relationships use words like "however" and "yet."
[0117] Semantic coherence assessment verifies the logical consistency and expressive consistency of the integrated medical record text. The assessment employs a multi-level verification mechanism: local coherence verification checks the semantic connections between adjacent sentences, while global coherence verification checks the overall thematic consistency and logical integrity of the text. Local coherence is measured by inter-sentence semantic similarity and the accuracy of referential resolution, with a similarity threshold set at 0.4 and a referential resolution accuracy requirement of over 85%. Global coherence is assessed through topic modeling and logical consistency analysis. The topic model uses the LDA algorithm to extract the text's topic distribution; a topic concentration higher than 0.6 is considered to indicate good topic consistency. Logical consistency analysis checks for self-contradictory statements in the text, using a knowledge graph-based reasoning verification method. The assessment results generate a coherence score ranging from 0 to 1; texts with scores higher than 0.8 are considered to have good semantic coherence.
[0118] In one optional implementation, identifying semantic conflicts and diagnostic-treatment contradictions between medical record fragments based on the multi-level scoring matrix, and tracing the inheritance of the semantically conflicting and diagnostic-treatment contradictory fragments in a multi-level TCM knowledge graph, includes:
[0119] Scoring elements for medical record fragment pairs are extracted from a multi-level scoring matrix. A two-level evaluation criterion is constructed based on semantic association threshold and syndrome-treatment correspondence threshold. By jointly judging the semantic expression similarity and syndrome-treatment rule fit of medical record fragment pairs in the multi-level scoring matrix, medical record fragments with semantic conflicts and syndrome-treatment contradictions are identified.
[0120] For the medical record fragments with semantic conflicts and contradictory treatments, a transmission tracing network is constructed in a multi-level TCM knowledge graph. The transmission tracing network maps the syndrome hierarchy relationship and treatment principle transformation process based on the chain of diagnostic rules between nodes, and determines the transmission path of the semantic conflicts and contradictory treatments at different levels.
[0121] The scoring element extraction process obtains quantitative feature data of medical record fragment pairs from a multi-level scoring matrix. The extraction module iterates through the non-zero elements of the matrix, and each element contains four fields: semantic expression similarity, syndrome-treatment pattern fit, confidence level, and timestamp. Semantic expression similarity is stored as a 32-bit floating-point number, ranging from 0 to 1, with precision maintained to four decimal places. The syndrome-treatment pattern fit also uses a 32-bit floating-point number, ranging from -1 to +1, where a positive value indicates consistency in syndrome-treatment logic, and a negative value indicates conflict. The extraction process employs a row-first scanning strategy, maintaining a memory buffer of 1024 matrix elements. Scoring elements are processed in batches when the buffer is full.
[0122] A two-tiered evaluation criterion is constructed using a semantic relevance threshold and a thesis-theory correspondence threshold. The semantic relevance threshold is set at 0.35, indicating semantic differences between fragment pairs below this threshold. The thesis-theory correspondence threshold is set at -0.25, indicating thesis-theory contradictions between fragment pairs below this threshold. The two-tiered evaluation criterion uses a logical AND operation to determine the results, confirming a significant conflict only when the semantic relevance is below 0.35 and the thesis-theory rule fit is below -0.25. A buffer mechanism is introduced into the evaluation criterion: semantic relevance between 0.35 and 0.45 is marked as semantically ambiguous, and the thesis-theory rule fit between -0.25 and -0.15 is marked as the thesis-theory questionable.
[0123] The system classifies conflict states by jointly determining semantic similarity and the fit between the principles of evidence and governance. The joint determination module outputs a conflict type identifier and a severity rating. Conflict types include four categories: semantic conflict, evidence-government contradiction, mixed conflict, and normal state. The determination logic adopts a decision tree structure. The root node determines whether the semantic relevance is below a threshold, and the left and right subtrees handle conflict and normal branches respectively. Severity is divided into three levels: mild, moderate, and severe. Conflicts with a semantic relevance below 0.2 or an evidence-government fit below -0.4 are rated as severe.
[0124] The inheritance tracing network is constructed within a multi-level TCM knowledge graph to establish theoretical origin paths for conflicting fragments. The network construction module starts with the knowledge nodes corresponding to the conflicting fragments and explores node relationships through breadth-first search, with a maximum traversal depth of 6 levels. The traversal process maintains a set of visited nodes and a path record table. The relationships between nodes include four types: theoretical inheritance, methodological evolution, syndrome correlation, and treatment transformation. The inheritance tracing network uses a multi-level graph structure for storage, recording direct relationships, indirect relationships, and inheritance root relationships respectively.
[0125] The dialectical rule chain mapping establishes a correspondence mechanism between the hierarchical relationship of syndromes and the transformation process of treatment principles. The mapping module maintains a rule base of 3200 dialectical rules, each rule adopting a triplet form of premise, reasoning process, and conclusion output. The mapping process uses pattern matching technology, supporting three modes: exact matching, fuzzy matching, and partial matching. The hierarchical relationship of syndromes is represented by a tree structure, divided into six levels from the etiology layer to the prescription layer. The transformation process of treatment principles is described using a state transition diagram, where state nodes represent treatment stages and transition edges represent treatment adjustment paths.
[0126] In one optional implementation, the complete medical record text is structurally split according to symptom characteristics, syndrome differentiation elements, and treatment plans, and a corresponding relationship is established with actual clinical efficacy data. The diagnosis and treatment paths and syndrome differentiation rules in the multi-level TCM knowledge graph are optimized based on an association rule mining algorithm, including:
[0127] The complete medical record text is structured and broken down according to symptom characteristics, syndrome differentiation elements and treatment plan, and a symptom-syndrome differentiation-treatment correlation matrix is constructed. The actual clinical efficacy data is written into the corresponding position of the symptom-syndrome differentiation-treatment correlation matrix.
[0128] The symptom-diagnosis-treatment association matrix is analyzed using an association rule mining algorithm. The support and confidence of each node combination are calculated. Node combinations whose support and confidence exceed the optimization threshold are marked as items to be optimized. The items to be optimized include effectiveness score and reliability score.
[0129] Based on the analysis results of the item to be optimized by the association rule mining algorithm, the diagnosis and treatment path and syndrome differentiation and treatment rules corresponding to the item to be optimized are located in the multi-level TCM knowledge graph. The effectiveness score and the reliability score are written into the diagnosis and treatment path and the syndrome differentiation and treatment rules respectively, thus completing the clinical verification and optimization update of the knowledge graph.
[0130] The complete medical record text is structurally segmented using rule-based text segmentation technology, parsing the content according to three dimensions: symptom characteristics, syndrome differentiation elements, and treatment plans. The segmentation module maintains a TCM terminology dictionary and a classification tag library. The terminology dictionary contains 3200 symptom terms, 1800 syndrome differentiation terms, and 2400 treatment terms, each associated with a semantic tag and classification identifier. Symptom characteristic segmentation identifies text fragments describing clinical manifestations in the medical record, including chief complaint, present illness history, and the four diagnostic methods (inspection, auscultation, olfaction, palpation, and olfaction). Named entity recognition algorithms are used to extract symptom names, severity descriptions, location information, and time features. Syndrome differentiation element segmentation extracts content reflecting pathogenesis analysis and syndrome judgment, including etiological analysis, lesion location attribution, disease nature judgment, and syndrome diagnosis. Dependency parsing is used to identify the logical relationships in syndrome differentiation. Treatment plan segmentation parses treatment-related information such as treatment method selection, prescription compatibility, and dosage adjustment, using medical knowledge graph matching technology to ensure extraction accuracy. The segmentation process employs a multi-round iterative strategy, processing one dimension of content in each round, and using contextual association analysis between rounds to avoid information duplication or omission. The splitting results are stored as structured data objects. The object fields include fragment identifier, content text, category label, confidence level, and association relationship. The fragment identifier is encoded using a 64-bit integer, the confidence level ranges from 0 to 1, and the association relationship records the logical dependencies between fragments.
[0131] A symptom-diagnosis-treatment association matrix is constructed using a three-dimensional tensor data structure. The first dimension represents the symptom feature index, the second dimension represents the diagnosis element index, and the third dimension represents the treatment plan index. The matrix dimensions are dynamically determined based on the splitting results, with a typical size of 500 symptom dimensions, 300 diagnosis dimensions, and 400 treatment dimensions, corresponding to a tensor size of 500×300×400. Matrix elements store association strength values, represented as 32-bit floating-point numbers, ranging from 0 to 1. Higher values indicate stronger association between the three elements. Association strength is calculated based on a weighted fusion of three factors: co-occurrence frequency, expert evaluation, and literature evidence. Co-occurrence frequency has a weight of 0.4, expert evaluation has a weight of 0.35, and literature evidence has a weight of 0.25. The matrix uses a sparse storage format, storing only non-zero elements and their coordinate indices, compressing storage space and improving access efficiency. The matrix supports both incremental and batch update modes. Incremental updates are suitable for real-time processing of single medical records, while batch updates are suitable for offline analysis of large-scale medical record data. The matrix construction process uses hash indexing technology to create a fast retrieval table for symptom, syndrome differentiation, and treatment elements, with a retrieval time complexity of constant level.
[0132] The clinical efficacy data writing process maps efficacy evaluation results to corresponding positions in the correlation matrix. The efficacy data package contains four indicators: efficiency, cure rate, adverse reaction rate, and recurrence rate. Each indicator is stored using 16-bit fixed-point numbers with precision maintained to two decimal places. Efficacy data sources include electronic medical record follow-up records, patient satisfaction surveys, and clinical trial reports. Data validity is ensured through a multi-source validation mechanism. The writing process employs a coordinate positioning algorithm to determine matrix coordinates based on the symptom-syndrome-treatment combination corresponding to the efficacy data, and then writes the efficacy indicators into the corresponding matrix elements. The writing operation supports overwrite mode and cumulative mode. Overwrite mode directly replaces the original values, while cumulative mode uses a weighted average to update values, with new data having a weight of 0.3 and historical data a weight of 0.7. An operation log is maintained during the data writing process, recording the writing time, data source, operation type, and preceding and following values for data traceability and error rollback. After writing is complete, a consistency check is triggered to verify the numerical range and logical relationships of matrix elements. Abnormal data is automatically marked, and processing suggestions are generated.
[0133] The association rule mining algorithm employs an improved Apriori algorithm to analyze the symptom-syndrome-treatment association matrix. The algorithm implementation consists of two stages: frequent itemset generation and rule extraction. Frequent itemset generation starts with a single item and expands layer by layer. Rule extraction calculates support and confidence based on frequent itemsets. The minimum support threshold is set to 0.15, and the minimum confidence threshold is set to 0.6. These thresholds are determined based on clinical expert experience and statistical analysis results. The algorithm uses a vertical data format to optimize storage and computational efficiency, converting the association matrix into a list of transactions. Each transaction records a symptom-syndrome-treatment combination and its association strength. Frequent itemset generation uses bit vector operations to accelerate set intersection operations; the bit vector length equals the total number of transactions, resulting in linear operation complexity. The rule extraction process generates association rules in the form of "symptom A and syndrome B lead to treatment C," typically ranging from 1000 to 5000 rules. The algorithm supports distributed computing, partitioning the association matrix and distributing it to multiple computing nodes for parallel processing. Nodes synchronize intermediate results through message passing.
[0134] Support and confidence scores quantify the statistical significance and predictive accuracy of association rules. Support calculates the frequency of a specific itemset across all transactions, calculated as the number of transactions containing that itemset divided by the total number of transactions, with a value ranging from 0 to 1. Confidence scores calculate conditional probability, representing the probability that the conclusion holds given the given premises. This is calculated by dividing the number of transactions containing both premises and conclusions by the number of transactions containing only premises. Floating-point arithmetic is used in the calculations, maintaining precision to four decimal places to avoid precision loss affecting subsequent analysis. Support and confidence calculations support a caching mechanism, temporarily storing the results in an in-memory hash table; duplicate calculations of the same itemset directly return the cached result. The calculation module also provides a lift metric, measuring the improvement of a rule relative to random cases; a lift greater than 1 indicates a positive correlation. The calculation results are sorted according to the combined score of support and confidence scores, calculated using a geometric mean to emphasize the balanced importance of the two metrics.
[0135] A rule selection and quality assessment mechanism was established to optimize threshold settings and the labeling of items to be optimized. The support optimization threshold was set at 0.25, and the confidence optimization threshold was set at 0.75. Node combinations that simultaneously meet both conditions were labeled as items to be optimized. Threshold settings considered both clinical applicability and rule reliability; thresholds that were too low resulted in too many rules that were difficult to apply, while thresholds that were too high missed valuable rules. The labeling process generated a candidate rule list, which was sorted in descending order of comprehensive score. The top 500 rules with the highest scores were selected as items to be optimized. The labeling process also considered the novelty and clinical significance of the rules; rules that conflicted with the existing knowledge graph received extra attention, and innovative rules were prioritized for labeling. The labeling results recorded the rule content, statistical indicators, labeling time, and evaluation status, supporting both manual review and automatic verification modes.
[0136] The efficacy and reliability scores quantify the clinical value and strength of evidence for the items to be optimized. The efficacy score is based on statistical analysis of treatment effects, comprehensively considering three dimensions: improved cure rate, degree of symptom relief, and shortened treatment time. The score ranges from 0 to 100, with higher scores indicating better clinical efficacy. The reliability score assesses the quality of evidence and applicability of the rules, considering three factors: sample size, study design quality, and consistency of results. The score also uses a 0-100 percentage system. The scoring method combines expert evaluation and statistical calculation, with expert evaluation having a weight of 0.4 and statistical calculation a weight of 0.6. The statistical calculation for the efficacy score is based on the comparison of the mean of efficacy data and significance tests, while the statistical calculation for the reliability score is based on sample distribution and confidence interval analysis. Scoring results retain the integer part to avoid misleading results from over-precision. The scoring process supports both batch calculation and real-time calculation modes. Batch calculation is suitable for large-scale rule evaluation, while real-time calculation is suitable for rapid evaluation of individual rules.
[0137] In a multi-level TCM knowledge graph, the location of diagnostic and treatment paths and syndrome differentiation rules employs a graph matching algorithm to identify the knowledge nodes corresponding to the items to be optimized. The localization algorithm is based on a comprehensive matching of semantic and structural similarity. Semantic similarity is calculated using the cosine distance of word vectors, while structural similarity is analyzed through graph structure isomorphism. An approximate matching strategy is used in the matching process, allowing for some differences to accommodate the diversity of knowledge representation. The similarity threshold is set to 0.8. The localization result corresponds to multiple candidate nodes, which are sorted according to their matching degree. The node with the highest matching degree is selected as the target location. A matching record table is maintained during the localization process, recording the correspondence between the items to be optimized and the knowledge nodes, the matching degree, and the localization time, for subsequent update tracking. The localization algorithm supports both fuzzy matching and exact matching modes. Fuzzy matching is suitable for newly discovered rules, while exact matching is suitable for updating known rules.
[0138] The effectiveness and reliability scores are written to update the corresponding diagnostic and treatment pathways and rules. The write operation uses a field update method, adding or modifying the effectiveness and reliability score fields in the target node. A version control mechanism is maintained during the write process, retaining historical score data and recording update times to support analysis of score change trends. The write operation triggers a local recalculation of the graph structure, updating the weights and connection strengths of relevant nodes to ensure the overall consistency of the knowledge graph. After the write is completed, a verification check is performed to ensure that the newly written score data conforms to the numerical range and logical constraints; abnormal data triggers a rollback operation. Knowledge graph updates support two strategies: incremental updates and full updates. Incremental updates only process changed nodes and edges, while full updates reconstruct the entire graph structure.
[0139] A second aspect of the present invention provides a human-computer collaborative intelligent writing system for medical documents, comprising:
[0140] Unit 1 is used to acquire patients' multimodal clinical data, perform cross-modal semantic alignment on the multimodal clinical data, and obtain a fused semantic vector by extracting the semantic features of the TCM domain from each modality of data and mapping them to a unified semantic representation space.
[0141] Unit 2 is used to perform reasoning and expansion on the fused semantic vector based on a pre-constructed multi-level TCM knowledge graph, and to filter combinations of diagnostic and treatment knowledge nodes that conform to the dialectical logic through a path constraint propagation mechanism to generate a structured dialectical framework.
[0142] Unit 3 is used to generate a set of candidate medical record fragments based on the semantic association strength and clinical applicability constraints of each knowledge node in the structured dialectical framework, and to construct a logical consistency verification graph of the set of candidate medical record fragments; by detecting semantic conflicts and syndrome-treatment contradictions between medical record fragments in the logical consistency verification graph, and using the syndrome differentiation rules in the multi-level TCM knowledge graph to replace or merge conflicting fragments, a logically self-consistent complete medical record text is generated.
[0143] Unit 4 is used to structurally split the complete medical record text according to symptom characteristics, syndrome differentiation elements and treatment plans, and establish a corresponding relationship with actual clinical efficacy data. Based on the association rule mining algorithm, it optimizes the diagnosis and treatment path and syndrome differentiation rules in the multi-level TCM knowledge graph to achieve dynamic optimization of the medical record generation model.
[0144] A third aspect of the present invention provides an electronic device, comprising:
[0145] processor;
[0146] Memory used to store processor-executable instructions;
[0147] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0148] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0149] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A human-computer collaborative intelligent writing method for medical documents, characterized in that: include: Acquire multimodal clinical data of patients, perform cross-modal semantic alignment on the multimodal clinical data, and obtain a fused semantic vector by extracting TCM domain semantic features from each modality of data and mapping them to a unified semantic representation space. Based on a pre-constructed multi-level TCM knowledge graph, the fused semantic vector is extended through reasoning, and a path constraint propagation mechanism is used to select combinations of diagnostic and treatment knowledge nodes that conform to the dialectical logic, thereby generating a structured dialectical framework. Based on the semantic association strength and clinical applicability constraints of each knowledge node in the structured dialectical framework, a set of candidate medical record fragments is generated, and a logical consistency verification graph of the candidate medical record fragment set is constructed. Semantic conflicts and contradictions in diagnosis and treatment are detected between medical record fragments in the logical consistency verification graph, and the dialectical rules in the multi-level TCM knowledge graph are used to replace or merge conflicting fragments to generate a logically consistent complete medical record text, including: The edge weight features are analyzed in the logical consistency verification graph to construct a multi-level scoring matrix that integrates semantic similarity and syndrome-treatment correlation. Based on the multi-level scoring matrix, semantic conflicts and syndrome-treatment contradictions between medical record fragments are identified, and the inheritance of the fragments with semantic conflicts and syndrome-treatment contradictions in the multi-level TCM knowledge graph is traced. Based on the syndrome differentiation rule chain in the multi-level TCM knowledge graph, the segments with semantic conflicts and syndrome-treatment contradictions are determined. The expression fit and syndrome-treatment consistency of the segments with the nodes in the syndrome differentiation rule chain are quantified by the multi-level scoring matrix. When the expression fit is lower than the fit threshold, it is classified as a semantic deepening segment. When the syndrome-treatment consistency deviates from the benchmark, it is classified as a rule reshaping segment. For the semantically deepened segments, segments that highly resonate with the dialectical rule chain are selected from the candidate medical record segment set for semantic transformation, and the scoring elements are updated in the multi-level scoring matrix; for the pattern reshaping segments, the segment level is adjusted based on the core pathogenesis nodes in the dialectical rule chain, and the adjusted segments are re-placed into the multi-level scoring matrix for verification. The validated medical record fragments are integrated according to the hierarchical sequence of the structured dialectical framework and semantic coherence is evaluated to generate a logically consistent complete medical record text. The complete medical record text is structurally broken down according to symptom characteristics, syndrome differentiation elements, and treatment plans, and a corresponding relationship is established with actual clinical efficacy data. Based on the association rule mining algorithm, the diagnosis and treatment paths and syndrome differentiation rules in the multi-level TCM knowledge graph are optimized to achieve dynamic optimization of the medical record generation model.
2. The method according to claim 1, characterized in that, Based on a pre-constructed multi-level TCM knowledge graph, the fused semantic vector is extended through reasoning, and a path constraint propagation mechanism is used to select combinations of diagnostic and treatment knowledge nodes that conform to the dialectical logic, generating a structured dialectical framework including: In a multi-level TCM knowledge graph, an initial set of syndrome nodes with a correlation degree exceeding a preset correlation threshold is located using a dynamic semantic matching strategy. Starting from the initial set of syndrome nodes, a multi-hop expansion is performed along the inter-level reasoning edges in the multi-level TCM knowledge graph. During the expansion process, the expansion path is pruned in real time according to the rules of dialectical logic constraints to generate a set of candidate reasoning paths. For each reasoning path in the candidate reasoning path set, the semantic consistency score between adjacent nodes and the evidence-theory correspondence score between cross-level nodes are calculated. The semantic consistency score and the evidence-theory correspondence score are bidirectionally propagated and cumulatively updated in the path through the path constraint propagation mechanism to obtain the comprehensive credibility score of each reasoning path. The candidate reasoning path set is hierarchically screened based on the comprehensive credibility score. The syndrome nodes, pathogenesis nodes, and treatment principle nodes in the effective reasoning paths that meet the screening conditions are extracted. The syndrome nodes, pathogenesis nodes, and treatment principle nodes are then organized in a structured manner according to the hierarchical order of TCM syndrome differentiation logic. Causal relationships and semantic dependencies between nodes are established to generate a structured dialectical framework.
3. The method according to claim 2, characterized in that, Starting from the initial set of syndrome nodes, a multi-hop expansion is performed along the inter-level reasoning edges in the multi-level TCM knowledge graph. During the expansion process, the expansion path is pruned in real time according to the rules of dialectical logic constraints, generating a set of candidate reasoning paths, including: The syndrome type identifier of each node is parsed from the initial syndrome node set, and the lower-level pathogenesis node is located in the multi-level TCM knowledge graph based on the syndrome type identifier. For the lower-level pathogenesis nodes, an extended state vector containing the trajectory of syndrome evolution is constructed. The extended state vector is input into the syndrome differentiation logic constraint rules to judge the rationality of syndrome mechanism transmission. When the judgment shows that there is a causal contradiction between syndrome and pathogenesis or that the transformation of disease nature violates the laws of traditional Chinese medicine theory, the extended path of the corresponding pathogenesis node is intelligently pruned. For the pathogenesis nodes identified through the rationality analysis of the mechanism of disease progression, a progressive analysis is carried out along the inter-level reasoning edge towards the treatment principle level and the prescription level. At each cross-level extension, the extended state vector is enriched to integrate the treatment method adaptation characteristics and prescription meridian tropism elements of the newly traversed nodes. The updated extended state vector is evaluated for treatment principle fit and prescription compatibility is verified through the dialectical logic constraint rules. When it is found that the treatment method does not match the pathogenesis or that the prescription compatibility is contradictory or antagonistic, the path extension is stopped. During the multi-hop expansion process, a path dynamic score is calculated for each retained path. The expansion priority sequence is adjusted based on the path dynamic score, and paths with scores exceeding the retention threshold are included in the candidate inference path set.
4. The method according to claim 1, characterized in that, Based on the semantic association strength and clinical applicability constraints of each knowledge node in the structured dialectical framework, a set of candidate medical record fragments is generated, and a logical consistency verification graph of the candidate medical record fragment set is constructed, including: Syndrome nodes, pathogenesis nodes, and treatment principle nodes are extracted from the structured dialectical framework, and the semantic association and hierarchical inheritance relationship between any two knowledge nodes are determined through multi-dimensional semantic feature analysis. Based on the core semantic features of each knowledge node, candidate expression templates with semantic expression fit higher than a preset fit threshold are identified in the pre-stored medical record text template library. Clinical applicability assessment criteria are constructed by combining the patient's age characteristics, physical constitution classification and disease evolution stage. The candidate expression templates are then screened and optimized and recombined at multiple levels to generate candidate medical record fragments corresponding to each knowledge node. The knowledge node pairs with direct reasoning relationships and their corresponding candidate medical record fragments in the structured dialectical framework are organized into a fragment pair set. The semantic consistency representation and the compatibility characteristics of the diagnosis and treatment logic between the two medical record fragments in each fragment pair are determined through semantic content association analysis and diagnosis and treatment logic deduction. Using each medical record fragment in the candidate medical record fragment set as a base node, and taking the semantic consistency representation and the diagnostic and treatment logic compatibility feature as the link, a logical consistency verification graph with a hierarchical reasoning structure is constructed.
5. The method according to claim 1, characterized in that, Based on the multi-level scoring matrix, semantic conflicts and contradictions in diagnosis and treatment between medical record fragments are identified. Tracing the inheritance of these semantically conflicting and contradictory fragments within a multi-level TCM knowledge graph includes: Scoring elements for medical record fragment pairs are extracted from a multi-level scoring matrix. A two-level evaluation criterion is constructed based on semantic association threshold and syndrome-treatment correspondence threshold. By jointly judging the semantic expression similarity and syndrome-treatment rule fit of medical record fragment pairs in the multi-level scoring matrix, medical record fragments with semantic conflicts and syndrome-treatment contradictions are identified. For the medical record fragments with semantic conflicts and contradictory treatments, a transmission tracing network is constructed in a multi-level TCM knowledge graph. The transmission tracing network maps the syndrome hierarchy relationship and treatment principle transformation process based on the syndrome differentiation rule chain between nodes, and determines the transmission path of the semantic conflicts and contradictory treatments at different levels.
6. The method according to claim 1, characterized in that, The complete medical record text is structurally broken down according to symptom characteristics, syndrome differentiation elements, and treatment plans, and a corresponding relationship is established with actual clinical efficacy data. The diagnosis and treatment paths and syndrome differentiation rules in the multi-level TCM knowledge graph are optimized based on association rule mining algorithms, including: The complete medical record text is structured and broken down according to symptom characteristics, syndrome differentiation elements and treatment plan, and a symptom-syndrome differentiation-treatment correlation matrix is constructed. The actual clinical efficacy data is written into the corresponding position of the symptom-syndrome differentiation-treatment correlation matrix. The symptom-diagnosis-treatment association matrix is analyzed using an association rule mining algorithm. The support and confidence of each node combination are calculated. Node combinations whose support and confidence exceed the optimization threshold are marked as items to be optimized. The items to be optimized include effectiveness score and reliability score. Based on the analysis results of the item to be optimized by the association rule mining algorithm, the diagnosis and treatment path and syndrome differentiation and treatment rules corresponding to the item to be optimized are located in the multi-level TCM knowledge graph. The effectiveness score and the reliability score are written into the diagnosis and treatment path and the syndrome differentiation and treatment rules respectively, thus completing the clinical verification and optimization update of the knowledge graph.
7. A human-computer collaborative intelligent writing system for medical documents, used to implement the method of any one of claims 1-6, characterized in that, include: Unit 1 is used to acquire patients' multimodal clinical data, perform cross-modal semantic alignment on the multimodal clinical data, and obtain a fused semantic vector by extracting the semantic features of the TCM domain from each modality of data and mapping them to a unified semantic representation space. Unit 2 is used to perform reasoning and expansion on the fused semantic vector based on a pre-constructed multi-level TCM knowledge graph, and to filter combinations of diagnostic and treatment knowledge nodes that conform to the dialectical logic through a path constraint propagation mechanism to generate a structured dialectical framework. Unit 3 is used to generate a set of candidate medical record fragments based on the semantic association strength and clinical applicability constraints of each knowledge node in the structured dialectical framework, and to construct a logical consistency verification graph of the set of candidate medical record fragments; by detecting semantic conflicts and syndrome-treatment contradictions between medical record fragments in the logical consistency verification graph, and using the syndrome differentiation rules in the multi-level TCM knowledge graph to replace or merge conflicting fragments, a logically self-consistent complete medical record text is generated. Unit 4 is used to structurally split the complete medical record text according to symptom characteristics, syndrome differentiation elements and treatment plans, and establish a corresponding relationship with actual clinical efficacy data. Based on the association rule mining algorithm, it optimizes the diagnosis and treatment path and syndrome differentiation rules in the multi-level TCM knowledge graph to achieve dynamic optimization of the medical record generation model.
8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Traditional Chinese medicine medical record generation method, device and equipment and storage medium
CN117594172A
Automatic medical record writing system and method based on multi-modal input and multi-agent driving
CN120708790A