A knowledge graph construction method and system based on traditional Chinese medicine classics
Through optical character recognition, multimodal feature extraction and causal reasoning engine, a dynamic dialectical reasoning model is built, which solves the problem of different term expressions and fragmentation of knowledge associations in traditional Chinese medicine classics, realizes real-time iteration of the traditional Chinese medicine knowledge graph and individualized diagnosis and treatment support, and promotes the modernization of traditional Chinese medicine.
Patent Information
- Application Number
- CN202510608712.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-05-13
AI Technical Summary
The existing technology is difficult to effectively solve the problems of different term expressions in traditional Chinese medicine classics, fragmented knowledge associations, limited semantic understanding, difficulty in integrating multi-source data, and lack of dynamic update of knowledge graphs, resulting in limitations in the modernization and intelligent application of traditional Chinese medicine.
Standardized corpus is generated through optical character recognition and multimodal feature extraction, and entity relationships are extracted based on language models and causal reasoning engines, and dynamic dialectical reasoning model is constructed, combining edge computing and privacy protection technology to realize real-time iteration of knowledge graphs and clinical efficacy feedback optimization.
It has realized the cross-age alignment and deep semantic integration of traditional Chinese medicine classics, supports millisecond deduction and multi-dimensional correction of individualized diagnosis and treatment plans, forms a dynamic knowledge network that runs through ancient and modern times, and supports the modern intelligent infrastructure of traditional Chinese medicine.
Smart Images

Figure CN120124733B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge graph construction, and in particular to a method and system for constructing a knowledge graph based on traditional Chinese medicine classics. Background Art
[0002] Traditional Chinese medicine (TCM) texts, including classic works such as the Yellow Emperor's Classic of Internal Medicine and Treatise on Febrile Diseases, serve as the theoretical foundation for a millennium-long tradition. These texts document systematic knowledge on the internal organs and meridians, etiology and pathogenesis, treatment principles and methods, and prescriptions. However, much of this knowledge exists in unstructured text, and suffers from issues such as terminology discrepancies (e.g., the semantic overlap between "liver qi stagnation" and "liver qi stagnation") and fragmented knowledge (e.g., the correspondence between prescriptions and symptoms is scattered across different texts).
[0003] With the penetration of artificial intelligence technology in the medical field, the demand for structured knowledge in scenarios such as TCM clinical decision support, personalized diagnosis and treatment, and new drug development is growing. The traditional model that relies on manual organization and experience inheritance can no longer meet the needs of efficient and precise applications. Therefore, the construction of a TCM knowledge graph is the core infrastructure for the modernization and intelligentization of TCM. It aims to achieve the digital reconstruction, semantic association, and intelligent application of ancient text knowledge through technical means, becoming a key breakthrough in promoting the development of TCM.
[0004] Some technologies currently attempt to address this problem. For example, entity recognition methods based on rule matching and dictionaries (such as the Traditional Chinese Medicine Language System (TCMLS), which matches entities using a preset dictionary) and relation extraction techniques combined with LSTM neural networks have achieved automated extraction of entities and relations to a certain extent, but they still have significant limitations.
[0005] First, insufficient terminology standardization makes knowledge integration difficult. For example, different texts lack a unified mapping mechanism for naming the same syndrome (such as "wind-heat cold" and "febrile disease"), requiring manual intervention for semantic alignment. Second, existing methods have limited depth of semantic understanding of ancient texts. For example, the LSTM model has a weak ability to capture long-range dependencies and complex semantic relationships (such as the multi-dimensional dialectical logic of "mixed deficiency and excess syndrome"), and does not fully consider the dynamic reasoning characteristics of the TCM theoretical system (such as the nonlinear transmission of the mutual generation and restraint relationship).
[0006] Furthermore, multi-source data integration technology is not yet mature. The heterogeneity of ancient texts, modern medical records, and pharmacopoeia databases makes it difficult to automatically resolve knowledge redundancies and conflicts (such as the contradiction between the ancient "Eighteen Antidotes" theory and modern pharmacological conclusions). More importantly, existing knowledge graphs are mostly limited to static storage, lacking dynamic update mechanisms (such as real-time iteration of diagnosis and treatment plans for emerging infectious diseases) and clinical feedback loops (such as the reverse correction of knowledge weights using patient efficacy data).
[0007] Therefore, there is an urgent need to design a knowledge graph construction method and system based on traditional Chinese medicine classics to solve the technical problems mentioned above. Summary of the Invention
[0008] Based on this, it is necessary to provide a knowledge graph construction method and system based on traditional Chinese medicine classics to address the above technical issues.
[0009] In a first aspect, the present invention provides a method for constructing a knowledge graph based on traditional Chinese medicine classics, comprising:
[0010] S1. Collect multi-source heterogeneous data of traditional Chinese medicine classics and perform standardization processing. Through optical character recognition and multimodal feature extraction, generate a standardized corpus aligned across time periods.
[0011] S2, based on the language model and causal reasoning engine, extracts entity relationships and dialectical rules from the standardized corpus to generate a dynamic dialectical reasoning model that supports nonlinear reasoning;
[0012] S3. Obtain the output of the dynamic dialectical reasoning model and build a redundant and dynamically calibrated TCM knowledge base by integrating and resolving conflicts among various TCM classics.
[0013] S4. Based on edge computing and privacy protection technologies, real-time iteration of knowledge graphs and optimization of clinical efficacy feedback are achieved.
[0014] Furthermore, we collected multi-source heterogeneous data from traditional Chinese medicine classics and standardized it. Through optical character recognition and multimodal feature extraction, we generated a standardized corpus aligned across time periods, including:
[0015] S11. Extract multi-source heterogeneous data from the ancient book digitization platform and hospital management information system, and perform grayscale correction, ink enhancement, and non-text area filtering according to data type. The multi-source heterogeneous data includes ancient book images, modern medical records, and pharmacopoeia database information.
[0016] S12. Deploy an optical character recognition engine, integrate the classical Chinese character library and domain dictionary, identify homophones and variant forms of characters, create a dynamic mapping chart, and output standardized text;
[0017] S13, extracting visual features of image data and time-frequency features of waveform data from the multi-source heterogeneous data, respectively, to obtain feature vectors of each component;
[0018] S14, automatically aligning terms in standardized text based on the cosine similarity algorithm;
[0019] S15. Use graph data to store the multi-dimensional relationship between classics content, physical sign data, and treatment methods and prescriptions, set up a spatiotemporal index structure, and realize distributed storage of processed and standardized text data.
[0020] Furthermore, based on the language model and causal reasoning engine, entity relationships and dialectical rules are extracted from the standardized corpus to generate a dynamic dialectical reasoning model that supports nonlinear reasoning, including:
[0021] S21. Generate a text vector sequence using a pre-trained language model, simulate the ambiguous context of homographs, calculate the semantic deviation between the context vector and the current question bank, adjust the entity type confidence weight, and generate an entity attribute matrix with a unified identifier.
[0022] S22. Based on the graph causal diffusion model, the generation-restraint relationship is encoded as a priori conditions of the graph network to construct the initial graph structure;
[0023] S23. Establish a reward function based on the generation-restraint relationship, use the proximal strategy optimization algorithm to train the dialectical strategy in a virtual environment, and generate a decision tree for the generation-restraint-organ-treatment chain reasoning;
[0024] S24. Based on the decision tree generation scheme, a dynamic dialectical reasoning model is constructed to solve the optimal solution and output a treatment path that includes suggestions for mutual generation and mutual restraint correction.
[0025] Furthermore, based on the graph causal diffusion model, the mutual generation and mutual restraint relationship is encoded as a priori conditions of the graph network, and the initial graph structure is constructed including:
[0026] S221. Convert the generation and restraint theory into graph structure constraints, setting the generation relationship between generation and restraint elements as positive weight connection and the restraint relationship as negative weight inhibition, and convert the discrete generation and restraint relationships into a continuous and differentiable probability distribution form to form a generation and restraint constraint matrix;
[0027] S222, dynamically weighted fusion of the generative constraint matrix and the statistical matrix extracted from the standardized corpus, and mapping the generative constraint relationship to the entity topology structure to generate an initial graph structure;
[0028] S223, introducing a Markov chain diffusion process into the initial graph structure, gradually injecting Gaussian noise according to preset noise scheduling parameters, and applying noise suppression to edge connections that conform to the Shengke theory;
[0029] S224. Set the theoretical conformity index to continuously monitor the theoretical deviation of the generated graph structure. When the theoretical deviation is greater than the preset threshold, the weight balancing mechanism is triggered to automatically adjust the fusion coefficient.
[0030] Furthermore, a reward function is established based on the generation-control relationship, and a proximal strategy optimization algorithm is used to train the dialectical strategy in a virtual environment. The output decision tree of the generation-control-viscera-treatment chain reasoning includes:
[0031] S231. Provide positive incentives for decisions that conform to the mutual growth path, impose penalty weights on operations that violate the mutual restraint principle, and use historical medical case data to verify the accuracy of efficacy prediction as a benchmark reference value;
[0032] S232. Construct a dynamic simulation environment and set the action space to select treatment methods and the attributes of mutual generation and mutual restraint;
[0033] S233, using the proximal strategy optimization algorithm to optimize the strategy network and evaluate the long-term benefits of the treatment strategy;
[0034] S234. Extract high-frequency treatment paths from the convergence strategy network, construct a multi-branch decision tree, and generate an interpretable decision map through decision tree pruning.
[0035] Furthermore, based on the decision tree generation scheme, a dynamic dialectical reasoning model is constructed. By solving the optimal solution, the treatment pathway containing the suggested corrections for the following factors is output:
[0036] S241. Construct a game framework that includes a physician agent, a bioinertial constraint agent, and a pharmacological verification agent, and set the strategy space and benefit dimension of each agent;
[0037] S242. Output a set of treatment plans that satisfy the benefits of the multi-agents through multi-agent iteration;
[0038] S243. Based on the evaluation results of the mutual generation and mutual restraint constraint agent, identify the nodes in the treatment plan that violate the mutual generation and mutual restraint principle, calculate the mutual generation and mutual restraint influencing factors, topologically sort the candidate paths, and retain the correction plans of mutual generation reinforcement and mutual restraint inhibition.
[0039] Furthermore, through multi-agent iteration, the output of the treatment plan set that satisfies the multi-agent benefits includes:
[0040] The strategies of each agent are encoded as mixed strategy vectors, and the equilibrium point is searched through a virtual game iterative algorithm. In each round of iteration, each agent updates its optimal response according to the opponent's strategy until the strategy change is less than the convergence threshold, and then outputs the equilibrium strategy combination; and a Pareto frontier screening mechanism is designed to retain a set of treatment plans that simultaneously meet the benefits of multiple agents.
[0041] Furthermore, the output results of the dynamic dialectical reasoning model are obtained, and through the knowledge integration and conflict resolution of various TCM classics, a redundant and dynamically calibrated TCM knowledge base is constructed, including:
[0042] S31. Align the treatment pathways with the standardized corpus at the entity level and establish bidirectional links for semantically overlapping entries, preserving the original contextual information of different classics.
[0043] S32. Detect conflicting entries in different classics, calculate the optimal solution using a Bayesian network, and prioritize solutions that have the highest frequency of clinical verification and comply with the Shengke rule constraints;
[0044] S33. Density clustering is performed on the symptom-syndrome relationship, representative items within the same cluster are retained, and the remaining items are converted to historical versions;
[0045] S34. Build a spatiotemporal sharded graph database storage structure as a TCM knowledge base, and introduce a composite index structure to achieve a full-link response of symptoms-syndromes-treatment methods-solutions.
[0046] Furthermore, based on edge computing and privacy protection technologies, the real-time iteration of knowledge graphs and optimization of clinical efficacy feedback are achieved, including:
[0047] S41. Build a central server and establish a communication channel with the edge computing node through smart contracts;
[0048] S42. Configure local training tasks on edge computing nodes, collect clinical multimodal data in real time for local model training, evaluate actual efficacy, and dynamically update the TCM knowledge base using incremental data;
[0049] S43. Use encryption technology to encrypt collected calculations to protect privacy data.
[0050] In a second aspect, the present invention further provides a knowledge graph construction system based on traditional Chinese medicine classics, comprising:
[0051] The data acquisition module is used to collect multi-source heterogeneous data from traditional Chinese medicine classics for standardization processing. Through optical character recognition and multimodal feature extraction, it generates a standardized corpus aligned across time periods.
[0052] The reasoning and dialectic module is used to extract entity relationships and dialectical rules from the standardized corpus based on the language model and causal reasoning engine, and generate a dynamic dialectical reasoning model that supports nonlinear reasoning;
[0053] The fusion and calibration module is used to obtain the output results of the dynamic dialectical reasoning model and build a redundant and dynamically calibrated TCM knowledge base through the fusion and conflict resolution of knowledge from various TCM classics.
[0054] Update the feedback module to achieve real-time iteration of the knowledge graph and optimization of clinical efficacy feedback based on edge computing and privacy protection technology.
[0055] The beneficial effects of the present invention are:
[0056] 1. Through the standardized processing of multi-source heterogeneous data and the dynamic knowledge fusion mechanism, the problems of cross-era evolution, regional expression differences, and multimodal feature fragmentation in the digitization of traditional Chinese medicine classics have been solved. Deep semantic alignment of text, image, and waveform data has been carried out to build a knowledge base system for real-time retrieval. This system not only fully preserves the classical theoretical framework of traditional Chinese medicine classics, but also organically integrates the characteristics of modern clinical diagnosis and treatment data, forming a dynamic knowledge network that connects the past and the present and verifies multiple sources. It enables the continuous evolution of the knowledge system while protecting the characteristics of academic schools, and provides a scalable intelligent infrastructure for the modernization of traditional Chinese medicine.
[0057] 2. By deeply integrating the theory of mutual generation and mutual restraint with modern causal reasoning technology, we can break through the linear reasoning limitations of traditional knowledge graphs; encode the transmission and transformation rules of the internal organs into the adaptive edge weights of the graph neural network, simulate the evolution path of complex syndromes such as "mixed deficiency and excess" and "mixed cold and heat", and realize the dynamic strategy generation and verification of the entire chain of "theory-method-prescription-drug" in a virtual environment, supporting millisecond-level deduction of individualized diagnosis and treatment plans and multi-dimensional correction suggestion output.
[0058] 3. Locally process sensitive data on edge computing nodes to achieve secure collaborative evolution of knowledge graphs between different edge computing nodes; collect multimodal efficacy data in real time through IoT terminals to provide high-quality data support for evidence-based medical research, forming a benign interactive closed loop of clinical practice and theoretical innovation. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0060] Figure 1 This is a flowchart of a method for constructing a knowledge graph based on TCM classics according to an embodiment of the present invention;
[0061] Figure 2 This is a system principle block diagram of a knowledge graph construction system based on traditional Chinese medicine classics according to an embodiment of the present invention.
[0062] Figure numbers: 1. Data acquisition module; 2. Reasoning and dialectics module; 3. Fusion and calibration module; 4. Update feedback module. DETAILED DESCRIPTION
[0063] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0064] See also Figure 1, provides a method for constructing a knowledge graph based on traditional Chinese medicine classics, including:
[0065] S1. Collect multi-source heterogeneous data of traditional Chinese medicine classics and perform standardization processing. Through optical character recognition and multimodal feature extraction, generate a standardized corpus aligned across time periods.
[0066] In the description of the present invention, S1 includes:
[0067] S11. Extract multi-source heterogeneous data from the ancient book digitization platform and hospital management information system, and perform grayscale correction, ink enhancement, and non-text area filtering according to the data type. The multi-source heterogeneous data includes ancient book images, modern medical case data, and pharmacopoeia database information.
[0068] S12. Deploy an optical character recognition engine, integrate the ancient Chinese character library and domain dictionary, identify homophones and variant characters, establish a dynamic mapping chart, and output standardized text.
[0069] Specifically, the deployed optical character recognition engine adopts a three-stage recognition framework:
[0070] In the first stage, the ResNet-18 directional classification model is used to determine text direction and segment text lines using vertical projection analysis. In the second stage, a library of classical Chinese characters and a dictionary for Traditional Chinese Medicine are integrated, and a multi-scale residual network is used to extract character stroke features. This network builds on the standard ResNet50 with a parallel multi-branch structure: ① 3×3 convolution kernels extract local details; ② 5×5 dilated convolutions capture macrostructure; ③ a self-attention mechanism enhances key stroke recognition. These three features are weightedly fused using channel-wise attention to output a character probability distribution. In the third stage, a generative adversarial network is constructed to address the problem of disambiguating homophones. The generator uses context across time periods to generate semantically ambiguous context fragments; the discriminator measures the semantic consistency of the recognition results with the ontology by calculating the KL divergence.
[0071] S13. Extract visual features of image data and time-frequency features of waveform data from multi-source heterogeneous data respectively to obtain feature vectors of each component.
[0072] Specifically, for image data, the EfficientNet-V2 model is used to extract visual feature vectors (the last layer of global average pooling output); the pulse diagnosis waveform data is decomposed into intrinsic mode functions through empirical mode decomposition, and the energy entropy and time-frequency domain statistics (mean, variance, kurtosis) of each component are calculated to form a feature vector.
[0073] S14. Automatically align terms in standardized text based on the cosine similarity algorithm.
[0074] Specifically, the automatic mapping and matching of high-frequency terms is achieved through cosine similarity calculation, and the similarity of terms in the text is evaluated pairwise. When the similarity value exceeds the preset threshold, the corresponding relationship between terms is automatically established, effectively solving the association problem of synonymous terms in different classics.
[0075] S15. Use graph data to store the multi-dimensional relationship between classics content, physical sign data, and treatment methods and prescriptions, set up a spatiotemporal index structure, and realize distributed storage of processed and standardized text data.
[0076] Specifically, graph database technology was used to model the multidimensional relationships between textual content, patient vital sign data, and treatment methods and prescriptions. By defining a four-dimensional index structure (the temporal dimension divides knowledge versions by dynasty code, the spatial dimension annotates regional medical schools, the modal dimension distinguishes textual classics, clinical data, and experimental research sources, and the semantic dimension links the Six-Jing Syndrome Differentiation classification system), multi-level organization and rapid location of knowledge were achieved.
[0077] The spatiotemporal index adopts a composite key design, combining the timestamp and geocoding to generate a unique partition identifier, and combines the topological characteristics of the graph structure to achieve efficient clustering storage of neighbor data.
[0078] S2. Based on the language model and causal reasoning engine, entity relationships and dialectical rules are extracted from the standardized corpus to generate a dynamic dialectical reasoning model that supports nonlinear reasoning.
[0079] In the description of the present invention, S2 includes:
[0080] S21. Generate a text vector sequence through the pre-trained language model, simulate the ambiguous context of homographs, calculate the semantic deviation between the context vector and the question bank, adjust the entity type confidence weight, and generate an entity attribute matrix with a unified identifier.
[0081] Specifically, to address the semantic ambiguity of homographs, a generative adversarial network framework was constructed. The generator simulates the polysemy of terms in different contexts, for example automatically generating differentiated contextual segments for the word "wind" when representing exogenous pathogenic factors and endogenous liver wind. The discriminator uses multi-level feature extraction to analyze the semantic consistency between the generated text and the ontology. It uses KL divergence to quantify the degree of semantic deviation and calculates a confidence decay factor, dynamically adjusting the weights assigned to entity types. Throughout this process, the generator and discriminator are continuously optimized through adversarial training, ultimately outputting a disambiguated entity attribute matrix.
[0082] S22. Based on the graph causal diffusion model, the mutual generation and mutual restraint relationship is encoded as the prior conditions of the graph network to construct the initial graph structure.
[0083] Among them, the present invention transforms the TCM theory of mutual restraint into computable graph structure constraints, and realizes the deep integration of theoretical rules and clinical practice through data-driven dynamic optimization, ensuring that the generated causal graph not only conforms to the core logic of mutual restraint (such as the inhibitory connection of liver wood restraining spleen earth) but also reflects the statistical laws of symptom co-occurrence.
[0084] In the description of the present invention, S22 includes:
[0085] S221. Convert the generation and restraint theory into graph structure constraints, set the generation relationship between generation and restraint elements as positive weight connection and the restraint relationship as negative weight inhibition, and convert the discrete generation and restraint relationship into a continuous and differentiable probability distribution form to form a generation and restraint constraint matrix.
[0086] Specifically, the TCM theory of mutual restraint is first converted into computable graph structure constraints. By defining the mutual generation relationship between the elements of generation and restraint as a positive weight connection (such as wood generates fire and is given an initial weight of +0.7) and the mutual restraint relationship as a negative weight inhibition (such as wood restrains earth and is set to -0.5), and standardizing the discrete generation and restraint relationships, it is converted into a continuous and differentiable probability distribution form, and a generation and restraint constraint matrix is generated as the theoretical prior basis.
[0087] S222. Dynamically weight the generative constraint matrix and the statistical matrix extracted from the standardized corpus, and map the generative constraint relationship to the entity topology structure to generate an initial graph structure.
[0088] Specifically, a dynamic weighted fusion of the Sheng-Guo constraint matrix and the symptom-syndrome co-occurrence statistical matrix extracted from a standardized corpus is performed. The Sheng-Guo relationship is mapped to an entity-level topological structure through tensor contraction operations to form an initial graph structure. During this process, the significance of symptom co-occurrence is screened using a chi-square test to ensure that the data-driven associations are statistically significant.
[0089] S223. A Markov chain diffusion process is introduced into the initial graph structure, Gaussian noise is gradually injected according to preset noise scheduling parameters, and noise suppression is applied to edge connections that conform to the Sheng-Kuang theory.
[0090] Specifically, after the initial graph structure is generated, a Markov chain diffusion process is introduced to gradually inject Gaussian noise, with the noise scheduling parameters set according to the cosine rule. A protection mechanism is implemented for high-confidence edges that conform to the Shengke theory, reducing the noise standard deviation by 30% to maintain the stability of the core theoretical structure.
[0091] S224. Set the theoretical conformity index to continuously monitor the theoretical deviation of the generated graph structure. When the theoretical deviation is greater than the preset threshold, the weight balancing mechanism is triggered to automatically adjust the fusion coefficient.
[0092] In the description of the present invention, the calculation formula of the theoretical deviation degree is:
[0093] ;
[0094] Where C t is the theoretical deviation degree at time t; W norm (i,j) is the edge e in the generative constraint matrix i,j Theoretical weight of W final (i,j) is the edge e in the final causal graph i,j The weight of ; E is the number of edges.
[0095] Furthermore, the rationality of the generated graph structure is continuously monitored through the degree of theoretical deviation, which calculates the mean absolute deviation of all edge weights from the theoretical value. When Ct > 0.2, a weight rebalancing mechanism is triggered, adjusting the fusion coefficients according to the sign function until the system converges to a steady-state equilibrium. This ensures that the causal graph achieves an optimal balance between theoretical conformity and data fitting error, supporting the precise modeling of complex pathological pathways such as "Liver Wood Riding Spleen Earth."
[0096] S23. Establish a reward function based on the generation-control relationship, use the proximal strategy optimization algorithm to train the dialectical strategy in a virtual environment, and generate a decision tree for the generation-control-organ-treatment chain reasoning.
[0097] In the description of the present invention, S23 includes:
[0098] S231. Provide positive incentives for decisions that conform to the mutual generation path, impose penalty weights on operations that violate the mutual restraint principle, and use historical medical case data to verify the accuracy of efficacy prediction as a benchmark reference value.
[0099] Specifically, the Traditional Chinese Medicine (TCM) theory of mutual generation and mutual restraint is quantified into a computable dynamic scoring mechanism. Treatment decisions that conform to the mutual generation path are assigned a positive incentive value of +0.8; those that violate the mutual restraint principle are penalized with a penalty weight of -0.6. A treatment efficacy prediction module is also introduced. Using a bidirectional LSTM model to analyze historical medical case data, it outputs a treatment efficacy probability value, which is converted into a treatment efficacy reward in the interval [-1, +1] using a linear mapping function. The final reward is the weighted sum of the theoretical conformity score and the predicted efficacy value. The theoretical rule weight coefficient λ is set to 1.2 to strengthen the guiding role of the mutual generation and mutual restraint logic.
[0100] S232. Construct a dynamic simulation environment and set the action space as the treatment method selection and the mutual generation and restraint attributes.
[0101] Specifically, a dynamic simulation environment was constructed to train syndrome differentiation strategies. A multidimensional state space was defined, encompassing visceral function indicators and sheng-ku balance coefficients, combined into a 15-dimensional state vector. The action space was divided into discrete treatment selection and continuous sheng-ku attribute manipulation. This environment supports dynamic deduction of pathogenesis pathways, such as "liver depression and spleen deficiency → soothing the liver and strengthening the spleen → updating sheng-ku state."
[0102] S233. Use the proximal strategy optimization algorithm to optimize the strategy network and evaluate the long-term benefits of the treatment strategy.
[0103] Specifically, the policy network receives inputs related to the patient's current state, including organ function indicators and the Sheng-Guo balance coefficient, and outputs a probability distribution for treatment options and the intensity of adjustments to their Sheng-Guo properties. During training, an experience replay mechanism is used to store historical interaction trajectories. Importance sampling techniques are used to compare the differences in action selection between the new and old strategies under the same conditions to calculate the direction of policy updates. A generalized advantage estimation method is also introduced to integrate immediate treatment effects with potential future benefits. A long-term benefit estimate is generated through the weighted accumulation of multi-step time difference errors.
[0104] S234. Extract high-frequency treatment paths from the convergence strategy network, construct a multi-branch decision tree, and generate an interpretable decision map through decision tree pruning.
[0105] Specifically, the pruning process utilizes a comprehensive cost-complexity assessment, eliminating redundant branches while preserving key decision paths, ensuring the simplicity and clinical interpretability of the tree structure. The resulting decision tree supports chain reasoning from the initial pathogenesis to the final treatment. Each internal node is labeled with the judgment logic and confidence level for the mutual generation and mutual restraint state, and leaf nodes are associated with specific treatment operations and their expected mutual generation and mutual restraint impact strength, forming a visual diagnosis and treatment map that combines theoretical rigor with practical guidance.
[0106] S24. Based on the decision tree generation scheme, a dynamic dialectical reasoning model is constructed to solve the optimal solution and output a treatment path that includes suggestions for mutual generation and mutual restraint correction.
[0107] In the description of the present invention, S24 includes:
[0108] S241. Construct a game framework that includes physician agents, life-control constraint agents, and pharmacological verification agents, and set the strategy space and benefit dimension of each agent.
[0109] Specifically, the strategy space of the physician agent is defined as a set of candidate treatment options extracted from the decision tree, and its payoff function integrates the predicted efficacy value and safety score; the mutual generation and mutual restraint constraint agent calculates the compliance of the mutual generation and mutual restraint rules by traversing the edge weights of the causal graph, and the payoff function is positively correlated with the matching ratio of the mutual generation path and negatively correlated with the mutual restraint conflict ratio; the pharmacology verification agent detects incompatibilities based on the pharmacopoeia database, and the payoff function is inversely proportional to the risk level.
[0110] S242. Through multi-agent iteration, output a set of treatment plans that meet the benefits of multi-agents.
[0111] Specifically, the strategy of each agent is encoded as a mixed strategy vector, and the equilibrium point is searched through a virtual game iterative algorithm. In each round of iteration, each agent updates its own optimal response according to the opponent's strategy until the strategy change amplitude is less than the convergence threshold, and then outputs the equilibrium strategy combination; and a Pareto frontier screening mechanism is designed to retain a set of treatment plans that simultaneously meet the benefits of multiple agents.
[0112] In each round of iteration, the physician agent calculates the expected therapeutic benefit based on the current Shengke constraints and pharmacological verification strategy, updates the probability of treatment plan selection through the mirror descent method, prioritizes improving the efficacy but is constrained by the Shengke rule penalty coefficient; the Shengke constraint agent adjusts the edge weight attenuation rate based on the Shengke influence value in the treatment plan; the pharmacological verification agent dynamically adjusts the rejection threshold based on historical risk data.
[0113] When the three-way strategy vectors meet convergence conditions, an equilibrium strategy combination is output. Then, through the Pareto front screening mechanism, a non-dominated sorting algorithm is used to retain non-inferior treatment options in terms of efficacy, rule compliance, and safety from the equilibrium solution set, while eliminating inferior solutions that are comprehensively surpassed by other options.
[0114] S243. Based on the evaluation results of the mutual generation and mutual restraint constraint agent, identify the nodes in the treatment plan that violate the mutual generation and mutual restraint principle, calculate the mutual generation and mutual restraint influencing factors, topologically sort the candidate paths, and retain the correction plans of mutual generation reinforcement and mutual restraint inhibition.
[0115] Specifically, based on the results of the mutual generation and mutual restraint constraint assessment, nodes in the treatment plan that violate the mutual generation and mutual restraint principle are identified. The mutual generation and mutual restraint influencing factors are calculated through gradient backpropagation. The offending edges are located in the causal graph, and the gradient signal is backpropagated along the treatment pathway to generate weight adjustment recommendations. Candidate pathways are topologically sorted and prioritized based on their mutual generation and mutual restraint ratios. Corrective solutions that simultaneously improve efficacy and meet theoretical compliance are retained.
[0116] S3. Obtain the output results of the dynamic dialectical reasoning model, and build a redundant and dynamically calibrated TCM knowledge base through knowledge integration and conflict resolution of various TCM classics.
[0117] In the description of the present invention, S3 includes:
[0118] S31. Perform entity-level alignment of treatment pathways with standardized corpora and establish bidirectional links for entries with semantic overlap, preserving the original contextual information of different classics.
[0119] Specifically, the treatment pathways output by the dynamic dialectical reasoning model were finely aligned with a standardized corpus. A term mapping table covering TCM classics was constructed. For each entity, its contextual window features in classical literature and co-occurrence statistics in modern medical records were extracted. Cross-modal semantic similarity was calculated, and bidirectional hyperlinks were established for entities with semantic overlap. Contextual information, such as chapter references and medical annotations from the original texts, was encoded as structured metadata.
[0120] S32. Detect conflicting entries in different classics, calculate the optimal solution through Bayesian network, and give priority to the solutions with the highest clinical verification frequency and compliance with the constraints of the mutual generation and mutual restraint rules.
[0121] Specifically, for conflicting items, by constructing a Bayesian network reasoning framework, three types of evidence nodes are defined, namely clinical verification frequency, compliance with the Shengke rule, and medical authority weight. The Markov chain Monte Carlo method is used to solve the posterior probability distribution. When the confidence level of the optimal solution of the conflicting item exceeds the preset threshold, the solution that meets the conditions of "high verification frequency and Shengke constraint deviation" is automatically retained, and the remaining conflicting items are transferred to the queue for verification and trigger the expert review process.
[0122] S33. Density clustering is performed on the symptom-syndrome relationship, representative items within the same cluster are retained, and the remaining items are converted to historical versions.
[0123] Specifically, to address the redundancy problem of symptom-syndrome relationships, the DBSCAN density clustering algorithm was used to group symptom-related data, and a symptom similarity matrix based on word vectors was constructed. Then, the efficacy dimensions of the treatment plans (such as efficacy labels such as clearing heat and removing dampness) were integrated to calculate the composite similarity. For redundant items in the same cluster (such as "clearing heat and detoxifying" and "purging fire and detoxifying"), representative items were screened using a comprehensive weight formula, and the version traceability information of the remaining items was retained when they were converted to historical versions.
[0124] S34. Deploy lightweight edge computing nodes to collect the improvement of physiological indicators of traditional Chinese medicine treatment plans in real time, aggregate encrypted gradients through the federated learning framework, and update the weights of knowledge base entries.
[0125] S34. Build a spatiotemporal sharded graph database storage structure as a TCM knowledge base, and introduce a composite index structure to achieve a full-link response of symptoms-syndromes-treatment methods-solutions.
[0126] S4. Based on edge computing and privacy protection technologies, real-time iteration of knowledge graphs and optimization of clinical efficacy feedback are achieved.
[0127] In the description of the present invention, S4 includes:
[0128] S41. Build a central server and establish a communication channel with the edge computing node through smart contracts.
[0129] Specifically, a central server is built as the global coordination hub. Based on the blockchain architecture, the server deploys smart contracts to manage the access and communication processes of edge computing nodes. The central server establishes a trusted connection with each edge computing node through a pre-set node authentication mechanism.
[0130] S42. Configure local training tasks on edge computing nodes, collect clinical multimodal data in real time for local model training, evaluate actual efficacy, and use incremental data to dynamically update the TCM knowledge base.
[0131] Specifically, local training tasks on edge computing nodes are deployed using containerization technology. Each node runs an independent Docker instance encapsulating a lightweight version of the dynamic dialectical reasoning model. Multimodal clinical data, including tongue images, pulse waveforms, and electronic medical record text, is collected in real time via medical IoT devices.
[0132] The actual efficacy index calculation combines short-term indicators (such as symptom relief rate) and long-term indicators (such as three-month recurrence rate) to generate a 0-1 standardized efficacy score through fuzzy comprehensive evaluation.
[0133] S43. Use encryption technology to encrypt collected calculations to protect privacy data.
[0134] See also Figure 2 , also provides a knowledge graph construction system based on traditional Chinese medicine classics, including:
[0135] Data acquisition module 1 is used to collect multi-source heterogeneous data of traditional Chinese medicine classics for standardization processing, and generate a standardized corpus aligned across eras through optical character recognition and multimodal feature extraction.
[0136] The reasoning and dialectics module 2 is used to extract entity relationships and dialectical rules from the standardized corpus based on the language model and causal reasoning engine, and generate a dynamic dialectical reasoning model that supports nonlinear reasoning.
[0137] The fusion and calibration module 3 is used to obtain the output results of the dynamic dialectical reasoning model, and to build a redundant and dynamically calibrated TCM knowledge base through the knowledge fusion and conflict resolution of various TCM classics.
[0138] Update feedback module 4, which is used to achieve real-time iteration of the knowledge graph and optimization of clinical efficacy feedback based on edge computing and privacy protection technology.
[0139] To sum up, with the help of the above technical solutions of the present invention, through the standardization processing of multi-source heterogeneous data and the dynamic knowledge fusion mechanism, the problems of cross-era evolution, regional expression differences, and multimodal feature fragmentation in the digitization process of traditional Chinese medicine classics are solved. Deep semantic alignment of text, image, and waveform data is performed to construct a real-time retrieval knowledge base system, which not only completely retains the classical theoretical framework of traditional Chinese medicine classics, but also organically integrates the characteristics of modern clinical diagnosis and treatment data, forming a dynamic knowledge network that connects ancient and modern times and has multi-source mutual verification, so as to realize the continuous evolution of the knowledge system while protecting the characteristics of academic schools, and provide a scalable intelligent infrastructure for the modernization of traditional Chinese medicine.
[0140] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
Claims
1. A method for constructing a knowledge graph based on traditional Chinese medicine classics, characterized in that: include: S1. Collect multi-source heterogeneous data of traditional Chinese medicine classics and perform standardization processing. Through optical character recognition and multimodal feature extraction, generate a standardized corpus aligned across time periods. S2, based on the language model and causal reasoning engine, extracts entity relationships and dialectical rules from the standardized corpus to generate a dynamic dialectical reasoning model that supports nonlinear reasoning; S3. Obtain the output of the dynamic dialectical reasoning model and build a redundant and dynamically calibrated TCM knowledge base by integrating and resolving conflicts among various TCM classics. S4. Based on edge computing and privacy protection technologies, it realizes real-time iteration of knowledge graphs and optimization of clinical efficacy feedback; The method of extracting entity relationships and dialectical rules from a standardized corpus based on a language model and a causal reasoning engine to generate a dynamic dialectical reasoning model that supports nonlinear reasoning includes: S21. Generate a text vector sequence using a pre-trained language model, simulate the ambiguous context of homographs, calculate the semantic deviation between the context vector and the current question bank, adjust the entity type confidence weight, and generate an entity attribute matrix with a unified identifier. S22. Based on the graph causal diffusion model, the generation-restraint relationship is encoded as a priori conditions of the graph network to construct the initial graph structure; S23. Establish a reward function based on the generation-restraint relationship, use the proximal strategy optimization algorithm to train the dialectical strategy in a virtual environment, and generate a decision tree for the generation-restraint-organ-treatment chain reasoning; S24. Based on the decision tree generation scheme, a dynamic dialectical reasoning model is constructed to find the optimal solution and output a treatment path including suggestions for synergy and correction; Based on the graph causal diffusion model, the generation-restraint relationship is encoded as a priori conditions of the graph network, and the initial graph structure is constructed, including: S221. Convert the generation and restraint theory into graph structure constraints, setting the generation relationship between generation and restraint elements as positive weight connection and the restraint relationship as negative weight inhibition, and convert the discrete generation and restraint relationships into a continuous and differentiable probability distribution form to form a generation and restraint constraint matrix; S222, dynamically weighted fusion of the generative constraint matrix and the statistical matrix extracted from the standardized corpus, and mapping the generative constraint relationship to the entity topology structure to generate an initial graph structure; S223, introducing a Markov chain diffusion process into the initial graph structure, gradually injecting Gaussian noise according to preset noise scheduling parameters, and applying noise suppression to edge connections that conform to the Shengke theory; S224. Set the theoretical conformity index to continuously monitor the theoretical deviation of the generated graph structure. When the theoretical deviation is greater than the preset threshold, the weight balancing mechanism is triggered to automatically adjust the fusion coefficient.
2. A method for constructing a knowledge graph based on TCM classics according to claim 1, characterized in that: The method of collecting multi-source heterogeneous data of traditional Chinese medicine classics and performing standardization processing, and generating a standardized corpus aligned across eras through optical character recognition and multimodal feature extraction includes: S11, extracting multi-source heterogeneous data from the ancient book digitization platform and the hospital management information system, and performing grayscale correction, ink enhancement, and non-text area filtering according to data type, wherein the multi-source heterogeneous data includes ancient book images, modern medical case data, and pharmacopoeia database information; S12. Deploy an optical character recognition engine, integrate the classical Chinese character library and domain dictionary, identify homophones and variant forms of characters, create a dynamic mapping chart, and output standardized text; S13, extracting visual features of image data and time-frequency features of waveform data from the multi-source heterogeneous data, respectively, to obtain feature vectors of each component; S14, automatically aligning terms in standardized text based on the cosine similarity algorithm; S15. Use graph data to store the multi-dimensional relationship between classics content, physical sign data, and treatment methods and prescriptions, set up a spatiotemporal index structure, and realize distributed storage of processed and standardized text data.
3. The method for constructing a knowledge graph based on TCM classics according to claim 2, characterized in that: The method of establishing a reward function based on the generation-restraint relationship, using a proximal strategy optimization algorithm to train a dialectical strategy in a virtual environment, and outputting a decision tree for the generation-restraint-organ-treatment chain reasoning includes: S231. Provide positive incentives for decisions that conform to the mutual growth path, impose penalty weights on operations that violate the mutual restraint principle, and use historical medical case data to verify the accuracy of efficacy prediction as a benchmark reference value; S232. Construct a dynamic simulation environment and set the action space to select treatment methods and the attributes of mutual generation and mutual restraint; S233, using the proximal strategy optimization algorithm to optimize the strategy network and evaluate the long-term benefits of the treatment strategy; S234. Extract high-frequency treatment paths from the convergence strategy network, construct a multi-branch decision tree, and generate an interpretable decision map through decision tree pruning.
4. The method for constructing a knowledge graph based on TCM classics according to claim 3, characterized in that: The decision tree-based generation scheme constructs a dynamic dialectical reasoning model, and by solving the optimal solution, outputs a treatment path containing suggestions for synergistic correction, including: S241. Construct a game framework that includes a physician agent, a bioinertial constraint agent, and a pharmacological verification agent, and set the strategy space and benefit dimension of each agent; S242. Output a set of treatment plans that satisfy the benefits of the multi-agents through multi-agent iteration; S243. Based on the evaluation results of the mutual generation and mutual restraint constraint agent, identify the nodes in the treatment plan that violate the mutual generation and mutual restraint principle, calculate the mutual generation and mutual restraint influencing factors, topologically sort the candidate paths, and retain the correction plans of mutual generation reinforcement and mutual restraint inhibition.
5. The method for constructing a knowledge graph based on TCM classics according to claim 4, characterized in that: The output of the treatment plan set that satisfies the multi-agent benefits through multi-agent iteration includes: The strategies of each agent are encoded as mixed strategy vectors, and the equilibrium point is searched through a virtual game iterative algorithm. In each round of iteration, each agent updates its optimal response according to the opponent's strategy until the strategy change is less than the convergence threshold, and then outputs the equilibrium strategy combination; and a Pareto frontier screening mechanism is designed to retain a set of treatment plans that simultaneously meet the benefits of multiple agents.
6. The method for constructing a knowledge graph based on TCM classics according to claim 1, characterized in that: The output of the dynamic dialectical reasoning model is obtained, and a redundant and dynamically calibrated TCM knowledge base is constructed through knowledge integration and conflict resolution of various TCM classics. The method includes: S31. Align the treatment pathways with the standardized corpus at the entity level and establish bidirectional links for semantically overlapping entries, preserving the original contextual information of different classics. S32. Detect conflicting entries in different classics, calculate the optimal solution using a Bayesian network, and prioritize solutions that have the highest frequency of clinical verification and comply with the Shengke rule constraints; S33. Density clustering is performed on the symptom-syndrome relationship, representative items within the same cluster are retained, and the remaining items are converted to historical versions; S34. Build a spatiotemporal sharded graph database storage structure as a TCM knowledge base, and introduce a composite index structure to achieve a full-link response of symptoms-syndromes-treatment methods-solutions.
7. The method for constructing a knowledge graph based on TCM classics according to claim 1, characterized in that: The real-time iteration of the knowledge graph and optimization of clinical efficacy feedback based on edge computing and privacy protection technologies include: S41. Build a central server and establish a communication channel with the edge computing node through smart contracts; S42. Configure local training tasks on edge computing nodes, collect clinical multimodal data in real time for local model training, evaluate actual efficacy, and dynamically update the TCM knowledge base using incremental data; S43. Use encryption technology to encrypt collected calculations to protect privacy data.
8. A knowledge graph construction system based on TCM classics, used to implement the knowledge graph construction method based on TCM classics according to any one of claims 1 to 7, characterized in that: include: The data acquisition module is used to collect multi-source heterogeneous data from traditional Chinese medicine classics for standardization processing. Through optical character recognition and multimodal feature extraction, it generates a standardized corpus aligned across time periods. The reasoning and dialectic module is used to extract entity relationships and dialectical rules from the standardized corpus based on the language model and causal reasoning engine, and generate a dynamic dialectical reasoning model that supports nonlinear reasoning; The fusion and calibration module is used to obtain the output results of the dynamic dialectical reasoning model and build a redundant and dynamically calibrated TCM knowledge base through the fusion and conflict resolution of knowledge from various TCM classics. Update the feedback module to achieve real-time iteration of the knowledge graph and optimization of clinical efficacy feedback based on edge computing and privacy protection technology.
Citation Information
Patent Citations
Construction method and system of traditional Chinese medicine knowledge graph, and storage medium
CN116049426A
Traditional Chinese medicine classical book knowledge base feedback correction method and system
CN117271796A