Knowledge graph construction method and system based on Chinese medicine classics
By standardizing the processing of traditional Chinese medicine classic data, extracting entity relationships and dialectical rules based on language models, fusion and conflict dissolution, and real-time iteration is achieved using edge computing, multiple technical limitations in the construction of traditional Chinese medicine classic knowledge graphs are solved, and deep semantic alignment and dynamic diagnosis and treatment plan generation are achieved.
Patent Information
- Application Number
- CN202510608712.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-13
AI Technical Summary
It is difficult for the existing technology to effectively build a knowledge graph of traditional Chinese medicine classics, especially in terms of term standardization, semantic understanding depth, multi-source data integration and dynamic updates.
By collecting multi-source heterogeneous data from medical classics for standardization, a standardized corpus that is aligned across eras is generated; a dynamic dialectical reasoning model that supports nonlinear reasoning is extracted based on the language model and the causal reasoning engine to generate a dynamic dialectical reasoning model that supports nonlinear reasoning; a de-redundant and dynamically calibrated traditional Chinese medicine knowledge base is constructed through knowledge fusion and conflict removal; and edge computing and privacy protection technology is used to achieve real-time iteration of the knowledge graph and clinical efficacy feedback optimization.
It realizes deep semantic alignment and real-time retrieval of traditional Chinese medicine classics knowledge, breaks through the linear reasoning limitations of traditional knowledge graphs, supports the dynamic generation and verification of individualized diagnosis and treatment plans, and realizes the secure collaborative evolution of knowledge graphs and the benign interactive closed loop of clinical practice and theoretical innovation through edge computing.
Smart Images

Figure CN120124733A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge graph construction, and in particular to a method and system for constructing a knowledge graph based on traditional Chinese medicine classics. Background Art
[0002] As a theoretical carrier passed down for thousands of years, traditional Chinese medicine classics include classic works such as "Huangdi Neijing" and "Treatise on Febrile Diseases", which record systematic knowledge such as zang-fu organs and meridians, etiology and pathogenesis, treatment principles and methods, and pharmaceutical formulas. However, most of this knowledge exists in the form of unstructured text, and there are problems such as differences in term expressions (such as semantic overlap between "liver qi stagnation" and "stagnation of liver qi") and fragmentation of knowledge associations (such as the correspondence between formulas and syndromes scattered in different literatures).
[0003] With the penetration of artificial intelligence technology in the medical field, the demand for structured knowledge in scenarios such as traditional Chinese medicine clinical decision support, personalized diagnosis and treatment, and new drug research and development is increasing day by day. The traditional model relying on manual collation and experience inheritance can no longer meet the application requirements of high efficiency and precision. Therefore, the construction of a traditional Chinese medicine knowledge graph is the core infrastructure for the modernization and intelligentization of traditional Chinese medicine, aiming to achieve digital reconstruction, semantic association, and intelligent application of ancient book knowledge through technical means, and become an important breakthrough for promoting the development of traditional Chinese medicine disciplines.
[0004] Currently, some technologies have tried to solve this problem. For example, entity recognition methods based on rule matching and dictionaries (such as the Traditional Chinese Medicine Language System TCMLS that matches entities through a preset dictionary), and relationship extraction technologies combined with LSTM neural networks. These technologies have achieved automatic extraction of entities and relationships to a certain extent, but there are still significant limitations.
[0005] Firstly, insufficient term standardization leads to difficulties in knowledge integration. For example, there is a lack of a unified mapping mechanism for the naming differences of the same syndrome in different classics (such as the difference between "wind-heat cold" and "warm disease"), and semantic alignment requires manual intervention. Secondly, the existing methods have limited depth of semantic understanding of ancient book texts. For example, the LSTM model has weak capabilities in capturing long-distance dependencies and complex semantic relationships (such as the multi-dimensional syndrome differentiation logic of "deficiency-excess complex syndrome"), and does not fully consider the dynamic reasoning characteristics in the traditional Chinese medicine theory system (such as the non-linear transmission of the generation and restriction relationship).
[0006] In addition, the multi-source data integration technology is not yet mature, and the heterogeneity of ancient book literature, modern medical records, and pharmacopoeia databases makes it difficult to automatically resolve knowledge redundancy and conflicts (such as the contradiction between the "eighteen incompatible medicaments" theory in ancient books and modern pharmacology conclusions). More critically, existing knowledge graphs are mostly limited to static storage, lacking a dynamic update mechanism (such as the real-time iteration of the diagnosis and treatment plan for sudden infectious diseases) and a clinical feedback loop (such as the reverse correction of knowledge weights by patient efficacy data).
[0007] Therefore, there is an urgent need to design a method and system for constructing a knowledge graph based on traditional Chinese medicine classics to solve the above-mentioned technical problems. Summary of the Invention
[0008] Based on this, it is necessary to provide a method and system for constructing a knowledge graph based on traditional Chinese medicine classics to address the above technical problems.
[0009] In the first aspect, the present invention provides a method for constructing a knowledge graph based on traditional Chinese medicine classics, including: S1. Collect multi-source heterogeneous data of traditional Chinese medicine classics for standardization processing, and generate a cross-era aligned standardized corpus through optical character recognition and multi-modal feature extraction; S2. Extract entity relationships and dialectical rules from the standardized corpus based on a language model and a causal reasoning engine, and generate a dynamic dialectical reasoning model that supports non-linear reasoning; S3. Obtain the output results of the dynamic dialectical reasoning model, and construct a redundant-free and dynamically calibrated traditional Chinese medicine knowledge base through knowledge fusion and conflict resolution of various traditional Chinese medicine classics; S4. Based on edge computing and privacy protection technologies, realize real-time iteration of the knowledge graph and optimization of clinical efficacy feedback.
[0010] Further, collecting multi-source heterogeneous data of traditional Chinese medicine classics for standardization processing, and generating a cross-era aligned standardized corpus through optical character recognition and multi-modal feature extraction includes: S11. Extract multi-source heterogeneous data from the ancient book digitization platform and the hospital management information system, and perform gray correction, ink enhancement, and non-text area filtering processing according to the data type. The multi-source heterogeneous data includes ancient book images, modern medical case data, and pharmacopoeia database information; S12. Deploy an optical character recognition engine, integrate a special ancient Chinese character library and a domain dictionary, establish a dynamic mapping chart by recognizing interchangeable characters and variant forms of Chinese characters, and output standardized text; S13. Extract the visual features of image data and the time-frequency features of waveform data in the multi-source heterogeneous data respectively to obtain the feature vectors of each component; S14. Automatically align the terms in the standardized text based on the cosine similarity algorithm; S15. Adopt a graph data structure to store the multi-dimensional relationships of classic content - physical signs data - treatment prescriptions, set a spatio-temporal index structure, and realize the distributed storage of the processed standardized text data.
[0011] Further, extracting entity relationships and dialectical rules from the standardized corpus based on a language model and a causal reasoning engine, and generating a dynamic dialectical reasoning model that supports non-linear reasoning includes: S21. Generate a text vector sequence through a pre-trained language model, simulate the ambiguous context of homographs, calculate the semantic deviation degree between the context vector and the semantic library of this question bank, adjust the confidence weight of entity types, and generate an entity attribute matrix with unified identifiers; S22. Based on the graph causal diffusion model, encode the generating and restricting relationships as prior conditions of the graph network, and construct an initial graph structure; S23. Establish a reward function according to the generating and restricting relationships, and use the proximal policy optimization algorithm to train the dialectical strategy in a virtual environment to generate a decision tree for the chain reasoning of generating-restricting - zang-fu organs - treatment methods; S24. Based on the decision tree generation scheme, construct a dynamic dialectical reasoning model, and by solving the optimal solution, output a treatment path containing generating-restricting correction suggestions.
[0012] Further, based on the graph causal diffusion model, encoding the generating and restricting relationships as prior conditions of the graph network, the construction of the initial graph structure includes: S221. Convert the generating and restricting theory into graph structure constraint conditions, set the generating relationship between generating and restricting elements as positive weight connection, the restricting relationship as negative weight suppression, and convert the discrete generating and restricting relationships into a continuous and differentiable probability distribution form to form a generating and restricting constraint matrix; S222. Dynamically and weightedly fuse the generating and restricting constraint matrix with the statistical matrix extracted from the standardized corpus, and map the generating and restricting relationships to the entity topological structure to generate an initial graph structure; S223. Introduce a Markov chain diffusion process into the initial graph structure, gradually inject Gaussian noise according to the preset noise scheduling parameters, and apply noise suppression to the edge connections that conform to the generating and restricting theory; S224. Set a theoretical compliance index to continuously monitor the degree of theoretical deviation of the generated graph structure. When the degree of theoretical deviation is greater than the preset threshold, trigger a weight balance mechanism to automatically adjust the fusion coefficient.
[0013] Further, establishing a reward function according to the generating and restricting relationships, and using the proximal policy optimization algorithm to train the dialectical strategy in a virtual environment, the output of the decision tree for the chain reasoning of generating-restricting - zang-fu organs - treatment methods includes: S231. Give positive incentives to the decisions that conform to the generating path, impose penalty weights on the operations that violate the restricting principle, and use the verification of the accuracy of curative effect prediction based on historical medical record data as a reference benchmark value; S232. Construct a dynamic simulation environment, and set the action space as treatment method selection and generating-restricting attributes; S233. Use the proximal policy optimization algorithm to optimize the policy network and evaluate the long-term benefits of the treatment strategy; S234. Extract high-frequency treatment paths from the converged policy network, construct a multi-branch decision tree, and generate an interpretable decision graph through decision tree pruning.
[0014] Further, based on the decision tree generation scheme, a dynamic dialectical reasoning model is constructed. By solving the optimal solution, the treatment path including the generating and restraining correction suggestions is output, including: S241. Construct a game framework including a physician agent, a generating and restraining constraint agent, and a pharmacological verification agent, and set the strategy space and benefit dimension of each agent; S242. Through multi-agent iteration, output a set of treatment plans that meet the benefits of multiple agents; S243. Based on the evaluation results of the generating and restraining constraint agent, identify the nodes in the treatment plan that violate the generating and restraining principles, calculate the generating and restraining influence factors, perform topological sorting on the candidate paths, and retain the correction plans of generating and strengthening, restraining and suppressing.
[0015] Further, through multi-agent iteration, the output of a set of treatment plans that meet the benefits of multiple agents includes: Encode the strategies of each agent into a mixed strategy vector, search for the equilibrium point through the virtual game iteration algorithm. In each round of iteration, each agent updates its own optimal response according to the opponent's strategy until the strategy change amplitude is less than the convergence threshold, and output the equilibrium strategy combination; and design a Pareto front screening mechanism to retain the set of treatment plans that meet the benefits of multiple agents at the same time.
[0016] Further, obtain the output results of the dynamic dialectical reasoning model, and construct a redundant-reduced and dynamically calibrated traditional Chinese medicine knowledge base through the knowledge fusion and conflict resolution of various traditional Chinese medicine classics, including: S31. Align the treatment path with the standardized corpus at the entity level, establish two-way links for the entries with semantic overlap, and retain the original context information of different classics; S32. Detect the conflicting entries in different classics, calculate the optimal solution through the Bayesian network, and preferentially retain the plan with the highest clinical verification frequency and meeting the constraints of the generating and restraining rules; S33. Perform density clustering on the symptom-syndrome relationship, retain the representative items within the same cluster, and convert the remaining entries into historical versions; S34. Build a spatio-temporal sharding graph database storage structure as the traditional Chinese medicine knowledge base, and introduce a composite index structure to achieve the full-link response of symptoms-syndromes-treatment methods-plans.
[0017] Further, based on edge computing and privacy protection technologies, realize the real-time iteration of the knowledge graph and the optimization of clinical efficacy feedback, including: S41. Build a central server and establish a communication channel with the edge computing nodes through smart contracts; S42. Configure local training tasks on the edge computing nodes, collect clinical multi-modal data in real time for local model training, evaluate the actual efficacy, and dynamically update the traditional Chinese medicine knowledge base using incremental data; S43. Encryption technology is adopted to encrypt the acquisition and calculation to achieve the protection of privacy data.
[0018] In a second aspect, the present invention also provides a knowledge graph construction system based on traditional Chinese medicine classics, including: A data acquisition module, which is used to collect and standardize multi-source heterogeneous data of traditional Chinese medicine classics, and generate a cross-era aligned standardized corpus through optical character recognition and multi-modal feature extraction; An inference and dialectics module, which is used to extract entity relationships and dialectics rules from the standardized corpus based on a language model and a causal inference engine, and generate a dynamic dialectical inference model that supports non-linear inference; A fusion and calibration module, which is used to obtain the output results of the dynamic dialectical inference model, and construct a redundant-free and dynamically calibrated traditional Chinese medicine knowledge base through knowledge fusion and conflict resolution of various traditional Chinese medicine classics; An update and feedback module, which is used to realize the real-time iteration of the knowledge graph and the optimization of clinical efficacy feedback based on edge computing and privacy protection technologies.
[0019] The beneficial effects of the present invention are as follows: 1. Through the standardization processing of multi-source heterogeneous data and the dynamic knowledge fusion mechanism, the problems existing in the digitalization process of traditional Chinese medicine classics, such as cross-era evolution, regional expression differences, and multi-modal feature fragmentation, are solved. The text, image, and waveform data are deeply semantically aligned, and a knowledge base system for real-time retrieval is constructed. It not only completely retains the classic theoretical framework of traditional Chinese medicine classics but also organically integrates the characteristics of modern clinical diagnosis and treatment data, forming a dynamic knowledge network that runs through ancient and modern times and is mutually verified by multiple sources, realizing the continuous evolution of the knowledge system while protecting the characteristics of academic schools, and providing an extensible intelligent infrastructure for the modernization of traditional Chinese medicine.
[0020] 2. By deeply integrating the theory of generation, restriction, over-restriction, and counter-restriction with modern causal inference technology, the linear inference limitation of the traditional knowledge graph is broken through; the law of zang-fu organ transmission and change is encoded as the adaptive edge weight of the graph neural network, simulating the evolution paths of complex syndromes such as "intermingled deficiency and excess" and "intermingled cold and heat", and realizing the generation and verification of the dynamic strategy of the whole chain of "theory-method-prescription-herb" in a virtual environment, supporting the millisecond-level deduction of individualized diagnosis and treatment plans and the output of multi-dimensional correction suggestions.
[0021] 3. By localizing the processing of sensitive data at the edge computing node, the safe collaborative evolution of the knowledge graph among different edge computing nodes is realized; through the real-time collection of multi-modal efficacy data by the Internet of Things terminal, high-quality data support is provided for evidence-based medicine research, forming a virtuous interaction closed-loop between clinical practice and theoretical innovation. Description of the Drawings
[0022] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings: Figure 1 is a flowchart of a method for constructing a knowledge graph based on traditional Chinese medicine classics according to an embodiment of the present invention; Figure 2 is a system principle block diagram of a system for constructing a knowledge graph based on traditional Chinese medicine classics according to an embodiment of the present invention.
[0023] Reference numerals in the drawings: 1, data acquisition module; 2, reasoning and dialectics module; 3, fusion and calibration module; 4, update and feedback module. Detailed implementation manners
[0024] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0025] Please refer to Figure 1 , which provides a method for constructing a knowledge graph based on traditional Chinese medicine classics, including: S1. Collect multi-source heterogeneous data of traditional Chinese medicine classics for standardization processing, and generate a cross-era aligned standardized corpus through optical character recognition and multi-modal feature extraction.
[0026] In the description of the present invention, S1 includes: S11. Extract multi-source heterogeneous data from the ancient book digitization platform and the hospital management information system, and perform gray correction, ink enhancement and non-text area filtering processing according to the data type. The multi-source heterogeneous data includes ancient book images, modern medical record data and pharmacopoeia database information.
[0027] S12. Deploy an optical character recognition engine, integrate a special ancient Chinese character library and a domain dictionary, establish a dynamic mapping chart by recognizing interchangeable characters and variant characters, and output standardized text.
[0028] Specifically, the deployed optical character recognition engine adopts a three-stage recognition framework: In the first stage, the ResNet-18 direction classification model is used to judge the text direction, and the text lines are segmented by combining the vertical projection analysis method. In the second stage, a special ancient Chinese character library and a dictionary in the field of traditional Chinese medicine are integrated, and a multi-scale residual network is used to extract the character stroke features. This network adds a parallel multi-branch structure on the basis of the standard ResNet50: ① A 3×3 convolutional kernel is used to extract local details; ② A 5×5 dilated convolution is used to capture the macroscopic structure; ③ The self-attention mechanism is used to strengthen the identification of key strokes. The three-way features are weighted and fused through channel attention and then the character probability distribution is output. In the third stage, an adversarial generation network is constructed to solve the problem of disambiguating interchangeable characters: the generator uses the context across epochs to generate context fragments containing semantic ambiguities; the discriminator measures the semantic consistency between the recognition result and the ontology library by calculating the KL divergence.
[0029] S13. Respectively extract the visual features of the image data and the time-frequency features of the waveform data in the multi-source heterogeneous data to obtain the feature vectors of each component.
[0030] Specifically, for the image data, the EfficientNet-V2 model is used to extract the visual feature vector (the output of the last global average pooling layer); the pulse diagnosis waveform data is decomposed into intrinsic mode functions by empirical mode decomposition, and the energy entropy and time-frequency domain statistics (mean, variance, kurtosis) of each component are calculated to form the feature vector.
[0031] S14. Based on the cosine similarity algorithm, automatically align the terms in the standardized text.
[0032] Specifically, the automatic mapping and matching of high-frequency terms is realized through cosine similarity calculation, and the pairwise similarity evaluation of the terms in the text is carried out. When the similarity value exceeds the preset threshold, the corresponding relationship between the terms is automatically established, effectively solving the problem of the association of synonymous and different-name terms between different ancient books.
[0033] S15. Use graph data to store the multi-dimensional relationship of the ancient book content - physical sign data - treatment prescriptions, set the spatio-temporal index structure, and realize the distributed storage of the processed standardized text data.
[0034] Specifically, graph database technology is used to model the multi-dimensional relationship between the ancient book content, the patient's physical sign data and the treatment prescriptions. By defining a four-dimensional index structure (the time dimension divides the knowledge version according to the dynasty coding, the space dimension marks the characteristics of the regional medical school, the modal dimension distinguishes the sources of text ancient books, clinical data and experimental research, and the semantic dimension is associated with the six-meridian syndrome differentiation classification system), the multi-level organization and rapid positioning of knowledge are realized.
[0035] The spatio-temporal index adopts a composite key design, combines the timestamp and the geographical coding to generate a unique partition identifier, and realizes the efficient clustering storage of neighboring data in combination with the topological features of the graph structure.
[0036] S2. Based on the language model and causal reasoning engine, extract entity relationships and dialectical rules from the standardized corpus to generate a dynamic dialectical reasoning model that supports non-linear reasoning.
[0037] In the description of the present invention, S2 includes: S21. Generate a text vector sequence through a pre-trained language model, simulate the ambiguous context of homographs, calculate the semantic deviation degree between the context vector and the topic library, adjust the confidence weight of entity types, and generate an entity attribute matrix with unified identifiers.
[0038] Specifically, for the semantic ambiguity problem of homographs, construct a generative adversarial network framework: the generator simulates the polysemous expressions of terms in different contexts, such as automatically generating different context fragments of "wind" when representing exogenous pathogenic factors and endogenous liver wind; the discriminator analyzes the semantic consistency between the generated text and the ontology library through multi-level feature extraction, quantifies the semantic deviation degree using KL divergence and calculates the confidence decay factor, and dynamically adjusts the weight distribution of entity types. In this process, the generator and the discriminator are continuously optimized through adversarial training, and finally an entity attribute matrix with ambiguity eliminated is output.
[0039] S22. Based on the graph causal diffusion model, encode the generating and restraining relationships as prior conditions of the graph network, and construct an initial graph structure.
[0040] Among them, the present invention transforms the traditional Chinese medicine generating and restraining theory into computable graph structure constraints, and realizes the deep integration of theoretical rules and clinical practice through data-driven dynamic optimization, ensuring that the generated causal graph not only conforms to the core logic of generating and restraining (such as the inhibitory connection of liver wood restraining spleen earth) but also can reflect the statistical law of symptom co-occurrence.
[0041] In the description of the present invention, S22 includes: S221. Convert the generating and restraining theory into graph structure constraint conditions, set the generating relationship between generating and restraining elements as positive weight connection, the restraining relationship as negative weight inhibition, and convert the discrete generating and restraining relationships into a continuous and differentiable probability distribution form to form a generating and restraining constraint matrix.
[0042] Specifically, first transform the traditional Chinese medicine generating and restraining theory into computable graph structure constraint conditions, define the generating relationship between generating and restraining elements as positive weight connection (such as wood generating fire with an initial weight of +0.7) and the restraining relationship as negative weight inhibition (such as wood restraining earth set as -0.5), and perform standardization processing on the discrete generating and restraining relationships to convert them into a continuous and differentiable probability distribution form, generating a generating and restraining constraint matrix as the theoretical prior basis.
[0043] S222. Dynamically and weightedly fuse the generating and restraining constraint matrix with the statistical matrix extracted from the standardized corpus, and map the generating and restraining relationships to the entity topological structure to generate an initial graph structure.
[0044] Specifically, the Shengke constraint matrix is dynamically weighted and fused with the symptom-syndrome co-occurrence statistical matrix extracted from the standardized corpus, and the Shengke relationship is mapped to the entity-level topological structure through tensor contraction operation to form an initial graph structure. In this process, the significance of symptom co-occurrence needs to be screened by chi-square test to ensure that the data-driven association is statistically significant.
[0045] S223. Introduce a Markov chain diffusion process into the initial graph structure, gradually inject Gaussian noise according to the preset noise scheduling parameter, and apply noise suppression to the edge connections that conform to the Shengke theory.
[0046] Specifically, after the initial graph structure is generated, a Markov chain diffusion process is introduced to gradually inject Gaussian noise, and the noise scheduling parameter is set according to the cosine rule. A protection mechanism with a 30% reduction in the noise standard deviation is implemented for the high-confidence edges that conform to the Shengke theory to maintain the stability of the core theory structure.
[0047] S224. Set a theoretical compliance index to continuously monitor the degree of theoretical deviation of the generated graph structure. When the degree of theoretical deviation is greater than the preset threshold, trigger the weight balance mechanism and automatically adjust the fusion coefficient.
[0048] In the description of the present invention, the calculation formula for the degree of theoretical deviation is: ; where C t is the degree of theoretical deviation at time t; W norm (i,j) is the theoretical weight of the edge e i,j in the Shengke constraint matrix; W final (i,j) is the weight of the edge e i,j in the final causal graph; and E is the number of edges.
[0049] In addition, the rationality of the generated graph structure is continuously monitored through the degree of theoretical deviation. This index calculates the average absolute deviation of all edge weights from the theoretical values. When Ct > 0.2, trigger the weight rebalancing mechanism and adjust the fusion coefficient in the direction of the sign function until the system converges to a steady-state balance again. Ensure that the causal graph reaches an optimal balance between theoretical compliance and data fitting error, and support the accurate modeling of complex pathogenesis paths such as "liver wood overacting on spleen earth".
[0050] S23. Establish a reward function according to the Shengke relationship, and use the proximal policy optimization algorithm to train the dialectical strategy in the virtual environment to generate a decision tree for the Shengke-zangfu-treatment chain reasoning.
[0051] In the description of the present invention, S23 includes: S231. Give positive incentives to decisions that conform to the mutual generation path, impose penalty weights on operations that violate the mutual restraint principle, and use historical medical case data to verify the accuracy of efficacy prediction as a benchmark reference value.
[0052] Specifically, the Chinese medicine theory of mutual restraint is quantified into a computable dynamic scoring mechanism. For treatment decisions that conform to the mutual generation path, a positive incentive value of +0.8 is given; for operations that violate the mutual restraint principle, a penalty weight of -0.6 is imposed. At the same time, an efficacy prediction module is introduced to analyze historical medical case data based on a bidirectional LSTM model, output the effectiveness probability value of the treatment plan, and convert it into an efficacy reward item in the interval [-1, +1] through a linear mapping function. The final reward value is the weighted sum of the theoretical compliance score and the efficacy prediction value, where the theoretical rule weight coefficient λ is set to 1.2 to strengthen the guiding role of the mutual generation and restraint logic.
[0053] S232. Construct a dynamic simulation environment and set the action space to select treatment methods and the attributes of mutual generation and mutual restraint.
[0054] Specifically, a dynamic simulation environment is constructed to train syndrome differentiation strategies, and a multidimensional state space is defined, including viscera function indicators and life-giving and restraining balance coefficients, which are combined into a 15-dimensional state vector. The action space is divided into discrete treatment method selection and continuous life-giving and restraining attribute operation. The environment supports the dynamic deduction of pathogenesis transmission paths such as "liver depression and spleen deficiency → soothing the liver and strengthening the spleen → life-giving and restraining state update".
[0055] S233. Use the proximal strategy optimization algorithm to optimize the strategy network and evaluate the long-term benefits of the treatment strategy.
[0056] Specifically, the strategy network receives the current patient status input, including the viscera function index and the life-destruction balance coefficient, and outputs the probability distribution of the treatment method selection and the adjustment strength of its life-destruction attribute. During the training process, the experience replay mechanism is used to store the historical interaction trajectory, and the importance sampling technology is used to compare the action selection differences of the new and old strategies under the same state to calculate the strategy update direction. The generalized advantage estimation method is introduced to integrate the immediate treatment effect and future potential benefits, and the long-term benefit evaluation value is generated by weighted accumulation of multi-step time difference errors.
[0057] S234. Extract high-frequency treatment paths from the convergence strategy network, construct a multi-branch decision tree, and generate an interpretable decision map through decision tree pruning.
[0058] Specifically, the pruning process uses a comprehensive cost-complexity evaluation to remove redundant branches while retaining key decision paths, ensuring the simplicity and clinical interpretability of the tree structure. The resulting decision tree supports chain reasoning from the initial pathogenesis state to the final treatment method. Each internal node is labeled with the judgment logic and confidence of the life-and-death state, and the leaf node is associated with the specific treatment method operation and its expected life-and-death impact intensity, forming a visual diagnosis and treatment map that is both theoretically rigorous and operationally instructive.
[0059] S24. Generate a solution based on the decision tree, construct a dynamic dialectical reasoning model, and output a treatment path containing the suggestions for generating and restraining corrections by solving the optimal solution.
[0060] In the description of the present invention, S24 includes: S241. Construct a game framework including a physician agent, a generating and restraining constraint agent, and a pharmacological verification agent, and set the strategy space and benefit dimension of each agent.
[0061] Specifically, the strategy space of the physician agent is defined as the set of candidate treatment solutions extracted from the decision tree, and its benefit function synthesizes the efficacy prediction value and the safety score; the generating and restraining constraint agent calculates the compliance degree of the generating and restraining rules by traversing the edge weights of the causal graph, and the benefit function is positively correlated with the matching ratio of the generating path and negatively correlated with the conflicting ratio of the restraining; the pharmacological verification agent detects the compatibility taboos based on the pharmacopoeia database, and the benefit function is inversely proportional to the risk level.
[0062] S242. Output a set of treatment solutions that meet the benefits of multiple agents through multi-agent iteration.
[0063] Specifically, encode the strategies of each agent into a mixed strategy vector, search for the equilibrium point through the virtual game iteration algorithm. In each round of iteration, each agent updates its own optimal response according to the opponent's strategy until the change range of the strategy is less than the convergence threshold, and output the equilibrium strategy combination; and design a Pareto front screening mechanism to retain the set of treatment solutions that simultaneously meet the benefits of multiple agents.
[0064] In each round of iteration, the physician agent calculates the expected efficacy benefit according to the current generating and restraining constraints and the pharmacological verification strategy, updates the treatment solution selection probability through the mirror descent method, preferentially improves the efficacy but is restricted by the generating and restraining rule penalty coefficient; the generating and restraining constraint agent adjusts the edge weight attenuation rate based on the generating and restraining influence value in the treatment solution; the pharmacological verification agent dynamically adjusts the veto threshold according to the historical risk data.
[0065] When the three-party strategy vectors meet the convergence conditions, output the equilibrium strategy combination. Subsequently, through the Pareto front screening mechanism, use the non-dominated sorting algorithm to retain the treatment solutions that are non-inferior in the three dimensions of efficacy, rule compliance, and safety from the equilibrium solution set, and eliminate the inferior solutions that are comprehensively surpassed by other solutions.
[0066] S243. Based on the evaluation results of the generating and restraining constraint agent, identify the nodes in the treatment solution that violate the generating and restraining principles, calculate the generating and restraining influence factors, perform topological sorting on the candidate paths, and retain the correction solutions that strengthen generation and inhibit restraint.
[0067] Specifically, based on the evaluation results of the generation-restriction constraints, identify the nodes in the treatment plan that violate the generation-restriction principle, and calculate the generation-restriction influence factor through gradient backpropagation: locate the violation edges in the causal graph, propagate the gradient signal backward along the treatment path, and generate weight adjustment suggestions. Perform topological sorting on the candidate paths, and perform priority sorting according to the promotion ratio of generation and the inhibition rate of restriction, and retain the modified plans that simultaneously meet the requirements of efficacy improvement and theoretical compliance.
[0068] S3. Obtain the output results of the dynamic dialectical reasoning model, and construct a redundant-free and dynamically calibrated traditional Chinese medicine knowledge base through the knowledge integration and conflict resolution of various traditional Chinese medicine classics.
[0069] In the description of the present invention, S3 includes: S31. Align the treatment path with the standardized corpus at the entity level, establish two-way links for the entries with semantic overlap, and retain the original context information of different classics.
[0070] Specifically, perform fine-grained alignment on the treatment path output by the dynamic dialectical reasoning model with the standardized corpus. Construct a term mapping table covering traditional Chinese medicine classics. For each entity, extract its context window features in classical literature and co-occurrence statistical features in modern medical records respectively, calculate the cross-era semantic similarity, establish two-way hyperlinks for the entities with semantic overlap, and encode the context information such as the chapter source and doctor's annotation of the original classic into structured metadata.
[0071] S32. Detect the conflicting entries in different classics, calculate the optimal solution through the Bayesian network, and preferentially retain the solution with the highest clinical verification frequency and meeting the generation-restriction rule constraints.
[0072] Specifically, for the conflicting entries, by constructing a Bayesian network inference framework, define three types of evidence nodes: clinical verification frequency, compliance with the generation-restriction rule, and doctor's authority weight. Use the Markov chain Monte Carlo method to solve the posterior probability distribution. When the confidence level of the optimal solution of the conflicting entry exceeds the preset threshold, automatically retain the solution that meets the conditions of "high verification frequency and deviation degree of generation-restriction constraint", and transfer the remaining conflicting entries to the pending verification queue and trigger the expert review process.
[0073] S33. Perform density clustering on the symptom-syndrome relationship, retain the representative items within the same cluster, and convert the remaining entries into historical versions.
[0074] Specifically, to address the redundancy issue of symptom-syndrome relationships, the DBSCAN density clustering algorithm is used to group symptom association data, construct a symptom similarity matrix based on word vectors, and then calculate the composite similarity by integrating the efficacy dimensions of treatment methods (such as efficacy labels like clearing heat and removing dampness). For redundant entries within the same cluster (such as "clearing heat and detoxifying" and "purging fire and detoxifying"), representative items are selected through a comprehensive weight formula, and version tracing information is retained when the remaining entries are converted to historical versions.
[0075] S34. Deploy lightweight edge computing nodes to collect the improvement degree of physiological indicators of traditional Chinese medicine treatment plans in real time, aggregate encrypted gradients through the federated learning framework, and update the weights of knowledge base entries.
[0076] S34. Build a spatio-temporal sharded graph database storage structure as the traditional Chinese medicine knowledge base, and introduce a composite index structure to achieve full-link response of symptoms-syndromes-treatment methods-plans.
[0077] S4. Based on edge computing and privacy protection technologies, achieve real-time iteration of the knowledge graph and optimization of clinical efficacy feedback.
[0078] In the description of the present invention, S4 includes: S41. Build a central server and establish a communication channel with edge computing nodes through smart contracts.
[0079] Specifically, build a central server as the global coordination center. The server deploys smart contracts based on the blockchain architecture to manage the access and communication processes of edge computing nodes. The central server establishes a trusted connection with each edge computing node through a preset node authentication mechanism.
[0080] S42. Configure local training tasks on edge computing nodes, collect clinical multimodal data in real time for local model training, evaluate the actual efficacy, and dynamically update the traditional Chinese medicine knowledge base using incremental data.
[0081] Specifically, containerization technology is used for the deployment of local training tasks on edge computing nodes. Each node runs an independent Docker instance to encapsulate a lightweight version of the dynamic dialectical reasoning model. Clinical multimodal data is collected in real time through medical Internet of Things devices, including tongue image, pulse waveform, and electronic medical record text, etc.
[0082] The calculation of the actual efficacy index integrates short-term indicators (such as symptom remission rate) and long-term indicators (such as three-month recurrence rate), and generates a 0-1 standardized efficacy score through fuzzy comprehensive evaluation.
[0083] S43. Use encryption technology to encrypt the collection and calculation to achieve privacy data protection.
[0084] Please refer to Figure 2, a knowledge graph construction system based on traditional Chinese medicine classics is also provided, including: A data collection module 1, which is used to collect multi-source heterogeneous data of traditional Chinese medicine classics for standardization processing, and generate a cross-era aligned standardized corpus through optical character recognition and multi-modal feature extraction.
[0085] An inference and dialectics module 2, which is used to extract entity relationships and dialectics rules from the standardized corpus based on a language model and a causal inference engine, and generate a dynamic dialectics inference model that supports non-linear inference.
[0086] A fusion and calibration module 3, which is used to obtain the output results of the dynamic dialectics inference model, and construct a redundant-free and dynamically calibrated traditional Chinese medicine knowledge base through knowledge fusion and conflict resolution of various traditional Chinese medicine classics.
[0087] An update and feedback module 4, which is used to realize the real-time iteration of the knowledge graph and the optimization of clinical efficacy feedback based on edge computing and privacy protection technologies.
[0088] In summary, by means of the above technical solutions of the present invention, through the standardization processing of multi-source heterogeneous data and the dynamic knowledge fusion mechanism, the problems existing in the digitalization process of traditional Chinese medicine classics, such as cross-era evolution, regional expression differences, and multi-modal feature fragmentation, are solved. Deep semantic alignment is performed on text, images, and waveform data to construct a knowledge base system for real-time retrieval, which not only completely retains the classic theoretical framework of traditional Chinese medicine classics but also organically integrates the characteristics of modern clinical diagnosis and treatment data, forming a dynamic knowledge network that connects the past and the present and mutually verifies multiple sources, realizing the continuous evolution of the knowledge system while protecting the characteristics of academic schools, and providing an extensible intelligent infrastructure for the modernization of traditional Chinese medicine.
[0089] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
Claims
1. A method for constructing a knowledge graph based on traditional Chinese medicine classics, characterized in that: include: S1. Collect multi-source heterogeneous data of traditional Chinese medicine classics for standardization processing, and generate a standardized corpus aligned across eras through optical character recognition and multimodal feature extraction; S2, based on the language model and causal reasoning engine, extract entity relationships and dialectical rules from the standardized corpus to generate a dynamic dialectical reasoning model that supports nonlinear reasoning; S3. Obtain the output results of the dynamic dialectical reasoning model, and build a redundant and dynamically calibrated TCM knowledge base through knowledge integration and conflict resolution of various TCM classics; S4. Based on edge computing and privacy protection technology, real-time iteration of knowledge graphs and optimization of clinical efficacy feedback are achieved.
2. A method for constructing a knowledge graph based on traditional Chinese medicine classics according to claim 1, characterized in that: The multi-source heterogeneous data of traditional Chinese medicine classics are collected and standardized, and a standardized corpus aligned across eras is generated through optical character recognition and multimodal feature extraction, including: S11, extracting multi-source heterogeneous data from the ancient book digitization platform and the hospital management information system, and performing grayscale correction, ink enhancement and non-text area filtering processing according to the data type, wherein the multi-source heterogeneous data includes ancient book images, modern medical case data and pharmacopoeia database information; S12. Deploy an optical character recognition engine, integrate the ancient Chinese character library and domain dictionary, identify homophones and variant characters, establish a dynamic mapping chart, and output standardized text; S13, respectively extracting visual features of image data and time-frequency features of waveform data from multi-source heterogeneous data to obtain feature vectors of each component; S14, automatically align terms in standardized text based on cosine similarity algorithm; S15. Use graph data to store the multi-dimensional relationship between classic content, physical sign data, and treatment methods and prescriptions, set up a spatiotemporal index structure, and realize distributed storage of processed and standardized text data.
3. A method for constructing a knowledge graph based on traditional Chinese medicine classics according to claim 1, characterized in that: The method of extracting entity relationships and dialectical rules from a standardized corpus based on a language model and a causal reasoning engine to generate a dynamic dialectical reasoning model supporting nonlinear reasoning includes: S21. Generate a text vector sequence through a pre-trained language model, simulate the ambiguous context of homographs, calculate the semantic deviation between the context vector and the question bank, adjust the entity type confidence weight, and generate an entity attribute matrix of a unified identifier; S22. Based on the graph causal diffusion model, the generation-restraint relationship is encoded as a priori conditions of the graph network to construct the initial graph structure; S23, establish a reward function based on the generation-control relationship, use the proximal strategy optimization algorithm to train the dialectical strategy in a virtual environment, and generate a decision tree of the generation-control-organ-treatment chain reasoning; S24. Based on the decision tree generation scheme, a dynamic dialectical reasoning model is constructed to output a treatment path containing suggestions for synergy and correction by solving the optimal solution.
4. A method for constructing a knowledge graph based on TCM classics according to claim 3, characterized in that: Based on the graph causal diffusion model, the generation-restraint relationship is encoded as a priori conditions of the graph network, and the initial graph structure is constructed, including: S221, converting the generation and restraint theory into graph structure constraints, setting the generation relationship between generation and restraint elements as positive weight connection and the restraint relationship as negative weight inhibition, and converting the discrete generation and restraint relationship into a continuous and differentiable probability distribution form to form a generation and restraint constraint matrix; S222, dynamically weighted fusion of the generative constraint matrix and the statistical matrix extracted from the standardized corpus, and mapping the generative constraint relationship to the entity topological structure to generate an initial graph structure; S223, introducing a Markov chain diffusion process into the initial graph structure, gradually injecting Gaussian noise according to preset noise scheduling parameters, and applying noise suppression to edge connections that conform to the Shengke theory; S224. Set the theoretical conformity index to continuously monitor the theoretical deviation of the generated graph structure. When the theoretical deviation is greater than the preset threshold, the weight balancing mechanism is triggered to automatically adjust the fusion coefficient.
5. The method for constructing a knowledge graph based on traditional Chinese medicine classics according to claim 3, characterized in that: The method of establishing a reward function according to the generation-restraint relationship, using a proximal strategy optimization algorithm to train a dialectical strategy in a virtual environment, and outputting a decision tree of the generation-restraint-organ-treatment chain reasoning includes: S231. Give positive incentives to decisions that conform to the mutual generation path, impose penalty weights on operations that violate the mutual restraint principle, and verify the accuracy of efficacy prediction based on historical medical case data as a benchmark reference value; S232, construct a dynamic simulation environment, and set the action space as the treatment method selection and the generation and restraint attributes; S233, optimize the strategy network using the proximal strategy optimization algorithm and evaluate the long-term benefits of the treatment strategy; S234. Extract high-frequency treatment paths from the convergence strategy network, construct a multi-branch decision tree, and generate an interpretable decision map through decision tree pruning.
6. A method for constructing a knowledge graph based on TCM classics according to claim 3, characterized in that: The decision tree-based generation scheme constructs a dynamic dialectical reasoning model, and by solving the optimal solution, outputs a treatment path containing the Shengke correction suggestions, including: S241. Construct a game framework including physician agents, life-control constraint agents and pharmacological verification agents, and set the strategy space and benefit dimension of each agent; S242. Output a set of treatment plans that satisfy the benefits of the multi-agents through multi-agent iteration; S243. Based on the evaluation results of the mutual generation and mutual restraint constraint agent, identify the nodes in the treatment plan that violate the mutual generation and mutual restraint principle, calculate the mutual generation and mutual restraint influencing factors, topologically sort the candidate paths, and retain the correction plans of mutual generation reinforcement and mutual restraint inhibition.
7. A method for constructing a knowledge graph based on TCM classics according to claim 6, characterized in that: The treatment plan set outputted through multi-agent iteration that satisfies the multi-agent benefits includes: The strategy of each agent is encoded as a mixed strategy vector, and the equilibrium point is searched through a virtual game iterative algorithm. In each round of iteration, each agent updates its optimal response according to the opponent's strategy until the strategy change is less than the convergence threshold, and then outputs the equilibrium strategy combination; and a Pareto frontier screening mechanism is designed to retain a set of treatment plans that simultaneously meet the benefits of multiple agents.
8. The method for constructing a knowledge graph based on traditional Chinese medicine classics according to claim 1, characterized in that: The output results of the dynamic dialectical reasoning model are obtained, and a redundant and dynamically calibrated TCM knowledge base is constructed through knowledge integration and conflict resolution of various TCM classics, including: S31. Align the treatment pathways with the standardized corpus at the entity level and establish bidirectional links for items with semantic overlap, preserving the original contextual information of different classics; S32. Detect conflicting entries in different classics, calculate the optimal solution through Bayesian network, and give priority to the solution with the highest clinical verification frequency and in compliance with the constraints of the Shengke rule; S33. Density clustering of symptom-syndrome relationships, retaining representative items within the same cluster, and converting the remaining items to historical versions; S34. Build a spatiotemporal sharded graph database storage structure as a TCM knowledge base, and introduce a composite index structure to achieve a full-link response of symptoms-syndromes-treatment methods-solutions.
9. The method for constructing a knowledge graph based on traditional Chinese medicine classics according to claim 1, characterized in that: The real-time iteration of the knowledge graph and optimization of clinical efficacy feedback based on edge computing and privacy protection technology include: S41. Build a central server and establish a communication channel with the edge computing node through smart contracts; S42. Configure local training tasks on edge computing nodes, collect clinical multimodal data in real time for local model training, evaluate actual efficacy, and dynamically update the TCM knowledge base using incremental data; S43. Use encryption technology to encrypt collected calculations to achieve privacy data protection.
10. A knowledge graph construction system based on TCM classics, used to implement the knowledge graph construction method based on TCM classics as claimed in any one of claims 1 to 9, characterized in that: include: The data collection module is used to collect multi-source heterogeneous data of traditional Chinese medicine classics for standardized processing, and generate a standardized corpus aligned across eras through optical character recognition and multimodal feature extraction; The reasoning and dialectic module is used to extract entity relationships and dialectical rules from the standardized corpus based on the language model and causal reasoning engine, and generate a dynamic dialectical reasoning model that supports nonlinear reasoning; The fusion and calibration module is used to obtain the output results of the dynamic dialectical reasoning model, and to build a redundant and dynamically calibrated TCM knowledge base through the knowledge fusion and conflict resolution of various TCM classics; Update the feedback module to achieve real-time iteration of the knowledge graph and optimization of clinical efficacy feedback based on edge computing and privacy protection technology.
Citation Information
Patent Citations
Construction method and system of traditional Chinese medicine knowledge graph, and storage medium
CN116049426A
Traditional Chinese medicine classical book knowledge base feedback correction method and system
CN117271796A
Traditional Chinese medicine knowledge graph construction method based on large language model
CN117493579A
Intelligent traditional Chinese medicine auxiliary diagnosis system based on knowledge graph retrieval enhanced generation
CN119170258A
Medical attribute knowledge graph construction method and apparatus, and device and medium
WO2021159733A1
Cited By
Project decision optimization control method, device and equipment based on knowledge graph
CN120509686A
Traditional Chinese medicine knowledge graph fusion system based on semantic alignment
CN120600343A
Intelligent report generation method and system
CN120893414A
Traditional Chinese medicine whole industry chain big data construction method and system
CN120952701A
Knowledge base dynamic construction method and system based on large model
CN121029732A