Traditional Chinese medicine electronic medical record natural language processing and knowledge graph construction method and system

By dividing and deeply analyzing multi-granular semantic units, a semantic parsing framework for TCM texts is constructed, which solves the problems of structuring and individualizing TCM electronic medical records, realizes the accurate processing of TCM diagnosis and treatment information and intelligent decision-making assistance, and improves the overall level of TCM systems.

CN121168618APending Publication Date: 2025-12-19SUZHOU TRADITIONAL CHINESE MEDICINE HOSPITAL
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511241025.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Traditional Chinese medicine (TCM) electronic medical records are mostly unstructured texts. Existing technologies struggle to accurately identify the unique semantics and logic of TCM. Traditional semantic analysis cannot form a complete chain of association between symptoms, syndrome types, and treatment methods. TCM knowledge graphs lack dynamic diagnostic and treatment processes and characterization of individual differences. Syndrome differentiation rules rely on manual summarization, which is inefficient. Existing intelligent systems lack dynamic data and individual support, resulting in low accuracy and insufficient practicality in syndrome differentiation.

Method used

By dividing and deeply analyzing multi-granular semantic units, a semantic parsing framework for TCM texts is constructed. Multi-dimensional semantic features are extracted and semantic conflicts are resolved. Explicit diagnostic rules and implicit reasoning paths are explored. A dynamic knowledge graph is constructed by combining medical record time series data. Individualized cognitive feature analysis and the inheritance of famous doctors' experience are carried out to form an intelligent system.

Benefits of technology

It has achieved the structured transformation of TCM electronic medical records, improved the ability to understand and process diagnostic and treatment information, enhanced the accuracy of the application of syndrome differentiation theory and the individualized reflection of the diagnosis and treatment process, promoted the precision and intelligence of TCM diagnosis and treatment, and improved the ability to inherit the experience of famous doctors and optimize the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121168618A_ABST
    Figure CN121168618A_ABST
Patent Text Reader

Abstract

The invention provides a traditional Chinese medicine electronic medical record natural language processing and knowledge graph construction method and system. The method belongs to the technical field of the crossing field of traditional Chinese medicine informatization, artificial intelligence natural language processing and knowledge engineering, and comprises the following steps: performing multi-granularity semantic unit division on the traditional Chinese medicine electronic medical record to generate traditional Chinese medicine text semantic unit data; constructing a multi-granularity semantic decoupling engine according to the traditional Chinese medicine text semantic unit data so as to construct a traditional Chinese medicine text semantic analysis framework; performing traditional Chinese medicine unstructured text deep analysis according to the traditional Chinese medicine text semantic analysis framework, and performing multi-dimensional semantic feature extraction to obtain traditional Chinese medicine text semantic feature data and medical record time sequence data; through multi-granularity semantic unit division and deep analysis, the unstructured text in the traditional Chinese medicine electronic medical record can be effectively converted into structured semantic data, and the ability to understand and process traditional Chinese medicine clinical information is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application provides a traditional Chinese medicine electronic medical record natural language processing and knowledge graph construction method and system, and belongs to the technical field of the cross field of traditional Chinese medicine informatization, artificial intelligence natural language processing and knowledge engineering. BACKGROUND

[0002] With the promotion of traditional Chinese medicine diagnosis and treatment informatization, electronic medical records have become the core carrier of traditional Chinese medicine clinical data, but there are still many bottlenecks in current processing and knowledge application: first, the traditional Chinese medicine electronic medical record is mostly unstructured text, and the term expression is heterogeneous (such as waist pain and knee pain, which refer to the same disease), so it is difficult for general natural language processing tools to accurately identify the specific semantics and logic of traditional Chinese medicine, and it is difficult to convert data into structured information; second, traditional semantic analysis only stays at a single granularity, and cannot form a complete correlation chain of symptoms, syndromes, treatment methods and prescriptions, which hinders the mining of syndrome differentiation rules; third, existing traditional Chinese medicine knowledge graphs are mostly static and focus on the commonness of groups, and lack of dynamic diagnosis and treatment process and individual difference description, which does not meet the needs of syndrome differentiation and treatment; fourth, the syndrome differentiation rules rely on artificial summary and fragmentation of famous doctor experience, and the efficiency of mining and inheritance is low; fifth, existing traditional Chinese medicine intelligent systems rely on static rules, lack of dynamic data and individual support, and have low accuracy and insufficient practicality in syndrome differentiation.

[0003] In summary, an integrated method is urgently needed to solve the problems of traditional Chinese medicine text structuring, semantic depth analysis, dynamic knowledge graph construction, individualized feature mining and reuse of famous doctor experience, and to break through the technical bottleneck to promote the intelligent development of traditional Chinese medicine diagnosis and treatment. SUMMARY

[0004] The application provides a traditional Chinese medicine electronic medical record natural language processing and knowledge graph construction method and system to solve the problems mentioned in the background art:

[0005] The traditional Chinese medicine electronic medical record natural language processing and knowledge graph construction method provided by the application comprises the following steps:

[0006] S1: performing multi-granularity semantic unit division on the traditional Chinese medicine electronic medical record to generate traditional Chinese medicine text semantic unit data; constructing a multi-granularity semantic decoupling engine according to the traditional Chinese medicine text semantic unit data, and thereby constructing a traditional Chinese medicine text semantic analysis framework;

[0007] S2: performing deep analysis on the traditional Chinese medicine unstructured text according to the traditional Chinese medicine text semantic analysis framework, and performing multi-dimensional semantic feature extraction to obtain traditional Chinese medicine text semantic feature data and medical record time sequence data; performing semantic conflict resolution processing on the traditional Chinese medicine text semantic feature data to generate standardized semantic feature data;

[0008] S3: According to the specification semantic feature data, the syndrome differentiation rule mining is carried out, and explicit syndrome differentiation rule data and implicit reasoning path data are obtained respectively; the rule credibility of the implicit reasoning path data is evaluated through the explicit syndrome differentiation rule data, and the reasoning path is optimized, and the optimized reasoning path data is obtained;

[0009] S4: The dynamic knowledge extraction processing is carried out on the medical record time sequence data to the traditional Chinese medicine diagnosis and treatment process, and the traditional Chinese medicine diagnosis and treatment dynamic knowledge data are generated; the knowledge graph is constructed through the optimized reasoning path data to the traditional Chinese medicine diagnosis and treatment dynamic knowledge data, and the dynamic knowledge graph containing the syndrome differentiation rule and the reasoning path is generated; the individualized cognitive feature data is generated through the individualized cognitive feature analysis of the dynamic knowledge graph;

[0010] S5: The individualized cognitive evolution model is constructed according to the individualized cognitive feature data, and the individualized cognitive evolution model is generated; the traditional Chinese medicine experience inheritance data is generated through the experience inheritance processing of famous doctors according to the individualized cognitive evolution model; and the system evolution optimization is carried out based on the traditional Chinese medicine experience inheritance data and the dynamic knowledge graph, and finally the traditional Chinese medicine intelligent system is formed.

[0011] The traditional Chinese medicine electronic medical record natural language processing and knowledge graph construction system provided by the application comprises:

[0012] One or more processors;

[0013] A memory for storing one or more programs,

[0014] When the one or more programs are executed by the one or more processors, the one or more processors realize the method described in any one of the above.

[0015] The application has the following advantages:

[0016] 1. Through multi-granularity semantic unit division and deep analysis, the unstructured text in the traditional Chinese medicine electronic medical record can be effectively converted into structured semantic data, and the understanding and processing ability of the traditional Chinese medicine clinical information is further improved.

[0017] 2. Through the combination of explicit syndrome differentiation rule data and implicit reasoning path data, not only the application of the traditional Chinese medicine syndrome differentiation theory can be enhanced, but also the reasoning accuracy in the diagnosis and treatment process can be improved through the optimization of the reasoning path.

[0018] 3. The medical record time sequence data is combined with the optimized reasoning path to construct a dynamic knowledge graph, which can accurately reflect the diagnosis and treatment process, individualized characteristics and evolution of diagnosis and treatment decision of the patient, and provide a more accurate treatment scheme.

[0019] 4. Based on individual cognitive feature data, a cognitive evolution model is constructed, which further promotes the accurate prediction of patient conditions and personalized treatment. For the inheritance and systematic optimization of famous doctor experience, it helps to improve the overall level of traditional Chinese medicine diagnosis and treatment.

[0020] 5. Through the combination of traditional Chinese medicine experience inheritance and dynamic knowledge graph, an intelligent traditional Chinese medicine system is finally formed, which provides accurate and intelligent auxiliary decision-making for clinical practice, and further promotes the combination and innovation of traditional Chinese medicine theory and practice.

[0021] 6. The method can make up for some deficiencies in traditional Chinese medicine diagnosis and treatment, realize knowledge inheritance, dynamic optimization and intelligent processing of traditional Chinese medicine through technical means, so that traditional Chinese medicine can play a greater role in the modern medical system. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 The method steps of the present application are shown in the figure. DETAILED DESCRIPTION

[0023] The preferred embodiments of the present application are described below in conjunction with the accompanying drawings, and it should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.

[0024] An embodiment of the present application is shown in Figure 1 The traditional Chinese medicine electronic medical record natural language processing and knowledge graph construction method comprises the following steps:

[0025] S1: Multi-granularity semantic unit division is performed on the traditional Chinese medicine electronic medical record to generate traditional Chinese medicine text semantic unit data; a multi-granularity semantic decoupling engine is constructed according to the traditional Chinese medicine text semantic unit data, so as to construct a traditional Chinese medicine text semantic analysis framework;

[0026] S2: Deep analysis of traditional Chinese medicine unstructured text is performed according to the traditional Chinese medicine text semantic analysis framework, and multi-dimensional semantic feature extraction is performed, so as to obtain traditional Chinese medicine text semantic feature data and medical record time sequence data; semantic conflict resolution processing is performed on the traditional Chinese medicine text semantic feature data to generate standardized semantic feature data;

[0027] S3: Syndrome differentiation rule mining is performed according to the standardized semantic feature data, and explicit syndrome differentiation rule data and implicit reasoning path data are obtained respectively; rule credibility evaluation is performed on the implicit reasoning path data through the explicit syndrome differentiation rule data, and reasoning path optimization is performed, so as to obtain optimized reasoning path data;

[0028] S4: dynamically extract knowledge from the medical record time series data to generate TCM diagnosis and treatment dynamic knowledge data; construct a knowledge graph based on the optimized reasoning path data to generate a dynamic knowledge graph containing diagnosis rules and reasoning paths; analyze the individualized cognitive characteristics of the dynamic knowledge graph to generate individualized cognitive characteristic data;

[0029] S5: build an individualized cognitive evolution model based on the individualized cognitive characteristic data to generate an individualized cognitive evolution model; process the inheritance of famous doctor's experience based on the individualized cognitive evolution model to generate TCM experience inheritance data; and optimize the system evolution based on the TCM experience inheritance data and the dynamic knowledge graph to ultimately form a TCM intelligent system.

[0030] The working principle and effects of the above technical solution are as follows: through multi-granularity semantic unit division and semantic decoupling engine, the TCM characteristic terms and complex semantic logic in the medical record can be accurately identified, avoiding the problems of term confusion and semantic loss in traditional processing methods, making the conversion of medical record text more in line with the actual TCM diagnosis and treatment, and improving the accuracy of TCM electronic medical record semantic analysis; in the past, unstructured medical records needed to be manually annotated and organized sentence by sentence, but this method reduces the manual intervention through automatic deep analysis and feature extraction, especially in large-scale medical record processing scenarios, which can significantly reduce the labor input and time consumption, and reduce the labor cost of TCM unstructured text processing; the implicit reasoning path is evaluated and optimized for reliability by the explicit diagnosis rule, eliminating the low clinical adaptability and illogical reasoning link, making the extracted diagnosis rules more in line with the TCM clinical diagnosis and treatment rules, providing more reliable basis for subsequent diagnosis and treatment assistance, and enhancing the reliability of the diagnosis rules and reasoning paths; in the process of dynamic knowledge graph construction, not only the entity attributes are supplemented, but also the relationship weights between entities are optimized, and the explicit diagnosis rules are integrated, avoiding the problems of traditional knowledge graph, such as heavy entity, light association, multiple redundancy, and few core, making the knowledge graph more complete and practical, and reducing the information loss and redundancy in knowledge graph construction; based on the extraction of dynamic diagnosis knowledge from medical record time series data and the analysis of individualized cognitive characteristics, the patient's condition evolution and scheme adjustment process from initial diagnosis to follow-up can be intuitively presented, breaking the limitations of traditional static knowledge presentation that cannot adapt to individual differences, and improving the dynamic and individualized presentation ability of TCM diagnosis and treatment knowledge; based on the system evolution of famous doctor's experience data and dynamic knowledge graph, the system's diagnosis accuracy and treatment recommendation rationality can be gradually improved by updating the knowledge content and rule weight without making radical adjustments to the system framework, reducing the technical threshold of long-term optimization of the system; the scattered medical record data is converted into structured semantic features, dynamic knowledge graphs and other reusable resources, which can be used in new drug research and development, diagnosis and treatment standard formulation, TCM teaching and other scenarios, making the originally one-time use of medical record data have long-term and multi-dimensional value, and improving the reuse value of TCM diagnosis and treatment data.

[0031] In one embodiment of the present invention, S1 includes:

[0032] S11: Obtain TCM electronic medical record text data from different sources, including outpatient medical records, inpatient medical records, and follow-up records, to form an original TCM electronic medical record dataset.

[0033] S12: Preprocess the original TCM electronic medical record dataset to generate preprocessed TCM electronic medical record text;

[0034] S13: Based on the preprocessed TCM electronic medical record text, semantic units are divided according to the multi-granularity levels of characters, words, phrases, sentences and paragraphs to generate TCM text semantic unit data;

[0035] S14: Based on the semantic unit data of TCM texts, combined with the TCM terminology system and natural language processing algorithms, a multi-granularity semantic decoupling engine is constructed based on semantic decoupling rules and processes.

[0036] S15: By integrating semantic parsing logic, domain knowledge association, and dynamic update mechanism through a multi-granularity semantic decoupling engine, a semantic parsing framework for TCM texts is constructed.

[0037] The working principle and effects of the above technical solution are as follows: By collecting medical record texts from different scenarios such as outpatient, inpatient, and follow-up visits, it covers the entire treatment cycle of patients from initial diagnosis to rehabilitation tracking, avoiding the problem of fragmented medical record information from a single source. This makes the raw dataset for subsequent processing more complete and more in line with the actual treatment process, improving the comprehensiveness of TCM electronic medical record data sources. The preprocessing stage can remove invalid characters, correct typos, and unify the format in medical records, reducing the interference of messy information on semantic parsing. The multi-granularity hierarchical division by characters, words, phrases, sentences, and paragraphs can accurately capture core terms such as Astragalus membranaceus and blood stasis (word level), and understand the complete semantics of Astragalus membranaceus 15g decoction (phrase level) and patient chest pain due to blood stasis (sentence level). This avoids the problem of traditional single-granularity division either missing key information or failing to grasp the core logic, enhancing the hierarchical sense and accuracy of TCM text semantic division. When constructing the multi-granularity semantic decoupling engine, it combines the TCM terminology system, such as for the complex cold and heat syndrome. This approach employs proprietary parsing rules for unique TCM expressions such as "superficial and substantial," resolving the difficulty of general NLP algorithms in accurately recognizing TCM-specific semantics. This makes semantic processing more aligned with TCM professional logic, reducing the compatibility issues between TCM terminology and general natural language processing. Integrating semantic parsing logic, domain knowledge association, and a dynamic update mechanism, it not only accurately parses existing medical records but also allows for rule supplementation through dynamic updates when encountering new TCM terms (such as emerging disease descriptions or improved prescription names). This eliminates the need for frequent framework reconstruction, reducing long-term maintenance costs and improving the practicality and scalability of the semantic parsing framework. Furthermore, by pre-constructing a standardized semantic parsing framework, subsequent processing of new unstructured TCM text eliminates the need to repeatedly build the parsing process; the framework can be directly invoked to complete entity recognition, relation extraction, and other operations. This improves processing efficiency and avoids inconsistencies in parsing results caused by inconsistent processes, enhancing the efficiency and stability of subsequent semantic parsing.

[0038] In one embodiment of the present invention, S13 includes:

[0039] S131: Based on the preprocessed TCM electronic medical record text, a TCM terminology word segmentation tool is used to scan the text character by character, identify and extract individual TCM characters and special symbols, and generate TCM text character-level semantic unit data.

[0040] S132: Based on character-level semantic unit data, combined with a vocabulary list in the field of traditional Chinese medicine and a bidirectional maximum matching algorithm, word combination recognition is performed on continuous characters to extract traditional Chinese medicine vocabulary units. The traditional Chinese medicine vocabulary units include symptom words, Chinese medicine names and prescription names, and generate traditional Chinese medicine text character-level semantic unit data.

[0041] S133: Based on word-level semantic unit data, through contextual semantic association analysis, words with close logical relationships are combined into phrase units, such as symptoms + degree, traditional Chinese medicine + dosage, to generate phrase-level semantic unit data for traditional Chinese medicine texts;

[0042] S134: Based on phrase-level semantic unit data, using punctuation marks as separators, and combining the characteristics of TCM sentence expression (e.g., superficial deficiency and internal excess, cold and heat mixed, where superficial deficiency and internal excess are opposing concepts, while cold and heat mixed is a comparison of two symptoms); continuous phrases are integrated into complete TCM diagnosis and treatment description statements to generate TCM text sentence-level semantic unit data.

[0043] S135: Based on sentence-level semantic unit data, sentences related to the same diagnosis and treatment stage or the same disease are clustered and integrated according to the content theme relevance to form paragraph units containing complete diagnosis and treatment information, thus generating TCM text paragraph-level semantic unit data.

[0044] S136: Integrate semantic unit data at the character, word, phrase, sentence, and paragraph levels, establish a multi-granularity semantic unit association index, and generate semantic unit data of TCM texts containing hierarchical relationships and semantic mappings.

[0045] The working principle and effects of the above technical solution are as follows: By using a traditional Chinese medicine term segmentation tool to scan and extract character-level units word by word, it can accurately capture unique traditional Chinese medicine characters such as "stasis", "phlegm", "astragalus", and special symbols commonly used in medical records such as "--" and "()", avoiding the problems of missing recognition of rare traditional Chinese medicine characters and incorrect deletion of key symbols by general character extraction tools, laying a precise foundation for subsequent semantic combination, and improving the accuracy of basic semantic recognition of traditional Chinese medicine texts; Combining the traditional Chinese medicine domain word list with the bidirectional maximum matching algorithm to extract word-level units, for example, it can accurately distinguish "Angelica sinensis" (the name of a traditional Chinese medicine) from "Angelica sinensis still" (a sentence component), and "Guizhi Decoction" (the name of a prescription) from "Ramulus Cinnamomi" (a single traditional Chinese medicine), reducing the situation of incorrect splitting and misidentification of traditional Chinese medicine terms by traditional segmentation algorithms, making the extraction of symptom words, traditional Chinese medicine names, and prescription names more in line with clinical practice, and reducing the error rate of traditional Chinese medicine core vocabulary recognition; By analyzing and combining phrase-level units through context, it can combine "headache" with "paroxysmal" to form "paroxysmal headache", and "Poria cocos" with "10g" to form "Poria cocos 10g", avoiding the problem of being unable to reflect key associated information such as symptom severity and traditional Chinese medicine dosage when extracting words alone, making semantic units closer to the expression habits of traditional Chinese medicine diagnosis and treatment, and enhancing the relevance and integrity of traditional Chinese medicine text semantics; Combining the expression characteristics of opposition and contrast in traditional Chinese medicine such as exterior deficiency and interior excess, cold and heat complexity to integrate sentence-level units, it can accurately sort out the logical relationship of complex sentences such as the patient presenting both exterior deficiency with spontaneous sweating and interior excess with constipation, avoiding the situation of chaotic logic splitting of unique traditional Chinese medicine syndrome differentiation expressions when splitting according to general punctuation marks, making the diagnosis and treatment description sentences more coherent, and reducing the logical discontinuity of traditional Chinese medicine sentence integration; Clustering to generate paragraph-level units according to the diagnosis and treatment stage (such as the first diagnosis and follow-up visit) or disease theme (such as cough and dizziness), it can integrate the description of symptoms, the conclusion of syndrome differentiation, and the prescription of traditional Chinese medicine at the first diagnosis scattered in different positions into the same paragraph, without having to search for associated information sentence by sentence from the whole medical record, making the acquisition of diagnosis and treatment information more efficient, and improving the aggregation efficiency of traditional Chinese medicine diagnosis and treatment information; Establishing an associated index to integrate semantic units at all levels. When different granularity information needs to be called later, for example, when wanting to check the character composition (character level), corresponding prescription (word level), and specific dosage (phrase level) of a certain traditional Chinese medicine, it can be quickly located through the index, avoiding the query difficulty caused by the scattered storage of units at all levels. At the same time, it can also trace the semantic evolution logic from characters to paragraphs, facilitating subsequent verification and adjustment, and enhancing the reusability and traceability of multi-granularity semantic units; Splitting and integrating semantic units in advance according to multiple granularities. When constructing a semantic decoupling engine and parsing the text later, there is no need to process the original text from scratch, and the already divided units at all levels can be directly called, reducing the repeated processing links, and also making the subsequent parsing process more focused on the core semantic association, reducing the technical implementation difficulty.

[0046] In one embodiment of the present invention, the S132 includes:

[0047] The semantic unit data of TCM text at the character level is initially matched with the constructed TCM domain lexicon, and the character sequences containing core characters of the TCM domain lexicon (such as stasis, phlegm, ginseng, and decoction) in the character-level units are selected to generate a set of character sequences to be identified.

[0048] Based on the set of character sequences to be identified, a bidirectional maximum matching algorithm is used to scan the character sequence from left to right. The character combinations are truncated according to the length of the longest word in the Chinese medicine vocabulary list. It is then determined whether they match Chinese medicine words in the vocabulary list to generate a set of left-side matching candidate words.

[0049] For the same set of character sequences to be identified, the bidirectional maximum matching algorithm is used to scan the character sequence from right to left. Similarly, the character combination is truncated according to the longest word length and compared with the vocabulary of traditional Chinese medicine to generate a set of right-side matching candidate words.

[0050] By comparing the left-side matching candidate vocabulary set and the right-side matching candidate vocabulary set, the matching length, vocabulary coverage and semantic rationality of the vocabulary in the two candidate sets are statistically analyzed. Vocabulary with longer matching length, higher coverage and conformity to the semantic logic of traditional Chinese medicine is selected first to generate a preliminary set of traditional Chinese medicine vocabulary units.

[0051] The initial set of TCM vocabulary units is classified and labeled. Based on the TCM vocabulary classification system, the vocabulary is labeled as symptom words (e.g., soreness and weakness of the waist and knees), Chinese medicine names (e.g., Angelica sinensis), and prescription names (e.g., Cinnamon Twig Decoction), thus generating a classified and labeled set of TCM vocabulary units.

[0052] Redundancy checks are performed on the TCM vocabulary unit set after classification and labeling, duplicate labeled vocabulary units are removed, and low-frequency words that are not recognized by the bidirectional maximum matching algorithm but conform to TCM semantics (such as zhengjia and laoli) are added to finally generate TCM text word-level semantic unit data.

[0053] The working principle and effect of the above technical solution are as follows: First, sequences containing core characters such as stasis, phlegm, ginseng, and decoction are screened, which can quickly identify text fragments that may contain TCM terms, avoiding blindly searching for words among a large number of non-core characters. This allows subsequent matching to focus more on TCM-related content, reducing wasted effort and improving the targeting of TCM term recognition. The bidirectional maximum matching algorithm scans and matches from both left and right directions. For example, when encountering words like "Dang Gui Wei" (Angelica sinensis tail), left scanning can match according to the longest word "Dang Gui Wei," and right scanning can avoid splitting it into "Dang" and "Gui Wei." Compared with unidirectional matching, this greatly reduces the chance of completely matching the end word. Breaking down traditional Chinese medicine (TCM) terminology into smaller parts allows for more complete vocabulary extraction and reduces the probability of errors in TCM terminology segmentation. Vocabulary selection is based on comparison of matching length, vocabulary coverage, and semantic logic. For example, when considering modifications to Guizhi Tang (Cinnamon Twig Decoction), the system prioritizes longer matches that better align with the semantics of the formula name and adjustment method, rather than simply selecting Guizhi (Cinnamon Twig). This avoids semantic bias caused by selecting words solely based on length, ensuring that the selected vocabulary is more relevant to the TCM diagnostic context and enhancing the rationality of TCM vocabulary selection. Symptom words, herbal names, and formula names are labeled according to a domain-specific vocabulary system; for example, "lower back and knee pain" is explicitly labeled as a symptom. Dang Gui (Angelica sinensis) is labeled as a Chinese herbal medicine, and Gui Zhi Tang (Cinnamon Twig Decoction) is labeled as a prescription. This allows for rapid differentiation of different types of terms during subsequent processing, eliminating the need for repeated attribute checks and avoiding subsequent parsing confusion caused by fuzzy classification. It also reduces the complexity of TCM terminology classification. During redundancy checks, low-frequency but important terms like "zhengjia" (abdominal masses) and "luoli" (scrofula) are added, preventing the algorithm from only recognizing high-frequency words and missing less common TCM terms. This allows some uncommon but clinically crucial terms to be included in word-level units, covering a more comprehensive range of TCM diagnostic and treatment scenarios and improving the recognition rate of low-frequency TCM terms. The final generated word-level data undergoes multiple rounds of processing. The accuracy and completeness of vocabulary are guaranteed through screening, matching, and verification. When building a semantic decoupling engine and parsing text, it is no longer necessary to spend a lot of effort correcting vocabulary errors, reducing rework, improving overall processing efficiency, and lowering the error correction cost of subsequent semantic processing. The classified and labeled vocabulary can be directly connected to subsequent semantic feature extraction. For example, when extracting symptom features, the labeled symptom vocabulary set can be directly called without having to screen from a massive vocabulary. This allows word-level data to quickly empower downstream processes, improve the smoothness of the overall process, and enhance the usability of TCM vocabulary units.

[0054] In one embodiment of the present invention, S135 includes:

[0055] Thematic keywords are extracted from sentence-level semantic unit data of TCM texts. By combining the TF-IDF algorithm with TCM domain word weight adjustment, keywords representing the diagnosis and treatment stage (e.g., initial diagnosis, follow-up, re-visit) or the core of the disease (e.g., cough, dizziness, diabetes) in each sentence are identified, and a sentence-keyword mapping table is generated.

[0056] Based on the sentence-keyword mapping table, a topic similarity assessment model is constructed to calculate the keyword overlap, semantic relevance, and diagnostic logic coherence between different sentences, and to generate a sentence topic similarity matrix.

[0057] Based on the sentence topic similarity matrix, a density clustering algorithm is used to group sentences with similarity higher than a set threshold (e.g., 85%) that revolve around the same diagnosis and treatment stage into one category, generating a cluster set of sentences for diagnosis and treatment stages (e.g., clustering of initial diagnosis descriptions and clustering of follow-up diagnosis efficacy).

[0058] For sentences not included in the diagnosis and treatment stage cluster set, secondary clustering is performed according to the core keywords of the disease. Sentences describing the etiology, symptoms, treatment, and prescriptions of the same disease are integrated into disease-themed cluster sets (e.g., coronary heart disease diagnosis and treatment cluster, chronic nephritis management cluster).

[0059] For each cluster set (including the diagnosis and treatment stage cluster set and the disease topic cluster set), the sentences are reordered according to the logical order of TCM diagnosis and treatment (e.g., symptom description -- syndrome differentiation and analysis -- treatment method establishment -- prescription -- efficacy feedback) to generate an ordered set of sentences;

[0060] Semantic connection optimization is performed on the ordered set of sentences, and transitional statements that conform to the expression habits of traditional Chinese medicine are added (e.g., based on the above symptoms, the diagnosis is..., the patient's feedback after taking the medicine...). Logical gaps between sentences are eliminated, forming paragraph units with complete structure and coherent information, and finally generating paragraph-level semantic unit data of traditional Chinese medicine text.

[0061] The working principle and effects of the above technical solution are as follows: By extracting keywords through TF-IDF combined with word weight adjustment in the field of Traditional Chinese Medicine (TCM), it can accurately grasp core information such as initial diagnosis and cough. For example, it will not misjudge a cold mentioned occasionally in the medical record as the core symptom, nor will it miss key expressions indicating the treatment stage, such as follow-up visits. This makes the topic tags of each sentence more closely match the actual treatment focus, improving the accuracy of TCM sentence topic positioning. The topic similarity model is used to calculate overlap, relevance, and logical coherence. For example, it quantifies the similarity between a patient's cough at the initial diagnosis and the initial prescription for cough suppression, avoiding bias caused by experience-based classification during manual judgment, and making sentence clustering more effective. This approach provides objective evidence, reduces the splitting of similar sentences and the merging of dissimilar sentences, and lowers the subjectivity of judging sentence thematic relevance. Sentences from the same treatment stage are clustered based on similarity thresholds; for example, initial symptoms, diagnosis, and initial prescription are all grouped into the initial diagnosis cluster. This eliminates the need to search through the entire medical record for information from different locations at the same stage, making subsequent review of the complete treatment process at a specific stage more efficient and avoiding confusion caused by scattered stage information, thus enhancing the aggregation of treatment stage information. Sentences not included in the stage clusters are secondary clustered based on the core symptoms; for example, the etiology of coronary heart disease, treatment methods for coronary heart disease, and management of chronic nephritis are clearly separated, preventing... The system avoids mixing diagnostic and treatment information for different symptoms within the same cluster, making each cluster more focused on a specific symptom and reducing cross-symptom information ambiguity. It also prioritizes sentences in the order of symptoms-syndrome differentiation-treatment-prescription-efficacy, for example, first presenting the patient's dizziness and fatigue (symptom), then the syndrome differentiation of qi and blood deficiency (syndrome differentiation), and finally the prescription of Bazhen Tang (prescription). This aligns with the clinical thinking habits of Traditional Chinese Medicine (TCM) and avoids logical breaks caused by the original disorganized sentence order, allowing readers to quickly understand the diagnostic and treatment process and improving the clarity of TCM diagnostic logic. Furthermore, it supplements the system with TCM-specific transitional phrases, such as those between syndrome differentiation and treatment. Based on the aforementioned symptoms, the diagnosis is qi and blood deficiency, therefore the treatment focuses on tonifying qi and nourishing blood. This makes the originally independent sentences connect more naturally, avoiding the abrupt transition from symptoms to prescriptions. The resulting paragraph units are more fluent to read and better conform to the expression logic of traditional Chinese medicine texts, enhancing the semantic coherence of paragraph units. The generated paragraph units have complete structures and clear logic. When constructing semantic parsing frameworks or extracting paragraph features, there is no need to spend time sorting out sentence order and logical relationships. Work can be carried out directly based on the well-organized paragraph units, reducing repetitive processing steps, improving overall semantic processing efficiency, and reducing the difficulty of subsequent paragraph-level semantic parsing.

[0062] In one embodiment of the present invention, S2 includes:

[0063] S21: Based on the semantic parsing framework of traditional Chinese medicine text, entity recognition, relation extraction and contextual semantic parsing are performed on unstructured TCM texts to complete deep parsing and generate text parsing results;

[0064] S22: Based on the text parsing results, perform multi-dimensional semantic feature extraction. The multi-dimensional features include symptom features, syndrome classification, treatment methods and time dimensions, to obtain TCM text semantic feature data and medical record time series data arranged in chronological order.

[0065] S23: Perform conflict detection on the semantic feature data of TCM texts, identify corresponding problems, including terminological ambiguity, descriptive contradictions and feature repetition, and generate semantic conflict data;

[0066] S24: Based on semantic conflict data, combined with the TCM standardized terminology database and clinical validation rules, semantic conflict resolution is performed to generate standardized semantic feature data.

[0067] The working principle and effects of the above technical solution are as follows: Through the combined operations of entity recognition, relationship extraction, and context semantic parsing, it can not only extract entities such as Astragalus membranaceus and kidney yang deficiency syndrome from chaotic medical record texts, but also clarify the association between kidney yang deficiency syndrome - the selection of Astragalus membranaceus, and can also understand the context logic of the alleviation of chills after the patient takes Astragalus membranaceus, avoiding the shallow problem of traditional parsing that only recognizes entities and does not understand associations, making the text parsing more in line with the logic of traditional Chinese medicine diagnosis and treatment, and improving the depth of parsing unstructured texts in traditional Chinese medicine; extracting features from four dimensions of symptoms, syndromes, treatment methods, and time, such as grasping both soreness and weakness in the waist and knees (symptoms), kidney yang deficiency (syndrome), remembering Jinkui Shenqi Pills (treatment method), and the first diagnosis in early May 2024 (time), without missing key information like single-dimensional extraction. When analyzing the evolution of the disease condition and tracking the treatment effect later, there is no need to go back and find data again, reducing the trouble of rework and lowering the omission rate of multi-dimensional semantic feature extraction; sorting the medical record information in chronological order can clearly present the complete timeline of the initial diagnosis symptoms - the adjustment of the syndrome type during the follow-up visit - the treatment effect during the follow-up visit. For example, it can be directly seen that the patient had dizziness in March (initial diagnosis) - the dizziness was relieved after taking traditional Chinese medicine in April (follow-up visit) - the dizziness disappeared in May (follow-up visit), avoiding the interruption of the diagnosis and treatment process caused by the chaotic time information in the original medical records, making it more convenient for subsequent dynamic knowledge extraction, and enhancing the usability of the medical record time series data; detecting conflicts to find problems such as term ambiguity (for example, "getting angry" refers to both excess heat and deficiency fire) and descriptive contradictions (saying no sweating at first and then mentioning spontaneous sweating later). Instead of waiting to discover data contradictions during subsequent diagnosis and treatment applications, early detection can reduce the辨证偏差 caused by data errors, making the semantic feature data more reliable and reducing the contradictions and ambiguities in traditional Chinese medicine semantic features; combining a standardized term library and clinical rules to resolve conflicts, such as unifying "pain in the lumbar spine" into the standard term "low back pain", and correcting the contradictory "no sweating / spontaneous sweating" to "spontaneous sweating after activity" through clinical verification, avoiding term confusion caused by different doctors' writing habits, enabling the medical record feature data from different sources to be uniformly connected to subsequent syndrome differentiation rule mining, reducing data incompatibility problems, and improving the standardization of traditional Chinese medicine semantic features; the generated standardized semantic feature data has no conflicts and meets the standards. When mining rules of symptoms - syndrome types from the data later, there is no need to spend a lot of time cleaning contradictory data and unifying terms, and it can directly focus on the core diagnosis and treatment association analysis, improving the efficiency of rule mining, and making the mined rules more in line with traditional Chinese medicine clinical standards, reducing the difficulty of subsequent syndrome differentiation rule mining; the feature data after in-depth parsing and standardization processing can not only be used for knowledge graph construction, but also directly for traditional Chinese medicine teaching (such as demonstrating standard terms and diagnosis and treatment logic), clinical quality control (such as checking whether the diagnosis and treatment conform to the norms), avoiding the inefficient problem of parsing the original unstructured text once at a time, enabling the data to play a role in multiple scenarios, and enhancing the reuse value of traditional Chinese medicine text data.

[0068] In one embodiment of the present invention, the S22 includes:

[0069] Based on the text parsing results, entity and attribute information related to symptoms are filtered out, and the location (e.g., waist and knee, chest and abdomen), nature (e.g., soreness, distending pain), frequency of occurrence (e.g., occasional, frequent) and accompanying symptoms (e.g., accompanied by dizziness, accompanied by fatigue) of symptoms are extracted to generate a symptom feature dataset.

[0070] Based on the syndrome description information in the text parsing results, combined with the TCM syndrome differentiation system (such as the Eight Principles of Syndrome Differentiation and the Zang-Fu Syndrome Differentiation), the core pathogenesis of the syndrome (such as "Qi deficiency" and "blood stasis"), the concurrent syndrome (such as "Qi deficiency combined with phlegm and dampness") and the diagnostic basis (such as "pale tongue with white coating and weak pulse are signs of Qi deficiency") are extracted to generate a syndrome classification feature dataset.

[0071] The diagnosis and treatment related content is separated from the text parsing results. The treatment methods (such as traditional Chinese medicine treatment and acupuncture treatment), specific plans (such as prescriptions: modified Guizhi Tang, acupoints: Zusanli and Guanyuan), dosage specifications (such as 15g of Astragalus membranaceus and acupuncture depth of 0.5 cun) and treatment course arrangements (such as 1 dose per day for 7 consecutive days) are extracted to generate a diagnosis and treatment method feature dataset.

[0072] The time information in the text parsing results is identified, and the occurrence time of the diagnosis and treatment event (e.g., initial visit on 2025-03-10, follow-up visit on 2025-03-17), the symptom onset time (e.g., cough appeared 3 days ago), and the treatment intervention time (e.g., medication started yesterday) are extracted to generate a time-stamped dataset.

[0073] By integrating symptom feature datasets, syndrome classification feature datasets, and treatment method feature datasets, semantic relationships between features are established (e.g., lower back and knee pain, pale tongue with white coating - kidney yang deficiency syndrome - treatment with Jin Kui Shen Qi Wan), and semantic feature data of traditional Chinese medicine text is generated.

[0074] Based on the time-stamped dataset, the diagnosis and treatment information (such as initial symptoms, follow-up syndrome adjustment, and follow-up efficacy) corresponding to different time nodes in the semantic feature data of TCM texts are sorted in chronological order, and the diagnosis and treatment logical connections between time nodes are supplemented (such as adjusting prescriptions during follow-up visits due to poor initial efficacy), generating time series data of medical records arranged in chronological order.

[0075] The working principle and effects of the above technical solution are as follows: Symptoms are extracted by breaking them down into location, nature, frequency of occurrence, and accompanying symptoms. For example, from a patient experiencing occasional lower back and knee pain accompanied by fatigue over the past week, the solution can accurately extract the location: lower back and knees; nature: pain; frequency: occasional; and accompanying symptoms: fatigue. Unlike traditional extraction methods that only focus on lower back and knee pain, this approach avoids missing crucial details. Subsequent diagnosis allows for more comprehensive reference to symptom information, improving the detail of TCM symptom feature extraction. Furthermore, it combines the Eight Principles and Zang-Fu (internal organs) diagnostic systems to extract syndrome characteristics. For instance, it clarifies that the core pathogenesis of Qi deficiency with phlegm-dampness is Qi deficiency, and the accompanying symptom is phlegm-dampness. It also indicates the diagnostic basis as pale tongue with white coating and slippery pulse, avoiding the problem of simply stating Qi deficiency without specifying the accompanying symptoms or lacking diagnostic support. This makes syndrome classification more clinically grounded and facilitates subsequent tracing. The diagnostic logic reduces the ambiguity of TCM syndrome classification; it extracts data from all dimensions, including type, treatment plan, dosage, and course of treatment. For example, it can completely extract the TCM treatment (type), the modified Guizhi Tang (treatment plan), the dosage of 15g of Astragalus membranaceus, and the duration of treatment (course of treatment). This prevents situations where dosage is missed, leading to incorrect prescriptions or unclear treatment durations. It allows diagnostic and treatment data to be directly linked to actual clinical applications, enhancing the completeness of diagnostic and treatment characteristics. It specifically extracts three types of time: diagnostic and treatment events, symptom onset, and treatment intervention. For example, it clearly distinguishes between the initial consultation on March 10, 2025 (diagnosis time), the cough three days ago (symptom time), and the medication taken yesterday (intervention time). This avoids the problem of time being interspersed in the original medical records, making it difficult to organize and facilitating subsequent processing. Time-based sorting lays a clear foundation, reducing the confusion in extracting time-related information. It establishes connections between symptoms, syndromes, and treatment methods; for example, clearly defining lower back and knee pain + pale tongue with white coating as corresponding to kidney yang deficiency syndrome, and then to Jin Kui Shen Qi Wan (a traditional Chinese medicine formula), avoids the need to spend time later finding which symptom corresponds to which syndrome and which medicine, as with storing three categories of data separately. This makes the feature data more meaningful for diagnosis and treatment guidance, improving the relevance of TCM semantic features. When sorting by time, it supplements the logical connections of diagnosis and treatment, such as explaining that the prescription was adjusted during the follow-up visit because the initial treatment was ineffective. This avoids the logical gaps that occur when information is simply listed by time without understanding why adjustments were made. It allows for a direct understanding of the reasons for the evolution of diagnosis and treatment from the initial visit to the follow-up visit. This also enables more accurate capture of diagnosis and treatment patterns during subsequent dynamic knowledge extraction, enhancing the time-based nature of medical records. The sequence is logically sound; the generated semantic feature data is clearly correlated and the time series logic is coherent. When performing semantic conflict detection, it can quickly determine whether a symptom contradicts the corresponding syndrome type. When performing syndrome differentiation rule mining, the correlation data of symptom-syndrome-treatment can be used directly without having to reintegrate information from different dimensions, which improves the smoothness of the overall process and reduces the connection cost of subsequent data applications. Each feature data can be traced back to the specific source of the original analysis result. For example, the syndrome differentiation basis of kidney yang deficiency can be traced back to the text description of pale tongue with white coating and weak pulse. The use of Jin Kui Shen Qi Wan can be traced back to the syndrome type of kidney yang deficiency. If data problems are found later, the source can be quickly located and corrected, reducing the trouble of troubleshooting data errors and improving the traceability of TCM diagnosis and treatment data.

[0076] In one embodiment of the present invention, S3 includes:

[0077] S31: Based on the standardized semantic feature data, the association rule mining algorithm is used to extract the direct correspondence between symptoms and syndrome types, and between syndrome types and treatment methods, to obtain explicit syndrome differentiation rule data;

[0078] S32: By using a deep learning model to mine the implicit associations in the standard semantic feature data, extract the indirect diagnosis and treatment logic chain, and generate implicit reasoning path data;

[0079] S33: Using explicit dialectical rule data as a benchmark, construct an evaluation system, which includes support, confidence and clinical conformity, to evaluate the rule credibility of implicit reasoning path data and generate path credibility evaluation results.

[0080] S34: Based on the path credibility assessment results, retain high credibility inference paths, correct or eliminate low credibility paths, complete the inference path optimization, and obtain optimized inference path data.

[0081] The working principle and effects of the above technical solution are as follows: The association rule mining algorithm directly captures the correspondence between symptoms and syndrome types, and between syndrome types and treatment methods. For example, it quickly identifies clear associations such as yellow and greasy tongue coating + bitter taste in the mouth – damp-heat accumulation in the spleen syndrome, and damp-heat accumulation in the spleen syndrome – heat-clearing and dampness-eliminating treatment. This eliminates the need for manual review of each medical record, significantly improving efficiency compared to traditional manual rule summarization. It also covers more medical record data, reduces human oversights, and improves the extraction efficiency of explicit TCM syndrome differentiation rules. Furthermore, it mines implicit associations through deep learning models. For instance, it can discover indirect chains such as long-term sleep deprivation – yin deficiency – night sweats – Baihe Gujin Tang (sleep deprivation is not a direct symptom, but it is indirectly associated with treatment methods and prescriptions through yin deficiency), avoiding the pitfalls of traditional methods. The problem of only seeing the direct relationship between symptoms and syndrome types without uncovering the deeper logic makes the diagnostic and treatment logic chain more complete and reduces the difficulty of uncovering the implicit diagnostic and treatment logic in traditional Chinese medicine. A three-dimensional evaluation is used: support (frequency of rule occurrence), confidence (accuracy of rules), and clinical conformity (whether it fits actual diagnosis and treatment). For example, a path with high support but low clinical conformity (such as headache - kidney yang deficiency syndrome) will be identified as low confidence, avoiding erroneous reasoning caused by only looking at data frequency without considering clinical rationality, making the paths more reliable and enhancing the credibility of implicit reasoning paths. Low-confidence paths are corrected or eliminated, such as deleting clinically rare paths like occasional nasal congestion - lung qi deficiency syndrome, and dry mouth - The vague description of dry mouth in Yin deficiency is supplemented by retaining nighttime dry mouth accompanied by night sweats, avoiding the inclusion of these unreliable paths in subsequent knowledge graph construction, which could lead to errors in subsequent syndrome differentiation recommendations, reducing application risk and minimizing interference from low-value reasoning paths in subsequent applications; explicit rules capture direct associations, implicit paths uncover indirect logic, and then high-value paths are optimized and integrated to form a two-layer rule system of direct + indirect. For example, it knows both that cough with yellow phlegm leads to lung heat syndrome and Sangju Yin (explicit), and that smoking leads to lung dryness, cough with little phlegm, and Sha Shen Mai Dong Tang (implicit), allowing syndrome differentiation rules to cover more diagnostic and treatment scenarios, not limited to a single association pattern, and improving the hierarchical nature of TCM syndrome differentiation rules; the optimized reasoning... The path data is both clinically logical and precise. When these paths are subsequently transformed into entity relationships in a knowledge graph, it is no longer necessary to spend a lot of time filtering out invalid associations. A complete link of symptom-pathogenesis-syndrome-treatment can be directly built using high-confidence paths, making the diagnostic and treatment logic of the knowledge graph more rigorous and improving the accuracy of subsequent knowledge graph construction. The explicit rules are clear and easy to understand, making it suitable for beginners to quickly master basic syndrome differentiation methods. The optimized implicit paths can reveal deeper diagnostic and treatment logic, such as why prolonged sitting is associated with blood stasis and then with Taohong Siwu Decoction. Beginners can understand the logic behind syndrome differentiation through these paths, rather than just memorizing rules, resulting in better learning outcomes and reducing the learning cost of syndrome differentiation for TCM beginners.Both explicit rules and optimized implicit paths align with clinical practice. When integrated into intelligent systems, they can provide doctors with comprehensive suggestions: what syndrome type corresponds to the symptoms, why it corresponds, and what treatment method should be used. This avoids the problem of traditional auxiliary tools only providing conclusions without supporting evidence, helping doctors make diagnostic decisions with greater confidence and enhancing the auxiliary value of TCM diagnostic decision-making.

[0082] In one embodiment of the present invention, step S4 includes:

[0083] S41: Based on medical record time series data, time series data mining technology is used to extract diagnosis and treatment information at different stages of diagnosis and treatment. The diagnosis and treatment information includes symptom changes, prescription adjustments and efficacy feedback information. Dynamic knowledge extraction processing is performed on the TCM diagnosis and treatment process to generate dynamic TCM diagnosis and treatment knowledge data.

[0084] S42: Using optimized reasoning path data as a framework, transform dynamic knowledge data of TCM diagnosis and treatment into entities and relationships between entities in a knowledge graph, and construct a preliminary dynamic knowledge graph;

[0085] S43: Supplement entity attributes and optimize relationship weights for the preliminary dynamic knowledge graph, incorporate explicit dialectical rule data, and generate a dynamic knowledge graph containing dialectical rules and reasoning paths;

[0086] S44: Based on dynamic knowledge graphs, clustering analysis and feature extraction algorithms are used to mine individual patient information, including diagnostic and treatment characteristics, physical characteristics, and disease progression patterns. Individualized cognitive feature analysis is performed on the dynamic knowledge graph to generate individualized cognitive feature data.

[0087] The working principle and effects of the above technical solution are as follows: By mining time-series data, it extracts symptom changes, prescription adjustments, and efficacy feedback at different stages. For example, it can clearly track the dynamic process of severe cough at the initial diagnosis – cough relief at the follow-up visit – cough disappearance during follow-up; and the initial use of Xiao Qinglong Tang at the initial diagnosis – reduction of ephedra dosage at the follow-up visit. This avoids the problem of traditional static knowledge extraction, which only focuses on the final treatment result and ignores the process evolution. It makes the extracted knowledge more aligned with the dynamic characteristics of TCM diagnosis and treatment, which involves adjustments based on symptoms, thus improving the dynamic capture ability of TCM diagnostic and treatment knowledge. Dynamic knowledge is transformed using optimized reasoning paths as a framework. For example, it follows a reasoning chain of symptom change – syndrome adjustment – ​​prescription optimization, transforming the cough relief – gradual dissipation of cold pathogens – Xiao Qinglong Tang into a formula. The reduction of ephedra is transformed into a symptom-syndrome-formula association in the knowledge graph, eliminating the need to build the graph logic from scratch. This avoids the disconnect between the graph structure and the diagnostic and treatment logic, making the initial graph more clinically instructive and reducing the difficulty of adapting the knowledge graph construction to the diagnostic and treatment logic. It supplements entity attributes (e.g., adding pungent and warm properties to ephedra and its lung meridian affinity, and adding exterior-releasing, cold-dispersing, lung-warming, and phlegm-resolving properties to Xiao Qinglong Tang), optimizes relationship weights (e.g., increasing the association weight of cold-induced lung binding syndrome - selection - Xiao Qinglong Tang), and incorporates explicit rules (e.g., aversion to cold and fever + no sweating + cough and wheezing - cold-induced lung binding syndrome). This avoids the problems of the graph only containing entities, lacking attributes, and having weak associations, making the graph more information-rich, logically rigorous, and enhancing its effectiveness. The information completeness of dynamic knowledge graphs: Clustering and feature extraction based on dynamic knowledge graphs can accurately locate the unique information of individual patients. For example, from a patient's cough worsening with each exposure to cold, pale tongue with white coating, and effectiveness of warming lung medicine, individualized characteristics such as a cold constitution, cough induced by cold, and sensitivity to warming lung treatment can be extracted. Unlike analyzing data directly without a graph, which fails to find the connection between individual characteristics and treatment logic, dynamic knowledge graphs make the mining results more targeted and reduce the blindness in individualized treatment feature mining. Dynamic knowledge graphs contain both group treatment patterns (such as general treatment methods for cold-induced lung syndrome) and record individual treatment details (such as specific additions and subtractions for a particular patient with cold-induced lung syndrome). When doctors encounter similar situations later... In case studies, both group rules and individual experiences can be referenced, avoiding the problem that high-quality clinical knowledge exists only in a single case and is difficult to reuse. It also provides an intuitive carrier for the inheritance of TCM experience, improving the reuse and inheritance value of TCM clinical knowledge. The generated individualized data can directly reflect the patient's clinical response patterns (e.g., patients are sensitive to astragalus, and dosage exceeding 15g can easily cause internal heat) and constitution-matching plans (e.g., patients with phlegm-dampness constitution should use drying Chinese medicine with caution). When constructing cognitive evolution models in the future, it is no longer necessary to sift through massive amounts of data for individual information. Personalized treatment plans can be directly formulated based on these characteristics, avoiding the problem of a one-size-fits-all approach and enhancing the clinical applicability of individualized cognitive characteristic data.In the initial knowledge graph stage, attribute supplementation, weight optimization, and rule integration are completed. Subsequent system evolution or experience updates only require adjustments to details within the existing framework (such as adding attributes of a certain Chinese herbal medicine or correcting the weight of a relationship), without reconstructing the knowledge graph structure. This reduces redundant development work and allows the knowledge graph to continuously adapt to new diagnostic and treatment needs, lowering the cost of subsequent optimization of the dynamic knowledge graph. Each entity and relationship in the dynamic knowledge graph can correspond to a specific diagnostic and treatment stage and time point. For example, the reduction of ephedra in Xiao Qinglong Decoction can be traced back to the follow-up visit and the background of cough relief; a patient's sensitivity to astragalus can be traced back to the record of "experiencing symptoms of internal heat after taking 20g of astragalus at the initial consultation." If diagnostic and treatment questions arise later, the entire process can be quickly traced back to investigate the cause of the problem, reducing medical risks and improving the traceability of the TCM diagnostic and treatment process.

[0088] In one embodiment of the present invention, S43 includes:

[0089] The entities in the preliminary dynamic knowledge graph are subjected to attribute missing detection. They are classified and sorted according to entity type (symptom, syndrome, Chinese medicine, prescription) to identify the missing attribute information of symptom entities. The attribute information includes the onset time and inducing factors, the properties, flavors, meridians and processing methods of missing Chinese medicine entities, and the formulation principles and decoction methods of missing prescription entities, generating a list of missing entity attributes.

[0090] Based on the list of missing entity attributes, the system associates authoritative databases in the field of traditional Chinese medicine (such as the Pharmacopoeia of the People's Republic of China and the Differential Diagnosis of Symptoms in Traditional Chinese Medicine) and historical medical records to extract corresponding supplementary attribute data, completes the attributes of entities, and generates a dynamic knowledge graph with improved attributes.

[0091] Initial weights are assigned to the relationships between entities in the dynamic knowledge graph after the attributes are improved (e.g., symptoms-cause-syndrome, syndrome-correspondence-treatment, treatment-selection-prescription). Basic weight values ​​are set according to the frequency of clinical occurrence and the degree of expert consensus, and a knowledge graph relationship set with initial weights is generated.

[0092] A semantic association strength algorithm is introduced to calculate the co-occurrence probability and treatment logic dependency of entity relationships in historical medical cases. Combined with the initial weight values, the algorithm is iteratively corrected to increase the weight of high clinical value relationships (e.g., damp-heat accumulation in the spleen syndrome - select - Yin Chen Hao Tang) and decrease the weight of marginal relationships (e.g., occasional headache - association - kidney yang deficiency syndrome), generating a dynamic knowledge graph with optimized relationship weights.

[0093] The explicit diagnostic rule data is structured and transformed, and the symptom combination-syndrome determination and syndrome-treatment recommendation rules are decomposed into triplet forms of condition entity-rule relation-result entity (e.g. [lower back and knee pain + aversion to cold and cold limbs]-[determined as kidney yang deficiency syndrome]), generating a rule triplet dataset;

[0094] The rule triple dataset is fused and matched with the dynamic knowledge graph with optimized relation weights. Dialectical rule association tags are added between corresponding entities in the knowledge graph to establish a mapping link between explicit dialectical rules and entity relationships. At the same time, the implicit relation links corresponding to the optimized reasoning path data are retained, and finally a dynamic knowledge graph containing dialectical rules and reasoning paths is generated.

[0095] The working principle and effects of the above technical solution are as follows: It sorts out the specific attributes missing from entities such as symptoms, traditional Chinese medicines, and prescriptions by type (e.g., the onset time of symptoms, the properties and meridian tropism of traditional Chinese medicines), giving attribute supplementation a clear target and avoiding omissions or errors in previous attribute supplementation methods. This makes the information of each entity more complete. For example, ephedra is not only a traditional Chinese medicine, but its properties of being pungent and warm, and tropism to the lung meridian, can be clearly defined, facilitating accurate subsequent application and improving the completeness of entity attributes. It also links to authoritative databases such as the Pharmacopoeia and historical medical records to supplement attributes, ensuring the professionalism and reliability of the supplemented data. For example, the principle of tonifying, relieving muscle tension, and harmonizing the Ying and Wei in the Guizhi Tang prescription is not added out of thin air, but based on authoritative evidence, avoiding misleading information caused by arbitrary fabrication of attribute information and enhancing the knowledge graph. The authority of entity information is ensured by first setting initial weights based on clinical frequency and expert consensus, and then using algorithms to calculate co-occurrence probability and logical dependency for correction. This results in higher weights for high-value relationships such as damp-heat accumulation in the spleen syndrome - Yin Chen Hao Tang, and lower weights for marginal relationships such as occasional headache - kidney yang deficiency syndrome. This avoids biases caused by assigning weights based on personal experience, making the importance of relationships more consistent with clinical reality and reducing the subjectivity of entity relationship weight setting. Rules are broken down into condition-relationship-outcome triplets (e.g., [lower back and knee pain + aversion to cold and cold limbs] -- [kidney yang deficiency syndrome]), which can quickly match with entity relationships in the graph, eliminating the need for repeated adjustments due to rule format incompatibility. This allows explicit rules to be smoothly integrated into the graph, enhancing the rule guidance capability of the graph and improving the integration of explicit diagnostic rules with the knowledge graph. The system achieves high fusion efficiency; it retains optimized implicit reasoning paths (such as indirect diagnostic logic) while also linking explicit diagnostic rules through tags, creating a complementary relationship. For example, it recognizes both the explicit rule of "cold evil binding the lungs - Xiao Qinglong Tang" and the implicit path of "cold catching a chill - worsening cough - cold evil binding the lungs," avoiding logical bias caused by focusing solely on rules or reasoning, and reducing the separation between explicit and implicit logic in the knowledge graph. The comprehensive attributes, reasonable weights, and fusion rules allow the graph to directly support clinical scenarios. For instance, when a doctor examines for kidney yang deficiency, they can see the corresponding Jin Kui Shen Qi Wan (high-weight relationship), refer to the diagnostic rule of "lower back and knee pain + aversion to cold and cold limbs - kidney yang deficiency," and understand the decoction method, eliminating the need to switch between multiple tools and improving efficiency. The convenience of this feature enhances the clinical applicability of the dynamic knowledge graph. Attributes are organized by type, relationships have weighted standards, and rules have a fixed format. When adding new entities (such as new prescriptions) or updating rules, a unified approach can be adopted. For example, adding attributes to a new prescription only requires referring to the missing list, without redesigning the process, reducing maintenance costs and lowering the difficulty of subsequent knowledge graph updates and maintenance. By distinguishing the importance of relationships through weights and integrating explicit and implicit logic, it can present the complex characteristics of TCM's treatment of the same disease in different ways and different diseases in the same way. For example, it can reflect the general rule of using Sijunzi Decoction for spleen deficiency syndrome, and also show the special case of adding Poria to Sijunzi Decoction when spleen deficiency is accompanied by dampness through weights. This makes the graph more in line with the flexibility of TCM diagnosis and treatment, and improves the knowledge graph's ability to express the complex diagnostic and treatment logic of TCM.

[0096] In one embodiment of the present invention, step S5 includes:

[0097] S51: Based on individualized cognitive characteristic data, combined with time-series prediction models and deep learning algorithms, construct an individualized cognitive evolution model that simulates the development of an individual's disease and the response to diagnosis and treatment;

[0098] S52: Match the individualized cognitive evolution model with the database of famous doctors' experience cases, extract the adaptation information that matches the individual characteristics, including the famous doctors' diagnosis and treatment ideas, medication experience and syndrome differentiation skills, and process the inheritance of famous doctors' experience according to the individualized cognitive evolution model to generate TCM experience inheritance data;

[0099] S53: Update the content and structure of the dynamic knowledge graph based on the data of traditional Chinese medicine experience inheritance, and iteratively optimize the system's diagnostic accuracy and recommendation rationality by combining reinforcement learning algorithms;

[0100] S54: Integrate and optimize the dynamic knowledge graph, individualized cognitive evolution model, and TCM experience inheritance data to construct a modular system with intelligent auxiliary diagnosis and treatment functions, and finally form a TCM intelligent system.

[0101] The working principle and effects of the above technical solution are as follows: The individualized cognitive evolution model can combine temporal prediction and deep learning to simulate individual characteristics such as the speed at which patients with Yang deficiency constitution relieve symptoms after taking warming Yang medicine, and the points where patients with phlegm-dampness constitution are prone to relapse. This allows doctors to predict the course of the disease in advance, avoiding the passive approach of traditional diagnosis and treatment, and improving the predictability of individual disease development. By matching the model with the experience of renowned doctors, it can accurately extract content suitable for individuals. For example, for patients with weak spleen and stomach and prone to diarrhea, it can match the experience of a famous doctor to use stir-fried Atractylodes macrocephala instead of raw Atractylodes macrocephala to reduce the laxative effect. This avoids the problem of the experience of renowned doctors being broad but difficult to implement, allowing the inherited experience to be directly connected with the specific situation of patients, enhancing the effectiveness of the model. This system enhances the targeted transmission of renowned physicians' experience; it continuously optimizes the knowledge graph and system functions through reinforcement learning. For example, if a recommended prescription for a certain syndrome type proves ineffective in 80% of cases, its weight is automatically lowered, gradually improving the accuracy of syndrome differentiation. This avoids outdated or erroneous recommendations caused by a static system after launch, making auxiliary diagnosis and treatment more reliable and reducing the error rate of the system's syndrome differentiation recommendations. By integrating the knowledge graph, evolutionary model, and experiential data, the system can simultaneously provide multi-dimensional support such as individual disease prediction, similar case references, and insights from renowned physicians' experiences. Doctors no longer need to switch between multiple tools; for example, during a consultation, they can simultaneously obtain information on the patient's potential symptoms for the following week and the approaches taken by a renowned physician in handling similar situations, improving clinical decision-making efficiency and enhancing the effectiveness of traditional Chinese medicine. The intelligent system boasts comprehensive auxiliary capabilities; its personalized model can address the unique characteristics of TCM, where different treatments are used for the same disease. For example, for two patients with a cold, the system will match different evolutionary paths and renowned doctors' experiences based on the differences in constitution—one with wind-heat and the other with wind-cold—avoiding a one-size-fits-all recommendation model. This makes the system more aligned with the practical needs of flexible TCM diagnosis and enhances its adaptability to complex diagnostic and treatment scenarios. The system's integrated renowned doctors' experience and dynamic knowledge can provide specific guidance, much like a mentor. For instance, when a young doctor encounters a complex case of cold and heat, the system will demonstrate how a renowned doctor differentiates between primary and secondary symptoms and adjusts the proportion of medications, eliminating the need for self-exploration from thick medical records, shortening the learning cycle, and reducing the difficulty of learning and growth for young doctors. By integrating scattered renowned doctors' experience and individual case data into the system, grassroots doctors can easily access high-quality resources. For example, when doctors in remote areas encounter difficult and complex cases, the system can quickly match relevant renowned doctors' experience and treatment logic, avoiding the waste problem of high-quality resources being concentrated in big cities and difficult to popularize, and improving the utilization efficiency of TCM diagnosis and treatment resources. After the modular system is systematized, the operation is simpler. Doctors do not need to master complex algorithm knowledge. They only need to input patient information to get clear auxiliary suggestions (such as recommending a famous doctor's experience of modifying Wendan Decoction based on the patient's constitution). This avoids the situation where the technology is complex and looks good but is not usable, allowing more doctors to actually use the system and lowering the application threshold of TCM intelligent system.

[0102] One embodiment of the present invention provides a natural language processing and knowledge graph construction system for traditional Chinese medicine electronic medical records, comprising:

[0103] One or more processors;

[0104] Memory, used to store one or more programs.

[0105] Wherein, when the one or more programs are executed by the one or more processors, the one or more processors are made to implement the method described in any one of the above.

[0106] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for natural language processing and knowledge graph construction of electronic medical records in Traditional Chinese Medicine, characterized in that, The method includes: S1: Divide TCM electronic medical records into multi-granularity semantic units to generate TCM text semantic unit data; construct a multi-granularity semantic decoupling engine based on the TCM text semantic unit data, thereby constructing a TCM text semantic parsing framework; S2: Based on the semantic parsing framework of TCM text, perform deep parsing of unstructured TCM text and extract multi-dimensional semantic features to obtain TCM text semantic feature data and medical record time series data; perform semantic conflict resolution on TCM text semantic feature data to generate standardized semantic feature data. S3: Based on the standard semantic feature data, dialectical rules are mined to obtain explicit dialectical rule data and implicit reasoning path data respectively; the implicit reasoning path data is evaluated for rule credibility using the explicit dialectical rule data, and the reasoning path is optimized to obtain optimized reasoning path data. S4: Dynamically extract and process knowledge from the TCM diagnosis and treatment process using time-series medical records to generate dynamic knowledge data of TCM diagnosis and treatment; construct a knowledge graph from the dynamic knowledge data of TCM diagnosis and treatment by optimizing reasoning path data to generate a dynamic knowledge graph; perform individualized cognitive feature analysis on the dynamic knowledge graph to generate individualized cognitive feature data. S5: Construct a cognitive evolution model based on individualized cognitive characteristic data to generate an individualized cognitive evolution model; process the inheritance of famous doctors' experience based on the individualized cognitive evolution model to generate TCM experience inheritance data; perform system evolution optimization based on TCM experience inheritance data and dynamic knowledge graph to finally form a TCM intelligent system.

2. The method for natural language processing and knowledge graph construction of TCM electronic medical records according to claim 1, characterized in that, S1 includes: S11: Obtain TCM electronic medical record text data from different sources to form the original TCM electronic medical record dataset; S12: Preprocess the original TCM electronic medical record dataset to generate preprocessed TCM electronic medical record text; S13: Based on the preprocessed TCM electronic medical record text, semantic units are divided according to the multi-granularity levels of characters, words, phrases, sentences and paragraphs to generate TCM text semantic unit data; S14: Based on the semantic unit data of TCM texts, combined with the TCM terminology system and natural language processing algorithms, a multi-granularity semantic decoupling engine is constructed based on semantic decoupling rules and processes. S15: By integrating semantic parsing logic, domain knowledge association, and dynamic update mechanism through a multi-granularity semantic decoupling engine, a semantic parsing framework for TCM texts is constructed.

3. The method for natural language processing and knowledge graph construction of TCM electronic medical records according to claim 2, characterized in that, S13 includes: S131: Based on the preprocessed TCM electronic medical record text, a TCM terminology word segmentation tool is used to scan the text character by character, identify and extract individual TCM characters and special symbols, and generate TCM text character-level semantic unit data. S132: Based on character-level semantic unit data, combined with a vocabulary list in the field of traditional Chinese medicine and a bidirectional maximum matching algorithm, word combination recognition is performed on continuous characters to extract traditional Chinese medicine vocabulary units and generate traditional Chinese medicine text character-level semantic unit data. S133: Based on word-level semantic unit data, through contextual semantic association analysis, words with close logical relationships are combined into phrase units to generate phrase-level semantic unit data for TCM texts; S134: Based on phrase-level semantic unit data, using punctuation marks as separators and combining the characteristics of TCM sentence expression, continuous phrases are integrated into complete TCM diagnosis and treatment description statements to generate TCM text sentence-level semantic unit data; S135: Based on sentence-level semantic unit data, sentences related to the same diagnosis and treatment stage or the same disease are clustered and integrated according to the content theme relevance to form paragraph units containing complete diagnosis and treatment information, thus generating TCM text paragraph-level semantic unit data. S136: Integrate semantic unit data at the character, word, phrase, sentence, and paragraph levels, establish a multi-granularity semantic unit association index, and generate semantic unit data of TCM texts containing hierarchical relationships and semantic mappings.

4. The method for natural language processing and knowledge graph construction of TCM electronic medical records according to claim 3, characterized in that, S132 includes: The semantic unit data of TCM text at the character level is initially matched with the constructed TCM domain lexicon, and character sequences containing core characters of the TCM domain lexicon are selected in the character-level units to generate a set of character sequences to be identified; Based on the set of character sequences to be identified, a bidirectional maximum matching algorithm is used to scan the character sequence from left to right. The character combinations are truncated according to the length of the longest word in the Chinese medicine vocabulary list. It is then determined whether they match Chinese medicine words in the vocabulary list to generate a set of left-side matching candidate words. For the same set of character sequences to be identified, the bidirectional maximum matching algorithm is used to scan the character sequence from right to left. Similarly, the character combination is truncated according to the longest word length and compared with the vocabulary of traditional Chinese medicine to generate a set of right-side matching candidate words. By comparing the left-side matching candidate vocabulary set and the right-side matching candidate vocabulary set, the matching length, vocabulary coverage and semantic rationality of the vocabulary in the two candidate sets are statistically analyzed. Vocabulary with longer matching length, higher coverage and conformity to the semantic logic of traditional Chinese medicine is selected first to generate a preliminary set of traditional Chinese medicine vocabulary units. The preliminary TCM vocabulary unit set is classified and labeled. Based on the TCM vocabulary classification system, the vocabulary is labeled as symptom words, Chinese medicine names, and prescription names, generating a classified and labeled TCM vocabulary unit set. Redundancy checks are performed on the TCM vocabulary unit set after classification and labeling, duplicate labeled vocabulary units are removed, and low-frequency words that are not recognized by the bidirectional maximum matching algorithm but conform to TCM semantics are added, finally generating TCM text word-level semantic unit data.

5. The method for natural language processing and knowledge graph construction of TCM electronic medical records according to claim 3, characterized in that, S135 includes: Thematic keywords are extracted from sentence-level semantic unit data of TCM texts. By combining the TF-IDF algorithm with TCM domain word weight adjustment, keywords representing the diagnosis and treatment stage or the core of the disease in each sentence are identified, and a sentence-keyword mapping table is generated. Based on the sentence-keyword mapping table, a topic similarity assessment model is constructed to calculate the keyword overlap, semantic relevance, and diagnostic logic coherence between different sentences, and to generate a sentence topic similarity matrix. Based on the sentence topic similarity matrix, a density clustering algorithm is used to group sentences with similarity higher than a set threshold and revolving around the same diagnosis and treatment stage into one category, generating a diagnosis and treatment stage sentence cluster set; For sentences not included in the diagnosis and treatment stage cluster set, secondary clustering is performed according to the core keywords of the disease, and sentences describing the etiology, symptoms, treatment and prescription of the same disease are integrated into disease theme cluster set; For each cluster set, the sentences are reordered according to the logical order of TCM diagnosis and treatment to generate an ordered set of sentences; Semantic connection optimization is performed on the ordered set of sentences, supplementing transitional sentences that conform to the expression habits of traditional Chinese medicine, eliminating logical gaps between sentences, forming paragraph units with complete structure and coherent information, and finally generating paragraph-level semantic unit data of traditional Chinese medicine text.

6. The method for natural language processing and knowledge graph construction of TCM electronic medical records according to claim 1, characterized in that, The S2 includes: S21: Based on the semantic parsing framework of traditional Chinese medicine text, entity recognition, relation extraction and contextual semantic parsing are performed on unstructured TCM texts to complete deep parsing and generate text parsing results; S22: Based on the text parsing results, multi-dimensional semantic features are extracted to obtain TCM text semantic feature data and medical record time series data arranged in chronological order; S23: Perform conflict detection on semantic feature data of TCM texts, identify corresponding problems, and generate semantic conflict data; S24: Based on semantic conflict data, combined with the TCM standardized terminology database and clinical validation rules, semantic conflict resolution is performed to generate standardized semantic feature data.

7. The method for natural language processing and knowledge graph construction of TCM electronic medical records according to claim 1, characterized in that, The S3 includes: S31: Based on the standardized semantic feature data, the association rule mining algorithm is used to extract the direct correspondence between symptoms and syndrome types, and between syndrome types and treatment methods, to obtain explicit syndrome differentiation rule data; S32: By using a deep learning model to mine the implicit associations in the standard semantic feature data, extract the indirect diagnosis and treatment logic chain, and generate implicit reasoning path data; S33: Using explicit dialectical rule data as a benchmark, construct an evaluation system, evaluate the rule credibility of implicit reasoning path data, and generate path credibility evaluation results; S34: Based on the path credibility assessment results, retain high credibility inference paths, correct or eliminate low credibility paths, complete the inference path optimization, and obtain optimized inference path data.

8. The method for natural language processing and knowledge graph construction of TCM electronic medical records according to claim 1, characterized in that, The S4 includes: S41: Based on medical record time series data, time series data mining technology is used to extract diagnosis and treatment information at different stages of diagnosis and treatment, and dynamic knowledge extraction processing is performed on the TCM diagnosis and treatment process to generate dynamic knowledge data of TCM diagnosis and treatment. S42: Using optimized reasoning path data as a framework, transform dynamic knowledge data of TCM diagnosis and treatment into entities and relationships between entities in a knowledge graph, and construct a preliminary dynamic knowledge graph; S43: Supplement entity attributes and optimize relationship weights for the preliminary dynamic knowledge graph, incorporate explicit dialectical rule data, and generate a dynamic knowledge graph containing dialectical rules and reasoning paths; S44: Based on dynamic knowledge graphs, clustering analysis and feature extraction algorithms are used to mine individual patient information, perform individualized cognitive feature analysis on dynamic knowledge graphs, and generate individualized cognitive feature data.

9. The method for natural language processing and knowledge graph construction of TCM electronic medical records according to claim 1, characterized in that, The S5 includes: S51: Based on individualized cognitive characteristic data, combined with time-series prediction models and deep learning algorithms, construct an individualized cognitive evolution model that simulates the development of an individual's disease and the response to diagnosis and treatment; S52: Match the individualized cognitive evolution model with the database of famous doctors' experience cases, extract the adaptation information of the individual characteristics, process the inheritance of famous doctors' experience according to the individualized cognitive evolution model, and generate TCM experience inheritance data. S53: Update the content and structure of the dynamic knowledge graph based on the data of traditional Chinese medicine experience inheritance, and iteratively optimize the system's diagnostic accuracy and recommendation rationality by combining reinforcement learning algorithms; S54: Integrate and optimize the dynamic knowledge graph, individualized cognitive evolution model, and TCM experience inheritance data to construct a modular system with intelligent auxiliary diagnosis and treatment functions, and finally form a TCM intelligent system.

10. A system for natural language processing and knowledge graph construction of electronic medical records in Traditional Chinese Medicine, including: One or more processors; Memory, used to store one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1 to 9.

Citation Information

Cited By

  • Medical document-oriented man-machine collaborative intelligent writing method and system

    CN121390080A

  • Traditional Chinese medicine voice input method and system based on semantic perception and dynamic dialectical reasoning

    CN122309716A

  • Traditional Chinese medicine voice input method and system based on semantic perception and dynamic dialectical reasoning

    CN122309716B