Traditional Chinese medicine ancient medical case implicit knowledge dominance and structuring processing method, device, medium and product

By digitally processing and semantically hierarchically annotating ancient Chinese medical records, and combining graph neural networks and a Chinese medicine ontology database, a knowledge graph of ancient Chinese medical records is constructed. This solves the problems of insufficient implicit knowledge mining and lack of structured annotation in ancient Chinese medical records, and achieves efficient knowledge explicitation and intelligent services.

CN122021822APending Publication Date: 2026-05-12上海信投智能科技股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
上海信投智能科技股份有限公司
Filing Date
2025-12-22
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Traditional Chinese medicine ancient medical records suffer from problems such as insufficient depth of implicit knowledge mining, complex language expression, lack of a unified structured annotation system, low level of intelligent knowledge service, and fragmented system processing flow, resulting in low processing accuracy and high maintenance costs.

Method used

By digitizing and preprocessing the original medical record data, using classical Chinese word segmentation, terminology standardization, and semantic hierarchical annotation, and combining graph neural networks and a TCM ontology database, a knowledge graph of ancient TCM medical records is constructed. This enables the explicit mining of implicit knowledge and knowledge fusion, forming a complete TCM knowledge base that supports intelligent question-and-answer interaction and high-precision knowledge retrieval.

Benefits of technology

It significantly improves the efficiency and accuracy of knowledge processing, enhances dynamic update capabilities, reduces long-term maintenance costs, realizes the explicit and structured processing of ancient Chinese medical records, and supports high-precision knowledge retrieval and clinical visualization-assisted applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122021822A_ABST
    Figure CN122021822A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of traditional Chinese medicine information, and discloses a traditional Chinese medicine ancient medical case implicit knowledge dominance and structuring processing method and device, a medium and a product. The method comprises the steps of obtaining a standardized medical case corpus according to original medical case data; determining a triple according to the standardized medical case corpus; the triple is used for mining implicit knowledge; obtaining a dominant knowledge set according to the triple; the dominant knowledge set is used for representing existing knowledge and implicit knowledge in a traditional Chinese medicine case; according to the dominant knowledge set, obtaining a traditional Chinese medicine ancient medical case knowledge graph; aiming at the traditional Chinese medicine ancient medical case knowledge graph, obtaining a fusion knowledge base; and realizing medical case knowledge service and visual application according to the fused knowledge base.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of traditional Chinese medicine information technology, and in particular to a method, device, medium and product for making implicit knowledge of ancient medical records of traditional Chinese medicine explicit and structured. Background Technology

[0002] Currently, in the field of traditional Chinese medicine, ancient medical records have high research and clinical application value. In recent years, the specific application of ancient medical records has roughly gone through several stages: digitization of ancient medical records, construction of knowledge bases and knowledge graphs based on digitization, and development of intelligent reasoning and question answering, in order to provide auxiliary support for traditional Chinese medicine diagnosis and treatment.

[0003] The inventors discovered that the related technologies suffer from at least the following technical problems: 1. Lack of reasoning mechanisms and insufficient depth of implicit knowledge mining, making it difficult to extract implicit empirical and regular knowledge such as etiology, pathogenesis, treatment principles, and compatibility rules from ancient medical case texts; 2. Non-standard and complex language expressions, resulting in low processing accuracy due to the large number of variant characters, dialects, archaic characters, and ambiguous expressions in ancient medical case texts; 3. Lack of a unified structured annotation system, low degree of automation in knowledge graph construction, and weak integration with traditional Chinese medicine theory; 4. Limited level of intelligence in knowledge services, with existing TCM intelligent auxiliary models exhibiting poor generalization ability and interpretability, providing only simple search results in the face of complex clinical problems; 5. Fragmented system processing flow, lack of incremental learning and dynamic update mechanisms, requiring high maintenance costs and reconstruction for data updates. Summary of the Invention

[0004] One objective of this application is to provide a method for making implicit knowledge of ancient Chinese medical records explicit and structured, at least to help solve the technical problems of difficulty in extracting empirical regular knowledge from medical record texts, lack of a unified structured annotation system, limited level of intelligence in knowledge services, and fragmented system processing flow.

[0005] To achieve the above objectives, some embodiments of this application provide the following aspects:

[0006] Firstly, some embodiments of this application provide a method for making implicit knowledge of ancient Chinese medical records explicit and structured, including: obtaining standardized medical record corpus based on original medical record data; determining triples based on the standardized medical record corpus; the triples being used to mine implicit knowledge; obtaining an explicit knowledge set based on the triples; the explicit knowledge set being used to represent existing knowledge and implicit knowledge in the Chinese medical records; obtaining a knowledge graph of ancient Chinese medical records based on the explicit knowledge set; obtaining a fused and complete Chinese medicine knowledge base based on the fused and complete Chinese medicine knowledge base; and realizing medical record knowledge services and visualization applications based on the fused and complete Chinese medicine knowledge base.

[0007] Secondly, some embodiments of this application also provide an electronic device, the electronic device comprising: one or more processors; and a memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method described above.

[0008] Thirdly, some embodiments of this application also provide a computer-readable medium having computer program instructions stored thereon, which can be executed by a processor to implement the method described above.

[0009] Fourthly, some embodiments of this application also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the method described above.

[0010] Compared with related technologies, the solution provided in this application addresses the problems of long processing cycles, high costs, and difficult updates in traditional medical record knowledge service systems. It proposes a comprehensive optimization method covering the entire process from medical record data collection and digitization, preprocessing and classical Chinese standardization, hierarchical semantic annotation and entity extraction, explicit mining of tacit knowledge, to graph construction and service application. By organically combining a fine-tuning model, a TCM ontology database, and a semantic annotation system with expert review, dynamic graph updates can be achieved. This significantly improves the efficiency and accuracy of single knowledge processing, and allows the system to dynamically follow data, feedback, and continuously evolve and improve, reducing long-term maintenance costs. Specifically, based on standardized medical record corpora, a multi-dimensional, multi-level labeling scheme is implemented to determine triples, providing a unified, fine-grained structured annotation standard for medical record knowledge. Through explicit mining of tacit knowledge, potential association patterns such as "etiology-treatment" and "syndrome type-prescription" not directly stated in medical records can be automatically learned and inferred. This transforms the tacit experience of renowned doctors into explicit, computable knowledge units, forming an explicit knowledge set and significantly improving the depth of knowledge discovery. By transforming and storing the explicit knowledge set, a knowledge graph of ancient Chinese medical records is constructed, reducing the reliance on manual annotation. Finally, knowledge fusion and semantic reasoning are performed to obtain a complete fused Chinese medicine knowledge base. This enables intelligent question-and-answer interaction and high-precision knowledge retrieval requests based on medical record knowledge. An interactive visualization front-end is built for the complete Chinese medicine knowledge base to realize clinical visualization auxiliary applications. Attached Figure Description

[0011] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0012] Figure 1An exemplary flowchart illustrating the method for making implicit knowledge of ancient Chinese medical records explicit and structured, provided in some embodiments of this application;

[0013] Figure 2 This is a schematic diagram of the structure of an electronic device provided in some embodiments of this application. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0015] The following terms are used in this document:

[0016] Implicit knowledge refers to knowledge in ancient Chinese medical records that is not directly expressed but is implied in the text, such as etiology, pathogenesis, treatment principles, and rules of prescription compatibility, which can only be revealed through reasoning and analysis.

[0017] Making tacit knowledge explicit involves transforming tacit knowledge into a clear, directly understandable, and usable form.

[0018] Structured knowledge organizes knowledge into standardized formats, such as databases, knowledge graphs, or ontology, making it easier for computers to process and apply it.

[0019] NLP, or Natural Language Processing, is used to process and analyze human language, including text segmentation, entity recognition, and relation extraction.

[0020] KG, Knowledge Graph, is a semantic network that represents knowledge through nodes (entities) and edges (relationships), supporting knowledge reasoning and querying.

[0021] Ancient medical records in traditional Chinese medicine refer to clinical cases recorded in ancient Chinese medicine literature. They are usually described in classical Chinese and include information such as symptoms, diagnosis, treatment and efficacy.

[0022] First Embodiment

[0023] The first embodiment of this application relates to a method for making implicit knowledge of ancient Chinese medical records explicit and structured, including:

[0024] Step S1: Obtain standardized medical record corpus based on the original medical record data;

[0025] Step S2: Based on the standardized medical record corpus, determine the triples; the triples are used to mine implicit knowledge.

[0026] Step S3: Based on the triples, obtain the explicit knowledge set; the explicit knowledge set is used to represent the existing knowledge and implicit knowledge in the TCM medical records.

[0027] Step S4: Based on the explicit knowledge set, obtain the knowledge graph of ancient Chinese medical records;

[0028] Step S5: Obtain the integrated and complete TCM knowledge base based on the TCM ancient medical case knowledge graph;

[0029] Step S6: Based on the integrated and complete TCM knowledge base, realize medical case knowledge services and visualization applications.

[0030] Specifically, the traditional approach involves scanning ancient Chinese medical records to obtain high-resolution images, and then using optical character recognition (OCR) to extract the text. The OCR results also require manual proofreading and data cleaning to remove errors and ensure text accuracy. Existing technologies also propose automatically annotating the generated electronic documents, recognizing and adding annotations or hyperlinks to prescription names, Chinese medicine names, literature titles, and disease names in the medical record text, thereby quickly building a structured electronic database of ancient Chinese medical records.

[0031] Before proceeding to step S1, medical record data is collected and digitized. For scanned copies of ancient medical records, text archives, and clinical records, OCR is used to recognize and output editable text. The recognition engine can be MinerU or VLM. Then, the editable text is cleaned, including formatting, noise removal, and correction of punctuation errors. Finally, the original medical record data is obtained. The original medical record data can be in text form and includes fields such as medical record number, source, time, physician information, and original medical record text.

[0032] For step S1, this step involves text preprocessing and classical Chinese standardization of the original medical record data to obtain standardized medical record corpus. Traditional Chinese medicine (TCM) medical record language has a long history and diverse expressions, often exhibiting numerous synonyms, near-synonyms, or polysemous words, as well as ambiguous expressions such as omissions and pronoun references. If these heterogeneous expressions are not uniformly normalized, it will lead to data sparsity and processing difficulties, reducing the accuracy of subsequent extraction and analysis. The text preprocessing and classical Chinese standardization operations can include: classical Chinese word segmentation, terminology standardization, terminology normalization, and denoising. Classical Chinese word segmentation uses a classical Chinese word segmentation model for sentence segmentation, such as a TCM word segmenter based on BERT; terminology standardization aims to standardize dialects, ancient characters, and variant characters, for example, unifying terms like "phlegm and fluid retention" and "blood stasis" into the same terminology; terminology normalization involves establishing a TCM terminology dictionary and using word vector semantic matching to normalize ambiguous terms; denoising removes redundant symbols and punctuation, retaining semantically relevant segments. The final standardized medical record corpus can be used in text format.

[0033] Regarding step S2, this step addresses issues such as ambiguous entity boundaries, nested confusion, and ambiguous roles in ancient medical records. It employs a hierarchical annotation and entity extraction method based on Traditional Chinese Medicine semantics, combined with an expert knowledge distillation strategy, to assign multi-level clinical semantic value to entities in the standardized medical record corpus, thereby obtaining complete hierarchical annotation results. Furthermore, triples are determined based on these complete hierarchical annotations. This step not only provides a unified, fine-grained, structured annotation standard for ancient medical record knowledge but also provides accurate and reliable foundational data for the implicit knowledge mining in the subsequent step S4.

[0034] Regarding step S3, addressing the difficulty of extracting implicit empirical knowledge such as etiology, pathogenesis, treatment principles, and compatibility rules from ancient medical case texts using traditional methods, this approach utilizes the triplet data determined in step S2 to perform explicit mining of implicit knowledge. In particular, by combining graph neural networks or rule-based reasoning with a TCM ontology database, it can automatically learn and infer potential correlation patterns such as "etiology-treatment" and "syndrome type-prescription" not directly stated in the medical case texts. This transforms implicit experience into explicit and computable knowledge units, resulting in an explicit knowledge set. This completes the implicit links in the TCM diagnosis and treatment knowledge chain and significantly improves the depth of knowledge discovery.

[0035] For step S4, this step focuses on the explicit knowledge set, performing knowledge structuring and graph construction. The non-standardized text in the explicit knowledge set is converted into standard text using the TCM ontology library; the standard text is stored in RDF or OWL format; a graph database is built using Neo4j or GraphDB, and indexes and query interfaces are established to obtain the TCM ancient medical case knowledge graph.

[0036] For step S5, this step involves knowledge fusion and semantic reasoning. Existing solutions typically employ a semi-automatic process, using algorithms to construct a knowledge graph from structured databases and supplementing unstructured documents with expert annotations to ultimately obtain an integrated TCM knowledge base. This step, however, targets the knowledge graph of ancient TCM medical records. It processes explicit synonyms through a thesaurus and implicit semantic similarities through semantic vector similarity calculations. This identifies and associates entities or diagnostic logics in different medical records that express different meanings but are similar, breaking down silos between medical records and achieving deep semantic fusion across records. Based on the fused knowledge graph network, a semantic indexing and query mechanism is constructed, directly supporting intelligent question-and-answer interaction and high-precision knowledge retrieval requests for medical record knowledge, resulting in a complete TCM knowledge base after semantic fusion.

[0037] For step S6, this step involves knowledge services and visualization applications. Based on the complete TCM knowledge base after semantic fusion, an interactive visualization front-end is built using a graph database engine such as Neo4j. Users can intuitively browse the complex relationship structure of "symptoms—syndromes—treatments—prescriptions" in medical records using a "point-line" network topology diagram, enabling subgraph exploration and path tracing. A natural language question-and-answer channel is constructed to parse user-input clinical questions, such as "How should cough with cold syndrome be treated?". Through semantic parsing, the question is mapped into a graph query statement, accurately retrieving the corresponding diagnostic and treatment logic subgraph from the knowledge base. Finally, multi-dimensional decision results are generated: based on the search results, three types of decision support information are output in a structured manner: one type is similar medical records used to trace historical classic cases as evidence; another type is recommended prescriptions used for optimal prescriptions and suggestions for addition and subtraction based on knowledge base reasoning; and the last type is empirical rules used to display relevant implicit compatibility rules and treatment principles.

[0038] Compared with related technologies, the solution provided in this application addresses the problems of long processing cycles, high costs, and difficult updates in traditional medical record knowledge service systems. It proposes a comprehensive optimization method covering the entire process from medical record data collection and digitization, preprocessing and classical Chinese standardization, hierarchical semantic annotation and entity extraction, explicit mining of implicit knowledge, to graph construction and service application. By organically combining standard systems such as fine-tuning models, TCM ontology databases, and semantic annotation systems with expert review, dynamic graph updates can be achieved. This significantly improves the efficiency and accuracy of single knowledge processing, and allows for dynamic data tracking, feedback, and continuous evolution and improvement, reducing long-term maintenance costs. Specifically, based on standardized medical record corpora, a multi-dimensional, multi-level labeling scheme is implemented to determine triples, providing a unified, fine-grained structured annotation standard for medical record knowledge. By mining tacit knowledge to make it explicit, we can automatically learn and infer potential correlation patterns such as "etiology-treatment" and "syndrome type-prescription" not directly stated in medical records. This transforms the tacit experience of renowned doctors into explicit, computable knowledge units, forming an explicit knowledge set and significantly improving the depth of knowledge discovery. Through transformation and storage operations, this explicit knowledge set is used to construct a knowledge graph of ancient Chinese medicine medical records, reducing reliance on manual annotation. Finally, knowledge fusion and semantic reasoning are performed to obtain a complete fused Chinese medicine knowledge base. This base supports intelligent question-and-answer interaction and high-precision knowledge retrieval requests based on medical record knowledge. An interactive visualization front-end is built for this complete Chinese medicine knowledge base, enabling clinical visualization-assisted applications.

[0039] Second Embodiment

[0040] The second embodiment of this application relates to a method for making implicit knowledge of ancient Chinese medical records explicit and structured. The second embodiment is an improvement on the first embodiment. The specific improvement is that it provides a method for performing hierarchical annotation and entity extraction of medical semantics.

[0041] Optionally, in some embodiments, the step S2, which involves determining the triples based on the standardized medical record corpus, may include:

[0042] Step S21: Establish a three-level dynamic semantic annotation system, which includes L0, L1 and L2 levels; wherein, L0 level is the basic entity layer, used to identify and annotate physically existing noun entities; L1 level is the clinical semantic layer, used to annotate L0 level entities to determine clinical roles; L2 level is the implicit attribute layer, used to link the extracted L0 and L1 level entities to an external TCM ontology database and obtain associated implicit attributes;

[0043] Step S22: Based on the three-level dynamic semantic annotation system, the L0 and L1 level annotation results are determined for the standardized medical record corpus execution model-expert collaborative annotation.

[0044] Step S23: Based on the annotation results of L0 and L1 levels, attach L2 level to obtain the complete three-level annotation results;

[0045] Step S24: Extract and determine triples based on the complete three-level annotation results.

[0046] Specifically, regarding step S21, the L0 basic entity layer is used to identify and extract physically existing nouns from standardized medical record corpora, clarifying entity boundaries. The entity types to be extracted, predefined in the Scheme corresponding to the L0 basic entity layer, can be: {Chinese medicine, acupoints, dosage, body parts, pulse description}. The L1 clinical semantic layer is used to determine the clinical role of entities in conjunction with the context of the medical record. The clinical role types predefined in the Schema corresponding to the L1 clinical semantic layer can be: {main symptom, concurrent symptom, present illness history, past medical history, prescription_principal drug, prescription_adjuvant drug}. Since the same drug has different roles in different positions, the same L0 entity may correspond to different L1 roles. The L2 latent attribute layer does not rely on the original text of the medical record for extraction, but instead obtains associated latent attributes as a structured reference by linking to the TCM ontology database to complete the core TCM features of the entity. The latent attribute classifications predefined in the Schema corresponding to the L2 latent attribute layer can be: {nature and flavor (cold / hot), meridian tropism, treatment method classification (sweating / vomiting / purging)}.

[0047] For step S23, without relying on the original medical record text, an external TCM ontology database is linked to structure the basic attribute information associated with L0-level basic entities and L1-level clinical roles from the ontology database, forming a complete three-level annotation result of "entity-role-attribute". For example, Cinnamon Twig [L0: Chinese Medicine]-[L1: Prescription_Chief Herb]-[L2: Properties: Warm and pungent / Meridian Tropism: Heart, Lung, Bladder / Treatment Method: Sweating Method].

[0048] For step S24, based on the relation rules defined in the TCM ontology, triples are extracted and determined from the complete three-level annotation results. This includes using entities corresponding to L0 and L1 levels as entity sources, clinical role associations at L1 level, and implicit attribute associations at L2 level as relation sources, and determining triples using two structures: factual triples (<head entity, relation, tail entity>) and attribute triples (<head entity, relation, attribute>). The TCM ontology defines entity relation rules, such as "entity-role association" and "entity-attribute association".

[0049] Current technologies lack unified standards and rules for data cleaning, structuring, and semantic annotation of TCM medical records. Although data standardization research exists, rules for entity extraction from medical records have not yet been established, and standard strategies for data quality evaluation and information preservation are lacking in the medical record processing workflow. This paper proposes a three-level dynamic semantic annotation system to identify entities in standardized medical record corpora and add corresponding clinical role annotations, clarifying the clinical functions of entities and supplementing their implicit TCM attributes, thus transforming isolated entity nouns into clinically meaningful knowledge units. Model-expert collaborative annotation ensures data quality, ultimately converting the relationships between the three-level dynamic semantic annotations into triples.

[0050] Optionally, in some embodiments, the step S22, which involves performing an expert collaborative annotation on the standardized medical record corpus to determine the annotation results for levels L0 and L1, may include:

[0051] Step S221: Pre-label L0 and L1 levels according to the fine-tuning model;

[0052] Step S222: Calculate the confidence entropy value for each pre-labeled entity;

[0053] Step S223: Determine the labeling results for L0 and L1 levels based on the confidence entropy value.

[0054] Specifically, regarding step S221, the purpose of fine-tuning the model is to achieve automated batch pre-annotation of L0 and L1 levels, which improves efficiency compared to manual annotation, and the model pre-annotation has higher accuracy, reducing the number of samples that need to be reviewed by experts in subsequent steps. The fine-tuned Qwen-TCM model can be used for pre-annotation.

[0055] In step S222, the purpose of calculating the confidence entropy value of each pre-labeled entity is to quantify the uncertainty of the fine-tuning model on the entity labeling results. The higher the entropy value, the greater the uncertainty, and the lower the entropy value, the more reliable the labeling results.

[0056] For step S223, entities with confidence entropy values ​​exceeding the threshold are prioritized for expert review, and the expert review results are received. Based on the expert review results, the L0 and L1 level annotation results are determined. Specifically, the confidence entropy values ​​of all pre-annotated entities are first traversed, and entities with confidence entropy values ​​greater than the threshold are selected and prioritized for expert review and correction. The expert review results are received, and the review results include the correct classification information of the entity to be updated, i.e., the basic category described by the L0 level entity and the clinical role corresponding to the L1 level entity. The annotation results are updated according to the expert review results, and the L0 and L1 level annotation results are finally determined. It should be noted that experts can access and correct the annotation of any entity at any stage, including entities that do not exceed the threshold. Therefore, through this step, correction can be performed only on confidence samples with confidence values ​​below the threshold. For entities with confidence entropy values ​​below the threshold, no expert intervention is required, and the pre-annotation results can be directly confirmed, greatly reducing manual costs.

[0057] Optionally, in some embodiments, the confidence entropy value is calculated as follows:

[0058]

[0059] in, This represents the confidence entropy value. The average entropy of an entity is the mean of the token-level entropy, reflecting the average level of uncertainty among all tokens in the entity. A token represents the smallest textual unit of an entity, that is, the single word or phrase fragment into which the entity is broken down in the L0 and L1 annotations. This represents the minimum confidence level of an entity, which is the minimum probability of the largest class among all tokens of that entity. It reflects the confidence level of the entity's most uncertain token. This represents the balance coefficient, and its value range is... The weighting of the entity's average entropy and minimum confidence level is used to adjust the proportion of the entity's weights. This represents a single entity to be computed in L0 and L1 levels.

[0060] Specifically, first calculate the token-level entropy, then calculate the entity average entropy, then calculate the entity minimum confidence level, and finally use the balance coefficient to calculate the confidence entropy value.

[0061] Minimum confidence level of entity The specific calculation method is as follows:

[0062]

[0063] in, Represents a single entity The corresponding token set Represents a single token in the set. This represents the category within an entity, such as the formula entity in layer L0. For "Guizhi Tang", , , , It can be .

[0064] Entity average entropy The specific calculation method is as follows:

[0065]

[0066] in, This represents the token-level entropy, which is a single token. The uncertainty measure; Represents a single entity The corresponding token set Represents a single token in the set. Represents a set The number of tokens in the system.

[0067] Token-level entropy The specific calculation method is as follows:

[0068]

[0069] in, Represents a single token in the set. Represents the category in an entity. Represents token Category The probability of.

[0070] By integrating token-level entropy and entity minimum confidence, this method effectively overcomes the error problem of traditional single average confidence algorithms in recognizing a single character in long entities. It can keenly capture local low-confidence segments and blurred boundaries within long entities in traditional Chinese medicine, significantly improving the extraction accuracy of complex entities. At the same time, by utilizing the dynamic adjustment mechanism of the balance coefficient, it achieves accurate positioning and push of high-risk samples. While ensuring that the model performance is close to the level of experts, it greatly reduces the labeling cost and workload of manual review, and optimizes the efficiency of human-machine collaboration.

[0071] Third Embodiment

[0072] The third embodiment of this application provides a method for making implicit knowledge in ancient Chinese medical records explicit and structured. The third embodiment is an improvement on the first and / or second embodiments, specifically in that it provides a method for making implicit knowledge explicit.

[0073] Optionally, in some embodiments, obtaining the explicit knowledge set based on the triples, i.e., step S3, may include:

[0074] Step S31: Construct an initial heterogeneous graph from the data in the triples; the initial heterogeneous graph is used to reflect the explicit associations of the triples, and the nodes include symptoms, prescriptions, and the gaps in the diagnosis and treatment links between them.

[0075] Step S32: Based on the initial heterogeneous graph, an enhanced atlas is obtained through a dual-channel inference mechanism; the enhanced atlas contains a complete diagnostic and treatment pathway.

[0076] Step S33: Apply a hierarchical community discovery algorithm to the enhanced graph to obtain an explicit knowledge set.

[0077] Specifically, for step S31, the data in the triples are used to construct an initial heterogeneous graph according to the rule of entities as nodes and relations as edges. The initial heterogeneous graph only reflects the explicit relationships that objectively exist in the triples. When the triples lack syndrome or treatment information, such as triples without intermediate relationships like "fever - appears in - patient" or "ephedra decoction - used in - patient", a diagnosis and treatment link gap will be formed where there is no effective connection between the symptom node and the prescription node. This diagnosis and treatment link gap is the core goal of subsequent implicit thinking chain completion.

[0078] Regarding step S32, the purpose of this step is to generate high-probability completion candidates for the diagnostic and treatment link gaps in the initial heterogeneous graph through a dual-channel reasoning mechanism, forming an enhanced graph containing candidate diagnostic and treatment paths. For example, medical records often state "the patient has fever and chills... (symptoms), prescribed Ephedra Decoction (prescription)," but the "diagnosis: wind-cold exterior syndrome" and "treatment: pungent and warm exterior-releasing method" are missing.

[0079] Regarding step S33, the hierarchical community algorithm follows the logic of first layering, then clustering, and then association. It can combine the Tree-Comm tree-shaped community paradigm with the hierarchical diagnosis and treatment logic of traditional Chinese medicine. First, the enhanced graph is split into several layers; then, for the node characteristics and domain rules of each layer, an appropriate clustering method is selected to obtain a single-level community; finally, strong associations between communities of different levels are mined to form a complete hierarchical knowledge system.

[0080] Implicit knowledge is difficult to identify in related technologies. Ancient Chinese medical records often contain implicit information such as the experience of renowned physicians, including the evolution of the underlying pathogenesis behind treatment methods and the compatibility strategies for prescriptions. This knowledge is extremely important for the inheritance of Chinese medicine, but it is not a direct record in the text; it requires deep reasoning to reveal. Existing NLP technologies focus on extracting surface information from texts and lack effective reasoning mechanisms, resulting in insufficient mining of the implicit diagnostic and treatment logic and empirical patterns in medical records. In this embodiment, implicit knowledge is mined through a dual-channel reasoning mechanism. A hierarchical community detection algorithm is used to perform pattern clustering and structured induction on the enhanced graph containing hypotheses, ultimately integrating high-frequency, strongly correlated patterns into an explicit knowledge set.

[0081] Optionally, in some embodiments, the step of obtaining the enhanced map based on the initial heterogeneous map through a dual-channel inference mechanism, i.e., step S32, can be applied for:

[0082] Step S321: Using the data-driven channel, calculate the first cosine similarity between the symptom vector and the potential syndrome vector in the initial heterogeneous graph, and / or calculate the second cosine similarity between the symptom vector and the treatment vector to obtain a list of potential hidden nodes.

[0083] Step S322: Using the rule channel and based on the TCM ontology database, execute the reasoning mechanism to filter the potential hidden node list and obtain the filtered node list.

[0084] Step S323: Take the intersection of the nodes filtered by the rule channel and the nodes generated by the data channel to determine the hidden nodes;

[0085] Step S324: Instantiate the hidden nodes and insert them into the tomographic diagnostic link of the initial heterogeneous graph to obtain the enhanced atlas.

[0086] Specifically, in step S321, the symptom set is transformed into an initial feature vector. The data-driven channel can utilize the GraphSage model to extract candidate symptom nodes and calculate the cosine similarity between the symptom vector and the symptom node, taking the Top-K as potential syndromes. For example, given the input symptoms {fever, chills, floating and tight pulse}, the similarity ranking is: wind-cold exterior excess (0.92) > wind-cold exterior deficiency (0.65) > wind-heat exterior syndrome (0.31), taking the Top-1, the potential syndrome is wind-cold exterior excess.

[0087] For step S322, the potential hidden node list is subjected to logical consistency screening, and the symptoms are mapped to micro-word logical expressions in the TCM ontology. The rule-driven channel traverses and matches the expressions based on the OWL reasoning mechanism of the TCM ontology. When the corresponding content is matched, a deterministic symptom result is output.

[0088] For step S323, compare the recommendation results of the data-driven channel with the derivation results of the rule-driven channel: if there is overlapping evidence, it is determined to be a hidden node.

[0089] For step S324, the determined hidden nodes obtained by dual-channel inference are structurally transformed to locate the fault links in the initial heterogeneous graph, complete the associated edges between nodes, and integrate to obtain the enhanced graph.

[0090] By using dual-channel reasoning results, diagnostic gaps in the graph are eliminated, forming a complete closed-loop path. The reasoning results are stored in the graph as exemplary nodes. As medical case data accumulates and the number of reasoning iterations increases, the nodes and related edges of the graph are continuously enriched, which is conducive to forming a self-iterative and self-improving TCM knowledge system.

[0091] Optionally, in some embodiments, applying a hierarchical community detection algorithm to the enhanced graph to obtain an explicit knowledge set, i.e., step S33 may include:

[0092] Step S331: Based on the enhanced graph, a hierarchical community algorithm is applied to construct a community graph with a multi-level structure; the community graph is used to display the association patterns and topic clustering features of TCM clinical knowledge.

[0093] Step S332: Discover hidden drug pairs from the bottom layer of the multi-level structure, discover stable prescription combinations from the middle layer of the multi-level structure, and discover macroscopic diagnosis and treatment patterns from the top layer of the structure.

[0094] Step S333 integrates the discovered implicit drug pairs, stable prescription groups, and macroscopic diagnosis and treatment models to obtain an explicit knowledge set.

[0095] Specifically, for steps S331-S333, the enhanced graph can be divided into 5 levels, including symptom layer, syndrome layer, treatment method layer, prescription layer, and traditional Chinese medicine layer. For nodes at different levels, differentiated community algorithms are selected to construct a tree-shaped community structure. The bottom community finds the node pairs with the highest connection weight and the most stable co-occurrence frequency as implicit drug pairs. The middle community groups drug pairs, symptoms, and syndromes together to form stable prescription groups. The top community clusters to mine macro-level medication styles and treatment patterns. The fragmented rules mined from the above three levels are uniformly encapsulated to form a structured explicit knowledge set. This set contains both micro-level information on how to combine drugs and macro-level treatment patterns, directly serving the subsequent clinical auxiliary decision-making system.

[0096] For example, assuming the hierarchical community detection algorithm constructs a community graph, and at the bottom layer it finds that "Ephedra" and "Cinnamon Twig" appear simultaneously, and "Apricot Kernel" and "Licorice" appear simultaneously, then the latent drug pairs [Ephedra-Cinnamon Twig] and [Apricot Kernel-Licorice] are extracted. At the middle layer it is found that [Ephedra-Cinnamon Twig] and [Apricot Kernel-Licorice] frequently appear together in the context of "aversion to cold, absence of sweating, and a floating and tight pulse," then the core formula group, Ephedra Decoction, is extracted as follows: Symptoms: Wind-cold exterior excess, Formula: [Ephedra, Cinnamon Twig, Apricot Kernel, Licorice]}. At the top layer it is found that the community containing the "Ephedra Decoction formula" is very close to the communities of "Cinnamon Twig Decoction" and "Major Blue Dragon Decoction," all belonging to a super-large community, then the macro-pattern of the Shanghan Lun - Taiyang Disease Chapter - sweating and exterior-releasing method is extracted.

[0097] By utilizing a hierarchical community discovery algorithm, a three-layer community graph was constructed from the bottom up. This hierarchical mining perfectly aligns with the theoretical system of "medicine-prescription-method" in Traditional Chinese Medicine (TCM), ensuring that the mined knowledge is no longer discrete fragments but a systematic and logically hierarchical knowledge network. Since the community discovery algorithm focuses on the topological density of the graph, it can automatically uncover the core medication logic and prescription ideas of physicians, discovering implicit compatibility patterns that are difficult to detect with the naked eye. This greatly enhances the clinical reference value of knowledge mining, transforming complex graph computation results into structured explicit knowledge sets, and translating obscure algorithmic clustering results into explicit knowledge sets readable by doctors, providing a high-quality structured knowledge collection.

[0098] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the scope of protection of this application.

[0099] Fourth embodiment

[0100] Some embodiments of this application also provide an electronic device. The electronic device can be various forms of digital computer, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, etc. The electronic device can also be various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices.

[0101] The electronic device includes: one or more processors; and a memory storing computer program instructions that, when executed, cause the processor to perform the steps of the methods provided in any one or more of the above embodiments. Figure 2An exemplary structural diagram of the electronic device is disclosed. The electronic device includes one or more processors 1101, a memory 1102, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components are interconnected via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations. The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0102] The electronic device may further include an input device 1103 and an output device 1104. The processor 1101, memory 1102, input device 1103 and output device 1104 may be connected by a bus or other means, as shown in the figure, which is connected by a bus.

[0103] Input device 1103 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device, such as a touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 1104 may include a display device, auxiliary lighting device (e.g., LED), and haptic feedback device (e.g., vibration motor). The display device may include, but is not limited to, a liquid crystal display, a light-emitting diode display, and a plasma display. In some embodiments, the display device may be a touch screen.

[0104] To provide interaction with the user, the electronic device can be a computer. The computer has: a display device (e.g., a cathode ray tube or LCD monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback); and input from the user can be received in any form (e.g., voice input or tactile input).

[0105] In this embodiment, a computer-readable medium stores a computer program / instructions that, when executed by a processor, implement the steps of the methods provided in any one or more of the above embodiments. This computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into that device. The aforementioned computer-readable medium carries one or more computer-readable instructions.

[0106] The memory 1102 can serve as a non-transitory computer-readable storage medium, used to store non-transitory software programs, non-transitory computer-executable programs, and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 1102, thereby implementing the program instructions / modules corresponding to the methods provided in any one or more of the embodiments described above in this application.

[0107] The memory 1102 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 1102 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 1102 may optionally include memory remotely located relative to the processor 1101, and these remote memories can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0108] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0109] Computer-readable media include permanent and non-permanent, removable and non-removable media, which can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technologies, read-only optical discs, digital versatile optical discs or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0110] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0111] In the above embodiments, all or part of the implementation can be achieved through software, hardware, firmware, or any combination thereof. For example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, magnetic or optical drives, floppy disks, and similar devices. In addition, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.

[0112] The computer program product provided in this application includes one or more computer programs / instructions. When executed by a processor, these computer programs / instructions generate, in whole or in part, the processes or functions described in this application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0113] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0114] The scope of this application is defined by the appended claims rather than the foregoing description, and is therefore intended to encompass all variations falling within the meaning and scope of equivalents of the claims. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device in software or hardware. Terms such as "first," "second," etc., are used only for distinguishing descriptions and do not indicate any particular order, nor should they be construed as indicating or implying relative importance.

[0115] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily made by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.

Claims

1. A method for making implicit knowledge in ancient Chinese medical records explicit and structured, characterized in that, include: Based on the original medical record data, a standardized medical record corpus was obtained; Based on the standardized medical record corpus, the triplet was determined; The triples are used to mine implicit knowledge; Based on the triples, an explicit knowledge set is obtained; the explicit knowledge set is used to represent the existing knowledge and implicit knowledge in traditional Chinese medicine medical records; Based on the explicit knowledge set, a knowledge graph of ancient Chinese medical records is obtained; Based on the aforementioned knowledge graph of ancient Chinese medical records, a complete integrated knowledge base of traditional Chinese medicine was obtained; Based on the integrated and complete TCM knowledge base, medical case knowledge services and visualization applications can be realized.

2. The method according to claim 1, characterized in that, The determination of triplets based on the standardized medical record corpus includes: A three-level dynamic semantic annotation system is established, which includes L0, L1 and L2 levels. Among them, L0 level is the basic entity layer, which is used to identify and annotate physically existing noun entities; L1 level is the clinical semantic layer, which is used to annotate L0 level entities to determine their clinical roles; L2 level is the implicit attribute layer, which is used to link the extracted L0 and L1 level entities to an external TCM ontology database and obtain the associated implicit attributes. Based on the three-level dynamic semantic annotation system, the annotation results of L0 and L1 levels were determined for the standardized medical record corpus execution model - expert collaborative annotation. Based on the annotation results of L0 and L1 levels, L2 level is attached to obtain the complete three-level annotation results; Based on the complete three-level annotation results, triples are extracted and determined; the triples include fact triples and attribute triples.

3. The method according to claim 2, characterized in that, The aforementioned standard medical record corpus execution model – expert collaborative annotation – determines the annotation results for L0 and L1 levels, including: Pre-labeling of L0 and L1 levels based on the fine-tuning model; Calculate the confidence entropy value for each pre-labeled entity; The labeling results for L0 and L1 levels are determined based on the confidence entropy value.

4. The method according to claim 3, wherein the confidence entropy value is calculated as follows: in, This represents the confidence entropy value. The average entropy of an entity is the mean of the token-level entropy, reflecting the average level of uncertainty of all tokens in an entity. A token represents the smallest text unit of an entity, that is, the single word or phrase fragment into which the entity is broken down in the labeled L0 and L1. This represents the minimum confidence level of an entity, which is the minimum probability of the largest category among all tokens of the entity, reflecting the confidence level of the entity's most uncertain token. This represents the balance coefficient, and its value range is... This is used to adjust the weighting of the entity's average entropy and minimum confidence level. This represents a single entity to be computed in L0 and L1.

5. The method according to claim 1, characterized in that, The explicit knowledge set obtained based on the triples includes: The data in the triples are used to construct an initial heterogeneous graph; the initial heterogeneous graph is used to reflect the explicit associations of the triples, and the nodes include symptoms, prescriptions, and the gaps in the diagnosis and treatment links between the two. Based on the initial heterogeneous graph, an enhanced atlas is obtained through a dual-channel inference mechanism; the enhanced atlas contains a complete diagnostic and treatment pathway. A hierarchical community discovery algorithm is applied to the enhanced knowledge graph to obtain an explicit knowledge set.

6. The method according to claim 5, characterized in that, The process of obtaining the enhanced graph based on the initial heterogeneous graph through a dual-channel inference mechanism includes: Using the data-driven channel, the first cosine similarity between the symptom vector and the potential syndrome vector in the initial heterogeneous graph is calculated, and / or the second cosine similarity between the symptom vector and the treatment vector is calculated to obtain a list of potential hidden nodes. Using rule channels and based on the TCM ontology database, an inference mechanism is executed to filter the potential hidden node list, resulting in a filtered node list. The hidden nodes are determined by taking the intersection of the nodes filtered by the rule channel and the nodes generated by the data channel. Instantiate hidden nodes and insert them into the tomographic diagnostic link of the initial heterogeneous graph to obtain an enhanced atlas.

7. The method according to claim 5, characterized in that, The application of a hierarchical community detection algorithm to the enhanced knowledge graph yields an explicit knowledge set, including: Based on the enhanced graph, a hierarchical community algorithm is applied to construct a community graph with a multi-level structure; the community graph is used to display the association patterns and topic clustering features of TCM clinical knowledge. Hidden drug pairs are extracted from the bottom layer of the multi-level structure, prescription groups are extracted from the middle layer of the multi-level structure, and macroscopic diagnosis and treatment patterns are extracted from the top layer of the structure. By integrating the discovered implicit drug pairs, prescription groups, and macroscopic diagnosis and treatment models, an explicit knowledge set is obtained.

8. An electronic device, characterized in that, The electronic device includes: Field-programmable gate arrays; and A memory storing configuration data, wherein the field-programmable gate array is configured by the configuration data to form hardware logic circuitry to perform the steps of the method as described in any one of claims 1 to 7.

9. A computer-readable medium storing configuration data thereon, characterized in that, When the configuration data is loaded into the field-programmable gate array (FPGA), the internal configuration of the FPGA forms a logic circuit, and the steps of the method as described in any one of claims 1 to 7 are executed.

10. A computer program product, comprising hardware description language code or a netlist, characterized in that, The hardware description language code or netlist is used to generate configuration data for configuring the field-programmable gate array to perform the steps of the method according to any one of claims 1 to 7.