Traditional Chinese medicine intelligent inquiry method and system based on knowledge graph and medical case enhanced RAG

By adopting RAG technology based on knowledge graph and medical case enhancement in the intelligent consultation system, the problem of complex diagnosis and treatment logic processing in traditional Chinese medicine is solved, the accuracy and personalization of diagnosis and treatment suggestions are improved, and the system's knowledge update and dynamic adaptability are realized.

CN120108694APending Publication Date: 2025-06-06JIANGSU UNIV
View PDF 0 Cites 12 Cited by

Patent Information

Application Number
CN202510163939.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06

Smart Images

  • Figure CN120108694A_ABST
    Figure CN120108694A_ABST
Patent Text Reader

Abstract

The invention discloses a traditional Chinese medicine intelligent inquiry method and system based on a knowledge graph and medical case enhancement RAG, and the method comprises the following steps: collecting traditional Chinese medicine related data from a plurality of sources, and carrying out the data cleaning, formatting and standardization processing, and constructing a knowledge graph containing traditional Chinese medicine symptoms, disease causes, prescriptions, and the mutual relation of the symptoms, the disease causes, the prescriptions; expressing a knowledge structure in the graph in a triple form; searching a local sub-graph related to query based on query keyword extraction and generalization by utilizing a Leiden community detection algorithm; a mixed retrieval strategy is adopted, global search and local search are combined, recalled contents are sorted and scored, global and local retrieval results are fused, and a large language model is uniformly output. The system aims at improving the intelligent level and individuation ability of traditional Chinese medicine diagnosis, firstly, a knowledge graph covering entities such as traditional Chinese medicine theories, symptoms, pathogenesis and prescriptions and relationships of the entities is constructed, mass traditional Chinese medicine data are systematically integrated, deep understanding and semantic association of traditional Chinese medicine complex diagnosis and treatment logic are ensured, and the traditional Chinese medicine diagnosis and treatment efficiency is improved. Through deep combination of the structured knowledge base and the generated model, the diagnosis and treatment accuracy is improved, and the model fine tuning and updating cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence and medical information, and in particular to a method and system for intelligent Chinese medicine consultation based on knowledge graph and medical record enhanced RAG. Background Art

[0002] As an important branch of traditional Chinese medicine, TCM has a diagnosis and treatment system that covers a complex theoretical framework including Yin and Yang, Five Elements, Zang-Fu and Meridians, Qi, Blood and Body Fluids. TCM diagnosis and treatment emphasizes individualized syndrome differentiation and treatment methods, relying on the doctor's rich experience and understanding of the complex relationship between causes and symptoms. However, with the rapid development of modern medical technology, how to combine TCM theory with information technology to improve the efficiency and accuracy of TCM diagnosis and treatment has become an important research direction. The complexity of TCM diagnosis and treatment methods makes the development of an intelligent consultation system suitable for TCM challenging.

[0003] At present, intelligent medical consultation systems use technologies such as natural language processing (NLP) and machine learning (ML) to simulate the interaction between doctors and patients and provide diagnosis and treatment recommendations. However, most of these systems are used in the field of Western medicine and usually rely on rule bases or simple retrieval technologies. This approach is inadequate when dealing with the complex diagnosis and treatment logic of traditional Chinese medicine. The construction and maintenance of rule bases is a heavy task that requires a lot of manual annotation and rule definition, and it is difficult to cover all possible causes and symptom combinations, especially the many-to-many mapping relationship emphasized by traditional Chinese medicine. In addition, the existing systems are not flexible enough to deal with individual differences and the needs of personalized diagnosis and treatment in traditional Chinese medicine, and it is difficult to provide accurate diagnosis and treatment recommendations.

[0004] In recent years, large language models (LLMs) based on deep learning have made significant progress in the field of natural language processing. These models are able to generate semantically rich texts through training on large-scale corpora and perform well in a variety of language tasks. However, due to the particularity of TCM theory, general language models show obvious deficiencies in processing TCM diagnosis and treatment logic. In order to make up for this deficiency, research on combining domain knowledge with language models has gradually emerged.

[0005] As an effective knowledge management and representation method, knowledge graph technology has gradually been applied to the field of traditional Chinese medicine. Knowledge graphs represent knowledge as nodes and edges through graph structures, which can intuitively display the complex relationships between symptoms, causes, and prescriptions of traditional Chinese medicine. This structured representation not only helps to improve the accuracy of diagnosis and treatment recommendations, but also enhances the interpretability of the system. The introduction of knowledge graphs enables the intelligent consultation system to not only rely on simple keyword matching when retrieving relevant information, but also generate more personalized and accurate diagnosis and treatment recommendations through semantic understanding and reasoning.

[0006] The application of retrieval-augmented generation (RAG) technology in intelligent medical consultation systems has gradually attracted attention. RAG technology improves the performance of the system in specific fields by combining knowledge retrieval and language generation. Under the RAG framework, the system first retrieves information related to the user input from the knowledge graph, and then generates diagnosis and treatment recommendations based on the retrieval results and language model. This method shows high flexibility and accuracy in dealing with the complex diagnosis and treatment logic of traditional Chinese medicine. However, with the continuous updating of clinical research and medical knowledge in the field of traditional Chinese medicine, how to achieve dynamic updating of knowledge has become a key issue. The introduction of dynamic graph technology provides the possibility for real-time updating of knowledge, ensuring that the system always provides diagnosis and treatment recommendations based on the latest medical research. Summary of the invention

[0007] The purpose of this section is to summarize some aspects of embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the specification abstract and the invention title of this application to avoid blurring the purpose of this section, the specification abstract and the invention title, and such simplifications or omissions cannot be used to limit the scope of the present invention.

[0008] In view of the above-mentioned problems existing in the existing TCM intelligent consultation method and system based on knowledge graph and medical record enhanced RAG, the present invention is proposed.

[0009] Therefore, the purpose of the present invention is to provide a method and system for intelligent diagnosis of traditional Chinese medicine based on knowledge graph and medical record enhanced RAG, which aims to integrate the knowledge of symptoms, causes, prescriptions, etc. in the field of traditional Chinese medicine, and realize accurate reasoning and personalized diagnosis and treatment suggestions for complex diagnosis and treatment logic of traditional Chinese medicine through dynamically updated medical record data and large language models. This system not only improves the accuracy and personalization level of diagnosis and treatment suggestions, but also enhances the system's ability in knowledge updating and dynamic adaptation, providing a new technical path for intelligent diagnosis and treatment of traditional Chinese medicine.

[0010] In order to solve the above technical problems, the present invention provides the following technical solutions: a TCM intelligent consultation method based on knowledge graph and medical record enhanced RAG, comprising the following steps:

[0011] Step 1: Collect TCM-related data from various sources, and after data cleaning, formatting and standardization, construct a knowledge graph containing TCM symptoms, causes, prescriptions and their relationships; represent the knowledge structure in the graph in the form of triples;

[0012] Step 2: Using the Leiden community detection algorithm, based on query keyword extraction and generalization, search for local subgraphs related to the query. The Leiden algorithm recursively narrows the search scope from top to bottom, identifies multiple subcommunities related to the user query and their associated entities, and extracts keywords from the query entered by the user to guide the subgraph retrieval process.

[0013] Step 3: Format the recalled local subgraph data into text and submit it to the big model for processing together with the question entered by the user to generate a response. The big model combines the contextual information in the knowledge graph to generate diagnosis and treatment recommendations.

[0014] Step 4: The back-end module uses the Django framework to process data and transmits the data to the front-end module through the API interface. The ORM function of the Django framework can automatically handle database operations. The specific steps include data reception, parsing and storage, transmitting data through the API interface, and supporting the transmission of multiple data formats to ensure the flexibility and scalability of the system.

[0015] As a preferred solution of the TCM intelligent consultation method based on knowledge graph and medical record enhanced RAG described in the present invention, the specific steps of graph construction in step 1 are as follows:

[0016] Step 1.1: The system collects data from professional data sets, authoritative data sources and real medical records in the field of traditional Chinese medicine. The data needs to be cleaned, denoised and standardized before constructing the knowledge graph. Data cleaning removes duplicate, erroneous and incomplete records in the data set; data denoising filters irrelevant noise data and retains valid information; data standardization converts data into a unified format and standard to ensure data consistency;

[0017] Step 1.2: Extract entities and relations in the field of traditional Chinese medicine from the preprocessed data;

[0018] Step 1.3: Build the extracted entities and relationships into a graph structure and store it in the Neo4j graph database.

[0019] As a preferred solution of the TCM intelligent consultation method based on knowledge graph and medical record enhanced RAG described in the present invention, the specific implementation method of the Leiden community detection algorithm in step 2 includes:

[0020] Step 2.1: Graph summary label description and preliminary concept extraction: By performing preliminary concept extraction on the entities and their relationships in the knowledge graph, the graph summary label is generated. The label describes the core features of each entity and relationship and its semantic role in the field of traditional Chinese medicine, such as the concepts of disease, symptoms, treatment methods, and drugs. Based on the graph summary label, the system divides the knowledge graph into multiple structured subgraphs, each of which represents a specific field or related topic in the graph.

[0021] Step 2.2: Recursive community detection and layer-by-layer narrowing of search scope: Use the Leiden community detection algorithm to recursively process the divided graph subgraphs; the Leiden algorithm narrows the search scope layer by layer from top to bottom, gradually refining it to a finer subgraph level;

[0022] Step 2.3: Multi-level entity association extraction and deep semantic analysis: Obtain entity-related information through the following steps:

[0023] Entity activation and information extraction: For each entity activated in the recall subgraph, the system extracts its detailed information, including the basic description of the entity, related symptoms, treatment methods, and drug effects;

[0024] Basic medical knowledge association: The system further extracts basic medical knowledge related to the entity. If the query is about a certain disease, the system not only extracts the symptoms and treatment methods of the disease, but also obtains medical knowledge about the cause, epidemiology, and related examinations of the disease.

[0025] Semantic associations between entities: The system identifies semantic associations between entities by analyzing the relationships between entities in the knowledge graph;

[0026] Top-K related entities and their contents: The system not only focuses on the query entities themselves, but also extracts the Top-K entities and their contents that are highly related to these entities; these Top-K entities represent other concepts or entities that are most relevant to the user's query.

[0027] As a preferred solution of the TCM intelligent consultation method based on knowledge graph and medical record enhanced RAG described in the present invention, the specific process of the Leiden algorithm includes:

[0028] S1, initialization phase: starting from the global structure of the knowledge graph, perform preliminary community division and identify relevant subgraph regions;

[0029] S2, recursive processing stage: According to the density of nodes and edges in each level, these areas are recursively divided more finely; through this top-down approach, the system can narrow the search scope at each level and screen out the most closely related subcommunities.

[0030] As a preferred solution of the TCM intelligent consultation method based on knowledge graph and medical record enhanced RAG described in the present invention, wherein: when constructing the TCM knowledge graph in the step 1, a hybrid strategy combining character separation and subject-based segmentation is used to segment semantic documents, and element extraction is performed with the help of an entity recognition method prompted by a large language model, to construct a knowledge graph that can accurately reflect the subject semantic structure and associated information, providing a solid knowledge foundation for TCM consultation. The specific implementation method of the hybrid strategy is:

[0031] First, use static characters to preliminarily divide the document into paragraphs;

[0032] Secondly, further topic segmentation is performed based on the semantic features of the text. By improving the proposition transfer algorithm, independent sentences are extracted from the original text, and paragraphs are converted into self-sufficient proposition units through semantic analysis.

[0033] During processing, the system evaluates each proposition through sequential analysis and decides to merge it with an existing block or generate a new block; it uses a sliding window technique to process five paragraphs at a time, keeping the focus on the topic by gradually removing and adding paragraphs;

[0034] Finally, a hard threshold is set to ensure that the length of each block does not exceed the context limit of the large language model. After the document is segmented, a knowledge graph is built on each independent data block to ensure that each graph can accurately reflect the semantic structure and related information of its corresponding topic.

[0035] As a preferred solution of the TCM intelligent consultation method based on knowledge graph and medical record enhanced RAG described in the present invention, the TCM intelligent consultation method based on knowledge graph and medical record enhanced RAG also includes an information filtering step, filtering the questions input by the user based on the BERT text filter, and the specific method of the information filtering step is:

[0036] Step A: define the set Q of all questions that can be input into the big model, the set of questions answered by the big model in a certain professional field is R, and the set of questions that generate professional answers is D. Obviously, Q>R>D;

[0037] Step B: Use the filter to make Q→R, ensuring that the question is within the range of R;

[0038] Step C, information filtering will ensure that the system answers questions within the system's capabilities to reduce the generation of hallucinations.

[0039] The intelligent TCM consultation system based on knowledge graph and medical record enhanced RAG includes: knowledge graph construction module, RAG technology application module, back-end module and front-end module;

[0040] The knowledge graph construction module is used to extract entities, relationships and attributes from TCM-related data, and construct them into a graph structure and store them in a graph database. The knowledge graph construction module includes but is not limited to data source selection, data preprocessing, and entity recognition units to ensure the accuracy and completeness of the knowledge graph;

[0041] The RAG technology application module includes a triple extraction unit, a subgraph recall unit and a subgraph context generation unit, which performs efficient information retrieval through the Leiden community detection algorithm to respond to user queries. The RAG technology application module also includes but is not limited to an information filtering unit to limit the scope of questions that the large model can answer;

[0042] The backend module is built using the Django framework and provides an API interface to implement data transmission and processing functions. The backend module also includes but is not limited to database design and data model definition units to ensure high performance and scalability of the system;

[0043] The front-end module is responsible for displaying information to users through page rendering and interacting with the back-end module. The front-end module also includes but is not limited to user interface design and interaction logic optimization unit to ensure a user-friendly interaction experience.

[0044] As a preferred solution of the TCM intelligent consultation system based on knowledge graph and medical case enhanced RAG described in the present invention, the data source of the knowledge graph construction module also includes real consultation medical cases, which are derived from real consultation medical cases provided by laboratory cooperative hospitals and are highly reliable and professional. The data set is constructed in the form of question-answer pairs, covering a wide range of TCM diagnosis and treatment scenarios, symptoms and treatment plans. These question-answer pairs come from actual clinical TCM consultation cases and have undergone strict medical review, collation and de-identification to ensure the accuracy and completeness of the data while complying with privacy standards.

[0045] Beneficial effects of the present invention:

[0046] 1. Improve the accuracy of diagnosis and treatment: Through the deep integration of knowledge graph and RAG technology, the system can accurately understand the user's symptoms and needs, generate more accurate diagnosis and treatment suggestions, and improve the accuracy and reliability of TCM consultation. The knowledge graph provides a rich background of TCM knowledge, and RAG technology combines the advantages of retrieval and generation, which can quickly retrieve relevant information in large-scale data and generate accurate diagnosis and treatment suggestions.

[0047] 2. Enhanced personalized services: The system can provide personalized diagnosis and treatment plans and drug recommendations based on individual differences and specific symptoms of users, meet the needs of different users, and improve user experience. Through user feedback, the system continuously optimizes the knowledge graph and RAG model to improve the personalization of diagnosis and treatment.

[0048] 3. Dynamic update of knowledge base: The dynamic update mechanism of the knowledge graph can absorb the latest TCM research results and clinical cases in real time, ensuring that the system always makes diagnoses and recommendations based on the latest knowledge, and maintaining the advancement and practicality of the system. The system regularly obtains the latest TCM research results and clinical cases from authoritative data sources and clinical cases, enters them into the knowledge graph, and automatically adjusts the nodes and relationship structures in the graph through dynamic graph technology.

[0049] 4. Reduce development costs: Compared with traditional fine-tuning methods, knowledge graph-enhanced RAG technology significantly reduces system development and operation costs, reduces dependence on large-scale annotated data and high-performance computing resources, and improves the system's scalability and application value. The knowledge graph provides a structured, searchable knowledge base, allowing the system to obtain accurate information through knowledge retrieval during consultation, without having to rely on the model to learn and generate all data end-to-end. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. Among them:

[0051] Figure 1 A schematic diagram of the process of the community summary of the present invention;

[0052] Figure 2 It is a schematic diagram of the multi-path search and recall process of the present invention;

[0053] Figure 3 A schematic diagram of the process of enhancing large model retrieval according to the present invention;

[0054] Figure 4 It is a schematic diagram of the system technical structure of the present invention. DETAILED DESCRIPTION

[0055] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0056] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0057] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0058] Secondly, the present invention is described in detail with reference to the schematic diagram. When describing the embodiments of the present invention in detail, for the sake of convenience, the cross-sectional diagrams showing the device structure will not be partially enlarged according to the general scale, and the schematic diagrams are only examples, which should not limit the scope of protection of the present invention. In addition, in actual production, the three-dimensional dimensions of length, width and depth should be included.

[0059] Reference Figure 1 - Figure 4 , provides a TCM intelligent consultation method based on knowledge graph and medical record enhanced RAG, including the following steps:

[0060] Step 1: Data construction and preprocessing

[0061] The implementation of this system requires the collection and organization of professional data sets to support the implementation of the system. The data sets and knowledge bases required by the system are constructed based on a variety of data, and data preprocessing is performed on these data. At present, the data sources of our system are mainly divided into the following three categories:

[0062] ① Based on the existing professional field data sets, directly collect the existing relevant data sets in the professional field, refer to their composition, and sort and filter out the required data. For the field of traditional Chinese medicine prescriptions, refer to the MedDialog, CBLUE, COMET, and CMeKG data sets to sort and construct relevant professional data.

[0063] ② Authoritative data, which is collected from professional books or authoritative websites. This part of the data comes from professional books and authoritative websites in related fields, and is used to build a knowledge base to provide professional knowledge support for the answers of the large model. For the field of traditional Chinese medicine prescriptions, a professional knowledge base of traditional Chinese medicine prescriptions is mainly built based on professional books such as prescriptions. At the same time, relevant data knowledge in the field of traditional Chinese medicine prescriptions is collected from professional authoritative websites such as NMPA (National Medical Products Administration), Yaorongyun-Traditional Chinese Medicine Database Group, TCMID Traditional Chinese Medicine Database, and Traditional Chinese Medicine Syndrome Association Database.

[0064] ③ Real medical records (the core data set of this system) are highly reliable and professional. This data set is constructed in the form of question-answer pairs, containing more than 20 million high-quality question-answer pairs, covering a wide range of TCM diagnosis and treatment scenarios, symptoms and treatment plans. These question-answer pairs come from actual clinical TCM consultation cases, and have undergone strict medical review, collation and de-identification to ensure the accuracy and completeness of the data while complying with privacy standards. Our team converted these clinical consultation records into a TCM knowledge base, and further constructed a refined TCM knowledge graph through data modeling and analysis.

[0065] Step 2: Information filtering

[0066] For question-and-answer in professional fields, the large language model does not need to answer questions in other fields. For this reason, this system adds a text filter based on BERT (Bidirectional Encoder Representations from Transformers) to filter questions to limit the range of questions that the large model can answer. Other models often produce subtle hallucinations and generate incorrect text when facing boundary or cross-cutting questions in professional fields. Assume that the set of all questions that can be input into the large model is Q, the set of questions that the large model can answer in a certain professional field is R, and the set of questions that can generate professional answers is D. Obviously, Q>R>D. Using fine-tuning to limit will make R→D, which will weaken the model's answering ability. Using the form of a filter to make Q→R will try to ensure that the questions asked are within the range of R. Although some data outside R will enter the large model, since the professional enhanced question-and-answer system designed by this system still retains a certain general ability, questions outside R can also be answered without professional verification. Information filtering will ensure that this system answers questions within the system's capabilities as much as possible to reduce the possibility of hallucinations.

[0067] Step 3: Construction of TCM Atlas

[0068] ①Semantic document segmentation

[0069] In order to effectively handle multiple topics or diverse content in large medical documents, this study first segmented them into data blocks that meet the contextual constraints of large language models (LLMs). Traditional segmentation methods based on token size or fixed characters often fail to accurately identify subtle changes in topics, which may lead to insufficient capture of contextual intent and affect the complete expression of semantics. To improve segmentation accuracy, our team proposed a hybrid strategy that combines character segmentation with topic-based segmentation. Specifically, the document is first preliminarily divided into paragraphs using static characters (such as line breaks). Subsequently, further topic segmentation is performed based on the semantic features of the text. We extract independent sentences from the original text by improving the proposition transfer algorithm, and transform paragraphs into self-sufficient propositional units through semantic analysis. During the processing, the system evaluates each proposition through sequential analysis and decides to merge it with an existing block or generate a new block. This decision process is based on the zero-shot method of LLM to ensure topic consistency. At the same time, to reduce the noise introduced by sequential processing, our team adopted a sliding window technique to process five paragraphs at a time, keeping the focus on the topic by gradually removing and adding paragraphs. In addition, a hard threshold is set to ensure that the length of each block does not exceed the contextual constraints of LLM. After completing the document segmentation, our team built a knowledge graph on each independent data block to ensure that each graph can accurately reflect the semantic structure and related information of its corresponding topic. This method not only improves the accuracy of segmentation, but also ensures the integrity and consistency of knowledge graph construction in downstream tasks.

[0070] ②Element extraction

[0071] In the process of identifying and extracting graph node instances from each source text block, an entity recognition method based on large language model (LLM) prompts was adopted. LLM identified all relevant entities in the text through prompts and extracted them in a structured manner. For each entity, LLM generates the name, type and description of the entity. The name can be the original text in the document or the derived terms commonly used in the medical context. These terms are carefully selected to ensure compliance with the terminology standards in the professional field of traditional Chinese medicine for subsequent processing. The type is selected by LLM from a predefined category table, and the description is the entity interpretation generated by LLM based on the document context to ensure semantic accuracy and consistency. To ensure the reliability of the extraction results and the generation quality of the model, we provide LLM with carefully designed examples to guide it to generate expected output. In the data structure of each entity, we attach a unique identifier (ID) so that its corresponding source document and specific paragraph can be accurately tracked. This ID is crucial in generating evidence-based responses in the subsequent stage, ensuring the traceability and consistency of information. At the same time, in order to improve the quality of entity extraction and reduce noise and variance, the extraction process is repeated multiple times. This iterative approach can prompt LLM to identify entities that may be initially overlooked. After each iteration, LLM decides whether to proceed to the next iteration to ensure dynamic adjustment and high efficiency of the extraction process.

[0072] Step 4: Community summary

[0073] In this system, the implementation of graph community summarization includes two key steps: community discovery and community summarization. In order to optimize the retrieval efficiency and improve the diagnosis and treatment accuracy of the system, Figure 1 As shown in the figure, the system uses the Leiden algorithm to discover communities and achieves efficient processing and reasoning of complex TCM diagnosis and treatment knowledge by making a comprehensive semantic summary of community subgraphs. Specifically, it includes the following contents:

[0074] ① Community discovery (based on Leiden algorithm)

[0075] The Leiden algorithm is a graph community discovery algorithm based on modularity optimization, which is particularly suitable for the partitioning and analysis of large-scale graph data. The algorithm generates a series of high-density subcommunities by gradually optimizing the node allocation in the graph, that is, the nodes in the graph are closely related, while the connection between communities is relatively weak. In the TCM consultation system, the knowledge graph covers entities such as Chinese medicinal materials, symptoms, causes, prescriptions, and their interrelationships. The Leiden algorithm is used to identify the close communities of these entities and relationships in order to perform more accurate consultation reasoning.

[0076] ② Modularity optimization and community division

[0077] The Leiden algorithm divides the graph structure by maximizing the modularity within the community. In the TCM knowledge graph, each entity (such as "Danggui" and "Spleen Deficiency") is a node, and the relationship between each entity (such as "treatment" and "compatibility") is an edge. The Leiden algorithm identifies highly associated node groups based on the frequency and intensity of interaction between these entities and relationships, forming multiple subcommunities. For example, in the knowledge graph of spleen deficiency, a subcommunity composed of entities such as "Spleen Deficiency", "Qi Deficiency", and "Spleen-Strengthening Drugs" may be formed, clearly showing the diagnosis and treatment system of the symptom.

[0078] ③Adaptive update and dynamic partitioning

[0079] Since the TCM knowledge graph is dynamically changing (for example, with the addition of new clinical studies and medical records), the improved Leiden algorithm has good adaptability and can automatically adjust the community division according to the incremental update of the graph structure to ensure that new knowledge can be integrated into the existing community in a timely manner. This dynamic division is crucial in the TCM consultation system, which can support the needs of continuous knowledge expansion and maintain the accuracy and stability of the community structure.

[0080] ④ Community Summary

[0081] After community discovery, the system needs to further summarize these subcommunities so that it can efficiently retrieve and generate treatment recommendations in the subsequent consultation process. The purpose of community summarization is to compress and summarize the entities and their relationships within the community at the semantic level, providing LLM with a simplified but information-rich subgraph representation. This step is assisted by combining prompt words with graph algorithms, allowing LLM to accurately identify community topics and generate highly relevant diagnostic content.

[0082] ⑤ Importance marking of graphic elements

[0083] To prevent LLM from failing due to information redundancy when processing large-scale subcommunities, the system uses graph algorithms such as PageRank to mark key nodes in the graph (such as Chinese medicines with important therapeutic effects or key symptoms) to ensure that key entities and relationships are retained when generating summaries. For example, in a community involving the treatment of "Qi deficiency", PageRank may highlight important nodes such as Angelica sinensis and Astragalus membranaceus, and ignore minor nodes on the edge, thereby optimizing the generation efficiency of LLM.

[0084] ⑥ Prompt word guidance and semantic summary

[0085] When performing community summaries, the system uses preset prompts to guide LLM to understand the structure of the community subgraph and generate a comprehensive summary based on the entity relationships within it. The prompts clearly instruct LLM how to handle entities and their relationships in the field of traditional Chinese medicine (such as drug compatibility, main diseases, etc.), and provide guidance through one-shot examples, so that LLM can generate accurate summaries more efficiently and help with retrieval during subsequent consultations. The challenge of this step is how to guide LLM to retain key community information as much as possible so that a more comprehensive community summary can be obtained during global searches. After multiple rounds of optimization by the team, the prompt template with excellent results is shown below:

[0086] ⑦ Streaming data acquisition and incremental reasoning

[0087] Since the data scale of community subgraphs is difficult to control and the community structure may be frequently updated (such as the addition of new medical cases or research results), the system adopts streaming data acquisition and incremental reasoning. Streaming processing ensures that the LLM does not exceed the context window limit, while incremental reasoning ensures that the system can quickly update the graph summary and accurately capture the latest Chinese medicine diagnosis information. In the core implementation of the graph community summary, the build_communities method of CommunityStore is referenced, and the adapter community_store_adapter provides implementation abstractions for different graph databases, mainly including the call entry discover_communities of the community discovery algorithm and the community details query entry get_community. Among them, the community summarizer community_summarizer is responsible for calling the large language model (LLM) to summarize the community subgraph, while the community metadata storage meta_store implements the storage and retrieval of community summaries based on the vector database.

[0088] Step 5: Multi-way search recall

[0089] Global search mainly relies on the community metadata storage built into the system. Figure 2As shown in the figure, after the community summary information of the graph is discovered by the Leiden algorithm, it is saved in the community metadata storage as the entry point for global retrieval. Unlike the traditional full-volume scan or MapReduce secondary summary strategy, the system greatly reduces the computational overhead and query latency by directly indexing and searching the community metadata during global search. Local search is achieved by traversing the subgraph of the TCM knowledge graph. The specific process is that the system extracts relevant keywords from the user's query, applies these keywords to the local knowledge graph subgraph, traverses the nodes and edges related to it, and returns the content that best matches the user's query. In order to improve the overall user experience and retrieval effect, this system adopts a hybrid retrieval strategy. By combining global search and local search, the system will call two paths at the same time when encountering user queries, and provide a more comprehensive response based on the retrieval results. This strategy can avoid the retrieval limitations of over-reliance on a single path, ensuring that the system can retrieve relevant content to the greatest extent even in complex or ambiguous query scenarios. In hybrid retrieval, the system calls global and local retrieval paths at the same time. The global search is based on the community information in the community metadata storage, calling the interface _community_store#search_communities for efficient global retrieval; while the local search extracts the keywords in the query through _keyword_extractor#extract, and combines _graph_store#explore to traverse and retrieve the knowledge graph subgraph. Through this parallel multi-way recall, the system can ensure that even if the global search does not hit, the local search can still play a role, and ultimately provide users with more comprehensive diagnostic support. After the multi-way recall is completed, the system organizes and scores the recalled content through the CommunitySummaryKnowledgeGraph (such as similar_search_with_scores), merges the global and local search results, and outputs them in a unified manner. This ensures that LLM can make full use of all retrieved relevant information when generating the final diagnosis and treatment recommendations, thereby improving the comprehensiveness and depth of diagnosis.

[0090] Among them, during use, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention is described in detail with reference to the preferred embodiments, ordinary technicians in the field should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention, which should be included in the scope of the claims of the present invention.

Claims

1. The intelligent TCM consultation method based on knowledge graph and medical record enhanced RAG is characterized by: The following steps are involved: Step 1: Collect TCM-related data from various sources, and after data cleaning, formatting and standardization, construct a knowledge graph containing TCM symptoms, causes, prescriptions and their relationships; represent the knowledge structure in the graph in the form of triples; Step 2: Based on the Leiden community detection algorithm, the query keyword extraction and generalization are used to search for local subgraphs related to the query. The Leiden algorithm recursively narrows the search scope from top to bottom, identifies multiple subcommunities related to the user query and their associated entities, and extracts keywords from the query entered by the user to guide the subgraph retrieval process; Step 3: Format the recalled local subgraph data into text and submit it to the big model for processing together with the question entered by the user to generate a response. The big model combines the contextual information in the knowledge graph to generate diagnosis and treatment recommendations. Step 4: The back-end module uses the Django framework to process data and transmits the data to the front-end module through the API interface. The ORM function of the Django framework can automatically handle database operations. The specific steps include data reception, parsing and storage, transmitting data through the API interface, and supporting the transmission of multiple data formats to ensure the flexibility and scalability of the system.

2. The intelligent TCM consultation method based on knowledge graph and medical record enhanced RAG according to claim 1, characterized in that: The specific steps of constructing the map in step 1 are as follows: Step 1.1: The system collects data from professional data sets, authoritative data sources and real medical records in the field of traditional Chinese medicine. The data needs to be cleaned, denoised and standardized before constructing the knowledge graph. Data cleaning removes duplicate, erroneous and incomplete records in the data set. Data denoising filters out irrelevant noise data and retains valid information; data standardization converts data into a unified format and standard to ensure data consistency; Step 1.2: Extract entities and relations in the field of traditional Chinese medicine from the preprocessed data; Step 1.3: Build the extracted entities and relationships into a graph structure and store it in the Neo4j graph database.

3. The intelligent TCM diagnosis method based on knowledge graph and medical record enhanced RAG according to claim 2 is characterized by: The specific implementation method of the Leiden community detection algorithm in step 2 includes: Step 2.1: Entity and relationship summary label description and preliminary concept extraction: By performing preliminary concept extraction on the entities and their relationships in the knowledge graph, summary labels of the entities and relationships are generated. The labels describe the core features of each entity and relationship and its semantic role in the field of traditional Chinese medicine, such as the concepts of diseases, symptoms, treatment methods, and drugs. Based on the summary labels of entities and relationships, the system divides the knowledge graph into multiple structured subgraphs, each of which represents a specific field or related topic in the graph. Step 2.2: Recursive community detection and layer-by-layer narrowing of search scope: Use the Leiden community detection algorithm to recursively process the divided graph subgraphs; the Leiden algorithm narrows the search scope layer by layer from top to bottom, gradually refining it to a finer subgraph level; Step 2.3: Multi-level entity association extraction and deep semantic analysis: Obtain entity-related information through the following steps: Entity activation and information extraction: For each entity activated in the recall subgraph, the system extracts its detailed information, including the basic description of the entity, related symptoms, treatment methods, and drug effects; Basic medical knowledge association: The system further extracts basic medical knowledge related to the entity. If the query is about a certain disease, the system not only extracts the symptoms and treatment methods of the disease, but also obtains medical knowledge about the cause, epidemiology, and related examinations of the disease. Semantic associations between entities: The system identifies semantic associations between entities by analyzing the relationships between entities in the knowledge graph; Top-K related entities and their contents: The system not only focuses on the query entities themselves, but also extracts the Top-K entities and their contents that are highly related to these entities; these Top-K entities represent other concepts or entities that are most relevant to the user's query.

4. The intelligent TCM diagnosis method based on knowledge graph and medical record enhanced RAG according to claim 3 is characterized by: The specific process of the Leiden algorithm includes: S1, initialization phase: starting from the global structure of the knowledge graph, perform preliminary community division and identify relevant subgraph regions; S2, recursive processing stage: According to the density of nodes and edges in each level, these areas are recursively divided more finely; through this top-down approach, the system can narrow the search scope at each level to facilitate the screening of the most closely related subcommunities.

5. The intelligent TCM consultation method based on knowledge graph and medical record enhanced RAG according to claim 1 is characterized by: When constructing the TCM knowledge graph in step 1, a hybrid strategy combining character separation and subject-based segmentation is used to segment semantic documents, and element extraction is performed with the help of entity recognition methods prompted by large language models, so as to construct a knowledge graph that can accurately reflect the subject semantic structure and related information, and provide a solid knowledge foundation for TCM consultation. The specific implementation method of the hybrid strategy is as follows: First, use static characters to preliminarily divide the document into paragraphs; Secondly, further topic segmentation is performed based on the semantic features of the text. By improving the proposition transfer algorithm, independent sentences are extracted from the original text, and paragraphs are converted into self-sufficient proposition units through semantic analysis. During processing, the system evaluates each proposition through sequential analysis and decides to merge it with an existing block or generate a new block; it uses a sliding window technique to process five paragraphs at a time, keeping the focus on the topic by gradually removing and adding paragraphs; Finally, a hard threshold is set to ensure that the length of each block does not exceed the context limit of the large language model. After the document is segmented, a knowledge graph is built on each independent data block to ensure that each graph can accurately reflect the semantic structure and related information of its corresponding topic.

6. The intelligent TCM diagnosis method based on knowledge graph and medical record enhanced RAG according to claim 1, characterized in that: The intelligent TCM consultation method based on knowledge graph and medical record enhanced RAG also includes an information filtering step, filtering the questions input by the user based on the BERT text filter. The specific method of the information filtering step is: Step A: Define the set Q of all questions that can be input into the big model. The set of questions answered by the big model in a certain professional field is R, and the set of questions that generate professional answers is D. Obviously, Q>R>D; Step B: Use the filter form to make Q→R, ensuring that the question asked is within the range of R; Step C, information filtering will ensure that the system answers questions within the system's capabilities to reduce the generation of hallucinations.

7. The intelligent TCM consultation system based on knowledge graph and medical records enhanced RAG is characterized by: include: Knowledge graph construction module, RAG technology application module, back-end module and front-end module; The knowledge graph construction module is used to extract entities, relationships and attributes from TCM-related data, and construct them into a graph structure and store them in a graph database. The knowledge graph construction module includes but is not limited to data source selection, data preprocessing, and entity recognition units to ensure the accuracy and completeness of the knowledge graph; The RAG technology application module includes a triple extraction unit, a subgraph recall unit and a subgraph context generation unit, and performs efficient information retrieval based on the Leiden community detection algorithm to respond to user queries. The RAG technology application module also includes but is not limited to an information filtering unit to limit the scope of questions that the large model can answer; The backend module is built using the Django framework and provides an API interface to implement data transmission and processing functions. The backend module also includes but is not limited to database design and data model definition units to ensure high performance and scalability of the system; The front-end module is responsible for displaying information to users through page rendering and interacting with the back-end module. The front-end module also includes but is not limited to user interface design and interaction logic optimization unit to ensure a user-friendly interaction experience.

8. The intelligent TCM consultation system based on knowledge graph and medical record enhanced RAG according to claim 7 is characterized by: The data source of the knowledge graph construction module also includes real medical records. The real medical records are derived from real medical records provided by the laboratory's cooperative hospitals and are highly reliable and professional. The dataset is constructed in the form of question-answer pairs, covering a wide range of TCM diagnosis and treatment scenarios, symptoms and treatment plans. These question-answer pairs come from actual clinical TCM consultation cases and have undergone strict medical review, organization and de-identification to ensure the accuracy and completeness of the data while complying with privacy standards.

Citation Information

Cited By

  • Pathology knowledge base construction method and device, electronic equipment and storage medium

    CN120409660A

  • Knowledge graph recall-based agent question and answer method, device, equipment and product

    CN120578749A

  • Fine-grained evaluation method and system for clinical diagnosis capability of large language model

    CN120878269A

  • Large model security management method based on organization isolation and authority control

    CN120930179A

  • Joint retrieval method and system based on knowledge enhancement, medium and terminal

    CN121166936A