Traditional Chinese medicine knowledge retrieval enhancement generation method and device
By constructing a TCM knowledge base and adopting a dual-path retrieval strategy, integrating information from the knowledge graph and clinical case database, and providing structured context for the large model, the knowledge limitations and illusion problems of LLM in TCM diagnosis are solved, and professional and accurate TCM diagnosis services are achieved.
Patent Information
- Application Number
- CN202510793414.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-26
AI Technical Summary
Existing large language models (LLMs) have knowledge limitations and hallucination problems when dealing with TCM problems, and are unable to provide professional, accurate and explainable TCM diagnosis services.
A traditional Chinese medicine knowledge base is constructed, including a traditional Chinese medicine knowledge graph and a clinical case vector database. A dual-path retrieval strategy is adopted to retrieve information from the knowledge graph and clinical case database, and organize it into a structured context as the input of the large model.
It effectively integrates TCM theoretical knowledge and practical experience to provide professional, accurate and explainable TCM diagnostic services, solving the knowledge limitations and illusion problems of LLM in dealing with TCM problems.
Smart Images

Figure CN120705315A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and device for enhancing the generation of traditional Chinese medicine knowledge retrieval. Background Art
[0002] In recent years, with the rapid development of artificial intelligence and natural language processing (NLP) technologies, large language models (LLMs) trained on massive corpora have been applied to a growing number of scenarios. For example, LLMs are being used in medical settings to address Traditional Chinese Medicine (TCM) issues. While general-purpose large models such as GPT and Claude typically possess strong contextual awareness and language understanding capabilities, they lack a strong grasp of TCM knowledge and, therefore, are unable to provide accurate diagnoses or recommendations consistent with TCM theory. Furthermore, LLMs are prone to hallucinations when addressing TCM issues, generating seemingly plausible content that is inconsistent with TCM theory and clinical facts.
[0003] In summary, LLMs have knowledge limitations and illusion problems when dealing with TCM issues, which leads to the inability to provide professional, accurate and explainable TCM diagnostic services. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method and device for enhancing the generation of traditional Chinese medicine knowledge retrieval, so as to solve the knowledge limitations and illusion problems existing in LLM when dealing with traditional Chinese medicine problems, so that LLM can provide professional, accurate and explainable traditional Chinese medicine diagnosis services.
[0005] The first aspect of the present invention provides a method for enhanced generation of traditional Chinese medicine knowledge retrieval, comprising the following steps: constructing a traditional Chinese medicine knowledge base, the traditional Chinese medicine knowledge base including a knowledge graph and a clinical case vector database in the field of traditional Chinese medicine; optimizing the user's original input to obtain an optimized query text; adopting a dual-path retrieval strategy, simultaneously performing knowledge graph retrieval and clinical case retrieval based on the knowledge graph and the clinical case vector database, retrieving and querying the most semantically relevant entities from the knowledge graph, and retrieving the cases most similar to the query from the clinical case database; and constructing a structured context based on the knowledge graph retrieval results and the clinical case retrieval results.
[0006] The second aspect of the present invention provides a TCM knowledge retrieval enhancement generation device, including: a knowledge base construction module, used to construct a TCM knowledge base, the TCM knowledge base includes a knowledge graph and a clinical case vector database in the field of TCM; a query text optimization module, used to optimize the user's original input to obtain an optimized query text; an information retrieval module, used to adopt a dual-path retrieval strategy, based on the knowledge graph and the clinical case vector database, to simultaneously perform knowledge graph retrieval and clinical case retrieval, retrieve and query the most semantically relevant entities from the knowledge graph, and retrieve the cases most similar to the query from the clinical case database; a context construction module, used to construct a structured context based on the knowledge graph retrieval results and the clinical case retrieval results.
[0007] The above-mentioned TCM knowledge retrieval enhancement generation method and device, after optimizing the user's original input to obtain the optimized query text, adopts a dual-path retrieval strategy, simultaneously retrieves information from the knowledge graph and clinical case vector database, and organizes the retrieved information into a structured context as the input of the large model reasoning. The structured context effectively integrates TCM theoretical knowledge and practical experience, provides comprehensive reference information for the large model, and can effectively solve the knowledge limitations and illusion problems that exist in LLM when dealing with TCM problems, enabling LLM to provide professional, accurate and explainable TCM diagnosis services. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 1 is a flow chart of a method for enhancing and generating traditional Chinese medicine knowledge retrieval in one embodiment of the present invention;
[0009] Figure 2 This is a flow chart of a method for constructing a knowledge graph in one embodiment of the present invention;
[0010] Figure 3 1 is a flow chart of a method for constructing a clinical medical record vector database in one embodiment of the present invention;
[0011] Figure 4 It is a structural diagram of a TCM knowledge retrieval enhancement generation device in one embodiment of the present invention. DETAILED DESCRIPTION
[0012] In order to enable those skilled in the art to more clearly understand the objectives, technical solutions and advantages of the present invention, the present invention is further described below with reference to the accompanying drawings and embodiments.
[0013] The enhanced generation method for traditional Chinese medicine knowledge retrieval provided by the present invention can be applied to LLM in scenarios such as intelligent health consulting services, primary medical auxiliary diagnosis, traditional Chinese medicine education and training, popularization of traditional Chinese medicine knowledge, telemedicine support, traditional Chinese medicine scientific research assistance, and health management platforms. LLM constructs a traditional Chinese medicine knowledge base including a knowledge graph in the field of traditional Chinese medicine and a clinical case vector database; optimizes the user's original input to obtain an optimized query text; adopts a dual-path retrieval strategy, and simultaneously performs knowledge graph retrieval and clinical case retrieval based on the knowledge graph and the clinical case vector database, retrieves and queries the most semantically relevant entities from the knowledge graph, and retrieves the cases most similar to the query from the clinical case database; constructs a structured context based on the knowledge graph retrieval results and the clinical case retrieval results as the input for large model reasoning. In the present invention, after optimizing the user's original input to obtain the optimized query text, a dual-path retrieval strategy is adopted to simultaneously retrieve information from the knowledge graph and the clinical case vector database, and the retrieved information is organized into a structured context as the input of the large model reasoning. The structured context effectively integrates the theoretical knowledge and practical experience of traditional Chinese medicine, provides comprehensive reference information for the large model, and can effectively solve the knowledge limitations and illusion problems that exist in LLM when dealing with traditional Chinese medicine problems, enabling LLM to provide professional, accurate and explainable traditional Chinese medicine diagnosis services.
[0014] like Figure 1 As shown, Figure 1 A schematic flow chart of a method for enhancing the generation of TCM knowledge retrieval provided by an embodiment of the present invention includes the following steps S10-S40:
[0015] S10: Build a TCM knowledge base, which includes a knowledge graph in the field of TCM and a clinical case vector database.
[0016] The goal of building a TCM knowledge base is to transform multi-source heterogeneous TCM data sources into a high-quality, structured knowledge base.
[0017] Multi-source heterogeneous TCM data sources are mainly divided into two categories:
[0018] (1) Classical textbooks and literature in the field of traditional Chinese medicine: including classic works of traditional Chinese medicine such as Huangdi Neijing and Shanghan Lun, as well as modern traditional Chinese medicine textbooks and academic papers. This type of data is usually characterized by strong theoretical and complete systems, but the formats are diverse and need to be processed uniformly.
[0019] (2) Clinical records: Case records from actual clinical practice, which have been strictly desensitized to ensure patient privacy. This type of data usually contains rich practical experience, but the structure and recording methods are diverse.
[0020] Among them, step S10, i.e., building a TCM knowledge base, includes the following steps:
[0021] S11: Construct a knowledge graph in the field of Traditional Chinese Medicine based on classic textbooks and literature in the field of Traditional Chinese Medicine.
[0022] S12: Construct a clinical medical record vector database based on clinical medical records.
[0023] For steps S11-S12, a knowledge graph in the field of traditional Chinese medicine is constructed based on classic textbooks and literature in the field of traditional Chinese medicine, which can provide data support for subsequent traditional Chinese medicine knowledge retrieval; a clinical case vector database is constructed based on clinical case records, which can provide data support for subsequent similar case retrieval; and furthermore, a data basis is provided for the subsequent structured contextual integration of traditional Chinese medicine theoretical knowledge and practical experience.
[0024] Figure 2 A flow chart of the knowledge graph construction method provided by the embodiment of the present invention is as follows: Figure 2 As shown, step S11, i.e., constructing a knowledge graph in the field of traditional Chinese medicine based on classic textbooks and literature in the field of traditional Chinese medicine, specifically includes the following steps:
[0025] S1101, Classic Textbooks and Literature Preprocessing:
[0026] The main goal of preprocessing classic textbooks and literature is to convert classic textbooks and literature in the field of traditional Chinese medicine into a structured and standardized data format for subsequent knowledge graph construction and retrieval. Specifically, the following steps are used to preprocess classic textbooks and literature in the field of traditional Chinese medicine:
[0027] Text cleaning: remove redundant information irrelevant to the core content, such as page numbers, table of contents, figure captions, and publication information;
[0028] Normalize paragraph separation: Ensure there are clear separators between paragraphs to facilitate subsequent text segmentation.
[0029] S1102, text segmentation:
[0030] Text segmentation of long texts is a key step in building a knowledge graph. Its goal is to divide long texts into semantically coherent and appropriately sized text blocks. The long texts are obtained by preprocessing classic textbooks and literature.
[0031] In order to solve the problem that traditional fixed-length chunking may cause semantic fragmentation, the present invention designs a dynamic chunking algorithm based on context judgment. The algorithm uses a queue Q to manage text paragraphs and designs a context judgment function C(p i ,p i-1 ) Determine the semantic relevance of adjacent paragraphs. In practice, this determination is performed using the semantic understanding capabilities of the LLM, using specially designed prompt words to guide the LLM in analyzing whether two paragraphs are in the same context.
[0032] For the paragraph collection in the queue, its total length needs to be controlled to not exceed the preset threshold:
[0033]
[0034] Among them, token is the basic unit of language model processing text, and the calculation method depends on the tokenizer used. The present invention adopts cl100k_base; L(Q) represents the total number of tokens in all paragraphs in queue Q, TokenCount(p i ) indicates paragraph p i The number of tokens, T threshold This is the preset upper limit for the total number of tokens, typically set to 1200, which is significantly smaller than the maximum context window of most LLMs. This is intended to ensure the stability and responsiveness of subsequent processing.
[0035] The core advantage of the dynamic chunking algorithm based on context judgment is that it can maintain semantic integrity, avoid semantic breaks that may be caused by traditional fixed-length chunking, and control the size of text chunks within the range that can be effectively processed by the large model.
[0036] The step S1102, that is, dividing the long text into blocks, specifically includes the following steps:
[0037] (1) Get the input long text as block input.
[0038] (2) Dividing the long text into paragraphs according to line break instructions. The line break instructions can be generated in response to line break characters (\r\n or \n) input by the user.
[0039] (3) Initialize an empty queue Q to temporarily store the paragraphs of the current text block.
[0040] (4) Traverse each paragraph in turn and obtain the token number of the current paragraph traversed.
[0041] (5) Determine whether the queue is empty. If so, execute step (6); if not, execute step (7).
[0042] (6) When the queue is empty, if the number of tokens in the current paragraph exceeds T threshold , then the current paragraph is truncated according to the number of tokens, with T threshold The current paragraph is divided into multiple sub-paragraphs with the upper limit of T. threshold Output multiple sub-segments one by one as independent text blocks, and output the token number not exceeding T threshold The sub-segment is added to the queue and jumps to step (8); if the number of tokens in the current paragraph does not exceed T threshold, then add the current paragraph to the queue and jump to step (8).
[0043] For step (6), if the number of tokens in the current paragraph exceeds T threshold , among the multiple sub-segments obtained by segmenting the current paragraph, except for the last sub-segment, the number of tokens is less than or equal to T threshold , the number of tokens in the remaining sub-segments is equal to T threshold When the number of tokens in the last sub-segment is equal to T threshold When the token number is T threshold The remaining sub-segments are output one by one as independent text blocks, and the number of tokens does not exceed T threshold The last sub-segment is added to the queue; when the number of tokens in the last sub-segment is equal to T threshold , all sub-segments will be output one by one as independent text blocks, and the queue will continue to remain an empty queue.
[0044] (7) When the queue is not empty, use the context judgment function C(p i ,p i-1 ) Determine whether the current paragraph is related to the last paragraph in the queue. If the context is relevant and the total number of tokens in the queue after the paragraph is added to the queue is L(Q)≤T threshold , then add the current paragraph to the queue and jump to step (8); if the context is relevant but the total number of tokens in the queue after the paragraph is added to the queue exceeds T threshold , or the context is irrelevant, the paragraph in the queue is output as an independent text block, and the queue is cleared and the process returns to step (5).
[0045] (8) If all paragraphs have been traversed, the paragraphs in the queue are output as an independent text block, ending the process of text segmentation for the long text; if all paragraphs have not been traversed, return to step (4) and continue traversing the paragraphs.
[0046] S1103, Entity and Relationship Extraction:
[0047] After the text is segmented, the next step is to extract entities and inter-entity relationships within each block of text, laying the foundation for knowledge graph construction. Entity and relationship extraction uses an LLM-based approach, using carefully designed prompt word templates to guide the LLM in identifying and extracting specialized concepts and their relationships within the text block.
[0048] The step S1103, i.e., extracting entities in the field of traditional Chinese medicine and relationships between entities from each text block, specifically includes the following steps:
[0049] (1) Define entity extraction function and relationship extraction function.
[0050] Specifically, for each text block T i , define the entity extraction function E and the relationship extraction function R:
[0051] E(T i )={e1,e2,…,e n}
[0052] R(T i )={r1,r2,…,r m}
[0053] Among them, each entity e j Represented by a triple:
[0054] e j =(name j ,description j ,{claim j1 ,claim j2 ,…,claim jk})
[0055] The meanings of the parameters in the formula are as follows:
[0056] name j : Entity names, such as "Huangqi", "Buqi", "Lung Deficiency" and other TCM concept names.
[0057] description j : Entity description, used to briefly define or explain the TCM concept.
[0058] {claim j1 ,claim j2 ,…,claim jk}: A set of claims or attributes related to the entity, such as "Astragalus is sweet and slightly warm in nature."
[0059] Each relation r l Represented by a triple:
[0060] r l =(subject l ,predicate l ,objective l}
[0061] subject l : The subject entity of the relationship, such as "Astragalus".
[0062] predicate l : Relationship type, such as "has efficacy", "treats", "belongs to", etc.
[0063] object l : The object entity of the relationship, such as "replenishing qi".
[0064] (2) Design a prompt word template.
[0065] To improve extraction quality, a few-shot learning strategy was employed, with 3-5 specific examples from the field of Traditional Chinese Medicine included in the prompt words as reference templates for LLM extraction. These examples should cover common TCM concept types (e.g., Chinese herbal medicine, syndrome type, prescription, etc.) and their relationship types (e.g., efficacy relationships, meridian relationships, composition relationships, etc.), to guide the LLM model to correctly identify similar patterns.
[0066] The prompt word template can also be adjusted according to the extraction effect. Specifically, by manually sampling and verifying the extraction results, when the extraction effect is not ideal, the prompt word template is adjusted to optimize the extraction effect to ensure the accuracy and completeness of entities and relationships.
[0067] (3) The prompt word template is used to guide the LLM to extract entities in the field of traditional Chinese medicine and the relationships between entities from each text block based on the entity extraction function and the relationship extraction function.
[0068] S1104, entity embedding generation:
[0069] For each entity e in the entity set extracted in step S1103 i , use a text embedding model that supports Chinese (such as Zhipu Qingyan's Embedding-3) to calculate the vector representation of its entity name and description:
[0070] v i =Emb(e i .name+":"+e i .description)
[0071] Among them, Emb represents the text embedding model, e i .name is the entity name, e i .description is the entity description, v i The generated entity vector representation (usually a 2048 or 1024-dimensional floating-point vector).
[0072] S1105. Build a preliminary knowledge graph:
[0073] We construct an initial knowledge graph G = (V, E) by using all extracted entities as nodes and relationships as edges, where V is the entity set containing all extracted TCM concepts, and E is the relationship set containing all extracted inter-entity relationships. The initial knowledge graph is initially stored in GraphML format.
[0074] S1106. Calculate the similarity matrix between entities:
[0075] For any two entities e i and e j , calculate its vector representation v i and v j The cosine similarity between:
[0076]
[0077] Among them, v i ·v j Represents vector dot product; ||v i || represents vector v i The Euclidean norm of ||v j || represents vector v j The Euclidean norm of the similarity value is in the range of [-1, 1], where a larger value indicates a more similar semantics.
[0078] S1107, HDBSCAN clustering:
[0079] Based on the similarity matrix, semantic clustering is performed using the HDBSCAN algorithm:
[0080] Clusters=HDBSCAN(sim,min_cluster_size,min_samples)
[0081] Among them, min_cluster_size is the minimum cluster size, usually set to 3, indicating that a cluster contains at least 3 entities; min_samples is the minimum number of core points, usually set to 2, used to determine density accessibility; sim is the similarity matrix between entities.
[0082] S1108. Select cluster representative entity:
[0083] For each cluster C k , select the entity most similar to the cluster center as the representative entity:
[0084]
[0085] The representative entity rep k It will serve as the central node of the cluster, representing a group of concepts with similar semantics.
[0086] S1109, Relationship Migration and Merging:
[0087] For cluster C k Each non-representative entity e in i, migrate all its related relations (incoming and outgoing edges) to the representative entity rep k , and merge relations of the same type. At the same time, the attributes and claims of non-represented entities are merged into the attribute set of the representative entity to retain complete semantic information.
[0088] S1110, build optimization map:
[0089] Based on the cluster representative entities and the migrated relations, an optimized knowledge graph G'=(V',E') is constructed, where V' contains all cluster representative entities and isolated entities that are not clustered, and E' contains the set of migrated and merged relations.
[0090] S1111, Atlas Storage:
[0091] The optimized knowledge graph is persistently stored in a graph database, such as Neo4j or ArangoDB, to support efficient graph query and reasoning.
[0092] S1112. Graph index creation:
[0093] An index is created for the vector representation of each node to support semantic-based similarity search and obtain the final knowledge graph.
[0094] The knowledge graph construction method in the embodiments of the present invention constructs an initial knowledge graph based on extracted entities and relationships, and optimizes it through semantic clustering. This effectively addresses the diversity of TCM terminology, reduces redundant nodes in the graph, and improves graph connectivity and retrieval efficiency. Furthermore, the clustering process preserves the semantic information of the original entities, ensuring knowledge integrity while reducing complexity.
[0095] Figure 3 A flow chart of a method for constructing a clinical case vector database provided by an embodiment of the present invention is shown as follows: Figure 3 As shown, step S12, i.e., constructing a clinical case vector database based on clinical case records, specifically includes the following steps:
[0096] S1201. Clinical medical record data preprocessing:
[0097] The goal of clinical medical record data preprocessing is to convert the original medical records into data objects in a standard format. Specifically, the collected clinical medical records are structured and converted into clinical medical records in JSON format.
[0098] The structured processing result contains two core fields:
[0099] (1) Context: Integrate the patient's basic information (such as age, gender, height, weight, medical history, etc.) and the main complaint;
[0100] (2) Diagnosis: Extract the diagnosis results of TCM doctors, including the diagnosis conclusion, basis, dialectical thinking and other reasoning processes.
[0101] An example of a clinical case record in JSON format is shown in the following table:
[0102]
[0103] S1202, context text conversion:
[0104] Convert the context part of clinical records in JSON format into structured context text to facilitate semantic understanding.
[0105] The following table shows an example of structured context text:
[0106]
[0107] S1203, Text Vectorization:
[0108] Calculate the embedding vector for the structured context text using a text embedding model consistent with the text embedding model used in the "Entity Embedding Generation" step in S1104:
[0109] vec_context=Emb(context_text)
[0110] Among them, context_text is the structured context text; vec_context is the generated embedding vector (usually a 2048 or 1024-dimensional floating-point vector); Emb is the selected text embedding model.
[0111] S1204. Vector database construction:
[0112] The generated embedding vectors and preprocessed clinical medical record data are stored in a vector database (such as Milvus) as shown in the following table to construct a clinical medical record vector database.
[0113]
[0114]
[0115] The vector field stores the embedded vector for similarity search; the metadata field stores the original context information and diagnosis results, which are returned as metadata of the retrieval results.
[0116] S1205, Index Construction:
[0117] Establish HNSW index for clinical case vector database to support fast nearest neighbor search.
[0118] S1206, Data persistence:
[0119] Complete the persistent storage of the clinical medical record vector database to ensure sustainable access and use of data.
[0120] Through the above steps S1201 to S1206, the clinical medical record data is converted into vector representation and stored in a modern vector database, and a clinical medical record vector database is constructed to provide efficient data support for subsequent similar case retrieval and LLM intelligent diagnosis.
[0121] S20: Optimize the user's original input to obtain an optimized query text.
[0122] The user's original input refers to the initial query information entering the system. In the TCM intelligent consultation scenario, the system integrates the following information sources as the user's original input:
[0123] (1.1) User current input: direct questions or descriptions entered by the user in the interactive interface, such as "I have been experiencing dizziness and tinnitus recently. What is the cause?", "What medicine can be taken together with Astragalus?", etc.
[0124] (1.2) Historical Conversation Records: If multiple rounds of conversation have already taken place, the system extracts key information from previous conversations to supplement the context and ensure semantic coherence and information integrity. Specifically, the system uses a dedicated prompt word to guide the large model to extract summaries from historical conversations.
[0125] (1.3) User basic information: includes the user's demographic characteristics and health information, mainly including age, gender, height and weight, permanent residence, medical history, current medication status, etc.
[0126] Step S20 is responsible for converting the user's original input into a more standardized and complete query text to improve the search quality. Step S20 uses LLM assistance to guide LLM to optimize the query through specially designed prompt words. Specifically, step S20 includes the following steps:
[0127] S21. Prompt word design:
[0128] Design structured prompt words, the elements of which are shown in the following table:
[0129]
[0130] S22. Information extraction and integration:
[0131] The structured prompts guide the LLM to extract key information from the original input. The extracted key information includes symptom descriptions and chief complaints, related physical manifestations, relevant information from past medical history, important information mentioned in historical conversations, and TCM terminology and its expression.
[0132] S23. Semantic Normalization:
[0133] The key information extracted by LLM is normalized at the semantic level, including converting colloquial expressions into professional terms (such as "red face" → "red face"), standardizing Chinese medicine terminology (such as "onset of heat" → "heat syndrome"), and supplementing key information that is implicit but not explicitly expressed.
[0134] S24. Query text generation:
[0135] Based on the key information extracted by LLM and semantically normalized, a concise and semantically complete query text is generated. This query text is the optimized query text.
[0136] Through steps S21 to S24 above, the user's original input can be optimized to obtain an optimized query text. This optimized query text retains the original semantics but is more standardized and precise. It has sufficient context, complete semantics, standardized entities, and clear intent, facilitating subsequent case retrieval and reasoning. The following table shows an example of an optimized query text obtained by optimizing the user's original input.
[0137]
[0138] S30: A dual-path retrieval strategy is adopted to simultaneously perform knowledge graph retrieval and clinical case retrieval based on the knowledge graph and clinical case vector database, retrieve and query the most semantically relevant entities from the knowledge graph, and retrieve the cases most similar to the query from the clinical case database.
[0139] After obtaining the optimized query text, a dual-path retrieval strategy is adopted to retrieve information from the knowledge graph and clinical case database simultaneously, balancing the reference value of theoretical knowledge and practical experience.
[0140] Specifically, the step S30 includes steps S31 to S32:
[0141] S31. Query vectorization:
[0142] Before performing a search, the optimized query text needs to be vectorized. The specific steps include S311 to S312:
[0143] S311. Perform necessary preprocessing on the optimized query text, including word segmentation, removal of stop words, etc.
[0144] S312. Perform vectorization processing on the optimized query text to generate a vector representation of the query text:
[0145] Q vec =Emb(Q′)
[0146] Among them, Q' is the optimized query text, Q vec is the vector representation of the query text, and Emb is the text embedding model. To ensure consistency in the vector space, the text embedding model used in this step is consistent with the text embedding model used in the "Entity Embedding Generation" step in S1104 and the "Text Vectorization" step in S1104.
[0147] S32, dual-path parallel retrieval: Based on the knowledge graph and clinical case vector database, knowledge graph retrieval and clinical case retrieval are performed simultaneously according to the vector representation of the query text:
[0148] (1) Knowledge graph retrieval: Retrieve and query the most semantically relevant entities from the knowledge graph:
[0149] E rel =TopK(KG,Q′,k1)
[0150] (2) Clinical medical record retrieval: Retrieve the cases most similar to the query from the clinical medical record database:
[0151] C rel =TopK(CaseDB,Q',k2)
[0152] Among them, KG is a knowledge graph, which contains entities and their relationships in the field of traditional Chinese medicine; CaseDB is a clinical case vector database; Q vec The vector representation of the query text; k1 is the threshold for knowledge graph entity retrieval, usually set to 10-20, indicating that the top k1 entities with the highest similarity are returned; k2 is the threshold for clinical case retrieval, usually set to 5-10, indicating that the top k2 cases with the highest similarity are returned; TopK means that in the retrieval process, the top K most similar entries are found in the clinical case vector database or knowledge graph based on cosine similarity; E rel C is a set of related entities retrieved from the knowledge graph, used to provide relevant knowledge references; rel It is a collection of retrieved related medical records, used to provide diagnosis and treatment plans for similar cases.
[0153] S40: Build structured context based on knowledge graph retrieval results and clinical case retrieval results.
[0154] Step S40 is responsible for organizing the retrieved information into a structured context for use by the large model. The context construction algorithm considers the relevance, completeness, and diversity of the information, ensuring that the large model can obtain comprehensive and high-quality reference information.
[0155] Specifically, the step S40 includes steps S41 to S42:
[0156] S41. Build a structured context. The building process can be expressed as:
[0157] Context=ContextBuilder(Q′,E rel ,C rel ,H,I)
[0158] Among them, Q' represents the optimized query text; E rel Represents the set of related entities retrieved from the knowledge graph; C rel Represents a collection of related medical records retrieved from the clinical medical record vector database; H represents historical conversation records; I represents basic user information; ContextBuilder represents a context construction function that organizes a structured context based on input information; Context represents the constructed structured context.
[0159] The structured context consists of the following parts:
[0160] (1) Optimized query text Q': contains the normalized representation of the user's current question, which serves as the core part of the context and guides the large model to understand the user's intention.
[0161] (2) Entity information E retrieved from the knowledge graph info : Entities retrieved from the knowledge graph and their related information:
[0162] E info ={(e.name,e.description,e.claims,e.relations)|e∈E rel}
[0163] Among them, e.name is the entity name, such as "Huangqi", "Liver Yang Hyperactivity", etc.; e.description is a brief description of the entity; e.claims is a set of claims or attributes related to the entity; e.relations is a set of relationships between the entity and other entities, including relationship types and related entities; E rel is the set of related entities retrieved from the knowledge graph.
[0164] (3) Medical record information C retrieved from the clinical medical record database info:Similar case information retrieved from the clinical case database:
[0165] C info ={(c.context,c.diagnosis)|c∈C rel}
[0166] Among them, c.context is the context information of the case, including the patient's basic information, symptom description, etc.; c.diagnosis is the diagnosis result and treatment plan of the case; C rel It is a collection of related cases retrieved from the clinical medical record database.
[0167] (4) Historical conversation record H: The historical conversation between the user and the system, which helps the big model understand the position of the current conversation in the entire interaction.
[0168] (5) User basic information I: Contains the user’s demographic characteristics and health information, including age, gender, height and weight, permanent residence, medical history, current medication status, etc., to help the model understand the user’s background information.
[0169] S42. Use a preset strategy to organize the component information of the structured context.
[0170] The preset strategies include a relevance priority strategy and an information block strategy.
[0171] Relevance priority strategy: put the information most relevant to the query in front and provide it to the large model for reference first.
[0172] Information segmentation strategy: Clearly separate information by category to improve the processing efficiency of large models, and use specific XML markup symbols to clearly distinguish different types of information.
[0173] Through steps S41 and S42 above, the information retrieved from the knowledge graph and clinical case vector database can be organized into a structured context. This structured context effectively integrates theoretical knowledge and practical experience, providing comprehensive reference information. This structured organization also improves the understanding and utilization efficiency of large models. The following table provides an example of a structured context.
[0174]
[0175] It can be seen that in the above scheme, after optimizing the user's original input to obtain the optimized query text, a dual-path retrieval strategy is adopted to simultaneously retrieve information from the knowledge graph and the clinical case vector database, and organize the retrieved information into a structured context as the input of the large model reasoning. The structured context effectively integrates the theoretical knowledge and practical experience of traditional Chinese medicine, provides comprehensive reference information for the large model, and can effectively solve the knowledge limitations and illusion problems of LLM when dealing with traditional Chinese medicine problems, enabling LLM to provide professional, accurate and explainable traditional Chinese medicine diagnosis services.
[0176] In one embodiment, a TCM knowledge retrieval enhancement generation device is provided, which corresponds to the TCM knowledge retrieval enhancement generation method in the above embodiment. Figure 4 As shown, the TCM knowledge retrieval enhancement generation device includes a knowledge base construction module 10, a query text optimization module 20, an information retrieval module 30, and a context construction module 40. The functional modules are described in detail as follows:
[0177] Knowledge base building module 10, for executing Figure 1 Step S10 in the TCM knowledge retrieval enhancement generation method in the illustrated embodiment is used to construct a TCM knowledge base, which includes a knowledge graph in the field of TCM and a clinical case vector database.
[0178] Query text optimization module 20, for executing Figure 1 Step S20 in the enhanced generation method for TCM knowledge retrieval in the illustrated embodiment is used to optimize the user's original input to obtain an optimized query text.
[0179] Information retrieval module 30, for executing Figure 1 Step S30 in the enhanced generation method for traditional Chinese medicine knowledge retrieval in the illustrated embodiment is used to adopt a dual-path retrieval strategy, simultaneously perform knowledge graph retrieval and clinical case retrieval based on the knowledge graph and the clinical case vector database, retrieve and query the most semantically relevant entities from the knowledge graph, and retrieve the cases most similar to the query from the clinical case database.
[0180] Context building module 40, used to execute Figure 1 Step S40 in the enhanced generation method for TCM knowledge retrieval in the illustrated embodiment is used to construct a structured context based on the knowledge graph retrieval results and the clinical case retrieval results.
[0181] The present invention provides a TCM knowledge retrieval enhancement generation device. After optimizing the user's original input to obtain an optimized query text, it adopts a dual-path retrieval strategy to simultaneously retrieve information from the knowledge graph and clinical case vector database, and organizes the retrieved information into a structured context as the input of the large model reasoning. The structured context effectively integrates TCM theoretical knowledge and practical experience, providing comprehensive reference information for the large model, which can effectively solve the knowledge limitations and illusion problems of LLM when dealing with TCM problems, enabling LLM to provide professional, accurate and explainable TCM diagnosis services.
[0182] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Those skilled in the art may make various equivalent changes and improvements based on the above embodiment. Any equivalent changes or modifications made within the scope of the claims shall fall within the scope of protection of the present invention.
Claims
1. A method for enhancing the generation of traditional Chinese medicine knowledge retrieval, characterized in that: The steps include: Build a TCM knowledge base, which includes a knowledge graph and a clinical case vector database in the field of TCM; Optimize the user's original input to obtain optimized query text; A dual-path retrieval strategy is adopted to simultaneously perform knowledge graph retrieval and clinical case retrieval based on the knowledge graph and clinical case vector database, retrieving and querying the most semantically relevant entities from the knowledge graph, and retrieving the cases most similar to the query from the clinical case database; Build structured context based on knowledge graph retrieval results and clinical case retrieval results.
2. The TCM knowledge retrieval enhancement generation method according to claim 1, characterized in that: The construction of the TCM knowledge base comprises the following steps: Construct a knowledge graph in the field of Traditional Chinese Medicine based on classic textbooks and literature in the field of Traditional Chinese Medicine; Construct a clinical case vector database based on clinical case records.
3. The TCM knowledge retrieval enhancement generation method according to claim 2, characterized in that: The construction of a knowledge graph in the field of traditional Chinese medicine based on classic textbooks and literature in the field of traditional Chinese medicine includes the following steps: Preprocessing of classic textbooks and literature: Preprocessing of classic textbooks and literature in the field of traditional Chinese medicine; Text segmentation: segment long text into chunks. The long text is the text obtained by pre-processing classic textbooks and literature. Entity and relationship extraction: Extract entities in the field of traditional Chinese medicine and relationships between entities from each text block; Entity embedding generation: For each entity in the entity set extracted in the entity and relationship extraction step, a text embedding model that supports Chinese is used to calculate the vector representation of its entity name and description; Construct a preliminary knowledge graph: Use all extracted entities as nodes and relationships as edges to construct an initial knowledge graph G = (V, E), where V is the entity set and E is the relationship set; Calculate the similarity matrix between entities: For any two entities, calculate the cosine similarity between the vector representations of the two entities; HDBSCAN clustering: Based on the similarity matrix, the HDBSCAN algorithm is used for semantic clustering; Select cluster representative entities: For each cluster C k , select the entity most similar to the cluster center as the representative entity; Relationship migration and merging: For cluster C k For each non-representative entity in the , migrate all its related relationships to the representative entity and merge the relationships of the same type. At the same time, merge the attributes and declarations of the non-representative entity into the attribute set of the representative entity; Constructing an optimized graph: Based on the cluster representative entities and the migrated relationships, construct an optimized knowledge graph G' = (V', E'), where V' contains all cluster representative entities and isolated entities that are not clustered, and E' contains the set of migrated and merged relationships. Graph storage: Persistently store the optimized knowledge graph in a graph database to support efficient graph query and reasoning; Graph indexing: Create an index for the vector representation of each node to support semantic-based similarity search and obtain the final knowledge graph.
4. The TCM knowledge retrieval enhancement generation method according to claim 3, characterized in that: The method of extracting entities in the field of traditional Chinese medicine and relationships between entities from each text block includes the following steps: Define entity extraction functions and relationship extraction functions; Design a prompt word template; The LLM is guided by the prompt word template to extract entities in the field of traditional Chinese medicine and relationships between entities from each text block based on the entity extraction function and the relationship extraction function.
5. The method for enhancing the generation of TCM knowledge retrieval according to claim 3, characterized in that: The method of dividing a long text into blocks comprises the following steps: (1) Get the input long text as block input; (2) Divide long text into paragraphs according to line break instructions; (3) Initialize an empty queue; (4) Traverse each paragraph in turn and obtain the token number of the current paragraph; (5) Determine whether the queue is empty. If so, execute step (6); if not, execute step (7); (6) When the queue is empty, if the number of tokens in the current paragraph exceeds T threshold , where T threshold If the preset upper limit of the total number of tokens is set, the current paragraph will be truncated according to the number of tokens, with T threshold The current paragraph is divided into multiple sub-paragraphs with the upper limit of T. threshold Output multiple sub-segments one by one as independent text blocks, and output the token number not exceeding T threshold The sub-segment is added to the queue and jumps to step (8); if the number of tokens in the current paragraph does not exceed T threshold , then add the current paragraph to the queue and jump to step (8); (7) When the queue is not empty, use the context judgment function C(p i ,p i-1 ) Determine whether the current paragraph is related to the last paragraph in the queue. If the context is relevant and the total number of tokens in the queue after the paragraph is added to the queue is L(Q)≤T threshold , then add the current paragraph to the queue and jump to step (8); if the context is relevant but the total number of tokens in the queue after the paragraph is added to the queue exceeds T threshold , or the context is irrelevant, then the paragraph in the queue is output as an independent text block, and the queue is cleared and the process returns to step (5); (8) If all paragraphs have been traversed, the paragraphs in the queue are output as an independent text block, ending the process of text segmentation for the long text; if all paragraphs have not been traversed, return to step (4) and continue traversing the paragraphs.
6. The method for enhancing the generation of TCM knowledge retrieval according to claim 2, characterized in that: The method of constructing a clinical medical record vector database based on clinical medical records includes the following steps: Clinical medical record data preprocessing: Structural processing of collected clinical medical records, converting original clinical medical records into clinical medical records in JSON format; Context text conversion: convert the context part in the clinical medical record in JSON format into structured context text; Text vectorization: Use the text embedding model to calculate the embedding vector for the structured context text: vec_context=Emb(context_text) Among them, context_text is the structured context text, vec_context is the generated embedding vector, and Emb is the text embedding model; Vector database construction: The generated embedded vectors and pre-processed clinical case data are stored in the vector database to construct the clinical case vector database; Index construction: Build an HNSW index for the clinical case vector database to support fast nearest neighbor search; Data persistence: Complete persistent storage of the clinical medical record vector database to ensure sustainable access and use of data.
7. The method for enhancing the generation of TCM knowledge retrieval according to claim 1, characterized in that: The original input includes the user's current input, historical conversation information, and basic user information. The optimizing process for the user's original input to obtain an optimized query text includes the following steps: Prompt word design: design structured prompt words; Information extraction and integration: Through the structured prompt words, the LLM is guided to extract key information from the original input. The extracted key information includes symptom descriptions and chief complaints, related physical manifestations, relevant information in the past medical history, important information mentioned in the historical conversation, and TCM professional terms and expressions; Semantic normalization: Normalize the key information extracted by LLM at the semantic level, including converting colloquial expressions into professional terms, standardizing TCM terminology, and supplementing key information that is implicit but not explicitly expressed; Query text generation: Based on the key information extracted by LLM and semantically normalized, a concise and semantically complete query text is generated. This query text is the optimized query text.
8. The method for enhancing the generation of TCM knowledge retrieval according to claim 1, characterized in that: The dual-path retrieval strategy is used to simultaneously perform knowledge graph retrieval and clinical case retrieval based on the knowledge graph and the clinical case vector database, retrieve and query the most semantically relevant entities from the knowledge graph, and retrieve the cases most similar to the query from the clinical case database, including the following steps: Query vectorization: Vectorize the optimized query text to generate a vector representation of the query text: Q vec =Emb(Q′) Among them, Q' is the optimized query text, Q vec is the vector representation of the query text, and Emb is the text embedding model; Dual-path parallel retrieval: Based on the knowledge graph and clinical case vector database, knowledge graph retrieval and clinical case retrieval are performed simultaneously based on the vector representation of the query text. The most semantically relevant entities are retrieved from the knowledge graph, and the cases most similar to the query are retrieved from the clinical case database. E rel NTopK(KG,Q′,k1) C rel =TopK(CaseDB,Q',k2) Among them, KG is a knowledge graph, which contains entities and their relationships in the field of traditional Chinese medicine; CaseDB is a clinical case vector database; Q vec The vector representation of the query text; k1 is the number threshold for knowledge graph entity retrieval, indicating that the top k1 entities with the highest similarity are returned; k2 is the number threshold for clinical case retrieval, indicating that the top k2 cases with the highest similarity are returned; TopK means that in the retrieval process, the top K most similar entries are found in the clinical case vector database or knowledge graph based on cosine similarity; E rel is the set of related entities retrieved from the knowledge graph; C rel A collection of relevant medical records retrieved.
9. The method for enhancing the generation of TCM knowledge retrieval according to claim 1, characterized in that: The step of constructing a structured context based on the knowledge graph retrieval results and the clinical case retrieval results includes the following steps: Build a structured context. The construction process can be expressed as: Context=ContextBuilder(Q′,E rel ,C rel ,H,I) Among them, Q' represents the optimized query text, E rel represents the set of related entities retrieved from the knowledge graph, C rel Represents the relevant medical record set retrieved from the clinical medical record vector database, H represents the historical conversation record; I represents the basic information of the user, ContextBuilder represents the context building function, and Context represents the constructed structured context; A preset strategy is used to organize the component information of the structured context, which includes the optimized query text, entity information retrieved from the knowledge graph, medical record information retrieved from the clinical medical record vector database, historical conversation records and basic user information.
10. A TCM knowledge retrieval enhancement generation device, characterized in that: include: The knowledge base construction module is used to build a TCM knowledge base, which includes a knowledge graph in the field of TCM and a clinical case vector database; The query text optimization module is used to optimize the user's original input to obtain the optimized query text; An information retrieval module, which uses a dual-path retrieval strategy to simultaneously perform knowledge graph retrieval and clinical case retrieval based on the knowledge graph and clinical case vector database, retrieve and query the most semantically relevant entities from the knowledge graph, and retrieve the cases most similar to the query from the clinical case database; The context construction module is used to build structured context based on the knowledge graph retrieval results and clinical case retrieval results.
Citation Information
Cited By
Organ transplantation clinical aid decision-making method, device and equipment based on AI large model, medium and product
CN121096601A
Medication guidance system and method based on knowledge graph and vector retrieval fusion
CN121388238A
Medication guidance system and method based on knowledge graph and vector retrieval fusion
CN121388238B
Medical database information management system and method
CN121434379A
Knowledge graph and large model dual-drive-based traditional Chinese medicine diagnosis and treatment cooperation method
CN122245648A