Novel large model traditional Chinese medicine course resource intelligent question answering method and system

CN120144706AActive Publication Date: 2025-06-13INST OF INFORMATION ON TRADITIONAL CHINESE MEDICINE CACMS

Patent Information

Application Number
CN202510212037.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The existing Chinese medicine course resource Q&A system has problems such as insufficient flexibility, limited coverage, limitations in semantic understanding, low interpretability and transparency, and potential "illusion" phenomena, and it is difficult to effectively deal with user complexity and innovation problems.

Method used

The new large-scale model traditional Chinese medicine course resource intelligent question-and-answer method is adopted to receive user questions through the large-scale model, embed and vector search, and combine Neo4j knowledge graph and Elasticsearch search engine to mine and filter entities, relationships and heteronouns, generate comprehensive solutions, and ensure the accuracy and reliability of the answers through verification mechanisms.

Benefits of technology

It improves the accuracy and professionalism of the answers, enhances the flexibility and coverage of the system, improves semantic understanding and response capabilities, ensures the timeliness and reliability of information, increases interpretability and transparency, and optimizes the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144706A_ABST
    Figure CN120144706A_ABST
Patent Text Reader

Abstract

The invention provides a novel large model traditional Chinese medicine course resource intelligent question answering method and system, and the method comprises the steps: enabling a large model to convert a user question into a question vector, and carrying out the similarity retrieval matching in a Fast vector library, and constructing a first reference answer; performing entity relationship and different noun mining on the question by the large model, performing knowledge graph filtering, analyzing the real intention of the question asked by the user, generating a corresponding Cypher statement according to the filtered entity and relationship, executing the Cypher statement in Neo4j, and constructing a second reference answer; performing fine word segmentation on the questions by the large model by using a predefined word segmentation device, executing full-text search in a pre-constructed Elasticsearch engine index kernel, sorting search results according to a characteristic sorting strategy, and constructing a third reference answer; and the large model carries out knowledge summarization and generation according to the first, second and third reference answers, strictly screens data reference sources through a verification mechanism, and generates the most accurate comprehensive answer. Intelligent question answering can be performed on traditional Chinese medicine course resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent question answering for traditional Chinese medicine course resources, and in particular to a novel intelligent question answering method and system for traditional Chinese medicine course resources based on a large model. Background Art

[0002] With the progress of the times and the continuous development and iteration of computer technology, the methods for using computers to answer questions about specific traditional Chinese medicine course resources are also constantly innovating. Currently, there are mainly the following two solutions: the question answering solution for traditional Chinese medicine course resources based on a rule base and the question answering solution for traditional Chinese medicine course resources based on the RAG large model.

[0003] The core idea of the first question answering solution for traditional Chinese medicine course resources based on a rule base is to carefully construct a detailed question-answer rule base for the selected course content by educational experts or teacher teams. However, this solution has many problems. For example, the flexibility is limited: the questions and answers in the rule base are predefined, and for unconventional or innovative questions raised by users, the system may not be able to give appropriate answers, and this limitation is particularly obvious when facing open-ended or discussion-type questions; the coverage is limited: even if the rule base is as detailed as possible, it is impossible to cover all possible questions, and the ways users ask questions vary greatly. Some questions may not match the appropriate answers due to different expressions, resulting in a poor user experience; the limitation of semantic understanding: although the system can perform semantic parsing on the user's questions, there may still be misunderstandings when dealing with complex or ambiguous questions, especially when users use non-standard languages or dialects, the system's understanding ability may be affected. The adaptability is poor: for rapidly developing disciplinary fields, such as information technology or medicine, the update speed of the rule base may not keep up with the update frequency of knowledge, thus affecting the timeliness and effectiveness of the system.

[0004] The second Q&A solution for traditional Chinese medicine (TCM) course resources based on the RAG large model adopts the advanced Retrieval-Augmented Generation (RAG) technology, aiming to provide an efficient, accurate, and in-depth Q&A mechanism for TCM course resources. However, this solution also has many problems, such as the limitations of semantic understanding: Although the RAG model can capture the meaning and context at the lexical level, its semantic understanding ability may still be insufficient when dealing with unconventional expressions, metaphors, or culture-specific content. For complex natural language phenomena, the performance of the model may be inferior to that of human experts. Low interpretability and transparency: The internal working principle of large language models is complex, resulting in the difficulty of explaining their decision-making process, which is particularly important in the educational scenario because teachers and students may hope to understand the logic and basis behind the answers. Potential "hallucination" phenomenon: Language models may produce the so-called "hallucinations", that is, generate seemingly reasonable but actually wrong information. In this case, the answers provided by the system may be misleading, especially when it refers to incompletely relevant or incorrect data. Summary of the Invention

[0005] In order to solve the above technical problems existing in the prior art, the present invention provides a new intelligent Q&A method for TCM course resources based on a large model, and the technical solution is as follows:

[0006] On the one hand, a new intelligent Q&A method for TCM course resources based on a large model is provided, and the method includes:

[0007] S1. The large model receives a question raised by the user for TCM course resources;

[0008] S2. The large model embeds the question into words, converts it into a question vector, performs similarity retrieval and matching of the question vector in the Fass vector library, and constructs a first reference answer according to the retrieval result;

[0009] S3. The large model mines entities, relationships, and synonyms of the question according to the dynamically constructed knowledge graph, positive and negative synonym tables, and characteristic positive and negative synonym detection algorithms, filters the mined entities, relationships, and synonyms of the knowledge graph, only retains the valid entities and relationships existing in the knowledge graph, analyzes the true intention of the user's question, determines the application scenario of graph query or graph calculation, automatically generates corresponding Cypher statements according to the filtered entities and relationships, executes them in Neo4j, combines the user's question and the execution result, generates a natural language answer, and constructs a second reference answer;

[0010] S4. The large model performs fine-grained word segmentation on the problem using a predefined word segmenter, performs full-text search in the pre-constructed Elasticsearch search engine index core, sorts the search results according to the sorting strategy of traditional Chinese medicine course resource characteristics, and generates answers based on multiple pieces of data with the highest scores in the search results to construct a third reference answer.

[0011] S5. The large model summarizes and generates knowledge based on the first, second, and third reference answers, and strictly screens data reference sources through a verification mechanism to generate the most accurate comprehensive answers.

[0012] Optionally, the Fass vector library stores chapter semi-structured data generated from traditional Chinese medicine course data resources, and the chapter semi-structured data is word-embedded by the large model and synchronously enters the Fass vector library.

[0013] Optionally, the dynamic construction process of the knowledge graph includes: a pre-construction process and a dynamic update process;

[0014] Among them, the pre-construction process includes:

[0015] Perform complete structuring on traditional Chinese medicine course data resources. The complete structuring includes two parts: automatic annotation by the large model and manual review. Among them, automatic annotation is to extract traditional Chinese medicine course data resources through large model prompts. The extraction content includes five parts: entities, entity types, entity attributes, relationships, and relationship attributes. After extraction, manual review and verification are carried out, and the parts that do not pass the review and verification are deleted. The completely structured data that passes the review and verification is imported into the neo4j graph database to form the knowledge graph;

[0016] The dynamic update process includes:

[0017] Perform real-time analysis, automatic annotation, and manual review on newly added traditional Chinese medicine course resources.

[0018] Optionally, the regular synonym and antonym list includes regular synonyms and antonyms of entities and relationships, and is used to discover regular synonyms and antonyms of entities and relationships;

[0019] The characteristic regular synonym and antonym detection algorithm is used to further calculate potential entity and relationship synonyms and antonyms. The algorithm steps are as follows:

[0020] Construct a feature set: For each entity term and relationship term Wi, construct a basic feature set Fwi. The basic feature set Fwi includes the corresponding structured attribute information, Fwi = {c1, c2, c3,...}, where c1, c2, c3 are the attribute lists;

[0021] Calculation construction: Intersection: The number of features shared by two terms |Fwa ∩ Fwb| = Fshare; Union: The total number of all features of two terms |Fwa ∪ Fwb| = Fall; Symmetric difference: The union of two term sets - intersection, Fall - Fshare = Fsym; Total number: FWa + FWb = Fnum;

[0022] Perform the calculation: Sim(Wa,Wb) = ((2 * Fshare) / Fnum) - (Fsym / Fall), where Sim(Wa,Wb) represents the similarity of two terms. The first part ((2 * Fshare) / Fnum) emphasizes the importance of common features. The more common features, the higher the similarity. The second part -(Fsym / Fall) slightly penalizes the mismatched part, making the result more balanced. When calculating the similarity, this part of the value will be subtracted, thus slightly reducing the similarity score, ensuring that even if two terms have many common features, the differences between them will be reflected in the scoring, and at the same time helping to avoid misclassification or overly high similarity scores caused by ignoring differences;

[0023] Similarity normalization: Normalize the similarity result of two terms: Sim(Wa,Wb) = max(0, min(1, Sim(Wa,Wb))), ensuring that the similarity score is between [0,1];

[0024] Result output: Output two terms with similarity scores higher than the preset threshold as positive synonyms.

[0025] Optionally, the tokenizer is HanLP. The entities in the knowledge graph are used as a custom vocabulary, the tokenization rules are set according to the entity types, and the tokenization weight values are dynamically adjusted according to the frequency of the entities appearing in the traditional Chinese medicine course resources.

[0026] Optionally, the pre - construction process of the Elasticsearch search engine index core is as follows:

[0027] Parse the outline directory of the traditional Chinese medicine course resources to generate semi - structured data in chapter body form;

[0028] Perform secondary verification and error correction on the semi - structured data in chapter body form through prompt words + large model to form new semi - structured information data;

[0029] Synchronize the new semi - structured information data to Elasticsearch to build an efficient indexed core.

[0030] Optionally, the sorting formula of the traditional Chinese medicine course resource feature sorting strategy is as follows:

[0031]

[0032] Among them: FScore represents the score of the retrieved document d; TF ti,d is the term frequency of the entity term t i in the document d; IDF ti is the inverse document frequency of the entity term t i The inverse document frequency = the total number of retrieved documents / the number of documents containing the entity term t i ; W ti is the word segmentation weight value corresponding to the frequency of occurrence of the entity term in the traditional Chinese medicine curriculum resources; dn is the total number of retrieved documents; R d is the timeliness factor, which represents the number of years since the document was stored in the database and is used to reflect the newness or oldness of the literature.

[0033] Optionally, the verification mechanism includes:

[0034] Performing relevance verification on the retrieved data reference source through large model prompts to determine whether the data reference source is truly strongly relevant to the question raised by the user.

[0035] Optionally, the comprehensive answer includes: the answer to the question raised by the user, as well as specific citations of the key literature fragments or chapters on which it is based, clearly marked, and a brief explanatory note is provided.

[0036] On the other hand, a new intelligent question - answering system for traditional Chinese medicine curriculum resources based on a large model is provided. The system includes:

[0037] A receiving module, which is used for the large model to receive questions raised by users regarding traditional Chinese medicine curriculum resources;

[0038] A first reference answer generation module, which is used for the large model to perform word embedding on the question, convert it into a question vector, perform similarity retrieval and matching of the question vector in the Fass vector library, and construct a first reference answer according to the retrieval result;

[0039] A second reference answer generation module, which is used for the large model to mine entities, relationships, and synonyms for the question according to the dynamically constructed knowledge graph, the positive - negative synonym table, and the characteristic positive - negative synonym detection algorithm, and perform knowledge graph filtering on the mined entities, relationships, and synonyms, only retaining the valid entities and relationships existing in the knowledge graph, analyzing the true intention of the user's question, determining the application scenario of graph query or graph calculation, automatically generating corresponding Cypher statements according to the filtered entities and relationships, executing them in Neo4j, and generating a natural - language answer in combination with the user's question and the execution result, and constructing a second reference answer;

[0040] The third reference answer generation module is used for the large model to perform fine-grained word segmentation on the question using a predefined word segmenter, perform full-text search in a pre-constructed Elasticsearch search engine index core, sort the search results according to the traditional Chinese medicine course resource feature sorting strategy, and generate a generative answer based on multiple pieces of data with the highest scores in the search results to construct the third reference answer;

[0041] The comprehensive answer generation module is used for the large model to summarize and generate knowledge based on the first, second, and third reference answers, and strictly screen data reference sources through a verification mechanism to generate the most accurate comprehensive answer.

[0042] On the other hand, an electronic device is provided. The electronic device includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned intelligent question-answering method for traditional Chinese medicine course resources of the new large model.

[0043] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned intelligent question-answering method for traditional Chinese medicine course resources of the new large model.

[0044] The beneficial effects brought by the technical solution provided by the present invention at least include:

[0045] Through the deep integration of a search engine, a knowledge graph based on Neo4j, vector retrieval, and large model technology, the present invention proposes an innovative intelligent question-answering method. This method cleverly integrates the technical advantages of the knowledge graph, search engine, and vector retrieval, realizes the efficient mining, precise analysis, and systematic collation of knowledge information of traditional Chinese medicine course resources. Therefore, for users' questions about specific traditional Chinese medicine course resources, professional and accurate answers can be provided, significantly improving the quality and efficiency of users' information acquisition, which is specifically manifested in the following aspects:

[0046] 1) Improve the accuracy and professionalism of answers

[0047] By constructing a knowledge graph based on Neo4j, the traditional Chinese medicine course resources are organized in a structured manner, enhancing the flexibility and adaptability of the system. At the same time, with the powerful natural language processing ability of the large model, it can more deeply understand the intention behind the user's query and provide more accurate and personalized answers. The introduction of the Elasticsearch full-text retrieval engine combined with a custom word segmenter optimizes the index creation process, improves the relevance of query results, and ensures the accuracy of information acquisition.

[0048] 2) Enhance the flexibility and coverage of the system

[0049] It solves the problem of insufficient flexibility in traditional rule-based Q&A solutions. By dynamically constructing a knowledge graph, it supports the adaptation and association discovery of new questions and is no longer limited to preset question patterns. The system can be continuously optimized as the course content is updated and new questions emerge, ensuring long-term effectiveness and adaptability, and covering more diverse user questioning methods.

[0050] 3) Improve semantic understanding and response ability

[0051] In response to the challenges posed by complex and ambiguous question expressions, as well as non-standard languages or dialects, advanced large model technologies are adopted to capture the nuances behind semantics, better understand and respond to the actual needs of users, significantly improving the system's performance in dealing with unconventional expressions, metaphors, or culture-specific content, and making it closer to the understanding level of human experts.

[0052] 4) Ensure the timeliness and reliability of information

[0053] A new mechanism for AI automatic annotation is implemented, enabling traditional Chinese medicine course resources to promptly reflect the latest academic progress, reducing the cost and time of manual maintenance. The application of the verification mechanism effectively reduces the occurrence of the "hallucination" phenomenon, ensuring that the provided answers are both accurate and reliable, and further enhancing the credibility of the system.

[0054] 5) Increase interpretability and transparency

[0055] Each generated answer will clearly mark the specific citation of the key literature fragments or sections on which it is based and provide a short explanatory note to help users understand the answer source and logic, especially suitable for the traditional Chinese medicine education scenario. This approach not only increases the transparency of the system but also makes it easier for learners to accept and trust the information provided.

[0056] 6) Optimize search efficiency and user experience

[0057] The special sorting strategy and algorithm optimization ensure high-efficient data access speed and enhance the speed experience of user queries. Combining the characteristics of traditional Chinese medicine course resources and the word segmentation weight values corresponding to Neo4j entities for search relevance sorting improves the hit rate, enabling users to quickly find the most relevant materials. Brief Description of the Drawings

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0059] Figure 1It is a flowchart of a novel large model intelligent question-answering method for traditional Chinese medicine course resources provided by an embodiment of the present invention;

[0060] Figure 2 It is a general block diagram of a novel large model intelligent question-answering method for traditional Chinese medicine course resources provided by an embodiment of the present invention;

[0061] Figure 3 It is a processing diagram of semi-structured data resources provided by an embodiment of the present invention;

[0062] Figure 4 It is a processing diagram of fully structured resources provided by an embodiment of the present invention;

[0063] Figure 5 It is a block diagram of a novel large model intelligent question-answering system for traditional Chinese medicine course resources provided by an embodiment of the present invention;

[0064] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific embodiments

[0065] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0066] An embodiment of the present invention provides a novel large model intelligent question-answering method for traditional Chinese medicine course resources. This method can be implemented by an electronic device, and the electronic device can be a terminal or a server. Figure 1 As shown in the flowchart of this method, Figure 2 As shown in the general block diagram of this method, the processing flow can include the following steps:

[0067] S1. The large model receives a question raised by the user regarding traditional Chinese medicine course resources;

[0068] The large model in the embodiment of the present invention can be a large language model (LLM).

[0069] S2. The large model performs word embedding on the question, converts it into a question vector, performs similarity retrieval and matching on the question vector in the Fass vector library, and constructs a first reference answer according to the retrieval result;

[0070] Optionally, the Fass vector library stores chapter semi-structured data generated from traditional Chinese medicine course data resources (by parsing the outline directory of traditional Chinese medicine course resources, chapter semi-structured data is generated, including chapter titles, authors, abstracts, contents, update times, etc.). The chapter semi-structured data is word-embedded by the large model and synchronously enters the Fass vector library, as Figure 3 shown.

[0071] S3. The large model mines entities, relationships, and synonyms for the question based on the dynamically constructed knowledge graph, positive and negative synonym list, and characteristic positive and negative synonym detection algorithm, and filters the mined entities, relationships, and synonyms through the knowledge graph, only retaining the valid entities and relationships existing in the knowledge graph, analyzes the true intention of the user's question, determines the application scenario of graph query or graph calculation, automatically generates corresponding Cypher statements according to the filtered entities and relationships, and executes them in Neo4j. Combining the user's question and the execution results, a natural language answer is generated to construct the second reference answer;

[0072] Optionally, the dynamic construction process of the knowledge graph includes: a pre-construction process and a dynamic update process;

[0073] Among them, the pre-construction process includes:

[0074] As Figure 4 shown, the traditional Chinese medicine course data resources are completely structured. The complete structuring process includes two parts: automatic annotation by the large model and manual review. Among them, automatic annotation is to extract the traditional Chinese medicine course data resources through the large model prompt words. The extracted content includes five parts: entities, entity types, entity attributes, relationships, and relationship attributes (entity: identify the key terms or concepts in the text; entity type: determine the category to which each entity belongs, such as medicinal materials, diseases, treatment methods, etc.; entity attribute: describe the specific characteristics or parameters of the entity; relationship: define the association between different entities; relationship attribute: refine the specific nature or conditions of the relationship). After extraction, manual review and verification are carried out, and the parts that do not pass the review and verification are deleted. The completely structured data that passes the review and verification is imported into the neo4j graph database to form the knowledge graph;

[0075] The dynamic update process includes:

[0076] Real-time analysis, automatic annotation, and manual review are carried out on the newly added traditional Chinese medicine course resources.

[0077] The update speed of the traditional rule base often fails to keep up with the rapid evolution of knowledge, which may weaken the timeliness and effectiveness of the system. Therefore, the embodiments of the present invention introduce large model technology to achieve real-time AI automatic annotation to ensure that the traditional Chinese medicine course resources can timely reflect the latest academic progress.

[0078] For example, the question raised by the user is:

[0079] What are the academic inheritances and influences of the Li family?

[0080] Then the prompt words are as follows:

[0081] Extract the entities or keywords of the following content, and the synonyms of the entities. The following content is:

[0082] What are the academic inheritances and influences of the Li family?

[0083] ---

[0084] The output format is as follows: ["entity1","entity2",...]

[0085] ---

[0086] Note: Only return the statements, without giving explanations or apologies. The output JSON format must be correct.

[0087] Optionally, the positive and negative synonym list (the positive and negative synonym list is pre-created in the embodiments of the present invention, and the format is [entity - entity synonym][relationship - relationship synonym]) includes the conventional synonyms of entities and relationships, and is used to mine the conventional synonyms of entities and relationships;

[0088] The special positive and negative synonym detection algorithm is used to further calculate the potential entity and relationship synonyms. The algorithm steps are as follows:

[0089] Construct a feature set: For each entity term and relationship term Wi, construct a basic feature set Fwi, and the basic feature set Fwi includes the corresponding structured attribute information, Fwi = {c1, c2, c3,...}, where c1, c2, c3 are the attribute lists;

[0090] Calculate and construct: Intersection: The number of features shared by two terms |Fwa ∩ Fwb| = Fshare; Union: The total number of all features of two terms |Fwa ∪ Fwb| = Fall; Symmetric difference: The union of two term sets - intersection, Fall - Fshare = Fsym; Total number: FWa + FWb = Fnum;

[0091] Execute the calculation: Sim(Wa, Wb) = ((2 * Fshare) / Fnum) - (Fsym / Fall), where Sim(Wa, Wb) represents the similarity of two terms. The first part ((2 * Fshare) / Fnum) emphasizes the importance of common features. The more common features there are, the higher the similarity. The second part - (Fsym / Fall) slightly penalizes the mismatched part, making the result more balanced. When calculating the similarity, this part of the value will be subtracted, thus slightly reducing the similarity score, ensuring that even if two terms have many common features, the differences between them will also be reflected in the scoring, and at the same time helping to avoid misclassification or overly high similarity scores caused by ignoring differences;

[0092] Similarity normalization: Normalize the similarity result of two terms: Sim(Wa, Wb) = max(0, min(1, Sim(Wa, Wb))), ensuring that the similarity score is between [0, 1];

[0093] Result output: Output two terms with similarity scores higher than the preset threshold as positive synonyms.

[0094] The similarity detection algorithm of the embodiments of the present invention not only considers the intersection and union of the terms themselves, but also combines the influence of the symmetric difference and the total quantity, thereby improving the accuracy and interpretability of the term similarity evaluation. For example, the features of "Astragalus membranaceus" are {qi deficiency, replenishing qi, Leguminosae, root, unique feature 1}, and the features of "Mianqi" are {qi deficiency, replenishing qi, Leguminosae, root, unique feature 2}. Then, both "Astragalus membranaceus" and "Mianqi" have 5 features, Fwa = 5, Fwb = 5, the intersection is {qi deficiency, replenishing qi, Leguminosae, root}, so Fshare = 4; the union is {qi deficiency, replenishing qi, Leguminosae, root, unique feature 1, unique feature 2}, so Fall = 6; Fsym = Fall - Fshare = 6 - 4 = 2; Fnum = FWa + Fwb = 5 + 5 = 10; the similarity Sim(Wa, Wb) = ((2 * Fshare) / Fnum) - (Fsym / Fall) = ((2 * 4) / 10) = 0.8 - (2 / 6) ≈ 0.467 > the preset threshold of 0.4, and the positive name is output: "Astragalus membranaceus", and the synonym is: "Mianqi".

[0095] The embodiments of the present invention use a large model to analyze the true intention of the user's question, determine the application scenario (graph query or graph calculation), automatically generate corresponding Cypher statements according to the filtered entities and relationships, and execute them in Neo4j to obtain the execution results. The specific examples are as follows:

[0096] Graph query:

[0097] Question: Which traditional Chinese medicines have the effect of clearing heat and detoxifying?

[0098] Cypher statement: MATCH(herb:Herb)-[:HAS_EFFECT]->(effect:Effect{name:"clearing heat and detoxifying"}) RETURN herb.name

[0099] Execution result: List the names of all traditional Chinese medicines with the effect of "clearing heat and detoxifying".

[0100] Graph calculation:

[0101] Question: Please find the shortest path between the two syndromes of "wind-heat cold" and "lung-heat cough", including the intermediate nodes (such as traditional Chinese medicines, other syndromes) and their relationship types.

[0102] Cypher statement: MATCH path = shortestPath((syndrome1:Syndrome{name: "Wind-Heat Cold"})-[:TREATS|CAUSES|LEADS_TO*]->(syndrome2:Syndrome{name: "Lung-Heat Cough"})) RETURN [node IN nodes(path)|node.name] AS PathNodes, [rel IN relationships(path)|type(rel)] AS PathRelationships

[0103] Execution result: {"PathNodes": ["Wind-Heat Cold", "Forsythia suspensa", "Lung-Heat Cough"], "PathRelationships": ["TREATS", "LEADS_TO"]}

[0104] In an embodiment of the present invention, the user's question and the execution result are then combined to generate a natural language answer, constructing a second reference answer.

[0105] S4. The large model performs fine-grained word segmentation on the question using a predefined word segmenter, performs full-text search in a pre-constructed Elasticsearch search engine index core, sorts the search results according to the sorting strategy of traditional Chinese medicine course resource characteristics, and generates an answer based on multiple data with the highest search result scores (such as the top three data with the highest scores), constructing a third reference answer;

[0106] Optionally, the word segmenter is HanLP. Entities in the knowledge graph are used as a custom word list, word segmentation rules are set according to entity types, and the word segmentation weight value is dynamically adjusted according to the frequency of entity occurrences in traditional Chinese medicine course resources.

[0107] This approach not only improves the accuracy of word segmentation but also enhances the understanding ability of specific terms in the field of traditional Chinese medicine.

[0108] Optionally, the process of pre-constructing the Elasticsearch search engine index core is as follows:

[0109] Parse the outline directory of traditional Chinese medicine course resources to generate chapter semi-structured data;

[0110] Perform secondary verification and error correction on the chapter semi-structured data through prompts + large model to form new semi-structured information data;

[0111] Synchronize the new semi-structured information data into Elasticsearch to build an efficient indexed core, as Figure 3 shown.

[0112] This core can effectively narrow the search scope and improve the hit rate, which is the basis for the entire system to quickly respond to user queries and ensures high-efficiency data access.

[0113] Optionally, the sorting formula of the sorting strategy for the characteristics of traditional Chinese medicine course resources is as follows:

[0114]

[0115] Where: FScore represents the score of the retrieved document d; TF ti,d is the term frequency of entity term t i in document d; IDF ti is the inverse document frequency of entity term t i The inverse document frequency = total number of retrieved documents / number of documents containing entity term t i ; W ti is the word segmentation weight value corresponding to the frequency of occurrence of the entity term in traditional Chinese medicine course resources; dn is the total number of retrieved documents; R d is the timeliness factor, which represents the number of years since the document was stored in the database and is used to reflect the newness and oldness of the literature.

[0116] S5. The large model summarizes and generates knowledge based on the first, second, and third reference answers, and strictly screens the data reference sources through a verification mechanism to generate the most accurate comprehensive answers.

[0117] This process is not just a simple summary of existing information, but rather a deep analysis and integration of information from different sources by the large model to extract key points, and adjust the angle and depth of the answer according to the specific situation of the user's question to ensure that the provided information is both comprehensive and targeted. At the same time, the large model will check the consistency and logic of the generated answer to avoid information conflicts or misleading, and finally customize the final answer text with appropriate term complexity and detailed explanations, enabling users to more easily understand complex traditional Chinese medicine concepts and theories. Finally, the entire answer will be optimized into a smooth and natural language expression form for easy reading and dissemination, ensuring that users can obtain satisfactory answers and benefit from them.

[0118] Optionally, the verification mechanism includes:

[0119] Verify the relevance of the retrieved data reference sources through the large model prompt words to determine whether the data reference sources are truly strongly relevant to the question raised by the user.

[0120] For example, the prompt words are as follows:

[0121] Determine whether the following content is relevant to the retrieval statement. The following content is:

[0122] {"zidingyi_indexpkey_score":"37.836456","Course section":"Four, <em>Li< / em> <em>Clan< / em> <em>'s< / em> <em>Academic< / em> <em>Inheritance< / em> <em>And< / em> <em>Influence< / em> <em>Of< / em> <em>Influence< / em> <em>Of< / em> ","Course content":"In the 'Treatment of All Diseases' section of 'Compendium of Materia Medica', many disease syndromes are recorded <em>The< / em> Drug application methods by application, application points <em>Have< / em> Shenque, Yongquan, Lao ;; ;; Gong, Yintang, etc. The summary list is as follows (Table 7-2). ;; ;; Acupoint application is very popular and has had a <em>Influence< / em> <em>Of< / em> great influence on later generations. In 'Li Yue Pian Wen', Wu Shiji in the Qing Dynasty recorded a lot of application content. Zhao <em>Learning< / em> Min dedicated a special section on 'Application Methods' in 'Chuan Ya Wai Bian' to promote this method. Modern clinical use of drug acupoint application in the treatment of ulcerative colitis, ulcerative colitis, bronchial asthma, constipation, etc. has achieved very good <em>Of< / em> curative effects, <em>Have< / em> widely <em>Of< / em> development prospects. Although the specific drug selection and acupoint selection <em>And< / em> <em>Li< / em> <em>Clan< / em> in the book are not the same, but <em>Have< / em> clear <em>Of< / em> <em>Inheritance< / em> <em>Of< / em> trajectory. This non-invasive and simple, convenient, inexpensive, and effective <em>Learning< / em> therapy has developed into a <em>Of< / em> branch <em>Learning< / em> marginal <em>Of< / em> discipline in acupuncture and moxibustion.","Relevance score":"37.836456"}

[0123] ---

[0124] The retrieval statement is:

[0125] ---

[0126] What are the academic inheritance and influence of Li?

[0127] ---

[0128] The output option can only be one of: {"isok": false} or {"isok": true}. True represents strong relevance, and false represents irrelevance.

[0129] ---

[0130] Note: Only return the statement without giving explanations or apologies. The output JSON format must be correct.

[0131] Optionally, the comprehensive solution includes: the answer to the question raised by the user, as well as specific references to the key literature fragments or chapters on which it is based, clearly marked, and a short explanatory note is provided.

[0132] After completing the comprehensive solution, the system may also provide some additional resources or suggestions, such as recommending relevant literature, course chapters, or other learning materials to encourage the user to further explore and learn.

[0133] As Figure 5 shown, an embodiment of the present invention also provides a new intelligent question-answering system for traditional Chinese medicine course resources in a large model. The system includes:

[0134] A receiving module 510 for the large model to receive questions raised by the user regarding traditional Chinese medicine course resources;

[0135] A first reference answer generation module 520 for the large model to perform word embedding on the question, convert it into a question vector, perform similarity retrieval and matching on the question vector in the Fass vector library, and construct a first reference answer according to the retrieval result;

[0136] A second reference answer generation module 530 for the large model to mine entities, relationships, and synonyms for the question according to the dynamically constructed knowledge graph, positive and negative synonym tables, and characteristic positive and negative synonym detection algorithms, and perform knowledge graph filtering on the mined entities, relationships, and synonyms, only retaining the valid entities and relationships existing in the knowledge graph, analyzing the true intention of the user's question, determining the application scenario of graph query or graph calculation, automatically generating corresponding Cypher statements according to the filtered entities and relationships, and executing them in Neo4j, and generating a natural language answer in combination with the user's question and the execution result to construct a second reference answer;

[0137] A third reference answer generation module 540 for the large model to perform fine-grained word segmentation on the question using a predefined word segmenter, perform full-text search in the pre-constructed Elasticsearch search engine index core, sort the search results according to the traditional Chinese medicine course resource characteristic sorting strategy, and generate an answer based on multiple data with the highest search result scores to construct a third reference answer;

[0138] The comprehensive answer generation module 550 is used for the large model to summarize and generate knowledge based on the first, second, and third reference answers, and strictly screen data reference sources through a verification mechanism to generate the most accurate comprehensive answer.

[0139] The intelligent Q&A system for traditional Chinese medicine course resources of the novel large model provided by the embodiment of the present invention has a functional structure corresponding to the intelligent Q&A method for traditional Chinese medicine course resources of the novel large model provided by the embodiment of the present invention, which will not be elaborated here.

[0140] Figure 6 It is a schematic structural diagram of an electronic device 600 provided by the embodiment of the present invention. The electronic device 600 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 601 and one or more memories 602. Among them, at least one instruction is stored in the memory 602, and the at least one instruction is loaded and executed by the processor 601 to implement the steps of the above-mentioned intelligent Q&A method for traditional Chinese medicine course resources of the novel large model.

[0141] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions, and the above instructions can be executed by a processor in a terminal to complete the above-mentioned intelligent Q&A method for traditional Chinese medicine course resources of the novel large model. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0142] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.

[0143] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A new large-scale intelligent question-answering method for traditional Chinese medicine course resources, characterized in that: The method comprises: S1. The big model receives questions raised by users regarding TCM course resources; S2, the large model embeds the question into words, converts it into a question vector, performs similarity search and matching on the question vector in the Fass vector library, and constructs a first reference answer based on the search results; S3. The large model mines entities, relationships and synonyms for the question based on the dynamically constructed knowledge graph, the table of synonyms and the characteristic synonym detection algorithm, and filters the mined entities, relationships and synonyms through the knowledge graph, retaining only valid entities and relationships existing in the knowledge graph, analyzing the real intention of the user's question, determining the application scenario of graph query or graph computing, automatically generating corresponding Cypher statements based on the filtered entities and relationships, and executing them in Neo4j, generating natural language answers based on the user's question and the execution results, and constructing a second reference answer; S4, the large model uses a predefined word segmenter to perform fine word segmentation on the question, performs a full-text search in a pre-built Elasticsearch search engine index core, sorts the search results according to the characteristic sorting strategy of traditional Chinese medicine course resources, and generates a generative answer based on multiple data with the highest search result scores to construct a third reference answer; S5. The large model summarizes and generates knowledge based on the first, second and third reference answers, and strictly screens data reference sources through a verification mechanism to generate the most accurate comprehensive answer.

2. The method according to claim 1, characterized in that The Fass vector library stores chapter-based semi-structured data generated from traditional Chinese medicine course data resources. The chapter-based semi-structured data is word-embedded by the large model and synchronously enters the Fass vector library.

3. The method according to claim 1, characterized in that The dynamic construction process of the knowledge graph includes: a pre-construction process and a dynamic update process; The pre-build process includes: The TCM course data resources are fully structured, and the fully structured processing includes two parts: large model automatic annotation and manual review. The automatic annotation is to extract the TCM course data resources through the large model prompt words. The extracted content is five parts: entity, entity type, entity attribute, relationship, and relationship attribute. After the extraction is completed, manual review and verification are performed, and the parts that fail the review and verification are deleted. The fully structured data that pass the review and verification are imported into the neo4j graph database to form the knowledge graph; The dynamic update process includes: Conduct real-time analysis, automatic annotation and manual review of newly added traditional Chinese medicine course resources.

4. The method according to claim 1, characterized in that: The formal synonym table includes regular synonyms of entities and relationships, and is used to mine regular synonyms of entities and relationships; The characteristic synonym detection algorithm is used to further calculate potential entity and relationship synonyms. The algorithm steps are as follows: Construct a feature set: for each entity term and relationship term Wi, construct a basic feature set Fwi, the basic feature set Fwi includes the corresponding structured attribute information, Fwi = {c1, c2, c3, ...}, c1, c2, c3 is an attribute list; Calculation construction: intersection: the number of features shared by two terms |Fwa∩Fwb| = Fshare; Union: the total number of all features of both terms |Fwa∪Fwb| = Fall; Symmetric difference: union-intersection of two term sets, Fall-Fshare=Fsym; total number: FWa+FWb=Fnum; Perform the calculation: Sim(Wa,Wb) = ((2*Fshare) / Fnum)-(Fsym / Fall), where Sim(Wa,Wb) represents the similarity between the two terms. The first part ((2*Fshare) / Fnum) emphasizes the importance of common features. The more common features, the higher the similarity. The second part - (Fsym / Fall) slightly penalizes the mismatched parts, making the result more balanced. When calculating the similarity, this part of the value will be subtracted, thereby slightly reducing the similarity score, ensuring that even if the two terms have many common features, the differences between them will be reflected in the score, and at the same time help avoid misclassification or too high similarity scores caused by ignoring differences; Similarity normalization: Normalize the similarity results of two terms: Sim(Wa,Wb)=max(0,min(1,Sim(Wa,Wb))) to ensure that the similarity score is between [0,1]; Result output: Two terms with similarity scores higher than the preset threshold are output as positive synonyms.

5. The method according to claim 1, characterized in that The word segmenter is HanLP, which takes the entities in the knowledge graph as a custom vocabulary, sets word segmentation rules according to entity types, and dynamically adjusts word segmentation weight values ​​according to the frequency of entities appearing in traditional Chinese medicine course resources.

6. The method according to claim 1, characterized in that The pre-construction process of the Elasticsearch search engine index core is as follows: Parse the outline catalog of TCM course resources to generate chapter-based semi-structured data; Performing secondary verification and error correction on the chapter body semi-structured data through prompt words + large model to form new semi-structured information data; The new semi-structured information data is synchronized to Elasticsearch to build an efficient indexing core.

7. The method according to claim 1, characterized in that The ranking formula of the TCM course resource feature ranking strategy is as follows: Among them: FScore represents the score of the hit document d; TF ti,d is the entity term t i The frequency of terms in document d; IDF ti is the entity term t i The inverse document frequency, where the inverse document frequency = the total number of hit documents / containing entity terms t i The number of documents; W ti is the word segmentation weight value corresponding to the frequency of entity terms appearing in TCM course resources; dn is the total number of hit documents; R d It is the timeliness factor, which indicates the year since the document was entered into the database, and is used to reflect the newness of the document.

8. The method according to claim 1, characterized in that The verification mechanism includes: The retrieved data reference sources are verified for relevance using large model prompt words to determine whether the data reference sources are truly strongly relevant to the questions raised by the user.

9. The method according to claim 1, characterized in that: The comprehensive answer includes: the answer to the question raised by the user, as well as specific citations of key literature fragments or chapters clearly marked as the basis, and a brief explanation is provided.

10. A new large-scale intelligent question-answering system for TCM course resources, characterized in that: The system comprises: The receiving module is used for the large model to receive questions raised by users regarding TCM course resources; A first reference answer generation module, which is used for the large model to embed the question into words, convert it into a question vector, perform similarity search and matching on the question vector in the Fass vector library, and construct a first reference answer based on the search results; The second reference answer generation module is used for the large model to mine entities, relationships and synonyms for the question according to the dynamically constructed knowledge graph, the table of synonyms and the characteristic synonym detection algorithm, and to filter the mined entities, relationships and synonyms through the knowledge graph, retaining only the valid entities and relationships existing in the knowledge graph, analyzing the real intention of the user's question, determining the application scenario of the graph query or graph calculation, automatically generating corresponding Cypher statements according to the filtered entities and relationships, and executing them in Neo4j, generating natural language answers in combination with the user's question and the execution results, and constructing the second reference answer; A third reference answer generation module is used for the large model to perform fine word segmentation on the question using a predefined word segmenter, perform a full-text search in a pre-built Elasticsearch search engine index core, sort the search results according to the characteristic sorting strategy of traditional Chinese medicine course resources, and generate answers based on multiple data with the highest search result scores to construct a third reference answer; The comprehensive answer generation module is used for the large model to summarize and generate knowledge based on the first, second and third reference answers, and strictly screen the data reference source through the verification mechanism to generate the most accurate comprehensive answer.

Citation Information

Patent Citations

  • Novel consulting method and device and electronic equipment

    CN112650857A

  • Information processing method and device, equipment, storage medium and program product

    CN116975437A

  • Law question answering system based on intention recognition and knowledge graph

    CN117891923A

  • Method and system for enhancing RAG questions and answers through mixed retrieval method

    CN118627625A

  • Hybrid retrieval method and system for RAG question-answering system

    CN118656482A

Cited By

  • Intelligent assurance knowledge question and answer method and system based on LLM and RAG technologies

    CN120744072A

  • Intelligent question and answer method and system based on traditional Chinese medicine classics meta-term engine

    CN121434337A

  • A smart question-answering method and system based on a meta-term engine of traditional Chinese medicine classics

    CN121434337B