Auxiliary question answering method and system based on traditional Chinese medicine ancient book genre inheritance

By constructing a knowledge graph and multi-task learning model for the inheritance of Chinese medicine schools, the problems of context fragmentation and character confusion in the ancient Chinese medicine question-answering system were solved, the systematic inheritance of Chinese medicine knowledge and accurate question-answering were achieved, and the efficiency and accuracy of the question-answering system were improved.

CN120705329APending Publication Date: 2025-09-26SUZHOU INTEGRATED TRADITIONAL CHINESE & WESTERN MEDICINE HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510790602.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing question-and-answer systems for ancient Chinese medicine books are usually based on knowledge bases built on information on physiology, pathology, nature, cognitive methods, and treatment methods. They are unable to effectively inherit Chinese medicine knowledge and have problems such as context fragmentation, character confusion, retrieval generalization, and difficulty in handling character relationships.

Method used

Build a knowledge graph based on the inheritance of TCM schools, use the pre-trained RoBERTa large model to segment and index ancient book materials, combine multi-task learning models for semantic routing and retrieval, generate accurate response data, and support medical-level literature retrieval and semantic consistency.

Benefits of technology

It achieves the rigorous and systematic inheritance of traditional Chinese medicine knowledge, improves the accuracy and consistency of questions and answers, supports original text positioning and systematic answers, and enhances the scalability of the system and the stability of responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705329A_ABST
    Figure CN120705329A_ABST
Patent Text Reader

Abstract

The invention discloses an auxiliary question answering method and system based on traditional Chinese medicine ancient book genre inheritance, and relates to the field of artificial intelligence. The method comprises the following steps: acquiring knowledge maps of different medical schools; nodes of the knowledge graph are doctors corresponding to medical schools, literature subsystems corresponding to the doctors are bound to the nodes, and the edge is the door-to-door inheritance relationship between the doctors; introducing a pre-trained RoBERTa large model into the literature subsystem to segment related ancient book data of a doctor corresponding to the literature subsystem, and constructing an index structure of a vector space; and obtaining question text data, analyzing the doctor's orientation from the question text data, performing semantic routing under the guidance of the knowledge graph, and generating response data of the question text data in combination with the index structure. The knowledge graph is constructed along inheritance veins of genres, and learning of a rigorous and systematic traditional Chinese medicine theory system is facilitated. And meanwhile, an exclusive literature retrieval engine is bound for each doctor node, so that the accuracy of an output result of the question and answer system can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to an auxiliary question-answering method and system based on the inheritance of ancient Chinese medical texts and schools. Background Art

[0002] As an integral part of traditional Chinese culture, Traditional Chinese Medicine (TCM) boasts a long and rich history, numerous schools of thought, and a vast collection of ancient texts. These texts contain a wealth of information on disease diagnosis, syndrome differentiation and treatment, drug preparation and use, medical records, treatment experience, and theoretical exploration, representing a valuable resource for the inheritance and development of TCM. With the changing times and accelerated modernization, the digitization of ancient TCM texts is an inevitable trend. By organizing and digitizing these texts, the efficiency of information utilization can be greatly improved.

[0003] With the development of large AI models, many institutions are developing RAG (Retrieval Augmentation Generation) systems to build knowledge-based question-answering systems. This involves first constructing a knowledge-based search engine and then using the search results as context for a large language model to answer questions. This approach allows TCM students to learn through self-service question-answering, thereby better helping them master TCM knowledge and rapidly improve their knowledge level.

[0004] However, the existing question-and-answer systems for ancient Chinese medicine books are usually based on knowledge bases established based on information on physiology, pathology, nature, cognitive methods, and treatment methods, which is not conducive to the inheritance of Chinese medicine knowledge. Summary of the Invention

[0005] Based on this, it is necessary to provide an auxiliary question-answering method and system based on the inheritance of ancient Chinese medical books and schools to address the above technical problems, so as to improve the efficiency of organizing, retrieving, question-answering and inheritance of Chinese medicine knowledge.

[0006] In a first aspect, the present application provides an auxiliary question-answering method and system based on the inheritance of ancient Chinese medical texts. The method includes:

[0007] Obtain knowledge graphs of different medical schools; the nodes of the knowledge graph represent the doctors of the corresponding medical schools, and each node is bound to the corresponding doctor's literature subsystem, and the edges represent the master-disciple inheritance relationship between doctors;

[0008] Introducing the pre-trained RoBERTa model into the literature subsystem to segment the ancient medical records corresponding to the literature subsystem and construct an index structure in the vector space;

[0009] Obtain question text data, parse medical references from the question text data, perform semantic routing under the guidance of the knowledge graph, and generate response data for the question text data in combination with the index structure.

[0010] In one embodiment, a pre-trained RoBERTa model is introduced into the literature subsystem to segment the ancient medical records corresponding to the literature subsystem. The index structure of the vector space is constructed as follows:

[0011] Construct a corpus containing knowledge unit annotations through pre-cutting;

[0012] Use the corpus to train the RoBERTa large model to learn the boundaries of knowledge units;

[0013] The ancient book data is fed into the trained RoBERTa model, which determines paragraph boundaries based on the learned knowledge unit boundaries and merges paragraphs into segments based on their semantic similarity to generate target knowledge units.

[0014] Construct an index structure based on the vectorized target knowledge units.

[0015] In one embodiment, obtaining question text data and parsing the doctor's direction from the question text data includes:

[0016] The pre-processed multi-task learning model identifies intent information in the question text data and constructs query conditions based on the intent information. The intent information includes doctors, diseases, drugs, ancient books, intent types, and whether it is in the field of traditional Chinese medicine.

[0017] Enter the query conditions into the RAG system to locate the doctor.

[0018] In one embodiment, the method further comprises:

[0019] Evaluate the response data, perform structured management on the response data and the question text data corresponding to the response data based on the evaluation results, and build semantic retrieval for the question text data;

[0020] For new question text data, if the semantic retrieval matches question text data with a similarity higher than a first threshold, the answer data corresponding to the matched question text data is reused as the answer data corresponding to the new question text data; if the semantic retrieval matches a similarity not higher than the first threshold and higher than the second threshold, the answer data corresponding to the matched question text data is injected as prompt information into the generation process of the answer data corresponding to the new question text data.

[0021] In one embodiment, the method further comprises:

[0022] Update, supplement or replace the question text data for constructing semantic retrieval.

[0023] In one embodiment, the method further comprises:

[0024] In the process of generating the response data, the target knowledge units related to the response data are sequentially stored in the cache stream;

[0025] After the response data generation is completed, the positioning information of the ancient book materials is extracted from the target knowledge unit and the positioning information is output.

[0026] Secondly, this application also provides an auxiliary question-answering method and system based on the inheritance of ancient Chinese medical texts. The system includes:

[0027] The knowledge base is used to obtain the knowledge graph of different medical schools. The nodes of the knowledge graph are doctors of the corresponding medical school, and each node is bound to the corresponding doctor's literature subsystem. The edges are the master-disciple inheritance relationships between doctors.

[0028] The retrieval module is used to introduce the pre-trained RoBERTa large model into the literature subsystem to segment the relevant ancient medical books in the literature subsystem and construct the index structure of the vector space;

[0029] The output module is used to obtain question text data, parse medical references from the question text data, perform semantic routing under the guidance of the knowledge graph, and generate response data for the question text data in combination with the index structure.

[0030] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the aforementioned auxiliary question-answering method and system based on the inheritance of ancient Chinese medical texts and schools.

[0031] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the aforementioned auxiliary question-answering method and system based on the inheritance of ancient Chinese medical texts and schools.

[0032] In a fifth aspect, the present application further provides a computer program product, comprising a computer program that, when executed by a processor, implements the steps of the aforementioned auxiliary question-answering method and system based on the inheritance of ancient Chinese medical texts and schools.

[0033] The aforementioned assisted question-answering method and system based on the inheritance of ancient Chinese medical texts and schools of thought comprises: obtaining a knowledge graph of different medical schools; the nodes of the knowledge graph represent the physicians of the corresponding medical schools, and each node is bound to the corresponding physician's literature subsystem, with edges representing the master-disciple inheritance relationship between physicians; introducing a pre-trained RoBERTa large model into the literature subsystem to segment the relevant ancient textual materials of the physicians in the literature subsystem and construct an index structure in the vector space; obtaining question text data, parsing the physician references from the question text data, and performing semantic routing guided by the knowledge graph, and generating response data based on the question text data in combination with the index structure. Constructing a knowledge graph along the inheritance lineage of the schools of thought helps to build a rigorous and systematic theoretical system of Chinese medicine and clearly grasp the internal connections between various knowledge points. Furthermore, binding a dedicated literature search engine to each physician node ensures semantic consistency, avoids interference from other physicians' opinions, generates more clearly directed response data, and improves the accuracy of the output results of the question-answering system. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 1 is a flow chart of an auxiliary question-answering method based on the inheritance of ancient Chinese medical texts and schools in one embodiment;

[0035] Figure 2 This is a relationship map of the medical practitioners of the Shicai School. DETAILED DESCRIPTION

[0036] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0037] The embodiment of the present application provides an auxiliary question-answering method and system based on the inheritance of ancient Chinese medical books. Figure 1 As shown, the following steps are included:

[0038] Step 102, obtain the knowledge graph of different medical schools; the nodes of the knowledge graph are the doctors of the corresponding medical schools, and each node is bound to the literature subsystem of the corresponding doctor, and the edges are the master-disciple inheritance relationship between doctors.

[0039] Based on the experience of experts, we construct corresponding knowledge graphs according to different schools of Chinese medicine. Each knowledge graph contains various doctors and their teacher-student relationships. Taking the Shicai School as an example, Figure 2 The figure shows the relationship map of the medical practitioners of this school. Each doctor (such as Ma Yuanyi and Shen Anbo) is a node in the map. The nodes are connected by edges such as "teacher-student" and "fellow student". Each node contains the representative works, medical records, annotations and other structured ancient book materials of the doctor.

[0040] Step 104: introduce the pre-trained RoBERTa large model into the document subsystem to segment the relevant ancient books and materials of the doctors in the document subsystem and construct an index structure of the vector space.

[0041] Different from the traditional unified knowledge base, this embodiment includes a dedicated literature search engine in the literature subsystem bound to each physician, which can search for ancient books and materials related to the physician in a targeted manner and support semantic search (vector search). In other embodiments, the literature subsystem can also support search algorithms such as keyword matching (BM25) based on semantic search.

[0042] By binding each node with a dedicated, search-capable literature subsystem, the retrieval granularity is refined, ensuring semantic consistency and avoiding interference from other medical experts. Furthermore, compared to traditional RAG systems that require reindexing when new knowledge is added, the node-level literature subsystem disclosed in this embodiment can be independently added and deleted, improving the system's scalability.

[0043] Step 106: obtain the question text data, parse the medical expert reference from the question text data, perform semantic routing under the guidance of the knowledge graph, and generate response data for the question text data in combination with the index structure.

[0044] After obtaining the question text data, the question text data is segmented, and then the segmented words are converted into vector form. In addition to the word vector, the entire question sentence needs to be represented as a vector. The sentence vector can be obtained by averaging, weighted summing, and other operations on the word vector. Based on the data in vector format, the medical direction is determined, and the specific node in the knowledge graph is located. Then, according to the semantics represented by the question text data, the optimal path is planned to guide the query to perform a targeted search in the knowledge graph, and then accurately locate the nodes and relationships that match the query semantics to obtain useful knowledge. At the same time, the document subsystem involving node binding will be used to search for relevant ancient books and materials. Based on the retrieved nodes, relationships, and ancient books and materials, reasoning and integration are performed to generate the final response data.

[0045] For example, when the question text data is "How does Ma Yuanyi treat diarrhea?", the knowledge graph locates the "Ma Yuanyi" node, calls its bound literature subsystem, retrieves "diarrhea" related content, and outputs the answer data.

[0046] For example, when the question text data is "How did the doctor before Shen Anbo treat diarrhea?", the knowledge graph locates "Shen Anbo" and traces back to all upstream nodes, searching in parallel in the bound documents respectively, summarizing the opinions of multiple doctors and generating systematic response data.

[0047] For example, when the question text data is "What is the relationship between Gu Shouzhi and You Zaijing?", the knowledge graph queries the path and edge type between the two, and generates structured character relationship descriptions such as "teacher-student relationship", "fellow students" or "citation relationship" as response data.

[0048] In other embodiments, the question text data cannot be directly parsed to identify a doctor, such as "How to treat diarrhea?" In this case, based on the key information "diarrhea" in the question, all collected ancient books and materials are first searched to find relevant information. This process can be combined with techniques such as inverted indexing to improve search efficiency. After finding an answer that matches the question text data, several medical experts can be located. Then, using the solution in step 106 of this embodiment, the treatment methods for diarrhea by various medical experts are searched and organized, generating response data with a lineage of medical traditions, facilitating analysis of the internal connections between the treatment methods of various medical experts.

[0049] Each school of Traditional Chinese Medicine (TCM) has its own unique core theories and academic insights. By studying a specific school, students can deeply grasp its core ideas and gain a more focused and in-depth understanding of TCM theory, avoiding the trap of superficial knowledge. Schools of thought often have a systematic theoretical framework, interconnected from basic theory to clinical application. Studying along the lines of a school's inheritance helps build a rigorous and systematic TCM theoretical system and clearly grasp the inherent connections between various knowledge points.

[0050] However, current mainstream RAG (Retrieval Augmented Generation) question-answering systems typically use a unified processing strategy when building knowledge bases. This strategy involves splitting all text into fixed-length chunks and storing them uniformly in the search engine, regardless of contextual factors such as author, genre, or era. This fragmented processing approach presents the following problems when processing complex, semantically coherent literature, such as ancient Chinese medical texts:

[0051] 1. Contextual fragmentation: the integrity of the physician’s thought cannot be preserved;

[0052] 2. Confusion of different medical experts: Mixed opinions can easily lead to cognitive bias;

[0053] 3. Generalized search: All questions are searched in the entire database, resulting in non-targeted answers.

[0054] 4. It is difficult to answer questions about the relationship between people: For example, non-documentary knowledge such as "the relationship between a certain doctor and a certain doctor" cannot be handled.

[0055] To overcome the above problems, this embodiment designs a "knowledge graph + precise search" system that integrates the inheritance structure of traditional Chinese medicine schools. Based on the inheritance relationship of traditional Chinese medicine schools through generations, this system combines modern search technology to achieve structured, personalized, and semantic search and question-answering of the content of ancient Chinese medical books. Compared with traditional RAG technology, this embodiment has the following technical advantages:

[0056] 1. Search granularity: Isolate medical-level documents with clear semantics;

[0057] 2. Targeted answers: Accurately identify the medical practitioners’ thoughts and ensure consistent content;

[0058] 3. Citations can be traced back to their source and support original text positioning;

[0059] 4. Character relationships: Directly query the knowledge graph and integrate it into the response data;

[0060] 5. System scalability: Node-level documents can be added and deleted independently.

[0061] In one embodiment, in step 104, a pre-trained RoBERTa large model is introduced into the document subsystem to segment the relevant ancient book materials of the medical practitioners corresponding to the document subsystem, and the index structure of the vector space is constructed, including: constructing a corpus containing knowledge unit annotations through pre-cutting; using the corpus to train the RoBERTa large model to learn the knowledge unit boundaries; inputting the ancient book materials into the trained RoBERTa large model, judging the paragraph boundaries based on the learned knowledge unit boundaries, and merging the paragraphs into segments according to the semantic similarity of the paragraphs to generate target knowledge units; and constructing an index structure based on the vectorized target knowledge units.

[0062] When using ancient texts for large-scale question-and-answer systems, the data volume is enormous and cannot be fed all at once. Therefore, the text must be segmented. Traditional RAG systems often segment documents according to fixed lengths, such as every 200 characters. This often results in knowledge interruptions, such as a complete sentence being cut off mid-sentence, negatively impacting subsequent large-scale question-and-answer systems. Some RAG systems also employ rules, such as stopping optimized document segmentation upon encountering a period, but the overall improvement is limited.

[0063] This example designs an ancient book fragment segmentation system based on the RoBERTa large model. It uses a neural network to intelligently identify knowledge units. This not only avoids the knowledge fragmentation that can occur when segmenting based on preset rules, but also merges multiple consecutive paragraphs with similar semantics into a single fragment (chunk) as the target knowledge unit. This creates a more reasonable index structure, and the retrieved results are more complete when subsequently using a large language model to enhance question answering.

[0064] In this example, a large language model such as GPT-4 is first used to construct a training corpus. Specifically, this model determines the natural semantic boundaries of each section of ancient text and labels it as "whether it is a complete knowledge unit." This process uses a manual review mechanism for quality control, thereby constructing a high-quality, annotated corpus.

[0065] Next, a sequence classification model based on RoBERTa is trained on the aforementioned corpus. The model's goal is to learn to determine whether each character or sentence represents a knowledge unit boundary. This model not only identifies the end of a natural paragraph but also semantically determines whether it should be merged with the previous paragraph. This judgment can be based on vector similarity or topic consistency. If the similarity or consistency is high, the paragraphs are considered semantically similar and merged into a single segment, serving as the target knowledge unit.

[0066] All target knowledge units generated in the end will be used to construct the vector index of the knowledge base for vector retrieval, providing a high-quality retrieval source for subsequent RAG question answering.

[0067] In one embodiment, in step 106, question text data is obtained, and parsing the doctor's direction from the question text data includes: a pre-multi-task learning model identifies intent information of the question text data, and constructs query conditions based on the intent information; wherein the intent information includes doctors, diseases, drugs, ancient books and materials, intent types, and whether it is in the field of traditional Chinese medicine; and the query conditions are input into the RAG system to locate doctors.

[0068] Traditional RAG-based question answering systems usually adopt the following process:

[0069] 1. Convert user queries into vector representations;

[0070] 2. Retrieve document fragments with similar vectors from the knowledge base, for example, through a vector index or Elasticsearch;

[0071] 3. The retrieved context is combined with the original question and input into the large language model to generate the final answer.

[0072] Although the above process has achieved certain results in common question-answering systems, it has obvious shortcomings when dealing with corpora with complex structures and unique language styles, such as ancient Chinese medical texts:

[0073] 1. Traditional Chinese Medicine Q&A scenarios are highly professional and context-dependent;

[0074] 2. Different medical schools may have different understandings of the same symptom;

[0075] 3. Entities such as drugs, prescriptions, ancient book names, and symptoms must be accurately identified;

[0076] 4. User intent is often implicit in complex sentences or metaphorical expressions.

[0077] Therefore, relying solely on shallow embedding representations for semantic matching cannot support high-quality question answering.

[0078] In order to gain a deeper understanding of user problems, this embodiment proposes a multi-task learning model as a pre-processor for the subsequent RAG system.

[0079] This model uses the BERT model as the foundation for semantic understanding. To better adapt the model to the linguistic characteristics of ancient TCM texts, a large-scale corpus of ancient TCM texts was constructed, including general classics such as the Yellow Emperor's Classic of Internal Medicine, Treatise on Cold Damage, and Compendium of Materia Medica, combined with all historical texts from specific TCM schools. Domain-adaptive pre-training was performed using the Masked Language Modeling (MLM) objective. After pre-training, a domain-specific language model, BERT, was obtained, capable of expressing TCM semantics.

[0080] To train the multi-task learning model, we manually constructed a high-quality sample set of Q&A questions from ancient Chinese medicine books, covering real user questions and typical medical inquiries. Using large language models such as GPT-4 as annotators, we performed multi-dimensional semantic understanding annotations on each question, including but not limited to:

[0081] Doctor identification (Doctor): such as "You Zaijing" and "Li Shicai";

[0082] Disease identification (Disease): such as "plague", "cough", "consumption";

[0083] Medicine identification: such as "Scutellaria baicalensis" and "Glycyrrhiza uralensis";

[0084] Ancient book positioning (Book): such as "Shou Shi Qing Bian";

[0085] Intent type: such as "seeking solution", "seeking source", "seeking explanation";

[0086] Whether it is the field of Traditional Chinese Medicine (Domain): Binary classification.

[0087] Each question is given a structured annotation label y={y1,y2,...,y n}, where each element represents the above recognition result, forming a training data set.

[0088] Based on the above BERT model, we perform multi-task fine-tuning and design the following learning objectives for the multi-task learning model:

[0089] Ltotal =L intent +L entity +L domain

[0090] Among them, L intent Represents multi-label classification loss (such as using Binary Cross Entropy); L entity represents the named entity recognition (NER) loss (such as CRF layer or token-level classification); L domain Indicates whether it is a binary classification loss in the field of traditional Chinese medicine;

[0091] The final model can output the following structured semantic representation:

[0092] {"doctor":"Ye Tianshi",

[0093] "disease":"typhoid fever",

[0094] "medicine":["Ginger","Cinnamon Twig"],

[0095] "book":"",

[0096] "intent":["seek treatment","seek source"],

[0097] "domain":"Traditional Chinese Medicine"}

[0098] After training the multi-task learning model, question text data is fed into the model to capture intent information, including information about doctors and diseases. This intent information is then structured into query conditions, creating more precise search queries and providing structured prompts to enhance large-scale model generation. The query conditions are then fed into the downstream RAG system as a query enhancement module, filtering or weighting the recall of ancient text materials, significantly improving the accuracy, expertise, and interpretability of the Q&A.

[0099] In one embodiment, the method also includes: evaluating the response data, performing structured management on the response data and the question text data corresponding to the response data based on the evaluation results, and constructing a semantic retrieval for the question text data; for new question text data, if the semantic retrieval matches question text data with a similarity higher than a first threshold, reusing the response data corresponding to the matched question text data as the new question text data corresponding response data; if the semantic retrieval matches a similarity not higher than the first threshold and higher than the second threshold, injecting the matched question text data corresponding response data as prompt information into the generation process of the new question text data corresponding response data.

[0100] Traditional RAG systems lack the ability to reuse excellent answers. They restart the search and large language model processing steps for each user's answer. In the end, the originally good answers may change due to the randomness of the large language model.

[0101] Therefore, this embodiment stores and indexes high-quality question-answer pairs and includes them in the memory library. High-quality question-answer pairs include those whose answer data has received positive feedback from users, those whose answer data has passed manual review or expert confirmation, and those whose answer data has been automatically assessed by the RAG system as having high accuracy, completeness, and interpretability.

[0102] These question-answer pairs are structured and managed, and a semantic index is built for the questions to support subsequent fast retrieval.

[0103] When presented with new question text, the system first searches its memory database for historical question text with similar semantics. If it finds a question with a similarity above a pre-set threshold, it reuses the corresponding historical answer data without having to re-run the search and generation process. This mechanism is particularly suitable for frequently asked questions, templated questions, or questions with minimal variability in user phrasing.

[0104] For questions that are similar but not identical to previous questions—that is, when the similarity is within a preset first threshold but above a preset second threshold—the historical response data can be injected as prompts into the large language model's generation process, guiding the model to optimize generation based on the historical response data. This approach to injecting answer prompts ensures generation flexibility while enhancing the stability and consistency of answers.

[0105] In one embodiment, the method further includes: updating, supplementing or replacing text data for constructing semantic retrieval.

[0106] Through the above operations, version management and quality labeling of historical response data stored in the memory library are supported. The labeling content includes labels such as "expert review", "user satisfaction", and "manual polishing". At the same time, it supports updating, supplementing or replacing historical response data stored in the memory library.

[0107] This embodiment can continuously improve the quality of historical response data by reflowing high-quality question-answer pairs for model fine-tuning, forming a closed-loop mechanism for knowledge enhancement, thereby achieving the goal of continuously optimizing model performance.

[0108] Compared with the traditional RAG system, the response data multiplexing mechanism proposed in this application has the following advantages:

[0109] 1. Response efficiency: directly return the answer by matching historical response data;

[0110] 2. Answer stability: The output is stable after hitting the cache;

[0111] 3. Cost control: A large number of questions can skip the generation process;

[0112] 4. Closed-loop capability: User feedback drives knowledge accumulation;

[0113] 5. Cold start: A high-frequency question and answer database can be quickly built.

[0114] In one embodiment, the method further includes: in the process of generating response data, storing the target knowledge units related to the response data in the cache stream in sequence; after completing the response data generation, extracting the positioning information of the ancient book materials from the target knowledge units and outputting the positioning information.

[0115] This embodiment dynamically generates references to the original text while streaming the answer. By locating ancient book materials, the answer data is supported by the text of the ancient book paragraphs. This can not only ensure the quality of the answer, but also facilitate subsequent verification of whether the answer of the large model is an illusion.

[0116] Specifically, during the knowledge graph construction phase (step 102), ancient Chinese medical texts are semantically segmented, numbered, and structured. During the retrieval phase (step 104), not only is the content of the relevant target knowledge units returned, but also the location information for each target knowledge unit, including the source book, chapter, page number, and paragraph number. This structured location information is retained during the answer generation phase for use by the citation mechanism.

[0117] When generating the response data in step 106, use a prompt to let the model output the answer and then output the target knowledge unit ID to which the answer belongs in real time. Create a cache stream of about 50 characters, so that the RAG system only outputs the response data when answering questions. When the target knowledge unit is involved, the corresponding target knowledge unit ID is cached first and not displayed on the front end. After the response data is generated, the referenced target knowledge unit is ready, and a post-processing is performed, that is, the corresponding positioning information is found according to the target knowledge unit ID, and only the positioning information is displayed on the front end, such as "Quoted from "Shou Shi Qing Bian" pages 5-7". Range annotation is supported when merging multiple paragraphs into a quotation.

[0118] The front-end supports multiple output formats for positioning information, such as hover viewing, automatic display at the end, etc. The output structure also supports JSON format, which facilitates further rendering and format control on the system front-end.

[0119] In one embodiment, a confidence scoring mechanism is used to assess the accuracy of the target knowledge units cited. For parts generated by the model but lacking clear support, the system will provide prompts such as "no original text support" or "not directly appearing in ancient books" to improve the credibility of the response data.

[0120] In summary, the present invention has designed a question-answering system based on school inheritance, which integrates medical inheritance into the answers, guides students to learn along the medical inheritance line, and then helps students systematically learn the theoretical system of traditional Chinese medicine, understand the development and evolution of traditional Chinese medicine theory, and grasp its essence. At the same time, traditional Chinese medicine knowledge is rich and complex. If there is no guidance from the inheritance line, learning can easily fall into fragmentation. With medical inheritance as a clue, doctors, works and academic viewpoints of various periods can be connected in series to form a complete knowledge system. Learning in this way can avoid viewing knowledge points in isolation, and make the knowledge learned interrelated and integrated.

[0121] In addition, the present invention also proposes a hierarchical knowledge graph construction scheme, which can improve the accuracy of the final question and answer. Combined with a sophisticated multi-task learning model, the overall response accuracy is improved to a certain extent; through sophisticated cutting of ancient book materials, the end-to-end accuracy is improved; by reusing excellent historical answers, the stability of the response system can be improved; and through the design of the original text citation mechanism, hallucinations can be reduced.

[0122] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0123] Based on the same inventive concept, the embodiment of the present application also provides an auxiliary question-answering method and system based on the inheritance of schools of thought of ancient Chinese medicine books for realizing the above-mentioned auxiliary question-answering method and system based on the inheritance of schools of thought of ancient Chinese medicine books. The implementation scheme for solving the problem provided by this system is similar to the implementation scheme recorded in the above-mentioned method, so the specific limitations in the embodiments of one or more auxiliary question-answering methods and systems based on the inheritance of schools of thought of ancient Chinese medicine books provided below can be referred to the limitations of the auxiliary question-answering methods and systems based on the inheritance of schools of thought of ancient Chinese medicine books above, and will not be repeated here.

[0124] In one embodiment, an auxiliary question-answering method and system based on the inheritance of schools of ancient Chinese medicine texts is provided, comprising: a knowledge base for obtaining a knowledge graph of different medical schools; nodes of the knowledge graph are doctors of corresponding medical schools, and each node is bound to a literature subsystem of the corresponding doctor, and edges are the master-disciple inheritance relationships between doctors;

[0125] The retrieval module is used to introduce the pre-trained RoBERTa large model into the literature subsystem to segment the relevant ancient medical books in the literature subsystem and construct the index structure of the vector space;

[0126] The output module is used to obtain question text data, parse medical references from the question text data, perform semantic routing under the guidance of the knowledge graph, and generate response data for the question text data in combination with the index structure.

[0127] Each module in the aforementioned assisted question-answering method and system based on the inheritance of ancient Chinese medical texts and schools can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0128] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in all the above method embodiments when executing the computer program.

[0129] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in all the above method embodiments are implemented.

[0130] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in all the above method embodiments when executed by a processor.

[0131] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.

[0132] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, etc., but are not limited to these.

[0133] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0134] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. An auxiliary question-answering method based on the inheritance of ancient Chinese medical books, characterized in that: The method comprises: Obtain a knowledge graph of different medical schools; the nodes of the knowledge graph are doctors of the corresponding medical school, and each node is bound to the corresponding doctor's literature subsystem, and the edges are the master-disciple inheritance relationships between doctors; Introducing a pre-trained RoBERTa large model into the literature subsystem to segment the relevant ancient medical books and materials corresponding to the literature subsystem, and constructing an index structure of the vector space; Obtain question text data, parse the medical expert reference from the question text data, perform semantic routing under the guidance of the knowledge graph, and generate response data for the question text data in combination with the index structure.

2. The method according to claim 1, characterized in that The pre-trained RoBERTa model is introduced into the literature subsystem to segment the relevant ancient medical books and materials of the literature subsystem, and the index structure of the vector space is constructed, including: Construct a corpus containing knowledge unit annotations through pre-cutting; Using the corpus to train the RoBERTa large model to learn knowledge unit boundaries; Input the ancient book data into the trained RoBERTa large model, determine the paragraph boundaries based on the learned knowledge unit boundaries, and merge the paragraphs into segments based on the semantic similarity of the paragraphs to generate the target knowledge units; The index structure is constructed based on the vectorized target knowledge unit.

3. The method according to claim 1, characterized in that The obtaining of question text data and parsing the doctor's direction from the question text data includes: The pre-multi-task learning model identifies intent information from the question text data and constructs query conditions based on the intent information; wherein the intent information includes doctors, diseases, drugs, ancient books, intent types, and whether it is in the field of traditional Chinese medicine; The query conditions are input into the RAG system to locate the doctor.

4. The method according to claim 1, wherein The method further comprises: Evaluating the response data, performing structured management on the response data and question text data corresponding to the response data according to the evaluation results, and constructing a semantic search for the question text data; For new question text data, if the semantic retrieval matches the question text data with a similarity higher than a first threshold, the answer data corresponding to the matched question text data is reused as the answer data corresponding to the new question text data; if the semantic retrieval matches the similarity not higher than the first threshold and higher than the second threshold, the answer data corresponding to the matched question text data is injected as prompt information into the generation process of the answer data corresponding to the new question text data.

5. The method according to claim 4, characterized in that The method further comprises: The question text data used to construct the semantic retrieval is updated, supplemented or replaced.

6. The method according to claim 2, characterized in that The method further comprises: In the process of generating the response data, the target knowledge units related to the response data are sequentially stored in a cache stream; After the response data is generated, the location information of the ancient book material is extracted from the target knowledge unit and the location information is output.

7. An auxiliary question-answering system based on the inheritance of ancient Chinese medical books and schools, characterized by: The system comprises: A knowledge base is used to obtain knowledge graphs of different medical schools; the nodes of the knowledge graph are doctors of the corresponding medical schools, and each node is bound to the corresponding doctor's literature subsystem, and the edges are the master-disciple inheritance relationships between doctors; A retrieval module is used to introduce a pre-trained RoBERTa large model into the document subsystem to segment the relevant ancient books and materials of the medical practitioners in the document subsystem and construct an index structure in the vector space; The output module is used to obtain question text data, parse medical references from the question text data, perform semantic routing under the guidance of the knowledge graph, and generate response data for the question text data in combination with the index structure.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.