Construction method of traditional Chinese medicine alert question-answering system based on large language model
By building a Chinese medicine pharmacovigilance Q&A system based on a large language model, using the PV-RAG model combined with BM25 and DPR models for search and text extraction, the shortcomings of the Chinese medicine pharmacovigilance system in understanding the semantics of Chinese medicine literature and providing personalized services are solved, and a high accuracy and reliability question-and-answer service is achieved.
Patent Information
- Application Number
- CN202510678342.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-27
AI Technical Summary
The existing Chinese medicine pharmacovigilance system relies on keyword matching, making it difficult to understand the complex semantics of Chinese medicine literature, resulting in inaccurate search results; large language models lack professional knowledge in the field of Chinese medicine, and are prone to hallucinations of answers; Chinese medicine knowledge is widely sourced, and it is difficult to integrate multi-source heterogeneous data, and lack a dynamic update mechanism; traditional systems cannot provide personalized services based on user background and needs.
A Chinese medicine pharmacovigilance question and answer system based on a large language model is built, a PV-RAG model is used to search with BM25 and DPR models, a MRI-BERT architecture is used for text extraction and reordering, query rewrite and answer generation is used based on a large language model, and personalized services are provided through the intelligent question and answer application.
It improves the accuracy and reliability of Chinese medicine pharmacovigilance Q&A, effectively solves the inaccurate answers of traditional systems in the field of traditional Chinese medicine, realizes the deep semantic understanding of Chinese medicine literature and the effective integration of multi-source data, and provides personalized Q&A services.
Smart Images

Figure CN120216657A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of natural language processing and pharmacovigilance, and particularly to a method for constructing a traditional Chinese medicine pharmacovigilance Q&A system based on a large language model. Background Art
[0002] In the field of traditional Chinese medicine, especially in the field of traditional Chinese medicine pharmacovigilance, efficient search and dialogue understanding are the keys to ensuring the safety of traditional Chinese medicine use. Traditional pharmacovigilance systems are often limited by keyword matching and surface-level retrieval methods, making it difficult to meet the complexity and professionalism requirements of traditional Chinese medicine literature. In addition, traditional Chinese medicine experts and ordinary users need to have natural and in-depth conversations with the system to more comprehensively understand and apply the knowledge in the literature.
[0003] Specifically, the retrieval methods of existing retrieval systems mainly rely on keyword matching, which makes the system very sensitive to the selection of accurate keywords. However, in the field of traditional Chinese medicine, traditional Chinese medicine terms may be very professional and complex, and it is difficult for traditional retrieval systems to understand the context relationships of these terms. At the same time, relying solely on surface keyword matching, traditional systems usually cannot understand the semantic structure of the text, which results in the system being unable to understand the complex context and the actual meaning of professional terms in traditional Chinese medicine literature. In addition, there may be multiple expressions in traditional Chinese medicine literature to describe similar concepts, and traditional systems may not be able to effectively capture and process this diversity. This limits the system's comprehensive understanding of the literature content and search effect, and affects the accuracy of search results.
[0004] Although some studies have used natural language processing technologies, such as word vector models, to improve the semantic understanding of literature retrieval and the ability to understand context relationships. However, these models constructed based on general languages usually cannot capture complex traditional Chinese medicine contexts and lack the ability to align traditional Chinese medicine texts, resulting in insufficient accuracy of search results. At the same time, traditional retrieval systems based solely on word vector model matching technology usually lack the ability of personalization and cannot customize search results according to the specific needs and backgrounds of users. This makes the user experience relatively fixed, less flexible and user-friendly.
[0005] In summary, existing retrieval systems rely on keyword matching, making it difficult to understand the complex semantics of traditional Chinese medicine literature, resulting in inaccurate search results; large language models lack professional knowledge in the field of traditional Chinese medicine and are prone to answer hallucinations; the sources of traditional Chinese medicine knowledge are extensive, the integration of multi-source heterogeneous data is difficult, and there is a lack of a dynamic update mechanism; traditional systems cannot provide personalized services according to the user's background and needs. Summary of the Invention
[0006] The present invention aims to solve at least one of the technical problems existing in the related technologies to a certain extent.
[0007] The object of the present invention is to provide a construction method of a traditional Chinese medicine drug safety monitoring Q&A system based on a large language model, to construct an intelligent Q&A system, improve the accuracy and reliability of traditional Chinese medicine drug safety monitoring Q&A, and provide intelligent support for the safety of traditional Chinese medicine use.
[0008] To achieve the above object, on the one hand, the present invention provides a construction method of a traditional Chinese medicine drug safety monitoring Q&A system based on a large language model, including:
[0009] Collect and organize traditional Chinese medicine drug-related data to establish a traditional Chinese medicine knowledge base;
[0010] Construct a retrieval augmented generation model PV-RAG for drug safety monitoring, and use the traditional Chinese medicine knowledge base as a training set to train and optimize the PV-RAG model; this PV-RAG model includes a retrieval module, a text extraction module, a query rewriting module, and a generation module;
[0011] Among them, the retrieval module of the PV-RAG model adopts a hybrid strategy of sparse and dense retrieval, combines the BM25 and DPR models, extracts keywords according to the user's query content, and retrieves document fragments related to the keywords from the traditional Chinese medicine knowledge base;
[0012] The text extraction module uses the MRI-BERT architecture to re-rank and judge the relevance of the retrieved document fragments, and extracts the text related to the user's query content;
[0013] The query rewriting module is based on the large language model to expand and rewrite the query content that fails the relevance judgment, and re-enter the expanded and rewritten query content into the retrieval module;
[0014] The generation module uses the large language model to generate answers to traditional Chinese medicine drug safety monitoring Q&A related to the user's query content based on the text extracted by the text extraction module;
[0015] Construct an intelligent Q&A application end, connect to the traditional Chinese medicine knowledge base and the trained PV-RAG model, and generate answers to traditional Chinese medicine drug safety monitoring according to the natural language questions input by the user.
[0016] A further preferred technical solution of the present invention is that the retrieval module of the PV-RAG model is constructed as:
[0017] Use the BM25 model to match the extracted keywords with the traditional Chinese medicine knowledge base for word frequency and inverse document frequency, and initially screen out relevant document fragments;
[0018] Use the DPR model to perform semantic encoding on the query content and the document fragments through the query encoder and the passage encoder, calculate the BERTScore relevance score, and supplement the semantically relevant document fragments not retrieved by the BM25 model;
[0019] Merge the retrieval results of the BM25 model and the DPR model to form a diverse set of candidate document fragments.
[0020] Preferably, the text extraction module of the PV-RAG model is constructed as:
[0021] Use the MRI-BERT architecture to re-rank and judge the relevance of candidate document fragments. By calculating the similarity score between the query content and the high-dimensional embedding vectors of the candidate document fragments, filter out the text highly relevant to the user's query content;
[0022] Input the query content that fails the relevance judgment into the query rewriting module of the PV-RAG model.
[0023] Preferably, the query rewriting module of the PV-RAG model is constructed as:
[0024] Based on the large language model, set the prompt words for synonym expansion, expand and rewrite the query content that fails the relevance judgment, and generate synonymous query content related to the original query content;
[0025] Input the rewritten synonymous query content into the retrieval module of the PV-RAG model again for a new round of retrieval.
[0026] Preferably, the generation module of the PV-RAG model is constructed as:
[0027] For the query content that passes the relevance judgment, by setting the prompt words for generating answers, use the large language model to generate the answers to the traditional Chinese medicine drug vigilance questions and answers based on the extracted relevant text.
[0028] Preferably, the method for collecting, organizing the data related to traditional Chinese medicine drugs and establishing the traditional Chinese medicine knowledge base is as follows:
[0029] Collect and organize the data of traditional Chinese medicine instructions, traditional Chinese medicine classics and clinical research literature;
[0030] Convert the collected and organized data into a document set and perform word segmentation processing as the traditional Chinese medicine knowledge base;
[0031] Divide the traditional Chinese medicine knowledge base into a training set, a validation set and a test set for the training and validation of the PV-RAG model.
[0032] Preferably, adopt a dynamic update mechanism to incorporate new data related to traditional Chinese medicine drugs into the traditional Chinese medicine knowledge base, and regularly use the updated traditional Chinese medicine knowledge base to retrain the PV-RAG model.
[0033] On the other hand, the present invention provides a non-transitory computer-readable storage medium storing computer instructions that cause a computer to execute the above-mentioned method for constructing a traditional Chinese medicine pharmacovigilance Q&A system based on a large language model.
[0034] On another aspect, the present invention provides an electronic device, comprising: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The processor calls the logic instructions in the memory to execute the above-mentioned method for constructing a traditional Chinese medicine pharmacovigilance Q&A system based on a large language model.
[0035] On yet another aspect, the present invention provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer executes the above-mentioned method for constructing a traditional Chinese medicine pharmacovigilance Q&A system based on a large language model.
[0036] Beneficial effects: The method for constructing a traditional Chinese medicine pharmacovigilance Q&A system based on a large language model of the present invention constructs a traditional Chinese medicine pharmacovigilance Q&A system based on a large language model and retrieval-augmented generation technology. Through a self-developed text embedding model, it realizes the deep semantic understanding of traditional Chinese medicine literature, effectively integrates multi-source data such as traditional Chinese medicine instructions, traditional Chinese medicine classics, and clinical research literature, and provides rich knowledge support for traditional Chinese medicine pharmacovigilance Q&A.
[0037] The present invention constructs a traditional Chinese medicine pharmacovigilance dialogue understanding system by using RAG technology and a large language model, which can parse the queries input by users in real time and provide personalized Q&A services, significantly improving the accuracy and reliability of traditional Chinese medicine pharmacovigilance Q&A, and effectively solving the problems of inaccurate answers and answer hallucinations existing in traditional large language models in the field of traditional Chinese medicine.
[0038] The present invention proposes a re-ranking model based on importance, which sorts the documents recalled by multiple parties through a unified measurement method, improving the accuracy and personalization degree of search results, and ensuring that users can quickly obtain the most relevant and valuable information.
[0039] The present invention adopts a dynamic update mechanism, which can timely incorporate new traditional Chinese medicine research results and clinical data into the knowledge base, maintaining the timeliness and accuracy of the model, enabling the system to better cope with the dynamic changes of traditional Chinese medicine knowledge, and adapting to the continuously updated traditional Chinese medicine research and clinical practice needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a processing flow chart of the method for constructing a traditional Chinese medicine pharmacovigilance Q&A system based on a large language model of the present invention.
[0041] Figure 2 The PV-RAG model framework diagram constructed for the present invention. Detailed implementation manners
[0042] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention, and they should not be construed as limiting the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts fall within the scope of protection of the present invention. In the description of the present invention, it should be understood that the terms used are only for the purpose of description and cannot be construed as indicating or implying relative importance.
[0043] The following is combined with Figure 1 - Figure 2 Describe the construction method of the traditional Chinese medicine drug safety question and answer system based on the large language model provided by the present invention.
[0044] Example 1: This example aims to provide a construction method of a traditional Chinese medicine drug safety question and answer system based on the large language model, so as to construct an innovative traditional Chinese medicine drug safety retrieval and dialogue understanding system, enabling the system to make full use of large language models (LLMs) and retrieval augmented generation (RAG) technologies to address the following challenges:
[0045] Necessity of semantic understanding: Traditional systems often can only perform surface-level grammar and keyword matching and cannot deeply understand the complex semantic structures in traditional Chinese medicine literature. Information such as the ingredients, efficacy, dosage and usage, and contraindications of traditional Chinese medicine is highly professional and complex, and it is difficult for traditional systems to accurately understand and process. This example introduces a large language model to enable the constructed system to better understand and process professional terms and complex contexts in the field of traditional Chinese medicine.
[0046] Retrieval efficiency and accuracy: Due to the large number of terms and concepts in the field of traditional Chinese medicine, traditional retrieval systems may produce a large number of irrelevant or duplicate results in large-scale literature knowledge bases. This example introduces RAG technology to improve the accuracy of retrieval. By performing advanced retrieval and then generation, relevant information can be provided more targeted to improve retrieval efficiency and the accuracy of retrieval results.
[0047] Demand for dialogue interaction: Compared with traditional static literature retrieval, traditional Chinese medicine experts and ordinary users hope to be able to have dynamic and natural conversations with the system. This example introduces dialogue generation technology to enable the constructed system to interact with users in a more user-friendly manner, answer questions and provide personalized information, enhancing the user experience.
[0048] Increasing personalized needs: Different users may have different requirements for the same traditional Chinese medicine. The system constructed in this embodiment, through personalized modeling, takes into account factors such as the user's professional background and historical queries to better meet the personalized search and dialogue needs of users and provide more accurate information services.
[0049] Based on the above challenges and requirements, the construction method of the traditional Chinese medicine pharmacovigilance Q&A system based on the large language model in this embodiment mainly constructs three parts of the system, including the construction of the traditional Chinese medicine knowledge base, the construction of the retrieval augmented generation model PV-RAG for pharmacovigilance, and the construction of the intelligent Q&A application end.
[0050] I. Construction of the traditional Chinese medicine knowledge base:
[0051] (1) Widely collect various types of traditional Chinese medicine-related data, including traditional Chinese medicine instructions, traditional Chinese medicine classics, clinical research literature, etc., and establish a traditional Chinese medicine knowledge base to ensure that the constructed system has a comprehensive and rich traditional Chinese medicine knowledge base.
[0052] Since the constructed system focuses on data related to the safety monitoring of traditional Chinese medicine, the data quality will directly affect the performance of the model in drug adverse reaction identification and medication risk assessment. To ensure the accuracy of the experimental results, this embodiment has carried out special designs from multiple dimensions such as the authority of data in the field of traditional Chinese medicine, the integrity of toxicity reaction records, and the special data segmentation strategy of traditional Chinese medicine.
[0053] First, the data mainly comes from the drug instructions publicly disclosed by authoritative parts, the drug instructions publicly available on the official websites of major drug manufacturers or the National Medical Products Administration (NMPA), the National Center for Adverse Drug Reactions Monitoring (ADR), the State Administration of Traditional Chinese Medicine, and the safety data of traditional Chinese medicine included in the Chinese Pharmacopoeia, including the safety instructions of traditional Chinese medicine decoction pieces, proprietary Chinese medicines, and classical prescriptions. The total number of drug instructions collected involves 13,639 drugs, ensuring the scientificity and professionalism of the research data. Secondly, in order to ensure the diversity of data types, the data of each drug collected includes multiple contents such as the name, ingredients, indications, usage and dosage, contraindications, precautions, and drug interactions of the drug.
[0054] These data sources are extensive and cover information such as the name, ingredients, efficacy, usage and dosage, contraindications, precautions, and adverse reactions of traditional Chinese medicine, providing a solid foundation for subsequent Q&A.
[0055] (2) Subsequently, the established traditional Chinese medicine knowledge base needs to be used for model training, so further processing of the traditional Chinese medicine knowledge base is also required.
[0056] Convert the collected drug data into a document collection, which contains the relevant drug description text content, then rearrange it, and select 10,000 pieces of data with detailed content descriptions as the initial collection; then use the ChineseRecrusiveTextSplitter tokenization tool to tokenize the data set to form a data set; subsequently, divide the data set into a training set, a validation set, and a test set by random sampling for training and validating the retrieval performance of the subsequent PV-RAG model. The specific data set division is as follows: The data set contains a total of 10,000 pieces of drug data, and the validation set and the test set contain 1,000 and 120 pieces of drug data information respectively. Both the validation set and the test set come with drug data that can be correctly answered, which is used to evaluate the accuracy of the generated answers of the retrieval enhancement module, and the rest are used as the training set.
[0057] Considering that there will be a large number of questions in the question-and-answer system, during the data preprocessing stage, the text of the data set also needs to be processed into question-and-answer pair data, including questions and standard answers, to help improve the model's accuracy in answering specific questions. The specific description of the question design is as follows:
[0058] Regarding the types of questions, in order to compare the characteristics and advantages of the PV-RAG model proposed in this embodiment with other large language models in the field of safe medication Q&A from multiple angles, this embodiment designs a set of multi-type question sets. The question types include four categories: single-choice, multiple-choice, judgment, and question-and-answer, covering a variety of task scenarios from simple selection to complex reasoning, aiming to compare and test the understanding and retrieval capabilities of each model for different types of questions and the ability to generate accurate answers. Through the combination of four types of questions, it can effectively cover various actual application scenarios that may be encountered in the drug Q&A system, ensuring the comprehensiveness and representativeness of the verification results.
[0059] Regarding the content of the questions, in order to cover as much medication safety knowledge as possible, the content design of the questions mainly includes four directions: population applicability questions, drug use taboo questions, precautions for drug use questions, and drug applicable symptoms questions. Population applicability questions explore which populations can safely use specific drugs and which populations will have taboo reactions when using drugs, etc.; drug use taboo questions are used to understand the contraindications and unsuitable situations of drugs; precautions for drug use questions focus on the matters that need special attention when using drugs; drug applicable symptoms questions ask which specific symptoms or diseases the drug is applicable to.
[0060] (3) Adopt a dynamic update mechanism to timely incorporate new Chinese medicine research results and clinical data into the Chinese medicine knowledge base, and regularly update and optimize the model to maintain the timeliness and accuracy of the model.
[0061] II. Construct a retrieval-enhanced generation model PV-RAG for pharmacovigilance:
[0062] The constructed PV-RAG includes a retrieval module, a text extraction module, a query rewriting module, and a generation module. The PV-RAG model is trained using the divided training set, validation set, and test set, and the model is fine-tuned with a large amount of traditional Chinese medicine pharmacovigilance Q&A data to optimize the model's parameters.
[0063] (1) Retrieval module: First, the BM25 model is used to perform word frequency and inverse document frequency matching on the retrieval keywords in the document collection to quickly screen out the initially relevant document fragments. Then, the DPR model is used to perform semantic encoding on the query and document fragments through the query encoder and the passage encoder, calculate the BERTScore correlation score, and further supplement the semantically relevant document fragments not retrieved by the BM25 model. Finally, the retrieval results of BM25 and DPR are merged to form a diverse set of candidate document fragments. The specific description of the retrieval module is as follows:
[0064] The dense retrieval model (DPR) constructs a query encoder and a passage encoder through the BERT model respectively, and converts the query and document fragments into dense vector representations. DPR uses a deep learning model to capture the semantic associations between texts, calculates the BERTScore between the query vector and the document fragment vector to measure their correlation. Compared with traditional sparse retrieval methods, DPR shows better performance on most datasets. However, in the tasks of the pharmacovigilance field, the performance of DPR will be affected by the shift of data distribution. Most of the data involved in the tasks of the pharmacovigilance field has been normalized, resulting in DPR sometimes being difficult to effectively capture the semantic features of the pharmacovigilance field. At this time, the traditional sparse retrieval method BM25 shows higher robustness in some cases due to its retrieval strategy based on word frequency and inverse document frequency (TF-IDF). BM25 directly performs retrieval based on word matching and can provide reliable retrieval results without semantic prior knowledge.
[0065] To make full use of the advantages of sparse retrieval and dense retrieval, the PV-RAG model adopts a hybrid retrieval method of BM25 and DPR in the retrieval stage, taking the union of their retrieval results to construct a more comprehensive candidate document set. First, BM25 is used to perform word frequency and inverse document frequency matching on the retrieval keywords in the document collection to quickly screen out the initially relevant document fragments. Then, DPR is used to perform semantic encoding on the query and document fragments through the query encoder and the passage encoder, calculate the BERTScore correlation score, and further supplement the semantically relevant document fragments not retrieved by the BM25 model. Finally, the retrieval results of BM25 and DPR are merged to form a diverse set of candidate document fragments.
[0066] In the information retrieval task, the union operation occurs between the results of sparse retrieval (BM25) and dense retrieval (DPR). Suppose the query is retrieved in the dataset , and BM25 and DPR respectively return a set of candidate document collections, denoted as and respectively. The union operation is expressed as:
[0067] ;
[0068] In the large model generation task, the input quality directly affects the accuracy of the generated result. The initial retrieval result may contain fragments unrelated to the query. If directly input into the generation model, it may cause the model to deviate from the user's intention, thus affecting the reliability of the generation. Therefore, re-ranking is performed after the retrieval stage to ensure that only fragments closely related to the query are input into the generation stage, which is also a necessary step to improve the task performance.
[0069] Similarity calculation is the core method to achieve this goal. It quantifies the semantic similarity between the query and the document fragments, and filters out the most relevant content. Especially in the field of pharmacovigilance, where the knowledge expression form is diverse, similarity calculation can help the model better capture semantic relationships. By introducing the guidance of the multi-head attention mechanism and the dynamic fusion strategy, a unified representation of the paragraph is finally formed. Finally, based on the relationship between the query and the paragraph representation, its similarity is calculated, and the calculation formula is:
[0070] ;
[0071] where is the th paragraph vector, is the multi-vector representation of the paragraph, and is the total paragraph vector.
[0072] (2) The text extraction module, whose main goal is to ensure the extraction of text fragments highly relevant to the user's query from the documents obtained in the retrieval stage. This process includes re-ranking, extracting, and utilizing the key information in the document fragments.
[0073] For the re-ranking work, since traditional BERT requires the concatenation of the query and the document as input when directly used for re-ranking, resulting in high computational overhead, while MR-BERT and its variants first independently encode the query and the document through a two-tower structure or a lightweight interaction module, and then calculate the fine-grained interaction, significantly reducing the computational cost. Therefore, in this embodiment, the MRI-BERT architecture is used for re-ranking.
[0074] For a given drug instruction dataset , where is the query text, is the th query text corresponding label, each is an integer variable of 0 or 1, that is . Define the objective of the re-ranking model as:
[0075] ;
[0076] where is the evaluation metric of the label and the prediction.
[0077] is the corresponding set of scores for the following documents. Based on this architecture, the query and document fragments are encoded as high-dimensional embedding vectors, and by calculating the similarity scores between these vectors, the relevance between the query and the document fragments can be effectively judged.
[0078] By using the MRI-BERT architecture to pre-train the drug instruction manual documents, the encoder can better understand the semantic relationship between the query and the positive example paragraphs, and has a greater advantage in obtaining information related to the expected results.
[0079] Among the documents obtained in the retrieval stage, there may be information irrelevant to the current problem, which has a negative impact on the processing process of the large model, resulting in the model being unable to focus on the core part of the problem, thus interfering with its understanding and extraction of key information. Therefore, ensuring relevance judgment in the retrieval stage to filter out irrelevant information and providing concise and relevant document inputs is crucial for improving the generation quality of the large model.
[0080] Since the GLM-4 model has better semantic understanding and efficient multilingual processing capabilities, this embodiment is based on the GLM-4-9B model and uses prompt engineering for relevance judgment. Set the input of the prompt words for relevance judgment as:
[0082] Judge whether the content of {input text 1} and {input text 2} is relevant and output {result}. If relevant, return "relevant"; if not relevant, return "not relevant".
[0083] {input text 1}={text1}
[0084] {input text 2}={text2}
[0085] {result}: ["relevant", "not relevant"]
[0087] Among them, text1 is the input question text, and text2 is the text found in the retrieval stage. Through this prompt, GLM-4-9B can stably return the judgment result of "relevant" or "irrelevant".
[0088] (3)Query rewriting module. For the content that fails the relevance judgment, it enters the query rewriting part. In this link, for the query that fails to obtain effective information in the first retrieval, words closely related to the words obtained from the previous keyword extraction are added, such as synonyms or abbreviations. By using a text filter to evaluate the effectiveness of the content, query rewriting based on synonym expansion is initiated, which helps to reduce the deviation of the user's intention caused by unnecessary query rewriting. The synonym expansion is completed using the GLM-4-9B model. For the synonyms generated by the GLM-4-9B model, it enters the retrieval part again for a new round of more targeted retrieval.
[0089] Set the prompt for synonym expansion as:
[0091] According to {input text 1}, perform synonym expansion on {input text 2} and output {result}.
[0092] {input text 1} = {text1}
[0093] {input text 2} = {text2}.
[0094] Among them, if {input text 2} contains n keywords, the format of {input text 2} is {keyword 1}+{keyword 2}+{Omit(n)}+{keyword n}. Among them, {Omit(n)}={Omit(n - 1)}+{keyword n - 1}, where n satisfies n≥3; when n = 2, {Omit(n)} = "NULL".
[0095] {result}={output text}. Among them, if {output text} contains m keywords, the format of {output text} is {new keyword 1}+{new keyword 2}+{Omit(m)}+{new keyword m}. Among them, {Omit(m)}={Omit(m - 1)}+{new keyword m - 1}, where m satisfies m≥3; when m = 2, {Omit(m)} = "NULL".
[0096] It is required that n + 2 ≤ m.
[0098] Among them, text1 is the input question text, and text2 is the keyword obtained before entering the retrieval link. Through this prompt, GLM-4-9B can stably return the new keyword after synonym expansion.
[0099] Taking text1 = "What adverse reactions may occur when using Chenlong Luoxin?", text2 = "Chenlong Luoxin, adverse reactions" as an example, the output is "Chenlong Luoxin, side effects, adverse reactions, risks".
[0100] Subsequently, the rewritten information is input into the retrieval controller for a more accurate and targeted secondary retrieval, thereby obtaining more relevant and valuable search results. This process can improve the information matching degree and retrieval efficiency, and more effectively meet the actual needs of the query.
[0101] (4) Generation module, using the GLM-4-9B model to generate accurate and concise Q&A answers for traditional Chinese medicine pharmacovigilance based on the extracted relevant text. By setting the prompt words for generating answers, the model can generate corresponding answers according to the input questions.
[0102] For the content judged by relevance, enter the generation model and use the GLM-4-9B large language model to generate answers. Set the prompt words for generating answers as:
[0103] Answer {input text2} based on {input text1} and output {result}.
[0105] {input text1} = {text1}
[0106] {input text2} = {text2}.
[0108] Among them, text1 is the relevant content obtained through the text extraction part, and text2 is the input question.
[0109] Taking the question "What are the precautions for using Kaidelin?" as an example, when using BERT-base-uncased as the Embedding model, the output result is:
[0110] 1. Penicillin skin test: Skin test is required before medication, and those with positive results are prohibited. Cross-allergy reaction: People allergic to one penicillin antibiotic may also be allergic to other penicillin antibiotics, and may also be allergic to penicillamine or cephalosporins.
[0111] 2. Situations for cautious use: History of allergic diseases: Patients with a history of allergic diseases such as asthma, eczema, hay fever, urticaria, etc. should use it with caution.
[0112] 3. Solution stability and concentration of the used solution: The higher the concentration of the ampicillin solution, the worse the stability, and the stability of the solution will decrease with the increase in temperature. Sensitizing substances: Sensitizing substances may be produced if the solution is left standing for too long, so it should be used as soon as possible and not left standing for a long time.
[0113] 4. Interference with diagnosis: During medication, false positives may occur when measuring urine glucose by the copper sulfate method, but it is not affected by the glucose oxidase method.
[0114] 5. High-dose use: When using high doses, serum sodium levels should be regularly monitored.
[0115] III. Building the intelligent Q&A application end:
[0116] As the application port of the traditional Chinese medicine drug safety warning Q&A system for users, the intelligent Q&A application end accesses the above-mentioned traditional Chinese medicine knowledge base and the trained PV-RAG model. For the content judged by relevance, the GLM-4-9B model is used to generate accurate and concise answers to traditional Chinese medicine drug safety warnings. By setting the prompt words for generating answers, the model can generate corresponding answers according to the input questions.
[0117] To better evaluate the performance of the constructed PV-RAG model in various Q&A scenarios, a comparative experiment is conducted on the PV-RAG constructed in this embodiment and the four large language models of Zhipu Qingyan, Wenxin Yiyan, iFlytek Spark, and Tongyi Qianwen, which are widely used in the Chinese field.
[0118] The question set for the experiment is asked according to the four directions and four question types designed above, and the same questions are successively asked on each large model. The Q&A process is as follows:
[0119] ① Input questions: The same questions are successively asked to each large model to ensure the consistency of the input questions.
[0120] ② Obtain model answers: Record the answer results of each model and compare them with the standard answers.
[0121] ③ Result recording: The answers of each model will be classified and saved according to different question types for subsequent analysis.
[0122] For the intelligent Q&A system for safe medication, the accuracy of the answers is the most intuitive indicator for evaluating its performance. Therefore, in the experiment, the accuracy rate (Accuracy ACC) of directly selecting answers for objective questions is used as the evaluation indicator for the performance of the large model.
[0123] In the multiple-choice questions and true / false questions of the objective questions, the answers are unique, and the model's answers are either correct or wrong. Therefore, the accuracy rate can be obtained by directly calculating the number of correct answers of the model. The calculation formula is:
[0124] ;
[0125] where nums represents the number of instances to be evaluated, and correct answers is the number of correct responses.
[0126] For multiple-choice questions, the answer may contain multiple options, and the accuracy rate is calculated using the method of inclusive matching. Assuming that each question has a full score of 100 points, if there are incorrect options in the model's answer, the score for that question is 0; if the model does not choose the wrong options but does not select all the correct options completely, the score for that question is calculated as follows:
[0127] ;
[0128] Among them, is the number of correct options selected, is the total number of correct options.
[0129] According to the score, the accuracy rate calculation formula for multiple-choice questions is:
[0130] ;
[0131] This method ensures a refined evaluation of the answers to multiple-choice questions and can reflect the model's ability to handle complex problems.
[0132] For subjective questions, considering the professionalism of each drug name in the answer and the diversity of text expressions in various adverse reactions and medication taboos, an artificial method is used to calculate two data items, the average accuracy rate and the average redundancy rate, as evaluation indicators for detecting the performance of large language models to ensure the comprehensiveness of the evaluation.
[0133] Taking the commonly used drug Amoxicillin and Clavulanate Potassium Chewable Tablets (Aiker) as an example, the adverse reactions marked in the drug instruction manual include: ① Gastrointestinal reactions such as nausea, vomiting, diarrhea, and pseudomembranous enteritis. ② Allergic reactions such as rash, drug fever, and asthma. ③ Anemia, thrombocytopenia, eosinophilia, etc. ④ Serum aminotransferase may increase slightly. ⑤ Superinfection caused by Candida or drug-resistant bacteria. ⑥ Occasionally, central nervous system symptoms such as excitement, anxiety, insomnia, dizziness, and abnormal behavior.
[0134] The following gives the answers generated by major language models to the question "What are the adverse reactions of using the drug Aiker?"
[0135] ① PV-RAG model
[0136] The adverse reactions of Aiker include: allergic reactions such as rash, drug fever, and asthma, anemia, thrombocytopenia, eosinophilia, slightly increased serum aminotransferase, gastrointestinal reactions such as nausea, vomiting, diarrhea, and pseudomembranous enteritis, and superinfection caused by Candida or drug-resistant bacteria.
[0137] ② Zhipu Qingyan ChatGLM-4
[0138] Skin reactions: including rashes, dryness, itching, etc. Sometimes severe skin reactions may occur, such as exfoliative dermatitis.
[0139] Digestive system reactions: such as diarrhea, nausea, vomiting, stomatitis, constipation, etc.
[0140] Liver problems: Abnormal liver function may occur, such as elevated transaminases and bilirubin.
[0141] Musculoskeletal symptoms: such as joint pain, muscle pain, etc.
[0142] Bone marrow suppression: May lead to decreased white blood cell, red blood cell and platelet counts, increasing the risk of infection and bleeding.
[0143] Respiratory system reactions: such as cough, dyspnea, etc.
[0144] Nervous system symptoms: such as headache, paresthesia, insomnia, depression, etc.
[0145] Heart problems: In rare cases, arrhythmia may occur.
[0146] During the treatment with Aikel, patients should closely monitor their own symptoms and seek medical attention in case of abnormalities.
[0147] ③ Wenxin Yiyan - 3.5
[0148] Gastrointestinal reactions: A few patients may experience mild gastrointestinal side effects such as nausea, vomiting and diarrhea. These side effects are usually mild, and taking the drug before meals may reduce discomfort.
[0149] Allergic reactions: Urticaria and measles-like rashes are occasionally seen. If urticaria or severe measles-like rashes occur, the use of the drug should be stopped immediately.
[0150] Abnormal liver function: Extremely rare cases may show temporary abnormal liver function. Therefore, during the use of this drug, patients should regularly monitor their liver function.
[0151] ④ iFlytek Spark - V4.0 Turbo
[0152] Hematological toxicity: May cause granulocytopenia and moderate anemia.
[0153] Peripheral neurotoxicity reactions: Usually manifested as decreased deep tendon reflexes, numbness is less common, and paresthesia may occur occasionally. Long-term use of the drug may lead to weakness in the lower limbs.
[0154] Digestive tract autonomic nervous system: Mainly constipation caused by intestinal paralysis, paralytic ileus is rare, and nausea and vomiting may occur occasionally.
[0155] Digestive system: May cause dyspnea or bronchospasm, and these symptoms may occur within a few minutes or hours after drug injection.
[0156] Other adverse reactions: Include progressive moderate hair loss, mandibular pain, local phlebitis, etc.
[0157] ⑤ Tongyi Qianwen - 2.5
[0158] If you are referring to the trade name or generic name of a specific drug, please provide more information or confirm the correct name of the drug so that an accurate answer can be given. For prescription drugs, the common adverse reactions are usually described in detail in the drug instruction manual, including possible side effects such as headache, nausea, rash, etc., with varying degrees of severity. For any drug, if adverse reactions occur, consult a doctor or pharmacist immediately.
[0159] As can be seen from the above answer examples, the answers given by different models vary greatly. The answers of some models go beyond the scope marked in the drug instruction manual, and their authenticity and reliability are in question, and there is a possibility of answer hallucination.
[0160] The accuracy of the answers to subjective questions is mainly measured by manual evaluation. Manual calculation and scoring of the answers generated by the model are carried out according to the professional knowledge of the drug instruction manual, including the correctness, completeness and relevance of the answers. The calculation method of the accuracy of a single question and answer is the ratio of the number of items in the answer generated by the large model that match the standard answer to the total number of items in the standard answer; the average accuracy is the average value of the accuracy of all questions and answers. This indicator can scientifically quantify the accuracy of the answers generated by the model and reflect the ability of the model to capture key information in the drug instruction manual Q&A task.
[0161] The data redundancy is used to evaluate the conciseness and information density of the answers generated by the model, and is measured by calculating the proportion of repeated or irrelevant information in the answers, aiming to avoid the model generating long but low-information answers. The calculation of redundancy is based on the ratio of the number of repeated or redundant items in the answers generated by the model to the total number of generated items. The lower the value of redundancy, the more concise and higher the information density of the answers generated by the model; on the contrary, the higher the value of redundancy, the more repeated or irrelevant information is included in the answers generated by the model, which may lead to information redundancy and reduced user reading efficiency.
[0162] Based on the above evaluation indicators, the constructed PV-RAG model is experimentally compared with various current mainstream large language models. The Q&A experiment comparison data of each model on objective questions and subjective questions are shown in Tables 1 and 2.
[0163] Table 1 Comparison of ACC data in objective question Q&A experiment
[0164]
[0165] Table 2 Comparison of subjective question answering experimental data
[0166]
[0167] Table 1 shows the question-answering performance of these five large models on three objective questions. In terms of the accuracy of multiple-choice questions and judgments, the PV-RAG model performs better, thanks to its strong contextual understanding ability and the reference basis provided by rich professional data, and shows good stability in answering accuracy. It can be seen that after having a high-quality knowledge base, the model can more accurately identify the most appropriate answer, and the improvement in retrieval accuracy is more significant compared to other large models. For multiple-choice questions that require reliance on multi-dimensional information for integration, more stringent requirements are placed on the comprehensive retrieval and analysis capabilities of the large model. From the experimental results, the accuracy of the PV-RAG model on multiple-choice questions is lower than that of single-choice and judgment questions, but it is still the highest compared to other large language models, showing good comprehensiveness.
[0168] Table 2 shows the comparison of question-answering accuracy and data redundancy obtained for the same subjective question on different large language models. From the experimental results, the PV-RAG and Wenxin Yiyan large models perform significantly better than other models in terms of average accuracy. This shows that RAG technology effectively enhances the retrieval ability, enabling it to generate more accurate and relevant answers in subjective question-answering tasks. At the same time, the average redundancy of PV-RAG is 0, indicating that its generated answers are more refined and avoid unnecessary information redundancy, which is a major advantage over other models. In contrast, Zhipu Qingyan's accuracy drops significantly without RAG, and its redundancy is high, indicating that in the absence of external knowledge retrieval enhancement, the model is prone to generate content that is irrelevant or repetitive to the question. Other large models, such as iFlytek Spark and Tongyi Qianwen, perform in the middle in both aspects. Although they can answer questions well, they still have many information redundancy problems.
[0169] During the experiment, all four big data models except PV-RAG showed hallucinations in the question-answering process, that is, the model would arbitrarily generate wrong answers without relevant information support. If this hallucination phenomenon occurs in the medication consultation question-answering scenario, it will cause serious misleading and pose a certain threat to the medication safety of ordinary people.
[0170] From the experimental results, in the case of having a rich dataset, the data retrieval enhancement ability of PV-RAG has obvious search accuracy and superiority compared with other large models. On the one hand, because it retrieves and dynamically selects different documents or information fragments, it can generate relatively accurate answers. On the other hand, it provides a traceability function for the information source during the retrieval process. By viewing the retrieved document fragments, users or developers can clearly see the data basis behind the generated content, which enhances the transparency and interpretability of the model and can effectively avoid the generation of a large amount of incorrect information due to the hallucination of large models. This experimental result further verifies the applicability and advantages of RAG in subjective question-and-answer tasks, especially in the pharmaceutical consultation scenario that requires precise retrieval and high-quality generation, which has important application value.
[0171] The intelligent question-and-answer system implemented by the present invention can provide users with efficient and accurate traditional Chinese medicine pharmacovigilance question-and-answer services through the collaborative work of the above modules. The system can not only understand the natural language input of users, but also provide accurate answers through PV-RAG technology, significantly improving the accuracy and reliability of traditional Chinese medicine medication safety question-and-answer.
[0172] In summary, the present invention constructs a PV-RAG large language model for the field of pharmacovigilance, combines cross-domain retrieval enhanced generation technology to implement a safe medication intelligent question-and-answer system, and provides an effective solution for effectively improving the reliability of safe medication question-and-answer. Experimental results show that PV-RAG shows higher accuracy and knowledge coverage in dealing with drug contraindication problems compared with the basic LLM model. Especially with the support of the dynamic retrieval and query rewriting mechanism, the query efficiency and answer quality are improved. The system shows strong adaptability and can handle the problem of diverse corpora. Especially in the question-and-answer of medication safety queries involving multiple factors, it shows higher accuracy and lower redundancy compared with other large language models, providing a certain reference basis for the wide application of large language model technology in the field of pharmacovigilance.
[0173] Embodiment 2: This embodiment provides a non-transitory computer-readable storage medium, on which computer instructions are stored. The computer instructions cause the computer to execute a method for constructing a traditional Chinese medicine pharmacovigilance question-and-answer system based on a large language model. The method includes the following steps:
[0174] Collect and organize traditional Chinese medicine drug-related data to establish a traditional Chinese medicine knowledge base;
[0175] Construct a retrieval enhanced generation model PV-RAG for pharmacovigilance, and use the traditional Chinese medicine knowledge base as a training set to train and optimize the PV-RAG model; the PV-RAG model includes a retrieval module, a text extraction module, a query rewriting module, and a generation module;
[0176] Among them, the retrieval module adopts a hybrid strategy of sparse and dense retrieval, combines the BM25 and DPR models, extracts keywords according to the user's query content, and retrieves document fragments related to the keywords from the traditional Chinese medicine knowledge base;
[0177] The text extraction module uses the MRI-BERT architecture to re-rank and judge the relevance of the retrieved document fragments, and extracts the text related to the user's query content;
[0178] The query rewriting module expands and rewrites the query content that fails the relevance judgment based on the large language model, and re-enters the expanded and rewritten query content into the retrieval module;
[0179] The generation module uses the large language model to generate answers to traditional Chinese medicine pharmacovigilance questions related to the user's query content based on the text extracted by the text extraction module;
[0180] Build an intelligent Q&A application terminal, connect it to the traditional Chinese medicine knowledge base and the trained PV-RAG model, and generate answers to traditional Chinese medicine pharmacovigilance according to the natural language questions input by the user.
[0181] Embodiment 3: This embodiment provides an electronic device, which may include: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The processor can call the logical instructions in the memory to execute the method for building a traditional Chinese medicine pharmacovigilance Q&A system based on the large language model. The method includes the following steps:
[0182] Collect and organize data related to traditional Chinese medicine, and establish a traditional Chinese medicine knowledge base;
[0183] Build a retrieval-enhanced generation model PV-RAG for pharmacovigilance, and use the traditional Chinese medicine knowledge base as a training set to train and optimize the PV-RAG model; this PV-RAG model includes a retrieval module, a text extraction module, a query rewriting module, and a generation module;
[0184] Among them, the retrieval module adopts a hybrid strategy of sparse and dense retrieval, combines the BM25 and DPR models, extracts keywords according to the user's query content, and retrieves document fragments related to the keywords from the traditional Chinese medicine knowledge base;
[0185] The text extraction module uses the MRI-BERT architecture to re-rank and judge the relevance of the retrieved document fragments, and extracts the text related to the user's query content;
[0186] The query rewriting module expands and rewrites the query content that fails the relevance judgment based on a large language model, and re-enters the expanded and rewritten query content into the retrieval module;
[0187] The generation module uses a large language model to generate answers to traditional Chinese medicine pharmacovigilance questions related to the user's query content based on the text extracted by the text extraction module;
[0188] Build an intelligent Q&A application terminal, connect it to the traditional Chinese medicine knowledge base and the trained PV-RAG model, and generate answers to traditional Chinese medicine pharmacovigilance according to the natural language questions input by the user.
[0189] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0190] Example 4: The computer program product provided in this example includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for constructing a traditional Chinese medicine pharmacovigilance Q&A system based on a large language model. The method includes the following steps:
[0191] Collect and organize data related to traditional Chinese medicine, and establish a traditional Chinese medicine knowledge base;
[0192] Construct a retrieval augmented generation model PV-RAG for pharmacovigilance, and use the traditional Chinese medicine knowledge base as a training set to train and optimize the PV-RAG model; this PV-RAG model includes a retrieval module, a text extraction module, a query rewriting module, and a generation module;
[0193] Among them, the retrieval module adopts a hybrid strategy of sparse and dense retrieval, combines the BM25 and DPR models, extracts keywords according to the user's query content, and retrieves document fragments related to the keywords from the traditional Chinese medicine knowledge base;
[0194] The text extraction module uses the MRI-BERT architecture to reorder and judge the relevance of the retrieved document fragments, and extracts the text related to the user's query content;
[0195] The query rewriting module expands and rewrites the query content that fails the relevance judgment based on the large language model, and re-enters the expanded and rewritten query content into the retrieval module;
[0196] The generation module uses the large language model to generate answers to traditional Chinese medicine drug vigilance questions and answers related to the user's query content based on the text extracted by the text extraction module;
[0197] Build an intelligent Q&A application end, connect it to the traditional Chinese medicine knowledge base and the trained PV-RAG model, and generate answers to traditional Chinese medicine drug vigilance according to the natural language questions input by the user.
[0198] The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0199] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to make a computer device (which can be a personal computer, server, or network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0200] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A construction method of a traditional Chinese medicine drug vigilance Q&A system based on a large language model, characterized in that, Including: Collect and organize data related to traditional Chinese medicine drugs, and establish a traditional Chinese medicine knowledge base; Construct a retrieval enhanced generation model PV-RAG for pharmacovigilance, and train and optimize the PV-RAG model with the traditional Chinese medicine knowledge base as the training set; the PV-RAG model includes a retrieval module, a text extraction module, a query rewriting module, and a generation module; Among them, the retrieval module adopts a hybrid strategy of sparse and dense retrieval, combines the BM25 and DPR models, extracts keywords according to the user's query content, and retrieves document fragments related to the keywords from the traditional Chinese medicine knowledge base; The text extraction module uses the MRI-BERT architecture to re-rank and judge the relevance of the retrieved document fragments, and extracts text related to the user's query content; The query rewriting module expands and rewrites the query content that fails the relevance judgment based on a large language model, and re-enters the expanded and rewritten query content into the retrieval module; The generation module uses a large language model to generate answers to traditional Chinese medicine pharmacovigilance questions related to the user's query content based on the text extracted by the text extraction module; Construct an intelligent Q&A application end, connect to the traditional Chinese medicine knowledge base and the trained PV-RAG model, and generate answers to traditional Chinese medicine pharmacovigilance according to the natural language questions input by the user.
2. The construction method of the traditional Chinese medicine pharmacovigilance Q&A system based on the large language model according to claim 1, characterized in that, The retrieval module of the PV-RAG model is constructed as: Use the BM25 model to match the extracted keywords with the traditional Chinese medicine knowledge base in terms of term frequency and inverse document frequency, and preliminarily screen out relevant document fragments; Use the DPR model to semantically encode the query content and the document fragments through the query encoder and the passage encoder, calculate the BERTScore relevance score, and supplement the semantically relevant document fragments not retrieved by the BM25 model; Merge the retrieval results of the BM25 model and the DPR model to form a diverse set of candidate document fragments.
3. The construction method of the traditional Chinese medicine pharmacovigilance Q&A system based on the large language model according to claim 2, wherein, The text extraction module of the PV-RAG model is constructed as: Use the MRI-BERT architecture to re-rank and judge the relevance of the candidate document fragments, and screen out the text highly relevant to the user's query content by calculating the similarity score between the high-dimensional embedding vectors of the query content and the candidate document fragments; Input the query content that fails the relevance judgment into the query rewriting module of the PV-RAG model.
4. The construction method of the traditional Chinese medicine pharmacovigilance Q&A system based on the large language model according to claim 1, characterized in that The query rewriting module of the PV-RAG model is constructed as: Based on a large language model, set prompt words for synonym expansion, expand and rewrite the query content that fails the relevance judgment, and generate synonymous query content related to the original query content; Re-enter the rewritten synonymous query content into the retrieval module of the PV-RAG model for a new round of retrieval.
5. The construction method of the traditional Chinese medicine pharmacovigilance Q&A system based on the large language model according to claim 1, characterized in that The generation module of the PV-RAG model is constructed as: For the query content that passes the relevance judgment, by setting prompt words for generating answers, use a large language model to generate answers to traditional Chinese medicine pharmacovigilance questions based on the extracted relevant text.
6. The construction method of the traditional Chinese medicine pharmacovigilance Q&A system based on the large language model according to claim 1, characterized in that, Collect and organize data related to traditional Chinese medicine drugs, and establish a traditional Chinese medicine knowledge base. The specific method is: Collect and organize data on traditional Chinese medicine instructions, traditional Chinese medicine classics, and clinical research literature; The collected and organized data are converted into a document collection and word segmentation is performed as a traditional Chinese medicine knowledge base; The traditional Chinese medicine knowledge base is divided into training set, validation set and test set for training and validation of the PV-RAG model.
7. The construction method of the traditional Chinese medicine pharmacovigilance Q&A system based on the large language model according to claim 6, characterized in that, A dynamic update mechanism is adopted to incorporate new TCM drug-related data into the TCM knowledge base, and the PV-RAG model is regularly trained using the updated TCM knowledge base.
8. A non-transitory computer-readable storage medium, characterized in that, Computer instructions are stored thereon, and the computer instructions enable the computer to execute the method for constructing a traditional Chinese medicine pharmacovigilance question-and-answer system based on a large language model as described in any one of claims 1 to 7.
9. An electronic device, characterized in that, include: A processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus, and the processor calls the logic instructions in the memory to execute the method for constructing a traditional Chinese medicine drug vigilance question and answer system based on a large language model as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a computer program, which is stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer executes the method for constructing a traditional Chinese medicine drug vigilance question and answer system based on a large language model as described in any one of claims 1 to 7.
Citation Information
Cited By
Multi-role configuration and effective judgment method based on semantic arbitration
CN120809299A
Chemical text attribute extraction method and system based on inverse reinforcement learning
CN121212393A
A Chemical Text Attribute Extraction Method and System Based on Inverse Reinforcement Learning
CN121212393B
Temporomandibular joint disease diagnosis method and system based on large language model
CN121354862A
Intelligent question and answer method and system based on traditional Chinese medicine classics meta-term engine
CN121434337A