RAG medical question answering method and system based on question rewriting and generation reverification
By introducing a RAG method of problem rewriting and generation and reverification in the medical question-and-answer system, the semantic difference between user spoken description and medical knowledge base is solved, and more accurate retrieval and diagnostic results are achieved, which significantly improves the reliability of the system.
Patent Information
- Application Number
- CN202510133469.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-13
AI Technical Summary
The existing medical Q&A system has semantic differences in processing user colloquial descriptions, resulting in unrelated or errors in the search results, and the generation module lacks an effective error correction mechanism, resulting in the diagnostic results that may be completely wrong.
The RAG medical question-and-answer method based on problem rewriting and generation reverification is adopted. The user's colloquial complaint text is converted into professional disease descriptions through the rewrite model, and the diagnostic results are retrieved and generated in the vector database. Finally, the accuracy of the diagnostic results is ensured through the secondary verification mechanism.
Effectively bridge the semantic gap between user spoken description and medical knowledge base, improve the accuracy of retrieval and the accuracy of diagnostic results, significantly reduce the rate of misdiagnosis, and enhance the overall reliability and security of the system.
Smart Images

Figure CN119993561A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of large-model medical intelligent question answering, and in particular to a RAG medical question answering method and system based on question re- and generation re-verification. Background Art
[0002] In recent years, large language models (LLMs) such as the GPT series have shown great potential in solving medical question-answering tasks. However, relying solely on the inherent knowledge obtained through pre-training within the model is often difficult to handle question-answering tasks in the medical field that require high accuracy and complexity. Although the traditional retrieval-augmented generation (RAG) method can supplement external knowledge, the retrieval results are biased due to the high inconsistency between user questions and the knowledge contained in the external knowledge base in the semantic space, which leads to incorrect diagnosis results. Moreover, when such errors occur, there is a lack of effective discrimination methods, which may ultimately lead to incorrect diagnosis results for patients.
[0003] There are currently two approaches to building medical question-answering systems: one based on fine-tuning of a large language model and the other based on retrieval-augmented generation (RAG).
[0004] The general knowledge learned during the pre-training of large language models has limitations when dealing with specific medical question-answering tasks. In order to improve its professionalism and accuracy in the medical field, model fine-tuning has become a common method. Model fine-tuning further trains the model on large-scale medical text datasets to adapt it to the terminology, context, and knowledge system in the medical field. This process can optimize the model's performance in medical question-answering tasks, enabling it to better understand and answer questions related to disease diagnosis, treatment plans, drug use, etc.
[0005] In the medical question-answering system, the RAG method builds an external medical knowledge base, such as a medical literature database, a clinical case library, etc., and uses a search engine to find information related to the question from the knowledge base. The generator then combines the retrieved external knowledge with the capabilities of the large language model to generate answers. The RAG method can introduce the latest medical knowledge in real time, making up for the lack of timeliness of the model's pre-trained knowledge.
[0006] In the current technical field, the Retrieval Augmented Generation (RAG) system is used as a method to expand and supplement the knowledge base of large language models. The RAG system provides dynamic supplementation capabilities for large language models by combining external real-time factual information with retrieval models, enabling them to better cope with the above problems. This technology is widely used in scenarios such as question-answering systems, medical diagnosis, legal consultation, and technical support that require instant access to authoritative information.
[0007] The core structure of the RAG system consists of two key modules, the retriever and the generator.
[0008] The main responsibility of the retriever is to find information related to the user input requirements from pre-built data stores. These data stores usually include structured and unstructured data, such as vector databases, document collections, or online resource libraries. The retriever relies on specific search algorithms and indexing mechanisms, such as traditional retrieval methods based on keyword matching, or vectorized retrieval techniques based on semantic similarity. In recent years, with the rapid development of deep learning, many advanced embedding models (such as BERT, Sentence-BERT, and GPT embedding models) have been used to generate high-dimensional semantic vectors for data blocks and queries, which significantly improves the performance of the retriever in complex semantic matching tasks. In addition, the performance of the retriever can be further improved by optimizing vector retrieval techniques (such as inner product, Euclidean distance, or cosine similarity calculation) and efficient index construction (such as inverted index, FAISS and other tools).
[0009] The generator module works on the basis of the retriever. It uses the retrieved information and the user's input to generate the final content. The generator usually takes the pre-trained large language model as the core and generates output content that meets the context requirements based on the retrieval results. The generator not only needs to integrate the explicit knowledge in the retrieval results, but also needs to use the implicit knowledge within the large language model to expand and supplement the information. In the medical field, the generator can combine the patient's input complaints with the retrieved medical literature to generate possible diagnosis results and treatment recommendations.
[0010] Deficiencies of existing technical solutions:
[0011] 1. In the prior art, user input is usually directly used as the input of the retriever. However, in the medical question-answering system, the user's description often has obvious colloquial characteristics and may lack professional medical knowledge. For example, a user may use "stomach pain" to describe stomach discomfort, while the more common terms in medical literature or knowledge bases may be "upper abdominal pain" or "peptic ulcer". This leads to a large semantic difference between the user input and the retrieval knowledge base. The knowledge base itself mostly comes from professional medical books, research papers, and medical guidelines, which contain a large number of complex professional terms and medical entities. This semantic gap directly leads to the inability of the retriever to accurately match the user's real needs, resulting in irrelevant or erroneous retrieval results. When these erroneous retrieval results are used in the generation stage, they will further cause downstream errors and reduce the overall reliability of the system.
[0012] 2. The generation module of the prior art is highly dependent on the output content of the retriever, that is, the generation results are mainly based on the retrieved citation information. When the content provided by the retriever is inconsistent with the user's actual question requirements or contains errors, the generation model lacks an effective error correction mechanism, resulting in the generated answers being completely wrong. For example, if the retriever mistakenly associates a description of abdominal discomfort with citations of unrelated diseases, the generation model will generate misleading diagnostic recommendations based on this erroneous information. This characteristic of error chain transmission significantly limits the applicability of existing RAG systems in medical scenarios, especially in medical question-and-answer systems, where erroneous information can have a serious impact on user decisions. Summary of the invention
[0013] In view of the deficiencies of the prior art, the present invention proposes a RAG medical question-answering method based on question rewriting and generation re-verification, the medical question-answering method comprising collecting professional medical data, performing block processing on the collected medical data, performing vectorization processing on the block-based medical data, constructing a chief complaint rewriting data set, training a rewriting model based on the constructed chief complaint rewriting data set, a retriever using the output of the rewriting model as input to perform retrieval and output K citation results, inputting the citation results and the patient's professional symptom description into a generator for analysis to generate a first disease diagnosis result, inputting the first disease diagnosis result into a RAG system for secondary verification, retrieving a more similar second disease diagnosis result and calculating a similarity matching score, and determining whether to output a diagnosis result according to the similarity matching score, specifically comprising:
[0014] Step 1: Collect professional medical data. The collected data is organized into a text set D = {d1, d2, ...d n};
[0015] Step 2: The collected professional medical data is divided into blocks. The block strategy includes logical division based on semantic content and fixed-length segmentation. Each block data is represented by C i =f(d i , L), and obtain the block set C = {C1, C2, ...C n};
[0016] Step 3: Use embedding technology to vectorize the medical data after segmentation, and transform each segment data C i Mapped to a high-dimensional vector E(C i )∈R d , and stored in the vector database, and the matching degree of different data blocks is evaluated through similarity calculation to establish a vector index;
[0017] Step 4: Construct a chief complaint rewriting database for subsequent training of the chief complaint rewriting model. The construction process includes collecting the spoken chief complaint texts of real patients U = {u1, u2, ...ui}, and rewrite the patient's colloquial complaint text into a standardized professional description of the disease through the designed prompt template and large model. i , get the main complaint rewriting data set T = {(u i , v i )};
[0018] Step 5: Based on the main complaint rewriting dataset constructed in step 4, a large language model with at least 7B parameters is fine-tuned to obtain a rewriting model;
[0019] Step 6: Input the patient’s chief complaint text into the trained rewriting model to generate the corresponding professional disease description v i ;
[0020] Step 7: Use the rewritten professional disease description i As the input of the retriever, the retrieval enhancement generation RAG method is used to perform a search operation in the vector database generated in step 3, and the inner product similarity calculation is used to find the vectors that match the professional disease description v. i The first k citations with the highest matching degree constitute the first citation set C k ;
[0021] Step 8: The first citation set C retrieved k The patient's main complaint text is input into the generator, which generates a potential first disease diagnosis result by comprehensively analyzing the contextual semantics of the input;
[0022] Step 9: Perform secondary verification on the first disease diagnosis result, input the first disease diagnosis result into the RAG system, retrieve the first k most relevant citations to form the second citation set C k ’ , and calculate the patient's main complaint text and the second citation set C k ’ When the similarity matching score S is greater than or equal to 0.6, the diagnosis result is output; if the similarity matching score S is less than 0.6, return to step 7 and re-perform citation retrieval and diagnosis generation.
[0023] The RAG medical question-answering system based on question rewriting and generation re-verification includes: a medical data processing module, a chief complaint rewriting data set construction module, a chief complaint text entry module, a question rewriting module, a retriever module, a generator module, a secondary verification module and a result output module, specifically including:
[0024] The medical data processing module is used to collect professional medical data, process the data in blocks, and perform vector processing to obtain a vector database;
[0025] The chief complaint rewriting dataset construction module is used to collect the spoken chief complaint texts of real patients, and rewrite the patients' spoken chief complaint texts into standardized professional symptom descriptions through the designed prompt template and large model, which are used for training the rewriting module;
[0026] The chief complaint text input module is used for patients to input their own symptoms and obtain the patient's chief complaint text;
[0027] The question rewriting module is used to convert the patient's chief complaint text into a standardized professional disease description;
[0028] A retriever module, used for taking the professional disease description as input, searching in the vector database and outputting a first citation result;
[0029] A generator module, used for inputting the first quotation set and the patient's chief complaint text into the generator for analysis to generate a first disease diagnosis result;
[0030] A secondary verification module is used to input the first disease diagnosis result into the RAG system to retrieve the second citation set, and calculate the similarity matching score between the second citation set and the patient's chief complaint. If the similarity matching score is greater than a preset threshold, the first disease diagnosis result is output to the result output module. If it is lower than the preset threshold, it is returned to the retriever module for regeneration;
[0031] The result output module is used to output the final disease diagnosis result.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] 1. Improve the accuracy of retrieval. By introducing advanced precise semantic matching and improved input processing technology, it can effectively bridge the semantic gap between user colloquial descriptions and the medical knowledge base. The system can automatically identify and convert non-professional terms in user input, such as accurately mapping "stomach pain" to "upper abdominal pain" or "peptic ulcer", thereby improving the match between the search engine and the knowledge base, ensuring that the retrieved results are more relevant and accurate.
[0034] 2. Improve the accuracy and credibility of the diagnosis results. By matching the diagnosis results finally generated by the model with the relevant disease information in the vector knowledge base and calculating the similarity between the generated results and the patient's complaints, the present invention can effectively verify the accuracy of the diagnosis. If the similarity score is higher than the preset threshold, it indicates that the generated diagnosis results are consistent with the professional standards in the medical knowledge base, thereby enhancing the reliability of the diagnosis results.
[0035] 3. Significantly reduce the misdiagnosis rate. The present invention can automatically verify the diagnosis results after they are generated through secondary verification and automatic error correction mechanism. When the similarity score is lower than the preset threshold, the system will automatically identify potential errors and trigger the regeneration process, avoiding the misleading of erroneous diagnosis results, thereby improving the overall robustness and security of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a processing flow chart of the medical question-answering method of the present invention;
[0037] Figure 2 It is a processing flow chart of the secondary verification of the present invention;
[0038] Figure 3 It is a structural schematic diagram of the medical question-answering system of the present invention. DETAILED DESCRIPTION
[0039] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessary confusion of the concept of the present invention.
[0040] The present invention uses a large language model to train a large amount of medical data to realize the process of converting the patient's colloquial chief complaint text into a standardized medical description. The process first obtains the patient's chief complaint information through text input. This information is usually unstructured, colloquial descriptions, and lacks professional terms or standardized expressions in the medical field, such as "I have had frequent headaches recently, especially at night." These input texts are sent to the trained rewriting model. The model recognizes key information such as symptoms, medical history background, and course of disease through a deep understanding of language, and normalizes it.
[0041] During the training process, the rewriting model learns how to map colloquial descriptions to standardized medical terms through large-scale medical case records, doctors' medical records, and standardized expressions in medical literature. Using deep learning technology, especially models based on the Transformer architecture, the model can capture the diverse expressions of different diseases and generate accurate medical terminology descriptions based on the context. The generated output text usually includes key information such as the symptoms, time of occurrence, and intensity of the disease. The format is more in line with doctors' recording habits and is not prone to ambiguity.
[0042] Through pre-training of large-scale medical texts, the model has strong reasoning capabilities, can handle complex medical terms and disease descriptions, and accurately rewrite them according to the input context. This automated standardized text generation converts colloquial language into professional medical terms, and maintains a high degree of consistency in the semantic space between the text input to the retrieval enhancement generation RAG system and the text in the vector database. It can effectively improve the accuracy of later retrieval.
[0043] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings. Figure 1 It is a processing flow chart of the method of the present invention.
[0044] The RAG medical question-and-answer method proposed in the present invention mainly includes: collecting professional medical data, dividing the collected medical data into blocks, vectorizing the divided medical data, constructing a chief complaint rewriting data set, training a rewriting model based on the constructed chief complaint rewriting data set, a retriever using the output of the rewriting model as input to retrieve and output K citation results, inputting the citation results and the standardized description of the patient into a generator for analysis to generate a first disease diagnosis result, inputting the first disease diagnosis result into a RAG system to retrieve a more similar second disease diagnosis result and calculating a similarity matching score, and deciding whether to output the diagnosis result based on the similarity matching score.
[0045] Step 1: Collect professional medical data. The collected data is organized into a text set D = {d1, d2, ...d n The collected medical data include classic medical books (such as Harrison's Internal Medicine), global medical guidelines (such as the diagnosis and treatment standards issued by the World Health Organization), cutting-edge medical papers (such as research results in the PubMed database), and anonymized real case data, as well as medical case records, doctors' medical records, and standard statements in medical literature. In order to provide a comprehensive and high-quality knowledge foundation for model training and reasoning, this method first collects large-scale medical data from authoritative sources. The diversity and authority of the data ensure the depth and breadth of knowledge, enabling the model to cover a wide range of scenarios from basic theory to clinical practice, laying the foundation for subsequent processing.
[0046] Step 2: The collected professional medical data is divided into blocks. The block strategy includes logical division based on semantic content and fixed-length segmentation. Each block is represented by C i =f(d i , L), where L is the maximum length of a single block, d i Represents the i-th document, and obtains the block set C = {C1, C2, ...C n};
[0047] The collected professional medical data may be too complex in its original form to be directly used for model processing. Therefore, the present invention performs a block operation on the data. Logical division based on semantic content, such as segmentation by chapter or theme, and fixed-length segmentation, such as 512 tokens (words). This block division method not only ensures the semantic coherence of the data, but also facilitates subsequent vectorization and retrieval processing.
[0048] Step 3: Use embedding technology to vectorize the medical data after segmentation, and transform each segment C i Mapped to a high-dimensional vector E(C i )∈R d , and stored in the vector database, and the matching degree of different data blocks is evaluated through similarity calculation to establish a vector index;
[0049] In order to achieve efficient knowledge retrieval, the block set is vectorized. Through the pre-trained embedding model, each block C i Mapped to a high-dimensional vector E(C i )∈R d These embedded vectors are stored in the vector database to support subsequent fast retrieval. The similarity calculation adopts the inner product method, through the formula To evaluate the matching degree between different data blocks. The establishment of this vector index provides technical support for the rapid search of relevant medical knowledge.
[0050] Step 4: Construct a chief complaint rewriting database for subsequent training of the chief complaint rewriting model. The construction process includes collecting the spoken chief complaint texts of real patients U = {u1, u2, ...u i}, and rewrite the patient's spoken complaint text into a standardized professional description of the disease through the designed prompt template and the use of the big model. i , get the main complaint rewriting data set T = {(u i , v i )};
[0051] For the colloquial complaints input by users, the present invention designs a complaint rewriting data set to train the subsequent rewriting model. The data set construction process includes. These conversion results T = {(u i , v i )} not only reflects the language characteristics in real scenarios, but also enhances the model's ability to transform natural language into professional medical terms, providing high-quality training samples for subsequent model fine-tuning.
[0052] Step 5: Based on the main complaint rewriting dataset constructed in step 4, a large language model with at least 7B parameters is fine-tuned to obtain a rewriting model;
[0053] The fine-tuning goal is to minimize the cross entropy loss The model is designed to enhance the model's ability to convert patients' colloquial descriptions into professional descriptions. The fine-tuned model is named "Rewrite Model" and has the function of standardizing the language of complaints in medical scenarios. This model significantly improves the efficiency and accuracy of subsequent knowledge retrieval and diagnosis generation.
[0054] Step 6: Input the patient’s chief complaint text into the trained rewriting model to generate the corresponding professional disease description v i ; Professional disease descriptions usually include key information such as disease symptoms, time of occurrence, intensity, etc. The format is more in line with doctors' recording habits and is less likely to cause ambiguity.
[0055] This conversion process is completed through a trained rewriting model. The output professional description is more in line with medical expression standards and provides a unified input format for subsequent processing.
[0056] Step 7: Use the rewritten professional disease description i As the input of the retriever, the retrieval-augmented generation (RAG) method is used to perform a search operation in the vector database generated in step 3. By calculating the inner product similarity, find the vectors that match the professional disease description v. i The first k citations with the highest matching degree constitute the first citation set C k ;
[0057] These citations cover the medical knowledge most relevant to the user's complaint, providing authoritative and diverse supporting information for subsequent diagnosis generation.
[0058] Step 8: The first citation set C retrieved k The patient's main complaint text is input into the generator, which generates a potential first disease diagnosis result D by comprehensively analyzing the contextual semantics of the input. diag ;
[0059] The generated diagnosis results are not only based on the user's description, but also make full use of information from authoritative medical citations, and are highly scientific and accurate.
[0060] Step 9: Perform secondary verification on the first disease diagnosis result (such as Figure 2 As shown in FIG. 1 ), the first disease diagnosis result is input into the RAG system, and the first k most relevant citations are retrieved to form a second citation set C k ’ , and calculate the patient's main complaint text and the second citation set C k ’ When the similarity matching score S is greater than or equal to 0.6, the diagnosis result is output.k ’ is the result of searching with disease results, C k is the result of searching for citations using the main complaint. If the similarity match score S is less than 0.6, return to step 7 and perform citation search and diagnosis generation again. Figure 2 The decision model in is the similarity matching process here.
[0061] To ensure the reliability of the diagnosis results, the present invention further verifies the generated diagnosis results. The mathematical expression of the similarity score is:
[0062]
[0063] The secondary verification mechanism proposed in the present invention matches the diagnosis result finally generated by the generator with the relevant disease information in the vector database, retrieves several citations, and calculates the similarity between the citations and the patient's chief complaint text. If the similarity score is higher than the preset threshold, the verification passes, indicating that the generated disease diagnosis result is credible; if the similarity score is lower than the threshold, the verification fails, which means that the generated diagnosis result may be inaccurate or wrong and needs to be regenerated.
[0064] like Figure 2 As shown in Figure 1, the process of the secondary verification mechanism is as follows: First, the generated diagnosis result is input into a vector knowledge base for matching. This step retrieves several citation information. Then, the system further matches the citation with the patient's chief complaint for similarity. Through this matching, the system can evaluate whether the generated diagnosis result is highly consistent with the patient's description. If it is consistent, the result is output, otherwise the result needs to be regenerated.
[0065] Unlike traditional retrieval methods, which usually retrieve related diseases based on symptoms, the secondary verification mechanism in this invention uses the generated diagnosis results to reversely search for symptom information and further match and verify it with the patient's symptoms. This process ensures whether the results retrieved in the first retrieval can accurately reflect the patient's actual condition and avoids misdiagnosis caused by model generation errors or incomplete knowledge base content.
[0066] Through this mechanism, secondary verification not only improves the accuracy of the model's diagnostic results, but also provides an effective means of verifying the results of the first search. Through similarity matching, the system can check whether the disease returned by the first search actually matches the patient's symptoms, thereby playing a role in error correction and optimization in the entire diagnostic process, ensuring the high reliability and accuracy of medical decisions.
[0067] The present invention also proposes a RAG medical question-answering system based on question rewriting and generation re-verification, such as Figure 3The medical question-answering system includes: a medical data processing module, a chief complaint rewriting data set construction module, a chief complaint text entry module, a question rewriting module, a retriever module, a generator module, a secondary verification module and a result output module, specifically:
[0068] The medical data processing module is used to collect professional medical data, process the data in blocks, and perform vector processing to obtain a vector database;
[0069] The chief complaint rewriting dataset construction module is used to collect the spoken chief complaint texts of real patients, and rewrite the patients' spoken chief complaint texts into standardized professional symptom descriptions through the designed prompt template and large model, which are used for training the rewriting module;
[0070] The chief complaint text input module is used for patients to input their own symptoms and obtain the patient's chief complaint text;
[0071] The question rewriting module is used to convert the patient's chief complaint text into a standardized professional disease description;
[0072] A retriever module, used for taking the professional disease description as input, searching in the vector database and outputting a first citation result;
[0073] A generator module, used for inputting the first quotation set and the patient's chief complaint text into the generator for analysis to generate a first disease diagnosis result;
[0074] A secondary verification module is used to input the first disease diagnosis result into the RAG system to retrieve the second citation set, and calculate the similarity matching score between the second citation set and the patient's chief complaint. If the similarity matching score is greater than a preset threshold, the first disease diagnosis result is output to the result output module. If it is lower than the preset threshold, it is returned to the retriever module for regeneration;
[0075] The result output module is used to output the final disease diagnosis result.
[0076] Combined with the patient's chief complaint rewriting and secondary verification mechanism, the present invention can significantly improve the accuracy and reliability of the medical diagnosis system. First, through the patient's chief complaint rewriting mechanism, the patient's colloquial, unstructured chief complaint text is converted into a standardized medical description, making the input more in line with medical expression standards and reducing the understanding deviation caused by language differences or non-standard expressions. This process provides a unified and accurate input format for subsequent medical data processing, knowledge retrieval and diagnosis generation, laying the foundation for the normal operation of the system.
[0077] On this basis, the secondary verification mechanism further improves the accuracy of the system. By matching the generated diagnosis results with the relevant disease information in the medical knowledge base, the system can verify the consistency between the diagnosis results and the patient's complaints, and determine whether the generated results meet the standard descriptions in the medical literature. If the similarity is lower than the set threshold, the system will automatically trigger the re-retrieval and diagnosis generation process to ensure that each diagnosis result is fully verified and avoid incorrect diagnosis output. This reverse verification mechanism not only ensures that the generated diagnosis results are more accurate, but also detects and corrects possible deviations in the first retrieval process, enhancing the reliability of the system.
[0078] Through the combination of patient complaint rewriting and secondary verification mechanism, the present invention has significant advantages in improving the accuracy and intelligence of medical question-answering systems and diagnostic tools. It effectively eliminates the semantic gap between the patient's colloquial expression and the medical knowledge base, and ensures the accuracy and authority of the final diagnosis results through multi-level verification, thereby providing clinicians with more reliable decision support and reducing the risk of misdiagnosis and missed diagnosis.
[0079] It should be noted that the above specific embodiments are exemplary, and those skilled in the art can come up with various solutions inspired by the disclosure of the present invention, and these solutions also belong to the disclosure scope of the present invention and fall within the protection scope of the present invention. Those skilled in the art should understand that the present invention description and its drawings are illustrative and do not constitute a limitation of the claims. The protection scope of the present invention is defined by the claims and their equivalents.
Claims
1. A RAG medical question-answering method based on question rewriting and generation revalidation, characterized in that: The medical question-answering method includes collecting professional medical data, processing the collected medical data in blocks, vectorizing the medical data after the blocks, constructing a chief complaint rewriting data set, training a rewriting model based on the constructed chief complaint rewriting data set, a retriever using the output of the rewriting model as input to retrieve and output K citation results to form a first citation result, and inputting the text of the patient's chief complaint into a generator for analysis to generate a first disease diagnosis result, inputting the first disease diagnosis result into a RAG system for secondary verification to generate a second citation set, calculating a similarity matching score between the second citation set and the patient's chief complaint, and determining whether to output a diagnosis result according to the similarity matching score, specifically including: Step 1: Collect professional medical data. The collected data is organized into a text set D = {d1, d2, ...d n }; Step 2: The collected professional medical data is divided into blocks. The block strategy includes logical division based on semantic content and fixed-length segmentation. Each block data is represented by C i =f(d i , L), and obtain the block set C = {C1, C2, ...C n }; Step 3: Use embedding technology to vectorize the medical data after segmentation, and transform each segment data C i Mapped to a high-dimensional vector E(C i )∈R d , and stored in the vector database, and the matching degree of different data blocks is evaluated through similarity calculation to establish a vector index; Step 4: Construct a chief complaint rewriting database for subsequent training of the chief complaint rewriting model. The construction process includes collecting the spoken chief complaint texts of real patients U = {u1, u2, ...u i }, and rewrite the patient's colloquial complaint text into a standardized professional description of the disease through the designed prompt template and large model. i , get the main complaint rewriting data set T = {(u i , v i )}; Step 5: Based on the main complaint rewriting dataset constructed in step 4, a large language model with at least 7B parameters is fine-tuned to obtain a rewriting model; Step 6: Input the patient’s chief complaint text into the trained rewriting model to generate the corresponding professional disease description v i ; Step 7: Use the rewritten professional disease description i As the input of the retriever, the retrieval enhancement generation RAG method is used to perform a search operation in the vector database generated in step 3, and the inner product similarity calculation is used to find the vectors that match the professional disease description v. i The first k citations with the highest matching degree constitute the first citation set C k ; Step 8: The first citation set C retrieved k The patient's main complaint text is input into the generator, which generates a potential first disease diagnosis result by comprehensively analyzing the contextual semantics of the input; Step 9: Perform secondary verification on the first disease diagnosis result, input the first disease diagnosis result into the RAG system, retrieve the first K most relevant citations to form the second citation set C k ’ , and calculate the patient's main complaint text and the second citation set C k ’ When the similarity matching score S is greater than or equal to 0.6, the diagnosis result is output; if the similarity matching score S is less than 0.6, return to step 7 and re-perform citation retrieval and diagnosis generation.
2. RAG medical question-answering system based on question rewriting and generation re-verification, characterized by: The question-answering system includes: a medical data processing module, a chief complaint rewriting data set construction module, a chief complaint text entry module, a question rewriting module, a retriever module, a generator module, a secondary verification module and a result output module, specifically including: The medical data processing module is used to collect professional medical data, process the data in blocks, and perform vector processing to obtain a vector database; The chief complaint rewriting dataset construction module is used to collect the spoken chief complaint texts of real patients, and rewrite the patients' spoken chief complaint texts into standardized professional symptom descriptions through the designed prompt template and large model, which are used for training the rewriting module; The chief complaint text input module is used for patients to input their own symptoms and obtain the patient's chief complaint text; The question rewriting module is used to convert the patient's chief complaint text into a standardized professional disease description; A retriever module, for taking the professional disease description as input, searching and outputting a first citation set in the vector database; A generator module, used for inputting the first quotation set and the patient's chief complaint text into the generator for analysis to generate a first disease diagnosis result; A secondary verification module is used to input the first disease diagnosis result into the RAG system to retrieve the second citation set, and calculate the similarity matching score between the second citation set and the patient's chief complaint. If the similarity matching score is greater than a preset threshold, the first disease diagnosis result is output to the result output module. If it is lower than the preset threshold, it is returned to the retriever module for regeneration; The result output module is used to output the final disease diagnosis result.
Citation Information
Cited By
Medical large model answer enhancement method, system, equipment and medium
CN120541192A
Rheumatism immunology department intelligent question and answer method based on multi-model cooperation
CN121880517A