Online medical question and answer dynamic retrieval enhancement generation method
By training the language model in Chinese medical dialogue data set and adopting dynamic retrieval enhancement generation method, the problem of large language models generating content bias and 'illusion' phenomena in medical question-and-answer system is solved, achieving more accurate and coherent content generation, and enhancing the model's responsiveness and adaptability.
Patent Information
- Application Number
- CN202510097306.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-22
AI Technical Summary
Large language models may produce ‘aliagnosis’ in medical question-and-answer systems, the generated content does not match actual medical knowledge, and in long conversations or multiple rounds of question-and-answer questions are likely to produce questions that the answer content deviates from the context.
By using the Chinese medical dialogue dataset to train the language model, a dynamic search enhancement generation method is adopted, including inputting user questions multiple times and generating candidate answers, establishing a subword generation probability dictionary, performing clustering and medical entity recognition, calculating the semantic entropy of the generated text, and determining whether to trigger the search to generate the final answer based on the semantic entropy value.
It effectively reduces the deviation and 'illusion' of content generated by large language models in medical Q&A systems, improves the accuracy and coherence of content generated, and enhances the model's responsiveness and adaptability on complex medical problems.
Smart Images

Figure BDA0005253203250000056 
Figure BDA0005253203250000062 
Figure BDA0005253203250000066
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large language models, and specifically relates to an online medical Q&A dynamic retrieval enhanced generation method. Background Art
[0002] With the rapid development of artificial intelligence technology, online medical Q&A systems have gradually become an important way for patients to obtain medical advice and health knowledge. However, when large language models generate answers to online medical consultations, they may encounter the "hallucination" problem, that is, the generated content does not conform to actual medical knowledge, and sometimes even provides information completely unrelated to the patient's question. Secondly, traditional Q&A systems are usually prone to the problem of the answer content deviating from the context in long conversations or multi-round Q&As, resulting in the generated content being less relevant or inaccurate. Summary of the Invention
[0003] In order to overcome the above deficiencies in technology, the present invention provides a method for reducing the deviation and "hallucination" phenomenon of the content generated by large language models in medical Q&A systems and improving the accuracy and coherence of the generated content. The technical solution adopted by the present invention to overcome its technical problems is as follows:
[0004] An online medical Q&A dynamic retrieval enhanced generation method, comprising:
[0005] a) Training a language model using a Chinese medical dialogue dataset to obtain a trained language model;
[0006] b) Inputting the user input question q into the trained language model N times, and each time the trained language model outputs a candidate answer, obtaining a candidate answer set A composed of N candidate answers num , A num = {A 1 , A 2 ,..., A i ,..., A N}, A i is the i-th candidate answer, i ∈ {1,..., N}, A i = {x 1 , x 2 ,..., x j ,..., x u}, x j is the j-th sub-word in the i-th candidate answer A i , j ∈ {1,..., u}, u is the number of sub-words in the i-th candidate answer A i ;
[0007] c) Establishing a sub-word generation probability dictionary DIC according to all sub-words of all N candidate answers;
[0008] d) Based on the i-th candidate answer A i obtain the clustering result C final , C final ={C 1 , C 2 ,..., C k ,..., C M}, where C k is the k-th clustering cluster, k ∈ {1,..., M}, and M is the number of clusters after clustering;
[0009] e) Obtain the representative medical element set according to the k-th clustering cluster C k
[0010] f) Calculate the generation probability P of the k-th clustering cluster C on the user input question q k ; k
[0011] g) Calculate the semantic entropy SE(q) of the generated text according to the generation probability P k ;
[0012] h) Trigger retrieval according to whether the semantic entropy SE(q) of the generated text exceeds the threshold, and generate the final answer according to the content of the retrieved knowledge base and the question q
[0013] Furthermore, in step a), the language model is the Qwen2.5-14B-Instruct model
[0014] Furthermore, step a) includes the following steps
[0015] a-1) Use the Chinese Medical Dialogue Dataset dataset as the fine-tuning dataset; a-2) Use LoRA to fine-tune the Qwen2.5-14B-Instruct model with the fine-tuning dataset to obtain the trained Qwen2.5-14B-Instruct model
[0016] Furthermore, the temperature parameter of the trained Qwen2.5-14B-Instruct model is set to T, 0.7 ≤ T ≤ 1.2
[0017] Furthermore, in step c), extract the logits corresponding to the j-th subword x i in the i-th candidate answer A j in the output layer of the Qwen2.5-14B-Instruct model, input the extracted logits into the softmax function, and output the generation probability p j of the j-th subword xj , the generation probabilities of all u sub-words in the i-th candidate answer A i constitute the generation probability set P i of the i-th candidate answer A i , P i = {p 1 , p 2 ,..., p j ,..., p u}, and the generation probability sets of all N candidate answers constitute the sub-word generation probability dictionary DIC, DIC = {P 1 , P 2 ,..., P i ,..., P N}.
[0018] Furthermore, step d) includes the following steps:
[0019] d-1) Input the i-th candidate answer A i into the Sentence-BERT model, and output the high-dimensional semantic vector V i , V i ∈ R d , where R is the real number space and d is the dimension of the high-dimensional semantic vector V i , and d takes the value of 768 or 1024;
[0020] d-2) Input the j-th candidate answer A j into the Sentence-BERT model, and output the high-dimensional semantic vector V j , V j ∈ R d , j ∈ {1,..., N};
[0021] d-3) Calculate the cosine similarity S i between the high-dimensional semantic vector V j and the high-dimensional semantic vector V i,j , and construct an N×N similarity matrix S from the cosine similarities between all high-dimensional semantic vectors. All the element values on the diagonal formed by the elements from the upper left corner to the lower right corner in the similarity matrix S are 1;
[0022] d-4) The cluster corresponding to the i-th candidate answer A i is C i , and the set of independent clusters corresponding to all N candidate answers is C, C = {C 1 , C 2 ,..., C i ,..., C N};
[0023] d-5) Find the maximum value S of all elements in the similarity matrix S except for all elements in the diagonal line formed by the elements from the upper left corner to the lower right corner w,r , maximum value S w,r is the wth candidate answer A w The high-dimensional semantic vector V w and the rth candidate answer A r The high-dimensional semantic vector V r The cosine similarity between , w∈{1,...,N}, r∈{1,...,N}, w≠r;
[0024] d-6) Through formula C wr =C w ∪C r ={A w ,A r} Calculate the wth candidate answer A w The corresponding cluster is C w and the rth candidate answer A r The corresponding cluster is C r The new cluster C after the merger wr , remove cluster C from the independent cluster set C w and Cluster C r And add a new cluster C wr , and obtain a new cluster set C′; d-7) by formula Calculate the new cluster C wr and the kth cluster C in the independent cluster set C k The similarity S wr,k , where S w,k is the wth candidate answer A w The high-dimensional semantic vector V w and the kth candidate answer A k The high-dimensional semantic vector V k The cosine similarity between, k∈{1,...,N}, k≠w, S r,k is the rth candidate answer A r The high-dimensional semantic vector V r and the kth candidate answer A k The high-dimensional semantic vector V k The cosine similarity between them, · represents the number of corresponding candidate answers in the cluster;
[0025] d-8) Establish a new similarity matrix S′ of (N-1)×(N-1), Where S wr,wr =1;
[0026] d-9) replacing the similarity matrix S in step c-5) with the new similarity matrix S′ and then repeating steps c-5) to c-8);
[0027] d - 10) Repeat step c - 9) O times until the minimum value among all elements except those on the diagonal formed by the elements from the upper - left corner to the lower - right corner in the newly obtained similarity matrix is greater than or equal to a threshold. This threshold is greater than or equal to 0.7 and less than or equal to 0.8. If so, this new similarity matrix is used as the final similarity matrix S of M×M. final , The final similarity matrix S final The k - th row or k - th column in represents the k - th cluster C k , k ∈ {1,..., M}, where A kh is the h - th candidate answer in the k - th cluster C k , h ∈ {1,..., C k}, C k is the number of candidate answers in the k - th cluster C k All M clusters constitute the clustering result C final , C final ={C 1 , C 2 ,..., C k ,..., C M}.
[0028] Furthermore, step e) includes the following steps:
[0029] e - 1) Use the CMeEE Chinese medical entity recognition dataset to train the MedBERT model to obtain the trained MedBERT model;
[0030] e - 2) Use the tokenizer of the trained MedBERT model to split the h - th candidate answer A k in the k - th cluster C kh into a sub - word sequence and convert each split sub - word into a unique identifier to obtain the identifier sequence A kh ′, A kh ′={x 1 ′, x 2 ′,..., x i ′,..., x t ′}, where x i ′ is the identifier of the i - th sub - word, i ∈ {1,..., t}, and t is the number of sub - words in the identifier sequence A kh ′;
[0031] e - 3) For the category labels dis, sym, ite, pro, dru in the CMeEE Chinese medical entity recognition dataset, establish the identifier sequence Akh The set of dis categories of ′ The set of sym categories The set of ite categories The set of pro categories The set of dru categories Input the identifier sequence A kh The identifier x of the i-th subword in ′ i ′ into the trained MedBERT model, and the probabilities that the i-th subword is the category label dis, category label sym, category label ite, category label pro, and category label dru in the CMeEE Chinese medical entity recognition dataset are output. Put this subword into the category set corresponding to the maximum probability value. After all t subwords are put into the corresponding category sets respectively, the h-th candidate answer A kh The set of medical elements E kh ,
[0032] e-4) Merge the same category sets in the sets of medical elements of all C k in the k-th clustering cluster C k candidate answers to obtain the classified set E k ,
[0033] is the i-th subword in the set , i ∈ {1,..., z 1}, z 1 is the number of subwords in the set ;
[0034] is the i-th subword in the set , i ∈ {1,..., z 2}, z 2 is the number of subwords in the set ;
[0035] is the i-th subword in the set , i ∈ {1,..., z 3}, z 3 is the number of subwords in the set ;
[0036] is the set the i-th subword in, where i ∈ {1,..., z 4}, z 4 is the number of subwords in the set ;
[0037] is the i-th subword in the set where i ∈ {1,..., z 5}, z 5 is the number of subwords in the set ;
[0038] e-5) For all z subwords in the set 1 , look up the generation probability corresponding to each subword in the subword generation probability dictionary DIC, and select the subword corresponding to the maximum generation probability among all z 1 generation probabilities as the core element of the set ; e-6) For all z subwords in the set 2 , look up the generation probability corresponding to each subword in the subword generation probability dictionary DIC, and select the subword corresponding to the maximum generation probability among all z 2 generation probabilities as the core element of the set ; e-7) For all z subwords in the set 3 , look up the generation probability corresponding to each subword in the subword generation probability dictionary DIC, and select the subword corresponding to the maximum generation probability among all z 3 generation probabilities as the core element of the set ; e-8) For all z subwords in the set 4 , look up the generation probability corresponding to each subword in the subword generation probability dictionary DIC, and select the subword corresponding to the maximum generation probability among all z 4 generation probabilities as the core element of the set ; e-9) For all z subwords in the set 5 , look up the generation probability corresponding to each subword in the subword generation probability dictionary DIC, and select the subword corresponding to the maximum generation probability among all z 5 generation probabilities as the core element of the set ; e-10) Establish the k-th clustering cluster C kSet of representative medical elements
[0039] Furthermore, step f) includes the following steps:
[0040] f-1) Sum the generation probabilities corresponding to the core elements the generation probabilities corresponding to the core elements the generation probabilities corresponding to the core elements the generation probabilities corresponding to the core elements the generation probabilities corresponding to the core elements to obtain the generation probability f-2) Calculate the generation probability P of the k-th clustering cluster C on the user input question q through the formula k k .
[0041] Furthermore, in step g), the semantic entropy SE(q) of the generated text is calculated through the formula
[0042] Furthermore, step h) includes the following steps:
[0043] h-1) Set the entropy threshold SE threshold , 0.5 ≤ SE threshold ≤ 0.8;
[0044] h-2) If the semantic entropy SE(q) of the generated text is less than the entropy threshold SE threshold then randomly select a candidate answer from the candidate answer set A num as the final answer to the user input question q;
[0045] h-3) If the semantic entropy SE(q) of the generated text is greater than or equal to the entropy threshold SE threshold then obtain a consultation question knowledge base D composed of n medical documents through the Elasticsearch search engine, D = {d 1 , d 2 ,..., d i ,..., d n}, d i is the i-th medical document, i ∈ {1,..., n}, retrieve l medical documents from the consultation question knowledge base D according to the user input question q using the Elasticsearch search engine to form a retrieval result set D′, splice the retrieval result set D′ with the user input question q and then input it into the trained Qwen2.5-14B-Instruct model, and output the answer A q as the final answer to the user input question q.
[0046] The beneficial effects of the present invention are as follows: It is applicable to a dynamic retrieval augmented generation system, especially in the field of medical consultation. This method has significant advantages in improving the accuracy of the model, optimizing the retrieval timing, and enhancing the quality of the generated text. By introducing semantic entropy and a dynamic threshold setting mechanism based on semantic entropy, the system determines whether to trigger retrieval according to the semantic entropy value of the generated content. This avoids excessive retrieval operations brought by the traditional static threshold method, improves the pertinence of retrieval, enhances the professionalism and accuracy of the answers, and reduces the possibility of the model generating "hallucinations" or irrelevant answers. By combining a medical domain-specific model (Med-BERT) for medical entity extraction, the present invention can accurately identify and classify medical elements in the generated text. Based on these medical elements, a more precise judgment basis can be provided for triggering the system's retrieval mechanism. It can effectively improve the response ability of the system to complex medical problems and enhance the adaptability of the model. Specific embodiments
[0047] The present invention will be further described below.
[0048] An online medical Q&A dynamic retrieval augmented generation method includes:
[0049] a) Training a language model using a Chinese medical dialogue dataset to obtain a trained language model.
[0050] b) Inputting the user's input question q into the trained language model N times, and each time the trained language model outputs a candidate answer, obtaining a candidate answer set A composed of N candidate answers num , A num = {A 1 , A 2 ,..., A i ,..., A N}, A i is the i-th candidate answer, i ∈ {1,..., N}, A i = {x 1 , x 2 , …, x j , …, x u}, x j is the j-th sub-word in the i-th candidate answer A i , j ∈ {1,…, u}, u is the number of sub-words in the i-th candidate answer A i .
[0051] c) Establishing a sub-word generation probability dictionary DIC according to all sub-words of all N candidate answers.
[0052] d) Obtaining a clustering result C based on the i-th candidate answer A i final , C final = {C 1 , C 2 , …, C k , …, C M}, where C k is the k-th clustering cluster, k ∈ {1, ..., M}, and M is the number of clusters after clustering.
[0053] e) Obtain the representative medical element set according to the k-th clustering cluster C k
[0054] f) Calculate the generation probability P of the k-th clustering cluster C k on the user input question q. k .
[0055] g) Calculate the semantic entropy SE(q) of the generated text according to the generation probability P k .
[0056] h) Trigger retrieval according to whether the semantic entropy SE(q) of the generated text exceeds the threshold, and generate the final answer according to the content of the retrieved knowledge base and the question q.
[0057] In an embodiment of the present invention, the language model in step a) is the Qwen2.5-14B-Instruct model. In an embodiment of the present invention, step a) includes the following steps:
[0058] a-1) Use the Chinese Medical Dialogue Dataset as the fine-tuning dataset.
[0059] a-2) Use the LoRA (Low-Rank Adaptation) to fine-tune the Qwen2.5-14B-Instruct model with the fine-tuning dataset to obtain the trained Qwen2.5-14B-Instruct model. In an embodiment of the present invention, the temperature parameter of the trained Qwen2.5-14B-Instruct model is set to T, 0.7 ≤ T ≤ 1.2.
[0060] During the generation process of candidate answers, the generation of candidate answers is gradually completed according to the sub-word sequence. Therefore, in an embodiment of the present invention, the j-th sub-word x i in the i-th candidate answer A jThe corresponding logits are input into the softmax function, and the output is the generation probability p of the j-th subword x. j of j , the generation probability set P of the i-th candidate answer A i is composed of the generation probabilities of all u subwords in i the i-th candidate answer A. i , P i = {p 1 , p 2 ,..., p j ,..., p u}, and the generation probability sets of all N candidate answers form the subword generation probability dictionary DIC, DIC = {P 1 , P 2 ,..., P i ,..., P N}.
[0061] In one embodiment of the present invention, step d) includes the following steps:
[0062] d-1) Input the i-th candidate answer A i into the Sentence-BERT model (SBERT), and the output is a high-dimensional semantic vector V i , V i ∈R d , where R is the real number space, d is the dimension of the high-dimensional semantic vector V i , and d takes a value of 768 or 1024, which depends on the version of the selected Sentence-BERT model. d-2) Input the j-th candidate answer A j into the Sentence-BERT model, and the output is a high-dimensional semantic vector V j , V j ∈R d , j ∈ {1,..., N}.
[0063] d-3) Calculate the cosine similarity S i between the high-dimensional semantic vector V j and the high-dimensional semantic vector V i,j , and construct an N×N similarity matrix S from the cosine similarities between all high-dimensional semantic vectors. All the element values in the diagonal formed by the elements from the upper left corner to the lower right corner in the similarity matrix S are 1 (i.e., the similarity of a candidate answer to itself is 1).
[0064] d-4) In order to classify the generated multiple candidate answers according to semantic similarity to form clustering clusters, a hierarchical clustering method is adopted. Specifically, the i-th candidate answer A iThe corresponding cluster is C i , the set of all independent clusters corresponding to all N candidate answers is C, C = {C 1 , C 2 ,..., C i ,..., C N}}.
[0065] d-5) Find the maximum value S of all elements in the similarity matrix S except for all elements on the diagonal formed by the elements from the upper left corner to the lower right corner w,r , the maximum value S w,r is the cosine similarity between the high-dimensional semantic vector V w of the w-th candidate answer A w and the high-dimensional semantic vector V r of the r-th candidate answer A r , w ∈ {1,..., N}, r ∈ {1,..., N}, w ≠ r.
[0066] d-6) Calculate the new cluster C wr = C w ∪ C r = {A w , A r} by the formula to calculate the new cluster C w corresponding to the w-th candidate answer A w and the new cluster C r corresponding to the r-th candidate answer A r after merging, remove the cluster C wr and the cluster C w from the set of independent clusters C and add the new cluster C r to obtain a new set of clusters C'. wr
[0067] d-7) Calculate the similarity S between the new cluster C wr and the k-th cluster C k in the set of independent clusters C by the formula wr,k , where S w,k is the cosine similarity between the high-dimensional semantic vector V w of the w-th candidate answer A w and the high-dimensional semantic vector V k of the k-th candidate answer A k , k ∈ {1,..., N}, k ≠ w, S r,k is the cosine similarity between the high-dimensional semantic vector V r of the r-th candidate answer A r and the high-dimensional semantic vector V k of the k-th candidate answer A k , |·| represents the number of corresponding candidate answers within the cluster.
[0068] d - 8) Establish a new similarity matrix S′ of size (N - 1)×(N - 1), where S wr,wr = 1. Since the clusters C w and C r in the similarity matrix S are replaced by a new cluster C wr , the dimension of the new similarity matrix S′ is reduced by 1.
[0069] d - 9) After replacing the similarity matrix S in step c - 5) with the new similarity matrix S′, repeat steps c - 5) to c - 8).
[0070] d - 10) Repeat step c - 9) O times until the minimum value among all elements of the new similarity matrix except those on the diagonal from the upper - left element to the lower - right element is greater than or equal to a threshold. This threshold is greater than or equal to 0.7 and less than or equal to 0.8. If so, this new similarity matrix is used as the final similarity matrix S of size M×M final , The k - th row or k - th column in the final similarity matrix S final represents the k - th clustering cluster C k , k ∈ {1,..., M}, where A kh is the h - th candidate answer in the k - th clustering cluster C k , h ∈ {1,..., |C k |}, |C k | is the number of candidate answers in the k - th clustering cluster C k . All M clustering clusters form the clustering result C final , C final = {C 1 , C 2 ,..., C k ,..., C M}.
[0071] In one embodiment of the present invention, step e) includes the following steps:
[0072] e - 1) Use the CMeEE Chinese medical entity recognition dataset to train the MedBERT model to obtain a trained MedBERT model.
[0073] e - 2) Use the tokenizer of the trained MedBERT model to split the h - th candidate answer A k in the k - th clustering cluster C kh into a sub - word sequence and convert each split sub - word into a unique identifier to obtain an identifier sequence A kh ′, Akh ′ = {x 1 ′, x 2 ′,..., x i ′,..., x t ′}, where x i ′ is the identifier of the i-th subword, i ∈ {1,..., t}, and t is the number of subwords in the identifier sequence A kh ′.
[0074] e-3) For the category labels dis (disease name), sym (symptom), ite (examination item), pro (treatment plan), and dru (drug recommendation) in the CMeEE Chinese medical entity recognition dataset, establish the dis category set, kh sym category set, ite category set, pro category set, dru category set of the identifier sequence A Input the identifier x kh ′ of the i-th subword in the identifier sequence A i ′ into the trained MedBERT model, and output the probabilities that the i-th subword is the category labels dis, sym, ite, pro, and dru in the CMeEE Chinese medical entity recognition dataset respectively. Put this subword into the category set corresponding to the maximum probability value. After all t subwords are put into the corresponding category sets, the medical element set E kh of the h-th candidate answer A kh ,
[0075] e-4) Merge the same category sets in the medical element sets of all C k candidate answers in the k-th clustering cluster C k to obtain the classified set E k ,
[0076] is the i-th subword in the set , i ∈ {1,..., z 1}, and z 1 is the number of subwords in the set ;
[0077] is the set The i-th subword in, where i ∈ {1,..., z 2}, z 2 is the number of subwords in the set ;
[0078] is the i-th subword in the set where i ∈ {1,..., z 3}, z 3 is the number of subwords in the set ;
[0079] is the i-th subword in the set where i ∈ {1,..., z 4}, z 4 is the number of subwords in the set ;
[0080] is the i-th subword in the set where i ∈ {1,..., z 5}, z 5 is the number of subwords in the set .
[0081] e-5) For all z subwords in the set, look up the generation probability corresponding to each subword in the subword generation probability dictionary DIC, and select the subword corresponding to the maximum generation probability among all z 1 generation probabilities as the core element of the set 1 ; e-6) For all z subwords in the set, look up the generation probability corresponding to each subword in the subword generation probability dictionary DIC, and select the subword corresponding to the maximum generation probability among all z 2 generation probabilities as the core element of the set 2 ; e-7) For all z subwords in the set, look up the generation probability corresponding to each subword in the subword generation probability dictionary DIC, and select the subword corresponding to the maximum generation probability among all z 3 generation probabilities as the core element of the set 3 ; e-8) For all z subwords in the set,4 The sub-words respectively look up the generation probability corresponding to each sub-word in the sub-word generation probability dictionary DIC, and select all z 4 Among the generation probabilities, select the sub-word corresponding to the maximum generation probability as the core element of the set e-9) Use the set All z in 5 The sub-words respectively look up the generation probability corresponding to each sub-word in the sub-word generation probability dictionary DIC, and select all z 5 Among the generation probabilities, select the sub-word corresponding to the maximum generation probability as the core element of the set e-10) Establish the representative medical element set of the k-th clustering cluster C k
[0082] In an embodiment of the present invention, step f) includes the following steps:
[0083] f-1) Sum the generation probabilities corresponding to the core elements The generation probabilities corresponding to the core elements The generation probabilities corresponding to the core elements The generation probabilities corresponding to the core elements The generation probabilities corresponding to the core elements The generation probabilities corresponding to the core elements to obtain the generation probability f-2) Calculate the generation probability P of the k-th clustering cluster C on the user input question q through the formula k k
[0084] In an embodiment of the present invention, in step g), the semantic entropy SE(q) of the generated text is calculated through the formula
[0085] In an embodiment of the present invention, step h) includes the following steps:
[0086] h-1) Set the entropy threshold SE threshold , 0.5 ≤ SE threshold ≤ 0.8.
[0087] h-2) If the semantic entropy SE(q) of the generated text is less than the entropy threshold SE threshold It means that the generation result of the model is relatively consistent, and the model has a high confidence in its answer. Then, randomly select a candidate answer from the candidate answer set A num as the final answer to the user input question q.
[0088] h-3) If the semantic entropy SE(q) of the generated text is greater than or equal to the entropy threshold SE threshold It is considered that the generation result of the model is scattered and the confidence is low. It may be necessary to retrieve external documents to obtain more accurate information. After triggering the retrieval, medical information related to the user's question is obtained from the medical knowledge base. At this time, an interrogation question knowledge base D composed of n medical literatures is obtained through the Elasticsearch search engine. D = {d 1 , d 2 ,..., d i ,..., d n}, where d i is the i-th medical literature, i ∈ {1,..., n}. According to the user input question q, l medical literatures are retrieved from the interrogation question knowledge base D through the Elasticsearch search engine to form a retrieval result set D'. After concatenating the retrieval result set D' with the user input question q, it is input into the trained Qwen2.5-14B-Instruct model, and the answer A q is obtained as the final answer to the user input question q.
[0089] As shown in Table 1, in order to verify the effectiveness of the online medical Q&A dynamic retrieval enhanced generation method of the present invention, we conducted experiments with different threshold SE threshold values, focusing on the balance between generation quality and retrieval efficiency. The experiment used the Qwen2.5-14B-Instruct model to fine-tune the Chinese medical dialogue dataset and evaluated BLEU, ROUGE, and retrieval ratio (Search Ratio) under different SE threshold values.
[0090] In the experiment, we set multiple SE threshold values and conducted detailed tests for each value. The experimental results show that when SE threshold is lower, the generated answers perform better in terms of BLEU and ROUGE scores, and the retrieval ratio is also higher. As SE threshold value increases, the quality of the generated answers and the retrieval efficiency show a downward trend.
[0091] Table 1 Different threshold SE threshold Experiment
[0092] <![CDATA[SE threshold > BLEU ROUGE(Recall) Search Ratio 0.50 0.48 0.62 16.9 0.55 0.49 0.64 15.8 0.60 0.51 0.66 14.2 0.65 0.52 0.68 13.0 0.70 0.53 0.70 12.3 0.75 0.55 0.73 11.5 0.80 0.54 0.72 10.8 0.85 0.52 0.69 9.5
[0093] In the experiment, when SE threshold = 0.75, the BLEU and ROUGE scores reached 0.55 and 0.73 respectively, and the retrieval ratio was 11.5. Compared with other SE threshold values, it showed higher generation quality and lower retrieval requirements. As SEthreshold As the value increases, although the generation quality slightly decreases, the retrieval ratio also decreases; while when the SE threshold value is low, although the generation quality improves, the retrieval ratio increases significantly, and the computational burden increases.
[0094] Therefore, by selecting SE threshold = 0.75, the generation quality and retrieval efficiency can be effectively balanced, and an accurate and efficient medical Q&A service can be provided in practical applications.
[0095] As shown in Table II, the comparison of the traditional RAG method and the method of the present invention in terms of generation quality uses common automatic evaluation metrics (such as BLEU, ROUGE, METEOR, etc.) to measure their generation performance, and the results are analyzed.
[0096] Table II Comparison of the traditional RAG method and the method of the present invention in terms of generation quality
[0097] Performance metric The method of the present invention Traditional RAG method BLEU 0.58 0.45 ROUGE(Recall) 0.74 0.60 METEOR 0.66 0.53
[0098] Compared with the traditional RAG method, the method of the present invention performs excellently in multiple performance metrics, especially in the balance between generation quality and retrieval efficiency. Specifically, the method of the present invention is significantly better than the traditional RAG method in evaluation metrics such as BLEU, ROUGE, and METEOR, with improvements of 0.13, 0.14, and 0.13 respectively, indicating that the system can generate higher-quality and more accurate answers.
[0099] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for dynamic retrieval enhancement generation of online medical questions and answers, characterized in that: include: a) Use the Chinese medical dialogue dataset to train the language model to obtain a trained language model; b) Input the user input question q into the trained language model N times, and each time the trained language model outputs a candidate answer, a candidate answer set A consisting of N candidate answers is obtained. num , A num ={A1,A2,...,A i ,...,A N }, A i is the i-th candidate answer, i∈{1,...,N}, A i ={x1,x2,…,x j ,…,x u }, x j is the i-th candidate answer A i The jth subword in, j∈{1,…,u}, u is the i-th candidate answer A i The number of neutron words; c) establishing a subword generation probability dictionary DIC based on all subwords of all N candidate answers; d) Based on the i-th candidate answer A i Get the clustering result C final , C final ={C1,C2,…,C k ,…,C M }, where C k is the kth cluster, k∈{1,...,M}, M is the number of clusters after clustering; e) According to the kth cluster C k Get a representative set of medical elements f) Based on a set of representative medical elements Calculate the kth cluster C k The generation probability P on the user input question q k ; g) According to the generation probability P k Calculate the semantic entropy SE(q) of the generated text; h) Trigger retrieval based on whether the semantic entropy SE(q) of the generated text exceeds a threshold, and generate the final answer based on the content of the retrieved knowledge base and question q.
2. The method for dynamic retrieval enhancement generation of online medical questions and answers according to claim 1, characterized in that: The language model in step a) is the Qwen2.5-14B-Instruct model.
3. The method for dynamic retrieval enhancement generation of online medical questions and answers according to claim 2, characterized in that: Step a) comprises the following steps: a-1) Use the Chinese Medical Dialogue Dataset as the fine-tuning dataset; a-2) Use LoRA to fine-tune the Qwen2.5-14B-Instruct model using the fine-tuning dataset to obtain the trained Qwen2.5-14B-Instruct model.
4. The method for dynamic retrieval enhancement and generation of online medical questions and answers according to claim 3, characterized in that: The temperature parameter of the trained Qwen2.5-14B-Instruct model is set to T, 0.7≤T≤1.
2.
5. The method for dynamic retrieval enhancement and generation of online medical questions and answers according to claim 2, characterized in that: In step c), the i-th candidate answer A is extracted from the output layer of the Qwen2.5-14B-Instruct model. i The jth subword x in j Corresponding logits, the extracted logits are input into the softmax function, and the output is the j-th subword x j The generation probability p j , the i-th candidate answer A i The generation probability of all u subwords in constitutes the i-th candidate answer A i The generation probability set P i , P i ={p1,p2,...,p j ,...,p u }, the generation probability set of all N candidate answers constitutes the subword generation probability dictionary DIC, DIC = {P1, P2, ..., P i ,...,P N }.
6. The method for dynamic retrieval enhancement generation of online medical questions and answers according to claim 5, characterized in that: Step d) comprises the following steps: d-1) Take the i-th candidate answer A i Input into the Sentence-BERT model and output a high-dimensional semantic vector V i , V i ∈R d , where R is the real number space and d is the high-dimensional semantic vector V i The dimension of d is 768 or 1024; d-2) Take the jth candidate answer A j Input into the Sentence-BERT model and output a high-dimensional semantic vector V j , V j ∈R d , j∈{1,...,N}; d-3) Calculate the high-dimensional semantic vector V i With the high-dimensional semantic vector V j The cosine similarity S between i,j , construct an N×N similarity matrix S based on the cosine similarity between all high-dimensional semantic vectors, In the similarity matrix S, all element values in the diagonal line from the upper left element to the lower right element are 1; d-4) The i-th candidate answer A i The corresponding cluster is C i , the independent cluster set corresponding to all N candidate answers is C, C = {C1, C2, ..., C i ,...,C N }; d-5) Find the maximum value S of all elements in the similarity matrix S except for all elements in the diagonal line formed by the elements from the upper left corner to the lower right corner w,r , maximum value S w,r is the wth candidate answer A w The high-dimensional semantic vector V w and the rth candidate answer A r The high-dimensional semantic vector V r The cosine similarity between , w∈{1,...,N}, r∈{1,...,N}, w≠r; d-6) Through formula C wr =C w ∪C r ={A w ,A r } Calculate the wth candidate answer A w The corresponding cluster is C w and the rth candidate answer A r The corresponding cluster is C r The new cluster C after the merger wr , remove cluster C from the independent cluster set C w and Cluster C r And add a new cluster C wr , get a new cluster set C′; d-7) Through the formula Calculate the new cluster C wr and the kth cluster C in the independent cluster set C k The similarity S wr,k , where S w,k is the wth candidate answer A w The high-dimensional semantic vector V w and the kth candidate answer A k The high-dimensional semantic vector V k The cosine similarity between, k∈{1,...,N}, k≠w, S r,k is the rth candidate answer A r The high-dimensional semantic vector V r and the kth candidate answer A k The high-dimensional semantic vector V k The cosine similarity between them, |·| represents the number of corresponding candidate answers in the cluster; d-8) Establish a new similarity matrix S′ of (N-1)×(N-1), Where S wr,wr =1; d-9) replacing the similarity matrix S in step c-5) with the new similarity matrix S′ and then repeating steps c-5) to c-8); d-10) Repeat step c-9) O times until the new similarity matrix is obtained. The minimum value of all elements except the elements in the diagonal formed by the elements from the upper left corner to the lower right corner is greater than or equal to the threshold, and the threshold is greater than or equal to 0.7 and less than or equal to 0.
8. If so, the new similarity matrix is used as the final similarity matrix S of M×M. final , The final similarity matrix S final The kth row or kth column in represents the kth cluster C k , k∈{1,...,M}, Among them A kh is the kth cluster C k The hth candidate answer in h∈{1,...,|C k |},|C k | is the kth cluster C k The number of candidate answers in the cluster, all M clusters form the clustering result C final , C final ={C1,C2,...,C k ,...,C M }。 7. The method for dynamic retrieval enhancement generation of online medical questions and answers according to claim 6, characterized in that: Step e) comprises the following steps: e-1) Use the CMeEE Chinese medical entity recognition dataset to train the MedBERT model to obtain the trained MedBERT model; e-2) Use the trained MedBERT model’s tokenizer to cluster the kth cluster C k The hth candidate answer A kh Split into subword sequences and convert each of the split subwords into a unique identifier to obtain an identifier sequence A kh ′,A kh ′={x1′,x2′,...,x i ′,...,x t ′}, where x i ′ is the identifier of the i-th subword, i∈{1,...,t}, t is the identifier sequence A kh The number of subwords in ′; e-3) Establish identifier sequences A for the category labels dis, sym, ite, pro, and dru in the CMeEE Chinese medical entity recognition dataset. kh ′’s dis category set sym category collection ite category collection pro category collection dru category collection The identifier sequence A kh The identifier x of the i-th subword in ′ i ′ is input into the trained MedBERT model, and the output is the probability that the i-th subword is the category label dis, category label sym, category label ite, category label pro, and category label dru in the CMeEE Chinese medical entity recognition dataset. The subword is placed in the category set corresponding to the category label with the maximum probability value. After all t subwords are placed in the corresponding category sets, the h-th candidate answer A is obtained. kh Medical element set E kh , e-4) Cluster the kth cluster C k All in |C k The same category sets in the medical element sets of | candidate answers are merged to obtain the classified set E k , For collection The i-th subword in, i∈{1,...,z1}, z1 is the set The number of neutron words; For collection The i-th subword in, i∈{1,...,z2}, z2 is the set The number of neutron words; For collection The i-th subword in, i∈{1,...,z3}, z3 is the set The number of neutron words; For collection The i-th subword in, i∈{1,...,z4}, z4 is the set The number of neutron words; For collection The i-th subword in, i∈{1,...,z5}, z5 is the set The number of neutron words; e-5) will gather All z1 subwords in the subword generation probability dictionary DIC are searched for the generation probability corresponding to each subword, and the subword corresponding to the maximum generation probability among all z1 generation probabilities is selected as the set The core elements e-6) will gather All z2 subwords in the subword generation probability dictionary DIC are searched for the generation probability corresponding to each subword, and the subword corresponding to the maximum generation probability among all z2 generation probabilities is selected as the set The core elements e-7) will gather All z3 subwords in the subword generation probability dictionary DIC are searched for the generation probability corresponding to each subword, and the subword corresponding to the maximum generation probability among all z3 generation probabilities is selected as the set The core elements e-8) will gather All z4 subwords in the subword generation probability dictionary DIC are searched for the generation probability corresponding to each subword, and the subword corresponding to the maximum generation probability among all z4 generation probabilities is selected as the set The core elements e-9) will gather All z5 subwords in the subword generation probability dictionary DIC are searched for the generation probability corresponding to each subword, and the subword corresponding to the maximum generation probability among all z5 generation probabilities is selected as the set The core elements e-10) Establish the kth cluster C k Representative medical elements collection 8. The method for dynamic retrieval enhancement generation of online medical questions and answers according to claim 7, characterized in that: Step f) comprises the following steps: f-1) Core elements Corresponding generation probability and core elements Corresponding generation probability and core elements Corresponding generation probability and core elements Corresponding generation probability and core elements The corresponding generation probabilities are summed to obtain the generation probability f-2) Through the formula Calculate the kth cluster C k The generation probability P on the user input question q k .
9. The method for dynamic retrieval enhancement and generation of online medical questions and answers according to claim 1, characterized in that: In step g), the formula The semantic entropy SE(q) of the generated text is calculated.
10. The method for dynamic retrieval enhancement generation of online medical questions and answers according to claim 3, characterized in that: Step h) comprises the following steps: h-1) Set the entropy threshold SE threshold , 0.5≤SE threshold ≤0.8; h-2) If the semantic entropy SE(q) of the generated text is less than the entropy threshold SE threshold Then from the candidate answer set A num Randomly select a candidate answer as the final answer to the user input question q; h-3) If the semantic entropy SE(q) of the generated text is greater than or equal to the entropy threshold SE threshold Then, the Elasticsearch search engine is used to obtain a medical question knowledge base D consisting of n medical documents, where D = {d1, d2, ..., d i ,...,d n }, d i For the i-th medical article, i∈{1,...,n}, according to the user input question q, the Elasticsearch search engine is used to retrieve l medical articles in the medical question knowledge base D to form a retrieval result set D′. The retrieval result set D′ is concatenated with the user input question q and input into the trained Qwen2.5-14B-Instruct model, and the output is the answer A q As the final answer to the user input question q.
Citation Information
Patent Citations
Text generation method for abstract extraction based on LDA and D2V
CN111159393A
Border examination legal question and answer dynamic retrieval enhancement generation method and system
CN118520093A
Summary generation apparatus, control method, and system
US20220114202A1
Cited By
Multi-round reasoning question answering method and system based on query graph driving
CN120687579A
Question and answer processing method, system and equipment based on three-dimensional entropy evaluation and medium
CN121009182A
A question and answer processing method, system, device and medium based on three-dimensional entropy evaluation
CN121009182B
Big language model illusion detection method and device and medium
CN121412716A
A hallucination detection method, device and medium of a large language model
CN121412716B