Method for constructing lightweight medical Q&A system
Through lightweight Transformer model and specific training strategies, the calculation complexity and accuracy of medical Q&A systems on low-computing devices are solved, and fast response and high accuracy are achieved, and suitable for VR and other devices.
Patent Information
- Application Number
- CN202510662894.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The existing medical Q&A system has limitations in terms of high computational complexity and large resource requirements, and it is difficult to effectively distinguish answers with similar semantics but different contexts, resulting in difficulty in deploying and insufficient accuracy on low-computing equipment.
Using a lightweight Transformer model, the model parameters are optimized to improve semantic distinction capabilities by inserting pooling layers into the model and simplifying the number of editor layers and attention heads, combining pre-stored answer library and specific negative sample generation strategies.
While achieving rapid response on low-computing devices, it improves semantic matching accuracy, controls the parameter quantity below 15MB, increases the inference speed by 3.5 times, and improves the accuracy to 0.7-0.8, and is adapted to embedded devices such as VR.
Smart Images

Figure CN120181141B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer medical health, and particularly relates to a method for constructing a lightweight medical question-answering system. Background Art
[0002] With the increasing maturity of artificial intelligence technology and its wide penetration in the medical field, medical question-answering systems have become a key technology for improving health education levels, optimizing remote diagnosis and treatment processes, and enabling virtual reality (VR) medical scenarios. Such systems can provide users with convenient and efficient medical information query services, and have important clinical and social values. However, current mainstream medical question-answering models often rely on deep learning frameworks with a large number of parameters, such as BERT (the number of parameters usually exceeds 100MB). Although such models perform well in semantic understanding, their inherent high computational complexity and huge resource requirements severely restrict their deployment and application on embedded devices with limited computing power, such as VR devices and mobile medical terminals.
[0003] In the field of medical health, the accuracy of question-answering systems is crucial. Inaccurate or irrelevant answers may lead users to have incorrect health cognitions and even pose potential health risks. Therefore, for medical question-answering systems, they not only need to have accurate semantic matching capabilities, but also must ensure fast response speeds to meet the needs of users in actual application scenarios.
[0004] In the prior art, research on model lightweighting mainly focuses on methods such as model pruning and knowledge distillation. These techniques can reduce the resource occupancy of the model to a certain extent, but often sacrifice the semantic representation ability of the model, resulting in a significant decline in performance when the model understands complex medical problems and matches relevant answers, and unable to ensure a high degree of relevance between the answer and the question context. In addition, traditional model training strategies usually use random negative samples for optimization. This method lacks pertinence and is difficult to effectively train the model to distinguish answers that are literally similar but have subtle differences in context semantics, thus limiting the application efficiency and accuracy of the system in specific medical fields. Summary of the Invention
[0005] In view of the above situation, the main objective of the present invention is to propose a method and system for constructing a lightweight medical question-answering system to solve the above technical problems.
[0006] The present invention proposes a method for constructing a lightweight medical question-answering system, and the method includes the following steps:
[0007] Step 1: Based on the Transformer model, simplify the number of encoder layers and the number of attention heads of the Transformer model, and insert a pooling layer at the output end of the encoder to obtain a lightweight Transformer model. Use the lightweight Transformer model as the base model and construct a pre-stored answer library. Given a training set, the medical Q&A pairs include medical questions and corresponding unique target answers;
[0008] Step 2: Use the unique target answer corresponding to the medical question as the positive sample, select answers from the training set that have relevant keywords but different context semantics from the medical question as hard negative samples, and randomly select several answers from the training set as simple negative samples;
[0009] Step 3: Input the training set into the base model to obtain question embedding vectors, positive sample embedding vectors, simple negative sample embedding vectors, and hard negative sample embedding vectors, and store the positive sample embedding vectors, simple negative sample embedding vectors, and hard negative sample embedding vectors in the pre-stored answer library;
[0010] Step 4: Construct a positive sample loss function according to the similarity between the question embedding vector and the positive sample embedding vector and the simple negative sample embedding vector, construct a hard negative sample loss function according to the similarity between the question embedding vector and the hard negative sample embedding vector, and weight and combine the positive sample loss function and the hard negative sample loss function to construct a combined loss function;
[0011] Step 5: Input the combined loss function into the Adam optimizer, set the weight decay strategy and learning strategy of the Adam optimizer, and optimize the parameters of the base model with the goal of maximizing the similarity between the question embedding vector and the target answer and minimizing the similarity between the question embedding vector and the non-target answer. After optimization, the trained model is obtained;
[0012] Step 6: Set up a temporary retrieval post-processing module at the output part of the trained model to obtain a lightweight medical Q&A system, and use the lightweight medical Q&A system for medical Q&A.
[0013] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0014] 1. After receiving a question raised by a user, the traditional medical Q&A model needs to calculate the semantic embedding vectors of the question and candidate answers in real time, and the computational complexity of this process is relatively high, especially when dealing with a large number of candidate answers. To solve this problem, the present invention proposes a strategy of pre-computing the semantic embeddings of all candidate answers and storing them. In the inference stage, the system only needs to calculate the semantic embedding of the question raised by the user, and then quickly match it with the pre-computed answer embeddings. In this way, the inference time complexity is reduced from to , where is the number of answers, is the embedding dimension, thus significantly improving the response speed of the system.
[0015] 2. The models in the prior art often have limited ability to distinguish answers with similar semantics but different context meanings. This is particularly important in the medical field because there is a high degree of similarity between many medical terms and concepts. To solve this problem, the present invention designs a hard negative sample generation mechanism and a positive sample weighted loss function. By introducing more challenging negative samples during the training process and applying higher weights to the positive samples, the present invention can effectively improve the model's ability to distinguish subtle semantic differences in the embedding space through mathematical optimization means.
[0016] 3. Traditional lightweight methods often sacrifice model performance, resulting in a decrease in the accuracy of semantic matching. The present invention successfully optimizes the semantic matching similarity from the baseline level of 0.4398 to a potential 0.7 - 0.8 in a resource-constrained environment by carefully innovating the lightweight of the Transformer model structure and combining a post-processing strategy of temporary retrieval, thus significantly improving the accuracy of question answering while ensuring the lightweight of the model.
[0017] 4. The core advantage of the present invention is that the number of model parameters is strictly controlled below 15MB, and at the same time, a training speed of up to 31it / s is achieved, enabling it to be successfully adapted to low-computing-power embedded devices such as VR.
[0018] The additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the embodiments of the present invention. Description of the Drawings
[0019] Figure 1 is the flowchart of the method for constructing a lightweight medical question answering system proposed by the present invention;
[0020] Figure 2 is the overall architecture diagram of the lightweight medical question answering system;
[0021] Figure 3 is the flowchart of the basic model optimization. Detailed Embodiments
[0022] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.
[0023] These and other aspects of the embodiments of the present invention will be apparent from the following description and the accompanying drawings. In these descriptions and drawings, specific embodiments of some embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0024] Please refer to Figures 1 to 3 , this embodiment provides a method for constructing a lightweight medical Q&A system, and the method includes the following steps:
[0025] Step 1: Based on the Transformer model, simplify the number of encoder layers and the number of attention heads of the Transformer model, and insert a pooling layer at the output end of the encoder to obtain a lightweight Transformer model. Use the lightweight Transformer model as the basic model and construct a pre-stored answer library. Given a training set, the medical Q&A pairs include medical questions and corresponding unique target answers;
[0026] In this step, a brand-new lightweight Transformer model is adopted, and its number of parameters is effectively reduced to about 15MB, and the computational complexity is also optimized to , thereby significantly improving the efficiency, where represents the hidden dimension, is the complexity calculation operator, indicating that the complexity is proportional to is proportional.
[0027] Embedding layer: Given an input token sequence of length n (length ), through a learnable embedding matrix (where represents the vocabulary university, the vocabulary size V is set to 21128, and the hidden dimension H is set to 384) map each token to an H-dimensional embedding vector:
[0028] ;
[0029] where represents the embedding vector of the th token, represents the embedding matrix, represents the one-hot encoded vector of the th token.
[0030] Encoder: This embodiment adopts a structure stacked by 3 layers of Transformer encoders, and each layer of encoder contains 6 attention heads.
[0031] Pooling layer: In order to convert the variable-length sequence representation output by the encoder into a fixed-length embedding vector, an average pooling layer is added after the output of the encoder in the present invention. Specifically, for the sequence representation output by the encoder , the pooling layer obtains a fixed-dimensional embedding vector by averaging all the vectors in the sequence dimension:
[0032] ;
[0033] wherein, represents the i th feature vector output by the encoder, represents the final embedding vector.
[0034] Step 2: Use the unique target answer corresponding to the medical problem as the positive sample, select answers from the training set that have relevant keywords but different context semantics from the medical problem as hard negative samples, and randomly select several answers from the training set as simple negative samples;
[0035] Step 3: Input the training set into the basic model to obtain the question embedding vector, positive sample embedding vector, simple negative sample embedding vector, and hard negative sample embedding vector, and store the positive sample embedding vector, simple negative sample embedding vector, and hard negative sample embedding vector in the pre-stored answer library;
[0036] Step 4: Construct a positive sample loss function according to the similarity between the question embedding vector and the positive sample embedding vector and the simple negative sample embedding vector respectively, construct a hard negative sample loss function according to the similarity between the question embedding vector and the hard negative sample embedding vector, and jointly construct a joint loss function by weighting the positive sample loss function and the hard negative sample loss function;
[0037] In this step, the cosine similarity is used for the similarity, and there is the following relational expression in the calculation process of the cosine similarity:
[0038] ;
[0039] wherein, represents the cosine similarity between the question and the target answer , represents the question embedding vector, represents the target answer embedding vector.
[0040] In this step, the relational expression of the positive sample loss function is as follows:
[0041] ;
[0042] wherein, Represents the positive sample loss function, Represents the similarity score between the question and the target answer, Represents a medical question, Represents the unique target answer corresponding to the medical question, Represents the question and the target answer similarity, Represents the set of simple negative samples, Represents the similarity between the question and the simple negative sample, Represents the total similarity between the positive sample and the simple negative sample, Represents the temperature coefficient.
[0043] In this step, the hard negative sample loss function relation is as follows:
[0044] ;
[0045] Among them, Represents the hard negative sample loss function, Represents the set of hard negative samples, Represents the similarity between the question and the hard negative sample, Represents the total similarity between the positive sample and the hard negative sample, Represents the temperature coefficient.
[0046] In this step, the weighted joint loss function relation is as follows:
[0047] ;
[0048] Among them, Represents the weighted joint loss function of the positive sample loss function and the hard negative sample loss function.
[0049] In the above solution, this embodiment proposes a method for screening out answers with keywords related to the question from the dataset (for example, for the question "Precautions for gynecological examinations", answers containing the keyword "gynecology" but with irrelevant context to the precautions can be screened out) but with different context semantics as hard negative samples. By introducing these more challenging negative samples during the training process, the model can be forced to learn more refined semantic features, thereby improving its discrimination ability.
[0050] To more effectively enhance the model's matching ability for the target answer, the present invention sets the loss weight of the positive sample to 2. By increasing the contribution of the positive sample in the loss function, the convergence of the model can be accelerated, and it can be made to pay more attention to learning the association between the question and the correct answer. It can effectively enhance the distinctiveness of the embedding vector and accelerate the convergence of the model. Experimental results show that: by this method, the loss value of the model can be reduced to an extremely low level of 0.00492.
[0051] Step 5: Input the combined loss function into the Adam optimizer, set the weight decay strategy and learning strategy of the Adam optimizer, and optimize the basic model parameters with the goal of maximizing the similarity between the problem embedding vector and the target answer and minimizing the similarity between the problem embedding vector and the non-target answer. After optimization, the trained model is obtained.
[0052] In this step, the optimization objective relationship is as follows:
[0053] ;
[0054] where, represents the model parameters, represents the set of positive sample pairs, represents the set of negative sample pairs, represents the medical problem, represents the non-target answer to the given problem, represents the cosine similarity between the problem and the non-target.
[0055] In this step, the learning strategy includes a warm-up phase and a cosine decay phase. The warm-up phase is defined as the first ten percent of the training steps, and in the warm-up phase, the learning rate linearly increases from 0 to the preset maximum learning rate; in the cosine decay phase, the learning rate decays according to the cosine function.
[0056] In the warm-up phase, the calculation process of the learning rate has the following relationship:
[0057] ;
[0058] where, represents the learning rate, represents the current training step, represents the total number of steps in the warm-up phase, represents the preset maximum learning rate;
[0059] In the cosine decay phase, the calculation process of the learning rate has the following relationship:
[0060] ;
[0061] where, represents the preset minimum learning rate, represents the total number of training steps, represents the pi.
[0062] Step 6: Set up a temporary retrieval post-processing module in the output part of the trained model to obtain a lightweight medical Q&A system, and use the lightweight medical Q&A system for medical Q&A.
[0063] In this step, the specific steps of using the lightweight medical Q&A system for medical Q&A are as follows:
[0064] When a user inputs a new medical question, first use the trained model to generate an embedding vector for the new medical question;
[0065] Calculate the similarity scores between the embedding vector of the new medical question and all the answer embedding vectors in the pre-stored answer library to obtain the similarity scores of all answers;
[0066] Use the temporary retrieval post-processing module to extract key medical terms or keywords from the new medical question;
[0067] Screen the top several candidate answers according to the magnitudes of the similarity scores of all answers, and then preferentially return the answers containing keywords from the aforementioned several candidate answers to obtain the final answer.
[0068] Taking 10,000 pairs of gynecological Q&A pairs as experimental data, the preprocessing process is as follows:
[0069] 1. Integrate the Q&A pair dataset: Screen the gynecological topics, ensure balanced classification, and delete invalid data.
[0070] 2. Term standardization: Unify terms such as "irregular menstruation", in combination with the dictionary and medical knowledge base.
[0071] 3. Remove stop words and irrelevant words: Delete words such as "excuse me" and "body" to highlight gynecological keywords.
[0072] 4. Length truncation: Limit to 50 tokens to adapt to the lightweight model.
[0073] 5. Positive samples and hard negative samples: Label positive samples and screen hard negative samples based on keywords and semantics.
[0074] In the experiment, for the question "How does a young girl have a gynecological examination" raised by the user, the system proposed by the present invention can successfully output the expected correct answer, and the similarity score of this answer reaches 0.4398. Among the top 10 most similar answers, the highest score reaches 0.9309. To further verify the generalization ability of the system, this embodiment additionally selects various types of medical questions in the test set for evaluation, such as "How to recover quickly from a cold" and "Dietary precautions for diabetic patients". For the question "How to recover quickly from a cold", the similarity score of the correct answer output by the system is 0.4521, and the highest score among the top 10 answers is 0.9156; for the question "Dietary precautions for diabetic patients", the similarity score of the correct answer is 0.4287, and the highest score among the top 10 answers is 0.9078.
[0075] In addition, this embodiment introduces Precision@K and Mean Reciprocal Rank (MRR) as evaluation metrics. On the test set, the Precision@1 of the system reached 82.5%, Precision@5 was 91.3%, and MRR was 0.876, indicating that the system can efficiently return correct answers in the top-k responses. To compare the performance of the present invention, this embodiment was compared with a baseline model (standard BERT, with 110MB of parameters). On the same test set, the Precision@1 of the standard BERT was only 76.8%, the MRR was 0.821, and the inference time was 0.32 seconds per question, while the inference time of the present invention was only 0.09 seconds per question, and the response speed was increased by about 3.5 times. At the same time, the number of parameters of the model of the present invention was controlled below 15MB, and compared with 110MB of the standard BERT, the resource occupancy was reduced by more than 86%.
[0076] The above experimental results show that: the present invention is not only superior to the baseline model in terms of semantic matching accuracy, but also shows significant advantages in terms of inference speed and resource occupancy, and can efficiently adapt to low-computing-power devices (such as VR devices), which fully demonstrates the application value and effectiveness of the present invention in actual medical Q&A scenarios.
[0077] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0078] It should be understood that each part of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following well-known technologies in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0079] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples.
[0080] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
Claims
1. A method for constructing a lightweight medical Q&A system, characterized in that The method includes the following steps: Step 1: Based on the Transformer model, simplify the number of encoder layers and the number of attention heads of the Transformer model, and insert a pooling layer at the output end of the encoder to obtain a lightweight Transformer model. Use the lightweight Transformer model as the base model, and construct a pre-stored answer library. Given a training set, the training set includes several medical Q&A pairs covering inspection, prevention, treatment, and symptoms. The medical Q&A pairs include medical questions and corresponding unique target answers; Step 2: Use the unique target answer corresponding to the medical question as the positive sample, select answers from the training set that have relevant keywords but different contextual semantics from the medical question as hard negative samples, and randomly select several answers from the training set as simple negative samples; Step 3: Input the training set into the base model to obtain question embedding vectors, positive sample embedding vectors, simple negative sample embedding vectors, and hard negative sample embedding vectors, and store the positive sample embedding vectors, simple negative sample embedding vectors, and hard negative sample embedding vectors in the pre-stored answer library; Step 4: Construct a positive sample loss function according to the similarity between the question embedding vector and the positive sample embedding vector and the simple negative sample embedding vector respectively, construct a hard negative sample loss function according to the similarity between the question embedding vector and the hard negative sample embedding vector, and jointly construct a combined loss function by weighting the positive sample loss function and the hard negative sample loss function; Step 5: Input the combined loss function into the Adam optimizer, set the weight decay strategy and learning strategy of the Adam optimizer, and optimize the parameters of the base model with the goal of maximizing the similarity between the question embedding vector and the target answer and minimizing the similarity between the question embedding vector and non-target answers. After optimization, obtain the trained model; Step 6: Set up a temporary retrieval post-processing module at the output part of the trained model to obtain a lightweight medical Q&A system, and use the lightweight medical Q&A system for medical Q&A; In the said Step 4, the relationship formula of the positive sample loss function is as follows: ; Among them, represents the positive sample loss function, represents the similarity score between the question and the target answer, represents the medical question, represents the unique target answer corresponding to the medical question, represents the question and the target answer similarity, represents the simple negative sample set, represents the similarity between the question and the simple negative sample, represents the total similarity of the positive sample and the simple negative sample, represents the temperature coefficient; The relationship formula of the hard negative sample loss function is as follows: ; Among them, represents the hard negative sample loss function, represents the set of hard negative samples, represents the similarity between the problem and the hard negative samples, represents the total similarity between the positive samples and the hard negative samples, represents the temperature coefficient; The following relationship formula exists in the process of jointly constructing the positive sample loss function and the hard negative sample loss function by weighting: ; Among them, represents the weighted joint loss function of the positive sample loss function and the hard negative sample loss function; In the said Step 5, the learning strategy includes a warm-up stage and a cosine decay stage. Among them, the warm-up stage is defined as the first ten percent of the training steps, and in the warm-up stage, the learning rate linearly increases from 0 to the preset maximum learning rate; in the cosine decay stage, the learning rate decays according to the cosine function; In the warm-up stage, the calculation process of the learning rate has the following relationship formula: ; Among them, represents the learning rate, represents the current training step number, represents the total number of steps in the warm-up stage, represents the preset maximum learning rate; In the cosine decay stage, the calculation process of the learning rate has the following relationship formula: ; Among them, represents the preset minimum learning rate, represents the total number of training steps, represents pi.
2. The method for constructing a lightweight medical Q&A system according to claim 1, wherein In the said Step 4, the cosine similarity is used for the similarity, and the calculation process of the cosine similarity has the following relationship formula: ; Among them, represents the problem and the target answer of the cosine similarity, represents the problem embedding vector, represents the target answer embedding vector.
3. The method for constructing a lightweight medical Q&A system according to claim 1, wherein, In the said Step 5, the optimization target relationship formula is as follows: ; Among them, represents model parameters, represents the set of positive sample pairs, represents the set of negative sample pairs, represents a medical problem, represents a non-target answer to a given problem, represents the cosine similarity between the problem and the non-target.
4. The method for constructing a lightweight medical Q&A system according to claim 1, wherein In the said Step 6, the specific steps of using the lightweight medical Q&A system for medical Q&A are as follows: When a user inputs a new medical question, first use the trained model to generate an embedding vector of the new medical question; Calculate the similarity scores between the embedding vector of the new medical problem and all the answer embedding vectors in the pre-stored answer library to obtain the similarity scores of all answers; Use the temporary retrieval post-processing module to extract key medical terms or keywords from the new medical problem; Screen the top several candidate answers according to the magnitudes of the similarity scores of all answers, and then preferentially return the answers containing keywords from the top several candidate answers to obtain the final answer.
Citation Information
Patent Citations
Binary Code Similarity Detection System Based on Hard Sample-aware Momentum Contrastive Learning
US20250110715A1
Automatic medical question answering method and apparatus, storage medium, and electronic device
WO2020034642A1