Multi-task large language model training method oriented to biomedical field
By building a multi-tasking large language model for the field of biomedical science and adopting context learning, instruction fine-tuning and thinking chain methods, the complexity and dynamic nature of text processing in the field of biomedical science is solved, and high-quality multi-tasking and entity recognition are achieved.
Patent Information
- Application Number
- CN202510227438.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Text processing in the field of biomedical science has complex professional terms, fast dynamic data changes, and high requirements for the accuracy and reliability of text processing, resulting in poor application of existing large language models in this field.
Using a multi-task large language model training method for the field of biomedical science, we construct instruction data sets for medical intelligent question-and-answer, medical report generation and biomedical information extraction tasks, and improve the generalization ability and adaptability of the model through context learning, instruction fine-tuning and thinking chain methods.
It has achieved significant improvement in the answer quality, generalization ability and entity position recognition ability of the model in multitasking in the field of biomedical science, and met the high standards for text processing in the field of biomedical science.
Smart Images

Figure CN120069086A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of large language model training, and relates to a multi-task large language model training method for the biomedical field. Background Art
[0002] In recent years, large language models (LLMs) have attracted much attention due to their excellent performance in multiple general natural language processing tasks. These models have demonstrated powerful generation capabilities, understanding capabilities, and few-shot / zero-shot learning capabilities through large-scale parameters and massive training data, and have made breakthrough progress in tasks such as text generation, question answering, and translation. However, due to its unique text characteristics and stringent application requirements, the biomedical field shows significant domain differences and technical bottlenecks in the direct application of LLMs, which have become important obstacles restricting its development in this field.
[0003] First of all, the biomedical field involves a large number of professional terms, proper nouns and their abbreviations, and these terms have a high dependence on professional background knowledge. In addition, the complex syntactic structures and uncommon vocabulary in biomedical texts further increase the difficulty of text processing. Secondly, due to the rapid development of biomedical research, a large number of new terms and knowledge points are constantly emerging, resulting in an obvious lag in the performance of traditional pre-trained models in dealing with new domain knowledge. This highly dynamic domain characteristic poses higher requirements on the generalization ability and adaptability of the model. More importantly, the biomedical field has extremely high requirements for the accuracy and reliability of text processing. In tasks such as clinical diagnosis, medical report generation, and intelligent question answering, model errors may directly affect medical decisions, thereby threatening the lives of patients. In addition, the model also needs to balance efficiency when processing biomedical data to meet the needs of large-scale data processing, while providing sufficient interpretability to ensure the credibility and traceability of prediction results.
[0004] In addition, although general large language models (LLMs) such as GPT-4, PaLM, and LLaMA perform excellently in a wide range of natural language processing tasks, they usually struggle to meet domain-specific requirements when directly applied to the biomedical field. This is because general models are mainly trained on open-domain data, which have significant differences from biomedical data in terms of corpus distribution, language patterns, and domain knowledge. Therefore, general models are prone to problems such as inaccurate semantic understanding, incorrect term translation, and knowledge gaps when processing biomedical texts, leading to a significant decline in model performance. Currently, research on this issue focuses on optimization through two directions: domain fine-tuning and dedicated model development. Some studies have significantly improved the performance of models in specific tasks by fine-tuning general models (such as PubMedBERT, BioBERT) on biomedical domain data. However, most of these methods are optimized for single tasks (such as question answering or text classification), and the multi-task generalization ability of the models has not been fully explored. In addition, models developed for single languages (such as mainly in English) have limited performance in multi-language tasks, while research and practice in the biomedical field often require cross-language processing capabilities, which also pose higher challenges for model development.
[0005] However, in the current application scenarios, there is usually a separate model for each task or each language. This approach not only consumes a large amount of resources but also has many inconveniences in deployment and use. Users need to frequently switch models among various requirements, increasing the operational complexity. In addition, small-parameter models have significant performance bottlenecks and poor reliability.
[0006] Therefore, developing multi-task large language models suitable for the biomedical field is not only an important direction for current technological development but also an urgent need to solve practical problems in the field of biomedical information processing. Such models not only need to be specifically fine-tuned on domain-specific data with strong professionalism but also need to fully consider the specific requirements of diverse tasks to improve the generality and adaptability of the models. In addition, by constructing bilingual or multi-lingual models with cross-language support, the multi-language communication needs in international medical research and practice can be effectively addressed. Summary of the Invention
[0007] To solve the problem of a single large language model achieving multi-task processing, in a first aspect, according to some embodiments of the present application, a training method for a multi-task large language model for the biomedical field, the model is used to automatically select and execute medical intelligent question answering, medical report generation, or biomedical information extraction tasks according to the task type of the input question;
[0008] The method includes
[0009] Construct a training set for the instruction dataset for medical intelligent question - answering tasks. The data in the instruction dataset has a first instruction format, and the first instruction format includes a data instruction format based on in - context learning and a data instruction format based on chain of thought.
[0010] Construct a training set for the second instruction dataset for medical report generation tasks. The training set data of the second instruction dataset has a second instruction format, and the second instruction format includes a data instruction format based on enhanced entity information embedding and a data instruction format based on enhanced spelling correction.
[0011] Construct a training set for the third instruction dataset for biomedical information extraction tasks. The training set data of the third instruction dataset has a third instruction format, and the third instruction format includes a data instruction format based on task description and type definition and a data instruction format based on special symbol marking.
[0012] Use the training set to train the large - language model.
[0013] According to the multi - task large - language model training method for the biomedical field in some embodiments of the present application, it further includes constructing a MOSS dataset for general - domain conversations, which is used to automatically select and execute general - domain conversations according to the task type of the input question.
[0014] According to the multi - task large - language model training method for the biomedical field in some embodiments of the present application, it further includes constructing a model self - awareness dataset, wherein the model self - awareness data is scattered and inserted into each of the training sets.
[0015] According to the multi - task large - language model training method for the biomedical field in some embodiments of the present application, wherein the large - language base model is a GLM - 4 - 9B model, and the hyperparameters of the large - language model are as follows:
[0016] Learning rate: 1e - 05
[0017] Gradient accumulation steps: 16
[0018] Batch size: 1
[0019] Data type: bf16
[0020] Random seed: 42.
[0021] According to the multi-task large language model training method for the biomedical field in some embodiments of the present application, wherein, constructing a training set of an instruction data set for a medical intelligent question-answering task, including
[0022] Construct a first training data in the training set by using the first instruction, a piece of data in the training set of the first data set, or by using the first instruction, a piece of data in the training set of the first data set, and examples of selected minority samples from the first data set. Obtain each piece of first training data corresponding to each piece of data in the training set of the first data set. The first training data is data in a context-based data instruction format;
[0023] Use the large language model to generate an inference corresponding to a piece of data in the training set of the first data set through the second instruction. Construct a second training data in the training set by using the second instruction, the piece of data, and the inference. Obtain each piece of second training data corresponding to each piece of data in the training set of the first data set. The second training data is data in a thought-chain-based data instruction format.
[0024] According to the multi-task large language model training method for the biomedical field in some embodiments of the present application, wherein:
[0025] The content of the first instruction includes "Answer in the format of 'Answer': option description";
[0026] The content of the second instruction includes "Show the reasoning process".
[0027] According to the multi-task large language model training method for the biomedical field in some embodiments of the present application, wherein, constructing a training set of a second instruction data set for a medical report generation task, including
[0028] Generate the entity category in a piece of data in the training set of the second data set. Construct a third training data in the training set by using the piece of data and the entity category. Obtain each piece of third training data corresponding to each piece of data in the training set of the second data set. The third training data is data in an entity information embedding-based data instruction format;
[0029] Use the large language model to generate a piece of data with spelling errors corrected for a piece of data in the training set of the second data set through the fourth instruction. Construct a fourth training data in the training set by using the fourth instruction, the piece of data, and the piece of data with spelling errors corrected. Obtain each piece of fourth training data corresponding to each piece of data in the training set of the second data set. The fourth training data is data in a spelling correction enhancement-based data instruction format.
[0030] According to the multi-task large language model training method for the biomedical field in some embodiments of the present application, wherein:
[0031] Generate the entity category in a piece of data in the training set of the second data set, including extracting the entity category using a medical entity extraction tool;
[0032] The content of the fourth indication includes "spelling error correction".
[0033] According to the multi-task large language model training method for the biomedical field in some embodiments of the present application, wherein, constructing the training set of the third instruction data set for the biomedical information extraction task includes
[0034] Construct a piece of the fifth training data in the training set by using the fifth indication and a piece of data in the training set of the third data set as a piece of data in the training set. Obtain each piece of the fifth training data corresponding to each piece of data in the training set of the third data set. The fifth training data is data in a data instruction format based on the task description and type definition; wherein, the content of the fifth indication includes the task description and type definition;
[0035] Generate multiple text sequences for a piece of data in the training set of the third data set, and each text sequence has an entity marked with a special symbol. The piece of data and the text sequences are constructed as a piece of the sixth training data in the training set. Obtain each piece of the sixth training data corresponding to each piece of data in the training set of the third data set. The sixth training data is data in a data instruction format based on the special symbol marking.
[0036] In a second aspect, an embodiment of the present application further provides an electronic device, which includes: one or more processors, a memory, and one or more programs; wherein, the one or more programs are stored in the memory, and the one or more programs include instructions. When the instructions are executed by the electronic device, the electronic device executes the first aspect and any possible technical solution of the first aspect.
[0037] Beneficial effects: By deeply researching the key technologies of large language models for the biomedical field, the present invention has achieved remarkable results, successfully constructed a multi-lingual and multi-task biomedical large model with excellent performance, covering multiple tasks such as multi-lingual biomedical intelligent question answering, doctor-patient dialogue, medical report generation, and biomedical information extraction.
[0038] Through technical means such as in-context learning, instruction fine-tuning, and chain-of-thought methods, the present invention effectively improves the answer quality and generalization ability of the model in Chinese medical question-answering tasks. In terms of medical report generation tasks, the present invention adopts optimization methods such as named entity information embedding and spelling correction, significantly improving the abstract generation performance of the model. At the same time, in terms of biomedical information extraction tasks, different instructions and decoding strategies based on large models are explored, which can not only accurately predict entity mentions but also indicate the specific positions of entities in the text.
[0039] As described above, the present invention integrates a large number of open-source Chinese and English datasets covering various tasks, and further enhances the generalization ability and entity position recognition ability of the model through carefully designed instruction templates and decoding strategies. In addition, the present invention also constructs role-playing data, multi-task question-answering data, and self-awareness data to further improve the comprehensive ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 Overall framework.
[0041] Figure 2 Zero-shot example.
[0042] Figure 3 Few-shot example.
[0043] Figure 4 Example of the instruction format of the MedQACoT training set.
[0044] Figure 5 Comparison of answers generated by two models.
[0045] Figure 6 Data sample after enhanced processing of entity information embedding.
[0046] Figure 7 Example of spelling mistakes in the HQS dataset.
[0047] Figure 8 Example of spelling mistake detection and correction in the HQS task.
[0048] Figure 9 Simple instruction.
[0049] Figure 10 Task description and type definition instruction.
[0050] Figure 11 Task description, type definition, and example instruction.
[0051] Figure 12 HTML decoding scheme.
[0052] Figure 13 Index decoding scheme.
[0053] Figure 14 Special symbol marking decoding scheme.
[0054] Figure 15 Instruction data format example.
[0055] Figure 16 Medical intelligent answer instruction data format example.
[0056] Figure 17 Self-awareness, general domain dialogue instruction data format example.
[0057] Figure 18 Self-awareness and dialogue ability test. Specific implementation manners
[0058] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0059] I. Overview With the rapid development of large language model technology, these models have made remarkable progress in a variety of general domain tasks, demonstrating strong generality and task adaptation capabilities. However, in the biological field, there are still unique challenges such as strong professionalism, rapid dynamic changes, and large data scale in biomedical texts. Therefore, the present invention aims to propose a multi-task fine-tuning method based on large language models for the special needs of the biomedical field, and construct a large language model system that can meet the high standards of the industry.
[0060] The core objective of the present invention is to develop a multi-task large language model specifically applicable to the biomedical field by integrating the research results of natural language processing technology, large language models, and biomedical informatics. This model can efficiently handle multiple task scenarios including biomedical intelligent question answering, doctor-patient dialogue, medical report generation, and information extraction.
[0061] The present invention covers three core tasks: medical question answering, medical report generation, and biomedical information extraction. In addition, considering that the present invention also needs to have a certain self-awareness ability and general domain dialogue ability, therefore, on the basis of task optimization, how to enable the model to clearly identify its own identity and have a certain basic dialogue ability will also be explored.
[0062] To ensure the best performance of each task, the present invention will prioritize the independent optimization of medical question answering, medical report generation, and biomedical information extraction tasks, develop independent optimization strategies for each task respectively, and explore the optimal training methods for each core task. After completing the optimization of the core tasks, a series of methods will be explored next to further enhance the self-awareness ability of the present invention and endow it with the dialogue ability in the general domain, enabling it to communicate naturally with users. After completing all the explorations, summarize these innovative task optimization methods and use the existing domain resources for integrated training to obtain the final biomedical large language model with multi-task bilingual capabilities. The overall framework of the present invention is as follows Figure 1 as shown
[0063] II. Technical Term Explanation
[0064] Large Language Models (LLMs for short) are artificial intelligence models built based on deep learning technology, aiming to understand and generate human language. These models usually contain tens of billions or even hundreds of billions of parameters and learn the patterns, grammar, and semantics of language through training on large-scale text data. The core architecture of large language models is mostly based on Transformer, which can capture the complexity and diversity of language and generate natural, fluent, and contextually understood text
[0065] III. Medical Intelligent Question Answering Task
[0066] The medical intelligent question answering task module uses methods such as in-context learning, instruction tuning, and chain-of-thought construction to gradually and deeply analyze and experiment with the performance of domestic large language models in Chinese medical question answering tasks. First, use prompt words to guide the model base of the Chat version to answer questions. Since the Chat version of the model has certain general capabilities, it can effectively answer Chinese medical-related questions through zero-shot learning or few-shot learning methods. Then, the QLoRA method is also used to perform efficient instruction tuning on multiple large language models to select the most suitable model for subsequent experiments. Furthermore, to improve the performance of models with relatively small numbers of parameters in medical question answering, the present invention uses the advanced large language model GLM-4-9B, leveraging its powerful generation ability to expand the MedQA dataset, constructing an instruction dataset with a chain of thought by generating medical explanations corresponding to the questions, thereby effectively improving the performance of the model in Chinese medical question answering. The specific implementation technical solutions and experimental results of this task module will be introduced in detail below
[0067] 3.1 In-Context Learning
[0068] In-context learning of large language models refers to, before the problem input, by providing a certain number of example samples, enabling the model to understand the pattern of the current task based on these examples and stimulate the knowledge inherent in the model itself, so as to answer questions more effectively. This method is also known as few-shot prompting learning, and the key lies in the design of prompt words and the selection of example data. Prompt words play a guiding role. They can clearly convey the goals and formats of tasks, thus stimulating the model's processing ability for specific tasks. Example data, on the other hand, plays the role of actual teaching. By showing specific instances of tasks, it helps the model better understand how to apply knowledge to solve problems. Just like when a student solves a math problem, looking at a few relevant example problems in advance can clarify the problem-solving steps and techniques, thus completing similar problems more smoothly.
[0069] When constructing prompt words for few-shot learning, the first step is to clearly define and describe the natural language processing task to be solved. In the present invention, the model is positioned as a "medical expert", and the task is to solve Chinese medical Q&A multiple-choice questions. Then, a small number of samples are carefully selected from the dataset related to the task as examples. These samples should be representative and cover various situations and variations that may occur in the task.
[0070] The present invention uses the CMB dataset in the in-context learning task. CMB contains Chinese medical Q&A multiple-choice questions of multiple categories. The present invention retains the data related to medicine, which are respectively physician examinations, nursing examinations, pharmacist examinations, medical technology examinations, professional knowledge examinations, and medical postgraduate entrance examinations, and at the same time screens and removes the data unrelated to medicine inside, such as subset data like political theory for postgraduate entrance examinations.
[0071] To ensure the effectiveness of the prompt words and prevent the model from being confused due to too many example samples, the number of selected samples is usually controlled between 3 and 5. After selecting the examples, in accordance with the principle of simplicity and clarity, the present invention organizes the selected examples into prompt content according to a specific template. After many attempts, it is decided to use "Question:...... Answer:......" as the template and arrange the examples in order to ensure clear logical relationships between the examples.
[0072] Finally, the test samples are added to the end of the prompt words and input into the large language model together with the examples for processing. The model generates prediction results for the test samples based on the prompt words and examples. The present invention uses 0-shot, 3-shot, and 5-shot (representing the number of example samples being 0, 3, and 5 respectively) to test the medical Q&A ability of the Chat model. Corresponding prompt word examples are as Figure 2 、 Figure 3 . shown. The experimental results are shown in Table 3.2.
[0073] 3.2 Exploration related to instruction tuning
[0074] The present invention constructs instruction data based on existing datasets and fine-tunes each model using the QLoRA method. The fine-tuned models can accurately output medical Q&A answers according to the instruction format, and finally evaluate the generated answer effects through scripts to explore the best way to construct the instruction dataset.
[0075] First, the present invention fine-tunes the model on the MedQA training set and evaluates the Q&A ability on the C-Eval medical-related subsets (seven subsets of clinical medicine, veterinary medicine, junior high school biology, physician qualification, basic medicine, plant protection, and senior high school biology respectively) and the MedQA test set. Using accuracy as the main evaluation index, it explores the performance of different models in Chinese medical Q&A tasks and the differences between models with different parameter scales. The results are shown in Table 3.3.
[0076] Then, the present invention fine-tunes the model using the training sets of three existing datasets (MedQA, CMExam, CMB) respectively and conducts performance evaluations on each test set to test the generalization ability of the model. At the same time, in order to enhance data diversity and enable the model to learn a wider range of patterns to improve its generalization ability in different tasks and scenarios, the present invention fuses the training sets of these three datasets to create a new training set Mix3. For the rigor of the experiment, duplicate data between the training sets is removed during the merging process, and the part that duplicates the test set is excluded to avoid data leakage. Fusing multiple datasets can also prevent the model from overly relying on the features of a specific dataset and enhance the universality of the model.
[0077] By fine-tuning using Mix3 and evaluating the model performance on the C-Eval, CMExam, CMB, and MedQA test sets, the present invention studies the impact of training data diversity on model performance, further reveals the relationship between data diversity and model performance, and provides an important reference for future model optimization. The results are shown in Table 3.4
[0078] 3.3 Constructing the Chain of Thought Dataset
[0079] In order to fully utilize the ability of large language models in reasoning and generation, the present invention combines the chain of thought method and constructs instruction data containing medical explanations during the fine-tuning process, enabling the model to follow the process of "reason first and then give the answer" when answering Chinese medical questions, thereby tapping its potential in medical Q&A.
[0080] To this end, the present invention constructs chain-of-thought data on the training set of MedQA. First, the present invention uses the system prompt of "As a medical expert, your task is to help users solve complex Chinese medical multiple-choice questions. For each question, show your reasoning process and select the best option.", consuming nearly five million tokens, and uses the GLM-4-9B model to generate corresponding medical answer explanations for 7,106 pieces of data in the training set of MedQA. After that, the present invention extracts the data answers using regular expressions, and manually identifies and modifies the answers generated by the model. Finally, by comparing with the correct answers, 6,221 pieces of data with correct reasoning are retained, forming a new chain-of-thought training set MedQACoT, the format of which is as Figure 4 shown, where the yellow module is the instruction, the light blue module is the Chinese medical question, the green module is the chain-of-thought reasoning process, and the dark blue module in the last row is the final answer inferred by the model.
[0081] 3.4 Experimental Results
[0082] During the training process, taking ChatGLM3-6B-Base as an example for parameter settings, the number of training rounds is 5, the learning rate is 2e-4, and a warm-up learning strategy is adopted with a warm-up step of 500, the rank of the QLoRA matrix is 64, the scaling factor is 16, and fp16 mixed precision is used. During the inference process, the temperature coefficient temperature is set to 0.35, and the sampling coefficient top_p is set to 0.9. Since all the data in the Chinese medical Q&A dataset used in the present invention are multiple-choice questions, the accuracy rate is used as the evaluation index in the experiment, and the calculation formula is: "Accuracy rate = Number of correct answers / Total number of questions".
[0083] The datasets used in the experiments of the present invention include the following main datasets:
[0084] (1) C-Eval: The C-Eval dataset consists of 13,948 multiple-choice questions, which cover a wide range of content from humanities to science and engineering, covering 52 different disciplines, and also include four difficulty levels: middle school, high school, university, and professional exams. All the questions in this dataset are single-choice questions. The present invention mainly uses the data subset related to Chinese medical Q&A in this Chinese LLM evaluation suite for model evaluation, including questions related to clinical medicine, veterinary medicine, junior high school biology, medical licensing, basic medicine, plant protection, and high school biology. The quantity of each category is shown in Table 3.1.
[0085] (2) MedQA: The MedQA dataset is the first free-form multiple-choice open Q&A dataset for medical questions, containing three languages: English, Simplified Chinese, and Traditional Chinese, with 12,723, 34,251, and 14,123 questions respectively. The questions in this dataset are multiple-choice questions, and the invention uses the Simplified Chinese part of the MedQA dataset for experiments..
[0086] (3) CMB: The CMB dataset is a comprehensive and multi-level Chinese medical evaluation dataset, containing 280,839 multiple-choice questions and 74 complex case consultation questions, covering all clinical medicine specialties and different professional levels. The goal of this dataset is to comprehensively evaluate the medical knowledge and clinical consultation ability of the model, covering multiple categories such as physician examinations, nursing examinations, pharmacist examinations, medical technician examinations, professional knowledge examinations, and medical postgraduate entrance examinations, and also including single and multiple-choice questions.
[0087] (4) CMExam: The questions in the CMExam dataset come from the National Medical Licensing Examination of China, and it contains more than 60,000 multiple-choice questions in total. This dataset is used for standardized and objective evaluation, and the test set also contains five additional annotation dimensions: disease group, clinical department, medical department, medical discipline, ability area, and question difficulty level. 85.24% of the questions in the dataset also contain medical explanations for some questions, and the difficulty levels are divided into 1 to 5, representing easy, manageable, medium, difficult, and extremely difficult respectively. The division of the training set, development set, and test set of each dataset is shown in Table 3.1.
[0088] Table 3.1 Dataset usage information
[0089]
[0090] The large language models used in this invention for research are mainly domestic open-source Chinese large language models, including the Tongyi Qianwen (Qwen) series of Alibaba Cloud, the Baichuan series of Hundred-Peak Intelligence, and the ChatGLM series of Zhipu.
[0091] (1) Comparison of question-answering performance based on in-context learning
[0092] To study the zero-shot learning ability of large language models and their performance in few-shot learning, this section conducted experiments on the test set of the CMB for the Chat version of the models. Zero-shot learning directly asks questions through prompts, while few-shot learning adds examples selected from the CMB training set to the prompts of zero-shot learning to form an example template. Table 3.2 shows the performance comparison of the three models on the CMB test set.
[0093] Table 3.2 Performance comparison between few-shot learning and zero-shot learning (%)
[0094]
[0095] The bold part is the highest value in this column. N-Shot represents the number of examples included in the prompt. 0 means there are no examples, and 3 and 5 represent having 3 and 5 examples respectively.
[0096] The results show that Qwen-7B-Chat performs significantly better than the other two models. Whether it is zero-shot learning or few-shot learning, its accuracy is higher than that of other models. In particular, Qwen-7B-Chat shows strong instruction-following ability in zero-shot learning, and its accuracy in zero-shot learning is even higher than that in few-shot learning. At the same time, both ChatGLM3-6B and Qwen-1.8B-Chat perform better in few-shot learning than in zero-shot learning. This indicates that few-shot learning can further improve the question-answering ability of the model by adding examples. Although 5-shot has two more examples in the prompt than 3-shot, it does not improve the model's accuracy.
[0097] (2) Comparison of the question-answering performance of different large language models after instruction fine-tuning
[0098] To test the performance of different domestic large language models in the Chinese medical question-answering task, in this section, domestic open-source large language models are fine-tuned on the MedQA training set, and then evaluated on the C-Eval medical-related subset and the MedQA test set respectively, and compared and analyzed. The test results of each model after fine-tuning training are shown in Table 3.3.
[0099] Analyzing the test results in Table 3.3, the following four conclusions can be drawn:
[0100] 1) Among models in the same series, models with larger parameter quantities generally perform better on the test set. For example, the 13B model in the Baichuan2 series has an average accuracy on C-Eval that is about 6% higher than that of the 7B model, and also shows an improvement of about 5% on the MedQA test set; similarly, the 7B model in the Qwen series is about 6% higher than the 1.8B model on the MedQA test set and nearly 10% higher on C-Eval. This shows that the larger the parameter quantity, the more content and knowledge the large language model can learn, so it can use its vast knowledge to better solve problems in the test and thus obtain a performance improvement.
[0101] Table 3.3 Fine-tuning test results of each model (%)
[0102]
[0103]
[0104] The underlined part is the highest value of the domestic large language model in this column, and the bold part is the highest value of all models in this column after considering GPT-3.5 and GPT-4.
[0105] 2) The performance of the Base and Chat versions of different large language models varies greatly. Taking Qwen-1.8B as an example, the Chat version performs better than the Base version on the C-Eval and MedQA test sets, with a gap of about 4%; for the ChatGLM3 series, the performance of the Base version is generally better than the Chat version; while the performance difference between the two versions of the Baichuan2-7B series is small. This difference may be closely related to the strategies and data used in the training process.
[0106] 3) A language model with a small number of parameters and a large number of parameters is not necessarily inferior to a language model with a large number of parameters. This does not conflict with conclusion (1). From the results in Table 2.7, we can see that ChatGLM3-6B-Base performs better than Baichuan2-13B, with an average accuracy of nearly 7% higher on the C-Eval test set. This phenomenon can be explained by the following reasons: First, the size of the parameters is not always directly positively correlated with the performance of the model on a specific task. ChatGLM3-6B-Base may have better generalization ability and can better learn the key information related to the task, so that it can answer the questions in the test set more accurately after fine-tuning. In addition, although models with a large number of parameters may have stronger learning ability, this may also cause them to learn some noise or irrelevant information that is not related to the task, resulting in overfitting or performance degradation. Models with a small number of parameters are usually more concise and can focus on the key information related to the task, thus avoiding such interference. In addition, the initialization and training process of the model may also affect the performance after fine-tuning. If a model with a large number of parameters is not fully initialized or fine-tuned, it may not be able to fully realize its potential. In contrast, ChatGLM3-6B-Base may perform better in specific tasks through better training strategies.
[0107] 4) Compared with GPT-3.5, the test results of ChatGLM3-6B-Base and Baichuan2-13B are generally better than GPT-3.5; however, compared with GPT-4, there is still a significant gap in domestic open source models. Except for the "High School Biology" questions and the MedQA test set, ChatGLM3-6B-Base performed slightly better than GPT-4, and the performance of other test sets was surpassed by GPT-4.
[0108] Table 3.4 Results of fine-tuning models with different training sets on different test sets (%)
[0109]
[0110] The underlined part indicates not considering the highest value in the column of test results trained on Mix3 and GPT-4 test results, and the bold part is the highest value in the column considering all test results.
[0111] Table 3.4 shows the results of the model evaluated on different test sets after fine-tuning on different training sets. The experimental results show that the model has good generalization ability after fine-tuning training. For the results of the C-Eval test set, it can be found that without considering GPT-4, the ChatGLM3-6B-Base fine-tuned using the CMExam dataset has the best performance. By examining the specific data analysis in this paper, it may be that the content correlation between these two data is relatively high. At the same time, the training set of CMExam has a certain data scale, and the model obtains relevant useful knowledge to solve the problems in the C-Eval dataset during the fine-tuning process, thus achieving good results.
[0112] By observing the underlined data, it can be analyzed that the generalization effect of fine-tuning on the CMB dataset is the best, which is to a certain extent due to the large amount of data in CMB. The model after fine-tuning on Mix3 obtained the highest value on most test sets. Except for being lower than GPT-4 in the C-Eval item, it is higher than GPT-4 in other items. This shows that the fused large-scale dataset contains more knowledge and enhances data diversity, which helps the model better adapt to various situations during the fine-tuning process. Therefore, using a large-scale dataset for fine-tuning can improve the stability of the model and enable it to maintain good performance in different scenarios.
[0113] (3) Comparison of Q&A performance based on the chain of thought method
[0114] To further explore the reasoning ability of large language models, this section uses the CMExam dataset and the self-constructed MedQACoT dataset with medical explanations to construct instruction data with the chain of thought, that is, <input, chain of thought, output>, fine-tunes the model, and tests it on the CMExam test set. The results are shown in Table 3.5 respectively.
[0115] Table 3.5 Results of chain of thought fine-tuning on the CMExam dataset (%)
[0116]
[0117] Among them, AO (answer only) indicates fine-tuning without the chain of thought, and CoT indicates fine-tuning using instruction data with the chain of thought. The bold part is the highest value in this column.
[0118] From the data in Table 3.5, it can be found that the performance of the model after adding the chain of thought did not improve significantly, and even the accuracy of some models decreased. The reason analysis shows that the number of parameters of the model is small, and the fine-tuned model may have reached the upper limit of its answering ability, and the chain of thought did not further stimulate its "emergent" ability. In the research of large language models, "emergence" refers to the fact that when the model scale reaches a certain threshold, its ability will experience a qualitative leap, showing excellent language understanding and reasoning abilities. Usually, only when the number of parameters of the model reaches 10B to 100B can the emergent ability be demonstrated.
[0119] By comparing the answering effects of the Chat version model, the present invention finds that for the ChatGLM3-6B-Base model fine-tuned with data with the chain of thought, its reasoning process is more in line with medical common sense, while the general dialogue model ChatGLM3-6B has a more serious "hallucination" problem in Chinese medical Q&A. For example Figure 5 as shown, the correct answer should be option D. Although the analysis of options A and D by ChatGLM3-6B during the answering process seems reasonable, compared with the answer of the fine-tuned ChatGLM3-6B-Base model, it is found that these analyses do not follow medical common sense and are actually fictional explanations by the model.
[0120] The reason for this phenomenon is that general dialogue models are usually pre-trained based on a large amount of cross-domain data to solve a wide range of natural language processing tasks. However, this broad generality may lead to unsatisfactory performance of the model in specific domains (such as medical Q&A). Through fine-tuning, the model can be trained for a specific domain on the basis of pre-training, learning professional terms, contexts and rules within the domain, thus significantly improving the model's performance in this domain. This targeted fine-tuning training can help the model better solve specific tasks and avoid inaccurate answers due to lack of professional knowledge.
[0121] IV. Medical report generation task
[0122] The medical report generation task module mainly focuses on the automatic summarization of medical texts, and conducts research on three types of task data: health problem summaries, medical radiology report generation, and doctor-patient dialogue summaries. Medical texts usually contain complex entity information and rich context relationships, and generating accurate summaries faces challenges. Therefore, the present invention proposes to optimize the performance of the model in the summary generation task through methods such as entity information embedding, spelling error detection and correction. The specific implementation technical solutions and experimental results are as follows.
[0123] 4.1 Entity information embedding enhanced generation method
[0124] In the medical report generation task, due to the long length of the dialogue data, the model often loses memory of some key information during the generation process, resulting in the omission of important content or inaccurate information in the summary. Therefore, in the medical dialogue summary generation task, it is crucial to incorporate entity information, which helps improve the model's ability to extract key information, reduce information forgetting, enhance the structure and readability of the summary, improve the model's understanding of the clinical background, and enhance the interpretability of the summary. By effectively utilizing entity information, the model can generate more accurate, comprehensive, and practical medical summaries, providing more valuable decision-making support for medical staff.
[0125] The present invention performs entity information embedding and conducts related experiments on the doctor-patient dialogue summary task (IMC-V2 dataset) and the medical radiology report generation task (RRS dataset) to explore the effectiveness of this method.
[0126] For the IMC-V2 dataset, each data in this dataset is a real doctor-patient dialogue, and the model needs to generate a medical summary based on the dialogue. The biological-related entities (symptoms, drugs, drug categories, examinations, operations, etc.) in the dialogue are extracted and added to the end of the dialogue in the format of "[entity category]: [entity content information]". For the RRS dataset, each data in this dataset is a radiology examination report, which is an English dataset. The present invention extracts the medical entities in the report through the open-source medical entity extraction tool MEDCAT and adds them to the end of the report content in the form of "[Contain Medical Entitys]:...". The processed data samples Figure 6 are shown as follows.
[0127] For the doctor-patient dialogue summary task, the present invention first directly performs QLoRA fine-tuning on the IMC-V2 dataset on the three models of Bio-ProphetNet, T5, and ChatGLM3 and tests it on its test set. The obtained results are used for subsequent experimental comparisons. After enhancing the IMC-V2 dataset through the entity information embedding method, the present invention performs fine-tuning and testing again on the ChatGLM3 model to explore the effectiveness of this method. At the same time, multiple experiments are conducted by changing the number of embedded entity information to explore the optimal strategy.
[0128] Similarly, for the medical radiology report generation task, the present invention performs QLoRA fine-tuning and testing on the RRS dataset on the three models of allenai / biomed_roberta_base, BERT-base, and Qwen2.5-7B-Instruct, and conducts experiments on Qwen2.5-7B-Instruct again after enhancing the dataset through entity embedding. The effectiveness of the method is demonstrated through the comparison of the experimental results.
[0129] 4.2 Enhanced Generation for Spelling Correction
[0130] In medical forums and online Q&A, users are often not experts, so spelling mistakes often exist in the dataset. Uncorrected spelling mistakes not only lead to mismatches between the source text and the abstract, but also produce incorrect outputs during the prediction process. Such mistakes are particularly prominent in string-matching based evaluation metrics (such as Rouge) because they disrupt n-gram matching. In addition, certain spelling mistakes may cause the model to deviate in semantic understanding, as Figure 7 shown.
[0131] For the above spelling correction solution, the present invention uses the GLM-4 model to detect spelling in the HQS training set, and by setting appropriate prompt words, the GLM-4 model is made to return the words to be modified, and the prompt words are as Figure 8 shown. Subsequently, manual verification and correction are carried out. This effectively reduces the spelling mistakes in the HQS dataset. The present invention fine-tunes the Llama-3.1 model using the original HQS dataset and the corrected HQS dataset, and tests it on the HQS test set, and compares the results to explore the effectiveness of the enhanced spelling correction solution.
[0132] 4.3 Experimental Results
[0133] The present invention conducts experiments on three tasks: Health Question Summary (HQS), Medical Radiology Report Generation (RRS), and Doctor-Patient Conversation Summary (IMCS-V2):
[0134] 1) Health Question Summary (HQS): This task aims to generate a concise question summary that covers the minimum information required to find the correct answer from the original question. The challenge of this task lies in how to accurately extract key information and refine it, avoiding over-simplification or omission of key information.
[0135] 2) Medical Radiology Report Generation (RRS): The radiology report summary aims to condense detailed examination analysis and results into a concise impression section, capturing the most prominent and actionable information in the study. The present invention uses the MIMIC-III radiology report dataset, which includes 7 anatomical regions (head, abdomen, chest, spine, neck, sinuses, pelvis) and two modalities (magnetic resonance imaging MRI and computed tomography CT). The data comes from patients in the intensive care unit of the Sheba Medical Center in Israel between 2001 and 2012.
[0136] 3) Doctor-Patient Conversation Summary (IMCS-V2): The goal of this task is to automatically generate a diagnosis and treatment report for doctor-patient conversations. The present invention uses the IMCS-V2 dataset constructed by the School of Big Data of Fudan University under the guidance of experts from the School of Medicine of Fudan University. This dataset covers real doctor-patient conversations with multi-level manual annotations and is applicable to tasks such as named entity recognition (NER), dialogue act analysis (DAC), and medical report generation (MRG). The generated medical reports cover six main aspects: chief complaint, present illness history, auxiliary examinations, past medical history, diagnosis, and suggestions. The statistical results of the dataset are shown in Table 4.1.
[0137] Table 4.1 Dataset Statistical Information
[0138]
[0139] The quality of the generated summary is measured by traditional summary evaluation metrics. The following metrics are mainly used:
[0140] 1) Rouge-1 is based on word-level overlap and evaluates the unigram overlap between the generated summary and the reference summary, focusing on measuring precision and recall.
[0141] 2) Rouge-2 evaluates the bigram overlap between the generated summary and the reference summary, calculates the matching degree of bigram phrases, and thus measures text similarity.
[0142] 3) ROUGE-L evaluates text similarity based on the longest common subsequence (LCS), considering precision and recall.
[0143] 4) BERTScore calculates the semantic similarity between the generated text and the reference text using the context embedding of BERT, thus supplementing the syntactic-level evaluation.
[0144] (1) Performance of Different Base Model Fine-tuning
[0145] The present invention performs QLoRA fine-tuning on the model by combining the Adam optimizer with a weight decay repair function. In the health problem summary task, the present invention selects three advanced open-source large models for comparison: Llama-3.1-7B-Instruct, Qwen2.5-7B-Instruct, and GLM-4-9B-Chat. The performance comparison results of the three are shown in Table 4.2, demonstrating the performance differences of each model in the task.
[0146] Table 4.2 Performance Comparison of Different Base Models in the HQS Task (Rouge-F1 as the main metric)
[0147]
[0148] (2) Spelling error detection and correction results
[0149] The present invention also fine-tunes the model for the Health Question Summary (HQS) task using the Adam optimizer with weight decay repair function and the QLoRA fine-tuning method, and presents the results of spelling error detection and correction. Table 4.3 shows the performance of the corrected model.
[0150] Table 4.3 Spelling check and correction in the HQS task
[0151]
[0152]
[0153] It can be seen from the results that after correcting the spelling errors, various indicators have been significantly improved. The correction of spelling errors helps to reduce the penalty in the string matching metric, and at the same time enhances the model's understanding of the text semantics, making it more accurate in the generation process.
[0154] (3) Model performance with entity information added
[0155] In the doctor-patient dialogue summary task, the present invention embeds entity information into the model for fine-tuning and compares the effects after adding different entities. Table 4.4 shows the performance results of the ChatGLM3 model without entity information embedded and with different entity information added. The results show that the introduction of entity information improves various evaluation indicators, proving the effectiveness of entity information in improving the prediction accuracy of the model. However, it is not the case that the more entity information, the better the model performance. Too much entity information may cause the model's attention to be scattered, affecting performance. Therefore, how to optimize the selection and use of entity information remains the key to further research.
[0156] In the medical radiology report generation (RRS) task, Table 4.5 shows the results after adding entity information in the report generation process. After introducing entity information, the recall rate has increased significantly, indicating that the model can pay attention to more key information and generate a more comprehensive summary. However, other indicators have decreased. After careful comparison, it is found that the model has paid attention to too much non-core content, resulting in too long generated text and having a negative impact on accuracy. Therefore, how to accurately screen and embed core entity information to balance the comprehensiveness and precision of the summary is the key to improving the model performance.
[0157] Table 4.4 Experimental results in the IMC-V2 doctor-patient dialogue summary task
[0158]
[0159] Table 4.5 Entity information embedding in the RRS task
[0160]
[0161] V. Biomedical Information Extraction Task
[0162] For the biological information extraction task module, the present invention mainly adopts a method of fine-tuning based on a large model for Chinese electronic medical record named entity recognition (CNER), which mainly includes the following steps:
[0163] In the instruction prompt design stage, three different types of instruction prompts are designed: simple instructions, task descriptions and type definitions, and task descriptions, type definitions and examples.
[0164] In the decoding scheme design stage, three decoding schemes are proposed: HTML scheme, index scheme and special symbol marking scheme.
[0165] In the instruction fine-tuning stage, based on the designed instruction prompts and decoding schemes, an adapted data set is constructed, and the large model is fine-tuned, and finally efficient named entity recognition of Chinese electronic medical records is achieved.
[0166] The specific implementation technical solutions and experimental results in the instruction prompt design and decoding scheme design stages are as follows.
[0167] 5.1 Instruction Prompt Design
[0168] When fine-tuning a large model, the instruction design in the data is crucial because it directly affects the model's performance and its ability to understand and execute tasks. Instructions, as the bridge between the model and the user or environment, carry task definitions, expected outputs, and relevant context information. The present invention adopts a progressive exploration step to explore the impact of different styles of instructions on the model's understanding ability in order to find the best instruction format. The exploration steps are divided into three key stages: simple instructions; task descriptions and type definitions; task descriptions, type definitions and examples. The complexity of these three design methods increases gradually.
[0169] First of all, the simple instruction only contains the sentence to be predicted, as Figure 9 shown. In this design, only the sentence to be predicted is input, aiming to compare with the subsequent complex instruction design, so as to analyze whether different instructions have an impact on the performance of the large model for named entity recognition and the degree of the impact.
[0170] Furthermore, in the task description and type definition instruction design, in addition to inputting the sentence to be predicted, the task requirements and entity type definitions are provided, and the specific definitions of the entities are clarified. The present invention helps the model to clarify the key points of the task in this way, guiding it to focus on the key information, so as to better understand the entity features and context relationships. As Figure 10 shows the format of the instruction as task description and type definition.
[0171] Furthermore, in addition to the task requirements and entity type definitions, the present invention also helps the model better understand the task by providing examples similar to the sentence to be predicted. Figure 11 It shows the instruction design format including task descriptions, type definitions, and examples. These examples are designed through scripts and use a vector database to retrieve the most similar content to the input text from the dataset to help the large model better identify entities and improve the accuracy and efficiency of named entity recognition.
[0172] 5.2 Decoding Scheme Design
[0173] In the Chinese Named Entity Recognition (CNER) task, designing an appropriate decoding scheme is one of the key issues. The goal of the CNER task is to identify specified entity information (such as diseases, drugs, clinical manifestations, etc.) in electronic medical record texts, providing a basis for medical information extraction and decision support. However, Chinese electronic medical records have characteristics such as varying text lengths, flexible and diverse language expressions, and contain a large number of professional terms and abbreviations, which all affect the accuracy and robustness of entity recognition. At the same time, the particularity of the medical field requires the CNER task to be able to understand the context and interrelationships of entities, thus increasing the complexity of the decoding process. In addition, the present invention not only focuses on identifying entities and their types but also aims to accurately obtain the position information of these entities in the text, making the decoding task even more challenging.
[0174] The present invention explores three decoding schemes: the HTML decoding scheme, the index decoding scheme, and the special symbol marking decoding scheme. The inputs of these schemes are kept consistent, all adopting the format of task descriptions and type definitions, to explore the impact of different decoding schemes on the CNER task effect.
[0175] First, the HTML decoding scheme emulates the format of HTML files, such as Figure 12 shown. This scheme guides the model to use tags with class attributes representing entity types and entity names to highlight the named entities in the sentence. Specifically, the <type:entity> tag represents the starting position of the entity, while the < / type:entity> tag represents the ending position of the entity. The marking forms of entity types and entity names help the model clarify the start and end of the entity, thereby improving the accuracy of entity recognition.
[0176] In addition, the structure of the index decoding scheme is as Figure 13As shown. The output of this solution first shows the original sequence of the sentence and the index of each character, and then outputs the entity, entity type, and start and end positions of the entity. By outputting each character together with its corresponding index, it can help the model better understand the exact position of the entity in the sentence. In addition, considering that there may be multiple entity types in the dataset and there may be nested relationships between different entities, when there are multiple entities of the same entity type, the entities will be separated by a semicolon (;).
[0177] Finally, the structure of the special symbol marking decoding solution is as Figure 14 shown. This solution transforms the CNER task into a text generation task through special symbol marking. Specifically, the model transforms the task into generating a text sequence, where special markers (such as [ and ]) are used to mark entities. If the input text does not contain entities, the output text will be the same as the input text; if the input text contains entities, special symbol marking is used to highlight these entities. In addition, to handle multiple entities and nested relationships between entities that may appear in the dataset, multiple entities of the same entity type in the design solution will be separated by a semicolon (;).
[0178] 5.3 Experimental Results
[0179] The dataset used in the research content of this part is the CMeEE-V2 dataset. This dataset is constructed according to the requirements of the fine-tuning framework and combined with different instruction schemes and decoding schemes.
[0180] CMeEE-V2 dataset: The core goal is to accurately identify and classify key entities in medical texts. Based on a predefined schema, this task extracts entities closely related to medical clinical practice from pure medical text documents and classifies them into nine categories, which cover diseases, clinical manifestations, medical procedures, medical devices, drugs, body parts, departments, and medical test items. The CMeEE-V2 dataset not only has many entity categories but also has a large number of nested entities, making it difficult to perform named entity recognition on it.
[0181] In named entity recognition, the commonly used evaluation metrics mainly include Precision, Recall, and F1 value. This paper mainly uses Micro-F1 as the evaluation metric.
[0182] (1) Performance Impact of Different Instruction Schemes
[0183] Table 5.1 shows the performance comparison of three instruction schemes when the decoding scheme uses special symbol marking and the Chatglm3-6B-base base model. From the results of strict metrics, the design of task description and type definition has the best effect, while only inputting the sentence to be predicted has the worst effect. The possible reason for this analysis is that a clear and accurate task description can help the model identify which key information it should focus on, and adjust its internal parameters and strategies accordingly. In the named entity recognition task, this guidance is particularly important because it requires the model to accurately identify and classify entities in the text. Through the task description, the model can better understand the core requirements of this task, and then more accurately identify entities in the subsequent processing.
[0184] However, compared with task description and type definition, the instruction method of adding examples additionally increases the workload but does not bring better results. This may be due to the particularity of the named entity recognition task. In the named entity recognition task, the model pays more attention to entity-level information rather than sentence-level input text. Therefore, the examples found at the sentence level may not exactly match the entities in the task sentence, which may instead have a negative impact on the model's learning. In addition, too many examples may also increase the computational burden on the model and reduce its processing speed.
[0185] Table 5.1 Performance Comparison of Different Instruction Schemes
[0186]
[0187] (2) Performance Impact of Different Decoding Schemes
[0188] On the basis of using task description and type definition as instruction prompts simultaneously, the performance of three different decoding schemes on the CMeEE-V2 dataset was compared, and the results are presented in Table 5.2 below. It can be observed from the results that from the perspective of strict evaluation metrics, the special symbol marking decoding scheme stands out among the three methods, showing significant advantages. In contrast, the HTML decoding scheme is slightly inferior, and its performance is not satisfactory.
[0189] Table 5.2 Performance Comparison of Different Decoding Schemes
[0190]
[0191] In response to this discovery, this paper analyzes that when the HTML decoding scheme processes nested entities, due to the complexity of its output structure, it may bring additional challenges to the model's learning process. When there is a nested relationship between entities, the HTML scheme tends to generate a more cumbersome and complex output, and this complexity undoubtedly increases the difficulty for the model to capture entity boundaries, thereby affecting its overall annotation effect.
[0192] In contrast, the special symbol tagging decoding scheme exhibits its unique advantages. By cleverly introducing specific tagging symbols around entities, this scheme provides clearer guidance for large models when generating annotated sequences. This design not only effectively bridges the gap between sequence tagging tasks and text generation tasks, but also simplifies the tagging process to a certain extent. More importantly, the special symbol tagging decoding scheme performs excellently when dealing with nested entities. It adopts a strategy of tagging only one entity per output sentence, and this simple and clear way greatly alleviates the complexity problems brought by nested entities, making the overall tagging process more efficient and accurate.
[0193] The index decoding scheme has relatively moderate performance. Its advantage is that it requires a shorter length of inference output from the large model and less time, so it can improve the efficiency of named entity recognition while maintaining relatively good performance.
[0194] (3) Comparison of performances of different bases
[0195] The difficulty in comparing the impacts of different bases and different model sizes on the named entity recognition task of Chinese electronic medical records based on large models lies in how to find the best balance among different bases (such as Qwen, Baichuan2, Chatglm, etc.) and different model sizes to improve the accuracy and efficiency of entity recognition. The key lies in comprehensively considering experimental design, dataset selection, and evaluation metrics. On the basis of controlling other influencing factors, with the instruction using task description and type definition, and the decoding using special symbol tagging, the invention compared the Qwen, Baichuan2, and Chatglm3 bases, and the results are shown in Table 5.3 below.
[0196] Table 5.3 Comparison of performances of different bases
[0197]
[0198] In the above table experiment, the named entity recognition performance of Qwen-7B-base, Baichuan2-7B-base, and Chatglm3-6B-base on the CMeEE-V2 dataset was studied and compared when only the base was changed while other conditions remained the same. The evaluation was based on the inference results, from which it can be seen that the Chatglm3-6B-base base performed better. In this paper, the Chatglm3-6B base with better performance was selected from the bases, and then the named entity recognition performance of the base and chat versions of Chatglm3-6B on the CMeEE-V2 dataset was compared. It can be seen that the base model performed better. The reasons are analyzed as follows: First, the base model had not undergone fine-tuning for specific tasks before participating in this experiment. This "untouched" state enabled the base model to view and adapt to the data from a relatively "pure" perspective when facing a new named entity recognition task. Therefore, when this study fine-tuned the base model on the CMeEE-V2 dataset, it could more easily capture the key features in the data, thus better meeting the requirements of the named entity recognition task.
[0199] VI. Self-awareness and general domain dialogue ability
[0200] In the process of exploring the application of large language models in the biomedical field, although they perform excellently in specific tasks (such as medical Q&A, report generation, information extraction, etc.), there are still problems with insufficient self-awareness. For example, existing models often cannot accurately identify their own roles, ability boundaries, or information credibility during the dialogue process, and may output incorrect or inappropriate content. In addition, the trained biomedical large models focus on specific tasks but lack basic general dialogue ability, resulting in limited performance in cross-domain communication or comprehensive Q&A. Generally speaking, when users use standard task format data, the model can output standard answers, but if users conduct ordinary conversations, the model is difficult to return normal content.
[0201] Therefore, the present invention explores a method to enhance the self-awareness and general domain dialogue ability of large language models, enabling them to accurately understand and define their own roles while performing biomedical tasks, improving the credibility of the output content, and supporting natural interaction with users.
[0202] 1. Self-awareness mechanism: By designing instruction data for role positioning, through the Q&A method, continuously ask the model identity inquiry questions to guide the model to think about its identity. The answers set in the data are the identity information set for the model by the present invention. The self-awareness data is dispersed into all training data to ensure the accurate role positioning of the model throughout the process.
[0203] 2. General domain dialogue ability: The present invention has collected and sorted out the currently most authoritative general domain dialogue dataset MOSS, and removed noise data such as websites, garbled characters, and non-standard pseudo-code in the dataset to ensure stability during model training; at the same time, self-cognitive information in the dataset is removed. The MOSS dataset hopes to set the cognitive positioning of the model as "MOSS agent", and including such self-cognitive information will affect the identity recognition ability of the present invention. After screening, about 400,000 high-quality data are retained as the training data for the general domain dialogue ability of the model, so that the model can reply well even when receiving inquiries about non-specific tasks.
[0204] The self-cognitive data and the processed MOSS data format are shown below.
[0205] VII. Fine-tuning of Multi-task Large Language Model for Biomedical Field
[0206] Through the exploration of the fine-tuning methods for the above different biomedical tasks, the present invention finally applied the method technologies of different tasks to the existing biomedical datasets, constructed the final version of the instruction data, and performed full-parameter instruction fine-tuning training on the GLM-4-9B base to obtain the final biomedical large model.
[0207] 7.1 Training Data
[0208] To make the most of the existing biomedical text resources, the present invention has mainly integrated comprehensive open-source datasets covering English and Chinese. The data sources mainly include the following three aspects: existing English and Chinese biomedical shared task datasets, data resources already used for training existing biomedical large models, and exercise contents on traditional Chinese medicine examination websites. Specific dataset information is shown in Table 7.1.
[0209] Table 7.1 Statistics of Training Set Information
[0210]
[0211] The dataset of the present invention covers multiple tasks, such as biomedical information extraction tasks (including named entity recognition, relation extraction), medical intelligent question answering tasks (fill-in-the-blank questions), and medical report generation tasks. Among them, the datasets used for the named entity recognition task are bc5cdr-chem, bc5cdr-dis, CmeEE, chemNER, NCBIdis; the datasets used for the relation extraction task are bc5cdr, biorelex, CMEie_v2, DDI_corpus; the datasets for the medical intelligent question answering tasks include choice_med_qa_en, choice_pubmed_qa.Choice_medqa_zh; the datasets used for the medical report generation task are HQS, RRS, IMCS_V2.
[0212] In order to construct a dataset suitable for instruction fine-tuning, the present invention designs a dedicated instruction template for each task's data according to the characteristics of each task and the strategies explored above.
[0213] For the instruction template of the question answering and dialogue tasks, the original question is directly used as the input of the model, and the corresponding answer is used as the output, aiming to help the model understand the core goal of the task.
[0214] The instruction template for the named entity recognition (NER) task adopts various decoding strategies designed by the present invention. It not only uses the scheme of directly answering entities but also designs a strategy based on special markers to clearly mark the positions of specific entities in the input, so as to improve the model's ability to recognize entity positions.
[0215] Furthermore, in order to enhance the comprehensive ability of the model, the present invention additionally generates role-playing data and multi-task question answering data. For different tasks, the model is made to play specific roles, such as "experienced natural language processing expert" or "information extraction expert". At the same time, n tasks (n ∈ [2, 10]) are randomly selected to construct a multi-task question answering dataset to simulate task interactions in complex scenarios.
[0216] Even further, to improve the model's performance in self-awareness scenarios, the present invention manually constructs approximately 1400 self-awareness data to ensure that when the model receives questions like "Who are you?", it can generate accurate answers according to the set roles. And to improve the model's general question answering ability, the present invention also adds a part of the MOSS dataset as a supplement.
[0217] All instruction data formats are as shown in Figures 15 - 17 shown.
[0218] 7.2 Model Training
[0219] The present invention aims to explore the potential of large language models (LLMs) in bilingual biomedical natural language processing tasks, and selects the current advanced GLM-4-9B model as the base model for full-parameter training. The ms-swift framework is mainly used for model training, and the link is as follows: https: / / github.com / modelscope / ms-swift?tab=readme-ov-file. The main reasons for selecting GLM-4-9B as the pre-trained model include the following points: 1) Excellent performance in bilingual tasks: GLM-4-9B performs outstandingly in both Chinese and English tasks and can meet the diverse requirements of bilingual biomedical tasks. 2) Moderate model size and high training efficiency: GLM-4-9B has a moderate parameter scale, has high training efficiency while ensuring task performance, and meets the limitations of existing computing resources. 3) Wide coverage of pre-training data: The pre-training dataset of GLM-4-9B covers the general data required for multiple tasks, providing a solid foundation for cross-task applications. Through training based on GLM-4-9B, it is expected to fully explore its potential capabilities in the bilingual biomedical field. Specific training parameters are shown in Table 7.2.
[0220] Table 7.2 Training Parameter Settings
[0221]
[0222] The following is an explanation of the role of each hyperparameter:
[0223] 1. Learning_rate (learning rate): The learning rate determines the magnitude of the adjustment of the model during each parameter update. A smaller learning rate (such as 1e-05) is usually used to avoid too fast convergence or oscillation, helping the model to find the optimal solution more stably. A small learning rate is suitable for training processes that require fine-tuning.
[0224] 2. Gradient_accumulation_steps (gradient accumulation steps): Gradient accumulation means calculating gradients on multiple small batches (mini-batch), but not immediately updating the model weights. Instead, after accumulating a certain number of small batches, a single weight update is performed. By setting it to 16, it means that gradients are accumulated every 16 small batches, and then a single weight update is performed. This can effectively use the video memory, and at the same time increase the batch size and improve the model performance.
[0225] 3. Batch_size (batch size): The batch size determines the number of training samples used in each iteration. A batch size of 1 is usually used for settings that require less video memory, or for tasks with longer sequence lengths (for example, long texts in NLP tasks). Although the training process will be slower, it can reduce the requirements for hardware resources.
[0226] 4. dtype (Data Type): The data type specifies how numerical values are represented. bf16 (bfloat16) is a floating-point representation format commonly used in deep learning to accelerate calculations and reduce memory usage. Compared to the standard float32, bf16 can maintain higher computational precision while reducing the demand for video memory, making it a common setting when training large models.
[0227] 5. seed (Random Seed): The random seed is used to initialize the random number generator, ensuring the repeatability of experiments. By fixing the seed, operations such as initialization and data shuffling during each training will yield the same results, facilitating the comparison and analysis of experimental results.
[0228] The settings of these hyperparameters take into account training stability, computational efficiency, and resource limitations, aiming to optimize performance and effects during the training process.
[0229] The final model training tasks include biomedical information extraction tasks (including named entity recognition NER, relation extraction RE), medical intelligent question answering tasks (fill-in-the-blank question answering QA), and medical report generation tasks. The experimental results on the medical information extraction tasks and medical intelligent question answering tasks are as follows:
[0230]
[0231] The test results of the model on the test set of the medical report generation task after training are as follows:
[0232]
[0233] The tasks of self-awareness and general domain dialogue are mainly used to stimulate the model's cognitive ability and general dialogue ability. Therefore, quantitative experimental analysis of performance is not conducted, but qualitative performance analysis is carried out through artificial dialogue tests. Examples of dialogue tests are as Figure 18 shown. It can be seen that the present invention has excellent performance in identity recognition and general domain dialogue ability.
[0234] VIII. Conclusions and Outlook
[0235] Through in-depth research on the key technologies of large language models for the biomedical field, the present invention has achieved remarkable results, successfully constructing a multi-lingual, multi-task biomedical large model with excellent performance, covering multiple tasks such as multi-lingual biomedical intelligent question answering, doctor-patient dialogue, medical report generation, and biomedical information extraction.
[0236] Through technical means such as in-context learning, instruction tuning, and chain-of-thought methods, the present invention effectively improves the answer quality and generalization ability of the model in Chinese medical question-and-answer tasks. In terms of medical report generation tasks, the present invention adopts optimization methods such as named entity information embedding and spelling correction, significantly improving the abstract generation performance of the model. At the same time, in terms of biomedical information extraction tasks, different instructions and decoding strategies based on large models are explored, which can not only accurately predict entity mentions but also mark the specific positions of entities in the text.
[0237] Finally, the present invention integrates a large number of open-source Chinese and English datasets covering various tasks, and through carefully designed instruction templates and decoding strategies, further enhances the generalization ability and entity position recognition ability of the model. In addition, the present invention also constructs role-playing data, multi-task question-and-answer data, and self-awareness data to further improve the comprehensive ability of the model.
[0238] Looking ahead, the present invention will continue to optimize and expand the application scope of biomedical large models, explore innovative technical methods to address the growing biomedical text data and increasingly complex application requirements. The present invention hopes to cooperate deeply with medical institutions to apply the model to the actual clinical environment and evaluate its effectiveness and feasibility in the real world. At the same time, during the model development and application process, ethical norms and privacy protection standards will be strictly followed to ensure the security and privacy of patient data and strengthen the security of the model. Through continuous technological innovation and interdisciplinary cooperation, the present invention expects to promote the development of biomedical large model technology and provide more powerful language processing support for biomedical research and clinical practice.
[0239] It should be noted that the terms "first", "second", etc. in the description, claims, and above-mentioned drawings of the present invention are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
Claims
1. A multi-task large language model training method for the biomedical field, characterized in that: The model is used to automatically select and execute medical intelligent question answering, medical report generation or biomedical information extraction tasks according to the task type of the input question; The method comprises Constructing a training set of an instruction data set for a medical intelligent question-answering task, wherein the instruction data set data has a first instruction format, and the first instruction format includes a data instruction format based on context learning and a data instruction format based on a thinking chain; Constructing a training set of a second instruction data set for a medical report generation task, wherein the training set data of the second instruction data set has a second instruction format, and the second instruction format includes a data instruction format enhanced based on entity information embedding and a data instruction format enhanced based on spelling correction; Constructing a training set of a third instruction data set for a biomedical information extraction task, wherein the training set data of the third instruction data set has a third instruction format, and the third instruction format includes a data instruction format based on a task description and a type definition and a data instruction format based on a special symbol mark; The large language model is trained using the training set.
2. The multi-task large language model training method for the biomedical field according to claim 1, characterized in that: It also includes building a MOSS dataset for general domain dialogues, which is used to automatically select and perform general domain dialogues based on the task type of the input question.
3. The multi-task large language model training method for the biomedical field according to claim 1 or 2, characterized in that: It also includes constructing a model self-cognition data set, wherein the model self-cognition data is dispersedly inserted into each of the training sets.
4. The multi-task large language model training method for the biomedical field according to claim 1, characterized in that: in, The large language base model is a GLM-4-9B model, wherein the hyper parameters of the large language model are as follows: Learning rate: 1e-05 Gradient accumulation steps (Gradient_accumulation_steps): 16 Batch size: 1 Data type (dtype): bf16 Random seeds:
42.
5. The multi-task large language model training method for the biomedical field according to claim 1, characterized in that: in, Construct a training set of instruction datasets for medical intelligent question answering tasks, including constructing a first indication, a piece of data in a training set of a first data set, or the first indication, a piece of data in a training set of a first data set, and a selected minority of samples from the first data set as a first training data in a training set, and obtaining each piece of first training data corresponding to each piece of data in the training set of the first data set, wherein the first training data is data in a context-based data instruction format; Using the large language model, an inference corresponding to a piece of data in the training set of the first data set is generated through a second indication, the second indication, the piece of data and the inference are constructed into a piece of data in the training set as a second training data in the training set, and each piece of second training data corresponding to each piece of data in the training set of the first data set is obtained, and the second training data is data in a data instruction format based on a thinking chain.
6. The multi-task large language model training method for the biomedical field according to claim 5, characterized in that: in: The content of the first instruction includes "format description of answering with "answer": option"; The content of the second instruction includes "demonstrating the reasoning process".
7. The multi-task large language model training method for the biomedical field according to claim 1, characterized in that: in, Construct a training set of the second instruction dataset for the medical report generation task, including Generate an entity category in a piece of data in the training set of the second data set, construct the piece of data and the entity category into a piece of third training data in the training set, and obtain each piece of third training data corresponding to each piece of data in the training set of the second data set, wherein the third training data is data in a data instruction format based on entity information embedding; Using the large language model, a piece of data after the spelling error of a piece of data in the training set of the second data set is corrected is generated through a fourth indication, the fourth indication, the piece of data and the piece of data after the spelling error is corrected are constructed as a fourth training data in the training set, and each piece of fourth training data corresponding to each piece of data in the training set of the second data set is obtained, and the fourth training data is data in a data instruction format enhanced based on spelling correction.
8. The multi-task large language model training method for the biomedical field according to claim 7, characterized in that: in: Generating an entity category in a piece of data in a training set of the second data set, comprising extracting the entity category using a medical entity extraction tool; The content of the fourth indication includes "spelling error correction".
9. The multi-task large language model training method for the biomedical field according to claim 1 or 7, characterized in that: in, Construct a training set of the third instruction dataset for biomedical information extraction tasks, including The fifth instruction and a piece of data in the training set of the third data set are constructed as a fifth training data in the training set, and each piece of fifth training data corresponding to each piece of data in the training set of the third data set is obtained, wherein the fifth training data is data in a data instruction format based on a task description and a type definition; wherein the content of the fifth instruction includes a task description and a type definition; Generate multiple text sequences of a piece of data in the training set of the third data set, and each text sequence has an entity marked by a special symbol, the piece of data and the text sequence are constructed as a sixth training data of the training set, and obtain each sixth training data corresponding to each piece of data in the training set of the third data set, and the sixth training data is data in a data instruction format based on special symbols.
10. An electronic device, comprising: One or more processors, a memory, and one or more programs; wherein the one or more programs are stored in the memory, and the one or more programs include instructions, which, when executed by the electronic device, enable the electronic device to execute any of the methods described in claims 1-6, 8.
Citation Information
Patent Citations
Outpatient service electronic medical record generation method based on Chinese medical big model
CN117253576A
Follow-up visit data acquisition method and system based on large language model and knowledge distillation
CN118352097A
Secondary training method of evaluation report entity extraction model and related equipment
CN118821733A
Model training method and device, equipment, storage medium and program product
CN119250153A
Visual language large model-based livestock and poultry parasitic disease diagnosis and treatment method
CN119274778A
Cited By
Webpage agent adaptive learning method and system based on self-cognition exploration
CN121257635A
Drug combination patent intelligent identification method and device based on large language model
CN121682540A