Large model fine adjustment method and device for medical college education scene
By fine-tuning the general large language model for medical school education scenarios, using specific data combination strategies and fine-tuning data sets, the problem of general large model performing poorly in medical education scenarios is solved, and the efficient functions of intelligent question-setting and intelligent question-answer are realized, providing comprehensive support for medical education.
Patent Information
- Application Number
- CN202510536211.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The existing general model performs poorly in medical education scenarios and cannot effectively adapt to the teaching needs of medical schools, especially in question generation and question answering.
By selecting a suitable general-purpose large language model, using the question-setting and question-answer fine-tuning data sets from medical education, combined with the general field dialogue data set, the model is supervised and fine-tuned by a variety of data combination strategies, and ultimately implementing intelligent question-setting and intelligent question-and-answer functions.
It improves the understanding and question-setting ability of general big models for medical knowledge, enhances the adaptability and practicality of the models, and provides comprehensive support and solutions for medical education.
Smart Images

Figure CN120067276A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence and medical education, and in particular to a large-scale model fine-tuning method and device for medical school education scenarios. Background Art
[0002] As knowledge continues to expand and update at an accelerating pace, medical students face the dual challenges of information overload and fragmented learning resources. Traditional learning methods are no longer able to meet their urgent need for efficient and accurate learning. At the same time, medical schools must employ modern methods to monitor and improve teaching quality in real time, ensuring that outcomes align with industry standards.
[0003] In the field of medical education, with the rapid development of information technology and artificial intelligence technology, big model technology is increasingly becoming a key force in promoting educational innovation and improving teaching quality. The current general big models, including the GPT series, Llama series, ChatGLM series and Qwen series, have performed well in many fields, but their performance in vertical fields still needs to be improved. Medical questions cover a wide range and are of different types, which cannot effectively adapt to the teaching scenarios of medical schools. Therefore, the present invention provides a fine-tuning big model method for medical school education scenarios, which uses the fine-tuned big model to realize intelligent question generation, intelligent question answering, and personalized assisted learning, etc., to provide comprehensive support and solutions for medical education. Summary of the Invention
[0004] The purpose of the present invention is to address the deficiencies of the existing technology and provide a large model fine-tuning method and device for medical school education scenarios.
[0005] The purpose of the present invention is achieved through the following technical solutions: a large model fine-tuning method for medical school education scenarios, comprising the following steps: (1) Select a suitable general large language model; (2) We obtained the question-setting fine-tuning dataset and the question-answering fine-tuning dataset from commonly used textbooks, tutorial books, exercise books, and high-quality medical record PDF files in medical school education, and collected open-source Chinese dialogue datasets in general fields to obtain the general field dialogue dataset. We combined the question-setting fine-tuning dataset, the question-answering fine-tuning dataset, and the general field dialogue dataset according to four data combination strategies to obtain the first fine-tuning dataset, the second fine-tuning dataset, the third fine-tuning dataset, and the fourth fine-tuning dataset. (3) Using the first fine-tuning dataset, the second fine-tuning dataset, the third fine-tuning dataset, and the fourth fine-tuning dataset, respectively, the selected general language model is fine-tuned to obtain the first fine-tuned large language model, the second fine-tuned large language model, the third fine-tuned large language model, and the fourth fine-tuned large language model; (4) Evaluate and iterate the first fine-tuned large language model, the second fine-tuned large language model, the third fine-tuned large language model, and the fourth fine-tuned large language model respectively, and obtain the fine-tuned large language model with the highest score as the optimal medical model, and use the data combination strategy corresponding to the fine-tuned large language model with the highest score as the optimal data combination strategy; (5) Through the optimal medical model, intelligent question-setting and intelligent question-answering functions are realized in the medical education auxiliary platform, providing comprehensive support and solutions for medical education.
[0006] Furthermore, the universal large language model includes a Qwen series, ChatGLM series, Llama series or Baichuan series large language model.
[0007] Furthermore, the step (2) specifically includes the following sub-steps: (2.1) Collect commonly used textbooks, tutorials, exercise books, and high-quality medical record PDF files for medical school education; (2.2) The collected PDF files of commonly used medical school textbooks, tutorials, exercise books, and high-quality medical records are converted into standard markdown text through the first data processing module; (2.3) Manually review the converted markdown text to ensure the correct title tree and text content; (2.4) Then, based on the reviewed markdown text, we extract question-answer pairs to obtain a fine-tuning dataset. The sub-step (2.4) specifically includes the following sub-steps: (2.4.1) Divide the reviewed markdown text into multiple data blocks according to the block strategy module to obtain a data block set; (2.4.2) Multiple calls to the open-source or closed-source general model generate a fixed number of definition explanation, medical knowledge comprehension, knowledge application, and practice evaluation questions based on the text corresponding to each data block in the data block set according to Bloom's hierarchy of cognition; The open source or closed source general large model includes ChatGPT, GPT-4, GPT-4o, Qwen2-72B-instruct or Claude 3 Opus model; (2.4.3) Sending the text corresponding to each data block and the corresponding question to an open-source or closed-source general model to generate a relevant answer corresponding to each question, thereby obtaining a first question-answer pair dataset, wherein each data item in the first question-answer pair dataset includes a question and a relevant answer corresponding to the text corresponding to the data block; (2.4.4) Input the text corresponding to each data block in the data block set into an open-source or closed-source general model, and generate relevant single-choice and multiple-choice questions based on the textbook content as a dataset for question fine-tuning; (2.5) Processing the first question-answer pair dataset according to the second data processing module to obtain a question-answer fine-tuning dataset; (2.6) Collect open-source Chinese dialogue datasets in general domains. The data contains rich instructions, multi-domain and multi-task problems, and obtain a general domain dialogue dataset; (2.7) Combine the question-setting fine-tuning dataset, the question-answering fine-tuning dataset, and the general domain dialogue dataset according to the four data combination strategies to obtain the first fine-tuning dataset, the second fine-tuning dataset, the third fine-tuning dataset, and the fourth fine-tuning dataset, specifically: The first data combination strategy is to sequentially combine the question-answering fine-tuning dataset and the question-setting fine-tuning dataset to obtain the first fine-tuning dataset; The second data combination strategy is to sequentially combine the question-answering fine-tuning dataset, the question-posing fine-tuning dataset, and the general domain dialogue dataset to obtain the second fine-tuning dataset; The third data combination strategy is to combine the question-answering fine-tuning dataset and the question-prompting fine-tuning dataset and randomly shuffle them to obtain the third fine-tuning dataset. The fourth data combination strategy is to combine the question-answering fine-tuning dataset, the question-posing fine-tuning dataset, and the general dialogue dataset, and randomly shuffle them to obtain the fourth fine-tuning dataset.
[0008] Furthermore, according to a large-scale model fine-tuning method for medical school education scenarios according to claim 3, it is characterized in that the first data processing module is used to build an open source high-quality PDF to markdown framework, open a closed-source high-quality PDF to markdown platform interface, and the closed-source high-quality PDF to markdown platform interface includes Miner-U, Textin and NoteGPT, converting PDF or image files into a standard markdown format and selecting the optimal converted markdown text.
[0009] Furthermore, the large model fine-tuning method for medical school education scenarios according to claim 3 is characterized in that the sub-step (2.4.1) specifically includes the following sub-steps: (2.4.1.1) Convert the reviewed Markdown text into a complete list of data blocks according to the third-level headings of the reviewed Markdown text in the block strategy module; (2.4.1.2) Given a fixed maximum character length threshold for a block as MaxLen; (2.4.1.3) Traverse each data block in the complete data block list in turn, first determine whether the character length of any data block is greater than MaxLen: If the character length of the data block is greater than MaxLen, divide the data block according to the segmentation mark and period until the character length of all divided data blocks is less than MaxLen; If the character length of the data block is less than MaxLen, combine the data block with similar data blocks until the character length of the combined data blocks is closest to MaxLen; (2.4.1.4) After the traversal is completed, all data blocks processed in sub-step (2.4.1.3) are sorted according to the original reading order of the reviewed markdown text to obtain a data block set.
[0010] Furthermore, the sub-step (2.5) specifically includes the following sub-steps: (2.5.1) Determine the character length of the answer to each piece of data in the first question-answer dataset using the second data processing module: If the character length of the answer is less than 50, delete the piece of data from the first question-answer dataset; If the character length of the answer is greater than 200 or the answer includes ordered or unordered list symbols, convert the answer to Markdown format using an open-source or closed-source universal model without changing the original content, thereby obtaining the converted answer, and then replace the original answer in the corresponding data with the converted answer; Otherwise, do not process the answer; After determining the character lengths of all relevant answers, a second question-answer pair dataset is obtained. (2.5.2) The second question-answer pair dataset is then converted into a question-answer fine-tuning dataset in the instruction-input-output format.
[0011] Furthermore, in step (3), supervised fine-tuning is performed by an efficient fine-tuning hyperparameter fine-tuning method, and the efficient fine-tuning hyperparameter fine-tuning method includes LoRA, QLoRA, full parameter fine-tuning, P-tuing and P-tuning V2 fine-tuning methods.
[0012] Furthermore, the step (4) specifically includes the following sub-steps: (4.1) Design four quantitative evaluation indicators for the question-generating task, namely, result parsing rate, question number accuracy, question completeness rate, and question repetition rate. Result parsing rate = number of questions generated by the fine-tuned model that can be correctly parsed into the specified format / total number of data in the question-generating evaluation set; question number accuracy = total number of correct questions in the generated results / total number of data in the question-generating evaluation set; question completeness rate = number of complete question elements in the generated results / total number of data in the question-generating evaluation set; question repetition rate = total number of repeated questions in the generated results / total number of data in the question-generating evaluation set. (4.2) Construct a medical question-answering evaluation dataset. This dataset includes textbooks for five-year and eight-year medical programs and final exam questions from schools. The dataset will be used to evaluate the medical question-answering capabilities of the first, second, third, and fourth fine-tuned large language models. (4.3) Select appropriate general-domain evaluation datasets to evaluate the general-domain capabilities of the first fine-tuned large language model, the second fine-tuned large language model, the third fine-tuned large language model, and the fourth fine-tuned large language model. The general-domain evaluation datasets include C-Eval, CMMLU, MMLU, and MT-Bench datasets. (4.4) Using an open-source or closed-source general-purpose large language model, evaluate the question-setting implications of the first, second, third, and fourth fine-tuned large language models. (4.5) The first fine-tuned large language model, the second fine-tuned large language model, the third fine-tuned large language model, and the fourth fine-tuned large language model are loaded for evaluation, and the fine-tuned large language model with the highest score is selected as the optimal medical model, and the data combination strategy corresponding to the fine-tuned large language model with the highest score is selected as the optimal data combination strategy.
[0013] The present invention also includes a large model fine-tuning device for medical school education scenarios, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used for the above-mentioned large model fine-tuning method for medical school education scenarios.
[0014] The present invention also includes a computer-readable storage medium on which a program is stored. When the program is executed by a processor, it implements the above-mentioned large model fine-tuning method for medical school education scenarios.
[0015] The beneficial effects of the present invention are: the method proposed in the present invention can help to quickly generate high-quality fine-tuning data in the field, improve the general large model's understanding of medical knowledge, improve the question-setting ability and efficiency, and improve the adaptability and practicality of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A flowchart of a large model fine-tuning method for medical school education scenarios; Figure 2 Flowchart of sub-step (2.4.1) in the embodiment; Figure 3 This is a structural diagram of a large-scale model fine-tuning device for medical school education scenarios. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to illustrate the present invention, rather than to represent all embodiments. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments of the present invention without creative work are within the scope of protection of the present invention.
[0018] Example 1: Figure 1 As shown, the present invention provides a large model fine-tuning method for medical school education scenarios, including the following steps: (1) Select a suitable general large language model. The general large language model includes the Qwen series, ChatGLM series, Llama series, or Baichuan series large language model.
[0019] (2) We obtained the question-setting fine-tuning dataset and the question-answering fine-tuning dataset from commonly used textbooks, tutorial books, exercise books, and high-quality medical record PDF files in medical school education, and collected open-source Chinese dialogue datasets in general fields to obtain the general field dialogue dataset. We combined the question-setting fine-tuning dataset, the question-answering fine-tuning dataset, and the general field dialogue dataset according to four data combination strategies to obtain the first fine-tuning dataset, the second fine-tuning dataset, the third fine-tuning dataset, and the fourth fine-tuning dataset.
[0020] The step (2) specifically includes the following sub-steps: (2.1) Data collection and conversion: Collect commonly used textbooks, tutorials, exercise books, and high-quality medical record PDF files for medical school education.
[0021] (2.2) Data preprocessing: The collected PDF files of commonly used medical school textbooks, tutorial books, exercise books, and high-quality medical records are converted into standard markdown text through the first data processing module.
[0022] The first data processing module is used to build an open source high-quality PDF to markdown framework, open a closed source high-quality PDF to markdown platform interface, and the closed source high-quality PDF to markdown platform interface includes Miner-U, Textin and NoteGPT, converting PDF or image files into a standard markdown format and selecting the optimal converted markdown text.
[0023] (2.3) Data review: Manually review the converted markdown text to ensure the correct title tree and text content.
[0024] (2.4) Then, based on the reviewed markdown text, question and answer pairs are extracted to obtain a question fine-tuning dataset.
[0025] The sub-step (2.4) specifically includes the following sub-steps: (2.4.1) According to the block strategy module, the reviewed markdown text is divided into multiple data blocks to obtain a data block set.
[0026] The sub-step (2.4.1) specifically includes the following sub-steps: (2.4.1.1) According to the block strategy module, the third-level headings of the reviewed markdown text are converted into a complete data block list based on the reviewed markdown text.
[0027] (2.4.1.2) Given a fixed maximum character length threshold for a chunk, MaxLen.
[0028] (2.4.1.3) Traverse each data block in the complete data block list in turn, and first determine whether the character length of any data block is greater than MaxLen: If the character length of the data block is greater than MaxLen, divide the data block according to the segmentation symbol and period until the character length of all divided data blocks is less than MaxLen; if the character length of the data block is less than MaxLen, combine the data block with similar data blocks until the character length of the combined data block is closest to MaxLen.
[0029] (2.4.1.4) After the traversal is completed, all data blocks processed in sub-step (2.4.1.3) are sorted according to the original reading order of the reviewed markdown text to obtain a data block set.
[0030] (2.4.2) Multiple calls to the open-source or closed-source general model are made to generate a fixed number of definition explanation, medical knowledge comprehension, knowledge application, and practice evaluation questions based on the text corresponding to each data block in the data block set according to Bloom's cognitive hierarchy.
[0031] The open source or closed source general large model has high language understanding and generation capabilities, contains rich world knowledge, and can perform well in many application scenarios, including the widely recognized GPT series models ChatGPT, GPT-4, GPT-4o, Qwen2-72B-instruct or Claude 3 Opus models.
[0032] (2.4.3) Sending the text corresponding to each data block and the corresponding question to an open-source or closed-source general large model to generate a relevant answer corresponding to each question, thereby obtaining a first question-answer pair dataset, wherein each data item in the first question-answer pair dataset includes a question and a relevant answer corresponding to the text corresponding to the data block.
[0033] (2.4.4) Input the text corresponding to each data block in the data block set into an open-source or closed-source general model, and generate relevant single-choice and multiple-choice questions based on the textbook content as a dataset for question fine-tuning.
[0034] (2.5) Question-answer pair post-processing: The first question-answer pair dataset is processed according to the second data processing module to obtain a question-answer fine-tuning dataset.
[0035] The sub-step (2.5) specifically includes the following sub-steps: (2.5.1) Data screening and format conversion: The second data processing module determines the character length of the relevant answer for each piece of data in the first question-answer pair dataset: if the character length of the relevant answer is less than 50, the data piece is deleted from the first question-answer pair dataset; if the character length of the relevant answer is greater than 200 or the relevant answer includes ordered or unordered list symbols, the format of the relevant answer is converted to markdown format using an open source or closed source general large model without changing the original content, to obtain the converted relevant answer, and then the original relevant answer in the corresponding data is replaced with the converted relevant answer; otherwise, no processing is performed.
[0036] After the character lengths of the relevant answers to all data are determined, a second question-answer pair dataset is obtained.
[0037] (2.5.2) The second question-answer pair dataset is then converted into a question-answer fine-tuning dataset in the instruction-input-output format.
[0038] (2.6) Collect general-domain dialogue datasets: Collect open-source Chinese dialogue datasets in general domains. The data contains rich instructions and multi-domain and multi-task problems. This is used to alleviate catastrophic forgetting after model fine-tuning and obtain general-domain dialogue datasets.
[0039] (2.7) Multi-strategy construction of multi-task datasets: The question-setting fine-tuning dataset, the question-answering fine-tuning dataset, and the general domain dialogue dataset are combined according to four data combination strategies to obtain the first fine-tuning dataset, the second fine-tuning dataset, the third fine-tuning dataset, and the fourth fine-tuning dataset. Specifically, The first data combination strategy is to sequentially combine the question-answering fine-tuning dataset and the question-setting fine-tuning dataset to obtain the first fine-tuning dataset; The second data combination strategy is to sequentially combine the question-answering fine-tuning dataset, the question-posing fine-tuning dataset, and the general domain dialogue dataset to obtain the second fine-tuning dataset; The third data combination strategy is to combine the question-answering fine-tuning dataset and the question-prompting fine-tuning dataset and randomly shuffle them to obtain the third fine-tuning dataset. The fourth data combination strategy is to combine the question-answering fine-tuning dataset, the question-posing fine-tuning dataset, and the general dialogue dataset, and randomly shuffle them to obtain the fourth fine-tuning dataset.
[0040] (3) The selected general language model is fine-tuned using the first fine-tuning dataset, the second fine-tuning dataset, the third fine-tuning dataset, and the fourth fine-tuning dataset, respectively, to obtain the first fine-tuned large language model, the second fine-tuned large language model, the third fine-tuned large language model, and the fourth fine-tuned large language model.
[0041] In step (3), the first fine-tuned large language model is obtained by performing supervised fine-tuning on the selected general language large model using the first fine-tuning dataset; the second fine-tuned large language model is obtained by performing supervised fine-tuning on the selected general language large model using the second fine-tuning dataset; the third fine-tuned large language model is obtained by performing supervised fine-tuning on the selected general language large model using the third fine-tuning dataset; and the fourth fine-tuned large language model is obtained by performing supervised fine-tuning on the selected general language large model using the fourth fine-tuning dataset.
[0042] In step (3), supervised fine-tuning is performed by an efficient fine-tuning hyperparameter fine-tuning method, wherein the efficient fine-tuning hyperparameter fine-tuning method includes LoRA, QLoRA, full parameter fine-tuning, P-tuning, and P-tuning V2 fine-tuning methods.
[0043] (4) Evaluate and iterate the first fine-tuned large language model, the second fine-tuned large language model, the third fine-tuned large language model, and the fourth fine-tuned large language model respectively, and obtain the fine-tuned large language model with the highest score as the optimal medical model, and use the data combination strategy corresponding to the fine-tuned large language model with the highest score as the optimal data combination strategy; The step (4) specifically includes the following sub-steps: (4.1) Four quantitative evaluation indicators are designed for the question-generating task, namely, result parsing rate, question number accuracy, question completeness rate, and question repetition rate. Result parsing rate = number of questions generated by the fine-tuned model that can be correctly parsed into the specified format / total number of data in the question-generating evaluation set; question number accuracy = total number of correct questions in the generated results / total number of data in the question-generating evaluation set; question completeness rate = number of complete question elements in the generated results / total number of data in the question-generating evaluation set; question repetition rate = total number of repeated questions in the generated results / total number of data in the question-generating evaluation set.
[0044] (4.2) Construct a medical question-answering evaluation dataset. Collect medical textbooks and final exams for five-year and eight-year medical programs as the medical question-answering evaluation dataset. Evaluate the medical question-answering capabilities of the first fine-tuned large language model, the second fine-tuned large language model, the third fine-tuned large language model, and the fourth fine-tuned large language model.
[0045] (4.3) Select appropriate general-domain evaluation datasets to evaluate the general-domain capabilities of the first fine-tuned large language model, the second fine-tuned large language model, the third fine-tuned large language model, and the fourth fine-tuned large language model. The general-domain evaluation datasets include C-Eval, CMMLU, MMLU, and MT-Bench datasets.
[0046] Examples of general domain evaluation datasets are shown in Table 1.
[0047] Table 1: Examples of general domain evaluation datasets
[0048] (4.4) Using an open-source or closed-source general-purpose large language model, evaluate the question-setting implications of the first fine-tuned large language model, the second fine-tuned large language model, the third fine-tuned large language model, and the fourth fine-tuned large language model.
[0049] (4.5) The first fine-tuned large language model, the second fine-tuned large language model, the third fine-tuned large language model, and the fourth fine-tuned large language model are loaded for evaluation, and the fine-tuned large language model with the highest score is selected as the optimal medical model, and the data combination strategy corresponding to the fine-tuned large language model with the highest score is selected as the optimal data combination strategy.
[0050] When evaluating the fine-tuned large language model, the output for generating three questions related to gastric cancer is shown in Table 2. The output for generating two multiple-choice questions about tracheitis is shown in Table 3.
[0051] Table 2: Output of generating three questions related to gastric cancer
[0052] Table 3: Output for generating two multiple-choice questions about tracheitis
[0053] (5) Through the optimal medical model, intelligent question-setting and intelligent question-answering functions are realized in the medical education auxiliary platform, providing comprehensive support and solutions for medical education.
[0054] Example 2: Step (2) specifically includes the following sub-steps: (2.1) Collect commonly used textbooks, tutorials, exercise books, and high-quality medical record PDF files for medical school education.
[0055] (2.2) The collected commonly used textbooks, tutorial books, exercise books, and high-quality medical record PDF files are converted into standard markdown text format through the first data processing module.
[0056] (2.3) Manually review the converted markdown text to ensure the correct title tree and text content.
[0057] (2.4) Then, based on the reviewed markdown text, question and answer pairs are extracted to obtain a question fine-tuning dataset.
[0058] The sub-step (2.4) specifically includes the following sub-steps: (2.4.1) According to the block strategy module, the reviewed markdown text is divided into multiple data blocks to obtain a data block set.
[0059] like Figure 2 As shown, the sub-step (2.4.1) specifically includes the following sub-steps: (2.4.1.1) According to the block strategy module, the third-level headings of the reviewed markdown text are converted into a complete data block list based on the reviewed markdown text.
[0060] (2.4.1.2) Given a fixed maximum character length threshold for a chunk, MaxLen.
[0061] (2.4.1.3) Traverse each data block in the complete data block list in turn, and first determine whether the character length of any data block is greater than MaxLen: If the character length of the data block is greater than MaxLen, divide the data block according to the segmentation symbol and period until the character length of all divided data blocks is less than MaxLen; if the character length of the data block is less than MaxLen, combine the data block with similar data blocks until the character length of the combined data block is closest to MaxLen.
[0062] (2.4.1.4) After the traversal is completed, all data blocks processed in sub-step (2.4.1.3) are sorted according to the original reading order of the reviewed markdown text to obtain a data block set.
[0063] (2.4.2) The GPT-4o model is called multiple times to generate a fixed number of definition explanation, medical knowledge comprehension, knowledge application, and practice evaluation questions based on the text corresponding to each data block in the data block set, according to Bloom's cognitive hierarchy.
[0064] (2.4.3) Send the text corresponding to each data block and the corresponding question to the GPT-4o model to generate the relevant answer corresponding to each question, thereby obtaining a first question-answer pair dataset, in which each data item in the first question-answer pair dataset includes the question and the relevant answer corresponding to the text corresponding to the data block.
[0065] The format of any generated data is: {"question":"the i-th question generated by the model","answer":"the i-th related answer generated by the model"}, where question represents the question corresponding to the text corresponding to each data block, and answer represents the related answer corresponding to the question.
[0066] (2.4.4) Input the text corresponding to each data block in the data block set into the GPT-4o model, and generate relevant single-choice and multiple-choice questions based on the textbook content as the question fine-tuning dataset.
[0067] (2.5) Processing the first question-answer pair dataset according to the second data processing module to obtain a question-answer fine-tuning dataset.
[0068] like Figure 3 As shown, the sub-step (2.5) specifically includes the following sub-steps: (2.5.1) Determine the character length of the answer to each piece of data in the first question-answer dataset using the second data processing module: If the character length of the answer is less than 50, delete the piece of data from the first question-answer dataset; If the character length of the answer is greater than 200 or the answer includes ordered or unordered list symbols, convert the answer to Markdown format using an open-source or closed-source universal model without changing the original content, thereby obtaining the converted answer, and then replace the original answer in the corresponding data with the converted answer; Otherwise, do not process the answer; After the character lengths of the relevant answers to all data are determined, a second question-answer pair dataset is obtained.
[0069] (2.5.2) The second question-answer pair dataset is then converted into a question-answer fine-tuning dataset in the instruction-input-output format.
[0070] (2.6) Collect the open-source Chinese dialogue dataset alpaca_data_zh_51k.json, which contains rich instructions, multi-domain and multi-task problems, and obtain a general-domain dialogue dataset; (2.7) The question-setting fine-tuning dataset, the question-answering fine-tuning dataset, and the general domain dialogue dataset are combined according to the four data combination strategies to obtain the first fine-tuning dataset, the second fine-tuning dataset, the third fine-tuning dataset, and the fourth fine-tuning dataset.
[0071] Example 3: This embodiment relates to a large model fine-tuning device for medical school education scenarios, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used for a large model fine-tuning method for medical school education scenarios in the above-mentioned Example 1; the device embodiment can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer.
[0072] like Figure 3 At the hardware level, the model watermark device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 Of course, in addition to software implementation, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0073] Improvements to a technology can be clearly categorized as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with technological advancements, many process flow improvements can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using physical hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can integrate a digital system onto a PLD through their own programming, eliminating the need for a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly performed using "logic compiler" software. This is similar to the software compiler used during program development. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0074] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the memory control logic. Those skilled in the art will also appreciate that, in addition to implementing the controller purely in computer-readable program code, the controller can also be implemented in the form of logic gates, switches, an application-specific integrated circuit, a programmable logic controller, an embedded microcontroller, etc. by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the means for implementing the various functions included therein can also be considered as structures within the hardware component. Alternatively, the means for implementing the various functions can be considered both a software module implementing the method and a structure within the hardware component.
[0075] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0076] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0077] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0078] The present invention may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.
[0079] Example 4: An embodiment of the present invention further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the method for fine-tuning a large model for medical school education scenarios in the above-mentioned embodiment 1 is implemented.
[0080] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A large model fine-tuning method for medical school education scenarios, characterized in that: The following steps are involved: (1) Select a suitable general large language model; (2) We obtained the question setting fine-tuning dataset and the question answering fine-tuning dataset from commonly used textbooks, tutorial books, exercise books, and high-quality medical record PDF files in medical school education, and collected open source Chinese dialogue datasets in general fields to obtain a general field dialogue dataset; The question fine-tuning dataset, the question-answering fine-tuning dataset, and the general domain dialogue dataset are combined according to four data combination strategies to obtain the first fine-tuning dataset, the second fine-tuning dataset, the third fine-tuning dataset, and the fourth fine-tuning dataset; (3) using the first fine-tuning dataset, the second fine-tuning dataset, the third fine-tuning dataset, and the fourth fine-tuning dataset to perform supervised fine-tuning on the selected general language model, respectively, to obtain a first fine-tuned large language model, a second fine-tuned large language model, a third fine-tuned large language model, and a fourth fine-tuned large language model; (4) Evaluate and iterate the first fine-tuned large language model, the second fine-tuned large language model, the third fine-tuned large language model, and the fourth fine-tuned large language model, respectively, and obtain the fine-tuned large language model with the highest score as the optimal medical model, and use the data combination strategy corresponding to the fine-tuned large language model with the highest score as the optimal data combination strategy; (5) Through the optimal medical model, intelligent question-setting and intelligent question-answering functions can be realized in the medical education auxiliary platform, providing comprehensive support and solutions for medical education.
2. According to claim 1, a large model fine-tuning method for medical school education scenarios is characterized in that: The universal large language model includes a Qwen series, a ChatGLM series, a Llama series or a Baichuan series large language model.
3. According to claim 1, a large model fine-tuning method for medical school education scenarios is characterized in that: The step (2) specifically includes the following sub-steps: (2.1) Collect commonly used textbooks, tutorials, exercise books, and high-quality medical record PDF files for medical school education; (2.2) The collected PDF files of commonly used textbooks, tutorial books, exercise books, and high-quality medical records for medical school education are converted into standard markdown text through the first data processing module; (2.3) Manually review the converted markdown text to ensure the correct title tree and text content; (2.4) Then, based on the reviewed markdown text, question and answer pairs are extracted to obtain a question fine-tuning dataset; The sub-step (2.4) specifically includes the following sub-steps: (2.4.1) Divide the reviewed markdown text into multiple data blocks according to the block strategy module to obtain a data block set; (2.4.2) Multiple calls to the open source or closed source general large model, according to Bloom's cognitive hierarchy, generate a fixed number of definition explanation, medical knowledge understanding, knowledge application, and practice evaluation questions based on the text corresponding to each data block in the data block set; The open source or closed source general large model includes ChatGPT, GPT-4, GPT-4o, Qwen2-72B-instruct or Claude 3 Opus model; (2.4.3) sending the text corresponding to each data block and the corresponding question to an open source or closed source general large model, generating a relevant answer corresponding to each question, and obtaining a first question-answer pair data set, wherein each data item in the first question-answer pair data set includes a question and a relevant answer corresponding to the text corresponding to the data block; (2.4.4) Input the text corresponding to each data block in the data block set into an open source or closed source general large model, and generate relevant single-choice questions and multiple-choice questions according to the textbook content as a dataset for question fine-tuning; (2.5) processing the first question-answer pair data set according to the second data processing module to obtain a question-answer fine-tuning data set; (2.6) Collect open-source Chinese dialogue datasets in general fields. The data contains rich instructions, multi-domain and multi-task problems, and obtain general-domain dialogue datasets; (2.7) The question-setting fine-tuning dataset, the question-answering fine-tuning dataset, and the general domain dialogue dataset are combined according to the four data combination strategies to obtain the first fine-tuning dataset, the second fine-tuning dataset, the third fine-tuning dataset, and the fourth fine-tuning dataset, which are as follows: The first data combination strategy is to sequentially combine the question-answering fine-tuning dataset and the question-posing fine-tuning dataset to obtain a first fine-tuning dataset; The second data combination strategy is to sequentially combine the question-answering fine-tuning dataset, the question-posing fine-tuning dataset, and the general domain dialogue dataset to obtain the second fine-tuning dataset; The third data combination strategy: combine the question-answer fine-tuning dataset and the question-setting fine-tuning dataset, and randomly shuffle them to obtain the third fine-tuning dataset; The fourth data combination strategy is to combine the question-answering fine-tuning dataset, the question-posing fine-tuning dataset, and the general dialogue dataset, and randomly shuffle them to obtain the fourth fine-tuning dataset.
4. According to claim 3, a large model fine-tuning method for medical school education scenarios is characterized in that: The first data processing module is used to build an open source high-quality PDF to markdown framework, open a closed source high-quality PDF to markdown platform interface, the closed source high-quality PDF to markdown platform interface includes Miner-U, Textin and NoteGPT, convert PDF or image files into standard markdown format, and select the optimal converted markdown text.
5. According to claim 3, a large model fine-tuning method for medical school education scenarios is characterized in that: The sub-step (2.4.1) specifically includes the following sub-steps: (2.4.1.1) According to the third-level headings of the reviewed markdown text in the block strategy module, convert the reviewed markdown text into a complete list of data blocks; (2.4.1.2) Given a fixed maximum character length threshold for a block as MaxLen; (2.4.1.3) Traverse each data block in the complete data block list in turn, and first determine whether the character length of any data block is greater than MaxLen: If the character length of the data block is greater than MaxLen, divide the data block according to the segmentation symbol and the period until the character length of all the divided data blocks is less than MaxLen; If the character length of the data block is less than MaxLen, combine the data block with the adjacent data blocks until the character length of the combined data block is closest to MaxLen; (2.4.1.4) After the traversal is completed, all data blocks processed in substep (2.4.1.3) are sorted according to the original reading order of the reviewed markdown text to obtain a data block set.
6. According to claim 4, a large model fine-tuning method for medical school education scenarios is characterized in that: The sub-step (2.5) specifically includes the following sub-steps: (2.5.1) Determine the character length of the relevant answer of each piece of data in the first question-answer pair data set according to the second data processing module: if the character length of the relevant answer is less than 50, delete the piece of data from the first question-answer pair data set; if the character length of the relevant answer is greater than 200 or the relevant answer includes ordered or unordered list symbols, convert the format of the relevant answer into markdown format without changing the original content using an open source or closed source general large model to obtain the converted relevant answer, and then replace the original relevant answer in the corresponding data with the converted relevant answer; otherwise, do not process; After the character lengths of the answers to all the data are determined, a second question-answer pair data set is obtained; (2.5.2) The second question-answer pair dataset is then converted into a question-answer fine-tuning dataset in the instruction-input-output format.
7. The method for fine-tuning a large model for medical school education scenarios according to claim 1 is characterized in that: In step (3), supervised fine-tuning is performed by an efficient fine-tuning hyperparameter fine-tuning method, and the efficient fine-tuning hyperparameter fine-tuning method includes LoRA, QLoRA, full parameter fine-tuning, P-tuing and P-tuning V2 fine-tuning methods.
8. The large model fine-tuning method for medical school education scenarios according to claim 1 is characterized in that: The step (4) specifically includes the following sub-steps: (4.1) Four quantitative evaluation indicators are designed for the question-setting task, namely, result parsing rate, question number accuracy, question completeness rate, and question repetition rate. Among them, result parsing rate = the number of questions generated by the fine-tuned model that can be correctly parsed into the specified format / the total number of question-setting evaluation set data; question number accuracy = the total number of correct questions in the generated results / the total number of question-setting evaluation set data; question completeness rate = the number of complete question elements in the generated results / the total number of question-setting evaluation set data; question repetition rate = the total number of repeated questions in the generated results / the total number of question-setting evaluation set data; (4.2) Construct a medical question-answering evaluation dataset, collect five-year and eight-year medical textbooks and school final exam questions as the medical question-answering evaluation dataset, and evaluate the medical question-answering capabilities of the first fine-tuned large language model, the second fine-tuned large language model, the third fine-tuned large language model, and the fourth fine-tuned large language model; (4.3) Select appropriate general domain evaluation datasets to evaluate the general domain capabilities of the first fine-tuned large language model, the second fine-tuned large language model, the third fine-tuned large language model, and the fourth fine-tuned large language model. The general domain evaluation datasets include C-Eval, CMMLU, MMLU, and MT-Bench datasets. (4.4) Using an open source or closed source general large model, evaluate the question content of the first fine-tuned large language model, the second fine-tuned large language model, the third fine-tuned large language model, and the fourth fine-tuned large language model; (4.5) The first fine-tuned large language model, the second fine-tuned large language model, the third fine-tuned large language model, and the fourth fine-tuned large language model are loaded for evaluation, and the fine-tuned large language model with the highest score is selected as the optimal medical model, and the data combination strategy corresponding to the fine-tuned large language model with the highest score is selected as the optimal data combination strategy.
9. A large model fine-tuning device for medical school education scenarios, characterized in that: It includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement a large model fine-tuning method for medical school education scenarios as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, a large model fine-tuning method for medical school education scenarios as described in any one of claims 1-8 is implemented.
Citation Information
Patent Citations
Data knowledge dual-drive intelligent medical dialogue system and method based on knowledge graph
CN117077786A
Education domain knowledge base search optimization method and device based on question generation
CN117540063A
Education large model tower type construction method based on multilevel experience learning
CN119202200A
Large model fine tuning training system
CN119443275A
Proportion determination method and device, electronic equipment and storage medium
CN119740656A
Cited By
Data set generation method and device, storage medium and electronic equipment
CN121278016A