Virtual diabetes mellitus standardized patient inquiry system and method based on large language model
Through the virtual diabetes standardized patient consultation system based on large language model, the problems of high cost, low efficiency and poor consistency of traditional standardized patients are solved, and efficient, convenient and consistent simulated consultation training for diabetes patients are achieved, improving the clinical skills training effect of medical students.
Patent Information
- Application Number
- CN202510135623.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional standardized patients are played by actors, which are costly, inefficient, and difficult to ensure consistency, cannot effectively respond to complex or non-standardized questions, and have weak contextual understanding.
A virtual diabetes standardized patient consultation system based on a large language model is adopted. Through user interaction modules, business processing modules, dialogue generation modules and data access modules, it simulates the symptoms, medical history and answers of real diabetes patients, and provides an efficient, convenient and consistent simulated patient solution.
It simulates the symptoms, medical history and answers of real diabetes patients for consultation and training, which improves the efficiency and consistency of clinical skills training and reduces costs.
Smart Images

Figure CN120108784A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical consultation systems, and more specifically, to a virtual diabetes standardized patient consultation system and method based on a large language model. Background Art
[0002] Diabetes is a common chronic disease that requires long-term management and treatment. In medical education and training, Standardized Patients (SP) are widely used to simulate real patients and help medical students and doctors improve their clinical skills. Traditional standardized patients are played by actors, which is costly, inefficient, and difficult to ensure consistency.
[0003] With the advancement of technology, informatization and artificial intelligence (AI) have gradually been introduced into the field of medical education to alleviate the limitations of real-life SPs. Traditional informatization methods mainly implement the functions of SPs through rule-driven simulation systems or templated dialogue systems. For example, virtual standardized patients (VSPs) based on simple scripts and programs can partially replace real-life SPs and provide fixed medical histories and physical signs. However, such systems are unable to cope with complex or non-standardized questions and have weak context understanding capabilities.
[0004] Large Language Model (LLM) performs well in natural language processing tasks and can generate fluent and natural text. Applying it to the virtual SP system can effectively solve the shortcomings of traditional standardized patients and provide an efficient, convenient and consistent simulated patient solution. Summary of the invention
[0005] In view of the technical problems existing in the prior art, the present invention provides a virtual diabetes standardized patient interview system and method based on a large language model. The system can simulate the symptoms, medical history and answers of real diabetes patients for doctors or medical students to conduct interview training.
[0006] According to a first aspect of the present invention, a virtual diabetes standardized patient consultation system based on a large language model is provided, comprising:
[0007] A user interaction module provides an intuitive and concise graphical interface through which users interact with the system;
[0008] The business processing module is responsible for managing the user registration and login process, including: strictly verifying the validity of the personal information entered during user registration; after the information is verified, encrypting the user's data and storing it in the database;
[0009] The dialogue generation module transmits natural language through text or voice and converts it into a standard input format before transmitting it to the system;
[0010] The data access module collects and saves the conversation records between users and the system for continuous optimization of the model's question-answering generation effect, and regularly evaluates the system's performance in clinical consultation simulations to ensure that the system can meet learning needs.
[0011] Based on the above technical solution, the present invention can also make the following improvements.
[0012] Optionally, the user interaction module provides a user-friendly consultation interface, including receiving user input and displaying system output, and the receiving user input supports text and voice input methods.
[0013] Optionally, the input of the user interaction module in the use phase is: text / voice input from the user, response content from the business processing module;
[0014] The output of the user interaction module in the use phase is: sending a user request to the business processing module, sending a conversation content update request or acquisition request to the data access module, feeding back the conversation content to the user, and system prompts.
[0015] Optionally, the business processing module is used to generate dialogues, generate requests to call the dialogue generation module and manage the access to user information and dialogue history, and ensure the normal operation of the system.
[0016] Optionally, the dialogue generation module receives the natural language in a standard format input by the user and returns the answer of the virtual diabetes standardized patient after processing it with a weight file, and generates a reasonable answer to the user input question based on a pre-trained large language model; in the dialogue generation module, the standardized patient role is simulated, and deep learning and reasoning of diabetes-related knowledge are performed through the large language model to form a suitable answer.
[0017] According to a second aspect of the present invention, a virtual diabetes standardized patient interview method based on a large language model is provided, comprising:
[0018] Construct a general dataset for pre-training and a standardized diabetes patient consultation dataset;
[0019] Constructing a pre-trained large language model and pre-processing the pre-trained data set; the attention mechanism of the pre-trained large language model adopts a ternary mask block algorithm, and the ternary mask block algorithm generates a ternary block mask matrix by pre-processing the mask matrix;
[0020] The pre-trained large language model is fine-tuned based on the constructed general dataset and the standardized diabetes patient consultation dataset.
[0021] Optionally, constructing a standardized diabetes patient consultation data set includes:
[0022] The collected conversations are organized into a standard Alpaca multi-turn conversation format dataset. The data preprocessing adopts the general approach of GPT. Before training, these predictions are processed by the word segmenter and then passed into the model for training;
[0023] We chose to build a Transformer-based large language model and fine-tune it to acquire medical knowledge so that it could understand patients’ symptom descriptions and medical histories.
[0024] Optionally, the standardized diabetes patient consultation dataset is represented by a diabetes patient conversation dataset collected in cooperation with an endocrinology department of a hospital, and the standardized diabetes patient consultation dataset covers background information of patients with the same disease but different symptoms.
[0025] Optionally, the building of a pre-trained large language model includes:
[0026] Collect Chinese data sets publicly available on the Internet for pre-training large language models, and collect conversations between doctors and patients to form fine-tuning data sets;
[0027] Preprocess the pre-training data set, including removing web page labels, advertising text, and garbled low-quality content;
[0028] Remove non-critical communication and courtesy language from the collected fine-tuning dataset, and anonymize all content that can reveal personally identifiable information;
[0029] The constructed large language model is trained using the pre-training dataset, and the pre-trained large language model is fine-tuned using the fine-tuning dataset to obtain a large language model for standardized diabetic patients.
[0030] Optionally, fine-tuning the pre-trained large language model based on the constructed general dataset and the standardized diabetes patient consultation dataset includes:
[0031] Select full fine-tuning and use the fine-tuning dataset to fine-tune the pre-trained large language model. In the full fine-tuning process, assuming that all parameters of the model are θ, the optimization objective is to minimize the loss function To adjust all parameters:
[0032]
[0033] Where η is the learning rate, is the loss function The gradient of the parameter θ.
[0034] Technical effects and advantages of the present invention:
[0035] The present invention provides a virtual diabetes standardized patient consultation system and method based on a large language model. The "ternary" mask block algorithm is used in the attention mechanism of the model, and the fine-tuning part is fine-tuned using a dialogue data set with a diabetic patient. The system can simulate the symptoms, medical history and answers of a real diabetic patient for doctors or medical students to conduct consultation training.
[0036] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A device module diagram of a virtual diabetes standardized patient consultation system based on a large language model provided by an embodiment of the present invention;
[0038] Figure 2 A flowchart of the steps of a virtual diabetes standardized patient consultation method based on a large language model provided by an embodiment of the present invention;
[0039] Figure 3 A schematic diagram of a large language model construction method for a virtual diabetes standardized patient consultation method based on a large language model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0040] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0041] It is understandable that based on the defects in the background technology, the embodiment of the present invention proposes a virtual diabetes standardized patient consultation system based on a large language model, specifically Figure 1 As shown, the system includes: a user interaction module, a business processing module, a dialogue generation module and a data access module; wherein,
[0042] The user interaction module is the front-end part of the system, providing an intuitive and concise graphical interface and being responsible for interacting with the user, including receiving user input, displaying system output and providing a user interface.
[0043] The user interaction module provides a user-friendly consultation interface that supports text and voice input, and the user can interact with the system efficiently through the interface;
[0044] The inputs of the user interaction module in the use phase are: inputs from users (text / voice, etc.) and response content from the business processing module.
[0045] The output of the user interaction module usage phase is: sending user requests to the business processing module, sending conversation content update or acquisition requests to the data access module, feedback to the user on the conversation content, and system prompts.
[0046] The business processing module is responsible for managing the user registration and login process, including: when the user registers, strictly verifying the validity of the personal information entered to ensure the integrity and security of the registration information; after the information passes the verification, the user's relevant data is encrypted and stored in the database to protect user privacy and data security;
[0047] The business processing module is the core part of the system. It is responsible for receiving user requests from the user interaction module, processing user login, registration, and identity authentication operations, and generating dialogue generation requests to call the dialogue generation module. At the same time, the business processing module also manages the access to user information and dialogue history, and ensures the normal operation of the system.
[0048] Inputs of the business processing module during the usage phase: user requests from the user interaction module, user information from the data access module, and conversation content;
[0049] The output of the business processing module usage phase: sending a dialogue generation request to the dialogue generation module, sending response content to the user interaction module, and sending or obtaining user information updates and dialogue content to the data storage module.
[0050] In the dialogue generation module, the user transmits natural language through text or voice and converts it into a standard input format before transmitting it to the system.
[0051] The dialogue generation module relies on a large language model that is fully fine-tuned after pre-training to generate natural, fluent and semantically accurate dialogue responses based on user input. The dialogue generation module not only focuses on the semantic matching of a single dialogue, but also deeply understands the logical relationship in the dialogue context. By introducing the context memory mechanism, the system can maintain consistency across dialogue dimensions and achieve a more personalized dialogue experience;
[0052] The dialogue generation module receives the natural language in a standard format input by the user and returns the answer of the virtual diabetes standardized patient after processing it with a weight file. The dialogue generation module is responsible for generating reasonable answers to the user input questions based on the pre-trained large language model. In this module, the standardized patient role is simulated, and the diabetes-related knowledge is deeply learned and reasoned through the large language model to form a suitable answer.
[0053] The large language model is fine-tuned in a manner to train diabetes-related consultation scenarios based on a general large language model.
[0054] Training phase: A large number of diabetes-related consultation datasets are used to fine-tune the large language model;
[0055] The input of the dialogue generation module in the use phase: dialogue generation request from the business processing module;
[0056] The output of the dialogue generation module usage phase: returns the answer content to the business processing module.
[0057] The data access module is used to collect and save the conversation records between the user and the system for continuous optimization of the question-answering generation effect of the model, improve its accuracy and fluency, and regularly evaluate the performance of the system in clinical consultation simulation to ensure that it can meet the learning needs of medical students.
[0058] The data access module is responsible for storing and querying various types of data in the system, including user information, conversation records, and system configuration, while ensuring the security and accessibility of the data. It ensures efficient storage, reading, filtering, and deletion of data, and manages the persistence of user information, conversation data, and other system-related data.
[0059] Training phase: Stores training data and preliminary conversation content to support the model's learning process.
[0060] Inputs of the data access module during the use phase: storage or query requests from the business processing module, and conversation content update or acquisition requests from the user interaction module;
[0061] The output of the data access module in the use phase: sending storage or query results to the business processing module and sending conversation content to the user interaction module.
[0062] According to the second aspect of the present invention, the embodiment of the present invention proposes a virtual diabetes standardized patient consultation method based on a large language model, specifically as follows: Figure 2 As shown, the method comprises the following steps:
[0063] Construct a general dataset for pre-training and a standardized diabetes patient consultation dataset;
[0064] Constructing a pre-trained large language model and pre-processing the pre-trained data set; the attention mechanism of the pre-trained large language model adopts a ternary mask block algorithm, and the ternary mask block algorithm generates a ternary block mask matrix by pre-processing the mask matrix;
[0065] The pre-trained large language model is fine-tuned based on the constructed general dataset and the standardized diabetes patient consultation dataset.
[0066] Specifically, the construction of a standardized diabetes patient consultation data set includes the following steps:
[0067] The collected conversations are organized into a standard Alpaca multi-turn conversation format dataset. Data preprocessing involves processing the expected data through a word segmenter before training and then passing it into the model for training.
[0068] We chose to build a large language model based on Transformer and fine-tune it to have medical knowledge and understand patients’ symptom descriptions and medical histories.
[0069] The diabetes standardized patient interview dataset is a diabetes patient dialogue dataset collected in cooperation with the endocrinology department of the Municipal Central Hospital. A diverse diabetes medical dialogue dataset is constructed, covering the background information of patients with the same disease but different symptoms, to ensure the diversity and comprehensiveness of the data.
[0070] It should be noted that the large language model proposes an improved "ternary" mask block algorithm in the attention mechanism part. The goal of this method is to generate a ternary block mask matrix by preprocessing the mask matrix, and only involve non-zero blocks in the mask in the calculation, thereby greatly reducing unnecessary calculations and memory accesses. Figure 3 As shown, building a pre-trained large language model includes:
[0071] Data acquisition: Public Chinese data sets collected on the Internet are used to pre-train the large language model, and then the conversations between doctors and patients are collected in cooperation with the endocrinology department of the hospital to be compiled into a fine-tuning data set.
[0072] The acquired data was preprocessed: the pre-training dataset was preprocessed to remove low-quality content such as HTML tags, advertising texts, and garbled characters; the fine-tuning dataset was then processed to remove non-critical communications and courtesy language. Secondly, in order to comply with ethical standards and protect patient privacy, all content that may reveal personal identity information (such as real name, address, contact information, etc.) was anonymized.
[0073] Pre-trained large language model: Use pre-training datasets to train self-built large language models.
[0074] Fine-tune the large language model: Fine-tune the pre-trained large language model using the fine-tuning dataset.
[0075] Large language model of standardized patients with diabetes: A large language model of standardized patients with diabetes can be obtained by fine-tuning the large language model.
[0076] The large language model is fine-tuned. Based on the general large language model, it is trained for diabetes-related consultation scenarios. Since the size of the pre-trained model is small and our dialogue system requires more accurate and effective answers, we choose full fine-tuning.
[0077] The fine-tuning of the pre-trained large language model based on the constructed general dataset and the standardized diabetes patient consultation dataset specifically includes:
[0078] In the full-scale fine-tuning process, assuming that all parameters of the model are θ, the optimization objective is to minimize the loss function To adjust all parameters:
[0079]
[0080] Among them, η is the learning rate (Learning Rate), is the loss function The gradient of the parameter θ.
[0081] Next, according to a specific embodiment, the “ternary” mask block algorithm used in the attention mechanism of the large language model is further described in detail.
[0082] The experimental operating system selected in the embodiment of the present invention is Linux Ubuntu 20.04, the processor is Intel(R) Xeon(R) Platinum 8352V CPU@2.10GHz, the memory size is 80GB, the GPU model is RTX 4090, and the video memory size is 24GB. The Python version used in the experiment is 3.8, the CUDA version is 11.8, and the PyTorch version is 2.0.0. The system deployment uses cloud servers to ensure high performance and high availability.
[0083] The common pre-training datasets include Chinese Wikipedia data, Chinese BaiduBaiKe data, C4_zh data, and Chinese Wudao open source data.
[0084] The pre-trained large language model of the embodiment of the present invention uses the open source ChatGLM2-6B word segmenter.
[0085] The parameters of the pre-trained large language model are shown in Table 1.
[0086] Table 1 shows the model pre-training parameter settings
[0087]
[0088] During the training phase, the pre-trained large language model is fine-tuned to optimize the model's ability to generate responses by using diabetes-related consultation data, standardized patient conversation data, etc. Through fine-tuning, the system is better able to return responses from virtual diabetes standardized patients.
[0089] During the use phase, the user inputs the diagnosis question through the user interaction module, the business processing module receives the request and calls the dialogue generation module to generate the answer, which is then sent back to the user interaction module by the business processing module for display. At the same time, the system will update or query the user information and dialogue records through the data access module.
[0090] The pre-trained large language model adopts the full fine-tuning method. The model training occupies 18GB of video memory. The training parameters are shown in Table 2.
[0091] Table 2 shows the model fine-tuning parameter settings;
[0092]
[0093] First, pre-training is performed through public Chinese data sets such as Chinese Wikipedia data, Chinese BaiduBaike data, C4_zh data, and Chinese Wudao open source data to acquire extensive language understanding capabilities.
[0094] Subsequently, fine-tuning training was performed by adding a medical dialogue dataset collected in cooperation with the Endocrinology Department of Xiangyang Central Hospital.
[0095] Secondly, the attention mechanism of the pre-trained large language model uses a ternary mask block algorithm to split the mask matrix into blocks of size B r ×B c of small pieces, where B r ×B c are the sizes of rows and columns of blocks respectively. For each block, the state of the block is determined according to the non-zero values in the mask matrix:
[0096] 0: This block skips calculation completely and does not participate in any calculation;
[0097] 1: This block performs standard calculations and is processed according to the original attention calculation formula;
[0098] 2: This block simplifies the calculation and reduces the computational complexity.
[0099] For each i,j block, if there are non-zero values in the block and not all are one, the block is marked as 1, indicating that the standard Flash-Attention calculation needs to be performed. If the elements in the block are all ones, the Softmax can be replaced by the low-rank approximation method. If the elements in the block are all zero, the calculation is skipped and does not participate in the calculation formula processing.
[0100] Assume there is a 4×4 mask matrix, which is divided into 2×2 blocks. The mask matrix (N=4) is as follows:
[0101]
[0102] Divide the mask matrix into 2×2 blocks to get the following small blocks:
[0103] Block 1: The status code is 1
[0104] Block 2: The status code is 1
[0105] Block 3: The status code is 0
[0106] Block 4: The status code is 2
[0107] According to the above rules, the generated ternary mask matrix block is
[0108] Specifically, given an N×N mask matrix and a block size B r and B c , d k is the sequence length or the dimension of the original mask. We generate a "ternary" mask matrix TripleBlockMask as an indicator, whose dimension is The definition is as follows:
[0109]
[0110] When calculating attention, the generated "ternary" block mask matrix (TripleBlockMask) can be used to skip the block whose corresponding value is zero, thereby reducing invalid calculations. Specifically, each block in the query matrix Q, key matrix K, value matrix V, and mask matrix mask is checked, and flash-attention calculations are performed on blocks whose corresponding entries in the "ternary" mask matrix are 1, and simplified calculations are performed on blocks whose corresponding entries are 2.
[0111] When the state of the block corresponding to the "ternary" mask is 2, a low-rank approximation method is used to replace the original standard operation, where
[0112] B r Indicates the block size in row direction.
[0113] B c Indicates the block size in column direction.
[0114] R represents an upper triangular matrix.
[0115] Q block Represents the current block of the query matrix, size B r ×d k .
[0116] K block Represents the current block of the key matrix, size B c ×d k .
[0117] Step 1: Calculate the dot product of the query matrix and the key matrix:
[0118] First, for the query matrix Q and the key matrix K, calculate their dot product and scale them:
[0119]
[0120] Step:2: Apply low-rank approximation (randomized SVD acceleration):
[0121] 1. Random Projection:
[0122] Generate a random Gaussian matrix Ω.
[0123]
[0124] Among them, Y is a low-dimensional approximate matrix.
[0125] 2.QR decomposition:
[0126] Perform QR decomposition on the matrix Y to obtain the orthogonal basis Q Y and the upper triangular matrix R
[0127] Q Y ,R=QR(Y),where R∈R k×k Among them, Q Y is the low-dimensional representation of the query, and R represents the remaining part.
[0128] 3. Projection to low-dimensional space
[0129] Then, the original score matrix S is calculated ij Projection into low-dimensional space.
[0130]
[0131] 4.SVD decomposition:
[0132]
[0133] 5. Reconstruct the approximate matrix:
[0134]
[0135] Step 3: Speed up Softmax calculation:
[0136] Exploiting low-rank structures in
[0137] 1. Precompute the scaling matrix:
[0138]
[0139] 2. Decomposition of Softmax:
[0140]
[0141] 3. Normalize by row:
[0142]
[0143] Step 4: Calculate the weighted sum:
[0144]
[0145] For blocks with state 1 in the ternary mask, the standard formula for calculating self-attention is executed:
[0146]
[0147] In this way, only the valid blocks in the mask are calculated, which greatly reduces the amount of calculation and memory access overhead. Especially when the mask matrix is very sparse, this optimization will significantly improve performance.
[0148] Therefore, the virtual diabetes standardized patient consultation system and method based on the large language model described in the embodiment of the present invention adopts the "ternary" mask block algorithm in the attention mechanism part of the large language model, and the fine-tuning part uses the dialogue data set with the diabetic patient for fine-tuning. This system can simulate the symptoms, medical history and answers of real diabetic patients for doctors or medical students to conduct consultation training.
[0149] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0150] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
[0151] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A virtual diabetes standardized patient consultation system based on a large language model, characterized by: include: A user interaction module is used to provide an intuitive and concise graphical interface through which users interact with the system; The business processing module is responsible for managing the user registration and login process, including: when the user registers, the validity of the personal information entered is verified; after the information is verified, the user's data is encrypted and stored in the database; The dialogue generation module transmits natural language through text or voice and converts it into a standard input format before transmitting it to the system; The data access module collects and saves the conversation records between users and the system for continuous optimization of the model's question-answering generation effect, and regularly evaluates the system's performance in clinical consultation simulations to ensure that the system can meet learning needs.
2. The virtual diabetes standardized patient consultation system based on a large language model according to claim 1, characterized in that: The user interaction module provides a user-friendly consultation interface, including receiving user input and displaying system output, wherein receiving user input supports text and voice input.
3. The virtual diabetes standardized patient consultation system based on a large language model according to claim 2 is characterized in that: The input of the user interaction module in the use phase is: text and voice input from the user, and response content from the business processing module; The output of the user interaction module in the use phase is: sending a user request to the business processing module, sending a conversation content update request or acquisition request to the data access module, feeding back the conversation content to the user, and system prompts.
4. The virtual diabetes standardized patient consultation system based on a large language model according to claim 1, characterized in that: The business processing module is used to generate dialogues, generate requests to call the dialogue generation module and manage the access of user information and dialogue history, and ensure the normal operation of the system.
5. The virtual diabetes standardized patient consultation system based on a large language model according to claim 1, characterized in that: The dialogue generation module receives the natural language in a standard format input by the user and returns the answer of the virtual diabetes standardized patient after processing it with a weight file, and generates a reasonable answer to the user input question based on the pre-trained large language model; in the dialogue generation module, the standardized patient role is simulated, and the diabetes-related knowledge is deeply learned and reasoned through the large language model to form a suitable answer.
6. A virtual diabetes standardized patient consultation method based on a large language model, used in the virtual diabetes standardized patient consultation system based on a large language model as claimed in claims 1 to 5, characterized in that: The method comprises the following steps: Construct a general dataset for pre-training and a standardized diabetes patient consultation dataset; Constructing a pre-trained large language model and pre-processing the pre-trained data set; the attention mechanism of the pre-trained large language model adopts a ternary mask block algorithm, and the ternary mask block algorithm generates a ternary block mask matrix by pre-processing the mask matrix; The pre-trained large language model is fine-tuned based on the constructed general dataset and the standardized diabetes patient consultation dataset.
7. The method for virtual diabetes standardized patient consultation based on a large language model according to claim 6, characterized in that: The construction of a standardized diabetes patient consultation data set includes: The collected conversations are organized into a standard Alpaca multi-turn conversation format dataset. Data preprocessing is done by processing the expected data through a word segmenter before training and then passing it into the model for training. We chose to build a Transformer-based large language model and fine-tune it to acquire medical knowledge so that it could understand patients’ symptom descriptions and medical histories.
8. The virtual diabetes standardized patient interview method based on a large language model according to claim 7, characterized in that: The diabetes standardized patient consultation dataset is a diabetes patient conversation dataset collected in cooperation with the endocrinology department of a hospital. The diabetes standardized patient consultation dataset covers background information of patients with the same disease but different symptoms.
9. The virtual diabetes standardized patient consultation system based on a large language model according to claim 6, characterized in that: The construction of the pre-trained large language model includes: Collect Chinese data sets publicly available on the Internet for pre-training large language models, and collect conversations between doctors and patients to form fine-tuning data sets; Preprocess the pre-training data set, including removing web page labels, advertising text, and garbled low-quality content; Remove non-critical communication and courtesy language from the collected fine-tuning dataset, and anonymize all content that can reveal personally identifiable information; The constructed large language model is trained using the pre-training dataset, and the pre-trained large language model is fine-tuned using the fine-tuning dataset to obtain a large language model for standardized diabetic patients.
10. The virtual diabetes standardized patient interview method based on a large language model according to claim 6, characterized in that: The fine-tuning of the pre-trained large language model based on the constructed general dataset and the standardized diabetes patient consultation dataset includes: Select full fine-tuning and use the fine-tuning dataset to fine-tune the pre-trained large language model. In the full fine-tuning process, assuming that all parameters of the model are θ, the optimization objective is to minimize the loss function To adjust all parameters: Where η is the learning rate, is the loss function The gradient of the parameter θ.