Medical record generation model training method and device and electronic equipment
By generating simulated recording data and combining real data, the problem of insufficient training data for the medical record generation model is solved, and the diversity and coverage of the model training data is significantly improved.
Patent Information
- Application Number
- CN202411974043.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-13
AI Technical Summary
Due to the scarcity of real medical record data and recording data, the training data of the medical record generation model is insufficient, especially the pairing of high-quality, real medical record data and recording data is very scarce.
By acquiring the input medical record data, determining the target medical record data similar to it is based on the first training data set, generating simulation recording data corresponding to the input medical record data, and using the input medical record data and simulation recording data as the second training data, the medical record generation model is trained in combination with the first training data set.
The data sets that can be used for training are significantly expanded, the diversity and coverage of model training data is enhanced, and the problem of insufficient training data in traditional medical record generation model training methods is overcome.
Smart Images

Figure CN119993355A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a medical record generation model training method, device and electronic equipment. Background Art
[0002] In clinical practice, medical record data usually involves a large amount of text descriptions, covering information such as patient medical history, signs, symptoms, and test results, while the corresponding audio recording data (such as the conversation between doctors and patients, the consultation process, etc.) is an important basis for generating medical record data and medical record reports. Real medical record data is usually generated by doctors during the actual diagnosis and treatment process, and the corresponding audio recording data (such as the recording of the consultation conversation between doctors and patients) is an important reference for automatic generation of medical records with the help of medical record generation models.
[0003] However, due to issues involving patient privacy and data protection, the collection of real medical record recording data is often strictly restricted. In addition, it is also very difficult to match recording data with medical record data, because it is necessary not only to ensure the accurate transcription of the recording content, but also to ensure the one-to-one correspondence between the recording content and the medical record document. Therefore, the training data for the medical record generation model is insufficient, especially the pairing of high-quality, real medical record data and recording data is very scarce. Summary of the invention
[0004] The embodiments of the present disclosure provide a medical record generation model training method, device and electronic device, aiming to solve the problems existing in the above-mentioned background technology.
[0005] In order to solve the above technical problems, the present disclosure is implemented as follows:
[0006] In a first aspect, an embodiment of the present disclosure provides a medical record generation model training method, the method comprising:
[0007] Obtain input medical record data;
[0008] Based on a first training data set, target medical record data similar to the input medical record data is determined, wherein each sample data in the first training data set is a data pair consisting of real recording data and corresponding real medical record data;
[0009] Generate simulated recording data corresponding to the input medical record data according to the input medical record data and the target medical record data;
[0010] The input medical record data and the simulated recording data are used as second training data, and the medical record generation model is trained based on the first training data set and the second training data.
[0011] Optionally, the method further comprises:
[0012] Acquire a plurality of preset medical record content selection tasks, wherein the plurality of medical record content selection tasks are pairwise combinations of target medical record content of the input medical record data and other medical record content except the target medical record content, and the medical record content selection tasks are used to instruct the medical record generation model to generate medical record content with the highest matching degree with the target medical record content;
[0013] Determining, from a plurality of candidate medical record contents corresponding to the other medical record contents, a target candidate medical record content having the highest matching degree with the target medical record content of the input medical record data;
[0014] The target medical record content, the multiple candidate medical record contents, and the target candidate medical record content are used as third training data.
[0015] Optionally, determining target medical record data similar to the input medical record data based on the first training data set includes:
[0016] Determine the semantic similarity between the semantic vector of the input medical record data and the semantic vector of each real medical record data in the first training data set;
[0017] The real medical record data corresponding to the highest semantic similarity is determined as the target medical record data similar to the input medical record data.
[0018] Optionally, generating the simulated recording data corresponding to the input medical record data according to the input medical record data and the target medical record data includes:
[0019] Determine the target recording data corresponding to the target medical record data;
[0020] Inputting the input medical record data and the target audio recording data into a text generation model;
[0021] Based on the first text prompt word, the text generation model generates simulated recording data corresponding to the input medical record data, and the first text prompt word is used to instruct the text generation model to rewrite the content in the target recording data into the content in the input medical record data.
[0022] Optionally, determining target medical record data similar to the input medical record data based on the first training data set includes:
[0023] Splitting the real recording data in the first training data set according to different context structures to obtain a plurality of split data;
[0024] Based on the real medical record data and the multiple split data, multiple target medical record data with the highest semantic similarity to the input medical record data are determined.
[0025] Optionally, generating the simulated recording data of the input medical record data according to the input medical record data and the target medical record data includes:
[0026] Determining, from the plurality of split data, a plurality of target split data corresponding one to one to the plurality of target medical record data;
[0027] According to different context structures, the plurality of target split data are randomly combined to obtain target combination data;
[0028] Based on the input medical record data and the target combination data, simulated recording data of the input medical record data is generated.
[0029] Optionally, the real recording data in the first training data set is split according to different context structures to obtain a plurality of split data, including:
[0030] Input the real recording data in the first training data set into text to generate a large model;
[0031] Based on the second text prompt word, multiple split data are generated by the text generation model, and the second text prompt word is used to instruct the text generation model to split the real recording data into multiple split data according to different context structures, and the context structures include greeting, chatting, consulting and farewell.
[0032] Optionally, the training of the medical record generation model based on the first training data set and the first training data includes:
[0033] The first training data set, the second training data, and the third training data are mixed in proportion to obtain a training data set, wherein the third training data accounts for the highest proportion and the second training data accounts for the lowest proportion;
[0034] The medical record generation model is trained based on the training data set.
[0035] In a second aspect, an embodiment of the present disclosure provides a medical record generation model training device, the device comprising:
[0036] An acquisition module, used to acquire input medical record data;
[0037] A determination module, configured to determine target medical record data similar to the input medical record data based on a first training data set, wherein each sample data in the first training data set is a data pair consisting of real recording data and corresponding real medical record data;
[0038] A generating module, used for generating simulated recording data corresponding to the input medical record data according to the input medical record data and the target medical record data;
[0039] A training module is used to use the input medical record data and the simulated recording data as second training data, and to train a medical record generation model based on the first training data set and the second training data.
[0040] Optionally, the device further comprises:
[0041] A task acquisition module, used to acquire a plurality of preset medical record content selection tasks, wherein the plurality of medical record content selection tasks are pairwise combinations of target medical record content of the input medical record data and other medical record content except the target medical record content, and the medical record content selection tasks are used to instruct the medical record generation model to generate medical record content with the highest matching degree with the target medical record content;
[0042] A similar medical record determination module, configured to determine, from among a plurality of candidate medical record contents corresponding to the other medical record contents, a target candidate medical record content having the highest matching degree with the target medical record content of the input medical record data;
[0043] The training data determination module is used to use the target medical record content, the multiple candidate medical record contents and the target candidate medical record content as third training data.
[0044] Optionally, the determining module includes:
[0045] A first determination submodule, used to determine the semantic similarity between the semantic vector of the input medical record data and the semantic vector of each real medical record data in the first training data set;
[0046] The first determination submodule is used to determine the real medical record data corresponding to the highest semantic similarity as the target medical record data similar to the input medical record data.
[0047] Optionally, the generating module includes:
[0048] A second determination submodule is used to determine the target recording data corresponding to the target medical record data;
[0049] A first input submodule, for inputting the input medical record data and the target recording data into a text generation model;
[0050] The first generation submodule is used to generate simulated recording data corresponding to the input medical record data through the text generation model based on a first text prompt word, and the first text prompt word is used to instruct the text generation model to rewrite the content in the target recording data into the content in the input medical record data.
[0051] Optionally, the determining module includes:
[0052] A splitting submodule, used for splitting the real recording data in the first training data set according to different context structures to obtain a plurality of split data;
[0053] The third determination submodule is used to determine a plurality of target medical record data having the highest semantic similarity with the input medical record data based on the real medical record data and the plurality of split data.
[0054] Optionally, the generating module includes:
[0055] A fourth determination submodule is used to determine, from the plurality of split data, a plurality of target split data corresponding one to one to the plurality of target medical record data;
[0056] A combination submodule, used for randomly combining the plurality of target split data according to different context structures to obtain target combination data;
[0057] The second generating submodule is used to generate simulated recording data of the input medical record data based on the input medical record data and the target combination data.
[0058] Optionally, the splitting submodule includes:
[0059] An input unit, used for inputting the real recording data in the first training data set into a text to generate a large model;
[0060] A splitting unit is used to generate multiple split data based on a second text prompt word through the text generation model, wherein the second text prompt word is used to instruct the text generation model to split the real recording data into multiple split data according to different context structures, and the context structures include greeting, chatting, asking questions and saying goodbye.
[0061] Optionally, the training module includes:
[0062] a mixing submodule, configured to mix the first training data set, the second training data, and the third training data in proportion to obtain a training data set, wherein the third training data accounts for the highest proportion and the second training data accounts for the lowest proportion;
[0063] A training submodule is used to train the medical record generation model based on the training data set.
[0064] In a third aspect, an embodiment of the present disclosure provides an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, steps of a medical record generation model training method are implemented.
[0065] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:
[0066] The present disclosure overcomes the problem of insufficient data caused by the scarcity of real medical record data and recording data by generating simulated recording data. Based on the input medical record data, the corresponding simulated recording data is generated, which significantly expands the data set that can be used for training. The present disclosure provides an alternative solution that greatly enhances the diversity and coverage of model training data. By generating simulated recording data that matches the input medical record data based on the input medical record data and combining it with the existing real data, the diversity of training data can be increased without relying on a large amount of newly collected data, thus overcoming the problem of insufficient training data in the traditional medical record generation model training method. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0068] Figure 1 It is a schematic diagram of the steps of a medical record generation model training method provided by an embodiment of the present disclosure;
[0069] Figure 2 It is a structural block diagram of a medical record generation model training device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0070] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0071] As the core of the intelligent medical record generation system, the medical record generation model relies on a large amount of audio recording data and corresponding medical record data as training data to realize the automatic generation of outpatient medical records. However, the relevant technology faces several significant defects in the application process, especially in the lack of training data. First, the pairing of real medical record data and corresponding audio recording data is very scarce. In actual medical scenarios, medical record data usually contains a lot of private information, and the acquisition process of audio recording transcription data is cumbersome and costly. Therefore, obtaining a large amount of high-quality training data has become one of the bottlenecks of the intelligent medical record generation model. Although the relevant technology has made progress in the research and application of medical record generation models, it still relies on a large amount of real medical records and audio recording data as the basis, and lacks effective solutions to deal with the problem of scarce training data. Secondly, traditional data enhancement methods usually cannot solve the problem of lack of training data. Although there are some technical solutions for data enhancement, most of these methods focus on the enhancement of text data or the simple expansion of existing data, and fail to effectively expand and enhance the specific needs of audio recording and medical record data in medical record generation tasks. Especially when real data is difficult to obtain, traditional data enhancement methods cannot effectively improve the performance and stability of the model. In view of this, the present disclosure proposes a solution for enhancing the training data of the medical record generation model. Specifically, the present disclosure generates simulated recording data based on the existing medical record data, thereby constructing a richer and more diverse training data set, which can not only effectively expand the amount of training data, but also improve the effect and stability of model training.
[0072] Figure 1 FIG. 1 is a schematic diagram of the steps of a medical record generation model training method provided by an embodiment of the present disclosure. Figure 1 As shown, the method includes:
[0073] Step S101, obtaining input medical record data.
[0074] The input medical record data is the medical record description text recorded by doctors or medical staff during the consultation process, including the patient's basic information, medical history, physical signs, diagnosis and treatment plan, etc. Due to technical limitations or privacy protection, the input medical record data is not equipped with corresponding audio transcription data. The form of the input medical record data can be shown in Table 1:
[0075]
[0076]
[0077] Table 1: Example of input medical record data
[0078] The input medical record data can be obtained from the electronic medical record system used by the hospital or clinic, or can be manually entered by the doctor.
[0079] Step S102: determining target medical record data similar to the input medical record data based on a first training data set, wherein each sample data in the first training data set is a data pair consisting of real recording data and corresponding real medical record data.
[0080] Each sample data in the first training data set is composed of real audio recording data and corresponding real medical record data. Specifically, each sample contains two main parts: real medical record data and real audio recording data. Real medical record data is the patient's medical record recorded by the doctor, including the patient's basic information, chief complaint, current medical history, past history, personal history, etc. The format is similar to the input medical record data. Real audio recording data is audio data recorded during the patient's visit. After being transcribed into text through speech recognition, it can provide information such as the tone and intonation behind the medical record description. The recording usually also contains parts such as small talk, doctor's questions, and patient answers that are not related to the medical record.
[0081] When determining the target medical record data similar to the input medical record data, cosine similarity can be used as a similarity measurement method, and the similarity is measured by calculating the angle between two text vectors. The value range is [0,1]. The closer the value is to 1, the higher the similarity of the two texts. Specifically, the input medical record data and the real medical record data are converted into vector representations. The bag of words model (Bag of Words), TF-IDF or Word2Vec can be used to convert the text into vectors. The cosine similarity between the vectors of the input medical record data and the target medical record data is calculated. The larger the value, the more similar it is. Some pre-trained large language models (such as BERT, GPT, RoBERTa, etc.) can also be used to generate semantic vectors of the input medical record data and the target medical record data, and then calculate the similarity between them. Use language models such as BERT to encode the input medical record data and the real medical record data respectively, generate their vector representations (i.e., semantic vectors), and calculate the cosine similarity of the two vectors to obtain their similarity scores. The real medical record data having the greatest similarity to the input medical record is determined as the target medical record data similar to the input medical record data.
[0082] Step S103, generating simulated recording data corresponding to the input medical record data according to the input medical record data and the target medical record data.
[0083] Based on the real recording data corresponding to the target medical record data and combined with the input medical record data, simulated recording data matching the input medical record is generated. Although the real recording data corresponding to the target medical record data does not completely correspond to the description of the input medical record, its content, tone, and expression can provide a reference for the generation of simulated recording data.
[0084] The simulated recording data is not simply transcribing the input medical record data into text, but generating a simulated voice that simulates the corresponding real recording data in the target medical record data in terms of tone, intonation, speaking speed and expression. Text-to-Speech (TTS) technology can be used to convert the text content in the input medical record data into audio output. In this process, it is necessary not only to ensure the accuracy of the voice, but also to try to make the voice performance consistent with the patient's emotional state, tone and other characteristics when describing the condition. For example, if the input medical record data describes that the patient has a headache and nausea, the generated simulated recording may need to reflect the patient's pain, the tone may be relatively low, and the speaking speed may be slightly slower. If it is a doctor asking a question, the tone should be calm and concerned.
[0085] Step S104: using the input medical record data and the simulated recording data as second training data, and training a medical record generation model based on the first training data set and the second training data.
[0086] The input medical record data and the generated simulated recording data are used as the second training data, and the medical record generation model is trained in combination with the first training data set. Each sample data in the second training data consists of the input medical record data and the corresponding simulated recording data. The goal of the medical record generation model is to learn the relationship between the input medical record data (such as text) and the simulated recording data (such as audio). The medical record generation model is a multimodal learning model that aims to handle the association between multiple input modalities (such as text and audio). During training, the medical record generation model will learn the mapping from text to speech, that is, given the medical record description text, the medical record generation model can generate the corresponding audio data (simulated recording). As well as the learning of tone, intonation and emotion, the medical record generation model should not only focus on the medical content of the medical record, but also capture the patient's emotions and tone, and accurately simulate real clinical communication. During the training process, the medical record generation model needs to adjust the parameters by optimizing the loss function so that the generated medical record description text and simulated recording data are closer to the real data.
[0087] Through training, the medical record generation model can generate synthetic simulated recording data by inputting only medical record description text without actual recording, and can also generate corresponding medical record description text through a piece of real recording data.
[0088] The present disclosure overcomes the problem of insufficient data caused by the scarcity of real medical record data and recording data by generating simulated recording data. Based on the input medical record data, the corresponding simulated recording data is generated, which significantly expands the data set that can be used for training. The present disclosure provides an alternative solution that greatly enhances the diversity and coverage of model training data. By generating simulated recording data that matches the input medical record data based on the input medical record data and combining it with the existing real data, the diversity of training data can be increased without relying on a large amount of newly collected data, thus overcoming the problem of insufficient training data in the traditional medical record generation model training method.
[0089] In an optional embodiment, the method further includes:
[0090] Step S201, obtaining a plurality of preset medical record content selection tasks, wherein the plurality of medical record content selection tasks are pairwise combinations of target medical record content of the input medical record data and other medical record content except the target medical record content, and the medical record content selection tasks are used to instruct the medical record generation model to generate medical record content with the highest matching degree with the target medical record content.
[0091] The purpose of the medical record content selection task is to provide different categories of medical record data content for the medical record generation model. The medical record content selection task combines different parts of the medical record description of the input medical record data to create a content matching task to help the medical record generation model understand which contents are most closely related. The target medical record content is the medical record content matched in the input medical record data, which can be any one of the chief complaint, current medical history, past medical history, and personal history; correspondingly, other medical record content is any one of the chief complaint, current medical history, past medical history, and personal history in the input medical record data except the target medical record content.
[0092] By combining the target medical record content with other medical record content in pairs, multiple medical record content selection tasks are formed. The goal of each medical record content selection task is to let the medical record generation model learn how to judge and select which contents have the highest match, that is, which contents are most closely related, and determine the part that best matches the target medical record content. For example, assuming that the target medical record content in the input medical record data is "chief complaint" (for example, "headache for 3 days, accompanied by nausea"), then "other medical record content" can be any one of the parts of "current medical history", "past medical history" and "personal history". By combining two by two, multiple selection tasks are generated:
[0093] Task 1: The target medical record content is "chief complaint", and the other medical record content is "history of present illness";
[0094] Task 2: The target medical record content is "chief complaint" and the other medical record content is "past medical history";
[0095] Task 3: The target medical record content is “chief complaint” and the other medical record content is “personal history”;
[0096] Task 4: The target medical record content is "history of current illness", and the other medical record content is "chief complaint"; and so on.
[0097] Step S202: Determine, from among the multiple candidate medical record contents corresponding to the other medical record contents, target candidate medical record contents that have the highest matching degree with the target medical record contents of the input medical record data.
[0098] The candidate medical record content is the medical record description corresponding to other medical record content in multiple pre-acquired real medical record data. For example, when the target medical record content is "chief complaint", the other medical record content is "current medical history". In this task, multiple candidate contents are current medical history 1, current medical history 2, current medical history 3, etc.
[0099] The matching degree here can be achieved by calculating the similarity between text contents. For example, as mentioned above, the cosine similarity converts the text content (such as the target medical record content and the candidate medical record content) into a vector representation, and then calculates the cosine angle between these vectors to measure the similarity. It can also be the semantic similarity of deep learning models such as BERT, which uses a pre-trained large language model to convert the target medical record content and the candidate medical record content into semantic vectors and calculate the distance or similarity between them. Once the target candidate medical record content with the highest match is selected, the target candidate medical record content will be used as part of the training data to optimize the medical record generation model, which is equivalent to giving a "correct answer" to help the medical record generation model understand how to generate the corresponding medical record description based on the target medical record content.
[0100] Step S203: Use the target medical record content, the multiple candidate medical record contents, and the target candidate medical record content as third training data.
[0101] For example, for the following medical record content selection task: target medical record content: chief complaint ("headache for 3 days, accompanied by nausea"); candidate medical record content: current medical history 1, current medical history 2; target candidate medical record content: current medical history 2. Based on this medical record selection task, the third training data is:
[0102] The training samples are composed as follows:
[0103] {
[0104] "target_content":"Main complaint (headache for 3 days, accompanied by nausea)",
[0105] "candidate_contents":[
[0106] "Current medical history 1: Headache has lasted for 4 days, accompanied by vomiting",
[0107] "Current medical history 2: Headache for 3 days, accompanied by nausea"
[0108] ],
[0109] "target_candidate_content":"Current medical history 2: Headache for 3 days, accompanied by nausea"
[0110] }; where target_content is the target medical record content, Candidate Content is the candidate medical record content, and Target Candidate Content / Label is the target candidate medical record content.
[0111] At this point, the medical record generation model will use the third training data constructed above for training. The input is: target medical record content (headache for 3 days, accompanied by nausea), candidate medical record content (current medical history 1, current medical history 2). The label is: target candidate medical record content (current medical history 2). The target candidate medical record content is the content that the model should generate, which serves as a supervisory signal to guide the learning of the medical record generation model. Through this correspondence between input and label, the medical record generation model will gradually learn how to select the most matching part from the candidate content based on the input target medical record content, thereby optimizing its generation capabilities.
[0112] The present disclosure helps the medical record generation model learn how to automatically select the most matching relevant medical record content based on the input target medical record content by constructing the third training data of the medical record generation model. This enables the medical record generation model to better understand the intrinsic connection between different medical record contents, thereby improving the accuracy and robustness of model generation. Ultimately, the medical record generation model can automatically generate appropriate medical record content that conforms to medical logic based on the actual input medical record description, thereby improving the work efficiency and accuracy of doctors in clinical applications, and promoting intelligent progress in the medical field.
[0113] In an optional embodiment, determining target medical record data similar to the input medical record data based on the first training data set includes: determining the semantic similarity between the semantic vector of the input medical record data and the semantic vector of each real medical record data in the first training data set; and determining the real medical record data corresponding to the highest semantic similarity as the target medical record data similar to the input medical record data.
[0114] The input medical record data is semantically represented, that is, the text content is converted into a digital form that the machine can understand and process. This digital form is called a semantic vector. Use a pre-trained large language model to obtain the semantic vector. After encoding the input medical record data (for example, text content such as chief complaint and current medical history), use models such as BERT to generate its semantic vector. This semantic vector is a multi-dimensional vector that can accurately represent the semantic information of the input text, not just the literal meaning. The semantic vector can be obtained through the output of the last layer of the model. Similarly, each real medical record data in the first training data set (including basic information of the medical record, chief complaint, current medical history, past medical history, etc.) also needs to be converted into a semantic vector. Similar to the above steps, a large language model can be used to encode the real medical record data to obtain its corresponding semantic vector. These semantic vectors will be compared with the semantic vectors of the input medical record data to calculate the similarity between them.
[0115] In order to measure the similarity between the input medical record data and each real medical record data in the first training data, cosine similarity can be used for comparison. Cosine similarity measures the angle between two vectors. The closer the value is to 1, the more similar the two vectors are; the closer the value is to 0, the less similar the two vectors are. The calculation formula is as follows:
[0116]
[0117] Where A and B are the semantic vectors of the input medical record data and the real medical record data, respectively. A·B represents the dot product of the vectors, and ||A|| and ||B|| are the vector moduli (i.e., the magnitude of the vector).
[0118] For each input medical record data and each real medical record data, their cosine similarity is calculated to obtain a similarity score, which reflects the semantic similarity between the two medical record descriptions. After calculating the semantic similarity of all real medical record data, the real medical record data with the highest semantic similarity to the input medical record data is selected as the target medical record data.
[0119] Through the above process, the input medical record data will eventually be matched with the real medical record data with the highest semantic similarity. This match is not only based on the literal content of the text, but also through deep semantic understanding to ensure that the selected target medical record data is as consistent as possible with the input medical record data in terms of semantics, emotion, and tone. This allows the patient's emotion and tone when describing his condition to be accurately simulated in the subsequent simulation recording generation process.
[0120] In an optional embodiment, the generating of simulated recorded data corresponding to the input medical record data based on the input medical record data and the target medical record data includes: determining the target recorded data corresponding to the target medical record data; inputting the input medical record data and the target recorded data into a text generation model; and generating the simulated recorded data corresponding to the input medical record data through the text generation model based on a first text prompt word, wherein the first text prompt word is used to instruct the text generation model to rewrite the content in the target recorded data into the content in the input medical record data.
[0121] Determine the target recording data corresponding to the target medical record data, that is, the real recording data corresponding to the target medical record data. The target recording data is the recording data that is consistent with the target medical record data, including the patient's voice description, the doctor's consultation, etc. Although the input medical record data and the target medical record data are not exactly the same, the recording of the target medical record data can provide information such as voice characteristics and emotional tendencies as the basis for generating simulated recordings.
[0122] The input medical record data and target recording data are used as input and processed by a large text generation model (such as GPT-4, GPT-3, etc.). Based on the contextual relationship between the input text information and the recording data, a new text is generated to simulate the language characteristics of the target medical record data and reflect the specific information in the input medical record data.
[0123] Use the first text prompt to guide the model to generate content that meets the requirements. The text prompt is an instruction to the text generation model, telling the model how to generate a simulated recording text that matches the input medical record data description based on the content of the target recording data. Although the target recording data describes a similar condition, its details may be different from the input medical record data. The prompt will require the model to modify the specific content in the target recording (such as the duration and location of the headache) to a description that matches the input medical record data.
[0124] The first text prompt word can be as follows:
[0125] You are a professional doctor. You are given a medical record and a real outpatient recording. Based on the medical record, you rewrite the corresponding content in the recording to match the description in the medical record, and keep the other content in the recording.
[0126] Medical records: ${input medical records}
[0127] Real outpatient recording: ${Real outpatient recording}
[0128] After the new text content is generated by the text generation model, the text is converted into simulated recording data. The simulated recording data is not just a simple text-to-speech conversion, but also needs to ensure that the emotional characteristics of the speech are consistent with the target recording data. Finally, the text generation model will output simulated recording data that conforms to the input medical record data. It contains both the medical information in the input medical record data and the emotions, tone and other characteristics in the target medical record data. The role of the simulated recording data is to simulate the real communication between patients and doctors, so as to provide real simulation data for medical training, auxiliary diagnosis, etc.
[0129] In an optional embodiment, the target medical record data similar to the input medical record data is determined based on the first training data set, including: splitting the real recording data in the first training data set according to different context structures to obtain multiple split data; based on the real medical record data and the multiple split data, determining multiple target medical record data with the highest semantic similarity to the input medical record data.
[0130] In the first training data set, the real recording data contains a complete medical record description, including the conversation between the patient and the doctor, the description of the condition, the doctor's questions, the patient's answers, etc. In order to make these recording data more targeted when generating simulated recordings, the present disclosure proposes to split the recording data according to the context structure.
[0131] In an optional implementation, the real recorded data in the first training data set is split according to different context structures to obtain multiple split data, including: inputting the real recorded data in the first training data set into a text generation model; based on a second text prompt word, generating multiple split data through the text generation model, the second text prompt word is used to instruct the text generation model to split the real recorded data into multiple split data according to different context structures, and the context structures include greeting, chatting, consulting and farewell.
[0132] In order to split the real recording data into different context structures, the second text prompt word is used to guide the model. The second text prompt word will help the text generation model recognize different contexts in the conversation and instruct it to split according to the specified context structure. The second text prompt word is as follows:
[0133] You are a professional doctor. Split the transcription result of the following real outpatient recording according to the structure of "greeting", "chatting", "consulting", and "goodbye", and output the split results in JSON format.
[0134] In the present disclosure, the split context structures include: greeting, which refers to the greeting and hello between doctors or patients at the beginning of the conversation, such as "Hello, how are you feeling today?"; small talk, which refers to non-medical content that appears in medical conversations, such as weather, daily life and other light conversations, such as "How is the weather recently?" or "Are you busy at work?"; consultation, which refers to the doctor asking questions about the patient's symptoms, medical history, etc. to understand the patient's condition, such as "How long did the headache last?", "Do you feel nauseous?"; goodbye, which refers to the farewell part between the doctor and the patient at the end of the conversation, such as "I hope you get well soon" and "See you next time". These context structures represent different situations between doctors and patients in the conversation. During the splitting process, the text generation model will segment the recording data according to these context structures to obtain multiple structured split data.
[0135] The split data will become multiple small sub-data blocks, also known as split data. Each split data represents a specific part of the recording. Each split data will be paired with the original real medical record data and provide more context and detail references for subsequent simulation recording generation.
[0136] According to the input medical record data and the split target recording data, multiple target medical record data with the highest semantic similarity to the input medical record data are found. The similarity matching method can still adopt the above cosine similarity calculation method, which will not be repeated here. The difference here is that multiple similar target medical records need to be output for the purpose of random combination of subsequent recording data.
[0137] By splitting the real recorded data in the first training data set, performing structured processing according to different contexts, topics and emotions, and combining the semantic similarity method, we can more accurately select target medical record data that is similar to the input medical record data, providing a richer, more diverse and more accurate data source for subsequent medical record generation model training, which helps to generate high-quality simulated recorded data.
[0138] In an optional embodiment, the method of generating simulated recorded data of the input medical record data based on the input medical record data and the target medical record data includes: determining, from the multiple split data, multiple target split data corresponding one to one to the multiple target medical record data; randomly combining the multiple target split data according to different context structures to obtain target combination data; and generating simulated recorded data of the input medical record data based on the input medical record data and the target combination data.
[0139] As mentioned above, the target medical record data is selected based on the real medical record data in the first training data set, and the multiple target split data are split parts of the target medical record data, which match the input medical record data. For example, if the input medical record data describes "headache lasts for 3 days", then the multiple target split data of the target medical record data that matches it correspond to different parts of the input medical record data.
[0140] After determining the target split data, these split data are randomly combined according to different context structures (greeting, consulting, chatting, and saying goodbye). For multiple split data in each target medical record data, they can be randomly combined according to different context structures to obtain the target combination data corresponding to the input medical record data. This random combination does not mean random splicing, but a reasonable combination based on the natural flow of the conversation situation and factors such as tone and intonation to ensure that the final generated simulation recording sounds both natural and in line with the normal pattern of medical conversation. For example, the "greeting" part (such as "Hello, how do you feel today?") can be combined with the "consultation" part (such as "How long did the headache last?"), and the "goodbye" part (such as "See you next time, I wish you a speedy recovery") is placed at the end of the conversation. The purpose of random combination is to increase the diversity of the generation process, to ensure that the simulation recording data not only has sufficient similarity, but also can show the natural changes of different conversation situations. For example, when a patient describes his symptoms, his tone, intonation, and speed may change, and this change can be achieved by reasonably combining different split data. Finally, based on the input medical record data and the target combination data, simulated recording data is generated. Similarly, based on the first text prompt word, the input medical record data and the target combination data can be input into the text to generate a large model to obtain simulated recording data of the input medical record data.
[0141] The above steps involve detailed context splitting, random combination, and generation of simulated recordings, so that the final output simulated recording data can not only faithfully reflect the content of the medical record data, but also reflect the natural conversation situation and tone changes, thereby improving the training quality of the model and the realism of the generated results.
[0142] In an optional embodiment, the training of the medical record generation model based on the first training data set and the first training data includes: mixing the first training data set, the second training data and the third training data in proportion to obtain a training data set, wherein the third training data accounts for the highest proportion and the second training data accounts for the lowest proportion; and training the medical record generation model based on the training data set.
[0143] The first training data set contains data pairs consisting of real recording data and real medical record data, and each sample includes a medical record description text and the corresponding recording data. The second training data is simulated recording data generated based on the input medical record data and the target medical record data. The third training data is data generated through the medical record content selection task. Due to the different characteristics and functions of these three types of training data, the proportion weights of the first training data set, the second training data and the third training data mixed in the training data set will affect the focus and effect of the medical record generation model during training. The proportion of the third training data (which can be 2) helps the medical record generation model learn to understand and generate different medical record contents from the input text, and is the basis for the medical record generation model to learn to generate medical record text. The proportion of the second training data is relatively low (which can be 0.5). Although it can improve the tone and intonation performance of the model when generating simulated recordings, its generation process depends on the existing real medical record data. Therefore, the diversity and quality of the second training data are not as good as the first training data set. The proportion of the first training data set can be set to 1 (or an appropriate lower proportion). The first training data set relies on real medical record data and recording data. The amount of data is limited, but it is still very important because it contains real clinical recording data, which can help the model learn the direct mapping relationship between text and speech.
[0144] The mixed training data set is used as the final training data of the medical record generation model, and is divided into multiple batches of training data according to the preset batch size to train the medical record generation model. The medical record generation model adopts a multimodal learning method and can process text and audio data at the same time. The medical record generation model will learn how to generate audio from text, how to generate appropriate medical record content, and how to capture features such as emotions and tone during the generation process. During the training process, the medical record generation model will generate corresponding recording data (audio) based on the input medical record data (text), and can also generate corresponding medical record data based on the input recording data. The medical record generation model updates the model parameters through the stochastic gradient descent method, in which the loss function is optimized according to the similarity between text and audio, taking into account both the accuracy of speech and emotional factors such as tone and intonation. By calculating the difference between the generated audio and the target audio (such as through indicators such as mean square error and cross entropy), the medical record generation model continuously optimizes its model parameters until the generated medical record description text and corresponding audio are as close to the ideal output as possible.
[0145] The following is a detailed description of how to train the case generation model based on the first training data set, the second training data, and the third training data. The first, second, and third training data are mixed in proportion (which can be 1:0.5:2) to obtain the final training data set. The training data set is divided into multiple batches, each batch containing a number of sample data. The case generation model starts training based on multiple batches of training data sets.
[0146] For the first training data, the case generation model learns the mapping from text to audio and captures tone and emotional features. For the second training data, the case generation model is trained on how to generate speech data that matches emotions and intonation (fits the actual context) by simulating speech. For the third training data, the case generation model learns how to select the most matching content from candidate medical records based on the medical record content selection task and generates accurate medical record description text. The loss function of the case generation model will consider the matching degree between text and audio as well as features such as tone and emotion. The loss function is optimized by calculating the difference between the generated audio and the target audio (such as mean square error or cross entropy). Use stochastic gradient descent or other optimization algorithms to update the model parameters to gradually improve the similarity between text and audio, as well as the accuracy of the generated text.
[0147] After training, the case generation model can be evaluated through the validation data set to check the quality of the medical record text and audio data it generates. By continuously optimizing the loss function, the training data set gradually learns to generate more accurate and natural medical record content and simulated recordings. It can be seen that even if the sample data is scarce, the case generation model can be trained by generating enhanced data.
[0148] By mixing the first training data set, the second training data set, and the third training data set in proportion, the medical record generation model can better learn the relationship between text and audio in medical record generation, while ensuring that the generated medical record description is consistent with clinical reality. Mixing different types of training data not only improves the diversity of training data, but also makes up for the shortcomings of insufficient traditional data sets, thereby improving the performance and adaptability of the medical record generation model.
[0149] Figure 2 is a structural block diagram of a medical record generation model training device provided by an embodiment of the present disclosure, such as Figure 2 As shown, the device comprises:
[0150] The acquisition module 301 is used to acquire input medical record data;
[0151] A determination module 302 is used to determine target medical record data similar to the input medical record data based on a first training data set, wherein each sample data in the first training data set is a data pair consisting of real recording data and corresponding real medical record data;
[0152] A generating module 303, configured to generate simulated recording data corresponding to the input medical record data according to the input medical record data and the target medical record data;
[0153] The training module 304 is used to use the input medical record data and the simulated recording data as second training data, and train the medical record generation model based on the first training data set and the second training data.
[0154] In an optional embodiment, the device further comprises:
[0155] A task acquisition module, used to acquire a plurality of preset medical record content selection tasks, wherein the plurality of medical record content selection tasks are pairwise combinations of target medical record content of the input medical record data and other medical record content except the target medical record content, and the medical record content selection tasks are used to instruct the medical record generation model to generate medical record content with the highest matching degree with the target medical record content;
[0156] A similar medical record determination module, configured to determine, from among a plurality of candidate medical record contents corresponding to the other medical record contents, a target candidate medical record content having the highest matching degree with the target medical record content of the input medical record data;
[0157] The training data determination module is used to use the target medical record content, the multiple candidate medical record contents and the target candidate medical record content as third training data.
[0158] In an optional implementation, the determining module includes:
[0159] A first determination submodule, used to determine the semantic similarity between the semantic vector of the input medical record data and the semantic vector of each real medical record data in the first training data set;
[0160] The first determination submodule is used to determine the real medical record data corresponding to the highest semantic similarity as the target medical record data similar to the input medical record data.
[0161] In an optional implementation, the generating module includes:
[0162] A second determination submodule is used to determine the target recording data corresponding to the target medical record data;
[0163] A first input submodule, for inputting the input medical record data and the target recording data into a text generation model;
[0164] The first generation submodule is used to generate simulated recording data corresponding to the input medical record data through the text generation model based on a first text prompt word, and the first text prompt word is used to instruct the text generation model to rewrite the content in the target recording data into the content in the input medical record data.
[0165] In an optional implementation, the determining module includes:
[0166] A splitting submodule, used for splitting the real recording data in the first training data set according to different context structures to obtain a plurality of split data;
[0167] The third determination submodule is used to determine a plurality of target medical record data having the highest semantic similarity with the input medical record data based on the real medical record data and the plurality of split data.
[0168] In an optional implementation, the generating module includes:
[0169] A fourth determination submodule is used to determine, from the plurality of split data, a plurality of target split data corresponding one to one to the plurality of target medical record data;
[0170] A combination submodule, used for randomly combining the plurality of target split data according to different context structures to obtain target combination data;
[0171] The second generating submodule is used to generate simulated recording data of the input medical record data based on the input medical record data and the target combination data.
[0172] In an optional implementation, the splitting submodule includes:
[0173] An input unit, used for inputting the real recording data in the first training data set into a text to generate a large model;
[0174] A splitting unit is used to generate multiple split data based on a second text prompt word through the text generation model, wherein the second text prompt word is used to instruct the text generation model to split the real recording data into multiple split data according to different context structures, and the context structures include greeting, chatting, asking questions and saying goodbye.
[0175] In an optional implementation, the training module includes:
[0176] a mixing submodule, configured to mix the first training data set, the second training data, and the third training data in proportion to obtain a training data set, wherein the third training data accounts for the highest proportion and the second training data accounts for the lowest proportion;
[0177] A training submodule is used to train the medical record generation model based on the training data set.
[0178] An embodiment of the present disclosure also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the computer program is executed by the processor, the various processes in the above-mentioned embodiment of a medical record generation model training method are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0179] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods and devices. Therefore, the embodiments of the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present disclosure may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0180] The embodiments of the present disclosure are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing terminal device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded into a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0181] Although the preferred embodiments of the present disclosure have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present disclosure.
[0182] Finally, it should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "include..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements. The above is a detailed introduction to a medical record generation model training method, device and electronic device provided by the present disclosure. The principles and implementation methods of the present disclosure are explained in this article using specific examples. The description of the above embodiments is only used to help understand the method and its core idea of the present disclosure; at the same time, for those of ordinary skill in the art, according to the idea of the present disclosure, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present disclosure.
Claims
1. A medical record generation model training method, characterized in that: The method comprises: Obtain input medical record data; Based on a first training data set, target medical record data similar to the input medical record data is determined, wherein each sample data in the first training data set is a data pair consisting of real recording data and corresponding real medical record data; Generate simulated recording data corresponding to the input medical record data according to the input medical record data and the target medical record data; The input medical record data and the simulated recording data are used as second training data, and the medical record generation model is trained based on the first training data set and the second training data.
2. The method according to claim 1, characterized in that The method further comprises: Acquire a plurality of preset medical record content selection tasks, wherein the plurality of medical record content selection tasks are pairwise combinations of target medical record content of the input medical record data and other medical record content except the target medical record content, and the medical record content selection tasks are used to instruct the medical record generation model to generate medical record content with the highest matching degree with the target medical record content; Determining, from a plurality of candidate medical record contents corresponding to the other medical record contents, a target candidate medical record content having the highest matching degree with the target medical record content of the input medical record data; The target medical record content, the multiple candidate medical record contents, and the target candidate medical record content are used as third training data.
3. The method according to claim 1, characterized in that The step of determining target medical record data similar to the input medical record data based on the first training data set includes: Determine the semantic similarity between the semantic vector of the input medical record data and the semantic vector of each real medical record data in the first training data set; The real medical record data corresponding to the highest semantic similarity is determined as the target medical record data similar to the input medical record data.
4. The method according to claim 3, characterized in that The step of generating simulated recording data corresponding to the input medical record data according to the input medical record data and the target medical record data includes: Determine the target recording data corresponding to the target medical record data; Inputting the input medical record data and the target audio recording data into a text generation model; Based on the first text prompt word, the text generation model generates simulated recording data corresponding to the input medical record data, and the first text prompt word is used to instruct the text generation model to rewrite the content in the target recording data into the content in the input medical record data.
5. The method according to claim 1, characterized in that The step of determining target medical record data similar to the input medical record data based on the first training data set includes: Splitting the real recording data in the first training data set according to different context structures to obtain a plurality of split data; Based on the real medical record data and the multiple split data, multiple target medical record data with the highest semantic similarity to the input medical record data are determined.
6. The method according to claim 5, characterized in that The step of generating simulated recording data of the input medical record data according to the input medical record data and the target medical record data includes: Determining, from the plurality of split data, a plurality of target split data corresponding one to one to the plurality of target medical record data; According to different context structures, the plurality of target split data are randomly combined to obtain target combination data; Based on the input medical record data and the target combination data, simulated recording data of the input medical record data is generated.
7. The method according to claim 5, characterized in that The real recording data in the first training data set is split according to different context structures to obtain a plurality of split data, including: Input the real recording data in the first training data set into text to generate a large model; Based on the second text prompt word, multiple split data are generated by the text generation model, and the second text prompt word is used to instruct the text generation model to split the real recording data into multiple split data according to different context structures, and the context structures include greeting, chatting, consulting and farewell.
8. The method according to claim 2, characterized in that: The training of the medical record generation model based on the first training data set and the first training data includes: The first training data set, the second training data, and the third training data are mixed in proportion to obtain a training data set, wherein the third training data accounts for the highest proportion and the second training data accounts for the lowest proportion; The medical record generation model is trained based on the training data set.
9. A medical record generation model training device, characterized in that: The device comprises: An acquisition module, used to acquire input medical record data; A determination module, configured to determine target medical record data similar to the input medical record data based on a first training data set, wherein each sample data in the first training data set is a data pair consisting of real recording data and corresponding real medical record data; A generating module, used for generating simulated recording data corresponding to the input medical record data according to the input medical record data and the target medical record data; A training module is used to use the input medical record data and the simulated recording data as second training data, and to train a medical record generation model based on the first training data set and the second training data.
10. An electronic device, characterized in that: It comprises a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement a medical record generation model training method as described in any one of claims 1-8.