Medical record generation model training method, medical record generation method and related equipment

By performing score screening and collaborative training of doctor-patient dialogue data, the generative model can more accurately identify and structure clinical entities, solve the problem of semantic faults in traditional medical record generation methods, and improve the accuracy and reliability of medical record generation.

CN120183592APending Publication Date: 2025-06-20PENG CHENG LAB
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510293863.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Traditional medical record generation methods cannot fully cover the dialogue information in all scenarios because the fixed preset medical record item matching rules cannot fully cover the dialogue information in all scenarios, resulting in the missing data item information or the filled-in data item information in the generated medical record, which is inaccurate, and has low accuracy and reliability.

Method used

By obtaining multiple doctor-patient conversation sample data and corresponding medical record labels, the conversation score of each conversation sample data is calculated, the target doctor-patient conversation sample data is selected, and the initial medical record generation model is collaboratively trained based on multiple language specification sample data, and the target doctor-patient conversation sample data is input to the medical record generation model one by one for data prediction processing, obtain the predicted medical record, and the model is updated based on the loss value of the predicted medical record and the target medical record label.

Benefits of technology

It improves the accuracy and reliability of the medical records generated based on doctor-patient dialogue data, can automatically identify clinical entities in the patient's oral description, and perform structured reorganization in accordance with medical standards, solve the semantic fault problems caused by traditional template filling, and achieve accurate extraction of key information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183592A_ABST
    Figure CN120183592A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a medical record generation model training method, a medical record generation method and related equipment. The method comprises the steps of obtaining multiple pieces of doctor-patient dialogue sample data and corresponding medical record labels; the dialogue score of each piece of doctor-patient dialogue sample data is calculated, target doctor-patient dialogue sample data is selected from the multiple pieces of doctor-patient dialogue sample data based on the dialogue scores, and a medical record label corresponding to the target doctor-patient dialogue sample data is a target medical record label; performing cooperative training on the initial medical record generation model based on the multiple pieces of language specification capability sample data to obtain a medical record generation model; inputting the target doctor-patient dialogue sample data into a medical record generation model one by one for data prediction processing to obtain a predicted medical record; and updating the medical record generation model based on the predicted medical record and the loss value of the target medical record label, the updated medical record generation model being used for performing medical record generation based on the doctor-patient dialogue data, and improving the accuracy and reliability of the generated medical record when generating the medical record according to the actual doctor-patient dialogue data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of model data processing, and particularly to a method for training a medical record generation model, a method for generating a medical record, and related devices. Background Art

[0002] With the continuous progress of medical technology and the increasing needs of patients, the traditional methods of medical consultation and medical record keeping are no longer able to meet the needs of modern medical services. Especially in the context of a shortage of medical resources and a large number of patients, in order to improve the efficiency of medical services and the accuracy of subsequent doctor diagnoses, how to automatically generate clear medical records based on doctor-patient communication information is a problem worthy of study.

[0003] In related technologies, usually, the dialogue information between doctors and patients during the medical consultation session is obtained, and then key information is screened from this dialogue information based on preset medical record item matching rules, and the screened key information is input into a preset medical record template to automatically generate the corresponding medical record. However, due to the fact that the actual doctor-patient communication methods are often restricted by multiple factors such as time, language, and knowledge level, the fixed preset medical record item matching rules cannot fully cover the data information that may appear in the dialogue information in all scenarios. For example, patients may not be able to describe their problems in professional terms, etc., resulting in the situation that the medical records generated in this way are prone to missing data item information or the filled data item information is inaccurate, leading to a relatively low accuracy and reliability of the generated medical records. Summary of the Invention

[0004] The embodiments of this application provide a method for training a medical record generation model, a method for generating a medical record, and related devices, which can improve the accuracy and reliability of the medical records generated based on doctor-patient dialogue data.

[0005] To achieve the above object, a first aspect of the embodiments of this application proposes a method for training a medical record generation model, and the method includes:

[0006] Obtain multiple doctor-patient dialogue sample data and corresponding medical record labels;

[0007] Calculate the dialogue score of each piece of the doctor-patient dialogue sample data, and select target doctor-patient dialogue sample data from the multiple pieces of doctor-patient dialogue sample data based on the dialogue score, and the medical record label corresponding to the target doctor-patient dialogue sample data is the target medical record label;

[0008] Co-train an initial medical record generation model based on multiple language specification ability sample data to obtain a medical record generation model;

[0009] Input the target doctor-patient dialogue sample data into the medical record generation model one by one for data prediction processing to obtain predicted medical records;

[0010] Update the medical record generation model based on the loss value of the predicted medical record and the target medical record label, where the updated medical record generation model is used to generate medical records based on doctor-patient dialogue data.

[0011] In some embodiments, the updating the medical record generation model based on the loss value of the predicted medical record and the target medical record label includes:

[0012] Calculate the cross-entropy loss value of the predicted medical record and the target medical record label based on the cross-entropy loss function;

[0013] Calculate the regularization term value based on the regularization term function and the weight matrix of the medical record generation model;

[0014] Accumulate the cross-entropy loss value and the regularization term value to obtain the loss value, and update the model parameters of the medical record generation model based on the loss value.

[0015] In some embodiments, the calculating the dialogue score of each doctor-patient dialogue sample data and selecting the target doctor-patient dialogue sample data from the multiple doctor-patient dialogue sample data based on the dialogue score includes:

[0016] Based on the character probability of each character in the doctor-patient dialogue sample data, perform logarithmic processing to obtain the character self-information score;

[0017] Average the character self-information scores in each doctor-patient dialogue sample data to obtain the dialogue score of each doctor-patient dialogue sample data;

[0018] Select the doctor-patient dialogue sample data with the dialogue score exceeding the preset score threshold from the multiple doctor-patient dialogue sample data as the target doctor-patient dialogue sample data.

[0019] In some embodiments, the language specification ability sample data includes abstract generation data, medical record structure data, language normalization data, and interrogation process data. The co-training the initial medical record generation model based on multiple language specification ability sample data to obtain the medical record generation model includes:

[0020] Calculate the abstract training gradient of the initial medical record generation model based on the abstract generation data;

[0021] Calculate the medical record structure training gradient of the initial medical record generation model based on the medical record structure data;

[0022] Calculate the language normalization training gradient of the initial medical record generation model based on the language normalization data;

[0023] Based on the medical interview process data, calculate the training gradient of the medical interview process for the initial medical record generation model;

[0024] Based on the summary training gradient, the medical record structure training gradient, the language normalization training gradient, and the medical interview process training gradient, update the initial model parameters of the initial medical record generation model.

[0025] In some embodiments, the updating the initial model parameters of the initial medical record generation model based on the summary training gradient, the medical record structure training gradient, the language normalization training gradient, and the medical interview process training gradient includes:

[0026] Calculate the gradient cosine similarity between every two training gradients, where the training gradients include the summary training gradient, the medical record structure training gradient, the language normalization training gradient, and the medical interview process training gradient;

[0027] Based on all the gradient cosine similarities and all the training gradients, obtain a combined training gradient;

[0028] Based on the combined training gradient, update the initial model parameters.

[0029] In some embodiments, the obtaining a combined training gradient based on all the gradient cosine similarities and all the training gradients includes:

[0030] When all the gradient cosine similarities are non - negative, accumulate the training gradients to obtain the combined training gradient;

[0031] When there is any one of the gradient cosine similarities that is negative, take the two training gradients corresponding to the negative gradient cosine similarity as the first adjustment gradient and the second adjustment gradient respectively;

[0032] Based on the first adjustment gradient and the second adjustment gradient, perform a projection process on the first adjustment gradient to obtain an updated first adjustment gradient, where the cosine similarity between the updated first adjustment gradient and the second adjustment gradient is non - negative;

[0033] Based on the cosine similarity between the first adjustment gradient and the updated first adjustment gradient, obtain a first projection weight;

[0034] Based on the first projection weight, perform a weighted process on the corresponding training gradient, and accumulate the weighted training gradients to obtain the combined training gradient.

[0035] The performing a projection process on the first adjustment gradient based on the first adjustment gradient and the second adjustment gradient to obtain an updated first adjustment gradient includes:

[0036] An adjustment gradient product is obtained based on the product of the first adjustment gradient and the second adjustment gradient;

[0037] Based on the ratio of the adjustment gradient product to the gradient norm of the second adjustment gradient, and then multiplying by the second adjustment gradient, a first projection difference is obtained;

[0038] Based on the difference between the first adjustment gradient and the first projection difference, the updated first adjustment gradient is obtained.

[0039] In some embodiments, the language specification ability sample data includes first batch sample data and second batch sample data. Updating the initial model parameters of the initial medical record generation model based on the summary training gradient, the medical record structure training gradient, the language normalization training gradient, and the consultation process training gradient further includes:

[0040] Updating the initial model parameters of the initial medical record generation model based on the training gradient corresponding to the first batch sample data to obtain first updated model parameters, where the training gradient includes the summary training gradient, the medical record structure training gradient, the language normalization training gradient, and the consultation process training gradient;

[0041] Updating the first updated model parameters of the initial medical record generation model based on the training gradient corresponding to the second batch sample data, where both the first batch sample data and the second batch sample data include at least one of the summary generation data, at least one of the medical record structure data, at least one of the language normalization data, and at least one of the consultation process data.

[0042] To achieve the above object, a second aspect of the embodiments of the present application proposes a medical record generation method, and the method includes:

[0043] Display a doctor-patient input prompt, and obtain doctor-patient dialogue data obtained based on the doctor-patient input prompt;

[0044] Based on the end flag in the doctor-patient dialogue data, input the doctor-patient dialogue data into the medical record generation model trained by the medical record generation model training method as described in the first aspect for data processing to generate a target medical record;

[0045] Display the target medical record.

[0046] To achieve the above object, a third aspect of the embodiments of the present application proposes a medical record generation model training device, and the device includes:

[0047] A sample data acquisition module for acquiring multiple doctor-patient dialogue sample data and corresponding medical record labels;

[0048] A target sample data screening module, configured to calculate the dialogue score of each piece of the doctor-patient dialogue sample data, and select target doctor-patient dialogue sample data from the multiple pieces of doctor-patient dialogue sample data based on the dialogue score, where the medical record label corresponding to the target doctor-patient dialogue sample data is a target medical record label;

[0049] An ability collaborative training module, configured to perform collaborative training on an initial medical record generation model based on multiple language specification ability sample data to obtain a medical record generation model;

[0050] A data processing module, configured to input the target doctor-patient dialogue sample data into the medical record generation model one by one for data prediction processing to obtain a predicted medical record;

[0051] A model update module, configured to update the medical record generation model based on the loss value of the predicted medical record and the target medical record label, and the updated medical record generation model is used to generate a medical record based on doctor-patient dialogue data.

[0052] To achieve the above object, a fourth aspect of the embodiments of the present application provides an electronic device, where the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the medical record generation model training method as described in the first aspect or the medical record generation method as described in the second aspect.

[0053] To achieve the above object, a fifth aspect of the embodiments of the present application provides a storage medium, where the storage medium is a computer-readable storage medium, the storage medium stores a computer program, and when the computer program is executed by a processor, it implements the medical record generation model training method as described in the first aspect or the medical record generation method as described in the second aspect.

[0054] The methods for training a medical record generation model, a method for generating a medical record, and related devices proposed in the embodiments of the present application include: First, obtain multiple doctor-patient dialogue sample data and corresponding medical record labels; Second, calculate the dialogue score of each doctor-patient dialogue sample data, and select target doctor-patient dialogue sample data from the multiple doctor-patient dialogue sample data based on the dialogue score. The medical record label corresponding to the target doctor-patient dialogue sample data is the target medical record label; Then, co-train the initial medical record generation model based on multiple language specification ability sample data to obtain a medical record generation model; Next, input the target doctor-patient dialogue sample data into the medical record generation model one by one for data prediction processing to obtain predicted medical records; Finally, update the medical record generation model based on the loss value between the predicted medical record and the target medical record label. The updated medical record generation model is used to generate medical records based on doctor-patient dialogue data. The embodiments of the present application pre-screen high-quality doctor-patient dialogue samples based on a dialogue score screening mechanism to ensure that the training data covers diverse expressions in real scenarios, and co-train the initial medical record generation model with language specification capabilities that can improve term standardization and medical record structure understanding, enabling the model to break through the preset rule limitations, automatically identify clinical entities in the patient's colloquial descriptions, and perform structured reorganization according to medical norms. Through the iterative optimization of the predicted medical record and the standard label, a deep semantic mapping from free dialogue to a standardized medical record is established, solving the semantic break problem caused by traditional template filling, and achieving accurate extraction of key information without omission, so that the trained medical record generation model has the ability to dynamically adapt to different expression habits and consultation scenarios, fundamentally overcoming the coverage blind spots of fixed rule systems, and thus improving the accuracy and reliability of the generated medical records when generating medical records based on actual doctor-patient dialogue data.

[0055] Other features and advantages of the present application will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the specification, the claims, and the drawings. Brief Description of the Drawings

[0056] Figure 1 is a flowchart of a method for training a medical record generation model provided by an embodiment of the present application.

[0057] Figure 2 is a schematic diagram of a doctor-patient dialogue sample data provided by another embodiment of the present application.

[0058] Figure 3 is a schematic diagram of adding an end symbol to doctor-patient dialogue sample data provided by another embodiment of the present application.

[0059] Figure 4 is a schematic diagram of a simple medical record generated by using an existing large model provided by another embodiment of the present application.

[0060] Figure 5 is Figure 1 The flowchart of step 102 in

[0061] Figure 6 is Figure 1 The flowchart of step 103 in

[0062] Figure 7 is Figure 6 The flowchart of step 605 in

[0063] Figure 8 is Figure 6 Another flowchart of step 605 in

[0064] Figure 9 is Figure 8 The flowchart of step 802 in

[0065] Figure 10 is Figure 9 The flowchart of step 903 in

[0066] Figure 11 is Figure 1 The flowchart of step 105 in

[0067] Figure 12 It is a schematic diagram of post-generation fine-tuning of medical records provided by another embodiment of the present application.

[0068] Figure 13 It is a schematic flowchart of the training and use of a medical record generation model provided by another embodiment of the present application.

[0069] Figure 14 It is a flowchart of a medical record generation method provided by another embodiment of the present application.

[0070] Figure 15 It is a schematic flowchart of the deployment and feedback iteration of a medical record generation model service provided by another embodiment of the present application.

[0071] Figure 16 It is a schematic structural diagram of a medical record generation model training device provided by another embodiment of the present application.

[0072] Figure 17 It is a schematic hardware structure diagram of an electronic device provided by another embodiment of the present application. Detailed implementation manners

[0073] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0074] It should be noted that although the functional modules are divided in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different module division from that in the device or a different order from that in the flowchart.

[0075] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0076] With the continuous progress of medical technology and the increasing needs of patients, the traditional methods of medical consultation and medical record keeping are no longer able to meet the needs of modern medical services. Especially in the context of a shortage of doctor resources and a large number of patients, in order to improve the efficiency of medical services and the diagnostic accuracy of subsequent doctors, how to improve the automatic generation of clear medical records based on doctor-patient communication information is a problem worthy of study.

[0077] In the related art, usually, the dialogue information between doctors and patients during the medical consultation session is obtained, and then key information is screened from this dialogue information based on preset medical record item matching rules, and the screened key information is input into a preset medical record template to automatically generate the corresponding medical record. However, due to the fact that the actual doctor-patient communication method is often restricted by multiple factors such as time, language, and knowledge level, the fixed preset medical record item matching rules cannot fully cover the data information that may appear in the dialogue information in all scenarios. For example, patients may not be able to describe their problems in professional terms, etc., resulting in the situation that the medical records generated in this way are prone to missing data item information or having inaccurate data item information filled in, leading to a relatively low accuracy and reliability of the generated medical records.

[0078] In order to improve the accuracy and reliability of the medical records generated based on doctor-patient dialogue data, the embodiments of this application pre-screen high-quality doctor-patient dialogue samples based on a dialogue score screening mechanism to ensure that the training data covers diverse expressions in real scenarios, and use the language specification ability that can improve term standardization and medical record structure understanding to co-train the initial medical record generation model so that the model can break through the preset rule limitations, automatically identify clinical entities in the patient's colloquial descriptions, and perform structured reorganization according to medical norms. Through the iterative optimization of the predicted medical records and standard labels, a deep semantic mapping from free dialogue to standardized medical records is established to solve the semantic break problem caused by traditional template filling, and achieve accurate extraction of key information without omission, so that the trained medical record generation model has the ability to dynamically adapt to different expression habits and medical consultation scenarios, fundamentally overcoming the coverage blind spots of the fixed rule system, and thus improving the accuracy and reliability of the generated medical records when generating medical records based on actual doctor-patient dialogue data.

[0079] The method for training a medical record generation model, the medical record generation method, and related devices provided by the embodiments of the present application will be further described below. First, the method for training a medical record generation model will be described. Refer to Figure 1 , which is an optional flowchart of the method for training a medical record generation model provided by the embodiments of the present application. Figure 1 The method in Figure 1 may include but is not limited to steps 101 to 105. At the same time, it can be understood that the order of steps 101 to 105 in this embodiment is not specifically limited, and the order of steps can be adjusted according to actual needs, or some steps can be reduced or added. The method for training a medical record generation model provided by the embodiments of the present application can be applied to any server or intelligent terminal connected to the hospital system, etc.

[0080] Step 101: Obtain multiple doctor-patient conversation sample data and corresponding medical record labels.

[0081] The following will describe step 101 in detail.

[0082] In some embodiments, in order to train a medical record generation model that can generate accurate medical records for actual doctor-patient conversation data, it is first necessary to obtain a large amount of doctor-patient conversation sample data and corresponding medical record labels when doctors and patients have conversations as training samples. The key among them is the chief complaint, current medical history, and diagnosis, which serve as both the basis for conversation generation and the basis for generating medical records. Refer to Figure 2 , which is a schematic diagram of a doctor-patient conversation sample data provided by the embodiments of the present application. As Figure 2 shown in

[0083] Refer to Figure 3 , which is a schematic diagram of adding an end symbol to the doctor-patient conversation sample data provided by the embodiments of the present application. As Figure 3 shown in <end>, <End of consultation>, when relevant symbols are detected, it will automatically exit the consultation process and enter the medical record generation stage as a calling tool.

[0084] Reference Figure 4 , is a schematic diagram of a simplified medical record generated using an existing large model provided in an embodiment of the present application. Figure 4 As shown in , we can directly use the existing industry-leading models such as Wenxin Yiyan and Tongyi Qianwen to generate medical records to obtain the following Figure 2 The medical records corresponding to the sample data of the doctor-patient dialogue shown in the figure. However, the medical records directly generated by this existing large model are prone to problems such as insufficient summarization ability, unclear structure of medical records, and poor entity normalization ability.

[0085] It is understandable that insufficient summarizing ability refers to the inability to extract important information related to the patient's condition from a longer doctor-patient conversation.

[0086] Not knowing the structure of medical records means that during the patient's visit to the hospital, the types of medical documents generated are diverse, covering dozens or even hundreds of types of documents such as admission records, discharge records, medical records, surgical records, and examination reports. When we give instructions to generate a specific type of document based on the content of the conversation, the large model may generate incorrect content because it does not know the type of document. In addition, different fields of different types of documents have writing specifications. For example, the writing specifications of the chief complaint in the admission record are: the chief complaint should be concise, generally not more than 20 words, and can derive the main diagnosis, or should be able to see the characteristics of the disease of the first diagnosis; the chief complaint only writes the disease that prompted the current hospitalization, and the symptoms and medical history of other diseases are written in the past history; the chief complaint should use symptomology terms, and in principle, it should not be replaced by diagnostic names or auxiliary examinations. Only when the diagnosis has been clearly diagnosed can the disease name be used, etc. If the model is not clear about these specifications, the quality of the generated results will be relatively poor.

[0087] Poor entity normalization capability means that in doctor-patient conversations, both patients and doctors are accustomed to using some easy-to-understand entity descriptions, such as stomachache.

[0088] In some embodiments, after obtaining multiple doctor-patient dialogue sample data to generate a doctor-patient consultation dialogue, post-processing is required. First, it is necessary to split the complete text dialogue data into dialogue data with roles, such as doctor role and patient role. The split roles can use regular expressions to distinguish the roles. Secondly, it is necessary to judge the doctor's replies with problems in the replies and questions and perform data filtering to filter out the training sample data with low reference value, thereby improving the training efficiency and reliability of the medical record generation model.

[0089] The evaluation process can be manual verification, that is, inviting medical students to review the conversations and record the scores. Since manual evaluation is too time-consuming and laborious, it is proposed to use self-information to evaluate the quality of the conversations and determine whether to retain the conversations as the training set, which is described in detail as follows.

[0090] Step 102: Calculate the conversation score for each doctor-patient conversation sample data, and select the target doctor-patient conversation sample data from multiple doctor-patient conversation sample data based on the conversation score.

[0091] The following is a detailed description of Step 102.

[0092] In some embodiments, using self-information to evaluate the quality of the conversation means calculating the conversation score for each doctor-patient conversation sample data through a language model, and then determining whether this sentence should be retained as the target doctor-patient conversation sample data based on the conversation score of each doctor-patient conversation sample data, and the medical record label corresponding to the target doctor-patient conversation sample data is the target medical record label. How to perform data screening will be further described below.

[0093] Refer to Figure 5 , calculating the conversation score for each doctor-patient conversation sample data, and selecting the target doctor-patient conversation sample data from multiple doctor-patient conversation sample data based on the conversation score, including the following steps 501 to 503.

[0094] Step 501: Based on the character probability of each character in the doctor-patient conversation sample data, perform logarithmic processing to obtain the character self-information score.

[0095] Step 502: Average the character self-information scores in each doctor-patient conversation sample data to obtain the conversation score for each doctor-patient conversation sample data.

[0096] Step 503: Select the doctor-patient conversation sample data with a conversation score exceeding the preset score threshold from multiple doctor-patient conversation sample data as the target doctor-patient conversation sample data.

[0097] The following is a detailed description of Steps 501 to 503.

[0098] In some embodiments, for each character in the doctor-patient conversation sample data, there should be a character probability of the next character as the current character, that is, P(x t |x0,x1,x2…x t-1 ) represents the character probability that the current character x t-1 appears when the previous characters x0,x1,x2…x t from the beginning to the previous character of the current character appear in the doctor-patient conversation sample data.

[0099] Next, take the logarithm of the character probability to obtain the self-information score of the current character x t as shown in the following formula (1).

[0100] I(x)= -log2P(x t |x0,x1,x2…x t-1 ) (1)

[0101] After that, for each doctor-patient dialogue sample data, calculate the self-information scores of all characters, and then take the average of the self-information scores of all characters in the doctor-patient dialogue sample data to obtain the dialogue score of the doctor-patient dialogue sample data. Then, select the doctor-patient dialogue sample data with a dialogue score exceeding the preset score threshold from multiple doctor-patient dialogue sample data as the target doctor-patient dialogue sample data.

[0102] That is, the doctor-patient dialogue score (the higher the score, the better) is the average of the scores of all characters in the dialogue. It can be understood that the preset score threshold can be a score value customized according to user needs, or a score value calculated according to the division ratio.

[0103] Self-information also uses the character prediction probability given by the language model during prediction to evaluate the probability of character occurrence, and evaluates the importance and the quality of the model's cognition based on this probability.

[0104] Through the above steps 501 to 503, calculate the dialogue self-information score using the logarithmic transformation of character-level probabilities, effectively capture abnormal expressions in doctor-patient dialogues, accurately identify high-quality dialogue samples with complete semantics and clear logic through the average self-information score of the overall dialogue information entropy, and then use the preset threshold screening mechanism to eliminate low-quality data containing invalid information from the source, ensuring that the training set strictly follows the professional norms of medical consultations. This data purification strategy enables the model to focus on learning the standard doctor-patient interaction mode, significantly improving the semantic coherence and clinical accuracy of the generated medical records, so as to solve problems such as term conversion errors and key information omissions caused by interference from low-quality data. Thus, when using the selected target doctor-patient corresponding sample data for training the medical record generation model in the future, the training efficiency and accuracy of the medical record generation model can be improved.

[0105] Step 103: Co-train the initial medical record generation model based on multiple language specification ability sample data to obtain the medical record generation model.

[0106] The following provides a detailed description of Step 103.

[0107] In some embodiments, in order to further solve the problems of the existing model such as insufficient summarization ability, unclear medical record structure, and poor entity normalization ability. Before training the model using the doctor-patient dialogue sample data, it is also necessary to first use the existing large model as the initial sample medical record generation model, and then co-train the initial sample medical record generation model using multiple language specification ability sample data to obtain a medical record generation model that can solve the problems of insufficient summarization ability, unclear medical record structure, and poor entity normalization ability, so as to improve the reliability of the generated model medical record generation model. The following will further describe how to co-train the initial medical record generation model using multiple language specification ability sample data.

[0108] Referring to Figure 6 , co-train the initial medical record generation model based on multiple language specification ability sample data to obtain a medical record generation model, including the following steps 601 to step 605.

[0109] Step 601: Calculate the summary training gradient of the initial medical record generation model based on the summary generation data.

[0110] Step 602: Calculate the medical record structure training gradient of the initial medical record generation model based on the medical record structure data.

[0111] Step 603: Calculate the language normalization training gradient of the initial medical record generation model based on the language normalization data.

[0112] Step 604: Calculate the interrogation process training gradient of the initial medical record generation model based on the interrogation process data.

[0113] Step 605: Update the initial model parameters of the initial medical record generation model based on the summary training gradient, medical record structure training gradient, language normalization training gradient, and interrogation process training gradient.

[0114] The following will describe steps 601 to 605 in detail.

[0115] In some embodiments, in order to solve the problems in the existing large model such as insufficient summarization ability, unclear medical record structure, and poor entity normalization ability.

[0116] Collect a large amount of publicly available summary generation data for training the summary generation task in the initial medical record generation model, mainly to improve the model's ability in this aspect through learning related tasks of these large amount of summary summarization. The summary generation data set includes: LCST large-scale Chinese short text summary data set, nlpcc2017 summary data, THUCNews data, SogouCS data, Chinese scientific literature csl summary data, and so on.

[0117] The medical record structure data is used to train the medical record structure understanding task in the initial medical record generation model. This task mainly collects a large number of real, high-quality medical records and splices the corresponding <medical record type, medical record content> of these medical records, and inputs them into the big model. The purpose is to allow the big model to see enough medical record data in the incremental pre-training stage and have a full understanding of the content and structure of the medical records. The specific data composition is in the form of <This is a medical record of type xx, the specific content of xx medical record>, where xx includes admission records, discharge records, first medical records, daily medical records, etc. The purpose is to improve the big model's ability to understand the medical record structure.

[0118] The entity normalization task in the initial medical record generation model is trained using language normalization data. In order to improve the standardization of entity expressions in the final generated medical record results, it is necessary to let the large model learn as many corresponding relationships between the colloquial expressions and the standardized expressions of entities as possible during the incremental pre-training stage. Therefore, through manual sorting and open source data sets, we have collected a large number of colloquial entities and their corresponding standardized entity pairs and combined them into the following data sets:<xx实体在病历中的规范化表述为,xx实体对应的规范化表述实体> .

[0119] The consultation process data is used to train the consultation dialogue training task in the initial medical record generation model to improve the fluency and logic of the consultation.

[0120] In some embodiments, the above four types of language standardization ability sample data can be used to train the initial medical record generation model one by one in batches and types to complete the corresponding training tasks, such as first using summary generation data to train the initial medical record generation model, then using medical record structure data to train the initial medical record generation model, and then using language normalization data to train the initial medical record generation model, and finally using consultation process data to train the initial medical record generation model.

[0121] However, this kind of sequential training is prone to the training results of the previous training tasks being affected by the later training tasks, that is, in the process of training the initial medical record generation model using the medical record structure data, the training results of the previous training using the summary generation data are affected. Therefore, the embodiment of the present application uses collaborative training to train these data together, thereby improving the summary extraction ability, format standardization ability, language normalization ability and fluency of the initial medical record generation model. The following will further describe how to perform collaborative training.

[0122] Reference Figure 7 , based on the summary training gradient, the medical record structure training gradient, the language normalization training gradient and the consultation process training gradient, the initial model parameters of the initial medical record generation model are updated, including the following steps 701 to 702.

[0123] Step 701: Update the initial model parameters of the initial medical record generation model based on the training gradients corresponding to the first batch of sample data to obtain the first updated model parameters.

[0124] Step 702: Update the first updated model parameters of the initial medical record generation model based on the training gradients corresponding to the second batch of sample data.

[0125] The following provides a detailed description of Steps 701 to 702.

[0126] In some embodiments, the language specification ability sample data includes multiple batches of ability sample data, such as the first batch of sample data and the second batch of sample data. Then, each batch of ability sample data includes at least one summary generation data, at least one medical record structure data, at least one language normalization data, and at least one interrogation process data.

[0127] Then, each batch of ability sample data is used to co-train the initial medical record generation model one by one. For example, update the initial model parameters of the initial medical record generation model based on the training gradients corresponding to the first batch of sample data to obtain the first updated model parameters, and then update the first updated model parameters of the initial medical record generation model based on the training gradients corresponding to the second batch of sample data.

[0128] The following will further describe the process of co-training the initial medical record generation model using each batch of ability sample data.

[0129] Refer to Figure 8 , based on the summary training gradient, medical record structure training gradient, language normalization training gradient, and interrogation process training gradient, updating the initial model parameters of the initial medical record generation model further includes the following Steps 801 to 803.

[0130] Step 801: Calculate the gradient cosine similarity between every two training gradients.

[0131] Step 802: Obtain the combined training gradient based on all the gradient cosine similarities and all the training gradients.

[0132] Step 803: Update the initial model parameters based on the combined training gradient.

[0133] The following provides a detailed description of Steps 801 to 803.

[0134] In some embodiments, first, use the relevant loss function to calculate the summary training gradient corresponding to the summary generation data, the medical record structure training gradient corresponding to the medical record structure data, the language normalization training gradient corresponding to the language normalization data, and the interrogation process training gradient corresponding to the interrogation process data.

[0135] It is understandable that the loss function can be a common loss function, such as the classification loss function of cross-entropy, to calculate the loss value corresponding to the initial medical record generation model for the sample data, and then calculate the training gradient value corresponding to the loss value through the chain rule, such as g n , where n ∈ {1, 2, 3, 4} corresponds to the abstract training gradient, the medical record structure training gradient, the language normalization training gradient, and the consultation process training gradient respectively.

[0136] During the process of co-training the initial medical record generation model with sample data of multiple capabilities, there may be conflicts in the gradient directions of the training gradients corresponding to different capability training data, resulting in a decrease in the overall performance. To solve this problem, the embodiments of the present application eliminate the gradient direction conflicts corresponding to different capability training data through projection, and then obtain the combined training gradient by weighting according to the cosine similarity of the gradient vectors before and after projection, as specifically described below.

[0137] First, calculate the gradient cosine similarity between every two training gradients (including the abstract training gradient, the medical record structure training gradient, the language normalization training gradient, and the consultation process training gradient) to determine whether there is a conflict, and then use all the gradient cosine similarities and all the training gradients to obtain the combined training gradient Φ(g1, g2, g3, g4) that does not conflict, so as to update the initial model parameters θ of the initial medical record generation model using the combined training gradient Φ(g1, g2, g3, g4) as shown in the following formula (2).

[0138] θ ← θ - η·Φ(g1, g2, g3, g4) (2)

[0139] Among them, η represents the learning rate, and Φ(g1, g2, g3, g4) represents the corresponding process of aggregating multiple training gradients based on conflict perception.

[0140] Next, how to generate the combined training gradient Φ(g1, g2, g3, g4) will be further described.

[0141] Refer to Figure 9 , and based on all the gradient cosine similarities and all the training gradients, the combined training gradient is obtained, including the following steps 901 to 905.

[0142] Step 901: When all the gradient cosine similarities are non-negative, accumulate the training gradients to obtain the combined training gradient.

[0143] Step 902: When there is any gradient cosine similarity that is negative, take the two training gradients corresponding to the negative gradient cosine similarity as the first adjustment gradient and the second adjustment gradient respectively.

[0144] Step 903: Based on the first adjustment gradient and the second adjustment gradient, perform a projection process on the first adjustment gradient to obtain an updated first adjustment gradient.

[0145] Step 904: Obtain a first projection weight based on the cosine similarity between the first adjustment gradient and the updated first adjustment gradient.

[0146] Step 905: Perform a weighted process on the corresponding training gradient based on the first projection weight, and accumulate the weighted training gradients to obtain a combined training gradient.

[0147] The following will describe steps 901 to 905 in detail.

[0148] In some embodiments, after obtaining the gradient cosine similarity between every two training gradients, when all gradient cosine similarities are non - negative, it indicates that there is no conflict between the training gradients corresponding to the four types of ability training data in this batch of ability training data. At this time, all the training gradients will be accumulated to obtain a combined training gradient as shown in the following formula (3).

[0149]

[0150] When there is any gradient cosine similarity that is negative, it indicates that there is a conflict between the two training gradients corresponding to the negative gradient cosine similarity. Therefore, the two ends of these conflicting training gradients are respectively used as the first adjustment gradient g i and the second adjustment gradient g j . Next, perform a projection process on these conflicting first adjustment gradients g i to obtain the corresponding updated first adjustment gradient g i ′, so that the cosine similarity between the updated first adjustment gradient g i ′ and the second adjustment gradient is positive, thus overcoming the problem of conflict between training gradients.

[0151] The following will further describe how to perform a projection process on the first adjustment gradient g i .

[0152] Referring to Figure 10 , based on the first adjustment gradient and the second adjustment gradient, performing a projection process on the first adjustment gradient to obtain an updated first adjustment gradient includes the following steps 1001 to 1003.

[0153] Step 1001: Based on the product of the first adjustment gradient and the second adjustment gradient, obtain an adjustment gradient product.

[0154] Step 1002: Based on the ratio of the gradient norm of the adjusted gradient product and the second adjusted gradient, and then multiplying by the second adjusted gradient, obtain the first projection difference.

[0155] Step 1003: Based on the difference between the first adjusted gradient and the first projection difference, obtain the updated first adjusted gradient.

[0156] The following provides a detailed description of Steps 1001 to 1003.

[0157] In some embodiments, first, based on the product of the first adjusted gradient g i and the second adjusted gradient g j , obtain the adjusted gradient product g i ·g j . Then, based on the ratio of the gradient norm ||g i ·g j and the second adjusted gradient g j || j || 2 , multiply by the second adjusted gradient to obtain the projection difference. Then, based on the difference between the first adjusted gradient and the projection difference, obtain the updated first adjusted gradient g i after projection adjustment, as shown in the following formula (4). i ′

[0158]

[0159] It can be understood that for the projection process, all training gradients are recorded in a list G = (g0, g1, g3, g4). Then, traverse the list G. For the training gradient g i , randomly shuffle G to get G'. Then, traverse G'. For the training gradient g j , calculate the cosine similarity (dot product) between the two training gradients (g i and g j ). If it is less than 0, project g i and then perform subsequent traversals. In fact, each training gradient will calculate this process with the other remaining 3 training gradients and sequentially eliminate conflicts with the remaining gradients through projection.

[0160] In some embodiments, after performing projection processing on all the first adjusted gradients g i corresponding to the training gradients with conflicts to obtain the updated first adjusted gradient g i ′, further obtain the updated projection weight <g i and the updated first adjusted gradient g i ′ based on the cosine similarity i , g i ′>, Next, based on the updated projection weight <g i , g i ′>, the corresponding training gradients (i.e., g n , n ∈ {1, 2, 3, 4}) are weighted, and the weighted training gradients are accumulated to obtain the combined training gradient as shown in the following formula (5).

[0161]

[0162] Among them, <·> represents the cosine similarity, and α is a hyperparameter used to control the influence of the conflict degree on the updated projection weight.

[0163] In each iteration, after obtaining the combined training gradient ψ(g0, g1,..., g N ), the update function (2) of the initial model parameter θ mentioned above is used to update the initial model parameter θ of the initial medical record generation model, thereby improving the reliability of co-training.

[0164] Through the above steps 801 to 803, steps 901 to 905, and steps 1001 to 1003, the cosine similarity between the training gradients of multiple ability training sample data in the training iteration process is calculated to accurately identify the gradient direction conflicts between the training gradients; when there are negatively correlated gradients, the projection elimination method is used to decompose the conflicting gradient vectors into orthogonal subspaces to eliminate the direction interference, and the effective components in each training gradient are retained through the cosine similarity weighting mechanism to ensure that the parameter update direction simultaneously satisfies the compatibility of training multiple ability training samples, so as to solve the problems of insufficient summarization ability, unclear medical record structure, and poor entity normalization ability existing in the existing large models, so that high-precision alignment of the feature space is achieved in the training of the initial medical record generation model using multiple ability training sample data, and the abstract extraction ability, format specification ability, language normalization ability, and fluency of the initial medical record generation model are improved simultaneously.

[0165] Step 104: Input the target doctor-patient dialogue sample data into the medical record generation model one by one for data prediction processing to obtain the predicted medical record.

[0166] The following is a detailed description of step 104.

[0167] In some embodiments, after training the initial medical record generation model using multiple batches of multiple ability sample data to obtain the case generation model, the selected target doctor-patient dialogue sample data is input into the medical record generation model one by one for data prediction processing to obtain the predicted medical record.

[0168] In the context of a language model, the true data distribution is the actual distribution of the next word in a sequence of words, while the model prediction distribution is the probability distribution that the model predicts for the next word. During the training phase, the model is optimized by maximizing the probability of predicting the correct next word. Specifically, for each word in a sequence, the model outputs a probability distribution indicating the probability that the next word is any word in the vocabulary.

[0169] Step 105: Update the medical record generation model based on the loss value between the predicted medical record and the target medical record label.

[0170] The following provides a detailed description of Step 105.

[0171] In some embodiments, after obtaining the predicted medical record corresponding to the target doctor-patient dialogue sample data, the loss value between the predicted medical record and the target medical record label corresponding to the target doctor-patient dialogue sample data is calculated to facilitate updating the case generation model using this loss value, thereby improving the accuracy and reliability of medical record generation, as described in detail below.

[0172] Refer to Figure 11 , and update the medical record generation model based on the loss value between the predicted medical record and the target medical record label, including the following steps 1101 to 1103.

[0173] Step 1101: Calculate the cross-entropy loss value between the predicted medical record and the target medical record label based on the cross-entropy loss function.

[0174] Step 1102: Calculate the regularization term value based on the regularization term function and the weight matrix of the medical record generation model.

[0175] Step 1103: Accumulate the cross-entropy loss value and the regularization term value to obtain the loss value, and update the model parameters of the medical record generation model based on the loss value.

[0176] The following provides a detailed description of Steps 1101 to 1103.

[0177] In some embodiments, the cross-entropy loss is used to train the case generation model. This loss function measures the difference between the probability distribution predicted by the model and the probability distribution of the true data. The cross-entropy loss calculates the difference between this probability distribution and the one-hot encoding of the true word, and its calculation formula is shown in the following formula (6).

[0178]

[0179] Among them, yi is the true distribution, usually a one-hot encoded vector, which is 1 at the position of the correct word and 0 at other positions; pi is the probability distribution predicted by the model. The cross-entropy loss calculates the negative logarithm of the predicted probability at the position of the true label. The goal of model training is to minimize this loss function, so that the predicted probability distribution of the model can be as close as possible to the distribution of the true data.

[0180] In practical applications, the cross-entropy loss will also be combined with regularization terms such as L1, L2 regularization or dropout and other techniques to prevent overfitting, and gradient descent or its variant algorithms are used to update the parameters of the model.

[0181] In this embodiment, after calculating the cross-entropy loss value L1 of the predicted medical record and the target medical record label based on the cross-entropy loss function (6), the regularization term value L2 will be further calculated based on the regularization term function and the weight matrix w of the medical record generation model. Then, the cross-entropy loss value L1 and the regularization term value L2 are accumulated to obtain the loss value LL, and the model parameters of the medical record generation model are updated based on the loss value. The regularization term function can be an L1 regularization function or an L2 regularization function.

[0182] Refer to Figure 12 , which is a schematic diagram of post-generation fine-tuning of medical records provided by an embodiment of the present application. As Figure 12 shown, in order to further improve the readability of the generated medical record and the user experience, a prompt is also added before the input information of the case generation model to help the large model evoke relevant capabilities, and a reply identifier is added before the generated medical record.

[0183] Refer to Figure 13 , which is a schematic diagram of the training and use process of a medical record generation model provided by an embodiment of the present application. As Figure 13 shown, the medical record generation model provided by the embodiment of the present application is divided into two major modules: offline training and online inference: the offline training side completes medical domain knowledge enhancement (including subtasks such as abstract extraction, medical record structure understanding, entity normalization, etc.) through an incremental pre-training framework, and constructs a generative model with structured output capabilities; the online inference side realizes an intelligent interaction closed-loop in the form of a dialogue flow - triggers the interrogation process through semantic analysis, dynamically filters invalid dialogues using a self-information evaluation mechanism, and finally generates a compliant electronic medical record based on the pre-trained model.

[0184] In addition, the embodiment of the present application also provides a method for generating medical records. Refer to Figure 14 , which is an optional flowchart of the method for generating medical records provided by the embodiment of the present application. Figure 14 The method in Figure 14 The order of steps 1401 to 1403 is not specifically limited and can be adjusted according to actual needs, or some steps can be reduced or added. The medical record generation method provided in the embodiments of this application can be applied to any server or intelligent terminal connected to the hospital system, etc.

[0185] Step 1401: Display the doctor-patient input prompt, and obtain the doctor-patient conversation data based on the doctor-patient input prompt.

[0186] Step 1402: Based on the end flag in the doctor-patient conversation data, input the doctor-patient conversation data into the medical record generation model trained by the medical record generation model training method for data processing to generate the target medical record.

[0187] Step 1403: Display the target medical record.

[0188] The following provides a detailed description of steps 1401 to 1403.

[0189] In some embodiments, after training the case generation model, in the actual application interface, display the doctor-patient input prompt (such as Figure 12 the "system prompt" shown in Figure 3 ), then obtain the doctor-patient conversation data input by the user according to this prompt, and based on the end flag in the doctor-patient conversation data (such as the "end of consultation" shown in

[0190] Figure 15 ), input the doctor-patient conversation data into the medical record generation model trained by the medical record generation model training method for data processing to generate an accurate and reliable target medical record and display it. Figure 15 Refer to

[0191] Figure 15 is a schematic diagram of the process of medical record generation model service deployment and feedback iteration provided by the embodiments of this application. As Figure 15 shown, the medical record generation model trained in this application can be deployed in a functional module associated with the hospital system, such as the auxiliary consultation function. The auxiliary consultation function can integrate the functions of multiple scenarios, and the large language model can be encapsulated and used, and can be used in the form of an API or in the form of an integrated system. Specifically, in the pre-consultation scenario, it helps doctors quickly understand the basic condition of patients before the patient's diagnosis and helps doctors quickly write medical records. At the same time, by using the consultation large model and accessing the APP and WeChat mini-programs, the ability to provide patients with pre-consultation before the diagnosis is provided. Specifically, patients have a waiting time before the consultation, and even when entering the face-to-face consultation with the doctor, there is only a short time for the consultation. Therefore, providing pre-consultation before the diagnosis helps patients understand their own conditions, helps doctors understand the patient's condition in advance, helps doctors write medical records, and saves doctors' writing time.

[0191] The medical record generation model training method, medical record generation method and related devices proposed in the embodiments of this application, the method includes: First, obtain multiple doctor-patient dialogue sample data and corresponding medical record labels; Secondly, based on the character probability of each character in the doctor-patient dialogue sample data, perform logarithmic processing to obtain the character self-information score, average the character self-information scores in each doctor-patient dialogue sample data to obtain the dialogue score of each doctor-patient dialogue sample data, and select the doctor-patient dialogue sample data with a dialogue score exceeding the preset score threshold from the multiple doctor-patient dialogue sample data as the target doctor-patient dialogue sample data, and the medical record label corresponding to the target doctor-patient dialogue sample data is the target medical record label; Then, based on the summary generation data, calculate the summary training gradient of the initial medical record generation model, based on the medical record structure data, calculate the medical record structure training gradient of the initial medical record generation model, based on the language normalization data, calculate the language normalization training gradient of the initial medical record generation model, based on the interrogation process data, calculate the interrogation process training gradient of the initial medical record generation model, calculate the gradient cosine similarity between every two training gradients, the training gradients include the summary training gradient, the medical record structure training gradient, the language normalization training gradient and the interrogation process training gradient. When all gradient cosine similarities are non-negative, accumulate the training gradients to obtain the combined training gradient. When there is any negative gradient cosine similarity, the two training gradients corresponding to the negative gradient cosine similarity are respectively used as the first adjustment gradient and the second adjustment gradient, based on the product of the first adjustment gradient and the second adjustment gradient, obtain the adjustment gradient product, based on the ratio of the adjustment gradient product to the gradient norm of the second adjustment gradient, and then multiply by the second adjustment gradient to obtain the first projection difference, based on the difference between the first adjustment gradient and the first projection difference, obtain the updated first adjustment gradient, the cosine similarity between the updated first adjustment gradient and the second adjustment gradient is non-negative, obtain the first projection weight based on the cosine similarity between the first adjustment gradient and the updated first adjustment gradient, perform weighted processing on the corresponding training gradient based on the first projection weight, and accumulate the weighted training gradients to obtain the combined training gradient, update the initial model parameters based on the combined training gradient; Next, input the target doctor-patient dialogue sample data into the medical record generation model one by one for data prediction processing to obtain the predicted medical record; Finally, based on the cross-entropy loss function, calculate the cross-entropy loss value between the predicted medical record and the target medical record label, based on the regularization term function and the weight matrix of the medical record generation model, calculate the regularization term value, accumulate the cross-entropy loss value and the regularization term value to obtain the loss value, and update the model parameters of the medical record generation model based on the loss value. The updated medical record generation model is used to generate medical records based on doctor-patient dialogue data.

[0192] Embodiments of the present application pre-screen high-quality doctor-patient dialogue samples based on a dialogue score screening mechanism to ensure that the training data covers diverse expressions in real scenarios. The initial medical record generation model is co-trained using a language specification ability that can improve term standardization and medical record structure understanding, enabling the model to break through the preset rule limitations, automatically identify clinical entities in the patient's colloquial descriptions, and perform structured reorganization according to medical norms. Through the iterative optimization of the predicted medical record and standard labels, a deep semantic mapping from free dialogue to standardized medical records is established to solve the semantic discontinuity problem caused by traditional template filling, achieving accurate extraction of key information without omission, so that the trained medical record generation model has the ability to dynamically adapt to different expression habits and consultation scenarios, fundamentally overcoming the coverage blind spots of fixed rule systems, and then improving the accuracy and reliability of the generated medical records when generating medical records based on actual doctor-patient dialogue data; in addition, the self-information score of the dialogue is calculated using the logarithmic transformation of the character-level probability to effectively capture abnormal expressions in the doctor-patient dialogue, and high-quality dialogue samples with complete semantics and clear logic are accurately identified through the average self-information score of the overall information entropy of the dialogue. Then, the preset threshold screening mechanism eliminates low-quality data containing invalid information from the source, ensuring that the training set strictly follows the professional norms of medical consultations. This data purification strategy enables the model to focus on learning the standardized doctor-patient interaction mode, significantly improving the semantic coherence and clinical accuracy of the generated medical records to solve problems such as term conversion errors and key information omissions caused by interference from low-quality data, so that when training the medical record generation model using the screened target doctor-patient corresponding sample data in the future, the training efficiency and accuracy of the medical record generation model can be improved; and, by calculating the cosine similarity between the training gradients of various ability training sample data during the training iteration process, the gradient direction conflict between the training gradients is accurately identified; when there are negatively correlated gradients, the projection elimination method is used to decompose the conflicting gradient vectors into orthogonal subspaces to eliminate direction interference, and the effective components in each training gradient are retained through the cosine similarity weighting mechanism to ensure that the parameter update direction simultaneously meets the compatibility of training various ability training samples, so as to solve the problems of insufficient summarization ability, unclear medical record structure, and poor entity normalization ability existing in existing large models, enabling high-precision alignment of the feature space in the training of the initial medical record generation model using various ability training sample data, and simultaneously improving the abstract extraction ability, format specification ability, language normalization ability, and fluency of the initial medical record generation model.

[0193] Embodiments of the present application also provide a medical record generation model training device that can implement the above-mentioned medical record generation model training method. Referring to Figure 16 , the device 1600 includes:

[0194] A sample data acquisition module 1610, configured to acquire multiple doctor-patient dialogue sample data and corresponding medical record labels;

[0195] The target sample data screening module 1620 is configured to calculate the dialogue score of each doctor-patient dialogue sample data, and select the target doctor-patient dialogue sample data from multiple doctor-patient dialogue sample data. The medical record label corresponding to the target doctor-patient dialogue sample data is the target medical record label;

[0196] The ability collaborative training module 1630 is configured to perform collaborative training on the initial medical record generation model based on multiple language specification ability sample data to obtain a medical record generation model;

[0197] The data processing module 1640 is configured to input the target doctor-patient dialogue sample data into the medical record generation model one by one for data prediction processing to obtain a predicted medical record;

[0198] The model update module 1650 is configured to update the medical record generation model based on the loss value between the predicted medical record and the target medical record label. The updated medical record generation model is used to generate medical records based on doctor-patient dialogue data.

[0199] In some embodiments, the model update module 1650 is further configured to:

[0200] Based on the cross-entropy loss function, calculate the cross-entropy loss value between the predicted medical record and the target medical record label;

[0201] Based on the regularization term function and the weight matrix of the medical record generation model, calculate the regularization term value;

[0202] Accumulate the cross-entropy loss value and the regularization term value to obtain the loss value, and update the model parameters of the medical record generation model based on the loss value.

[0203] In some embodiments, the target sample data screening module 1620 is further configured to:

[0204] Based on the character probability of each character in the doctor-patient dialogue sample data, perform logarithmic processing to obtain the character self-information score;

[0205] Average the character self-information scores in each doctor-patient dialogue sample data to obtain the dialogue score of each doctor-patient dialogue sample data;

[0206] Select the doctor-patient dialogue sample data with a dialogue score exceeding a preset score threshold from multiple doctor-patient dialogue sample data as the target doctor-patient dialogue sample data.

[0207] In some embodiments, the ability collaborative training module 1630 is further configured to:

[0208] Based on the summary generation data, calculate the summary training gradient of the initial medical record generation model;

[0209] Based on the medical record structure data, calculate the medical record structure training gradient of the initial medical record generation model;

[0210] Based on the language normalization data, calculate the language normalization training gradient of the initial medical record generation model;

[0211] Based on the consultation process data, calculate the consultation process training gradient of the initial medical record generation model;

[0212] Based on the summary training gradient, the medical record structure training gradient, the language normalization training gradient, and the consultation process training gradient, update the initial model parameters of the initial medical record generation model.

[0213] In some embodiments, the ability collaborative training module 1630 is further configured to:

[0214] Calculate the gradient cosine similarity between every two training gradients, where the training gradients include the summary training gradient, the medical record structure training gradient, the language normalization training gradient, and the consultation process training gradient;

[0215] Based on all the gradient cosine similarities and all the training gradients, obtain the combined training gradient;

[0216] Update the initial model parameters based on the combined training gradient.

[0217] In some embodiments, the ability collaborative training module 1630 is further configured to:

[0218] When all the gradient cosine similarities are non - negative, accumulate the training gradients to obtain the combined training gradient;

[0219] When there is any negative gradient cosine similarity, take the two training gradients corresponding to the negative gradient cosine similarity as the first adjustment gradient and the second adjustment gradient respectively;

[0220] Based on the first adjustment gradient and the second adjustment gradient, perform projection processing on the first adjustment gradient to obtain the updated first adjustment gradient, and the cosine similarity between the updated first adjustment gradient and the second adjustment gradient is non - negative;

[0221] Based on the cosine similarity between the first adjustment gradient and the updated first adjustment gradient, obtain the first projection weight;

[0222] Based on the first projection weight, perform weighted processing on the corresponding training gradient, and accumulate the weighted training gradients to obtain the combined training gradient.

[0223] In some embodiments, the ability collaborative training module 1630 is further configured to:

[0224] Based on the product of the first adjustment gradient and the second adjustment gradient, obtain the adjustment gradient product;

[0225] Based on the ratio of the gradient norm of the adjusted gradient product and the second adjusted gradient, and then multiplying by the second adjusted gradient, the first projection difference is obtained;

[0226] Based on the difference between the first adjusted gradient and the first projection difference, the first adjusted gradient is updated.

[0227] In some embodiments, the ability collaborative training module 1630 is further configured to:

[0228] Based on the training gradients corresponding to the first batch of sample data, the initial model parameters of the initial medical record generation model are updated to obtain the first updated model parameters, and the training gradients include the abstract training gradient, the medical record structure training gradient, the language normalization training gradient, and the interrogation process training gradient;

[0229] Based on the training gradients corresponding to the second batch of sample data, the first updated model parameters of the initial medical record generation model are updated, and both the first batch of sample data and the second batch of sample data include at least one abstract generation data, at least one medical record structure data, at least one language normalization data, and at least one interrogation process data.

[0230] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, the specific implementation manners of the medical record generation model training device are basically the same as those of the above medical record generation model training method, and will not be elaborated here.

[0231] In the embodiments of the present application, the medical record generation model training device selects high-quality doctor-patient dialogue samples based on a dialogue score screening mechanism to ensure that the training data covers diverse expressions in real scenarios. It co-trains the initial medical record generation model with a language specification ability that can improve term standardization and medical record structure understanding, enabling the model to break through the preset rule limitations, automatically identify clinical entities in the patient's colloquial descriptions, and perform structured recombination according to medical norms. Through the iterative optimization of the predicted medical record and standard labels, a deep semantic mapping from free dialogue to standardized medical record is established, solving the semantic discontinuity problem caused by traditional template filling, and achieving accurate extraction of key information without omission. As a result, the trained medical record generation model has the ability to dynamically adapt to different expression habits and consultation scenarios, fundamentally overcoming the coverage blind spots of fixed rule systems, and thus improving the accuracy and reliability of the generated medical records when generating medical records based on actual doctor-patient dialogue data. In addition, the self-information score of the dialogue is calculated using the logarithmic transformation of the character-level probability to effectively capture abnormal expressions in the doctor-patient dialogue. High-quality dialogue samples with complete semantics and clear logic are accurately identified through the average self-information score of the overall information entropy of the dialogue. Then, the preset threshold screening mechanism eliminates low-quality data containing invalid information from the source, ensuring that the training set strictly follows the professional norms of medical consultations. This data purification strategy enables the model to focus on learning the standardized doctor-patient interaction mode, significantly improving the semantic coherence and clinical accuracy of the generated medical records, and solving problems such as term conversion errors and key information omission caused by interference from low-quality data. Thus, when using the screened target doctor-patient corresponding sample data for training the medical record generation model subsequently, the training efficiency and accuracy of the medical record generation model can be improved. Moreover, the cosine similarity between the training gradients of various ability training sample data during the training iteration process is calculated to accurately identify the gradient direction conflicts between the training gradients. When there are negatively correlated gradients, the projection elimination method is used to decompose the conflicting gradient vectors into orthogonal subspaces to eliminate direction interference, and the effective components in each training gradient are retained through the cosine similarity weighting mechanism, ensuring that the parameter update direction satisfies the compatibility of training multiple ability training samples simultaneously, so as to solve the problems of insufficient summarization ability, unclear medical record structure, and poor entity normalization ability existing in existing large models, enabling high-precision alignment of the feature space in the training of the initial medical record generation model using various ability training sample data, and simultaneously improving the abstract extraction ability, format specification ability, language normalization ability, and fluency of the initial medical record generation model.

[0232] The embodiments of the present application also provide an electronic device, including:

[0233] At least one memory;

[0234] At least one processor;

[0235] At least one program;

[0236] The program is stored in a memory, and a processor executes the at least one program to implement the medical record generation model training method described above in the embodiments of the present application. The electronic device may be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA for short), an in-vehicle computer, etc.

[0237] Please refer to Figure 17 , Figure 17 which schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:

[0238] A processor 1701, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0239] A memory 1702, which can be implemented in forms such as a ROM (Read Only Memory), a static storage device, a dynamic storage device, or a RAM (Random Access Memory). The memory 1702 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1702 and are called by the processor 1701 to execute the medical record generation model training method of the embodiments of the present application;

[0240] An input / output interface 1703, which is used to implement information input and output;

[0241] A communication interface 1704, which is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0242] A bus 1705, which transmits information between various components of the device (such as the processor 1701, the memory 1702, the input / output interface 1703, and the communication interface 1704);

[0243] Among them, the processor 1701, the memory 1702, the input / output interface 1703, and the communication interface 1704 are communicatively connected to each other inside the device through the bus 1705.

[0244] The embodiments of the present application also provide a storage medium, which is a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned medical record generation model training method is implemented.

[0245] As a non-transitory computer-readable storage medium, a memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0246] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0247] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0248] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0249] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0250] In the description of this application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0251] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0252] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0253] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0254] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0255] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0256] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.< / end>

Claims

1. A medical record generation model training method, characterized in that: The method comprises: Obtain multiple doctor-patient conversation sample data and corresponding medical record labels; Calculate the conversation score of each of the doctor-patient conversation sample data, and select target doctor-patient conversation sample data from the multiple doctor-patient conversation sample data based on the conversation score, and the medical record label corresponding to the target doctor-patient conversation sample data is the target medical record label; Based on multiple language standard ability sample data, the initial medical record generation model is collaboratively trained to obtain a medical record generation model; Inputting the target doctor-patient conversation sample data into the medical record generation model one by one for data prediction processing to obtain a predicted medical record; Based on the loss values ​​of the predicted medical records and the target medical record labels, the medical record generation model is updated, and the updated medical record generation model is used to generate medical records based on doctor-patient conversation data.

2. The medical record generation model training method according to claim 1, characterized in that: The updating of the medical record generation model based on the loss value of the predicted medical record and the target medical record label includes: Based on the cross entropy loss function, the cross entropy loss value of the predicted medical record and the target medical record label is calculated; Based on the regularization term function and the weight matrix of the medical record generation model, a regularization term value is calculated; The cross entropy loss value and the regularization term value are accumulated to obtain the loss value, and the model parameters of the medical record generation model are updated based on the loss value.

3. The medical record generation model training method according to claim 1, characterized in that: The calculating of the dialogue score of each of the doctor-patient dialogue sample data, and selecting target doctor-patient dialogue sample data from the plurality of doctor-patient dialogue sample data based on the dialogue score, comprises: Based on the character probability of each character in the doctor-patient dialogue sample data, logarithmic processing is performed to obtain a character self-information score; Averaging the character self-information scores in each piece of doctor-patient dialogue sample data to obtain the dialogue score for each piece of doctor-patient dialogue sample data; The doctor-patient dialogue sample data whose dialogue score exceeds a preset score threshold is selected from the multiple doctor-patient dialogue sample data as the target doctor-patient dialogue sample data.

4. The medical record generation model training method according to claim 1, characterized in that: The language standardization capability sample data includes summary generation data, medical record structure data, language normalization data, and consultation process data. The initial medical record generation model is collaboratively trained based on multiple language standardization capability sample data to obtain a medical record generation model, including: Based on the summary generation data, a summary training gradient of an initial medical record generation model is calculated; Based on the medical record structure data, calculating the medical record structure training gradient of the initial medical record generation model; Based on the language normalization data, calculating the language normalization training gradient of the initial medical record generation model; Based on the consultation process data, calculating the consultation process training gradient of the initial medical record generation model; Based on the summary training gradient, the medical record structure training gradient, the language normalization training gradient, and the consultation process training gradient, the initial model parameters of the initial medical record generation model are updated.

5. The medical record generation model training method according to claim 4, characterized in that: The updating of the initial model parameters of the initial medical record generation model based on the summary training gradient, the medical record structure training gradient, the language normalization training gradient, and the consultation process training gradient includes: Calculating the gradient cosine similarity between every two training gradients, the training gradients including the summary training gradient, the medical record structure training gradient, the language normalization training gradient, and the consultation process training gradient; Obtaining a combined training gradient based on all of the gradient cosine similarities and all of the training gradients; The initial model parameters are updated based on the combined training gradient.

6. The medical record generation model training method according to claim 5, characterized in that: The step of obtaining a combined training gradient based on all the gradient cosine similarities and all the training gradients includes: When all the gradient cosine similarities are non-negative, accumulating the training gradients to obtain the combined training gradient; When any one of the gradient cosine similarities is a negative value, two training gradients corresponding to the negative gradient cosine similarity are used as a first adjustment gradient and a second adjustment gradient respectively; Based on the first adjustment gradient and the second adjustment gradient, performing a projection process on the first adjustment gradient to obtain an updated first adjustment gradient, wherein the cosine similarity between the updated first adjustment gradient and the second adjustment gradient is non-negative; Obtaining a first projection weight based on the cosine similarity of the first adjustment gradient and the updated first adjustment gradient; The corresponding training gradients are weighted based on the first projection weights, and the weighted training gradients are accumulated to obtain the combined training gradients.

7. The medical record generation model training method according to claim 6, characterized in that: The projecting the first adjustment gradient based on the first adjustment gradient and the second adjustment gradient to obtain an updated first adjustment gradient includes: Obtaining an adjustment gradient product based on a product of the first adjustment gradient and the second adjustment gradient; A first projection difference is obtained by multiplying the ratio of the adjusted gradient product to the gradient norm of the second adjusted gradient by the second adjusted gradient; The updated first adjustment gradient is obtained based on a difference between the first adjustment gradient and the first projection difference.

8. The medical record generation model training method according to claim 4, characterized in that: The language standardization ability sample data includes a first batch of sample data and a second batch of sample data, and the updating of the initial model parameters of the initial medical record generation model based on the summary training gradient, the medical record structure training gradient, the language normalization training gradient, and the consultation process training gradient also includes: Based on the training gradient corresponding to the first batch of sample data, the initial model parameters of the initial medical record generation model are updated to obtain first updated model parameters, wherein the training gradient includes the summary training gradient, the medical record structure training gradient, the language normalization training gradient, and the consultation process training gradient; The first updated model parameters of the initial medical record generation model are updated based on the training gradient corresponding to the second batch of sample data, and the first batch of sample data and the second batch of sample data both include at least one summary generation data, at least one medical record structure data, at least one language normalization data and at least one consultation process data.

9. A method for generating medical records, characterized in that: The method comprises: Displaying the doctor-patient input prompt, and acquiring the doctor-patient dialogue data based on the doctor-patient input prompt; Based on the end mark in the doctor-patient conversation data, inputting the doctor-patient conversation data into the medical record generation model trained by the medical record generation model training method according to claim 1 for data processing to generate a target medical record; The target medical record is displayed.

10. A medical record generation model training device, characterized in that: The device comprises: The sample data acquisition module is used to obtain multiple doctor-patient dialogue sample data and corresponding medical record labels; A target sample data screening module is used to calculate the dialogue score of each of the doctor-patient dialogue sample data, and select target doctor-patient dialogue sample data from the multiple doctor-patient dialogue sample data based on the dialogue score, and the medical record label corresponding to the target doctor-patient dialogue sample data is the target medical record label; A capability collaborative training module, used to collaboratively train the initial medical record generation model based on multiple language standard capability sample data to obtain a medical record generation model; A data processing module, used to input the target doctor-patient conversation sample data into the medical record generation model one by one to perform data prediction processing to obtain a predicted medical record; A model updating module is used to update the medical record generation model based on the loss values ​​of the predicted medical record and the target medical record label, and the updated medical record generation model is used to generate medical records based on doctor-patient conversation data.

11. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the medical record generation model training method described in any one of claims 1 to 8 or the medical record generation method described in claim 9 is implemented.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the medical record generation model training method according to any one of claims 1 to 8 or the medical record generation method according to claim 9.

Citation Information

Cited By

  • Medical record generation method based on large model and generation method of large model of medical record

    CN117747036A

  • Medical record generation method based on large model, and method for generating large model of medical record

    CN117747036B

  • Method and device for training large language model

    CN120509491A

  • Method and device for training large language model

    CN120509491B

  • Electronic medical record generation method, device and system, electronic equipment and storage medium

    CN121905406A