Medical record generation method based on large model, and method for generating large model of medical record

By filtering doctor-patient conversation data through character scores and training large models, the problem of insufficient quality of medical record generation in online consultation services was solved, and the generation of high-quality medical records was achieved.

CN117747036BActive Publication Date: 2025-09-19BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311767107.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-09-19
Estimated Expiration
2043-12-20

AI Technical Summary

Technical Problem

Existing medical record generation technology has difficulty in generating high-quality medical records in online consultation services, especially when processing doctor-patient conversation data, where there is a problem of useless or interfering information affecting the generation quality.

Method used

By obtaining the doctor-patient conversation data to be processed, calculating the score of each character and filtering it, obtaining the filtered data and inputting it into the big model, combining the prompt information to generate medical records, and improving the quality of medical record generation by pre-training and update training of the big model.

Benefits of technology

It improves the quality of medical record generation, reduces the impact of useless or interfering information, enhances the large model's ability to understand the medical record structure, and improves the accuracy and reliability of generated medical records.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117747036B_ABST
    Figure CN117747036B_ABST
Patent Text Reader

Abstract

This disclosure proposes a method for generating medical records based on a large model, and a method for generating a large model of medical records. The method relates to the fields of computer technology, particularly to artificial intelligence technologies such as natural language processing, deep learning, and large models, and can be applied to scenarios such as clinical medical records, medical research, medical insurance, and telemedicine. The specific scheme is as follows: first, obtain the first consultation data to be processed, then process the first consultation data to obtain a first score for each character in the first consultation data, then filter the first consultation data based on the first score of each character to obtain filtered data, and finally input the filtered data and prompt information into the large model to obtain the first medical record output by the large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, in particular to the fields of artificial intelligence technology such as natural language processing, deep learning, and large models, and specifically to a method for generating medical records based on large models and a method for generating large models of medical records. Background Art

[0002] Medical record generation technology is a crucial technology in the healthcare field, aiming to solve the problem of automatically generating and managing medical records. With the widespread use of online consultation services and the collaborative development of healthcare systems, the requirements for medical record generation technology are becoming increasingly stringent. Summary of the Invention

[0003] The present disclosure aims to solve one of the technical problems in the related art at least to a certain extent.

[0004] The first embodiment of the present disclosure provides a method for generating medical records based on a large model, comprising:

[0005] Obtaining first medical consultation data to be processed;

[0006] processing the first medical inquiry data to obtain a first score for each character in the first medical inquiry data;

[0007] filtering the first medical inquiry data according to the first score of each character to obtain filtered data;

[0008] The filtered data and prompt information are input into the large model to obtain the first medical record output by the large model.

[0009] A second embodiment of the present disclosure provides a method for generating a large model for generating medical records, comprising:

[0010] Acquire a first data set and a second data set, wherein the first data set includes at least one training data associated with a non-medical record generation task, and the second data set includes a plurality of second medical consultation data and associated second medical records;

[0011] Pre-training an initial large model based on the first data set to obtain a reference large model;

[0012] The reference large model is updated and trained based on the second data set to obtain an updated large model.

[0013] The third embodiment of the present disclosure provides a medical record generation device based on a large model, comprising:

[0014] A first acquisition module, configured to acquire first medical inquiry data to be processed;

[0015] a first processing module, configured to process the first medical inquiry data to obtain a first score for each character in the first medical inquiry data;

[0016] a filtering module, configured to filter the first medical inquiry data according to the first score of each character to obtain filtered data;

[0017] The second processing module is used to input the filtered data and prompt information into the large model to obtain the first medical record output by the large model.

[0018] A fourth embodiment of the present disclosure provides a device for generating a large model of medical records, comprising:

[0019] A second acquisition module is configured to acquire a first data set and a second data set, wherein the first data set includes at least one training data associated with a non-medical record generation task, and the second data set includes a plurality of second medical consultation data and associated second medical records;

[0020] A first training module is configured to pre-train an initial large model based on the first data set to obtain a reference large model;

[0021] The second training module is used to update the reference large model based on the second data set to obtain an updated large model.

[0022] The fifth aspect embodiment of the present disclosure proposes a computer device, including: a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements the medical record generation method based on the large model proposed in the first aspect embodiment of the present disclosure and the method for generating a large model for generating medical records proposed in the second aspect embodiment of the present disclosure.

[0023] The sixth aspect embodiment of the present disclosure proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the medical record generation method based on a large model proposed in the first aspect embodiment of the present disclosure and the method for generating a large model for generating medical records proposed in the second aspect embodiment of the present disclosure.

[0024] The seventh aspect embodiment of the present disclosure proposes a computer program product, including a computer program. When the computer program is executed by a processor, it implements the medical record generation method based on the big model proposed in the first aspect embodiment of the present disclosure and the method for generating a big model for generating medical records proposed in the second aspect embodiment of the present disclosure.

[0025] The method for generating medical records based on a large model and the method for generating a large model of medical records provided by the present disclosure have the following beneficial effects:

[0026] In the disclosed embodiment, after obtaining the first medical inquiry data to be processed, the first medical inquiry data is first processed to obtain a first score for each character in the first medical inquiry data. Then, based on the first score of each character, the first medical inquiry data is filtered to obtain filtered data. Finally, the filtered data and prompt information are input into the large model to obtain the first medical record output by the large model. The large model can be obtained by the following training method: by obtaining a first data set and a second data set, and pre-training an initial large model based on the first data set to obtain a reference large model, and then updating and training the reference large model based on the second data set to obtain an updated large model. As a result, the quality of the medical records generated by the large model is improved.

[0027] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0029] Figure 1 A flowchart of a method for generating medical records based on a large model according to an embodiment of the present disclosure;

[0030] Figure 2 A flowchart of a method for generating medical records based on a large model according to an embodiment of the present disclosure;

[0031] Figure 3 A flowchart of a method for generating medical records based on a large model according to an embodiment of the present disclosure;

[0032] Figure 4 A flowchart of a method for generating medical records based on a large model according to an embodiment of the present disclosure;

[0033] Figure 5 A flowchart of a method for generating medical records based on a large model according to an embodiment of the present disclosure;

[0034] Figure 6 A flowchart of a method for generating medical records based on a large model according to an embodiment of the present disclosure;

[0035] Figure 7 A flowchart of a method for generating a large model for generating medical records provided by an embodiment of the present disclosure;

[0036] Figure 8 A flowchart of a method for generating a large model for generating medical records provided by an embodiment of the present disclosure;

[0037] Figure 9 A schematic diagram of the process of large-model-based medical record generation and large-model incremental pre-training provided in an embodiment of the present disclosure;

[0038] Figure 10 A schematic structural diagram of a medical record generation device based on a large model provided in an embodiment of the present disclosure;

[0039] Figure 11 A schematic diagram of the structure of a device for generating a large model of medical records provided by an embodiment of the present disclosure;

[0040] Figure 12 A block diagram of an exemplary computer device suitable for implementing embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0041] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0042] The present disclosure relates to artificial intelligence technology fields such as natural language processing, deep learning, and large models.

[0043] Artificial Intelligence (AI) is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence.

[0044] Natural Language Processing (NLP) is an interdisciplinary subject in the fields of computer science, artificial intelligence, and linguistics. It mainly studies how to enable computers to understand, process, generate, and simulate human language, so as to achieve the ability to have natural conversations with humans.

[0045] Deep learning involves learning the inherent patterns and representational hierarchies of sample data. The information gained from this learning process is highly helpful in interpreting data such as text, images, and sounds. The ultimate goal of deep learning is to enable machines to have the same analytical and learning capabilities as humans, enabling them to recognize data such as text, images, and sounds.

[0046] The large model can also be called the Foundation Model. The model extracts knowledge from billions of corpora or images, learns, and then produces a large model with billions of parameters.

[0047] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0048] The following describes a method for generating a medical record based on a large model and a method for generating a large model of a medical record according to an embodiment of the present disclosure with reference to the accompanying drawings.

[0049] Figure 1 A flowchart of a method for generating medical records based on a large model provided in one embodiment of the present disclosure.

[0050] like Figure 1 As shown, the medical record generation method based on the large model may include the following steps:

[0051] Step 101: Acquire first medical inquiry data to be processed.

[0052] Among them, the first consultation data can be doctor-patient conversation data.

[0053] It should be noted that when the first medical consultation data is doctor-patient conversation data, it may be online doctor-patient conversation data, or it may be offline doctor-patient conversation data, which may include at least one doctor-patient conversation segment, etc., and this disclosure does not limit this.

[0054] Step 102: Process the first medical inquiry data to obtain a first score for each character in the first medical inquiry data.

[0055] The first score may be the self-information score of each character in the first medical inquiry data.

[0056] In the present disclosure, after obtaining the first medical inquiry data to be processed, each character in the first medical inquiry data can be scored using a language model to obtain a first score corresponding to each character. The formula for calculating the first score I(x) of each character is:

[0057] I(x)=-log2P(x t x0,x1,x2…x t-1 ) (1)

[0058] Among them, x t Indicates the t+1th character; x0, x1, x2…x t-1 represents all characters from the first character of the first medical consultation data to the t+1th character; P(x t |x o ,x1,x2…x t-1 ) means the t+1th character is x t probability.

[0059] Step 103: Filter the first medical inquiry data according to the first score of each character to obtain filtered data.

[0060] In the present disclosure, after obtaining the first score of each character in the first medical inquiry data, in order to obtain characters in the first medical inquiry data that are more closely associated with the condition, thereby improving the quality of the first medical inquiry data, the first medical inquiry data can be filtered based on the first score of each character to obtain filtered data.

[0061] In some possible implementation forms, characters in the first medical inquiry data whose first scores are lower than a first score threshold may be deleted to obtain filtered data.

[0062] The first score threshold is a first score critical value used to determine whether to delete any character in the first medical inquiry data. It can be preset and is not limited in the present disclosure.

[0063] In the present disclosure, when the first score corresponding to any character in the first medical consultation data is lower than the first score threshold, it can be considered that the degree of correlation between the any character and the condition in the first medical consultation data is low. In order to improve the quality of the first medical consultation data, the any character can be deleted to obtain filtered data.

[0064] In some possible implementation forms, the first scores of each character contained in each sentence of the first medical consultation data can be first merged to obtain the second score of the sentence, and then the sentences in the first medical consultation data whose second scores are lower than the second score threshold can be deleted to obtain filtered data.

[0065] The second score threshold is a second score critical value used to determine whether to delete any sentence in the first medical inquiry data. It can be preset and is not limited in the present disclosure.

[0066] It should be noted that when the first scores of each character contained in each sentence of the first medical consultation data are fused, the second score of the corresponding sentence can be obtained by calculating the average of the first scores of all characters contained in each sentence in the first medical consultation data. This disclosure does not limit this.

[0067] In the present disclosure, when the second score of a certain sentence in the first medical consultation data is lower than the second score threshold, it can be considered that the sentence has a low degree of correlation with the condition in the first medical consultation data. In order to improve the quality of the first medical consultation data, the sentence can be deleted to obtain filtered data.

[0068] Step 104: input the filtered data and prompt information into the large model to obtain the first medical record output by the large model.

[0069] The prompt information is information used to prompt the medical record type corresponding to the filtered data.

[0070] In the present disclosure, after obtaining the filtered data, in order to improve the big model's ability to understand the medical record structure, thereby improving the quality of medical records generated by the big model, the filtered data and prompt information can be spliced ​​and input into the big model to obtain the first medical record output by the big model.

[0071] In the disclosed embodiment, after obtaining the first medical inquiry data to be processed, the first medical inquiry data is first processed to obtain a first score for each character in the first medical inquiry data. The first medical inquiry data is then filtered based on the first score for each character to obtain filtered data. Finally, the filtered data and prompt information are input into the macro model to obtain a first medical record output by the macro model. Thus, by filtering the medical inquiry data based on the medical inquiry data itself and then inputting the filtered data and prompt information into the macro model to obtain the medical record, the impact of useless or interfering information on the macro model is reduced, thereby improving the quality of the medical record generated by the macro model.

[0072] Figure 2 A flowchart of a method for generating medical records based on a large model provided in one embodiment of the present disclosure.

[0073] like Figure 2 As shown, the medical record generation method based on the large model may include the following steps:

[0074] Step 201: Acquire first medical inquiry data to be processed.

[0075] The specific implementation of step 201 can refer to the detailed description in other embodiments of the present disclosure and will not be described in detail here.

[0076] Step 202: predict the (i+1)th predicted character based on the first i characters in the first medical inquiry data, where i is a positive integer.

[0077] The predicted characters are characters predicted by the language model.

[0078] In the present disclosure, after obtaining the first medical inquiry data, the i+1th character may be predicted based on the contents of the first i characters in the first medical inquiry data through a language model.

[0079] Step 203 : Determine a first score for the i+1th character based on the difference between the i+1th actual character in the first medical inquiry data and the i+1th predicted character.

[0080] The actual characters are the characters contained in the first medical inquiry data.

[0081] In the present disclosure, after obtaining the i+1th predicted character, a first score of the i+1th character may be determined based on the difference between the i+1th actual character in the first medical inquiry data and the i+1th predicted character.

[0082] It should be noted that the difference between the i+1th actual character and the i+1th predicted character can be determined by the cosine distance between the vectors corresponding to the two characters, or can also be determined by the similarity between the two characters, etc., and this disclosure does not limit this.

[0083] Step 204 : Filter the first medical inquiry data according to the first score of each character to obtain filtered data.

[0084] Step 205: Input the filtered data and prompt information into the large model to obtain the first medical record output by the large model.

[0085] The specific implementation of steps 204 to 205 can refer to the detailed descriptions in other embodiments of the present disclosure and will not be described in detail here.

[0086] In the disclosed embodiment, after obtaining the first medical inquiry data to be processed, the first predicted character (i+1) is predicted based on the first i characters in the first medical inquiry data. Then, based on the difference between the actual character (i+1) in the first medical inquiry data and the predicted character (i+1), a first score for the character (i+1) is determined. The first medical inquiry data is then filtered based on the first score of each character to obtain filtered data. Finally, the filtered data and prompt information are input into the large model to obtain the first medical record output by the large model. Thus, by determining the score for the character (i+1) based on the difference between the predicted character (i+1) and the actual character (i+1) in the first medical inquiry data, and filtering the medical inquiry data based on the character scores, the quality of the medical inquiry data is improved, providing conditions for improving the quality of the medical records generated by the large model.

[0087] Figure 3 A flowchart of a method for generating medical records based on a large model provided in one embodiment of the present disclosure.

[0088] like Figure 3 As shown, the medical record generation method based on the large model may include the following steps:

[0089] Step 301: Acquire the first medical inquiry data to be processed.

[0090] The specific implementation of step 301 can be referred to the detailed description in other embodiments of the present disclosure, and will not be described in detail here.

[0091] Step 302: Based on the first i characters in the first medical inquiry data, determine the candidate character sequence corresponding to the i+1th character and the probability value of each candidate character, where i is a positive integer.

[0092] The candidate character sequence corresponding to the (i+1)th character is a sequence consisting of possible candidate characters for the (i+1)th character in the first medical inquiry data.

[0093] The probability value of a candidate character is the probability that the i+1th character is the candidate character.

[0094] Step 303: When the (i+1)th actual character in the first medical inquiry data is any candidate character in the candidate character sequence, determine a first score for the (i+1)th actual character according to the probability value of any candidate character.

[0095] In the present disclosure, when the i+1th actual character in the first medical inquiry data is any candidate character in the candidate character sequence, the first score of the i+1th actual character can be determined based on the probability value of this any candidate character and formula (1).

[0096] Step 304: Filter the first medical inquiry data according to the first score of each character to obtain filtered data.

[0097] Step 305: Input the filtered data and prompt information into the big model to obtain the first medical record output by the big model.

[0098] The specific implementation of steps 304 to 305 can refer to the detailed descriptions in other embodiments of the present disclosure and will not be described in detail here.

[0099] In the disclosed embodiment, after obtaining the first medical inquiry data to be processed, the candidate character sequence corresponding to the i+1th character and the probability value of each candidate character are first determined based on the first i characters in the first medical inquiry data. If the i+1th actual character in the first medical inquiry data is any candidate character in the candidate character sequence, the first score of the i+1th actual character is determined based on the probability value of any candidate character. Then, based on the first score of each character, the first medical inquiry data is filtered to obtain the filtered data. Finally, the filtered data and prompt information are input into the large model to obtain the first medical record output by the large model. Thus, if any character is any candidate character, the score of any character is determined based on the probability value of each candidate character corresponding to the any character, and the medical inquiry data is filtered based on the character score, thereby providing conditions for improving the accuracy of the first medical inquiry data.

[0100] Figure 4A flowchart of a method for generating medical records based on a large model provided in one embodiment of the present disclosure.

[0101] like Figure 4 As shown, the medical record generation method based on the large model may include the following steps:

[0102] Step 401: Acquire the first medical inquiry data to be processed.

[0103] Step 402: Based on the first i characters in the first medical inquiry data, determine the candidate character sequence corresponding to the i+1th character and the probability value of each candidate character, where i is a positive integer.

[0104] The specific implementation of steps 401 to 402 can refer to the detailed descriptions in other embodiments of the present disclosure and will not be described in detail here.

[0105] Step 403 : When the (i+1)th actual character in the first medical inquiry data is different from each candidate character in the candidate character sequence, determining a first score of the (i+1)th actual character to be a specified value.

[0106] The designated value may be preset, for example, 0, etc., which is not limited in the present disclosure.

[0107] In the present disclosure, when the i+1th actual character in the first medical consultation data is different from each candidate character in the candidate character sequence, it can be considered that the i+1th actual character has a low degree of correlation with the condition in the first medical consultation data. At this time, the first score of the i+1th actual character can be determined as a specified value, such as 0, etc.

[0108] Step 404: Filter the first medical inquiry data according to the first score of each character to obtain filtered data.

[0109] Step 405: Input the filtered data and prompt information into the large model to obtain the first medical record output by the large model.

[0110] The specific implementation of steps 404 to 405 can refer to the detailed descriptions in other embodiments of the present disclosure and will not be described in detail here.

[0111] In the disclosed embodiment, after obtaining the first medical inquiry data to be processed, the candidate character sequence corresponding to the i+1th character and the probability value of each candidate character are first determined based on the first i characters in the first medical inquiry data. If the i+1th actual character in the first medical inquiry data is different from each candidate character in the candidate character sequence, the first score of the i+1th actual character is determined to be a specified value. Then, based on the first score of each character, the first medical inquiry data is filtered to obtain the filtered data. Finally, the filtered data and prompt information are input into the large model to obtain the first medical record output by the large model. Thus, if any character in the first medical inquiry data is not any candidate character, by directly determining the score of this character as the specified value, it provides conditions for improving the quality of the medical record output by the large model.

[0112] Figure 5 A flowchart of a method for generating medical records based on a large model provided in one embodiment of the present disclosure.

[0113] like Figure 5 As shown, the medical record generation method based on the large model may include the following steps:

[0114] Step 501: Acquire the first medical inquiry data to be processed.

[0115] The specific implementation of step 501 can refer to the detailed descriptions in other embodiments of the present disclosure and will not be described in detail here.

[0116] Step 502: Determine the similarity between the first medical inquiry data and reference words associated with medical records of different types of structures.

[0117] In the present disclosure, after obtaining the first medical inquiry data, the similarity between the first medical inquiry data and the reference words may be determined based on the reference words associated with different types of structural medical records.

[0118] It should be noted that, depending on the type and structure of the medical records, the associated reference words may be the same or different, and this disclosure does not limit this.

[0119] Step 503: Determine the target type to which the first medical record belongs based on the similarity.

[0120] The target type may be any of the following: admission record, discharge record, medical record, operation record, test report, etc., which is not limited in this disclosure.

[0121] In the present disclosure, when determining the target type to which the first medical record belongs based on the similarity between the first consultation data and the reference words of medical records of different types of structures, the medical record structure type corresponding to the reference words with higher similarity can be determined as the target type to which the first medical record belongs.

[0122] Step 504: Obtain prompt information associated with the target type.

[0123] In the present disclosure, after determining the target type of the first medical record, prompt information associated with the target type can be obtained. For example, when the target type of the first medical record is a medical record, the associated prompt information can be "This is a medical record of the medical record type," etc., which is not limited in the present disclosure.

[0124] Step 505: Process the first medical inquiry data to obtain a first score for each character in the first medical inquiry data.

[0125] Step 506: Filter the first medical inquiry data according to the first score of each character to obtain filtered data.

[0126] Step 507: Input the filtered data and prompt information into the large model to obtain the first medical record output by the large model.

[0127] The specific implementation of steps 505 to 507 can refer to the detailed descriptions in other embodiments of the present disclosure and will not be described in detail here.

[0128] In the disclosed embodiment, after obtaining the first medical inquiry data to be processed, the similarity between the first medical inquiry data and reference words associated with different types of structural medical records is first determined. Based on the similarity, the target type to which the first medical record belongs is determined, and prompt information associated with the target type is obtained. The first medical inquiry data is then processed to obtain a first score for each character in the first medical inquiry data. The first medical inquiry data is then filtered based on the first score of each character to obtain filtered data. Finally, the filtered data and prompt information are input into the large model to obtain the first medical record output by the large model. Thus, by determining the target type to which the first medical record belongs and the associated prompt information based on the similarity between the first medical inquiry data and reference words associated with different types of structural medical records, and inputting the filtered data and prompt information into the large model, the large model's ability to understand the medical record structure is improved, thereby improving the quality of the medical records generated by the large model.

[0129] Figure 6 A flowchart of a method for generating medical records based on a large model provided in one embodiment of the present disclosure.

[0130] like Figure 6 As shown, the medical record generation method based on the large model may include the following steps:

[0131] Step 601: Acquire the first medical inquiry data to be processed.

[0132] Step 602: Process the first medical inquiry data to obtain a first score for each character in the first medical inquiry data.

[0133] Step 603: Filter the first medical inquiry data according to the first score of each character to obtain filtered data.

[0134] The specific implementation of steps 601 to 603 can be referred to the detailed descriptions in other embodiments of the present disclosure, and will not be described in detail here.

[0135] Step 604: Receive a prompt information selection instruction, wherein the selection instruction includes a target type associated with the prompt information to be selected.

[0136] In the present disclosure, after receiving the prompt information selection instruction, the target type associated with the prompt information to be selected can be determined.

[0137] Step 605: Obtain prompt information associated with the target type.

[0138] Step 606: Input the filtered data and prompt information into the big model to obtain the first medical record output by the big model.

[0139] The specific implementation of steps 605 to 606 can refer to the detailed descriptions in other embodiments of the present disclosure and will not be described in detail here.

[0140] In the disclosed embodiment, after obtaining the first medical inquiry data to be processed, the first medical inquiry data is first processed to obtain a first score for each character in the first medical inquiry data. The first medical inquiry data is then filtered based on the first score for each character to obtain filtered data. A prompt information selection instruction is then received to obtain prompt information associated with the target type. Finally, the filtered data and prompt information are input into the large model to obtain a first medical record output by the large model. Thus, by determining the prompt information based on the received prompt information selection instruction and inputting the filtered data and prompt information into the large model, the large model's ability to understand the medical record structure is improved, thereby improving the quality of the medical records generated by the large model.

[0141] Figure 7 A flowchart of a method for generating a large model for generating medical records provided in one embodiment of the present disclosure.

[0142] like Figure 7 As shown, the method for generating a large model for generating medical records may include the following steps:

[0143] Step 701: Acquire a first data set and a second data set, wherein the first data set includes at least one training data associated with a non-medical record generation task, and the second data set includes a plurality of second consultation data and associated second medical records.

[0144] Among them, the first data set is a data set used for incremental pre-training of the large model.

[0145] In some possible implementations, non-medical record generation tasks may include: summary generation tasks, medical record type understanding tasks, and entity normalization tasks.

[0146] Among them, the second data set is a data set used to update the large model after incremental pre-training.

[0147] In this disclosure, when the non-medical record generation task is a summary generation task, in order to improve the summarization and induction capabilities of the large model, a large number of publicly available summary generation datasets can be obtained as training data associated with the summary generation task. For example, the training data associated with the summary generation task can include any summary dataset, such as the Large-Scale Chinese Short Text (LCST) dataset and the Chinese Scientific Literature (CSL) summary dataset, and this disclosure does not limit this.

[0148] When the non-medical record generation task is a medical record type understanding task, in order to improve the large model's ability to understand the medical record structure, a large number of high-quality medical records can be obtained, and the <medical record type, medical record content> corresponding to these medical records can be spliced ​​as training data associated with the medical record type understanding task.

[0149] When the non-medical record generation task is an entity normalization task, in order to improve the standardization of entity expressions in medical records generated by the large model, a data set can be formed by obtaining a large number of colloquial entity expressions and their corresponding standardized entity expressions. For example, the composed data set can be <The standardized expression of ×× entity in the medical record is..., the standardized expression entity corresponding to ×× entity>, which serves as training data associated with the entity normalization task.

[0150] In the present disclosure, pre-training the model based on at least one non-medical record generation task can improve the quality of medical records generated by the large model.

[0151] Step 702: Pre-train the initial large model based on the first data set to obtain a reference large model.

[0152] The reference large model is a large model obtained after pre-training the large model.

[0153] In the present disclosure, after obtaining the first data set, the training data in the first data set can be first spliced, and then the spliced ​​training data can be input into the large model, and the large model can be corrected, thereby achieving pre-training of the large model, that is, incremental pre-training.

[0154] Step 703: Update and train the reference large model based on the second data set to obtain an updated large model.

[0155] Among them, update training is supervised fine-tuning training for large models.

[0156] In the present disclosure, after obtaining the second data set, supervised fine-tuning training can be performed on the reference large model based on the second data set to obtain an updated large model.

[0157] In the disclosed embodiment, after obtaining the first and second datasets, the initial large model is first pre-trained based on the first dataset to obtain a reference large model. The reference large model is then updated and trained based on the second dataset to obtain an updated large model. Thus, by incrementally pre-training the large model based on training data associated with the summary generation task, the medical record structure understanding task, and the entity normalization task, and then performing supervised fine-tuning training on the pre-trained large model, the accuracy and reliability of the large model for generating medical records are improved.

[0158] Figure 8 A flowchart of a method for generating a large model for generating medical records provided in one embodiment of the present disclosure.

[0159] like Figure 8 As shown, the method for generating a large model for generating medical records may include the following steps:

[0160] Step 801 , obtaining a first data set and a second data set, wherein the first data set includes at least one training data associated with a non-medical record generation task, and the second data set includes a plurality of second medical consultation data and associated second medical records.

[0161] Step 802: pre-train the initial large model based on the first data set to obtain a reference large model.

[0162] The specific implementation of steps 801 to 802 can refer to the detailed descriptions in other embodiments of the present disclosure and will not be described in detail here.

[0163] Step 803: Determine prompt information according to the type of the second medical record.

[0164] It should be noted that the prompt information associated with medical records of different types of structures in the present disclosure may be pre-set. After obtaining the second data set, the corresponding prompt information may be determined directly based on the type structure of the second medical record.

[0165] Step 804: input the prompt information and the second medical inquiry data into the reference macro model to obtain a third medical record generated by the reference macro model.

[0166] In the present disclosure, after obtaining the prompt information corresponding to the second medical record, in order to update and train the reference large model, the prompt information and the corresponding second medical consultation data can be input into the reference large model to obtain the third medical record generated by the reference large model.

[0167] Step 805 : Modify the reference macro model according to the difference between the third medical record and the second medical record to obtain an updated macro model.

[0168] In the present disclosure, after obtaining the third medical record generated by the reference large model, the loss value can be calculated based on the difference between the third medical record and the second medical record, and then the reference large model can be corrected based on the loss value to obtain an updated large model.

[0169] The following combination Figure 9 The method for generating medical records based on a large model and the method for generating a large model for generating medical records provided in an embodiment of the present disclosure are described. Figure 9 A schematic diagram of the process of large-model-based medical record generation and large-model incremental pre-training provided in an embodiment of the present disclosure.

[0170] like Figure 9 As shown, Figure 9 Taking the online reasoning part as an example, the medical record generation method based on the large model provided by the embodiment of the present disclosure is illustrated.

[0171] First, the doctor-patient conversation data is input into the language model. Then, the language model is used to perform content filtering on the conversation data based on self-information to improve the quality of the conversation data input into the large model.

[0172] like Figure 9 As shown, Figure 9 Taking the offline training section in the middle as an example, the method for generating a large model for generating medical records provided by the embodiments of the present disclosure is illustrated. The large model is incrementally pre-trained through the summary generation task (original text, summary summary result), the medical record structure understanding task (medical record document type, medical record document content), and the entity normalization task (entity colloquial expression, entity standardized expression). Then, the large model is supervised fine-tuned based on the "doctor-patient dialogue, medical record generation result" task, thereby improving the quality of medical records generated by the large model.

[0173] Finally, the filtered data and prompt information are input into the trained large model to obtain the medical record generation results.

[0174] In the disclosed embodiment, after obtaining the first and second data sets, the initial large model is first pre-trained based on the first data set to obtain a reference large model. Then, based on the type of the second medical record, prompt information is determined. The prompt information and the second medical consultation data are then input into the reference large model to obtain a third medical record generated by the reference large model. Finally, based on the difference between the third medical record and the second medical record, the reference large model is modified to obtain an updated large model. Thus, by incrementally pre-training the large model based on the first data set and then updating the incrementally pre-trained large model based on the second data set to obtain an updated large model, the accuracy of the large model used to generate medical records is improved, providing conditions for improving the quality of medical records generated by the large model.

[0175] In order to implement the above embodiments, the present disclosure also proposes a medical record generation device based on a large model.

[0176] Figure 10 A schematic structural diagram of a large model-based medical record generation device provided in an embodiment of the present disclosure.

[0177] like Figure 10 As shown, the medical record generation device 1000 based on a large model includes: a first acquisition module 1001 , a first processing module 1002 , a filtering module 1003 , and a second processing module 1004 .

[0178] A first acquisition module 1001 is used to acquire first medical inquiry data to be processed;

[0179] A first processing module 1002 is configured to process the first medical inquiry data to obtain a first score for each character in the first medical inquiry data;

[0180] A filtering module 1003 is configured to filter the first medical inquiry data according to the first score of each character to obtain filtered data;

[0181] The second processing module 1004 is used to input the filtered data and prompt information into the large model to obtain the first medical record output by the large model.

[0182] In a possible implementation of the present disclosure, the first processing module 1001 is specifically configured to:

[0183] Based on the first i characters in the first medical consultation data, predict the i+1th predicted character, where i is a positive integer;

[0184] A first score for the i+1th character is determined based on a difference between the i+1th actual character in the first medical inquiry data and the i+1th predicted character.

[0185] In a possible implementation of the present disclosure, the first processing module 1001 is specifically configured to:

[0186] Based on the first i characters in the first medical inquiry data, determine the candidate character sequence corresponding to the i+1th character and the probability value of each candidate character, where i is a positive integer;

[0187] When the (i+1)th actual character in the first medical inquiry data is any candidate character in the candidate character sequence, a first score of the (i+1)th actual character is determined according to the probability value of any candidate character.

[0188] In a possible implementation of the present disclosure, after determining the candidate character sequence corresponding to the (i+1)th character and the probability value of each candidate character, the first processing module 1001 is further configured to:

[0189] When the (i+1)th actual character in the first medical inquiry data is different from each candidate character in the candidate character sequence, the first score of the (i+1)th actual character is determined to be a specified value.

[0190] In a possible implementation of the present disclosure, the filtering module 1003 is specifically configured to:

[0191] The characters in the first medical inquiry data whose first scores are lower than the first score threshold are deleted to obtain filtered data.

[0192] In a possible implementation of the present disclosure, the filtering module 1003 is specifically configured to:

[0193] fusing the first scores of the characters contained in each sentence of the first medical inquiry data to obtain a second score of the sentence;

[0194] Sentences with second scores lower than a second score threshold in the first medical inquiry data are deleted to obtain filtered data.

[0195] In a possible implementation of the present disclosure, before inputting the filtered data and prompt information into the large model, the second processing module 1004 is further configured to:

[0196] determining similarities between the first consultation data and reference words associated with medical records of different types of structures;

[0197] According to the similarity, determine the target type to which the first medical record belongs;

[0198] Gets the hint information associated with the target type.

[0199] In a possible implementation of the present disclosure, before inputting the filtered data and prompt information into the large model, the second processing module 1004 is further configured to:

[0200] receiving a prompt information selection instruction, wherein the selection instruction includes a target type associated with the prompt information to be selected;

[0201] Gets the hint information associated with the target type.

[0202] The functions and specific implementation principles of the above modules in the embodiments of the present disclosure can be referred to the above method embodiments and will not be repeated here.

[0203] In the disclosed embodiment, after obtaining the first medical inquiry data to be processed, the first medical inquiry data is first processed to obtain a first score for each character in the first medical inquiry data. The first medical inquiry data is then filtered based on the first score for each character to obtain filtered data. Finally, the filtered data and prompt information are input into the macro model to obtain a first medical record output by the macro model. Thus, by filtering the medical inquiry data based on the medical inquiry data itself and then inputting the filtered data and prompt information into the macro model to obtain the medical record, the impact of useless or interfering information on the macro model is reduced, thereby improving the quality of the medical record generated by the macro model.

[0204] In order to implement the above embodiments, the present disclosure also proposes a device for generating a large model of medical records.

[0205] Figure 11 A schematic diagram of the structure of a device for generating a large model of medical records provided in an embodiment of the present disclosure.

[0206] like Figure 11 As shown, the device 1100 for generating a large model of medical records includes: a second acquisition module 1101, a first training module 1102, and a second training module 1103.

[0207] A second acquisition module 1101 is configured to acquire a first data set and a second data set, wherein the first data set includes at least one training data associated with a non-medical record generation task, and the second data set includes a plurality of second medical consultation data and associated second medical records;

[0208] A first training module 1102 is configured to pre-train an initial large model based on a first data set to obtain a reference large model;

[0209] The second training module 1103 is used to update the reference large model based on the second data set to obtain an updated large model.

[0210] In a possible implementation of the present disclosure, non-medical record generation tasks include: summary generation tasks, medical record type understanding tasks, and entity normalization tasks.

[0211] In a possible implementation of the present disclosure, the second training module 1103 is specifically configured to:

[0212] Determine prompt information according to the type of the second medical record;

[0213] Inputting the prompt information and the second medical consultation data into the reference large model to obtain a third medical record generated by the reference large model;

[0214] According to the difference between the third medical record and the second medical record, the reference large model is modified to obtain an updated large model.

[0215] The functions and specific implementation principles of the above modules in the embodiments of the present disclosure can be referred to the above method embodiments and will not be repeated here.

[0216] In the disclosed embodiment, after obtaining the first and second datasets, the initial large model is first pre-trained based on the first dataset to obtain a reference large model. The reference large model is then updated and trained based on the second dataset to obtain an updated large model. Thus, by incrementally pre-training the large model based on training data associated with the summary generation task, the medical record structure understanding task, and the entity normalization task, and then performing supervised fine-tuning training on the pre-trained large model, the accuracy and reliability of the large model for generating medical records are improved.

[0217] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0218] Figure 12 A schematic block diagram of an example electronic device 1200 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0219] like Figure 12As shown, the device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. Various programs and data required for the operation of the device 1200 can also be stored in the RAM 1203. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.

[0220] Various components in device 1200 are connected to I / O interface 1205, including an input unit 1206, such as a keyboard and mouse; an output unit 1207, such as various types of displays and speakers; a storage unit 1208, such as a magnetic disk and optical disk; and a communication unit 1209, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 1209 allows device 1200 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0221] The computing unit 1201 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1201 performs the various methods and processes described above, such as the medical record generation method based on the large model. For example, in some embodiments, the medical record generation method based on the large model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into the RAM 1203 and executed by the computing unit 1201, one or more steps of the medical record generation method based on the large model described above can be performed. Alternatively, in other embodiments, the computing unit 1201 may be configured to execute the large model-based medical record generation method in any other appropriate manner (for example, by means of firmware).

[0222] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0223] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0224] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0225] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0226] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0227] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.

[0228] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0229] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present disclosure, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined. In the description of the present disclosure, the words "if" and "if" used can be interpreted as "at the time of" or "when" or "in response to a determination" or "under the circumstances of".

[0230] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for generating a large model for generating medical records, comprising: Obtaining a first data set and a second data set, wherein the first data set includes training data associated with at least one non-medical record generation task, and the second data set includes a plurality of second consultation data and associated second medical records; the non-medical record generation tasks include: a summary generation task, a medical record type understanding task, and an entity normalization task; splicing the training data in the first data set, inputting the spliced ​​training data into a large model, and modifying the large model to obtain a reference large model; The reference large model is updated and trained based on the second data set to obtain an updated large model.

2. The method according to claim 1, wherein The updating and training of the reference large model based on the second data set to obtain an updated large model includes: determining prompt information according to the type of the second medical record; Inputting the prompt information and the second medical inquiry data into the reference large model to obtain a third medical record generated by the reference large model; According to the difference between the third medical record and the second medical record, the reference large model is modified to obtain the updated large model.

3. A method for generating medical records based on a large model, comprising: Obtaining the first medical consultation data to be processed; processing the first medical inquiry data to obtain a first score for each character in the first medical inquiry data; filtering the first medical inquiry data according to the first score of each character to obtain filtered data; The filtered data and prompt information are input into a large model generated by the method according to any one of claims 1-2 to obtain a first medical record output by the large model.

4. The method according to claim 3, wherein: The processing of the first medical inquiry data to obtain a first score for each character in the first medical inquiry data includes: Based on the first i characters in the first medical inquiry data, predict the i+1th predicted character, where i is a positive integer; A first score for the i+1th character is determined according to a difference between the i+1th actual character in the first medical inquiry data and the i+1th predicted character.

5. The method according to claim 3, wherein: The processing of the first medical inquiry data to obtain a first score for each character in the first medical inquiry data includes: Based on the first i characters in the first medical inquiry data, determine a candidate character sequence corresponding to the i+1th character and a probability value of each candidate character, where i is a positive integer; When the (i+1)th actual character in the first medical inquiry data is any candidate character in the candidate character sequence, a first score of the (i+1)th actual character is determined according to the probability value of the any candidate character.

6. The method according to claim 5, wherein: After determining the candidate character sequence corresponding to the (i+1)th character and the probability value of each candidate character, the method further includes: When the (i+1)th actual character in the first medical inquiry data is different from each candidate character in the candidate character sequence, a first score of the (i+1)th actual character is determined to be a specified value.

7. The method of claim 3, wherein: The filtering of the first medical inquiry data according to the first score of each character to obtain filtered data includes: The characters in the first medical inquiry data whose first scores are lower than a first score threshold are deleted to obtain filtered data.

8. The method of claim 3, wherein: The filtering of the first medical inquiry data according to the first score of each character to obtain filtered data includes: fusing the first scores of the characters contained in each sentence of the first medical inquiry data to obtain a second score of the sentence; The sentences in the first medical inquiry data whose second scores are lower than the second score threshold are deleted to obtain filtered data.

9. The method according to any one of claims 3 to 8, wherein: Before inputting the filtered data and prompt information into the large model, the method further includes: Determining similarities between the first medical inquiry data and reference words associated with medical records of different types of structures; determining, based on the similarity, a target type to which the first medical record belongs; Gets the hint information associated with the target type.

10. The method according to any one of claims 3 to 8, wherein: Before inputting the filtered data and prompt information into the large model, the method further includes: receiving a prompt information selection instruction, wherein the selection instruction includes a target type associated with the prompt information to be selected; Gets the hint information associated with the target type.

11. A device for generating a large model of medical records, comprising: A second acquisition module is configured to acquire a first data set and a second data set, wherein the first data set includes training data associated with at least one non-medical record generation task, and the second data set includes a plurality of second consultation data and associated second medical records; the non-medical record generation tasks include: a summary generation task, a medical record type understanding task, and an entity normalization task; A first training module is configured to splice the training data in the first data set, input the spliced ​​training data into a large model, and modify the large model to obtain a reference large model; The second training module is used to update the reference large model based on the second data set to obtain an updated large model.

12. The device according to claim 11, wherein The second training module is specifically used to: determining prompt information according to the type of the second medical record; Inputting the prompt information and the second medical inquiry data into the reference large model to obtain a third medical record generated by the reference large model; According to the difference between the third medical record and the second medical record, the reference large model is modified to obtain the updated large model.

13. A medical record generation device based on a large model, comprising: A first acquisition module, configured to acquire first medical inquiry data to be processed; a first processing module, configured to process the first medical inquiry data to obtain a first score for each character in the first medical inquiry data; a filtering module, configured to filter the first medical inquiry data according to the first score of each character to obtain filtered data; The second processing module is used to input the filtered data and prompt information into a large model generated by the device according to any one of claims 11 to 12 to obtain a first medical record output by the large model.

14. The apparatus of claim 13, wherein: The first processing module is specifically configured to: Based on the first i characters in the first medical inquiry data, predict the i+1th predicted character, where i is a positive integer; A first score for the i+1th character is determined according to a difference between the i+1th actual character in the first medical inquiry data and the i+1th predicted character.

15. The apparatus of claim 13, wherein: The first processing module is specifically configured to: Based on the first i characters in the first medical inquiry data, determine a candidate character sequence corresponding to the i+1th character and a probability value of each candidate character, where i is a positive integer; When the (i+1)th actual character in the first medical inquiry data is any candidate character in the candidate character sequence, a first score of the (i+1)th actual character is determined according to the probability value of the any candidate character.

16. The apparatus of claim 15, wherein: After determining the candidate character sequence corresponding to the (i+1)th character and the probability value of each candidate character, the first processing module is further configured to: When the (i+1)th actual character in the first medical inquiry data is different from each candidate character in the candidate character sequence, a first score of the (i+1)th actual character is determined to be a specified value.

17. The apparatus of claim 13, wherein: The filtering module is specifically used for: The characters in the first medical inquiry data whose first scores are lower than a first score threshold are deleted to obtain filtered data.

18. The apparatus of claim 13, wherein: The filtering module is specifically used for: fusing the first scores of the characters contained in each sentence of the first medical inquiry data to obtain a second score of the sentence; The sentences in the first medical inquiry data whose second scores are lower than the second score threshold are deleted to obtain filtered data.

19. The device according to any one of claims 13 to 18, wherein: Before inputting the filtered data and prompt information into the large model, the second processing module is further used to: Determining similarities between the first medical inquiry data and reference words associated with medical records of different types of structures; determining, based on the similarity, a target type to which the first medical record belongs; Gets the hint information associated with the target type.

20. The device according to any one of claims 13 to 18, wherein: Before inputting the filtered data and prompt information into the large model, the second processing module is further used to: receiving a prompt information selection instruction, wherein the selection instruction includes a target type associated with the prompt information to be selected; Gets the hint information associated with the target type.

21. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that may be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.

22. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.

23. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Cleaning method and device for automatic voice recognition technology and electronic equipment

    CN114627875A

  • Outpatient service electronic medical record generation method based on Chinese medical big model

    CN117253576A

  • Medical record generation model training method, medical record generation method and related equipment

    CN120183592A