Template-based medical record generation method, device, equipment and medium
Patent Information
- Application Number
- CN202611043189.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-25
AI Technical Summary
每当新医院或新科室提出需求,需重新准备数据、构建模型,导致工作量成倍增加
本发明实施例提供的一种基于模板挖掘的专科病历生成方法、装置、设备及介质,所述方法基于病历要素抽取、汇总合并、归纳模板的系统化流程,从存量病历数据中自动挖掘可复用的病历要素和模板架构,形成标准化的EBNF形式化模板语言,实现了病历生成能力的跨科室、跨医院复用,极大的降低了新需求的接入成本。
Smart Images

Figure CN122822196A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of large language model technology, and in particular to a method, apparatus, equipment and medium for generating specialist medical records based on template mining. Background Technology
[0002] A Large Language Model (LLM) is a deep learning language model trained on a large corpus. It possesses capabilities such as natural language understanding, generation, and reasoning, and can be used for tasks such as text analysis, information extraction, and content generation. Using language models to generate AI medical records is currently a fairly common technical approach.
[0003] Existing AI-powered medical record generation systems primarily operate on a "hospital-department" basis, building specific prompts or training dedicated models. Whenever a new hospital or department requests a new system, data must be prepared and models rebuilt, significantly increasing the workload. While there are numerous similar medical record structures and elements across different hospitals and departments, the lack of a unified, standardized mechanism for data preservation and reuse results in fragmented data and model assets and severe duplication of effort.
[0004] On the other hand, the generated medical records are inconsistent with the medical record templates in the Hospital Information System (HIS). Different departments and different diseases have independent medical record templates in the HIS, but the medical records generated by the existing AI system are still mainly based on a general structure. The output content is inconsistent with the template chapter structure of the HIS, making it impossible to write back directly or requiring doctors to make extensive modifications, which reduces the efficiency of use.
[0005] Furthermore, the system lacks an automatic optimization mechanism based on doctor feedback. Doctors' modifications to medical records and their preferences cannot be effectively fed back into the model, resulting in the model's inability to continuously improve and its long-term effectiveness being difficult to enhance.
[0006] Therefore, there is an urgent need for a new method for generating specialist medical records based on template mining. Summary of the Invention
[0007] In view of the above problems, the present invention is proposed to provide a method, apparatus, device and medium for generating specialist medical records based on template mining that overcomes or at least partially solves the above problems.
[0008] Other features and advantages of the invention will become apparent from the following detailed description, or may be learned in part by practice of the invention.
[0009] According to a first aspect of the present invention, a method for generating specialist medical records based on template mining is provided, comprising: The existing medical record data of the target specialty is obtained. After semantic analysis of the existing medical record data using a large language model, multiple medical record elements are extracted and the medical record elements are summarized to form a standardized medical record element library. Based on the medical record element library and the existing medical record data, the structure of the medical records is analyzed using a large language model to generate formalized medical record templates, which are then stored in the medical record template library. Obtain current patient information, and match and select a target medical record template from the medical record template library based on the current patient information; Using the current patient information as input, and combining it with the target medical record template, a large language model is invoked to generate a target medical record that conforms to the structure of the target medical record template.
[0010] In some embodiments of the present invention, the semantic analysis of the existing medical record data using a large language model includes: using a Few-Shot prompt word template to guide the large language model to perform semantic analysis on the existing medical record data, so as to extract multiple medical record elements. Each extracted medical record element is represented as a tuple containing semantic tags and attribute parameters. The Few-Shot prompt word template includes task instructions, a predefined medical record element category system, and multiple annotation examples.
[0011] In some embodiments of the present invention, the generation of formalized medical record templates includes: Using a large language model, the combination patterns, arrangement order, and occurrence constraints of each medical record element in the existing medical record data are analyzed and summarized. The combined patterns, permutations, and occurrence constraints are defined as one or more template rules using the extended Backus-Naur paradigm to form the medical record template.
[0012] In some embodiments of the present invention, after forming the medical record template, the method further includes: The generated medical record template is automatically validated to generate a validation result. The automatic validation includes grammatical validity checks and logical conflict detection. If the verification result fails, the medical record template is manually reviewed, and the medical record template is either edited again or regenerated using a large language model.
[0013] In some embodiments of the present invention, obtaining current patient information and matching and selecting a target medical record template from the medical record template library based on the current patient information includes: Obtain current patient information, and extract the patient's department and pre-diagnosis results from the current patient information; In response to the API call request of the Hospital Information System (HIS), the currently selected medical record template in the HIS is matched as the target medical record template; If the API call request is not obtained, a precise match is made from the medical record template library based on the patient's department and the disease name with the highest confidence in the pre-diagnosis results, and the matched medical record template is used as the target medical record template. If an exact match fails, the semantic similarity between the Top-N candidate diseases and the candidate templates in the pre-identified disease results is calculated, and the medical record template with a similarity exceeding a preset threshold is selected as the target medical record template. If semantic matching fails, the basic general template of the department to which the patient belongs is selected as the target medical record template.
[0014] In some embodiments of the present invention, before matching and selecting a target medical record template from the medical record template library, the method further includes: In response to the user's first configuration command, a basic template is selected from one or more existing medical record templates, and existing medical record elements or newly added custom medical record elements are selected from the medical record element library. The medical record elements are combined to form a new medical record template, and the new medical record template is stored in the medical record template library. Or / and in response to a second configuration command from the user, receive the uploaded medical record sample document, extract medical record elements and medical record structure from the medical record sample document using a large language model, and pre-generate a new medical record template based on the mining results, which is then stored in the medical record template library after user confirmation.
[0015] In some embodiments of the present invention, after generating the target medical record, the method further includes: Based on the medical record element library and medical record template library, multiple patient profiles and simulated consultation dialogues are generated through a large language model, and the patient profiles and simulated consultation dialogues are combined into medical record generation training data. The synthesized medical record training data was used to perform supervised fine-tuning of the SFT of the base large language model; Collect modification operations and / or preference feedback on the target medical records to construct preference data pairs containing preferred and unpreferred samples; Using the aforementioned preference data pairs, the base large language model, after supervised fine-tuning of SFT, is subjected to direct preference optimization (DPO) to obtain an optimized medical record generation model. This optimized model is used to generate target medical records that conform to the target medical record template structure.
[0016] According to a second aspect of the present invention, a specialist medical record generation apparatus based on template mining is provided, the apparatus comprising: The medical record element extraction module is used to acquire existing medical record data of the target specialty, and extract multiple medical record elements after semantic analysis of the existing medical record data using a large language model. The medical record elements are then summarized to form a standardized medical record element library. The medical record template generation module is used to analyze the structure of medical records using a large language model based on the medical record element library and the existing medical record data, generate formalized medical record templates, and store them in the medical record template library. The medical record matching module is used to obtain current patient information and match and select a target medical record template from the medical record template library based on the current patient information. The target medical record generation module is used to take the current patient information as input, combine it with the target medical record template, and call the large language model to generate a target medical record that conforms to the structure of the target medical record template.
[0017] According to a third aspect of the present invention, a computer device is provided, including a processor and a memory, the memory storing computer program instructions executable by the processor, wherein when the processor executes the computer program instructions, it implements the instructions as described in any of the above methods.
[0018] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, wherein computer program instructions are stored therein, the computer program instructions being loaded and executed by a processor to perform the operations performed by the method described in any of the preceding claims.
[0019] The technical solutions provided in the embodiments of the present invention have at least the following technical effects or advantages: This invention provides a method, apparatus, device, and medium for generating specialized medical records based on template mining. The method is based on a systematic process of extracting, summarizing, and merging medical record elements and incorporating templates. It automatically mines reusable medical record elements and template architectures from existing medical record data to form a standardized EBNF formal template language. This enables cross-departmental and cross-hospital reuse of medical record generation capabilities and greatly reduces the access cost for new requirements.
[0020] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A flowchart illustrating a method for generating specialist medical records based on template mining, provided in an embodiment of the present invention; Figure 2 A schematic diagram of the principle structure of a specialty medical record generation device based on template mining provided in an embodiment of the present invention; Figure 3 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0023] Exemplary embodiments of this application will now be described in more detail with reference to the accompanying drawings.
[0024] The accompanying drawings illustrate various structural schematics according to embodiments of this application. These drawings are not to scale, and some details have been enlarged for clarity, and some details may have been omitted. The shapes of the various regions and layers shown in the drawings, as well as their relative sizes and positional relationships, are merely exemplary and may deviate from reality due to manufacturing tolerances or technical limitations. Furthermore, those skilled in the art can design regions / layers with different shapes, sizes, and relative positions as needed.
[0025] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. In the context of this application, similar or identical parts may be represented by the same or similar reference numerals.
[0026] To better understand the above technical solutions, the following will describe the above technical solutions in detail with reference to specific implementation methods. It should be understood that the embodiments of this application and the specific features in the embodiments are detailed descriptions of the technical solutions of the present invention, rather than limitations on the technical solutions of the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0027] Figure 1 This is a flowchart illustrating a method for generating specialist medical records based on template mining, as provided in an embodiment of the present invention. Figure 1As shown, the template mining-based specialist medical record generation method includes the following steps: S1. Obtain the existing medical record data of the target specialty, and extract multiple medical record elements by performing semantic analysis on the existing medical record data using a large language model. Summarize the medical record elements to form a standardized medical record element library. The target specialty is determined based on the target department. By obtaining the existing medical record data of the target specialty as input to the large language model, the semantic analysis of the existing medical record data using the large language model in this embodiment of the invention includes: using Few-Shot prompt word templates to guide the large language model to perform semantic analysis on the existing medical record data in order to extract multiple medical record elements.
[0028] Each medical record element extracted in this embodiment of the invention is represented as a tuple containing semantic tags and attribute parameters, such as <semantic tag>(parameter 1: type, parameter 2: type, ..., parameter n: type). The "main symptom (symptom, time:time)" in the chief complaint represents a description of the main symptom that includes the symptom and its duration. The "onset (triggering factor, symptom, time:time)" in the present medical history represents a description of the onset of the present medical history that includes the triggering factor, symptoms, and duration.
[0029] The Few-Shot prompt template includes task instructions, a predefined medical record element category system, and multiple annotation examples. The task instructions specify the output format (tuple form). The predefined medical record element category system is a clinically expert-confirmed system used as a reference for identification. The annotation examples are high-quality manually annotated examples that cover different departments and typical scenarios prone to ambiguity, allowing the large language model to learn consistent extraction standards. The Few-Shot prompt template effectively improves the accuracy of early identification of medical record elements. The prompts in the task instructions clearly define format boundaries and require the model to mark the position of each element in the original text, facilitating subsequent manual verification. The Few-Shot examples (predefined medical record element category system and annotation examples) pre-select easily confused scenarios reviewed by clinical experts as demonstrations.
[0030] In this embodiment of the invention, the mining results in the cold start stage are based on manual review (i.e., offline mining of medical record templates in steps S1-S2), and in the hot update stage (i.e., online template configuration and medical record generation in steps S3-S4), there is a medical record element review node. The newly mined results will be compared with the existing database to exclude duplicate or conflicting medical record elements.
[0031] After extracting all medical record elements, this embodiment of the invention summarizes the medical record elements extracted from all medical record data, classifies and merges them according to element name, semantic meaning and parameter attributes, disambiguates and integrates elements with semantic repetition or ambiguity, and finally forms a standardized medical record element library that is unified throughout the hospital or department.
[0032] In other embodiments of the present invention, the mining of medical record elements can be carried out not only through large language models, but also through rule-based NLP (Natural Language Processing) methods, knowledge graph-based extraction methods, or a combination of multiple methods, depending on the actual application requirements.
[0033] The large language model can use various open-source or other commercial large language models with Decoder-Only architectures, such as the Qwen series, DeepSeek series, Llama series, and ChatGLM series. The large language model can be deployed on cloud servers, local servers, or edge computing devices, and can also adopt different deployment architectures such as public cloud, private cloud, or hybrid cloud, which can be flexibly selected according to actual application needs.
[0034] S2. Based on the medical record element library and the existing medical record data, analyze the structure of the medical records using a large language model, generate formalized medical record templates, and store them in the medical record template library. The generation of formalized medical record templates in this embodiment of the invention includes: using a large language model to analyze and summarize the combination patterns, arrangement order, and occurrence constraints of each medical record element in the existing medical record data; and using an extended Backus-Naur paradigm to define the combination patterns, arrangement order, and occurrence constraints as one or more template rules to form the medical record template.
[0035] Based on the compiled medical record element database and the existing stock of medical record data, the system again uses a large language model to analyze the organizational structure, combination patterns, arrangement order, and occurrence constraints of each element in the medical record. Then, the system uses Extended Backus-Naur Normal Form (EBNF) as a formal grammar to define the above structural relationships (such as order, optionality, repetition, etc.) as machine-readable (formalized) medical record templates. For example, the chief complaint template is: "( {main symptom(symptom, time:time)} ({accompanying symptoms(symptom, time:time)} ) )"; the present illness template is: "( {onset (cause, symptom, time:time)},{main symptom(symptom, time:time, location, characteristics?, severity?)},( {positive accompanying symptoms(symptom, time:time, location, characteristics?, severity?)} ),( {negative symptoms} ). ( {treatment experience(plan, date:date?, result)} ). {general situation( Mental state? Physical condition? Diet? Sleep? Bowel and bladder function?
[0036] The EBNF can also be replaced with JSON Schema, XML Schema, custom DSL (domain-specific language), or other structured description languages to define the template's organization and constraint rules.
[0037] After the medical record template is generated, the embodiments of the present invention further include: automatically verifying the generated medical record template and generating a verification result, wherein the automatic verification includes grammatical validity checking and logical conflict detection; if the verification result fails, the medical record template is manually reviewed and the medical record template is edited again or regenerated using a large language model. This invention uses an EBNF parser (such as Python's Lark) at the grammatical level to check basic rules such as bracket matching and operation legality. At the logical level, it uses a conflict detection algorithm to check whether the same element is simultaneously marked as required and prohibited, and whether there is a forced circular dependency. If a problem is found during automatic verification, the medical record template will be automatically marked as pending review and the configuration personnel will be prompted to conduct a manual review. The configuration personnel can then edit the medical record template or regenerate it using a large language model. For example, the review interface of the configuration personnel will display an EBNF syntax tree diagram automatically rendered by the system, as well as a preview of a medical record sample generated based on the template. The configuration personnel will judge from three perspectives: logical completeness, reasonable combination, and distinguishability of existing templates. If a problem is found, it can be directly corrected using editing tools or the large language model LLM can be requested to regenerate it. Once the manual review is passed, the template will be stored in the medical record template library and published.
[0038] Medical record templates that have passed automatic verification and manual review will be officially added to the medical record template library. The medical record template library can serve as a basic general template or a special template for a specific department, providing medical record templates for subsequent online configuration and medical record generation.
[0039] S3. Obtain the current patient information, and match and select the target medical record template from the medical record template library based on the current patient information; Before matching and selecting a target medical record template from the medical record template library, this embodiment of the invention includes: responding to a first configuration instruction from the user, selecting a basic template from one or more existing medical record templates, and selecting existing medical record elements or adding custom medical record elements (providing element names, parameters, definitions, and examples) from the medical record element library, combining the medical record elements to form a new medical record template, and storing the new medical record template in the medical record template library; or / and responding to a second configuration instruction from the user, receiving an uploaded medical record sample document, extracting medical record elements and medical record structure from the medical record sample document using a large language model, and pre-generating a new medical record template based on the mining results, and storing it in the medical record template library after user confirmation.
[0040] This invention obtains current patient information by having the patient complete a pre-consultation questionnaire, and performs pre-diagnosis based on the questionnaire information using a large language model. The system automatically selects the diagnosis result with the highest confidence as the pre-diagnosis result. The confidence is derived from the normalized result of the logarithmic probability of the token corresponding to the disease name output by the model. The array is sorted from high to low confidence. The model outputs Top-N candidate diseases and their respective confidence scores (floating-point numbers from 0 to 1) in structured JSON format. The model selects the first one as the optimal diagnosis result, and the Top-N list is reserved as alternatives. After obtaining the current patient information, a target medical record template is matched for the patient from the existing medical record template library.
[0041] This embodiment of the invention obtains current patient information and selects a target medical record template from the medical record template library based on the current patient information, including: obtaining current patient information and extracting the patient's department and pre-diagnosis results from the current patient information; responding to the API (interface) call request of the Hospital Information System (HIS) and matching the currently selected medical record template in the HIS as the target medical record template; if the API call request is not obtained, performing precise matching from the medical record template library based on the patient's department and the highest confidence disease name in the pre-diagnosis results, and using the matched medical record template as the target medical record template; if precise matching fails, calculating the semantic similarity (e.g., cosine similarity) between the Top-N candidate diseases in the pre-diagnosis results and the candidate templates, and selecting the medical record template with a similarity exceeding a preset threshold as the target medical record template; if semantic matching fails, selecting the basic general template of the patient's department as the target medical record template.
[0042] S4. Using the current patient information as input, and combining it with the target medical record template, call the large language model to generate a target medical record that conforms to the structure of the target medical record template.
[0043] The current patient information is used as input, which may also include consultation dialogue data (e.g., obtained through real-time automatic speech recognition (ASR) speech-to-text transcription, or other methods such as text input, video consultation transcription, and intelligent consultation robot dialogue). Combined with the selected target medical record template, a large language model is invoked to generate a target medical record that conforms to the structure of the target medical record template. The target medical record may include chapters such as chief complaint, present illness, and past medical history. In other embodiments of the present invention, the patient's past medical record information (previous consultation history and time information) can be additionally associated. The past information and the content of the current consultation are integrated through the large language model to generate a complete medical record containing the evolution of the condition.
[0044] The generated target medical record is unstructured text and will be reassembled according to the chapter structure of the template. The data elements of each chapter correspond to the medical record element slots defined in the template. The system provides slot view (structured) and text view (plain text) for doctors to edit. The target medical record is displayed in real time on the interface in draft form according to the chapters of the target medical record template for the doctor currently consulting. The doctor can modify the target medical record and / or provide feedback on preferences, and define the target medical record as a preferred sample and a poorly selected sample based on the modification operation and / or preference feedback. For example, the original target medical record before the modification operation and / or preference feedback is used as a poorly selected sample, and the target medical record after the modification operation is used as a preferred sample; or a preferred sample is defined as one with few or no modifications and good preference feedback (likes), and a poorly selected sample is defined as one with many or large modifications and poor preference feedback (dislikes).
[0045] After the doctor confirms the results of the target medical record, the system synchronously writes the medical record results back to the hospital information system (HIS) in a structured manner according to the HIS mapping rules configured on the management end, so that the generated results are almost ready for use.
[0046] To address the issue of insufficient training data, this embodiment of the invention utilizes the existing medical record element library and medical record template library to synthesize high-quality training data. After generating the target medical record, this embodiment further includes: generating multiple patient profiles and simulated consultation dialogues based on the medical record element library and medical record template library using a large language model; synthesizing the patient profiles and simulated consultation dialogues into medical record generation training data; using the synthesized medical record generation training data to perform supervised fine-tuning (SFT) on the base large language model; collecting modification operations and / or preference feedback on the target medical record to construct preference data pairs containing preferred and unpreferred samples; and using the preference data pairs to perform direct preference optimization (DPO) on the base large language model after supervised fine-tuning (SFT) to obtain an optimized medical record generation model, which is used to generate target medical records that conform to the target medical record template structure.
[0047] This invention generates a patient profile containing multiple clinical features, and then, based on a predefined consultation script, uses elements from a target medical record template as guidance to enable the Large Language Model (LLM) to generate a simulated consultation dialogue that conforms to the script structure. Finally, based on the simulated consultation dialogue and the template, a standardized target medical record template is synthesized as a training sample to perform supervised fine-tuning of the base Large Language Model (SFT).
[0048] In this embodiment of the invention, the simulated consultation dialogue is generated according to a predefined consultation script structure, constrained by a patient profile (age, gender, main symptoms, duration, summary of past medical history, etc.). The script covers a standard consultation node sequence such as "patient's chief complaint → doctor's follow-up questions → patient description → treatment experience → general condition". During generation, the various medical record elements involved in the target medical record template are retrieved from the medical record element library as guides, constraining the LLM to cover all necessary information points in the dialogue. For the same patient profile, multiple sets of variations are generated using different random seeds, and synonym replacement and sentence structure variations are added to enhance diversity.
[0049] The synthetic data undergoes multiple quality checks before model training, including coverage, consistency, and common sense verification. For example, data with less than 20% coverage of essential elements is discarded; Rouge-L comparisons are performed between the synthesized dialogue input to the model and the original template, with scores below 0.6 indicating low quality; common sense verification is performed using a medical knowledge graph; and the data is reviewed by clinical experts. When the synthetic data meets the requirements for coverage, consistency, and common sense verification, it is used as the training set for model training.
[0050] The underlying large language model is, for example, a Decoder-Only architecture model such as Qwen3-8B; SFT (Supervised Fine-Tuning) is performed using synthetic training data and real labeled data to enable the model to learn the ability to generate medical records according to a template format; feedback generated by doctors in actual work (modifications to generated medical records, likes / dislikes) is used to optimize the model. The system analyzes doctors' modifications through a difference algorithm and automatically determines the modification type (such as element correction, description optimization). For example, the version before modification is used as the inferior sample and the version after modification is used as the superior sample to construct preference data pairs, which are used to trigger DPO training so that the model gradually gets closer to doctors' clinical preferences; for example, an online DPO training can be triggered when the number of preference data pairs reaches a preset threshold (100).
[0051] In other embodiments of the present invention, the training method of the base large language model, in addition to SFT+DPO, can also be replaced by other preference alignment methods such as RLHF (reinforcement learning based on human feedback), PPO (proximal policy optimization), and KTO (Kahneman-Tversky optimization) for model optimization.
[0052] In other embodiments of the present invention, an automated quality improvement closed loop based on Badcase analysis is constructed. The system automatically analyzes problems (Badcases) in the generated target medical records using a large language model, categorizing them as element problems, specification problems, or format problems. The system performs multiple checks (e.g., before optimization is triggered, it checks whether format problems actually violate template rules, whether specification problems are supported by a medical knowledge graph, and whether any missed elements actually appeared in the original dialogue; failure in any one of these steps results in downgrading to manual review) to prevent misjudgment. Manual review is conducted through random sampling; for example, 5 Badcases are randomly selected for review out of every 50 processed Badcases. If the accuracy rate of the random sampling is below 80%, automatic optimization is paused, and a full manual review is initiated. Alternatively, an impact report is automatically generated before implementing optimization measures, estimating the positive and negative impacts on existing indicators. Negative impacts exceeding a threshold require manual confirmation before implementation. After confirming the problem, the system automatically triggers corresponding optimization measures based on the problem type. For example, for element problems, additional configurations are added or DPO optimization is triggered; for format problems, the template is automatically corrected. The entire process forms a closed loop of "evaluation -> attribution -> optimization -> re-evaluation," continuously improving the generated quality.
[0053] The template mining-based specialist medical record generation method described in this invention has the following advantages compared to existing technologies: 1. Based on a systematic process of extracting, summarizing, and merging medical record elements and incorporating templates, reusable medical record elements and template architectures are automatically mined from existing medical record data to form a standardized EBNF formal template language. This enables cross-departmental and cross-hospital reuse of medical record generation capabilities, greatly reducing the access cost for new requirements. 2. During the cold start phase, basic templates are automatically discovered. During the hot update phase, two template configuration methods are supported: direct configuration and automatic discovery. This allows a single model to dynamically adapt to medical record types of different departments and diseases, achieving out-of-the-box and quick access capabilities. 3. Combining two training methods, SFT supervised fine-tuning and DPO direct preference optimization, preference data is automatically generated based on doctor usage feedback (modification records, likes and dislikes, etc.) and drives continuous model optimization, realizing the model's self-evolution capability from repeated manual debugging to intelligent dynamic adaptation; 4. By using the mapping mechanism between the target medical record template and the hospital information system (HIS) template, the medical record structure generated by the large language model is consistent with the template format in the hospital information system (HIS), supports structured write-back, reduces the manual modification cost after doctors synchronize, and achieves the effect of being almost usable immediately after generation. 5. An automatic Badcase attribution mechanism was introduced, which uses a large language model to automatically identify element problems, standardization problems, and format problems in the generated medical records, and automatically triggers corresponding optimization measures according to the problem type, thereby achieving continuous and automated improvement of the quality of medical record generation.
[0054] Based on the above embodiments, as a supplement to the above... Figure 1 The present invention provides an embodiment of a specialty medical record generation device based on template mining, which is similar to the implementation of the method shown. Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices, see reference. Figure 2 As shown, the template mining-based specialist medical record generation device includes: The medical record element extraction module 100 is used to acquire the existing medical record data of the target specialty, extract multiple medical record elements after semantic analysis of the existing medical record data using a large language model, and summarize the medical record elements to form a standardized medical record element library. The medical record template generation module 200 is used to analyze the structure of the medical record based on the medical record element library and the existing medical record data, generate a formalized medical record template, and store it in the medical record template library. The medical record matching module 300 is used to obtain current patient information and match and select a target medical record template from the medical record template library based on the current patient information. The target medical record generation module 400 is used to take the current patient information as input, combine it with the target medical record template, and call a large language model to generate a target medical record that conforms to the structure of the target medical record template.
[0055] The modules in the aforementioned template mining-based specialist medical record generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0056] The template mining-based specialist medical record generation device described in this embodiment of the invention can execute the template mining-based specialist medical record generation method provided in the above embodiments. The template mining-based specialist medical record generation device has the corresponding functional steps and beneficial effects of the template mining-based specialist medical record generation method described in the above embodiments. For details, please refer to the embodiments of the template mining-based specialist medical record generation method described above. The embodiments of the present invention will not be repeated here.
[0057] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 3 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface, such as a network interface card, is used for wired or wireless communication with external terminals. Wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a template mining-based method for generating specialist medical records. The display unit of the computer device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0058] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0059] In one exemplary embodiment, a chip is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the method in any of the above embodiments.
[0060] In one exemplary embodiment, a network interface card is provided, including a chip as described in any of the above embodiments and multiple interfaces, wherein the chip communicates externally through the interfaces.
[0061] In one embodiment, a computer device is also provided, including a processor, a chip in any of the above embodiments, or a network interface card in any of the above embodiments, wherein the chip or the network interface card is used to schedule packets to the processor or the chip or the network interface card itself for processing, and the processor is used to process the packets scheduled by the chip or the network interface card.
[0062] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0063] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0064] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory may include read-only memory (Read-Only Memory). Only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application may include at least one of relational databases and non-relational databases. Non-relational databases may include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the various embodiments provided in this application may be general-purpose processors, central processing units, graphics processors, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited thereto.
[0065] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0066] Similarly, it should be understood that, for the purpose of simplification and aiding understanding of one or more aspects of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof in the description of exemplary embodiments of the invention above. Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and it should be noted that the above embodiments are illustrative of the invention and not restrictive, and that alternative embodiments can be devised by those skilled in the art without departing from its scope.
Claims
1. A method for generating specialist medical records based on template mining, characterized in that, include: The existing medical record data of the target specialty is obtained. After semantic analysis of the existing medical record data using a large language model, multiple medical record elements are extracted and the medical record elements are summarized to form a standardized medical record element library. Based on the medical record element library and the existing medical record data, the structure of the medical records is analyzed using a large language model to generate formalized medical record templates, which are then stored in the medical record template library. Obtain current patient information, and match and select a target medical record template from the medical record template library based on the current patient information; Using the current patient information as input, and combining it with the target medical record template, a large language model is invoked to generate a target medical record that conforms to the structure of the target medical record template.
2. The method for generating specialist medical records based on template mining according to claim 1, characterized in that, The semantic analysis of the existing medical record data using a large language model includes: using a Few-Shot prompt word template to guide the large language model to perform semantic analysis on the existing medical record data, so as to extract multiple medical record elements. Each extracted medical record element is represented as a tuple containing semantic tags and attribute parameters. The Few-Shot prompt word template includes task instructions, a predefined medical record element category system, and multiple annotation examples.
3. The method for generating specialist medical records based on template mining according to claim 1, characterized in that, The generated formalized medical record template includes: Using a large language model, the combination patterns, arrangement order, and occurrence constraints of each medical record element in the existing medical record data are analyzed and summarized. The combined patterns, permutations, and occurrence constraints are defined as one or more template rules using the extended Backus-Naur paradigm to form the medical record template.
4. The method for generating specialist medical records based on template mining according to claim 3, characterized in that, After generating the medical record template, the method further includes: The generated medical record template is automatically validated to generate a validation result. The automatic validation includes grammatical validity checks and logical conflict detection. If the verification result fails, the medical record template is manually reviewed, and the medical record template is either edited again or regenerated using a large language model.
5. The method for generating specialist medical records based on template mining according to claim 1, characterized in that, The step of obtaining current patient information and matching and selecting a target medical record template from the medical record template library based on the current patient information includes: Obtain current patient information, and extract the patient's department and pre-diagnosis results from the current patient information; In response to the API call request of the Hospital Information System (HIS), the currently selected medical record template in the HIS is matched as the target medical record template; If the API call request is not obtained, a precise match is made from the medical record template library based on the patient's department and the disease name with the highest confidence in the pre-diagnosis results, and the matched medical record template is used as the target medical record template. If an exact match fails, the semantic similarity between the Top-N candidate diseases and the candidate templates in the pre-identified disease results is calculated, and the medical record template with a similarity exceeding a preset threshold is selected as the target medical record template. If semantic matching fails, the basic general template of the department to which the patient belongs is selected as the target medical record template.
6. The method for generating specialist medical records based on template mining according to claim 1, characterized in that, Before matching and selecting a target medical record template from the medical record template library, the method further includes: In response to the user's first configuration command, a basic template is selected from one or more existing medical record templates, and existing medical record elements or newly added custom medical record elements are selected from the medical record element library. The medical record elements are combined to form a new medical record template, and the new medical record template is stored in the medical record template library. Or / and in response to a second configuration command from the user, receive the uploaded medical record sample document, extract medical record elements and medical record structure from the medical record sample document using a large language model, and pre-generate a new medical record template based on the mining results, which is then stored in the medical record template library after user confirmation.
7. The method for generating specialist medical records based on template mining according to claim 1, characterized in that, After generating the target medical record, the method further includes: Based on the medical record element library and medical record template library, multiple patient profiles and simulated consultation dialogues are generated through a large language model, and the patient profiles and simulated consultation dialogues are combined into medical record generation training data. The synthesized medical record training data was used to perform supervised fine-tuning of the SFT of the base large language model; Collect modification operations and / or preference feedback on the target medical records to construct preference data pairs containing preferred and unpreferred samples; Using the aforementioned preference data pairs, the base large language model, after supervised fine-tuning of SFT, is subjected to direct preference optimization (DPO) to obtain an optimized medical record generation model. This optimized model is used to generate target medical records that conform to the target medical record template structure.
8. A specialist medical record generation device based on template mining, applied to the method described in any one of claims 1-7, characterized in that, The device includes: The medical record element extraction module is used to obtain the existing medical record data of the target specialty, and after performing semantic analysis on the existing medical record data using a large language model, extract multiple medical record elements, and summarize the medical record elements to form a standardized medical record element library. The medical record template generation module is used to analyze the structure of medical records using a large language model based on the medical record element library and the existing medical record data, generate formalized medical record templates, and store them in the medical record template library. The medical record matching module is used to obtain current patient information and match and select a target medical record template from the medical record template library based on the current patient information. The target medical record generation module is used to take the current patient information as input, combine it with the target medical record template, and call the large language model to generate a target medical record that conforms to the structure of the target medical record template.
9. A computer device comprising a processor and a memory, characterized in that, The memory stores computer program instructions that can be executed by the processor, and when the processor executes the computer program instructions, it implements the instructions of the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that are loaded and executed by a processor to perform the operations described in any one of claims 1-7.