System and method for extracting electronic medical record information
By employing a closed-loop mechanism of hierarchical prompts and logical checks, the inaccuracies and logical inconsistencies in electronic medical record information extraction are resolved, generating high-quality, standard-compliant structured medical record data that supports clinical diagnosis and treatment as well as scientific research analysis.
Patent Information
- Application Number
- CN202511516247.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-23
AI Technical Summary
Existing technologies for processing electronic medical records suffer from inaccurate and non-standard information extraction, inconsistent logic, and a lack of effective verification mechanisms, making it difficult to meet the stringent requirements of clinical applications.
A prompt word generation module with a hierarchical structure is used to guide the large language model to extract information, and a medical knowledge base is combined to perform terminology correction and logical verification, forming a closed-loop correction process to ensure the accuracy and logical consistency of the output electronic medical record information.
It significantly improves the accuracy and standardization of information extraction, enhances the clinical usability and system integration of electronic medical records, and can generate structured data that meets industry standards.
Smart Images

Figure CN120994654A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a system and method for extracting electronic medical record information. BACKGROUND
[0002] With the development of medical informatization, electronic medical records have become the core data source for clinical diagnosis and treatment and scientific research analysis. However, a large amount of current electronic medical record content, such as transcribed records of doctor-patient conversations and doctor-entered medical history records, exists in the form of unstructured free text. The expression of these texts is diverse, and the information is scattered, making it difficult to be directly structured and utilized by computer systems.
[0003] In order to extract valuable medical information from unstructured text, existing technologies have made various explorations. Early technologies mainly used rule-based and medical dictionary-based methods to extract information through pre-set regular expressions and keyword matching. Although this method is simple to implement, its accuracy and recall rate are not ideal when faced with complex sentence patterns and varied colloquial expressions, and the construction and maintenance of the rule base is extremely costly. Subsequently, traditional machine learning and deep learning-based methods were introduced, such as conditional random field models, recurrent neural networks, or transformer models. These methods learn to identify entities and relationships from text by training on a large amount of labeled data, and their performance is superior to rule-based methods to some extent, but their effectiveness is heavily dependent on the size of high-quality labeled data, and the models themselves lack deep medical knowledge and logical reasoning capabilities, making it difficult to process long texts with complex temporal relationships and effectively filter out noise information such as idle chatter and mistakes in doctor-patient conversations.
[0004] In recent years, large language models have shown great potential in information extraction due to their strong natural language understanding and generation capabilities. Some studies have begun to explore the use of prompt engineering to guide large language models to directly extract structured information from medical records. For example, by constructing instruction templates, the model is required to extract specific fields in the medical record, and combined with a medical knowledge base, the extracted medical terms are standardized and corrected. However, this method still has obvious defects: first, simple prompts cannot accurately constrain the output of the model, resulting in unstable result formats and information hallucinations or omissions; second, the model itself cannot guarantee the absolute consistency of the extracted information in terms of temporal sequence and medical logic, for example, it may confuse a symptom that occurred three days ago with a symptom that occurred one day ago in the timeline, or incorrectly categorize chronic disease information that should belong to "past medical history" into the "chief complaint" of the current visit; finally, existing solutions are linear extraction processes that lack effective verification and correction mechanisms for the extracted results, resulting in insufficient reliability of the output information and making it difficult to meet the stringent requirements of clinical applications. SUMMARY
[0005] The purpose of the present application is to provide an electronic medical record information extraction system and method, aiming to solve the problems of inaccurate, non-standard, inconsistent logic and lack of effective correction mechanism in the prior art when processing free text medical records, so as to improve the clinical usability of the finally generated electronic medical records.
[0006] To achieve the above-mentioned purpose, the present application provides an electronic medical record information extraction system, comprising: A data acquisition module is configured to acquire multi-module medical text data, wherein the multi-module medical text data comprises doctor-patient conversation voice transcription text, free text medical record and OCR-recognized imaging and test report. A prompt word generation module is configured to generate prompt words containing hierarchical structure for a preset medical information extraction task. A model inference module is configured to combine the prompt words containing hierarchical structure and the preprocessed multi-module medical text data, and input them into a preset large language model to obtain a preliminary extraction result containing candidate medical information. A term correction module is configured to call a preset medical knowledge base to compare and standardize the medical terms in the preliminary extraction result. A logic verification module is configured to perform logic verification on the extraction result corrected by the term correction module, wherein the logic verification comprises at least one of time sequence relationship verification and cross-field information consistency verification to generate final structured medical record information.
[0007] Further, the prompt words containing hierarchical structure comprise: a system layer for defining global output format and constraints, a task layer for setting extraction priority and strategy, a field layer for constructing templates and constraints for specific medical fields, and a verification layer for guiding the large language model to perform self-checking.
[0008] Further, the logic verification performed by the logic verification module comprises at least one of the following operations: performing time sequence relationship verification to sort the sequence of symptom occurrence according to the time expressions in the original text, and converting relative time expressions into standardized absolute time stamps; performing cross-field information consistency verification to eliminate redundant symptoms or semantically similar symptoms.
[0009] Further, the logic verification module is further configured to generate an error correction prompt word when an error or inconsistency is found in the logic verification, and control the model inference module to input the error correction prompt word and the original text and context information into the large language model again, instructing it to perform targeted correction to form a closed-loop correction process.
[0010] Further, the system further comprises: A text preprocessing module configured between the data acquisition module and the model inference module, configured to preprocess the medical text data to remove casual conversation, noise and non-medical related content.
[0011] Further, it further comprises a structured output module configured to convert the final structured medical record information into a preset data format and interface to a hospital information system or an electronic medical record system; the preset data format includes JSON, XML, FHIR or HL7.
[0012] To achieve the above-mentioned purpose, the present application also provides a method for extracting electronic medical record information, which is applied to the system for extracting electronic medical record information according to any one of the above-mentioned embodiments, and the method comprises: Acquiring multi-module medical text data; the multi-module medical text data includes doctor-patient conversation voice transcription text, free text medical record and OCR-recognized imaging and laboratory reports; Generating prompt words containing hierarchical structure for a preset medical information extraction task; Combining the prompt words containing hierarchical structure with the preprocessed multi-module medical text data, and inputting them into a preset large language model to obtain a preliminary extraction result containing candidate medical information; Calling a preset medical knowledge base to compare and correct medical terms in the preliminary extraction result; Performing logical verification on the extraction result corrected by the term correction module, the logical verification including at least one of time sequence relationship verification and cross-field information consistency verification to generate final structured medical record information.
[0013] Further, after generating prompt words containing hierarchical structure for a preset medical information extraction task, combining the prompt words containing hierarchical structure with the preprocessed multi-module medical text data, and inputting them into a preset large language model to obtain a preliminary extraction result containing candidate medical information, the method further comprises: Preprocessing the medical text data to remove non-medical related content.
[0014] Compared with the prior art, the present application has the following beneficial effects: 1. Improve the accuracy and standardization of information extraction: By using prompt words containing a hierarchical structure (system layer, task layer, field layer, and verification layer), the output behavior of large language models can be accurately constrained, effectively reducing the problems of model hallucinations, errors, or unstable formats. Combined with the term correction module, medical terms are standardized to ensure that the extraction results meet the clinical terminology standards, thereby significantly improving the accuracy and structuredness of the information.
[0015] 2. Enhance logical consistency and clinical reasonableness: Through the logical verification module, the time sequence relationship (such as the sequence of symptom occurrence, relative time standardization) and cross-field information consistency (such as redundant symptom resolution) are verified, which can effectively avoid logical contradictions such as timeline confusion and field classification errors, ensuring that the generated electronic medical record has consistency and traceability in medical logic, meeting the stringent requirements of clinical decision-making.
[0016] 3. Compatible with multi-source data and efficient denoising: The system can integrate medical and patient conversation transcripts, free text medical records, OCR reports, and other multi-module medical data, and remove non-medical noise such as casual conversation and errors through the text preprocessing module, thereby focusing on key medical information and improving data utilization efficiency and extraction efficiency.
[0017] 4. Improve clinical usability and system integration: The structured medical record information generated can be converted into standard data formats such as JSON and FHIR, seamlessly integrating with hospital information systems or electronic medical record systems, directly supporting clinical diagnosis and treatment, scientific research analysis, and other applications, solving the pain point of non-structured text being difficult to directly utilize, and enhancing the practicality and promotional value of the system.
[0018] In summary, through the synergistic effect of multi-level prompt word design, term standardization, logical verification, and closed-loop correction, the application fundamentally solves the problems of inaccurate, non-standard, and inconsistent logic in existing technologies, and provides a high-reliability and high-availability solution for the automatic generation of electronic medical records.
[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the application. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0021] Figure 1is a system composition schematic diagram of electronic medical record information extraction according to an exemplary embodiment; Figure 2 is a method flow schematic diagram of electronic medical record information extraction according to an exemplary embodiment. DETAILED DESCRIPTION
[0022] In order to make the purposes, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described in detail below. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the scope of protection of the present application.
[0023] In the specific implementation, please refer to Figure 1 , Figure 1 is a system composition schematic diagram of electronic medical record information extraction according to an exemplary embodiment, the system comprises: A data acquisition module 10 is configured to acquire multi-module medical text data, wherein the multi-module medical text data comprises doctor-patient conversation voice transcription text, free text medical record, and OCR-recognized imaging and test report. A prompt word generation module 20 is configured to generate prompt words containing hierarchical structure for a preset medical information extraction task. A model inference module 30 is configured to combine the prompt words containing hierarchical structure and the preprocessed multi-module medical text data, and input them into a preset large language model to obtain a preliminary extraction result containing candidate medical information. A term correction module 40 is configured to call a preset medical knowledge base to compare and standardize the medical terms in the preliminary extraction result. A logic verification module 50 is configured to perform logic verification on the extraction result corrected by the term correction module, wherein the logic verification comprises at least one of time sequence relationship verification and cross-field information consistency verification, so as to generate final structured medical record information.
[0024] In specific implementation, the present application combines guided extraction with multi-level verification, especially introduces a feedback correction mechanism after logical verification failure, forms a closed-loop process of "extraction-verification-correction", which can effectively overcome the randomness and illusion of large language model output, significantly improve the accuracy and coherence of the extracted medical record information in time sequence and medical logic. Secondly, by combining medical knowledge base for term correction, it can convert oral and non-standard medical expressions into unified medical terms, improving the quality and usability of structured data. In addition, the present application can automatically filter noise information, process complex context and time relationship, and correct logical errors automatically, with high automation, effectively reducing the workload of doctors and improving the efficiency of medical record recording. Finally, the present application can output structured data conforming to industry standards, enhancing the traceability, reusability and cross-system sharing ability of medical record data, providing a high-quality data foundation for clinical decision support, medical research and other downstream applications.
[0025] Embodiment 1 Figure 2 is a method flow diagram of electronic medical record information extraction according to an exemplary embodiment. The purpose of this flow is to convert an unstructured medical text (for example, a voice transcription record of doctor-patient conversation) into a structured, standardized and logically consistent electronic medical record data. The functions of each module and the execution steps of the method will be described in detail below in conjunction with a specific application scenario.
[0026] Suppose the system receives a doctor-patient conversation text transcribed by voice recognition technology: "Doctor, I have been coughing for the past two days, and I have a fever. Last night, the highest temperature was 38.9℃, and I felt a little short of breath." Firstly, in the medical data acquisition step S1, the data acquisition module 10 is responsible for receiving the text. It can be understood that the data acquisition module 10 can be configured with multiple interfaces, such as an interface for receiving real-time voice stream and calling external voice recognition services to convert to text, or an interface for receiving text messages pushed by the hospital information system, or an interface for receiving user-uploaded medical record photos. In this scenario, the data acquisition module 10 receives the medical text in the form of the above string.
[0027] Subsequently, the collected text is sent to the text preprocessing module to perform the text preprocessing step. The main function of this module is to purify the input text and remove noise irrelevant to the medical information extraction task. For the input text "Doctor, I have been coughing for two days and have a fever. Last night, the temperature was 38.9°C, and I feel a little short of breath.", the text preprocessing module will perform a series of operations, including but not limited to: 1. Identifying and removing honorifics such as "Doctor"; 2. Based on the pre-set common error mapping table, correcting the errors or homophones that may be introduced by speech recognition; 3. Performing sentence segmentation and punctuation completion to ensure the integrity of the text syntax structure. After processing, the output of the pure text is: "I have been coughing for two days and have a fever. Last night, the temperature was 38.9°C, and I feel a little short of breath." It should be noted that this step helps to reduce the complexity of subsequent large language model processing, thereby improving the accuracy of extraction.
[0028] After preprocessing, the process enters the prompt word generation step S20, which is performed by the prompt word engineering and hierarchical extraction module 20. This step is a key link to guide the large language model to perform accurate extraction. The module generates a prompt word containing structured field definitions and extraction rules for the pre-set medical information extraction task (e.g., generating an outpatient initial diagnosis medical record). This prompt word is not a simple instruction, but a structured text template. For example, module 20 will generate the following prompt word: {"task": "Extract structured medical record information from the provided patient self-report text.","output_schema":{"chief complaint": "string / / The patient's most significant discomfort during this visit, concisely summarized.","present illness": {"onset time":"string / / The relative or absolute time of the first appearance of the symptom.","symptom list": [{"symptom name": "string","symptom description": "string / / Details including the nature, severity, frequency, and duration of the symptom.","relevant time": "string / / The time when the specific symptom appeared or worsened."}]},"vital signs": {"temperature": {"value": "float","unit": "string / / Degrees Celsius.","measurement time": "string"}},"past medical history": "string / / The patient's past medical history, return 'none' if none.","allergy history": "string / / The patient's history of drug or food allergies, return 'none' if none."},"extraction_rules": ["Return results strictly according to the JSON format of 'output_schema'.", "Please parse and output all time-related fields if possible.", "For the 'body temperature' field, if the temperature value mentioned in the text is higher than 37.1 degrees Celsius, it should be judged as 'fever' and reflected in the 'symptom list'.", "If a field is not mentioned in the text, return null."]} The above prompts clearly define the fields to be extracted (such as chief complaint, present medical history, etc.), the data type and hierarchical structure of each field, and attach specific extraction rules.
[0029] Accordingly, the process proceeds to the large language model inference step S3. The model inference module 30 combines the prompt words generated in the previous step with the preprocessed medical text into a complete input and calls a pre-configured large language model service. This large language model can be a general-purpose large model based on a transformer architecture, or a specialized model fine-tuned with medical data. An example of the content input to the model is as follows: [System Instruction]: Please process the user-provided text according to the following task definition, output format, and rules. [Task Definition, Output Format, and Rules]: {...the prompt words in the above JSON format...} [User-provided Text]: "I've been coughing and have a fever for the past two days. I took my temperature last night, and it was as high as 38.9℃. I feel a bit short of breath." After understanding these instructions and context, the large language model performs inference and generates a preliminary extraction result containing candidate medical information. This result follows the format requirements of the prompt words, but its content may contain irregularities or logical inconsistencies. For example, the model might output: { "Chief Complaint": "Cough, Fever", "Present Illness": { "Onset Time": "These past two days", "Symptom List": [ { "Symptom Name": "Cough", "Symptom Description": "Constant cough", "Related Time": "These past two days"}, { "Symptom Name": "Fever", "Symptom Description": "Highest temperature 38.9℃", "Related Time": "Last night"}, { "Symptom Name": "Shortness of breath", "Symptom Description": "Feeling a little short of breath", "Related Time": null} ]}, "Vital Signs": { "Temperature": { "Value": 38.9, "Unit": "℃", "Measurement Time": "Last night"}}, "Past Medical History": null, "Allergy History": null} Next, the process enters the medical knowledge base correction step S4, where the medical knowledge base correction module 40 processes the preliminary extraction results. The main task of this module is terminology standardization. Internally, the module connects to one or more pre-defined medical knowledge bases, such as a database containing a large number of colloquial expressions and their correspondences with standard medical terms. This database can be stored in key-value pairs, such as {"shortness of breath": "difficulty breathing", "lack of energy": "fatigue", "diarrhea": "diarrhea"}. The medical knowledge base correction module 40 iterates through the "symptom name" and "symptom description" fields in the preliminary extraction results, replacing the identified colloquial expression "shortness of breath" with the standard medical term "difficulty breathing". The corrected result then becomes: { ... "Symptom List": [ ..., { "Symptom Name": "Difficulty Breathing", "Symptom Description": "Feeling a little short of breath", "Related Time": null} ], ...} This step greatly improves the standardization and professionalism of the final generated medical records, laying the foundation for subsequent data analysis and system integration.
[0030] Next, the logical consistency and chronological order verification module 50 performs the logical consistency verification step S5, which is a core step in ensuring the accuracy and reliability of medical record information. This module performs at least two verifications on the results after terminology correction: chronological order relationship verification and cross-field information consistency verification.
[0031] For temporal relationship verification, module 50 parses all time-related fields, such as “two days ago” in “onset time” and “last night” in “associated time”. It converts these relative time expressions into standardized absolute timestamps based on the current date (assuming the operation is performed on August 17, 2025). For example, “two days ago” is parsed as the start of the onset date range, i.e., “2025-08-15”; “last night” is parsed as “2025-08-16”. Module 50 also verifies whether the logical order of event occurrences is reasonable based on these timestamps, e.g., the fever time (2025-08-16) is later than or equal to the first onset time (2025-08-15), which is a valid logical relationship.
[0032] For cross-field information consistency verification, module 50 checks whether there are contradictions or redundancies between different fields. In this example, it verifies whether the rule in the prompt is correctly executed: it checks that the temperature in “vital signs” is 38.9°C, which is higher than 37.1°C, and that there is indeed a “fever” symptom in the “symptom list” in “present illness history”, confirming that the information is consistent. In addition, it checks whether there is a simple repetition of symptoms in “chief complaint” and “present illness history”, and may integrate them. In this example, the verification module confirms that the logic is basically consistent.
[0033] In this embodiment, since all verifications pass, the process smoothly enters the final step of storing structured medical record information. The structured output and storage module is responsible for converting the verified and logically rigorous medical record information into a pre-defined final data format. This format can be a common JSON or XML, or a format that meets the medical industry standard, such as Fast Healthcare Interoperability Resources or Health Level Seven Standard. For example, the final generated resource may contain a Condition resource (for recording cough, fever, and difficulty breathing) and an Observation resource (for recording a temperature of 38.9°C), and is associated by reference to the same Encounter resource. Finally, this structured data is pushed to the hospital information system or electronic medical record system for storage and utilization.
[0034] Through the above steps, the system and method of this embodiment successfully convert a piece of colloquial, unstructured conversation into high-quality structured medical record data containing accurate timelines, standardized terminology, and logically consistent information.
[0035] Embodiment 2 This embodiment, based on Embodiment 1, focuses on demonstrating how the system achieves closed-loop correction through a feedback correction mechanism when errors are detected during logical verification. A judgment step is added after the logical consistency verification step. If the verification fails, the step of generating error correction prompts is triggered, and the system re-enters the large language model inference step.
[0036] Let's take a different medical text scenario as an example. The data acquisition module 10 receives a text message: "I've been dizzy for three days and have a history of high blood pressure." After text preprocessing, the text is fed into subsequent steps. In the initial large language model inference step, because the large language model may not perfectly distinguish between acute and chronic issues in the medical record, it may produce a logically flawed preliminary extraction result. For example, the model might consider "hypertension" as both a problem in the current visit and a past medical history, resulting in the following output: { "Chief Complaint": "Dizziness, hypertension", "Present Illness": { "Onset Time": "Three days ago", "Symptom List": [ { "Symptom Name": "Dizziness", "Symptom Description": "Dizziness for three days", "Related Time": "Three days ago"} ]}, "Past Medical History": "History of Hypertension", ...} After the medical knowledge base correction step (in this case, there are no obvious terms that need to be corrected), the result is sent to the logical consistency and time sequence verification module for verification (step S5). This module has a series of deeper medical logic rules built in, one of the key rules being: "The chief complaint field should record the acute or major problem that caused this visit, while long-term chronic diseases should be recorded in the past medical history field and should not appear in the chief complaint." When performing cross-field information consistency checks, the verification module 50 detected that the term "hypertension" appeared in both the "chief complaint" and "past medical history" fields. According to the built-in rule base, "hypertension" is a chronic disease, and its appearance in the "chief complaint" field constitutes a logical conflict.
[0037] At this point, in the judgment step, the verification result is determined to be "No" (i.e., verification failed), thus triggering the feedback correction process. The process then proceeds to the step of generating error correction prompts. The logical consistency and time sequence verification module generates a targeted error correction prompt based on the specific error type detected. This error correction prompt is no longer a general extraction instruction, but a specific correction instruction. For example, the system will generate the following error correction information, which will then be packaged into a new prompt by the prompt generation module: {"context": {"original_text": "I have been dizzy for three days, and I have a history of high blood pressure.", "previous_attempt": {"output": { "chief complaint": "dizziness, high blood pressure", "past medical history": "history of high blood pressure",...}, "error_found": "Logical inconsistency: the chronic condition 'high blood pressure' is incorrectly included in the 'chief complaint' field."}, "correction_instruction": "Please correct the errors in the 'previous_attempt' above. According to medical writing standards, the 'chief complaint' field should not contain a history of chronic conditions. Please remove 'high blood pressure' from the 'chief complaint' field and generate a new complete JSON output containing only the correct information."} The information extraction system now initiates a second request to the large language model service, a correction request. The model inference module 30 sends this new prompt containing the context and explicit correction instruction to the large language model.
[0038] Upon receiving this correction request, the large language model focuses on the pointed-out error. Due to the specific instruction, the model can easily understand and perform the correction. It re-examines the original text and the previous output, then generates a corrected result. This process can be seen as a kind of "self-reflection and correction" of the model. The corrected output is: { "chief complaint": "dizziness", "present illness history": { "time of onset": "three days ago", "symptom list": [ { "symptom name": "dizziness", "symptom description": "dizziness for three days", "related time": "three days ago"} ]}, "past medical history": "history of high blood pressure",...} This corrected result is again sent to the medical knowledge base correction module 40 and the logic verification module 50. This time, when the logic verification module 50 performs the verification again, it finds that "high blood pressure" only appears in the "past medical history" field, and the "chief complaint" field content is correct, the logical inconsistency has been eliminated. Therefore, in the judgment step, the verification result is "yes", and the process continues.
[0039] Finally, the structured output and storage module formats and stores or pushes this closed-loop corrected and logically rigorous medical record information to the downstream system.
[0040] Through this "extraction-verification-feedback-correction" closed-loop mechanism, the system of the present application can actively find and correct the logical errors that the large language model may have made in the initial reasoning, thereby greatly improving the accuracy and clinical usability of the final output result, achieving significant improvement compared to the existing linear extraction process.
[0041] Embodiment 3 This embodiment elaborates on how the prompt generation module 20 constructs and applies complex hierarchical prompts to achieve fine-grained control over the large language model's reasoning process. As an optional implementation, the hierarchical prompt structure is one of the key technologies to improve extraction quality and cope with complex medical scenarios.
[0042] According to a specific implementation scheme of the present application, the prompt can be designed to contain four levels: the system level, the task level, the field level, and the verification level. The four levels work together to set a complete set of behavior standards and workflow for the large language model.
[0043] Take a more complex patient self-reporting scenario as an example: "I have had diabetes for ten years and have been taking metformin. Yesterday I started a fever and a cough, and I took some cephalexin, but it didn't work." When this text enters the prompt generation step, the prompt hierarchical module 30 will construct a structured prompt containing four levels as follows: System level prompt: This layer is responsible for defining global, bottom-level behavior constraints to ensure the format and basic specifications of the output. It usually serves as the initial instruction for interacting with the large language model. For example: "system_prompt": "You are a professional medical information extraction AI assistant. Your task is to extract information from the provided text according to user instructions. All your outputs must strictly follow the JSON format specified by the user. All time-related strings must be parsed and converted to 'YYYY-MM-DD' format. For fields that cannot be found in the text, you must return null values instead of omitting the field." This layer of instructions ensures that the "skeleton" of the model output remains stable and predictable regardless of the task.
[0044] Task level prompt: This layer sets the overall goal and strategy for this specific extraction task, helping the model understand the context and determine the priority of information extraction. For example: "task_prompt": "The current task is to process a medical record about a respiratory tract infection in an outpatient setting. Please focus on the symptoms related to this episode (present illness) and the recent medication history. At the same time, you must pay attention to strictly distinguishing these acute problems from the patient's chronic medical history (past medical history) and their long-term medication." This layer of instructions guides the model to focus on acute information such as "fever, cough, cephalexin", while being vigilant not to confuse it with "diabetes, metformin".
[0045] Field level prompt: This layer is the most specific, as it defines the name and data type for each medical field to be extracted, and may also include specific extraction templates and constraints containing medical knowledge. For example: The json file "field_definitions": [{"name": "Present Illness History","description": "Records symptoms and events directly related to this visit.","structure": { "Symptoms": "string[]", "Start Time":"string", "Related Medications": "string[]"}}, {"name": "Past Medical History", "description": "Records the patient's long-term, chronic disease history."} "structure": { "Disease Name": "string", "Disease Course": "string", "Long-Term Medication": "string[]"}, "rules": ["For example: hypertension, diabetes, coronary heart disease, etc."]},{"name": "Allergy history","description": "Record the patient's known drug or food allergy history." "rules": ["If the text mentions antibiotics such as 'cephalosporins', special attention should be paid to whether there are any allergy descriptions."]}] At the field level, examples are added to "Patient History" to help the model better identify chronic diseases; context-sensitive dynamic rules are added to "Allergy History" to enhance risk identification capabilities.
[0046] Validation Prompt: As an innovation of this application, this layer embeds instructions within the prompts to guide the large language model to perform a preliminary "self-check." This encourages the model to perform a pre-validation of logical consistency before generating the final output. For example: "validation_prompt": "Before generating the final JSON output, please perform the following self-check: 1. Confirm that the relevant medications (such as cephalosporins) in the present medical history are logically related to the reported symptoms (such as fever and cough).
[0047] 2. Confirm that the 'long-term medication' (such as metformin) in the 'past medical history' matches the recorded chronic disease (such as diabetes).
[0048] 3. Check for any discrepancies between the information in the 'present illness' and 'past medical history' sections. When these four levels of prompts are integrated and input into the large language model inference module 30 along with the original text, the model performs a highly structured and controlled inference process. It first understands the global format requirements (system layer), then clarifies the task focus (task layer), then matches and extracts information from the text according to detailed field definitions and rules (field layer), and finally performs an internal logical review based on the instructions of the validation layer before output.
[0049] Therefore, for the input text "I've had diabetes for ten years and have been taking metformin. Yesterday I started having a fever and cough, and I took some cephalosporin, but it didn't work.", a model that follows this hierarchical prompt word structure is more likely to directly generate a high-quality, logically clear preliminary result, such as: { "Chief Complaint": "Fever, Cough", "Present Illness": { "Symptoms": ["Fever", "Cough"], "Start Time": "2025-08-16", "Related Medications": ["Cephalexin"]}, "Past Medical History": { "Disease Name": "Diabetes Mellitus", "Duration of Disease": "Ten Years", "Long-Term Medications": ["Metformin"]}, "Allergy History": null} This result has quite accurately separated acute and chronic information and associated drugs with corresponding medical histories. Even if the result still has minor flaws, the processing burden on the subsequent medical knowledge base correction module 40 and logic verification module 50 has been greatly reduced.
[0050] By using this layered prompting word approach to enhance the application, the system in this application can avoid common errors to the greatest extent possible during the reasoning stage, making the entire information extraction process more efficient and accurate, and the output data richer and more intelligent.
[0051] It is understood that the same or similar parts in the above embodiments can be referred to each other, and the contents not described in detail in some embodiments can be referred to the same or similar contents in other embodiments.
[0052] It should be noted that in the description of this application, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "multiple" or "more" means at least two.
[0053] It should be understood that when an element is referred to as being "on" or "connected to" another element, it can be directly on or connected to the other element or intervening elements can be present. In addition, the term "connected" as used herein can include wirelessly connected. Also, the term "on" as used herein can include "directly on" and "indirectly on" when used in the context of interlayers.
[0054] Any process or method described in flow chart form or otherwise described herein can be understood as representing a module, segment, or portion of code that includes one or more executable instructions for implementing specific logical functions or steps in the process, and the various embodiments of the application can include additional or fewer steps performing the described functions in the illustrated or discussed order, including substantially simultaneous execution of the functions, or in reverse order as appropriate, as would be understood by one of ordinary skill in the art.
[0055] It should be understood that each of the elements of the present application can be realized in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be realized in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if realized in hardware, as in another embodiment, any of the following technologies known in the art or a combination thereof can be used: discrete logic circuitry having logic gates for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.
[0056] Those of ordinary skill in the art can understand that all or part of the steps carried out by the above-mentioned embodiments can be completed by programs instructing relevant hardware, and the programs can be stored in a computer readable storage medium, and when executed, include one or a combination of steps of the method embodiments.
[0057] In addition, each functional unit in each embodiment of the present application can be integrated into one processing module, or each unit can exist physically separately, or two or more units can be integrated into one module. The above-mentioned integrated module can be realized in the form of hardware or in the form of a software functional module. When the integrated module is realized in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer readable storage medium.
[0058] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0059] In the description of the specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are contained in at least one embodiment or example of the present application. In the specification, the illustrative description of the above terms does not necessarily mean the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0060] Although the embodiments of the present application have been shown and described above, it is understood that the above-described embodiments are exemplary, and cannot be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above-described embodiments within the scope of the present application.
Claims
1. A system for extracting electronic medical record information, characterized in that, The system includes: The data acquisition module is used to collect medical text data from multiple modules, including: doctor-patient dialogue speech-to-text, free text medical records, and OCR-recognized imaging and laboratory reports. The prompt word generation module is used to generate prompt words with a hierarchical structure for preset medical information extraction tasks; The model reasoning module is used to combine the hierarchical prompt words with the preprocessed medical text data from the multiple modules and input them into a preset large language model to obtain preliminary extraction results containing candidate medical information. The terminology correction module is used to call a preset medical knowledge base to compare and standardize the medical terms in the preliminary extraction results. The logical verification module is used to perform logical verification on the extraction results after correction by the terminology correction module. The logical verification includes at least one of time sequence relationship verification and cross-field information consistency verification to generate the final structured medical record information.
2. The system according to claim 1, characterized in that, The hierarchical prompts include: The system layer defines the global output format and constraints; the task layer sets the extraction priority and strategy; the field layer builds templates and constraints for specific medical fields; and the validation layer guides the large language model to perform self-checks.
3. The system according to claim 1, characterized in that, The logical verification performed by the logical verification module includes at least one of the following operations: Perform time sequence validation to sort the chronological order of symptoms based on the time expressions in the original text, and convert relative time expressions into standardized absolute timestamps; Perform cross-field information consistency checks to redundancy resolution of symptoms of duplication or semantic similarity.
4. The system according to claim 1, characterized in that, The logic verification module is also used to: when the logic verification finds an error or inconsistency, generate an error correction prompt word, and control the model inference module to re-input the error correction prompt word, the original text, and the context information into the large language model, instructing it to make targeted corrections to form a closed-loop correction process.
5. The system according to claim 1, characterized in that, The system also includes: A text preprocessing module, configured between the data acquisition module and the model inference module, is used to preprocess the medical text data to remove chatter, noise, and non-medical content.
6. The system according to claim 1, characterized in that, Also includes: The structured output module is used to convert the final structured medical record information into a preset data format and connect it to the hospital information system or electronic medical record system; the preset data format includes JSON, XML, FHIR or HL7.
7. A method for extracting electronic medical record information, applied to the electronic medical record information extraction system according to any one of claims 1-6, characterized in that, The method includes: Collect medical text data from multiple modules; the medical text data from multiple modules includes: doctor-patient dialogue speech-to-text, free text medical records and OCR-recognized imaging and laboratory reports; Generate prompts with a hierarchical structure for the preset medical information extraction task; The hierarchical prompt words are combined with the preprocessed multi-module medical text data and then input into a preset large language model to obtain preliminary extraction results containing candidate medical information. A preset medical knowledge base is invoked to compare and standardize the medical terms in the preliminary extraction results; Logical verification is performed on the extraction results after correction by the terminology correction module. The logical verification includes at least one of time sequence relationship verification and cross-field information consistency verification to generate the final structured medical record information.
8. The method according to claim 7, characterized in that, Before generating hierarchical prompt words for a preset medical information extraction task, combining these hierarchical prompt words with preprocessed multi-module medical text data, and inputting them into a preset large language model to obtain preliminary extraction results containing candidate medical information, the process further includes: The medical text data is preprocessed to remove non-medical related content.
Citation Information
Patent Citations
Outpatient service electronic medical record generation method based on Chinese medical big model
CN117253576A
Medical large model question answering system based on medical records
CN118016324A
Ophthalmology information extraction system based on electronic medical record
CN119626573A
Medical history identification method and system based on large model and expert strategy
CN120032782A
Electronic medical record intelligent evaluation method based on complex quality control indexes
CN120564931A
Cited By
Prompt word optimization method, device and equipment based on AI, RPA, LLM and AI Agent
CN121525881A
Intelligent knowledge extraction and structuring method for modern assembly type historical literature
CN121580967A
Medical ward round record generation method and system based on large language model
CN122091061A
Clinical research information batch extraction system and extraction method
CN122117190A
A clinical research information batch extraction system and extraction method
CN122117190B