Electronic medical record analysis method, system, equipment and medium
By constructing a basic standard knowledge base of electronic medical records and a large language model combined with a split dictionary, the interpretability and hardware requirements of electronic medical record analysis in the existing technology are solved, and efficient and accurate structured data conversion is achieved to adapt to the standardization and standardized management of medical data.
Patent Information
- Application Number
- CN202510133583.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-07-04
AI Technical Summary
When the existing technology relies on deep learning models for electronic medical record analysis, there are problems such as poor interpretability, difficult to guarantee the quality of the training set, high hardware configuration requirements, long response and loss of computing power for complex text processing tasks.
By constructing a basic standard knowledge base for electronic medical records, identifying data element tags in combination with large language models, and using split dictionary to split text to generate structured data.
It realizes efficient analysis of electronic medical records from unstructured and semi-structured text to structured data, improves analysis efficiency and accuracy, reduces implementation difficulty, supports the analysis of multiple medical records types, and can be updated according to medical specifications.
Smart Images

Figure CN120257972A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical data governance, and more specifically relates to a method, system, device and medium for parsing electronic medical records. Background Art
[0002] An electronic medical record (EMR, Electronic Medical Record) refers to a system in which doctors record patient information, diagnosis results, treatment processes and medical files during medical treatment. Electronic medical records entered in free text form need to be structurally parsed in order to extract data elements that can be further used, such as chief complaints, medical histories, diagnoses, etc.
[0003] Currently, the mainstream parsing method is based on artificial intelligence technologies such as natural language processing (NLP, Natural Language Processing), and deep learning models are used to extract various data elements from medical record texts. However, relying solely on natural language processing technology for electronic medical record parsing has the defect of poor interpretability. At the same time, deep learning models need to annotate a large number of medical record samples to obtain a training set, and manual annotation is prone to annotation errors, so the quality of the training set is difficult to guarantee; moreover, a large amount of data support is a necessary basis for the accuracy of deep learning models, and personal data with extremely strong privacy such as electronic medical records is particularly difficult to collect and process, which is not feasible for enterprises and individuals.
[0004] With the emergence of large language models (LLMs, Large Language Model), a new idea has been provided for the parsing of electronic medical records. Large models have excellent text understanding and generation capabilities, so there are more and more products that directly use large models for the structuring of electronic medical records.
[0005] Although LLM large models have significant advantages in text processing capabilities, they have high requirements for hardware configuration, and ordinary servers are difficult to support frequent model call requests; moreover, the response time of large models is relatively long, and the more complex the text, the more significant the disadvantage in terms of response time. Finally, with the accumulation of the number of medical records, complex text processing tasks will continuously consume the computing power of large models, resulting in a substantial increase in the probability of subsequent understanding ambiguities and even text hallucinations. Summary of the Invention
[0006] Aiming at the above problems, the purpose of the present invention is to provide a method, system, device and medium for parsing electronic medical records. By using a large language model to identify the data element tags of electronic medical records and then combining knowledge base rules to establish a splitting dictionary, the parsing conversion of electronic medical records from unstructured and semi-structured texts to structured data is realized. The present invention can reduce the implementation difficulty while ensuring the degree of intelligence of parsing and improving the efficiency of electronic medical record parsing.
[0007] In order to achieve the above object, the present invention is implemented through the following technical solutions: In a first aspect, an embodiment of the present application provides an electronic medical record parsing method, comprising: Based on the existing text format and structure of electronic medical records, a knowledge base of basic specifications for electronic medical records is constructed; Obtaining an electronic medical record to be parsed, and identifying the medical record type of the parsed electronic medical record based on the title of the electronic medical record; According to the medical record type, a large language model is used to extract labels from the electronic medical record text to be parsed and generate a label list; According to the tag list, the basic standard knowledge base of electronic medical records is traversed to extract the standard tag names and corresponding splitting regular expressions, and generate a splitting dictionary; The electronic medical record text to be parsed is split using a split dictionary, the text is split into a list of sentences or paragraphs, the tag name and the tag content of the tag are extracted, a parsing table is generated, and it is stored in the database.
[0008] In an optional implementation, the basic standard knowledge base of electronic medical records includes: medical record types, standard tag names, and corresponding splitting regular expressions; The types of medical records include: admission records, discharge records, and first medical records; The standard label names include: chief complaint, personal history, current medical history, past medical history, family history, marital history, treatment process information, differential diagnosis information, and physician signature.
[0009] In an optional implementation, the obtaining of the electronic medical record to be parsed and identifying the medical record type of the parsed electronic medical record based on the title of the electronic medical record include: Obtain the electronic medical record to be parsed, and divide it into medical record text with title and medical record text without title; For medical record texts with titles, determine the medical record type based on the medical record title; For untitled medical record texts, the medical record texts are input into a preset electronic medical record classifier, and the medical record texts are classified into preset medical record types by extracting the semantics and structure of the medical record texts.
[0010] In an optional implementation, the method of extracting labels from the electronic medical record text to be parsed using a large language model according to the medical record type and generating a label list includes: According to the medical record type, a large language model is used to extract labels from the electronic medical record text to be parsed. The standard label names in the electronic medical record text are extracted using the prompt words of the standard label names, and duplicate standard label names are deleted through iterative deduplication. Generate a tag list based on the proposed standard tag names.
[0011] In an optional implementation, the method of traversing the basic standard knowledge base of electronic medical records according to the tag list, extracting the standard tag name and the corresponding splitting regular expression, and generating a splitting dictionary includes: According to each standard tag name in the tag list, find the same standard tag name by traversing the basic specification knowledge base of electronic medical records, and extract the corresponding splitting regular expression; Based on the traversal results, a split dictionary is constructed with the standard tag name as the key and the corresponding split regular expression as the value.
[0012] In an optional embodiment, the electronic medical record text to be parsed is split using a split dictionary, the text is split into a list of sentences or paragraphs, the tag name and the tag content of the tag are extracted, a parsing table is generated, and stored in a database, including: Traverse the electronic medical record text to be parsed according to each value in the split dictionary to find the matching string; Replace the matched string with the key value corresponding to the value value, so as to replace the original tag name in the electronic medical record text to be parsed with the standard tag name, and enter the split breakpoint mark at the replacement point of the text; According to the breakpoint mark of split in the electronic medical record text to be parsed, the corresponding text is split into a list of sentences or paragraphs; Convert each sentence or paragraph into a preset standard format; the preset standard format includes a standard tag name and tag content; Generate a parsing table in json format based on the converted sentence or paragraph list; The parsing table is stored in the database through the SQL statement.
[0013] In an optional implementation, the large language model adopts the Qwen-72B lightweight model, and the database adopts an Oracle database, a MySQL database, or an openGauss database.
[0014] In a second aspect, the present application also provides an electronic medical record parsing system, including: A knowledge base construction module is used to build a basic specification knowledge base of electronic medical records based on the text format and structure of existing electronic medical records; A medical record classification module is used to obtain the electronic medical record to be parsed and identify the medical record type of the parsed electronic medical record based on the title of the electronic medical record; The label list generation module is used to extract labels from the electronic medical record text to be parsed based on the medical record type and generate a label list using a large language model; The split dictionary construction module is used to traverse the basic standard knowledge base of electronic medical records according to the label list, extract the standard label name and the corresponding split regular expression, and generate a split dictionary; The splitting and parsing module is used to split the electronic medical record text to be parsed using the splitting dictionary, split the text into sentences or paragraph lists, extract the tag name and the tag content of the tag, generate a parsing table, and store it in the database.
[0015] In a third aspect, an embodiment of the present application further provides an electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of the electronic medical record parsing method as described in any one of the above items are implemented.
[0016] In a fourth aspect, an embodiment of the present application further provides a storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the electronic medical record parsing method as described in any one of the above items are implemented.
[0017] It can be seen from the above technical solutions that the present invention has the following advantages: In the electronic medical record parsing method provided in this application, the data element labels of the electronic medical record are identified through a large model, and then a decomposition dictionary is established in combination with the knowledge base rules to achieve structured management of the electronic medical record. This method combines the advantages of the intelligence of natural language processing technology and the strong controllability of knowledge base logic, improves the efficiency of electronic medical record parsing, and reduces the difficulty of implementation.
[0018] This application achieves the standardization and normalization of electronic medical records by building a basic knowledge base of electronic medical records, which helps to unify electronic medical record data from different sources and formats, and provides a basis for subsequent data analysis and utilization.
[0019] This application uses a large language model to extract labels and combines a split dictionary to split the electronic medical record text, which can efficiently parse the electronic medical record and extract key information, greatly improving the efficiency of data processing and reducing the need for manual intervention.
[0020] This application supports the parsing of multiple medical record types, including admission records, discharge records, first medical records, etc., and can add new medical record types and tags as needed. At the same time, the basic specification knowledge base of electronic medical records can also be continuously updated and improved to adapt to new medical standards and data formats.
[0021] For untitled medical record texts, this application uses a preset electronic medical record classifier to perform intelligent classification, thereby improving the accuracy and intelligence level of data processing.
[0022] This application realizes storing the parsed data in a database in a standard format, facilitating subsequent data query, analysis, and utilization. At the same time, the storage of data also realizes traceability, which helps to control and improve the medical quality.
[0023] This application supports storing the parsed data in multiple databases such as Oracle database, MySQL database, or openGauss database, improving the data compatibility and flexibility. Brief Description of the Drawings
[0024] To more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0025] Figure 1 It is a schematic flowchart of an electronic medical record parsing method provided by this application.
[0026] Figure 2 It is a schematic flowchart of another electronic medical record parsing method provided by this application.
[0027] Figure 3 It is a schematic structural diagram of an electronic medical record parsing system provided by this application.
[0028] Figure 4 It is a schematic structural diagram of an electronic device provided by this application. Detailed Embodiments
[0029] In the following, the specific steps of the electronic medical record parsing method will be described in detail, and various embodiments of the present disclosure will be described more comprehensively. The present disclosure can have various embodiments and adjustments and changes can be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but the present disclosure should be understood to cover all adjustments, equivalents, and / or alternative solutions falling within the spirit and scope of the various embodiments of the present disclosure.
[0030] In the following, the term "comprising" or "may comprise" that can be used in various embodiments of the present disclosure indicates the presence of the disclosed functions, operations, or elements, and does not limit the addition of one or more functions, operations, or elements. Further, as used in various embodiments of the present disclosure, the terms "comprising", "having" and their cognates are only intended to indicate a specific feature, number, step, operation, element, component, or combination of the foregoing items, and should not be construed as precluding the existence or addition of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing items.
[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0032] Please refer to Figure 1 The following is a method flowchart of an electronic medical record parsing method in a specific embodiment. The method includes: S1: Based on the text format and structure of existing electronic medical records, construct a knowledge base of basic electronic medical record specifications.
[0033] Exemplarily, the implementation of this method first requires establishing a knowledge base of basic electronic medical record specifications. The knowledge base of basic electronic medical record specifications analyzes common medical record text formats and structures, summarizes patterns and rules therefrom, and formulates and records parsing rules based on the analysis results, including the positioning of information paragraphs and the extraction methods of common data elements, etc. After accumulation and precipitation, a knowledge base of basic electronic medical record specifications is formed.
[0034] Specifically, the knowledge base of basic electronic medical record specifications is used to record medical record types, standard tag names, and corresponding splitting regular expressions. Medical record types include: admission record, discharge record, and initial course record. Standard tag names include data element names such as chief complaint, five histories, preliminary diagnosis, and physician signature. For example, information such as chief complaint, personal history, current history, past history, family history, marriage and childbearing history, diagnosis and treatment process information, differential diagnosis information, and physician signature.
[0035] The structure of the knowledge base of basic electronic medical record specifications is shown in Table 1 below: Table 1: Reference table for the structure of the knowledge base of basic electronic medical record specifications
[0036] In order to more specifically explain the content of the basic standard knowledge base of electronic medical records, taking the admission record as an example, the specific content of the basic standard knowledge base of electronic medical records is shown in Table 2 below: Table 2: Content reference table of the basic standard knowledge base of electronic medical records
[0037] S2: Obtain the electronic medical record to be parsed, and identify the medical record type of the parsed electronic medical record based on the title of the electronic medical record.
[0038] For example, before parsing medical records, the medical record text must first be intelligently cleaned to identify the medical record type (admission record, discharge record, first medical record, etc.) as a basis for subsequent calls to specific terminology knowledge bases and parsing rules.
[0039] First, the electronic medical records to be parsed are obtained and divided into titled medical record texts and untitled medical record texts.
[0040] For medical record texts with titles, such as those in XML, HTML and other formats, the data elements are usually relatively regular, so the medical records are classified by identifying the medical record titles to determine the medical record type; for untitled medical record texts, the electronic medical record text data set is used to train an electronic medical record classifier to intelligently process the complex semantics and structure in the medical record text and classify the medical records into the predetermined medical record types in the training set.
[0041] S3: Based on the medical record type, a large language model is used to extract labels from the electronic medical record text to be parsed and generate a label list.
[0042] For example, after identifying the medical record type, a large language model is used to extract labels for the electronic medical record text to be parsed based on the medical record type, and the names of data elements such as the complaint, preliminary diagnosis, and physician's signature in the medical record are identified as the basis for subsequent large-scale parsing and stored in the form of a label list.
[0043] Specifically, the standard tag names in the electronic medical record text are first extracted using the standard tag names as prompt words, and repeated standard tag names are deleted through iterative deduplication; then a tag list is generated based on the proposed standard tag names. Among them, the large language model can use the Qwen-72B lightweight model.
[0044] S4: According to the tag list, traverse the basic specification knowledge base of electronic medical records, extract the standard tag names and corresponding splitting regular expressions, and generate a splitting dictionary.
[0045] Exemplarily, based on the tag list, traverse the Knowledge Base of the Basic Specifications for Electronic Medical Records, extract the standard name of the tag as the key, and extract the splitting regular expression of this tag accumulated in the knowledge base as the value to form a splitting dictionary.
[0046] Specifically, first, for each standard tag name in the tag list, traverse the Knowledge Base of the Basic Specifications for Electronic Medical Records to find the same standard tag name, and extract the corresponding splitting regular expression.
[0047] Then, based on the traversal results, construct a splitting dictionary with the standard tag name as the key and the corresponding splitting regular expression as the value.
[0048] S5: Use the splitting dictionary to split the electronic medical record text to be parsed, split the text into a list of sentences or paragraphs, extract the tag name and the tag content of the tag, generate a parsing table, and store it in the database.
[0049] Exemplarily, when the electronic medical record text to be parsed is unstructured medical record text, in this step, parse and split the medical record text according to the splitting dictionary, generate a parsing table, and store it in the database.
[0050] First, traverse the electronic medical record text to be parsed according to each value in the splitting dictionary to find the matching string.
[0051] If a matching string can be found, replace the matching string with the key value corresponding to the value to replace the original tag name in the electronic medical record text to be parsed with the standard tag name, that is, map the original tag name in the text to the standard name, and mark a split breakpoint at the replacement position in the text.
[0052] Then, for the split breakpoint identifier of split in the electronic medical record text to be parsed, split the corresponding text into a list of sentences or paragraphs; the format of each sentence or each paragraph is "key + original content".
[0053] Next, convert each sentence or each paragraph into a preset standard format; the preset standard format includes the standard tag name and the tag content. Specifically, process each sentence in the list of split sentences or paragraphs one by one, split at the key, where the key is the standard tag name, and the original text after the key is the tag content.
[0054] Finally, generate a parsing table in json format based on the converted list of sentences or paragraphs; and store the parsing table in the database through the database insertion SQL statement. The database can use Oracle database, MySQL database or openGauss database.
[0055] In this embodiment, through the comprehensive application of means such as the construction of an electronic medical record basic specification knowledge base, large language model technology, and intelligent splitting and standardized storage, the efficient and accurate parsing of electronic medical record texts is achieved. It can not only intelligently identify the types of medical records, use large language models to deeply mine and extract standard tag names, but also accurately split medical record texts into sentences or paragraphs by constructing a splitting dictionary, and at the same time extract and standardize the storage of tag names and tag contents. This process not only significantly improves the accuracy and efficiency of electronic medical record parsing, but also effectively promotes the standardized and normalized management of medical data, providing strong support for the improvement of medical quality, the development of medical informatization, and the in-depth mining and utilization of medical data.
[0056] Furthermore, as a refinement and extension of the specific implementation method of the above embodiment, in order to fully illustrate the specific implementation process in this embodiment, another method for parsing electronic medical records is provided.
[0057] Refer to Figure 2 As shown, this embodiment discloses a method for parsing electronic medical records, which specifically includes the following steps: 1. Construct a medical record index.
[0058] Generally, all electronic medical records under the same medical institution are stored in one table, which includes medical record texts containing key medical information such as admission records, discharge records, and surgical records, and also stores various electronic documents such as notices, consent forms, and evaluation forms, which are generally quite mixed. Therefore, the first thing to solve is to be able to distinguish different types of medical records.
[0059] The storage logics of EMR systems from different manufacturers for electronic medical records are different. In the problem of medical record classification, it can be divided into two major categories: with index and without index. For electronic medical record tables with indexes, the types of medical records can be identified from specific fields in the database; for electronic medical record tables without indexes, a classifier trained using deep learning algorithms is pre-employed to classify the medical records and generate index fields.
[0060] 2. Randomly sample the types of medical records to be parsed.
[0061] According to the index, randomly sample a certain type of medical record text to form a sample pool. For example, if the target is to parse all admission records of a certain medical institution, first filter out all admission record medical records in the electronic medical record table through the index, and then use a third-party database connection package in Python to execute SQL for random sampling, and randomly select some admission record medical record texts as samples for the subsequent large model to identify tags. Among them, the random function is selected according to the type of the source database. For example, the dbms_random function is used for Oracle databases, and the RANK function can be used for MySQL databases.
[0062] The purpose of sampling is to improve the parsing accuracy as much as possible while ensuring the system efficiency. Since the next step uses a large LLM to analyze and extract data elements from medical records, directly calling the large model to identify all medical record texts under a medical institution will be very time-consuming. According to the complexity of medical record texts, the response time of the large model ranges from 5s to 50s for a single medical record, and the amount of electronic medical record data under a medical institution is usually in the millions, and the number of medical records in individual hospitals can reach tens of millions or even hundreds of millions. Therefore, sampling is needed to avoid full-table traversal. Secondly, sampling can avoid over-consuming the computing power of the large model and release the hardware pressure at the same time. Generally, the more samples are taken, the better the parsing effect of the electronic medical records of this medical institution will be.
[0063] 3. Use the large LLM to identify labels for the sample medical records and generate a label list.
[0064] For each sample medical record, first call the large LLM to identify the data element labels for each medical record sample and generate a label list; then iterate through the label list. If a new label is extracted from the next medical record sample, it is incorporated into the label list.
[0065] For example, use the Qwen-72B lightweight model with the prompt "You are an expert in extracting electronic medical record labels. Extract the electronic medical record-related entity labels from the given electronic medical record, and give priority to extracting the labels in the knowledge base and return them in the form of a list".
[0066] For example, for the following sample: Admission record name: Patient name\nPlace of birth: X Province X City\nGender: Male\nOccupation: Resident\nAge: 74 years old\nDate of admission: 2014-10-26 09:26\nNationality: Han\nRecord date: 2014-10-26 13:40\nMarriage: Married\nMedical history narrator: Patient's main complaint: Reducible mass in the left inguinal area for 1 month\nCurrent medical history: The patient reported that a reducible mass appeared in the left inguinal area about 1 month ago without obvious cause, about the size of an egg, without obvious pain and tenderness, no nausea and vomiting, no obvious changes in stool, no urination difficulties, no fever. The mass appeared after standing for a long time and moving, and it can be reduced by itself after lying down and resting, and can also be reduced by hand. The patient did not take any treatment, and the tumor did not grow significantly in the past 3 days. There was no history of incarceration, and no examination or treatment was performed outside. The patient came to the hospital for surgical treatment. After outpatient examination, he was admitted to our department with "left inguinal hernia". \nSince the onset of the disease, the patient has been conscious, mentally healthy, eating and sleeping well, urinating normally, and his weight has not changed. \nPast history: no history of infectious diseases such as hepatitis and tuberculosis and their close contact. No history of hypertension or diabetes. There was a history of chronic bronchitis for 5 years, a history of coronary heart disease for more than 1 month, and no history of surgical trauma. There was no history of drug or food allergies and no history of blood transfusion. Vaccination history is unknown. \nPersonal history: Place of birth, no history of long-term residence in other places, no history of long-term residence in epidemic areas, regular life, no smoking and drinking habits, no history of contact with toxic substances, dust and radioactive substances, no history of traveling, and no history of major mental trauma. \nMarriage and childbearing history: Married at the right age, with 1 son and 3 daughters, and the spouse and children are all healthy. \nFamily history: Denies any family history of genetic diseases and infectious diseases. \nThe above contents are recorded as true. The patient or the client signs: \n\nPhysical examination T36.4℃P: 80 times / minR: 20 times / minBP: 100 / 60 \nElderly male, conscious, mentally sound, normal development, moderate nutrition, self-positioning, and cooperative in physical examination. There is no yellowing of the skin and mucous membranes, no bleeding spots or rashes, and no palpable swelling of superficial lymph nodes throughout the body. There is no deformity of the head, no edema of the eyelids, no pale conjunctiva, no yellowing of the sclera, and the pupils on both sides are equal in size and round, and there is light reflex. There is no deformity of the ears and nose, no cyanosis of the lips, and no enlargement or suppuration of the tonsils. The neck is symmetrical on both sides, the neck is soft, the trachea is in the middle, the thyroid gland is not enlarged, and there is no distended neck veins. The chest is symmetrical on both sides, the respiratory movement is symmetrical on both sides, the breath sounds of both lungs are clear, and no dry or wet rales are heard. There is no abnormal pulsation or bulge in the precordial area, the cardiac dullness boundary is not enlarged, the heart rate is 80 beats / min, the rhythm is regular, and no pathological murmurs are heard in the auscultation area of each valve. The abdomen is examined by a specialist. There is no deformity of the spine and limbs, and the movement is normal. The anus, rectum and external genitalia were not examined. The abdominal wall reflex and tendon reflex are normal, and the Babinski sign and meningeal irritation sign are negative. \nSpecialist examination: The abdomen is flat, and a soft mass of about 3×5 cm can be palpated in the left inguinal area. It has not fallen into the scrotum, and there is no obvious tenderness. It can be retracted when lying flat.Press on the internal ring orifice and cough forcefully. The mass no longer protrudes. A soft mass about 2×3 cm in size can be palpated in the right inguinal region. It does not descend into the scrotum, has no obvious tenderness, and can be reduced when lying flat and pressing on it. Press on the internal ring orifice and cough forcefully. The mass no longer protrudes. The bilateral testes can be normally palpated.\nAuxiliary examinations: None for the time being\nPreliminary diagnosis: 1. Bilateral indirect inguinal hernia\n 2. Chronic bronchitis\n 3. Coronary heart disease\n\nSignature: Physician's name. Calling the large model will return: ['Name', 'Place of birth', 'Gender', 'Occupation', 'Age', 'Date of admission', 'Ethnic group', 'Date of record', 'Marital status', 'Person providing medical history', 'Chief complaint', 'Current medical history', 'Past medical history', 'Personal history', 'Reproductive history', 'Family history', 'Physical examination', 'Auxiliary examinations', 'Preliminary diagnosis', 'Signature'] 4. Combine with the construction of the basic specification knowledge base of electronic medical records to build a splitting dictionary.
[0067] Referring to Table 1, call the basic specification knowledge base of electronic medical records, perform regular matching on the above tags in the DATA_ELEMENT_REGEXP field. If there is a match, extract this DATA_ELEMENT_REGEXP as the value, and the corresponding DATA_ELEMENT as the key value to build a splitting dictionary.
[0068] For example, for the above tag list, the constructed splitting dictionary is: {'Personal history': ['Per\s*son\s*al\s* his\s*tory'], 'Chief complaint': ['Ch\s*ief\s* com\s*plaint'], 'Physical examination': ['Ph\s*y\s*si\s*cal\s* ex\s*am\s*in\s*ation'], 'Date of admission': ['Da\s*te\s* of\s* ad\s*mis\s*sion'], 'Place of birth': ['Pl\s*ace\s* of\s* bir\s*th'], 'Preliminary diagnosis': ['Pr\s*el\s*im\s*in\s*ary\s* di\s*ag\s*no\s*sis'], 'Name': ['Na\s*me'], 'Marital status': ['Mar\s*ital\s* sta\s*tus'], 'Reproductive history': ['Re\s*pro\s*duc\s*tive\s* his\s*tory'], 'Family history': ['Fa\s*mi\s*ly\s* his\s*tory'], 'Age': ['Age'], 'Gender': ['Gen\s*der'], 'Past medical history': ['Pas\s*t me\s*di\s*cal\s* his\s*tory'], 'Ethnic group': ['Eth\s*nic\s* gro\s*up'], 'Current medical history': ['Cur\s*rent\s* me\s*di\s*cal\s* his\s*tory'], 'Date of record': ['(Da\s*te\s* of\s*)* re\s*cor\s*d'], 'Person providing medical history': ['Per\s*son\s* pr\s*o\s*vi\s*ding\s* me\s*di\s*cal\s* his\s*tory'], 'Occupation': ['Occ\s*up\s*ation'], 'Auxiliary examinations': ['Aux\s*il\s*iar\s*y\s* ex\s*am\s*in\s*ations (results)*']} 5. Use the above splitting dictionary to split the sample medical records and perform parsing tests.
[0069] First, traverse all the splitting regular expressions of the tags in the medical sample records, that is, the list of values of the splitting dictionary. If a match is found, replace the matching string with the key value corresponding to the regular expression, that is, map the original tag name in the text to the standard name, and insert a split breakpoint identifier at the replacement position in the text.
[0070] Then, traverse the text processed in the above steps to identify the split identifiers throughout the text, and perform splitting at the identifiers to split the text into a list of sentences or paragraphs. The format of each sentence or paragraph is "key + original content".
[0071] Finally, process each sentence or paragraph in the split list of sentences or paragraphs. Split at the key, where the key is the tag name, and the original text after the key is the tag content, and store it in a parsed table in JSON format.
[0072] At this time, determine the test result according to the parsed table. If the test result is not satisfactory, return to the sampling step and resample and iterate.
[0073] 6. Automatically generate a parsing script that can batch parse medical records.
[0074] Based on step S5 of the electronic medical record parsing method provided in the foregoing embodiment, the following script can be written in Python: from collections import defaultdict import re result = defaultdict(str) key_base_lst = [(item, k) for k, v in key_base_dict.items() for item in v] keys, keys_2 = [item[0] for item in key_base_lst], [item[1] for item in key_base_lst] for key, key2 in key_base_lst: if re.search(key + '[::\s]', content): content = re.sub(key + '[::\s]+','split' + key2 + '\t', content, 1) if key2 in ('Physical examination', 'Auxiliary examination'): if re.search('%s[::]?' % key, content) is not None: content = re.sub('%s[::]?' % key,'split' + key2 + '\t', content, 1) deal_str = content.split('split') for line in deal_str: for key in keys_2: if key + '\t' in line: result[key] = (result[key] + '+' + line.replace(key + '\t', '').strip() if key in result.keys() else line.replace(key + '\t', '').strip()) Break Among them, key_base_lst is the splitting dictionary, that is, the execution result of the previous step.
[0075] In this step, on the basis of step S5 of the electronic medical record parsing method provided in the foregoing embodiment, a read database and a database storage logic are added to form a parsing script base. The parameters passed to the read database function are the connection method of the electronic medical record source table library and the sql statement for retrieving electronic medical record data. The parameters passed to the database storage function are the connection method of the parsing library and the sql statement for storing the parsed data into the database. Among them, Python provides connection packages for multiple databases such as cx_Oracle, pymysql, and py_opengauss. The corresponding function packages are called to connect to the database, and then the sql statement is executed for read database or database storage operations.
[0076] The constructed splitting dictionary is passed as a parameter to the main part of the parsing function of the script base to combine an electronic medical record parsing script that can be used for the entire medical institution.
[0077] The generated parsing script (.py) can be directly called on the server and run in the background to automatically perform the electronic medical record parsing work of a medical institution. Parsing by the python script can maximize the parsing efficiency, and the parsing speed of a single medical record can be as fast as milliseconds at most.
[0078] Exemplarily, the parsing effect of the sample shown above is as follows: {'Name': 'Patient's name', 'Place of birth': 'X province X city', 'Gender': 'Male', 'Occupation': 'Resident', 'Age': '74 years old', 'Date of admission': '2014-10-26 09:26', 'Ethnicity': 'Han nationality', 'Date of medical record': '2014-10-26 13:40', 'Marriage': 'Married', 'Narrator of medical history': 'Patient', 'Chief complaint': 'Reducible mass in the left inguinal area for 1 month', 'Current medical history': 'The patient reported that a reducible mass about the size of an egg appeared in the left inguinal area without obvious cause about 1 month ago, with no obvious pain or tenderness, no nausea or vomiting, no obvious changes in stool, no dysuria, and no fever. The mass appeared after standing for a long time and moving. It can be reduced by itself after lying down and resting, and it can also be reduced by hand. The patient did not take any treatment. The mass did not grow significantly in the past 3 days. There was no history of incarceration. No examination or treatment was performed outside. The patient came to the hospital for surgical treatment. After outpatient examination, he was admitted to our department with "left inguinal hernia". Since the onset of the disease, the patient has been conscious, mentally healthy, eating and sleeping well, urinating normally, and weight has not changed. ', 'Past history': 'No history of infectious diseases such as hepatitis and tuberculosis and their close contact. No history of hypertension or diabetes. There was a history of chronic bronchitis for 5 years and coronary heart disease for more than 1 month. There was no history of surgical trauma. There was no history of drug or food allergies or blood transfusion. Vaccination history is unknown. ', 'Personal history': 'Birthplace, no history of long-term residence in other places, no history of long-term residence in epidemic areas, regular life, no smoking and drinking habits, no history of contact with toxic substances, dust and radioactive substances, no history of prostitution, and no history of major mental trauma. ', 'Marital and childbearing history': 'Married at the right age, with 1 son and 3 daughters, and the spouse and children are all healthy. ', 'Family history': 'Deny any family history of genetic diseases and infectious diseases. The above contents are recorded as true and signed by the patient or the client', 'Physical examination': 'T36.4℃\u3000P: 80 times / min\u3000R: 20 times / min\u3000BP: 100 / 60 \u3000Elderly male, clear consciousness, good spirit, normal development, moderate nutrition, self-positioning, and cooperation in physical examination. There is no yellowing of the skin and mucous membranes, no bleeding spots or rashes, and no palpable enlargement of superficial lymph nodes throughout the body. There is no deformity of the head, no edema of the eyelids, no pallor of the conjunctiva, no yellowing of the sclera, and the pupils on both sides are equal in size and round, with light reflex. There is no deformity of the ears and nose, no cyanosis of the lips, and no enlargement or suppuration of the tonsils. The neck is symmetrical on both sides, the neck is soft, the trachea is centered, the thyroid gland is not enlarged, and there is no distended neck vein. The chest is symmetrical on both sides, the respiratory movement is symmetrical on both sides, the breath sounds of both lungs are clear, and no dry or wet rales are heard. There is no abnormal pulsation or bulge in the precordial area, the heart dullness boundary is not enlarged, the heart rate is 80 beats / min, the rhythm is regular, and no pathological murmurs are heard in the auscultation area of each valve. The abdomen is examined by a specialist. There is no deformity of the spine and limbs, and the movements are normal. The anus, rectum and external genitalia were not examined.The abdominal wall reflex and tendon reflex are normal, and the Babinski sign and meningeal irritation sign are negative. Special examination: The abdomen is flat. A soft mass about 3×5 cm in size can be felt in the left inguinal region, which has not descended into the scrotum, has no obvious tenderness, and can be reduced when lying flat and pressed. When the internal ring orifice is compressed and coughing forcefully, the mass no longer protrudes. A soft mass about 2×3 cm in size can be felt in the right inguinal region, which has not descended into the scrotum, has no obvious tenderness, and can be reduced when lying flat and pressed. When the internal ring orifice is compressed and coughing forcefully, the mass no longer protrudes. The bilateral testicles can be felt normally.', 'Auxiliary examination': 'None available', 'Initial diagnosis': '1. Bilateral indirect inguinal hernia 2. Chronic bronchitis 3. Coronary heart disease', 'Physician's signature': 'Physician's name'}. 7. Parse and store the electronic medical records into the database.
[0079] Connect the parsing script to a server that can connect to the database, execute the parsing script, and parse and store the electronic medical record table under the medical institution.
[0080] For example, the server calls the parsing script to parse the electronic medical record table under the medical institution, and generates a parsing table to store and save the parsed structured data. The present invention uses a vertical table to store the parsed data, and the complete parsing result of this medical record can be found by screening the medical record ID.
[0081] Taking the table creation in Oracle as an example, the structure of the parsing table of this method is shown in Table 3 below, and the generated parsing table is shown in Table 4 below.
[0082] Table 3: Reference table for the structure of the parsing table
[0083] Table 4: Parsing table
[0084] This embodiment discloses an electronic medical record parsing method. By comprehensively applying means such as the construction of an electronic medical record basic specification knowledge base, large language model technology, and intelligent splitting and standardized storage, it realizes the efficient and accurate parsing of electronic medical record texts. It can not only intelligently identify the medical record type, use the large language model to deeply mine and extract standard tag names, but also accurately split the medical record text into sentences or paragraphs by constructing a splitting dictionary, and at the same time extract and standardize the storage of tag names and tag contents. This process not only significantly improves the accuracy and efficiency of electronic medical record parsing, but also effectively promotes the standardized and normalized management of medical data, providing strong support for the improvement of medical quality, the development of medical informatization, and the in-depth mining and utilization of medical data.
[0085] Such as Figure 3As shown below, the following is an embodiment of an electronic medical record parsing system provided by the present disclosure. This system and the electronic medical record parsing methods of the above embodiments belong to the same inventive concept. For the details not described in detail in the embodiment of the electronic medical record parsing system, reference can be made to the embodiments of the above electronic medical record parsing methods.
[0086] An electronic medical record parsing system includes: a knowledge base construction module, a medical record classification module, a tag list generation module, a splitting dictionary construction module, and a splitting and parsing module.
[0087] The knowledge base construction module is used to construct a knowledge base of basic electronic medical record specifications based on the text format and structure of existing electronic medical records.
[0088] The medical record classification module is used to obtain the electronic medical record to be parsed and identify the medical record type of the parsed electronic medical record based on the title of the electronic medical record.
[0089] The tag list generation module is used to extract tags from the electronic medical record text to be parsed using a large language model according to the medical record type and generate a tag list.
[0090] The splitting dictionary construction module is used to traverse the knowledge base of basic electronic medical record specifications according to the tag list, extract the standard tag names and corresponding splitting regular expressions, and generate a splitting dictionary.
[0091] The splitting and parsing module is used to split the electronic medical record text to be parsed using the splitting dictionary, split the text into a list of sentences or paragraphs, extract the tag names and tag contents of the tags, generate a parsing table, and store it in the database.
[0092] The electronic medical record parsing system provided by this embodiment significantly improves the accuracy, efficiency, and standardization level of electronic medical record processing by constructing a knowledge base of basic electronic medical record specifications, applying large language model technology for intelligent parsing and tag extraction, and combining precise text splitting and standardized storage means, laying a solid foundation for the efficient management and in-depth application of medical data.
[0093] Figure 4 The hardware structure diagram of an electronic device for implementing each embodiment of the present invention.
[0094] The electronic medical record parsing method provided by the embodiments of this application can be applied to an electronic device. Those skilled in the art can understand that the structure of the electronic device involved in the embodiments of the present invention does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the electronic device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of this application described herein and / or claimed.
[0095] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, a key, a camera, a display screen, and a SIM card interface, etc.
[0096] The processor may include one or more processing units. For example, the processor may include a central processing unit (CPU), etc., an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0097] Among them, the processor may be the nerve center and command center of the electronic device. The controller can generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching instructions and executing instructions.
[0098] A memory can also be set in the processor for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can hold the instructions or data that the processor has just used or recycled. If the processor needs to use the instruction or data again, it can directly call it from this memory. This avoids repeated accesses, reduces the waiting time of the processor, and thus improves the system efficiency.
[0099] The external memory interface can be used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor through the external memory interface to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0100] The internal memory can be used to store computer-executable program code, and the computer-executable program code includes instructions. The processor executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory. The internal memory can include a program storage area and a data storage area. The internal memory can include a high-speed random access memory and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0101] The wireless communication function of the electronic device can be implemented through an antenna, a wireless communication module, a modem processor, a baseband processor, etc.
[0102] The wireless communication module can provide wireless communication solutions applied to the electronic device, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSSs), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.
[0103] The electronic device can implement audio functions, etc., through an audio module, a speaker, a receiver, a microphone, a headphone interface, an application processor, etc.
[0104] The electronic device can implement a shooting function through an ISP, a camera, a video codec, a GPU, a display screen, an application processor, etc.
[0105] An electronic device can implement a display function through a GPU, a display screen, an application processor, etc.
[0106] The GPU is a microprocessor for image processing, connecting the display screen and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor may include one or more GPUs, which execute program instructions to generate or change display information.
[0107] The display screen is used to display images, videos, etc. The display screen includes a display panel.
[0108] The above electronic device realizes the efficient and intelligent parsing of electronic medical record texts by comprehensively applying the construction of the basic specification knowledge base of electronic medical records and advanced large language model technology. At the same time, by using the medical record types, standard tag names, and splitting regular expressions in the knowledge base, combined with the deep semantic understanding and tag extraction capabilities of the large language model, it can accurately identify and extract key information in the electronic medical record, accurately split the medical record text into sentences or paragraphs, and mark the corresponding tag names and tag contents. The above electronic device achieves the beneficial effect of greatly improving the accuracy and efficiency of electronic medical record parsing.
[0109] In the storage medium provided by this application, there is a program product that can implement the electronic medical record parsing method.
[0110] The electronic medical record parsing method includes: constructing a basic specification knowledge base of electronic medical records based on the text format and structure of existing electronic medical records; obtaining the electronic medical record to be parsed, and identifying the medical record type of the parsed electronic medical record based on the title of the electronic medical record; according to the medical record type, using the large language model to perform tag extraction on the electronic medical record text to be parsed to generate a tag list; according to the tag list, traversing the basic specification knowledge base of electronic medical records to extract standard tag names and corresponding splitting regular expressions to generate a splitting dictionary; using the splitting dictionary to split the electronic medical record text to be parsed, splitting the text into a list of sentences or paragraphs, extracting the tag names and the tag contents of the tags, generating a parsing table, and storing it in the database.
[0111] In some possible implementation manners, the electronic medical record parsing method of the present disclosure can be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section above of this specification.
[0112] The storage medium of the present disclosure may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0113] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An electronic medical record parsing method, characterized in that, include: Based on the existing text format and structure of electronic medical records, a knowledge base of basic specifications for electronic medical records is constructed; Obtaining an electronic medical record to be parsed, and identifying the medical record type of the parsed electronic medical record based on the title of the electronic medical record; According to the medical record type, a large language model is used to extract labels from the electronic medical record text to be parsed and generate a label list; According to the tag list, the basic standard knowledge base of electronic medical records is traversed to extract the standard tag names and corresponding splitting regular expressions, and generate a splitting dictionary; The electronic medical record text to be parsed is split using a split dictionary, the text is split into a list of sentences or paragraphs, the tag name and the tag content of the tag are extracted, a parsing table is generated, and it is stored in the database.
2. The electronic medical record parsing method according to claim 1, wherein The electronic medical record basic specification knowledge base includes: medical record type, standard tag name and corresponding split regular expression; The types of medical records include: admission records, discharge records, and first medical records; The standard label names include: chief complaint, personal history, current medical history, past medical history, family history, marital history, treatment process information, differential diagnosis information, and physician signature.
3. The electronic medical record parsing method according to claim 2, characterized in that The step of obtaining the electronic medical record to be parsed and identifying the medical record type of the parsed electronic medical record based on the title of the electronic medical record includes: Obtain the electronic medical record to be parsed, and divide it into medical record text with title and medical record text without title; For medical record texts with titles, determine the medical record type based on the medical record title; For untitled medical record texts, the medical record texts are input into a preset electronic medical record classifier, and the medical record texts are classified into preset medical record types by extracting the semantics and structure of the medical record texts.
4. The electronic medical record parsing method according to claim 3, wherein According to the medical record type, the large language model is used to extract labels from the electronic medical record text to be parsed, and a label list is generated, including: According to the medical record type, a large language model is used to extract labels from the electronic medical record text to be parsed. The standard label names in the electronic medical record text are extracted using the prompt words of the standard label names, and duplicate standard label names are deleted through iterative deduplication. Generate a tag list based on the proposed standard tag names.
5. The electronic medical record parsing method according to claim 4, wherein, According to the label list, the basic standard knowledge base of electronic medical records is traversed to extract the standard label name and the corresponding split regular expression, and generate a split dictionary, including: According to each standard tag name in the tag list, find the same standard tag name by traversing the basic specification knowledge base of electronic medical records, and extract the corresponding splitting regular expression; Based on the traversal results, a split dictionary is constructed with the standard tag name as the key and the corresponding split regular expression as the value.
6. The electronic medical record parsing method according to claim 5, characterized in that, The electronic medical record text to be parsed is split using a split dictionary, the text is split into a sentence or paragraph list, the tag name and the tag content of the tag are extracted, a parsing table is generated, and the parsing table is stored in a database, including: Traverse the electronic medical record text to be parsed according to each value in the split dictionary to find the matching string; Replace the matched string with the key value corresponding to the value value, so as to replace the original tag name in the electronic medical record text to be parsed with the standard tag name, and enter the split breakpoint mark at the replacement point of the text; According to the breakpoint mark of split in the electronic medical record text to be parsed, the corresponding text is split into a list of sentences or paragraphs; Convert each sentence or paragraph into a preset standard format; the preset standard format includes a standard tag name and tag content; Generate a parsing table in json format based on the converted sentence or paragraph list; The parsing table is stored in the database through the SQL statement.
7. The electronic medical record parsing method according to claim 6, wherein The large language model adopts the Qwen-72B lightweight model, and the database adopts an Oracle database, a MySQL database or an openGauss database.
8. An electronic medical record parsing system, characterized in that, The system adopts the electronic medical record parsing method as claimed in any one of claims 1 to 7; The system comprises: A knowledge base construction module is used to build a basic specification knowledge base of electronic medical records based on the text format and structure of existing electronic medical records; A medical record classification module is used to obtain the electronic medical record to be parsed and identify the medical record type of the parsed electronic medical record based on the title of the electronic medical record; The label list generation module is used to extract labels from the electronic medical record text to be parsed based on the medical record type and generate a label list using a large language model; The split dictionary building module is used to traverse the basic standard knowledge base of electronic medical records according to the label list, extract the standard label name and the corresponding split regular expression, and generate a split dictionary; The splitting and parsing module is used to split the electronic medical record text to be parsed using the splitting dictionary, split the text into sentences or paragraph lists, extract the tag name and the tag content of the tag, generate a parsing table, and store it in the database.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, the steps of the electronic medical record parsing method according to any one of claims 1 to 7 are implemented.
10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the electronic medical record parsing method according to any one of claims 1 to 7 are implemented.