Method, System, Device and Storage Medium for Extracting Electronic Medical Record Information
Through the electronic medical record information extraction model, the key information in the electronic medical record is automatically extracted and structured, which solves the problem of low diagnosis efficiency of doctors and realizes rapid acquisition and display of structured information.
Patent Information
- Application Number
- CN202510258525.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-05
AI Technical Summary
In the prior art, electronic medical records have a large amount of information, and doctors need to read them for a long time during diagnosis, resulting in inefficient diagnosis.
Through the electronic medical record information extraction model, the Word2Vec model and the normalized model extract text and picture features, combined with BERT, LSTM, and CNN models for feature recognition, the multi-channel self-attention model is used to obtain entity information, and structured information is generated through the structured model, and key information is automatically extracted and displayed.
Significantly shorten the diagnosis time of doctors, improve diagnosis efficiency, reduce the time for manual review of electronic medical records, and improve the efficiency of information extraction and display.
Smart Images

Figure CN119740558B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the technical field of data processing, and in particular, to a method, system, device, and storage medium for extracting electronic medical record information. Background Art
[0002] With the continuous development of electronic technology, it has become very common in various fields to perform analysis and operations on the data in the field through model training.
[0003] In the medical field, in order to ensure the accuracy of patient diagnosis, electronic medical records have long been popular. Based on their characteristics of being easy to store and not easily lost, electronic medical records can retain all the diagnostic process information of each patient in detail. For any patient, in order to ensure the accuracy of diagnosis, each diagnosis may require the support of multiple examination data. Therefore, when the patient finally faces the doctor for diagnosis, there may already be a large amount of data in the patient's electronic medical record, and there are many pages that need the doctor to continuously click on the display terminal (such as a computer) to view. The doctor needs to view for a certain period of time to make a diagnosis decision. Summary of the Invention
[0004] In view of the above-mentioned defects or deficiencies in the related art, it is desirable to provide a method, system, device, and medium for extracting electronic medical record information, which can solve the problem that electronic medical records often have a large amount of content, and it takes a long time to view the electronic medical record manually to make a diagnosis decision, thereby reducing the diagnosis efficiency. Through the electronic medical record information extraction method of the present application, the doctor's diagnosis time can be greatly shortened, and the doctor's diagnosis efficiency can be improved.
[0005] In a first aspect, a method for extracting electronic medical record information is provided, and the method includes:
[0006] Extracting first text features in the first electronic medical record information through a first sub-model in the electronic medical record information extraction model, and extracting first picture features in the first electronic medical record information through the first sub-model, where the first sub-model is a Word2Vec model and a normalization model;
[0007] Processing the text features in the first electronic medical record information and the image features in the first electronic medical record information through a second sub-model in the electronic medical record information extraction model, identifying second text features in the first text features and second picture features in the first picture features, where the second sub-model is a multi-task model, the multi-task model includes a BERT model, an LSTM model, and a CNN model, the second text features are part or all of the text features in the first text features, and the second picture features are part or all of the picture features in the first picture features;
[0008] The third sub-model in the electronic medical record information extraction model is used to identify the second text feature and the second picture feature, and obtain the entity information corresponding to the first electronic medical record information. The entity information includes: user information, user status, and conclusive information of the user status. The third sub-model includes a multi-channel self-attention model;
[0009] The fourth sub-model in the electronic medical record information extraction model processes the entity information according to a predefined structured format to generate structured information, which is used to display the key information in the entity information according to a preset format. The fourth sub-model is a structured model.
[0010] In this application, the Word2Vec model and the normalization model in the electronic medical record information extraction model are used to extract the first text feature in the first electronic medical record information, and the first picture feature in the first electronic medical record information is extracted through the above-mentioned first sub-model; then, the multi-task model in the electronic medical record information extraction model is used to process the text feature in the first electronic medical record information and the image feature in the first electronic medical record information to identify the second text feature in the first text feature and the second picture feature in the first picture feature (the multi-task model includes a BERT model, an LSTM model, and a CNN model). The second text feature is part or all of the text features in the first text feature, and the second picture feature is part or all of the picture features in the first picture feature; then, the multi-channel self-attention model in the electronic medical record information extraction model is used to identify the second text feature and the second picture feature, and obtain the entity information corresponding to the first electronic medical record information. The entity information includes: user information, user status, and conclusive information of the user status; finally, the structured model in the electronic medical record information extraction model processes the entity information according to a predefined structured format to generate structured information, which is used to display the key information in the entity information according to a preset format. In this way, the electronic medical record information extraction model can automatically extract all the valid information in the electronic medical record and make structural adjustments, and finally display the structural information for the user, without the user having to view all the electronic medical records by themselves. In this way, the efficiency of doctors viewing electronic medical records can be greatly improved, and the time for viewing electronic medical records can be shortened.
[0011] In a second aspect, an electronic medical record information extraction system is provided, and the system includes:
[0012] An extraction unit is configured to extract the first text feature from the first electronic medical record information through a first sub-model in an electronic medical record information extraction model, and extract the first image feature from the first electronic medical record information through the first sub-model. The first sub-model is a Word2Vec model and a normalization model;
[0013] An execution unit is configured to process the text feature in the first electronic medical record information and the image feature in the first electronic medical record information through a second sub-model in the electronic medical record information extraction model, and identify a second text feature in the first text feature and a second image feature in the first image feature. The second sub-model is a multi-task model, and the multi-task model includes a BERT model, an LSTM model, and a CNN model. The second text feature is part or all of the text features in the first text feature, and the second image feature is part or all of the image features in the first image feature;
[0014] An acquisition unit is configured to identify the second text feature and the second image feature through a third sub-model in the electronic medical record information extraction model, and acquire entity information corresponding to the first electronic medical record information. The entity information includes: user information, user status, and conclusive information of the user status. The third sub-model includes a multi-channel self-attention model;
[0015] A generation unit is configured to process the entity information through a fourth sub-model in the electronic medical record information extraction model according to a predefined structured format, and generate structured information. The structured information is used to display key information in the entity information according to a preset format. The fourth sub-model is a structured model.
[0016] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the first aspect above is implemented.
[0017] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. The computer program, when executed by a processor, implements the method described in the first aspect above.
[0018] In a fifth aspect, a computer program product is provided. The computer program product contains instructions that, when run by a processor, implement the method described in the first aspect above.
[0019] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Other features, objects, and advantages of the present application will become more apparent from the following detailed description of non-limiting embodiments read in conjunction with the accompanying drawings:
[0021] Figure 1 It is a schematic flowchart of a method for extracting electronic medical record information provided by an embodiment of the present application;
[0022] Figure 2 It shows a schematic diagram of the application process of an electronic medical record information extraction model provided by an embodiment of the present application;
[0023] Figure 3 It is a schematic diagram of the training process of an electronic medical record information extraction model provided by an embodiment of the present application;
[0024] Figure 4 It is a schematic flowchart of the process for obtaining structured information provided by an embodiment of the present application;
[0025] Figure 5 It is a schematic diagram of the structure of a resource access device provided by an embodiment of the present application;
[0026] Figure 6 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed Embodiments
[0027] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the relevant invention and not for limiting the invention. Additionally, it should be noted that for the sake of description, only parts related to the invention are shown in the drawings.
[0028] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.
[0029] Figure 1 It is a schematic flowchart of a method for extracting electronic medical record information provided by an embodiment of the present application. As Figure 1 shown, the method includes the following steps 301 to step 304:
[0030] Step 301: Extract the first text feature in the first electronic medical record information through the first sub-model in the electronic medical record information extraction model, and extract the first picture feature in the first electronic medical record information through the first sub-model.
[0031] In the embodiment of the present application, the first sub-model is a Word2Vec model and a normalization model.
[0032] It can be understood that in the embodiments of the present application, the first electronic medical record information may include text information or picture information. Therefore, in order to ensure the integrity of the extraction of the first electronic medical record information, it is necessary to extract both the text features and the picture features of the first electronic medical record information. Information
[0033] In the embodiments of the present application, the Word2Vec model can be used to identify the text information in the first electronic medical record and further extract the first text features in the first electronic medical record, and the normalization model can be used to identify the picture information in the first electronic medical record and further extract the first picture features in the first electronic medical record.
[0034] In the embodiments of the present application, the above-mentioned first text features are the feature information after vectorizing the text information in the first electronic medical record; correspondingly, the above-mentioned first picture features are the feature information after vectorizing the picture information in the first electronic medical record information.
[0035] Step 302: Process the text features in the first electronic medical record information and the image features in the first electronic medical record information through the second sub-model in the above-mentioned electronic medical record information extraction model, and identify the second text features in the above-mentioned first text features and the second picture features in the above-mentioned first picture features.
[0036] In the embodiments of the present application, the above-mentioned second sub-model is a multi-task model.
[0037] Exemplarily, the multi-task model includes a BERT model, an LSTM model, and a CNN model.
[0038] In the embodiments of the present application, the second text features are part or all of the text features in the first text features, and the second picture features are part or all of the picture features in the first picture features;
[0039] It can be understood that the above-mentioned first text features are all the text features in the first electronic medical record information. Correspondingly, the first picture features are also all the image features in the first electronic medical record information. For the present application, its purpose is to extract the effective features or key features in the first electronic medical record information. The effective features or key features may be part or all of the features in the first electronic medical record information. Based on this, the second text features are part or all of the first text features. Correspondingly, the second picture features are part or all of the first picture features.
[0040] In the embodiments of the present application, the above-mentioned BERT model, LSTM model, and CNN model have been trained and can be used to identify and extract key features.
[0041] In the embodiments of the present application, the above BERT model and LSTM model can be used to identify and extract the second text feature from the first text feature; the above CNN model can be used to identify and extract the second image feature from the first image feature.
[0042] Step 303: Identify the above second text feature and the above second image feature through the third sub-model in the above electronic medical record information extraction model to obtain the entity information corresponding to the above first electronic medical record information.
[0043] In the embodiments of the present application, the above entity information includes: user information, user status, and conclusive information of the user status.
[0044] In the embodiments of the present application, the above third sub-model includes a multi-channel self-attention model.
[0045] In the embodiments of the present application, after obtaining the key information in the above first electronic medical record information, that is, the second text feature and the second image feature, it is necessary to further clarify the association relationship between different key information, so as to enhance the readability of the key information presented subsequently. Based on this, it is necessary to integrate the above second text feature and the second image feature to finally obtain the entity information corresponding to the above first electronic medical record.
[0046] It can be understood that the user information in the above entity information can be the personal information of the user itself, the above user status can be the inspection process information of the user, such as the user's index information, etc., and the conclusive information of the above user status is the conclusive information corresponding to the user status.
[0047] Furthermore, the above entity information can include the past information in the first electronic medical record, or the past information and the current information, where the current information includes user information and user status information.
[0048] Step 304: Process the above entity information through the fourth sub-model in the above electronic medical record information extraction model according to a predefined structured format to generate structured information.
[0049] In the embodiments of the present application, the above structured information is used to display the key information in the entity information according to a preset format.
[0050] In the embodiments of the present application, the above fourth sub-model is a structured model.
[0051] It can be understood that after generating the above entity information, it is necessary to process the entity information through a structured model according to a predefined structured format, so that the information finally presented to the user, such as a doctor, is structured information with strong readability and unified structure, thereby improving the efficiency of the user reading the first electronic medical record and obtaining effective information, and shortening the subsequent judgment time.
[0052] In the method provided by the embodiment of the present application, the Word2Vec model and the normalization model in the electronic medical record information extraction model are used to extract the first text features in the first electronic medical record information, and the first sub-model is used to extract the first image features in the first electronic medical record information; then, the multitask model in the electronic medical record information extraction model is used to process the text features in the first electronic medical record information and the image features in the first electronic medical record information, and the second text features in the first text features and the second image features in the first image features are identified (the multitask model includes a BERT model, an LSTM model, and a CNN model), the second text features are part or all of the text features in the first text features, and the second image features are part or all of the image features in the first image features; then, the multi-channel self-attention model in the electronic medical record information extraction model is used to identify the second text features and the second image features, and the entity information corresponding to the first electronic medical record information is obtained, and the entity information includes: user information, user status, and conclusive information of the user status; finally, the structured model in the electronic medical record information extraction model processes the entity information according to a predefined structured format to generate structured information, and the structured information is used to display the key information in the entity information according to a preset format. In this way, the electronic medical record information extraction model can automatically extract all the effective information in the electronic medical record and make structural adjustments, and finally display the structural information for the user, without the user having to view all the electronic medical records by themselves. In this way, the efficiency of doctors viewing electronic medical records can be greatly improved, and the time for viewing electronic medical records can be shortened.
[0053] In another embodiment of the present application, a specific implementation manner for preprocessing the initial electronic medical record information is further provided. Exemplarily, before the "extracting the first text features in the first electronic medical record information by the first sub-model in the electronic medical record information extraction model" mentioned above, the implementation manner further includes: preprocessing the initial electronic medical record information according to a preset preprocessing manner to obtain the first electronic medical record information.
[0054] Exemplarily, the preset preprocessing manner includes: information word segmentation, part-of-speech tagging, grayscale processing, and binarization processing.
[0055] It can be understood that after obtaining the initial electronic medical record information, it is necessary to preprocess the initial electronic medical record information first, so that the first text features and the first image features can be clearly and quickly extracted in the subsequent extraction process.
[0056] For example, for the text information of the initial electronic medical record information, the text information should be pre-processed by information segmentation and part-of-speech tagging. After the initial electronic medical record information is segmented and part-of-speech tagged, a vocabulary list can be generated.
[0057] Exemplarily, for the image information of the initial electronic medical record information, the image information should be pre-processed by grayscale processing and binarization processing.
[0058] Exemplarily, in the above-mentioned process of preprocessing the initial electronic medical record information according to a preset method, it is necessary to collect and purify the initial electronic medical record information, and eliminate irrelevant information and noise, so that in the subsequent process of extracting the first text feature, the first text feature required in the text information and the first image feature required in the image information can be more accurately extracted.
[0059] In another embodiment of the present application, a specific implementation method for extracting the first text feature is also provided. Exemplarily, the "extracting the first text feature in the first electronic medical record information through the first sub-model in the electronic medical record information extraction model" mentioned above includes: identifying the first text information in the first electronic medical record information through the Word2Vec model; obtaining all lexical features in the first text information, and processing the lexical features into lexical vectors, obtaining the association between different lexical vectors; generating the first text feature according to the lexical vector and the association between different lexical vectors.
[0060] Exemplarily, the first text feature is feature information formed by a text vector.
[0061] Exemplarily, the first text feature vector is generated by a trained Word2Vec model. The first text feature vector may be a word vector, which is used to capture the relationship between words by mapping each word to a vector space.
[0062] Exemplarily, the above Word2Vec model includes the calculation formula of the Skip-Gram model.
[0063] In the related art, the calculation formula of the Skip-Gram model is shown in formula (1):
[0064] (1)
[0065] in, is the vector of the target word, is the vector of the context word, is the total number of words in the vocabulary.
[0066] In the embodiments of the present application, the above Word2Vec model is trained with a large amount of corpora in the medical field. Therefore, a large number of word vectors adapted to the medical field are introduced into the Word2Vec model in the embodiments of the present application, making the word vectors closer to the context of the recognized first electronic medical record information and improving the model's ability to understand medical terms.
[0067] In the embodiments of the present application, the calculation formula (2) of the Skip-Gram model in the embodiments of the present application is:
[0068] (2)
[0069] Where is the vector of the target word, is the vector of the context word, is the domain-adapted bias term, is the total number of words in the vocabulary.
[0070] In another embodiment of the present application, a specific implementation manner for extracting the first picture feature is also provided. Exemplarily, the "extracting the first picture feature in the above first electronic medical record information through the second sub-model" mentioned above specifically includes: identifying all picture information in the above first electronic medical record information through the above normalization processing model; obtaining the illumination brightness information of each picture information, and the above illumination brightness information includes the illumination brightness information of multiple local regions in each picture information; performing normalization processing on the illumination brightness information of each above picture information to generate the first picture feature.
[0071] Exemplarily, the above one picture information includes the above multiple local regions
[0072] Exemplarily, the above first picture feature includes multiple first sub-picture features in the above first electronic medical record information, and one first sub-picture feature corresponds to one picture information.
[0073] Exemplarily, the normalization model can be used to perform normalization processing on the image information, thereby reducing the influence of illumination and contrast on the recognition process during the process of recognizing the images in the first electronic medical record.
[0074] Exemplarily, in the embodiments of the present application, during the process of performing normalization processing on the image information, local normalization processing for the image information needs to be considered. The reason is that for the non-uniform illumination problem that may exist in the image information in the first electronic medical record information, using the mean and standard deviation of local regions for normalization can more effectively process the illumination changes in the image information.
[0075] The normalization processing calculation formula of the normalization model in the related art is as formula (3) below
[0076] (3)
[0077] Wherein, is the original image, is the mean value of the image, is the standard deviation of the image.
[0078] The improved normalization calculation formula in the embodiment of the present application is as follows, Formula (4)
[0079] (4)
[0080] Wherein is the original image, is the mean value of the local area of the image, is the standard deviation of the local area of the image.
[0081] In another embodiment of the present application, a specific implementation manner for recognizing the second text feature and the second picture feature is further provided. Exemplarily, the "processing the first text feature and the first image feature in the first electronic medical record information through the second sub-model in the above-mentioned electronic medical record information extraction model to recognize the second text feature in the first text feature and the second picture feature in the first picture feature" involved above specifically includes: extracting the text sequence vector in the first text feature through the BERT model and the LSTM model, and processing the first text feature according to the text sequence vector to generate the second text feature; extracting the text sequence vector in the first picture feature through the CNN model, and processing the first picture feature according to the text sequence vector to generate the second picture feature.
[0082] Exemplarily, in the embodiment of the present application, the extracted first text feature and first picture feature need to be further processed to identify the key features therein.
[0083] Exemplarily, in the process of extracting the second text feature from the first text feature, the BERT model is used to encode the text data and capture the context relationship and semantic information in the text. Through its Transformer architecture, BERT can understand the meaning of words in different contexts, thereby extracting rich text features. The LSTM model is then used to process sequence data, such as time series data, to capture long-term dependence relationships. LSTM is particularly suitable for processing and predicting data based on time series. For example, we use BERT to identify the disease name and drug name in the first text feature, while LSTM can be used to analyze the treatment process and the time series of the disease development in the first text feature. By combining in this way, the model can more comprehensively understand and process the information content of the first electronic medical record information.
[0084] In the related art, the calculation formula of the BERT model is as shown in the following formula (5):
[0085] (5)
[0086] Wherein, is the input word vector sequence, is the core architecture of the BERT model.
[0087] In the embodiments of the present application, the above BERT model can be specifically fine-tuned for its applicable scenarios in the medical field, that is, continuously training the BERT model using electronic medical record data to fine-tune it, so that it can better capture the language features in the medical field. The calculation formula of the BERT model after improvement and adjustment is formula (6):
[0088] (6)
[0089] Wherein is the input word vector sequence, is a parameter specific to the medical field.
[0090] For the image part, the CNN model is used.
[0091] In the related art, the calculation formula of the CNN model is as shown in the following formula (7):
[0092] (7)
[0093] Wherein, is the input image, is the convolution kernel, represents the convolution operation, b is the bias term, and σ is the activation function.
[0094] Furthermore, in the embodiments of the present application, in order to characterize different components of the second text feature and enhance the model's ability to identify key information, an attention mechanism can also be introduced.
[0095] In the related art, the calculation formula of the attention mechanism is as shown in the following formula (8):
[0096] (8)
[0097] Wherein, is the query matrix, is the key matrix, is the value matrix, is the dimension of the key vector.
[0098] In the embodiments of the present application, it is necessary to adjust the above multi-channel self-attention mechanism. By applying the attention mechanism to different channels and outputting a two-dimensional weight matrix, different components of the sentence are characterized, enhancing the model's ability to recognize key information. The calculation formula of the improved attention mechanism in the embodiments of the present application is shown in the following formula (9):
[0099] (9)
[0100] Where is the query matrix, is the key matrix, is the value matrix, is the interaction matrix between channels, is the dimension of the key vector.
[0101] In another embodiment of the present application, a specific implementation manner for generating structured information is also provided. Exemplarily, the "processing the above entity information in a predefined structured format through the fourth sub-model in the above electronic medical record information extraction model to generate structured information" mentioned above specifically includes: through the BiLSTM-CRF model in the structured model, identifying the above entity information and extracting the entity relationships in the entity information, and determining the entity boundaries of the above entity information; according to the above entity information, the entity relationships in the above entity information, and the entity boundaries, determining the key information in the above entity information and generating structured information.
[0102] Exemplarily, the above entity boundary is the boundary of the entity, that is, the start and end positions of the entity in the electronic medical record information.
[0103] In the embodiments of the present application, during the process of generating the above structured information, entity recognition and entity relationship arrangement are required. During the process of entity recognition, the process of named entity recognition and relationship extraction is required, and the BiLSTM-CRF model is used to complete this process.
[0104] In the related art, the BiLSTM-CRF model is used. The BiLSTM-CRF model enhances the feature expression ability of characters at different distances in the Transformer network based on weight-based positional embedding, improving the recognition accuracy of the model for entity boundaries.
[0105] The calculation formulas of the BiLSTM-CRF model are shown in the following formulas (13) and (14):
[0106] (13)
[0107] (14)
[0108] Where is the input sequence, is the label sequence, is the transition matrix, is the emission matrix.
[0109] In the embodiment of the present application, for the application scenario of electronic medical records, the BiLSTM-CRF model will be improved. The calculation formula of the improved BiLSTM-CRF model is shown in Formulas (15) and (16):
[0110] (15)
[0111] (16)
[0112] where is the input sequence, is the label sequence, is the transition matrix, is the emission matrix, is the position embedding based on weights.
[0113] In one example, for the above first electronic medical record information, the process of determining entity information, extracting entity relationships, and determining the entity boundaries of the entity information is as follows:
[0114] (1) Entity recognition: Use the BiLSTM-CRF (Bidirectional Long Short-Term Memory Network - Conditional Random Field) model to identify medical entities in the text, such as disease names, drug names, treatment methods, etc., and confirm the entity boundaries. As shown in the following table, this table is used to display the recognition of key information:
[0115]
[0116] (2) Relationship extraction: After identifying the entities, the model further analyzes the relationships between the entities, such as the association between "patient" and "disease", and the connection between "drug" and "indication". This relationship extraction process is determined by analyzing the context and grammatical structure between the entities.
[0117] (3) Structured information generation: Organize the identified entities and their boundaries, as well as the relationships between the entities, in a predefined structured format (such as JSON or XML).
[0118] The structured information is as follows (in JSON format):
[0119] Among them, the content included in the above structured information is described as follows:
[0120] PatientInfo: Contains basic information of the patient, such as name, gender, and age.
[0121] Chief Complaint: Record the patient's main symptoms and problems.
[0122] Medical History: Provide a summary of the patient's medical history, including past medical history and long-term medications.
[0123] Test Results: Record in detail the patient's test results, including blood routine and blood glucose levels.
[0124] Diagnosis: Clearly define the patient's diagnosis results, including disease type and severity.
[0125] Treatment Plan: List the treatment plan, including drug names, dosages, and treatment cycles.
[0126]
[0127] In this example, entities such as "Zhang San" (patient name), "drug-induced liver injury" (diagnosis), and "Dictamnus dasycarpus Turcz." (drug name) are identified and the entity boundaries are marked. The relationships between entities, such as the impact of "Dictamnus dasycarpus Turcz." on "abnormal liver function", are also extracted and structured. In this way, the unstructured electronic medical record data is converted into structured information, which is convenient for storage, retrieval, and further data analysis.
[0128] It should be noted that in the embodiments of this application, both named entity recognition (NER) and relation extraction use the BiLSTM-CRF model
[0129] It can be understood that in the embodiments of this application, entity information can also be extracted from the first text information by using named entity recognition (NER) technology. Key regions are extracted from images by using object detection or image segmentation technology. The extracted information is organized according to a predefined structured format, such as JSON, XML, etc., and then relation extraction technology is used to identify the associations between entities, so as to obtain the association relationships between different word vectors, and structured processing can also be performed.
[0130] As Figure 2 shown, Figure 2 the application process of the electronic medical record information extraction model is shown. This application process is trained in the manner shown in the figure.
[0131] In another embodiment of the present application, a specific implementation manner for generating an electronic medical record information processing model is further provided. Exemplarily, the implementation manner includes: obtaining multiple groups of original electronic medical records and target data corresponding to the multiple groups of original electronic medical records; inputting the multiple groups of original electronic medical records into an initial electronic medical record information extraction model to generate test data corresponding to the original electronic medical records; comparing the multiple groups of test data with the multiple groups of target data, and determining the difference between the multiple groups of test data and the multiple groups of target data;
[0132] In the case where the above difference is greater than a preset threshold, adjusting the initial electronic medical record information extraction model according to the above difference, so that the difference between the test data generated by the electronic medical record information processing model and the target data is less than the preset threshold.
[0133] Exemplarily, the above target data includes: target structured information, and the structured information corresponding to the multiple groups of original electronic medical records is structured information in a target format corresponding one by one to each sub-medical record in the multiple groups of original electronic medical records, and the multiple groups of original electronic medical records include data information of at least two users.
[0134] Exemplarily, the above test data includes test structured information
[0135] Exemplarily, the above initial electronic medical record information extraction model includes an initial Word2Vec model, a normalization model, a BERT model, an LSTM model, a CNN model, a multi-channel self-attention model, and a structured model.
[0136] It can be understood that in the embodiment of the present application, the initial electronic medical record information extraction model needs to be trained multiple times through the target data and the original data corresponding to multiple groups of original electronic medical records, and the difference between each training result and the target result is compared, and this difference can be represented by the above difference. It is not until the difference is less than the preset threshold that an electronic medical record information extraction model that can be actually put into use can be obtained.
[0137] Furthermore, in the embodiment of the present application, by using the labeled data, that is, the target data to train the model, methods such as cross-validation can be used to evaluate and optimize the model performance. Technologies such as transfer learning are used to improve the generalization ability of the model.
[0138] It can be understood that in the embodiment of the present application, the above electronic medical record information extraction model includes multiple sub-models, such as a first sub-model, a second sub-model, a third sub-model, and a fourth sub-model. Among them, the above four sub-models are as described in the foregoing content, and the training process of the electronic medical record information extraction model can be completed by means of operations research optimization.
[0139] In the embodiments of the present invention, operations research optimization is used in the process of training an electronic medical record information extraction model.
[0140] It can be understood that operations research optimization is mainly used to solve decision-making problems in the information extraction process. These problems can be formalized as an optimization problem, such as using linear programming or integer programming to optimize the information extraction strategy, which includes a series of decision variables, objective functions, and constraint conditions. Feeding back the results of operations research optimization into the model training process forms an iterative loop to continuously optimize the model performance. Operations research optimization may be closely integrated with the training process of deep learning models, guiding various aspects of model training through mathematical programming methods, thereby improving the model's performance in electronic medical record information extraction and structuring tasks, more precisely controlling the model's learning process, and ensuring that the model can effectively generalize to new datasets while meeting specific performance metrics.
[0141] The operations research optimization process mainly includes the following aspects:
[0142] 1. Feature selection and weight assignment: Operations research optimization is used to determine the weights of different features (such as text features, image features, etc.) in the model to improve the accuracy of information extraction. This involves defining decision variables, where each feature can be assigned a weight value, and determining these weights by optimizing the objective function to achieve feature selection.
[0143] 2. Model parameter adjustment: Select the optimal model parameter configuration through operations research methods. For example, in a multi-task learning model, integer programming can be used to balance the processing intensity of text and image data to improve the model performance.
[0144] 3. Information extraction strategy optimization: Use linear programming or integer programming to optimize the information extraction strategy to ensure that the key information extracted is both accurate and comprehensive. This includes setting the objective function and constraint conditions to optimize the information extraction process.
[0145] 4. Resource allocation: When dealing with large-scale electronic medical record data, operations research can help optimize the allocation of computing resources. For example, reasonably allocate computing nodes and storage resources in a cloud environment to improve the processing efficiency.
[0146] The operations research optimization process is described as follows. The operations research optimization process includes linear programming and integer programming:
[0147] In the related art, the calculation formula for the optimization objective of the linear programming process is shown as formula (10) below:
[0148] (10)
[0149] Where, is the cost vector, is the decision variable vector, is the constraint matrix, is the constraint value vector.
[0150] For the penalty term of the combined category constraint matrix of linear programming and integer programming, aiming at the characteristics of the unbalanced distribution of entity relationships in electronic medical records, the penalty term based on the category constraint matrix is derived, and the model is learned and trained together with the loss function. Based on this, in the embodiments of the present application, the improved linear programming calculation formula is shown as the following formula (11):
[0151] (11)
[0152] Wherein, is the cost vector, is the decision variable vector, is the constraint matrix, is the constraint value vector. is the penalty term coefficient, is the penalty term based on the category constraint matrix.
[0153] (2) Integer programming:
[0154] The calculation formula of the optimization objective is shown as the following formula (12):
[0155] subject to quad (12)
[0156] Wherein, is the integer decision variable vector.
[0157] Such as Figure 3 shown, Figure 3 shows the training process of the electronic medical record information extraction model. In this process, model parameters need to be set in advance and trained in the way shown in the figure.
[0158] The electronic medical record information extraction method in the embodiments of the present application is summarized as follows. The electronic medical record information extraction method proposed in the embodiments of the present application and the electronic medical record information extraction model used in this method mainly achieve a significant improvement in the extraction effect of electronic medical record information through the combination of deep learning and operations research optimization techniques. In the actual application process of this model, first, in the data preprocessing stage, basic word segmentation, part-of-speech tagging, and grayscale and binarization processing of images are performed. At the same time, a domain-adaptive Word2Vec model is introduced to generate word vectors closer to the medical context, and local normalization processing is used to more effectively cope with the illumination changes in images. Specifically, this model is composed of multiple sub-models and belongs to a multi-task learning model. This model combines BERT and LSTM to process text data, uses CNN to process image data, and introduces a multi-channel self-attention mechanism to enhance the model's ability to identify key information. In addition, by integrating operations research optimization techniques, the information extraction process is optimized by introducing a penalty term based on a category constraint matrix, especially performing well in dealing with the problem of unbalanced distribution of entity relationships in electronic medical records. In the information extraction and structuring stage, this method uses an improved BiLSTM-CRF model, and enhances the recognition accuracy of entity boundaries by introducing weight-based position embeddings. These optimization measures not only improve the accuracy and efficiency of the model, but also enhance the model's ability to handle the complexity and diversity of electronic medical record data, making the present invention innovative and practical in the field of electronic medical record information extraction technology.
[0159] The entire process of obtaining structured information described above is summarized through a flowchart as follows, as Figure 4 shown:
[0160] When the electronic medical record information extraction model obtains the electronic medical record, it will first preprocess the data of the original electronic medical record (S101). This data preprocessing mainly performs word segmentation, part-of-speech tagging on unstructured data such as medical record records, inspection reports, and medication records in the electronic medical record, as well as grayscale and binarization processing of images, to obtain the processed first electronic medical record. Then, feature extraction is performed (S102). Specifically, models such as Word2Vec are used to extract the text features of word vectors in the text, and normalization models, such as CNN models, are used to extract image features. Then, the text features and image features are fused through a multi-task learning model (S103). Specifically, the BERT model and LSTM model are used to process the text features, and the CNN model is used to process the image features, to obtain the key information in the text features and the key information in the image features. Then, through the multi-channel self-attention model of the attention mechanism, it automatically identifies and focuses on the key information in the above text features and the key information in the image features (S104). The BiLSTM-CRF model is used to identify entity information (S105) from the text, such as disease names, drug names, corresponding treatment plans, etc., and perform relationship extraction (S106) to identify the relationships between entities. Finally, according to predefined structured formats, such as JSON and XML, the extracted information and relationships are organized into structured information, and the final structured information is output (S107). These information can be used for further data analysis, clinical decision support and other applications.
[0161] It can be understood that in the above process, the model will identify the key medical record information in the first electronic medical record.
[0162] The following describes the application process of the present application in specific application scenarios as follows:
[0163] In actual application, assume there is an original electronic medical record of Zhang San, which records the basic information, chief complaint, medical history summary, examination results, diagnosis and corresponding treatment plan of patient Zhang San. The original electronic medical record contains key information such as Zhang San's current symptoms, medication records, examination and test value results, and clinical diagnosis and treatment activities.
[0164] Its information includes:
[0165] 1. Patient information: Zhang San, male, 45 years old;
[0166] 2. Chief complaint: Weight loss of 5 kg in the past month, accompanied by intermittent fever;
[0167] 3. Medical history summary: The patient has a history of diabetes and has been taking hypoglycemic drugs for a long time;
[0168] 4. Examination results: Blood routine shows a high white blood cell count, and blood sugar is 8.5 mmol / L.
[0169] 5. Diagnosis: Drug-induced liver injury, acute, hepatocyte injury type, RUCAM score: 10 points.
[0170] For the original electronic medical record processing flow, multiple different sub-models in the electronic medical record information extraction model need to be used, including the Word2Vec model, normalization model, BERT model, LSTM model, CNN model, multi-channel self-attention model, structured model, etc. The specific processing process includes:
[0171] 1. Data preprocessing
[0172] 1) Segment and tag the part-of-speech of the text information in the original electronic medical record, and generate word vectors using the domain-adapted Word2Vec model.
[0173] 2) Perform local normalization processing on the image information (such as X-ray films, MRI scans) to improve the stability of image features.
[0174] 2. Extract features through the model and construct the model
[0175] 1) Use the BERT model to process the text information data and extract text features (i.e., the above-mentioned first text features).
[0176] 2) Use the CNN model to process the image information data and extract image features (i.e., the above-mentioned first image features).
[0177] 3) Combine the text features and image features, and further identify the key text features (i.e., the above-mentioned second text features) and key image features (i.e., the above-mentioned second image features) according to the pre-set multi-task learning model.
[0178] For example, the key text features and key image features identified from the above information are:
[0179] Patient information: 1) Name: Zhang San
[0180] 2) Gender: Male
[0181] 3) Age: 45 years old
[0182] Chief complaint: 1) Weight loss of 5 kg
[0183] 2) Intermittent fever
[0184] Medical history summary: 1) History of diabetes
[0185] 2) Long-term use of hypoglycemic drugs
[0186] Examination results: 1) Blood routine: High white blood cell count
[0187] 2) Blood glucose: 8.5 mmol / L
[0188] Diagnosis: 1) Drug-induced liver injury, acute, hepatocyte injury type
[0189] 2) RUCAM score: 10 points
[0190] Treatment plan: 1) Insulin injection, once a day
[0191] Antibiotic treatment for 7 consecutive days
[0192] 3. Apply the attention mechanism
[0193] By using the multi-channel self-attention mechanism, identify the entity information in the above key text features and key graph features.
[0194] In this example, three types of recognition results need to be carried out, namely medical record information recognition, drug information recognition, and clinical information recognition. Specifically, medical record information recognition includes: identifying the patient's basic information, chief complaint, and medical history summary; drug information recognition includes: identifying the hypoglycemic drugs and antibiotics used by the patient; corresponding plan recognition includes: identifying the corresponding plan for the patient's symptoms
[0195] 4. Information extraction and structuring
[0196] 1) Use the BiLSTM-CRF model for named entity recognition to confirm entity information.
[0197] 2) Conduct relationship extraction to identify the relationships between entities.
[0198] 3) Organize the extracted information and relationships into structured data according to a predefined structured format.
[0199] In this example, the structured data is the "structured information is as follows (JSON format)" in the foregoing content
[0200] It should be noted that for the above electronic medical record information extraction model, pre-training is also required. The training process can include the process of operational research optimization. Specifically, this operational research optimization can optimize the information extraction process through the linear programming and integer programming models of the training model to ensure the accuracy and integrity of the information extracted by the model.
[0201] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the training rule determination method described in the embodiment of the present application. For example, it can execute Figure 1 each step of the method shown
[0202] An embodiment of the present application provides a computer program product, which includes instructions that, when run by a processor, implement Figure 1 each step of the method shown.
[0203] It should be noted that although the operations of the method of the present invention are described in a specific order in the drawings, this does not require or imply that these operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired result.
[0204] Figure 5 It is a block diagram of an electronic medical record information extraction system according to an embodiment of the present application. Refer to Figure 5 , the system includes an extraction unit 601, an execution unit 602, an acquisition unit 603, and a generation unit 604.
[0205] The extraction unit 601 is configured to extract the first text feature in the first electronic medical record information through the first sub-model in the electronic medical record information extraction model, and extract the first picture feature in the first electronic medical record information through the first sub-model. The first sub-model is a Word2Vec model and a normalization model;
[0206] The execution unit 602 is configured to process the text feature in the first electronic medical record information and the image feature in the first electronic medical record information through the second sub-model in the electronic medical record information extraction model, and identify the second text feature in the first text feature and the second picture feature in the first picture feature. The second sub-model is a multi-task model, and the multi-task model includes a BERT model, an LSTM model, and a CNN model. The second text feature is part or all of the text features in the first text feature, and the second picture feature is part or all of the picture features in the first picture feature;
[0207] The acquisition unit 603 is configured to identify the second text feature and the second picture feature through the third sub-model in the electronic medical record information extraction model, and acquire the entity information corresponding to the first electronic medical record information. The entity information includes: user information, user status, and conclusive information of the user status. The third sub-model includes a multi-channel self-attention model;
[0208] The generation unit 604 is configured to process the entity information in a predefined structured format through the fourth sub-model in the electronic medical record information extraction model, and generate structured information. The structured information is used to display the key information in the entity information in a preset format. The fourth sub-model is a structured model.
[0209] In one embodiment, the execution unit 602 is further configured to perform data preprocessing on the initial electronic medical record information according to a preset preprocessing method to obtain the first electronic medical record information; the preset preprocessing method includes: information word segmentation, part-of-speech tagging, grayscale processing, and binarization processing.
[0210] In one embodiment, the extraction unit 601 is specifically configured to identify the first text information in the first electronic medical record information through a Word2Vec model; obtain all lexical features in the first text information, process the lexical features into lexical vectors, and obtain the association relationships between different lexical vectors; generate first text features according to the lexical vectors and the association relationships between different lexical vectors; wherein, the first text features are feature information composed of text vectors.
[0211] In one embodiment, the extraction unit 601 is specifically configured to identify all the picture information in the first electronic medical record information through the normalization processing model; obtain the illumination brightness information of each picture information, where the illumination brightness information includes the illumination brightness information of multiple local regions in each picture information, and one picture information includes the multiple local regions; perform normalization processing on the illumination brightness information of each picture information to generate first picture features, where the first picture features include multiple first sub-picture features in the first electronic medical record information, and one first sub-picture feature corresponds to one picture information.
[0212] In one embodiment, the acquisition unit 603 is specifically configured to extract the text sequence vectors in the first text features through a BERT model and an LSTM model, and process the first text features according to the text sequence vectors to generate second text features; extract the text sequence vectors in the first picture features through a CNN model, and process the first picture features according to the text sequence vectors to generate second picture features.
[0213] In one embodiment, the production unit 604 is specifically configured to identify the entity information and extract the entity relationships in the entity information through a BiLSTM-CRF model in a structured model, and determine the entity boundaries of the entity information; determine the key information in the entity information according to the entity information, the entity relationships in the entity information, and the entity boundaries, and generate structured information.
[0214] In one embodiment, the execution unit 602 is further configured to obtain multiple groups of original electronic medical records and target data corresponding to the multiple groups of original electronic medical records. The target data includes target structured information. The structured information corresponding to the multiple groups of original electronic medical records is structured information in a target format that corresponds one-to-one to each sub-medical record in the multiple groups of original electronic medical records. The multiple groups of original electronic medical records include data information of at least two users. Input the multiple groups of original electronic medical records into an initial electronic medical record information extraction model to generate test data corresponding to the original electronic medical records. The test data includes test structured information. Compare the multiple groups of test data with the multiple groups of target data and determine the difference between the multiple groups of test data and the multiple groups of target data. In the case where the difference is greater than a preset threshold, adjust the initial electronic medical record information extraction model according to the difference so that the difference between the test data generated by the electronic medical record information processing model and the target data is less than the preset threshold. The initial electronic medical record information extraction model includes an initial Word2Vec model, a normalization model, a BERT model, an LSTM model, a CNN model, a multi-channel self-attention model, and a structured model.
[0215] It should be understood that the various units described in the electronic medical record information extraction system correspond to the respective steps in the method described in the drawings. Thus, the operations and features described above for the method also apply to the electronic medical record information extraction system and the units included therein, and will not be repeated here. The electronic medical record information extraction system can be pre-implemented in a browser or other secure application of a computer device, or can be loaded into the browser or its secure application of the computer device by means such as downloading. The corresponding units in the electronic medical record information extraction system can cooperate with the units in the computer device to implement the solutions of the embodiments of the present application.
[0216] Regarding the several modules or units mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0217] It should be noted that for the details not disclosed in the electronic medical record information extraction system of the embodiments of the present application, please refer to the details disclosed in the above embodiments of the present application, and will not be repeated here.
[0218] The following refers to Figure 6 , Figure 6 which shows a schematic structural diagram of a computer device suitable for implementing the embodiments of the present application. As Figure 6As shown, computer system 1700 includes a central processing unit (CPU) 1701, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1702 or programs loaded from a storage section 1708 into a random access memory (RAM) 1703. In the RAM 1703, various programs and data required for the operation instructions of the system are also stored. The CPU 1701, ROM 1702, and RAM 1703 are connected to each other via a bus 1704. An input / output (I / O) interface 1705 is also connected to the bus 1704.
[0219] The following components are connected to the I / O interface 1705; an input section 1706 including a keyboard, a mouse, etc.; an output section 1707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1708 including a hard disk, etc.; and a communication section 1709 including a network interface card such as a LAN card, a modem, etc. The communication section 1709 performs communication processing via a network such as the Internet. A drive 1710 is also connected to the I / O interface 1705 as needed. A removable medium 1711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1710 as needed so that a computer program read from it can be installed into the storage section 1708 as needed.
[0220] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart Figure 1 can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1709, and / or installed from the removable medium 1711. When the computer program is executed by a central processing unit (CPU) 1701, the above functions defined in the system of the present application are executed.
[0221] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution unit, apparatus, or device. In this application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution unit, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0222] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operation instructions of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the foregoing module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two connected blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for executing the specified functions or operation instructions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0223] The units or modules involved in the embodiments described in this application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. For example, it can be described as: a processor includes a first charging module, a second charging module, and a sending module. Among them, the names of these units or modules do not, in some cases, limit the units or modules themselves.
[0224] As another aspect, this application also provides a computer-readable storage medium. The computer-readable storage medium can be included in the electronic device described in the above embodiments, or can exist alone without being assembled into the electronic device. The above computer-readable storage medium stores one or more programs, and when the above programs are executed by one or more processors, they are used to perform the method for extracting electronic medical record information described in this application.
[0225] The above description is only a preferred embodiment of this application and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of disclosure involved in this application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the foregoing disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in this application.
Claims
1. A method for extracting electronic medical record information, characterized in that, Including: Extracting the first text feature from the first electronic medical record information through the first sub-model in the electronic medical record information extraction model, and extracting the first image feature from the first electronic medical record information through the first sub-model, where the first sub-model is a Word2Vec model and a normalization model; Processing the text feature in the first electronic medical record information and the image feature in the first electronic medical record information through the second sub-model in the electronic medical record information extraction model, identifying the second text feature in the first text feature and the second image feature in the first image feature, where the second sub-model is a multi-task model, and the multi-task model includes a BERT model, an LSTM model, and a CNN model, the second text feature is part or all of the text features in the first text feature, and the second image feature is part or all of the image features in the first image feature; Identifying the second text feature and the second image feature through the third sub-model in the electronic medical record information extraction model, and obtaining the entity information corresponding to the first electronic medical record information, where the entity information includes: user information, user status, and conclusive information of the user status, and the third sub-model includes a multi-channel self-attention model; Processing the entity information through the fourth sub-model in the electronic medical record information extraction model according to a predefined structured format to generate structured information, where the structured information is used to display the key information in the entity information according to a preset format, and the fourth sub-model is a structured model, and where, according to the entity information, the entity relationship in the entity information, and the entity boundary, the key information in the entity information is determined and the structured information is generated.
2. The method according to claim 1, wherein Before extracting the first text feature from the first electronic medical record information through the first sub-model in the electronic medical record information extraction model, the method includes: Performing data preprocessing on the initial electronic medical record information according to a preset preprocessing method to obtain the first electronic medical record information; The preset preprocessing method includes: information word segmentation, part-of-speech tagging, grayscale processing, and binarization processing.
3. The method according to claim 1, wherein The extracting the first text feature from the first electronic medical record information through the first sub-model in the electronic medical record information extraction model includes: Identifying the first text information in the first electronic medical record information through the Word2Vec model; Obtaining all the lexical features in the first text information, processing the lexical features into lexical vectors, and obtaining the association relationship between different lexical vectors; Generating the first text feature according to the lexical vectors and the association relationship between different lexical vectors; Wherein, the first text feature is feature information composed of text vectors.
4. The method according to claim 1, wherein The extracting the first image feature from the first electronic medical record information through the second sub-model includes: Identifying all the image information in the first electronic medical record information through the normalization processing model; Obtaining the illumination brightness information of each image information, where the illumination brightness information includes the illumination brightness information of multiple local regions in each image information, and one image information includes the multiple local regions; Normalize the illumination brightness information for each of the picture information to generate first picture features, where the first picture features include a plurality of first sub-picture features in the first electronic medical record information, and one first sub-picture feature corresponds to one picture information.
5. The method according to claim 1, wherein The processing of the first text features in the first electronic medical record information and the image features in the first electronic medical record information by the second sub-model in the electronic medical record information extraction model to identify the second text features in the first text features and the second picture features in the first picture features includes: Extract the text sequence vector in the first text features through a BERT model and an LSTM model, and process the first text features according to the text sequence vector to generate second text features; Extract the text sequence vector in the first picture features through a CNN model, and process the first picture features according to the text sequence vector to generate second picture features.
6. The method according to claim 1, wherein The processing of the entity information by the fourth sub-model in the electronic medical record information extraction model according to a predefined structured format to generate structured information includes: Through the BiLSTM-CRF model in the structured model, identify the entity information, extract the entity relationships in the entity information, and determine the entity boundaries of the entity information.
7. The method according to claim 1, characterized in that The method further includes: Obtain multiple groups of original electronic medical records and the target data corresponding to the multiple groups of original electronic medical records, where the target data includes: target structured information, and the structured information corresponding to the multiple groups of original electronic medical records is structured information in a target format that corresponds one by one to each sub-medical record in the multiple groups of original electronic medical records, and the multiple groups of original electronic medical records include data information of at least two users; Input the multiple groups of original electronic medical records into an initial electronic medical record information extraction model to generate test data corresponding to the original electronic medical records, where the test data includes test structured information; Compare the multiple groups of test data with the multiple groups of target data, and determine the difference between the multiple groups of test data and the multiple groups of target data; In the case where the difference is greater than a preset threshold, adjust the initial electronic medical record information extraction model according to the difference, so that the difference between the test data generated by the electronic medical record information processing model and the target data is less than the preset threshold. The initial electronic medical record information extraction model includes an initial Word2Vec model, a normalization model, a BERT model, an LSTM model, a CNN model, a multi-channel self-attention model, and a structured model.
8. An electronic medical record information extraction system, characterized in that, Includes: An extraction unit for extracting the first text features in the first electronic medical record information through the first sub-model in the electronic medical record information extraction model, and extracting the first picture features in the first electronic medical record information through the first sub-model, where the first sub-model is a Word2Vec model and a normalization model; An execution unit, configured to process the text features in the first electronic medical record information and the image features in the first electronic medical record information through a second sub-model in the electronic medical record information extraction model, to identify second text features in the first text features and second image features in the first picture features, where the second sub-model is a multi-task model, and the multi-task model includes a BERT model, an LSTM model, and a CNN model; the second text features are partial text features or all text features in the first text features, and the second image features are partial image features or all image features in the first picture features; An acquisition unit, configured to identify the second text features and the second image features through a third sub-model in the electronic medical record information extraction model, to obtain entity information corresponding to the first electronic medical record information, where the entity information includes: user information, user status, and conclusive information of the user status; the third sub-model includes a multi-channel self-attention model; A generation unit, configured to process the entity information in a predefined structured format through a fourth sub-model in the electronic medical record information extraction model, to generate structured information, where the structured information is used to display key information in the entity information in a preset format; the fourth sub-model is a structured model, and where, according to the entity information, the entity relationships in the entity information, and entity boundaries, key information in the entity information is determined, and the structured information is generated.
9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the method described in any one of claims 1-7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, the method described in any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Information extraction method, device and equipment and computer readable storage medium
CN116912842A
Electronic medical record feature extraction method based on natural language processing
CN119361058A