Information extraction method, terminal device and readable storage medium for medical record data
Through the character-based entity object recognition and traversal relationship extraction model, combined with vocabulary information and remote supervision, the problem of low accuracy in medical record data information extraction is solved, and more efficient entity object recognition and relationship extraction is achieved.
Patent Information
- Application Number
- CN202111438121.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-29
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-11-29
AI Technical Summary
In the existing medical record data information extraction methods, there is a problem of low accuracy in entity recognition and relationship extraction, especially character-based entity recognition methods lose vocabulary information, while vocabulary-based methods rely on entity recognition results, resulting in cumulative errors.
Character-based entity object recognition method is adopted, entity objects are marked through position encoding, and vocabulary information is introduced. The structure is marked by a stacked pointer, combined with a traversal relationship extraction model, the subject object is randomly extracted and the object object and its relationship are predicted. Remote supervision and conditional layer standardized structure optimization extraction process are adopted.
It effectively improves the performance of Chinese entity object recognition, reduces error rate and complexity, and improves the accuracy of medical record data information extraction.
Smart Images

Figure CN114220505B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of data processing, and particularly relates to a method for extracting information from medical record data, a terminal device, and a computer-readable storage medium. Background Art
[0002] The processing and structuring of clinical medical data represented by electronic medical record data have always been a major difficulty in the development of medical informatization. In the field of medical informatization, information extraction is an important step in medical record parsing and structuring, and entity recognition and entity relationship extraction are the core tasks of information extraction.
[0003] Among them, due to errors in Chinese word segmentation in entity recognition, character-based entity recognition methods are usually better than vocabulary-based entity recognition methods, which can avoid errors during word segmentation. However, character-based entity recognition methods are prone to losing lexical information in the text, resulting in low entity recognition accuracy. And currently, entity relationship extraction heavily depends on the results of entity extraction, and is prone to the problem of error accumulation, resulting in low accuracy of information extraction.
[0004] In summary, the current information extraction of medical record data has the problem of low extraction accuracy. Summary of the Invention
[0005] In view of this, embodiments of this application provide a method for extracting information from medical record data, a terminal device, and a computer-readable storage medium to solve the problem of low extraction accuracy in the current information extraction of medical record data.
[0006] In a first aspect, embodiments of this application provide a method for extracting information from medical record data, including:
[0007] Identifying all entity objects from a medical record statement and labeling all the entity objects with position encodings;
[0008] Randomly extracting a subject object from the entity objects, and extracting an object object corresponding to the subject object and the relationship between the subject object and the object object based on the subject object, until all entity objects are traversed to obtain the extraction results of all entity objects.
[0009] Optionally, the identifying all entity objects from a medical record statement and labeling all the entity objects with position encodings includes:
[0010] Constructing a head position encoding and a tail position encoding for each character;
[0011] Inputting the medical record statement labeled with the head position encoding and the tail position encoding into a language representation model for entity recognition to determine all entity objects in the medical record statement.
[0012] Optionally, randomly extracting a subject object from the entity objects, and extracting an object object corresponding to the subject object and the relationship between the subject object and the object object based on the subject object until all entity objects are traversed to obtain the extraction results of all entity objects, including:
[0013] Randomly extracting a subject object from the entity objects;
[0014] Extracting the object object corresponding to the subject object through a traversal relationship extraction model;
[0015] Predicting the relationship between the subject object and the object object according to the subject object and the object object;
[0016] Taking the object object as the subject object and repeating the operation of predicting the object object corresponding to the subject object and the relationship between the subject object and the object object through the traversal relationship extraction model until the extraction results of all entity objects are obtained.
[0017] Optionally, randomly extracting a subject object from the entity objects, and extracting an object object corresponding to the subject object and the relationship between the subject object and the object object based on the subject object until all entity objects are traversed to obtain the extraction results of all entity objects, including:
[0018] Randomly extracting a subject object from the entity objects;
[0019] Predicting the object object corresponding to the subject object and the relationship between the subject object and the object object through a traversal relationship extraction model;
[0020] Taking the object object as the subject object and repeating the operation of predicting the object object corresponding to the subject object and the relationship between the subject object and the object object through the traversal relationship extraction model until the extraction results of all entity objects are obtained.
[0021] Optionally, the traversal relationship extraction model includes a first multi-head attention mechanism layer, a second multi-head attention mechanism layer, a first normalization layer, a feed-forward neural network, a point cloud dynamic graph convolutional neural network, and a second normalization layer.
[0022] Optionally, after randomly extracting a subject object from the entity objects, and extracting an object object corresponding to the subject object and the relationship between the subject object and the object object based on the subject object until all entity objects are traversed to obtain the extraction results of all entity objects, it further includes: supervising the extraction results based on distant supervision.
[0023] Optionally, before identifying all entity objects from the medical record statements and annotating all the entity objects with position encoding, it further includes: segmenting the medical record statements from the medical record data.
[0024] In a second aspect, an embodiment of the present application provides a terminal device, including:
[0025] An entity recognition module, configured to identify all entity objects from the medical record statements and annotate all the entity objects with position encoding;
[0026] A relationship extraction module, configured to randomly extract a subject object from the entity objects, and based on the subject object, extract an object object corresponding to the subject object and the relationship between the subject object and the object object until all entity objects are traversed to obtain the extraction results of all entity objects.
[0027] In a third aspect, an embodiment of the present application provides a terminal device, the terminal device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, it implements the method as described in the first aspect or any optional manner of the first aspect.
[0028] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method as described in the first aspect or any optional manner of the first aspect.
[0029] In a fifth aspect, an embodiment of the present application provides a computer program product, when the computer program product runs on a terminal device, it causes the terminal device to execute the method as described in the first aspect or any optional manner of the first aspect.
[0030] Implementing an information extraction method and terminal device, computer-readable storage medium, and computer program product for medical record data provided by the embodiments of the present application has the following beneficial effects:
[0031] By using a character-based method for entity object recognition, introducing lexical information on the basis of characters, annotating each entity object with a position pointer, and using a cascaded pointer annotation as the basic structure, it can solve the problems of multiple relationships and entity overlaps of entity pairs, effectively improve the performance of Chinese entity object recognition, and adopt a subject-aware joint scheme based on traversal for relationship extraction (that is, by randomly extracting a subject object and predicting the corresponding object object and the relationship between the two), which can effectively reduce the error rate and complexity, improve the extraction accuracy, and solve the problem of low extraction accuracy in the current information extraction of medical record data. Description of the Drawings
[0032] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for use in the embodiments or the description of the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0033] Figure 1 is a schematic flowchart of a method for extracting information from medical record data provided by an embodiment of the present application;
[0034] Figure 2 is a schematic diagram of the process of entity object recognition provided by an embodiment of the present application;
[0035] Figure 3 is a schematic diagram of the architecture of a traversal relationship extraction model provided by an embodiment of the present application;
[0036] Figure 4 is a schematic flowchart of a method for extracting information from medical record data provided by another embodiment of the present application;
[0037] Figure 5 is a schematic diagram of the scenario of the method for extracting information from medical record data provided by an embodiment of the present application;
[0038] Figure 6 is a schematic diagram of the structure of a terminal device provided by an embodiment of the present application;
[0039] Figure 7 is a schematic diagram of the structure of a terminal device provided by another embodiment of the present application;
[0040] Figure 8 is a schematic diagram of the structure of a computer-readable storage medium provided by an embodiment of the present application. Detailed implementation manners
[0041] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, the detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0042] It should be understood that the term "and / or" as used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. Additionally, in the description of the specification and appended claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0043] It should also be understood that the reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but rather mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0044] It should be noted that the medical record data mentioned in the embodiments of this application mainly refers to electronic medical record data, with Chinese medical record texts as the processing objects. The purpose is to extract the entity objects contained in the Chinese medical record texts and the relationships between the entity objects. It should be noted that the entity objects contained in the medical record texts include but are not limited to diseases, types, onset sites, examinations, treatments, etc. The relationships between the above-mentioned entity objects include but are not limited to various relationships such as etiology, drugs, treatment methods, onset sites, etc. Exemplarily, the medical record text is "Chemotherapy for malignant tumors can affect the oral mucosa and cause oral mucositis". Among them, "malignant tumor", "oral mucosa", and "oral mucositis" are all entity objects in this medical record text; and the relationship between "oral mucosa" and "oral mucositis" can be the relationship of the onset site.
[0045] Currently, in entity recognition, a character-based NER system can be used to identify the entity objects in medical record sentences. However, the NER system does not pay attention to lexical information and is prone to recognition errors. And currently, entity relationship extraction generally includes two categories: pipeline extraction and joint extraction. Pipeline relationship extraction divides relationship extraction into two parts: entity recognition and relationship prediction. This extraction method highly depends on the results of entity recognition, is prone to cumulative errors, and at the same time does not consider the correlation between the two parts, and will bring redundant information into the process of relationship extraction, resulting in a high error rate problem. While joint relationship extraction has problems such as entity overlap and non-single entity relationships.
[0046] To solve the above problems, an embodiment of the present application proposes a method for extracting information from medical record data. The method uses a character-based approach for entity object recognition, introduces lexical information on the basis of characters, marks each entity object with a position pointer, and uses a cascaded pointer annotation as the basic structure, which can solve the problems of multiple relationships and entity overlaps existing in entity pairs, effectively improve the performance of Chinese entity object recognition, and uses a subject-aware joint scheme based on traversal for relation extraction (that is, randomly extracting a subject object and predicting the corresponding object and the relationship between the two), which can effectively reduce the error rate and complexity, improve the extraction accuracy, and solve the problem of low extraction accuracy existing in the current information extraction of medical record data.
[0047] The following will provide a detailed description of the method for extracting information from medical record data, the terminal device, and the computer-readable storage medium provided by the embodiments of the present application:
[0048] Please refer to Figure 1 , Figure 1 FIG. is a schematic flowchart of a method for extracting information from medical record data provided by an embodiment of the present application. In the embodiment of the present application, the execution subject of the above method for extracting information from medical record data may be a terminal device. The above terminal device includes, but is not limited to, devices with computing capabilities such as mobile phones, tablet computers, desktop computers, servers, etc.
[0049] Specifically, as Figure 1 shown, the above method for extracting information from medical record data may include S11 to S12, which are described in detail as follows:
[0050] S11: Identify all entity objects from the medical record sentences and mark all entity objects with position encoding.
[0051] In the embodiment of the present application, by adding the lexical position of the sentence text to the head and tail of each character in each medical record sentence, head position encoding and tail position encoding are constructed for each character and each word, and then based on the head position encoding and tail position encoding of each character, the head position encoding and tail position encoding of the corresponding word can be determined. Determining the positions of each character and each word based on the head position encoding and tail position encoding, and obtaining the interaction relationship between each character and the corresponding word can effectively avoid the problem of repeated introduction of entity objects.
[0052] Exemplarily, as Figure 2As shown, the character "急" can match the word "新冠肺炎"; the character "支" can match the two words "桥管" and "兄弟炎". The head position of the character "急" is coded as 1, and the tail position is coded as 1; the head position of the character "性" is coded as 2, and the tail position is coded as 2, and the corresponding word "新冠肺炎" is coded as 1 and 2. The head position of the character "支" is coded as 3, and the tail position is also coded as 3; the head position of the character "气" is coded as 4, and the tail position is also coded as 4; the head position of the character "管" is coded as 5, and the tail position is also coded as 5; the head position of "炎" is coded as 6, and the tail position is also coded as 6. Therefore, the head position of the word "桥管" is coded as 3, and the tail position is coded as 5; the head position of the word "兄弟炎" is coded as 3, and the tail position is coded as 6.
[0053] In an embodiment of the present application, after position encoding of each character in the medical record sentence and determining the relative position encoding of each word, entity recognition can be performed based on the language representation model, and the entity objects in each medical record sentence can be identified through the language representation model.
[0054] In the embodiment of the present application, entity recognition can be implemented based on the BERT (Bidirectional Encoder Representation from Transformers) language representation model. It should be noted that the BERT language representation model is a pre-trained language representation model. It no longer uses the traditional unidirectional language model or the shallow splicing method of two unidirectional language models for pre-training as in the past, but uses a new masked language model (MLM) to generate deep bidirectional language representation. It should be noted that the embodiment of the present application can also use other types of language representation models to implement entity recognition, such as the XLNet model, the REALM model, etc.
[0055] It should also be noted that the identified entity objects are also distinguished based on the position code to avoid the problem of repeated introduction of entity objects.
[0056] It should be noted that the embodiments of the present application can specifically use the FT-BERT language representation model for entity recognition, wherein the FT-BERT language representation model is an applicable neural network model obtained by pre-training the BERT model on an unlabeled Chinese clinical corpus, and the model can utilize unlabeled domain-specific knowledge.
[0057] Please continue reading Figure 2, after pointer annotation is performed on each character of the medical record statement and then input into the FT-BERT language representation model for processing, the recognition results of entity objects can be input. For example, it is recognized that "bronchitis" is a disease, "acute" is a type, and "bronchus" is the location of onset, etc.
[0058] When recognizing entity objects, if the start position index is 1 and the end position index is 1, it indicates that the word is an entity object, and its attributes are related to the start position index and the end position index. For example, for "bronchitis", which corresponds to a disease, the position where the start position index is 1 is the position where the character "zhi" is located, and the position where the end position index is 1 is the position where the character "yan" is located, and the remaining positions are all 0.
[0059] Based on this, in an embodiment of the present application, the above S11 may include the following steps:
[0060] Construct head position encoding and tail position encoding for each character;
[0061] Input the medical record statement annotated with head position encoding and tail position encoding into the language representation model for entity recognition to determine all entity objects in the medical record statement.
[0062] S12: Randomly extract a subject object from the entity objects, and based on the subject object, extract the object object corresponding to the subject object and the relationship between the subject object and the object object until all entity objects are traversed to obtain the extraction results of all entity objects.
[0063] In an embodiment of the present application, based on the annotated entity objects, a random entity object is extracted from multiple entity objects as the subject object, and then the object object corresponding to the subject object is extracted through a traversal relationship extraction model. Then, based on the subject object and the object object, the relationship between the subject object and the object object is predicted to form a triple (the triple is the extraction result). First, the corresponding object object is predicted through the subject object, and then the object object is used as the subject object to further predict the next object object, and so on until the annotation ends.
[0064] It should be noted that randomly extracting an entity object from multiple entity objects as the subject object can also be achieved through the above traversal relationship extraction model.
[0065] It should be noted that the above traversal relationship extraction model may further include the language representation model described in S11, that is, embedding the language representation model into the above traversal relationship extraction model, and extracting entity objects, corresponding object objects, and the relationship between the entity object (subject object) and the object object through this traversal relationship extraction model.
[0066] In practical applications, an entity object can be randomly selected as the subject object, and then its corresponding object object and the relationship between the subject object and the object object can be directly predicted to form a triple. That is, the corresponding object object and the relationship between the subject object and the object object are combined as the prediction object for combined prediction. Moreover, after the prediction of a subject object is completed, the embodiments of the present application will also determine the object object corresponding to each entity object and the relationship between the subject object and the object object based on traversal-based subject perception. The above traversal-based subject perception method can first randomly select an entity object as the subject object, then predict its corresponding object object, and then input the subject object and the object object to predict the relationship between the subject object and the object object.
[0067] In an embodiment of the present application, the above traversal relationship extraction can be implemented based on a traversal-based relationship extraction model, and the above traversal-based relationship extraction model can be implemented based on an existing relationship extraction neural network, only adding a traversal process.
[0068] In an embodiment of the present application, please refer to Figure 3 , Figure 3 which shows a schematic structural diagram of a traversal-based relationship extraction model provided by the embodiments of the present application. As Figure 3 shown, the above traversal-based relationship extraction model includes a first multi-head attention mechanism layer (Multi-Head Attention1), a second multi-head attention mechanism layer (Multi-Head Attention2), a first normalization layer (Add&Norm1), a feed-forward neural network, a point cloud dynamic graph convolutional neural network (DGCNN), and a second normalization layer (Add&Norm2).
[0069] In the embodiments of the present application, through two multi-head attention mechanism layers, the first multi-head attention mechanism layer and the second multi-head attention mechanism layer are connected in parallel, so that the extracted underlying features can notice more comprehensive position information, grammatical information, and rare words. And a point cloud dynamic graph convolutional neural network is added after the feed-forward neural network, increasing the dilation width and expanding the receptive field of view, so that when performing convolutional operations, the data in the middle of the dilation width will be skipped, so that the same-sized convolutional kernel can obtain a wider input matrix data, improving the processing accuracy.
[0070] Please refer to Figure 4 , in an embodiment of the present application, the information extraction method for the above medical record data may further include the following steps:
[0071] S13: Supervise the extraction result based on remote supervision.
[0072] In the embodiments of the present application, in order to improve the accuracy of relation extraction, the extraction results can be supervised based on distant supervision. The above distant supervision can be achieved by forming a knowledge base with the triples in the training set. When processing new medical record sentences, search through the above knowledge base to obtain some candidate triples for the medical record sentence, and then use the candidate triples as features and input them into the above traversal relation extraction model. First, form a 0 / 1 vector similar to the annotation structure with all the entity objects obtained by distant supervision, then splice it into the encoded vector sequence, and then perform the prediction of the subject object; then form a 0 / 1 vector similar to the annotation structure with all the object objects and the corresponding relations obtained by distant supervision, splice it into the encoded vector sequence, and then perform the prediction of the object object and the corresponding relation, so as to realize the supervision of the extraction results.
[0073] It should be noted that when training the traversal relation extraction model, when constructing the distant supervision features, the triples of the current training sample itself should be excluded first, that is, only the triples of other samples can be used to generate the distant supervision results of the current sample, so as to effectively improve the accuracy of the extraction results.
[0074] In another embodiment of the present application, the generation behavior of the traversal relation extraction model can also be adjusted based on the conditional layer normalization structure. It should be noted that the process of adjusting the generation behavior of the model based on the conditional layer normalization structure can be implemented with reference to the existing Conditional Layer Normalization, and the present application will not elaborate on this.
[0075] In order to further describe that the information extraction method for medical record data provided by the embodiments of the present application can effectively extract entity objects and the relationships between entity objects, Figure 5 Fig. shows a schematic scenario diagram of the information extraction method for medical record data provided by the embodiments of the present application. As Figure 5 shown, taking "Chemotherapy for malignant tumors can affect the oral mucosa and cause oral mucositis" as an example, input it into FT-BERT to identify entity objects, and then trigger DGCNN-BERT based on the subject object to predict the object object and the corresponding relation, and finally output the extraction result based on the conditional layer normalization structure and distant supervision. It can be Figure 5 seen that the identified entity objects include "malignant tumor", "oral mucosa" and "oral mucositis". Randomly select an entity object as the subject object (for example, "oral mucositis" is selected). At this time, the relationship between "oral mucositis" and "malignant tumor" (the predicted object object) and the relationship between "oral mucositis" and "oral mucosa" (the predicted other object object) can be obtained, that is, the relationship between "oral mucositis" and "malignant tumor" is the cause, and the relationship between "oral mucositis" and "oral mucosa" is the site of onset.
[0076] In another embodiment of the present application, the information extraction method for the above medical record data may further include the following steps:
[0077] Segment medical record sentences from the medical record data.
[0078] In an embodiment of the present application, the above medical record data may be an electronic medical record text, and medical record sentences are segmented based on punctuation marks in the electronic medical record text. Specifically, it may be segmented based on "。".
[0079] As can be seen above, the information extraction method for medical record data provided by the embodiments of the present application uses a character-based method for entity object recognition, introduces lexical information on the basis of characters, marks each entity object through a position pointer, and uses a cascaded pointer annotation as the basic structure, which can solve the problems of multiple relationships and entity overlaps of entity pairs, effectively improve the performance of Chinese entity object recognition, and uses a subject-aware joint scheme based on traversal for relationship extraction (that is, randomly extracting a subject object and predicting the corresponding object and the relationship between the two), which can effectively reduce the error rate and complexity, improve the extraction accuracy, and solve the problem of low extraction accuracy in the current information extraction of medical record data.
[0080] It should be understood that the magnitudes of the sequence numbers of the above steps do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0081] Based on the information extraction method for medical record data provided by the above embodiments, the embodiments of the present invention further provide an embodiment of a terminal device for implementing the above method embodiments.
[0082] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a terminal device provided by an embodiment of the present application. In an embodiment of the present application, each unit included in the terminal device is used to execute Figure 1 the corresponding steps in the embodiment. Specifically, please refer to Figure 1 and Figure 1 the relevant descriptions in the corresponding embodiments. For the sake of convenience of description, only the parts related to this embodiment are shown. As Figure 6 shown, the terminal device 60 includes: an entity recognition module 61 and a relationship extraction module 62.
[0083] The entity recognition module 61 is used to identify all entity objects from the medical record sentences and mark all the entity objects through position encoding.
[0084] The relation extraction module 62 is used to randomly extract a subject object from the entity objects, and based on the subject object, extract an object object corresponding to the subject object and the relation between the subject object and the object object, until all entity objects are traversed to obtain the extraction results of all entity objects.
[0085] Optionally, the entity recognition module 61 is specifically used for:
[0086] Construct a head position encoding and a tail position encoding for each character;
[0087] Input the medical record statement marked with the head position encoding and the tail position encoding into the language representation model for entity recognition to determine all entity objects in the medical record statement.
[0088] Optionally, the above-mentioned relation extraction module 62 is specifically used for:
[0089] Randomly extract a subject object from the entity objects;
[0090] Extract the object object corresponding to the subject object through the traversal relation extraction model;
[0091] Predict the relation between the subject object and the object object according to the subject object and the object object;
[0092] Take the object object as the subject object and repeat the operation of predicting the object object corresponding to the subject object and the relation between the subject object and the object object through the traversal relation extraction model until the extraction results of all entity objects are obtained.
[0093] Optionally, the above-mentioned relation extraction module 62 is specifically further used for:
[0094] Randomly extract a subject object from the entity objects;
[0095] Predict the object object corresponding to the subject object and the relation between the subject object and the object object through the traversal relation extraction model;
[0096] Take the object object as the subject object and repeat the operation of predicting the object object corresponding to the subject object and the relation between the subject object and the object object through the traversal relation extraction model until the extraction results of all entity objects are obtained.
[0097] Optionally, the above-mentioned traversal relation extraction model includes a first multi-head attention mechanism layer, a second multi-head attention mechanism layer, a first normalization layer, a feed-forward neural network, a point cloud dynamic graph convolutional neural network, and a second normalization layer.
[0098] Optionally, the above terminal device 60 may further include a remote supervision module and a statement segmentation module, where:
[0099] The remote supervision module is used to supervise the extraction result based on remote supervision.
[0100] The statement segmentation module is used to segment medical record statements according to medical record data.
[0101] It should be noted that the information interaction, execution process, etc. between the above modules / units, due to being based on the same concept as the method embodiment of the present application, for their specific functions and the technical effects brought, please refer to the method embodiment part specifically, and will not be elaborated here.
[0102] Therefore, the terminal device provided by the embodiment of the present application can also identify entity objects in a character-based manner, introduce lexical information on the basis of characters, mark each entity object through a position pointer, and use the cascaded pointer annotation as the basic structure, which can solve the problems of multiple relationships and entity overlaps of entity pairs, effectively improve the performance of Chinese entity object recognition, and adopt a subject-aware joint scheme based on traversal for relationship extraction (that is, by randomly extracting the subject object and predicting the corresponding object and the relationship between the two), which can effectively reduce the error rate and complexity, improve the extraction accuracy, and solve the problem of low extraction accuracy in the current information extraction of medical record data.
[0103] Figure 7 It is a schematic structural diagram of a terminal device provided by another embodiment of the present application. As Figure 7 shown, the terminal device 7 provided by this embodiment includes: a processor 70, a memory 71, and a computer program 72 stored in the memory 71 and executable on the processor 70, such as an information extraction program for medical record data. When the processor 70 executes the computer program 72, it implements the steps in the above-mentioned method embodiments for information extraction of each medical record data, such as Figure 1 S11 - S12 shown. Alternatively, when the processor 70 executes the computer program 72, it implements the functions of each module / unit in the above-mentioned terminal device embodiments, such as Figure 6 the functions of the units 61 - 62 shown.
[0104] Exemplarily, the computer program 72 can be divided into one or more modules / units, and the one or more modules / units are stored in the memory 71 and executed by the processor 70 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 72 in the terminal device 7. For example, the computer program 72 can be divided into each unit / module, and for the specific functions of each unit / module, please refer toFigure 6 The relevant descriptions in the corresponding embodiments are not elaborated here.
[0105] The terminal device may include, but is not limited to, a processor 70 and a memory 71. Those skilled in the art can understand that Figure 7 These are merely examples of the terminal device 7 and do not constitute a limitation on the terminal device 7. It may include more or fewer components than shown in the figure, or combine certain components, or have different components. For example, the terminal device may further include an input / output device, a network access device, a bus, etc.
[0106] The so-called processor 70 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0107] The memory 71 may be an internal storage unit of the terminal device 7, such as the hard disk or memory of the terminal device 7. The memory 71 may also be an external storage device of the terminal device 7, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 7. Further, the memory 71 may also include both the internal storage unit and the external storage device of the terminal device 7. The memory 71 is used to store the computer program and other programs and data required by the terminal device. The memory 71 may also be used to temporarily store data that has been output or will be output.
[0108] The embodiments of the present application also provide a computer-readable storage medium. Please refer to Figure 8 , Figure 8 FIG. is a schematic structural diagram of a computer-readable storage medium provided by the embodiments of the present application. As Figure 8 shown, a computer program 81 is stored in the computer-readable storage medium 8. When the computer program 81 is executed by a processor, the above-mentioned information extraction method for medical record data can be realized.
[0109] An embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, it enables the terminal device to execute and implement the information extraction method for the above-mentioned medical record data.
[0110] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the terminal device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.
[0111] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0112] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0113] The above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application and should all be included in the protection scope of the present application.
Claims
1. An information extraction method for medical record data, characterized in that, Including: All entity objects are identified from the medical record statements in a character-based manner, and all the entity objects are labeled by position encoding; by adding the lexical position of the text of each sentence at the beginning and end of each character of each medical record sentence, head position encoding and tail position encoding are constructed for each character and each word, and then the head position encoding and tail position encoding of the word matching the character can be determined according to the head position encoding and tail position encoding of each character; Entity recognition is performed based on a language representation model, and the entity objects in each medical record sentence are identified through the FT-BERT language representation model; the identified entity objects are also distinguished based on position encoding; among them, the FT-BERT language representation model is an applicable neural network model obtained by pre-training the BERT model on an unlabeled Chinese clinical corpus, and this model can utilize unlabeled domain-specific knowledge; Subject objects are randomly extracted from the entity objects, and corresponding object objects and the relationships between the subject objects and the object objects are extracted based on the subject objects until all entity objects are traversed to obtain the extraction results of all entity objects; based on the labeled entity objects, an entity object is randomly extracted from multiple entity objects as the subject object, and then the object object corresponding to the subject object is extracted through the traversal relationship extraction model DGCNN-BERT, and then the relationship between the subject object and the object object is predicted according to the subject object and the object object to form a triple; the corresponding object object is predicted first through the subject object, and then the object object is used as the subject object to further predict the next object object, and so on until the annotation ends; the traversal relationship extraction model also includes the above-mentioned language representation model, that is, the language representation model is embedded into the traversal relationship extraction model, and entity objects, corresponding object objects and the relationships between the entity objects (i.e., subject objects) and the object objects are extracted through this traversal relationship extraction model; The method further includes: supervising the extraction results based on distant supervision, specifically, all entity objects obtained by distant supervision are formed into a 0 / 1 vector similar to the annotation structure, and then spliced into the encoding vector sequence, and then the prediction of the subject object is performed; then all object objects obtained by distant supervision and the corresponding relationships are also formed into a 0 / 1 vector similar to the annotation structure, and after being spliced into the encoding vector sequence, the prediction of the object object and the corresponding relationship is performed, thereby realizing the supervision of the extraction results.
2. The information extraction method for medical record data according to claim 1, wherein The traversal relationship extraction model includes a first multi-head attention mechanism layer, a second multi-head attention mechanism layer, a first normalization layer, a feed-forward neural network, a point cloud dynamic graph convolutional neural network, and a second normalization layer.
3. The information extraction method for medical record data according to any one of claims 1 to 2, characterized in that, Before all entity objects are identified from the medical record statements and all the entity objects are labeled by position encoding, it further includes: Segmenting the medical record statements according to the medical record data.
4. A terminal device, characterized in that, Including: The entity recognition module is used to recognize all entity objects from medical record statements in a character-based manner and label all the entity objects through position encoding; by adding the lexical position of the text of each sentence at the beginning and end of each character in each medical record sentence, head position encoding and tail position encoding are constructed for each character and each word, and then based on the head position encoding and tail position encoding of each character, the head position encoding and tail position encoding of the corresponding word can be determined; Entity recognition is performed based on a language representation model, and entity objects in each medical record sentence are recognized through the FT-BERT language representation model; the recognized entity objects are also distinguished based on position encoding; among them, the FT-BERT language representation model is a neural network model that can be applied obtained by pre-training the BERT model on an unlabeled Chinese clinical corpus, and this model can utilize unlabeled domain-specific knowledge; The relation extraction module is used to randomly extract a subject object from the entity objects, and based on the subject object, extract the corresponding object object and the relationship between the subject object and the object object until all entity objects are traversed to obtain the extraction results of all entity objects; based on the labeled entity objects, a random entity object is extracted from multiple entity objects as the subject object, and then the corresponding object object of the subject object is extracted through a traversal relation extraction model, and then the relationship between the subject object and the object object is predicted according to the subject object and the object object to form a triple; first, the corresponding object object is predicted through the subject object, and then the object object is used as the subject object to further predict the next object object, and so on until the annotation ends; the traversal relation extraction model also includes the above-mentioned language representation model, that is, the language representation model is embedded in the traversal relation extraction model, and the entity object, the corresponding object object, and the relationship between the entity object (i.e., the subject object) and the object object are extracted through this traversal relation extraction model; The remote supervision module is used to supervise the extraction results based on remote supervision. Specifically, it is used to form a 0 / 1 vector similar to the annotation structure from all entity objects obtained by remote supervision, and then splice it into the encoded vector sequence, and then perform the prediction of the subject object; then all object objects obtained by remote supervision and the corresponding relationships are also formed into a 0 / 1 vector similar to the annotation structure, and after splicing into the encoded vector sequence, the prediction of the object object and the corresponding relationship is performed, thereby realizing the supervision of the extraction results.
5. A terminal device, characterized in that, The terminal device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in any one of claims 1 to 3 is implemented.
6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
Method and device for extracting entity relationship in text, equipment and storage medium
CN113704392A