Entity Relationship Joint Extraction System and Method Based on Hybrid Feature Representation
By using a hybrid feature representation and multi-head self-attention mechanism in the knowledge extraction system, the problem of identifying multiple overlapping relationship triplets in industrial text data is solved, which significantly improves the accuracy and performance of knowledge extraction.
Patent Information
- Application Number
- CN202210202416.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-03
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-03-03
AI Technical Summary
The prior art is difficult to effectively deal with the problem of multiple overlapping relationship triplets in industrial text data, resulting in poor knowledge extraction performance.
An entity relationship joint extraction system based on mixed feature representation is adopted, which includes feature extraction module, feature fusion module, model building module and joint identification module. Fusion of character-level and word-level features through maximum pooling operations, a bidirectional LSTM model with attention mechanism is used for encoding, and a pre-trained BERT model is combined for relationship and tail entity recognition.
Through mixed feature representation and multi-head self-attention mechanism, the entities and relationships in industrial text data are effectively identified, which significantly improves the accuracy and performance of the knowledge extraction model, especially when dealing with overlapping triples of complex relationships.
Smart Images

Figure CN114595338B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of knowledge extraction, and particularly relates to an entity relation joint extraction system and method based on hybrid feature representation. Background Art
[0002] In recent years, pre-trained language models such as BERT and GPT have become very popular and achieved great success in various natural language understanding tasks, such as knowledge extraction, sentiment analysis, question answering, and language reasoning.
[0003] Although the method of fine-tuning pre-trained models has achieved great success in both the fields of named entity recognition and relation extraction, there will be a large number of nested entities and overlapping relation triples in some actual scenarios. Directly applying the fine-tuning pre-trained model to extract them, its performance is not perfect. The early relation-entity extraction research adopted a pipeline method, which first identified all entities in a sentence and then classified the relations for each entity pair. This method is prone to error propagation problems because early errors cannot be corrected later.
[0004] To solve this problem, joint learning methods for entities and relations have been successively proposed in the prior art. However, most methods cannot effectively handle the scenario where a sentence contains multiple mutually overlapping relation triples. Recently, span-based methods have been proposed and applied to named entity recognition, effectively solving the entity nesting problem. Its essence is to predict the start and end positions of entities and identify various types of entities through a combination method. However, its model is very likely to decode incorrect entities or non-entities. Therefore, how to effectively handle the scenario where a sentence contains multiple mutually overlapping relation triples has become a key issue in knowledge extraction. Summary of the Invention
[0005] In view of this, the present invention proposes an entity relation joint extraction system and method based on hybrid feature representation, which is used to solve the problem that multiple mutually overlapping relation triples cannot be effectively processed when performing knowledge extraction on industrial text data.
[0006] In the first aspect of the present invention, an entity relation joint extraction system based on hybrid feature representation is disclosed. The system includes:
[0007] A feature extraction module: used to extract character-level feature vectors and word-level feature vectors from industrial text data;
[0008] A feature fusion module: used to fuse the character-level feature vectors and word-level feature vectors using max pooling operation to generate hybrid feature vectors;
[0009] Model construction module: used to construct an entity relationship joint extraction model based on a bidirectional LSTM encoder, a head entity recognition unit, an entity type classification unit, and a relationship-tail entity recognition unit;
[0010] Joint recognition module: used to input the mixed feature vector into the entity relationship joint extraction model to identify all entities and relationships in the industrial text data.
[0011] Based on the above technical solutions, preferably, the feature extraction module is specifically used for:
[0012] Extract character-level feature vectors from industrial text data based on a CNN model. At the same time, use a Chinese word segmenter to segment the industrial text data, match the segmented words with external dictionary information and external knowledge bases, and obtain word-level feature vectors through the Word2Vec model.
[0013] Based on the above technical solutions, preferably, in the model construction module, the bidirectional LSTM encoder is a bidirectional LSTM model with an attention mechanism, which is used to encode the input mixed feature vector, extract the dependencies between long-distance named entities in the industrial text data text, and at the same time extract the correlations between characters, between characters and named entities, and between entity character positions in the industrial text data.
[0014] Based on the above technical solutions, preferably, in the model construction module, the head entity recognition unit includes two identical first binary classifiers, which are used to label the encoded mixed feature vector output by the bidirectional LSTM encoder. Each label is assigned a binary identifier to detect the start position and end position of the entity respectively. The start position and end position of the entity generate multiple entity feature vectors.
[0015] Based on the above technical solutions, preferably, in the model construction module, the entity type classification unit is used to splice each entity feature vector with the encoded mixed feature vector as the input, classify the entity through the probability output of Softmax, and set a probability threshold for entity filtering to remove entities and non-entities below the probability threshold, and retain entities greater than or equal to the probability threshold as head entities.
[0016] Based on the above technical solutions, preferably, in the model construction module, the relation-tail entity recognition unit regards the recognition of the relation and the tail entity as a machine reading comprehension task, obtains the description information of the relation through prior knowledge, splices the description information of the relation and the head entity as the question of the machine reading comprehension task, uses the encoded mixed feature vector as the passage of the machine reading comprehension task, embeds it into the pre-trained BERT model in the way of reading comprehension, and identifies the tail entity corresponding to the description information of the relation and the head entity through two second binary classifiers;
[0017] In the pre-trained BERT model, the multi-head self-attention mechanism is used to capture the interaction information between tokens, provide prior knowledge for industrial text data, and capture context semantic feature information during the training process, so as to eliminate the ambiguity of homophones and express semantic and syntactic patterns.
[0018] Based on the above technical solutions, preferably, in the relation-tail entity recognition unit, the second binary classifier outputs multiple start indices and multiple end indices for a given context and a specific query, and supports extracting all relevant entities according to the query.
[0019] In the second aspect of the present invention, an entity relation joint extraction method based on hybrid feature representation is disclosed, and the method includes:
[0020] S1. Extract character-level feature vectors and word-level feature vectors from industrial text data;
[0021] S2. Use max pooling operation to fuse the character-level feature vectors and word-level feature vectors to generate mixed feature vectors;
[0022] S3. Encode the input mixed feature vectors through a bidirectional LSTM model with an attention mechanism;
[0023] S4. Mark the encoded mixed feature vector h output by the bidirectional LSTM encoder through two identical first binary classifiers, and assign a binary identifier to each mark to respectively detect the start position and end position of the entity, generating multiple entity feature vectors; N S5. Splice each entity feature vector with the encoded mixed feature vector respectively, classify the entity through the probability output of Softmax, and perform entity filtering to retain the high-probability entities and their types as head entities;
[0024]
[0025] S6. Identify the relationship and the tail entity as a machine reading comprehension task. Use the pre-trained BERT model to encode two sentences, where the description information of the relationship and the head entity are concatenated as the question, and the encoded hybrid feature vector is used as the passage, and identify the overlapping triples with complex relationships through two second binary classifiers.
[0026] In the third aspect of the present invention, an electronic device is disclosed, including: at least one processor, at least one memory, a communication interface, and a bus;
[0027] Wherein, the processor, the memory, and the communication interface complete communication with each other through the bus;
[0028] The memory stores program instructions executable by the processor, and the processor invokes the program instructions to implement the method described in the second aspect of the present invention.
[0029] In the fourth aspect of the present invention, a computer-readable storage medium is disclosed. The computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to implement the method described in the second aspect of the present invention.
[0030] The present invention has the following beneficial effects compared with the prior art:
[0031] 1) The hybrid feature vector of the present invention integrates character-level information and word-level information, where the character-level feature vector provides morphological feature information; the word-level feature vector embedding combined with external dictionary information and external knowledge base provides boundary feature information. The hybrid feature vector enriches the hybrid feature information and improves the performance of entity boundary recognition.
[0032] 2) The present invention encodes the input hybrid feature vector through a bidirectional LSTM model with an attention mechanism, and performs head entity recognition, entity type classification and filtering, and relationship-tail entity recognition respectively based on the encoded hybrid feature vector, and finally realizes the recognition of overlapping triples with complex relationships. The present invention makes full use of character-word level, temporal structure, context embedding and other feature information, enriches the hybrid feature representation, integrates information at multiple granularity levels, reduces the weight of noise information, and at the same time, with the help of the self-attention mechanism, effectively captures the importance of different information in the text, eliminates the ambiguity of homophones, and significantly improves the accuracy of the joint extraction model. Description of the Drawings
[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0034] Figure 1 Schematic structural diagram of an entity relationship joint extraction system based on hybrid feature representation proposed by the present invention;
[0035] Figure 2 Schematic principle diagram of an entity relationship joint extraction system based on hybrid feature representation proposed by the present invention;
[0036] Figure 3 Schematic diagram of the bidirectional LSTM model with attention mechanism of the present invention is shown. Specific implementation manners
[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in combination with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0038] The present invention proposes an entity relationship joint extraction system based on hybrid feature representation, Figure 1 Schematic structural diagram of an entity relationship joint extraction system based on hybrid feature representation proposed by the present invention, the system includes a feature extraction module 10, a feature fusion module 20, a model construction module 30, and a joint recognition module 40.
[0039] Figure 2 Schematic principle diagram of an entity relationship joint extraction system based on hybrid feature representation proposed by the present invention, the following will be combined with Figure 1 、 Figure 2 Specifically illustrate the system principle of the present invention.
[0040] Feature extraction module 10: used to extract character-level feature vectors and word-level feature vectors from industrial text data, including a character-level feature extraction unit 101 and a word-level feature extraction unit 102.
[0041] The character-level feature extraction unit 101 extracts character-level feature vectors from industrial text data based on a CNN model to construct a text character-level vector representation. Meanwhile, the word-level feature extraction unit 102 uses a Chinese word segmenter to segment the industrial text data, matches the segmented words with external dictionary information and external knowledge bases, and obtains word-level feature vectors through a Word2Vec model to construct a text word-level vector representation.
[0042] Feature fusion module 20: It is used to fuse the character-level feature vectors and word-level feature vectors using max pooling operation to generate hybrid feature vectors.
[0043] The present invention fuses the character-level feature vectors and word-level feature vectors to construct a hybrid feature representation and generate hybrid feature vectors. Among them, the character-level vector representation provides morphological feature information, such as prefixes and suffixes of words, and the word-level vector embedding combining external dictionaries and domain knowledge bases provides boundary feature information. The hybrid feature vectors enrich the character feature information and can effectively solve the problem of polysemy.
[0044] Figure 2 The part of constructing the hybrid feature representation at the bottom schematically shows the feature extraction and feature fusion process of a certain text data, including the character-level feature vectors and the word-level feature vectors which are fused through max pooling operation. Among them, e1 is fused from the character-level feature vector and the word-level feature vector ; e2 is fused from the character-level feature vector and the word-level feature vector ; the fusion of other feature vectors is as shown in the part of constructing the hybrid feature representation in Figure 2 and the final fusion results of each one keep the same dimension. The fused feature vectors are combined to form hybrid feature vectors.
[0045] Model construction module 30: It is used to construct an entity relationship joint extraction model based on a bidirectional LSTM encoder 301, a head entity recognition unit 302, an entity type classification unit 303, and a relation-tail entity recognition unit 304;
[0046] The bidirectional LSTM encoder 301 is a bidirectional LSTM (Bi-LSTM, Bidirectional Long Short Term Memory) model with an attention mechanism, which is used to encode the input hybrid feature vectors and output the encoded hybrid feature vectors h N , Figure 3The figure shows a schematic diagram of the bidirectional LSTM model with an attention mechanism according to the present invention. The bidirectional LSTM model can further depict the dependencies between long-distance named entities in the text. To further capture the correlations between characters in the text, between characters and named entities, and between entity-character positions, a multi-head self-attention mechanism is developed in the Bi-LSTM layer, which can strengthen the dependencies between characters and words while improving the overall operating efficiency of the model.
[0047] The bidirectional LSTM encoder of the present invention adds an attention mechanism on the basis of the bidirectional LSTM model. On the one hand, it can effectively capture the information features within a specific time range and enhance the weights of key features in the text. On the other hand, it can effectively capture the global semantic information features in the text, further enrich the mixed feature representation, while reducing the cumulative error of semantic information transmission between layers and enhancing the correlations between entities in the text.
[0048] The head entity recognition unit 302 includes two identical first binary classifiers, which are used to label the encoded mixed feature vectors output by the bidirectional LSTM encoder, as Figure 2 shown. Each label is assigned a binary identifier to detect the start position and end position of the entity respectively, and k entity feature vectors are generated based on the start position and end position of the entity. And the encoded mixed feature vector h N is respectively concatenated with each entity feature vector to obtain
[0049] The entity type classification unit 303 is used to take each entity feature vector concatenated with the encoded mixed feature vector as input, classify the entity through the probability output of Softmax, and set a probability threshold for entity filtering to remove entities and non-entities below the probability threshold, and retain high-probability entities and types greater than or equal to the probability threshold as head entities. Taking the Agnews news dataset as an example, the types in the entity include: Sports, Business, World, Sci / Tech, and the Softmax layer outputs the probabilities that the entity belongs to these types. Suppose the probability threshold is set to 0.5. If the probabilities output by Softmax are 0.5, 0.2, 0.1, 0.2 respectively, it is considered that the probability of belonging to the first category belongs to a high-probability entity. If the probabilities output by Softmax are 0.3, 0.2, 0.2, 0.3, it is considered that the probability of belonging to the first category belongs to a low-probability entity or some non-entities.
[0050] The relation-tail entity recognition unit 304 regards the recognition of relations and tail entities as a machine reading comprehension task, that is, obtains the description information of relations through prior knowledge, splices the description information of relations and the head entity as the question of the machine reading comprehension task, uses the encoded hybrid feature vector as the passage of the machine reading comprehension task, embeds it into the pre-trained BERT model in the way of reading comprehension, and identifies the tail entity corresponding to the input relation description information and head entity through two second binary classifiers, so as to realize the recognition of overlapping triples with complex relations.
[0051] The description information R1,..., R of the relation n is artificially defined according to prior knowledge. For example, a relation such as "belong to" can be defined as:
[0052] part of: part of, belong to something, including, pertain, appertain, be classified.
[0053] The pre-trained BERT model is pre-trained on a large amount of data, which can provide prior knowledge for the text. At the same time, the model will capture more context semantic feature information during the training process. In the pre-trained BERT model, the multi-head self-attention mechanism is used to capture the interaction information between tokens, and provide the embedding of context semantic feature information and the prior knowledge in the pre-trained large-scale language model, so as to eliminate the ambiguity of homonyms and express semantic and syntactic patterns.
[0054] Among them, the second binary classifier outputs multiple start indices and multiple end indices for a given context and a specific query, supporting the extraction of all relevant entities according to the query.
[0055] The joint recognition module 40: is used to input the fused hybrid feature vector into the entity-relation joint extraction model to recognize all entities and relations in the industrial text data.
[0056] Specifically, the hybrid feature vector fused by the feature fusion module 20 is input into the entity-relation joint extraction model constructed by the model construction module 30 to capture the hidden features between them to recognize all entities and relations in the text, recognize overlapping triples, and solve the problem of polysemy.
[0057] The present invention makes full use of character-word level, temporal structure, context embedding and other feature information, enriches the hybrid feature representation, and at the same time, with the help of the multi-head self-attention mechanism, effectively identifies the boundaries of important entities, and significantly improves the accuracy and performance of the joint extraction model.
[0058] The parameters in the convolutional neural network, Word2Vec word embedding model, and bidirectional long short-term memory network provided by the present invention, the length of the input sentence in the BERT model, and the probability threshold in entity filtering can be set according to actual needs or device limitations and other factors.
[0059] Corresponding to the above system embodiment, the present invention also proposes a method for jointly extracting entity relationships based on hybrid feature representation, and the method includes:
[0060] S1. Extract character-level feature vectors and word-level feature vectors from industrial text data;
[0061] S2. Use max pooling operation to fuse the character-level feature vectors and word-level feature vectors to generate hybrid feature vectors;
[0062] S3. Encode the input hybrid feature vectors through a bidirectional LSTM model with an attention mechanism;
[0063] S4. Mark the encoded hybrid feature vectors output by the bidirectional LSTM encoder through two identical first binary classifiers, and assign a binary identifier to each mark to respectively detect the start position and end position of the entity, generating multiple entity feature vectors;
[0064] S5. Concatenate each entity feature vector with the encoded hybrid feature vectors respectively, classify the entities through the probability output of Softmax, and perform entity filtering, retaining entities and types greater than or equal to the probability threshold as head entities;
[0065] S6. Regard the recognition of relationships and tail entities as a machine reading comprehension task, use the pre-trained BERT model to encode two sentences with the description information of the relationship and the head entity concatenated as the question and the encoded hybrid feature representation as the paragraph, and identify the tail entity through two second binary classifiers, so as to realize the recognition of overlapping triples with complex relationships.
[0066] The above system embodiments and method embodiments correspond one by one. For the brief description of the method embodiments, please refer to the system embodiments.
[0067] The present invention also discloses an electronic device, including: at least one processor, at least one memory, a communication interface, and a bus; wherein, the processor, the memory, and the communication interface complete communication with each other through the bus; the memory stores program instructions executable by the processor, and the processor calls the program instructions to implement the method described above in the present invention.
[0068] The present invention also discloses a computer-readable storage medium storing computer instructions, which cause the computer to implement all or part of the steps of the method described in the embodiments of the present invention. The storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0069] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be distributed over multiple network units. Those of ordinary skill in the art can, without creative effort, select some or all of the modules according to actual needs to achieve the purpose of the solution of this embodiment.
[0070] The above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An entity relation joint extraction system based on hybrid feature representation, characterized in that, The system includes: A feature extraction module: used to extract character-level feature vectors and word-level feature vectors from industrial text data; A feature fusion module: used to fuse the character-level feature vectors and word-level feature vectors using max pooling operation to generate hybrid feature vectors; A model construction module: used to construct an entity relationship joint extraction model based on a bidirectional LSTM encoder, a head entity recognition unit, an entity type classification unit, and a relationship-tail entity recognition unit; In the model construction module, the bidirectional LSTM encoder is a bidirectional LSTM model with an attention mechanism, used to encode the input hybrid feature vectors, extract the dependencies between long-distance named entities in the industrial text data text, and at the same time extract the correlations between characters, between characters and named entities, and between entity character positions in the industrial text data; In the model construction module, the head entity recognition unit includes two identical first binary classifiers, used to label the encoded hybrid feature vectors output by the bidirectional LSTM encoder, and each label is assigned a binary identifier to respectively detect the start position and end position of the entity, and generate multiple entity feature vectors based on the start position and end position of the entity; In the model construction module, the relationship-tail entity recognition unit regards the recognition of relationships and tail entities as a machine reading comprehension task, obtains the description information of the relationship through prior knowledge, concatenates the description information of the relationship and the head entity as the question of the machine reading comprehension task, takes the encoded hybrid feature vectors as the passage of the machine reading comprehension task, embeds it into the pre-trained BERT model in the way of reading comprehension, and identifies the tail entity corresponding to the input relationship description information and head entity through two second binary classifiers; In the pre-trained BERT model, a multi-head self-attention mechanism is used to capture the interaction information between tokens, provide prior knowledge for industrial text data, and at the same time capture context semantic feature information during the training process, so as to eliminate the ambiguity of homophones and express semantic and syntactic patterns; In the relationship-tail entity recognition unit, the second binary classifier outputs multiple start position indices and multiple end position indices for a given context and a specific query, supporting the extraction of all relevant entities according to the query; A joint recognition module: used to input the hybrid feature vectors into the entity relationship joint extraction model to identify all entities and relationships in the industrial text data.
2. The entity relation joint extraction system based on hybrid feature representation according to claim 1, characterized in that, Specifically, the feature extraction module is used for: Extracting character-level feature vectors from industrial text data based on a CNN model, and at the same time using a Chinese word segmenter to segment the industrial text data, matching the segmented words with external dictionary information and external knowledge bases, and obtaining word-level feature vectors through a Word2Vec model.
3. The entity relation joint extraction system based on hybrid feature representation according to claim 1, characterized in that, In the model construction module, the entity type classification unit is used to splice each entity feature vector with the encoded hybrid feature vector as the input, classify the entities through the probability output of Softmax, and set a probability threshold for entity filtering to remove entities and non-entities below the probability threshold, and retain entities greater than or equal to the probability threshold as head entities.
4. An entity relation joint extraction method based on hybrid feature representation, which is implemented based on the entity relation joint extraction system based on hybrid feature representation according to any one of claims 1 to 3, characterized in that, The method includes: S1. Extract character-level feature vectors and word-level feature vectors from industrial text data; S2. Use max pooling operation to fuse the character-level feature vectors and word-level feature vectors to generate hybrid feature vectors; S3. Encode the input hybrid feature vectors through a bidirectional LSTM model with an attention mechanism; S4. Mark the encoded hybrid feature vector h output by the bidirectional LSTM encoder through two identical first binary classifiers, and assign a binary identifier to each mark to respectively detect the start position and the end position of the entity, generating multiple entity feature vectors; N S5. Splice each entity feature vector with the encoded hybrid feature vector, classify the entities through the probability output of Softmax, and set a probability threshold for entity filtering, and retain entities greater than or equal to the probability threshold as head entities; S6. Regard the recognition of relationships and tail entities as a machine reading comprehension task, use a pre-trained BERT model to encode two sentences with the description information of the relationship and the head entity spliced as the question and the encoded hybrid feature vector as the passage, and implement tail entity recognition through two second binary classifiers.
5. An electronic device, characterized in that, It includes: At least one processor, at least one memory, a communication interface, and a bus; Wherein, the processor, the memory, and the communication interface complete mutual communication through the bus; The memory stores program instructions executable by the processor, and the processor calls the program instructions to implement the method as described in claim 4.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to implement the method as described in claim 4.
Citation Information
Patent Citations
Knowledge graph completion method and device, storage medium and electronic equipment
CN112836064A
Entity relationship recognition method and system based on improved graph attention network
CN113010683A
Chinese entity relationship joint extraction method
CN113128229A