A Method and Device for Joint Extraction of Text Entity Relationships in the Field of Retired Electromechanical Products
The segment attention fusion mechanism enhances entity relationship extraction in retired machine-electrical products by addressing data heterogeneity and context utilization, improving accuracy in real-world applications.
Patent Information
- Application Number
- CN202210760882.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-06-30
AI Technical Summary
There is a problem of overlapping entity relationships in the field of retired electromechanical products, and the existing models fail to make full use of context information, resulting in poor results in joint extraction of entity relationships.
The entity relationship joint extraction model based on the segmented attention fusion mechanism is adopted, and the semantic and grammatical features are extracted through the combination of word embedding layer, entity recognition layer and relationship classification layer, and the BERT model and graph attention neural network are used to extract semantic and grammatical features, combined with the pooled attention mechanism and segmented attention fusion mechanism to alleviate the problems of entity nesting and relationship overlap.
It improves the accuracy of joint extraction of entity relationships, effectively alleviates the problems of entity nesting and relationship overlap, and makes full use of text context feature information.
Smart Images

Figure CN115221278B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of joint extraction of text entity relationships in the field of retired electromechanical products, and particularly to a method and device for joint extraction of text entity relationships in the field of retired electromechanical products. Background Art
[0002] In recent years, with the rapid development of the manufacturing industry, the scrapping of a large number of electromechanical products has become a major challenge for the development of China's green economy, and the construction of the reverse logistics network for retired electromechanical products has attracted extensive attention from the country and society. In the reverse logistics process of retired electromechanical products, each link is not isolated, and product design and manufacturing information will have an important impact on the recycling and remanufacturing processes. However, in the entire closed-loop supply chain of production - sales - service - recycling, electromechanical products flow among various enterprises, and there are problems such as large spatio-temporal spans, inconsistent data sources, diverse data structures, and ambiguous data in the full life cycle information of electromechanical products. Recycling institutions often cannot accurately obtain the most valuable information, resulting in the loss of recycling value. Knowledge graphs provide a good solution for the establishment of a unified data model.
[0003] In the process of constructing a knowledge graph for retired electromechanical products, entity recognition and relationship extraction are important steps to obtain structured information. Entity relationship extraction methods are divided into pipeline methods and joint extraction methods. The pipeline method regards entity recognition and relationship extraction as independent tasks and performs them sequentially. The operation is relatively flexible, but it ignores the connection between the two tasks, easily causing information redundancy and leading to the problem of error accumulation. The joint extraction method models the two tasks together and makes good use of the correlation between the two tasks. Existing entity recognition and relationship extraction mostly adopt the joint extraction method.
[0004] Although great progress has been made in current entity relationship joint extraction, there is still much room for improvement in its extraction effect, especially for the field of retired electromechanical products with less application. First of all, the semantic and context information of sentence texts will have varying degrees of influence on both entity recognition and relationship extraction. The current model's way of using context information is relatively single and can be further improved to enhance the entity relationship joint extraction effect. Secondly, there are many entity relationship overlap phenomena in the texts of the retired electromechanical product field, that is, an entity has relationships with multiple entities. If such situations are not considered, it will greatly limit the performance of the entity relationship joint extraction model in the retired electromechanical product field.
[0005] Therefore, how to solve the problem of entity relationship overlap in retired electromechanical products, make full use of context information features, and improve the entity relationship joint extraction effect is a huge challenge faced in the process of joint extraction of text entity relationships in the field of retired electromechanical products. Summary of the Invention
[0006] To solve the problem of overlapping entity relationships in retired electromechanical products in the prior art, an embodiment of the present invention provides a method and device for jointly extracting text entity relationships in the field of retired electromechanical products. The technical solution is as follows:
[0007] On the one hand, a method for jointly extracting text entity relationships in the field of retired electromechanical products is provided. This method is implemented by a device for jointly extracting text entity relationships in the field of retired electromechanical products, and the method includes:
[0008] S1. Obtain historical text data of retired electromechanical products, and preprocess the historical text data of retired electromechanical products to obtain sample data;
[0009] S2. Construct an initial entity relationship joint extraction model based on a segmented attention fusion mechanism;
[0010] S3. Divide the sample data into a training set, a validation set, and a test set. Tune the initial entity relationship joint extraction model based on the validation set, train the tuned initial entity relationship joint extraction model based on the training set, and test the trained initial entity relationship joint extraction model based on the test set to obtain an entity relationship joint extraction model;
[0011] S4. Based on the entity relationship joint extraction model and target text data, perform entity relationship extraction on the target text data, where the target text data is the text data of retired electromechanical products to be extracted.
[0012] Optionally, the entity relationship joint extraction model is divided into three parts: a word embedding layer, an entity recognition layer, and a relationship classification layer;
[0013] The S4 of performing entity relationship extraction on the target text data based on the entity relationship joint extraction model and target text data includes:
[0014] S41. Input the target text data into the word embedding layer to obtain a sentence feature vector that fuses semantic features and syntactic features. Input the target text data and the sentence feature vector into the entity recognition layer, and input the sentence feature vector into the relationship classification layer;
[0015] S42. Based on the entity recognition layer, segment the target text data into text fragments. Based on the sentence feature vector and the text fragments, determine the category of the text fragments. Determine the text fragments identified as entities as entity fragments, and input the entity fragments into the relationship classification layer;
[0016] S43. Combine entity fragments in pairs based on the relationship classification layer to obtain a candidate set of entity pairs. Based on the candidate set of entity pairs and the sentence feature vector, obtain the relationship category between entity fragments, and then complete the entity relationship extraction of the target text data.
[0017] Optionally, the word embedding layer includes a BERT model and a graph attention neural network.
[0018] The step of inputting the target text data into the word embedding layer in S41 to obtain a sentence feature vector that fuses semantic features and syntactic features includes:
[0019] S411. Input the target text data into the BERT model for encoding to obtain the sentence semantic feature vector in the target text data.
[0020] S412. Perform dependency parsing on the target text, and update the feature vector through the graph attention neural network in combination with the sentence semantic feature vector to obtain the sentence syntactic feature vector.
[0021] S413. Concatenate the sentence semantic feature vector and the sentence syntactic feature vector to obtain a sentence feature vector that fuses semantic features and syntactic features.
[0022] Optionally, the entity recognition layer includes a pooling attention mechanism, a fully connected network, and a softmax function.
[0023] The step of determining the category of the text fragment based on the sentence feature vector and the text fragment in the entity recognition layer in S42 includes:
[0024] S421. Based on the sentence feature vector and the text fragment, obtain a fragment feature vector, and perform average pooling operation on the fragment feature vector to obtain a fragment semantic feature vector.
[0025] S422. Input the fragment feature vector and the sentence feature vector into the pooling attention mechanism to determine the fragment context feature vector.
[0026] S423. Concatenate the fragment semantic feature vector, the fragment context feature vector, and the fragment length feature vector to obtain a feature vector.
[0027] S424. Input the feature vector into the fully connected network and the softmax function to determine the probability of the text fragment in each entity category and non-entity, and determine the category corresponding to the maximum probability as the category of the text fragment.
[0028] Optionally, the step of inputting the fragment feature vector and the sentence feature vector into the pooling attention mechanism in S422 to determine the fragment context feature vector includes:
[0029] S4221. Use the fragment feature vector as the query vector, use the sentence feature vector as the key vector and value vector, and obtain the correlation coefficient between each word in the text fragment and each word in the sentence through the product of the query vector and the key vector;
[0030] S4222. Through the max-pooling operation, determine the correlation coefficient between the text fragment and each word in the sentence, and after scaling through the softmax function, determine the correlation coefficient matrix between the text fragment and each word in the sentence;
[0031] S4223. Determine the fragment context feature vector by multiplying the correlation coefficient matrix with the sentence feature vector.
[0032] Optionally, the relationship classification layer in S43 combines entity fragments in pairs to obtain a candidate set of entity pairs, and based on the candidate set of entity pairs and the sentence feature vector, obtains the relationship category between entity fragments, including:
[0033] S431. Based on the relationship classification layer, combine entity fragments in pairs and add them to the candidate set of entity pairs. Any entity pair in the candidate set of entity pairs includes a first entity fragment and a second entity fragment;
[0034] S432. For any entity pair in the candidate set of entity pairs, according to the positions of the first entity fragment and the second entity fragment in the entity pair in the target text data, divide the target text data into five temporary fragments. The five temporary fragments are, in order, the left temporary fragment, the first entity fragment, the middle temporary fragment, the second entity fragment, and the right temporary fragment; according to the five temporary fragments and the sentence feature vector, determine the first entity feature vector corresponding to the first entity fragment, the second entity feature vector corresponding to the second entity fragment, the left temporary text vector corresponding to the left temporary fragment, the middle temporary text vector corresponding to the middle temporary fragment, and the right temporary text vector corresponding to the right temporary fragment;
[0035] S433. Adopt a segmented attention fusion mechanism to perform feature extraction and fusion on the first entity feature vector, the second entity feature vector, the left temporary text vector, the middle temporary text vector, and the right temporary text vector, and obtain the updated feature vectors of the five temporary fragments;
[0036] S434. Input the updated feature vectors of the five temporary fragments into a fully connected network and the softmax function to obtain the probabilities of each entity pair in each relationship category and the non-existence of a relationship, and determine the relationship category between the two entity fragments in the entity pair as the relationship category corresponding to the maximum probability.
[0037] Optionally, in S433, a segmented attention fusion mechanism is adopted to perform feature extraction and fusion on the first entity feature vector, the second entity feature vector, the left temporary text vector, the middle temporary text vector, and the right temporary text vector, and five updated feature vectors of the temporary segments are obtained, including:
[0038] S4331. Perform average pooling operations on the first entity feature vector, the second entity feature vector, the left temporary text vector, the middle temporary text vector, and the right temporary text vector to determine the first entity pooling feature vector corresponding to the first entity feature vector, the second entity pooling feature vector corresponding to the second entity feature vector, the left pooling feature vector corresponding to the left temporary text vector, the middle pooling feature vector corresponding to the middle temporary text vector, and the right pooling feature vector corresponding to the right temporary text vector, and splice the five pooling feature vectors obtained to get a spliced feature vector;
[0039] S4332. Take the spliced feature vector as the query vector, key vector, and value vector, calculate the correlation between every two of the five temporary segments through the product of the query vector and the key vector, and after scaling through the softmax function, multiply it by the value vector to obtain five updated feature vectors of the temporary segments.
[0040] On the other hand, a device for jointly extracting text entity relationships in the field of retired electromechanical products is provided. This device is applied to the method for jointly extracting text entity relationships in the field of retired electromechanical products, and this device includes:
[0041] An acquisition module, configured to acquire historical text data of retired electromechanical products, and preprocess the historical text data of retired electromechanical products to obtain sample data;
[0042] A construction module, configured to construct an initial entity relationship joint extraction model based on a segmented attention fusion mechanism;
[0043] A training module, configured to divide the sample data into a training set, a validation set, and a test set, optimize the initial entity relationship joint extraction model based on the validation set, train the optimized initial entity relationship joint extraction model based on the training set, and test the trained initial entity relationship joint extraction model based on the test set to obtain an entity relationship joint extraction model;
[0044] An extraction module, configured to perform entity relationship extraction on the target text data based on the entity relationship joint extraction model and the target text data, where the target text data is text data of retired electromechanical products to be extracted.
[0045] Optionally, the entity relationship joint extraction model is divided into three parts: a word embedding layer, an entity recognition layer, and a relationship classification layer;
[0046] The extraction module includes a word embedding sub-module, an entity recognition sub-module, and a relationship classification sub-module; where:
[0047] The word embedding sub-module is used to input the target text data into a word embedding layer to obtain a sentence feature vector that fuses semantic features and syntactic features, input the target text data and the sentence feature vector into an entity recognition layer, and input the sentence feature vector into a relationship classification layer;
[0048] The entity recognition sub-module is used to segment the target text data into text fragments based on the entity recognition layer, determine the category of the text fragments based on the sentence feature vector and the text fragments, determine the text fragments recognized as entities as entity fragments, and input the entity fragments into the relationship classification layer;
[0049] The relationship classification sub-module is used to combine entity fragments in pairs based on the relationship classification layer to obtain a candidate set of entity pairs, and obtain the relationship category between entity fragments based on the candidate set of entity pairs and the sentence feature vector, thereby completing the entity relationship extraction of the target text data.
[0050] Optionally, the word embedding layer includes a BERT model and a graph attention neural network;
[0051] The word embedding sub-module is further used for:
[0052] S411: Input the target text data into the BERT model for encoding to obtain the sentence semantic feature vector in the target text data;
[0053] S412: Perform dependency parsing on the target text, and update the feature vector through the graph attention neural network in combination with the sentence semantic feature vector to obtain the sentence syntactic feature vector;
[0054] S413: Concatenate the sentence semantic feature vector and the sentence syntactic feature vector to obtain a sentence feature vector that fuses semantic features and syntactic features.
[0055] Optionally, the entity recognition layer includes a pooling attention mechanism, a fully connected network, and a softmax function;
[0056] The entity recognition sub-module is further used for:
[0057] S421: Obtain a fragment feature vector based on the sentence feature vector and the text fragment, perform an average pooling operation on the fragment feature vector to obtain a fragment semantic feature vector;
[0058] S422. Input the fragment feature vector and the sentence feature vector into a pooling attention mechanism to determine a fragment context feature vector;
[0059] S423. Concatenate the fragment semantic feature vector, the fragment context feature vector, and the fragment length feature vector to obtain a feature vector;
[0060] S424. Input the feature vector into a fully connected network and a softmax function to determine the probabilities of the text fragment for each entity category and non-entity, and determine the category corresponding to the maximum probability as the category of the text fragment.
[0061] Optionally, the entity recognition sub-module is further configured to:
[0062] S4221. Use the fragment feature vector as a query vector, the sentence feature vector as a key vector and a value vector, and obtain the correlation coefficients between each word in the text fragment and each word in the sentence through the product of the query vector and the key vector;
[0063] S4222. Determine the correlation coefficients between the text fragment and each word in the sentence through a max-pooling operation, and scale them through a softmax function to determine the correlation coefficient matrix between the text fragment and each word in the sentence;
[0064] S4223. Determine the fragment context feature vector as the product of the correlation coefficient matrix and the sentence feature vector.
[0065] Optionally, the relationship classification sub-module is further configured to:
[0066] S431. Based on a relationship classification layer, combine entity fragments in pairs and add them to an entity pair candidate set, where any entity pair in the entity pair candidate set includes a first entity fragment and a second entity fragment;
[0067] S432. For any entity pair in the entity pair candidate set, according to the positions of the first entity fragment and the second entity fragment in the entity pair in the target text data, divide the target text data into five temporary fragments. The five temporary fragments are, in order, a left temporary fragment, a first entity fragment, a middle temporary fragment, a second entity fragment, and a right temporary fragment; according to the five temporary fragments and the sentence feature vector, determine a first entity feature vector corresponding to the first entity fragment, a second entity feature vector corresponding to the second entity fragment, a left temporary text vector corresponding to the left temporary fragment, a middle temporary text vector corresponding to the middle temporary fragment, and a right temporary text vector corresponding to the right temporary fragment;
[0068] S433. Adopt a segmented attention fusion mechanism to perform feature extraction and fusion on the first entity feature vector, the second entity feature vector, the left temporary text vector, the middle temporary text vector, and the right temporary text vector to obtain the updated feature vectors of the five temporary segments;
[0069] S434. Input the updated feature vectors of the five temporary segments into a fully connected network and a softmax function to obtain the probabilities of each entity pair in each relationship category and the absence of a relationship, and determine the relationship category between the two entity segments in the entity pair as the relationship category corresponding to the maximum probability.
[0070] Optionally, the relationship classification sub-module is further configured to:
[0071] S4331. Perform average pooling operations on the first entity feature vector, the second entity feature vector, the left temporary text vector, the middle temporary text vector, and the right temporary text vector to determine the first entity pooling feature vector corresponding to the first entity feature vector, the second entity pooling feature vector corresponding to the second entity feature vector, the left pooling feature vector corresponding to the left temporary text vector, the middle pooling feature vector corresponding to the middle temporary text vector, and the right pooling feature vector corresponding to the right temporary text vector, and splice the five pooling feature vectors obtained to obtain a spliced feature vector;
[0072] S4332. Use the spliced feature vector as the query vector, key vector, and value vector, calculate the correlation between every two of the five temporary segments through the product of the query vector and the key vector, and after scaling through the softmax function, multiply it by the value vector to obtain the updated feature vectors of the five temporary segments.
[0073] On the other hand, an electronic device is provided. The electronic device includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned method for jointly extracting text entity relationships in the field of retired electromechanical products.
[0074] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned method for jointly extracting text entity relationships in the field of retired electromechanical products.
[0075] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include:
[0076] In the embodiments of the present invention, aiming at the problems of text data entity recognition and relationship extraction in the field of retired electromechanical products, a joint entity relationship extraction model for the field of retired electromechanical products based on a segmented attention fusion mechanism is established. Considering the problem of error accumulation in the pipeline method, a joint extraction method is adopted for entity recognition and relationship extraction. There are problems of entity nesting and relationship overlap in the text data of retired electromechanical products, especially the phenomenon that one entity corresponds to multiple entities with relationships. In the entity recognition layer of this model, a fragment arrangement method is used for modeling, and in the relationship classification layer, all entities are combined in pairs to judge the relationships between entities, effectively alleviating the problems of entity nesting and relationship overlap. Secondly, a pooling attention mechanism is designed in the entity recognition layer of the model to extract fragment context feature information, while a segmented attention fusion mechanism is designed in the relationship classification layer to extract and fuse the text context feature information of the two entities, making full use of the text context features and effectively improving the accuracy of joint entity relationship extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0078] Figure 1 is a flowchart of a method for jointly extracting text entity relationships in the field of retired electromechanical products provided by an embodiment of the present invention;
[0079] Figure 2 is a flowchart of a method for constructing and training a joint entity relationship extraction model provided by an embodiment of the present invention;
[0080] Figure 3 is a flowchart of a method for extracting entity relationships based on a joint entity relationship extraction model provided by an embodiment of the present invention;
[0081] Figure 4 is a schematic diagram of the framework of a joint entity relationship extraction model based on a segmented attention fusion mechanism provided by an embodiment of the present invention;
[0082] Figure 5 is a block diagram of a device for jointly extracting text entity relationships in the field of retired electromechanical products provided by an embodiment of the present invention;
[0083] Figure 6 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0084] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0085] An embodiment of the present invention provides a method for jointly extracting text entity relationships in the field of retired electromechanical products. This method can be implemented by an electronic device, which can be a terminal or a server, such as Figure 1 As shown in the flowchart of the method for jointly extracting text entity relationships in the field of retired electromechanical products, the processing flow of this method can include the following steps:
[0086] S11. Obtain historical text data of retired electromechanical products, preprocess the historical text data of retired electromechanical products to obtain sample data.
[0087] S12. Construct an initial entity relationship joint extraction model based on a segmented attention fusion mechanism.
[0088] S13. Divide the sample data into a training set, a validation set, and a test set. Tune the initial entity relationship joint extraction model based on the validation set, train the tuned initial entity relationship joint extraction model based on the training set, and test the trained initial entity relationship joint extraction model based on the test set to obtain an entity relationship joint extraction model.
[0089] S14. Based on the entity relationship joint extraction model and the target text data, perform entity relationship extraction on the target text data, where the target text data is the text data of retired electromechanical products to be extracted.
[0090] Optionally, the entity relationship joint extraction model is divided into three parts: a word embedding layer, an entity recognition layer, and a relationship classification layer.
[0091] The performing entity relationship extraction on the target text data based on the entity relationship joint extraction model and the target text data in S14 includes:
[0092] S141. Input the target text data into the word embedding layer to obtain a sentence feature vector that fuses semantic features and syntactic features. Input the target text data and the sentence feature vector into the entity recognition layer, and input the sentence feature vector into the relationship classification layer.
[0093] S142. Based on the entity recognition layer, segment the target text data into text fragments. Based on the sentence feature vector and the text fragments, determine the category of the text fragments. Determine the text fragments identified as entities as entity fragments, and input the entity fragments into the relationship classification layer.
[0094] S143. Based on the relationship classification layer, combine entity fragments in pairs to obtain a candidate set of entity pairs. Based on the candidate set of entity pairs and the sentence feature vector, obtain the relationship category between entity fragments, and then complete the entity relationship extraction of the target text data.
[0095] Optionally, the word embedding layer includes a BERT model and a graph attention neural network.
[0096] Inputting the target text data into the word embedding layer in S141 to obtain a sentence feature vector that fuses semantic features and syntactic features includes:
[0097] S1411. Input the target text data into the BERT model for encoding to obtain the sentence semantic feature vector in the target text data.
[0098] S1412. Perform dependency analysis on the target text, and update the feature vector through the graph attention neural network in combination with the sentence semantic feature vector to obtain the sentence syntactic feature vector.
[0099] S1413. Concatenate the sentence semantic feature vector and the sentence syntactic feature vector to obtain a sentence feature vector that fuses semantic features and syntactic features.
[0100] Optionally, the entity recognition layer includes a pooling attention mechanism, a fully connected network, and a softmax function.
[0101] Determining the category of the text fragment based on the sentence feature vector and the text fragment in S142 includes:
[0102] S1421. Based on the sentence feature vector and the text fragment, obtain a fragment feature vector, and perform average pooling operation on the fragment feature vector to obtain a fragment semantic feature vector.
[0103] S1422. Input the fragment feature vector and the sentence feature vector into the pooling attention mechanism to determine the fragment context feature vector.
[0104] S1423. Concatenate the fragment semantic feature vector, the fragment context feature vector, and the fragment length feature vector to obtain a feature vector.
[0105] S1424. Input the feature vector into the fully connected network and the softmax function to determine the probability of the text fragment in each entity category and non-entity, and determine the category corresponding to the maximum probability as the category of the text fragment.
[0106] Optionally, inputting the fragment feature vector and the sentence feature vector into the pooling attention mechanism in S1422 to determine the fragment context feature vector includes:
[0107] S14221. Use the fragment feature vector as the query vector, the sentence feature vector as the key vector and value vector, and obtain the correlation coefficient between each word in the text fragment and each word in the sentence through the product of the query vector and the key vector.
[0108] S14222. Through the max-pooling operation, determine the correlation coefficient between the text fragment and each word in the sentence, and after scaling through the softmax function, determine the correlation coefficient matrix between the text fragment and each word in the sentence.
[0109] S14223. Determine the fragment context feature vector as the product of the correlation coefficient matrix and the sentence feature vector.
[0110] Optionally, in S143, the relationship classification layer combines entity fragments in pairs to obtain a candidate set of entity pairs, and based on the candidate set of entity pairs and the sentence feature vector, obtains the relationship category between entity fragments, including:
[0111] S1431. Based on the relationship classification layer, combine entity fragments in pairs and add them to the candidate set of entity pairs. Any entity pair in the candidate set of entity pairs includes a first entity fragment and a second entity fragment.
[0112] S1432. For any entity pair in the candidate set of entity pairs, according to the positions of the first entity fragment and the second entity fragment in the entity pair in the target text data, divide the target text data into five temporary fragments. The five temporary fragments are, in order, the left temporary fragment, the first entity fragment, the middle temporary fragment, the second entity fragment, and the right temporary fragment. According to the five temporary fragments and the sentence feature vector, determine the first entity feature vector corresponding to the first entity fragment, the second entity feature vector corresponding to the second entity fragment, the left temporary text vector corresponding to the left temporary fragment, the middle temporary text vector corresponding to the middle temporary fragment, and the right temporary text vector corresponding to the right temporary fragment.
[0113] S1433. Adopt a segmented attention fusion mechanism to perform feature extraction and fusion on the first entity feature vector, the second entity feature vector, the left temporary text vector, the middle temporary text vector, and the right temporary text vector to obtain the updated feature vectors of the five temporary fragments.
[0114] S1434. Input the updated feature vectors of the five temporary fragments into a fully connected network and the softmax function to obtain the probabilities of each entity pair in each relationship category and the absence of a relationship, and determine the relationship category between the two entity fragments in the entity pair as the relationship category corresponding to the maximum probability.
[0115] Optionally, in S1433, a segmented attention fusion mechanism is adopted to perform feature extraction and fusion on the first entity feature vector, the second entity feature vector, the left temporary text vector, the middle temporary text vector, and the right temporary text vector, obtaining the updated feature vectors of the five temporary segments, including:
[0116] S14331. Perform average pooling operations on the first entity feature vector, the second entity feature vector, the left temporary text vector, the middle temporary text vector, and the right temporary text vector to determine the first entity pooling feature vector corresponding to the first entity feature vector, the second entity pooling feature vector corresponding to the second entity feature vector, the left pooling feature vector corresponding to the left temporary text vector, the middle pooling feature vector corresponding to the middle temporary text vector, and the right pooling feature vector corresponding to the right temporary text vector, and splice the obtained five pooling feature vectors to obtain a spliced feature vector.
[0117] S14332. Use the spliced feature vector as the query vector, key vector, and value vector, calculate the correlation between every two of the five temporary segments through the product of the query vector and the key vector, and after scaling by the softmax function, multiply it with the value vector to obtain the updated feature vectors of the five temporary segments.
[0118] In the embodiments of the present invention, aiming at the problems of entity recognition and relationship extraction in the text data of the field of retired electromechanical products, a joint entity relationship extraction model for the field of retired electromechanical products based on a segmented attention fusion mechanism is established. Considering the problem of error accumulation in the pipeline method, a joint extraction method is adopted for entity recognition and relationship extraction. There are problems of entity nesting and relationship overlap in the text data of retired electromechanical products, especially the phenomenon that one entity corresponds to multiple entities with relationships. In the entity recognition layer of this model, a fragment arrangement method is adopted for modeling, and in the relationship classification layer, all entities are pairwise combined to judge the relationships between entities, effectively alleviating the problems of entity nesting and relationship overlap. Secondly, a pooling attention mechanism is designed in the entity recognition layer of the model to extract fragment context feature information, and a segmented attention fusion mechanism is designed in the relationship classification layer to extract and fuse the text context feature information of the two entities, making full use of the text context features and effectively improving the accuracy of joint entity relationship extraction.
[0119] The embodiments of the present invention provide a method for constructing and training a joint entity relationship extraction model, which can be implemented by an electronic device. The electronic device can be a terminal or a server, such as Figure 2 shown in the flowchart of the method for constructing and training a joint entity relationship extraction model. The processing flow of this method can include the following steps:
[0120] S21. Obtain historical text data of retired electromechanical products, and preprocess the historical text data of retired electromechanical products to obtain sample data.
[0121] In a feasible implementation, the text data of retired electromechanical products is mostly in the form of documents. For document data, data cleaning is required to remove the noise in the document data. According to the characteristics of the field of retired electromechanical products, entity and relationship annotation are performed on the text data of retired electromechanical products to construct a data set, which is the sample data. The specific process can be divided into the following steps S221 - S223:
[0122] S211. The obtained text data of retired electromechanical products needs to be cleaned to remove noise data.
[0123] S212. According to the characteristics of the text field of retired electromechanical products, define the entity category as L = {l1, l2, ……, l p}, the number of entity categories is p, define the relationship category as R = {r1, r2, ……, r q}, and the number of relationship categories is q. Annotate the entities and relationships in the text data of retired electromechanical products according to the entity category and relationship category.
[0124] S213. Split the document into sentences, construct a data set, and the maximum length of the sentence is 512.
[0125] S22. Construct an initial entity - relationship joint extraction model based on the segment - level attention fusion mechanism.
[0126] In a feasible implementation, the initial entity - relationship joint extraction model is divided into three parts: a word embedding layer, an entity recognition layer, and a relationship classification layer. Among them, the word embedding layer is used to encode the input text data of retired electromechanical products to obtain word vectors (also called sentence feature vectors in the present invention) that fuse semantic features and syntactic features, and then input the word vectors into the entity recognition layer. The entity recognition layer is used to segment the text of retired electromechanical products to obtain multiple text segments, and use a segment classifier to judge the segment category of the text segments by combining the semantic features, segment length features, and the context features of the sentences where the segments are located. When it is determined that the segment category of the text segment is an entity, then determine that the text segment is an entity segment, and combine the entity segments in pairs to form entity pairs. The relationship classification layer classifies the relationship between two entities based on the segment - level attention fusion mechanism in combination with the entity pairs and their context features. The loss function of this initial entity - relationship joint extraction model is the sum of the entity recognition loss function and the relationship classification.
[0127] Its specific construction process can be divided into the following steps S221 - S224:
[0128] S221. Given the sentence X = {x1, x2, ……, x n} in the data set of retired electromechanical products, x iLet the i-th word in the sentence be \(w_i\), and \(n\) be the length of the sentence. The word embedding layer encodes the input sentence \(X\) into word vectors \(E\), which integrate the semantic and syntactic features of the sentence.
[0129] S222. The entity recognition layer segments the text \(X\) of retired electromechanical products to obtain segments \(s = \{x i , x i+1 , \cdots, x i+k \}\), where the length of the segment is \(k + 1\). It extracts and fuses the semantic, length, and context features of the segments and concatenates them to obtain the feature vector \(x s \). The feature vector \(x s \) passes through the segment classifier to obtain the probabilities on each entity category and non-entity. The category corresponding to the maximum probability is the category of the segment.
[0130] S223. The relation classification layer combines entity segments in pairs and adds them to the candidate set of entity pairs. For entity segments \(s1\) and \(s2\), it extracts and fuses the context features of the entity segments based on the segmental attention fusion mechanism to obtain the feature vector \(x r \). After passing through the relation classifier, it obtains the probabilities of the two entity segments on each relation category and the absence of a relation. The relation category corresponding to the maximum probability is the relation category between the two entity segments.
[0131] S224. The entity recognition layer uses the multi-class cross-entropy loss function to obtain the loss function \(L NER \), and the relation classification layer uses the multi-class cross-entropy loss function to obtain the loss function \(L RE \). The sum of the entity recognition loss function \(L NER \) and the relation classification loss function \(L RE \) gives the loss function \(L\) of the joint model.
[0132] S23. Divide the sample data into a training set, a validation set, and a test set. Tune the initial entity-relation joint extraction model based on the validation set, train the tuned initial entity-relation joint extraction model based on the training set, and test the trained initial entity-relation joint extraction model based on the test set to obtain the entity-relation joint extraction model.
[0133] In a feasible implementation, the training process can be divided into the following steps S231 - S235:
[0134] S231. Process the text dataset of retired electromechanical products into the input format of the model, and then divide the text dataset of retired electromechanical products into a training set, a validation set, and a test set according to the ratio of 6:2:2. If the dataset is large, the proportions of the validation set and the test set can be appropriately reduced.
[0135] S232. Initialize the model hyperparameters, set the initial learning rate learning_rate, data batch size batch_size, and number of training epochs epoch. The data batch size refers to the number of input data each time.
[0136] S233. Input the validation set into the model, and adjust the hyperparameters in combination with the change of the loss function during the iteration process to obtain a set of relatively optimal hyperparameter combinations.
[0137] In a feasible implementation, the loss function of the entity recognition layer adopts the multi-class cross-entropy loss function, which is represented by the following formula (1):
[0138]
[0139] where p is the number of entity label categories, is the probability that the segment s is the i-th label, and y i is the one-hot vector distribution of the entity labels in the sample.
[0140] The loss function of the relationship classification layer adopts the multi-class cross-entropy loss function, which is represented by the following formula (2):
[0141]
[0142] where q is the number of relationship categories, is the probability that the relationship between two segments is the j-th category, and y j is the one-hot vector distribution of the relationship category labels in the sample.
[0143] The loss function of the entity relationship joint extraction model is the sum of the two-task loss functions, which is represented by the following formula (3):
[0144] L = L NER + L RE (3)
[0145] S234. Let the hyperparameters of the model be equal to the adjusted hyperparameter values, input the training set into the model, and save the model after training.
[0146] S235. Call the saved model, input the test set into the model, calculate the accuracy, recall, and F1 metric values of entity recognition according to the output data, calculate the accuracy, recall, and F1 metric values of relationship extraction according to the output data, and evaluate the model performance.
[0147] In a feasible implementation, step S235 may further specifically include the following steps S2351 - S2353:
[0148] S2351. Call the saved model, input the test set into the model to obtain output data, process the data to obtain the entities contained in each sentence of the retired electromechanical product text data and the relationships between the entities.
[0149] S2352. Calculate the accuracy Precision of entity recognition according to the output data NER , recall Recall NER and F1 metric value F1 NER , evaluate the entity recognition effect of the model, which is represented by the following formulas (4), (5), and (6):
[0150]
[0151]
[0152]
[0153] Among them, TP NER is the number of entities predicted correctly among the entities predicted by entity recognition, FP NER is the number of entities predicted incorrectly among the entities predicted by entity recognition, FN NER is the number of entities not predicted by entity recognition.
[0154] S2353. Calculate the accuracy Precision of relation extraction according to the output data RE , recall Recall RE and F1 metric value F1 RE , evaluate the relation extraction effect of the model, which is represented by the following formulas (7), (8), and (9):
[0155]
[0156]
[0157]
[0158] Among them, TP RE is the number of relations predicted correctly among the relations predicted by relation classification. Here, a correct relation means not only the correct relation category, but also the correct two entities and their entity categories. FN RE is the number of relations predicted incorrectly among the relations predicted by relation classification, FN RE is the number of relations not predicted by relation classification.
[0159] In the embodiments of the present invention, for the problem of text data entity recognition and relationship extraction in the field of retired electromechanical products, a joint entity relationship extraction model for the field of retired electromechanical products based on a segmented attention fusion mechanism is established. Considering the problem of error accumulation in the pipeline method, a joint extraction method is adopted for entity recognition and relationship extraction. There are problems of entity nesting and relationship overlap in the text data of retired electromechanical products, especially the phenomenon that one entity corresponds to multiple entities with relationships. In the entity recognition layer of this model, a fragment arrangement method is used for modeling, and in the relationship classification layer, all entities are combined in pairs to judge the relationships between entities, effectively alleviating the problems of entity nesting and relationship overlap. Secondly, a pooling attention mechanism is designed in the entity recognition layer of the model to extract fragment context feature information, and a segmented attention fusion mechanism is designed in the relationship classification layer to extract and fuse the text context feature information of the two entities, making full use of the text context features and effectively improving the accuracy of joint entity relationship extraction.
[0160] Embodiments of the present invention provide a method for entity relationship extraction based on a joint entity relationship extraction model. This method can be implemented by an electronic device, which can be a terminal or a server. In the embodiments of the present invention, the joint entity relationship extraction model is divided into three parts: a word embedding layer, an entity recognition layer, and a relationship classification layer, as Figure 3 shown in a flowchart of a method for entity relationship extraction based on a joint entity relationship extraction model, as Figure 4 shown in a framework schematic diagram of a joint entity relationship extraction model based on a segmented attention fusion mechanism. The processing flow of this method can include the following steps:
[0161] S31. Obtain target text data.
[0162] In a feasible implementation, obtain the sentence X = {x1, x2, ……, x n} in the retired electromechanical product dataset, where x i is the i-th word in the sentence, n is the length of the sentence, and use the sentence X as the target text data. Define the entity relationship triple (Entity1, Relation, Entity2), where Entity1 and Entity2 are entities extracted from the sentence X, and Relation is the relationship between the entity pair Entity1 and Entity2. The purpose of the joint entity relationship extraction task is to identify all entity relationship triples from the text.
[0163] S32. Input the target text data into the word embedding layer to obtain a sentence feature vector that fuses semantic and syntactic features. Input the target text data and the sentence feature vector into the entity recognition layer, and input the sentence feature vector into the relationship classification layer.
[0164] Among them, the word embedding layer includes a BERT model and a graph attention neural network.
[0165] In a feasible implementation, step S32 may specifically include the following steps S321 - S323:
[0166] S321. Input the target text data into the BERT model for encoding to obtain the sentence semantic feature vectors in the target text data.
[0167] In a feasible implementation, input sentence X into the BERT model to obtain the semantic feature vector H of the text = {h1, h2, ……, h n}.
[0168] S322. Perform dependency parsing on the target text, and update the feature vectors through the graph attention neural network in combination with the sentence semantic feature vectors to obtain the sentence syntactic feature vectors.
[0169] In a feasible implementation, generate the dependency relationship matrix A between each word in the text of retired electromechanical products through a dependency parser n×n , and input the dependency relationship matrix A n×n and the text semantic feature vector H into the graph attention neural network to obtain the text syntactic feature vector G of the retired electromechanical products = {g1, g2, ……, g n}.
[0170] S323. Concatenate the sentence semantic feature vectors and the sentence syntactic feature vectors to obtain the sentence feature vectors that fuse semantic and syntactic features.
[0171] In a feasible implementation, concatenate the semantic feature vectors h i and the syntactic feature vectors g i of each word in the text sequence of retired electromechanical products to obtain the final output feature vector e i of the word, which is represented by the following formula (10):
[0172] e i = [g i ; h i (10)
[0173] The output feature vectors of the text sequence of retired electromechanical products in the word embedding layer are E = {e1, e2, ……, e n}.
[0174] S33. Based on the entity recognition layer, segment the target text data into text fragments, determine the categories of the text fragments based on the sentence feature vectors and the text fragments, determine the text fragments recognized as entities as entity fragments, and input the entity fragments into the relationship classification layer.
[0175] Among them, the entity recognition layer includes a pooling attention mechanism, a fully connected network, and a softmax function.
[0176] In a feasible implementation manner, step S33 may specifically include the following steps S331-S334:
[0177] S331. Based on the sentence feature vector and the text fragment, obtain the fragment feature vector, and perform an average pooling operation on the fragment feature vector to obtain the fragment semantic feature vector.
[0178] In a feasible implementation manner, the text is segmented to obtain all possible fragments of lengths 1, 2, ……, and it is judged whether the category of the fragment s = {x i , x i+1 , ……, x i+k}, and the fragment length is k + 1.
[0179] The semantic feature vector span of the fragment s is obtained by performing an average pooling operation on the fragment vector e(s) = {e i , e i+1 , ……, e i+k} output by the corresponding word embedding layer, and is represented by the following formula (11):
[0180] span = Avgpooling(e i , e i+1 , ……, e i+k ) (11)
[0181] S332. Input the fragment feature vector and the sentence feature vector into the pooling attention mechanism to determine the fragment context feature vector.
[0182] In a feasible implementation manner, the fragment vector e(s) and the sentence feature vector E are used to obtain the fragment context feature vector c s .
[0183] Optionally, step S332 may include the following steps S3321-S3323:
[0184] S3321. Use the fragment feature vector as the query vector, use the sentence feature vector as the key vector and the value vector, and obtain the correlation coefficient between each word in the text fragment and each word in the sentence through the product of the query vector and the key vector.
[0185] In a feasible implementation manner, use the fragment semantic feature vector e(s) as the query vector of the attention mechanism, use the sentence semantic feature vector E as the key vector and the value vector, and are represented by the following formulas (12)(13)(14):
[0186]
[0187]
[0188]
[0189] Among them, Q s has a dimension of (k + 1)·d k , K s has a dimension of n·d k , V s has a dimension of n·d V .
[0190] S3322. Determine the correlation coefficient between the text segment and each word in the sentence through a max pooling operation, and scale it through the softmax function to determine the correlation coefficient matrix between the text segment and each word in the sentence.
[0191] In a feasible implementation, the weight coefficient is obtained through the max pooling operation of the product of the query vector and the key vector, and the attention weight att of the segment for each word in the sentence is obtained by scaling through the softmax function, which is represented by the following formula (15):
[0192]
[0193] S3323. Determine the segment context feature vector as the product of the correlation coefficient matrix and the sentence feature vector.
[0194] In a feasible implementation, the product of the weight vector att and the value vector gives the segment context feature vector c s , which is represented by the following formula (16):
[0195] c s = att·V s (16)
[0196] S333. Concatenate the segment semantic feature vector, the segment context feature vector, and the segment length feature vector to obtain a feature vector.
[0197] In a feasible implementation, the corresponding segment length feature vector w is found from the given embedding matrix W k+1 , which contains vectors corresponding to segments of lengths 1, 2, ……, and the embedding matrix W can be learned through backpropagation.
[0198] Concatenate the segment semantic feature vector span, the segment context feature vector c s and the segment length feature vector w k+1 of the three-part feature vectors to obtain the final feature vector x of the retired electromechanical product text segment s .
[0199] S334. Input the feature vector into the fully connected network and the softmax function to determine the probabilities of the text segment for each entity category and non-entity, and determine the category of the text segment as the category corresponding to the maximum probability.
[0200] In a feasible implementation, this model uses the fully connected network as the segment classifier, and inputs the feature vector x s into the fully connected network and the softmax function to obtain the probabilities of this segment mapped to each entity category and non-entity The category corresponding to the maximum probability among them is the category of the text segment s of this retired electromechanical product, which is represented by the following formula (17):
[0201]
[0202] S34. Based on the relationship classification layer, combine entity segments in pairs to obtain a candidate set of entity pairs. Based on the candidate set of entity pairs and the sentence feature vector, obtain the relationship category between entity segments, and then complete the entity relationship extraction of the target text data.
[0203] In a feasible implementation, step S34 can be divided into the following steps S341 - S344:
[0204] S341. Based on the relationship classification layer, combine entity segments in pairs and add them to the candidate set of entity pairs. Any entity pair in the candidate set of entity pairs includes a first entity segment and a second entity segment.
[0205] In a feasible implementation, for the segments classified as entities, exhaustively combine all pairs of entity segments in the text of the retired electromechanical product and add them to the candidate set of entity pairs. For the entity pairs s1 and s2 (which can be respectively called the first entity segment and the second entity segment) in the retired electromechanical product text S, judge the relationship between entities.
[0206] S342. For any entity pair in the candidate set of entity pairs, according to the positions of the first entity segment and the second entity segment in the entity pair in the target text data, divide the target text data into five temporary segments. The five temporary segments are, in order, the left temporary segment, the first entity segment, the middle temporary segment, the second entity segment, and the right temporary segment. According to the five temporary segments and the sentence feature vector, determine the first entity feature vector corresponding to the first entity segment, the second entity feature vector corresponding to the second entity segment, the left temporary text vector corresponding to the left temporary segment, the middle temporary text vector corresponding to the middle temporary segment, and the right temporary text vector corresponding to the right temporary segment.
[0207] In a feasible implementation manner, the input text of retired electromechanical products is segmented into five segments of text according to the positions of the entities. Combining with the text feature vector E of the retired electromechanical products output by the word embedding layer, the feature vector e(s1) of entity 1, the feature vector e(s2) of entity 2, and the feature vector f of the text on the left side of the two entities are obtained. left , the feature vector f of the text between the two entities middle , the feature vector f of the text on the right side of the two entities right , but when the two entities overlap, the feature vector of the text between the two entities is the text feature vector of the overlapping part of the entities.
[0208] S343. Adopt a segmented attention fusion mechanism to perform feature extraction and fusion on the first entity feature vector, the second entity feature vector, the left temporary text vector, the middle temporary text vector, and the right temporary text vector, and obtain the updated feature vectors of the five temporary segments.
[0209] In a feasible implementation manner, step S343 can be specifically divided into the following steps S3431 - S3432:
[0210] S3431. Perform average pooling operations on the first entity feature vector, the second entity feature vector, the left temporary text vector, the middle temporary text vector, and the right temporary text vector to determine the first entity pooling feature vector corresponding to the first entity feature vector, the second entity pooling feature vector corresponding to the second entity feature vector, the left pooling feature vector corresponding to the left temporary text vector, the middle pooling feature vector corresponding to the middle temporary text vector, and the right pooling feature vector corresponding to the right temporary text vector, and splice the five pooling feature vectors obtained to obtain a spliced feature vector.
[0211] In a feasible implementation manner, since the lengths of the five segments of text are not fixed, this model uses average pooling operations to extract and fuse the internal word features of the five segments of text and then splices them to obtain a feature vector T, which is represented by the following formulas (18)(19)(20)(21)(22)(23):
[0212] t left = Avgpooling(f left ) (18)
[0213] t s1 = Avgpooling(e(s1)) (19)
[0214] t middle = Avgpooling(f middle ) (20)
[0215] t s2 = Avgpooling(e(s2)) (21)
[0216] t right = Avgpooling(f right ) (22)
[0217] T = [t left ; t s1 ; t middle ; t s2 ; t right (23)
[0218] S3432. Use the concatenated feature vector as the query vector, key vector, and value vector. Calculate the correlation between every two of the five temporary segments by multiplying the query vector and the key vector. After scaling by the softmax function, multiply it with the value vector to obtain the updated feature vectors of the five temporary segments.
[0219] In a feasible implementation, the self-attention mechanism is adopted to extract and fuse features among five texts. This model generates a query vector Q, a key vector K, and a value vector V based on the feature vector T, which are represented by the following equations (24), (25), and (26):
[0220] Q = TW Q (24)
[0221] K = TW K (25)
[0222] V = TW V (26)
[0223] Among them, the dimension of Q is 5 × d k , the dimension of K is 5 × d k , and the dimension of V is 5 × d V .
[0224] Obtain the weight coefficients between every two of the five texts by multiplying the query vector and the key vector, which represents the correlation between the five texts. Then scale the weight coefficients by the softmax function and multiply it with the value vector V to obtain the fused feature vector x r , which is represented by the following equation (27):
[0225]
[0226] S344. Input the updated feature vectors of the five temporary segments into the fully connected network and the softmax function to obtain the probabilities of each entity pair in each relationship category and the absence of a relationship. Determine the relationship category between the two entity segments in the entity pair as the relationship category corresponding to the maximum probability.
[0227] In a feasible implementation manner, a fully connected network is used as a relation classifier, and the fused feature vector x r is input into the fully connected network and the softmax function to obtain the probabilities of mapping the entity pair to each relation category and the non-existence of the relation The relation category corresponding to the maximum probability among them is the relation category between entity pairs in the text of retired electromechanical products, which is represented by the following formula (28):
[0228]
[0229] In the embodiment of the present invention, aiming at the problems of entity recognition and relation extraction of text data in the field of retired electromechanical products, a joint entity relation extraction model for the field of retired electromechanical products based on a segmented attention fusion mechanism is established. Considering the problem of error accumulation in the pipeline method, a joint extraction method is adopted for entity recognition and relation extraction. There are problems of entity nesting and relation overlap in the text data of retired electromechanical products, especially the phenomenon that one entity corresponds to multiple entity existence relations. In the entity recognition layer of this model, a fragment arrangement method is adopted for modeling, and in the relation classification layer, all entities are combined in pairs to judge the relations between entities, effectively alleviating the problems of entity nesting and relation overlap. Secondly, a pooling attention mechanism is designed in the entity recognition layer of the model to extract fragment context feature information, and a segmented attention fusion mechanism is designed in the relation classification layer to extract and fuse the text context feature information of the two entities, making full use of the text context features and effectively improving the accuracy of joint entity relation extraction.
[0230] Figure 5 is a block diagram 500 of a device for jointly extracting text entity relations in the field of retired electromechanical products shown according to an exemplary embodiment. Referring to Figure 5 , the device includes:
[0231] An acquisition module 510, configured to acquire historical text data of retired electromechanical products, and preprocess the historical text data of retired electromechanical products to obtain sample data;
[0232] A construction module 520, configured to construct an initial joint entity relation extraction model based on a segmented attention fusion mechanism;
[0233] A training module 530, configured to divide the sample data into a training set, a validation set, and a test set, optimize the initial joint entity relation extraction model based on the validation set, train the optimized initial joint entity relation extraction model based on the training set, and test the trained initial joint entity relation extraction model based on the test set to obtain a joint entity relation extraction model;
[0234] An extraction module 540 is configured to perform entity-relationship extraction on the target text data based on the entity-relationship joint extraction model and the target text data, where the target text data is text data of retired electromechanical products to be extracted.
[0235] Optionally, the entity-relationship joint extraction model is divided into three parts: a word embedding layer, an entity recognition layer, and a relationship classification layer;
[0236] The extraction module 540 includes a word embedding sub-module 5401, an entity recognition sub-module 5402, and a relationship classification sub-module 5403; where:
[0237] The word embedding sub-module 5401 is configured to input the target text data into the word embedding layer to obtain a sentence feature vector that fuses semantic features and syntactic features, input the target text data and the sentence feature vector into the entity recognition layer, and input the sentence feature vector into the relationship classification layer;
[0238] The entity recognition sub-module 5402 is configured to segment the target text data into text fragments based on the entity recognition layer, determine the category of the text fragments based on the sentence feature vector and the text fragments, determine the text fragments recognized as entities as entity fragments, and input the entity fragments into the relationship classification layer;
[0239] The relationship classification sub-module 5403 is configured to combine entity fragments pairwise based on the relationship classification layer to obtain a set of entity pair candidates, and obtain the relationship category between entity fragments based on the set of entity pair candidates and the sentence feature vector, thereby completing the entity-relationship extraction of the target text data.
[0240] Optionally, the word embedding layer includes a BERT model and a graph attention neural network;
[0241] The word embedding sub-module 5401 is further configured to:
[0242] S411. Input the target text data into the BERT model for encoding to obtain the sentence semantic feature vector in the target text data;
[0243] S412. Perform dependency analysis on the target text, and update the feature vector through the graph attention neural network in combination with the sentence semantic feature vector to obtain the sentence syntactic feature vector;
[0244] S413. Concatenate the sentence semantic feature vector and the sentence syntactic feature vector to obtain a sentence feature vector that fuses semantic features and syntactic features.
[0245] Optionally, the entity recognition layer includes a pooling attention mechanism, a fully connected network, and a softmax function;
[0246] The entity recognition sub-module 5402 is further configured to:
[0247] S421: Obtain a segment feature vector based on the sentence feature vector and the text segment, perform an average pooling operation on the segment feature vector to obtain a segment semantic feature vector;
[0248] S422: Input the segment feature vector and the sentence feature vector into a pooling attention mechanism to determine a segment context feature vector;
[0249] S423: Concatenate the segment semantic feature vector, the segment context feature vector, and the segment length feature vector to obtain a feature vector;
[0250] S424: Input the feature vector into a fully connected network and a softmax function to determine the probabilities of the text segment for each entity category and non-entity, and determine the category corresponding to the maximum probability as the category of the text segment.
[0251] Optionally, the entity recognition sub-module 5402 is further configured to:
[0252] S4221: Use the segment feature vector as a query vector, use the sentence feature vector as a key vector and a value vector, and obtain the correlation coefficient between each word in the text segment and each word in the sentence through the product of the query vector and the key vector;
[0253] S4222: Determine the correlation coefficient between the text segment and each word in the sentence through a max pooling operation, and scale it through a softmax function to determine the correlation coefficient matrix between the text segment and each word in the sentence;
[0254] S4223: Determine the product of the correlation coefficient matrix and the sentence feature vector as the segment context feature vector.
[0255] Optionally, the relationship classification sub-module 5403 is further configured to:
[0256] S431: Combine entity segments in pairs based on the relationship classification layer and add them to the entity pair candidate set, where any entity pair in the entity pair candidate set includes a first entity segment and a second entity segment;
[0257] S432. For any entity pair in the candidate set of entity pairs, divide the target text data into five temporary segments according to the positions of the first entity segment and the second entity segment in the entity pair in the target text data. The five temporary segments are, in order, the left temporary segment, the first entity segment, the middle temporary segment, the second entity segment, and the right temporary segment; determine the first entity feature vector corresponding to the first entity segment, the second entity feature vector corresponding to the second entity segment, the left temporary text vector corresponding to the left temporary segment, the middle temporary text vector corresponding to the middle temporary segment, and the right temporary text vector corresponding to the right temporary segment according to the five temporary segments and the sentence feature vector;
[0258] S433. Adopt a segmented attention fusion mechanism to perform feature extraction and fusion on the first entity feature vector, the second entity feature vector, the left temporary text vector, the middle temporary text vector, and the right temporary text vector to obtain the updated feature vectors of the five temporary segments;
[0259] S434. Input the updated feature vectors of the five temporary segments into a fully connected network and a softmax function to obtain the probabilities of each entity pair in each relationship category and the absence of a relationship, and determine the relationship category between the two entity segments in the entity pair as the relationship category corresponding to the maximum probability.
[0260] Optionally, the relationship classification sub-module 5403 is further configured to:
[0261] S4331. Perform average pooling operations on the first entity feature vector, the second entity feature vector, the left temporary text vector, the middle temporary text vector, and the right temporary text vector to determine the first entity pooling feature vector corresponding to the first entity feature vector, the second entity pooling feature vector corresponding to the second entity feature vector, the left pooling feature vector corresponding to the left temporary text vector, the middle pooling feature vector corresponding to the middle temporary text vector, and the right pooling feature vector corresponding to the right temporary text vector, and splice the five pooled feature vectors obtained to get a spliced feature vector;
[0262] S4332. Use the spliced feature vector as the query vector, key vector, and value vector, calculate the correlation between every two of the five temporary segments through the product of the query vector and the key vector, and after scaling by the softmax function, multiply by the value vector to obtain the updated feature vectors of the five temporary segments.
[0263] In the embodiments of the present invention, aiming at the problems of text data entity recognition and relationship extraction in the field of retired electromechanical products, a joint entity relationship extraction model for the field of retired electromechanical products based on a segmented attention fusion mechanism is established. Considering the problem of error accumulation in the pipeline method, a joint extraction method is adopted for entity recognition and relationship extraction. There are problems of entity nesting and relationship overlap in the text data of retired electromechanical products, especially the phenomenon that one entity corresponds to multiple entities with relationships. In the entity recognition layer of this model, a fragment arrangement method is used for modeling, and in the relationship classification layer, all entities are combined in pairs to judge the relationships between entities, effectively alleviating the problems of entity nesting and relationship overlap. Secondly, a pooling attention mechanism is designed in the entity recognition layer of the model to extract fragment context feature information, and a segmented attention fusion mechanism is designed in the relationship classification layer to extract and fuse the text context feature information of the two entities, making full use of the text context features and effectively improving the accuracy of joint entity relationship extraction.
[0264] Figure 6 FIG. 600 is a schematic structural diagram of an electronic device 600 provided by an embodiment of the present invention. The electronic device 600 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 601 and one or more memories 602. Among them, at least one instruction is stored in the memory 602, and at least one instruction is loaded and executed by the processor 601 to implement the steps of the above-mentioned method for jointly extracting text entity relationships in the field of retired electromechanical products.
[0265] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions, and the above instructions can be executed by a processor in a terminal to complete the above-mentioned method for jointly extracting text entity relationships in the field of retired electromechanical products. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0266] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0267] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for jointly extracting text entity relationships in the field of retired electromechanical products, characterized in that, The method includes: S1. Obtain the text data of historical retired electromechanical products, preprocess the text data of historical retired electromechanical products to obtain sample data; S2. Construct an initial entity relationship joint extraction model based on a segmented attention fusion mechanism; S3. Divide the sample data into a training set, a validation set, and a test set, optimize the initial entity relationship joint extraction model based on the validation set, train the optimized initial entity relationship joint extraction model based on the training set, and test the trained initial entity relationship joint extraction model based on the test set to obtain an entity relationship joint extraction model; S4. Based on the entity relationship joint extraction model and the target text data, perform entity relationship extraction on the target text data, where the target text data is the text data of retired electromechanical products to be extracted; Among them, the entity relationship joint extraction model is divided into three parts: a word embedding layer, an entity recognition layer, and a relationship classification layer; The S4's performing entity relationship extraction on the target text data based on the entity relationship joint extraction model and the target text data includes: S41. Input the target text data into the word embedding layer to obtain a sentence feature vector that fuses semantic features and syntactic features, input the target text data and the sentence feature vector into the entity recognition layer, and input the sentence feature vector into the relationship classification layer; S42. Based on the entity recognition layer, divide the target text data into text segments, determine the category of the text segments based on the sentence feature vector and the text segments, determine the text segments identified as entities as entity segments, and input the entity segments into the relationship classification layer; S43. Based on the relationship classification layer, combine entity segments in pairs to obtain a set of entity pair candidates, and obtain the relationship category between entity segments based on the set of entity pair candidates and the sentence feature vector, thereby completing the entity relationship extraction of the target text data; Among them, the S43's combining entity segments in pairs based on the relationship classification layer to obtain a set of entity pair candidates, and obtaining the relationship category between entity segments based on the set of entity pair candidates and the sentence feature vector includes: S431. Based on the relationship classification layer, combine entity segments in pairs and add them to the set of entity pair candidates, and any entity pair in the set of entity pair candidates includes a first entity segment and a second entity segment; S432. For any entity pair in the candidate set of entity pairs, according to the positions of the first entity segment and the second entity segment in the entity pair in the target text data, the target text data is segmented into five temporary segments, which are, in order, the left temporary segment, the first entity segment, the middle temporary segment, the second entity segment, and the right temporary segment; according to the five temporary segments and the sentence feature vector, determine the first entity feature vector corresponding to the first entity segment, the second entity feature vector corresponding to the second entity segment, the left temporary text vector corresponding to the left temporary segment, the middle temporary text vector corresponding to the middle temporary segment, and the right temporary text vector corresponding to the right temporary segment; S433. Adopt a segmented attention fusion mechanism to perform feature extraction and fusion on the first entity feature vector, the second entity feature vector, the left temporary text vector, the middle temporary text vector, and the right temporary text vector to obtain the updated feature vectors of the five temporary segments; S434. Input the updated feature vectors of the five temporary segments into a fully connected network and a softmax function to obtain the probabilities of each entity pair in each relationship category and the absence of a relationship, and determine the relationship category between the two entity segments in the entity pair as the relationship category corresponding to the maximum probability.
2. The method according to claim 1, wherein The word embedding layer includes a BERT model and a graph attention neural network; The step in S41 of inputting the target text data into the word embedding layer to obtain a sentence feature vector that fuses semantic features and syntactic features includes: S411. Input the target text data into the BERT model for encoding to obtain the sentence semantic feature vector in the target text data; S412. Perform dependency parsing on the target text, and update the feature vector through the graph attention neural network in combination with the sentence semantic feature vector to obtain the sentence syntactic feature vector; S413. Concatenate the sentence semantic feature vector and the sentence syntactic feature vector to obtain a sentence feature vector that fuses semantic features and syntactic features.
3. The method according to claim 1, characterized in that The entity recognition layer includes a pooling attention mechanism, a fully connected network, and a softmax function; The step in S42 of determining the category of the text segment based on the sentence feature vector and the text segment includes: S421. Based on the sentence feature vector and the text segment, obtain a segment feature vector, and perform average pooling operation on the segment feature vector to obtain a segment semantic feature vector; S422. Input the segment feature vector and the sentence feature vector into the pooling attention mechanism to determine the segment context feature vector; S423. Concatenate the segment semantic feature vector, the segment context feature vector, and the segment length feature vector to obtain a feature vector; S424. Input the feature vector into a fully connected network and a softmax function to determine the probabilities of the text segment in each entity category and non-entity, and determine the category corresponding to the maximum probability as the category of the text segment.
4. The method according to claim 3, characterized in that, Inputting the fragment feature vector and the sentence feature vector into the pooling attention mechanism in S422 to determine the fragment context feature vector includes: S4221: Using the fragment feature vector as the query vector, the sentence feature vector as the key vector and the value vector, and obtaining the correlation coefficient between each word in the text fragment and each word in the sentence through the product of the query vector and the key vector; S4222: Determining the correlation coefficient between the text fragment and each word in the sentence through max pooling operation, scaling it through the softmax function, and determining the correlation coefficient matrix between the text fragment and each word in the sentence; S4223: Determining the fragment context feature vector by multiplying the correlation coefficient matrix with the sentence feature vector.
5. The method according to claim 1, characterized in that Adopting the segmented attention fusion mechanism in S433 to perform feature extraction and fusion on the first entity feature vector, the second entity feature vector, the left temporary text vector, the middle temporary text vector, and the right temporary text vector to obtain five updated feature vectors of the temporary fragments, including: S4331: Performing average pooling operation on the first entity feature vector, the second entity feature vector, the left temporary text vector, the middle temporary text vector, and the right temporary text vector to determine the first entity pooling feature vector corresponding to the first entity feature vector, the second entity pooling feature vector corresponding to the second entity feature vector, the left pooling feature vector corresponding to the left temporary text vector, the middle pooling feature vector corresponding to the middle temporary text vector, and the right pooling feature vector corresponding to the right temporary text vector, and splicing the five pooling feature vectors obtained to get the spliced feature vector; S4332: Using the spliced feature vector as the query vector, the key vector, and the value vector, calculating the correlation between the five temporary fragments pairwise through the product of the query vector and the key vector, and after scaling through the softmax function, multiplying it with the value vector to obtain five updated feature vectors of the temporary fragments.
6. An apparatus for jointly extracting text entity relationships in the field of retired electromechanical products, characterized in that, The text entity relationship joint extraction device in the field of retired electromechanical products is used to implement the text entity relationship joint extraction method in the field of retired electromechanical products. The device includes: An acquisition module, configured to acquire historical retired electromechanical product text data, and preprocess the historical retired electromechanical product text data to obtain sample data; A construction module, configured to construct an initial entity relationship joint extraction model based on the segmented attention fusion mechanism; A training module, configured to divide the sample data into a training set, a validation set, and a test set, optimize the initial entity relationship joint extraction model based on the validation set, train the optimized initial entity relationship joint extraction model based on the training set, and test the trained initial entity relationship joint extraction model based on the test set to obtain the entity relationship joint extraction model; An extraction module, configured to perform entity relationship extraction on the target text data based on the entity relationship joint extraction model and the target text data, where the target text data is the retired electromechanical product text data to be extracted; Among them, the entity relationship joint extraction model is divided into three parts: a word embedding layer, an entity recognition layer, and a relationship classification layer; The extraction module includes a word embedding sub-module, an entity recognition sub-module, and a relationship classification sub-module; among them: The word embedding sub-module is used to input the target text data into the word embedding layer to obtain a sentence feature vector that fuses semantic features and syntactic features, input the target text data and the sentence feature vector into the entity recognition layer, and input the sentence feature vector into the relationship classification layer; The entity recognition sub-module is used to segment the target text data into text segments based on the entity recognition layer, determine the category of the text segment based on the sentence feature vector and the text segment, determine the text segment recognized as an entity as an entity segment, and input the entity segment into the relationship classification layer; the relationship classification sub-module is used to combine entity segments in pairs based on the relationship classification layer to obtain a candidate set of entity pairs, and obtain the relationship category between entity segments based on the candidate set of entity pairs and the sentence feature vector, thereby completing the entity relationship extraction of the target text data; Among them, the relationship classification sub-module 5403 is further used for: S431. Combine entity segments in pairs based on the relationship classification layer and add them to the candidate set of entity pairs. Any entity pair in the candidate set of entity pairs includes a first entity segment and a second entity segment; S432. For any entity pair in the candidate set of entity pairs, divide the target text data into five temporary segments according to the positions of the first entity segment and the second entity segment in the target text data. The five temporary segments are, in order, a left temporary segment, a first entity segment, a middle temporary segment, a second entity segment, and a right temporary segment; determine the first entity feature vector corresponding to the first entity segment, the second entity feature vector corresponding to the second entity segment, the left temporary text vector corresponding to the left temporary segment, the middle temporary text vector corresponding to the middle temporary segment, and the right temporary text vector corresponding to the right temporary segment according to the five temporary segments and the sentence feature vector; S433. Adopt a segmented attention fusion mechanism to perform feature extraction and fusion on the first entity feature vector, the second entity feature vector, the left temporary text vector, the middle temporary text vector, and the right temporary text vector to obtain the updated feature vectors of the five temporary segments; S434. Input the updated feature vectors of the five temporary segments into a fully connected network and a softmax function to obtain the probability of each entity pair in each relationship category and the absence of a relationship, and determine the relationship category corresponding to the maximum probability as the relationship category between the two entity segments in the entity pair.
7. The device according to claim 6, characterized in that The word embedding layer includes a BERT model and a graph attention neural network; The word embedding sub-module is further used for: S411. Input the target text data into the BERT model for encoding to obtain the sentence semantic feature vector in the target text data; S412. Perform dependency parsing on the target text, and update the feature vector by combining the sentence semantic feature vector through a graph attention neural network to obtain a sentence syntactic feature vector; S413. Concatenate the sentence semantic feature vector and the sentence syntactic feature vector to obtain a sentence feature vector that fuses semantic and syntactic features.
Citation Information
Patent Citations
Entity relationship joint extraction method and device, equipment and medium
CN111666427A
Food safety relation extraction method based on BERT and improved PCNN
CN113821571A