Text entity joint relation extraction model training method and device and computer readable storage medium
By using multi-head attention mechanism and two-way annotation methods in the text entity joint relationship extraction model, dynamically fusion and processing features in the entity recognition and relationship extraction tasks, the problem of limited representation and learning ability of existing models is solved, and better model performance is achieved.
Patent Information
- Application Number
- CN202510590579.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The existing text entity joint relationship extraction model fails to fully explore the potential correlation between entity recognition tasks and relationship extraction tasks during the training process, resulting in limited representation and learning ability of the model.
The multi-head attention mechanism is used to dynamically fuse the original feature representation of each word element in the entity recognition task and the relationship extraction task, generate the first fusion feature and the second fusion feature, and entity recognition and relationship allocation of the recognition method and the double affine network structure based on the baseline model through the bidirectional annotated entity, and finally build the total loss function to optimize the model.
By enhancing the modeling ability of the model to model interactions between subtasks, the model's representation and learning ability is improved, and the generated text entity joint relationship extraction model shows good performance indicators.
Smart Images

Figure CN120218072A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of computer text recognition, and in particular, to a training method, device, and computer-readable storage medium for a text entity joint relationship extraction model. Background Art
[0002] A text entity joint relationship extraction system is a natural language processing (NLP) system that aims to identify entities (entity recognition task) from text and determine the relationships between entities (relationship extraction task). It is usually applied to text information extraction, knowledge graph construction, etc. Currently, text entity joint relationship extraction systems usually adopt a text entity joint relationship extraction model obtained through training.
[0003] There is a conventional text entity joint relationship extraction model which is a joint relationship extraction model based on a bidirectional extraction framework. Its main training process includes: using a pre-trained language model to encode the text to extract context-related word vector representations; based on a bidirectional annotation-based entity pair recognition method, combined with the fused feature representations, entity recognition is performed in the head entity-tail entity and tail entity-head entity directions respectively to extract all entity pairs with potential relationships from the input text; using a bi-affine module to assign relationships to all entity pairs to assign corresponding relationship types; and separately constructing loss functions for entity recognition and relationship extraction, and adopting a joint training method to obtain a training model by minimizing the loss functions.
[0004] However, the inventors found in specific implementations that although the above method uses two parallel networks to perform entity recognition tasks and relationship extraction tasks respectively, the two tasks interact through shared input features and each encodes corresponding specific representations. However, this method does not fully explore the potential correlation between the two tasks, resulting in limited representation and learning capabilities of the trained model. Summary of the Invention
[0005] The technical problem to be solved by the embodiments of the present invention is to provide a training method for a text entity joint relationship extraction model, and the trained model can effectively improve the representation and learning capabilities.
[0006] The further technical problem to be solved by the embodiments of the present invention is to provide a training device for a text entity joint relationship extraction model, and the trained model can effectively improve the representation and learning capabilities.
[0007] The further technical problem to be solved by the embodiments of the present invention is to provide a computer-readable storage medium for storing a computer program whose trained model can effectively improve the representation and learning capabilities.
[0008] To solve the above technical problems, the embodiments of the present invention provide the following technical solutions: A training method for a text entity joint relationship extraction model, including the following steps: Encode the input text to respectively obtain the first original feature representation of each token included in the input text in the entity recognition task and the second original feature representation in the relationship extraction task; Based on the multi-head attention mechanism, dynamically fuse the first original feature representation and the second original feature representation of each token during the entity recognition task processing to obtain the first fusion feature corresponding to each token, and fuse the first original feature representation and the second original feature representation of each token during the relationship extraction task processing to obtain the second fusion feature corresponding to each token; Based on the entity pair recognition method with bidirectional annotation, combine the first fusion feature and the second fusion feature to perform entity recognition from two complementary directions, so as to extract all entity pairs with potential relationships from the input text; Adopt a bi-affine network structure based on a baseline model to assign relationships to each pair of the entity pairs, so as to calculate and obtain the relationship probability corresponding to the entity pairs; and Combine the relationship probabilities to respectively construct the first loss function for entity recognition and the second loss function for relationship assignment, combine the first loss function and the second loss function to obtain the total loss function, and minimize the total loss function to output the text entity joint relationship extraction model.
[0009] Further, the encoding of the input text to respectively obtain the first original feature representation of each token included in the input text in the entity recognition task and the second original feature representation in the relationship extraction task specifically includes: Adopt a pre-trained language model as the encoder to extract the word vectors of the input text; and Generate the first original feature representation based on the word vectors of the input text and the first trainable projection matrix for the entity recognition task, and generate the second original feature representation based on the word vectors of the input text and the second trainable projection matrix for the relationship extraction task.
[0010] Further, the dynamic fusion of the first original feature representation and the second original feature representation of each token based on the multi-head attention mechanism to respectively obtain the corresponding fusion feature representation specifically includes: Based on the multi-layer perceptron network model, the first attention score and the second attention score corresponding to the entity recognition feature and the relation extraction feature of each token in the entity recognition task are extracted by combining the first original feature representation and the second original feature representation. Based on the multi-layer perceptron network model, the third attention score and the fourth attention score corresponding to the entity recognition feature and the relation extraction feature of each token in the relation extraction task are extracted by combining the first original feature representation and the second original feature representation; Normalize the first attention score and the second attention score to obtain the first normalized weight matrix for the entity recognition task, and normalize the third attention score and the fourth attention score to obtain the second normalized weight matrix for the relation extraction task; Based on the first normalized weight matrix, the first original feature representation and the second original feature representation are weighted and fused to obtain a single entity recognition fusion feature corresponding to the entity recognition task in a single subspace. Based on the second normalized weight matrix, the first original feature representation and the second original feature representation are weighted and fused to obtain a single relation extraction fusion feature in a single subspace; and Based on the first trainable weight matrix, the outputs of all attention heads are weighted and averaged by combining the single entity recognition fusion feature, and the Dropout operation is applied to obtain the first fusion feature in the entity recognition task. Based on the second trainable weight matrix, the outputs of all attention heads are weighted and averaged by combining the single relation extraction fusion feature, and the Dropout operation is applied to obtain the second fusion feature in the relation extraction task.
[0011] Further, the entity pair recognition method based on bidirectional annotation combines the first fusion feature and the second fusion feature to perform entity recognition from two complementary directions, so as to extract all entity pairs with potential relationships from the input text, specifically including: Multiply the first fusion feature by the first trainable parameter matrix and the second trainable parameter matrix respectively to obtain the first context feature representation for head entity extraction and the second context feature representation for tail entity extraction; Using the first head entity annotator, based on the pointer network method, combining the first context feature representation for the head entity, the third trainable parameter matrix and the fourth trainable parameter matrix, calculate the first head entity start probability and the first head entity end probability at the start position of each token in the input text as the head entity. Using the first tail entity annotator, based on the pointer network method, combining the second context feature representation for the tail entity, the fifth trainable parameter matrix and the sixth trainable parameter matrix, calculate the first tail entity start probability and the first tail entity end probability at the start position of each token in the input text as the tail entity; Determine the feature representation of the head entity in the input text based on the first head entity start probability and the first head entity end probability of each token, and determine the feature representation of each tail entity in the input text based on the first tail entity start probability and the first tail entity end probability of each token; Use the second context feature for the tail entity as the input feature, and the feature representation of the head entity as the conditional function to input the first CLN network model to calculate and generate the first text encoding representation that fuses the head entity features. Use the first context feature for the head entity as the input feature, and the feature representation of the tail entity as the conditional function to input the second CLN network model to calculate and generate the second text encoding representation that fuses the tail entity features; Use the second tail entity annotator to calculate and obtain the second tail entity start probability at the start position of each token as the tail entity and the second tail entity end probability at the end position based on the pointer network method in combination with the first text encoding representation, the fifth trainable parameter matrix, and the sixth trainable parameter matrix. Use the second head entity annotator to calculate and obtain the second head entity start probability at the start position of each token as the head entity and the second head entity end probability at the end position based on the pointer network method in combination with the second text encoding representation, the third trainable parameter matrix, and the fourth trainable parameter matrix; and Extract the head entity and the tail entity with potential relationships as a pair of entity pairs based on the first head entity start probability, the first head entity end probability, the second tail entity start probability, and the second tail entity end probability of each token in the input text. Extract the head entity and the tail entity with potential relationships as a pair of entity pairs based on the first tail entity start probability, the first tail entity end probability, the second head entity start probability, and the second head entity end probability of each token in the input text.
[0012] Further, the step of using the dual-affine network structure based on the baseline model to assign relationships to each pair of the entity pairs to calculate and obtain the relationship probability corresponding to the entity pairs specifically includes: Use the dual-affine network structure based on the baseline model to calculate the feature representations of the head entity and the tail entity in the entity pair respectively; and Calculate the relationship probability corresponding to the entity pair based on the feature representations of the head entity and the tail entity in the entity pair and the trainable relationship parameter matrix.
[0013] Further, the first loss function is expressed as: ; The second loss function is expressed as: ; Wherein, ; p represents the relationship probability, and t represents the true label. It represents the position marker indicating the start or end of the head entity or the tail entity in the entity pair. They respectively represent the first tail entity tagger, the first head entity tagger, the second tail entity tagger, and the second head entity tagger. g represents the number of tokens included in the input text, and G is a predefined set of relationships.
[0014] Furthermore, the total loss function is expressed as: , where α and β are function weights for flexible adjustment according to the importance of the current dataset task.
[0015] Furthermore, the following parameters are corrected to minimize the total loss function: the pre-trained language model, the first trainable projection matrix, the second trainable projection matrix, the first trainable weight matrix, the second trainable weight matrix, the first trainable parameter matrix, the second trainable parameter matrix, the third trainable parameter matrix, the fourth trainable parameter matrix, the fifth trainable parameter matrix, the sixth trainable parameter matrix, the first CLN network model, the second CLN network model, and the trainable relationship parameter matrix.
[0016] On the other hand, to solve the above further technical problems, an embodiment of the present invention further provides the following technical solution: A training device for a text entity joint relationship extraction model, the device includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the training method of the text entity joint relationship extraction model as described in any one of the above.
[0017] On yet another hand, to solve the above further technical problems, an embodiment of the present invention further provides the following technical solution: A computer-readable storage medium, the computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the training method of the text entity joint relationship extraction model as described in any one of the above.
[0018] After adopting the above technical solution, the embodiments of the present invention have at least the following beneficial effects: After encoding the input text in the embodiments of the present invention to obtain the first original feature representation of each token in the entity recognition task and the second original feature representation in the relation extraction task, the original feature representations of each token in the entity recognition task and the relation extraction task are dynamically fused based on the multi-head attention mechanism to respectively obtain the corresponding first fusion feature and the second fusion feature, capturing the fine-grained feature interaction between entities and relations in multiple subspaces, dynamically adjusting the contribution of each subspace feature, thereby enhancing the model's ability to model the interaction between subtasks, and improving the model's representation and learning ability; further, after extracting all entity pairs with potential relations from the input text based on the entity pair recognition method with bidirectional annotation, a bi-affine network structure based on the baseline model is used to assign relations to each pair of the entity pairs, and finally, the first loss function for entity recognition and the second loss function for the relation assignment are constructed by combining the relation probabilities. The final text entity joint relation extraction model generated by minimizing the combined first loss function and the second loss function has good performance indicators. Description of the Drawings
[0019] Figure 1 It is a flowchart of the steps of an optional embodiment of the training method of the text entity joint relation extraction model of the present invention.
[0020] Figure 2 It is a schematic diagram of feature fusion in step S2 of an optional embodiment of the training method of the text entity joint relation extraction model of the present invention.
[0021] Figure 3 It is a schematic diagram of the CLN network model of an optional embodiment of the training method of the text entity joint relation extraction model of the present invention.
[0022] Figure 4 It is a principle block diagram of an optional embodiment of the training device of the text entity joint relation extraction model of the present invention.
[0023] Figure 5 It is a functional module diagram of an optional embodiment of the training device of the text entity joint relation extraction model of the present invention. Detailed Description of the Embodiments
[0024] The following further describes the present application in detail with reference to the drawings and specific embodiments. It should be understood that the following illustrative embodiments and descriptions are only used to explain the present invention and are not intended to limit the present invention. Moreover, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0025] As Figure 1As shown in the figure, an alternative embodiment of the present invention provides a training method for a text entity joint relationship extraction model, including the following steps: S1: Encode the input text to respectively obtain the first original feature representation of each token included in the input text in the entity recognition task and the second original feature representation in the relationship extraction task; S2: Dynamically fuse the first original feature representation and the second original feature representation of each token during the entity recognition task processing based on the multi-head attention mechanism to obtain the first fused feature corresponding to each token, and fuse the first original feature representation and the second original feature representation of each token during the relationship extraction task processing to obtain the second fused feature corresponding to each token; S3: Based on the entity pair recognition method with bidirectional annotation, combine the first fused feature and the second fused feature to perform entity recognition from two complementary directions, so as to extract all entity pairs with potential relationships from the input text; S4: Use a bi-affine network structure based on a baseline model to assign relationships to each pair of the entity pairs, so as to calculate and obtain the relationship probability corresponding to the entity pairs; and S5: Combine the relationship probabilities to respectively construct the first loss function for entity recognition and the second loss function for relationship assignment, combine the first loss function and the second loss function to obtain the total loss function, and minimize the total loss function to output the text entity joint relationship extraction model.
[0026] After the embodiments of the present invention encode the input text to obtain the first original feature representation of each token in the entity recognition task and the second original feature representation in the relationship extraction task, the original feature representations of each token in the entity recognition task and the relationship extraction task are dynamically fused based on the multi-head attention mechanism to respectively obtain the corresponding first fused feature and second fused feature, capturing the fine-grained feature interaction between entities and relationships in multiple subspaces, dynamically adjusting the contribution of each subspace feature, thereby enhancing the model's ability to model the interaction between subtasks, and improving the model's representation and learning ability; further, after extracting all entity pairs with potential relationships from the input text based on the entity pair recognition method with bidirectional annotation, use a bi-affine network structure based on a baseline model to assign relationships to each pair of the entity pairs, and finally combine the relationship probabilities to construct the first loss function for entity recognition and the second loss function for relationship assignment. By minimizing the total loss function obtained by combining the first loss function and the second loss function, the finally generated text entity joint relationship extraction model has good performance indicators.
[0027] In specific implementation, in step S3, it is specifically a bidirectional entity pair recognition method based on conditional layer normalization that combines the first fusion feature and the second fusion feature to perform entity recognition from two complementary directions respectively, so as to extract all entity pairs with potential relationships from the input text.
[0028] In an alternative embodiment of the present invention, step S1 specifically includes: Using a pre-trained language model as an encoder to extract word vectors of the input text, where the input text is represented as , where x CLS and x SEP respectively represent the start token and the segmentation token of the input text, and the word vectors are represented as: , where , b represents the batch size, l represents the sequence length of the input text, and d represents the hidden layer dimension; and Generating the first original feature representation based on the word vectors of the input text and the first trainable projection matrix for the entity recognition task, and generating the second original feature representation based on the word vectors of the input text and the second trainable projection matrix for the relation extraction task, where The first original feature representation is: , The second original feature representation is: , and respectively represent the first trainable projection matrix and the second trainable projection matrix, and are the corresponding bias terms, and e and r respectively represent the reference annotations for the entity recognition task and the relation extraction task.
[0029] In this embodiment, a pre-trained language model is used as an encoder to extract word vectors of the input text, and then the original feature representations corresponding to the entity recognition task and the relation extraction task are generated, with high encoding efficiency. In specific implementation, the selection of the pre-trained language model can be determined according to the dataset of the corresponding task, so as to improve the compatibility of the model. For example: when it is a Chinese dataset, the pre-trained language model can select the RoBERTa-wwm-ext model.
[0030] In an alternative embodiment of the present invention, step S2 specifically includes: S21: Based on the multi-layer perceptron network model, combine the first original feature representation and the second original feature representation to extract the first attention score and the second attention score corresponding to the entity recognition feature and the relation extraction feature of each token in the entity recognition task respectively, and based on the multi-layer perceptron network model, combine the first original feature representation and the second original feature representation to extract the third attention score and the fourth attention score corresponding to the entity recognition feature and the relation extraction feature of each token in the relation extraction task; Among them, the first attention score and the second attention score of the i-th token are respectively expressed as: , , the third attention score and the fourth attention score of the i-th token are respectively expressed as: , ; S22: Normalize the first attention score and the second attention score to obtain the first normalized weight matrix for the entity recognition task, and normalize the third attention score and the fourth attention score to obtain the second normalized weight matrix for the relation extraction task; Among them, the first normalized weight matrix and the second normalized weight matrix are respectively expressed as: , ; S23: Based on the first normalized weight matrix, weighted fuse the first original feature representation and the second original feature representation to obtain the single entity recognition fusion feature corresponding to the entity recognition task under a single subspace, and based on the second normalized weight matrix, weighted fuse the first original feature representation and the second original feature representation to obtain the single relation extraction fusion feature under a single subspace; Among them, the single entity recognition fusion feature and the single relation extraction fusion feature are respectively expressed as: , , Among them, […, 0:1] and […, 1:2] represent slicing operations, used to extract the first column or the second column from the last dimension of the corresponding normalized weight matrix, and k represents the k-th attention head; S24: Based on the first trainable weight matrix, perform weighted average on the outputs of all attention heads in combination with the single entity recognition fusion feature and apply the Dropout operation to obtain the first fusion feature in the entity recognition task, and based on the second trainable weight matrix, perform weighted average on the outputs of all attention heads in combination with the single relation extraction fusion feature and apply the Dropout operation to obtain the second fusion feature in the relation extraction task; Among them, the first trainable weight matrix and the second trainable weight matrix are respectively expressed as: , , The initial values of the first trainable weight matrix and the second trainable weight matrix are all - one matrices; The first fused feature is: , The second fused feature is: , Among them, , , and respectively represent the weight ratios corresponding to the k - th attention head, and H represents the total number of attention heads.
[0031] In this embodiment, based on the multi - head attention mechanism, fine - grained feature interactions between entities and relationships are captured in multiple sub - spaces. Through the weighted fusion method, the contributions of features in each sub - space are dynamically adjusted, thereby enhancing the model's ability to model interactions between subtasks and improving the model's representation and learning ability.
[0032] In steps S21 - S24, as Figure 2 shown, two symmetric and independent feature fusion units (Feature Fusion) can be used for calculation to obtain the fused features corresponding to the entity recognition task and the relation extraction task respectively. Among them, in step S24, to improve the generalization ability of the model and reduce overfitting, after respectively performing weighted averaging on the outputs of all attention heads for the single entity recognition fused feature and the single relation extraction fused feature, a Dropout operation needs to be applied.
[0033] In an alternative embodiment of the present invention, step S3 specifically includes: S31: Multiply the first fused feature by the first trainable parameter matrix and the second trainable parameter matrix respectively to obtain a first context feature representation for head entity extraction and a second context feature representation for tail entity extraction; Among them, the first context feature representation and the second context feature representation are respectively expressed as: , , W s and W o respectively represent the first trainable parameter matrix and the second trainable parameter matrix, b is the corresponding bias term, and s and o are the referential annotations for the head entity and the tail entity respectively; S32: Use the first head entity annotator to calculate the first head entity start probability and the first head entity end probability of each token in the input text as the start position of the head entity and the end position of the head entity respectively based on the pointer network method combined with the first context feature representation for the head entity, the third trainable parameter matrix, and the fourth trainable parameter matrix. Use the first tail entity annotator to calculate the first tail entity start probability and the first tail entity end probability of each token in the input text as the start position of the tail entity and the end position of the tail entity respectively based on the pointer network method combined with the second context feature representation for the tail entity, the fifth trainable parameter matrix, and the sixth trainable parameter matrix; Among them, the first head entity start probability and the first head entity end probability of the i-th token are respectively expressed as: , , and respectively represent the third trainable parameter matrix and the fourth trainable parameter matrix. The first tail entity start probability and the first tail entity end probability of the i-th token are respectively expressed as: , , and respectively represent the fifth trainable parameter matrix and the sixth trainable parameter matrix. σ is the sigmoid activation function, sta represents the reference annotation of the start position of the head entity or the tail entity, and end represents the reference annotation of the end position of the head entity or the tail entity; S33: Determine the feature representation of the head entity in the input text based on the first head entity start probability and the first head entity end probability of each token, and determine the feature representation of each tail entity in the input text based on the first tail entity start probability and the first tail entity end probability of each token; Among them, the feature representation of the head entity in the input text is , and the feature representation of each tail entity in the input text , where represents the vector span representation of a head entity from the start position to the end position, represents the vector span representation of a tail entity from the start position to the end position, and maxpool represents the max pooling operation; S34: Use the second context feature for the tail entity as the input feature and the feature representation of the head entity as the conditional function to input the first CLN network model to calculate and generate the first text encoding representation that fuses the head entity features. Use the first context feature for the head entity as the input feature and the feature representation of the tail entity as the conditional function to input the second CLN network model to calculate and generate the second text encoding representation that fuses the tail entity features; Among them, the first text encoding representation that fuses the head entity features is The second text encoding representing the fused tail entity feature is It can be understood that the first CLN network model and the second CLN network model have the same network structure; S35: Using the second tail entity annotator, based on the pointer network method, combine the first text encoding representation, the fifth trainable parameter matrix, and the sixth trainable parameter matrix to calculate the second tail entity start probability for each token as the start position of the tail entity and the second tail entity end probability for the end position. Use the second head entity annotator, based on the pointer network method, combine the second text encoding representation, the third trainable parameter matrix, and the fourth trainable parameter matrix to calculate the second head entity start probability for each token as the start position of the head entity and the second head entity end probability for the end position; and wherein, the second tail entity start probability and the second tail entity end probability of the i-th token are respectively expressed as: , , the second head entity start probability and the second head entity end probability of the i-th token are respectively expressed as , ; S36: Based on the first head entity start probability and the first head entity end probability of each token in the input text, as well as the second tail entity start probability and the second tail entity end probability, extract the head entity and the tail entity with potential relationships as a pair of entity pairs. Based on the first tail entity start probability and the first tail entity end probability of each token in the input text, as well as the second head entity start probability and the second head entity end probability, extract the head entity and the tail entity with potential relationships as a pair of entity pairs.
[0034] In this embodiment, in step S31, since the entity recognition task needs to extract entity pairs in two opposite directions, and there are slight differences in entity extraction between the two directions, therefore, the fused features corresponding to the entity recognition task are respectively multiplied by the trainable and different first parameter matrix and the second parameter matrix, so as to obtain the context feature representations for head entity and tail entity extraction; In steps S32 - S35, since the calculation processes for the two extraction directions are similar, the description is presented for the head entity - tail entity extraction direction. Since it is based on the pointer network method, it identifies the head entity by annotating the start and end positions of the head entity in the text sequence. The pointer network uses a simple 1 / 0 annotation method. First, it calculates the probability of each token being the start and end positions of the entity through the sigmoid activation function. If the probability is greater than the preset threshold, the token position is marked as 1; otherwise, it is marked as 0. For example, if the start position of the head entity of the 7th token in the input text is marked as 1 and the end position of the tail entity of the 9th token is marked as 1, then the 7th to 9th tokens form a head entity. After determining the head entity, for the annotation of the tail entity, the entity pair extraction is regarded as two interrelated processes, that is, the extraction of the tail entity is guided by the previously extracted head entity. After obtaining the feature representation of the head entity, by introducing the CLN (Conditional Layer Normalization) network model, the CLN network model is a technology based on layer normalization (LN). By introducing conditional information, the model can dynamically adjust its normalization process according to different input conditions, thereby improving its performance in conditional generation tasks, such as Figure 3 As shown, in the CLN network model, the input feature x first undergoes standardization processing. Subsequently, a conditional function c is introduced in the normalization structure, and the conditional function c is transformed into the same dimension as the scaling factor γ and the offset β through a linear transformation, so that the transformation process of each feature can be dynamically adjusted according to the input conditional information. The specific calculation process of the CLN network model is as follows: (Formula 1); (Formula 2); (Formula 3); (Formula 4); Among them, μ and σ respectively represent the mean and standard deviation of the second context feature representation h of the tail entity extraction o of, W γ and W γ are the scaling matrix and the offset matrix respectively, which need to be corrected when minimizing the total loss function, b γ and b γ are the corresponding bias terms. In addition, in the above formulas, x is the second context feature representation h of the tail entity extraction o , and c is the feature representation of the head entity , by introducing the feature representation of the head entity into the context feature representation of the tail entity extraction, it is possible to flexibly adjust their fusion method according to the context information of different entity and text features, making the feature fusion process more refined, thereby improving the accuracy and stability of the relation extraction task.
[0035] In an alternative embodiment of the present invention, step S4 specifically includes: S41: Calculate the feature representations of the head entity and the tail entity in the entity pair respectively by using a bi-affine network structure (Biaffine Model) based on a baseline model; wherein, the entity pair is represented as (s_k, o_j), where s_k represents the head entity in the entity pair, and o_j represents the tail entity in the entity pair; The feature representation of the head entity in the entity pair is: ; The feature representation of the tail entity in the entity pair is: ; S42: Calculate the relation probability corresponding to the entity pair based on the feature representations of the head entity and the tail entity in the entity pair and the trainable relation parameter matrix; wherein, the relation probability is represented as: , wherein, T represents the transpose matrix, represents the trainable relation parameter matrix for maintaining the y-th relation.
[0036] In this embodiment, the bi-affine network structure based on the baseline model adopted can accurately model the features of the relation by allocating a relation parameter matrix for each relation compared with the traditional linear layer; secondly, it can output the corresponding relation probability for each relation, thus facilitating the mining of the interaction relationship between two entities and improving the expression ability of the trained model.
[0037] In an alternative embodiment of the present invention, the first loss function is represented as: ; The second loss function is represented as: ; wherein, ; p represents the relation probability, t represents the true label, represents the position marker for the start or end of the head entity or the tail entity in the entity pair, respectively represent the first tail entity annotator, the first head entity annotator, the second tail entity annotator, and the second head entity annotator, g represents the number of tokens included in the input text, and G is a predefined set of relationships.
[0038] In this embodiment, since when performing entity recognition, the pointer network annotation method is used for entity annotation, a total of four annotators are included (i.e., the first tail entity annotator, the first head entity annotator, the second tail entity annotator, and the second head entity annotator). Each annotator labels the start and end positions of the entity as 1, and the remaining positions as 0. Therefore, the entity recognition task can be regarded as a binary classification problem, and the binary cross-entropy loss function can be used to train the entity recognition task; for the relationship assignment task, the relationship assignment of entity pairs can be regarded as a multi-classification problem, and the cross-entropy loss function can be used to train this module. Based on the above principle, the first loss function and the second loss function can be constructed accordingly.
[0039] In an optional embodiment of the present invention, the total loss function is expressed as: , where α and β are function weights used to flexibly adjust according to the importance of the current dataset task. In this embodiment, by setting two different function weights, the first loss function and the second loss function are combined to form the total loss function, so that the function weights can be flexibly adjusted according to the importance of the current dataset task. For example: usually when the entity recognition task and the relationship extraction task have the same difficulty or importance, α = β = 1 can be set; while when the entity recognition task has a higher difficulty or importance relative to the relationship extraction task, α = 2 and β = 1 can be set.
[0040] In an optional embodiment of the present invention, the following parameters are corrected to minimize the total loss function: the pre-trained language model, the first trainable projection matrix, the second trainable projection matrix, the first trainable weight matrix, the second trainable weight matrix, the first trainable parameter matrix, the second trainable parameter matrix, the third trainable parameter matrix, the fourth trainable parameter matrix, the fifth trainable parameter matrix, the sixth trainable parameter matrix, the first CLN network model, the second CLN network model, and the trainable relationship parameter matrix. In this embodiment, the parameters of the pre-trained language model need to be corrected to adapt to the corresponding downstream tasks, and the parameters of each matrix or network model for building the model also need to be updated accordingly to minimize the total loss function and improve the learning and expression ability of the model.
[0041] On the other hand, as Figure 4As shown in the figure, an alternative embodiment of the present invention further provides a training device 1 for a text entity joint relationship extraction model. The device 1 includes a processor 10, a memory 12, and a computer program stored in the memory 12 and configured to be executed by the processor 10. When the processor 10 executes the computer program, the training method of the text entity joint relationship extraction model described in any of the above embodiments is implemented.
[0042] Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units are stored in the memory 12 and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the training device 1 for the text entity joint relationship extraction model. For example, the computer program can be divided into Figure 5 The functional modules in the training device 1 for the text entity joint relationship extraction model, where the text encoding module 31, the task feature interaction module 32, the entity recognition module 33, the relationship assignment module 34, and the loss function construction and optimization module 35 respectively execute the above steps S1 - step S5.
[0043] The training device 1 for the text entity joint relationship extraction model can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The training device 1 for the text entity joint relationship extraction model may include, but is not limited to, a processor 10 and a memory 12. Those skilled in the art can understand that the schematic diagram is only an example of the training device 1 for the text entity joint relationship extraction model, and does not constitute a limitation on the training device 1 for the text entity joint relationship extraction model. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the training device 1 for the text entity joint relationship extraction model may further include input / output devices, network access devices, a bus, etc.
[0044] The processor 10 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor 10 is the control center of the training device 1 of the text entity joint relationship extraction model, and connects various parts of the training device 1 of the entire text entity joint relationship extraction model through various interfaces and lines.
[0045] The memory 12 can be used to store the computer programs and / or modules. The processor 10 realizes various functions of the training device 1 of the text entity joint relationship extraction model by running or executing the computer programs and / or modules stored in the memory 12, and by calling the data stored in the memory 12. The memory 12 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a graphic recognition function, a graphic stacking function, etc.); the data storage area can store data created according to the use of the control device (such as graphic data, etc.). In addition, the memory 12 may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0046] If the functions described in the embodiments of the present invention are implemented in the form of software function modules or units and sold or used as independent products, they can be stored in a storage medium readable by a computing device. Based on such an understanding, to implement all or part of the processes in the above-described method embodiments, the present invention can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 10, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0047] In another aspect, an alternative embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the training method of the text entity association relationship extraction model as described in any of the above embodiments.
[0048] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.
[0049] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the claims of the present invention. All of these are within the protection scope of the present invention.
Claims
1. A training method for a text entity joint relationship extraction model, characterized in that: The method comprises the following steps: Encoding the input text to obtain a first original feature representation of each word contained in the input text in the entity recognition task and a second original feature representation in the relationship extraction task; Based on the multi-head attention mechanism, the first original feature representation and the second original feature representation of each word are dynamically fused in the process of entity recognition task processing to obtain the first fused feature corresponding to each word, and the first original feature representation and the second original feature representation of each word are fused in the process of relation extraction task processing to obtain the second fused feature corresponding to each word; The entity pair recognition method based on bidirectional annotation combines the first fusion feature and the second fusion feature to perform entity recognition from two complementary directions respectively, so as to extract all entity pairs with potential relationships from the input text; Using a bi-affine network structure based on a baseline model to perform relationship assignment on each pair of entities, so as to calculate the relationship probability corresponding to the entity pair; and A first loss function for entity recognition and a second loss function for relationship assignment are constructed respectively in combination with the relationship probability, a total loss function is obtained by combining the first loss function and the second loss function, and the total loss function is minimized to output a text entity joint relationship extraction model.
2. The training method for extracting a text entity joint relationship model according to claim 1, wherein: The encoding of the input text to obtain a first original feature representation of each word contained in the input text in the entity recognition task and a second original feature representation in the relationship extraction task specifically includes: Using a pre-trained language model as an encoder to extract word vectors of the input text; and The first original feature representation is generated based on the word vector of the input text and a first trainable projection matrix of the entity recognition task, and the second original feature representation is generated based on the word vector of the input text and a second trainable projection matrix of the relationship extraction task.
3. The training method for the text entity joint relationship extraction model as claimed in claim 2, characterized in that: The method of dynamically fusing the first original feature representation and the second original feature representation of each word in the entity recognition task processing based on the multi-head attention mechanism to obtain the first fused feature corresponding to each word, and fusing the first original feature representation and the second original feature representation of each word in the relationship extraction task processing to obtain the second fused feature corresponding to each word specifically includes: Extracting the first attention score and the second attention score corresponding to the entity recognition feature and the relationship extraction feature of each word in the entity recognition task based on the multi-layer perceptron network model combined with the first original feature representation and the second original feature representation, and extracting the third attention score and the fourth attention score corresponding to the entity recognition feature and the relationship extraction feature of each word in the relationship extraction task based on the multi-layer perceptron network model combined with the first original feature representation and the second original feature representation; Normalizing the first attention score and the second attention score to obtain a first normalized weight matrix for an entity recognition task, and normalizing the third attention score and the fourth attention score to obtain a second normalized weight matrix for a relationship extraction task; weightedly fusing the first original feature representation and the second original feature representation based on the first normalized weight matrix to obtain a single entity recognition fused feature corresponding to the entity recognition task in a single subspace, and weightedly fusing the first original feature representation and the second original feature representation based on the second normalized weight matrix to obtain a single relationship extraction fused feature in a single subspace; and Based on the first trainable weight matrix combined with the single entity recognition fusion feature, the outputs of all attention heads are weighted averaged and the Dropout operation is applied to obtain the first fusion feature in the entity recognition task. Based on the second trainable weight matrix combined with the single relationship extraction fusion feature, the outputs of all attention heads are weighted averaged and the Dropout operation is applied to obtain the second fusion feature in the relationship extraction task.
4. The training method for the text entity joint relationship extraction model as claimed in claim 3, characterized in that: The entity pair recognition method based on bidirectional annotation combines the first fusion feature and the second fusion feature to perform entity recognition from two complementary directions respectively, so as to extract all entity pairs with potential relationships from the input text, specifically including: Multiplying the first fused feature with a first trainable parameter matrix and a second trainable parameter matrix respectively to obtain a first context feature representation for head entity extraction and a second context feature representation for tail entity extraction; A first head entity tagger is used based on a pointer network method in combination with a first context feature representation for a head entity, a third trainable parameter matrix, and a fourth trainable parameter matrix to calculate and obtain a first head entity start probability for each word element in the input text being a head entity start position and a first head entity end probability for the end position; a first tail entity tagger is used based on a pointer network method in combination with a second context feature representation for a tail entity, a fifth trainable parameter matrix, and a sixth trainable parameter matrix to calculate and obtain a first tail entity start probability for each word element in the input text being a tail entity start position and a first tail entity end probability for the end position; Determine a feature representation of a head entity in the input text based on the first head entity start probability and the first head entity end probability of each word-gram, and determine a feature representation of each tail entity in the input text based on the first tail entity start probability and the first tail entity end probability of each word-gram; Input the second context feature for the tail entity as an input feature and the feature representation of the head entity as a conditional function into the first CLN network model to calculate and generate a first text encoding representation that fuses the head entity feature; input the first context feature for the head entity as an input feature and the feature representation of the tail entity as a conditional function into the second CLN network model to calculate and generate a second text encoding representation that fuses the tail entity feature; A second tail entity tagger is used based on a pointer network method in combination with the first text encoding representation, the fifth trainable parameter matrix and the sixth trainable parameter matrix to calculate the second tail entity start probability of each word element being the start position of the tail entity and the second tail entity end probability of the end position, and a second head entity tagger is used based on a pointer network method in combination with the second text encoding representation, the third trainable parameter matrix and the fourth trainable parameter matrix to calculate the second head entity start probability of each word element being the start position of the head entity and the second head entity end probability of the end position; and Based on the first head entity start probability and the first head entity end probability and the second tail entity start probability and the second tail entity end probability of each word in the input text, the head entity and the tail entity with potential relationship are extracted as a pair of entity pairs; based on the first tail entity start probability and the first tail entity end probability and the second head entity start probability and the second head entity end probability of each word in the input text, the head entity and the tail entity with potential relationship are extracted as a pair of entity pairs.
5. The training method for extracting a text entity joint relationship model according to claim 4, characterized in that: The step of using a bi-affine network structure based on a baseline model to perform relationship assignment on each pair of entities to calculate and obtain the relationship probability corresponding to the entity pairs specifically includes: Using a dual affine network structure based on a baseline model to respectively calculate feature representations of the head entity and the tail entity in the entity pair; and The relationship probability corresponding to the entity pair is calculated based on the feature representation of the head entity and the tail entity in the entity pair and the trainable relationship parameter matrix.
6. The training method for the text entity joint relationship extraction model as claimed in claim 5, characterized in that: The first loss function is expressed as: ; The second loss function is expressed as: ; in, ; p represents the relationship probability, t represents the true label, A position marker indicating the beginning or end of the head entity or the tail entity in the entity pair, They respectively represent the first tail entity tagger, the first head entity tagger, the second tail entity tagger, and the second head entity tagger, g represents the number of word units contained in the input text, and G is a predefined relationship set.
7. The training method for the text entity joint relationship extraction model as claimed in claim 6, characterized in that: The total loss function is expressed as: , where α and β are function weights used to flexibly adjust according to the importance of the current dataset task.
8. The training method for extracting a text entity joint relationship model as claimed in claim 5, characterized in that: The total loss function is minimized by modifying the following parameters: the pre-trained language model, the first trainable projection matrix, the second trainable projection matrix, the first trainable weight matrix, the second trainable weight matrix, the first trainable parameter matrix, the second trainable parameter matrix, the third trainable parameter matrix, the fourth trainable parameter matrix, the fifth trainable parameter matrix, the sixth trainable parameter matrix, the first CLN network model, the second CLN network model, and the trainable relationship parameter matrix.
9. A training device for a text entity joint relationship extraction model, characterized in that: The device includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the training method of the text entity joint relationship extraction model as described in any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the training method of the text entity joint relationship extraction model as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Entity relationship extraction method and system based on deep learning
CN117574902A
Entity relationship joint extraction method and system for data discovery
CN117743475A
Named entity recognition model based on multi-task learning and attention mechanism
CN118114667A
Entity relationship identification method, device, and readable storage medium
WO2023134069A1