Single-stage Joint Entity and Relationship Extraction Method and System Based on Enhanced Sequence Labeling Strategy

By converting entity relationship extraction into sequence labeling task, the combination strategy of BERT model and fully connected neural network is used to solve the problems of nested entities, exposure deviation, redundant calculation and overlapping relationship in joint entity relationship extraction, and the extraction effect and efficiency are improved.

CN115310445BActive Publication Date: 2025-06-24Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210846389.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-19
Publication Date
2025-06-24
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

The prior art has problems such as nested entities, exposure deviation, redundant calculation and overlapping relationships in joint entity relationship extraction, which are difficult to effectively solve.

Method used

Based on the enhanced sequence annotation strategy, the entity relationship extraction task is transformed into the sequence annotation task, the BERT model is used for text encoding, and the label mapping and entity correlation matrix interaction is achieved by combining the fully connected neural network to build a combined loss function for training.

Benefits of technology

It improves the extraction effect of entities and relationships, reduces the amount of model parameters, improves the single-sentence inference speed and F1 value, and significantly improves the extraction performance in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115310445B_ABST
    Figure CN115310445B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of information extraction, and particularly relates to a single-stage joint entity relation extraction method and system based on an enhanced sequence annotation strategy. First, an entity relation extraction model is constructed and trained. The entity relation extraction model includes an encoder for encoding an input text sequence to output a corresponding word vector representation, a labeling component and an entity correlation matrix for performing label mapping on the word vector representation, and a decoder for decoding the label mapping result to extract relevant entity relation triples. In label mapping, the labeling component is used to label the word vector representation with a combined label composed of an entity position, the position of a word in the entity, and a relation type, and the entity correlation matrix is used to enhance the information interaction between the combined labels. Then, the target text sequence to be extracted is input into the trained entity relation extraction model, and the trained entity relation extraction model is used to output the relevant entity triples of the target text sequence, improving the relation entity extraction effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of information extraction, and particularly relates to a single-stage joint entity relation extraction method and system based on an enhanced sequence annotation strategy. Background Art

[0002] In the early stage, the pipeline method was adopted for entity relation extraction, that is, first, a named entity recognition model was used to extract entities in the text, and then a relation classification model was used to predict the relations between candidate entity pairs. Although this method is flexible, simple and easy to implement, and the two subtasks can use independent data sets, there are problems such as error propagation, lack of interaction between the two subtasks, and increased redundant calculations. To solve these problems, subsequent research proposed a joint entity relation extraction method, that is, an end-to-end model based on a neural network was used to simultaneously extract the entities and relations existing in the text. By designing reasonable annotation strategies, vector fusion methods and decoding methods, the interaction between the two subtasks was continuously enhanced, and the extraction effect of the model was continuously improved, achieving better performance compared with the pipeline method.

[0003] In recent years, significant progress has been made in the research on joint entity relation extraction. However, there are still four major challenges as follows: (1) Entity nesting problem. This refers to the situation where one or more other entities are contained within an entity. For example, "Henan Museum" is an entity of the type of organization name, and "Henan" within "Henan Museum" is also an entity of the type of place name. (2) Exposure bias problem. This means that the inputs of each component in the training stage and the inference stage of the model are inconsistent. For example, although models such as CasRel and PRGC can encode entities and relations simultaneously, they degenerate into a pipeline manner in the decoding stage. In the training stage, the inputs of each component come from the true labels, while in the inference stage, the inputs of each component come from the prediction results of the previous component. If the prediction result of the previous component is incorrect, it will lead to error accumulation. (3) Redundant calculation problem. For example, models such as CasRel, TPLinker, and OneRel usually need to pre-define multiple relations and create a matrix for each relation in the training stage. In the inference stage, regardless of whether a certain relation or certain relations exist in the text, all pre-defined relation matrices need to be traversed to extract all entity relation triples, resulting in a redundant calculation problem. Moreover, the more pre-defined relations there are, the longer the inference time will be, the more memory will be occupied, and more computing resources will be consumed. (4) Relation overlap problem. According to the overlap degree of entities in the entity relation triple, sentences can be divided into three types: normal (Normal), entity pair overlap (EntityPairOverlap, EPO), and single entity overlap (SingleEntityOverlap, SEO). If there are no overlapping entities in all entity relation triples in a sentence, this sentence belongs to the normal type; if there are multiple different relations between the same entity pair in a sentence, this sentence belongs to the entity pair overlap type; if an entity in a sentence exists in multiple entity relation triples, this sentence belongs to the single entity overlap type. Summary of the Invention

[0004] Therefore, aiming at the technical problems that the existing joint entity relation extraction in the prior art cannot simultaneously solve nested entities, exposure bias, redundant calculation, and overlapping relations, etc., the present invention provides a single-stage joint entity relation extraction method and system based on an enhanced sequence annotation strategy, which transforms the joint entity relation extraction task into a sequence annotation task to improve the effect of entity extraction.

[0005] According to the design solution provided by the present invention, a single-stage joint entity relation extraction method based on an enhanced sequence annotation strategy is provided, including the following content:

[0006] Construct an entity relationship extraction model and train it. Among them, the entity relationship extraction model includes an encoder for encoding the input text sequence to output the corresponding word vector representation, a labeling component and an entity correlation matrix for label mapping of the word vector representation, and a decoder for decoding the label mapping result to extract relevant entity relationship triples; in label mapping, the labeling component is used to label the word vector representation with a combined label composed of entity positions, positions of words in entities, and relationship types, and the entity correlation matrix is used to enhance the information interaction between combined labels.

[0007] Input the target text sequence to be extracted into the trained entity relationship extraction model, and use the trained entity relationship extraction model to output the relevant entity triples of the target text sequence.

[0008] As the single-stage joint entity relationship extraction method based on the enhanced sequence annotation strategy in the present invention, further, the entity relationship extraction model uses the BERT model structure as the encoder to obtain the word vector representation of the input text sequence. And in the BERT model, first, the input text sequence is converted into a to-be-encoded embedding vector composed of word embedding vectors, segmentation embedding vectors, and position embedding vectors; then the to-be-encoded embedding vector is input into the BERT model for encoding.

[0009] As the single-stage joint entity relationship extraction method based on the enhanced sequence annotation strategy in the present invention, further, a fully connected neural network is used in the entity relationship extraction model to implement the combined label annotation of the labeling component, convert the prediction of each word label in the word vector representation into a multi-label classification problem, use sigmoid as the activation function to obtain the prediction probability of each word belonging to the combined label, and obtain the label mapping corresponding to the word according to a preset probability threshold.

[0010] As the single-stage joint entity relationship extraction method based on the enhanced sequence annotation strategy in the present invention, further, the calculation process of the prediction probability of the combined label of each word is expressed as: p i = sigmoid(W s x i + b s ), where R is the number of predefined entity relationships, W s (g) represents the trainable weight matrix of the network, x i represents the word vector representation of the i-th word, b s represents the trainable bias constant of the network.

[0011] As the single-stage joint entity relation extraction method based on the enhanced sequence annotation strategy of the present invention, further, a fully-connected neural network is used in the entity relation extraction model to realize the combined label information interaction of the entity correlation matrix, sigmoid is used as the activation function to obtain the correlation probability between the combined label and the start words of the head entity and the tail entity, and the corresponding combined label mapping is obtained according to the preset correlation probability threshold.

[0012] As the single-stage joint entity relation extraction method based on the enhanced sequence annotation strategy of the present invention, further, the calculation process of the correlation probability that the combined label is the start word of the head entity and the start word of the tail entity is expressed as: p is,js = sigmoid(W m [x is ; x js +b m ), where W m (g) represents the trainable weight matrix of the network, x is represents the word vector representation of the start word of the i-th head entity, x js represents the word vector representation of the start word of the j-th tail entity, and b m represents the trainable bias constant of the network.

[0013] As the single-stage joint entity relation extraction method based on the enhanced sequence annotation strategy of the present invention, further, the decoder in the entity relation extraction model first decodes the head entity and the tail entity with relations according to the label mapping of the annotation component to find the combined label according to the label index; then, generates entity relation triples by combining the head entities and the tail entities with the same relation in pairs, and decodes the combination of the start words of the head entity and the start words of the tail entity with relations according to the combined label mapping result of the entity correlation matrix; finally, matches the decoded output of the annotation component label mapping and the decoded output of the combined label mapping of the entity correlation matrix, and retains the entity relation triples with relations.

[0014] As the single-stage joint entity relation extraction method based on the enhanced sequence annotation strategy of the present invention, further, a combined loss function composed of an annotation component loss function and an entity correlation matrix loss function is constructed, and the entity relation extraction model is trained using four datasets of NYT, NYT*, WebNLG, and WebNLG*, and the annotation component and the entity correlation matrix share the encoded output of the encoder during the training process.

[0015] As the single-stage joint entity relation extraction method based on the enhanced sequence annotation strategy of the present invention, further, the combined loss function is expressed as: Among them, N represents the length of the input text sequence, R represents the number of predefined relationships, M represents the maximum length of the input text sequence, and y i,j represents the true label, p i,j and p is,js represent the output probabilities of each element in the entity-related matrix in the enhanced sequence annotation component.

[0016] Furthermore, the present invention also provides a single-stage joint entity relationship extraction system based on an enhanced sequence annotation strategy, including: a model training module and a target extraction module, where,

[0017] The model training module is used to build and train an entity relationship extraction model. The entity relationship extraction model includes an encoder for encoding the input text sequence to output the corresponding word vector representation, a labeling component and an entity-related matrix for label mapping of the word vector representation, and a decoder for decoding the label mapping result to extract relevant entity relationship triples; in the label mapping, the labeling component is used to label the word vector representation with a combined label composed of the entity position, the position of the word in the entity, and the relationship type, and the entity-related matrix is used to enhance the information interaction between the combined labels;

[0018] The target extraction module is used to input the target text sequence to be extracted into the trained entity relationship extraction model, and use the trained entity relationship extraction model to output the relevant entity triples of the target text sequence.

[0019] Advantages of the present invention:

[0020] In this case, the enhanced sequence labeling strategy is used to map the entity position, the position of the word in the entity, and the relationship type as a combined label for word label mapping, converting the joint entity relationship extraction task into a sequence annotation task, solving the technical problems such as nested entities, exposure bias, redundant calculation, and overlapping relationships in the existing technology for joint entity relationship extraction, improving the extraction effect of entities and relationships in the text sequence, and facilitating practical scenario applications. Further experimental data shows that compared with the early joint entity relationship extraction models, the performance of the solution in this case is significantly improved. Even compared with advanced models such as CasRel and TPLinker, the model parameter size in the solution of this case can be reduced by 3.23 - 5.36MB, the single-sentence inference speed is increased by 2 - 4.2 times, and the F1 value is increased by 0.5% - 2.1%. Description of the Drawings

[0021] Figure 1 Schematic diagram of the single-stage joint entity relationship extraction process based on the enhanced sequence annotation strategy in the embodiment;

[0022] Figure 2 Schematic diagram of the entity relationship extraction model architecture in the embodiment;

[0023] Figure 3 Schematic diagram of the comparison results in complex scenarios in the embodiments;

[0024] Figure 4 Schematic diagram of the change curve of the F1 value when the model in the embodiments is trained on the WebNLG* dataset;

[0025] Figure 5 Example of the enhanced sequence annotation strategy in the embodiments. Detailed implementation manners

[0026] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and technical solutions.

[0027] In the embodiments of the present invention, refer to Figure 1 as shown, a single-stage joint entity relation extraction method based on an enhanced sequence annotation strategy is provided, including:

[0028] S101. Construct an entity relation extraction model and train it. Among them, the entity relation extraction model includes an encoder for encoding the input text sequence to output the corresponding word vector representation, a labeling component and an entity correlation matrix for performing label mapping on the word vector representation, and a decoder for decoding the label mapping result to extract relevant entity relation triples; in the label mapping, the labeling component is used to label the word vector representation with a combined label composed of the entity position, the position of the word in the entity, and the relation type, and the entity correlation matrix is used to enhance the information interaction between the combined labels;

[0029] S102. Input the target text sequence to be extracted into the trained entity relation extraction model, and use the trained entity relation extraction model to output the relevant entity triples of the target text sequence.

[0030] Using the enhanced sequence labeling strategy, the entity position, the position of the word in the entity, and the relation type are used as combined labels for word label mapping, and the joint entity relation extraction task is transformed into a sequence annotation task. Refer to Figure 2As shown, the text sequence is input into the text sequence encoder to obtain the word vector representation of each word; then the word vector representations of all words are simultaneously input into the sequence labeling component and the entity correlation matrix to obtain the mapping labels of the prediction results; in the decoding module, the mapping labels of the prediction results of the sequence labeling component and the entity correlation matrix are decoded to obtain the head entity and the tail entity with relationships and the relevant start word combinations, and then the head entities and tail entities with the same relationship are combined in pairs to generate entity relationship triples; finally, the relevant entity relationship triples are retained and the irrelevant ones are deleted to obtain the final extraction result, thereby solving the technical problems such as nested entities, exposure bias, redundant calculation and overlapping relationships that cannot be solved simultaneously in the existing joint entity relationship extraction, and improving the entity extraction effect in the text sequence.

[0031] In a piece of text, a word may belong to different types of entities, and the same entity may also participate in entity relationship triples of different relationship types. Since the traditional sequence labeling strategy can only label a word with one label, it cannot solve the problems of nested entities and overlapping relationships. In view of this problem, in the embodiments of this case, using an enhanced sequence labeling strategy, each word is labeled with a combined label of entity position (i.e., head entity or tail entity), relationship type, and the position of the word in the entity (i.e., start word, internal word, or non-entity word), and the order and meaning of each label are listed in Table 1. Thus, it can be calculated that when the length of the input text sequence is N and the number of predefined relationship types is R, the total number of labels for each word is 4×R + 1, and the total number of labels to be labeled for this text sequence is N×(4×R + 1).

[0032] Table 1 Order and meaning of labels in the enhanced sequence labeling component

[0033]

[0034] In the embodiments of this case, further, the entity relationship extraction model uses the BERT model structure as the encoder to obtain the word vector representation of the input text sequence. And in the BERT model, first, the input text sequence is converted into an embedding vector to be encoded composed of word embedding vectors, segment embedding vectors, and position embedding vectors; then the embedding vector to be encoded is input into the BERT model for encoding.

[0035] Using BERT as the text sequence encoder to obtain the word vector representation of the input text sequence. Given an input text sequence e = [e1, e2, …, e N , where e i represents the i-th word in the input text sequence, it is necessary to first convert e into the embedding vector t = [t1, t2, …, t N required by the BERT model input. This vector is composed of the word embedding vector W T, the segmentation embedding vector W S and the position embedding vector W P are added together, and the calculation formula is t = W T + W S + W P . Then, the embedding vector t is input into the BERT model for encoding, and its output vector x = [x1, x2, …, x N is the word vector representation of the input text sequence, where x i represents the word vector representation of the i-th word. The calculation formula can be expressed as x = BERT(t).

[0036] Furthermore, in the embodiments of this case, the entity relationship extraction model uses a fully connected neural network to implement the combined label annotation of the annotation components, converts the prediction of each word label in the word vector representation into a multi-label classification problem, uses sigmoid as the activation function to obtain the prediction probability of the combined label to which each word belongs, and obtains the label mapping corresponding to the word according to a preset probability threshold.

[0037] Converts the label prediction of each word into a multi-label classification problem, rather than the traditional multi-classification problem. First, the enhanced sequence marking component is implemented using a fully connected neural network, and the activation function is sigmoid. Second, the word vector representation of each word is input into the component, and the prediction probability of the label to which each word belongs is output. The calculation formula can be expressed as p i = sigmoid(W s x i + b s ). Where R is the number of predefined relationships, W s (g) represents the trainable weight matrix, x i represents the word vector representation of the i-th word, and b s represents the trainable bias constant. Finally, if the prediction probability of the label to which each word belongs exceeds the set threshold, the mapping result is 1, otherwise it is 0.

[0038] Furthermore, in the embodiments of this case, the entity relationship extraction model uses a fully connected neural network to implement the combined label information interaction of the entity correlation matrix, uses sigmoid as the activation function to obtain the correlation probability between the combined label as the starting word of the head entity and the starting word of the tail entity, and obtains the corresponding combined label mapping according to a preset correlation probability threshold.

[0039] An entity - related matrix is introduced to enhance the interaction between the starting word of the head entity and the starting word of the tail entity, reducing the output of meaningless entity - relation triples. Assuming that the maximum input text sequence length of the model is M, the dimension of the entity - related matrix is [M, M]. First, the entity - related matrix is implemented using a fully - connected neural network, and the sigmoid activation function is used. Second, the word - vector representation of each word is input into the matrix, and the correlation probability between the starting word of the head entity and the starting word of the tail entity is output. The calculation formula can be expressed as p is,js = sigmoid(W m [x is ; x js +b m ). Among them, W m (g) represents the trainable weight matrix, x is represents the word - vector representation of the i - th starting word of the head entity, x js represents the word - vector representation of the j - th starting word of the tail entity, and b m represents the trainable bias constant. Finally, if the correlation probability between the starting word of the head entity and the starting word of the tail entity exceeds the set threshold, the mapping result is 1; otherwise, it is 0.

[0040] Furthermore, in the embodiments of this case, for the entity - relation extraction model decoder, first, the head entity and the tail entity with relations are decoded according to the label mapping of the annotation component to find the combined label according to the label index; then, the head entity and the tail entity with the same relation are combined pairwise to generate entity - relation triples, and the combination of the starting word of the head entity and the starting word of the tail entity with relations is decoded according to the combination - label mapping result of the entity - related matrix; finally, the decoded output of the annotation - component label mapping and the decoded output of the combination - label mapping of the entity - related matrix are matched, and the entity - relation triples with relations are retained.

[0041] The decoder consists of a decoded enhanced - sequence annotation component and a decoded entity - related matrix. First, according to the output mapping result of the enhanced - sequence marking component, the head entity and the tail entity with relations are decoded. The decoded entity can find the label index of the internal words of the head entity according to the label index of the starting word of the head entity, +R, find the label index of the starting word of the tail entity, +2×R, and find the label index of the internal words of the tail entity, +3×R. The corresponding entity - decoding algorithm is shown in Algorithm 1 as follows:

[0042] Algorithm 1 Entity - decoding algorithm

[0043] Input: text sequence T, text - sequence label - prediction result tags, starting - word label index id, predefined relation, number R

[0044] Output: entity list D

[0045]

[0046] Then, the head entities and tail entities with the same relationship can be combined in pairs to generate entity-relationship triples. According to the output mapping result of the entity correlation matrix, the combination of the starting words of the head entity and the starting words of the tail entity with a relationship can be decoded. The corresponding algorithm is shown in Algorithm 2:

[0047] Algorithm 2 Entity-Relationship Triple Decoding Algorithm

[0048] Input: text sequence T, text sequence label prediction result tags, entity correlation matrix label prediction result G, predefined number of relationships R

[0049] Output: list of entity-relationship triples S

[0050]

[0051] Finally, the decoding results of the enhanced sequence labeling component and the decoding results of the entity correlation matrix are matched, and the entity-relationship triples with relationships are retained, while the meaningless entity-relationship triples are deleted, so as to obtain the final extraction result.

[0052] Furthermore, in the embodiments of this case, a combined loss function composed of an annotation component loss function and an entity correlation matrix loss function is constructed, and the entity-relationship extraction model is trained using four datasets, namely NYT, NYT*, WebNLG, and WebNLG*. During the training process, the annotation component and the entity correlation matrix share the encoding output of the encoder.

[0053] The model is trained in a joint learning manner. During the model training, the enhanced sequence labeling component and the entity correlation matrix share the encoding results of the text sequence encoder to optimize the combined loss function. Therefore, the combined loss function consists of two parts: the loss function of the enhanced sequence annotation component and the loss function of the entity correlation matrix. The calculation formulas are as follows respectively:

[0054]

[0055]

[0056] Among them, N represents the length of an input text sequence, R represents the number of predefined relationships, M represents the maximum length of the input text sequence, and y i,j represents the true label, p i,j and p is,js represent the output probabilities of each element in the enhanced sequence annotation component and the entity correlation matrix. The final overall loss function can be as follows:

[0057]

[0058] Furthermore, based on the above method, an embodiment of the present invention further provides a single-stage joint entity relation extraction system based on an enhanced sequence annotation strategy, including: a model training module and a target extraction module, where,

[0059] The model training module is used to construct and train an entity relation extraction model. The entity relation extraction model includes an encoder for encoding an input text sequence to output a corresponding word vector representation, a labeling component and an entity correlation matrix for performing label mapping on the word vector representation, and a decoder for decoding the label mapping result to extract relevant entity relation triples; in label mapping, the labeling component is used to label the word vector representation with a combined label composed of entity positions, positions of words in the entity, and relation types, and the entity correlation matrix is used to enhance the information interaction between the combined labels;

[0060] The target extraction module is used to input a target text sequence to be extracted into the trained entity relation extraction model, and use the trained entity relation extraction model to output relevant entity triples of the target text sequence.

[0061] To verify the effectiveness of the solution in this case, further explanation is given below in combination with experimental data:

[0062] Experiments were conducted on different relationship overlap types existing in sentences on the NYT* and WebNLG* datasets, and the results are as Figure 3 shown. At the same time, experiments were conducted on the number of different entity relation triples existing in sentences, and the results are listed in Table 2.

[0063] Table 2 Comparison results of F1 values of the model of the present invention and the baseline model under the condition of the number of different entity relation triples existing in sentences

[0064]

[0065] From Figure 3 and Table 2, it can be seen that: the single-stage joint entity relation extraction based on the enhanced sequence annotation strategy in the solution of this case has significantly better entity relation extraction ability than traditional joint extraction models in complex scenarios such as overlapping relations and multi-entity relation triples, and even when the number of triples included in the sentence increases continuously, it can basically maintain stable extraction performance.

[0066] The model effectively improves the extraction effect by introducing an entity correlation matrix to enhance the interaction between the head entity and the tail entity of the present invention. To prove the effectiveness of this component, two groups of ablation experiments were designed. The first group of ablation experiments was to remove the entity correlation matrix, and the second group of ablation experiments was to replace the entity correlation matrix with a relation prediction component. The experimental results are listed in Table 3.

[0067] Table 3 Ablation Experiment Results

[0068]

[0069] From the results of the first group of ablation experiments, it can be found that introducing the entity-related matrix can significantly improve the precision, with an increase of 3.6% and 3.7% on the NYT* / NYT datasets respectively, and an increase of 0.5% and 0.7% on the WebNLG* / WebNLG datasets respectively. This shows that only using the sequence annotation component will generate a large number of meaningless candidate entity relation triples, and the entity-related matrix can play a good auxiliary role. However, due to the sparsity of the entity-related matrix, while deleting meaningless candidate entity relation triples, some correct entity relation triples are also deleted, so the recall rate is relatively low.

[0070] Figure 4 Show the change of F1 value when the model of this case, the CasRel model and the TPLinker model are trained on the WebNLG* dataset. Table 4 shows the statistical information of the number of parameters of these three models and the single-sentence inference time under different batch sizes.

[0071] Table 4 Computational Efficiency Analysis Results

[0072]

[0073] From Figure 4 it can be seen that the model of this case has approached convergence at the 14th round, significantly earlier than the CasRel model and the TPLinker model, greatly reducing the training cost. From Table 4, it can be seen that the model of this case has the smallest number of parameters among the three models, but the fastest single-sentence inference speed.

[0074] See Figure 2 As shown, when the input text sequence is "Zhang Sanwu wrote and directed Kung Fu", and the predefined relations are "director" and "starring", then there are a total of 9 labels for each word. The detailed labels are listed on the Figure 5 left side. Among them, the first row is the non-entity label, the second row to the fifth row are the head entity labels, and the sixth row to the ninth row are the tail entity labels. This text contains two entity relation triples: (Zhang Sanwu, director, Kung Fu) and (Zhang Sanwu, starring, Kung Fu). The annotation results of the first entity relation triple are marked by light gray squares in Figure 2 and 4 , and the second entity relation triple is marked by dark gray squares in Figure 2 and 4 . The remaining non-entity words are marked by squares in Figure 2 .

[0075] The mapping result of the enhanced sequence annotation component is asFigure 2 As shown in (c) of ,

[0076] First, traversing the head entity tags marked as 1 can decode the head entities with relationships (director, Zhang Sanwu) and (lead actor, Zhang Sanwu), and traversing the tail entity tags marked as 1 can decode the tail entities with relationships (director, "Kung Fu"), (director, Kung Fu), and (lead actor, Kung Fu). Then, pairwise combining the head entities and tail entities with the relationship of "director" generates entity relationship triples (Zhang Sanwu, director, "Kung Fu") and (Zhang Sanwu, director, Kung Fu), and pairwise combining the head entities and tail entities with the relationship of "lead actor" generates entity relationship triples (Zhang Sanwu, lead actor, Kung Fu).

[0076] The mapping result of the entity correlation matrix is as Figure 2 shown in (d) of . Figure 2 Traversing all elements marked as 1 can decode the relevant start word combinations (Zhang, Gong). Finally, matching the decoding results of the enhanced sequence tagging component and the decoding results of the entity correlation matrix, retaining the entity relationship triples with relationships and deleting the meaningless entity relationship triples, thus obtaining the final extraction result. The detailed process is shown in Algorithm 2. As Figure 2 shown in the decoding module of ,

[0077] the start word combination of (Zhang Sanwu, director, "Kung Fu") is (Zhang, "), which is not a relevant start word combination, so the result is deleted; the start word combination of (Zhang Sanwu, director, Kung Fu) is (Zhang, Gong), which is a relevant start word combination, so the result is retained; the start word combination of (Zhang Sanwu, lead actor, Kung Fu) is (Zhang, Gong), which is a relevant start word combination, so the result is retained. Therefore, the final entity relationship extraction results are (Zhang Sanwu, director, Kung Fu) and (Zhang Sanwu, lead actor, Kung Fu).

[0077] Unless otherwise specifically stated, the relative steps, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0078] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0079] The units and method steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation is not considered to exceed the scope of the present invention.

[0080] Those of ordinary skill in the art can understand that all or part of the steps in the above method can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disc, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, each module / unit in the above embodiments can be implemented in the form of hardware or in the form of a software functional module. The present invention is not limited to any specific form of the combination of hardware and software.

[0081] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, and are not intended to limit them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A single-stage joint entity and relation extraction method based on an enhanced sequence annotation strategy, characterized in that It includes the following content: Construct an entity relation extraction model and train it. Among them, the entity relation extraction model includes an encoder for encoding the input text sequence to output the corresponding word vector representation, a labeling component and an entity correlation matrix for label mapping of the word vector representation, and a decoder for decoding the label mapping result to extract relevant entity relation triples; in label mapping, the labeling component is used to label the word vector representation with a combined label composed of entity position, word position in the entity, and relation type, and the entity correlation matrix is used to enhance the information interaction between combined labels; the entity relation extraction model uses the BERT model structure as the encoder to obtain the word vector representation of the input text sequence. In the BERT model, first, the input text sequence is converted into a to-be-encoded embedding vector composed of word embedding vectors, segment embedding vectors, and position embedding vectors; then the to-be-encoded embedding vector is input into the BERT model for encoding; and in the entity relation extraction model, a fully connected neural network is used to implement the combined label labeling of the labeling component, converting the prediction of each word label in the word vector representation into a multi-label classification problem, using sigmoid as the activation function to obtain the prediction probability of each word belonging to the combined label, and obtaining the label mapping corresponding to the word according to a preset probability threshold; in the entity relation extraction model, a fully connected neural network is used to implement the combined label information interaction of the entity correlation matrix, using sigmoid as the activation function to obtain the correlation probability between the combined label as the start word of the head entity and the start word of the tail entity, and obtaining the corresponding combined label mapping according to a preset correlation probability threshold; the decoder in the entity relation extraction model first decodes the head entity and the tail entity with a relationship according to the label mapping of the labeling component to find the combined label according to the label index; then, the head entities and tail entities with the same relationship are combined in pairs to generate entity relation triples, and the combined label mapping result of the entity correlation matrix is used to decode the combination of the start word of the head entity and the start word of the tail entity with a relationship; finally, the decoding output of the label mapping of the labeling component and the decoding output of the combined label mapping of the entity correlation matrix are matched, and the entity relation triples with a relationship are retained; Input the target text sequence to be extracted into the trained entity relation extraction model, and use the trained entity relation extraction model to output the relevant entity triples of the target text sequence.

2. The single-stage joint entity relation extraction method based on the enhanced sequence annotation strategy according to claim 1, wherein The calculation process of the predicted probability of each word's combined tag is expressed as: p i = sigmoid(W s x i + b s ), where, R is the number of predefined entity relationships, W s () represents the trainable weight matrix of the network, x i represents the word vector representation of the i-th word, b s represents the trainable bias constant of the network.

3. The single-stage joint entity-relation extraction method based on the enhanced sequence annotation strategy according to claim 1, characterized in that The calculation process of the relevant probability combining the starting word of the head entity and the starting word of the tail entity is expressed as: p is,js = sigmoid(W m [x is ; x js + b m ), where W m () represents the trainable weight matrix of the network, x is represents the word vector representation of the i-th starting word of the head entity, x js represents the word vector representation of the j-th starting word of the tail entity, and b m represents the trainable bias constant of the network.

4. The single-stage joint entity relation extraction method based on an enhanced sequence annotation strategy according to claim 3, characterized in that Construct a combined loss function composed of a labeling component loss function and an entity correlation matrix loss function, and use four datasets, namely NYT, NYT*, WebNLG, and WebNLG*, to train the entity relation extraction model. During the training process, the labeling component and the entity correlation matrix share the encoding output of the encoder.

5. The single-stage joint entity-relation extraction method based on the enhanced sequence annotation strategy according to claim 4, characterized in that The combined loss function is expressed as: where N represents the length of the input text sequence, R represents the number of predefined relationships, M represents the maximum length of the input text sequence, y i,j represents the true label, p i,j and p is,js represent the output probabilities of each element in the entity-related matrix in the augmented sequence annotation component.

6. A single-stage joint entity and relation extraction system based on an enhanced sequence annotation strategy, characterized in that It includes: a model training module and a target extraction module, where, A model training module for constructing and training an entity relationship extraction model. The entity relationship extraction model includes an encoder for encoding an input text sequence to output corresponding word vector representations, an annotation component and an entity correlation matrix for label mapping of the word vector representations, and a decoder for decoding the label mapping results to extract relevant entity relationship triples. In the label mapping, the annotation component is used to annotate the word vector representations with combined labels composed of entity positions, positions of words in entities, and relationship types, and the entity correlation matrix is used to enhance the information interaction between the combined labels. The entity relationship extraction model uses the BERT model structure as the encoder to obtain the word vector representations of the input text sequence. In the BERT model, first, the input text sequence is converted into a to-be-encoded embedding vector composed of word embedding vectors, segment embedding vectors, and position embedding vectors. Then, the to-be-encoded embedding vector is input into the BERT model for encoding. In the entity relationship extraction model, a fully connected neural network is used to implement the combined label annotation of the annotation component, converting the prediction of each word label in the word vector representations into a multi-label classification problem, using sigmoid as the activation function to obtain the prediction probabilities of each word belonging to the combined labels, and obtaining the label mapping corresponding to the words according to a preset probability threshold. In the entity relationship extraction model, a fully connected neural network is used to implement the information interaction of the combined labels of the entity correlation matrix, using sigmoid as the activation function to obtain the correlation probabilities between the combined labels that are the start words of the head entity and the start words of the tail entity, and obtaining the corresponding combined label mapping according to a preset correlation probability threshold. The decoder in the entity relationship extraction model first decodes the head entity and the tail entity with relationships according to the label mapping of the annotation component to find the combined labels according to the label indices. Then, the head entities and the tail entities with the same relationship are combined pairwise to generate entity relationship triples, and the combined labels of the start words of the head entity and the start words of the tail entity with relationships are decoded according to the combined label mapping results of the entity correlation matrix. Finally, the decoded outputs of the label mapping of the annotation component and the decoded outputs of the combined label mapping of the entity correlation matrix are matched, and the entity relationship triples with relationships are retained. A target extraction module for inputting a target text sequence to be extracted into the trained entity relationship extraction model and using the trained entity relationship extraction model to output the relevant entity triples of the target text sequence.