Entity Relationship Extraction Method and System Based on Cross-Granularity Cross-Attention Fusion

By adopting a cross-grained cross-attention fusion method in relation extraction, combined with the Bert model and the cross-attention mechanism, the problem of dealing with overlapping entities and relationship scenarios in the prior art is solved, and the accuracy and efficiency of relationship extraction are improved.

CN115391556BActive Publication Date: 2025-05-30UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211032476.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-26
Publication Date
2025-05-30
Estimated Expiration
2042-08-26

AI Technical Summary

Technical Problem

The existing relationship extraction method is difficult to effectively deal with scenes containing overlapping entities and overlapping relationships in sentences, and ignores the differences in span characteristics and the fusion of global information, which affects the extraction accuracy of relational triplets.

Method used

The entity relationship extraction method based on cross-grained cross-attention fusion is adopted. By constructing a sentence semantic information representation model based on Bert, combining the cross-grained global information representation, and combining the entity span representation, assisting entity and type detection and relationship prediction, and outputting entity relationship triples.

Benefits of technology

Effectively learning semantic information related to entities and relationships from sentences improves the accuracy of relationship extraction, can more accurately model the intrinsic connections between entities and relationships, and reduces the accumulation of errors caused by error propagation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115391556B_ABST
    Figure CN115391556B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of relation extraction, and specifically discloses a method and system for entity relation extraction based on cross-granularity cross-attention fusion. The method of the present invention represents sentences as word vectors at the span granularity and token granularity respectively; through cross-attention, it deeply fuses the information of the two granularities of span and token, enhancing the cross-granularity global information representation of entities and relations; in addition, in the joint extraction task, based on different mapping relationships, feature representations for entity detection and relation extraction subtasks are respectively defined, which can capture task-specific context semantic information and can effectively model entity relation triples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of relation extraction, and in particular to an entity relation extraction method and system based on cross-granularity cross-attention fusion. Background Art

[0002] Relation extraction is a subtask in the field of information extraction, which refers to classifying entity pairs into a set of known relations using documents containing mentions of entity pairs (subject, object). Extracting relation triples (subject, relation, object) from natural language text is a key step in constructing a large-scale knowledge graph. Most traditional relation extraction models are based on pattern matching and statistics, requiring a large amount of manual operations. Recently, neural network-based models can more efficiently extract semantic features by learning representations to replace manually constructed features and have achieved considerable success in relation extraction tasks. In neural network-based models, early work adopted a pipeline method, regarding entity recognition and relation extraction as two independent processes. First, all entities in a sentence are recognized, and then relation classification is performed on each entity pair. This method often causes error propagation. Therefore, the academic community has proposed a joint extraction model, jointly modeling the internal connection between entities and relations, alleviating the problem of error accumulation.

[0003] Most existing relation extraction methods adopt sequence labeling and cannot handle scenarios where a sentence contains overlapping entities and overlapping relations at the same time. Recent work has tried to use a span-based joint extraction pattern to solve this problem. However, this method uses the original representation for all spans and global information during entity recognition, ignoring the difference in span features under different categories and the effective integration of spans and global information. When performing relation classification, only the intermediate information of entity pairs is used, ignoring the local context information of relation triples, which affects the accuracy of relation triple extraction. Therefore, this patent proposes an entity relation joint extraction method and system based on cross-granularity cross-attention fusion, effectively learning semantic information related to entities and relations from sentences and improving the extraction accuracy. Summary of the Invention

[0004] To solve the problems existing in the prior art, the present invention provides an entity relation extraction method and system based on cross-granularity cross-attention fusion. Starting from the coarse-grained span representation and the fine-grained token representation, using the cross-attention mechanism, a novel relation extraction model structure is proposed, which can effectively learn the span semantic features and global representation representing entities and relations, realize the extraction of relation triples, and solve the problems mentioned in the above background art.

[0005] To achieve the above object, the present invention provides the following technical solution: An entity relation extraction method based on cross-granularity cross-attention fusion, comprising the following steps:

[0006] Step S10: Construct a sentence semantic information representation model based on Bert. For each sentence in the text, establish a span-based semantic information representation and a token-based semantic information representation;

[0007] Step S20: Establish a linear mapping layer to map the span and token semantic representations into a coarse-grained entity span semantic representation and a fine-grained token semantic representation respectively;

[0008] Step S30: Construct a cross-attention mechanism representation model that fuses entity span and token semantics, generate a cross-granularity global information representation, and combine the entity span representation to assist entity and type detection, and output the type of each entity span;

[0009] Step S40: Filter out entity spans of the none type, perform a linear mapping on the remaining entity spans to generate a span representation for relation prediction;

[0010] Step S50: Use the cross-attention mechanism to fuse the relation span and token semantic representations as the global information representation for relation extraction, combine the span pair representation for relation prediction, predict the relation type, and output the entity relation triples in the text to achieve relation extraction.

[0011] Preferably, the specific steps of Step S10 are as follows:

[0012] Step S101, decompose the text to be subjected to relation extraction into multiple sentences, and the original representation of each sentence is S I =[t 1 , t 2 , t 3 ,..., t l , l is the number of words in the sentence, and the original sentence generates sentence token representations through pre-trained Bert and sentence span representations where n is the number of tokens after word segmentation of the sentence, m is the number of spans in the sentence, h cls and e cls are the global representations at the token granularity and span granularity respectively, and each entity span is represented as the maximum pooling of tokens is the dimension of the Bert word vector.

[0013] Preferably, the specific steps of Step S20 are as follows:

[0014] Step S201, represent the entity span through the entity span mapping function as The token is represented by the token mapping function as The mapping function is as follows:

[0015]

[0016] in, b s , b t are all learnable parameters,

[0017] Preferably, the specific steps of step S30 are as follows:

[0018] Step S301, given the width k of each span, generate a representation of the width k

[0019] Step S302: using the global representation of entity span granularity and token granularity respectively Information is exchanged between the sentence representation at the token granularity and the sentence representation at the entity span granularity, and then merged with the original global representation at the entity span granularity and the original global representation at the token granularity. The calculation formula is as follows:

[0020]

[0021] in, is a learnable parameter;

[0022] Step S303, connect the entity span representation, span width representation and entity global representation for classification, and the calculation formula is as follows:

[0023]

[0024] in, Indicates concat connection, W s , b s is a learnable parameter. Add the none type to the original entity type and select The largest entity type is used as the predicted type for the current entity.

[0025] Preferably, the specific steps of step S40 are as follows:

[0026] Step S401, filter out all entities of type none, leaving m′ entities;

[0027] Step S402: linearly map the remaining entities to generate a span representation of the relationship

[0028] in, br is a learnable parameter,

[0029] Preferably, the specific steps of step S50 are as follows:

[0030] Step S501, respectively use the global representations of the relation span granularity and the token granularity Exchange information between the sentence representation at the token granularity and the sentence representation at the relation span granularity, and then project back to its own granularity to blend with its own global representation. The calculation formula is as follows:

[0031]

[0032]

[0033] Wherein, is a learnable parameter; is the final global representation;

[0034] Step S502, connect the relation span representation, the span width representation, and the relation global representation for classification. The calculation formula is as follows:

[0035]

[0036] Wherein, p represents the number of entity pairs participating in relation prediction, and rn represents the number of relations; b r is a learnable parameter. When classifying each pair of entities, a matrix will be maintained for each relation to accurately model the features of the relation. According to the scores of each relation type judge the relation of the current entity pair, generate entity-relation triples, and realize relation extraction.

[0037] In addition, to achieve the above object, the present invention also provides the following technical solution: An entity relation extraction system based on cross-granularity cross-attention fusion, the extraction system includes:

[0038] An entity detection module based on entity spans and the global representation after cross-attention fusion: used to construct a sentence semantic information representation model based on Bert according to the input sentence to be extracted, establish a semantic information representation based on spans and a semantic information representation based on tokens, and through a linear mapping layer, map the above semantic representations into a coarse-grained entity span semantic representation and a fine-grained token semantic representation respectively. Using the cross-attention mechanism, the semantic representation that fuses entity spans and tokens is used as the global information representation for entity detection, and combined with the entity span representation, it assists in entity and type detection, and outputs the type of each entity span;

[0039] Relationship extraction module based on relationship spans and global representation after cross-attention fusion: Filter all spans predicted as the none type, leaving a set of entity spans to form real triple entity pairs, and linearly map these entity pairs to generate relationship span representations. Then, use the cross-attention mechanism again to take the semantic representation that fuses the relationship span and tokens as the global information representation for relationship extraction, and combine the representation of the relationship span pair to predict the relationship type, outputting the entity relationship triples in the text to achieve relationship extraction.

[0040] The beneficial effects of the present invention are as follows:

[0041] 1) The present invention takes the sentence to be subjected to relationship extraction as the research object, and realizes the deep fusion of information at two granularities by establishing a cross-attention fusion model based on span and token granularities, enhancing the cross-granularity global information representation of entities and relationships; and on this basis, combines linear mapping to generate feature representations for two different subtasks of entity detection and relationship extraction, enabling entity representations and sentence representations to be learned according to the tasks, improving the accuracy of subsequent classification, establishing a more accurate and comprehensive entity relationship representation, and realizing entity category detection and relationship triple extraction.

[0042] 2) The input embedding of the present invention is composed of the max-pooling representation of entity spans, width representation, and global representation after cross-attention, and the use of max-pooling makes entity features more prominent. The method of the present invention jointly models the internal relationship between entities and relationships, alleviates the problem of error accumulation caused by error propagation, is more convenient and concise, and has high efficiency. Brief Description of the Drawings

[0043] Figure 1 It is a schematic flow chart of the method steps of the present invention. Detailed Embodiments

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0045] Embodiment 1

[0046] Please refer to Figure 1 , the present invention provides a technical solution: an entity relationship extraction method based on span and cross-attention fusion, and the specific steps of the relationship extraction method are as follows:

[0047] Step 1: According to the sentences to be extracted as required by the input, construct a sentence semantic information representation model based on Bert, and establish a semantic information representation based on span and a semantic information representation based on token;

[0048] Step 1-1, decompose the text for which relation extraction is to be performed into multiple sentences, and the original representation of each sentence is S I =[t 1 , t 2 , t 3 ,..., t l , where l is the number of words in the sentence, and the original sentence generates sentence token representations through pre-trained Bert and sentence span representations where n is the number of tokens after word segmentation of the sentence, m is the number of spans in the sentence, h cls and e cls ; are the global representations at the token granularity and span granularity respectively, and each entity span is represented as the max pooling of tokens where d is the word vector dimension 768 defined by Bert, and the length range of the span is 1-10, that is, all tokens in the sentence with length ranges are traversed as spans, and then 100 are sampled from them. The samples include the real entities in the sentence and random negative samples.

[0049] Step 2: Establish a linear mapping layer to map the original semantic representations of spans and tokens into coarse-grained entity span semantic representations and fine-grained token semantic representations respectively;

[0050] Step 2-1, represent the entity span through the entity span mapping function as represent the token through the token mapping function as The mapping function is as follows:

[0051]

[0052] where, b s , b t are all learnable parameters,

[0053] Step 3: Construct a cross-attention mechanism representation model that fuses entity span and token semantics, generate cross-granularity global information representations, and combine entity span representations (including span length representations) to assist entity and type detection, and output the type of each entity span;

[0054] Step 3-1: Given the width k of each span, generate a representation of width k

[0055] Step 3-2: Respectively utilize the global representations at the entity span granularity and token granularity Exchange information between the sentence representation at the token granularity and the sentence representation at the entity span granularity, and then blend it with the original global representation at the entity span granularity and the original global representation at the token granularity. The calculation formula is as follows:

[0056]

[0057] Among them, is a learnable parameter;

[0058] Step 3-3: Concatenate the entity span representation, the span width representation, and the entity global representation for classification. The calculation formula is as follows:

[0059]

[0060] Among them, represents concat connection, W s , b s are learnable parameters, softmax is the activation function. By calculating the scores of the entity types to which each span belongs, the class with the highest score is used as the final type prediction of the span. Spans that do not belong to any type are defined as the none type (i.e., the predefined type of negative samples), and spans of this type will not participate in the relationship prediction. Add the none type to the original entity types, and select the entity type that makes the largest as the predicted type of the current entity.

[0061] Step 4: Filter all spans assigned to the none type, leaving a set of entity spans to form the real triple entity pairs, and perform a linear mapping on these entity pairs to generate span representations for relationship prediction;

[0062] Step 4-1: Filter out all entities assigned to the none type, leaving m' entities;

[0063] Step 4-2: Perform a linear mapping on the remaining entities (all entities that are not none) to generate relationship span representations

[0064]

[0065] Among them, b r is a learnable parameter,

[0066] Step 5: Use the cross-attention mechanism to take the semantic representation that fuses the relation span and tokens as the global information representation for relation extraction, and combine it with the span pair representation for relation prediction to predict the relation type;

[0067] Step 5-1: Respectively use the global representations at the relation span granularity and token granularity Exchange information between the sentence representation at the token granularity and the sentence representation at the relation span granularity, and then project back to its own granularity to blend with its own global representation. The calculation formula is the same as that in Step S302, except that the entity span representation in S302 is replaced with the relation span representation, that is, use to replace Use s r (.) to replace s e (.), and finally calculate the global representation using cross-attention The specific calculation formula is as follows:

[0068]

[0069] Among them, is a learnable parameter; is the final global representation;

[0070] Step 5-2: Concatenate the relation span representation, the span width representation, and the relation global representation for classification. The calculation formula is as follows:

[0071]

[0072] ( (represents concat connection)

[0073]

[0074] Among them, p represents the number of entity pairs participating in relation prediction, and rn represents the number of relations; b r is a learnable parameter, sigmoid is the activation function. When classifying each pair of entities, a matrix is maintained for each relation, which can accurately model the features of the relation. According to the score of each relation type, judge the relation of the current entity pair, generate the entity relation triple, that is, (entity 1, relation type, entity 2), and realize relation extraction.

[0075] Example 2

[0076] A joint entity relation extraction system based on cross-granularity cross-attention fusion includes an entity detection module based on entity spans and global representations after cross-attention fusion, and a relation extraction module based on relation spans and global representations after cross-attention fusion.

[0077] Entity detection module based on entity spans and global representations after cross-attention fusion: For the input sentence to be extracted, construct a sentence semantic information representation model based on Bert, establish span-based semantic information representation and token-based semantic information representation. Through a linear mapping layer, map the above semantic representations into coarse-grained entity span semantic representations and fine-grained token semantic representations respectively. Use the cross-attention mechanism to take the semantic representation that fuses entity spans and tokens as the global information representation for entity detection, and combine the entity span representation (including span length representation) to assist in entity and type detection, and output the type of each entity span.

[0078] Relation extraction module based on relation spans and global representations after cross-attention fusion: Filter all spans predicted as the none type, leaving a set of entity spans to form real triple entity pairs, and perform a linear mapping on these entity pairs to generate relation span representations. Again, use the cross-attention mechanism to take the semantic representation that fuses relation spans and tokens as the global information representation for relation extraction, and combine the representation of the relation span pair to predict the relation type, output the entity relation triples in the text, and achieve relation extraction.

[0079] The method of the present invention takes the sentence to be relation-extracted as the research object. By establishing a cross-attention fusion model based on span and token granularities, it realizes the deep fusion of information at two granularities, enhances the cross-granularity global information representation of entities and relations; and on this basis, combines linear mapping to generate feature representations for two different subtasks of entity detection and relation extraction, establishes a more accurate and comprehensive entity relation representation, and realizes entity category detection and relation triple extraction.

[0080] Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An entity relation extraction method based on cross-granularity cross-attention fusion, characterized in that, it includes the following steps: Step S10: Construct a sentence semantic information representation model based on Bert. For each sentence in the text, establish a span-based semantic information representation and a token-based semantic information representation; Step S20: Establish a linear mapping layer to map the span and token semantic representations into a coarse-grained entity span semantic representation and a fine-grained token semantic representation respectively; Step S30: Construct a cross-attention mechanism representation model that fuses entity span and token semantics, generate a cross-granularity global information representation, and combine the entity span representation to assist entity and type detection, and output the type of each entity span; The specific steps are as follows: Step S301, given the width k of each span, generate a representation with width k Step S302, respectively use the global representations at the entity span granularity and the token granularity Exchange information between the sentence representation at the token granularity and the sentence representation at the entity span granularity, and then blend it with the original global representation at the entity span granularity and the original global representation at the token granularity. The calculation formula is as follows: k 1 = W k1 h cross , v 1 = W v1 h cross k 2 = W k2 s e cross , v 2 = W v2 s e cross wherein, are learnable parameters; Step S303, connect the entity span representation, the span width representation, and the entity global representation for classification, and the calculation formula is as follows: Among them, represents concat connection, W s , b s are learnable parameters. Add the none type on the basis of the original entity types, and select the entity type that makes the largest as the predicted type of the current entity; Step S40: Filter out entity spans of the none type, and perform a linear mapping on the remaining entity spans to generate a span representation for relation prediction; Step S50: Use the cross-attention mechanism to fuse the relation span and token semantic representations as the global information representation for relation extraction, combine the span pair representation for relation prediction, predict the relation type, and output the entity relation triple in the text to achieve relation extraction.

2. The entity relation extraction method based on cross-granularity cross-attention fusion according to claim 1, characterized in that: The specific steps of Step S10 are as follows: Step S101: Decompose the text for which relationship extraction is to be performed into multiple sentences. The original representation of each sentence is S I = [t 1 , t 2 , t 3 ,..., t l , where l is the number of words in the sentence. The original sentence is used to generate sentence token representations and sentence span representations through pre-trained Bert and sentence span representations where n is the number of tokens after word segmentation of the sentence, m is the number of spans in the sentence, h cls and e cls are the global representations at the token granularity and span granularity respectively. Each entity span is represented as the max pooling of tokens d is the dimension of the Bert word vector.

3. The entity relation extraction method based on cross-granularity cross-attention fusion according to claim 1, characterized in that: The specific steps of Step S20 are as follows: Step S201, represent the entity span through the entity span mapping function as represent the token through the token mapping function as The mapping function is as follows: Among them, b s , b t are all learnable parameters, 4. The entity relation extraction method based on cross-granularity cross-attention fusion according to claim 1, characterized in that: The specific steps of Step S40 are as follows: Step S401, filter out all entities of the none type, and there are m' remaining entities; Step S402, perform a linear mapping on the remaining entities to generate a relation span representation Among them, b r is a learnable parameter, 5. The entity relation extraction method based on cross-granularity cross-attention fusion according to claim 1, characterized in that: The specific steps of Step S50 are as follows: Step S501, respectively utilize the global representations at the relation span granularity and the token granularity Exchange information between the sentence representation at the token granularity and the sentence representation at the relation span granularity, and then project back to its own granularity to blend with its own global representation. The calculation formula is as follows: k 3 = h cross ,v 3 = W v3 h cross k 4 = W k4 s r cross , v 4 = W v4 s r cross Among them, are learnable parameters, is the final global representation; Step S502, connect the relation span representation, the span width representation, and the relation global representation for classification, and the calculation formula is as follows: Among them, p represents the number of entity pairs participating in relationship prediction, and rn represents the number of relationships. b r is a learnable parameter. When classifying each pair of entities, a matrix is maintained for each relationship to accurately model the characteristics of the relationship. According to the scores of each relationship type judge the relationship of the current entity pair, generate entity-relationship triples, and realize relationship extraction.

6. A relation extraction system for the entity relation extraction method based on cross-granularity cross-attention fusion according to any one of claims 1 to 5, characterized in that: The extraction system includes: Entity Detection Module Based on Entity Spans and Global Representations Blended by Cross-Attention: For the input sentence to be extracted, construct a sentence semantic information representation model based on Bert, establish semantic information representations based on spans and tokens, and through a linear mapping layer, map the above semantic representations into a coarse-grained entity span semantic representation and a fine-grained token semantic representation respectively. Utilize the cross-attention mechanism to take the semantic representation that fuses entity spans and tokens as the global information representation for entity detection, and combine it with the entity span representation to assist in entity and type detection, and output the type of each entity span. Relation Extraction Module Based on Relation Spans and Global Representations Blended by Cross-Attention: Filter all spans predicted as the none type, leaving a set of entity spans to form real triple entity pairs, and perform a linear mapping on these entity pairs to generate relation span representations. Once again, utilize the cross-attention mechanism to take the semantic representation that fuses relation spans and tokens as the global information representation for relation extraction, and combine it with the representation of relation span pairs to predict the relation type, and output the entity relation triples in the text to achieve relation extraction.

Citation Information

Patent Citations

  • Method for extracting chapter relation by fusing multi-level information extraction and noise reduction

    CN113435190A

  • Method for identifying chapter-level event roles based on knowledge graph information guidance

    CN114880434A