DocRE method based on three-perspective attention and entity-relation correlation

By optimizing document-level relationship extraction through a three-angle attention mechanism and an adaptive loss function, the problems of insufficient multi-hop reasoning and cross-perspective semantic modeling in existing methods are solved, achieving more efficient document-level relationship extraction and accurate identification of long-tail relationships.

CN120407806BActive Publication Date: 2025-09-09SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510899504.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-09-09
Estimated Expiration
2045-07-01

AI Technical Summary

Technical Problem

Existing document-level relationship extraction methods have low multi-hop reasoning efficiency and distorted information transmission in long document scenarios. They lack dynamic attention integration and cross-perspective semantic modeling, resulting in unbalanced relationship label distribution and insufficient model generalization ability.

Method used

A three-angle attention mechanism is adopted, including entity pair-context attention, entity pair-evidence attention and inter-entity pair attention, combined with graph attention network and adaptive loss function, to dynamically integrate multi-dimensional information, optimize the generalization ability of long-tail relations and collaborative modeling of evidence information.

Benefits of technology

It improves the accuracy and efficiency of document-level relationship extraction, enhances the collaborative modeling capabilities of multi-hop reasoning and evidence information, and improves the recognition accuracy of long-tail relationships and the training stability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407806B_ABST
    Figure CN120407806B_ABST
Patent Text Reader

Abstract

The present invention specifically discloses a DocRE method based on three-angle attention and entity-relationship relevance, which belongs to the field of natural language processing technology. The method first encodes the input document into a pre-trained language model PLM to obtain an embedding representation and a multi-head attention matrix; then designs three mechanisms: entity pair-context attention, entity pair-evidence attention, and entity pair-inter-attention, and performs feature alignment and weighted fusion on the three types of attention information through an attention transfer module to generate a unified entity pair semantic representation; then constructs an entity-relationship co-occurrence graph, uses a graph attention network to achieve node feature aggregation and propagation, and obtains entity-relationship co-occurrence features; finally, the fused semantic features are spliced ​​with the co-occurrence features for classification prediction of entity pair relationships. This method effectively improves the cross-sentence reasoning ability and long-distance entity modeling effect, and enhances the model's extraction accuracy and generalization ability for complex relationships.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to a DocRE method based on three-angle attention and entity-relation relevance. Background Art

[0002] In recent years, with the explosive growth of information and the increasing demand for large-scale text data processing, document-level relation extraction (DocRE) has become a core research hotspot in the field of natural language processing. Current technical approaches can be divided into two main categories: one is graph neural network-based methods, which explicitly model entity interactions by constructing semantic graph structures such as dependency syntactic graphs and coreference graphs. However, these methods rely heavily on rules or heuristic strategies to construct graph structures, making it difficult to adaptively handle complex contextual dependencies across sentences and paragraphs in documents. In particular, they suffer from prominent issues such as inefficient multi-hop reasoning and information distortion in long document scenarios. The other is pre-trained language model methods based on the Transformer architecture (such as BERT and RoBERTa), which implicitly model global semantic dependencies using the self-attention mechanism. While this simplifies the feature engineering process, it is significantly limited in its ability to model relationships between distant entities due to the length of the positional encoding. Furthermore, the lack of explicit modeling of the reasoning process leads to weak model interpretability. In addition, DocRE tasks generally face the problem of extremely unbalanced distribution of relationship labels. Existing solutions mostly use static weighting strategies such as cross entropy, Focal Loss, and other static loss functions or negative sampling strategies to alleviate this contradiction. However, it is difficult to dynamically distinguish negative samples of different difficulty levels. In multi-label scenarios, it is also difficult to adapt to the sparsity differences of different relationship categories, which in turn affects the model's generalization ability and training stability on long-tail relationships.

[0003] Although document-level relationship extraction methods based on graph neural networks and Transformer have become mature, most of them are optimized based on constructing semantic graphs between entities or introducing self-attention mechanisms, but there are still the following technical bottlenecks: First, existing methods generally lack the ability to collaboratively model multi-hop reasoning processes with contextual semantics and evidence information, making it difficult to achieve effective linkage between cross-sentence evidence and reasoning paths; second, static fusion strategies are often adopted, lacking the ability to dynamically align and integrate multiple attention semantics, and unable to adaptively adjust the weight distribution of different attentions according to input features; third, most of them fail to integrate the comprehensive modeling of relationships from multiple perspectives such as semantic space and relation space, resulting in limited adequacy of relationship representation and difficulty in capturing the deep correlation characteristics between entities and relationships.

[0004] The above problems affect the system performance and practicality in complex relationship extraction tasks. It is urgent to carry out systematic innovation in dimensions such as multi-hop reasoning mechanism, dynamic attention migration and cross-perspective semantic modeling to break through the existing technical bottlenecks. Summary of the Invention

[0005] The purpose of this invention is to propose a DocRE method based on three-angle attention and entity-relationship correlation to solve the above-mentioned problems existing in the prior art, improve the accuracy and efficiency of document-level relationship extraction, and protect research and development results.

[0006] To achieve the above objectives, the present invention proposes a DocRE method based on three-angle attention and entity-relation relevance, which includes the following steps:

[0007] Step S1: Input the document into the pre-trained language model PLM to obtain the token embedding representation and multi-head attention matrix;

[0008] Step S2: construct three attention mechanisms: entity pair-context attention E2CA, entity pair-evidence attention E2EV, and entity pair-inter-attention E2EA based on the multi-head attention matrix of the last layer of the pre-trained language model;

[0009] Step S3: Design an attention transfer module AT to perform feature alignment and weighted fusion on the information of the three attention mechanisms E2CA, E2EV, and E2EA to generate a unified entity pair semantic representation;

[0010] Step S4: construct an entity-relationship co-occurrence graph, introduce a graph attention network (GAT) based on the entity-relationship co-occurrence graph to aggregate and propagate features of nodes in the graph, and obtain entity-relationship co-occurrence features;

[0011] Step S5: combining the fused semantic features with the co-occurrence features as a basis for determining entity pair relationships;

[0012] Step S6: loss function design, including adaptive focal loss and loss function guided by evidence information;

[0013] Step S7: perform relationship classification prediction on each entity pair and output the final relationship category result.

[0014] Preferably, in step S2, the entity pair attention E2EA adopts an axial attention mechanism to decompose the attention matrix output by the PLM into horizontal and vertical axis directions to calculate the attention distribution respectively, and mine the multi-hop paths and non-local dependency information between entities.

[0015] Preferably, in step S2, entity pair-context attention E2CA uses entity pair embedding as query Query and context embedding as key / value pair Key / Value, and captures local semantic associations through attention calculation.

[0016] Preferably, in step S2, the entity pair-evidence attention E2EV extracts the contextual position of the entity pair attention based on the token-level attention output of PLM, and maps it to the sentence level to construct sentence-level evidence attention.

[0017] Preferably, in step S3, the specific operations of feature alignment and weighted fusion are: first, the features of the three attention types E2CA, E2EV and E2EA are spliced, and then through multiple layers of linear transformation and nonlinear activation function, the weight distribution of different attention types in the final fusion representation is learned to generate a unified entity pair semantic representation.

[0018] Preferably, in step S4, the step of constructing an entity-relationship co-occurrence graph includes: based on the training corpus statistics of the co-occurrence information of entities and relationship labels, constructing an entity-relationship co-occurrence matrix, with entities and relationship labels as nodes and co-occurrence relationships as edges; the weight of the edge is determined by the conditional probability The calculation formula is as follows:

[0019] ;

[0020] in, is the entity in the entity-relationship co-occurrence matrix With relationship tags The number of co-occurrences of .

[0021] Preferably, in step S5, the calculation formula of the adaptive focus loss is as follows:

[0022] ;

[0023] ;

[0024] ;

[0025] in, represents the loss of the relation extraction task, 、 Indicates that the entity Next predicted relationship category The probability of 、 Indicates that the entity The above prediction is the threshold class The probability of Represents entity pairs For relationships The logit value of Represents entity pairs For the threshold class The logit value of represents the negative class subset, is the control factor of adaptive focus loss, For entity pairs For the threshold class The logit value of For entity pairs For relationships The logit value of is the threshold category, is the positive subset;

[0026] The calculation formula of the loss function guided by evidence information is as follows:

[0027] ;

[0028] in, is the KL divergence loss, is the artificial evidence distribution, is the evidence distribution of the model prediction, represents the loss of the evidence retrieval task;

[0029] The total loss function The formula is as follows:

[0030] ;

[0031] in, is an adjustable balance factor.

[0032] Therefore, this paper proposes a DocRE method based on three-angle attention and entity-relation correlation, which has the following beneficial effects:

[0033] (1) Multi-dimensional attention mechanism enhances semantic modeling capabilities: Through entity pair attention (E2EA) to mine multi-hop paths and non-local dependencies, entity pair-context attention (E2CA) to focus on local semantic associations, and entity pair-evidence attention (E2EV) to locate supporting evidence, combined with the attention transfer module (AT) to dynamically integrate three types of information, improve long-distance relationship modeling and evidence integration capabilities, and solve the problem of insufficient collaborative modeling of multi-hop reasoning and contextual semantics and evidence information in existing methods.

[0034] (2) Entity-relationship co-occurrence graph optimizes the generalization ability of long-tail relationships: Constructing an entity-relationship co-occurrence graph and propagating features through the graph attention network (GAT), using conditional probability to quantify the co-occurrence correlation between entities and relationships, explicitly introducing semantic prior knowledge, effectively alleviating the category imbalance problem, and improving the long-tail relationship recognition accuracy and training stability.

[0035] (3) Adaptive loss function optimizes training and reasoning performance: Adaptive focus loss dynamically distinguishes the difficulty of negative samples, KL divergence loss guides the consistency of evidence distribution, and the total loss function balances relationship classification and evidence modeling, improving the generalization ability and interpretability of the model in long-tail relationships and cross-paragraph reasoning scenarios, and improving the shortcomings of existing static loss functions.

[0036] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a flowchart of a DocRE method based on three-angle attention and entity-relationship relevance of the present invention;

[0038] Figure 2 This is the architecture diagram of the three-angle attention mechanism in the present invention;

[0039] Figure 3 Schematic diagram of the structure of the attention transfer module in the present invention;

[0040] Figure 4 Schematic diagram for constructing the entity-relationship co-occurrence graph in the present invention. DETAILED DESCRIPTION

[0041] To make the technical solutions, advantages, and objectives of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below. The described embodiments are part of the embodiments of the present invention, not all of them. Based on the described embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0042] Unless otherwise defined, technical or scientific terms used in the present invention shall have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs.

[0043] Example 1

[0044] like Figure 1 As shown in FIG, a flowchart of a DocRE method based on three-angle attention and entity-relationship correlation of the present invention is shown, and the specific steps are as follows:

[0045] S1. Input the document to the pre-trained language model PLM to obtain the token embedding representation and multi-head attention matrix;

[0046] S2, such as Figure 2 As shown in Figure 1, three attention mechanisms, namely entity pair-context attention E2CA, entity pair-evidence attention E2EV, and entity pair-inter-attention E2EA, are constructed based on the multi-head attention matrix of the last layer of the pre-trained language model.

[0047] Entity-to-Entity Attention (E2EA) uses an axial attention mechanism to decompose the attention matrix output by PLM into horizontal and vertical axes to calculate the attention distribution respectively, and mine multi-hop paths and non-local dependencies between entities.

[0048] Entity-Context Attention (E2CA) uses entity pair embeddings as queries and context embeddings as key / value pairs. It uses an attention mechanism to focus on local context fragments that are highly relevant to the entity pair, capturing their direct semantic connections.

[0049] Entity Pair-Evidence Attention E2EV is based on the token-level attention output of PLM, extracts the contextual position of entity pair attention, and maps it to the sentence level, constructs sentence-level evidence attention, guides the model to focus on potential supporting sentences, and enhances the modeling ability of evidence required for relational reasoning.

[0050] S3, such as Figure 3 As shown in the figure, an attention transfer module AT is designed to perform feature alignment and weighted fusion on the information of the three attention mechanisms E2CA, E2EV and E2EA to generate a unified entity pair semantic representation. The specific operation is: first, the features of the three attention mechanisms E2CA, E2EV and E2EA are spliced ​​together, and then through multiple layers of linear transformation and nonlinear activation function, the weight distribution of different attention types in the final fused representation is learned.

[0051] S4, such as Figure 4 As shown in the figure, an entity-relationship co-occurrence graph is constructed, and a graph attention network GAT is introduced based on the entity-relationship co-occurrence graph to aggregate and propagate features of nodes in the graph to obtain entity-relationship co-occurrence features;

[0052] The DREEAM model is used on the Re-DocRED dataset to obtain the association distribution between entities and relationships; then, an entity-relationship co-occurrence matrix is ​​constructed based on this distribution. ; Where m is the number of entities and n=97 is the number of relationship labels.

[0053] The nodes in the graph represent entities and different relationship labels (such as "Located in", "Part of", etc.), and the edges represent the co-occurrence relationship between entities and relationship labels. The weight of the edge is determined by the conditional probability Decision, the probability that a given entity In case of occurrence, the relationship tag The probability of occurrence is calculated as follows:

[0054] ;

[0055] in, is the entity in the entity-relationship co-occurrence matrix With relationship tags The number of co-occurrences of .

[0056] S5. Combine the fused semantic features and co-occurrence features as the basis for determining entity pair relationships;

[0057] S6. Loss function design, including adaptive focal loss and loss function guided by evidence information;

[0058] The adaptive focus loss is calculated as follows:

[0059] ;

[0060] ;

[0061] ;

[0062] The entity The corresponding label space is divided into positive class subsets and negative subset .in, A set of labels indicating the existence of a true relationship between entity pairs, Represents all relationship labels that are irrelevant to the entity pair; in addition, a learnable threshold class (TH) is introduced in the label space to serve as the discriminant boundary between positive and negative classes; Indicates that the entity Next predicted relationship category The probability of In short, ,in Represents entity pairs For relationships The logit value of Represents entity pairs For the threshold class The logit value of Indicates that the entity The next prediction is the threshold class The probability of Abbreviated as , Represents entity pairs For the threshold class The logit value of Represents entity pairs For relationships The logit value of is the control factor of adaptive focus loss, represents the loss of the relation extraction task.

[0063] The calculation formula of the loss function guided by evidence information is as follows:

[0064] ;

[0065] in, is the KL divergence loss, is the artificial evidence distribution, Predict the distribution for the model;

[0066] The total loss function The formula is as follows:

[0067] ;

[0068] in, is an adjustable balance factor.

[0069] S7. Perform relationship classification prediction for each entity pair and output the final relationship category result.

[0070] It is worth noting that the contents not elaborated in detail in the present invention are all prior art and are well known to those skilled in the art.

[0071] Therefore, the present invention provides a DocRE method based on three-angle attention and entity-relationship correlation. Through the three-angle attention mechanism, attention transfer module, entity-relationship co-occurrence graph modeling and adaptive loss function, multi-hop reasoning and evidence information collaborative modeling, dynamic integration of multi-source attention and cross-perspective relationship representation are realized, thereby improving the accuracy of document-level relationship extraction, the generalization ability of long-tail relationships and the interpretability of the model.

[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A DocRE method based on three-perspective attention and entity-relation correlation, characterized by: Here are the steps: Step S1: Input the document into the pre-trained language model PLM to obtain the token embedding representation and multi-head attention matrix; Step S2: construct three attention mechanisms: entity pair-context attention E2CA, entity pair-evidence attention E2EV, and entity pair-inter-attention E2EA based on the multi-head attention matrix of the last layer of the pre-trained language model; Step S3: Design an attention transfer module AT to perform feature alignment and weighted fusion on the information of the three attention mechanisms E2CA, E2EV, and E2EA to generate a unified entity pair semantic representation; Step S4: construct an entity-relationship co-occurrence graph, introduce a graph attention network (GAT) based on the entity-relationship co-occurrence graph to aggregate and propagate features of nodes in the graph, and obtain entity-relationship co-occurrence features; Step S5: combining the fused semantic features with the co-occurrence features as a basis for determining entity pair relationships; Step S6: loss function design, including adaptive focal loss and loss function guided by evidence information; Step S7: perform relationship classification prediction on each entity pair and output the final relationship category result; In step S4, the step of constructing the entity-relationship co-occurrence graph includes: based on the training corpus, counting the co-occurrence information of entities and relationship labels, constructing an entity-relationship co-occurrence matrix, with entities and relationship labels as nodes and co-occurrence relationships as edges; the weight of the edge is determined by the conditional probability The calculation formula is as follows: ; in, is the entity in the entity-relationship co-occurrence matrix With relationship tags The number of co-occurrences of In step S6, the adaptive focus loss is calculated as follows: ; ; ; in, represents the loss of the relation extraction task, 、 Indicates that the entity Next predicted relationship category The probability of 、 Indicates that the entity The upper prediction is the threshold class The probability of Represents entity pairs For relationships The logit value of Represents entity pairs For the threshold class The logit value of represents the negative class subset, is the control factor of adaptive focus loss, For entity pairs For the threshold class The logit value of For entity pairs For relationships The logit value of is the threshold category, is the positive subset; The calculation formula of the loss function guided by evidence information is as follows: ; in, is the KL divergence loss, is the artificial evidence distribution, is the evidence distribution of the model prediction, represents the loss of the evidence retrieval task; The total loss function The formula is as follows: ; in, is an adjustable balance factor.

2. The DocRE method based on three-angle attention and entity-relationship relevance according to claim 1, characterized in that: In step S2, the entity pair attention E2EA adopts the axial attention mechanism to decompose the attention matrix output by PLM into horizontal and vertical axis directions to calculate the attention distribution respectively, and mine the multi-hop path and non-local dependency information between entities.

3. The DocRE method based on three-angle attention and entity-relationship relevance according to claim 1, characterized in that: In step S2, entity pair-context attention (E2CA) uses entity pair embedding as query and context embedding as key / value pair to capture local semantic associations through attention calculation.

4. The DocRE method based on three-angle attention and entity-relationship relevance according to claim 1, characterized in that: In step S2, entity pair-evidence attention E2EV extracts the contextual position of entity pair attention based on the token-level attention output of PLM, maps it to the sentence level, and constructs sentence-level evidence attention.

5. The DocRE method based on three-angle attention and entity-relationship relevance according to claim 1, characterized in that: In step S3, the specific operations of feature alignment and weighted fusion are as follows: first, the features of the three attention types E2CA, E2EV and E2EA are spliced ​​together, and then through multiple layers of linear transformation and nonlinear activation functions, the weight distribution of different attention types in the final fusion representation is learned to generate a unified entity pair semantic representation.

Citation Information

Patent Citations

  • Dense object image detection method based on multiple regression and adaptive focus loss

    CN115272652A

  • Document-level relation extraction method based on natural language processing

    CN119938927A