Document-level relation extraction method based on natural language processing
By adopting multi-grained feature fusion, axial attention mechanism and evidence enhancement module methods in document-level relationship extraction, the problem that traditional methods are difficult to capture complex entity relationships is solved, the accuracy and adaptability of relationship extraction is improved, and higher quality data support is provided for related downstream tasks.
Patent Information
- Application Number
- CN202510024751.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional sentence-level relationship extraction methods are difficult to capture complex entity relationships across sentences or segments, and the context information of entities in long texts may be scattered, increasing the difficulty of information extraction.
A document-level relationship extraction method is adopted for multi-grained feature fusion, axial attention mechanism and evidence enhancement modules. Through the multi-grained feature fusion module combining local and global features, the axial attention mechanism captures richer contextual relationships, and the evidence enhancement module generates more reliable relationship representations.
It improves the accuracy and adaptability of relationship extraction in long texts, and provides a higher quality data foundation for downstream tasks such as intelligent question-and-answer, knowledge graph construction and information retrieval.
Smart Images

Figure CN119938927A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing. Specifically, the present invention designs a method for extracting entity relationships from long text documents. Background Art
[0002] Document-level Relation Extraction (DocRE) is a key technology in natural language processing, which is used to identify and extract the relationship between entities from long documents. With the rapid development of information technology and the surge in the number of Internet users, massive semi-structured and unstructured data are rapidly accumulated on the Internet. These data often exist in the form of free text, which contains rich potential information. In order to extract valuable knowledge from these data, relation extraction technology came into being.
[0003] Traditional sentence-level relationship extraction methods are limited to single short sentences and have difficulty capturing complex entity relationships across sentences or paragraphs. With the growth of text information and the increasing demand for knowledge graph construction in fields such as enterprises and scientific research, document-level relationship extraction has gradually become one of the key technologies for realizing knowledge systems. In practical applications, document-level relationship extraction faces challenges. The contextual information of an entity may exist in multiple paragraphs or even the entire document, and may contain interference information, which increases the difficulty of information extraction. At the same time, the complexity of the relationship requires the relationship between different entity pairs to be judged in combination with contextual information, which places high demands on the accuracy and generalization ability of the model.
[0004] This paper proposes an innovative document-level relationship extraction method that integrates advanced technical means such as multi-granularity feature fusion, axial attention mechanism and evidence enhancement module to meet the challenges in long text relationship extraction. The multi-granularity feature fusion module can combine local and global features to capture the full range of semantic information of the document; the axial attention mechanism can capture richer contextual relationships through a multi-level attention mechanism; the evidence enhancement module generates a more reliable relationship representation by dynamically analyzing the relationship between entity pairs and their context, providing more accurate support for document-level relationship extraction. While this method improves the accuracy and adaptability of relationship extraction in long texts, it also provides a higher quality data foundation for downstream tasks such as intelligent question answering, knowledge graph construction and information retrieval. Summary of the invention
[0005] In view of the existing problems, this paper proposes a document-level relationship extraction method based on multi-granularity feature fusion, axis attention mechanism and evidence enhancement. This paper aims to improve the performance of document-level relationship extraction. The specific steps are as follows:
[0006] Step 1: Use the encoder to encode the text content, including word embedding, position encoding, and segment encoding, and synthesize the input representation.
[0007] Step 1.1: Split the input text into words or subword units, and use the embedding matrix to map each word or subword into a vector representation to obtain an embedding representation containing rich semantic information;
[0008] Step 1.2: Add positional encoding to word embeddings to provide sequential information. Sine and cosine functions are used and the resulting vector is added to the word embedding to ensure that the encoder can use the sequential information to process inputs at different positions.
[0009] Step 1.3: When processing sentence pair tasks, use segment encoding to distinguish different sentences, add word embedding, position encoding, and segment encoding to get the final representation of each word.
[0010] Step 2: Through multi-granular feature extraction and fusion strategy, local segment features and global semantic features are combined to form feature representation.
[0011] Step 2.1: Extract n-gram features through convolutional neural network (CNN) to capture local contextual relationships and obtain local segment features.
[0012] Step 2.2: Capture the overall semantics of the text, use BERT to generate global features, and combine local segment features and global semantic features to form the final feature representation.
[0013] Step 3: Perform axial attention mechanism calculation on the text, and obtain the neighborhood relationship information of entity pairs layer by layer through self-attention calculation along the horizontal and vertical axes.
[0014] Step 3.1: Concatenate the entity representations corresponding to a specific relationship type to form an adjacency matrix for entity pairs.
[0015] Step 3.2: Map the arranged entity embeddings to the hidden state space through linear layers and nonlinear activation functions, obtain the representations of the head entity and the tail entity respectively, and perform structured representations of entity pairs under different relationship types.
[0016] Step 3.3: Apply bilinear function and sigmoid activation function to calculate the relationship probability g (s,o,r) .
[0017] Step 4: Extract contextual evidence information of entity pairs through the evidence enhancement module to further enrich the semantic representation of entity pairs.
[0018] Step 4.1: Extract contextual information fragments that may support the target relation from the document, such as sentences, sub-paragraphs, etc. containing the target entity pair or relation.
[0019] Step 4.2: Fuse the extracted evidence information with the features of the target entity pair.
[0020] Step 5: Classify the fused features through the relation classification model to generate relation labels for entity pairs.
[0021] Step 5.1: Introduce the threshold class TH to learn a specific threshold for each entity pair, and dynamically determine the threshold according to the characteristics of each pair of entities to overcome the limitations of the traditional global fixed threshold.
[0022] Step 5.2: Design a special loss function that takes into account the existence of the threshold class TH and is based on the standard categorical cross entropy loss function.
[0023] Step 5.3: Use the hyperparameter λ to balance the relation loss (Relation Extraction, RE) and evidence loss (Evidence Retrieval, ER).
[0024] Step 5.4: Calculate the score of each possible relationship type, and select the relationship type with the highest score as the final relationship output based on the maximum score and the threshold. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The attached figure is a diagram of the model structure. DETAILED DESCRIPTION
[0026] The following describes an implementation of document-level relationship extraction based on natural language processing with reference to the accompanying drawings.
[0027] As shown in the accompanying drawings, the present invention proposes a document-level relationship extraction method, which realizes document-level relationship extraction through multi-step encoding and feature processing, and improves the text comprehension ability and the accuracy of relationship classification. The encoder is used to preprocess and encode the text, generate an input representation including word embedding, position encoding and segment encoding, and integrate the text information into the input structure required by the model. A multi-granularity feature fusion strategy is adopted to combine local fragment features with global semantic features, so that the model can carefully understand the details and overall context in the text. The semantic hierarchical relationship is captured through the axial attention mechanism, so that the model pays attention to the multi-level connections between entity pairs at different levels, and improves the model's relationship extraction performance for complex and long texts. An evidence enhancement module is designed to obtain richer contextual information by strengthening the representation of entities. Through the relationship classification method of adaptive threshold, the mapped relationship embedding vector is generated, the entity relationship is classified, and the final extraction result is output.
[0028] Step 1: Use the encoder to encode the text content, including word embedding, position encoding, and segment encoding, and synthesize the input representation.
[0029] Step 1.1: Given a document d = [x1, x2, ..., x L ], where L is the number of tokens in the document. A special token “*” is inserted before and after each entity mention to highlight the entity during the encoding process. The encoder obtains the embedding representation H = [h1, h2, ..., h L ], mapping each word or subword into a vector space of fixed dimension.
[0030] Step 1.2: Add positional encoding to word embedding to provide order information, using sine and cosine functions, calculated as follows:
[0031]
[0032] Among them, PE represents position embedding, pos represents the position of the word in the sentence, and d model is the dimension of the word embedding, 2i is an even dimension, and 2i+1 is an odd dimension. The generated vector is added to the word embedding to ensure that the encoder can use the sequential information to process the input at different positions.
[0033] Step 1.3: When dealing with sentence pair tasks, use segment encoding to distinguish different sentences and help the model understand the relationship between sentences. Add word embedding, position encoding, and segment encoding to get the final representation of each word.
[0034] Step 2: Through multi-granular feature extraction and fusion strategy, local segment features and global semantic features are combined to form feature representation.
[0035] Step 2.1: Extract n-gram features through convolutional neural network (CNN) to capture local contextual relationships. The convolution operation can be expressed as:
[0036] H t =CNN(t)=max_pool(ReLU(W t X+b t )) (3)
[0037] Among them, W t is the convolution kernel, X is the input sequence, b t is the bias term. The activated features are pooled to obtain the local segment features h n :
[0038] h n =MaxPool(H t ) (4)
[0039] Step 2.2: Capture the overall semantics of the text and use BERT to generate global features h clsCombining entity information with global features, a fusion representation is obtained through a weighted mechanism:
[0040] h' cls =Softmax(W cls [e s ;e o ;h cls ]) (5)
[0041] Combine local segment features and global semantic features to form the final feature representation:
[0042] H = Concat(h n ,h' cls ) (6)
[0043] Step 3: Use the axis attention mechanism to decompose the text sequence into multiple axes, effectively handle long-distance dependencies, and improve the model's ability to model long-distance dependencies.
[0044] Step 3.1: Concatenate the entity representations corresponding to a specific relationship type to form an adjacency matrix for entity pairs.
[0045] Step 3.2: Map the arranged entity embeddings to the hidden state space through linear layers and nonlinear activation functions, obtain the representations of the head entity and the tail entity respectively, and perform structured representations of entity pairs under different relationship types. The details are as follows:
[0046]
[0047] in These are all learnable parameters in the model.
[0048] Step 3.3: To reduce the number of parameters of the bilinear classifier, the embedding dimension is divided into k groups of equal size. In each group, the bilinear function and sigmoid activation function are applied to calculate the relationship probability.
[0049]
[0050]
[0051] in, is a learnable parameter, g (s,o,r) It is an entity pair The probability of reaching relation r.
[0052] Step 3.4: After obtaining the embedding of entity pairs under different relation types, use axial attention to enhance the axial neighborhood information of each entity pair. Axis attention is calculated by self-attention along the horizontal and vertical axes, and a residual connection is added for each calculation along the axis.
[0053]
[0054] where q (s,o,r) , k (s,o,r) , v (s,o,r) W q g (s,o,r) , W k g (s,o,r) , W v g (s,o,r) The linear transformation results in is the learnable weight matrix of the model.
[0055] Step 4: Extract contextual evidence information of entity pairs through the evidence enhancement module to further enrich the semantic representation of entity pairs.
[0056] Step 4.1: Capture the s , e o ) Related context c (s,o) , calculate its contextual features based on the pre-trained encoded attention matrix A:
[0057]
[0058] c (s,o) =H T q (s,o) (17)
[0059] in is the Hadamard product, A s is for all the head entities in the document s The relevant attention, take e s The average of the mention-level attention.
[0060] Step 4.2: Analyze the sentence to determine whether it is a specific entity pair (e h , e t ) is the evidence sentence. Using the LogSumExp pooling technique, all tokens in the sentence are integrated into a sentence embedding s n To capture information related to the relationship between entity pairs.
[0061]
[0062] Step 4.3: Use a bilinear function to evaluate the relationship between sentence embedding and context embedding to measure the importance of the sentence to entity pair. The formula is as follows:
[0063] P(s n |e h , e t )=σ(s n Wv c h;t +b v ) (19)
[0064] Among them, W v and b v is a learnable parameter and σ is the activation function.
[0065] Step 4.4: Use binary cross entropy as the objective function to train the model to predict whether a sentence is an evidence sentence. The formula is as follows:
[0066]
[0067] Among them, y n is an evidence label, when s n ∈V h;t 1 if the value is 0, otherwise it is 0.
[0068] Step 5: Classify the fused features through the relation classification model to generate relation labels for entity pairs.
[0069] Step 5.1: Introduce the threshold class TH to learn a specific threshold for each entity pair, and dynamically determine the threshold according to the characteristics of each pair of entities to overcome the limitations of the traditional global fixed threshold.
[0070] Step 5.2: Design a special loss function. This loss function takes into account the existence of the threshold class TH and is designed based on the standard classification cross entropy loss function. The loss function can be decomposed into two parts, and the calculation formula is as follows:
[0071]
[0072] Step 5.3: Use the hyperparameter λ to balance the relation extraction loss (RE) and evidence retrieval (ER). The loss function is calculated as follows:
[0073] L=L RE +λL ER (twenty two)
[0074] Step 5.4: Calculate the score of each possible relationship type, and select the relationship type with the highest score as the final relationship output based on the maximum score and the threshold.
Claims
1. A document-level relationship extraction method based on natural language processing, characterized in that: Extracting entity relationships from a document, the method comprises the following steps: Text input encoding: Use an encoder to encode the text content, including word embedding, position encoding, and segment encoding, and synthesize the input representation; Multi-granularity feature fusion: Through multi-granularity feature extraction and fusion strategies, local segment features and global semantic features are combined to form feature representation; Axis attention mechanism: Use the axis attention mechanism to decompose the text sequence into multiple axes, effectively handle long-distance dependencies, and improve the model's ability to model long-distance dependencies; Evidence enhancement module: The evidence enhancement module extracts contextual evidence information of entity pairs to further enrich the semantic representation of entity pairs; Relation classification: The fused features are classified through the relation classification model to generate relation labels for entity pairs.
2. The document-level relationship extraction method based on natural language processing according to claim 1 is characterized in that: The method includes converting text content into vector encoding through an encoder, and the method is as follows: given a document d = [x1, x2, ..., x L ], where L is the number of tokens in the document. A special token "*" is inserted before and after each entity mention to highlight the entity during the encoding process. The encoder obtains the embedding representation H = [h1, h2, ..., h L ], mapping each word or subword into a vector space of fixed dimension; Positional encoding is added to word embeddings to provide sequential information. Sine and cosine functions are used and the resulting vector is added to the word embedding to ensure that the encoder can use sequential information to process inputs at different positions. When processing sentence pair tasks, segment encoding is used to distinguish different sentences, and the word embedding, position encoding and segment encoding are added together to obtain the final representation of each word.
3. The document-level relationship extraction method based on natural language processing according to claim 1 is characterized in that: The invention comprises a multi-granularity feature extraction and fusion strategy, combining local segment features and global semantic features to form a feature representation. The method is specifically as follows: extracting n-gram features through a convolutional neural network (CNN), capturing local contextual relationships, and obtaining local segment features; capturing the overall semantics of the text, using BERT to generate global features, and combining local segment features and global semantic features to form a final feature representation.
4. The document-level relationship extraction method based on natural language processing according to claim 1 is characterized in that: The method includes obtaining neighborhood relationship information of entity pairs layer by layer through self-attention calculation along the horizontal axis and the vertical axis. The method is specifically as follows: concatenating entity representations corresponding to specific relationship types to form an adjacency matrix of entity pairs; mapping the arranged entity embeddings to the hidden state space through a linear layer and a nonlinear activation function, respectively obtaining representations of the head entity and the tail entity, and performing structured representations of entity pairs under different relationship types; applying a bilinear function and a sigmoid activation function to calculate the relationship probability g (s,o,r) .
5. The document-level relationship extraction method based on natural language processing according to claim 1 is characterized in that: The method comprises utilizing an evidence enhancement module to obtain deeper semantic relationships by aggregating entity information within a multi-hop neighborhood. The method is specifically as follows: extracting context information fragments that may support a target relationship from a document, such as sentences, sub-paragraphs, etc. containing a target entity pair or relationship; and fusing the extracted evidence information with the features of the target entity pair.
6. The document-level relationship extraction method based on natural language processing according to claim 1 is characterized in that: The inclusion relation classification adopts an adaptive threshold mechanism to adaptively classify entity pairs according to different relation types. The method is specifically as follows: a threshold class TH is introduced to learn a specific threshold for each entity pair, and the threshold is dynamically determined according to the characteristics of each pair of entities to overcome the limitations of the traditional global fixed threshold; a special loss function is designed, considering the existence of the threshold class TH, and is designed based on the standard classification cross entropy loss function; a hyperparameter λ is used to balance the relation loss (Relation Extraction, RE) and evidence loss (Evidence Retrieval, ER); the score of each possible relation type is calculated, and the relation type with the highest score is selected as the final relation output based on the maximum score and the threshold judgment.
Citation Information
Cited By
DocRE method based on three-angle attention and entity-relation correlation
CN120407806A
DocRE method based on three-perspective attention and entity-relation correlation
CN120407806B
Document-level saliency detection method and related device
CN121542806A