A Document-Level Relation Extraction Method Based on Entity Structure Encoding and Two-Classifications

By introducing entity structure information encoding and two-classification methods in document-level relationship extraction, the problem of existing methods ignoring entity structure information and semantic dependencies is solved, and a more accurate and efficient document-level relationship extraction effect is achieved, providing powerful triple support for knowledge graph construction.

CN117076676BActive Publication Date: 2025-06-24UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311060015.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-22
Publication Date
2025-06-24
Estimated Expiration
2043-08-22

AI Technical Summary

Technical Problem

Existing document-level relationship extraction methods ignore structural information and semantic dependencies between entities, making it difficult to extract some potential triples.

Method used

Using the method based on entity structure coding and two-class classification, the self-attention mechanism of entity structure information is introduced into the pre-trained language model, the semantic capture of entity embedding is enhanced, and preliminary and quadratic classification is performed through bilinear classification functions to improve the accuracy of document-level relationship extraction.

Benefits of technology

It effectively reduces the overhead of manual extraction, provides triple support for the construction of large-scale knowledge graphs, and improves the extraction effect of potential triple in documents, especially when processing complex documents, showing stronger semantic understanding capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117076676B_ABST
    Figure CN117076676B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of natural language processing, and specifically provides a document-level relation extraction method based on entity structure encoding and two-stage classification to solve the problem that some potential triples are difficult to extract due to the neglect of the structural information between entities and the semantic dependence relationship between entity pairs in document corpora. The present invention designs a new relation extraction framework, and adopts two-stage classification to extract simple triples and potential triples respectively; after encoding the document entities, pre-classify the spliced entity pairs, and extract the easily classifiable simple triples based on an improved adaptive threshold loss function; then use the pre-classification result as auxiliary inference information to enhance the entity representation and perform the second classification, which can effectively improve the extraction effect of potential triples in the document; in summary, the present invention can automatically extract the relationship between specified entities according to the input document.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of natural language processing, and specifically provides a document-level relation extraction method based on entity structure encoding and two-stage classification. Background Art

[0002] Information extraction is an important research content in natural language processing, which is used to extract valuable information from unstructured or semi-structured texts; as an important branch of information extraction, the goal of relation extraction is to determine whether there is a relationship between two given entities in the text and what the type of the relationship is, so as to generate a triple in the form of "head entity - relation - tail entity", which is convenient for constructing a large-scale knowledge graph. According to different extraction corpora, relation extraction can be divided into sentence-level relation extraction and document-level relation extraction. These two types of methods respectively judge the relationship between two given entities from texts of different lengths to form triple knowledge; the obtained knowledge can be used to construct a knowledge graph, or provide knowledge support for an intelligent question-answering system, or for an information retrieval system, and has wide application value.

[0003] At present, although great progress has been made in sentence-level relation extraction methods, in practical applications, a large number of relations exist in document-level corpora, such as Wikipedia articles or Baidu Baike articles; conventional sentence-level methods cannot take a document as input and lack the ability to jointly infer entity relations from multiple sentences. Compared with sentence-level relation extraction, document-level relation extraction has more difficulties. On the one hand, a document contains a large number of entities and relations, and the entities are scattered, so cross-sentence extraction is required; on the other hand, there are many reasoning phenomena in document corpora, and to fully mine the triples in the document, it is necessary to consider modeling the reasoning phenomena. Existing document-level relation extraction methods can be divided into sequence-based methods, graph-based methods, and pre-trained language model-based methods. These methods model long texts in different ways to capture the semantic dependency relationships between entities, but none of them introduce more research on the reasoning phenomena in the document-level relation extraction task, especially ignoring the fact that the prediction difficulty between entity pairs is different. The relationships between some simple entity pairs can be extracted only by template matching. However, there are also some relationships between entity pairs that need to be extracted by logical reasoning, and predicting the relationships between them often depends on the assistance of other simple triples; moreover, most pre-trained language model-based methods do not consider the structural information between entities, and ignoring the dependent structural relationships between entities is not conducive to mining deep semantic relationships. Summary of the Invention

[0004] In view of the many problems existing in the background art, the purpose of the present invention is to provide a document-level relation extraction method based on entity structure encoding and two-stage classification, so as to solve the problem that some potential triples are difficult to extract due to the neglect of the structural information between entities and the semantic dependence relationship between entity pairs in the document corpus.

[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A document-level relation extraction method based on entity structure encoding and two-stage classification includes: a training process and a prediction process, and the specific steps are as follows:

[0007] A. Training process;

[0008] Step 1: Perform data preprocessing on the training data, and label the entities and each entity mention in the document by means of manual annotation or named entity recognition tools;

[0009] Step 2: Encode the entity structure information, add a self-attention learning mechanism for entity structure information to the pre-trained language model, enhance the semantic capture of entity word embeddings through the entity structure information, and obtain mention embeddings and entity embeddings;

[0010] Step 3: On the basis of the entity embeddings, use the context information features to enhance the hidden features of the head and tail entities in the entity pair, adopt a bilinear classification function as the pre-classifier and use an adaptive threshold loss function to complete the preliminary training. The pre-classifier makes a preliminary prediction on all entity pairs to obtain a set S of pre-classified entity pairs:

[0011] Step 4: Integrate the pre-classified entity pair information with the original entity embedding representation as inference information to enhance the entity embeddings:

[0012] Step 5: On the basis of the enhanced entity embeddings, adopt a bilinear classification function as the secondary classifier and use the cross-entropy loss function to complete the preliminary training. The pre-classifier performs secondary inference on all entity pairs to obtain a set S of secondary inference entity pairs hard ;

[0013] B. The prediction process is specifically as follows:

[0014] Perform the same data preprocessing on the document to be processed, use a named entity recognition tool to obtain the entities in the document to be processed, and use the trained pre-classifier and secondary classifier for pre-classification and secondary inference respectively to obtain a set S' of pre-classified entity pairs and a set S' of secondary inference entity pairs hard , and the two are integrated into the document-level relation S all : S all = S' ∪ S' hard .

[0015] Furthermore, the specific process of step 2 is as follows:

[0016] Step 21: Encode the entity structure information to construct the entity structure information matrix M;

[0017] Step 22: Add the entity structure self-attention mechanism to the Transformer encoder and use the entity structure information matrix M to guide the encoding;

[0018] Step 23: Use the BERT encoder to encode the document;

[0019] Use the tokenizer to tokenize the document to obtain a word sequence; add tags <m s > and <m e > before and after each mention of each entity to highlight the entity mention, and input the sequence into the BERT encoder to obtain the word embeddings of each word, and the set forms the embedding H of the document;

[0020] Step 24: Obtain the mention embedding and entity embedding according to the document encoding:

[0021] Take the tag <m s > before each mention as the embedding of this mention, and then obtain the set of mention embeddings of each entity denote the j-th mention of entity e i , denote the number of all mentions of entity e i ; for each entity, use the logsumexp function to aggregate all mention embeddings as the entity embedding:

[0022] Furthermore, the specific process of step 3 is as follows:

[0023] Step 31: Obtain the context information features;

[0024] Step 32: Based on the original entity embedding and context information, use the feed-forward neural network to obtain the head and tail entity hidden features;

[0025] Step 33: Use the enhanced head and tail entity hidden features as the input, use the bilinear function to form a pre-classifier and use the adaptive threshold loss function to complete the preliminary training. The pre-classifier outputs the predicted logits of all entity pairs and the logit of the threshold class TH;

[0026] Step 34: Sort the predicted logits of all entity pairs and the logit of the threshold class TH, and select the entity pairs with the predicted logit greater than the logit of the threshold class TH as the high-confidence entity pairs to form the pre-classified entity pair set S.

[0027] Further, the specific process of step 4 is as follows:

[0028] Step 41: Calculate the information aggregation weight w of the entity pair j ;

[0029] Calculate the probability p for each entity pair j :

[0030]

[0031] where logit j represents the predicted logit of the j-th entity pair, and max(logit j ) represents the maximum value of the predicted logits of all entity pairs;

[0032] Then calculate the information aggregation weight w j :

[0033]

[0034] Step 42: Aggregate entity pair information to enhance entity representation;

[0035] For entity e i , assume the set S i is the subset of all entity pairs containing entity e i ; for each entity pair therein, splice the embeddings of the head entity and the tail entity in the entity pair to obtain the inference path embedding of the entity pair, and aggregate all the inference path embeddings of the entity pair subset S j to obtain the inference information f i of entity e i : i :

[0036]

[0037] where represents the inference path embedding of the j-th entity pair in the set S i , and represents the size of the set S i ;

[0038] Use the inference information f i to expand the original entity embedding to obtain the enhanced entity embedding:

[0039] h′ i= tanh(W i h i +W f f i )

[0040] where W i and W fCorrespondingly represent the feature transformation weight matrix.

[0041] Furthermore, the specific process of step 5 is as follows:

[0042] Step 51: Use the bilinear classification function as the secondary classifier, and recalculate the probability P s , e o ) of the existence relationship for all entity pairs (e haasr (e s , e o ):

[0043] P hasr (e s , e o ) = sigmoid(z′ s W h z′ o + b n )

[0044] z′ s = tanh(W s h′ s )

[0045] z′ o = tanh(W o h′ o )

[0046] Among them, h′ s represents the enhanced entity embedding of the head entity, and h′ o represents the enhanced entity embedding of the tail entity; W h , W s , W o correspondingly represent the feature transformation weight matrix, and b n represents the bias term;

[0047] Step 52: Use the cross-entropy loss function as the secondary inference loss function L h , and optimize the secondary classifier;

[0048] L h = ∑(y (s,o) P hasr (e s , e o ) + (1 - y (s,o) )(1 - P hasr (e s , e o )));

[0049] Among them, y (s,o) represents the label.

[0050] Based on the above technical solutions, the beneficial effects of the present invention are as follows:

[0051] 1) The present invention can automatically extract the relationships between specified entities according to the input document, reduce the cost of manual extraction, and provide triple support for the construction of large-scale knowledge graphs.

[0052] 2) The present invention designs a new relationship extraction framework, which uses a two-classification method to extract simple triples and potential triples respectively; after encoding the document entities, the concatenated entity pairs are classified for the first time, and the simple triples that are easy to classify are extracted based on an improved adaptive threshold loss function; then the pre-classification results are used as auxiliary inference information to enhance the entity representation and perform the second classification, which can effectively improve the extraction effect of potential triples in the document.

[0053] 3) The present invention uses entity structure information to assist in guiding the encoder to learn the dependency relationships between entities, so that the encoding process pays more attention to the context information useful for the current entity relationship extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a schematic diagram of the training process of the document-level relationship extraction method based on entity structure encoding and two-classification in the present invention.

[0055] Figure 2 It is a schematic diagram of the principle of enhancing entity encoding using the entity structure information matrix in the present invention.

[0056] Figure 3 It is a schematic diagram of the principle of using a sliding window to process documents with lengths exceeding the input limit in the present invention.

[0057] Figure 4 It is a schematic diagram of the principle of entity feature enhancement in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] To make the objectives, technical solutions and beneficial effects of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and embodiments.

[0059] Embodiment:

[0060] This embodiment aims to propose a document-level relationship extraction method based on entity structure encoding and two-classification, to make up for the problems of existing methods that ignore the structural information between entities and are difficult to extract potential triples; its process is as Figure 1As shown, it includes a training process and a prediction process. The main steps of the training process are as follows: 1. Obtain document data samples and perform preprocessing, label the entities in the document, the mentions of the entities, and the relationships between the entities; 2. Use the entity structure information to guide the pre-trained language model to encode the document and obtain the embedded representations of the entities; 3. Concatenate the entity embeddings for any entity pair to obtain the entity pair embedded representation, and use the bilinear function to pre-classify the entity pair; 4. Obtain high-confidence triple information based on the classification results in step 3 and generate inference information; 5. Integrate the inference information with the original entity information to enhance the entity representation and perform a second classification; 6. Calculate the loss function and train the model.

[0061] A. The training process is as Figure 1 shown, and the specific implementation process is as follows:

[0062] Step 1. Training set acquisition and data preprocessing:

[0063] Step 11. Crawl the document content of multiple topic pages in the wiki encyclopedia, and use manual annotation or named entity recognition tools to label the entities in the article and the mentions of each entity;

[0064] An open-source named entity recognition tool such as Stanford CoreNLP can be used to process the document to obtain the named entities in each sentence, mark the same named entity that appears multiple times, and record the sentence number where it is located and its start and end positions in the sentence;

[0065] Step 12. Define the relationship types to be extracted, and label the entity pairs containing the relevant relationship types in the document as samples;

[0066] Step 2. Use the encoder to encode the document and obtain the embedded representations of the mentions and entities, as Figure 2 shown;

[0067] Step 21. Encode the entity structure information and construct an entity structure information matrix;

[0068] For a document sequence with n words, construct a two-dimensional entity structure information matrix M ∈ R n*n , initialize the matrix with 0, and assign 1 to the corresponding positions in the matrix according to the structural dependencies between the entities;

[0069] Specifically, if the i-th word tokeni and the j-th word token j satisfy any of the following conditions, then let M ij be 1:

[0070] 1) token i and tokenj The same mention belonging to the same entity,

[0071] 2)token i and token j different mentions belonging to the same entity,

[0072] 3)token i and token j mentions belonging to different entities but appearing in the same sentence;

[0073] Step 22, Improve the pre-trained language model by adding an entity structure self-attention mechanism to the Transformer encoder;

[0074] For each layer of the Transformer encoder, the original self-attention coefficient is calculated as follows:

[0075]

[0076] The self-attention coefficient of the added entity structure self-attention mechanism is as follows:

[0077]

[0078] where Q1, Q2, K1, K2 are trainable parameter matrices, and d k is the dimension of the feature; represents the original self-attention coefficient, represents the self-attention coefficient of the added entity structure self-attention mechanism;

[0079] The final attention coefficient is the sum of the two, that is

[0080] Step 23, Encode the document using the BERT encoder of the improved pre-trained language model:

[0081] Use a tokenization tool to tokenize the document to obtain a sequence of n words For each mention of each entity, add tags <m s > and <m e > before and after to highlight the entity mention, and input the sequence into the BERT encoder to obtain the word embeddings of each word, and the set forms the embedding H of the document;

[0082] Step 24, Obtain mention embeddings and entity embeddings according to the document encoding:

[0083] Take the tag <m s > before the position of each mention as the embedding of this mention, so that the word embeddings of different mentions of all entities are obtained, that is, there is a set of mention embeddings for each entity Denote the j-th mention of entity e i , and denote the number of all mentions of entity e i ; Considering the situation of coreference reasoning in the document, inferring the relationship between entities may require referring to the occurrence of each mention; Use the logsumexp function to aggregate the embeddings of all mentions of each entity as the embedding of the entity:

[0084]

[0085] Step 3: Classify entity pairs for the first time and obtain high-confidence triple information, that is, the set S of pre-classified entity pairs:

[0086] Step 31: Calculate context information for subsequent enhancement of head entity features and tail entity features;

[0087] When classifying the relationship between entity pairs, not only consider the path information between entities, but the document context where the entities are located also plays an important role; Use the multi-head attention matrix A ∈ R H×l×l of the last layer of the BERT encoder to calculate the context information that the head and tail entities jointly focus on. For the j-th word, A ijk represents the attention coefficient of this word to the k-th word in the i-th attention head; In step 24, the present invention uses the token <m s > before the position of each mention as the embedding of this mention. Similarly, use the attention matrix of this token as the attention matrix of this mention; For an entity with multiple mentions, use the average of the attention matrices of all mentions as the attention matrix of the entity, thereby obtaining the attention matrices of the head entity and the tail entity;

[0088] Multiply and normalize the attention matrices of the head and tail entities to obtain the attention matrix of the words that the head and tail entities jointly focus on, and finally multiply it by the context embedding to obtain the context information embedding c (h,t) ;

[0089] c (h,t) = Ha (h,t)

[0090]

[0091] where A h represents the attention matrix of the head entity, A t represents the attention matrix of the tail entity, H represents the embedding of the document; 1 T represents the transpose of the unit row matrix;

[0092] Step 32: Use a feed-forward neural network to obtain entity hidden features;

[0093] The original embedding h of the head entity s and the original embedding h of the tail entity o and the context information embedding c (h,t) are respectively input into the feed-forward neural network, activated by the tanh function, and the enhanced head entity feature z s and the tail entity feature z o are obtained:

[0094] z s = tanh(W s h s + W c c (h,t) )

[0095] z o = tanh(W o h o + W c c (h,t) )

[0096] where W s , W o , W c correspond to the feature transformation weight matrices, and are all trainable parameters;

[0097] Step 33: Use the groupbilinear function as a pre-classifier to classify the relationship of the entity pair (e s , e o ), where e s represents the head entity and e o represents the tail entity; the preliminary classifier outputs the predicted logits of all entity pairs and the logits of the additional threshold class TH;

[0098] First, the head entity feature z s and the tail entity feature z o are respectively divided into k groups of the same size, bilinear operations are performed respectively, and then summed, and the sigmoid function is used for classification; specifically expressed as:

[0099]

[0100]

[0101]

[0102] where represents the feature transformation weight matrix of the i-th pair of head and tail entity features for the relationship r, and W r i ∈ R d / k×d / k, b r represents the bias term and represents the i-th group of the head entity feature z s and the tail entity feature z o ; By classifying in this way, the model parameters can be effectively reduced from d 2 to d 2 / k. The classification result P(r|e s , e o ) represents the predicted probability of the output entity pair (e s , e o ) for each relationship r;

[0103] Train the pre-classifier using the adaptive threshold loss function;

[0104] In addition to outputting the predicted logits of all entity pairs, the model additionally outputs a logit for the threshold class TH: logit TH . Based on all logits, calculate the positive sample loss function L1 and the negative sample loss function L2 respectively, and the sum of the two constitutes the pre-classification loss function L pre :

[0105]

[0106]

[0107] where P T represents the set of positive sample labels, and N T represents the set of negative sample labels; L1 is the positive sample loss function, which drives the probability of each positive sample label to be greater than the threshold class TH; L2 is the negative sample loss function, which drives the probability of the threshold class TH to be greater than the probabilities of all negative sample labels;

[0108] To enable the model to have a certain classification ability at the beginning, first train the model using the pre-classification loss function for 5 epochs, and then train the model together with the subsequent secondary inference loss function;

[0109] Step 34, Obtain high-confidence entity pair information;

[0110] Calculate the output probabilities logit of all entity pairs including the threshold class TH, then sort the logit, and select the entity pairs with logit greater than that of the threshold class TH as the pre-classification results; if the logit of TH is the maximum value, it means that no relationship is predicted for this entity pair during the pre-classification stage; after the pre-classification is completed, all entity pairs with existing relationships can be obtained, forming a simple triple set S, and each entity pair expresses at least one relationship;

[0111] Step 4: Integrate the high-confidence entity pair information with the original entity embedding representation to enhance the original entity representation with the inference information, as shown in Figure 4 ;

[0112] Step 41: Calculate the information aggregation weight w of the entity pair j ;

[0113] Based on the rule that entity pairs with higher confidence should have higher weights, calculate the probability of the existence of a relationship for each entity pair:

[0114]

[0115] where logit j represents the predicted logit of the j-th entity pair, and max(logit j ) represents the maximum value of the predicted logits of all entity pairs;

[0116] Then use the softmax function to calculate the weight:

[0117]

[0118] Step 42: Aggregate the entity pair information to enhance the entity representation;

[0119] After pre-classification is completed, obtain the set S of pre-classified entity pairs. For entity e i , assume that the set is the subset of all entity pairs containing entity e i ; For each entity pair among them, concatenate the embeddings of the head entity and the tail entity in the entity pair to obtain the inference path embedding of the entity pair, and aggregate all the inference path embeddings of the entity pair subset S j to obtain the inference information f i of entity e i : i :

[0120]

[0121] where represents the inference path embedding of the j-th entity pair in set S i , and represents the size of set S i .

[0122] After obtaining the inference information f i , use this inference path information to expand the original entity embedding to obtain the enhanced entity embedding:

[0123] h′ i= tanh(W i h i +Wf f i )

[0124] Among them, W i 、W f correspond to the feature transformation weight matrices, which are trainable parameters;

[0125] Step 5: Perform the second classification, re-predict all samples to mine potential relationships, calculate the cross-entropy loss function, and train the classifier using the stochastic gradient descent method. Finally, fuse the result set with the pre-classified result set;

[0126] Step 51: Perform the second classification and re-predict the difficult triplets;

[0127] For the entity pair (e s , e o ), use the bilinear classification function and the sigmoid function to form a secondary classifier, and re-calculate the probability P hasr (e s , e o ) of its existence relationship:

[0128] P hasr (e s , e o ) = sigmoid(z′ s W h z′ o + b n )

[0129] z′ s = tanh(W s h′ s )

[0130] z′ o = tanh(W o h′ o )

[0131] Among them, h′ s represents the enhanced entity embedding of the head entity, and h′ o represents the enhanced entity embedding of the tail entity; W s , W o , W h represent the corresponding feature transformation weight matrices, which are trainable parameters; b n represents the bias term;

[0132] Step 52: Use the cross-entropy loss function as the secondary inference loss function to optimize the secondary classifier;

[0133]

[0134] Among them, Lh is the quadratic inference loss function; y (s,o) represents the label. If there is a relationship r between the entity pair (e s , e o ), then y (s,o) is 1, otherwise it is 0;

[0135] In this stage, the pre-classifier and the quadratic classifier are jointly optimized;

[0136] B. The prediction process is specifically as follows:

[0137] Use the trained document-level relation extraction model to extract entity relations in the actual document text. First, preprocess the document. After using the named entity recognition tool to obtain the entities in the document, perform pre-classification and quadratic classification respectively to obtain the simple triple set S′ and the potential triple set S′ hard , and the two are fused as the output of the model, that is: S all : S all = S′ ∪ S′ hard .

[0138] As described above, it is only the specific implementation manner of the present invention. Any feature disclosed in this specification, unless specifically described, can be replaced by other equivalent or similar-purpose alternative features; all the features disclosed, or all the steps in all the methods or processes, except for the mutually exclusive features and / or steps, can be combined in any way.

Claims

1. A document-level relation extraction method based on entity structure encoding and two-stage classification, comprising: The training process and the prediction process are as follows: A. Training process; Step 1: Perform data preprocessing on the training data, and label the entities and each entity mention in the document by means of manual annotation or named entity recognition tools; Step 2: Encode the entity structure information. Add a self-attention learning mechanism for entity structure information to the pre-trained language model, enhance the semantic capture of entity word embeddings through the entity structure information, and obtain mention embeddings and entity embeddings; Step 3: Based on the entity embeddings, use the context information features to enhance the hidden features of the head and tail entities in the entity pair. Use a bilinear classification function as a pre-classifier and an adaptive threshold loss function to complete the initial training. The pre-classifier makes a preliminary prediction on all entity pairs to obtain a set S of pre-classified entity pairs: Step 4: Integrate the pre-classified entity pair information with the original entity embedding representation as inference information to enhance the entity embeddings; Step 5. On the basis of enhancing entity embeddings, use a bilinear classification function as the secondary classifier and use the cross-entropy loss function to complete the preliminary training. The pre-classifier performs secondary inference on all entity pairs to obtain the secondary inference entity pair set S hard ; The specific process is as follows: Step 51: Use a bilinear classification function as the secondary classifier, and recalculate the probability P s , e o ) of the existence relationship for all entity pairs (e hasr , e s , e o ): P hasr (e s ,e o ) = sigmoid(z′ s W h z′ o + b n ) z′ s =tanh(W s h′ s ) z′ o =tanh(W o h′ o ) Among them, h' s represents the enhanced entity embedding of the head entity, and h' o represents the enhanced entity embedding of the tail entity; W h and W s and W o correspondingly represent the feature transformation weight matrices, and b n represents the bias term; Step 52: Use the cross-entropy loss function as the secondary inference loss function L h , and optimize the secondary classifier; L h = ∑(y (s,o) P hasr (e s , e o )) + (1 - y (s,o) )(1 - P hasr (e s , e o ))); where y (s,o) represents a label; B. The prediction process is specifically as follows: Perform the same data preprocessing on the document to be processed, use a named entity recognition tool to obtain the entities in the document to be processed, and use the trained pre-classifier and secondary classifier for pre-classification and secondary inference to obtain the pre-classification entity pair set S′ and the secondary inference entity pair set S′ respectively hard , and the two are fused into the document-level relationship S all : S all = S′ ∪ S′ hard .

2. The method for document-level relation extraction based on entity structure encoding and two-stage classification according to claim 1, wherein The specific process of Step 2 is: Step 21: Encode the entity structure information and construct an entity structure information matrix M; Step 22: Add an entity structure self-attention mechanism to the Transformer encoder and use the entity structure information matrix M to guide the encoding; Step 23: Use the BERT encoder to encode the document; Use a tokenization tool to tokenize the document to obtain a sequence of words; add tags <m s > and <m e > before and after each mention of each entity to highlight the entity mentions, and input the sequence into the BERT encoder to obtain the word embeddings of each word, and the set constitutes the embedding H of the document; Step 24: Obtain mention embeddings and entity embeddings according to the document encoding: Take the tag <m before each mention s > as the embedding of that mention, and thus obtain a set of mention embeddings for each entity Denote the j-th mention of entity e i ; Denote the number of all mentions of entity e i ; For each entity, aggregate all mention embeddings using the logsumexp function as the entity embedding:

3. The method for document-level relation extraction based on entity structure encoding and two-stage classification according to claim 1, characterized in that, The specific process of Step 3 is: Step 31: Obtain context information features; Step 32: Based on the original entity embeddings and context information, use a feed-forward neural network to obtain the hidden features of the head and tail entities; Step 33: Use the enhanced hidden features of the head and tail entities as the input, use a bilinear function to form a pre-classifier and an adaptive threshold loss function to complete the initial training. The pre-classifier outputs the prediction logits of all entity pairs and the logits of the threshold class TH; Step 34: Sort the prediction logits of all entity pairs and the logits of the threshold class TH, and select the entity pairs with prediction logits greater than the logits of the threshold class TH as high-confidence entity pairs to form a set S of pre-classified entity pairs.

4. The method for document-level relation extraction based on entity structure encoding and two classifications according to claim 1, wherein The specific process of Step 4 is: Step 41, calculate the information aggregation weight w of the entity pair j ; Calculate the probability p for each entity pair j : where, logit j represents the predicted logit of the j-th entity pair, and max(logit j ) represents the maximum value of the predicted logits of all entity pairs; Furthermore, calculate the information aggregation weight w j : Step 42: Aggregate entity pair information to enhance the entity representation; For entity e i , assume that set S i is the subset of entity pairs that all contain entity e i ; for each entity pair among them, splice the embeddings of the head entity and the tail entity in the entity pair to obtain the inference path embedding of the entity pair, and aggregate all the inference path embeddings of the entity pair subset S j according to the weight w i to obtain the inference information f i of entity e i : Among them, represents the inference path embedding of the j-th entity pair in the set S i , and represents the size of the set S i ; Utilize the inference information f i Expand the original entity embedding to obtain an enhanced entity embedding: h′ i = tanh(W i h i + W f f i ) Among them, W i and W f correspondingly represent the feature transformation weight matrices.

Citation Information

Patent Citations

  • Network document relationship extraction method and system

    CN115357775A

  • System and method for relation extraction with adaptive thresholding and localized context pooling

    US20220121822A1