A Document-Level Relation Extraction Method Based on Long-Tail Data Distribution

Through the data augmentation and comparative learning pre-training method based on relationship encoding, the problem of long-tail data distribution in document-level relationship extraction is solved, the prediction accuracy of tail relationships is improved and the calculation efficiency is optimized, which is suitable for document-level relationship extraction tasks.

CN114861645BActive Publication Date: 2025-07-08ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210469592.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-28
Publication Date
2025-07-08
Estimated Expiration
2042-04-28

AI Technical Summary

Technical Problem

When the existing document-level relationship extraction method based on deep learning is trained on long-tail data distribution, the model overfits the head relationship prediction, and cannot accurately predict complex and rare tail relationships. The traditional data augmentation method has low computational efficiency and affects the head relationship performance.

Method used

Using the data augmentation method based on relationship encoding and the comparative learning pre-training method, a new ternary vector group is generated by random perturbing entity pairs, and a model training method combining adaptive threshold and momentum update is improved to improve the prediction accuracy of the tail relationship and optimize the calculation efficiency.

Benefits of technology

The prediction accuracy of the document-level relationship extraction model in tail relationship type is improved, and the overall performance of the model in long-tail data distribution is improved, especially in scenarios with limited hardware resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114861645B_ABST
    Figure CN114861645B_ABST
Patent Text Reader

Abstract

The present invention discloses a document-level relation extraction method based on long-tail data distribution, belonging to the fields of information extraction and machine learning. It includes document preprocessing, document encoding, relation encoding, data augmentation, and relation prediction. In terms of data augmentation, for the set of labeled triple vectors, the present invention randomly selects or presets the relation types that need to be augmented, designs a mask vector, and perturbs the pooled context representation in the original triple vectors to be augmented with data to generate new triple vectors; it can effectively improve the accuracy of the document-level relation extraction model in predicting tail relation types. At the same time, compared with the traditional text-based data augmentation method, the present invention does not require an additional text encoding process, improving the computational efficiency of model training. In addition, the contrastive learning pre-training framework proposed by the present invention can effectively improve the accuracy of document-level relation extraction in the long-tail data distribution scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of information extraction and machine learning, and particularly to a document-level relation extraction method based on long-tail data distribution. Background Art

[0002] Relation extraction plays a crucial role in information extraction, which aims to predict the relationships between entities in text. Early relation extraction work mainly focused on sentence-level relation extraction, that is, predicting entity relationships from a single sentence. With the development of deep learning technology, neural relation extraction methods have achieved good results in sentence-level relation extraction. Recently, the research of relation extraction has developed to document-level relation extraction, which is a more practical and challenging scenario than sentence-level relation extraction.

[0003] In the document-level relation extraction task, the relationship patterns between entity pairs across different sentences are often more complex, and the distances between these entity pairs are also relatively long. Therefore, document-level relation extraction requires the model to find relevant contexts and perform reasoning across sentences, rather than memorizing simple entity relation patterns in a single sentence. In addition, in document-level relation extraction, multiple entity pairs coexist in a single document, and each entity may be mentioned more than once in a sentence. Therefore, document-level relation extraction also requires the model to extract the relationships of multiple entity pairs from a single document at one time. To solve the above problems, in recent years, document-level relation extraction methods based on deep learning have achieved preliminary results.

[0004] However, the existing document-level relation extraction methods based on deep learning ignore an important problem, that is, the long-tail data distribution. The long-tail distribution is a common phenomenon in real-world data. In the document-level relation extraction task, the long-tail distribution phenomenon has also been observed. In the most commonly used document-level relation extraction dataset DocRED: the 7 most frequently occurring entity relation types account for 55.12% (there are a total of 96 entity relation types in this dataset); there are 60 entity relation types with an occurrence frequency of less than 200 times. Therefore, directly training the document-level relation extraction method based on deep learning on long-tail data will result in the model achieving good performance in predicting head relations (frequently occurring entity relations), but the insufficient fitting on tail relations (entity relations with low occurrence frequencies) leads to a low prediction accuracy. Although the overall performance of the document-level relation extraction method depends largely on the performance of head relations because they are in the majority, the prediction failure on tail relations is an important problem in the real-world document-level relation extraction scenario, which will cause the model to only be able to predict some simple and common entity relations, but not accurately predict complex and rare entity relations.

[0005] Data Augmentation is a common strategy to address the long-tail problem. Nevertheless, it is not easy to effectively apply data augmentation in document-level relation extraction. Ordinary data augmentation operations on documents, including randomly deleting or replacing text, require the document-level relation extraction model to perform an additional encoding process on the entire document, which results in very low computational efficiency because a document usually contains many sentences and words. In addition, tail relations and head relations may coexist in a document. Therefore, if the above ordinary text augmentation method is used to augment tail relations, the head relations will also be augmented, which may cause the model to overfit on head relations, thus reducing the model's prediction performance on head relations. Summary of the Invention

[0006] In view of the problems existing in the above-mentioned prior art, the present invention proposes a document-level relation extraction method based on long-tail data distribution.

[0007] A document-level relation extraction method based on long-tail data distribution includes the following steps:

[0008] Step 1: Document preprocessing

[0009] Annotate all entities in the given document and annotate special characters at the entity boundaries as a mention of the entity in the document;

[0010] Step 2: Document encoding

[0011] Take the preprocessed document as the input of the pre-trained Transformer model to obtain the context semantic representation of all characters in the document as vector encoding, and obtain the self-attention matrix between entities;

[0012] Step 3: Relation encoding

[0013] Traverse entity pairs formed by pairwise entities; according to the document encoding results, calculate the vector representation of each entity in the document and the pooled context representation of the entity pair to form a triple vector group; in the model training stage, it is necessary to label the relationship label of each entity pair and execute Step 4; in the actual prediction stage, directly execute Step 5;

[0014] Step 4: Data augmentation

[0015] For the set of labeled triple vector groups, randomly select or preset the relation types that need to be augmented, design a mask vector, perturb the pooled context representation in the original triple vector groups to be data-augmented to generate new triple vector groups; use the original triple vector group set and the triple vector group set obtained by data augmentation as the training set to train and obtain a document-level relation extraction model;

[0016] Step 5: Relation prediction

[0017] Preprocess, encode the document, and encode the relationships for the given document using the methods in steps 1 - 3. Use the trained document - level relationship extraction model to predict the relationships for the obtained triple - vector groups, and output the entity pairs with valid relationships and their belonging relationships.

[0018] Furthermore, the pre - trained Transformer model uses the BERT model.

[0019] Furthermore, step 2 is specifically as follows:

[0020] Input the annotated entities and the mentioned documents into the BERT model to obtain the context semantic representations H of all characters in the document and the self - attention matrix A.

[0021] Furthermore, step 3 is specifically as follows:

[0022] 3.1) Traverse entity pairs formed by pairwise entities;

[0023] 3.2) According to the document encoding results, calculate the vector representations of each entity in the document; the first entity in the entity pair is called the head entity \(e_h\), and the head - entity vector is denoted as \(e\) h , the second entity in the entity pair is called the tail entity \(e_t\), and the tail - entity vector is denoted as \(e\) t ;

[0024] 3.3) Calculate the pooled context representation of the entity pair:

[0025] For the entity pair \((e_h,e_t)\), calculate the pooled context representation \(c\) h,t ;

[0026] 3.4) For the entity pair \((e_h,e_t)\), its triple - vector group representation is \(T\) h,t =(e h ,c h,t ,e t ), and obtain the set of all triple - vector group representations \(\varepsilon\) represents the entity set.

[0027] Furthermore, step 4 is specifically as follows:

[0028] 4.1) Set the set of relationship types for which data augmentation is required \(R\) is the set of all relationship types;

[0029] 4.2) Given an entity pair \((e_h,e_t)\), if its relationship from index the original triple - vector group representation \((e h ,ch,t , e t );

[0030] First, randomly generate a mask vector p, where each dimension of the mask vector is generated by a Bernoulli distribution with parameter p; then perform a dot product of the mask vector p and A h,t to apply a masking operation to it;

[0031] Furthermore, obtain the perturbed context representation vector c' h,t ;

[0032] Generate a new triple vector group representation (e h , c' h,t , e t ); Take the union of all the triple vector group representations generated by perturbation and the original triple vector group representation set as the training set

[0033] Furthermore, in step 4.2), set the number of times α of data augmentation, and at the same time generate α random masks, which can generate α perturbed context representation vectors That is, generate α new triple vector group representations.

[0034] Furthermore, when using the training set to train the document-level relation extraction model, adopt a model training method based on an adaptive threshold, specifically:

[0035] For any triple vector group representation First, apply two linear transformations with Tanh activation functions to fuse the context representation vector c h,t , the head entity vector e h and the tail entity representation vector e t , and the results after linear transformation are h and t;

[0036] After that, split the linearly transformed vectors h and t into k groups, and use grouped bilinear layers to calculate the score for each relation type r ;

[0037] In the training stage, by introducing the threshold class relation TH, dynamically learn the threshold θ of each entity pair h,t .

[0038] In the inference stage, the threshold of the valid relation score is score TH . If the predicted score for a certain relation is higher than score TH , then take it as a valid relation and output the corresponding triple vector group (e_h, r, e_t), and finally obtain the predicted triple set

[0039] Further, before performing model training based on the adaptive threshold, it also includes a process of contrastive learning pre-training for the document-level relation extraction model, specifically:

[0040] For the documents in the pre-training dataset, obtain the set of triple vector representations of all entity pairs according to the methods in steps 1-4 As the training set, for one of the triple vector representations Use two linear transformations to fuse the triple vector representations to obtain the linearly transformed vectors h and t;

[0041] Next, use a multi-layer perceptron with a ReLU activation function to fuse the relation representation x finally used in the contrastive learning stage;

[0042] Maintain a relation representation queue Q for each relation r r , in order to maintain the consistency of the relation representations in Q r , use the momentum update model to encode the positive and negative samples in the contrastive learning, and the original model is updated through backpropagation, and the momentum update model is updated through the momentum hyperparameter;

[0043] Next, input the document into in the same process as obtaining the relation representation x to get x′, and push all the relation representation sets into the relation representation queues Q according to their relation labels r The rule is: if the relation r holds between the entity pair (e_h, e_t), then x′ will be pushed into Q r ; finally, obtain the positive and negative relation representation sets of x from the queue;

[0044] For the relation representation x corresponding to the entity pair (e_h, e_t) used in the contrastive learning stage, calculate the contrastive learning loss.

[0045] The beneficial effects of the present invention are:

[0046] (1) The data augmentation method based on relation encoding proposed by the present invention can effectively improve the accuracy of predicting the tail relation type of the document-level relation extraction model. At the same time, compared with the traditional text-based data augmentation method, the present invention does not require an additional text encoding process, improving the computational efficiency of model training.

[0047] (2) The contrastive learning pre-training method for the document-level relation extraction model proposed by the present invention further improves the accuracy of predicting the tail relation type of the document-level relation extraction model, especially suitable for the scenario where the currently manually labeled data is limited. Description of the Drawings

[0048] Figure 1 is a schematic diagram of a document-level relation extraction method based on long-tail data distribution proposed by the present invention;

[0049] Figure 2 is a schematic diagram of a model pre-training method based on contrastive learning proposed by the present invention;

[0050] Figure 3 is a schematic diagram of the data distribution of the document-level relation extraction dataset DocRED shown in this embodiment. Detailed Description of the Invention

[0051] In order to further explain the technical details and advantages of the present invention, this embodiment will start with specific examples to show the details of the above steps.

[0052] A document-level relation extraction method based on long-tail data distribution proposed by the present invention can extract all existing relation triple sets from a given document.

[0053] For example, given a document where l is the length of the word sequence of the document. There are n entities annotated in the document, and this entity set is denoted as where each entity e_i has m mentions, and this set is denoted as Each mention is marked at the left and right boundary positions in the document. In addition, all possible relation type sets are also predefined in advance.

[0054] The goal of the document-level relation extraction model is: for a document with annotated entities and mentions, extract the triple set where each triple (e_h, r, e_t) extracted by the model can be interpreted as the relation r existing in the entity pair composed of the head entity e_h and the tail entity e_t.

[0055] In this embodiment, the document-level relation extraction model is divided into four parts: a document encoding module, a relation encoding module, a data augmentation module, and a relation prediction module. Combined with Figure 1 as shown, the calculation processes corresponding to each module are described respectively.

[0056] (1) Document Encoding Module.

[0057] The present invention uses a large-scale pre-trained Transformer model to obtain document vector encodings. The pre-trained Transformer model can obtain vector representations of the context semantics of text and has become the most commonly used and effective text encoding model in deep learning-based information extraction methods; the present invention uses the most common BERT model in the pre-trained Transformer model. In addition, in order to obtain entity relationship encodings, the present invention adds special "*" characters to the boundaries of all entities in the input text. Finally, by inputting the document into the BERT model, vector encodings can be obtained for each character in the document.

[0058] In a specific implementation of the present invention, the document with labeled entities and mentions is input into the BERT model to obtain the context semantic representations H of all characters in the document and the self-attention matrix A; which is expressed as:

[0059]

[0060] where Ptr(.) represents the pre-trained BERT model, H is the word vector output by the last layer of the BERT model, the matrix dimension is l×d, d is the dimension of the word vector, and l is the length of the word sequence of the document; A is the self-attention matrix in the last layer of the BERT model, and the matrix dimension is l×l×h, where h is the number of self-attention heads in the multi-head attention mechanism of the pre-trained language model.

[0061] In the art, using the BERT model as a text feature extraction network is the most commonly used method in the information extraction field in recent years, and it has been proven that using the BERT model to extract text information can greatly improve the performance of downstream tasks. In addition, the self-attention mechanism in the BERT model can help model the relationships between various characters in long texts, can help filter out information that is not important for the task, and help the model pay more attention to the most critical information.

[0062] (2) Relationship encoding module.

[0063] In this step, according to the document vector encoding obtained in the above-mentioned document encoding step, the entity pair relationship is represented as a triple vector group, namely, the head entity vector (head entity representation), the relation context representation vector (relation context representation), and the tail entity vector (tail eneity representation). Among them, the head entity vector / tail entity vector is obtained by taking the vector average of the vector encodings of the "*" characters on all the left boundaries of the head / tail entity in the document. The relation context representation vector is obtained by using the self-attention mechanism (self-attention) inside the Transformer model, that is, by obtaining the attention weights of the current head entity and tail entity to all other characters in the document, and further obtaining the weighted average of the vector representations of all characters in the document. After the processing of this step, the triple vector group representation of the relationship can be generated for all entity pairs in the document.

[0064] In a specific implementation of the present invention, according to the context semantic representation H of all characters in the document obtained during the document encoding process, and all mentions of each entity marked in the document First, calculate the vector representation m corresponding to each mention m_ij of each entity ij , and this vector representation is the word vector of the "*" character on the left boundary of this mention, obtained by indexing the H matrix;

[0065] For the vector representation m corresponding to each mention of each entity ij Use the pooling method of logsumexp to obtain the vector representation e of each entity i , and the calculation formula is:

[0066]

[0067] In the process of relation extraction, each entity in the entity set will form an entity pair with another entity for relation prediction. The first entity in the entity pair is called the head entity e_h, and the head entity vector is denoted as e h , the second entity in the entity pair is called the tail entity e_t, and the tail entity vector is denoted as e t .

[0068] The document-level relation extraction task requires the model to capture the dependencies between entities, mentions, and context words, and filter out unnecessary context information from long documents. The self-attention matrix A ∈ R obtained from BERT l×l×h has implicitly modeled the dependencies between entities, mentions, and context words, and can be used to obtain meaningful pooled context representations.

[0069] Given an entity pair \((e_h, e_t)\in\varepsilon\times\varepsilon\), the pooled context representation \(c\) of the entity pair can be obtained through the following two equations h,t :

[0070]

[0071] A h,t = A h * A t

[0072] where 1\in\mathbb{R} l×l ; A h,t is the product of the attention scores of all words in the document by the head entity \(e_h\) and the tail entity \(e_t\); A is the attention score of the head entity \(e_h\) for all words in the document h is the attention score of the tail entity \(e_t\) for all words in the document t t is the attention score of the tail entity \(e_t\) for all words in the document .

[0073] Taking A h as an example, it is obtained by averaging the attention scores A m_hj of all entity mentions \(m_{hj}\), specifically: Similar to obtaining the vector representation \(m\) corresponding to the context mention \(m_{ij}\), the present invention obtains the attention score A ij of the mention for other words in the document through the indexed self-attention matrix A m_hj . In addition, note that before performing the indexing, all attention heads in the self-attention matrix are first averaged, and the resulting matrix has a dimension of \(l\times l\). A t can also be calculated according to the same process.

[0074] Finally, for the entity pair \((e_h, e_t)\), a three-element vector representation \(T\) can be formed h,t = (e h , c h,t , e t ); where \(T\) h,t contains all the information for relation prediction and forms the basis for the data augmentation and contrastive learning methods in the present invention.

[0075] (3) Data Augmentation Module Based on Relation Encoding.

[0076] To address the long-tail problem in document-level relation extraction, the present invention proposes a data augmentation mechanism based on relation encoding (entity pair relation representation) to increase the frequency of tail relation types, so that the model can learn more fully in the prediction of tail relation types and improve the performance of the model under the long-tail data distribution.

[0077] The goal of this step is to perform data augmentation for tail relations, i.e., relations with lower occurrence frequencies, to increase the occurrence frequencies of tail relations. First, the user can define the set of relation types for which data augmentation is to be performed. Generally, the relation types for which data augmentation is needed are all tail relations. According to the triple vector representations of all entity-pair relations obtained in the above step of obtaining entity-pair relation representations, traverse all the triple vector representations. If, during the traversal, it is found that the relation type of the current triple vector representation belongs to the predefined set of relation types for which data augmentation is to be performed, then data augmentation is performed on the current triple vector representation. The specific algorithm for data augmentation is to keep the head entity vector and the tail entity vector in the current triple vector unchanged, and by applying random perturbations to the attention weights of the head entity and the tail entity to all other characters in the document in the step of obtaining entity-pair relation representations, thereby perturbing the relation context representation vector to generate a new triple vector representation. The newly generated triple vector representation and the original triple vector representation have the same relation type label. The above data augmentation method does not require re-encoding of the document, and when augmenting tail relations, it does not augment head relations at the same time, thus avoiding overfitting of head relations.

[0078] In a specific implementation of the present invention, after the relation encoding in step (2), the set of triple vector representations of all entity pairs is denoted as The set of relation types for which data enhancement is to be performed can be manually selected according to actual needs

[0079] Given an entity pair (e_h, e_t), if its relation First, from index the original triple vector representation (e h , c h,t , e t ). Since the pooled context representation c h,t encodes the context information for relation inference, and slight perturbations to the context do not affect the result of relation prediction. Therefore, the present invention generates a data-augmented triple vector representation by adding a small perturbation to c h,t .

[0080] Specifically, first randomly generate a mask vector Each dimension of this mask vector is generated by a Bernoulli distribution with parameter p; then multiply this mask vector p by the A h,t corresponding to the original triple vector calculated in the above step to perform a masking operation on it. The formula is:

[0081] A′ h,t = p * A h,t

[0082] Among them, A′ h,t is the masked attention score.

[0083] Applying a random mask to the attention score can be interpreted as randomly filtering out some context information. Since the attention score at the masked position is set to 0, the degree of perturbation can be controlled by setting an appropriate p.

[0084] The perturbed context representation vector c′ h,t is calculated as follows:

[0085]

[0086] Among them, the superscript T represents transpose.

[0087] In this embodiment, by setting the number of times α of data augmentation, α random masks can be generated simultaneously, and α perturbed context representation vectors can be generated through the same perturbation method as above Finally, a set of all triple vector representations generated through perturbation can be obtained Combining the set of triple vector groups of all original entity pairs and the set of triple vector groups generated by the random perturbation method together, a training set after data augmentation is obtained

[0088]

[0089] (4) Relationship prediction module: Relationship prediction and model training and inference based on an adaptive threshold.

[0090] In this step, according to the set of all triple vector representations after data augmentation obtained in the above steps, relationship type prediction is performed. First, the present invention uses two linear layers and activation functions to fuse the head entity vector and the context representation vector and the tail entity vector and the context representation vector respectively. Finally, the present invention uses a grouped bilinear model to calculate scores for all relationship types according to the two fused vectors obtained above.

[0091] In a specific implementation of the present invention, based on the set of triple vector groups of all entity pairs after data augmentation a relationship prediction module is trained to predict the relationship between each entity pair.

[0092] For any triple vector representation the present invention first applies two linear transformations with Tanh activation functions to fuse the context representation vector c h,t , the head entity vector e hThe tail entity representation vector e t , and the result after linear transformation is:

[0093] h = tanh(W h ·e h + W c1 ·c h,t )

[0094] t = tanh(W t ·e t + W c2 ·c h,t )

[0095] where W h , W t , W c1 , W c2 ∈R d×d are the trainable parameters of the model, and h, t represent the vectors after linear transformation.

[0096] After that, the vectors h, t after linear transformation are sliced into k groups, and a grouped bilinear layer is used to calculate the score for each relationship type :

[0097]

[0098] where is the bilinear layer parameter of the i-th group, h i and t i are the results of the i-th group after slicing, and score r is the predicted score that the triple vector representation belongs to the relationship r.

[0099] In the training stage, the present invention applies an adaptive threshold loss function, and dynamically learns the threshold θ of each entity pair by introducing a threshold class relationship TH h,t .

[0100]

[0101] where is the set of all valid relationships between the entity pair (e_h, e_t). When there is no relationship between the entity pair, is empty. is the set of all invalid relationships between the entity pair (e_h, e_t); score TH is the predicted score that the triple vector representation belongs to the threshold class relationship TH.

[0102] During the training phase, the model uses the backpropagation algorithm to calculate the gradients of the adaptive threshold loss function and update the model. Eventually, when the loss function converges, it can be used for relation prediction inference.

[0103] During the inference phase, the threshold θ of the valid relation score is score TH , if the predicted score for a certain relation is higher than score TH , then it is regarded as a valid relation, and the corresponding triple vector group (e_h, r, e_t) is output, and finally the set of predicted triples is obtained In addition, during the inference phase, no data augmentation operation is required.

[0104] In the second aspect of the present invention, a contrastive learning pre-training method based on the above-mentioned relation encoding data augmentation method is proposed, which is used to utilize the set of triple vector groups of all entity pairs after data augmentation to perform contrastive learning training on the relation prediction module. This method further improves the prediction performance of the document-level relation extraction model under the long-tail data distribution, especially the accuracy on the tail relations.

[0105] The present invention uses the proposed contrastive learning framework to pre-train the document-level relation extraction model obtained by the remote supervision method.

[0106] Under the setting of the text-level relation extraction task, the present invention believes that semantically similar samples should be entity pairs with the same relation r, including the original triple vector group and the vector group after data augmentation. However, only a few entity pairs have the same relation in the same document, especially for the tail relation types. Increasing the batch size of the model, that is, increasing the number of documents processed simultaneously, can partially alleviate this problem, but it requires a very large GPU memory for training, which cannot be obtained in scenarios with limited hardware resources. Based on this, the present invention changes the MOCO framework in classical contrastive learning to a setting more suitable for the document-level relation extraction task, named MoCo-DocRE.

[0107] The MoCo-DocRE framework proposed in the present invention can perform contrastive learning without using a large batch. It maintains a relation representation queue Q for each relation r r , which contains q relation representations in the previous batch for each relation , so that the relation representations encoded in the previous batch can be reused during the training process.

[0108] The following combines the attached Figure 2 to illustrate the contrastive learning framework proposed by the present invention.

[0109] The purpose of this step is to unify the triple vector representations after data augmentation and the original triple vector representations in the vector space, so that the triple vector representations of entity pairs with different relationship types are dispersed in the space, enabling the model to obtain good representation ability and distinguish different relationship types, especially the augmented tail relationship types. The present invention combines the commonly used MOCO framework in contrastive learning and designs a new contrastive learning method based on the above data augmentation method to make it more suitable for the relationship extraction scenario and improve the performance of the model on long-tail data.

[0110] In a specific implementation of the present invention, for the documents in the pre-training dataset First, perform the aforementioned document encoding, relationship encoding, and data augmentation to obtain the set T of triple vector representations of all entity pairs. For one of the triple vector representations Use two linear transformations to fuse the triple vector representation to obtain the linearly transformed vectors h and t.

[0111] Next, use a multi-layer perceptron with a ReLU activation function to fuse the relationship representation finally used in the contrastive learning stage:

[0112] x = relu(W2(W1[h:t] + b1) + b2)

[0113] where, [:] represents the vector concatenation operation, and are the trainable model parameters in the pre-training stage, d r is the dimension of the relationship representation x used for contrastive learning; x is the relationship representation finally used in the contrastive learning stage.

[0114] To maintain the consistency of the relationship representations in Q r use a momentum update model to encode the positive and negative samples in contrastive learning. The original model is updated through backpropagation, and the momentum update model is updated by the following formula:

[0115]

[0116] where, m is the momentum hyperparameter that can control the evolution speed of.

[0117] Next, input the document into in the same process as obtaining the relationship representation x to get x′, and push all the relationship representation sets into the relationship representation queue Q according to their relationship labels rAmong them, the rule is: if the relationship r holds between the entity pair (e_h, e_t), then x′ will be pushed into Q r . Finally, the positive and negative relationship representation sets of x can be obtained from the queue, that is:

[0118]

[0119]

[0120] Among them, is the positive relationship representation set of x, is the negative relationship representation set of x.

[0121] For the relationship representation x corresponding to the entity pair (e_h, e_t) used in the contrast learning stage, the contrast learning loss is expressed as:

[0122]

[0123] Among them, τ is the temperature hyperparameter, the superscript T represents transpose, x + is the relationship representation in the positive relationship representation set, x - is the relationship representation in the negative relationship representation set. Before calculating the loss, x, x + , x - have undergone L2 regularization.

[0124] In the contrast learning pre-training stage, the above loss function is used to train the model until the loss function converges. After the pre-training is completed, the model is further trained based on an adaptive threshold, and finally a trained document-level relationship extraction model is obtained for actual relationship prediction.

[0125] To verify the effectiveness of the method proposed in the present invention, in this embodiment, a series of experiments are carried out on the DocRED dataset. DocRED is the most commonly used document-level relationship extraction dataset, and its data distribution is as Figure 3 shown. The 7 most frequently occurring entity relationship types account for 55.12% (there are a total of 96 entity relationship types in this dataset); there are 60 entity relationship types with an occurrence frequency of less than 200 times.

[0126] The experimental results are shown in the following table.

[0127] Method Macro Macro@500 Macro@200 Macro@100 ATLOP 39.54 35.20 26.82 18.76 ERA 40.55 36.21 28.51 20.50 ERACL 41.34 37.13 29.43 22.31

[0128] All experimental results in the table are averaged after three repetitions. Among them, ATLOP is the model with the best performance in the DocRED dataset, ERA represents the model trained using the data augmentation method based on relation encoding, and ERACL represents the model pre-trained using contrastive learning. The Macro indicator represents the average F1 value of the model on all relation types, and the Macro@500 / 200 / 100 indicator represents the average F1 value of the model on relation types with an occurrence frequency of less than 500 / 200 / 100. The experimental results show that the data augmentation method based on relation encoding and the contrastive learning pre-training framework proposed in the present invention can significantly improve the prediction accuracy of the document-level relation extraction model on the tail relation types. The lower the frequency of the relation type, the more significant the improvement.

[0129] The above examples are only specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples, and many variations are possible. All variations that can be directly derived or associated with the contents disclosed by a person skilled in the art should be considered as the protection scope of the present invention.

Claims

1. A document-level relation extraction method based on long-tail data distribution, characterized in that The following steps are involved: Step 1: Document preprocessing Annotate all entities in a given document and annotate entity boundaries with special characters as a mention of the entity in the document; Step 2: Document encoding Use the preprocessed document as the input of the pre-trained Transformer model, obtain the contextual semantic representation of all characters in the document as vector encoding, and obtain the self-attention matrix between entities; Step 3: Relational encoding Traverse each pair of entities to form entity pairs; according to the document encoding results, calculate the vector representation of each entity in the document and the pooled context representation of the entity pair to form a three-element vector group; in the model training stage, it is necessary to annotate the relationship label of each entity pair and execute step 4; in the actual prediction stage, directly execute step 5; Step 4: Data Augmentation For the set of labeled triple vector groups, randomly select or preset the relationship type that needs to be augmented, design a mask vector, perturb the pooled context representation in the original triple vector group to be augmented, and generate a new triple vector group; use the original triple vector group set and the triple vector group set obtained by data augmentation as training sets to train a document-level relationship extraction model; Step 5: Relationship Prediction Use the methods in steps 1-3 to preprocess, encode documents, and encode relations for the given document. Use the trained document-level relation extraction model to predict relations for the obtained triple vector group, and output entity pairs with valid relations and their relations.

2. The method for document-level relation extraction based on long-tail data distribution according to claim 1, wherein The pre-trained Transformer model adopts the BERT model.

3. The method for document-level relation extraction based on long-tail data distribution according to claim 2, wherein The step 2 is specifically as follows: Input the annotated entities and mentioned documents into the BERT model to obtain the contextual semantic representation H of all characters in the document and the self-attention matrix A; expressed as: Among them, represents a document with a word sequence length of l, and w l represents the l-th character in the document; Ptr(.) represents the pre-trained BERT model, H is the word vector output by the last layer of the BERT model, which is the context semantic representation of all characters in the document; A is the self-attention matrix in the last layer of the BERT model.

4. The method for document-level relation extraction based on long-tail data distribution according to claim 1, wherein The step 3 is specifically as follows: 3.1) Traverse each pair of entities to form entity pairs; 3.2) Based on the document encoding results, calculate the vector representation of each entity in the document: Among them, e i represents the vector representation of the i-th entity, m ij represents the vector representation of the i-th entity's j-th mention in the document, that is, the word vector corresponding to the special character at the left boundary of the mention, obtained by indexing the context semantic representations of all characters in the document in step 2; m represents the number of times the i-th entity is mentioned in the document; The first entity in the entity pair is called the head entity \(e_h\), and the head entity vector is denoted as \(e\). h The second entity in the entity pair is called the tail entity \(e_t\), and the tail entity vector is denoted as \(e\). t ; 3.3) Compute the pooled context representation of entity pairs: For the entity pair (e_h, e_t), the pooled context representation c of this entity pair is obtained through the following two equations h,t : A h,t = A h * A t Among them, A h,t is the product of the attention scores of the head entity \(e_h\) and the tail entity \(e_t\) for all words in the document; A h ∈R l×1 is the attention score of the head entity \(e_h\) for all words in the document, and A t ∈R l×1 is the attention score of the tail entity \(e_t\) for all words in the document, where \(H\) is the context semantic representation of all characters in the document, and \(1\in R\) l×l ; 3.4) For the entity pair (e_h, e_t), its triple vector representation is T h,t =(e h , c h,t , e t ), obtaining the set of all triple vector representations ε represents the entity set.

5. The method for document-level relation extraction based on long-tail data distribution according to claim 1, wherein, The step 4 is specifically as follows: 4.1) Set the set of relationship types for which data augmentation is required R is the set of all relationship types; 4.2) Given an entity pair (e_h, e_t), if its relationship from index the original triple vector group representation (e h , c h,t , e t ); First, a mask vector p is randomly generated, and each dimension of the mask vector is generated by a Bernoulli distribution with parameter p; then the mask vector p is dot-multiplied with A h,t to perform a masking operation on it, and the formula is: A′ h,t = p * A h,t Among them, A h,t is the product of the attention scores of the head entity e_h and the tail entity e_t for all words in the document, and A ′ h,t is the masked attention score; The context representation vector c′ after perturbation h,t is calculated as follows: Among them, the superscript T represents transposition; Generate a new ternary vector group representation (e h , c′ h,t , e t ); Take the union of the set of all ternary vector group representations generated after perturbation and the set of original ternary vector group representations as the training set Among them, ε represents the entity set.

6. The method for document-level relation extraction based on long-tail data distribution according to claim 5, wherein In step 4.2), set the number of times α of data augmentation, and at the same time generate α random masks, which can generate α perturbed context representation vectors That is, generate α new triple vector group representations.

7. The method for document-level relation extraction based on long-tail data distribution according to claim 5, wherein Using the training set When training a document-level relation extraction model, an adaptive threshold-based model training method is adopted, specifically as follows: For any triple vector group representation First, apply two linear transformations with Tanh activation functions to fuse the context representation vector c h,t , the head entity vector e h and the tail entity representation vector e t . The result after the linear transformation is: h = tanh(W h ·e h +W c1 ·c h,t ) t = tanh(W t ·e t +W c2 ·c h,t ) Among them, W h , W t , W c1 , W c2 ∈R d×d are the trainable parameters of the model, and h and t represent the vectors after linear transformation; After that, the linearly transformed vectors h and t are sliced into k groups, and grouped bilinear layers are used to calculate the scores for each relation type : Among them, W r i are the bilinear layer parameters of the i-th group, h i and t i are the results of the i-th group after segmentation, score r is the prediction score of the triple vector group representation belonging to the relationship r; In the training phase, the threshold θ of each entity pair is dynamically learned by introducing the threshold class relationship TH h,t , and the loss function is as follows: Among them, is the set of all valid relationships between entity pairs (e_h, e_t). When there is no relationship between entity pairs, is empty; is the set of all invalid relationships between entity pairs (e_h, e_t); score TH is the triple vector representation belongs to the predicted score of the threshold class relationship TH.

8. The method for document-level relation extraction based on long-tail data distribution according to claim 7, wherein Before training the model based on the adaptive threshold, the document-level relationship extraction model is also pre-trained by comparative learning, specifically: For the documents in the pre-training dataset, obtain the set of triple vector representations of all entity pairs according to the method in Steps 1-4 As the training set, for one of the triple vector representations Use two linear transformations to fuse the triple vector representation to obtain the linearly transformed vectors h and t; Next, a multi-layer perceptron with ReLU activation function is used to fuse the relational representations that are finally used in the contrastive learning phase: x=relu(W2(W1[h:t]+b1)+b2) where [:] represents the vector concatenation operation, W1 and W2 are model parameters that can be trained in the pre-training stage, and d r is the dimension of the relational representation x for contrastive learning; x is the relational representation finally used in the contrastive learning stage; Maintain a relational representation queue Q for each relation r r , to maintain the consistency of the relational representations in Q r , use a momentum update model to encode positive and negative samples in contrastive learning, and the original model is updated through backpropagation, and the momentum update model is updated by the following formula: where m is a momentum hyperparameter used to control the evolution speed; Next, input the document into in the same process as obtaining the relationship representation x to obtain x ′ , and push all relationship representation sets into the relationship representation queue Q according to their relationship tags, where the rule is: if the relationship r holds between the entity pair (e_h, e_t), then x r will be pushed into Q ′ ; finally, obtain the positive and negative relationship representation sets of x from the queue; ′ r ​ For the relation representation x corresponding to the entity pair (e_h, e_t) used in the contrastive learning phase, the contrastive learning loss is expressed as: where τ is the temperature hyperparameter, the superscript T represents transpose, and x + is the relation representation in the positive relation representation set, and x - is the relation representation in the negative relation representation set, is the positive relation representation set of x, is the negative relation representation set of x.

9. The method for document-level relation extraction based on long-tail data distribution according to claim 7, wherein During the inference stage, the threshold of the valid relationship score is score TH , if the predicted score for a certain relationship is higher than score TH , then it is regarded as a valid relationship, and the corresponding triple vector (e_h, r, e_t) is output, and finally the set of predicted triples is obtained

Citation Information

Patent Citations

  • Cross-class ontology integration for language modeling

    US20210294970A1

  • System for entity and evidence-guided relation prediction and method of using the same

    US20220067278A1