Entity and Evidence-Guided Relationship Prediction System and Method of Using the Same

By generating entity-guided input sequences and using the internal attention probability of pre-trained language models for joint training, the problem of cross-sentence relationship extraction in the existing technology is solved, and more efficient relationship and evidence prediction effects are achieved.

CN114266246BActive Publication Date: 2025-07-18BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110985298.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-25
Filing Date
2021-08-25
Publication Date
2025-07-18
Estimated Expiration
2041-08-25

AI Technical Summary

Technical Problem

Existing neural models and related databases only consider intra-sentence relationships, making it difficult to effectively solve the problem of cross-sentence relationship extraction between entities in text.

Method used

By generating entity-guided input sequences, combining the internal attention probability of the pre-trained language model, joint training of entity and evidence statements is performed, multiple bilinear layers are used for relationship and evidence prediction, and the parameters of the language model are updated to improve the accuracy of relationship extraction.

Benefits of technology

It significantly improves the accuracy and evidence prediction ability of cross-sentence relationship extraction, achieving advanced performance on DocRED datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114266246B_ABST
    Figure CN114266246B_ABST
Patent Text Reader

Abstract

Multi-task prediction system and method. The system includes a computing device. The computing device has a processor and a storage device storing computer-executable code. The computer-executable code is configured to: provide a head entity and a document containing the head entity; process the head entity and the document through a language model to obtain a head extraction corresponding to the head entity, a tail extraction corresponding to a tail entity in the document, and a statement extraction corresponding to a statement in the document; use a first bilinear layer to predict a head-tail relationship between the head extraction and the tail extraction; use a second bilinear layer to combine the statement extraction and a relationship vector corresponding to the predicted head-tail relationship to obtain a statement-relationship combination; and use a third bilinear layer to predict an evidence statement supporting the head-tail relationship based on the statement-relationship combination and the attention extracted from the language model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference

[0002] In the description of the present disclosure, some references are cited and discussed, which may include patents, patent applications, and various publications. The citation and / or discussion of such references are provided only to clarify the description of the present disclosure and do not admit that any such reference is "prior art" of the disclosure described herein. All references cited and discussed in this specification are incorporated herein by reference in their entirety to the same extent as if each reference was individually incorporated by reference. Technical Field

[0003] The present disclosure generally relates to relation extraction, and more particularly, to entity and evidence guided relation extraction (E2GRE). Background Art

[0004] The background description provided herein is to present the context of the present disclosure generally. The work of the current inventors, to the extent described in this background section, and aspects that may not conform to the description of the prior art at the time of filing the application, are not expressly or implicitly admitted as prior art of the present disclosure.

[0005] Relation extraction (RE) extracts the relations between entity pairs in plain text and is an important task in natural language processing (NLP). It has downstream applications for many other NLP tasks, such as knowledge graph construction, information retrieval, question answering, and dialogue systems. RE can be implemented by manually constructed patterns, bootstrapping methods, supervised methods, distant supervision, and unsupervised methods. These methods generally involve using neural models to learn relations. Neural models for RE have made progress. However, these neural models and related databases only consider intra-sentence relations.

[0006] Therefore, there is an unresolved need in the art to address the above-mentioned deficiencies and drawbacks. Summary of the Invention

[0007] In some aspects, the present disclosure relates to a system. In some embodiments, the system includes a computing device, the computing device includes a processor and a storage device storing computer-executable code. The computer-executable code is configured, when executed by the processor, to:

[0008] Provide a head entity and a document containing the head entity;

[0009] Process the head entity and the document through a language model to obtain a head extraction corresponding to the head entity, a tail extraction corresponding to the tail entity in the document, and a sentence extraction corresponding to the sentences in the document;

[0010] Use a first bilinear layer to predict the head entity-tail entity relationship between the head extraction and the tail extraction;

[0011] Use a second bilinear layer to combine the sentence extraction and the relationship vectors corresponding to the head extraction and the tail extraction to obtain a sentence-relationship combination;

[0012] Based on the sentence-relationship combination and the attention extracted from the language model, use a third bilinear layer to predict evidence sentences from the document, where the evidence sentences support the head-tail relationship; and

[0013] Based on the predicted head entity-tail entity relationship, the predicted evidence sentences, and the labels of the document, update the parameters of the language model, the first bilinear layer, the second bilinear layer, and the third bilinear layer, where the labels of the document include the true head entity-tail entity relationship and the true evidence sentences.

[0014] In some embodiments, multiple annotated documents are used to train the language model, the first bilinear layer, the second bilinear layer, and the third bilinear layer. At least one of the annotated documents has E entities, and the at least one annotated document is extended to E samples. Each of the E samples includes the at least one annotated document and a head entity corresponding to one of the E entities, where E is a positive integer.

[0015] In some embodiments, the computer-executable code is configured to update the parameters based on a loss function, where the loss function is defined by L RE is the relationship prediction loss, is the sentence prediction loss, and λ1 is a weight factor that has a value equal to or greater than 0.

[0016] In some embodiments, the language model includes at least one of the following: Generative Pretrained Model GPT, GPT-2, Bidirectional Encoder Representations from Transformers BERT, Robustly Optimized BERT Approach roBERTa, and Reparameterized Transformer-XL Network XLnet.

[0017] In some embodiments, the computer-executable code is configured to extract the attention from the last 2 to 5 layers of the language model. In some embodiments, the computer-executable code is configured to extract the attention from the last 3 layers of the language model.

[0018] In some embodiments, the first bilinear layer is defined by where is the predicted value of the i-th relation among multiple relations between the head entity h and the k-th tail entity t, δ represents the sigmoid function, W k is the learned weight of the first bilinear layer, and b i is the bias of the first bilinear layer. In some embodiments, the second bilinear layer is defined by i where is the predicted probability that the j-th statement s in the document is a supporting statement for the i-th relation r j with respect to the i-th relation r, i and and are the learnable parameters of the second bilinear layer with respect to the i-th relation. In some embodiments, the third bilinear layer is defined by where is the predicted probability that the j-th statement in the document is a supporting statement for the i-th relation with respect to the k-th tail entity, δ represents the sigmoid function, is the learned weight of the third bilinear layer, and

[0019] is the bias of the third bilinear layer.

[0020] In some embodiments, after training, the language model, the first bilinear layer, the second bilinear layer, and the third bilinear layer are configured to provide relation prediction and evidence prediction for a query entry, where the query entry has a query head entity and a query document including the query head entity.

[0021] In some aspects, the present disclosure relates to a method. In some embodiments, the method includes:

[0022] providing, by a computing device, a head entity and a document including the head entity;

[0023] processing, by a language model stored in the computing device, the head entity and the document to obtain a head extraction corresponding to the head entity, a tail extraction corresponding to a tail entity in the document, and a statement extraction corresponding to a statement in the document;

[0024] predicting, by a first bilinear layer stored in the computing device, a head-tail relationship between the head extraction and the tail extraction;

[0025] The second bilinear layer stored in the computing device combines the statement extraction and the relationship vectors corresponding to the head extraction and the tail extraction to obtain a statement-relationship combination;

[0026] Based on the statement-relationship combination and the attention extracted from the language model, a third bilinear layer stored in the computing device predicts evidence statements from the document, where the evidence statements support the head-tail relationship; and

[0027] Based on the predicted head entity-tail entity relationship, the predicted evidence statements, and the labels of the document, update the parameters of the language model, the first bilinear layer, the second bilinear layer, and the third bilinear layer, where the labels of the document include the true head entity-tail entity relationship and the true evidence statements.

[0028] In some embodiments, the step of updating the parameters is performed based on a loss function, which is defined by L RE is the relationship prediction loss, is the statement prediction loss, λ1 is a weight factor, and the weight factor has a value equal to or greater than 0.

[0029] In some embodiments, the language model includes at least one of the following: the Generative Pretrained Model GPT, GPT-2, the Bidirectional Encoder Representations from Transformers BERT, the Robustly Optimized BERT Approach roBERTa, and the Reparameterized Transformer-XL Network XLnet.

[0030] In some embodiments, the attention is extracted from the last 3 layers of the language model.

[0031] In some embodiments, where the first bilinear layer is defined by is the predicted value of the i-th relationship among multiple relationships between the head entity h and the k-th tail entity t, δ represents the sigmoid function, W k is the learned weight of the first bilinear layer, and b i is the bias of the first bilinear layer; where the second bilinear layer is defined by i is the predicted probability of the j-th statement s in the document j is the i-th relationship r i of the supporting statement, and are the learnable parameters of the second bilinear layer with respect to the i-th relationship; and where the third bilinear layer is defined by ​​​​ is the predicted probability that the j-th statement in the document is a supporting statement for the i-th relation regarding the k-th tail entity, δ represents the sigmoid function, are the learned weights of the third bilinear layer, is the bias of the third bilinear layer.

[0032] In some embodiments, the method further includes, after being trained: providing relation prediction and evidence prediction for a query entry, where the query entry has a query head entity and a query document including the query head entity.

[0033] In some aspects, the present disclosure relates to a non-transitory computer-readable medium storing computer-executable code, which, when executed by a processor of a computing device, is configured to execute the above method.

[0034] These and other aspects of the present disclosure will become apparent from the following description of the preferred embodiments in conjunction with the accompanying drawings and their captions, although changes and modifications may be made therein without departing from the spirit and scope of the novel concepts of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The drawings illustrate one or more embodiments of the present disclosure and are used in conjunction with the specification to explain the principles of the present disclosure. Wherever possible, the same reference numerals are used to refer to the same or similar elements of the embodiments.

[0036] Figure 1 Schematically shows an evidence-guided multi-task learning framework according to some embodiments of the present disclosure.

[0037] Figure 2 Schematically shows an evidence-guided relation extraction system according to some embodiments of the present disclosure.

[0038] Figure 3 Shows an example of an entry from the DocRED database.

[0039] Figure 4 Schematically shows a process of training evidence-guided relation extraction according to some embodiments of the present disclosure.

[0040] Figure 5 Schematically shows a process of relation prediction according to some embodiments of the present disclosure.

[0041] Figure 6 Shows the leading public leaderboard numbers on DocRED, where our E2GRE method uses RoBERTa-large.

[0042] Figure 7Shows the relationship extraction results on the supervision settings for DocRED, where BERT-base is used as the pre-trained language model, and comparisons are made with E2GRE and other published models on the validation set.

[0043] Figure 8 Shows the ablation study of entity-guided RE and evidence-guided RE, where BERT + joint training is the BERT baseline combined with joint training for RE and evidence prediction, and the results are evaluated on the validation set.

[0044] Figure 9 Shows the ablation study of different numbers of attention probability layers for evidence prediction from BERT. The results are evaluated on the development set.

[0045] Figure 10 Shows the baseline BERT attention heatmap on the tokenized document of the DocRED example.

[0046] Figure 11 Shows the attention heatmap of E2GRE on the tokenized document of the DocRED example. Detailed implementation

[0047] The present disclosure is described in more detail in the following examples, which are only intended to be illustrative, and many modifications and variations will be apparent to those skilled in the art. Various embodiments of the present disclosure are now described in detail. Referring to the accompanying drawings, the same numbers indicate the same components. Unless the context clearly dictates otherwise, the meanings of "a", "an", and "the" used in the description herein and in the following claims include plural references. In addition, as used in the description herein and in the following claims, the meaning of "in..." includes "in..." and "on...", unless the context clearly dictates otherwise. Also, the specification may use headings or subheadings for the convenience of the reader without affecting the scope of the invention. In addition, some terms used in this specification are defined more specifically below.

[0048] The terms used in this specification generally have their ordinary meanings in the art, in the context of the present disclosure, and in the particular context in which each term is used. Some of the terms used to describe the present disclosure are discussed below or elsewhere in the specification to provide additional guidance to the practitioner regarding the description of the present disclosure. It is understood that the same thing can be said in more than one way. Accordingly, alternative languages and synonyms may be used for any one or more of the terms discussed herein, and no special significance is to be attached to whether or not a term is elaborated or discussed herein. Synonyms for certain terms are provided. The use of one or more synonyms does not exclude the use of other synonyms. The use of examples anywhere in this specification, including examples of any of the terms discussed herein, is illustrative only and in no way limits the scope and meaning of the present disclosure or of any exemplary term. Similarly, the present disclosure is not limited to the various embodiments given in this specification.

[0049] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It should also be understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the relevant art and the context of this disclosure, and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0050] As used herein, the term "module" may refer to a portion of an application specific integrated circuit (ASIC) or include an application specific integrated circuit (ASIC), an electronic circuit, a combinatorial logic circuit, a field programmable gate array (FPGA), a processor (shared, dedicated, or a group of processors) that executes code, other suitable hardware components that provide the described functionality, or a combination of some or all of the foregoing, such as in a system on a chip. The term "module" may include a memory (shared, dedicated, or a group of processors) that stores code executed by the processor.

[0051] The term "code" as used herein may include software, firmware, and / or microcode, and may refer to a program, a routine, a function, a class, and / or an object. The term "shared" as used above means that portions of or all of the code from multiple modules may be executed using a single (shared) processor. Additionally, portions of or all of the code from multiple modules may be stored in a single (shared) memory. The term "group" as used above means that portions of or all of the code from a single module may be executed using a group of processors. Additionally, a group of memories may be used to store some or all of the code from a single module.

[0052] As used herein, the term "interface" generally refers to a communication tool or device at the interaction point between components for performing data communication between components. Generally, interfaces can be applied at the hardware and software levels and can be unidirectional or bidirectional interfaces. Examples of physical hardware interfaces can include electrical connectors, buses, ports, cables, terminals, and other I / O devices or components. Components that communicate with an interface can be, for example, multiple components of a computer system or peripheral devices.

[0053] This disclosure relates to computer systems. As shown, computer components can include physical hardware components, which are shown as solid blocks, and virtual software components, which are shown as dashed blocks. Those of ordinary skill in the art will understand that, unless otherwise stated, these computer components can be implemented in forms including but not limited to software, firmware, or hardware components or combinations thereof.

[0054] The apparatuses, systems, and methods described herein can be implemented by one or more computer programs executed by one or more processors. The computer programs include processor-executable instructions stored on a non-transitory tangible computer-readable medium. The computer programs can also include stored data. Non-limiting examples of non-transitory tangible computer-readable media include non-volatile memory, magnetic storage, and optical storage.

[0055] The present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, in which embodiments of the present disclosure are shown. However, the present disclosure can be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.

[0056] In some aspects, the present disclosure relates to a joint training framework E2GRE. In some embodiments, first, the present disclosure introduces an entity-guided sequence as an input to a pre-trained language model (LM), such as Bidirectional Encoder Representaition from Transformers (BERT) and robustly optimized BERT approach (roBERTa). The entity-guided sequence helps the LM focus on the document regions related to the entity. Second, the present disclosure guides the fine-tuning of the pre-trained LM by using the internal attention probabilities of the pre-trained LM as additional features for evidence prediction. Thus, the disclosed method encourages the pre-trained LM to focus on entities and supporting / evidence statements. In some embodiments, the disclosed E2GRE is evaluated on DocRED, where DocRED is a recently released large-scale dataset for relation extraction. E2GRE is able to achieve advanced results on the public leaderboard in all metrics, indicating that E2GRE is both effective and collaborative in relation extraction and evidence prediction.

[0057] Specifically, in E2GRE, for each entity in the document, the present disclosure generates a new input sequence by appending the entity to the beginning of the document and then feeding it to the pre-trained LM. Thus, for each document with N e entities, the present disclosure generates N e entity-guided input sequences for training. By introducing these new training inputs, the present disclosure encourages the pre-trained LM to focus on the entities appended to the beginning of the document. The present disclosure further exploits the pre-trained LM by directly using the internal attention probabilities as additional features for evidence prediction. The joint training of relation extraction and evidence prediction helps the model locate the correct semantics required for relation extraction. Both strategies leverage the pre-trained LM to fully utilize the pre-trained LM in our task. The main contributions of the E2GRE method include: (1) For each document, the present disclosure generates multiple new inputs to feed into the pre-trained language model: the present disclosure concatenates each entity with the document and provides it as an input sequence to the LM. This allows for fine-tuning of the internal representation from the pre-trained LM to be guided by the entity. (2) The present disclosure also uses the internal BERT attention probabilities as additional features for evidence prediction. This allows for fine-tuning of the internal representation from the pre-trained LM to also be guided by the evidence / supporting statements.

[0058] Each of these strategies significantly improves the performance, and by combining them, the present disclosure is able to achieve advanced results on the DocRED leaderboard.

[0059] Figure 1 Schematically shows an E2GRE framework according to some embodiments of the present disclosure. As Figure 1 shown, the framework 100 includes a language model 108, a first bilinear layer 118, a second bilinear layer 126, a third bilinear layer 130, and an update module 134. The language model 108, the first bilinear layer 118, the second bilinear layer 126, the third bilinear layer 130, and the update module 134 are also collectively referred to as the model or the E2GRE model. A sample 102 can be provided to the language model 108, and the sample 102 can be a training sample or a query sample. In some embodiments, the present disclosure organizes the sample 102 by attaching the head entity 104 to the beginning of the document or context 106. The head entity 104 is also included in the document 106. In some embodiments, the present disclosure can also organize the sample into the head entity 104, the tail entity, and the document 106. However, it is advantageous not to define the tail entity when preparing the sample, which makes the preparation of the sample and the operation of the framework more efficient. In some embodiments, the sample is prepared into the following sequence: "[CLS]" + H + "[SEP]" + D + "[SEP]", where [CLS] is a class token placed at the beginning of the entity-guided input sample, [SEP] is a separator, H is the token of the first mention of a single entity, and D is the document token. There are a total of N e entities in the document D, including the head entity H and N e -1 tail entities, where N e is a positive integer.

[0060] In some embodiments, the language model 108 is BERT. BERT can only process sequences with a maximum length of 512 tokens. Due to this limitation, if the length of the training input is greater than 512, the present disclosure uses a sliding window method on the document. If the input sequence length is greater than 512, the embodiment decomposes the input sequence into two sequences or two documents. In some embodiments, these embodiments can divide a longer passage into more windows. The first sequence is the original input sequence, with a maximum of 512 tokens. The second sequence has the same format as the first sequence, with an offset added in the document so that the second sequence can reach the end. This looks like "[CLS]" + H + "[SEP]" + D[offset:end] + "[SEP]". The embodiment combines these two input sequences in our model by averaging the token embeddings and the BERT attention probabilities (where the embeddings and the BERT attention probabilities are calculated twice in the model).

[0061] In some embodiments, when the sample 102 is a training sample, the training sample 102 further includes labels for the tail entity, the relation, and the evidence statements supporting the relation. For example, these labels can be used by the second bilinear layer 126 to retrieve the relation vectors corresponding to the head and tail entities, or by the update module 134 to calculate the loss function corresponding to the training sample, the predicted relation, and the predicted evidence statements.

[0062] The sample 102 is input into the language model 108. In some embodiments, the language model 108 is a pre-trained language model (LM). LMs are extremely powerful tools that have emerged in recent years. The most recent pre-trained LMs are transformer-based and trained using large amounts of data. In some embodiments, the language model 108 can be any type of language model, such as the Generative Pre-training Model (GPT), GPT-2, Bidirectional Encoder Representations from Transformers (BERT), Robustly Optimized BERT approach (RoBERTa), and re-parameterized Transformer-XL network (XLNet).

[0063] The language model 108 processes the input 102 and produces an output sequence 112. The head entity 115 can be extracted from the output sequence 112 through head extraction 114, and the tail entity 117 can be extracted from the output sequence 112 through tail extraction 116. The head extraction 114 averages the embeddings on the concatenated head entity tokens to obtain the head entity embedding h. The tail extraction 116 extracts a set of tail entity embeddings from the output sequence. For the k-th tail entity embedding t k , the tail extraction 116 locates the indices of the tokens of the k-th tail entity and averages the output embeddings of BERT at these indices to obtain t k .

[0064] After obtaining the head entity embedding and all the tail entity embeddings in the entity-guided sequence, where 1 ≤ k ≤ N e -1 and are real numbers with d dimensions, the embodiment uses a first bilinear layer 118 with a sigmoid activation function to predict the probability of the i-th relation between the head entity h and the k-th tail entity t k , denoted by , as follows:

[0065]

[0066] where δ is the sigmoid function, the T in h T is the transpose, W i and b iis the learnable parameter corresponding to the i-th relation, where 1 ≤ i ≤ N r , N r is a positive integer representing the total number of relations.

[0067] The language model 108 such as BERT can also be fine-tuned by the following multi-label cross-entropy loss:

[0068]

[0069] where y ik is the true value or label of the i-th relation in the training sample, is the predicted value of the i-th relation regarding the head entity h and the k-th tail entity t k .

[0070] During the inference process, the goal of relation extraction is to predict the relations for each pair of head / tail entities in the document. For a given entity-guided input sequence "[CLS]" + entity + "[SEP]" + document + "[SEP]", the output of the model is a set of N e - 1 relation predictions. In this embodiment, the predictions of each sequence generated from different head entities in the same document are combined to obtain all relation predictions for the document.

[0071] The output of the first bilinear 118 is the relation prediction 120. In some embodiments, when the predicted value is equal to or greater than 0.5, the relation is defined as the predicted relation. In some embodiments, when the predicted values of all relations are less than 0.5, it indicates that there is no relation between the corresponding head entity and tail entity. In some embodiments, the relations predicted by relation extraction are used to query the correct relation vectors during inference / testing.

[0072] Reference Figure 1 , the statement extraction 122 and the obtained relation prediction 120 can be used together to predict the evidence / support statements. The evidence statements contain supporting facts, which are important for the model to predict the correct relation between the head entity and the tail entity. Therefore, the evidence prediction task is a good auxiliary task for relation extraction, providing interpretability for the multi-task model.

[0073] The goal of evidence prediction is to predict whether a given statement is evidence / support for a given relation. Given a statement s, the embodiment first obtains the statement embedding by averaging the embeddings of all words in the statement s These embeddings are the statement extraction 122 calculated from the output of the language model 108. At the same time, for the i-th relation (1 ≤ i ≤ N r ), the embodiment defines the vector r i ∈ R d as the relation embedding. These relation embeddings r are randomly initializedi an OR-relation vector 124 and learn these relation embeddings r from the model i an OR-relation vector 124. Here, R d is a real number with d dimensions.

[0074] Then, the embodiment employs a second bilinear layer 126 that uses the statement embedding 122 and the relation embedding 124. Specifically, the second bilinear layer 126 with a sigmoid activation function is used to predict the probability that the j-th statement s j is the supporting statement for a given i-th relation r i as follows:

[0075]

[0076] where s j and r i represent the embeddings of the j-th statement and the i-th relation respectively, is the j-th statement s in the document j is the predicted probability that it is the supporting statement for the i-th relation r i and and are the learnable parameters of the second bilinear layer with respect to the i-th relation. In some embodiments, the learnable parameters and are called weights, and the learnable parameters and are called biases.

[0077] Finally, assuming there are N s statements in a given context, the embodiment defines the evidence prediction loss under a given relation i as follows:

[0078]

[0079] where, when the statement j is an evidence statement for inferring the i-th relation, Note that in the training phase, the model uses the embedding of the true relation in Equation (3). In the testing phase, the model uses the embedding of the relation predicted by the relation extraction model in Equation (1).

[0080] The internal attention probabilities of a language model 108 such as BERT can also be used to fine-tune the model. In some embodiments, the BERT attention probabilities determine which locations in the document the BERT model will focus on. Thus, these attention probabilities can guide the language model 108 to focus on relevant regions in the document for relation extraction. In some embodiments, the present disclosure has found that regions with higher attention values typically come from supporting statements. Thus, in some embodiments, these attention probabilities contribute to evidence prediction. For each head h and tail t k , the present disclosure uses the attention probabilities extracted from the last l internal BERT layers for evidence prediction.

[0081] In some embodiments, let be the query, let be the key of the multi-head self-attention layer, N h is the number of attention heads described in Vaswani et al., 2017 (incorporated herein by reference in its entirety), L is the length of the entity-guided input sequence, and d is the embedding dimension. The present disclosure first extracts the output of the multi-head self-attention (MHSA, Multi-Headed Self-Attention) from a given layer in BERT, as follows (attention extraction 128 in Figure 1 ):

[0082]

[0083]

[0084] A = Concat(Att-head i ,..., Att-head n ) (7)

[0085] For each head h and tail t k , some embodiments of the present disclosure extract the attention probabilities corresponding to the head and tail tokens to assist with the relation. Specifically, the embodiments concatenate the MHSAs of the last l BERT layers extracted by equation (7) to form an attention probability tensor:

[0086] Then, the embodiments calculate the attention probability representation for each statement under a given head-tail entity pair as follows.

[0087] 1. The embodiments first apply a max pooling layer along the attention head dimension (i.e., the second dimension) on . The maximum value helps to show the locations that a particular attention head might be looking at. After that, the embodiments apply an average pooling on the last l layers. The embodiments obtain

[0088] 2. Then, according to the start and end positions in the document [line 483, page 5 - please define "what"], the embodiment extracts the attention probability tensor from the head entity marker and the tail entity marker. The embodiment averages the attention probabilities of all tokens of the head and tail embeddings to obtain

[0089] 3. Finally, the embodiment generates a statement representation by averaging the attention of each token in the given statement in the document, obtaining to obtain

[0090] to obtain the attention probability α sk After that, the embodiment combines α sk with the evidence prediction result of statement s from equation (3) to form a new statement representation, and feeds it into a bilinear layer with sigmoid for evidence statement prediction, as follows:

[0091]

[0092] where is the fused representation vector of the statement embedding and the relation embedding for a given head / tail entity pair.

[0093] Finally, the embodiment defines the evidence prediction loss under a given relation i based on the attention probability representation, as follows:

[0094]

[0095] where, is the j - th value of calculated by equation (8).

[0096] The embodiment combines the relation extraction loss and the attention - probability - guided evidence prediction loss as the final objective function for joint training:

[0097]

[0098] where L RE is the relation prediction loss, is the evidence prediction loss, λ1≥0 is a weight factor used to trade - off between the two losses, and the two losses are data - related. In other words, the present disclosure does not use the loss functions of equations (2) and (9), but uses the loss function of equation (10), which is the combination of equations (2) and (9), for the entire model.

[0099] Figure 2 Schematically shows an E2GRE system according to some embodiments of the present disclosure. E2GRE performs relation prediction and evidence prediction. AsFigure 2 As shown, system 200 includes computing device 210. In some embodiments, computing device 210 can be a server computer, a cluster, a cloud computer, a general-purpose computer, a headless computer, or a dedicated computer, and computing device 210 provides relationship prediction and evidence prediction. Computing device 210 can include, but is not limited to, processor 212, memory 214, and storage device 216. In some embodiments, computing device 210 can include other hardware components and software components (not shown) to perform its corresponding tasks. Examples of these hardware and software components can include, but are not limited to, other required memories, interfaces, buses, input / output (I / O) modules or devices, network interfaces, and peripherals.

[0100] Processor 212 can be a central processing unit (CPU), which is configured to control the operation of computing device 210. Processor 212 can execute the operating system (OS) of computing device 210 or other applications. In some embodiments, computing device 210 can have multiple CPUs as processors, such as two CPUs, four CPUs, eight CPUs, or any suitable number of CPUs.

[0101] Memory 214 can be volatile memory, such as random access memory (RAM), for storing data and information during the operation of computing device 210. In some embodiments, memory 214 can be a volatile memory array. In some embodiments, computing device 210 can run on more than one memory 214. In some embodiments, computing device 210 can also include a graphics card to assist processor 212 and memory 214 in image processing and display.

[0102] Storage device 216 is a non-volatile data storage medium for storing the operating system (not shown) and other applications of computing device 210. Examples of storage device 216 can include non-volatile memory, such as flash memory, memory cards, USB drives, hard disk drives, floppy disks, optical drives, solid state drives, or any other type of data storage device. In some embodiments, computing device 210 can have multiple storage devices 216, which can be the same storage device or different types of storage devices, and the applications of computing device 210 can be stored in one or more storage devices 216 of computing device 210.

[0103] In this embodiment, processor 212, memory 214, and storage device 216 are components of computing device 210, such as a server computing device. In other embodiments, computing device 210 can be a distributed computing device, and processor 212, memory 214, and storage device 216 are shared resources from multiple computing devices in a predefined area.

[0104] The storage device 216 particularly includes a multi-task prediction application 218 and training data 240. In some embodiments, the E2GRE application 218 is also named as the E2GRE model, including model weights that can be trained using the training data, and the model can be used to make predictions using the well-trained model weights. The training data 240 is optional for the computing device 210 as long as the E2GRE application 218 can access the training data stored in other devices.

[0105] As Figure 2 shown, the E2GRE application 218 includes a data preparation module 220, a language model 222, a relationship prediction module 224, a statement-relationship combination module 226, an attention extraction module 228, an evidence prediction module 230, and an update module 232. In some embodiments, the E2GRE application 218 may include other applications or modules necessary for the operation of the E2GRE application 218, such as an interface for a user to input queries to the E2GRE application 218. It should be noted that the modules 220-232 are all implemented by computer-executable code or instructions, data tables or databases, or a combination of hardware and software, which together constitute an application. In some embodiments, each module may also include sub-modules. Alternatively, some modules may be combined into a stack. In other embodiments, some modules may be implemented as circuits rather than executable code. In some embodiments, the modules may also be collectively referred to as a model, and the model can be trained using the training data, and after training, the model can be used to make predictions.

[0106] The data preparation module 220 is used to prepare training samples or query data and send the prepared training samples or query data to the language module 222. The prepared training samples input to the language module 222 include a head entity and a document or context containing the head entity. In some embodiments, the format of the input training samples is "[CLS]" + head entity + "[SEP]" + document + "[SEP]". The head entity can be a single word or several words, such as the head entity "New York City" or the head entity "Beijing". The prepared query data or query samples have the same format as the training samples. Although only the head entity and the document in each training sample are used as the input to the language model 222, the label of the tail entity, the tail entity, the relationship between the head entity and the tail entity, and the evidence statements can all be used to calculate the loss function and backpropagate during the training process to optimize the parameters of the model. Here, the model corresponds to the E2GRE application 218.

[0107] In some embodiments, an example of the training dataset may be a document containing multiple entities and multiple relationships between the entities. The data preparation module 220 is configured to split the example into multiple training samples. For example, if the example includes a document with 20 entities, the data preparation module 220 may be configured to provide 20 training samples. Each training sample includes one of the 20 entities as the head entity among the 20 entities, the document, the relationship between the head entity and the corresponding tail entity, and the evidence statement in the document that supports the relationship. If a head entity has no relationship with any other entity, the sample may also be provided as a negative sample. Each of several samples can be used separately in the training process. The present disclosure improves the training efficiency by expanding one example into multiple training samples.

[0108] Figure 3 Shows a DocRED sample for training. As Figure 3 shown, the document includes seven statements, with the head entity being "The Legend of Zelda" and the tail entity being "Link". The relationship between the head entity and the tail entity is "Publisher". This relationship is supported by statements 0, 3, and 4.

[0109] Continuing to refer to Figure 2 , in some embodiments, the data preparation module 220 may prepare the input sample in the form of "[CLS]" + head entity + "[SEP]" + tail entity + "[SEP]" + document + "[SEP]", rather than in the form of "[CLS]" + head entity + "[SEP]" + document + "[SEP]". However, samples without a defined tail entity are preferred. By only defining the head entity and the document as the input, the training process is accelerated and the model parameters converge faster. In some embodiments, the data preparation module 220 is configured to store the prepared training samples in the training data 240.

[0110] The language model 222 is configured to process a training sample to obtain an output sequence after receiving one of the training samples from the data preparation module 220, and provide the output sequence to the relation prediction module 224 and the statement-relation combination module 226. In some embodiments, the output sequence is a plurality of vectors that correspond to the head entity, words in the document, and the [CLS] or [SEP] tokens. The vectors are also referred to as embeddings, and these embeddings are contextually aware. The language model 222 can be any pre-trained language model, such as GPT (Radford et al., 2018, Improving language understanding by generative pre-training), GPT-2 (Radford et al., 2019, Language models are unsupervised multitask learners), BERT (Devlin et al., 2018, BERT: pre-training of deep bidirectional transformers for language understanding), RoBERTa (Liu et al., 2019, RoBERTa: a robustly optimized BERT pretraining approach), and XLNet (Yang et al., 2019, XLNet: generalized autoregressive pretraining for language understanding), the cited references are incorporated herein by reference in their entirety. In some embodiments, the language model 222 is BERT-based, where L = 12, H = 768, and A = 12. L is the number of stacked encoders, H is the hidden size, and A is the number of heads in the multi-head attention layer. In some embodiments, the language model 222 is BERT large, where L = 24, H = 1024, and A = 16. The vectors represent the features of words, such as the meaning of the words and the positions of the words.

[0111] The relationship prediction module 224 is configured to, when an output sequence is available, extract a head entity from the output sequence, extract a tail entity from the output sequence, perform a first bilinear analysis on the vectors of the extracted head entity and tail entity to obtain a relationship prediction, and provide the relationship prediction to the statement-relationship combination module 226 and the update module 232. In some embodiments, the extracted head entity is a vector. When the head entity contains multiple words and the representation of the head entity in the output sequence is multiple vectors, the average of these vectors is taken as the representation of the extracted head entity. In some embodiments, for example, a Named Entity Recognizer may be used to perform the extraction of the tail entity. The extracted tail entity is a vector. Since there may be multiple tail entities in the document, there are multiple tail entity representations. These representations are vectors, and the average of the multiple vectors in the output sequence is obtained to get the extracted tail entity. Therefore, the vector of the tail entity is determined by its text and its position. In some embodiments, the extracted head entity is excluded from the extracted tail entity. In some embodiments, when extracting the head entity and the tail entity, the relationship prediction module 224 is configured to perform the relationship prediction using the above equation (1). The extracted head entity and the extracted tail entity form multiple head entity-tail entity pairs, and a relationship prediction is performed for each pair. For each head entity-tail entity pair, the bilinear analysis will provide a result for each relationship. If the result for any relationship regarding the head entity-tail entity pair is in the range of 0 to 1. In some embodiments, when the result is equal to or greater than 0.5, it indicates that the head entity-tail entity pair has a corresponding relationship. Otherwise, the head entity-tail entity pair does not have a corresponding relationship. In some embodiments, if the results for more than one relationship regarding the head entity-tail entity pair are equal to or greater than 0.5, it indicates that the head entity-tail entity pair has more than one relationship. If all the relationship results for the head entity-tail entity pair are less than 0.5, it means that there is no relationship between the head entity and the tail entity. During training, the relationship prediction module 224 does not need to provide the predicted relationship between the head entity and the tail entity to the statement-relationship combination module 226, because the statement-relationship combination module 226 will use the labels of the training samples to determine whether the head entity and the tail entity are related, and if so, what the relationship is. Instead, during actual prediction, the relationship prediction module 224 will provide the predicted relationship to the statement-relationship combination module 226. Specifically, the predicted / true relationship is an index used to query which correct relationship vector to use.

[0112] The statement-relation combination module 226 is configured to, when an output sequence is available, extract statements from the output sequence, provide relation vectors, combine the extracted statements with the corresponding relation vectors through a second bilinear layer, and provide the combination to the evidence prediction module 230. The document includes multiple statements, and the output sequence includes vectors of words in the statements. Generally, one word corresponds to one vector, but a long word or a special word can also be split into two or more vectors for representation. Statement extraction is performed by averaging the word vectors corresponding to the words in the statement to obtain a vector of the statement. The relation vectors correspond to the relations defined in the model, and the values of the relation vectors can be randomly initialized at the beginning of the training process. Then, the values of the relation vectors can be updated during the training process. Each relation has a corresponding relation vector. During the training process, when a head entity-tail entity pair is analyzed and predicted with a relation, the true relation corresponding to the head entity-tail entity pair can be obtained from the training data labels, and the relation vector corresponding to the true relation is selected for the combination. In the true prediction, when a head entity-tail entity pair is analyzed and predicted to have a relation, the predicted relation corresponds to a relation vector, and the relation vector corresponding to the predicted relation is selected for the combination. After selecting the relation vector, the statement-relation combination module 226 is configured to combine the statement extraction and the selected relation vector using the above equation (3). For example, if there are 4 statements and the relation vector corresponding to the selected true relation or predicted relation has 100 dimensions, the combination will be 4 100-dimensional statement vectors. After combination, the statement-relation combination module 226 is further configured to send the combination to the evidence prediction module 230.

[0113] The attention extraction module 228 is configured to, after the operation of the language model 222, extract attention probabilities from the last l layers of the language model 222 and send the extracted attention probabilities to the evidence prediction module 230. In some embodiments, the language model 222 is BERT base and l is in the range of 2 to 5. In one embodiment, the language model 222 is BERT base and l is 3. In some embodiments, depending on the specific language model used, l can be other values. In some embodiments, the present disclosure uses the attention probability values from the BERT multi-head attention layer as the attention extraction.

[0114] The evidence prediction module 230 is configured to, after receiving the combination of the statement extraction and the relation vectors from the statement-relation combination module 226 and the extracted attention probabilities from the attention extraction module 228, perform a third bilinear analysis on the combination and the extracted attention probabilities to obtain an evidence prediction and provide the evidence prediction to the update module 232. The evidence prediction provides the statement most relevant to the corresponding relation. In some embodiments, the third bilinear analysis is performed using equation (9).

[0115] The update module 232 is used to calculate a loss function using the relationship prediction result, the evidence prediction result, and the training data labels of the true relationship and the true evidence statement when the relationship prediction module 224 obtains a relationship prediction and the evidence prediction module 230 obtains an evidence prediction, and perform backpropagation to update the model parameters of the language model 222, the relationship prediction module 224, the statement-relationship combination module 226, the attention extraction module 228, and the evidence prediction module 230. In some embodiments, the loss function is in the form of equation (12). In some embodiments, after the model parameters are updated, the training process using the same training samples can be performed again with the new parameters. In other words, for each training sample, there may be multiple rounds of training to obtain optimized model parameters. For a training sample, the repeated multiple rounds of training can be ended after a predetermined number of rounds or after the model parameters converge.

[0116] Figure 4 Schematically shows a training process for multi-task prediction according to some embodiments of the present disclosure. In some embodiments, the training process is implemented by Figure 2 the server computing device shown. It should be noted in particular that unless otherwise specified in the present disclosure, the steps of the training process or method can be arranged in different orders, and thus are not limited to Figure 4 the order shown.

[0117] As Figure 4 shown, in step 402, the data preparation module 220 prepares training samples and sends the training samples to the language model 222. The training samples input to the language model 222 are in the form of a head entity and a document, such as "[CLS]" + head entity + "[SEP]" + document + "[SEP]". In addition, the labels of the tail entity, the relationship, and the evidence statement in the training samples are available for the update module 232 to calculate the loss function and perform backpropagation. In some embodiments, the data preparation module 220 can prepare the input samples in the form of a head entity, a tail entity, and a document, rather than in the form of a head entity and a document. However, training samples without defining the tail entity are preferred.

[0118] In step 404, after receiving the input training samples, the language model 222 processes the input training samples to obtain an output sequence and provides the output sequence to the relationship prediction module 224 and the statement-relationship combination module 226. The input samples are in text format, and the output sequence is in vector format. The output sequence has vectors corresponding to the head entity, vectors corresponding to the words in the document, and vectors corresponding to the [CLS] or [SEP] tokens, indicating the start of the sample, the end of the sample, the separation between the head entity and the context, and the separation between the statements in the context. The language model 222 can be any suitable pre-trained language model, such as GPT, GPT-2, BERT, roBERTa, and XLNet. Continuing to refer toFigure 1 The language model has multiple layers, and the last k layers 110 are used for attention extraction.

[0119] In step 406, the relation prediction module 224 extracts the head entity from the output sequence and extracts the tail entity from the output sequence. The extracted head entity and tail entity are vectors. Since the head entity is defined in the input sample and placed after the [CLS] token, the head entity can be directly extracted from the output sequence. The relation prediction module 224 also extracts multiple tail entities from the context part of the output sequence. In some embodiments, the tail entity can also be provided when preparing the training sample. In some embodiments, the relation prediction module 224 identifies and classifies key elements from the context as tail entities. For example, the tail entity can correspond to a location, a company, a person's name, a date, and a time. In some embodiments, a named entity recognizer is used to extract the tail entity. In some embodiments, the extracted head entity is excluded from the extracted tail entities.

[0120] In step 408, after extracting the head entity and the tail entity, the relation prediction module 224 predicts the relation between the extracted head entity and each extracted tail entity, provides the head-entity-tail-entity pair to the statement-relation combination module 226, and provides the predicted relation to the update module 232. In some embodiments, the relation prediction module 224 performs pairwise relation prediction between the head entity and each tail entity. For each head-entity-tail-entity pair, the relation prediction module 224 uses a bilinear layer to generate a predicted value for each relation. The total number of relations can be different. For example, 97 relations are listed in DocRED. When the value of one of the multiple relations of the head-entity-tail-entity pair is equal to or greater than 0.5, the head-entity-tail-entity is defined as having that relation. In some embodiments, a head-entity-tail-entity pair can have more than one relation. When none of the multiple relations of the head-entity-tail-entity pair has a value equal to or greater than 0.5, it is defined that the head-entity-tail-entity has no relation. In some embodiments, multiple relations can include a relation of "no relation", which is given when the head-entity-tail-entity pair has no relation. In some embodiments, equation (1) is used to perform relation prediction. Note that the relation prediction module 224 does not have to provide the predicted relation to the statement-relation combination module 226 during training, but needs to provide the predicted relation to the statement-relation combination module 226 during prediction.

[0121] In step 410, when the output sequence is available, the statement-relation combination module 226 extracts statements from the output sequence. Specifically, the context includes multiple statements, and each word in the statement is represented by one or several vectors in the output sequence. For each statement in the output sequence, the word vectors corresponding to the words in the statement are averaged to obtain a vector, which is called a statement vector. Therefore, the extracted statements include multiple statement vectors, and each statement vector represents a statement in the document.

[0122] In step 412, the statement-relation combination module 226 provides a relation vector, combines the extracted statement with one of the relation vectors, and sends the combination to the evidence prediction module 230. The relation vectors correspond to all the relations to be analyzed in the model, and one of the relation vectors used in the combination corresponds to the labeled relation of the head entity-tail entity pair in step 408. In some embodiments, the statement-relation combination is performed using equation (3). Note that the statement-relation combination module 226 obtains the relations between entities from the training samples, that is, from the true relation labels during training. For DocRED, the number of relation vectors can be 97. At the beginning of the training process, the values of the relation vectors can be randomly generated and updated during the subsequent training process. In some embodiments, the values of the relation vectors can also be stored in the training data 240. The statement-relation combination module 226 uses the second bilinear layer 126 to perform the combination of the statement vector and the relation vector corresponding to the true relation.

[0123] In step 414, after executing the language model 222, the attention extraction module 228 extracts the attention probabilities from the last l layers of the language model 222 and provides the attention probabilities to the evidence prediction module 230. In some embodiments, the value of l depends on the complexity of the language model 222 and the complexity of the training data. In some embodiments, when the language model 222 is a BERTbase model, l ranges from 2 to 5. In one embodiment, l is 3.

[0124] In step 416, after receiving the combination of the statement extraction and the relation vector from the statement-relation combination module 226 and the attention probabilities from the attention extraction module 228, the evidence prediction module 230 uses the third bilinear layer 130 to predict the evidence statements that support the relation. In some embodiments, the evidence prediction is performed using equation (9).

[0125] In step 418, when the relation prediction and the evidence prediction are completed, the update module 232 calculates the loss function using the prediction and the training data labels and updates the parameters of the model based on the loss function. In some embodiments, the loss function is calculated using equation (10).

[0126] In step 420, after updating the model parameters, the training steps 402-418 can be performed again using the same training samples to optimize the model parameters. The training iteration using the same samples can be completed after a predetermined number of iterations, or until the model parameters converge. After training with one training sample, the above training steps 402-420 are repeated using another training sample. When the parameters of the model converge after training with a set of samples, the model is fine-tuned and ready for prediction. In some embodiments, the training process can also be stopped when a predetermined number of rounds of the training process are executed.

[0127] Figure 5 Schematically shows a method for relationship prediction and evidence prediction according to some embodiments of the present disclosure. In some embodiments, the prediction process is implemented by Figure 2 the server computing device shown. It should be noted in particular that, unless otherwise specified in the present disclosure, the steps of the training process or method can be arranged in a different order and are thus not limited to Figure 5 the order shown.

[0128] As Figure 5 shown, method 500 includes steps 502-516, which are substantially the same as steps 402-416. However, the query sample includes the head entity and the context, but does not include the tail entity label, the relationship label, and the evidence statement label. Further, in step 412, the combination of the extracted sequence and the relationship vector uses the true labeled relationship corresponding to the head entity-tail entity pair analyzed in step 408 in the training sample, while in step 512, the relationship vector corresponding to the relationship predicted in step 508 is used in the combination.

[0129] In some embodiments, when the user only needs relationship prediction and does not need evidence prediction, the model can also be used to only operate the relationship prediction part, making the prediction speed faster. Although the evidence prediction including steps 510-516 is not necessary, a well-trained model still considers the contribution of evidence prediction through the language model parameters. Specifically, the relationship prediction still considers the statements that it should assign more weights to, so the relationship prediction is more accurate than not considering evidence prediction.

[0130] In some embodiments, the model can also define the head entity, the tail entity, and the context for the sample, rather than defining the head entity and the context for the sample. However, this consumes more computing resources, so the training samples and query samples formatted with the head entity and the context are more preferred than the samples formatted with the head entity, the tail entity, and the context.

[0131] In certain aspects, the present disclosure relates to a non-transitory computer-readable medium storing computer-executable code. In some embodiments, the computer-executable code can be software stored in the storage device 216 as described above. The computer-executable code, when executed, can execute one of the above methods.

[0132] Experiments were conducted to demonstrate the advantages of the embodiments of the present disclosure. The dataset used in the experiments is DocRED. DocRED is a large document-level dataset for relation extraction and evidence statement prediction tasks. It consists of 5,053 documents, 132,375 entities, and 56,354 relations mined from Wikipedia articles. For each (head, tail) entity pair, there are 97 different relation types as candidate relation types to be predicted, where the first relation type is the "NA" relation (i.e., no relation) between the two entities, and the rest correspond to a WikiData relation name. Each head / tail pair containing a valid relation also includes a set of supporting evidence statements. We split the data into training / validation / test following the same settings in (Yao et al., 2019) for model evaluation for fair comparison. The number of documents in training / validation / test is 3,000 / 1,000 / 1,000 respectively.

[0133] The dataset is evaluated using the relation extraction metric RE F1 and the evidence Evi F1. There are also cases where facts with relations appear in the validation and training sets, so we also evaluate Ign RE F1, where these facts with relations are removed.

[0134] The experimental settings are as follows. First is the hyperparameter setting. The configuration of the BERT-base model follows the settings in (Devlin et al., 2019). The learning rate is set to 1e -5 , λ1 is set to 1e -4 , the hidden dimension of the relation vector is set to 108, and the internal attention probabilities are extracted from the last three BERT layers.

[0135] Most of the experiments are conducted by fine-tuning the BERT-base model. This implementation is based on the PyTorch (Paszke et al., 2017) implementation of BERT. Our model is run for 60 iterations (epochs) on a single V100 GPU, resulting in approximately one day of training. The DocRED baseline and our E2GRE model have 115 million parameters.

[0136] Our model is compared with the following four published baseline models.

[0137] 1. Context-Aware BiLSTM. Yao et al., 2019 introduced the original baseline for DocRED in their paper. They used a context-aware BiLSTM (+ additional features such as entity type, co-reference, and distance) to encode the documents. Then the head and tail entities are extracted for relation extraction.

[0138] 2. BERT Two-Step Method. Wang et al., (2019) introduced fine-tuning BERT in a two-step process, where the model first predicts the NA relation and then predicts the remaining relations.

[0139] 3. HIN. Tang et al., (2020) introduced the use of a hierarchical inference network to help aggregate information from entities to statements and further to the document level for semantic reasoning of the entire document.

[0140] 4. BERT+LSR. Nan et al., (2020) introduced the use of an induced latent graph structure to help learn how information should flow between entities and statements in a document.

[0141] As Figure 6 shown in Table 1, our method E2GRE is the current state-of-the-art model on the DocRED public leaderboard.

[0142] Figure 7 Table 2 compares our method with the baseline models. It can be seen from Table 2 that our E2GRE method is not only competitive compared with the previous best method on the development set, but also has the following advantages compared with the previous models: (1) Compared with the HIN model and the BERT+LSR model, our method is more intuitive and simpler in design. In addition, our method provides interpretable relation extraction with supporting evidence prediction. (2) Our method is also superior to all other models in terms of the Ign RE F1 metric. This shows that our model does not memorize the relation facts between entities, but examines the relevant regions in the document to generate correct relation extraction.

[0143] Compared with the original BERT baseline, our training time is slightly longer due to multiple new entity-guided input sequences. We investigated the idea of generating new sequences based on each head-tail entity pair, but this method would scale quadratically with the number of entities in the document. Using our entity-guided method can achieve a balance between performance and training time.

[0144] We also conducted ablation studies. Figure 8Table 3 shows an ablation study of the effectiveness of our method for entity-guided and evidence-guided training. The baseline here is the joint training model for relation extraction and evidence prediction when using BERT-base. We see that entity-guided BERT improves by 2.5% over this baseline, and evidence-guided training further improves the method by 1.7%. This shows that both parts of our method are important for the entire E2GRE method. Compared with this baseline, our E2GRE method not only obtains an improvement in relation extraction F1, but also obtains a significant improvement in evidence prediction. This further shows that our evidence-guided fine-tuning method is effective, and evidence-guided joint training is more helpful for relation extraction.

[0145] We also conducted experiments to analyze the impact of the number of BERT layers used to obtain attention probability values. As Figure 9 shown in Table 4, it can be seen that using more layers is not necessarily better for relation extraction. A possible reason may be that the BERT model encodes more syntactic information in the intermediate layers (Clark et al., 2019).

[0146] Figure 3 shows an example in our model's validation set. In the Figure 3 example, the relationship between "The Legend of Zelda" and "Link" depends on information from multiple sentences in the given document.

[0147] Figure 10 shows the attention heatmap for simply applying BERT for relation extraction. This heatmap shows the attention for each word obtained from "The Legend of Zelda" and "Link". It can be seen that the model is able to locate the relevant regions of "Link" and "Legend of Zelda series", but the attention values for the rest of the document are very small. Therefore, the model has difficulty extracting information from the document to generate correct relation predictions.

[0148] In contrast, Figure 11 shows that our E2GRE model highlights the evidence sentences, especially in the regions where the model finds relevant information. Phrases related to "Link" and "The Legend of Zelda series" are assigned higher weights. The words connecting these phrases (such as "protagonist" or "involves") also have high weights. In addition, the scale of the attention probabilities of E2GRE is also much larger than that of the baseline. All these phrases and connecting words are located in the evidence sentences, enabling our model to also perform better in evidence prediction.

[0149] In summary, to more effectively expand the pre-trained LM for document-level RE, a new entity and evidence-guided relation extraction (E2GRE) method is provided. First, new entity-guided sequences are generated to feed into the LM, enabling the model to focus on relevant regions in the document. Then, the internal attention extracted from the last l layers is used to help guide the LM to focus on relevant regions in the document. Our E2GRE method improves the performance in RE and evidence prediction on the DocRED dataset and achieves advanced performance on the DocRED public leaderboard.

[0150] In some embodiments, our idea of using attention-guided multi-task learning is also combined with other NLP tasks with evidence statements. In some embodiments, our method is combined with graph-based models for NLP tasks. In some embodiments, our method is also combined with graph neural networks.

[0151] The foregoing description of the exemplary embodiments of the present disclosure is presented for purposes of illustration and description only and is not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings.

[0152] The embodiments are chosen and described in order to explain the principles of the present disclosure and its practical applications, so that others skilled in the art can utilize the present disclosure and various embodiments and various modifications suitable for the specific purposes contemplated. Alternative embodiments will become apparent to those skilled in the art to which the present disclosure pertains without departing from the spirit and scope of the present disclosure. Accordingly, the scope of the present disclosure is defined by the appended claims rather than the foregoing description and the exemplary embodiments described therein.

[0153] References (incorporated herein by reference in their entirety):

[0154] 1. Christoph Alt, Marc Hubner, and Leonhard Hennig, Improving relation extraction by pre-trained language representations, 2019, arXiv:1906.03088.

[0155] 2. Livio Baldini Soares, Nicholas FitzGerald, Jeffrey Ling, and Tom Kwiatkowski, Matching the blanks: Distributional similarity for relation learning, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019, 2895 - 2905.

[0156] 3. Razvan Bunescu and Raymond Mooney, A shortest path dependency kernel for relation extraction, Proceedings of Human Language Technology Conference and Conference on Empirical Methods in Natural Language Processing,.2005, 724 - 731.

[0157] 4. Rui Cai, Xiaodong Zhang, and Houfeng Wang, Bidirectional recurrent convolutional neural network for relation classification, Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics,.2016, 1:756 - 765.

[0158] 5. Fenia Christopoulou, Makoto Miwa, and Sophia Ananiadou, Connecting the dots: Document-level neural relation extraction with edge-oriented graphs, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2019, 4924 - 4935.

[0159] 6. Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning, What does BERT look at? an analysis of BERT’s attention, 2019, ArXiv, abs / 1906.04341.

[0160] 7. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, BERT: Pre-training of deep bidirectional transformers for language understanding, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2019a, 1:4171 - 4186.

[0161] 8. Markus Eberts and Adrian Ulges, Span-based joint entity and relation extraction with transformer pre-training, 2019, arXiv:1909.07755.

[0162] 9.Zhijiang Guo,Yan Zhang,and Wei Lu,Attention guided grapheonvolutional networks for relation extraction,Proceedings of the 57th AnnualMeeting of the Assoeiation for Computational Linguistics,.2019,241-251.

[0163] 10.Xu Han,Pengfei Yu,Zhiyuan Liu,Maosong Sun,and Peng Li,Hierarehicalrelation extraction with coarse-to-fine grained attention,Proceedings of the2018 Conference on Empirical Methods in Natural Language Processing,2018,2236-2245.

[0164] 11.Iris Hendrickx,Su Nam Kim,Zornitsa Kozareva,Preslav Nakov,DiarmuidO Seaghdha,Sebastian Pado,Marco Pennacchiotti,Lorenza Romano,and StanSzpakowicz,SemEval-2010 task 8:Multi-way classification of semantic relationsbetween pairs of nominals,Proceedings of the 5th International Workshop onSemantic Evaluation,2010,33-38.

[0165] 12. Robin Jia, Cliff Wong, and Hoifung Poon, Document-level n-ary relation extraction with multiscale representation learning, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,. 2019, 1: 3693-3704.

[0166] 13. Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy, Spanbert: Improving pre-training by representing and predicting spans, 2019, arXiv:1907.10529.

[0167] 14. Jiao Li, Yueping Sun, Robin J. Johnson, Daniela Sciaky, Chih-Hsuan Wei, Robert Leaman, Allan Peter Davis, Carolyn J. Mattingly, Thomas C. Wiegers, and Zhiyong Lu, BioCreative V CDR task corpus: a resource for chemical disease relation extraction, Database, 2016, doi:10.1093 / database / baw068.

[0168] 15. Guoshun Nan, Zhijiang Guo, Ivan Sekulic, and Wei Lu, Reasoning with latent structure refinement for document-level relation extraction, 2020, arXiv:2005.06312.

[0169] 16. Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, RoBERTa: A robustly optimized BERT pre-training approach, 2019, arXiv:1907.11692.

[0170] 17. Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer, Automatic differentiation in pytorch, 31st Conference on Neural Information Processing Systems (NIPS 2017), 2017.

[0171] 18. Nanyun Peng, Hoifung Poon, Chris Quirk, Kristina Toutanova, and Wentau Yih, Cross-sentence N-ary relation extraction with graph LSTMs, Transactions of the Association for Computational Linguistics, 2017, 5: 101-116.

[0172] 19. Chris Quirk and Hoifung Poon, Distant supervision for relation extraction beyond the sentence boundary, Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, 2017, 1: 1171-1182.

[0173] 20. Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever, Language models are unsupervised multitask learners, 2019.

[0174] 21. Linfeng Song, Yue Zhang, Zhiguo Wang, and Daniel Gildea, N-ary relation extraction using graph state LSTM, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2018, 2226-2235.

[0175] 22. Hengzhu Tang, Yanan Cao, Zhenyu Zhang, Jiangxia Cao, Fang Fang, Shi Wang, and Pengfei Yin, HIN: hierarchical inference network for document-level relation extraction, 2020, arXiv:2003.12754.

[0176] 23. Bayu Distiawan Trisedya, Gerhard Weikum, Jianzhong Qi, and Rui Zhang, Neural relation extraction for knowledge base enrichment, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019, 229-240.

[0177] 24. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, Attention is all you need, Proceedings of the 31st International Conference on Neural Information Processing Systems, 2017, 6000 - 6010.

[0178] 25. David Wadden, Ulme Wennberg, Yi Luan, and Hannaneh Hajishirzi, Entity, relation, and event extraction with contextualized span representations, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing(EMNLP - IJCNLP), 2019, 5783 - 5788.

[0179] 26. Hong Wang, Christfried Focke, Rob Sylvester, Nilesh Mishra, and William Wang, Fine - tune BERT for DocRED with two - step process, 2019, arXiv:1909.11898.

[0180] 27. Linlin Wang, Zhu Cao, Gerard de Melo, and Zhiyuan Liu, Relation classification via multi-level attention CNNs, Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, 2016, 1:1298-1307.

[0181] 28. Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le, XLNET: Generalized autoregressive pretraining for language understanding, Advances in Neural Information Processing Systems 32 (NIPS2019), 2019, 5754-5764.

[0182] 29. Yuan Yao, Deming Ye, Peng Li, Xu Han, Yankai Lin, Zhenghao Liu, Zhiyuan Liu, Lixin Huang, Jie Zhou, and Maosong Sun, DocRED: A large-scale document-level relation extraction dataset, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019, 764-777.

[0183] 30. Tom Young, Erik Cambria, Iti Chaturvedi, Hao Zhou, Subham Biswas and Minlie Huang, Augmenting end-to-end dialog systems with commonsense knowledge, thirty-Second AAAI Conference on Artificial Intelligence, 2018, 4970-4977.

[0184] 31. Mo Yu, Wenpeng Yin, Kazi Saidul Hasan, Cicero dos Santos, Bing Xiang, and Bowen Zhou, Improved neural relation detection for knowledge base question answering, Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, 2017, 1:571-581.

[0185] 32. Dmitry Zelenko, Chinatsu Aone, and Anthony Richardella, Kernel methods for relation extraction, Journal of Machine Learning Research,.2003, 3:1083-1106.

[0186] 33. Daojian Zeng, Kang Liu, Siwei Lai, Guangyou Zhou, and Jun Zhao, Relation classification via convolutional deep neural network, Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers, 2014, 2335-2344.

[0187] 34. Yuhao Zhang, Victor Zhong, Danqi Chen, Gabor Angeli, and Christopher D. Manning, "Position-aware attention and supervised data improve slot filling", Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (EMNLP2017), 2017, 35 - 45.

[0188] 35. Yi Zhao, Huaiyu Wan, Jianwei Gao, and Youfang Lin, "Improving relation classification by entity pair graph", Proceedings of The Eleventh Asian Conference on Machine Learning, Proceedings of Machine Learning Research, 2019, 10l: 1156 - 1171.

Claims

1. An entity and evidence-guided relationship prediction system, comprising a computing device, the computing device including a processor and a storage device storing computer-executable code, wherein, The computer-executable code is configured, when executed by a processor, to: provide a head entity and a document containing the head entity; process the head entity and the document through a language model to obtain a head extraction corresponding to the head entity, a tail extraction corresponding to a tail entity in the document, and a statement extraction corresponding to a statement in the document; predict a head entity-tail entity relationship between the head extraction and the tail extraction using a first bilinear layer; combine the statement extraction and a relationship vector corresponding to the head extraction and the tail extraction using a second bilinear layer to obtain a statement-relationship combination; predict evidence statements from the document using a third bilinear layer based on the statement-relationship combination and attention extracted from the language model, wherein the evidence statements support the head-tail relationship; and update parameters of the language model, the first bilinear layer, the second bilinear layer, and the third bilinear layer based on the predicted head entity-tail entity relationship, the predicted evidence statements, and a label of the document, wherein the label of the document includes a true head entity-tail entity relationship and true evidence statements.

2. The system according to claim 1, wherein a plurality of annotated documents are used to train the language model, the first bilinear layer, the second bilinear layer, and the third bilinear layer, at least one of the annotated documents has E entities, the at least one annotated document is extended to E samples, and each of the E samples includes the at least one annotated document and a head entity corresponding to one of the E entities, where E is a positive integer.

3. The system according to claim 1, wherein the computer-executable code is configured to update the parameters based on a loss function, the loss function being defined by defined, is a relationship prediction loss, is a statement prediction loss, is a weight factor, the weight factor having a value equal to or greater than 0.

4. The system according to claim 1, wherein the language model includes at least one of the following: Generative Pretrained Transformer GPT, GPT-2, Bidirectional Encoder Representations from Transformers BERT, Robustly Optimized BERT Approach roBERTa, and Reparameterized Transformer-XL Network XLnet.

5. The system according to claim 1, wherein, The computer-executable code is configured to extract the attention from the last 2 to 5 layers of the language model.

6. The system according to claim 5, wherein, The computer-executable code is configured to extract the attention from the last 3 layers of the language model.

7. The system according to claim 1, wherein the first bilinear layer is defined by and is the predicted value of the i-th relationship among multiple relationships between the head entity h and the k-th tail entity , represents the sigmoid function, are the learned weights of the first bilinear layer, is the bias of the first bilinear layer.

8. The system according to claim 7, wherein the second bilinear layer is defined by Define, , Is the j-th statement in the document Is the prediction probability of the supporting statement for the i-th relationship , And Are the learnable parameters of the second bilinear layer with respect to the i-th relationship.

9. The system according to claim 8, wherein the third bilinear layer is defined by Define, is the predicted probability that the j-th statement in the document is a supporting statement for the i-th relation regarding the k-th tail entity, represents the sigmoid function, is the learned weight of the third bilinear layer, is the bias of the third bilinear layer.

10. The system according to claim 1, wherein after training, the language model, the first bilinear layer, the second bilinear layer, and the third bilinear layer are configured to provide a relationship prediction and an evidence prediction for a query entry, wherein the query entry has a query head entity and a query document containing the query head entity.

11. The system according to claim 1, wherein, The computer-executable code is further configured to provide the tail entity of the document.

12. An entity and evidence-guided relationship prediction method, comprising: providing, by a computing device, a head entity and a document containing the head entity; processing, by a language model stored in the computing device, the head entity and the document to obtain a head extraction corresponding to the head entity, a tail extraction corresponding to a tail entity in the document, and a statement extraction corresponding to a statement in the document; predicting, by a first bilinear layer stored in the computing device, a head-tail relationship between the head extraction and the tail extraction; The second bilinear layer stored in the computing device combines the statement extraction and the relationship vectors corresponding to the head extraction and the tail extraction to obtain a statement-relationship combination; Based on the statement-relationship combination and the attention extracted from the language model, a third bilinear layer stored in the computing device predicts evidence statements from the document, where the evidence statements support the head-tail relationship; and Based on the predicted head entity-tail entity relationship, the predicted evidence statements, and the label of the document, update the parameters of the language model, the first bilinear layer, the second bilinear layer, and the third bilinear layer, where the label of the document includes the true head entity-tail entity relationship and the true evidence statements.

13. The method according to claim 12, wherein the step of updating the parameters is performed based on a loss function, the loss function being defined by defined, is a relationship prediction loss, is a statement prediction loss, is a weight factor, the weight factor having a value equal to or greater than 0.

14. The method according to claim 12, wherein the language model includes at least one of the following: Generative Pretrained Model GPT, GPT-2, Bidirectional Encoder Representations from Transformers BERT, Robustly Optimized BERT Approach roBERTa, and Reparameterized Transformer-XL Network XLNet.

15. The method according to claim 12, extracting the attention from the last 3 layers of the language model.

16. The method according to claim 12, wherein the first bilinear layer is defined by and is the predicted value of the i-th relationship among multiple relationships between the head entity h and the k-th tail entity , represents the sigmoid function, is the learning weight of the first bilinear layer, and is the bias of the first bilinear layer; wherein the second bilinear layer is defined by and , is the j-th statement in the document is the predicted probability of the supporting statement for the i-th relation , and are learnable parameters of the second bilinear layer with respect to the i-th relation; and wherein the third bilinear layer is defined by and is the predicted probability that the j-th statement in the document is a supporting statement for the i-th relation regarding the k-th tail entity, denotes the sigmoid function, are the learned weights of the third bilinear layer, is the bias of the third bilinear layer.

17. The method according to claim 12, further comprising, after training: providing relationship prediction and evidence prediction for a query entry, where the query entry has a query head entity and a query document including the query head entity.

18. A non-transitory computer-readable medium storing computer-executable code, wherein, When the computer-executable code is executed by a processor of an active computing device, it is configured to: Provide a head entity and a document including the head entity; Process the head entity and the document through a language model to obtain a head extraction corresponding to the head entity, a tail extraction corresponding to a tail entity in the document, and a statement extraction corresponding to a statement in the document; Use a first bilinear layer to predict a head entity-tail entity relationship between the head extraction and the tail extraction; Use a second bilinear layer to combine the statement extraction and the relationship vectors corresponding to the head extraction and the tail extraction to obtain a statement-relationship combination; Based on the statement-relationship combination and the attention extracted from the language model, use a third bilinear layer to predict evidence statements from the document, where the evidence statements support the head-tail relationship; and Based on the predicted head entity-tail entity relationship, the predicted evidence statements, and the label of the document, update the parameters of the language model, the first bilinear layer, the second bilinear layer, and the third bilinear layer, where the label of the document includes the true head entity-tail entity relationship and the true evidence statements.

19. The non-transitory computer-readable medium according to claim 18, wherein the computer-executable code is configured to update the parameters based on a loss function, the loss function being defined by defined, is a relationship prediction loss, is a statement prediction loss, is a weight factor, the weight factor having a value equal to or greater than 0.

20. The non-transitory computer-readable medium according to claim 18, wherein the first bilinear layer is defined by and is the predicted value of the i-th relationship among multiple relationships between the head entity h and the k-th tail entity , represents the sigmoid function, are the learned weights of the first bilinear layer, is the bias of the first bilinear layer; wherein the second bilinear layer is defined by and , is the j-th statement in the document is the predicted probability of the supporting statement for the i-th relation ; and and are learnable parameters of the second bilinear layer with respect to the i-th relation; and wherein the third bilinear layer is defined by as is the predicted probability that the j-th statement in the document is a supporting statement for the i-th relation regarding the k-th tail entity, denotes the sigmoid function, are the learned weights of the third bilinear layer, is the bias of the third bilinear layer.

Citation Information

Patent Citations

  • General text mining method and system

    CN106156035A

  • Inter-drug relationship extraction method based on attention of multiple entities and improved pre-training language model

    CN111078889A