Chinese Entity Relation Extraction Method Based on Gating Mechanism and Graph Attention Network
By introducing a gated mechanism and graph attention network in the Chinese entity relationship extraction technology, combining the Chinese BERT model and the mask self-attention mechanism, the problem of relying on artificial features and high time complexity in the existing technology is solved, and a more efficient and accurate Chinese entity relationship extraction is achieved.
Patent Information
- Application Number
- CN202210281501.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-21
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-03-21
AI Technical Summary
The existing Chinese entity relationship extraction technology has the problem of relying on manual design characteristics and high time complexity, and it is difficult to effectively extract entity relationships in Chinese texts.
Using a method based on gating mechanism and graph attention network, the Chinese BERT pre-trained model is used to convert text into vector form, and combining the global information gating mechanism and mask self-attention mechanism to extract entities and relationship characteristics in sentences.
It improves the accuracy and robustness of Chinese entity relationship extraction, can extract core semantics more effectively in sentences, and enhances the ability to classify relationships.
Smart Images

Figure CN114722820B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a Chinese entity relation extraction method based on a gating mechanism and a graph attention network, and belongs to the field of information extraction for Chinese. Background Art
[0002] In recent years, the Internet has witnessed booming development, with a large amount of information flooding the network, and humanity has entered the big data era. In this era of massive information, how to quickly obtain important information has become an urgent problem to be solved. The information extraction technology (IE) was proposed to serve this problem. As a technology to liberate human labor, its purpose is to automatically and efficiently extract specific and valuable information from semi-structured or unstructured texts and store this information in a reasonable structure in a storage medium. Named Entity Recognition (NER), Relation Extraction (RE), and Event Extraction are important subtasks of information extraction.
[0003] Relation extraction methods can be divided into methods based on traditional machine learning and methods based on deep learning. Machine learning-based methods include two methods based on feature vectors and kernel functions. The former relies too much on manually designed features and natural language analysis tools, and the latter has too high time complexity. Deep learning-based methods propose an end-to-end relation classification method, which automatically extracts classification features of sentences relying on the learning ability of neural networks and is a hot research direction of current relation extraction technology.
[0004] The entity relation extraction technology for Chinese has great research value. The professional text data in various industries in China and the text information on the Chinese Internet are an astronomical figure that is difficult to estimate. It is difficult to complete the integration and sorting of these data by relying on pure manual methods. Moreover, data is being generated on the network all the time. To achieve the goal of sustainable development, only by applying this efficient computer technology to the field of Chinese information can all information be properly applied. Summary of the Invention
[0005] Object of the Invention: Aiming at the technical problems existing in the prior art, a Chinese entity relation extraction method based on a gating mechanism and a graph attention network is provided.
[0006] Technical Solution: A Chinese entity relation extraction method based on a gating mechanism and a graph attention network, the specific steps are as follows:
[0007] Step 1) Use the Chinese BERT pre-trained model to convert the text into a vector form recognizable by a machine;
[0008] Step 2) Concatenate the entity embeddings for classification in the sentence behind each word embedding, adopt a global information gating mechanism, calculate the gating vector, and achieve the enhancement of entity semantics of the word embedding;
[0009] Step 3) Conduct dependency syntactic analysis on the text to obtain a dependency syntactic tree, construct an adjacency matrix, a dependency type matrix, and a dependency direction matrix, use masked self-attention to obtain an attention weight matrix, and then extract features from the text sentence in the graph attention network;
[0010] Step 4) From the output of the graph attention network, obtain the representation vectors of the two entities and the sentence, convert the representation vectors to the classification space through a multi-layer perceptron, and input them into the classifier to complete the relationship classification.
[0011] Furthermore, using the Chinese BERT pre-trained model in step 1) to convert the text into a vector form recognizable by the machine, that is, text to word vectors; includes the following processes:
[0012] 1-1) Split the sentence s into a character sequence, and then call the BERT pre-trained model to vectorize the character sequence to form a character vector sequence {c 1 , c 2 , …};
[0013] 1-2) Use an off-the-shelf natural language processing tool to tokenize the sentence to obtain a word sequence;
[0014] 1-3) Utilize the character vector sequence {c 1 , c 2 , …} in the first step to initialize the word sequence as a word vector sequence {v 1 , v 2 , …}, and the rule is that the word vector is the average of the sum of the vectors of the characters it contains.
[0015] Furthermore, the enhancement of entity semantics of the word embedding in step 2) refers to enhancing the entity semantics of the word embedding obtained by the BERT model conversion, specifically including the following steps:
[0016] 2-1) Concatenate the embedding representations of the head entity and the tail entity, and then use a feed-forward network to fuse the semantic information of the two entities, and its process is formulated as:
[0017] v e = tanh(W e [v h , v t + b e )
[0018] where, v h and vt The word embeddings corresponding to the head and tail entities respectively is a learnable parameter matrix for the linear transformation of concatenating entity embeddings is a bias term, tanh is the hyperbolic tangent function, v e is the entity embedding that fuses the information of the head and tail entities
[0019] 2-2) Concatenate the fused entity embedding with the word vectors of each word in the sentence to initially obtain the candidate word vectors enhanced by the entity embedding. At the same time, sum and average these candidate word vectors to obtain the supervision vector that fuses the global information and the entity embedding The above process can be expressed as:
[0020]
[0021]
[0022] where represents the candidate word vector of the i-th word, and n is the number of words in sentence s
[0023] 2-3) Taking the supervision vector and the candidate word vector as inputs, output the gating vector corresponding to each candidate word vector:
[0024]
[0025] where is a parameter matrix to be trained (d g = d w + d e ) is the bias term, and the operator ⊙ represents element-wise multiplication of the two vectors at both ends. The output range of the sigmoid function is (0, 1)
[0026] 2-4) Calculate the word embedding representation of the i-th word after entity embedding enhancement Multiply the candidate word vector of this word with the corresponding gating vector element-wise. The calculation process is as follows:
[0027]
[0028] Furthermore, in order to extract features from the sentence using the graph attention network, first, it is necessary to construct the adjacency matrix, dependency type matrix, and dependency direction matrix based on the dependency syntactic tree of the sentence, then calculate the attention transfer weights using the masked self-attention mechanism, and finally update the node features using the adjacent aggregation mechanism
[0029] The construction rules of the adjacency matrix are as follows: Suppose there are n nodes in the dependency syntax tree, then an n×n adjacency matrix A can be used to represent the dependency syntax tree; when there is a dependency edge between node i and node j, the elements a i,j and a j,i in A are 1, otherwise 0. At this time, A is an undirected graph; in particular, in the adjacency matrix converted from the dependency tree, each node has a self-loop edge, that is, all a i,i are equal to 1. The construction methods of the dependency type matrix T and the dependency direction matrix D are similar. The size of the dependency type matrix T is n×n. If the dependency type between node i and node j is nsubj, then the value of the element t i,j in T is type_to_id_mapping(nsubj), where type_to_id_mapping represents the mapping from the dependency type to a numerical value. The size of the dependency direction matrix D is also n×n, and its element values are only 1 and -1. The element d i,j =1 indicates that the corresponding dependency edge is forward, that is, there is a dependency edge i→j. Conversely, d j,i =-1 indicates that the dependency edge is backward.
[0030] The graph attention network adopts a masked self-attention mechanism, which calculates a brand-new attention weight matrix for information transmission at each layer. This weight matrix not only completely preserves the structural information of the original dependency syntax tree but also assigns different weights to the dependency edges. The word vector sequence enhanced by entity embedding is and is also the initial input of the graph attention network The symbol represents the hidden layer vector of the i-th word at layer l. In an L-layer graph attention network, at layer l, a masked self-attention is used to calculate an n×n adjacency matrix P (l) (the same size as A), which represents the weight of the dependency edge between node i and node j. The calculation process is as follows:
[0031]
[0032] where a i,j represents the weight in the original adjacency matrix, which only has two values, 0 and 1. When a i,j =0, the corresponding value calculated by self-attention is When a i,j =1, represents the embedding vector of the dependency type corresponding to the dependency type t i,j . fun(·) is an attention function used to calculate the importance value of the dependency edge between two nodes, that is, e i,j . The specific details of this attention function are shown in the following formula:
[0033]
[0034] Among them, LeakyRelu is the activation function, and [·] represents the concatenation operation of vectors. f(d i,j ) represents a function related to the direction of the dependency edge. If the dependency edge i→j is forward, the forward parameter matrix is used; otherwise, the backward parameter matrix is used. The details are shown below:
[0035]
[0036] Among them, correspond to the forward and backward learnable weight matrices respectively.
[0037] Combining the attention weight matrix of this layer and the network input, calculate the hidden layer vector of each node in the graph attention network. The process is represented by the following formula:
[0038]
[0039] Among them, W (l) represents the learnable weight matrix of the l-th layer graph attention network, and b (l) is the bias term.
[0040] Furthermore, in step 4), the max-pooling operation is used to obtain the representation vectors of the two entities in the sentence and the representation vector (h sent ) expressing the semantics of the whole sentence. Then, the three vectors are concatenated to obtain the output vector h full ; the output vector cannot be directly used for relation classification and needs to be input into a multi-layer perceptron to be transformed into the classification space. The process is as follows:
[0041] o = M d h full + b
[0042] Among them, is the learnable weight matrix, which changes the dimension of the relation feature vector into the classification space, is the bias vector, and |R| is the number of predefined relation types.
[0043] Then, o is input into the softmax classifier to obtain the normalized probability distribution of the entity pair relation classification. The probability that the sentence s is classified as the true relation r is:
[0044]
[0045] Among them, h and t are the head and tail entities included in the sentence s, r represents the true relation of the entity pair, and o k represents the k-th element of the vector o.
[0046] Optimize the parameters of the entity relation extraction model. Optionally, use the Stochastic Gradient Descent (SGD) method. The objective function is the commonly used Cross Entropy loss function in classification tasks, and its definition is as follows:
[0047]
[0048] where θ represents the training parameters of the model, |B| represents the number of instances in a training batch, s i is the i-th sentence in the training batch B, and r i corresponds to the true relation of the entity pair <h i , t i [[ID=15>>.
[0049] A computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above computer program, it implements the Chinese entity relation extraction method based on the gating mechanism and the graph attention network as described above.
[0050] A computer-readable storage medium stores a computer program that executes the Chinese entity relation extraction method based on the gating mechanism and the graph attention network as described above.
[0051] Advantages of the present invention:
[0052] 1) The described relation extraction method uses the BERT model pre-trained with a large-scale corpus to complete the transformation from text to vector, and fine-tunes the BERT model during the training process to make the word embeddings contain more context semantics. Using this vector as the input of the subsequent module enables the network model to have stronger representation ability;
[0053] 2) The described relation extraction method attaches great importance to the semantics of the two entities themselves. Since relation extraction is the prediction of the relation between two entities, the semantics of the entities have an important impact on the prediction result. Although existing methods highlight the role of entities in sentences through relative position features and other means, the mining of entity semantics is not deep enough. Therefore, the present invention proposes a global information gating mechanism to embed entity information into all word embedding representations, realizing the enhancement of the word embedding representation. In this way, the robustness of the method can be enhanced;
[0054] 3) The described relation extraction method has the ability to extract the core semantics of sentences. The graph attention network has a stronger and more reasonable neighbor aggregation ability compared to the graph convolutional network. The graph convolutional network can only use hard pruning to highlight the core semantics of each node, but the artificially defined pruning strategy will cause huge semantic errors under complex context conditions. The graph attention network dynamically learns the information transfer weight matrix through the self-attention mechanism, which belongs to a soft pruning method, assigning different weights to each dependency edge, that is, giving different adjacent nodes different information preferences. This method introduces a dependency type matrix and a dependency direction matrix, and calculates a more reasonable attention weight matrix using the type information and direction information of the dependency edges, enhancing the ability to extract the core semantics of sentences. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a flowchart of the method according to an embodiment of the present invention;
[0056] Figure 2 It is a framework diagram of the Chinese entity relation extraction model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The present invention will be further clarified below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification of the present invention by those skilled in the art all fall within the scope defined by the appended claims of this application.
[0058] This embodiment takes predicting the relation type between "company" and "chair" in the Chinese sentence "This company produces plastic chairs" as an example to illustrate the method execution process.
[0059] As Figure 1 shown, a Chinese entity relation extraction method based on a gating mechanism and a graph attention network, its specific implementation process includes the following steps:
[0060] Step 1) Use the Chinese BERT pre-trained model to convert the text into a vector form recognizable by a machine:
[0061] 1-1) Split the sentence s into a character sequence, and then call the BERT pre-trained model to vectorize the character sequence to form a character vector sequence {c 1 , c 2 ,...};
[0062] 1-2) Adopt an existing natural language processing tool, such as jieba, to tokenize the sentence to obtain a word sequence;
[0063] 1-3) Utilize the character vector sequence {c 1 , c 2, …}, initialize the word sequence as a sequence of word vectors {v 1 , v 2 , …}, and the rule is that the word vector is the average of the vectors of the characters it contains.
[0064] Step 2) Concatenate the entity embedding behind each word embedding, and adopt a global information gating mechanism to calculate the gating vector to achieve the enhancement of the entity semantics of the word embedding.
[0065] 2-1) Concatenate the embedding representations of the head entity and the tail entity, and then use a feed-forward network to fuse the semantic information of the two entities. The process is formulated as:
[0066] v e = tanh(W e [v h , v t + b e )
[0067] where v h and v t correspond to the word embeddings of the head and tail entities respectively, is a learnable parameter matrix for the linear transformation of concatenating entity embeddings, is a bias term, tanh is the hyperbolic tangent function, and v e is the entity embedding that fuses the information of the head and tail entities.
[0068] 2-2) Concatenate the fused entity embedding with the word vectors of each word in the sentence to initially obtain the candidate word vectors enhanced by the entity embedding. At the same time, sum and average these candidate word vectors to obtain the supervision vector that fuses global information and entity embedding The above process can be expressed as:
[0069]
[0070]
[0071] where represents the candidate word vector of the i-th word, and n is the number of words in sentence s.
[0072] 2-3) Use the supervision vector and candidate word vectors as inputs to output the gating vector corresponding to each candidate word vector:
[0073]
[0074] where is a parameter matrix to be trained, (d g = d w + d e) is the bias term. The operator ⊙ represents element-wise multiplication of the two vectors at both ends. The output range of the sigmoid function is (0, 1).
[0075] 2-4) Calculate the word embedding representation of the i-th word after entity embedding enhancement Multiply the candidate word vector of the word by the corresponding gating vector element-wise. The calculation process is as follows:
[0076]
[0077] Step 3) Perform dependency syntactic analysis on the text to obtain a dependency syntax tree, construct an adjacency matrix, a dependency type matrix, and a dependency direction matrix, and use masked self-attention to obtain an attention weight matrix, and then perform feature extraction on the text in the graph attention network.
[0078] 3-1) Construct an adjacency matrix A according to the dependency syntax tree of the sentence. Assume there are n nodes in the dependency syntax tree, then an n×n adjacency matrix A can be used to represent the dependency syntax tree. When there is a dependency edge between node i and node j, the elements a i,j and a j,i in A are 1, otherwise 0. At this time, A is an undirected graph. In particular, in the adjacency matrix converted from the dependency tree, each node has a self-loop edge, that is, a i,i = 1.
[0079] Construct a dependency type matrix T and a dependency direction matrix D according to the dependency type and dependency direction information provided by the dependency syntax tree. The size of the dependency type matrix T is n×n. If the dependency type between node i and node j is nsubj, the value of the element t i,j in T is type_to_id_mapping(nsubj). The size of the dependency direction matrix D is also n×n. If there is a dependency edge i→j, the element d i,j = 1 indicates that the dependency edge is forward, and conversely, d j,i = -1 indicates that the dependency edge is backward.
[0080] For example, the dependency parsing result of the text sentence ["this", "company", "produces", "plastic", "chairs"] is [("det", 2), ("nsubj", 3), ("ROOT", 0), ("amod", 5), ("dobj", 3)], and the corresponding adjacency matrix is [[1, 1, 0, 0, 0], [1, 1, 1, 0, 0], [0, 1, 1, 0, 1], [0, 0, 0, 1, 1], [0, 0, 1, 1, 1]]. According to the construction rules of the dependency direction matrix, the constructed dependency direction matrix is [[1, -1, 0, 0, 0], [1, 1, -1, 0, 0], [0, 1, 1, 0, 1], [0, 0, 0, 1, -1], [0, 0, -1, 1, 1]]. To obtain the dependency type matrix, numerical mapping of the dependency types is required, and the following mapping table can be used to complete this process:
[0081] Dependency type Mapping id Dependency type Mapping id Self-loop 1 iobj 21 advcl 2 cc:preconj 22 discourse 3 parataxis 23 cc 4 det:predet 24 dep 5 conj 25 cop 6 nsubj 26 csubjpass 7 acl:relcl 27 compound 8 nmod 28 aux 9 csubj 29 compound:prt 10 nsubjpass 30 dobj 11 acl 31 case 12 nmod:tmod 32 punct 13 ROOT 33 amod 14 mark 34 appos 15 mwe 35 advmod 16 nmod:npmod 36 det 17 ccomp 37 auxpass 18 neg 38 expl 19 nummod 39 xcomp 20 Nmod:poss 40
[0082] The dependency type matrix constructed using the above mapping rules is [[1, 17, 0, 0, 0], [17, 1, 26, 0, 0], [0, 26, 1, 0, 11], [0, 0, 0, 1, 14], [0, 0, 11, 14, 1]].
[0083] 3 - 2) The graph attention network adopts a masked self - attention mechanism, and calculates a brand - new attention weight matrix for information transfer at each layer. This weight matrix not only completely preserves the structural information of the original dependency syntax tree, but also assigns different weights to the dependency edges. The word vector sequence enhanced by entity embedding is which is also the initial input of the graph attention network The symbol represents the hidden layer vector of the i - th word at the l - th layer. In a graph attention network with L layers, the l - th layer will use masked self - attention to calculate an n×n adjacency matrix P (l) (the same size as A), representing the weight of the dependency edge between node i and node j, and its calculation process is as follows:
[0084]
[0085] where, a i,j represents the weight in the original adjacency matrix, which has only two values, 0 and 1. When a i,j = 0, it corresponds to the obtained by self - attention calculation. When a i,j = 1, represents the dependency type t i,jThe embedding vector corresponding to the dependent type, and fun(·) is an attention function used to calculate the importance value of the dependent edge between two nodes, that is, e i,j , and the specific details of this attention function are shown in the following formula:
[0086]
[0087] where LeakyRelu is the activation function, [·] represents the concatenation operation of vectors, and f(d i,j ) represents a function related to the direction of the dependent edge. If the dependent edge i→j is forward, the forward parameter matrix is used; otherwise, the reverse parameter matrix is used. The details are shown below:
[0088]
[0089] where, correspond to the forward and reverse learnable weight matrices respectively.
[0090] 3-3) Combine the attention weight matrix of this layer and the network input to calculate the hidden layer vector of each node in the graph attention network. The process is shown in the following formula:
[0091]
[0092] where, W (l) represents the learnable weight matrix of the l-th layer of the graph attention network, and b (l) is the bias term.
[0093] Step 4) Obtain the representation vectors of the two entities and the sentence from the output of the graph attention network, convert the vector to the classification space through a multi-layer perceptron, and input it into the classifier to complete the relationship classification.
[0094] Use the max pooling operation to obtain the representation vectors of the two entities in the sentence and the representation vector (h sent ) expressing the semantics of the whole sentence, and then concatenate the three vectors to obtain the output vector h full ; The output vector cannot be directly used for relationship classification and needs to be input into a multi-layer perceptron to be converted to the classification space. The process is as follows:
[0095] o = M d h full + b
[0096] where, is the learnable weight matrix, which changes the dimension of the relationship feature vector to the classification space,
[0097] is the bias vector, and |R| is the number of predefined relationship types.
[0098] Then, input o into the softmax classifier to obtain the normalized probability distribution for entity pair relationship classification. The probability that the sentence s is classified as the true relationship r is:
[0099]
[0100] where h and t are the head and tail entities included in the sentence s, r represents the true relationship of the entity pair, and o k represents the k-th element of the vector o.
[0101] Obviously, those skilled in the art should understand that each step of the above Chinese entity relationship extraction method based on the gating mechanism and the graph attention network in the embodiments of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module for implementation. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.
Claims
1. A Chinese entity relation extraction method based on a gating mechanism and a graph attention network, characterized in that, it includes the following steps: Step 1) Use the Chinese BERT pre-trained model to convert the text into a vector form that can be recognized by a machine; Step 2) Concatenate the entity embeddings behind each word embedding, and adopt a global information gating mechanism to calculate the gating vector to achieve the enhancement of the entity semantics of the word embedding; Step 3) Perform dependency syntactic analysis on the text to obtain a dependency syntax tree, construct an adjacency matrix, a dependency type matrix, and a dependency direction matrix, use masked self-attention to obtain the attention weight matrix, and then perform feature extraction on the text sentence in the graph attention network; Step 4) From the output of the graph attention network, obtain the representation vectors of two entities and the sentence, convert the representation vectors to the classification space through a multi-layer perceptron, and input them into the classifier to complete the relation classification; The enhancement of the entity semantics of the word embedding in the said Step 2 refers to enhancing the entity semantics of the word embedding obtained by the BERT model conversion, which specifically includes the following steps: 2-1) Concatenate the embedding representations of the head entity and the tail entity, and then use a feed-forward network to fuse the semantic information of the two entities, and its process is formulated as: v e = tanh(W e [v h ,v t + b e ) Among them, v h and v t correspond to the word embeddings of the head and tail entities respectively, is a learnable parameter matrix for the linear transformation of splicing entity embeddings, is a bias term, tanh is the hyperbolic tangent function, and ve is the entity embedding that fuses the information of the head and tail entities; 2-2) Concatenate the fused entity embedding with the word vectors of each word in the sentence to initially obtain the candidate word vectors with enhanced entity embedding; at the same time, sum and average these candidate word vectors to obtain the supervision vector s that fuses the global information and the entity embedding. The above process can be expressed as: Among them, represents the candidate word vector of the i-th word, and n is the number of words in sentence s; 2-3) Use the supervision vector and the candidate word vectors as inputs to output the gating vector corresponding to each candidate word vector: Among them, is a parameter matrix to be trained, (d g = d w + d e ) is the bias term. The operator ⊙ represents element-wise multiplication of the two end vectors. The output range of the sigmoid function is (0, 1); 2-4) Calculate the word embedding representation of the $i$-th word after entity embedding enhancement Multiply the candidate word vector of this word and the corresponding gating vector element-wise. The calculation process is as follows: In the said Step 3, an adjacency matrix, a dependency type matrix, and a dependency direction matrix are constructed according to the dependency syntax tree of the sentence, and a masked self-attention mechanism is used to calculate the attention transfer weight, and then feature extraction is performed on the text in the graph attention network. The specific steps include: 3-1) Construct the adjacency matrix A according to the dependency syntactic tree of the sentence. Suppose there are n nodes on the dependency syntactic tree, then an n×n adjacency matrix A can be used to represent the dependency syntactic tree; when there is a dependency edge between node i and node j, the elements a i,j and a j,i in A are 1, otherwise 0. At this time, A is an undirected graph; in particular, in the adjacency matrix converted from the dependency tree, each node has a self-loop edge, that is, a i,i = 1; According to the dependency type and dependency direction information provided by the dependency syntax tree, the construction method of the dependency type matrix T and the dependency direction matrix D; the size of the dependency type matrix T is n×n. If the dependency type between node i and node j is nsubj, then the element t i,j in T has a value of type_to_id_mapping(nsubj), where type_to_id_mapping represents the mapping from the dependency type to a numerical value; the size of the dependency direction matrix D is also n×n. If there is a dependency edge i→j, then the element d i,j = 1 indicates that this dependency edge is forward, and conversely, d j,i = -1 indicates that the dependency edge is backward; 3-2) The graph attention network adopts a masked self-attention mechanism to calculate a brand-new attention weight matrix for information transmission at each layer. This weight matrix not only completely preserves the structural information of the original dependency syntactic tree but also assigns different weights to the dependency edges; the sequence of word vectors enhanced by entity embedding is which is also the initial input of the graph attention network symbol represents the hidden layer vector of the i-th word at layer l; in an L-layer graph attention network, at layer l, a masked self-attention is used to calculate an n×n adjacency matrix P (l) , represents the weight of the dependency edge between node i and node j, and its calculation process is as follows: Among them, a i,j represents the weight in the original adjacency matrix, which only takes two values, 0 and 1. When a i,j = 0, it corresponds to the When a i,j = 1, represents the embedding vector of the corresponding dependency type t i,j . fun(·) is an attention function used to calculate the importance value of the dependency edge between two nodes, that is, e i,j . The specific details of this attention function are shown in the following formula: Among them, LeakyRelu is the activation function, [·] represents the concatenation operation of vectors, and f(d i,j ) represents a function related to the direction of the dependency edge. If the dependency edge i→j is forward, the forward parameter matrix is used; otherwise, the reverse parameter matrix is used. The details are shown below: Among them, correspond to the forward and backward learnable weight matrices respectively; 3-3) Combine the attention weight matrix of this layer and the network input to calculate the hidden layer vector of each node in the graph attention network, and its process is expressed as follows: Among them, W (l) represents the learnable weight matrix of the l-th layer graph attention network, and b (l) is the bias term.
2. The Chinese entity relation extraction method based on a gating mechanism and a graph attention network according to claim 1, characterized in that, the use of the Chinese BERT pre-trained model in the said Step 1 to convert the text into a vector form that can be recognized by a machine, that is, text to word vectors; includes the following process: 1-1) Split the sentence s into a sequence of characters, and then call the BERT pre-trained model to vectorize the character sequence, forming a sequence of character vectors {c 1 , c 2 , …}; 1-2) Use a ready-made natural language processing tool to segment the sentence to obtain a word sequence; 1-3) Using the character vector sequence {c 1 , c 2 , …} of the first step, initialize the word sequence as the word vector sequence {v 1 , v 2 , …}, and the rule is that the word vector is the average of the sum of the vectors of the characters it contains.
3. The Chinese entity relation extraction method based on a gating mechanism and a graph attention network according to claim 1, characterized in that, In step 4), the maximum pooling operation is used to obtain the representation vectors of the two entities in the sentence and the representation vector (h sent ) that expresses the semantics of the entire sentence. Then, the three vectors are concatenated to obtain the output vector h full ; The output vector cannot be directly used for relation classification and needs to be input into a multi-layer perceptron to be transformed into the classification space. The process is as follows: o = M d h full + b Among them, is a learnable weight matrix that changes the dimension of the relational feature vector into the classification space, is a bias vector, and |R| is the number of predefined relation types; Then input o into the softmax classifier to obtain the normalized probability distribution of the entity pair relation classification, and the probability that the sentence s is classified as the true relation r is: Among them, h and t are the head and tail entities contained in sentence s, r represents the true relationship of this entity pair, and o k represents the k-th element of vector o.
4. The Chinese entity relation extraction method based on a gating mechanism and a graph attention network according to claim 1, characterized in that, Optimize the parameters of the entity relation extraction model, adopt the stochastic gradient descent method, and the objective function is the cross-entropy loss function commonly used in the classification task, and its definition is as follows: where θ represents the training parameters of the model, |B| represents the number of instances in a training batch, si is the i-th sentence instance in the training batch B, and r i corresponds to the entity pair <h i , t i > for the true relationship.
5. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the Chinese entity relation extraction method based on the gating mechanism and the graph attention network as described in any one of claims 1-4.
6. A computer-readable storage medium, characterized in that: the computer-readable storage medium stores a computer program for executing the Chinese entity relation extraction method based on the gating mechanism and the graph attention network as described in any one of claims 1-4.
Citation Information
Patent Citations
Entity relationship extraction method and device based on syntactic tree and graph attention mechanism
CN113255320A
Method for extracting chapter relation by fusing multi-level information extraction and noise reduction
CN113435190A