A Relationship Extraction Method for a Knowledge Graph Automatic Construction System
By combining the multi-head attention mechanism and the weighted dependency matrix, using graph convolution network to extract the relationship and extract information from different dimensions in the model, the performance and generalization problems of the existing models when dealing with natural language complexity are solved, and the excellent effect on the public data set is achieved.
Patent Information
- Application Number
- CN202111133794.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-27
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-09-27
AI Technical Summary
When dealing with the complexity of natural language, existing relationship extraction models have problems such as artificial noise, limited performance and weak generalization. Especially when using syntactic dependency structures, there is artificial noise in the pruning method, while soft pruning destroys the dependency structure.
The multi-head attention mechanism and the weighted dependency matrix are adopted. By combining the word vector and the dependency matrix, the graph convolution network is used to extract information of different dimensions of the text, and relationship prediction is performed through the feedforward neural network.
Excellent results were achieved on the public data set of relationship extraction, improving model performance, reducing time costs, and effectively utilizing the rich information in the syntactic dependency structure.
Smart Images

Figure CN113901758B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing and artificial intelligence, and particularly relates to a relation extraction method for an automatic knowledge graph construction system. Background Art
[0002] Relation extraction is a key subtask in the field of natural language processing and an important part of the information extraction task. The purpose of relation extraction is to extract the relation information between entities from unstructured text. By combining with the named entity recognition task, triples in the form of <subject, predicate (relation), object> required for constructing a knowledge graph can be generated.
[0003] Traditional relation extraction methods mainly analyze the text by using linguistic knowledge, and use statistical and rule-based methods to perform text matching and relation extraction by manually designing extraction rules or kernel functions. However, due to the complexity of natural language, the relation extraction model based on artificial rules can no longer meet the performance requirements of people. The model often introduces artificial noise, with limited performance and weak generalization ability.
[0004] With the rapid development of neural networks and deep learning, researchers have begun to introduce neural networks into the relation extraction task. Neural networks and deep learning methods can effectively fit and extract text features by simulating the working principle of brain neurons, breaking the limitations of manually designed rules. Existing neural network-based relation extraction models are mainly divided into sequence-based models and dependency-based models.
[0005] The sequence-based model encodes the word sequence in a sentence, and uses a convolutional neural network to extract the distance position features of words in the sequence relative to the entity. As a time series model, the recurrent neural network is more sensitive to the relationship between distant entity pairs. Combining with the convolutional neural network can effectively alleviate the problem that it is difficult to obtain the relationship information between distant words in the text. However, the sequence-based model focuses on the word sequence and ignores the overall syntactic structure information of the sentence.
[0006] Compared with sequence-based models, dependency-based models can effectively utilize the syntactic structure features of sentences and capture other implicit long-distance syntactic relationships. Dependency-based models usually convert a sentence into a dependency tree according to the dependency relationships between words, and then further convert the dependency tree into a corresponding dependency adjacency matrix to participate in the training of the neural network, capturing implicit long-distance syntactic relationships and multi-hop relationships through the dependency relationships between each word. Since the dependency structure is usually a graph structure, graph convolutional neural networks have also been introduced into the dependency-based relation extraction model. The current main work focuses on how to perform more effective pruning on the dependency tree, pruning information irrelevant to relation extraction to improve the model performance. However, rule-based pruning also has artificial noise, while soft pruning based on the attention mechanism destroys the original dependency structure and cannot fully utilize the rich information contained in the dependency matrix. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide a relation extraction method for a knowledge graph automatic construction system, which uses the multi-head attention mechanism and the weighted dependency matrix to parallelly obtain key information in different dimensions of the text, and has achieved excellent results on the public dataset of relation extraction.
[0008] The present invention provides a relation extraction method for a knowledge graph automatic construction system, including the following steps:
[0009] Step S1: Embed each word in the text using a pre-trained word vector dictionary, and convert the part-of-speech tagging information and named entity recognition information of each word into vector representations and concatenate them with the vector representation of the word itself to obtain vector x. i ;
[0010] Step S2: Perform bidirectional long short-term memory network operations on vector x. i Concatenate the forward operation result and the backward operation result to obtain vector h'. t ;
[0011] Step S3: Construct a syntactic dependency tree A through the syntactic dependency structure of the text, set a learnable weight variable D, use A to construct a dependency adjacency matrix and one-hot the matrix values, and multiply them bit by bit with the weight variable D to obtain a weighted dependency matrix A'.
[0012] Step S4: Obtain the feature representation matrices Q and K of the text through vector h'. t Use the multi-head attention mechanism to obtain k attention matrices of the text. After linear dimensionality reduction, obtain matrix A''.
[0013] Step S5: Use matrix A' and matrix A'' as the inputs of a graph convolutional module with different numbers of graph convolutional network layers to perform graph convolutional operations, and respectively obtain matrices. and matrix Obtain matrix H after linear dimensionality reduction output ;
[0014] Step S6: Obtain the feature representation matrix h of the sentence from matrix H output and the feature representation matrices of two entities sent and and Use a feedforward neural network to obtain the relational feature representation matrix h relation , and finally perform relational prediction through the softmax function to obtain the final classification result.
[0015] As a further technical solution of the present invention, in step S1, the vector wherein, the vector w i is the word vector of the word itself, and the vectors and the vector are the word vectors of the part-of-speech tagging information and named entity recognition information of the word respectively, and a concatenation operation is performed.
[0016] Further, in step S2, calculate the hidden state vector h of vector x i in a certain direction at time t, and the formula is as follows t :
[0017] I t =σ(x t W xi +h t-1 W hi +b i )
[0018] F t =σ(x t W xf +h t-1 W hf +b f )
[0019] O t =σ(x t W xo +h t-1 W ho +b o )
[0020]
[0021] h t =O t ⊙tanh(C t )
[0022] wherein, x tis the input at time t, σ is the sigmoid activation function, tanh is the hyperbolic tangent activation function, W xi 、W xf 、W xo and W xc are the weight parameter matrices of x t in the input gate, forget gate, output gate, and memory cell respectively, W hi 、W hf 、W ho and W hc are the weight parameter matrices of h t in the input gate, forget gate, output gate, and memory cell respectively, b i 、b f 、b o and b c are the bias parameters of the input gate, forget gate, output gate, and memory cell respectively, I t 、F t 、O t 、 and C t are the outputs of the input gate, forget gate, output gate, candidate memory cell, and memory cell at time t respectively, ⊙ is element-wise multiplication of matrices;
[0023] Concatenate the forward output and the of the backward output to obtain the output h′ t is
[0024] Furthermore, the calculation formula for the weighted dependency matrix A′ in step 3 is
[0025] A′ = φ(onehot(A)·D)
[0026] φ(x) = max(x, 0);
[0027] where A is the original dependency tree, onehot is the one-hot operation, φ is the ReLU activation function, and max is the maximum value.
[0028] Furthermore, the formula for the attention matrix in step S4 is
[0029]
[0030] where k is the number of multi-head attention heads, Q and K are the feature representations obtained by the text through steps S1 and S2, and are the weight parameter matrices, d is the input dimension, softmax is the normalized exponential function, and the k attention matrices are concatenated and then reduced in dimension through a linear layer to obtain A″, and the formula is
[0031]
[0032] Among them, W A and b A are the weight parameter matrix and bias parameter of the linear transformation layer.
[0033] Furthermore, the formula for calculating the result of each graph convolution module in step S5 is
[0034]
[0035] output o = W o [input 0 ; GCN(input 0 );..; GCN(output i-1 )]
[0036] input c = W c [input i-1 ; output 0 ;...; output N )
[0037]
[0038] Among them, in the graph convolution network of layer L, the set of initial input feature representations is The node i of the l-th layer receives as input and outputs W (l) is the graph convolution network weight parameter matrix, b (l) is the graph convolution network bias parameter, N is the number of graph convolution layers of the previous sub-module, M is the number of sub-modules, W o , W c , W f are all linear transformation layer weight parameter matrices.
[0039] Furthermore, the two graph convolution modules obtained in step S5 respectively generate and After splicing and linear dimensionality reduction, H output is obtained.
[0040] Furthermore, in step S6, the formula for calculating the final relationship feature is
[0041]
[0042] Among them, FFNN represents the calculation of the feedforward neural network.
[0043] The advantages of the present invention are as follows:
[0044] 1. The present invention encodes the text using word vectors pre-trained on a large-scale dictionary and a bidirectional long short-term memory network to obtain an initial vector representation of the text. The initial vector representation already contains part of the text feature information, and this vector is used as the input to the subsequent neural network model.
[0045] 2. The present invention uses a multi-head attention mechanism to obtain multiple attention matrices of the text, and each attention matrix is obtained from different important parts of the text. The multi-head attention mechanism can obtain important information other than the syntactic dependency information of the text, and effective feature extraction is performed through the graph convolutional network module.
[0046] 3. The present invention uses a weighted dependency matrix to obtain the syntactic dependency structure information of the text, assigns learnable weights to each relationship category, and through neural network iterative update, transforms the dependency matrix from a 0-1 matrix into a weighted matrix, enabling it to express more and more accurate syntactic structure information, and feature extraction is performed through the graph convolutional network module.
[0047] 4. The multi-head attention mechanism and the weighted dependency matrix in the invention respectively extract key information in different dimensions of the text, and can perform parallel computing at the same time, reducing the time cost while improving the performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a schematic diagram of the relationship extraction model of the present invention,
[0049] Figure 2 is a schematic diagram of the process of constructing a weighted dependency matrix in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0050] Please refer to Figure 1 , this embodiment provides a relationship extraction method for a knowledge graph automatic construction system of the present invention, which uses a multi-head attention mechanism and a weighted dependency matrix to parallelly obtain key information in different dimensions of the text. The specific steps are as follows:
[0051] Step 1: Perform initial word embedding on each word in the original text through a pre-trained word vector dictionary to obtain the vector representation w i of each word, where i is the i-th word in the text. Additionally, the part-of-speech tagging information and named entity recognition information of each word are also transformed into vector representations, obtaining vectors and which are concatenated with the vector representation of the word itself, and finally obtaining as the final word embedding vector representation of each word.
[0052] Step 2: For the vector representation x of each word obtained in Step 1 i perform bidirectional long short-term memory network operations to encode the forward and backward sequences of the sentence, and obtain the hidden state vector H = [h1, h2,..., h n , where n is the number of words in the sentence. At time t, the hidden state vector h t in a certain direction is calculated as follows:
[0053] I t = σ(x t W xi + h t-1 W hi + b i ) (1)
[0054] F t = σ(x t W xf + h t-1 W hf + b f ) (2)
[0055] O t = σ(x t W xo + h t-1 W ho + b o ) (3)
[0056]
[0057]
[0058] h t = O t ⊙ tanh(C t ) (8)
[0059] where x t represents the input at time t, σ represents the sigmoid activation function, tanh represents the hyperbolic tangent activation function, W xi , W xf , W xo and W xc represent the weight parameter matrices of x t in the input gate, forget gate, output gate, and memory cell respectively, W hi , W hf , W ho and W ho represent the weight parameter matrices of h t in the input gate, forget gate, output gate, and memory cell respectively, b i , b f , bo and b c represent the bias parameters of the input gate, forget gate, output gate, and memory cell respectively. I t , F t , O t , and C t represent the outputs of the input gate, forget gate, output gate, candidate memory cell, and memory cell at time t respectively. ⊙ represents element-wise multiplication of matrices. At time t, the final output h′ t is obtained by concatenating the forward output and the backward output, and the calculation formula is as follows:
[0060]
[0061] Step 3: Construct a syntactic dependency tree based on the syntactic dependency structure of the text. If the sentence contains n words, the dependency tree has n nodes, and the dependency tree can be transformed into an n×n dependency adjacency matrix A. If there is a dependency relationship between word a and word b, then A ab = 1, otherwise A ab = 0. Set the learnable weight variable D = [d1, d2,..., d Q , where Q is the number of relationship categories in the dataset, and d q is the weight of the relationship category with index q, defaulting to 1. First, we replace the values outside the main diagonal of the dependency matrix with the indices of the corresponding relationship categories in the weight variable D. For index q, construct a one-hot vector r q = [0,..., 0, 1, 0,...0], where r q [q] = 1 and the rest of the values are 0, so that the weight variable D can participate in the neural network calculation through element-wise matrix multiplication to achieve parameter update, and then obtain the weight scalar through matrix summation to keep the shape of the dependency matrix constant. For the original dependency tree A, the formula for constructing the adjacency matrix A′ is as follows:
[0062] A′ = φ(onehot(A)·D) (10)
[0063] φ(x) = max(x, 0) (11) where onehot represents the one-hot operation, φ represents the ReLU activation function, and max represents taking the maximum value. Specifically as Figure 2 shown.
[0064] Step 4: Apply the multi-head attention mechanism directly to the text to obtain k (k is the number of multi-head attention heads) attention matrices of the same size as the dependency matrix A, and the calculation formula is as follows:
[0065]
[0066] Among them, Q and K are the feature representations obtained by the text through steps 1 and 2, and is the weight parameter matrix, d represents the input dimension, and softmax represents the normalized exponential function. After obtaining k matrices they are concatenated and then dimensionally reduced through a linear layer to obtain A″ as the input of the graph convolution module. The calculation formula is as follows:
[0067]
[0068] where W A and b A are the weight parameter matrix and bias parameter of the linear transformation layer.
[0069] Step 5: Use A′ and A″ obtained in steps 3 and 4 as the input of the graph convolution module. For each graph convolution module, use graph convolution networks with different depths as sub-modules, and use dense connections inside the sub-modules to obtain the output output of each graph convolution layer (l) , which is concatenated and used as the input of the next sub-module. Finally, the outputs of all sub-modules are densely connected to obtain In the graph convolution network of layer L, the initial input feature representation set is The node i of the l-th layer receives as the input and outputs The calculation formula is as follows:
[0070]
[0071] output o = W o [input 0 ; GCN(input 0 );..; GCN(ooutput i-1 )] (15)
[0072] input c = W c [input i-1 ; output 0 ;...; output N (16
[0073]
[0074] where W (l) represents the weight parameter matrix of the graph convolution network, b (l)Denote the bias parameters of the graph convolutional network, N is the number of graph convolutional layers of the previous sub-module, M is the number of sub-modules, and W o 、W c 、W f are all weight parameter matrices of the linear transformation layer. In formula (17), for each calculated input i (i≥1), a dropout operation is performed to randomly discard neurons. Figure 1 The two graph convolutional modules in and generate output respectively. After splicing and linear dimensionality reduction, H
[0075] Step 6: Obtain the feature representations h output of the sentence, the feature representations sent of the two entities, and and from H relation obtained in Step 5 respectively. Use a feed-forward neural network to obtain the final relational feature representation h
[0076]
[0077] where FFNN represents the calculation of the feed-forward neural network.
[0078] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above specific embodiments, and the above specific embodiments and the descriptions in the specification are only for further illustrating the principles of the present invention. Without departing from the spirit scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the claims and their equivalents.
Claims
1. A method for relation extraction for an automatic knowledge graph construction system, characterized in that, including the following steps Step S1: Embed each word in the text using a pre-trained word vector dictionary, and convert the part-of-speech tagging information and named entity recognition information of each word into vector representations and concatenate them with the vector representation of the word itself to obtain vector x i ; Step S2: Perform a bidirectional long short-term memory network operation on the vector x i and splice the results of the forward operation and the backward operation to obtain the vector h' t ; Step S3: Construct a syntactic dependency tree A through the syntactic dependency structure of the text, set a learnable weight variable D, use A to construct a dependency adjacency matrix and one-hot encode the matrix values, and multiply them bitwise with the weight variable D to obtain a weighted dependency matrix A'; Step S4: Obtain the feature representation matrices Q and K of the text through the vector h' t Obtain the feature representation matrices Q and K of the text, and linearly reduce the dimensionality of the k attention matrices of the multi-head attention mechanism text to obtain the matrix A″; Step S5: Use matrix A' and matrix A'' as the inputs of graph convolution modules with different numbers of graph convolution network layers to perform graph convolution operations, and respectively obtain matrix and matrix After linear dimensionality reduction, obtain matrix H output ; Step S6: Obtain the feature representation matrix h of the sentence from the matrix H output and the feature representation matrices of the two entities sent and Use a feedforward neural network to obtain the relational feature representation matrix h relation , and finally perform relational prediction through the softmax function to obtain the final classification result; In the above step S3, the calculation formula of the weighted dependency matrix A' is A’ = φ(onehot(A)·D) φ(x) = max(x, 0); where A is the original dependency tree, onehot is the one-hot encoding operation, φ is the ReLU activation function, and max is the operation of taking the maximum value; The attention matrix in step S4 has the formula Among them, k is the number of multi-head attention heads, Q and K are the feature representations obtained by the text through steps S1 and S2, W i Q and are weight parameter matrices, d is the input dimension, softmax is the normalized exponential function, and after splicing k attention matrices it is reduced in dimension through a linear layer to obtain A″, and the formula is Among them, W A and b A are the weight parameter matrix and bias parameter of the linear transformation layer; In the above step S6, the formula for calculating the final relationship feature is where FFNN represents the calculation of a feedforward neural network.
2. The method for relation extraction for an automatic knowledge graph construction system according to claim 1, characterized in that, In the step S1, the vector where the vector w i is the word vector of the word itself, and the vectors and the vector are the word vectors of the part-of-speech tagging information and named entity recognition information of the word respectively, and a concatenation operation is performed.
3. The method for relation extraction for an automatic knowledge graph construction system according to claim 1, characterized in that, In the step S2, for the vector x i the hidden state vector h in a certain direction at time t t is calculated, and the formula is as follows I t = σ(x t W xi +h T-1 W hi +b i ) F t = σ(x t W xf + h T-1 W hf + b f ) O t = σ(x t W xo + h T-1 W ho + b o ) h t = O t ⊙tanh(C t ) where, x t is the input at time t, σ is the sigmoid activation function, tanh is the hyperbolic tangent activation function, W xi , W xf , W xo and W xc are the weight parameter matrices of x t at the input gate, forget gate, output gate, and memory cell respectively, W hi , W hf , W ho and W hc are the weight parameter matrices of h t at the input gate, forget gate, output gate, and memory cell respectively, b i , b f , b o and b c are the bias parameters of the input gate, forget gate, output gate, and memory cell respectively, I t , F t , O t , and C t are the outputs of the input gate, forget gate, output gate, candidate memory cell, and memory cell at time t respectively; ⊙ is element-wise multiplication of matrices; Forward output and backward output are concatenated to obtain the obtained output h′ t as 4. The method for relation extraction for an automatic knowledge graph construction system according to claim 1, characterized in that, In the above step S5, the formula for calculating the result of each graph convolution module is output o = W o [input 0 ; GCN(input 0 ));..; GCN(output i-1 )] input c = W c [input i-1 ; output 0 ;...; output N Among them, in the graph convolutional network of layer L, the initial input feature representation set is Node i in the l-th layer receives as input and outputs W (l) is the weight parameter matrix of the graph convolutional network, b (l) is the bias parameter of the graph convolutional network, N is the number of graph convolutional layers in the previous sub-module, M is the number of sub-modules, W o 、W c 、W f are all weight parameter matrices of the linear transformation layer.
5. The relation extraction method of a knowledge graph-oriented automatic construction system according to claim 1, characterized in that The two graph convolution modules obtained in step S5 respectively generate and which, after splicing and linear dimensionality reduction, result in H output .