Open domain information extraction algorithm based on graph neural network and relation discrimination

By combining pre-trained models and graph neural network methods, the problem of traditional open domain information extraction methods being poor in complex texts is solved, and better information extraction results are achieved, which are suitable for question-and-answer systems and knowledge graph construction.

CN120492629APending Publication Date: 2025-08-15CHENGDU QUANTUM MATRIX TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410722793.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-06-05
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Traditional open domain information extraction methods have limited effects when facing complex texts and cannot make full use of text-dependent information.

Method used

Combining the pre-trained model and graph neural network, entity relationship extraction is performed through the global pointer network, multi-grained relationship dependency graph is constructed, context encoding is used for BERT, entity relationship graph embedding is performed based on the graph attention network, and SPO triple extraction is performed through the dual affine network.

Benefits of technology

It improves the information extraction ability in complex text scenarios, can better utilize text-dependent knowledge, and is suitable for question-and-answer systems and knowledge graph construction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_4
    Figure SMS_4
  • Figure SMS_6
    Figure SMS_6
Patent Text Reader

Abstract

The invention provides an open domain information extraction algorithm based on a graph neural network and relation discrimination. The method comprises the following steps of: extracting an entity relationship in a text and finishing construction of an entity relationship graph; performing text embedding representation and context coding on the input text; fusing the entity relationship to perform graph information embedding; and decoding is completed by using a double affine network, and an SPO triple extraction result is obtained. The invention aims to better solve the problem of open domain information extraction. According to the algorithm, technologies such as a pre-training model and a graph neural network are combined, meanwhile, the thought of relation judgment is combined, entity relation information is embedded into the network in a topological graph structure, and therefore the model can better utilize dependency knowledge in a text. Therefore, the problem that a traditional open domain information extraction method is limited in effect when facing complex texts and the defect that text dependent information cannot be fully utilized are overcome. According to the method, the open domain information extraction requirement in a complex scene can be met, and a technical basis can be provided for downstream tasks such as a question answering system and knowledge graph construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention proposes an open domain information extraction algorithm based on graph neural network and relationship discrimination, which belongs to the field of natural language processing. Background Art

[0002] A pre-trained model is a deep learning model. Unlike traditional deep learning models, its most notable feature is that it is pre-trained on large amounts of data. This type of training often takes a long time and requires significant computing resources to learn the common knowledge across all the data. A trained model typically requires only a small amount of data labeling and fine-tuning before it can be applied to various downstream tasks. In the field of natural language processing, common pre-trained models include BERT, ELMo, and GPT. These models are largely based on the Transformer architecture and attention mechanism, learn from large amounts of text data, and possess strong language comprehension capabilities. They perform well on various tasks and have been applied in multiple fields.

[0003] Graph neural networks are deep learning networks specifically designed to process graph data structures. Due to the unique properties of graph data, including topology, sparsity, and locality, traditional machine learning and deep learning methods are not well suited for processing graph data. In graph neural networks, the connectivity between nodes and edges is highly flexible. Each node can connect to a varying number and type of neighbors, forming graphs with diverse topologies. This allows graph neural networks to better transfer node features, capturing both global and local information and leveraging relationships between entities within graph data. Graph neural networks also have a variety of applications in natural language processing tasks, including text classification, sentiment analysis, entity relationship extraction, and grammatical parsing. Summary of the Invention

[0004] This paper proposes an open domain information extraction algorithm based on graph neural network and relationship discrimination. The overall calculation structure of the algorithm is shown in the figure Figure 1 As shown in Figure 2, the present invention aims to better solve the problem of open-domain information extraction. By combining pre-trained models with graph neural networks and incorporating the concept of relationship identification, the algorithm embeds entity relationship information into the network using a topological graph structure, enabling the model to better utilize dependency knowledge in the text. This addresses the limited effectiveness of traditional open-domain information extraction methods for complex text and their inability to fully utilize text dependency information.

[0005] The technical solutions of the present invention are as follows:

[0006] First, entity relationship triples are extracted from the text and a relationship graph is constructed. This paper innovatively proposes an entity relationship extraction method based on a global pointer network to efficiently extract entity relationship triples from the input text. It also designs a multi-granularity relationship dependency graph that combines contextual encoding features with entity relationship features through entity relationship links and word character connections.

[0007] After completing the entity relationship diagram construction in step one, the text input algorithm needs to be embedded and contextualized. This step primarily addresses the discrepancy between the input of the deep learning network being a numeric vector and the original dataset being textual, so the text representation must first be vectorized. The calculations involved in this process primarily include calculating the embedding layer and generating positional embeddings. After obtaining the embedding vectors, the BERT model is used for contextual encoding, allowing the word vectors to learn contextual meaning and enhance their features, allowing the model's output vectors to be used for downstream tasks.

[0008] Next, we need to use a graph neural network to embed the entity relationship graph. This step uses a graph attention network as the graph encoder. The graph attention network's ability to dynamically weight nodes while aggregating information from adjacent nodes, combined with the context-sensitive encoding representation from the previous step, gives the model greater adaptability and flexibility.

[0009] Finally, a bi-affine network is used to complete decoding and extract the SPO triples. This step combines the graph embedding results from the previous step to complete the SPO triple extraction and decoding process. In this step, the entity extraction task is considered as the problem of identifying the head and tail indexes, assigning a type to the interval formed by the head and tail, and using a bi-affine network to solve the problem of hidden layer vector information interaction.

[0010] The beneficial result of the present invention is that an open domain information extraction algorithm based on graph neural networks and relationship discrimination is proposed. This invention combines pre-trained models and graph neural networks and other technologies, while incorporating the idea of relationship discrimination, and embeds entity relationship information into the network in the structure of a topological graph, so that the model can better utilize the dependency knowledge in the text. This solves the problem that traditional open domain information extraction methods have limited effectiveness when facing complex texts and the disadvantage of not being able to fully utilize text dependency information. The present invention can meet the needs of open domain information extraction in complex scenarios and provide a technical foundation for downstream tasks such as question-answering systems and knowledge graph construction.

[0011] Description of the accompanying drawings and tables

[0012] Figure 1 It is the overall structure diagram of the algorithm of the present invention;

[0013] Figure 2Predict graph for global pointer network;

[0014] Figure 3 MRDG diagram proposed by the present invention

[0015] Figure 4 Transformer model schematic Specific implementation methods

[0016] Step 1: Entity Relationship Diagram Construction

[0017] In order to obtain better contextual embedding features and acquire multi-hop dependency knowledge, an entity relationship extraction model is designed in this step to extract entity relationship information from the input text. The global pointer network model adopts the design idea of the attention mechanism and uses entities as the basic unit for judgment. Its prediction method is as follows Figure 2 As shown. Specifically, for a text sequence S of length n, assuming that only one sequence label needs to be identified, and assuming that each entity to be identified is a continuous segment of the sequence, with unlimited length and can be nested within each other (there is an intersection between two entities), it can be seen that the sequence has a total of n(n+1) / 2 different continuous candidate entity subsequences, which contain all possible answers. Therefore, the task is to select the true answer from these n(n+1) / 2 "candidate entities", which is a "n(n+1) / 2 choose m" multi-label classification problem. Furthermore, when the original problem is expanded to require the identification of K entity types, the target problem will be transformed into K "n(n+1) / 2 choose m" multi-label classification problems.

[0018] Assume that the original input is a text sequence S of length n, and the hidden vector sequence after context encoding is [h1, h2, ..., h n ], for the vector sequence used by the αth type of entity, the key-value matrix corresponding to this category can be obtained through linear transformation. The specific operations are as follows:

[0019] q i,α =W q,α h i +b q,α

[0020] k i,α =W k,α h i +b k,α

[0021] Where W q,α and W k,α is the key-value weight matrix of the model under category α; b q,α and b k,α is the key-value bias matrix under category α.

[0022] Through this operation, we can get the transformed sequence vector sequence [q 1,α ,q 2,α ,…,q n,α ] and [k 1,α , k 2,α ,…,k n,α ], which are the vector sequences used to identify the αth type of entity. At this point, the score function for each position can be defined as follows:

[0023]

[0024] Where s α (i, j) represents the score of the continuous segment from i to j being an entity of type α.

[0025] For the sequence S=[w1,w2,…w n ] is a continuous substring S consisting of the i-th to j-th elements [wi:wj] The model uses q i,α With k j,α The inner product of is the score of the entity of type α. After combining the rotation position encoding, the original formula is as follows. α (i, j) injects relative position information.

[0026]

[0027] In the specific implementation, since S(s h , s t )、S(o h , o t ) is used to identify the entities corresponding to subject and object, which is equivalent to the NER task with two entity types, which can be completed with a global pointer network; as for S(s h , o h |p) and S(s t , o t |p), which are used to identify the predicate as p (s h , o h ) for and (s t , o t ) Yes, unlike NER, there is no need for s h ≤o h or s t ≤o t The constraints are also completed using a global pointer network, but the default lower triangular mask matrix of the global pointer network needs to be removed.

[0028] Finally, in order to complete the training process of the model, the loss function needs to be designed. As mentioned above, in the global pointer network, for an entity recognition problem of type α, it is converted into a classification problem of "n(n+1) / 2 select m". A naive idea is to regard the result as n(n+1) / 2 binary classifications to calculate the loss function. However, in actual use, n is often very large, and the extraction results m of each sentence are not many. This method will bring extremely serious category imbalance problems. Therefore, multi-label cross entropy loss is used here. Instead of turning multi-label classification into multiple binary classification problems, it turns it into a pairwise comparison of the target category score and the non-target category score, and with the help of the good properties of the logarithmic operation, the weight of each item is automatically balanced. The general form of the loss is as follows:

[0029]

[0030] In the formula are the sets of positive and negative categories respectively.

[0031] After completing the entity relationship extraction, it is necessary to construct a multi-grid relationship dependency graph (MRDG) based on the extraction results, such as Figure 2 As shown. The MRDG graph uses two types of undirected edges, relation dependency edges and character segmentation edges, to connect entity nodes and character nodes respectively. In order to avoid error accumulation in the word segmentation stage, in the MRDG graph, this paper does not keep the word segmentation results as word nodes, but uses the entities in the entity relationship as word nodes, and uses the relationship between entities as undirected dependency edges between word nodes. At the same time, each character in the sentence is added to the MRDG graph as a character node. The word node has a segmentation link connected to the character nodes that make up the word. In this way, when this paper predicts the character node, the embedding information of the word node will be fused with the entity relationship obtained by this paper in the previous step for calculation. Ultimately, the MRDG graph combines context encoding features and entity relationship features through its entity relationship links and word character connections to help the model make better decisions.

[0032] Step 2: Text Embedding and Context Encoding

[0033] After completing the MRDG graph construction in step 1, the text needs to be input into the BERT model for word embedding and context encoding. The main structure of BERT is composed of multiple Transformer encoders stacked layer by layer to form a deep network. Its structure is as follows Figure 3 This deep structure enables the model to learn more abstract and complex semantic representations, thereby achieving better performance in various natural language processing tasks.

[0034] Compared with the original Transformer structure, BERT has made some modifications to adapt to pre-training and downstream task fine-tuning. First, the original Transformer consists of two parts, an encoder and a decoder, and is mainly used for sequence-to-sequence tasks. The encoder processes the input sequence, while the decoder is used to generate the output sequence. However, BERT does not use the decoding layer in the Transformer because it is a pre-trained language model and does not need to perform the task of generating the next word. Second, BERT does not use the position encoding in the original Transformer, but instead conveys position information by introducing position embedding in the input embedding stage. These adjustments make BERT more suitable for its pre-training and fine-tuning goals. First, the multi-head attention sublayer is the most important part of feature extraction in the Transformer, and the multi-head attention is composed of a self-attention mechanism. In the self-attention mechanism, three vectors Q, K, and V are usually used to represent words, where Q is the query vector, K represents the attributes of the word itself, and V represents the information contained in the word. Q, K, and V are calculated by multiplying the input and transformation matrices, as shown below:

[0035] Q=W Q X

[0036] K=W K X

[0037] V=W V X

[0038] Where W Q , W K , W V is the weight matrix with different weights; X is the embedding vector of the input.

[0039] Then, the scaled dot product self-attention mechanism is used to calculate the attention weights of Q and K. Since the V vector represents the information contained in the word, the obtained attention weight is multiplied by V to obtain the final representation. The calculation formula is as follows:

[0040]

[0041] Where Attention(·) represents the dot product attention operation; d k is the dimension of vectors Q and K. In order to avoid the result of multiplying vectors Q and K being too large, which will cause the softmax gradient to disappear, the product is reduced by dividing the dimensions of the two.

[0042] The scaled dot-product self-attention mechanism mentioned above is only one component of the multi-head attention mechanism. To better extract features of varying importance and enhance the model's learning capabilities, the model is further divided into multiple heads, forming multiple independent subspaces. This allows the model to focus on information of different dimensions. The calculation process of the multi-head attention mechanism involves calculating the attention score for each sub-head. The specific calculation method can be referred to in the formula below. This design enables the model to more comprehensively capture information of different dimensions, improving its ability to learn and express multi-level and multi-scale features.

[0043]

[0044] In the formula, head i Represents the calculation result of the i-th attention mechanism; Head i Query, Key and Value transformation matrix in W O is the output transformation matrix.

[0045] Then the small matrices of each head are concatenated into a large matrix to obtain the final representation. The formula is as follows:

[0046] MultiHead(Q,K,V)=concat(head1,head2,...,head n )W O

[0047] After the multi-head self-attention mechanism sublayer, it is followed by the feedforward neural network sublayer (FFN). This sublayer is used to perform nonlinear transformations on the encoded output of the attention sublayer, helping the model to more fully learn the complex relationships between inputs. This structure is applied independently at each position, allowing the Transformer to more effectively capture the features of different positions in the sequence. Typically, FFN consists of two linear layers (fully connected layers) and an activation function. If the input is represented as a dimension of d model The tensor, denoted as X, represents the hidden layer dimension of the model. FFN first passes through the first fully connected layer, then applies the activation function to the result, and finally passes it through another fully connected layer. This design further improves the model's ability to capture features at different positions in the sequence. The calculation process is as follows:

[0048] FFN(X)=ReLU(XW1+b1)W2+b2

[0049] Where W1 is the weight matrix of the first linear layer; b1 is the bias of the first linear layer; W2 is the weight matrix of the second linear layer; b2 is the bias of the second linear layer.

[0050] Furthermore, each sublayer is followed by layer normalization and residual chaining. The tensor X computed at each layer is calculated using the following formula. This structure allows the model to better maintain gradient flow between sublayers, facilitating more stable and efficient training and learning of complex features.

[0051] Y = LayerNorm(X + Sub(X))

[0052] Where LayerNorm(·) is the layer normalization operation, and the calculation formula of LayerNorm(·) is as follows:

[0053]

[0054] where γ and β represent learnable scaling and translation factors; μ and σ are the mean and standard deviation of x in the last dimension, respectively.

[0055] For the input vector embedding matrix A, the BERT model is used for context encoding, which is expressed as follows:

[0056] [h1, h2, ..., h n ]=BERT([α1,α2,…,α n ])

[0057] Where h i The hidden vector representing the context-encoded word unit at position i. Because the BERT encoder has massive text features and a unique bidirectional structure, it enables word vectors to learn contextual meaning, enhancing the features of word vectors. The model's output vectors can then be used for downstream tasks.

[0058] Step 3: Entity Relationship Diagram Embedding

[0059] Graph neural networks (GNNs) play an important role in modeling graph-structured data. Their core mechanism is to effectively extract the multi-level dependencies within the graph structure by integrating information surrounding each node. Among the many GNN variants, this paper uses the Graph Attention Network (GAT) as a graph encoder. GAT is unique in that it dynamically weights nodes while aggregating information about neighboring nodes. This feature gives the model greater adaptability and flexibility, enabling it to be adapted and optimized according to the specific needs of different tasks.

[0060] Specifically, Represents a graph, where is the vertex set, ε is the edge set. In the MRDG graph of this paper, the vertex set express By all the byte nodes and character nodes , ε contains all the relationship dependency edges and character segment edges. represents the node embedding of the i-th node in the GAT layer of layer l, where d is the dimension of the node embedding. This paper uses the hidden state output of the last layer of the context encoding layer to initialize the node embedding, and the formula can be obtained as follows:

[0061]

[0062] Where, Represents all neighbors of node i that have character segment edges with node i. The initialization vector of the entity node is calculated by average pooling.

[0063] Then, the vector of the middle layer is input into the multi-head attention for calculation. Let the number of attention heads be H, and for each attention head h, calculate its attention score The formula is as follows:

[0064]

[0065] Where, and Represents the features of node i and node j at layer l-1; It is a linear projection matrix, which transforms the input into d h dimension; || is the concatenation operator; is the learnable attention weight vector.

[0066] Then, this paper uses the softmax function to calculate the normalized attention weight, and the calculation formula is as follows:

[0067]

[0068] Where, Represents the set of neighbors of node i. This paper uses normalized attention weights to linearly combine the features of node neighbors. The output feature formula of node i on the attention head h is as follows:

[0069]

[0070] Where, It is d h dimensional output features.

[0071] Connecting all the outputs of H attention heads, we get the complete node feature representation formula as follows:

[0072]

[0073] Where, It is d h×H-dimensional vector. For the sake of convenience, this paper chooses the same dimension for the input and output representation of the GAT layer, that is, d = d h ×H.

[0074] By stacking L GAT layers, each node can collect information from its L-hop neighbors. Finally, the node features of the last GAT layer are retained. The triplet is extracted as the input of the decoding layer.

[0075] Step 4: Bi-affine decoding and output

[0076] Based on the previous discussion, this paper models predicate extraction as a span classification problem. Existing methods simply concatenate the head and tail features of an entity, without any information exchange between each position. We need a method that allows the hidden vectors at the head and tail of an entity to interact with each other. This goal can be achieved using a biaffine architecture. The key idea is to view the entity extraction task as a problem of identifying the head and tail indices, while also assigning a type to the span formed by the head and tail.

[0077] The specific encoding structure uses a dual affine network to capture the syntactic information between each span in the text. In the predicate extraction stage, for the candidate span <c i ,…,c j > 1≤i≤j≤N , assuming the input is the encoded text sequence matrix The head and tail hidden layer vectors are fused through a dual affine network to obtain the decoding input formula for each span as follows:

[0078]

[0079] Where s i , e i and are the indicators of the start and end of the i-th span; and is the hidden vector output by the last layer of GAT network; U m is a d×c×d tensor; W m It is a c×2d matrix. In the model of this chapter, c=2, indicating a classification model. represents vector concatenation operation; b m is the bias.

[0080] The vector r needs to satisfy s i ≤e i The constraint indicates that the starting point of the entity is before its end point. Under this constraint, a score is provided for all possible spaces that may constitute the extraction result. This paper assigns a span result to each span space and predicts the probability of it becoming an extraction result as follows:

[0081] P = softmax(r i )

[0082] In the parameter extraction stage, given each predicate span obtained in the predicate extraction stage, we extract its corresponding subject and object. We use relative position embedding as an additional input feature f i =p i +u i To indicate the predicate position. For subject and object, different classification matrices are constructed, and finally the predicted parameter probability is obtained. On the subject, the prediction result of the span starting point is as follows:

[0083] P sub_start =softmax(W sub_start U)

[0084] The prediction formula for the subject span endpoint is as follows:

[0085] P sub_end =softmax(W sub_end U)

[0086] Where, and is the weight matrix; P sub_start and P sub_end are the start and end probability distributions over sentence characters.

[0087] A similar method is used to obtain the results for the object. The prediction result of the span starting point is as follows:

[0088] P obj_start =softmax(W obj_start U)

[0089] The prediction formula for the object span endpoint is as follows:

[0090] P obj_end =softmax(W obj_end U) where is the weight matrix; P obj_start and P obj_end are the start and end probability distributions over sentence characters.

[0091] During training, the predicate extraction model and the parameter extraction model are optimized independently. The predicate extraction model is trained on the target predicate of the sentence. During inference, the predicate obtained by the predicate extraction model is fed into the parameter extraction model. Ultimately, we need to combine the results of these two stages to obtain the complete fact triple.

Claims

1. An open domain information extraction algorithm based on graph neural network and relation discrimination Step 1: Extracting entity relationship triples and constructing a relationship graph. This paper innovatively proposes an entity relationship extraction method based on a global pointer network to efficiently extract entity relationship triples from input text. It also designs a multi-granularity relationship dependency graph that combines contextual encoding features with entity relationship features through entity relationship links and word character connections. Step 2: Text Embedding and Context Encoding. After completing the entity-relationship graph in Step 1, the text needs to be input into the algorithm for word embedding and contextual representation. This step primarily addresses the discrepancy between the input of the deep learning network being a numeric vector and the original dataset being text. Therefore, vector embedding of the text representation is necessary. This process primarily involves calculating the embedding layer and generating positional embeddings. After obtaining the embedding vector, it is also necessary to use the BERT model for context encoding so that the word vector can learn the contextual meaning and enhance the characteristics of the word vector, so that the output vector of the model can be used for downstream tasks. Step 3: Entity Relationship Graph Embedding. After completing the text encoding in Step 2, the entity relationship graph is embedded using the graph neural network in Step 1. This step uses a graph attention network as the graph encoder. The graph attention network's ability to dynamically weight nodes while aggregating information from adjacent nodes, combined with the context-sensitive encoding representation in the previous step, gives the model greater adaptability and flexibility. Step 4: Bi-affine decoding and output. In this step, a bi-affine network is used to complete decoding and output, obtaining the SPO triple extraction results. This step requires combining the graph embedding results from step 3 to complete the SPO triple extraction and decoding process. In this step, the entity extraction task is considered as the problem of identifying the head and tail indexes, assigning a type to the interval formed by the head and tail, and using a bi-affine network to solve the problem of hidden layer vector information interaction.

2. The method according to claim 1, wherein: Step 1 proposes an entity relationship extraction model and a method for constructing a multi-granularity relational dependency graph. During the entity relationship extraction phase, a global pointer-based entity relationship extraction model is designed to extract triples from the input text. Subsequently, a multi-granularity relational dependency graph is constructed using entities as word nodes and relationships between entities as undirected dependency edges between word nodes. Contextual encoding features and entity relationship features are combined through entity relationship links and word character connections.

3. The method according to claim 1, wherein: Step 2 proposes a text embedding and context encoding method. In this step, the input text is embedded and encoded using the BERT model. Through the multi-layer Transformer encoder, the text vector can learn more abstract and complex semantic representations through multi-layer interactions.

4. The method according to claim 1, wherein: Step 3 proposes a method for embedding entity relationship graphs. In this step, it is necessary to combine the multi-granularity relationship dependency graph constructed in step 1 and the context encoding results in step 2. By using a graph attention network to aggregate information from adjacent nodes and assign dynamic weights to nodes, the topological information of the relationship graph is integrated into the text representation.

5. The method according to claim 1, wherein: Step 4 proposes a bi-affine decoding and output method. In this step, the present invention formulates the open information extraction task as a span classification problem. A bi-affine structure is used to capture the syntactic information of each interval in the text, enabling head-tail information interaction. The spans formed by the head and tail intervals are classified.