A Fact Verification Method Based on Graph Neural Network and Reinforcement Learning
The RERG framework enhances fact verification in table-type data by using graph neural networks and reinforcement learning to identify and aggregate key evidence, addressing inefficiencies in existing methods and improving accuracy and interpretability.
Patent Information
- Application Number
- CN202211085132.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-06
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-09-06
AI Technical Summary
In the factual verification of tabular data, the graph neural network based on analytical statements considers fewer features and can only passively aggregate neighbor node information. The pre-trained model based on table perception is difficult to meet the characteristics of multi-hop inference, resulting in poor detection results.
RERG, a reinforcement learning framework based on graph neural networks, uses multi-grained feature representation and reinforcement learning-driven node selection, combined with self-attention mechanism and secondary update strategy, to enhance the information aggregation ability of graph neural networks, and simulate the selection of key evidence in human reasoning.
It improves the effect and interpretability of fact verification, adapts to a variety of semi-structured tables, and improves the accuracy of detection and F1 value.
Smart Images

Figure CN115511082B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a fact verification method based on graph neural network and reinforcement learning, which can be used to verify whether a statement conforms to the facts contained in tabular data, and belongs to the fields of Internet and artificial intelligence technologies. Background Art
[0002] With the rapid development of Internet technology, people can more conveniently access information from various online media platforms through intelligent devices such as computers and mobile phones. However, the lack of quality management mechanism for the release of network information makes the network information chaotic and complex, with potential safety hazards, which will prevent users from obtaining correct network resources. Therefore, people need some technology to detect the information content, so as to filter out false information in advance.
[0003] The task of fact verification for tabular data aims to verify whether a statement conforms to the facts contained in the tabular data. Existing work mainly analyzes a statement sentence through logical form, and then processes the parsed statement representation by using a graph neural network. Or design a table-aware pre-trained language model to process the "statement-table" pair. However, the graph neural network constructed based on the method of parsing statements considers fewer features and can only passively aggregate the information of neighbor nodes in each graph reasoning step. And the table-aware pre-trained model is difficult to meet the multi-hop reasoning characteristics of the fact verification task, so the detection effect is not good. In order to improve the performance of fact detection for tabular data, the present invention deeply explores the graph structure features of the "statement-text" pair, and proposes a fact verification framework based on graph neural network and reinforcement learning (Reinforced Evidence Reasoning framework with Graph neural network, RERG), which simulates the behavior of focusing on certain words at each step in the human reasoning process. Specifically, based on the pre-trained model RoBERTa, RERG first uses a Transformer-based graph neural network to represent multi-granularity features; then, designs a monitoring node, and connects the monitoring node with some potential key nodes through the reward feedback of reinforcement learning on each graph neural network layer. In this way, RERG can use the monitoring node to predict the label of the statement text and aggregate key information through multiple layers of graph neural network. In addition, the secondary update method added after the attention mechanism can enhance the information aggregation ability of each layer of the graph neural network. Summary of the Invention
[0004] To solve the problems and deficiencies in the prior art, the present invention proposes a fact verification method based on graph neural networks and reinforcement learning. This method utilizes a fact verification framework that can select potential key evidence during the reasoning process, and obtains better features of the "statement-table" pair through multi-granularity features, thereby improving the effect of fact verification and the interpretability of the selected words. The present invention improves the graph neural network and performs secondary information aggregation after calculating the attention weights of each graph neural network layer.
[0005] To achieve the above object, the technical solution of the present invention is as follows: A fact verification method based on graph neural networks and reinforcement learning, comprising the following steps:
[0006] Step 1: Obtain the table text representation. Using the "statement-table" pair as input, learn its vector representation by RoBERTa enhanced by the MultiNLI corpus.
[0007] Step 2: Graph construction and node initialization. The purpose of this step is to provide a structure graph for learning for the graph neural network layer according to the co-occurrence relationship of rows, columns, numbers, and words between the extracted statement sentences and the target table, as well as the syntactic dependency structure tree of the statement. Subsequently, use bidirectional LSTM (BiLSTM) and multi-layer perceptron (MLP) to complete the initialization of the nodes.
[0008] Step 3: Apply a Transformer-based graph neural network to encode the graph. First, encode different types of edges into independent vectors through the graph neural network, and then improve the original graph neural network by introducing node information and edge information into the self-attention mechanism. Use the improved self-attention mechanism to aggregate the node information to complete the update of the node information and the encoding of the graph data.
[0009] Step 4: Reinforcement learning-driven node selection. This step introduces a monitoring node into the graph. According to the reward feedback of reinforcement learning, select appropriate valid evidence nodes on each graph neural network layer. Through the information aggregation of multiple graph neural network layers, the monitoring node can capture various potential key evidences for final verification.
[0010] Step 5: Secondary update and training and testing of the model. To eliminate the influence of the uneven distribution of the number of graph node neighbors on node information aggregation, this step introduces a fusion layer to enhance information aggregation. First, use the sigmoid function to adaptively assign a weight to the nodes, then calculate the neighbor information to promote information aggregation, and finally fuse the neighbor information obtained in the previous step with the node information to complete the secondary update of the node information. The model uses relevant tables to predict the label distribution on conditional statements and minimizes the cross-entropy loss for training. Use accuracy and F1 value metrics to evaluate the model and test the performance of the model.
[0011] A matter-of-fact verification method based on graph neural network and reinforcement learning, the method includes a table text input learning layer, a graph construction layer, a graph neural network layer based on Transformer, a key node selection layer driven by reinforcement learning, and a secondary update layer. Compared with the prior art, the advantages of the present invention are as follows:
[0012] (1) The reinforcement evidence reasoning framework RERG based on graph neural network adopted by the present invention only depends on the language and structural features of statements and target semi-structured tables.
[0013] (2) The present invention adopts a key node selection strategy driven by reinforcement learning to guide the selection of key evidence for each graph neural network layer, and adopts a secondary update strategy to enhance the information aggregation ability of neighbor nodes.
[0014] (3) The present invention is applicable to various semi-structured tables, and the proposed graph feature extraction method can extract tables including database type tables, key-value pair type tables, and tables in PDF format included in academic papers.
[0015] (4) The present invention has updateability, and the involved pre-trained model and reinforcement learning algorithm can be replaced with better models, such as the DeBERTa pre-trained model and the Actor-Critic algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a flowchart of the method according to an embodiment of the present invention.
[0017] Figure 2 It is an overall model diagram according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] In order to deepen the understanding and recognition of the present invention, the present invention will be further clarified below in conjunction with specific embodiments.
[0019] Embodiment 1: A matter verification method based on graph neural network and reinforcement learning. This method first obtains the text feature representation of the table; then, according to the co-occurrence relationship of rows, columns, numbers, and words between the extracted statement sentence and the target table, as well as the syntactic dependency structure tree of the statement, a structure graph for learning is constructed for the graph neural network. Subsequently, BiLSTM and MLP are used to complete the initialization of the graph nodes; then, different types of edges are encoded as independent vectors by the graph neural network, and the original graph neural network is improved by introducing node information and edge information into the self-attention mechanism. The improved self-attention mechanism is used to aggregate the node information to complete the update of the node information and the encoding of the graph data; next, a monitoring node is introduced into the graph. According to the reward feedback of reinforcement learning, appropriate valid evidence nodes are selected on each graph neural network layer. Through the aggregation of information in multiple graph neural network layers, the monitoring node can capture various potential key evidences for final verification; finally, the sigmoid function is used to adaptively assign a weight to the nodes, then the neighbor information is calculated to promote information aggregation, and finally the neighbor information obtained in the previous step is fused with the node information to complete the secondary update of the node information. The model uses the relevant table to predict the label distribution on the conditional statement and minimizes the cross-entropy loss for training, and uses the accuracy and F1 value metrics to evaluate the model and test the performance of the model. For the specific model, see Figure 2 , and the detailed implementation steps are as follows:
[0020] Step 1, Obtain the table text representation. To obtain the vector representation of the "statement-table" pair, this step first concatenates the statement sentence and the table content and adds the [CLS] identifier at the head; then the long sequence is input into the RoBERTa pre-trained language model enhanced by MultiNLI to obtain a meaningful feature representation. The calculation formula for obtaining the representation of each token is shown in Formula (1).
[0021]
[0022] Among them, is the representation matrix of sample X i d model represents the dimension of each token, and n represents the length of the nth "statement-table" sequence.
[0023] Step 2, Graph construction and node initialization. First, extract the co-occurrence relationship of rows, columns, numbers, and words between the statement sentence and the target table, as well as the syntactic dependency structure tree of the statement. Specifically, given a statement sentence and its related table, the construction rules of the structure graph are as follows:
[0024] (1) Row - column relationship: Cells in the same row usually describe different attributes of the same record, while cells in the same column reflect the same attributes of different records. For a key - value table, each row has only two components, namely the key and the value. The present invention defines a value relationship to represent words in the same value component.
[0025] (2) Numerical relationship: Since there are often greater - than, less - than or equal numerical relationships in cells containing numbers and dates, extracting these relationships is beneficial for numerical reasoning of these cells.
[0026] (3) Word co - occurrence relationship: Intuitively, the same words and their related statements co - existing in the table are crucial for verifying the correctness of the table. For example, through Figure 2 the word "date" in the statement in [ ], the relevant cells in the table can be conveniently found to further verify the date in the statement.
[0027] (4) Statement relationship: The syntactic dependency tree represents the linguistic relationship between different tokens and can reveal the dependency relationship between the words in the statement.
[0028] (5) Monitoring node: Since there are some irrelevant nodes in the graph of the statement, we add a monitoring node to select some possible key nodes for final verification, avoiding pooling all nodes to obtain the graph representation.
[0029] Subsequently, the extracted graph not only retains the internal structure of the statement and the table, but also explores the numerical relationship characteristics between the nodes. When exporting the token representation E i , the next step is to obtain the node - level representation of the graph because a node in a graph may contain multiple tokens. To this end, we apply a BiLSTM on top of the token representation and then use an MLP to obtain the integrated representation of each node.
[0030] Step 3: Apply a Transformer - based graph neural network to encode the graph. To better utilize the heterogeneous graph where some nodes have multiple relationships, a Transformer - based graph neural network is used to encode the graph data, which can convert the relationship label (i.e., the type id of the edge) into a distinguishable vector. All nodes form an adjacency matrix Αnn, where n represents the number of nodes, and each element a ij ∈Ann represents a triple <node i , node j , edge ij >, i ≤ n, j ≤ n. By improving the self - attention mechanism, the edge vector information is added to the process of calculating the dot product of the query and key vectors to achieve adjacent node information aggregation. The relationship between node i and node j can be modeled as:
[0031]
[0032] Among them, is the parameter matrix, d' = d model / n head is the size of each attention head. The edges between the i-th node and the j-th node are respectively represented as vectors and They are respectively related to the key vector and the value vector of the self-attention mechanism.
[0033] Then, for the i-th node, the information from neighbor nodes can be aggregated through a weighted sum of the transformed node representations:
[0034]
[0035] where node' i is the updated node i , is a parameter matrix, and α ij is obtained by applying the softmax function on node j . When multiple head self-attention layers are connected, the vector information obtained from independent parameter spaces can be fused to obtain deeper representations. The number of multi-head attentions is h = d model / d'.
[0036] Step 4, Reinforcement Learning-driven Node Selection. Since not all nodes are beneficial for verifying the statement, potential key nodes must be determined in the graph. Similar to the human reasoning process, the model needs to focus on partial key evidence at each step. For this purpose, we propose a monitoring node in the graph and formulate a reinforcement learning process to select valid evidence nodes for the monitoring node. The Agent of this reinforcement learning will interact with the environment at several discrete time steps. The reinforcement learning-based node selection takes the representation of each node and the adjacency matrix as the input for each state. Then, the Agent applies a policy network to take actions, that is, to determine whether each node is selected. Rewards will be obtained from the actions, and then the policy network is optimized. In a specific example, each graph reasoning layer is a step, and the policy network performs node selection for the monitoring node at each step. The following explains the environmental settings of the reinforcement learning:
[0037] State: At each decision step t, the state is the concatenation of these two parts: the representation of each node in the t-th graph reasoning layer the current adjacency matrix state
[0038] Action: For each node node iActions are binary. Indicates that the monitoring node is connected to node i with an edge weight of 1, Indicates that the weight of the connected edge is 0. The behavior result at each time step t is used to update the neighbors of the monitoring node in the adjacency matrix.
[0039] Reward: The reward is generated when all inference layers are traversed. Since the last inference layer is used for verification in our framework, we use the negative cross-entropy loss as the reward R for the current batch of training data, as well as the penalty term R p , to prevent too many nodes from being selected. R p is calculated as shown in Equation (4):
[0040]
[0041] where λ is a hyperparameter used to control the ratio of the nodes connected by the monitoring node.
[0042] Policy network: We establish a policy network with two MLPs to support the behavior of the monitoring node, and then use a softmax function to expand the behavior probability to the interval (0,1). The structure of the policy network is shown in Equations (5) and (6):
[0043]
[0044]
[0045] where W and b are the parameters in the policy network. Set
[0046] The policy network is optimized using the reinforcement learning algorithm, and the parameters are updated through the policy gradient algorithm. The specific optimization strategy is shown in Equation (7):
[0047]
[0048] where R(s,a) is the state-action function, θ is the parameter of the policy network, B is the number of samples in a batch, T is the number of steps in an episode (i.e., the number of inference layers), R(s it ,a it ) = R·R p . The parameter θ is updated after each round of iteration.
[0049] Generally speaking, the monitoring node aggregates potential key node information on each layer. It can be regarded as the representation of the overall graph because it can aggregate various information from different nodes through multi-layer calculations. In this way, the monitor node can be used for final verification instead of pooling all nodes.
[0050] Step 5, perform secondary update and model training and testing. To eliminate the impact of uneven distribution of the number of graph node neighbors on node information aggregation, a fusion layer is introduced in this step to enhance information aggregation. Specifically, we apply a sigmoid function to adaptively assign a weight to node j:
[0051] β j = sigmoid(M a node j + b) (8)
[0052] where M a is the trainable parameter matrix and b is the bias. Then we calculate the neighbor information to promote information aggregation:
[0053]
[0054] where node″ i is the aggregated information of node i, C i is the neighbor of node i, and M e is a transformation matrix. Finally, the neighbor information is fused with the node information as follows:
[0055] node i = ReLU(M f node i + node″ i + b f ) (10)
[0056] where M f is the weight matrix and b f is the bias vector.
[0057] Finally, we use the relevant table to predict the label distribution on the conditional statement and minimize the cross-entropy loss L to train our model. Specifically as follows:
[0058] P(Y|S i ,T i ) = softmax(M c Node monitor ) (11)
[0059] L = CrossEntropy(Q(Y), P(Y|S i ,T i)) (12)
[0060] In the above two equations, represents the weight matrix for classification, and n label represents the number of label categories in the relevant dataset, while Q(Y) represents the distribution of the true labels. We use the accuracy and F1-score metrics to test and evaluate the model and examine its performance.
[0061] Based on the same inventive concept, a method for fact verification based on a graph neural network and reinforcement learning according to the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements the above-described method for fact verification of tabular data based on a graph neural network and reinforcement learning for tabular fact verification.
[0062] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principles of the present invention. It should be understood that the embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification of the present invention fall within the scope defined by the claims of this application.
Claims
1. A fact verification method based on graph neural network and reinforcement learning, characterized in that, The method includes the following steps: Step 1: Obtain the tabular text representation; Step 2: Graph construction and node initialization; Step 3: Apply a Transformer-based graph neural network to encode the graph; Step 4: Node selection driven by reinforcement learning; Step 5: Perform secondary update and training and testing of the model; Among them, Step 3: Apply a Transformer-based graph neural network to encode the graph, specifically as follows: (1) Use a Transformer-based graph neural network to encode the graph data, which can convert the relationship label, i.e., the edge type id, into a distinguishable vector; (2) Add edge vector information to the process of calculating the dot product of query and key vectors by improving the self-attention mechanism to achieve adjacent node information aggregation; (3) Connect multiple head self-attention layers to fuse the vector information obtained from independent parameter spaces to obtain a deeper representation; Step 4: Node selection driven by reinforcement learning, specifically as follows: (1) Design a monitoring node in the graph and formulate the reinforcement learning process as selecting effective evidence nodes for the monitoring node. The Agent of this reinforcement learning will interact with the environment at several discrete time steps. The node selection based on reinforcement learning takes the representation of each node and the adjacency matrix as the input of each state; (2) The Agent applies a policy network to sample operations, that is, to determine whether each node is selected. The reward will be obtained from the actions, and then the policy network is optimized. In a specific example, each graph inference layer is a step, and the policy network performs node selection on the monitoring node at each step; (3) Use a reinforcement learning algorithm to optimize the policy network and update the parameters through the policy gradient algorithm.
2. The fact verification method based on graph neural network and reinforcement learning according to claim 1, characterized in that Step 1: Obtain the tabular text representation, specifically as follows: (1) In this step, first concatenate the statement sentence and the table content and add the [CLS] identifier at the head; (2) Then input the long sequence into the pre-trained language model RoBERTa enhanced by the MultiNLI corpus to obtain a meaningful feature representation.
3. The fact verification method based on graph neural network and reinforcement learning according to claim 1, characterized in that Step 2: Graph construction and node initialization, specifically as follows: (1) Extract the co-occurrence relationships of rows, columns, numbers, and words between the statement sentence and the target table, as well as the syntactic dependency structure tree of the statement; (2) When the labeled representation E in step 1 is obtained i At this time, the next step is to obtain the node-level representation of the graph. Since a node of a graph contains multiple tokens, for this purpose, a BiLSTM model is applied on top of the token representation, and then an MLP layer is used to obtain the final representation of each node.
4. The fact verification method based on graph neural network and reinforcement learning according to claim 1, characterized in that, Step 5: Perform secondary update and training and testing of the model, specifically as follows: (1) To eliminate the influence of the uneven distribution of the number of neighbors of graph nodes on node information aggregation, this step introduces a fusion layer to enhance information aggregation; (2) Apply a sigmoid function to adaptively assign a weight to node j; then, calculate the neighbor information to promote information aggregation; finally, fuse the neighbor information with the node information; (3) Minimize the cross-entropy loss L to train the model; (4) Use accuracy and F1 value metrics to test and evaluate the model and examine the performance of the model.