A table question answering method based on TaPas model and graph attention network
By combining the BERT model and graph attention network in table question-and-answer processing, directly predicting the answers to questions rather than generating SQL statements, the complex natural language problem processing and SQL statement generation problems in the prior art are solved, and higher accuracy and lower labeling costs are achieved.
Patent Information
- Application Number
- CN202211563273.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-07
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-12-07
AI Technical Summary
Existing table question and answer technologies are difficult to effectively deal with complex natural language problems, and the process of generating SQL statements is complex and error-prone. Deep learning models have challenges in generating statements that strictly conform to SQL syntax.
The table question-and-answer processing method based on TaPas model and graph attention network is adopted, and the table is filtered through the BERT model, unnecessary columns are removed, and the graph attention network is used to fuse the feature vectors extracted by the TAPAS pre-trained model to directly predict the answers to the question without generating intermediate SQL statements.
It improves the accuracy of table question and answer processing, alleviates the difficulties of deep learning models in generating complex SQL statements, and reduces the difficulty and cost of data annotation.
Smart Images

Figure CN115794871B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of natural language processing and table question answering technology, and in particular, relates to a table question answering processing method based on a TaPas model and a graph attention network. Background Art
[0002] Data is usually stored in relational databases in the form of tables. If users want to obtain specific data, they need to master a database query language, such as SQL. In order for users to quickly find the data they want, the table question answering task is proposed. The table question answering task requires the model to get the answer the user wants based on natural language questions and tables.
[0003] Existing table question answering tasks usually adopt the method of converting natural language questions into SQL statements, which can be roughly divided into the following two types:
[0004] 1. One is the sequence-to-sequence conversion model represented by seq2sql. Early work used this method, which completely relied on the model to predict each character of the SQL statement, and then input the generated SQL statement into the database execution engine for execution to get the answer;
[0005] 2. Another method is to set up a SQL statement template and divide the SQL statement into different parts. The model is responsible for predicting what needs to be filled in each place, and then input the generated SQL statement into the database execution engine for execution to get the answer;
[0006] However, the above implementation method has the following disadvantages:
[0007] 1. The SQL statements generated by the deep learning model are only an intermediate step in retrieving answers, and it is very difficult for the deep learning model to generate such SQL statements. It requires that the statements generated by the model must strictly comply with the SQL syntax and cannot have any errors, otherwise it will be an SQL statement that cannot be run. As the user's natural language questions become more and more complex and difficult, the corresponding SQL statements will become more and more complex and flexible, and the method of using preset templates will become ineffective, which poses a great challenge to the deep learning model;
[0008] 2. Collecting such a data set is time-consuming and costly, and the workers who annotate the data set need to master SQL syntax. For more complex problems, the workers spend more time writing correct SQL statements, and the error rate is higher.
[0009] In order to solve the above problems, some scholars have proposed to use the model to directly predict the answer to the question based on the question and the table, without generating the intermediate SQL statement. TAPAS is a famous model of this method, but it also has several problems:
[0010] 1. After TAPAS was pre-trained on a large number of table-text pairs, only two fully connected layers were used during fine-tuning, which cannot fully utilize the capabilities of the TAPAS pre-trained model.
[0011] 2. TAPAS also has an obvious limitation, which is the text length limit for the input question form, which is generally limited to 512 symbols. Summary of the invention
[0012] In view of the above problems, the present invention proposes a table question answering processing method based on the TaPas model and the graph attention network. According to the characteristics of the table question answering task, the graph attention neural network is mainly used to utilize and fuse the feature vectors extracted by the TAPAS pre-training model, and before the table is input into the TAPAS model, the Bert model is used to filter the table based on the question to remove unnecessary columns, so that TAPAS can process larger tables.
[0013] The technical solution of the present invention is:
[0014] The table question answering processing method based on the TaPas model and the graph attention network includes the following steps:
[0015] S1. Filter the tables used for question and answer based on the questions to be processed, specifically:
[0016] The BERT model is used to judge the problem using N columns of the table and the probability of each column being related to the problem. The first N columns are retained in descending order of probability, and the remaining columns are deleted to obtain the filtered table.
[0017] S2, input the question and the filtered table into the pre-trained TaPas model, and output the representation vector of each symbol of the natural language question and the table;
[0018] S3. Build a graph attention network, specifically: represent each cell in the table with a node, each table header with a node, and the question with a total node, where each node is connected to itself, the node representing the question is connected to the node of each cell, the node of each column header is connected to the node of each cell in the same column, and the nodes of each cell in each row are connected to each other;
[0019] S4. Get the answer through the graph attention network.
[0020] Furthermore, the specific implementation method of step S1 is as follows: for each column of the table, we concatenate the column name, data type, and table name of each column into a string, and then combine it with the natural language question into a column-question pair. A column-question pair has the same input format as the sentence pair classification task of BERT, so that BERT's capabilities can be fully utilized. We need to use the representation vector of the special symbol [CLS] because the representation vector of [CLS] combines the comprehensive information of the natural language question and the column. Our snapshot model uses different fully connected layers to calculate the number of columns N related to the natural language question and the probability of each column being related to the natural language question based on the representation vector of [CLS]. The top N columns are retained according to the probability ranking, and the other columns are deleted.
[0021] The specific implementation method of step S4 is as follows: after fusing the feature vector extracted by the TAPAS pre-trained model through the graph attention network, a fully connected layer is used to calculate the probability that each table cell is part of the answer for the feature vector of the cell. When the probability exceeds the set threshold, the table cell is considered to be part of the answer; a fully connected layer is used to calculate which aggregation operator should be used for this problem for the feature vector of [CLS]. According to the predicted aggregation operator and table cell, the answer can be obtained.
[0022] The beneficial effect of the present invention is that compared with the traditional method, the present invention effectively improves the accuracy of form question and answer processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 are the cells and aggregation operators for which model predictions are chosen.
[0024] Figure 2 It is a model framework diagram.
[0025] Figure 3 This is a diagram of the snapshot module filter table.
[0026] Figure 4 This is how the graph is constructed. DETAILED DESCRIPTION
[0027] The present invention is described in detail below with reference to the accompanying drawings.
[0028] like Figure 1 As shown, the method of the present invention is full supervision in weak supervision, that is, the data set needs to provide aggregation operators and selected table cells as training labels. Compared with the complete SQL query statement, this supervision method can alleviate the difficulty of labeling the data set.
[0029] The model designed by the present invention mainly consists of three modules, such as Figure 2As shown in the figure, the first module is a column-based table snapshot model. This model uses the Bert model to determine how many columns of the table will be used in a natural language question, as well as the probability of each column being related to the question. Finally, the table will retain the columns with the highest probability, and the other columns will be deleted. The second module is the TAPAS encoder. Based on Bert, TAPAS adds encoding for tables and performs mask pre-training tasks on millions of table-text pairs. In this article, the TAPAS encoder takes natural language questions and tables filtered by the snapshot model as input, outputs representation vectors of natural language questions and each symbol of the table, and then passes these representation vectors to the third module. The third module is the graph attention network. We connect natural language questions, table cells, and table headers according to specific relationships to construct a graph. The initial representation vectors of the nodes of this graph come from the representation vectors learned by the TAPAS encoder of the second module. The graph attention network is used for learning on this graph, and then a fully connected layer is used to calculate the probability that each symbol is part of the answer. When the probability exceeds the set threshold, the table cell is considered to be part of the answer. Another fully connected layer is used to calculate which aggregation operator should be used for this question based on the feature vector of [CLS]. The answer can be obtained based on the predicted aggregation operator and table cell.
[0030] Example:
[0031] This example specifically includes the following steps:
[0032] S1. Using the BERT model to judge the problem will use N columns of the table, as well as the probability of each column being related to the problem. The first n columns are retained in descending order of probability, and the remaining columns are deleted to obtain the filtered table.
[0033] Specifically, given a table and a natural language question, assuming that the natural language question is q and the columns of the table are c 1 , c 2 , ..., c k , then, for each column in the table, a string tuple is formed (in, refers to the data type of the i-th column, Refers to the table name of the table to which the column belongs, and the concat function is a concatenation function). Input the string tuple (column-question input pair) into the Bert model in the form of [CLS], x 1 , x 2 , ..., x m , [SEP], y 1 ,y 2 , ..., y n , [SEP] (where x iRepresents the i-th token of the column string, y i Represents the i-th token of the natural language question string), and then we can get the feature representation extracted by the Bert pre-training model:
[0034]
[0035] where h [CLS] After the Bert pre-trained model extracts the feature representation of the column-question pair for each column by integrating the comprehensive information of the question and the column, (1) a fully connected layer is used to calculate C i The probability associated with the problem: P(c i ∈R q |q)=sigmoid(w r ·h [CLS] )(where w r Refers to the fully connected layer used to calculate the probability that the column is related to the question, sigmoid is the activation function, R q refers to the set of columns related to the problem); (2) use another fully connected layer to use C i The information prediction problem uses the probability that the number of columns in the table is n: P(n|c i ,q)=softmax(W n ·h [CLS] ) n (W n Refers to C-based i The information calculation problem will use the probability of the number of columns n in the table. The softmax function is a normalized exponential function, and the n with the largest probability is selected; (3) Finally, the prediction problem involves the number of columns in the table: (4) Based on the probability calculated in the first step, take the first columns.
[0036] S2, input the question and the filtered table into the pre-trained TaPas model, and output the representation vector of each symbol of the natural language question and the table;
[0037] S3. Build the table and questions into a graph, connecting them as follows Figure 3As shown in the figure, specifically: each cell in the table is represented by a node, each table header is also represented by a node, and the problem is represented by a total node, where each node is connected to itself, the node representing the problem is connected to the node of each cell, the node of the table header of each column is connected to the node of each cell in the same column, and the nodes of each cell in each row are connected to each other; an initial vector is assigned to each node in the graph, which is the feature vector corresponding to each node in S2 after TAPAS learning, and then the graph attention network GAT is used on this graph for learning and fusion. In order to learn different features of the data, we used 4 graph attention networks for multi-head feature aggregation, and then spliced the feature vectors learned by 4 different graph attention networks.
[0038] S4. After fusing the feature vectors extracted by the TAPAS model through the graph attention network, a fully connected layer is used to calculate the probability that each table cell is part of the answer for the feature vector of the cell. When the probability exceeds the set threshold, the table cell is considered to be part of the answer. A fully connected layer is used to calculate the feature vector of [CLS] which aggregation operator has the highest probability of being used for this question, and the aggregation operator is used. According to the predicted aggregation operator, the corresponding aggregation operation is performed on the predicted table cell to obtain the answer the user wants.
Claims
1. Tabular question answering method based on TaPas model and graph attention network, It is characterized in that The following steps are involved: S1. Filter the tables used for question and answer based on the questions to be processed, specifically: The BERT model is used to determine how many columns of the table will be used for the question, as well as the probability of each column being related to the question. The first n columns are retained in descending order of probability, and the remaining columns are deleted to obtain the filtered table. S2, input the question and the filtered table into the pre-trained TaPas model, and output the representation vector of each symbol of the natural language question and the table; S3. Build a graph attention network, specifically: represent each cell in the table with a node, each table header with a node, and the question with a total node, where each node is connected to itself, the node representing the question is connected to the node of each cell, the node of each column header is connected to the node of each cell in the same column, and the nodes of each cell in each row are connected to each other; Assign an initial vector to each node in the graph. The initial vector is the feature vector corresponding to each node in S2 after TAPAS learning. Then use the graph attention network GAT to learn and fuse on this graph. S4. After fusing the feature vectors extracted by the TAPAS model through the graph attention network, a fully connected layer is used to calculate the probability that each table cell is part of the answer for the feature vector of the cell. When the probability exceeds the set threshold, the table cell is considered to be part of the answer. A fully connected layer is used to calculate the probability of which aggregation operator is the most likely to be used for this question, and the aggregation operator is used. According to the predicted aggregation operator, the corresponding aggregation operation is performed on the predicted table cell to get the answer the user wants.
2. According to claim 1, the table question answering method based on the TaPas model and the graph attention network, It is characterized in that The specific method of step S1 is: Define the natural language question as q and the columns of the table as c 1 , c 2 , ..., c k , for each column in the table, a tuple of strings is formed in Refers to the data type of the i-th column, Refers to the table name of the table to which the i-th column belongs. The concat function is a concatenation function that inputs a string tuple into the Bert model in the form of [CLS]. 1 , x 2 , ..., x m , [SEP], y 1 ,y 2 , ..., y n , [SEP], where x i Represents the i-th token of the column string, y i Represents the i-th token of the natural language question string, and then we can get the feature representation extracted by the Bert pre-training model: where h [CLS] After integrating the comprehensive information of questions and columns, the Bert pre-trained model extracts the feature representation of the column-question pair for each column: S11. Use a fully connected layer to calculate C i Probability associated with the question: P(c i ∈R q |q)=sigmoid(w r ·h [CLS] ) Among them, w r Refers to the fully connected layer used to calculate the probability that the column is related to the question, sigmoid is the activation function, R q Refers to the set of columns related to the question; S12, use another fully connected layer, using C i The information prediction problem uses the probability that the number of columns in the table is n: P(n|c i ,q)=softmax(W n ·h [CLS] ) n Among them, W n Refers to C-based i The information calculation problem uses a fully connected layer of probability of the number of columns n in the table. The softmax function is a normalized exponential function, and the n with the largest probability is selected. S13. Considering all the columns, the prediction problem will involve the number of columns in the table: S14. According to the probability calculated in the first step, take the first columns.
3. According to claim 2, the table question answering processing method based on the TaPas model and the graph attention network, It is characterized in that In S3, four graph attention networks are used for multi-head feature aggregation, and then the feature vectors learned by the four graph attention networks are concatenated.
Citation Information
Patent Citations
Data processing method and system based on TaPas model and storage medium
CN114218214A
Text table data query method based on questions and answers
CN115062070A