A Table Open-Domain Question Answering Method Based on Text-to-SQL

By using Text-to-SQL and execution guidance methods in open-domain question-and-answer in table open-domain question-and-answer, the question sentences are converted into SQL and executed on the table, and the back-execution results are used as the basis for similarity calculation, the problem of unsatisfactory table retrieval in the prior art is solved, and higher retrieval accuracy and question-and-answer accuracy are achieved.

CN115563248BActive Publication Date: 2025-05-27UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211226678.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-09
Publication Date
2025-05-27
Estimated Expiration
2042-10-09

AI Technical Summary

Technical Problem

The existing open-domain question-and-answer method of tables does not optimize the table content during the search stage, resulting in unsatisfactory search results and missing a large amount of data stored in relational databases that can only be accessed by structured languages.

Method used

Using a method based on Text-to-SQL and execution bootstrap, the question is converted into SQL logical form when extracting answers, and SQL is executed on the table, and the back-execution results are used as the basis for similarity calculation to improve the accuracy of table retrieval.

Benefits of technology

Through the use of Text-to-SQL model and the implementation of the boot traceback strategy, the accuracy of table retrieval is improved, thereby improving the accuracy of the results of the entire open domain question and answer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115563248B_ABST
    Figure CN115563248B_ABST
Patent Text Reader

Abstract

The present invention belongs to the fields of natural language processing and question-answering tasks, and relates to a table open-domain question-answering method based on Text-to-SQL. First, a retriever is used to preliminarily screen relevant tables from a table corpus to obtain a table pool, and then the top k tables are sorted according to similarity scores as subsequent inputs; when extracting answers, a deep learning Text-to-SQL model is used to convert the question into a standardized logical form such as SQL in combination with the question and table schema information, execute the SQL on the table, and determine whether an error occurs in the execution result; and use this result as one of the relevance bases to trace back to the table reordering and incorporate it into a new round of similarity calculation. The present invention uses the execution result of the Text-to-SQL model as the similarity sorting basis for table retrieval, making the retrieved tables more accurate, and thus improving the result accuracy of the entire open-domain question-answering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of natural language processing and question-answering tasks, and relates to a table open-domain question-answering method based on Text-to-SQL. Background Art

[0002] Compared with closed-domain question answering, the answers in open-domain question answering are not limited to a given reading material, but exist in a large corpus. As an important way to store information, tables exist in large numbers in web services and relational databases. Compared with the free text form, tables store a large amount of information and the content is more specific, which is an important information source for open-domain question answering. The table open-domain question-answering task changes the retrieval object from text to table, and studies how to obtain the information that users are interested in from a large number of tables, which can play an important role in search engines and artificial customer service.

[0003] Currently, the mainstream method for open-domain question answering is a two-stage framework: the retrieval-reading model, which divides open-domain question answering into two stages: the purpose of the retrieval stage is to find text segments related to the question from thousands of texts, and the reading stage extracts answers from these relevant contents. For table open-domain question answering, the reading stage can be regarded as a closed-domain table question answering.

[0004] In terms of retrieval, the traditional BM25 method and the deep learning-based DPR method are both for free text and do not make special optimizations for tables. Some scholars have proposed DTR. Compared with DPR, it uses Tapas to replace Bert as the encoder and encodes the table structure. However, the retrieval task is strongly related to the table content, and only encoding the table structure results in unsatisfactory table retrieval effects.

[0005] Great research progress has been made in closed-domain table question answering. There are two main methods. One is to encode the structure and content of the table and select the target cells as answers, such as Tapas; the other is based on semantic parsing, which converts the question described in natural language into a logical form that can be executed on the table, such as SQL, and then executes the statement to obtain the corresponding answer to the question. In the current research on open-domain table question answering, the first method mentioned above is mainly used, that is, directly extracting answers using a weakly supervised method, which misses a large amount of data stored in relational databases that requires structured language to access. At the same time, the models of the closed-domain Text-to-SQL task have achieved an accuracy rate of more than 90% on large datasets such as WikiSQL, and table retrieval has become the bottleneck for improving the accuracy of the entire open-domain question answering. Summary of the Invention

[0006] In response to the above problems, the present invention provides a method for table open domain question answering based on Text-to-SQL and execution guidance. When extracting answers, a deep learning Text-to-SQL model is used to convert the question into a standardized logical form such as SQL in combination with the question and table pattern information, and SQL is executed on the table to determine whether an error occurs in the execution result; and this result is used as one of the correlation bases to trace back to the table retrieval stage and integrated into a new round of similarity calculation to improve the accuracy of table retrieval.

[0007] The technical solution of the present invention is:

[0008] The process of open domain question answering is divided into: table retrieval stage and answer extraction stage;

[0009] The table retrieval stage is based on the input question and the given table corpus, and includes three sub-stages:

[0010] a1. Table preprocessing. Use tools such as SQLite to convert the table corpus into a db format so that SQL statements can be executed on any table. If the table itself exists in a relational database, skip the preprocessing stage; flatten the table by row to form the table content in text form.

[0011] a2. Carry out preliminary search. Encode the table content and load it into the index library; use the retriever to calculate the original similarity between the question and each table in the table corpus; sort the original similarity from large to small loss, and select the top N tables to form a table pool.

[0012] a3. Table reordering: Select the top-k tables from the table pool according to the similarity score to enter the subsequent stage.

[0013] The answer extraction phase has the following steps:

[0014] b1. Semantic parsing: For each table, the question and table pattern information are input into the Text-to-SQL deep learning model to obtain the corresponding SQL statement.

[0015] b2. Execute the obtained SQL semantics on the table to obtain candidate answers; the answer may be the content of a single cell of the table, the content of a group of cells, or the aggregate result of a group of cells.

[0016] The table reordering, i.e. the method of selecting the top-k tables from the table pool, depends on the execution of the subsequent answer extraction stage, which is called execution-guided backtracking. The specific method is:

[0017] After obtaining the table pool, initialize the execution of guided backtracking. Set the maximum limit of the execution of guided backtracking to X, and the current backtracking count to x = 0; K is a pre-set value, K <= N, and initially set k = K.

[0018] At the end of the table retrieval in the initial situation, directly use the top top-k tables with the highest original similarity in the table pool to enter the answer extraction stage. After that, whenever the execution stage of the answer extraction phase ends, that is, after step b2 is executed, the following conditional judgment is made:

[0019] If the backtracking count x < X, then judge whether execution errors occur for each table. If no execution error occurs, then set res EG = 1. If an execution error occurs, then set res EG = 0; Backtrack to the re-ranking stage of table retrieval and calculate the new similarity score according to the following formula:

[0020] sim withEG = (1 - α)·sim origin / maxsim origin + α·res EG

[0021] where sim origin is the original similarity score in the retrieval stage, maxsim origin is the maximum original similarity score of all tables in the retrieval stage, and sim withEG is the new similarity score;

[0022] Re-sort the N tables in the table pool according to the new similarity score sim withEG from largest to smallest to obtain the top top-k tables with the highest new similarity score;

[0023] Update the value of k according to certain rules, such as decreasing or halving; increase the backtracking count x by 1; Select the top top-k tables with the highest new similarity score and enter the answer extraction stage again.

[0024] If the backtracking count x >= X, then stop backtracking, and use the candidate answer corresponding to the table with the highest similarity ranking and no execution error as the final answer.

[0025] The beneficial effect of the present invention is that: by using the Text-to-SQL model in the answer extraction stage, after executing the SQL, the execution result is backtracked and used as one of the bases for the similarity ranking in the retrieval stage, making the retrieved tables more accurate, and thus improving the result accuracy of the entire open-domain question answering. Brief Description of the Drawings

[0026] Figure 1 is a schematic diagram of the two-stage model for open-domain question answering.

[0027] Figure 2 It is a model flow chart for table open-domain question answering based on Text-to-SQL and execution guidance.

[0028] Figure 3 It is a comparison of experimental results with and without using the execution-guided backtracking strategy. Specific implementation manners

[0029] The present invention will be described in detail below with reference to the accompanying drawings.

[0030] In the present invention, the two-stage model for table open-domain question answering is as Figure 1 . The model flow for table open-domain question answering based on Text-to-SQL and execution guidance is as Figure 2 , which involves two deep learning models: a retriever and a Text-to-SQL model. Before the actual question answering task, the corresponding model training needs to be completed in combination with the annotated table corpus data.

[0031] Preprocessing of the table. The table corpus is converted into a db form using tools such as SQLite, so that SQL statements can be executed on any table. If the table already exists in a relational database, the preprocessing stage is skipped.

[0032] Flattening and expansion of the table. Before building the index library, all tables in the corpus are processed into a form of continuous text. The title, column names, and row contents of the table are concatenated in the following form. If there is no table title, it is filled with spaces.

[0033] Table title|Table column names|First row content|Second row content|…|nth row content

[0034] During the concatenation process, delimiters are added to the concatenated text after the table conversion. Use "," to represent the separation of each cell in a row, and "." to represent the separation between rows. The delimiters enable the retriever to learn the structure of the table to a certain extent.

[0035] Perform preliminary retrieval. The goal of the retriever is to retrieve N tables T 1 , T 2 ,..., T N from a large table corpus of tens of thousands, in descending order of scores according to the relevance to the question q, as the table pool related to the question sentence. Since this is a preliminary retrieval stage, N is generally a relatively large number. Considering the efficiency of subsequent screening, N = 200 is set.

[0036] The retriever can adopt traditional methods such as BM25. Load the preprocessed table into the ElasticSearch index library, set the built-in similarity algorithm of ElasticSearch to BM25, set the algorithm parameters as k = 1.2 and b = 0.75, and specify the number of query results to be N. Then it can be used to implement the preliminary retrieval of the table. The BM25 method does not require model training.

[0037] The retriever can also use the deep learning dense retriever DPR. DPR is a dual encoder based on Bert. The two encoders respectively encode the question q nl and the table T i into vectors v q and The vector lengths are the same as the encoding length of BERT, which is d = 728. After fine-tuning, DPR will convert the question and the table into vector forms and ensure that the inner product of semantically related question-table pairs is larger than that of other unrelated pairs, and it can better capture semantic similarity, while traditional methods such as BM25 are more sensitive to keywords.

[0038] During the implementation process, a DPR retriever is trained based on the tiled table corpus. In the training process, 1 non-gold table with the highest BM25 score is selected as a hard example, batch_size = 32 is set, and the In-Batch negative training method is used. After training is completed, use one of the dual encoders of DPR: the table encoder to perform similar processing on all tables in the corpus to obtain the encoded table vectors, and load them into the dense vector retrieval library FASSI. During the preliminary retrieval, use one of the dual encoders of DPR: the question encoder to encode the question into a vector and perform a search in the FASSI library.

[0039] For the target question, the retriever returns the top N tables with higher original similarity as the table pool.

[0040] Perform the initialization of guided backtracking. Set the maximum limit number of times for guided backtracking to X, the current backtracking number to x = 0; K is a pre-set value, K <= N, and initially set k = K. In a simple implementation, X = 1, K = N = 200 can be set.

[0041] Table re-ranking. Select the top top-k tables with the highest similarity scores and enter the answer extraction stage. If backtracking has not been performed, directly use the top top-k tables with the highest original similarity in the table pool to enter the answer extraction stage. If backtracking has been performed, re-rank the N tables in the table pool according to the new similarity score sim withEG from large to small and use this as the basis for selecting the top-k tables.

[0042] Semantic parsing. A Text-to-SQL deep learning model exemplified by HydraNet is introduced as the semantic parser for the questions. A HydraNet model is trained based on the WikiSQL dataset, using batch_size = 64, learn_rate = 6*10 -6 , and the base model uses Roberta. During the table reordering process, for each table in the table pool, the question and the table schema information are first input into the HydraNet model to parse out the SQL statement corresponding to the question.

[0043] Execute the SQL statement on the table and count whether an error occurs during the execution.

[0044] Judge whether the current backtracking count x is greater than or equal to the maximum limit X for guiding backtracking. If not, perform the execution-guided backtracking. Otherwise, stop backtracking, and use the candidate answer corresponding to the table with no execution error and the highest similarity ranking as the final answer.

[0045] The specific method for performing the execution-guided backtracking is as follows: Convert the results executed on the candidate table into additional parameters according to different types: res EG , when no execution error occurs, res EG = 1, when an execution error occurs, res EG = 0. In reality, the expected result is often not empty when asking questions, so an empty result can be regarded as a special execution error.

[0046] Calculate a new similarity score based on the following function,

[0047] sim withEG = (1 - α)·sim origin / maxsim origin +α·res EG

[0048] where sim origin is the original similarity score in the retrieval stage, maxsim origin is the maximum original similarity score of all tables in the retrieval stage, sim withEG is the new similarity score; the coefficient α is used to measure the importance of the execution result, and the original retrieval correlation score and the execution result are linearly summed to obtain a new similarity estimation score. Different coefficient α values are taken according to the actual situation of the table corpus; in this implementation process, α = 0.9 is taken.

[0049] Update the k value, using a decreasing or halving strategy; in this implementation, directly set k = 1. Increment the x value, and then backtrack to the table reordering stage.

Claims

1. A table open-domain question answering method based on Text-to-SQL, characterized in that, it includes: a table retrieval stage and an answer extraction stage; The table retrieval stage is based on the input question and the given table corpus, and includes three sub-stages: a1. Table preprocessing: Convert the format of the table corpus so that SQL statements can be executed on any table; Tile the table row by row to form the table content in text form; a2. Initial retrieval stage: Use the retriever to calculate the original similarity of each table in the table corpus; Sort according to the loss from large to small of the original similarity, and select the top N tables to form a table pool; a3. Re-ranking stage: Select the top top-k tables from the table pool; The input of the answer extraction stage is the top top-k tables obtained in the table retrieval stage, and the process is divided into two sub-stages: b1. Semantic parsing stage: For each table, input the question and the table schema information into the Text-to-SQL deep learning model to obtain the corresponding SQL statement; b2. Execution stage: Execute the obtained SQL semantics on the table to obtain candidate answers; The answer is one of the content of a single cell of the table, the content of a group of cells, or the aggregation result of a group of cells; Among them, the re-ranking stage, that is, the method of selecting the top top-k tables from the table pool, depends on the execution situation of the subsequent answer extraction stage. The specific method is: Set the maximum limit number of times for executing guided backtracking to X, and the initial backtracking number is x = 0; K is a pre-set value, K <= N, and initially set k = K; In the initial situation, directly use the top top-k tables with the highest original similarity in the table pool to enter the answer extraction stage. After that, every time step b2 is executed, the following conditional judgment is made: If the backtracking number x >= X, stop backtracking, and use the candidate answer corresponding to the table with the highest similarity ranking without execution error as the final answer; If the number of backtracking times x < X, then judge whether an execution error has occurred for each table. If no execution error has occurred, then set res EG = 1. If an execution error has occurred, then set res EG = 0; go back to the reordering stage of table retrieval and calculate the new similarity score: sim withEG = (1 - α)·sim origin / maxsim origin + α·res EG where sim origin is the original similarity score in the retrieval stage, maxsim origin is the maximum original similarity score of all tables in the retrieval stage, and sim withEG is the new similarity score; Re - order the N tables in the table pool according to the new similarity score sim withEG to obtain the top - k tables with the highest new similarity scores, sorted from largest to smallest; Update the k value in a decreasing or halving manner, increase the backtracking number x by 1, and select the new top-k tables with the highest similarity score to enter the answer extraction stage again.

Citation Information

Patent Citations

  • System and method for transferable natural language interface

    CA3135717A1

  • Geographic knowledge question-answering system based on ontology semantic similarity

    CN110659357A