Table question and answer data generation method and device based on large language model and medium

Through the table question and answer data generation method based on large language model, the data set is rewritten and enhanced, and the existing data set cannot cover complex scenarios is solved, high-quality and diverse data generation is achieved, and the performance of the table question and answer system is improved.

CN120162408APending Publication Date: 2025-06-17ZHEJIANG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510230936.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing table question and answer dataset cannot effectively cover large-scale and high-complex tables in real scenarios, resulting in weak generalization capabilities of the model and lack of diversity, making it difficult to solve complex problems.

Method used

The table question and answer data generation method based on large language models is adopted to improve the diversity and quality of the data by obtaining seed data sets, rewriting questions, sampling or amplifying the tabular data, and performing quality inspection and enhancement.

Benefits of technology

The generated data sets have higher diversity and quality in question and tabular data, covering different table lengths and complex questions, improving the overall performance of the table question and answer system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162408A_ABST
    Figure CN120162408A_ABST
Patent Text Reader

Abstract

The invention discloses a table question and answer data generation method and device based on a large language model and a medium. The method comprises the steps that a table question and answer data set is obtained to serve as a seed data set; for each iteration generated by the table data, sampling a piece of table data from the seed data set; filling the table data and the problem rewriting direction into a cue word template, and rewriting an original problem of the table data through a large language model to obtain a rewritten problem; sampling or amplifying the table data; inputting the rewriting question and the sampled or amplified table data into a large language model to generate a model response, and taking the model response as a rewriting answer corresponding to the rewriting question; performing quality inspection on the rewritten answer; taking the rewriting answers and the rewriting questions passing the quality inspection as updated table data; and enhancing the updated table data, adding the enhanced table data into the seed data set of the next iteration, and performing iteration to obtain a table question and answer data generation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of large language models and table question answering, and in particular, relates to a method, device, and medium for generating table question answering data based on a large language model. Background Art

[0002] Table question answering is an important task in the field of natural language processing, which aims to answer natural language questions raised by users from structured table data. However, the tables involved in current table question answering datasets are usually small, with simple structure and limited data volume, which is often far from the large-scale and complex tables in real scenarios. For example, tables in reality usually contain a large number of rows and columns, and have complex hierarchical structures, foreign key dependencies, and various column types. Existing datasets often cannot cover these complex scenarios, resulting in weak generalization ability of the model when processing large-scale and highly complex tables.

[0003] In addition, current table question answering datasets often lack diversity, especially in terms of table length, question difficulty, etc. This makes the performance of most table question answering models limited by the scale and diversity of existing datasets, making it difficult to solve many table question answering problems in real scenarios. Therefore, there is an urgent need for a method to generate high-quality table question data with higher diversity, richer data types, and covering different table lengths and complex question difficulties.

[0004] Large language models have been widely used to generate synthetic data and enhance training sets due to their powerful text understanding and generation capabilities. In particular, they have made significant progress in data generation in the fields of mathematics and programming. However, traditional synthetic data methods based on large language models are not adapted to tabular question answering data, making it difficult to synthesize high-quality and diverse tabular question answering data. Summary of the invention

[0005] To solve these technical problems, the present invention proposes a method, device, and medium for generating tabular question and answer data based on a large language model.

[0006] In a first aspect, an embodiment of the present invention provides a method for generating tabular question-answer data based on a large language model, the method comprising the following steps:

[0007] Get the table question answering dataset as a seed dataset;

[0008] For each iteration of tabular data generation, a tabular data is sampled from the seed data set; a question rewriting direction is set, the tabular data and the question rewriting direction are filled into the prompt word template, and the original question of the tabular data is rewritten through the large language model to obtain a rewritten question;

[0009] Sampling or amplifying the tabular data;

[0010] Input the rewriting problem and the tabular data after sampling or amplification into a large language model to generate a model response, and use this model response as the rewritten answer corresponding to the rewriting problem;

[0011] Conduct quality inspection on the rewritten answer; use the rewritten answer that passes the quality inspection and the rewriting problem as the updated tabular data; enhance the updated tabular data, and add the enhanced tabular data to the seed data set for the next iteration;

[0012] Obtain the generation result of tabular Q&A data according to the seed data sets of multiple iterations.

[0013] In a second aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor, and the memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned method for generating tabular Q&A data based on a large language model.

[0014] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned method for generating tabular Q&A data based on a large language model is implemented.

[0015] In a fourth aspect, an embodiment of the present invention provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the above-mentioned method for generating tabular Q&A data based on a large language model is implemented.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0017] The present invention provides a method for generating tabular Q&A data based on a large language model. The present invention makes full use of the powerful semantic understanding and generation capabilities of the large language model to rewrite the data in the existing seed data set, and cooperates with operations such as amplification, sampling or enhancement of the table, so that the synthetic data set has better diversity and quality in terms of questions and tabular data compared with the seed data set, and covers different table lengths and complex question difficulties, laying a foundation for subsequent optimization of model training and improvement of the overall performance of the tabular Q&A system. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1Schematic flowchart of the method for generating table question-answering data based on a large language model provided by an embodiment of the present invention;

[0020] Figure 2 Schematic diagram of constructing a directed acyclic graph and topological sorting in table sampling considering foreign key constraints;

[0021] Figure 3 Schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0023] As Figure 1 shown, an embodiment of the present invention provides a method for generating table question-answering data based on a large language model. The method includes the following steps:

[0024] Step S1, obtain a table question-answering data set as a seed data set.

[0025] Further, the table question-answering data set is selected from the Spider Text-to-SQL data set, the BIRD Text-to-SQL data set, the TabFact table fact-checking data set, and the WikiTQ open-domain table question-answering data set.

[0026] Step S2, for each iteration of table data generation, sample a table data from the seed data set; set the question rewriting direction, fill the table data and the question rewriting direction into the prompt template, and rewrite the original question of the table data through a large language model to obtain a rewritten question.

[0027] Further, in this example, by setting the question rewriting direction, a large language model is used to select a question rewriting direction to modify the question part of the table data; the question rewriting direction includes changing query conditions, adding query conditions, adding external knowledge, adding reasoning steps, and / or changing question types.

[0028] Among them, changing query conditions means that by modifying the query conditions in the original question, the query range or standard changes, and this change is usually reflected in the adjustment of restrictive conditions such as numerical values, time, or categories;

[0029] Adding query conditions means adding additional query conditions on the basis of the original question to make the query result more accurate or refined;

[0030] Adding external knowledge: According to the column information in the table, introduce formulas or knowledge rules in a specific field, and modify the question to clearly require using this formula or knowledge rule to generate the question;

[0031] Adding reasoning steps: Increase the logical complexity of the question, so that the model needs to perform more steps of reasoning to obtain the answer;

[0032] Changing the question type: There are many types of lattice question-and-answer questions. For example, Text-to-SQL questions that generate SQL given text, fact-checking questions that give a factual statement and require the model to judge whether the statement is supported by data in the table, open-domain table question-and-answer questions that require the model to directly give the result based on the context and world knowledge, etc. Different question types have different answering methods. By changing the question type, the diversity of the dataset questions is increased.

[0033] It should be noted that in practical applications, those skilled in the art can also add other rewriting directions according to needs, which will not be elaborated here.

[0034] Exemplarily, the prompt word template for applying the large language selection question rewriting direction and generating rewritten questions is:

[0035] Table 1: Prompt word template for rewritten questions

[0036]

[0037]

[0038] Step S3, sample or amplify the table data.

[0039] It should be noted that in this example, in order to obtain tables of various sizes and increase the diversity of table sizes, it is necessary to randomly perform amplification or sampling operations on the tables of the current data to obtain new tables. Table sampling is relatively simple, and table amplification is generated with the help of a large language model.

[0040] To narrow the gap between tabular tasks and natural language generation tasks, this example converts the data rows in a table into natural language expressions. Specifically, for the value A in a certain column of the table (such as column a), this example converts it into a short sentence like "a is A". Then, the short sentences for each row are combined together to form a coherent discourse. This format is provided as the input to the large language model and is also the format requirement for the table output by the large language model. Considering the length limit of the input context by the large language model, this example only samples the first n rows of data in the table. In addition, for non-categorical numerical columns, this example also provides statistical information such as the average value of the column as meta-information of the table data; the meta-information also includes the table name, table description, and special generation requirements, etc. After obtaining the output of the large language model, this example will convert the generated natural language sentences back into table format and add the generated new table data to the original table.

[0041] Exemplarily, the prompt template for applying the large language model to augment the table is:

[0042] Table 2: Prompt Template for Table Augmentation

[0043]

[0044] Furthermore, when the table data is a non-relational database table, the process of augmenting the table data includes:

[0045] Set the prompt template for table augmentation;

[0046] Obtain the table meta-information corresponding to the table data;

[0047] Fill the table data and table meta-information into the prompt template for table augmentation, and apply the large language model to output the augmented table data.

[0048] It should be noted that for relational database tables, foreign key constraint conditions need to be considered during the process of addition and deletion. Foreign keys are used in relational databases to model the association relationships between tables to ensure data consistency and integrity. Exemplarily, as Figure 2 shown, the "User ID" column in the "Borrowing Record" table is a foreign key of the "User" table, and the "User" table is referenced by the foreign key of the "Borrowing Record" table. Among them, the "Borrowing Record" table is the child table and the "User" table is the parent table. This foreign key relationship means that the values in the "User ID" column of the "Borrowing Record" table must be found in the "User ID" column of the "User" table. Exemplarily, for example, for the following Table 3 and Table 4:

[0049] Table 3: Example 1 of "User" Table

[0050] User ID User Name E-mail 0001 Xiaoming ming@163.com 0002 Xiaohong hong@qq.com 0003 Zhang San san@outlook.com

[0051] Table 4: Example 1 of "Borrowing Record" Representation

[0052] Record ID User ID Book ID Borrowing Start Time Latest Return Time 0001 0001 0001 2019.02.19 2025.02.19 0002 0004 0002 2020.04.04 2026.04.04

[0053] Tables 3 and 4 are illegal tables that violate the foreign key constraint because "0004" in the "User ID" column of the "Borrowing Record" table does not exist in the "User ID" of the "User" table. Corresponding to the real situation, a person who has not registered in the borrowing system has a borrowing record, which belongs to problematic data. To ensure that the foreign key constraint is not violated during the process of expanding and sampling the relational database tables, additional processing is required.

[0054] Further, when the table data is a relational database table, consider the foreign key constraint condition to sample the table data. The specific process includes:

[0055] Take each table as a node in the graph (Graph), traverse the foreign keys, and establish a directed edge from the child table to the parent table, indicating that the child table depends on the parent table, thus obtaining a directed acyclic graph;

[0056] Perform a topological sort on the graph formed by the tables, so that the parent table is always ranked before the child table to ensure that the delete operation conforms to the foreign key constraint. Specifically, it includes: ① Calculate the in-degree for each node, indicating how many dependency relationships the node has; ② Add all nodes with an in-degree of 0 to the queue; ③ Take out nodes from the queue in turn, put them into the sorting result, and subtract 1 from the in-degree of all its adjacent nodes; if the in-degree of an adjacent node becomes 0, add it to the queue; ④ Repeat this process until the queue is empty. Because it is a directed acyclic graph, the topological sort will definitely succeed.

[0057] Traverse the tables in the order of the topological sort; if the current table has no foreign keys, directly sample the table; if the current table has foreign key constraints, check the parent tables pointed to by the foreign keys in turn, and then sample the table. For example, assume that column a of table A depends on column b of table B, then it is necessary to find the data in column a that does not exist in column b and delete the rows where these data are located to clear the invalid references, and then perform sampling.

[0058] Further, when the table data is a relational database table, consider the foreign key constraint condition to expand the table data. The specific process includes:

[0059] Take each table as a node, traverse the foreign keys, and establish a directed edge from the child table to the parent table, thus obtaining a directed acyclic graph;

[0060] Perform a topological sort on the graph formed by the tables, and then perform reverse processing so that the child table is always ranked before the parent table;

[0061] Traverse the table in topological sorting order; if the current table is not referenced by a foreign key, directly amplify the table; if the current table is referenced by a foreign key, it is necessary to check each child table where the foreign key comes from one by one, and then amplify the table;

[0062] Among them, the process of amplifying the table includes: setting the table amplification prompt word template; obtaining the table meta information corresponding to the table data; filling the table data and table meta information into the table amplification prompt word template, and applying the large language model to output the amplified table data. For example, assuming that column a of table A depends on column b of table B, it is necessary to ensure that the data in column a appears in column b. When generating data, ensure that the rows containing these data (by adding requirements in the table meta information of table 2) to avoid invalid references.

[0063] Exemplarily, for non-relational database tables:

[0064] Table 5: "Automobile" Representation Example 1

[0065] Car ID Car Name On-sale Time 001 BYD Dolphin 2021.07 002 Tesla Model S 2021.06 003 XPENG P7 2020.04

[0066] Possible results of the sampling operation:

[0067] Table 6: Sampled "Automobile" Representation Example 1

[0068] Car ID Car Name On-sale Time 001 BYD Dolphin 2021.07 002 Tesla Model S 2021.06

[0069] For the amplification operation, this example uses the following prompt words:

[0070] Table 7: "Automobile" Representation Example 1 Amplification Prompt Words

[0071]

[0072]

[0073] Assume that the obtained output is "The automobile ID is 004, the automobile name is NIO ES8, and the release time is December 2018.", then the amplified table is:

[0074] Table 8: Amplified "Automobile" Representation Example 1

[0075] Car ID Car Name On-sale Time 001 BYD Dolphin 2021.07 002 Tesla Model S 2021.06 003 XPENG P7 2020.04 004 NIO ES8 2018.12

[0076] For relational database tables, assume that there is now the following table that satisfies the foreign key constraint as Figure 2 shown:

[0077] Table 9: "User" Representation Example 2

[0078] User ID User Name E-mail 0001 Xiaoming ming@163.com 0002 Xiaohong hong@qq.com 0003 Zhang San san@outlook.com

[0079] Table 10: Example 2 of "Book" Representation

[0080] Book ID Title Author 0001 Dream of the Red Chamber Cao Xueqin 0002 To Live Yu Hua 0003 Les Misérables Victor Hugo 0004 One Hundred Years of Solitude Gabriel García Márquez

[0081] Table 11: Example 2 of "Borrowing Record" Representation

[0082]

[0083]

[0084] And the foreign key dependencies are as Figure 2 shown, that is, the "User ID" column in the "Borrowing Record" table is a foreign key of the "User" table; the "Book ID" column in the "Borrowing Record" table is a foreign key of the "Book" table. First, the list [User, Book, Borrowing Record] is obtained through "constructing a directed acyclic graph" and "topological sorting".

[0085] For the table sampling considering foreign key constraint conditions, sampling is performed in the order of the list. Since there are no foreign keys in the "User" table and the "Book" table, direct sampling is performed. Assume the result is:

[0086] Table 12: Sampled Example 2 of "User" Representation

[0087] User ID User Name E-mail 0001 Xiaoming ming@163.com 0003 Zhang San san@outlook.com

[0088] Table 13: Sampled Example 2 of "Book" Representation

[0089] Book ID Title Author 0001 Dream of the Red Chamber Cao Xueqin 0002 To Live Yu Hua 0003 Les Misérables Victor Hugo

[0090] Process the "Borrowing Record" table in order. Since there are foreign keys in the "Borrowing Record" table, it is necessary to check the parent tables pointed to by the foreign keys in turn and delete the data that does not meet the foreign key constraints, that is, delete the data where the value of the "User ID" column is "0002" and the value of the "Book ID" column is "0004", and get:

[0091] Table 14: Processed Example 2 of "Borrowing Record" Representation

[0092] Record ID User ID Book ID Borrowing Start Time Latest Return Time 0001 0001 0003 2019.02.19 2025.02.19 0003 0003 0001 2021.05.06 2027.05.06

[0093] For the processed borrowing records, sampling is performed again. Assume the result is:

[0094] Table 15: Sampled Example 2 of "Borrowing Record" Representation

[0095] Record ID User ID Book ID Borrowing Start Time Latest Return Time 0001 0001 0003 2019.02.19 2025.02.19

[0096] Then Table 12, Table 13 and Table 15 are the sampled results.

[0097] For the table augmentation considering foreign key constraints, first reverse the topological sort list to get [Borrowing Record, Book, User], and then perform augmentation according to the new list order. Since the "Borrowing Record" table has no foreign keys, it is directly augmented. Assume the result is:

[0098] Table 16: Example 2 of the augmented "Borrowing Record" table

[0099] Record ID User ID Book ID Borrowing Start Time Latest Return Time 0001 0001 0003 2019.02.19 2025.02.19 0002 0002 0002 2020.04.04 2026.04.04 0003 0003 0001 2021.05.06 2027.05.06 0004 0003 0004 2022.06.07 2028.06.07 0005 0004 0005 2023.09.12 2029.09.12

[0100] Process the "Book" table in order. Since the "Book" table is referenced by the foreign key of the "Borrowing Record" table and "0005" does not appear in the "Book ID" of the "Book" table, it needs to be added to the table meta-information. The corresponding augmentation prompt is:

[0101] Table 17: Augmentation prompt for Example 1 of the "Car" table

[0102]

[0103] Assume the output obtained is "The book ID is 0005, the title is *Ordinary World*, and the author is Lu Yao.", then the augmented table is:

[0104] Table 18: Example 2 of the augmented "Book" table

[0105]

[0106]

[0107] Similarly, the augmented "User" table can be obtained:

[0108] Table 19: Example 2 of the augmented "User" table

[0109] User ID User Name E-mail 0001 Xiaoming ming@163.com 0002 Xiaohong hong@qq.com 0003 Zhang San san@outlook.com 0004 Zhang Xiao zhang@126.com

[0110] Then Tables 16, 18, and 19 are the augmented results.

[0111] Step S4, input the rewritten question and the sampled or augmented table data into the large language model to generate a model response, and use this model response as the rewritten answer corresponding to the rewritten question.

[0112] Step S5, perform quality inspection on the rewritten answer; use the rewritten answer that passes the quality inspection and the rewritten question as the updated table data; enhance the updated table data and add the enhanced table data to the seed dataset for the next iteration.

[0113] Furthermore, the process of performing quality inspection on the rewritten answer includes:

[0114] The quality of the rewritten answers is inspected using regular expression filtering rules, SQL execution filtering rules, or large language model filtering rules.

[0115] Among them, regular expression filtering means: the regular expression filters out the data for which the model refuses to answer questions.

[0116] SQL execution filtering means: for Text-to-SQL questions, execute the SQL code, and filter out this data if it cannot be executed properly or times out.

[0117] Large language model filtering means: use the large language model to compare the new question with the original question. If there are problems such as the new question has less improvement compared to the original question, the new question directly copies the original question, the new question itself is unreasonable or too simple, or the new answer does not match the new question, then filter out this data.

[0118] Exemplarily, the prompt template for asking the large language model to judge data quality:

[0119] Table 20: Prompt template for judging data quality

[0120]

[0121]

[0122] Furthermore, the enhancement of the updated tabular data includes: enhancing the diversity of the table without changing the semantics of the tabular data. In this example, the following enhancement methods can be adopted:

[0123] Fuzzifying the table column names: including converting the column names to pinyin, pinyin abbreviations, other languages, synonym replacement, etc. If it is a Text-to-SQL question, pay attention to the synchronous modification of the SQL answers;

[0124] Enhancing the tabular text information: If there are text columns with non-categorical features in the table, rewrite some of the text through the large language model;

[0125] Transposing the table: For non-Text-to-SQL questions, the table can be transposed;

[0126] Shuffling the table: Randomly shuffle the order of the rows or columns of the table;

[0127] Converting the table numbers: For non-Text-to-SQL questions, convert the numbers in non-categorical columns, which can be converted to English expressions, Chinese expressions, Chinese capital expressions, Roman numerals, scientific notation, etc.;

[0128] Table format conversion: For non-Text-to-SQL problems, convert the table representation format into one of the formats such as markdown, HTML, Latex, etc.

[0129] Step S6: Obtain the result of generating table Q&A data according to the seed data set after multiple iterations.

[0130] As Figure 3 shown, an embodiment of the present application provides an electronic device, which includes a memory 101 for storing one or more programs; a processor 102. When the one or more programs are executed by the processor 102, the method according to any one of the above first aspects is implemented.

[0131] It further includes a communication interface 103, and the memory 101, the processor 102, and the communication interface 103 are directly or indirectly electrically connected to each other to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The memory 101 can be used to store software programs and modules, and the processor 102 executes various functional applications and data processing by executing the software programs and modules stored in the memory 101. The communication interface 103 can be used for signaling or data communication with other node devices.

[0132] Among them, the memory 101 can be, but is not limited to, a random access memory 101 (Random Access Memory, RAM), a read-only memory 101 (Read Only Memory, ROM), a programmable read-only memory 101 (Programmable Read-Only Memory, PROM), an erasable read-only memory 101 (Erasable Programmable Read-Only Memory, EPROM), an electrically erasable read-only memory 101 (Electric Erasable Programmable Read-Only Memory, EEPROM), etc.

[0133] The processor 102 may be an integrated circuit chip with signal processing capabilities. The processor 102 may be a general-purpose processor 102, including a central processing unit 102 (CPU), a network processor 102 (NP), etc.; it may also be a digital signal processor 102 (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0134] In the embodiments provided in this application, it should be understood that the disclosed methods and systems may also be implemented in other ways. The method and system embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the methods, systems, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

[0135] In addition, in each embodiment of this application, the various functional modules may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0136] On the other hand, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor 102, the method according to any one of the above first aspects is implemented. If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory 101 (ROM, Read-Only Memory), a random access memory 101 (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0137] After considering the specification and practicing the content disclosed herein, those skilled in the art will readily think of other implementation manners of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and the embodiments are only regarded as exemplary.

[0138] It should be understood that the present application is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. A method for generating tabular question-answering data based on a large language model, characterized in that: The method comprises the following steps: Get the table question answering dataset as a seed dataset; For each iteration of tabular data generation, a tabular data is sampled from the seed data set; a question rewriting direction is set, the tabular data and the question rewriting direction are filled into the prompt word template, and the original question of the tabular data is rewritten through the large language model to obtain a rewritten question; Sampling or amplifying the tabular data; Inputting the rewritten question and the sampled or augmented table data into the large language model to generate a model response, and using the model response as a rewritten answer corresponding to the rewritten question; Performing a quality check on the rewritten answers; using the rewritten answers and rewritten questions that have passed the quality check as updated table data; enhancing the updated table data, and adding the enhanced table data to the seed data set of the next iteration; Based on multiple iterations of the seed dataset, the tabular question-answering data generation results are obtained.

2. The method for generating tabular question-answer data based on a large language model according to claim 1, characterized in that: The tabular question answering dataset is selected from a Spider dataset, a BIRD dataset, a TabFact dataset, and a WikiTQ dataset.

3. The method for generating tabular question-answer data based on a large language model according to claim 1, characterized in that: The question rewriting direction includes changing the query conditions, adding query conditions, adding external knowledge, adding reasoning steps and / or changing the question type.

4. The method for generating tabular question-answer data based on a large language model according to claim 1, characterized in that: When the table data is a non-relational database table, the process of expanding the table data includes: Set up a template for table expansion prompts; Obtaining table meta information corresponding to the table data; The table data and table meta-information are filled into a table augmentation prompt word template, and the augmented table data is output using a large language model.

5. The method for generating tabular question-answer data based on a large language model according to claim 1, characterized in that: When the table data is a relational database table, the table data is sampled by considering the foreign key constraint condition. The specific process includes: Take each table as a node, traverse the foreign keys, and establish directed edges from the child table to the parent table to obtain a directed acyclic graph; Topologically sort each table so that the parent table always comes before the child table; Traverse the table in the order of topological sorting; if the current table has no foreign key, sample the table directly; if the current table has a foreign key constraint, check the parent table pointed to by the foreign key in turn, and then sample the table.

6. The method for generating tabular question-answer data based on a large language model according to claim 1, characterized in that: When the table data is a relational database table, the table data is expanded by considering the foreign key constraint condition. The specific process includes: Take each table as a node, traverse the foreign keys, and establish directed edges from the child table to the parent table to obtain a directed acyclic graph; Topologically sort each table so that the parent table always comes before the child table; Traverse the table in the order of topological sorting; if the current table is not referenced by a foreign key, directly expand the table; if the current table is referenced by a foreign key, check the sub-tables of the foreign key source one by one, and then expand the table; The process of expanding the table includes: setting a table expansion prompt word template; obtaining table meta information corresponding to the table data; filling the table data and table meta information into the table expansion prompt word template, and applying a large language model to output the expanded table data.

7. The method for generating tabular question-answer data based on a large language model according to claim 1, characterized in that: The process of quality checking rephrased answers includes: The rewritten answers are quality checked through regular expression filtering rules, SQL execution filtering rules, or large language model filtering rules.

8. An electronic device, comprising a memory and a processor, characterized in that: The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the method for generating tabular question and answer data based on a large language model as described in any one of claims 1 to 7 above.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for generating tabular question and answer data based on a large language model as described in any one of claims 1 to 7 is implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the method for generating tabular question and answer data based on a large language model as described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Demand information feedback method and device, electronic equipment and storage medium

    CN120973840A