Table question and answer data construction method and device, computer equipment and storage medium
By obtaining seed question data and performing structural, semantic, character, and depth expansion based on user feature information, we generate target tabular question and answer data. This solves the problem of inaccurate answers in LLM in tabular question and answer, improves the iteration speed and accuracy of large language models, and reduces labor costs.
Patent Information
- Application Number
- CN202510823047.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-10
AI Technical Summary
When different users use the Large Language Model (LLM) to answer questions in a table, the answers are less accurate due to the ambiguity in the intent of the question language, and it is difficult to efficiently collect a large number of user questions for optimization.
By obtaining seed question data, determining the expansion method based on user feature information, generating initial extended question data, and generating a variety of extended question data through structural, semantic, character, noise and depth expansion methods, combined with the usage requirements of the initial extended question data, determining the target table question and answer data.
It achieves automatic and rapid generation of initial extended question data close to actual scenarios, optimizes the iteration speed of large language models, reduces labor costs, and improves the accuracy of answers and iteration efficiency.
Smart Images

Figure CN120763285A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of large language models, and in particular to a method, apparatus, computer equipment, and storage medium for constructing tabular question-and-answer data. Background Art
[0002] Datasheet Q&A plays a crucial role in today's information processing and decision support. It primarily supports Q&A for data stored in Excel and CSV files, and for data stored in databases. Datasheet Q&A uses natural language, using a Large Language Model (LLM) to convert natural language into Python or SQL code, which is then used to retrieve answers from the stored data.
[0003] However, the intentions of different users' question language when using LLM are vague, resulting in low accuracy of the answers obtained by LLM from the stored data based on the question language. Summary of the Invention
[0004] Based on this, it is necessary to provide a table question and answer data construction method, device, computer equipment and storage medium that can improve the accuracy of LLM in obtaining answers from stored data to address the above technical problems.
[0005] In a first aspect, the present application provides a method for constructing tabular question-and-answer data, comprising:
[0006] Obtain seed question data; the seed question data is data that the question-answering model in the large language model can determine the answer to the seed question data from the original table data;
[0007] determining an expansion method based on user feature information when using the question-answering model, and generating initial expanded question data based on the expansion method and the seed question data;
[0008] Target table question and answer data is determined according to the initial extended question data and the requirements for use of the initial extended question data.
[0009] In one embodiment, the expansion method includes a structural expansion method; generating initial expanded question data according to the expansion method and the seed question data includes:
[0010] The structure of the seed question data is expanded based on the structure expansion method to obtain first expanded question data with a structure similar to that of the seed question data; the initial expanded question data includes the first expanded question data.
[0011] In one embodiment, the expansion method includes a semantic expansion method; generating initial expanded question data according to the expansion method and the seed question data includes:
[0012] The semantics of the seed question data are expanded based on the semantic expansion method to obtain second expanded question data with the same semantics as the seed question data; the initial expanded question data includes the second expanded question data.
[0013] In one embodiment, the expansion method includes a character expansion method; and generating the initial expansion question data according to the expansion method and the seed question data includes:
[0014] Based on the character extension method, erroneous characters are randomly added to the seed question data to extend the seed question data and generate third extended question data; the initial extended question data includes the third extended question data.
[0015] In one embodiment, the expansion method includes a noise expansion method; and generating the initial expanded question data according to the expansion method and the seed question data includes:
[0016] Noise data is randomly added to the seed question data based on the noise expansion method to expand the seed question data and generate fourth extended question data; the initial extended question data includes the fourth extended question data.
[0017] In one embodiment, the expansion method includes a depth expansion method; generating initial expansion question data according to the expansion method and the seed question data includes:
[0018] The seed question data is expanded based on the depth expansion method to increase the question depth of the seed question data, and fifth expanded question data is generated; the initial expanded question data includes the fifth expanded question data.
[0019] In one embodiment, determining target table question and answer data based on the initial extended question data and the requirements for use of the initial extended question data includes:
[0020] Deduplication processing is performed on the initial extended question data to obtain intermediate extended question data;
[0021] When the requirement to be used for the initial extended question data is to optimize the large language model, obtaining data information of the original table data;
[0022] Detecting the correlation between the data information and the intermediate extended question data to determine target extended question data that meets the correlation requirement;
[0023] The target table question and answer data is determined according to the target extended question data and the original table data.
[0024] In a second aspect, the present application also provides a device for constructing table question and answer data, comprising:
[0025] An acquisition module is used to acquire seed question data; the seed question data is data that the question-answering model in the large language model can determine the answer to the seed question data from the original table data;
[0026] a generation module, configured to determine an expansion method based on user characteristic information when using the question-answering model, and generate initial expanded question data based on the expansion method and the seed question data;
[0027] A determination module is used to determine target table question and answer data based on the initial extended question data and the requirements for use of the initial extended question data.
[0028] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the method steps provided in the first aspect when executing the computer program.
[0029] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the method steps provided in the first aspect when the computer program is executed by a processor.
[0030] In a fifth aspect, the present application also provides a computer program product, comprising a computer program, which implements the method steps provided in the first aspect when executed by a processor.
[0031] The above-mentioned table question and answer data construction method, device, computer equipment and storage medium obtain seed question data, determine the expansion method according to the user feature information when using the question and answer model, and generate initial expanded question data according to the expansion method and the seed question data, and determine the target table question and answer data according to the initial expanded question data and the requirements for the use of the initial expanded question data; the seed question data is the data by which the question and answer model in the large language model can determine the answer to the seed question data from the original table data. In the embodiment of the present application, the expansion method is determined based on the user feature information, and with the help of the large language model and the expansion method, the initial expanded question data close to the actual scenario can be automatically and quickly generated to obtain the target table question and answer data that can optimize the large language model, thereby reducing labor costs based on improving the iteration speed of the large language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.
[0033] Figure 1 This is a diagram of an application environment of a method for constructing table question-and-answer data in one embodiment;
[0034] Figure 2 Schematic diagram of a flow chart of a method for constructing tabular question-and-answer data in one embodiment;
[0035] Figure 3 Schematic diagram of a flow chart of a method for determining target table question and answer data in one embodiment;
[0036] Figure 4 A flowchart of a method for constructing tabular question-and-answer data in another embodiment;
[0037] Figure 5 It is a structural block diagram of a device for constructing table question and answer data in one embodiment. DETAILED DESCRIPTION
[0038] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0039] Datasheet Q&A plays a crucial role in today's information processing and decision support. Datasheet Q&A is primarily categorized into Q&A for data stored in Excel or CSV files and Q&A for data stored in databases. Datasheet Q&A uses natural language. LLMs are used to convert natural language into Python or SQL code, which is then used to retrieve answers from the stored data. For example, for data stored in Excel or CSV files, LLMs are typically used to generate Python code, which is then executed to retrieve answers from the stored data. For data stored in databases, LLMs are typically used to generate SQL code for the corresponding database system, which is then executed to retrieve answers from the stored data. Both approaches require dynamically constructed prompts to guide the LLM in generating code. Prompts guide the LLM in generating specific output text. In natural language processing (NLP), prompts are often used to provide context, explain tasks, or pose questions, helping the LLM generate relevant and meaningful responses. Specifically, prompts are typically constructed based on the user's question, data headers, and sample data. NLP is a discipline that studies how to enable computers to understand, analyze, and generate human language. It involves processing text and speech data to achieve tasks such as text classification, sentiment analysis, machine translation, and question-answering systems.
[0040] At present, the intentions of different users' question language when using LLM are vague, resulting in low accuracy of the answers obtained by LLM from the stored data based on the question language. Therefore, a large number of questions need to be obtained to optimize LLM, but it is difficult to collect a large number of user questions. If a professional annotation team is used, firstly, the cost is high, and secondly, because the questions are constructed with tasks in mind, they will be too deliberate, and the questions finally constructed may deviate from the actual user questions. Therefore, this application proposes a method for constructing tabular question and answer data to solve the above technical problems.
[0041] The table question and answer data construction method provided in the embodiment of the present application can be applied to Figure 1 The application environment shown in FIG. The application environment includes a computer device, which may be a server, and its internal structure diagram may be as shown in FIG. Figure 1As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store relevant data for constructing table question and answer data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for constructing table question and answer data is implemented. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.
[0042] Those skilled in the art will understand that Figure 1 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0043] In an exemplary embodiment, Figure 2 As shown, a method for constructing tabular question-answer data is provided, which is applied to Figure 1 The computer device in the example is used to illustrate, including the following S201 to S203.
[0044] S201, obtaining seed question data; the seed question data is data that the question-answering model in the large language model can use to determine the answer to the seed question data from the original table data.
[0045] LLM is a statistical model used to predict the probability distribution of the next word or character in a given context. LLMs can be built by learning from large amounts of text data and are able to capture the patterns and structure of language. LLMs are widely used in many natural language processing (NLP) tasks, such as machine translation, speech recognition, and automatic summarization. LLMs can be used to generate coherent and fluent sentences, complete user input to provide better suggestions, and assess whether a given sentence conforms to grammatical and semantic rules.
[0046] In this embodiment of the present application, seed question data is data that the question-answering model within the large language model can use to determine the answer to the seed question data from the original table data. If the original table data is as shown in Table 1, the seed question data could be, for example, "Who has the highest total score, and which class?", "What is the average score for each class?", "What is the difference between the highest and lowest scores?", etc. This seed question data will subsequently be expanded to generate additional extended question data.
[0047] Table 1
[0048]
[0049] S202, determining an expansion method based on user feature information when using the question-answering model, and generating initial expanded question data based on the expansion method and the seed question data.
[0050] In related art, when a user asks "Who has the highest total score and which class is he from", the constructed prompt words are usually as follows:
[0051] "I have a data table with the following headers: student number, name, class, and total score. Here are some sample data:
[0052] <Data>
[0053] ("Student ID": 1, "Name": "Li Lei", "Class": "Class 1", "Total Score": 527),
[0054] (Student ID: 2, Name: Han Meimei, Class: Class 1, Total Score: 642),
[0055] < / 数据>
[0056] I have a question: Who has the highest total score and which class is he from?
[0057] Please help me generate a Python code based on the sample data above to get the corresponding answer."
[0058] Although LLM already has relatively powerful code generation capabilities, it is currently mainly used in the field of auxiliary programming. The code generation in the auxiliary programming field is significantly different from the code generation in the field of table question and answer. The most core difference is that in the field of auxiliary programming, the context given to LLM is usually code, and the characteristics of code are strong logic and no ambiguity. Table question and answer uses natural language and information related to table data as context to give to LLM. The most significant feature of natural language is that it is not precise enough, and the questions asked by different users may themselves be ambiguous in intent. In this case, directly using LLM to generate code for table data analysis will produce a lot of wrong answers. Therefore, the embodiment of the present application summarizes the user feature information when users use the table question and answer model for data analysis, and obtains the user feature information of user questions in these scenarios, including: user expressions are flexible and diverse; user expressions are more arbitrary and less standardized; some advanced users ask more in-depth questions; some users ask more concise questions, etc.
[0059] In the embodiments of the present application, different expansion methods are determined for different user feature information. For example, for the characteristic of "user expressions are flexible and diverse", the corresponding expansion methods are determined to be structural expansion and semantic expansion, so as to construct different expressions of the same seed question data; for the characteristic of "user expressions are relatively arbitrary", the corresponding expansion methods are determined to be character expansion and noise expansion; for the characteristic of "some advanced users ask more in-depth questions", the corresponding expansion method is determined to be depth expansion, so as to obtain more in-depth question data than the seed question data; for the characteristic of "some users ask more concise questions", the corresponding expansion method is determined to be simplified character expansion.
[0060] In an embodiment of the present application, a large language model and different expansion methods are used to expand the seed question data, which may include the following implementation methods.
[0061] In one possible implementation, the structure of the seed question data is expanded based on a structural expansion method to obtain first expanded question data with a structure similar to the seed question data. The semantics of the seed question data is expanded based on a semantic expansion method to obtain second expanded question data with the same semantics as the seed question data. Error characters are randomly added to the seed question data based on a character expansion method to expand the seed question data and generate third expanded question data. Noise data is randomly added to the seed question data based on a noise expansion method to expand the seed question data and generate fourth expanded question data. The seed question data is expanded based on a depth expansion method to increase the question depth of the seed question data and generate fifth expanded question data; the initial expanded question data includes the first expanded question data, the second expanded question data, the third expanded question data, the fourth expanded question data, and the fifth expanded question data.
[0062] In another possible implementation, some characters in the seed question data can be removed based on the simplified character expansion method to simplify the expression of the seed question data and obtain the sixth extended question data; each feature corresponds to an expansion method, for example, the initial extended question data can include the first extended question data, the third extended question data, the fifth extended question data and the sixth extended question data.
[0063] S203: Determine target table question and answer data according to the initial extended question data and the requirements for use of the initial extended question data.
[0064] In an embodiment of the present application, if the initial expanded question data is to be used for optimizing a large language model, the correlation between the initial expanded question data and the original table data can be obtained, and the correlation can be used to determine whether the original table data can answer the questions in the initial expanded question data. The initial expanded question data that can be answered by the original table data is used as the target expanded question data, and the target expanded question data, the seed question data, and the original table data are used as the target table question and answer data.
[0065] In one possible implementation, the initial extended question data may be deduplicated to obtain intermediate extended question data. A correlation between the intermediate extended question data and the original table data is obtained, and the correlation is used to determine whether the original table data can answer the questions in the intermediate extended question data. The intermediate extended question data that can be answered by the original table data is used as the target extended question data, and the target extended question data, the seed question data, and the original table data are used as the target table question-and-answer data.
[0066] In another possible implementation, if the initial extended question data is to be used for building a classification model, there is no need to determine the correlation between the initial extended question data and the original table data. Alternatively, there is no need to determine the correlation between the intermediate extended question data and the original table data.
[0067] In the above-mentioned table question and answer data construction method, seed question data is obtained, an expansion method is determined based on the user feature information when using the question and answer model, and initial expanded question data is generated based on the expansion method and the seed question data. Target table question and answer data is determined based on the initial expanded question data and the requirements for the initial expanded question data to be used; the seed question data is data that the question and answer model in the large language model can use to determine the answer to the seed question data from the original table data. In an embodiment of the present application, the expansion method is determined based on user feature information, and with the help of the large language model and the expansion method, initial expanded question data close to the actual scenario can be automatically and quickly generated to obtain target table question and answer data that can optimize the large language model, thereby reducing labor costs based on improving the iteration speed of the large language model.
[0068] In one embodiment, the initial extended question data is generated according to the extension method and the seed question data, including the following five methods:
[0069] The first method: the structure of the seed question data is expanded based on the structure expansion method to obtain first expanded question data with a structure similar to that of the seed question data; the initial expanded question data includes the first expanded question data.
[0070] In the embodiment of the present application, since some question data can be classified into the same category of question data, mainly because they have the same grammatical structure, but the grammatical structure cannot be exhaustively enumerated, similar prompt words can be used to construct "similarly structured" question data, and the prompt words are as follows:
[0071] "Help me create three different statements based on the structure of the following sentence. Include only the return statement, no other descriptive content, and use line breaks to separate the sentences:
[0072] [Seed problem data]".
[0073] Assume that the seed question data input into the above prompt word is "What is the difference between the months with the highest profits from 2018 to 2021?" The first expanded question data constructed is as follows:
[0074] What was the difference in annual growth rates between 2008 and 2012, the years with the fastest growth rates?
[0075] How much did the average increase in employees’ best performance months from 2017 to 2020?
[0076] How much difference does the highest earning quarter make from 2017 to 2020 each year?
[0077] As can be seen, the seed question data only focuses on "profit", while the first extended question data constructed based on the large language model pays attention to "growth rate", "employee performance" and "revenue".
[0078] The second way is to expand the semantics of the seed question data based on the semantic expansion method to obtain second extended question data with the same semantics as the seed question data; the initial extended question data includes the second extended question data.
[0079] In the embodiments of the present application, since different users have different expressions for the same question, i.e., have the same keywords, the question data can be constructed using the following prompt words, and the prompt words are as follows:
[0080] "Help me create 3 different sentences based on the semantics of the following sentence. Only return the sentences, do not include other descriptive content, and use a newline character to separate different sentences:
[0081]
Seed question data
[0082] Suppose the seed question data input into the above prompt words is "How much difference does the highest earning quarter make from 2017 to 2020 each year?", and the second extended question data constructed is as follows:
[0083] How much difference does the highest earning quarter make from 2017 to 2021 each year?
[0084] How much difference does the highest earning quarter make from 2018 to 2021 each year?
[0085] How much difference does the highest earning quarter make from 2018 to 2021 each year?
[0086] As can be seen, the second extended question data has the same semantics as the seed question data, both of which focus on the profit difference between the highest earning months, only the form of expression is different.
[0087] The third way is to randomly add error characters to the seed question data based on the character expansion method to expand the seed question data and generate third extended question data; the initial extended question data includes the third extended question data.
[0088] In the embodiment of this application, since users often ask questions containing typos when using the question-answering model, directly using the question-answering model to generate code will bring the typos into the query conditions, resulting in incorrect results. To avoid this, it is usually necessary to collect such data in advance and perform targeted optimization. The following prompt words can be used to construct such question data, which are as follows:
[0089] "Help me create two sentences for the following sentence by adding a few typos to make the large language model more robust. Return only the sentence, without any other descriptive content, and use line breaks to separate different sentences:
[0090] [Seed problem data]".
[0091] Assume that the seed question data entered into the above prompt is "Which city has the highest GDP and what percentage of the province it accounts for?" The third extended question data constructed using this method is as follows:
[0092] Which city has the highest GDP and what is its proportion to the entire province?
[0093] Which city has the highest GDP and what proportion does it account for in the province?
[0094] Among them, "战" and "列" are typos introduced to simulate incorrect characters input by users in actual scenarios.
[0095] The fourth method: randomly adding noise data to the seed question data based on the noise expansion method to expand the seed question data and generate fourth extended question data; the initial extended question data includes the fourth extended question data.
[0096] In the embodiment of the present application, since users may not necessarily use very standard sentences to ask questions when actually using the question-answering model, there may be problems such as sentences that do not strictly comply with grammatical specifications or inversion. This generation method mainly solves such problems and can use the following prompt word construction. The prompt words are as follows:
[0097] "Help me create two sentences for the following sentence by adding a small amount of noise data to make my large language model more robust. Return only the sentence, without any other descriptive content, and use line breaks to separate different sentences:
[0098] [Seed problem data]".
[0099] Assume that the seed question data input into the above prompt is "Which city has the highest GDP and what percentage of the province it accounts for?" The fourth extended question data constructed is as follows:
[0100] Which city has the highest GDP level, and what proportion does this figure account for in the province's GDP?
[0101] Which city has the highest GDP and what percentage of the province does it account for?
[0102] The fifth method: expanding the seed question data based on the depth expansion method to increase the question depth of the seed question data and generate the fifth expanded question data; the initial expanded question data includes the fifth expanded question data.
[0103] In the embodiment of the present application, some senior analysts usually have a deeper analysis of the problem and ask more in-depth questions. The following prompt words can be used to construct such question data. The prompt words are as follows:
[0104] “Based on the following sentence, create five more in-depth questions to help me better understand the data. Return only the sentence, without any other descriptive content. Use line breaks to separate the sentences:
[0105] [Seed problem data]".
[0106] Assume that the seed question data input into the above prompt word is "What is the name of the product with a critical order priority on April 8, 2020?" The fifth extended question data constructed is as follows:
[0107] Are there any duplicate orders for products with a critical order priority on April 8, 2020? What is the duplication rate?
[0108] Compared with April 8, 2020, how did the number of critical priority orders increase or decrease in the week before and after?
[0109] What is the frequency distribution of product names in Critical priority orders on April 8, 2020?
[0110] Are there any specific customer group preferences for products in the Critical priority order on April 8, 2020?
[0111] Which regions are these critical priority orders shipped to?
[0112] It can be seen that the seed question data only focuses on the product name, while the fifth extended question data constructed based on the deep expansion method further focuses on the product's customer base, shipping location, and growth status.
[0113] In an embodiment of the present application, the structure of the seed question data is expanded based on a structural expansion method to obtain first expanded question data with a structure similar to the seed question data, the semantics of the seed question data is expanded based on a semantic expansion method to obtain second expanded question data with the same semantics as the seed question data, erroneous characters are randomly added to the seed question data based on a character expansion method to expand the seed question data to generate third expanded question data, noise data is randomly added to the seed question data based on a noise expansion method to expand the seed question data to generate fourth expanded question data, and the seed question data is expanded based on a depth expansion method to increase the question depth of the seed question data to generate fifth expanded question data. In an embodiment of the present application, a variety of expansion methods are used to expand the seed question data, which can quickly construct a problem data set close to the actual scenario without the need for labeling personnel to access, thereby discovering problems to be solved earlier and faster, reducing labor costs, and improving the iteration efficiency of large language models.
[0114] Figure 3 FIG. 1 is a flow chart of a method for determining target table question-answer data in one embodiment. Figure 3 As shown, the embodiment of the present application relates to a possible implementation method of determining target table question and answer data based on initial extended question data and the requirements for use of the initial extended question data, including the following steps:
[0115] S301 , performing deduplication processing on initial extended question data to obtain intermediate extended question data.
[0116] In an embodiment of the present application, due to the sampling characteristics of a large language model, initial expanded question data generated using different expansion methods may overlap. Therefore, all generated initial expanded question data are uniformly deduplicated to obtain intermediate expanded question data.
[0117] S302 : When the requirement to be used for the initial extended question data is to optimize a large language model, data information of the original table data is obtained.
[0118] Optionally, the large language model may include the above-mentioned question-answering model, or a code generation model, etc.
[0119] If the initial extended question data is used to optimize a large language model, it is necessary to ensure that the generated intermediate extended question data is related to the original table data, that is, it can be answered using the original table data. Therefore, in this embodiment of the application, the data information of the original table data can be obtained through Python methods. Optionally, the data information can include field names, column names, types, etc.
[0120] S303, detecting the relevance of the data information and the intermediate extended question data, and determining target extended question data meeting the relevance requirement.
[0121] In the embodiment of the application, the relevance of the data information and each intermediate extended question data is detected by inputting the data information and the intermediate extended question data into the large language model, and target extended question data meeting the relevance requirement is determined. The relevance requirement can be that the large language model outputs "yes".
[0122] For example, the relevance of the data information and each intermediate extended question data is obtained by using the large language model, and the constructed prompt word is as follows:
[0123] "def build_prompt (intermediate extended question data, data information):
[0124] prompt_tmpl = """
[0125] You are a senior data analyst.
[0126] I now have a sample data (corresponding to the original table data) and a question (corresponding to the intermediate extended question data), your task is to combine the column name (content in <columns>< / columns> ), data type (content in <dtypes>< / dtypes> ), and data typical value (content in <values>< / values> ), to determine whether the question (content in <question>< / question> ) can be answered by the information provided by the sample data.
[0127] <columns>
[0128]
List name
[0129] < / columns>
[0130] <dtypes>
[0131] Data type
[0132] < / dtypes>
[0133] <values>
[0134]
Typical data values
[0135] < / values>
[0136] <question>
[0137]
Intermediate extended problem data
[0138] < / question>
[0139] Only reply "yes" or "no", do not include any descriptive content and additional information, and the result does not include quotation marks.
[0140] S304, determining target table question and answer data according to the target extended question data and the original table data.
[0141] In the embodiment of the application, the target extended question data and the original table data are used as the target table question and answer data; or the target extended question data, the seed question data and the original table data are used as the target table question and answer data.
[0142] In the embodiments of the present application, the initial extended question data is de-duplicated to obtain intermediate extended question data. In the case where the to-be-used demand of the initial extended question data is to optimize a large language model, data information of the original table data is obtained, the relevance of the data information and the intermediate extended question data is detected, target extended question data meeting the relevance requirement is determined, and target table question and answer data is determined according to the target extended question data and the original table data. In the embodiments of the present application, the relevance of the data information and the intermediate extended question data is detected, and the target table question and answer data is determined based on the relevance detection result, thereby improving the accuracy of the determination of the target table question and answer data.
[0143] Figure 4 For another embodiment of the table question and answer data construction method, as shown in Figure 4 the flowchart, the following steps are included: obtaining seed question data, expanding the structure of the seed question data based on a structure expansion manner to obtain first extended question data similar to the structure of the seed question data (i.e., expanding using a structure similar manner); expanding the semantics of the seed question data based on a semantic expansion manner to obtain second extended question data identical to the semantics of the seed question data (i.e., expanding using a semantic identical manner); adding error characters to the seed question data randomly based on a character expansion manner to expand the seed question data and generate third extended question data (i.e., expanding using a manner containing a wrong character); adding noise data to the seed question data randomly based on a noise expansion manner to expand the seed question data and generate fourth extended question data (i.e., expanding using a manner of increasing noise); expanding the seed question data based on a depth expansion manner to increase the depth of the seed question data and generate fifth extended question data (i.e., expanding using a manner of a deeper question); de-duplicating the initial extended question data to obtain intermediate extended question data; in the case where the to-be-used demand of the initial extended question data is to optimize a large language model, obtaining data information of the original table data; detecting the relevance of the data information and the intermediate extended question data, determining target extended question data meeting the relevance requirement; determining target table question and answer data according to the target extended question data and the original table data; the seed question data is data from which a question and answer model in a large language model can determine the answer of the seed question data from the original table data; the initial extended question data includes the first extended question data, the second extended question data, the third extended question data, the fourth extended question data, and the fifth extended question data.
[0144] In an embodiment of the present application, seed question data is obtained, an expansion method is determined based on user feature information when using the question-answering model, and initial expanded question data is generated based on the expansion method and the seed question data. Target table question-answering data is determined based on the initial expanded question data and the requirements for the initial expanded question data to be used. The seed question data is data that enables the question-answering model in the large language model to determine the answer to the seed question data from the original table data. In an embodiment of the present application, an expansion method is determined based on user feature information, and with the help of the large language model and the expansion method, initial expanded question data close to the actual scenario can be automatically and quickly generated to obtain target table question-answering data that can optimize the large language model, thereby reducing labor costs based on improving the iteration speed of the large language model.
[0145] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0146] Based on the same inventive concept, embodiments of the present application also provide a table question and answer data construction device for implementing the table question and answer data construction method involved above. The implementation solution provided by this device is similar to the implementation solution described in the above method. Therefore, the specific limitations of one or more embodiments of the table question and answer data construction device provided below can be found in the above limitations of the table question and answer data construction method, and will not be repeated here.
[0147] In an exemplary embodiment, Figure 5 As shown, a table question and answer data construction device is provided, including: an acquisition module 11, a generation module 12 and a determination module 13, wherein:
[0148] An acquisition module 11 is configured to acquire seed question data; the seed question data is data that the question-answering model in the large language model can use to determine the answer to the seed question data from the original table data;
[0149] A generating module 12 is configured to determine an expansion method based on user characteristic information when using the question-answering model, and to generate initial expanded question data based on the expansion method and the seed question data;
[0150] The determination module 13 is configured to determine the target table question and answer data according to the initial extended question data and the requirements for use of the initial extended question data.
[0151] In one embodiment, the generating module 12 is specifically configured to expand the structure of the seed question data based on a structure expansion method to obtain first expanded question data having a structure similar to that of the seed question data; the initial expanded question data includes the first expanded question data.
[0152] In one embodiment, the generating module 12 is specifically configured to expand the semantics of the seed question data based on a semantic expansion method to obtain second expanded question data having the same semantics as the seed question data; the initial expanded question data includes the second expanded question data.
[0153] In one embodiment, the generating module 12 is specifically configured to randomly add erroneous characters to the seed question data based on a character expansion method to expand the seed question data and generate third extended question data; the initial extended question data includes the third extended question data.
[0154] In one embodiment, the generating module 12 is specifically configured to randomly add noise data to the seed question data based on a noise expansion method to expand the seed question data and generate fourth extended question data; the initial extended question data includes the fourth extended question data.
[0155] In one embodiment, the generating module 12 is specifically configured to expand the seed question data based on a depth expansion method to increase the question depth of the seed question data and generate fifth expanded question data; the initial expanded question data includes the fifth expanded question data.
[0156] In one embodiment, the determination module 13 is specifically used to deduplicate the initial extended question data to obtain intermediate extended question data; when the requirement for the initial extended question data to be used is to optimize the large language model, obtain the data information of the original table data; detect the correlation between the data information and the intermediate extended question data to determine the target extended question data that meets the correlation requirements; determine the target table question and answer data based on the target extended question data and the original table data.
[0157] Each module in the aforementioned form question-and-answer data construction device may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0158] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps of any of the above method embodiments when executing the computer program.
[0159] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above method embodiments are implemented.
[0160] In one embodiment, a computer program product is provided, comprising a computer program, which implements the steps of any of the above method embodiments when executed by a processor.
[0161] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0162] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.
[0163] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0164] The above-described embodiments are merely illustrative of several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.
Claims
1. A method for constructing tabular question-answer data, characterized in that: The method comprises: Obtain seed question data; the seed question data is data that the question-answering model in the large language model can determine the answer to the seed question data from the original table data; determining an expansion method based on user feature information when using the question-answering model, and generating initial expanded question data based on the expansion method and the seed question data; Target table question and answer data is determined according to the initial extended question data and the requirements for use of the initial extended question data.
2. The method according to claim 1, characterized in that The expansion method includes a structural expansion method; generating initial expanded question data according to the expansion method and the seed question data includes: The structure of the seed question data is expanded based on the structure expansion method to obtain first expanded question data with a structure similar to that of the seed question data; the initial expanded question data includes the first expanded question data.
3. The method according to claim 1, characterized in that The expansion method includes a semantic expansion method; generating initial expanded question data according to the expansion method and the seed question data includes: The semantics of the seed question data are expanded based on the semantic expansion method to obtain second expanded question data with the same semantics as the seed question data; the initial expanded question data includes the second expanded question data.
4. The method according to claim 1, wherein The expansion mode includes a character expansion mode; and generating initial expansion question data according to the expansion mode and the seed question data includes: Randomly adding erroneous characters to the seed question data based on the character extension method to extend the seed question data and generate third extended question data; the initial extended question data includes the third extended question data.
5. The method according to claim 1, wherein The expansion method includes a noise expansion method; generating initial expanded question data according to the expansion method and the seed question data includes: Noise data is randomly added to the seed question data based on the noise expansion method to expand the seed question data and generate fourth extended question data; the initial extended question data includes the fourth extended question data.
6. The method according to claim 1, characterized in that The expansion method includes a deep expansion method; generating initial expansion question data according to the expansion method and the seed question data includes: The seed question data is expanded based on the depth expansion method to increase the question depth of the seed question data, and fifth expanded question data is generated; the initial expanded question data includes the fifth expanded question data.
7. The method according to claim 1, characterized in that The step of determining target table question-and-answer data based on the initial extended question data and the requirements for use of the initial extended question data includes: Deduplication processing is performed on the initial extended question data to obtain intermediate extended question data; When the requirement to be used for the initial extended question data is to optimize the large language model, obtaining data information of the original table data; Detecting the correlation between the data information and the intermediate extended question data to determine target extended question data that meets the correlation requirement; The target table question and answer data is determined according to the target extended question data and the original table data.
8. A device for constructing table question and answer data, characterized in that: The device comprises: An acquisition module is used to acquire seed question data; the seed question data is data that the question-answering model in the large language model can determine the answer to the seed question data from the original table data; a generation module, configured to determine an expansion method based on user characteristic information when using the question-answering model, and generate initial expanded question data based on the expansion method and the seed question data; A determination module is used to determine target table question and answer data based on the initial extended question data and the requirements for use of the initial extended question data.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.