Large model training method and system for table question and answer tasks

By introducing supervised fine-tuning, mirror inversion training and reward feedback optimization fine-tuning training in large language models, we optimize the table question and answer tasks, which solves the shortcomings in accuracy and data use of existing models, and achieves more efficient and accurate table data processing.

CN119990321APending Publication Date: 2025-05-13INSTITUTE OF COMPUTING INNOVATION ZHEJIANG UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510120065.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-25
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing large language models have low accuracy in the field of table question and answers, insufficient training data, and lack targeted optimization training methods, resulting in poor performance when dealing with multi-dimensional information fusion and accurate answers.

Method used

A large-model optimization training method for table Q&A tasks is proposed, including supervised fine-tuning, mirror inversion training and reward feedback optimization fine-tuning training. By building a table Q&A reward evaluation system, designing a comprehensive reward evaluation mechanism based on the code execution situation, and optimizing the executability and accuracy of the code generated by the model.

Benefits of technology

It significantly improves the accuracy and efficiency of the table Q&A big model, reduces manual intervention, improves the efficiency of automated training, and achieves more efficient and accurate table data analysis and query processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990321A_ABST
    Figure CN119990321A_ABST
Patent Text Reader

Abstract

The invention discloses a training method and system for a table question and answer large model. The table question and answer task refers to that according to provided table data such as csv files, excel files and database db data, table-related problems such as data query, data statistical analysis and visualization are proposed for the table data, and the question and answer task of answers can be executed through Python or SQL codes. The invention provides a big language model enhanced training method for a table question and answer task in combination with the characteristics of the table question and answer field, on the basis of an existing big language model, a special data set related to the table question and answer task is constructed, and a coincidence reward feedback system combined with table question and answer is designed; and in combination with a reinforcement learning training strategy of the mirror image model, the table data question and answer ability of the large language model is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology and relates to a large model training method and system for form question answering tasks. Background Art

[0002] With the rapid development of large language models, the emergence of models such as ChatGPT, Claude, LLaMA and Qwen has brought huge industry opportunities to the field of AI applications. The table question answering task aims to generate solutions to problems and corresponding data query or data analysis codes based on the input table data and user questions, covering NL2SQL and NL2PYTHON tasks, that is, generating SQL or Python code through questions to query and analyze data. In industries such as finance, medical care, and supply chain management that require extremely high data analysis accuracy, advanced table question answering systems will become a key tool to improve efficiency and reduce decision-making risks.

[0003] However, the development of large language models in the field of tabular question answering still faces many challenges, especially in processing multi-dimensional information fusion and accurate answers. Therefore, achieving deep support and efficient application of large language models for tabular data and improving the accuracy of output results are key issues in promoting the further development of AI in the field of tabular data question answering.

[0004] In the prior art, CN118886519A discloses a large model training method, including a three-stage process of data pre-training, fine-tuning and preference optimization. Among them, the method pre-trains the initial language processing model through the original data set, then uses the extended data set for fine-tuning, and finally uses the preference data set to adjust the preference based on reward feedback. However, this method is only applicable to general large language model training methods, and fails to optimize the training process in a targeted manner in combination with task scenarios (tabular data question-and-answer tasks). The preference alignment stage relies on manually constructed data sets, cannot automatically iterate, and requires manual intervention;

[0005] In addition, CN117972070A discloses an optimization method for large-model table question answering, which improves the ability to understand table semantics through an automated corpus generation scheme based on template design and rule formulation, and uses technologies such as Lora fine-tuning and prompt learning to drive specific tasks of large models under limited resource conditions. However, this method mainly focuses on data set construction and partial parameter fine-tuning. The training method described in this patent is also a conventional large language model training method. The innovations are mainly concentrated in table corpus generation and serialized semantic analysis, and there is a lack of further breakthroughs in the training method of the large language model itself. Summary of the invention

[0006] The invention aims to solve the deficiencies in the prior art, mainly targeting the low accuracy and insufficient training data of the current table question and answer big model system, and proposes a big model optimization training method and system specifically for table question and answer. The table question and answer big model trained by the present invention has significant improvements in answer accuracy and code executability, while reducing manual intervention and significantly improving the efficiency of automated training, thereby achieving more efficient and accurate table data analysis and query processing.

[0007] A large model training method for a table question answering task in the present invention comprises the following steps:

[0008] Supervised fine-tuning training from the basic large model to the tabular large model, including collecting tabular data and related task instructions from multiple fields, building a supervised fine-tuning dataset, and fine-tuning the basic large language model to generate a structured large language model enhanced by the tabular task;

[0009] Supervised fine-tuning training of the mirror-reversal large model, by reversing the input and output of the supervised fine-tuning dataset, training the mirror-reversal large model to improve the model's adaptability to diverse inputs;

[0010] Build a table question-answering reward evaluation system, design a comprehensive reward evaluation mechanism based on the code execution situation, parse and execute the code generated by the model, and give different scores based on the execution results;

[0011] Reward feedback optimization fine-tuning training constructs seed input data, uses a structured large language model to generate output, and then uses the mirror-reversed large model to sample new input data, construct an Input-Output-Pairs dataset, and use a reward feedback mechanism to evaluate and optimize the model.

[0012] A large model training system for table question answering tasks in the present invention includes:

[0013] Data collection module, used to collect tabular data and related task instructions in multiple fields;

[0014] The supervised fine-tuning module is used to construct a supervised fine-tuning dataset and perform supervised fine-tuning on the basic large language model;

[0015] The mirror reversal training module is used to train the mirror reversal model by reversing input and output;

[0016] The reward evaluation module is used to parse and execute the code generated by the model and give scores based on the execution results;

[0017] The optimization and fine-tuning module is used to evaluate and optimize the model through a reward feedback mechanism.

[0018] The beneficial effects of the present invention are:

[0019] Improve the accuracy and efficiency of large-scale table question answering models: Through supervised fine-tuning, mirror reversal training, and reinforcement learning algorithms, significantly improve the accuracy and automation efficiency of table question answering, data analysis, and data verification tasks, and reduce reliance on human intervention.

[0020] Solve the problem of insufficient training data: Use prompt generation and mirror reversal technology to automatically generate diverse data samples, expand the coverage of the data set, and ensure the wide applicability of the model in different fields and scenarios.

[0021] Enhance code generation and execution capabilities: Integrate Python and SQL task data, optimize the executability and accuracy of code generation through a reward evaluation mechanism, and ensure that the code output in data processing and query tasks is correct.

[0022] Improve the efficiency of automated iterative optimization: Through the reward feedback optimization process and the new input data generated by mirror reversal, the model can continuously improve itself, greatly improving the processing power and output effect of table question answering tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0024] Figure 1 This is the training process of the table question-answering large model of the present invention. DETAILED DESCRIPTION

[0025] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0026] like Figure 1 As shown, the embodiment of the present application provides a training method for a table question answering large model, comprising the following steps:

[0027] Step 1: Supervised fine-tuning training from the basic large model to the tabular large model

[0028] First, during the data collection and preparation stage, data from multiple fields are widely collected to ensure data diversity and coverage.

[0029] In some embodiments, the data sources include common tabular data types in industries such as finance, medical care, supply chain management, and e-commerce. These tabular data have various structures and formats, such as CSV files, Excel tables, database tables, and JSON format files.

[0030] In order to further improve the model's ability to process tabular data, this application uses the prompt generation capability of the large language model to automatically generate instruction data for a variety of table-related tasks. These task instructions include but are not limited to: table questions and answers (such as extracting quarterly sales from a sales data table), table fact verification (such as verifying whether inventory data exceeds a certain threshold), table reordering (such as sorting data by date or sales), table data completion (filling missing values ​​based on existing information), multi-table connection relationship extraction (integrating data from multiple tables, such as connecting customer and order information) and table format conversion (converting table data from Excel format to JSON format).

[0031] In order to enhance the model's ability in code generation, this application also collects a large amount of instruction data related to Python and SQL. For example, Python data processing tasks may involve data cleaning and visualization, such as using df['total_sales'] = df['units_sold'] * df['price_per_unit'] for data calculation; SQL query tasks include extracting specific information from the database, such as SELECT SUM(revenue) FROM sales_data WHERE year = 2023.

[0032] In addition, some task instructions related to mathematics and logical reasoning are collected to improve the model's capabilities in data calculation and logical reasoning, such as calculating inventory cycles or making sales forecasts. In order to improve the model's performance in multi-round dialogues and contextual understanding, this application also generates multi-round interaction instruction data to train the model's contextual understanding capabilities in continuous dialogues. For example, a user may ask multiple related questions in succession, and the model needs to be able to remember and reasonably adjust previous answers when answering. At the same time, some general dialogue instructions are collected to ensure that the model can also demonstrate good natural language understanding capabilities in free dialogue scenarios in open fields.

[0033] Next, by integrating the above collected data in a certain proportion, a supervised fine-tuning dataset (SFT-Train-Data) is constructed. For example, the composition of the dataset may include: 60% table tasks (such as table question answering, data verification, reordering, data completion), 20% programming tasks (such as Python and SQL code generation), 10% mathematical and logical reasoning tasks, and 10% multi-round interaction and general dialogue tasks. This combination can ensure a balanced improvement in the model's capabilities in multiple aspects such as table data processing, programming generation, reasoning analysis, and dialogue understanding.

[0034] In the supervised fine-tuning stage, the basic large language model (such as QWEN or LLaMA) is fine-tuned using the constructed SFT-Train-Data dataset. During the training process, the weights of the model are adjusted through supervised learning so that it can better understand the semantics of tabular data and generate corresponding code and analysis results. After fine-tuning, the generated structured large language model enhanced by table tasks (LLM-Table-SFT) shows higher accuracy and efficiency in tabular data processing, SQL / Python code generation, and complex data query tasks.

[0035] Step 2: Supervised fine-tuning training of the mirror-flipped large model

[0036] This step aims to build a model that can automatically generate and expand input data by training the large mirror-reversed model (LLM-Table-Reverse-SFT) to improve its adaptability to diverse inputs.

[0037] The SFT-Train-Data dataset generated in step 1 is processed by reversing its input and output. This means that the output originally used to answer the question is used as the new input, and the question itself is used as the new output. For example, suppose the original dataset contains the following task instructions:

[0038] Input: What is the sum of the sales in rows 2 and 4 of the table?

[0039] Output: The total sales amount is $3,000.

[0040] In the mirror-reversed dataset, this instruction would be reversed to:

[0041] Input (after reversal): The sum of sales is $3,000.

[0042] Output (after reversal): What is the sum of the sales in rows 2 and 4 of the table?

[0043] Through such a reversal operation, a new training data set is constructed to train a large mirror reversal model (LLM-Table-Reverse-SFT). The goal of this model is to predict possible inputs based on the outputs, so that it can "reverse reasoning" the source of the problem. This process helps the model understand the intrinsic connection between the data output and the original problem. In this reversal process, it allows the generation of diverse problem statements without requiring strict matching of input and output. By introducing a certain amount of randomness and generalization strategies, the large mirror reversal model can generate a variety of possible problems to expand the diversity of data samples.

[0044] Based on the reversed dataset, a special mirror reversal model (LLM-Table-Reverse-SFT) is trained, whose core task is to infer possible inputs based on the outputs. The training of this model does not require the generated questions to be exactly the same as the original inputs, but focuses on whether the generated inputs are semantically reasonable. The trained LLM-Table-Reverse-SFT can be used to automatically generate more diverse sample inputs (Input), as an important part of the self-feedback optimization closed-loop training.

[0045] Step 3: Build a table question and answer reward evaluation system

[0046] When designing the reward model, traditional large language models are mainly evaluated for general question-answering scenarios, so they perform poorly when handling tabular question-answering tasks. In particular, the code output generated in tabular question-answering tasks often contains complex grammar and logic requirements, and the existing reward model is difficult to accurately capture the subtle errors, resulting in inaccurate scoring.

[0047] Therefore, based on the general question-answering reward model, this application designs a comprehensive reward evaluation mechanism in combination with the code execution situation, so that the scoring of the model output is more accurate.

[0048] In some embodiments, first, the code generated by the model is parsed, and the code is actually executed by the code executor. Different scores are given according to different execution results. The final reward score is the result of the comprehensive RM score and the code execution score.

[0049] Furthermore, based on the actual performance of code execution, the code execution is divided into several levels and scored respectively. The specific scoring rules are as follows:

[0050] Code execution Score Not executable 0 Executable, with warning message 1 Executable, no warning message, execution result is empty 2 Executable, no warning message, execution result is not empty 3

[0051] Final reward score calculation formula:

[0052]

[0053] Where α is 0.5, σ is the sigmoid function, rm represents the model output score, and exec_score is the execution score, which is between 0 and 3.

[0054] Step 4: Reward feedback optimization fine-tuning training

[0055] After the supervised fine-tuning training in step 1, although the LLM-Table-SFT model has shown significant improvement in structured table tasks, there are still some problems due to the uneven distribution of training data, poor quality of some data, and lack of sufficient complex and difficult samples. For example, the generated code may not be executed correctly, the code output may be incomplete, key steps may be omitted, or the generated answers may not be comprehensive and accurate enough. Therefore, it is necessary to further optimize the model through a reward-feedback-based reinforcement learning algorithm (such as PPO or GRPO) to improve the accuracy, completeness and reliability of its output. Specifically:

[0056] Build input seed data (Input-Seed-Data): Build a batch of small-scale input seed data sets (Input-Seed-Data), which do not require pre-labeled output results. When building, ensure the diversity of input data to cover as many problem scenarios as possible. The purpose of these seed input data is to simulate a variety of questions that users may ask in real application scenarios without manually providing corresponding answers.

[0057] Generate output using LLM-Table-SFT: Input the above-constructed Input-Seed-Data in batches into the LLM-Table-SFT model that has been trained in step 1 to generate preliminary outputs. The goal of this stage is to allow the LLM-Table-SFT model to generate the most logical and semantic preliminary answer based on the input.

[0058] Sampling new input data for the mirror-reversal large model: The Outputs generated by the LLM-Table-SFT model are input into the LLM-Table-Reverse-SFT mirror-reversal model trained in step 2 to generate new input data (Inputs). This process is equivalent to "reverse reasoning". Through this mirror-reversal method, more diverse input questions are generated to increase the breadth of the model's training data.

[0059] Construct Input-Output-Pairs dataset: After the above steps, the initial Input-Seed-Data is combined with the Inputs generated by mirror reversal to obtain a batch of new Input-Output-Pairs datasets. These datasets contain the answers generated by the model and the input data generated by reverse.

[0060] Next, the generated Input-Output-Pairs dataset is evaluated using a reward feedback mechanism. A reward score is assigned to each sample using the reward evaluation method defined in step 3. Based on the evaluation score results, a reinforcement learning algorithm (such as PPO or GRPO) is used to adjust and optimize the parameters of the LLM-Table-SFT model. The goal of reward feedback is to encourage the model to generate higher-scoring outputs, i.e., answers with executable code, no warnings, and accurate results.

[0061] The updated model is applied again to the data generated by the new round of sampling, and new input data is continuously generated through the mirror-inverted large model for cyclic evaluation and reinforcement learning. Specifically, a batch of newly sampled Inputs data will be fed into the latest optimized LLM-Table-SFT model to generate new Outputs, and mirror-inversion will be used again to generate more input data. In each round of iteration, the model parameters are continuously updated according to the reward feedback mechanism. Through this continuous automated iterative optimization process, the model gradually improves its understanding ability and output effect of the table question-answering task, and ultimately generates more accurate, complete and reliable answers. This process does not require a lot of manual intervention, which greatly improves training efficiency and model performance.

[0062] Based on the same concept as the above method, this application also proposes a large model training system for table question answering tasks, including:

[0063] Data collection module, used to collect tabular data and related task instructions in multiple fields;

[0064] The supervised fine-tuning module is used to construct a supervised fine-tuning dataset and perform supervised fine-tuning on the basic large language model;

[0065] The mirror reversal training module is used to train the mirror reversal model by reversing input and output;

[0066] The reward evaluation module is used to parse and execute the code generated by the model and give scores based on the execution results;

[0067] The optimization and fine-tuning module is used to evaluate and optimize the model through a reward feedback mechanism.

[0068] Furthermore, reinforcement learning (RL) is a machine learning method that focuses on allowing the agent to learn how to maximize its long-term reward by interacting with the environment. In this framework, the agent makes a decision at each step, and the environment gives feedback based on the decision and advances the system to a new state. The goal of the agent is to learn a policy that maximizes its accumulated reward throughout the task.

[0069] Reinforcement learning is particularly useful in some scenarios, such as games, robot control, recommendation systems, and text generation in natural language processing. In reinforcement learning, commonly used algorithms include policy optimization algorithms, of which PPO and GRPO are two common policy optimization methods. Specifically:

[0070] PPO (Proximal Policy Optimization) is a deep reinforcement learning algorithm used to solve the instability and high variance problems in traditional policy gradient methods. PPO makes the model more stable and efficient during training by limiting the amplitude of policy updates. The core idea of ​​PPO is to constrain the policy when updating it to prevent instability caused by too fast policy updates. PPO introduces the Clipped Probability Ratio mechanism to control the update amplitude of the policy. The PPO algorithm has the following advantages: by limiting the amplitude of policy updates, the instability of training is reduced; there is no need to calculate the second-order derivative, which is more efficient than other policy optimization algorithms such as TRPO (Trust Region Policy Optimization); due to its simplicity and efficiency, PPO is widely used in reinforcement learning tasks such as game AI, autonomous driving, text generation, etc.

[0071] GRPO (Generalized Reward Policy Optimization) is an emerging reinforcement learning algorithm that aims to expand the traditional PPO method to adapt to more complex reward mechanisms. GRPO introduces the concept of generalized rewards and handles different types of reward signals more flexibly, thereby improving performance in specific tasks.

[0072] GRPO is an extension of PPO, which allows it to handle reward signals from a variety of different sources. GRPO's policy update is also based on probability ratios, but a wider range of indicators are introduced in the reward function to ensure that the output generated by the model reaches a higher standard in multiple dimensions. GRPO is able to handle multiple different reward signals, so it performs well in tasks that require optimizing multiple objectives at the same time.

[0073] Compared with the existing technology, this application introduces a multi-level evaluation system that combines scenario-based code execution feedback with a general reward mechanism, innovatively adopts the input data sampling method of the mirror-inverted large model, and combines it with an automatic iterative optimization process for continuous training and adjustment. This improvement not only improves the generalization ability of the model, but also significantly improves the accuracy and robustness of the output, making it more effective in dealing with complex data query and analysis tasks. In addition, this method reduces manual intervention through an automated feedback iteration mechanism, thereby achieving a better balance between efficiency and accuracy.

[0074] Embodiment: This embodiment proposes a complete training process, which aims to enable the large language model (LLM) to better understand and process various types of tabular data and generate corresponding analysis results and codes.

[0075] (1) Fine-tuning model training system based on table question answering tasks

[0076] Collection and preparation of data sets. During the data preparation phase, various types of tabular data are widely collected from multiple industries (including finance, healthcare, supply chain management, and e-commerce). These data exist in different formats, including CSV files, Excel tables, database tables, and files in JSON format. These tabular data may include sales data, inventory records, financial statements, etc. In order for the model to understand these tabular data, this embodiment automatically generates a series of instruction tasks. These task examples include:

[0077] Table questions: For example, "What will be the sales in the third quarter of 2023?".

[0078] Data validation: Verify whether a piece of data meets a certain threshold, such as "Is the inventory less than 100 pieces?".

[0079] Data sorting: such as "sort the data by sales from high to low".

[0080] Data completion: such as "fill in the missing date records in the table".

[0081] Multi-table join: Integrate multiple tables, such as "join customer information with order information to generate a complete report."

[0082] To further enhance the model’s capabilities in code generation and processing, task data related to Python and SQL are also collected.

[0083] Supervised fine-tuning training integrates the collected data into a supervised fine-tuning dataset (SFT-Train-Data). The dataset is composed of 60% table tasks, 20% programming tasks (Python and SQL), 10% mathematics and logical reasoning tasks, and 10% multi-round interactive dialogue tasks. This combination ensures that the model performs well in table question answering, code generation, logical reasoning, and dialogue understanding. In the fine-tuning stage, the pre-trained QWEN or LLaMA language model is used as the basis for training through supervised learning methods. The training process includes inputting questions and expected answers, and adjusting parameters through the model to better understand the requirements of table data and related tasks. After fine-tuning, the generated model LLM-Table-SFT can significantly improve in processing table data, generating code, and answering complex questions.

[0084] (2) Training of the mirror-reversal large model

[0085] In order to enhance the adaptability of the model in different scenarios, this embodiment further optimizes the model by means of input-output reversal. The core idea of ​​the mirror reversal training is to achieve the purpose of automatically generating samples for the input data by reversing the input and output.

[0086] Data reversal processing, assuming that there are the following tasks in the previous supervised fine-tuning dataset:

[0087] Input: What is the sum of the sales in rows 2 and 4 of the table?

[0088] Output: The total sales amount is $3,000.

[0089] In the mirror-reversal dataset, the above task will be adjusted as follows:

[0090] Input (after reversing): Sum of sales is $3,000.

[0091] Output (after reversing): What is the sum of the sales in rows 2 and 4 of the table?

[0092] Through this reverse training, the model can not only answer questions, but also deduce questions from known answers. This ability is very useful in practical applications, such as helping users generate more diverse question descriptions. The reverse model (LLM-Table-Reverse-SFT) can also generate diverse questions through "reverse reasoning" to expand the coverage and diversity of training data.

[0093] (3) Construction of the table question and answer reward evaluation system

[0094] In order to evaluate the performance of the model on the tabular task, a reward evaluation mechanism is designed to quantify the accuracy of the model output and the executableness of the code generation. This helps to further optimize the performance of the model in actual tasks.

[0095] Code generation and execution evaluation, when the model generates code, it is parsed and actually run using the code parser and executor. Based on the execution results, this embodiment designs the following scoring rules:

[0096] Not executable: Score 0.

[0097] Executable with warnings: score 1.

[0098] Executable, no warnings, but empty results: score 2.

[0099] Executable, no warning messages, and the result is not empty: the score is 3.

[0100] Comprehensive score calculation, this embodiment uses the following formula to calculate the final reward score:

[0101] Reward=0.5×Sigmoid(RM+Exec)

[0102] Among them, RM is the semantic score of the model output, and Exec is the code execution score. By comprehensively considering the correctness and executability of the model output, this embodiment ensures that the model can not only give the correct answer, but also generate code that can actually run.

[0103] (4) Reward feedback optimization fine-tuning training

[0104] After the above training phase is completed, the model has a certain ability to process tables. However, in order to further improve its accuracy and output quality, reinforcement learning technology is introduced for optimization.

[0105] Seed Dataset Construction and Model Output Generation: First, a batch of seed input data (Input-Seed-Data) was constructed, which does not require manual annotation output. Its purpose is to simulate various questions that real users may ask. Then, this data is input into the fine-tuned LLM-Table-SFT model to generate preliminary answers. Next, the mirror reversal model is used to generate new input questions to expand the diversity of the dataset.

[0106] Reinforcement learning and reward feedback optimization. In the reinforcement learning stage, the Proximal Policy Optimization (PPO) algorithm is used to further optimize the model. Through the reward feedback mechanism, the generated input-output pairs are scored and the model parameters are adjusted to improve the model's answer accuracy and code generation capabilities. The entire process is an automated closed-loop system: by continuously generating new questions and new answers, the model's learning ability is gradually improved, and ultimately more accurate form question answering and data processing are achieved.

[0107] Through the above implementation process, this embodiment greatly improves the efficiency and accuracy of the table question-and-answer large model in processing actual tasks. For example, in data analysis in the financial industry, the use of the optimized model can accurately answer complex financial questions, such as "What was the net profit growth rate last quarter?" and generate the corresponding Python code to calculate the results. In addition, in the inventory management of e-commerce platforms, the model can automatically predict inventory demand and generate SQL queries to support decision-making. These capabilities significantly lower the threshold for data query and analysis and improve the data processing efficiency of enterprises.

[0108] This implementation improves the reward evaluation mechanism, combines scenario-based code execution feedback and a general reward system, and innovatively introduces a mirror-reversal large model sampling method for input data, an automated feedback iteration mechanism, and reduces manual intervention, thereby achieving a better balance between efficiency and accuracy, and improving the model training efficiency and result accuracy of the table question answering task. Through the optimization of the above technologies, this embodiment aims to promote the depth and breadth of the application of large language models in the field of table question answering, and provide more accurate and efficient solutions for data analysis and decision-making in various industries.

[0109] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A large model training method for table question answering tasks, characterized in that: The following steps are involved: Supervised fine-tuning training from the basic large model to the tabular large model, including collecting tabular data and related task instructions from multiple fields, building a supervised fine-tuning dataset, and fine-tuning the basic large language model to generate a structured large language model enhanced by the tabular task; Supervised fine-tuning training of the mirror-reversal large model, by reversing the input and output of the supervised fine-tuning dataset, training the mirror-reversal large model to improve the model's adaptability to diverse inputs; Build a table question-answering reward evaluation system, design a comprehensive reward evaluation mechanism based on the code execution situation, parse and execute the code generated by the model, and give different scores based on the execution results; Reward feedback optimization fine-tuning training constructs seed input data, uses a structured large language model to generate output, and then uses the mirror-reversed large model to sample new input data, construct an Input-Output-Pairs dataset, and use a reward feedback mechanism to evaluate and optimize the model.

2. The large model training method according to claim 1, characterized in that: In the supervised fine-tuning training step, data sources include common tabular data types in the finance, medical, supply chain management, and e-commerce industries, as well as instruction data related to Python and SQL.

3. The large model training method according to claim 1 or 2, characterized in that: In the supervised fine-tuning training step of the mirror reversal large model, a new training data set is constructed by reversing the input and output, and the mirror reversal large model LLM-Table-Reverse-SFT is trained to generate more diverse input problems.

4. The large model training method according to claim 3, characterized in that: In the step of constructing the tabular question-and-answer reward evaluation system, the code execution is divided into several levels according to the actual performance of the code execution, and each level is scored separately. The final reward score is the result of the comprehensive model output score and the code execution score.

5. The large model training method according to claim 1, characterized in that: In the reward feedback optimization fine-tuning training step, the model is optimized using the PPO or GRPO algorithm to improve the accuracy, completeness and reliability of the model output.

6. A large model training system for table question answering tasks, characterized in that: include: Data collection module, used to collect tabular data and related task instructions in multiple fields; The supervised fine-tuning module is used to construct a supervised fine-tuning dataset and perform supervised fine-tuning on the basic large language model; The mirror reversal training module is used to train the mirror reversal model by reversing input and output; The reward evaluation module is used to parse and execute the code generated by the model and give scores based on the execution results; The optimization and fine-tuning module is used to evaluate and optimize the model through a reward feedback mechanism.

7. The large model training system according to claim 6, characterized in that: The data collected by the data collection module includes common table data types in the finance, medical, supply chain management, and e-commerce industries, as well as instruction data related to Python and SQL.

8. The large model training system according to claim 6 or 7, characterized in that: The mirror reversal training module builds a new training data set by reversing input and output, and trains the mirror reversal large model to generate more diverse input problems.

9. The large model training system according to claim 8, characterized in that: The reward evaluation module divides the code execution into several levels according to the actual performance of the code execution and scores them respectively. The final reward score is the result of the comprehensive model output score and the code execution score.

10. The large model training system according to claim 6, characterized in that: The optimization and fine-tuning module uses the PPO or GRPO algorithm to optimize the model to improve the accuracy, completeness and reliability of the model output.

Citation Information

Cited By

  • Large language model data set construction method and large language model enhancement method

    CN120653734A

  • Table question and answer task capability enhancement processing method of large language model

    CN121598963A

  • A table question and answer task capability enhancement processing method of a large language model

    CN121598963B