Structured data continuous pre-training method and device based on reference model
Through the structured data based on the reference model, the high-loss token is trained in a targeted manner, the problems of low efficiency and insufficient accuracy in the processing of tabular data are solved, and more efficient and accurate tabular data processing is achieved.
Patent Information
- Application Number
- CN202510237268.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-02
- Publication Date
- 2025-06-17
AI Technical Summary
Existing pre-trained models are difficult to effectively utilize the structural characteristics of tables when processing table data, resulting in low training efficiency and waste of computing resources, and it is difficult to balance efficiency and precision.
The structured data based on the reference model is used to continue pre-training method, and each token in the pre-trained data set is scored through the reference model. Only tokens with higher losses are trained in a targeted manner to optimize the performance of the base model in the structured data question and answer tasks.
It significantly improves the performance of the model on tabular data, avoids unnecessary computing overhead, improves the accuracy and speed of training, and maintains the universality of the model.
Smart Images

Figure CN120162586A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of large language models, and specifically relates to a method and device for continued pre-training of structured data based on a reference model. Background Art
[0002] With the rapid development of artificial intelligence technology, models based on natural language processing, especially large-scale pre-trained language models (LLMs), have achieved remarkable results in multiple fields. However, these models are usually trained on unstructured natural language data, and the row and column characteristics of tabular data make it significantly different from traditional natural language data. The characteristics of tabular data such as column structure, numerical relationships, and cell logical dependencies pose great challenges to existing pre-trained models.
[0003] Although existing continued pre-training methods can help models adapt to tabular data, these methods usually use a global loss function to train all tokens, and it is difficult to avoid inefficient training samples during the training process, which also leads to waste of computing resources. In addition, for tabular question answering tasks based on traditional methods, it is usually difficult to achieve a good balance between efficiency and accuracy. Therefore, it is particularly important to develop a continued pre-training method that can more efficiently and accurately focus on high-loss tokens.
[0004] Based on the above problems, the present invention proposes a large model continued pre-training technology for structured data based on a reference model, aiming to significantly improve the performance of the model on tabular data by selectively training high-loss tokens under the guidance of a better model, while avoiding unnecessary computational overhead. Summary of the Invention
[0005] The core idea of the present invention is to propose a method for continued pre-training of structured data (tabular data) based on a reference model (RSTM). By using the reference model to score the tokens in the pre-training dataset and only targeting the training of those tokens with higher losses, the performance of large-scale language models in structured data (tabular data) question answering tasks is improved.
[0006] According to one aspect of the embodiments of the present application, a method for continued pre-training of structured data based on a reference model is provided, including the following steps:
[0007] Obtain a reference model, where the reference model is an existing large-scale pre-trained model or a high-quality reference model trained according to domain data;
[0008] Use the reference model to perform inference and scoring on each token in the pre-training dataset, and calculate the reference loss LRM of each token;
[0009] Use the base model to perform inference and scoring on each token in the pre-training dataset, and calculate the base loss L of each token;
[0010] For each token in the pre-training dataset, calculate its difference loss ΔL, where the difference loss is the difference between the base loss L and the reference loss LRM;
[0011] According to the difference loss, select the tokens with higher losses, and the proportion of the selected tokens is the preset k%;
[0012] Apply a loss function to train on the selected high-loss tokens to optimize the performance of the base model in the structured data question-answering task.
[0013] According to another aspect of the embodiments of the present application, there is also provided a device for continuing pre-training of structured data based on a reference model, including:
[0014] A reference model acquisition module for acquiring a reference model, where the reference model is an existing large-scale pre-trained model or a high-quality reference model trained according to domain data;
[0015] A loss calculation module for using the reference model to perform inference and scoring on each token in the pre-training dataset, calculating the reference loss LRM of each token, and using the base model to perform inference and scoring on each token in the pre-training dataset, calculating the base loss L of each token;
[0016] A high-loss token selection module for calculating the difference loss ΔL of each token and selecting the tokens with higher losses according to the difference loss;
[0017] A training module for applying a loss function to train on the selected high-loss tokens to optimize the performance of the base model in the structured data question-answering task.
[0018] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for continuing pre-training of structured data based on a reference model.
[0019] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the method for continuing pre-training of structured data based on a reference model.
[0020] The beneficial effects of the present invention are:
[0021] Efficient training optimization: By calculating the loss of tokens offline and selectively training high-loss tokens, it focuses on the parts that contribute significantly to model improvement. This can effectively filter out the interference of irrelevant and dirty data, avoid unnecessary calculations, and improve the accuracy and speed of training.
[0022] Improve the ability to process tabular data: Compared with the traditional method of involving all data in training, the code for tabular analysis and processing is specifically optimized, enabling the model to pay more attention to detailed information when processing structured data, and improving the passing rate of code generation and the accuracy of answers in tabular Q&A tasks.
[0023] Without sacrificing generality: Through the reference model and selective training methods of the present invention, the general ability of the original large-scale pre-trained model is not affected, and the capabilities of the original model in other natural language tasks are maintained, while being significantly improved in tabular data Q&A tasks.
[0024] Flexible applicability: The technical solution of the present invention can be combined with any large-scale pre-trained model, with strong adaptability. Whether using open-source pre-trained models or custom domain-specific models, it can be flexibly applied.
[0025] Save computing resources: By reducing the training of low-loss tokens, the amount of calculation and memory occupancy are reduced. Especially in large-scale datasets and complex tasks, this method can significantly reduce the training time and consumption of computing resources, ensuring the maximum utilization of resources.
[0026] The present invention can significantly improve the application effect of large-scale language models in the field of structured data processing. Especially in tasks such as automated Q&A of tabular data, Python, and SQL generation, it can provide higher-precision and more efficient solutions. Brief Description of the Drawings
[0027] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments
[0028] In order to enable those skilled in the art of this technology to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0029] By selectively training those tokens with high losses, the training of irrelevant tokens and loss calculation are reduced, enabling the model to align with the capabilities of a larger model faster during training and obtain the essence information in the training set. Through this method, it converges faster than using all tokens (saving training time and resources), and at the same time, the model effect is better. It can be optimized for specific tabular data tasks without losing the general capabilities of the model, thereby improving the accuracy and efficiency of the model in tabular data question-answering tasks.
[0030] As Figure 1 shown, the structured data continued pre-training method (RSTM) based on a reference model provided in this application is implemented through the following steps:
[0031] S1. Obtain a reference model
[0032] In this step, first, a reference model needs to be selected to guide the model for targeted optimization during training. The reference model is used to evaluate the loss of each token in the pre-training dataset. The selection of the reference model is crucial for the training effect.
[0033] In one embodiment:
[0034] S1.1 Select a reference model: An existing open-source large-scale pre-trained model can be selected as the reference model. For example, in this sample, Qwen2.5-72B is selected as the reference model. The Qwen2.5-72B model is a large-scale pre-trained language model with strong general capabilities and rich context understanding capabilities. Through evaluation, it is found that the effect of this model on tabular data tasks is far better than the selected base model Qwen2.5-7B and is suitable for evaluating the token loss in tabular data tasks.
[0035] Furthermore, it can further include:
[0036] S1.2 Train the reference model: If there is no suitable reference model, a reference model can also be custom-trained on its own domain data.
[0037] Select a high-quality dataset related to the domain. This high-quality dataset is selected from the full dataset, has the same distribution as the full domain dataset, and at the same time, the data volume of the high-quality dataset is much smaller than that of the full dataset.
[0038] Suppose a high-quality dataset containing structured data (such as tables, SQL code, etc.) is used to train a Qwen2.5-7B as a reference model. The standard cross-entropy loss function is adopted to train the Qwen2.5-7B model on the high-quality dataset. After training, the Qwen2.5-7B model can effectively understand the characteristics of structured data and can significantly improve the ability to answer questions about tabular data and the accuracy of corresponding code writing.
[0039] S2. Calculate the loss of each token
[0040] In this step, the reference model (Qwen2.5-72B) is used to evaluate the loss of each token in the pre-training dataset. By calculating the loss value of each token, it is possible to identify which tokens are crucial for the task and which tokens have less impact on the task.
[0041] In one embodiment:
[0042] S2.1 Calculate the reference loss (LRM):
[0043] Use the reference model (Qwen2.5-72B) to evaluate each token in the pre-training dataset offline. For each token denoted as t i , use the reference model to calculate its predicted probability P(t i |data), and calculate the reference loss value (LRM):
[0044] LBM(t i ) = -log(P(t i |data))
[0045] This loss value reflects the degree of understanding of each token by the reference model and its importance in the current context.
[0046] S2.2 Calculate the base model loss (L):
[0047] Combined with the comparison in resources and actual use, Qwen2.5-7B is selected as the training base model. The Qwen2.5-7B model is a relatively small language model. Through comparison experiments, it is suitable as a basic training model for structured data both in Chinese performance and understanding ability.
[0048] Use the base model (Qwen2.5-7B) to evaluate each token in the pre-training dataset offline. For each token t i , calculate its predicted probability P(t i |data), and calculate the current loss value (L):
[0049] L(t i ) = -log(P(t i |data))
[0050] This loss value reflects the understanding degree of the current base model for each token and its importance in the current context.
[0051] S3. Select high-loss tokens
[0052] The purpose of this step is to mark the high-loss tokens in the training dataset according to the scoring of the reference model (i.e., the loss of the tokens). Subsequently, only those tokens with higher losses are trained, so as to quickly align with the reference model and improve the task performance of the model.
[0053] In one embodiment:
[0054] S3.1 Calculate the difference loss (LΔ): For each token t in the pre-training dataset i , calculate the difference loss (LΔ), that is, the difference between the loss of the current base model to be trained and the loss of the reference model:
[0055] LΔ(t i ) = L(t i ) - LRM(t i )
[0056] Where L(t i ) is the loss of the current training model for token t i , and LRM(t i ) is the loss of the reference model.
[0057] S3.2 Select tokens with higher losses: According to the calculated difference loss (LΔ), select tokens with higher losses for further training.
[0058] In this embodiment, a selection ratio k% is defined, and this ratio determines which tokens are to be selected for continued pre-training. Usually, k% can be adjusted in combination with the loss distribution, device resources, and experimental performance. It can be defined that k is 60, that is, select the tokens with the top 60% of the losses for training.
[0059] S3.3 Sort and select the top-k% tokens: Sort all tokens according to the difference loss LΔ, and select the top-k% tokens with higher difference losses for continued training. For example, we select the top 60% of the tokens with the largest difference losses for training, so as to ensure that the model focuses on learning those parts with greater learning value.
[0060] S4. Training the Structured Data Large Language Model
[0061] This step focuses on the tokens with higher losses in the training set to ensure that the model can accurately learn on tabular data.
[0062] In one embodiment:
[0063] S4.1 Optimize high-loss tokens: Apply the loss function only to those tokens with higher differential losses for training. During the training process, by only training those tokens that have a significant impact on the task and ignoring those tokens with lower losses and less impact on the task, the interference of irrelevant training is reduced.
[0064] S4.2 Training process: During the continued pre-training process, by only focusing on the tokens with higher losses, the Qwen2.5-7B model can converge quickly during the training process, improve the model's understanding ability of specific data, and approach the reference model. At the same time, since the low-loss tokens in the training set are ignored, computing resources are saved and the training efficiency is significantly improved.
[0065] According to an embodiment of the present application, there is also provided a device for continued pre-training of structured data based on a reference model, including:
[0066] A reference model acquisition module for acquiring a reference model, where the reference model is an existing large-scale pre-trained model or a high-quality reference model trained based on domain data;
[0067] A loss calculation module for using the reference model to infer and score each token in the pre-training dataset, calculating the reference loss LRM of each token, and using the base model to infer and score each token in the pre-training dataset, calculating the base loss L of each token;
[0068] A high-loss token selection module for calculating the differential loss ΔL of each token and selecting the tokens with higher losses according to the differential loss;
[0069] A training module for applying the loss function to train on the selected high-loss tokens to optimize the performance of the base model in the structured data question answering task.
[0070] According to an embodiment of the present application, there is also provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for continued pre-training of structured data based on a reference model.
[0071] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and this program can be stored in a computer-readable storage medium. The storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, etc.
[0072] According to an embodiment of the present application, there is also provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements a method for continuing pre-training of structured data based on a reference model.
[0073] This electronic device can be any one of the electronic devices in the electronic device group. Optionally, in this embodiment, the above electronic device can also be replaced with a terminal device such as a mobile terminal.
[0074] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A method for continuous pre-training of structured data based on a reference model, characterized in that: The following steps are involved: Obtain a reference model, where the reference model is an existing large-scale pre-trained model, or a high-quality reference model trained based on domain data; Use the reference model to infer and score each token in the pre-training dataset, and calculate the reference loss LRM of each token; Use the base model to infer and score each token in the pre-training dataset and calculate the base loss L for each token; For each token in the pre-training dataset, calculate its difference loss ΔL, which is the difference between the base loss L and the reference loss LRM; According to the difference loss, a token with a higher loss is selected, and the ratio of the selected token is a preset k%; The loss function is applied on the selected high-loss tokens for training to optimize the performance of the base model in the structured data question answering task.
2. The method for continuous pre-training of structured data based on a reference model according to claim 1, characterized in that: The reference model is the Qwen2.5-72B model, and the base model is the Qwen2.5-7B model.
3. The method for continuous pre-training of structured data based on a reference model according to claim 1, characterized in that: The reference model is obtained by training on a domain-related high-quality dataset, and the amount of data in the high-quality dataset is smaller than that of the full dataset.
4. A method for continued pre-training of structured data based on a reference model according to claim 1, 2 or 3, characterized in that: The reference loss LRM is calculated as follows: LRM(t i )=-log(P(t i |data)) where t i represents the i-th token, p() represents the conditional probability function, and data represents the data context.
5. A method for continued pre-training of structured data based on a reference model according to claim 1, 2 or 3, characterized in that: The base loss L is calculated as follows: L(t i )=-log(P(t i |data)) where t i represents the i-th token, p() represents the conditional probability function, and data represents the data set.
6. The method for continued pre-training of structured data based on a reference model according to claim 1, characterized in that: The selection ratio k% is adjusted according to loss distribution, equipment resources and experimental performance.
7. A structured data continued pre-training device based on a reference model, characterized in that: include: A reference model acquisition module is used to acquire a reference model, where the reference model is an existing large-scale pre-trained model or a high-quality reference model trained based on domain data; A loss calculation module, used to use the reference model to infer and score each token in the pre-training dataset, calculate the reference loss LRM of each token, and use the base model to infer and score each token in the pre-training dataset, and calculate the base loss L of each token; A high-loss token selection module is used to calculate the difference loss ΔL of each token and select a token with a higher loss according to the difference loss; A training module is used to apply a loss function on selected high-loss tokens for training to optimize the performance of the base model in the structured data question answering task.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, a method for continued pre-training of structured data based on a reference model as described in any one of claims 1 to 6 is implemented.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method for continuing pre-training of structured data based on a reference model as described in any one of claims 1 to 6 is implemented.