Method and device for improving fine tuning effect of large model based on iterative data enhancement strategy
Through iterative data augmentation strategy to generate more data samples in a limited environment, it solves the problem that traditional data augmentation methods are difficult to generate diverse and representative samples, and improves the training and generalization capabilities of the model.
Patent Information
- Application Number
- CN202510160819.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-13
AI Technical Summary
In a data-limited environment, traditional data augmentation methods are difficult to generate more representative and diverse samples, resulting in insufficient training and generalization capabilities of the model.
The iterative data enhancement strategy is adopted to generate new enhancement samples through multiple iterations. Each iteration combines the performance and error feedback of the model, and dynamically adjusts the data enhancement strategy to generate more targeted samples.
Through iterative data enhancement, more data samples can be generated when data is limited, the training sample size and generalization ability of the model can be improved, and the model can be dynamically optimized, which is especially suitable for the problem of uneven data distribution.
Smart Images

Figure CN119988979A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of large computer natural language models, and in particular to a method and device for improving the fine-tuning effect of large models based on an iterative data enhancement strategy. Background Art
[0002] In the field of modern machine learning and deep learning, the training effect of the model is highly dependent on the quantity and quality of data. Especially with the emergence of deep neural networks, a large amount of high-quality data has become particularly important because these models require massive amounts of data to capture complex features and patterns. However, in many practical applications, obtaining sufficient labeled data is not only expensive, but also time-consuming and even difficult to achieve in some areas.
[0003] To solve this problem, data augmentation has emerged as an effective technology that transforms existing data and generates new samples to increase the diversity and quantity of data. Traditional data augmentation methods mainly include simple operations such as image rotation, scaling, shearing, and flipping, which are used in the field of image processing; in the field of text processing, data augmentation can be achieved by replacing synonyms, deleting or adding words, etc. Although these methods can expand the amount of data and improve the generalization ability of the model to a certain extent, they may have limited effects in practical applications, especially when the original data samples are very limited, and more representative and diverse samples cannot be generated.
[0004] In order to further improve the effect of data augmentation, iterative data augmentation has gradually become a research hotspot. This method continuously generates new augmented samples through multiple iterations. Each iteration combines the performance of the model in the previous round of training and error feedback to dynamically adjust the data augmentation strategy. Through such a feedback mechanism, iterative data augmentation can not only provide more samples for the model, but also generate more targeted samples to help the model better learn complex features.
[0005] The core concept of iterative data enhancement is "adaptive enhancement" - that is, gradually adjusting the generated samples according to the needs and performance of the model, so that data enhancement is deeply coupled with the model training process. This method is particularly suitable for problems with uneven data distribution or scenarios where the model needs to capture subtle differences in high-dimensional feature space. For example, in the fields of medical imaging, autonomous driving, financial data, etc., due to the scarcity of data samples or the difficulty of labeling, iterative data enhancement can effectively improve the performance of the model.
[0006] In general, iterative data augmentation is not only a data expansion technology, but also a dynamic training strategy that adapts to the model. It continuously adjusts and optimizes the way data samples are generated, so that the model converges to better parameters faster, and finally obtains better generalization ability and robustness on the test data. This technology provides new solutions for scenarios with insufficient or unbalanced data, and also promotes the development of high-performance models. Summary of the invention
[0007] In order to solve the problems existing in the background technology, the present invention provides a method and device for improving the fine-tuning effect of a large model based on an iterative data enhancement strategy, which is used to improve the adaptability of the model by generating specific synthetic data in a low-data environment, reduce dependence on a large amount of labeled data, and improve the performance of a large language model (LLM).
[0008] According to one aspect of an embodiment of the present application, a method for improving the fine-tuning effect of a large model based on an iterative data enhancement strategy is provided, comprising the following steps:
[0009] Generate questions in various fields, generate corresponding answers based on the questions, label each piece of data, and obtain the original seed data;
[0010] Enhance the seed data to obtain enhanced data;
[0011] Checking the enhanced data to obtain checked data;
[0012] Use the checked data as a training set to fine-tune the model and obtain a trained model;
[0013] Perform inference on the trained model to obtain the inference result;
[0014] Conduct a comprehensive evaluation of the inference results;
[0015] Determine whether to continue training based on the evaluation results. If necessary, enhance the data with evaluation scores below the threshold, and return to the data enhancement step until the proportion of samples with evaluation scores below the threshold is lower than the preset ratio, and then end the training.
[0016] According to another aspect of an embodiment of the present application, a device for improving a large model fine-tuning effect based on an iterative data enhancement strategy is provided, comprising:
[0017] The seed data generation module is used to generate questions in various fields, generate corresponding answers based on the questions, label each piece of data, and obtain the original seed data;
[0018] A data enhancement module is used to enhance the seed data to obtain enhanced data;
[0019] A data checking module is used to check the enhanced data to obtain checked data;
[0020] The model fine-tuning module is used to use the inspected data as a training set to fine-tune the model and obtain a trained model;
[0021] Model inference module, used to infer the trained model and obtain the inference results;
[0022] Result evaluation module, used to comprehensively evaluate the reasoning results;
[0023] The iterative judgment module is used to determine whether training needs to continue based on the evaluation results. If necessary, the data with evaluation scores below the threshold will be enhanced and returned to the data enhancement module until the proportion of samples with evaluation scores below the threshold is lower than the preset proportion, and the training is terminated.
[0024] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method for improving the fine-tuning effect of a large model based on an iterative data enhancement strategy is implemented.
[0025] The beneficial effects of the present invention are:
[0026] Expand the amount of data. When data is limited, generate more data samples through iterative enhancement to increase the size of the model's training samples and make up for the lack of data.
[0027] Improve the generalization ability of the model. By continuously generating diverse data samples, the model can better cope with different variations in the data and improve its performance on real data sets.
[0028] Dynamically optimize the model. In each iteration, samples are generated in a targeted manner based on the model’s current weaknesses or error feedback, thus improving the model’s learning efficiency in each iteration.
[0029] To deal with imbalanced data, through iterative data enhancement, we can generate data samples of minority classes, balance the category distribution, and improve the model's prediction accuracy for minority classes. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 A flowchart of a method for improving large model fine-tuning effects based on an iterative data enhancement strategy provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0031] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0032] The present invention will be further described below in conjunction with the accompanying drawings and specific implementations.
[0033] like Figure 1 As shown, the basic concept of this application:
[0034] 1. Artificially generate a small amount of seed data in various fields;
[0035] 2. Perform data enhancement on the seed data by using customized rules (noun replacement, language conversion, etc.) and large model prompt words, and check the quality of the enhanced data;
[0036] 3. Fine-tune the model with the enhanced data; use the LLaMA-Factory framework to fine-tune the model, and use the latest Qwen2.5-7B as the base model.
[0037] 4. Use the fine-tuned model to infer the training dataset;
[0038] 5. Comprehensively evaluate the results of reasoning;
[0039] 6. Determine whether to end training based on the evaluation results; if the proportion of samples with comprehensive evaluation scores below the threshold is too high and exceeds 5%, the data with inference evaluation scores below the threshold will be enhanced, and the entire process will return to the second step for repeated iterations; if the proportion of samples with comprehensive evaluation scores below the threshold is low and is less than 5%, the training will be ended.
[0040] Specifically:
[0041] (1) Generate seed data
[0042] A small amount of high-quality seed data is manually generated for the application field of the model. These seed data cover multiple sub-fields to ensure the diversity and representativeness of the content. The seed data will serve as the basic sample set for subsequent data enhancement and model fine-tuning.
[0043] In an optional embodiment:
[0044] (1.1) Generate questions in various fields; for example, in the field of table question answering, questions related to the table can be asked based on the table. These questions can be answered by directly observing the table, or by writing code to solve the problem by looking at the table. The generated questions should be diverse and cover various fields and types of questions.
[0045] (1.2) Generate corresponding answers based on the questions;
[0046] (1.3) Label each piece of data, including domain, difficulty, etc., and obtain the original seed data D0.
[0047] (2) Data enhancement and quality check
[0048] The seed data D0 is enhanced, and the amount of data enhancement is determined according to the proportion of data in different fields and the difficulty of the data, and the enhanced data D is obtained. aug .
[0049] In an optional embodiment, data enhancement is divided into the following two methods:
[0050] 1. Enhance the question, but keep the answer the same as before, to ensure the semantics and content consistency between the enhanced data and the seed data;
[0051] 2. Use a large model to randomly enhance questions and answers. Both questions and answers will be enhanced to ensure that the enhanced questions and answers correspond to each other.
[0052] further:
[0053] 2.1 Enhance the problem, including:
[0054] 2.1.1 Noun replacement: By replacing key nouns in the seed data, the data content is adjusted to different expressions to increase data diversity while maintaining semantic consistency.
[0055] 2.1.2 Language conversion: Use multilingual translation technology to perform two-way or multi-directional language translation on the seed data, and then translate the translated content back into the source language to generate samples with similar semantics but different expressions.
[0056] 2.1.3 Random Noise Insertion: Insert a small amount of random noise into the data, such as minor spelling errors or misuse of punctuation, to enhance the robustness of the model, especially to perform better when faced with real data with more noise.
[0057] 2.1.4 Template filling: Use predefined templates to fill in sentences. For example, for classification tasks, you can generate sentences like "In XXX situation, how to deal with XXX". Through diversified template filling, the model can adapt to different context structures.
[0058] 2.1.5 Word order adjustment: Adjust the word order of a sentence without changing the original meaning of the sentence. This method is especially suitable for languages with flexible word order. This method can break the fixed structure of the sentence and make the model more robust to diverse expressions.
[0059] 2.2 Enhancements to questions and answers, including:
[0060] Big model prompts: Generate diverse samples with domain characteristics with the help of a big model, using prompts such as “generate more concise synonyms” or “generate extended content based on a given context”.
[0061] After generating new enhanced samples, the large model is used to check the quality of questions and answers of the enhanced data, and the data samples with poor quality are filtered out to obtain the checked data D check .
[0062] (3) Model fine-tuning
[0063] Use the checked data D check Use this as a training set to fine-tune the model.
[0064] Furthermore, this embodiment uses the LLaMA-Factory framework to fine-tune the model, selects the latest Qwen2.5-7B as the underlying model, and sets the learning rate of the trained model to 1e -6 Avoid overfitting. The distributed training function of this framework ensures the training efficiency and quality of the model. During the fine-tuning process, the performance of the model on the enhanced dataset is continuously monitored to achieve real-time optimization in the adjustment of model parameters.
[0065] (4) Model inference training data
[0066] The fine-tuned model will infer the training data set and output the prediction results. For all training data sets, different temperature and top_k parameters are set, and one sample is repeated for 5 times to obtain 5 different results. This process is used to verify the model's understanding and learning ability of enhanced data, and provide feedback data for subsequent evaluation and optimization.
[0067] (5) Comprehensive evaluation of the inference results
[0068] For code-based data (sql, python, etc.), the code part is extracted using regular expressions, and whether the code can be successfully executed is used as an evaluation indicator.
[0069] For other common data, the difference between the output result and the label is used as the evaluation indicator, including perplexity, the matching degree between the answer and the question, security and other aspects.
[0070] Then, the scores of each dimension are weighted according to a certain weight to obtain a comprehensive score. In this embodiment, since each sample is inferred 5 times, the average of the 5 comprehensive scores is taken as the final score of the sample to measure the performance of the model on the current data.
[0071] (6) Iterative judgment
[0072] The evaluation score is used to determine whether training needs to continue. If the proportion of samples with evaluation scores below the preset threshold exceeds 5%, it means that the model has not learned enough about certain features. At this time, the samples with reasoning errors are returned to the second step of data enhancement for re-enhancement, and the new enhanced samples are added to the training data set. The entire process is iterated from the second to the fifth step until the proportion of overall sample evaluation scores below the set threshold does not exceed 5%, and the training ends.
[0073] According to an embodiment of the present application, a device for improving the fine-tuning effect of a large model based on an iterative data enhancement strategy is also provided, including:
[0074] The seed data generation module is used to generate questions in various fields, generate corresponding answers based on the questions, label each piece of data, and obtain the original seed data;
[0075] A data enhancement module is used to enhance the seed data to obtain enhanced data;
[0076] A data checking module is used to check the enhanced data to obtain checked data;
[0077] The model fine-tuning module is used to use the inspected data as a training set to fine-tune the model and obtain a trained model;
[0078] Model inference module, used to infer the trained model and obtain the inference results;
[0079] Result evaluation module, used to comprehensively evaluate the reasoning results;
[0080] The iterative judgment module is used to determine whether training needs to continue based on the evaluation results. If necessary, the data with evaluation scores below the threshold will be enhanced and returned to the data enhancement module until the proportion of samples with evaluation scores below the threshold is lower than the preset proportion, and the training is terminated.
[0081] According to an embodiment of the present application, a computer-readable storage medium is also provided, on which a computer program is stored. When the program is executed by a processor, a method for improving the fine-tuning effect of a large model based on an iterative data enhancement strategy is implemented.
[0082] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0083] According to an embodiment of the present application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a method for improving the fine-tuning effect of a large model based on an iterative data enhancement strategy is implemented.
[0084] The electronic device may be any electronic device in the electronic device group. Optionally, in this embodiment, the electronic device may also be replaced by a terminal device such as a mobile terminal.
[0085] In summary, the core technology of this application is to perform iterative data enhancement on difficult samples, error-prone samples, and unbalanced samples during model training, and to perform targeted data enhancement on model weaknesses.
[0086] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for improving the fine-tuning effect of a large model based on an iterative data enhancement strategy, characterized in that: The following steps are involved: Generate questions in various fields, generate corresponding answers based on the questions, label each piece of data, and obtain the original seed data; Enhance the seed data to obtain enhanced data; Checking the enhanced data to obtain checked data; Use the checked data as a training set to fine-tune the model and obtain a trained model; Perform inference on the trained model to obtain the inference result; Conduct a comprehensive evaluation of the inference results; Determine whether to continue training based on the evaluation results. If necessary, enhance the data with evaluation scores below the threshold, and return to the data enhancement step until the proportion of samples with evaluation scores below the threshold is lower than the preset ratio, and then end the training.
2. The method according to claim 1, characterized in that The generating of questions in various fields includes raising relevant questions based on the table in the table question and answer field, and the question types include finding the answer directly by observing the table and solving the problem by looking at the table and writing code.
3. The method according to claim 1, characterized in that The step of enhancing the seed data includes determining the amount of data enhancement according to the proportion of data in different fields and the difficulty of the data.
4. The method according to claim 3, characterized in that The seed data enhancement is divided into enhancing the question alone and enhancing the question and the answer together.
5. The method according to claim 4, characterized in that The individual enhancements to the questions include noun replacement, language conversion, random noise insertion, template filling and word order adjustment.
6. The method according to claim 1, characterized in that The method of fine-tuning the model using the inspected data as a training set includes fine-tuning the model using the LLaMA-Factory framework and selecting Qwen2.5-7B as the underlying model.
7. The method according to claim 1 or 6, characterized in that: The reasoning of the trained model includes setting different temperature and top_k parameters, and repeating the reasoning of a sample multiple times to obtain multiple different results.
8. The method according to claim 1, characterized in that The comprehensive evaluation of the inference results includes extracting the code part using a regular method for code-related data, and using whether the code can be successfully executed as an evaluation indicator; for other general data, using the difference between the output result and the label as an evaluation indicator.
9. A device for improving the fine-tuning effect of a large model based on an iterative data enhancement strategy, characterized in that: include: The seed data generation module is used to generate questions in various fields, generate corresponding answers based on the questions, label each piece of data, and obtain the original seed data; A data enhancement module is used to enhance the seed data to obtain enhanced data; A data checking module is used to check the enhanced data to obtain checked data; The model fine-tuning module is used to use the inspected data as a training set to fine-tune the model and obtain a trained model; Model inference module, used to infer the trained model and obtain the inference results; Result evaluation module, used to comprehensively evaluate the reasoning results; The iterative judgment module is used to determine whether training needs to continue based on the evaluation results. If necessary, the data with evaluation scores below the threshold will be enhanced and returned to the data enhancement module until the proportion of samples with evaluation scores below the threshold is lower than the preset proportion, and the training is terminated.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, a method for improving the fine-tuning effect of a large model based on an iterative data enhancement strategy as described in any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Automatic scoring method and system based on comparison error sample mining and model fine tuning
CN121073406A
Automatic scoring method and system based on comparative error sample mining and model fine-tuning
CN121073406B