Automatic fine-tuning data screening and correcting method in large language model field

Through automated data screening and correction methods, confidence estimation and LLM are used to generate alternative responses, which solves the problems of inconsistent data quality and high cost in existing technologies and improves the performance of large language models in specific fields.

CN120632309APending Publication Date: 2025-09-12GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510936269.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In the field of large language models, existing technologies rely on manual or advanced LLMs for data quality screening and correction, resulting in low efficiency, high cost, and difficulty in ensuring consistency and accuracy. This makes it difficult to generate high-quality data in specific fields, affecting model performance.

Method used

The BSDetector tool is used to automatically filter low-quality data through confidence estimation, and the fine-tuned LLM is used to generate alternative responses. The confidence of the basic LLM is combined for automatic correction to form a high-quality data set, and the model performance is improved through iterative optimization.

Benefits of technology

It achieves efficient and accurate data screening and correction, improves the quality of data sets, enhances the adaptability and performance of models in specific fields, and reduces computing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632309A_ABST
    Figure CN120632309A_ABST
Patent Text Reader

Abstract

The invention provides an automatic fine-tuning data screening and correcting method in the field of large language models, and belongs to the field of artificial intelligence. Comprising the following steps: an automatic filtering stage is used for identifying and removing low-quality data pairs so as to ensure that a data set used in a subsequent fine tuning process is as high as possible in quality; the auto-correction phase is aimed at auto-correcting those data pairs identified as low quality but likely to generate a better response by LLM. And repeating the process, namely performing LLM fine tuning by using the updated data set again to form a continuously optimized cycle until a satisfactory performance level is achieved. According to the method, the overall quality of the data set is remarkably improved. Not only is the problem of learning sample reduction caused by direct discarding of data reduced, but also valuable information is reserved, and the adaptability and performance of the model are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention provides a method for automatically screening and correcting fine-tuning data in the field of a large language model, belonging to the field of artificial intelligence. Background Art

[0002] Large language models (LLMs) have become a core technology in the field of artificial intelligence (AI) due to their outstanding performance in natural language generation tasks. Pre-training on massive amounts of text allows LLMs to capture complex language patterns, but their performance in specific domains or specialized tasks requires further optimization through instruction fine-tuning. The core of instruction fine-tuning is to use high-quality paired datasets (input instructions, target responses) to supervise the pre-trained model to enhance its task adaptability. However, this process is highly dependent on data quality: noisy data (such as incorrectly labeled, poorly formatted, and logically incoherent responses) can lead to biased model learning and generate inaccurate, irrelevant, or malformed outputs. While existing research has made progress in fine-tuning algorithms, the critical role of data quality has yet to be systematically addressed. In particular, in real-world scenarios, large-scale instruction datasets are often generated through crowdsourcing, log scraping, or automation, which inevitably introduces noise. Traditional data cleaning methods (such as manual review and rule-based filtering) are costly and difficult to scale, becoming a bottleneck restricting model performance.

[0003] Current data correction methods have significant limitations. First, existing technologies often rely on manual annotation or strong assumptions, making them inefficient for processing high-dimensional, complex data for text generation tasks. For example, filtering methods based on manual scoring or rule engines struggle to capture subtle semantic errors (such as logical inconsistencies and contextual disconnects) and are unable to adapt to diverse task requirements. Second, some automated solutions require the use of more powerful LLMs (such as GPT-4) for data evaluation or correction. This not only increases computational costs but also raises the technical barriers, limiting their applicability in resource-constrained scenarios (such as specialized models for specific verticals). Furthermore, existing methods often handle data filtering and correction independently, lacking systematic integration. For example, directly deleting low-quality data can lead to information loss, while blindly correcting data can introduce biases within the model itself, creating a vicious cycle. These issues make it difficult for existing data curation techniques to balance data quality and scale, ultimately impacting the robustness and generalization capabilities of fine-tuned models. Therefore, a model-independent, externally resource-independent, automated data curation framework is urgently needed that can accurately identify and correct noisy data while preserving valid information, thereby providing high-quality training sets for LLM fine-tuning and breaking through current technical bottlenecks.

[0004] One existing technique is an alignment method based on a small amount of high-quality data (Lu K, Yuan H, Yuan Z, et al. #InsTag:Instruction Tagging for Analyzing Supervised Fine-tuning of LargeLanguage Models[C] / / The Twelfth International Conference on LearningRepresentations.). This method collects high-quality question-answer pairs from online community forums. This data is screened to ensure representativeness and diversity of questions and answers. Secondly, 250 examples are manually compiled, covering a variety of task types, focusing on task diversity and consistency of response styles to simulate real user interactions. The collected data is quality-controlled to remove content that does not meet requirements, such as being too short, too long, or containing sensitive information. Furthermore, the data is formatted and cleaned to make it suitable for model training. The collected data is divided into training, development, and test sets. The training set contains 1,000 examples, while the development and test sets are used for model evaluation and validation. Ultimately, a model is trained that performs on par with or better than GPT-4.

[0005] This existing technical solution has significant deficiencies in screening low-quality data for models, and mainly relies on manual screening and correction of large amounts of data. Although this method can ensure the quality of data to a certain extent, its limitations are particularly obvious when faced with large-scale data sets. First, manual screening is difficult to maintain consistency and accuracy. Due to the existence of human factors, it is impossible to have inconsistent standards, resulting in uneven quality of the final data set. Moreover, when processing large amounts of data in multiple fields, manual screening is not only inefficient but also costly, because it requires operators to have a broad professional knowledge background in order to accurately identify and eliminate data that does not meet the requirements.

[0006] The second existing technique is a teacher-model-based data modification and iterative optimization model algorithm (Lee N, Wattanawong T, Kim S, et al. LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement [C] / / Findings of the Association for Computational Linguistics ACL2024.2024:6498-6526). First, in the preparation phase, the specific task type and required output format must be determined, and an initial seed dataset must be prepared. This step lays the foundation for subsequent model training. Next, in the base model fine-tuning phase, the selected base student model undergoes preliminary training, adjusting hyperparameter settings to suit the specific task requirements. The next key steps are error analysis and data generation. In this phase, the teacher model is used to generate targeted synthetic data based on the student model's errors on the validation set. To ensure the quality of the generated data, regular expressions are applied to ensure the correct output format, and the ROUGE metric is used to remove highly similar data points, maintaining a certain level of diversity while focusing on task relevance. Next comes the data modification phase, where the screened, high-quality synthetic data is added to the original seed dataset to form an enhanced training set, and the student model is fine-tuned again. This process not only expands the amount of training data but also specifically fills the knowledge gaps of the student model. To achieve optimal results, the entire process uses an iterative optimization approach, repeating the process of error analysis, data generation, and data enhancement. After each iteration, the performance of the student model is evaluated and the generation strategy is adjusted accordingly. In addition, it is worth noting that in each iteration, only a small number of new data points are generated for a single error example, rather than a large number of them all at once, effectively improving model performance.

[0007] This technical approach has certain limitations in data modification and augmentation, primarily relying on more advanced LLMs for implementation. While these advanced models excel on general tasks, even the most advanced LLMs currently available may struggle to generate high-quality responses that meet the needs for domain-specific tasks. This is because domain-specific tasks often require deep expertise and precise data processing capabilities, while existing LLMs may lack the necessary expertise or be unable to accurately understand certain specific contexts. Furthermore, invoking more powerful LLMs often requires more computing resources, which not only increases computational costs but also places higher hardware requirements. Summary of the Invention

[0008] The present invention provides a method for automatically screening and correcting fine-tuning data in a large language model field, and aims to solve the following technical problems:

[0009] To address the issue of data quality screening, existing methods mainly rely on manual screening or semi-automatic evaluation and screening of very large models (GPT-4), which is not only time-consuming and labor-intensive, but also difficult to ensure consistency and accuracy. Existing automatic data screening technologies fail to provide a comprehensive and reliable framework to improve datasets and their trained model outputs, especially without a stronger LLM as an auxiliary tool. This results in the fact that even if advanced fine-tuning algorithms are used, it is difficult to achieve ideal model performance if the input data quality is not high. Therefore, solving the technical problem of how to automatically, efficiently and accurately screen high-quality data is the key to improving the performance of large language models.

[0010] Another important technical challenge is how to effectively correct data samples that can be improved. Traditionally, this approach has relied on more powerful language models to generate alternative answers, but this is not feasible for domain-specific tasks, as even the most advanced LLMs currently available may not be able to generate better responses tailored to specific domain needs. Furthermore, directly editing data can introduce new biases or amplify flaws in existing model outputs, especially in the absence of sufficient confidence assessment mechanisms. Therefore, how to correct data that can be improved, retain valuable information, and minimize the reduction in learning samples caused by directly discarding data while maintaining domain expertise is another important technical issue that needs to be addressed.

[0011] The complete technical solution provided by the present invention:

[0012] The method for automatically screening and correcting fine-tuning data for large language models includes the following steps:

[0013] (1) Automatic filtering;

[0014] The automatic filtering stage is used to identify and remove low-quality data pairs to ensure that the dataset used in the subsequent fine-tuning process is of the highest possible quality.

[0015] It includes the following sub-steps:

[0016] ① Initialization:

[0017] Get an instruction tuning data set, where x i Represents prompt input, y i Represents the corresponding target response. qwen2.5-7b-instruct is used as the base model (without any task-specific fine-tuning).

[0018] ②Calculate the confidence score:

[0019] For each data pair (x i ,y i), use the BSDetector tool to estimate the confidence score of its quality. BSDetector works as follows:

[0020] By increasing the diversity of sampling techniques (such as temperature sampling, chain thinking, etc.), LLM can be based on prompt x i Generate multiple candidate responses.

[0021] Use natural language inference techniques to evaluate these candidate responses and the target response y i This step can be viewed as a measure of the semantic consistency between the observed

[0022] Directly let the LLM self-evaluate the quality of the target response, asking the LLM how it feels about the response it gave i level of confidence.

[0023] Combining the results of the above two methods, we get the comprehensive confidence score c i , which takes into account both randomness and knowledge uncertainty.

[0024] ③Filter data:

[0025] The confidence score c calculated based on the calculation i Filter the data set: set a predefined threshold γ, retain the data pairs above the threshold γ, and form a new data set F = {(x i ,y i )|c i >γ}.

[0026] ④Instruction fine-tuning:

[0027] The LLM is fine-tuned using the selected high-confidence dataset F to improve the performance of the model in a specific field and achieve initial fine-tuning, which is called LLM′.

[0028] (2) Automatic correction;

[0029] The auto-correction stage aims to automatically correct data pairs that are identified as low quality but have the potential to generate better responses through LLM.

[0030] It includes the following sub-steps:

[0031] ① Generate candidate responses using the fine-tuned LLM:

[0032] For each hint x in the original dataset i , using the LLM′ fine-tuned in the Auto-Filter stage to generate a new candidate response y′ i .

[0033] ②Evaluate the quality difference between new and old responses:

[0034] Use the base pre-trained LLM (i.e., the state before task-specific fine-tuning) as the critic to compare the newly generated responses y′ i and the response y in the original dataset i This step involves asking the base LLM which response is better and getting the LLM's confidence score for its preferred choice. The BSDetector method is used here to get the confidence score to determine whether the newly generated response is better than the original response.

[0035] ③ Set the confidence threshold and decide whether to replace the response:

[0036] Set a confidence threshold η, if the base LLM believes that the newly generated response y′ i Than the original response y i Better, and the confidence estimate of this preference choice exceeds the set threshold η, then y′ is used in the data set i Replace the original y i .

[0037] ④ Use the modified dataset for further fine-tuning:

[0038] The LLM is fine-tuned again using a dataset that contains some replacements with new responses in order to obtain a model with better performance.

[0039] (3) Iterative improvement;

[0040] Repeat the above process, i.e., use the updated dataset to fine-tune the LLM again, forming a continuous optimization cycle until a satisfactory performance level is achieved.

[0041] The beneficial effects brought about by the technical solution of the present invention are:

[0042] This paper designs an automated data screening method that uses confidence estimates derived from LLM to identify and remove low-quality data samples, significantly improving the overall quality of the dataset. Compared to manual screening or semi-automatic evaluation and screening based on very large models (such as GPT-4), this method not only addresses the data quality and consistency issues of existing technologies, but also provides higher-quality input data for fine-tuning algorithms, thereby improving the performance of the final model.

[0043] This paper designs a method for automatically correcting data samples that can be improved and generating alternative answers based on a fine-tuned LLM. This method confidently corrects data that can be improved while maintaining domain expertise, avoiding the introduction of new biases or amplification of flaws in existing model outputs. This method not only mitigates the problem of reduced learning samples caused by simply discarding data, but also preserves valuable information, enhancing the model's adaptability and performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 This is a flow chart of automatic data filtering of the present invention;

[0045] Figure 2 is a flow chart of the data modification process of the present invention;

[0046] Figure 3 Flowchart of the present invention. DETAILED DESCRIPTION

[0047] The specific technical solutions of the present invention will be described with reference to the accompanying drawings. Figure 3 As shown, the following process is included:

[0048] 1. Automatic data filtering

[0049] The data filtering process of the present invention is as follows Figure 1 As shown in the figure, we first calculate the execution confidence of different question-answer pairs. If the confidence exceeds the set threshold, we add a new training dataset. If it is below the threshold, we add a pending modified dataset. Finally, we use the new training dataset to fine-tune the LLM.

[0050] 2. Data correction process

[0051] The data correction process of the present invention is as follows Figure 2 As shown in the figure, first, the newly fine-tuned LLM is used to generate new responses to the data to be corrected. The confidence of the new and old questions and answers is calculated. If the new confidence is greater than the old confidence and exceeds the set threshold, the new training dataset is added. If the new confidence is lower than the threshold, the data is directly deleted. Finally, the new training dataset is used to iteratively update the data.

[0052] For the data screening part, an alternative approach is to use a multi-stage data filtering strategy that combines rule-based preprocessing with machine learning-based post-processing. First, the raw data is preliminarily screened by setting a series of clear quality control rules (for example, length restrictions, keyword matching, etc.). Then, a trained classification model is used to further evaluate the preliminarily screened data to identify and remove low-quality data samples. However, this method requires additional rule design and model training work, and can be adapted to different application scenarios by adjusting the rules and optimizing the model.

[0053] A viable alternative to data correction is an iterative data refinement process. Specifically, an existing fine-tuned model is used to generate preliminary answers. The answers are then revised based on feedback from multiple annotators collected through a crowdsourcing platform. Finally, the model is rerun to verify that the revised answers outperform the original versions. This approach is time-consuming and labor-intensive.

Claims

1. A method for automatically screening and correcting fine-tuning data in a large language model domain, characterized by: The following steps are involved: (1) Automatic filtering; The automatic filtering stage is used to identify and remove low-quality data pairs to ensure that the dataset used in the subsequent fine-tuning process is as high-quality as possible; (2) Automatic correction; The auto-correction stage aims to automatically correct those data pairs that are identified as low-quality but have the potential to generate better responses through LLM; (3) Iterative improvement; Repeat the above process, i.e., use the updated dataset to fine-tune the LLM again, forming a continuous optimization cycle until a satisfactory performance level is achieved.

2. The method for automatically screening and correcting domain fine-tuning data for a large language model according to claim 1, characterized in that: (1) Automatic filtering specifically includes the following sub-steps: ① Initialization: Get an instruction tuning data set, where x i Represents prompt input, y i represents the corresponding target response; at the same time, qwen2.5-7b-instruct is used as the base model; ②Calculate the confidence score: For each data pair (x i ,y i ), use the BSDetector tool to estimate the confidence score of its quality; BSDetector works as follows: By increasing the diversity sampling technique, LLM is based on the hint x i generating multiple candidate responses; Use natural language inference techniques to evaluate these candidate responses and the target response y i Semantic consistency between Directly let the LLM self-evaluate the quality of the target response, asking the LLM how it feels about the response it gave i level of confidence; Combining the results of the above two methods, we get the comprehensive confidence score c i ,This score takes into account both randomness and knowledge uncertainty; ③Filter data: The confidence score c calculated based on the calculation i Filter the data set: set a predefined threshold γ, retain the data pairs above the threshold γ, and form a new data set F = {(x i ,y i )|c i >γ}; ④Instruction fine-tuning: The LLM is fine-tuned using the selected high-confidence dataset F to improve the performance of the model in a specific field and achieve initial fine-tuning, which is called LLM′.

3. The method for automatically screening and correcting domain fine-tuning data for a large language model according to claim 1, characterized in that: (2) Automatic correction specifically includes the following sub-steps: ① Generate candidate responses using the fine-tuned LLM: For each hint x in the original dataset i , using the LLM′ fine-tuned in the Auto-Filter stage to generate a new candidate response y′ i ; ②Evaluate the quality difference between new and old responses: Use the base pre-trained LLM as a critic to compare the newly generated responses y′ i and the response y in the original dataset i The quality of the generated response is obtained by using the BSDetector method to obtain a confidence score to determine whether the newly generated response is better than the original response. ③ Set the confidence threshold and decide whether to replace the response: Set a confidence threshold η, if the base LLM believes that the newly generated response y′ i Than the original response y i Better, and the confidence estimate of this preference choice exceeds the set threshold η, then y′ is used in the data set i Replace the original y i ; ④ Use the modified dataset for further fine-tuning: The LLM is fine-tuned again using a dataset that contains some replacements with new responses in order to obtain a model with better performance.