Big language fine-tuning data screening method, system and device based on anti-fact data enhancement and storage medium
By generating pseudo-responses and counterfactual questions through counterfactual data augmentation, and combining large language model validation and determinant point process filtering, the problems of noise and duplicate data in large language model instruction fine-tuning are solved. This enables efficient filtering of high-quality and diverse datasets, improves model performance, and reduces costs.
Patent Information
- Application Number
- CN202510756049.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-07
- Publication Date
- 2025-10-31
AI Technical Summary
During the instruction fine-tuning process of existing large language models, the datasets often contain noise and duplicate data. Existing data filtering methods cannot effectively filter out high-quality and diverse datasets, resulting in high computational costs and affecting model performance.
We employ counterfactual data augmentation methods, generating pseudo-responses through a weak language model, verifying the correctness of pseudo-answers using a large language model, designing a masked natural language reasoning task to generate counterfactual questions, and filtering high-quality and diverse datasets through counterfactual data quality metrics and determinant point processes.
It effectively selects high-quality and diverse datasets, improves model performance in instruction fine-tuning, and reduces computational costs.
Smart Images

Figure CN120873113A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power plant fault database data management technology, specifically involving a method, system, device and storage medium for filtering large language fine-tuning data based on counterfactual data enhancement. Background Technology
[0002] In the fine-tuning of instructions in large language models, the quality and diversity of the training data are crucial. However, current instruction fine-tuning datasets often contain spurious relevance data, and mainstream data filtering methods cannot adequately consider the relationship between the original data and counterfactual data. Furthermore, these methods measure data quality and diversity separately, leading to increased computational costs.
[0003] With the rapid development of generative artificial intelligence, large language models have achieved remarkable success in the field of natural language processing. Instruction tuning, as a crucial training stage, can improve the performance of large language models on downstream tasks. Existing commonly used large-scale instruction fine-tuning datasets may contain noisy and repetitive data. Therefore, data filtering methods are needed to select the most valuable datasets for fine-tuning large language models from a large amount of data. Some existing data filtering schemes include:
[0004] 1) Based on indicators, such as perplexity, the average uncertainty of a large model in predicting the next element in a sequence is used to measure the quality of the data, and data with high uncertainty is selected.
[0005] Limitations: It only focuses on the uncertainty of the data for large models and does not fully consider the possibility of spurious correlations in the dataset, which affects model performance.
[0006] 2) Based on closed-source large language models, such as ALPAGASUS, design manual prompt word templates, and let the closed-source large language model output the quality score of each fine-tuned data, and filter the data with high quality scores;
[0007] Disadvantages: Calling closed-source large models requires high costs;
[0008] 3) Methods based on small language models, such as IterSelectTune, train a classifier based on the BERT model and fit the judgment of GPT-4 to select high-quality instruction data;
[0009] Disadvantages: It requires the prior construction of a training set using a closed-source large model before training, which is costly and inflexible, and the capabilities of the small model cannot be fully aligned with those of the large model. Summary of the Invention
[0010] The first objective of this invention is to provide a method for filtering large language fine-tuning data based on counterfactual data augmentation, addressing the aforementioned problems.
[0011] Therefore, the above-mentioned objective of the present invention is achieved through the following technical solution:
[0012] A method for filtering large language fine-tuning data based on counterfactual data augmentation includes the following steps:
[0013] S1. Counterfactual data enhancement, which includes the following sub-steps:
[0014] S11, Generate a pseudo response;
[0015] S12. Verify the correctness of the false answer;
[0016] S13. Generate counterfactual questions;
[0017] S14, Question-Answer Verification;
[0018] S2. Data filtering, which includes the following sub-steps:
[0019] S21. Establish quality metrics for counterfactual data;
[0020] S22, Determinant point process filtering.
[0021] While adopting the above technical solutions, the present invention may also adopt or combine the following technical solutions:
[0022] As a preferred technical solution of the present invention: in step S11, a pseudo response is generated for the original instruction through a weak language model.
[0023] As a preferred technical solution of the present invention: in step S12, the correctness of the pseudo-answer is verified by a closed-source large language model, and low-difficulty data is discarded.
[0024] As a preferred technical solution of the present invention: in step S13, a masked natural language reasoning task prompt template is designed, and a counterfactual question is generated based on the pseudo-answer using a large model.
[0025] As a preferred technical solution of the present invention: in step S14, an answer verification task is constructed to filter new questions that are semantically consistent and correct.
[0026] As a preferred technical solution of the present invention: In step S21, the Counterfactual data quality metric, CounterScore, is determined by the following formula:
[0027] CounterScore = IFD(i o )×IFD(i cf )×(1-sim(i o i cf ))
[0028] In the formula, IFD represents the instruction following difficulty score, i o Represents the original data, i cf Representing the corresponding counterfactual data, sim is the cosine similarity calculation function.
[0029] As a preferred embodiment of the present invention, step S22 further includes the following sub-steps:
[0030] Deterministic point processes are defined by a positive semidefinite kernel matrix K. For a point process P, a subset of... The probability of selection is:
[0031]
[0032] In the formula, P represents a point process, Y is a set containing all data, A is a subset of Y, and K is a positive semi-definite kernel matrix used to define the point process. A Determinant is obtained from the coordinate index K of the element in A;
[0033] Using L-ensemble theory, the probability of selecting subset A is rewritten by defining the DPP through a real symmetric positive semidefinite matrix L as:
[0034]
[0035] In the formula, L is a real symmetric and positive semi-definite kernel matrix. A I is obtained from the coordinate index L of the elements in A, where I is the identity matrix;
[0036] The parameterized representation of matrix L is as follows:
[0037]
[0038] In the formula, q i It is the quality score of data point i, φ i It is the vector representation of data point i;
[0039] To balance data quality and diversity, a tradeoff hyperparameter λ∈[0,1] is introduced, and the modified log probability is:
[0040]
[0041] In the formula, λ is a hyperparameter balancing data quality and diversity, and S A It is a matrix representing the Euclidean distance between data in set A;
[0042] By employing maximum a posteriori inference, selecting high-quality and diverse datasets, and using a fast greedy algorithm, the time complexity is O(m). 2The final subset A is a subset that simultaneously possesses both high quality and high diversity.
[0043] The second objective of this invention is to provide a large language fine-tuning data filtering system based on counterfactual data augmentation, comprising the following modules:
[0044] The counterfactual data enhancement module is used to generate counterfactual samples and mitigate spurious relevance.
[0045] The data filtering module is used to filter out high-quality and diverse subsets of data, balancing data quality and diversity.
[0046] A third objective of this invention is to provide an electronic device comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus.
[0047] The memory is used to store computer programs;
[0048] A processor for executing a computer program stored in memory to implement the steps of the large language fine-tuning data filtering method based on counterfactual data augmentation as described above.
[0049] Another objective of this invention is to provide a non-volatile storage medium storing an executable program, which, when executed by a processor, implements the steps of the large language fine-tuning data filtering method based on counterfactual data augmentation as described above.
[0050] Compared with existing technologies, the present invention has the following advantages: the present invention, through counterfactual scores and deterministic point processes, can simultaneously consider the quality and diversity of data, and select a better subset of data; the present invention, through the generation and verification of counterfactual data, can effectively improve the quality of counterfactual data and enhance the model's performance in instruction fine-tuning; the present invention, through the rapid implementation of deterministic point processes, can select high-quality data at a lower computational cost. Attached Figure Description
[0051] Figure 1 The flowchart shows the large language fine-tuning data filtering method based on counterfactual data enhancement provided by this invention. Detailed Implementation
[0052] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0053] like Figure 1 As shown, a method for filtering large language fine-tuning data based on counterfactual data augmentation specifically includes the following steps:
[0054] S1. Counterfactual data enhancement, which includes the following sub-steps:
[0055] S11. Generate pseudo-responses: Use a weak language model (such as Qwen-2.5-0.5B) to generate pseudo-responses for the original instructions. Specifically, for each piece of data to be filtered in the original dataset, which contains two parts: a question and an answer, input the question into the weak language model to obtain the pseudo-response output by the model. For example, the question is: What language was the world's first printer used to print? The original answer is: The world's first printer was invented in Kutenberg, Germany, and it printed Latin, the religious language that was widely used at the time. The pseudo-response is: The world's first printer was invented in Kutenberg, Germany, so it printed German.
[0056] S12. Verify the correctness of pseudo-answers. A Large Language Model (LLM, specifically DeepSeekV3 in this invention) is used as the judge to verify the correctness of pseudo-answers. If the pseudo-answer is correct, the data point is discarded because if a weak language model can answer correctly, the difficulty level is low, and data augmentation for it has limited value. Specifically, for each question-answer pair and its corresponding pseudo-response in the original dataset, all three are input into the Large Language Model. Through cue word engineering, the Large Language Model acts as the judge to determine whether there is a deviation between the pseudo-response and the original answer. If there is a deviation, it means the model lacks the ability to correctly answer the question, and the model needs this data for training to improve performance. If there is no deviation, it means the model has the ability to answer the question, the question-answer pair is low in difficulty, and it is redundant data, so it is discarded. For example, in the above example, the pseudo-response is clearly incorrect; therefore, we will retain this data.
[0057] S13. For the filtered training data, the powerful DeepseekV3 is used to generate counterfactual questions. A Mask-NLI style prompt template is designed, using the original instruction-response pair (q,a) and the pseudo-response a' as examples of in-context learning. The large language model fills the mask based on the pseudo-response to generate new counterfactual questions, and the large model generates counterfactual questions based on the pseudo-answers. For example, the following is an example style prompt template:
[0058] You are a question generation expert. Given a model answer, you must generate a new counterfactual question based on the answer. Note that you should not directly copy the question from the example. Example: Question: [MASK], Model Answer: The world's first printer was invented in Kutenberg, Germany, and it printed Latin, the widely used religious language at the time. New Counterfactual Question Output by the Model: What language was the world's first printer used to print? End of example. Question: [MASK], Model Answer: The world's first printer was invented in Kutenberg, Germany, so it printed German. New Counterfactual Question:
[0059] Based on the prompts, the model will output a new counterfactual question: If the world's first printer was not used for religious purposes, which language might it have printed first? S14, Question-Answer Verification: Introducing a checking module, the verification process is constructed as an Answer Verification (AV) task. Given an answer and a newly generated counterfactual question, the large language model evaluates whether the response correctly resolves the counterfactual question, discarding irrelevant or incorrect responses and filtering for semantically consistent and correct new questions. For example, in the above example, a new counterfactual question and response pair is obtained: New counterfactual question: If the world's first printer was not used for religious purposes, which language might it have printed first? New response: The world's first printer was invented in Kutenberg, Germany, so it printed German. Using cue word engineering, the large language model acts as the judge, verifying the answer to this new counterfactual question and response pair.
[0060] Here is an answer verification prompt template: You are an answer verification expert. Given a question and a model answer, please determine whether the model answer correctly answers the question. If it is correct, output "Yes"; if it is incorrect or irrelevant, output "No". Please output the judgment first, then the explanation. Question: Model Answer:
[0061] The new counterfactual question and response pairs are populated into the corresponding positions in the prompt template, input into the large model, and the judgment result is obtained. Irrelevant or incorrect responses are discarded, and new counterfactual question and response pairs that are semantically consistent and correct are retained.
[0062] S2. Data filtering, given a dataset:
[0063] D = {(x1,y1),(x2,y2),…,(x n ,y n )}
[0064] In the formula, the instruction and input (Instruction,[Input]) is x, and the model response is y.
[0065] The goal of instruction fine-tuning is to minimize model M. θ Loss function:
[0066]
[0067] In the formula: M represents the model, M θ This represents the model parameters, where n represents the number of data points, and x represents the model parameters. i Indicates the model input, y i This represents the model output.
[0068] Select a high-quality subset containing m data points using the data selection method π. The goal is to maximize the evaluation metric Q, i.e.
[0069]
[0070] In the formula: π represents the data selection method, Q represents the evaluation index, argmax represents the input value that makes the function reach its maximum value, m represents the number of selected data, and H represents the high-quality data set.
[0071] Specifically, it includes the following sub-steps:
[0072] S21. Establish the Counterfactual data quality metric, CounterScore, determined by the following formula:
[0073] CounterScore = IFD(i o )×IFD(i cf )×(1-sim(i o i cf ))
[0074] In the formula, IFD represents the instruction following difficulty score, i o Represents the original data, i cf Representing the corresponding counterfactual data, sim is the cosine similarity calculation function.
[0075] This indicator takes into account the quality of the raw data and counterfactual data, as well as the semantic differences between them.
[0076] S22. Use a data filtering method based on determinantal point processes (DPP), employing CounterScore-based determinantal point processes (CS-DPP) for data selection. A determinantal point process is defined by a positive semi-definite kernel matrix K. For a point process P, a subset... The probability of selection is:
[0077]
[0078] In the formula, P represents a point process, Y is a set containing all data, A is a subset of Y, and K is a positive semi-definite kernel matrix used to define the point process. A Det is obtained from the coordinate index K of the element in A, where det represents the determinant.
[0079] Using L-ensemble theory, the probability of selecting subset A is rewritten by defining the DPP through a real symmetric positive semidefinite matrix L as:
[0080]
[0081] In the formula, L is a real symmetric and positive semi-definite kernel matrix. A I is obtained from the coordinate index L of the elements in A, where I is the identity matrix.
[0082] The parameterized representation of matrix L is as follows:
[0083]
[0084] In the formula, q i It is the quality score of data point i, φ i It is the vector representation of data point i;
[0085] To balance data quality and diversity, a tradeoff hyperparameter λ∈[0,1] is introduced, and the modified log probability is:
[0086]
[0087] In the formula, λ is a hyperparameter balancing data quality and diversity, and S A It is a matrix representing the Euclidean distance between data in set A.
[0088] By employing maximum a posteriori inference, selecting high-quality and diverse datasets, and using a fast greedy algorithm, the time complexity is O(m). 2 The final subset A is a subset that simultaneously possesses both high quality and high diversity.
[0089] This invention also provides a large language fine-tuning data filtering system based on counterfactual data augmentation, comprising the following modules:
[0090] The counterfactual data enhancement module is used to generate counterfactual samples and mitigate spurious relevance.
[0091] The data filtering module is used to filter out high-quality and diverse subsets of data, balancing data quality and diversity.
[0092] The present invention also provides an electronic device, which includes a processor, a communication interface, a memory, and a communication bus. The processor, the communication interface, and the memory communicate with each other via the communication bus.
[0093] Memory, used to store computer programs;
[0094] A processor is used to execute computer programs stored in memory to implement the steps of the large language fine-tuning data filtering method based on counterfactual data augmentation as described above.
[0095] The present invention also provides a non-volatile storage medium storing an executable program, which, when executed by a processor, implements the steps of the large language fine-tuning data filtering method based on counterfactual data augmentation as described above.
[0096] The technical solution of the present invention has been described in conjunction with the specific experimental procedures shown in the accompanying drawings. However, the scope of protection of the present invention is not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions resulting from such changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A method for filtering large language fine-tuning data based on counterfactual data augmentation, characterized in that, Includes the following steps: S1. Counterfactual data enhancement, which includes the following sub-steps: S11, Generate a pseudo response; S12. Verify the correctness of the false answer; S13. Generate counterfactual questions; S14, Question-Answer Verification; S2. Data filtering, which includes the following sub-steps: S21. Establish quality metrics for counterfactual data; S22, Determinant point process filtering.
2. The method according to claim 1, characterized in that: In step S11, a pseudo-response is generated for the original instruction using a weak language model.
3. The method according to claim 1, characterized in that: In step S12, the correctness of the pseudo-answer is verified by a closed-source large language model, and low-difficulty data is discarded.
4. The method according to claim 1, characterized in that: In step S13, a masked natural language reasoning task prompt template is designed, and a large model is used to generate counterfactual questions based on pseudo-answers.
5. The method according to claim 1, characterized in that: In step S14, an answer verification task is constructed to filter out new questions that are semantically consistent and correct.
6. The method according to claim 1, characterized in that: In step S21, the Counterscore, a quality metric for counterfactual data, is determined using the following formula: CounterScore=IFD(i o )×IFD(i cf )×(1-sim(i o ,i cf )) In the formula, IFD represents the instruction following difficulty score, i o Represents the original data, i cf Representing the corresponding counterfactual data, sim is the cosine similarity calculation function.
7. The method according to claim 1, characterized in that: Step S22 also includes the following sub-steps: Determinant point processes are defined by a positive semidefinite kernel matrix K. For a point process P, a subset of... The probability of selection is: In the formula, P represents a point process, Y is a set containing all data, A is a subset of Y, and K is a positive semi-definite kernel matrix used to define the point process. A Determinant is obtained from the coordinate index K of the element in A; Using L-ensemble theory, the probability of selecting subset A is rewritten by defining the DPP through a real symmetric positive semidefinite matrix L as: In the formula, L is a real symmetric and positive semi-definite kernel matrix. A I is obtained from the coordinate index L of the elements in A, where I is the identity matrix; The parameterized representation of matrix L is as follows: In the formula, q i It is the quality score of data point i, φ i It is the vector representation of data point i; To balance data quality and diversity, a tradeoff hyperparameter λ∈[0,1] is introduced, and the modified log probability is: In the formula, λ is a hyperparameter balancing data quality and diversity, and S A It is a matrix representing the Euclidean distance between data in set A; By employing maximum a posteriori inference, selecting high-quality and diverse datasets, and using a fast greedy algorithm, the time complexity is O(m). 2 The final subset A is a subset that simultaneously possesses both high quality and high diversity.
8. A large language fine-tuning data filtering system based on counterfactual data augmentation, characterized in that, Includes the following modules: The counterfactual data enhancement module is used to generate counterfactual samples and mitigate spurious relevance. The data filtering module is used to filter out high-quality and diverse subsets of data, balancing data quality and diversity.
9. An electronic device, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus, characterized in that: The memory is used to store computer programs; A processor for executing a computer program stored in memory to implement the steps of the large language fine-tuning data filtering method based on counterfactual data augmentation as described in any one of claims 1-7.
10. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores an executable program, which, when executed by a processor, implements the steps of the large language fine-tuning data filtering method based on counterfactual data augmentation as described in any one of claims 1-7.