An instruction fine-tuning dataset generation method, an electronic device, and a storage medium
By cross-validating task instruction seed sets and real-world materials to generate high-quality instruction fine-tuning datasets, the problems of high cost and low efficiency in existing technologies are solved. This enables the generation of complex and diverse datasets and enhances the comprehensive application capabilities of large language models.
Patent Information
- Application Number
- CN202511733406.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-08-04
- Estimated Expiration
- 2045-11-24
AI Technical Summary
Existing methods for generating instruction fine-tuning datasets are costly and inefficient, making it difficult to scale up the generation of datasets that are complex, diverse, and of high quality. They also have significant shortcomings, especially in complex domains such as cybersecurity situation or scenario analysis.
By acquiring a set of task instruction seeds, generating materials and questions, and performing inference calculations, cross-validating with real-world materials, employing Few-shot Prompting technology and information entropy enhancement, and utilizing multi-model inference paths to ensure answer consistency, a high-quality instruction fine-tuning dataset is finally generated.
It enables efficient and large-scale generation of instruction fine-tuning datasets with both complexity and diversity, significantly improving the comprehensive application capabilities of large language models in complex scenarios and domains, and meeting the needs of complex tasks.
Smart Images

Figure CN121579077B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and in particular to a method for generating instruction fine-tuning datasets, an electronic device, and a storage medium. Background Technology
[0002] In recent years, large language models (LLMs) have made significant progress. However, their fundamental pre-training objective is essentially sequential "next word prediction," which means that the original LLMs themselves lack the ability to directly follow complex human instructions or perform task-specific dialogues. To bridge this gap, instruction tuning techniques have emerged. This technique teaches pre-trained models to understand and execute human intentions by performing supervised fine-tuning (SFT) on a dataset consisting of (instruction, output) pairs, enabling them to handle a wide range of tasks from simple question answering to complex report generation. Therefore, constructing high-quality, large-scale instruction tuning datasets has become crucial for improving the capabilities of LLMs.
[0003] Currently, the methods for constructing instruction fine-tuning datasets mainly fall into two categories: traditional manual construction methods and automated synthetic data generation techniques. While these methods can solve certain problems, they also have several drawbacks. First, manual construction is costly. Domain experts manually write and label data, ensuring quality, but the process is time-consuming, labor-intensive, and extremely expensive, and difficult to scale up, failing to meet the large data demands of model iteration. Second, the quality of existing datasets varies greatly. Publicly available general datasets contain few samples relevant to complex cognitive tasks. Data generated through simple automated methods (such as synonym replacement or sentence rewriting of existing text) often has simple logic and insufficient depth, easily leading to models learning superficial patterns rather than true reasoning ability—that is, "overfitting" to the data format rather than the task itself. Third, the task dimensions are limited. Existing datasets typically only cover a few task types, lacking comprehensive evaluation and training of the model's overall capabilities, particularly showing significant shortcomings in core capabilities of complex tasks such as time computation, multi-point information integration, statistical analysis, and event trend judgment.
[0004] Therefore, how to generate instruction fine-tuning datasets that are complex, diverse, and of high quality at low cost, high efficiency, and on a large scale to meet the needs of complex fields such as network security situation or situation analysis is an urgent technical problem to be solved. Summary of the Invention
[0005] To address the aforementioned technical problems, the technical solution adopted by this invention is as follows: A method for generating instruction fine-tuning datasets, applied to large language models, includes the following steps: Step S01: Obtain the task instruction seed set S = (S1, S2, ..., S...) i S j ), respectively proceed to steps S02 and S03; where i = 1, 2, ..., j; j is the number of instruction templates contained in S; S i This is the i-th instruction template; each instruction template corresponds to a task type; each instruction template is used to generate the materials and / or problems corresponding to its corresponding task type; Step S02: Obtain at least one material and its corresponding question generated based on S; perform reasoning calculations on the material and its corresponding question to obtain the answer corresponding to each material and its corresponding question; Step S03: Calculate at least one real-world material based on S to obtain at least one question; perform reasoning calculations on the real-world material and its corresponding question to obtain the answer corresponding to each real-world material and its corresponding question; Step S04: Cross-validate the answers generated in steps S02 and S03, and determine the answers that pass the validation as the target answers; Step S05: Store each target answer, its corresponding question, and its corresponding material or real-world material into the instruction fine-tuning dataset.
[0006] Furthermore, obtaining at least one material and its corresponding problem generated based on S includes: By processing S using the Few-shot Prompting technique, we obtain the corresponding material and problem set W = (W1, W2, ..., W...). i ,…,W j ); where W i According to S i The generated material and problem subset; W i =(W i,1 W i,2 ,…,W i,a ,…,W i,b ); a = 1, 2, ..., b; b is S i Corresponding preset materials and number of questions; W i,a For W i The a-th material and question; W i,a =(C i,a Q i,a ); C i,a For W i The a-th material in the middle; Q i,a Ci,a The corresponding question.
[0007] Furthermore, the method also includes: Obtain the material set C = (C1, C2, ..., C) corresponding to S. i C j ); where C i For the i-th subset of materials; C i =(C i,1 C i,2 C i,a C i,b ); Analysis C i,a The information units contained therein, resulting in C i,a The corresponding set of information units X = (X1, X2, ..., X...) h , ..., X f ); where h = 1, 2, ..., f; f is C i,a The number of information units contained; X h C i,a The corresponding h-th information unit; According to X, C i,a Add at least one interference information to obtain S. i,a To improve C i,a Information entropy H; the interference information is the C i,a The corresponding background information, and / or information similar to but unrelated to any information unit in X; where H = ;P(X) h ) is X h In C i,a The probability of it appearing in the middle.
[0008] Furthermore, the step of reasoning and calculating on the materials and their corresponding questions to obtain the answer for each material and its corresponding question includes: For the S i,a and Q i,a The answer A obtained through reasoning and calculation i,a To obtain the answer set A = (A1, A2, ..., A...) corresponding to S. i A j ); where A i A is a subset of the i-th answer; i = (A i,1 A i,2 A i,a A i,b ).
[0009] Furthermore, the calculation based on S on at least one real-world material yields at least one problem, including: Obtain a real-world material set R = (R1, R2, ..., R...) g , ..., R k ); where g = 1, 2, ..., k; k is the preset quantity of materials to be acquired; R g To obtain the g-th material; The R is processed by an information validity filter to obtain a high-quality material set G = (G1, G2, ..., G...). v , ..., G z ); where v = 1, 2, ..., z; z is the quantity of high-quality material obtained; G v This is the vth high-quality material; Based on the given S, G is processed using a Few-shot method to obtain the problem set U = (U1, U2, ..., U...) corresponding to G. v , ..., U z ); where U v For G v The corresponding subset of problems; when according to G v When it is impossible to obtain a problem for any task type corresponding to S, U v Marked as 0; when according to G v When at least one task type corresponding to S can be obtained, U v =(U v,1 U v,2 , ..., U v,e , ..., U v,r e = 1, 2, ..., r; r is U v Number of issues included; U v,e For U v The e-th question included.
[0010] Furthermore, the step of reasoning and calculating based on the real-world materials and their corresponding questions to obtain the answer for each real-world material and its corresponding question includes: For the G v and U v,e The answer E obtained through reasoning and calculation v,e To obtain the answer set E = (E1, E2, ..., E...) corresponding to G. v , ..., E z ); where E v For the v-th answer subset; when U v When marked as 0, E v Marked as 0; when U v When not marked as 0, E v=(E v,1 E v,2 , ..., E v,e , ..., E v,r ).
[0011] Further, step S04 includes: Based on different models or different inference paths based on the same model, the S i,a and Q i,a Through reasoning and calculation, several first-pass answers are obtained; A is determined. i,a Whether the consistency with the aforementioned first verification answers meets the specified threshold; if it does, then A is... i,a This has been identified as the target answer. Based on different models or different inference paths based on the same model, the G v and U v,e By performing reasoning and calculations, several second verification answers are obtained; E is determined. v,e Whether the consistency with the aforementioned second verification answers meets the specified threshold; if it does, then E... v,e This has been identified as the target answer.
[0012] Furthermore, when A is determined i,a E v,e For the target answer, step S05 includes: S i,a Q i,a and A i,a Store the instruction fine-tuning dataset in triplet format; G v U v,e and E v,e Store the instruction fine-tuning dataset in triplet format.
[0013] A non-transitory computer-readable storage medium storing at least one instruction or at least one program segment, the at least one instruction or the at least one program segment being loaded and executed by a processor to implement the aforementioned method.
[0014] An electronic device includes a processor and the aforementioned non-transitory computer-readable storage medium.
[0015] The present invention has at least the following beneficial effects: This invention is based on a set of task instruction seeds encompassing multiple core capabilities such as time calculation, statistical analysis, inductive reasoning, and event understanding. It guides a large-scale oracle model to generate materials and / or questions based on instruction templates. This includes generating materials with complex plots and corresponding questions from scratch, and generating corresponding instruction questions from materials based on the real world. Then, a large-scale language model is used to perform in-depth analysis of the "material-question" pair, simulating the thought process of human experts to deduce answers. The generated answers are then automatically cross-validated to ensure their accuracy, consistency, and logical coherence, ultimately selecting a high-quality instruction fine-tuning dataset. This invention can generate high-quality, diverse instruction fine-tuning datasets in parallel and on a large scale, meeting the needs of complex tasks such as analyzing development trends or situations in a specific field, thus significantly improving the comprehensive application capabilities of large-scale language models in complex scenarios and domains. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A flowchart of a method for generating instruction fine-tuning datasets provided in an embodiment of the present invention; Figure 2 This is a structural diagram of an instruction fine-tuning dataset generation system provided in an embodiment of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] This invention provides an embodiment of a method for generating instruction fine-tuning datasets, applied to large language models, such as... Figure 1 As shown, it includes the following steps: Step S01: Obtain the task instruction seed set S = (S1, S2, ..., S...) i S j ), respectively proceed to steps S02 and S03; where i = 1, 2, ..., j; j is the number of instruction templates contained in S; S iThis is the i-th instruction template; each instruction template corresponds to a task type; each instruction template is used to generate materials and / or questions corresponding to its task type; the task types include, but are not limited to, time calculation, statistical analysis, inductive summarization, and event understanding; Step S02: Obtain at least one material and its corresponding question generated based on S; perform reasoning calculations on the material and its corresponding question to obtain the answer corresponding to each material and its corresponding question; Step S03: Calculate at least one real-world material based on S to obtain at least one question; perform reasoning calculations on the real-world material and its corresponding question to obtain the answer corresponding to each real-world material and its corresponding question; Step S04: Cross-validate the answers generated in steps S02 and S03, and determine the answers that pass the validation as the target answers; Step S05: Store each target answer, its corresponding question, and its corresponding material or real-world material into the instruction fine-tuning dataset.
[0020] Figure 1 The embodiment includes two parallel data generation pipelines: a command-based synthetic material generation pipeline corresponding to step S02, and a command generation pipeline based on real materials corresponding to step S03. Based on a set of task command seeds encompassing multiple core capabilities such as time calculation, statistical analysis, inductive reasoning, and event understanding, the pipeline guides a large-scale oracle model to generate materials and / or questions according to command templates. This includes generating materials with complex plots and their corresponding questions from scratch, and generating corresponding command questions from materials based on the real world. Then, a large-scale language model is used to perform in-depth analysis of the "material-question" pair, simulating the thought process of human experts to deduce answers. The generated answers are then automatically cross-validated to ensure accuracy, consistency, and logic, ultimately selecting a high-quality command fine-tuning dataset. Figure 1 The embodiments described above can generate high-quality, diverse instruction fine-tuning datasets in parallel and on a large scale, meeting the needs of complex tasks such as analyzing the development trends or situations in a certain field, thereby significantly improving the comprehensive application capabilities of large language models in complex scenarios and domains.
[0021] Preferably, obtaining at least one material and its corresponding problem generated based on S includes: By processing S using the Few-shot Prompting technique, we obtain the corresponding material and problem set W = (W1, W2, ..., W...). i ,…,W j ); where W i According to S iThe generated material and problem subset; W i =(W i,1 W i,2 ,…,W i,a ,…,W i,b ); a = 1, 2, ..., b; b is S i Corresponding preset materials and number of questions; W i,a For W i The a-th material and question; W i,a =(C i,a Q i,a ); C i,a For W i The a-th material in the middle; Q i,a C i,a The corresponding question.
[0022] The above-mentioned preferred scheme utilizes a large language model and, through Few-shot Prompting technology, generates a batch of materials and corresponding questions based on the instruction templates contained in S. Examples of prompts for some task types are shown below.
[0023] 1. Generating Prompts for Time Calculation Questions: You are an information extraction and problem generation expert. Your task is to generate problems involving complex materials and "time calculations." Please strictly adhere to the following requirements when generating output: [Task Description] Please write a descriptive passage of no less than 500 words. The passage should involve a person, organization, device, or animal performing multiple events at different times (at least 3 time points, involving more than 4 behaviors). These events should have a clear chronological order and changes in location or context. The background can include, but is not limited to, economic activities, disaster relief, historical events, etc.
[0024] Please make the material as detailed as possible, such as by adding background information, story information, etc. to expand the content.
[0025] Following the material, please formulate a "time calculation" question based on its content. This question should involve calculations of a certain duration of stay, the interval between two events, etc., and should not provide a direct answer.
[0026] Output Format Please strictly use the following list JSON format for output: [ { "content": "(This is the detailed material you generate. It should be rich in content, have a clear background, and be logically complete.)" "question": "(This is your generated time calculation question; do not include the answer.)" }, ] [Output Example] (Generate output examples based on specific training requirements) [ { (Generate sample content, such as a detailed description of the 2025 National Day holiday.) Question: How many days is the National Day holiday in 2025? }, ] Please start generating five samples with different backgrounds, each sample containing only one set of content and question.
[0027] 2. Summary / Inductive Prompt: You are an information extraction and question generation expert. Your task is to generate complex materials, summarize their content, and formulate questions. Please strictly adhere to the following requirements when generating output: [Task Description] Please write a descriptive passage. The passage should describe a subject (such as a person, organization, equipment, vehicle, etc.) engaging in a series of activities or experiencing a series of events (involving at least three different locations or objects) under different locations, times, or conditions. These activities or events should contain key information points that can be summarized. Background information may include, but is not limited to, economic activities, disaster relief, historical events, etc.
[0028] Please make the material as detailed as possible, including background information, specific details, relevant people, environmental descriptions, etc., so that the material is sufficient to support subsequent summarization.
[0029] Following the provided material, please formulate a summary question based on its content. This question should require you to extract, integrate, and list multiple key information points from the material (e.g., locations, people, items, event characteristics, etc.). The question can appropriately summarize the material or use words that are not entirely identical to the material but have a similar meaning.
[0030] Output Format Please strictly use the following list JSON format for output: [ { "content":"(This is the detailed material you generated)", "question":"(This is the inductive / summary question you generated)" }, ] [Output Example] (Generate output examples based on specific training requirements) [ { (Generate sample content, such as a detailed description of the 2025 National Day holiday.) Question: "What diverse and enriching activities were held across the country during the 2025 National Day holiday?" }, ] Please start generating five samples with different backgrounds, each sample containing only one set of content and question.
[0031] Preferably, the method further includes: Obtain the material set C = (C1, C2, ..., C) corresponding to S. i C j ); where C i For the i-th subset of materials; C i =(C i,1 C i,2 C i,a C i,b ); Analysis C i,a The information units contained therein, resulting in C i,a The corresponding set of information units X = (X1, X2, ..., X...) h , ..., X f ); where h = 1, 2, ..., f; f is C i,a The number of information units contained; X h C i,a The corresponding h-th information unit; According to X, C i,a Add at least one interference information to obtain S. i,a To improve C i,a Information entropy H; the interference information is the C i,a The corresponding background information, and / or information similar to but unrelated to any information unit in X; where H = ;P(X) h ) is X h In C i,a The probability of it appearing in the middle.
[0032] The above-mentioned preferred scheme enhances the generated material to increase its complexity and challenge to the model's capabilities. The concept of information entropy is introduced to quantify information complexity. By adding background information and similar but irrelevant interference information, the information entropy of the material is increased. Adding effective interference information makes the probability distribution more uniform, thereby increasing information entropy and forcing the model to perform deeper semantic understanding rather than shallow matching. Example prompts are as follows: You are a professional content expansion expert. Please enrich and expand the following material to make it more detailed and vivid: Expansion requirements: 1. Add appropriate background information, historical context, or prerequisites. 2. Supplement with necessary details, character traits, or event sequence. 3. Add relevant causal relationship analysis, outcome analysis, or impact assessment. 4. Maintain the coherence and logic of the content. 5. The expanded content should be 50%-100% richer than the original content. Important limitations: - Strictly maintain the core facts and key information of raw materials - Do not change or deviate from the original direction of the problem. - Expanded content must be closely related to the original material; avoid introducing irrelevant topics. - Output the complete and rich post-processing material directly, without including explanations, analyses, or other additional content. Raw materials: {{content}} Original question: {{question}} Please do not output prefixes or suffixes like "(enriched materials:)". Directly output the complete enriched material text.
[0033] Preferably, the step of performing reasoning calculations on the materials and their corresponding questions to obtain the answer for each material and its corresponding question includes: For the S i,a and Q i,a The answer A obtained through reasoning and calculation i,a To obtain the answer set A = (A1, A2, ..., A...) corresponding to S. i A j ); where A i A is a subset of the i-th answer; i = (A i,1 A i,2 A i,a A i,b ).
[0034] The above-mentioned preferred solution inputs the obtained materials and questions back into the large language model, and after detailed thinking and reasoning, generates the answer.
[0035] Preferably, the step of calculating at least one problem based on at least one real-world material according to S includes: Obtain a real-world material set R = (R1, R2, ..., R...) g , ..., R k ); where g = 1, 2, ..., k; k is the preset quantity of materials to be acquired; R g To obtain the g-th material; The R is processed by an information validity filter to obtain a high-quality material set G = (G1, G2, ..., G...). v , ..., G z ); where v = 1, 2, ..., z; z is the quantity of high-quality material obtained; G v This is the vth high-quality material; Based on the given S, G is processed using a Few-shot method to obtain the problem set U = (U1, U2, ..., U...) corresponding to G. v , ..., U z ); where U v For G v The corresponding subset of problems; when according to G v When it is impossible to obtain a problem for any task type corresponding to S, U v Marked as 0; when according to G v When at least one task type corresponding to S can be obtained, U v =(U v,1 U v,2 , ..., U v,e , ..., U v,r e = 1, 2, ..., r; r is U v Number of issues included; U v,e For U v The e-th question included.
[0036] The preferred approach described above first acquires a batch of real-world materials, then filters and selects these materials, removing those with insufficient information, overly simplistic content, or unsuitable for generating complex analytical questions, resulting in a high-quality material set. For the selected high-quality materials, questions are generated using a Few-shot method based on the task quality seed set. In this process, the basic task types can be expanded to more than ten, adding more advanced cognitive tasks such as causal relationship analysis, comparative analysis, semantic understanding, and counterfactual reasoning to further enrich the diversity of the dataset. An example of a prompt for generating questions based on real-world materials is shown below: You are an expert in information extraction and question generation. I will provide you with some material. Based on the characteristics of the material, please freely choose the most suitable question type and generate a targeted question.
[0037] Based on the content of the material, you can choose the most suitable type of question from the following options: 1. **Time Calculation Problems:** These problems involve calculating a certain period of time, the time interval between two events, etc.
[0038] 2. **Statistical Analysis Questions:** These questions require quantitative statistics, proportional comparisons, or trend summaries of different categories, groups, events, and states appearing in the material.
[0039] 3. **Summarization and Conclusion Questions:** These questions require integrating multiple key pieces of information from the material and extracting lists, key points, or common characteristics, such as the locations, people, items, and event attributes involved.
[0040] 4. **Event Comprehension Questions:** These questions require you to understand the background, development process, causes, results, characteristics, or potential impacts of an event or phenomenon as a whole.
[0041] 5. **Information Extraction Questions:** These questions require directly extracting factual information explicitly described in the material, such as a character's specific actions, statements, time points, locations, quantities, etc.
[0042] 6. **Causal Relationship Questions:** These questions require identifying the causal chain described in the material and asking about the causes and consequences or motivations behind the events.
[0043] 7. **Comparative Analysis Questions:** These questions require analyzing the differences or similarities between two or more objects, time periods, or states mentioned in the material.
[0044] 8. **Inference and Judgment Questions:** These questions, based on explicit information in the material, require reasonable reasoning to arrive at an answer, such as relationship judgments.
[0045] 9. **Opinion and Attitude Questions:** These questions address the stance, attitude, or opinion expressed by the individuals, organizations, or groups mentioned in the material.
[0046] 10. **Semantic Comprehension Questions:** These questions test your understanding of the language, descriptions, or expressions used in the material, such as metaphors, allusions, and rhetorical meanings.
[0047] Please strictly follow the following requirements when generating questions: - Based on the content of the material, independently determine the most suitable question type. - Questions should be clear and specific, and can be answered directly based on the materials provided. - Information not mentioned in the material or settings that exceed the scope of reasonable inference are not allowed. - Do not include suggestive words or answer content in the question. - If the material information is insufficient to support any type of question, return "None" and mark the type as 0. The materials are as follows: {content} Please return the results in JSON format, as follows: {{ "question": "The content of the question you generated", "type": The number (1-10) corresponding to the problem type. }} If you cannot determine the appropriate question, please return to: {{ "question": "None", "type": 0 }} Please prioritize generating time calculation problems and statistical analysis problems.
[0048] Preferably, the step of reasoning and calculating based on the real-world materials and their corresponding questions to obtain the answer for each real-world material and its corresponding question includes: For the G v and U v,e The answer E obtained through reasoning and calculation v,e To obtain the answer set E = (E1, E2, ..., E...) corresponding to G. v , ..., E z ); where E v For the v-th answer subset; when U v When marked as 0, E v Marked as 0; when U v When not marked as 0, E v =(E v,1 E v,2 , ..., E v,e , ..., E v,r ).
[0049] The above-mentioned preferred solution inputs the materials and the corresponding generated questions back into the large language model, and after detailed thinking and reasoning, generates the answer.
[0050] Preferably, step S04 includes: Based on different models or different inference paths based on the same model, the S i,a and Q i,aThrough reasoning and calculation, several first-pass answers are obtained; A is determined. i,a Whether the consistency with the aforementioned first verification answers meets the specified threshold; if it does, then A is... i,a This has been identified as the target answer. Based on different models or different inference paths based on the same model, the G v and U v,e By performing reasoning and calculations, several second verification answers are obtained; E is determined. v,e Whether the consistency with the aforementioned second verification answers meets the specified threshold; if it does, then E... v,e This has been identified as the target answer.
[0051] To ensure data quality, the above preferred approach employs multiple different models or uses different inference paths (e.g., Chain-of-Thought, Tree-of-Thought) on the same model to answer the same question. A consistency threshold is set. A data point is only accepted if the consistency of answers from different sources exceeds this threshold. The quality score for a data point is Q. score It can be represented as: Q score (D i = w1·Consistency(A1,A2,...)+w2·Complexity(M,Q); where Consistency is the consistency measure of the answer, Complexity is the complexity assessment of the material and the question, and w1 and w2 are weighting coefficients.
[0052] Preferably, when A is determined i,a E v,e For the target answer, step S05 includes: S i,a Q i,a and A i,a Store the instruction fine-tuning dataset in triplet format; G v U v,e and E v,e Store the instruction fine-tuning dataset in triplet format.
[0053] Accordingly, the present invention provides an embodiment of an instruction fine-tuning dataset generation system, such as... Figure 2 As shown, it includes: Instruction-Driven Generation Module 21: This module is based on a predefined "task instruction seed set," which covers various core capabilities such as time calculation, statistical analysis, inductive summarization, and event understanding. This module adopts a Few-shot learning paradigm to guide the large language model to generate data based on instructions; Dual-path data source processing module 22: This module is used to execute two paths in parallel. Path 1 (Synthesis): Directly generate materials with complex plots and corresponding problems from scratch based on task instructions. This path includes a key material quality improvement step, which increases the complexity and realism of the materials by actively injecting background and distracting information; Path Two (Real): Starting from unprocessed real materials, firstly, a high-value material suitable for generating complex questions is selected through an information validity filter, and then corresponding instruction questions are generated based on these materials; Answer Reasoning and Generation Module 23: Utilize a large language model to conduct in-depth analysis of the generated "material-question" pairs, simulate the thinking process of human experts, and derive logically rigorous and informationally accurate answers; Cross-validation and quality control module 24: Performs multi-dimensional and automated cross-validation on the generated question-answer pairs to ensure the accuracy, consistency and logic of the data, and ultimately selects high-quality datasets.
[0054] Through the collaborative work of the above modules, this system is able to produce high-quality, diverse instruction fine-tuning datasets in parallel and on a large scale.
[0055] Figure 2 The specific implementation process of the system embodiment and Figure 1 The method embodiments described are similar, so the system embodiments are described in a relatively simple way. For relevant details, please refer to the foregoing method embodiments.
[0056] Compared with existing technologies, Figure 1 , Figure 2 The embodiments described above have the following significant advantages and beneficial effects: First, it is highly automated and scalable. From material generation and problem construction to answer reasoning and quality inspection, it minimizes human intervention and can efficiently and massively produce datasets, solving the problems of high cost and low efficiency of traditional methods.
[0057] Secondly, the data is of high quality and challenging. Through a unique "material quality enhancement" step and a rigorous "cross-validation" mechanism, the accuracy, logic, and complexity of the generated data are ensured. In particular, by actively adding interference information to synthetic materials and screening real materials, the dataset can effectively evaluate and train the model's robust reasoning ability in noisy environments, avoiding shallow shortcuts in model learning.
[0058] Third, the task dimensions are rich and controllable. Based on a customizable "task instruction seed set," it can generate questions covering multiple core abilities and can be flexibly expanded to more cognitive dimensions. This design makes the generated dataset highly diverse and balanced in terms of task types, enabling more comprehensive ability training and evaluation of the model.
[0059] Fourth, the generation path operates on a dual-track approach, combining creativity and realism. It uniquely integrates synthetic and real data generation paths. The synthetic path ensures that specific, extreme, or rare logical challenge scenarios can be created as needed; the real data path ensures that the data closely aligns with real-world topics and contexts. The two paths complement each other, greatly enhancing the breadth and depth of the final dataset.
[0060] Embodiments of the present invention also provide a non-transitory computer-readable storage medium that can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a method in the method embodiments, wherein the at least one instruction or the at least one program is loaded and executed by the processor to implement the method provided in the above embodiments.
[0061] Embodiments of the present invention also provide an electronic device, including a processor and the aforementioned non-transitory computer-readable storage medium.
[0062] Embodiments of the present invention also provide a computer program product including program code, which, when the program product is run on an electronic device, causes the electronic device to perform the steps of the methods described above in various exemplary embodiments of the present invention.
[0063] While specific embodiments of the invention have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of the invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the invention. The scope of this invention is defined by the appended claims.
Claims
1. An instruction fine-tuning dataset generation method applied to a large language model, characterized in that, The method includes the following steps: Step S01: Obtain the task instruction seed set S = (S1, S2, ..., S...) i S j ), respectively proceed to steps S02 and S03; where i = 1, 2, ..., j; j is the number of instruction templates contained in S; S i This is the i-th instruction template; each instruction template corresponds to a task type; each instruction template is used to generate the materials and questions corresponding to its task type; Step S02: Obtain at least one material and its corresponding problem generated based on S, and obtain the set of materials and problems corresponding to S, W = (W1, W2, ..., W...). i ,…,W j ); where W i According to S i The generated material and problem subset; W i =(W i,1 W i,2 ,…,W i,a ,…,W i,b ); a = 1, 2, ..., b; b is S i Corresponding preset materials and number of questions; W i,a For W i The a-th material and question; W i,a =(C i,a Q i,a ); C i,a For W i The a-th material in the middle; Q i,a C i,a The corresponding problem is: to obtain the material set C = (C1, C2, ..., C...) corresponding to S. i C j ); where C i For the i-th subset of materials; C i =(C i,1 C i,2 C i,a C i,b ); Analysis C i,a The information units contained therein, resulting in C i,a The corresponding set of information units X = (X1, X2, ..., X...) h , ..., X f ); where h = 1, 2, ..., f; f is C i,a The number of information units contained; X h C i,a The corresponding h-th information unit; according to X to C i,a Add at least one interference information to obtain S. i,a To improve C i,a Information entropy H; for the S i,a and Q i,a The answer A is obtained through reasoning and calculation. i,a To obtain the answer set A = (A1, A2, ..., A...) corresponding to S. i A j ); where A i A is a subset of the i-th answer; i = (A i,1 A i,2 A,..., A i,a A,..., A i,b ); Step S03: Obtain a set of materials based on the real world, R = (R1, R2, ..., R...). g , ..., R k ); where g = 1, 2, ..., k; k is the preset quantity of materials to be acquired; R g To obtain the g-th material; process R through an information validity filter to obtain a high-quality material set G = (G1, G2, ..., G...). v , ..., G z ); where v = 1, 2, ..., z; z is the quantity of high-quality material obtained; G v For the v-th high-quality material; obtain the problem set U = (U1, U2, ..., U...) corresponding to G. v , ..., U z ); where U v For G v The corresponding subset of problems; when according to G v When it is impossible to obtain a problem for any task type corresponding to S, U v Marked as 0; when according to G v When at least one task type corresponding to S can be obtained, U v =(U v,1 U v,2 , ..., U v,e , ..., U v,r e = 1, 2, ..., r; r is U v Number of issues included; U v,e For U v The e-th question included in G; v and U v,e The answer E is obtained through reasoning and calculation. v,e To obtain the answer set E = (E1, E2, ..., E...) corresponding to G. v , ..., E z ); where E v This is a subset of the v-th answer; Step S04: Based on different models or different inference paths based on the same model, perform the following steps on the S... i,a and Q i,a Through reasoning and calculation, several first-pass answers are obtained; A is determined. i,a Whether the consistency with the aforementioned first verification answers meets the specified threshold; if it does, then A is... i,a The target answer is determined; based on different models or different reasoning paths based on the same model, the G is... v and U v,e By performing reasoning and calculations, several second verification answers are obtained; E is determined. v,e Whether the consistency with the aforementioned second verification answers meets the specified threshold; if it does, then E... v,e The target answer is determined; when A is determined... i,a E v,e If the target answer is found, proceed to step S05; Step S05: Place S i,a Q i,a and A i,a Store the instruction fine-tuning dataset in triplet format; store G v U v,e and E v,e Store the instruction fine-tuning dataset in triplet format.
2. The method of claim 1, wherein, The step of obtaining at least one material and its corresponding problem generated based on S, to obtain the set W of materials and problems corresponding to S, includes: By processing S using the Few-shot Prompting technique, we obtain the corresponding material and problem set W.
3. The method of claim 2, wherein, The interference information is C. i,a The corresponding background information, and / or information similar to but unrelated to any information unit in X; wherein, ;P(X) h ) is X h In C i,a The probability of it appearing in the middle.
4. The method according to any one of claims 1 to 3, characterized in that, The process of obtaining the problem set U corresponding to G includes: Based on the S, the G is processed using a Few-shot method to obtain the problem set U corresponding to G.
5. The method of claim 4, wherein, When U v When marked as 0, E v Marked as 0; when U v When not marked as 0, E v =(E v,1 E v,2 , ..., E v,e , ..., E v,r ).
6. A non-transitory computer-readable storage medium storing at least one instruction or at least one program segment, said at least one instruction or said at least one program segment being loaded and executed by a processor to implement the method of any one of claims 1-5.
7. An electronic device, comprising: Includes a processor and the non-transitory computer-readable storage medium of claim 6.