Instruction fine tuning data set generation method, electronic equipment and storage medium
By using task instruction seed sets and multi-path data generation methods, the problems of high cost and low efficiency in generating instruction fine-tuning datasets are solved, resulting in high-quality and diverse datasets that enhance the application capabilities of large language models in complex scenarios.
Patent Information
- Application Number
- CN202511733406.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2045-11-24
AI Technical Summary
Existing methods for generating instruction fine-tuning datasets are costly, inefficient, and of inconsistent quality, making them difficult to scale up and lacking in complexity and diversity, thus failing to meet the needs of complex domains.
By acquiring a set of task instruction seeds, generating materials and problems, and performing inference calculations and cross-validation, combining synthetic materials and real material paths, and utilizing a large language model for in-depth analysis and automated cross-validation, a high-quality and diverse instruction fine-tuning dataset is generated.
It enables low-cost, high-efficiency, and large-scale generation of instruction fine-tuning datasets with both complexity and diversity, significantly improving the comprehensive application capabilities of large language models in complex scenarios and domains.
Smart Images

Figure CN121579077A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, in particular to an instruction fine-tuning data set generation method, an electronic device and a storage medium. BACKGROUND
[0002] In recent years, Large Language Models (LLMs) have made significant progress, but their basic pre-training goal is essentially a sequential "next word prediction", which makes the original LLM itself not capable of directly following complex human instructions or performing specific task dialogues. To bridge this gap, Instruction Tuning technology has emerged. This technology supervises the fine-tuning (SFT) of pre-trained models on a dataset composed of (instruction, output) pairs, thereby teaching the model to understand and execute human intentions, enabling it to handle a wide range of tasks from simple question answering to complex report generation. Therefore, building high-quality, large-scale Instruction Tuning datasets is crucial to improving LLM capabilities.
[0003] Currently, the construction methods of Instruction Tuning datasets mainly fall into two categories: the first is the traditional manual construction method, and the second is the automatic synthetic data generation technology. Although these two methods can solve certain problems, they also have many problems. First, manual construction is costly, as it requires domain experts to manually write and annotate data. Although the quality is guaranteed, the process is time-consuming and labor-intensive, with high costs and difficulty in large-scale production, which cannot meet the demand for large amounts of data for model iteration. Second, the quality of existing datasets is uneven, with few samples related to complex cognitive tasks in publicly available general datasets. Data generated by simple automated methods (such as synonym replacement or sentence rewriting on existing text) often have simple logic and insufficient depth, which can easily lead to the model learning surface patterns rather than real reasoning ability, i.e., "overfitting" to the data format rather than the task itself. Third, the task dimension is single, and existing datasets usually only cover a few task types, lacking comprehensive evaluation and training of the model's overall ability, especially in complex task core abilities such as time calculation, multi-point information integration, statistical analysis, and event trend judgment.
[0004] Therefore, how to generate Instruction Tuning datasets with complexity, diversity, and high quality at low cost, high efficiency, and large scale to meet the needs of complex fields such as network security situation or situation analysis in a certain field is a technical problem that needs to be solved urgently. SUMMARY
[0005] To solve the above technical problems, the technical solution adopted by the present application is: A method for generating instruction fine-tuning datasets, applied to large language models, includes the following steps: Step S01: Obtain the task instruction seed set S = (S1, S2, ..., S...) i S j ), respectively proceed to steps S02 and S03; where i = 1, 2, ..., j; j is the number of instruction templates contained in S; S i This is the i-th instruction template; each instruction template corresponds to a task type; each instruction template is used to generate the materials and / or problems corresponding to its corresponding task type; Step S02: Obtain at least one material and its corresponding question generated based on S; perform reasoning calculations on the material and its corresponding question to obtain the answer corresponding to each material and its corresponding question; Step S03: Calculate at least one real-world material based on S to obtain at least one question; perform reasoning calculations on the real-world material and its corresponding question to obtain the answer corresponding to each real-world material and its corresponding question; Step S04: Cross-validate the answers generated in steps S02 and S03, and determine the answers that pass the validation as the target answers; Step S05: Store each target answer, its corresponding question, and its corresponding material or real-world material into the instruction fine-tuning dataset.
[0006] Furthermore, obtaining at least one material and its corresponding problem generated based on S includes: By processing S using the Few-shot Prompting technique, we obtain the corresponding material and problem set W = (W1, W2, ..., W...). i ,…,W j ); where W i According to S i The generated materials and problem subset; W i =(W i,1 W i,2 ,…,W i,a ,…,W i,b ); a = 1, 2, ..., b; b is S i Corresponding preset materials and number of questions; W i,a For W i The a-th material and question; W i,a =(C i,a Q i,a ); C i,a For W i The a-th material in the middle; Q i,a Ci,a The corresponding question.
[0007] Furthermore, the method also includes: Obtain the material set C = (C1, C2, ..., C) corresponding to S. i C j ); where C i For the i-th subset of materials; C i =(C i,1 C i,2 C i,a C i,b ); Analysis C i,a The information units contained therein, resulting in C i,a The corresponding set of information units X = (X1, X2, ..., X...) h , ..., X f ); where h = 1, 2, ..., f; f is C i,a The number of information units contained; X h C i,a The corresponding h-th information unit; According to X, C i,a Add at least one interference information to obtain S. i,a To improve C i,a Information entropy H; the interference information is the C i,a The corresponding background information, and / or information similar to but unrelated to any information unit in X; where H = ;P(X) h ) is X h In C i,a The probability of it appearing in the middle.
[0008] Furthermore, the step of reasoning and calculating on the materials and their corresponding questions to obtain the answer for each material and its corresponding question includes: For the S i,a and Q i,a The answer A obtained through reasoning and calculation i,a To obtain the answer set A = (A1, A2, ..., A...) corresponding to S. i A j ); where A i A is a subset of the i-th answer; i = (A i,1 A i,2 A i,a A i,b ).
[0009] Furthermore, the calculation based on S on at least one real-world material to obtain at least one problem includes: Obtain a real-world material set R = (R1, R2, ..., R...) g , ..., R k ); where g = 1, 2, ..., k; k is the preset quantity of materials to be acquired; R g To obtain the g-th material; The R is processed by an information validity filter to obtain a high-quality material set G = (G1, G2, ..., G...). v , ..., G z ); where v = 1, 2, ..., z; z is the quantity of high-quality material obtained; G v It is the vth high-quality material; Based on the given S, G is processed using a Few-shot method to obtain the problem set U = (U1, U2, ..., U...) corresponding to G. v , ..., U z ); where U v For G v The corresponding subset of problems; when according to G v When it is impossible to obtain any task type corresponding to S, U v Marked as 0; when according to G v When U can obtain a problem corresponding to at least one task type of S, v =(U v,1 U v,2 , ..., U v,e , ..., U v,r e = 1, 2, ..., r; r is U v Number of issues included; U v,e For U v The e-th question included.
[0010] Furthermore, the step of reasoning and calculating based on the real-world materials and their corresponding questions to obtain the answer for each real-world material and its corresponding question includes: For the G v and U v,e The answer E obtained through reasoning and calculation v,e To obtain the answer set E = (E1, E2, ..., E...) corresponding to G. v , ..., E z ); where E v For the v-th answer subset; when U v When marked as 0, E v Marked as 0; when U v When not marked as 0, E v= (E v,1 , E v,2 , …, E v,e , …, E v,r ).
[0011] Further, the step S04 comprises: based on different models or based on different inference paths of the same model, performing inference calculation on the S i,a and Q i,a to obtain a plurality of first check answers; determining whether the consistency of A i,a with the plurality of first check answers meets a specified threshold, and if so, determining A i,a as the target answer; based on different models or based on different inference paths of the same model, performing inference calculation on the G v and U v,e to obtain a plurality of second check answers; determining whether the consistency of E v,e with the plurality of second check answers meets a specified threshold, and if so, determining E v,e as the target answer.
[0012] Further, when A i,a and E v,e are determined as the target answers, the step S05 comprises: storing S i,a , Q i,a and A i,a in the instruction fine-tuning data set in a triple format; storing G v , U v,e and E v,e in the instruction fine-tuning data set in a triple format.
[0013] A non-transitory computer readable storage medium, the storage medium storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by a processor to implement the foregoing method.
[0014] An electronic device comprising a processor and the foregoing non-transitory computer readable storage medium.
[0015] The present application has at least the following beneficial effects: The present application is based on a task instruction seed set covering multiple core capabilities such as time calculation, statistical analysis, induction and summary, and event understanding, guiding the large prophecy model to generate materials and / or questions according to the instruction template, including generating materials containing complex plots and corresponding questions from scratch, and generating corresponding instruction questions from real-world materials, then using a large language model to analyze the "material-question" in depth, simulating the thinking process of human experts, deducing the answers, and then automatically cross- verifying the generated answers to ensure the accuracy, consistency and logic of the answers, and finally screening out high-quality instruction fine-tuning data sets. The present application can generate high-quality and diversified instruction fine-tuning data sets in parallel and on a large scale to meet the complex task requirements of certain field development trends or situation analysis, thereby significantly improving the comprehensive application ability of large language models in complex scenarios and fields. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0017] Figure 1 A flow chart of an instruction fine-tuning data set generation method provided by an embodiment of the present application is shown in the figure. Figure 2 A system structure diagram of an instruction fine-tuning data set generation system provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present application will be described clearly and completely with reference to the drawings of the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0019] The present application provides an instruction fine-tuning data set generation method embodiment applied to a large language model, such as Figure 1 As shown in the figure, the method comprises the following steps: Step S01: obtaining a task instruction seed set S=(S1, S2, …, S i , …, S j ); respectively entering step S02 and step S03; wherein i=1, 2, …, j; j is the number of instruction templates contained in S; S iThe i-th instruction template; each instruction template corresponds to a task type; each instruction template is used to generate materials and / or questions corresponding to the task type corresponding to the instruction template; the task type includes but is not limited to time calculation, statistical analysis, induction summary, event understanding; Step S02: obtaining at least one material generated based on S and the question corresponding to the material; performing reasoning calculation on the material and the question corresponding to the material to obtain an answer corresponding to each material and the question corresponding to the material; Step S03: performing calculation on at least one real-world-based material according to S to obtain at least one question; performing reasoning calculation on the real-world-based material and the question corresponding to the material to obtain an answer corresponding to each real-world-based material and the question corresponding to the material; Step S04: cross-checking the answers generated in steps S02 and S03, and determining the answers passing the check as target answers; Step S05: storing each target answer, the question corresponding to the target answer, and the material or real-world-based material corresponding to the target answer in an instruction fine-tuning data set.
[0020] Figure 1 The embodiment includes two parallel data generation pipelines, a step S02 corresponding instruction-based synthetic material generation pipeline, and a step S03 corresponding real material-based instruction generation pipeline, based on a task instruction seed set covering time calculation, statistical analysis, induction summary, event understanding and other core capabilities, guiding the large prophecy model to generate materials and / or questions according to the instruction template, including generating materials containing complex plots and the corresponding questions from scratch, and generating corresponding instruction questions from real-world-based materials, then using a large language model to perform in-depth analysis on the "material-question", simulate the thinking process of human experts, deduce the answers, and then automatically cross-verify the generated answers to ensure the accuracy, consistency and logic of the answers, and finally screen out high-quality instruction fine-tuning data sets. Figure 1 The embodiment can generate high-quality and diversified instruction fine-tuning data sets in parallel and on a large scale, meet the needs of complex tasks such as certain field development trends or situation analysis, and significantly improve the comprehensive application ability of large language models in complex scenarios and fields.
[0021] Preferably, the obtaining at least one material generated based on S and the question corresponding to the material comprises: processing S by Few-shot Prompting technology to obtain a material and question set W=(W1, W2, …, W i , …, W j ) corresponding to S; i wherein W iThe generated material and problem subset; W i =(W i,1 W i,2 ,…,W i,a ,…,W i,b ); a = 1, 2, ..., b; b is S i Corresponding preset materials and number of questions; W i,a For W i The a-th material and question; W i,a =(C i,a Q i,a ); C i,a For W i The a-th material in the middle; Q i,a C i,a The corresponding question.
[0022] The above-mentioned preferred scheme utilizes a large language model and, through Few-shot Prompting technology, generates a batch of materials and corresponding questions based on the instruction templates contained in S. Examples of prompts for some task types are shown below.
[0023] 1. Generating Prompts for Time Calculation Questions: You are an information extraction and problem generation expert. Your task is to generate problems involving complex materials and "time calculations." Please strictly adhere to the following requirements when generating output: [Task Description] Please write a descriptive passage of no less than 500 words. The passage should involve a person, organization, device, or animal performing multiple events at different times (at least 3 time points, involving more than 4 behaviors). These events should have a clear chronological order and changes in location or context. The background can include, but is not limited to, economic activities, disaster relief, historical events, etc.
[0024] Please make the material as detailed as possible, such as by adding background information, story information, etc. to expand the content.
[0025] Following the material, please formulate a "time calculation" question based on its content. This question should involve calculations of a certain duration of stay, the interval between two events, etc., and should not provide a direct answer.
[0026] Output Format Please strictly use the following list JSON format for output: [ { "content": "(This is the detailed material you generate. It should be rich in content, have a clear background, and be logically complete.)" "question": "(This is your generated time calculation question; do not include the answer.)" }, ] [Output Example] (Generate output examples based on specific training requirements) [ { (Generate sample content, such as a detailed description of the 2025 National Day holiday.) Question: How many days is the National Day holiday in 2025? }, ] Please start generating five samples with different backgrounds, each sample containing only one set of content and question.
[0027] 2. Summary / Inductive Prompt: You are an information extraction and question generation expert. Your task is to generate complex materials, summarize their content, and formulate questions. Please strictly adhere to the following requirements when generating output: [Task Description] Please write a descriptive passage. The passage should describe a subject (such as a person, organization, equipment, vehicle, etc.) engaging in a series of activities or experiencing a series of events (involving at least three different locations or objects) under different locations, times, or conditions. These activities or events should contain key information points that can be summarized. Background information may include, but is not limited to, economic activities, disaster relief, historical events, etc.
[0028] Please make the material as detailed as possible, including background information, specific details, relevant people, environmental descriptions, etc., so that the material is sufficient to support subsequent summarization.
[0029] Following the provided material, please formulate a summary question based on its content. This question should require you to extract, integrate, and list multiple key information points from the material (e.g., locations, people, items, event characteristics, etc.). The question can appropriately summarize the material or use words that are not entirely identical to the material but have a similar meaning.
[0030] Output Format Please strictly use the following list JSON format for output: [ { "content":"(This is the detailed material you generated)", "question":"(This is the inductive / summary question you generated)" }, ] Output Example (Generate output examples based on specific training requirements) [ { (Generate sample content, such as a detailed description of the 2025 National Day holiday.) Question: "What diverse and enriching activities were held across the country during the 2025 National Day holiday?" }, ] Please start generating five samples with different backgrounds, each sample containing only one set of content and question.
[0031] Preferably, the method further includes: Obtain the material set C = (C1, C2, ..., C) corresponding to S. i C j ); where C i For the i-th subset of materials; C i =(C i,1 C i,2 C i,a C i,b ); Analysis C i,a The information units contained therein, resulting in C i,a The corresponding set of information units X = (X1, X2, ..., X...) h , ..., X f ); where h = 1, 2, ..., f; f is C i,a The number of information units contained; X h C i,a The corresponding h-th information unit; According to X, C i,a Add at least one interference information to obtain S. i,a To improve C i,a Information entropy H; the interference information is the C i,a Corresponding background information, and / or information similar to but unrelated to any information unit in X; where H = ;P(X) h ) is X h In C i,a The probability of it appearing in the middle.
[0032] The above preferred solution enhances the generated material to increase its complexity and challenge the model's capabilities. The concept of information entropy is introduced to quantify the complexity of information, and by increasing background information, similar but irrelevant interference information, etc., the information entropy of the material is increased. Increasing the effective interference information will make the probability distribution more uniform, thereby increasing the information entropy, forcing the model to perform deeper semantic understanding rather than shallow matching. The prompt instruction example is as follows: You are a professional content expansion expert. Please enrich and expand the following material reasonably, making it more detailed and vivid: Expansion requirements: 1. Add appropriate background information, historical environment, or prerequisite conditions 2. Supplement necessary details, character features, or event processes 3. Add relevant causal relationships, result analysis, or impact assessment 4. Maintain the coherence and logic of the content 5. The expanded content should be 50%-100% richer than the original material Important restrictions: - Strictly maintain the core facts and key information of the original material - Do not change or deviate from the investigation direction of the original problem - The expanded content must be closely related to the original material, do not introduce unrelated topics - Directly output the complete and enriched material, do not include explanations, analyses, or other additional content Original material: {{content}} Original problem: {{question}} Please do not output prefixes and suffixes such as (Enriched material:) and directly output the complete and enriched material text.
[0033] Preferably, the reasoning calculation on the material and its corresponding question is performed to obtain the answer corresponding to each material and its corresponding question, including: The reasoning calculation on the S i,a and Q i,a is performed to obtain the answer A i,a , so that the answer set A=(A1,A2,…,A i ,…,A j ) corresponding to S is obtained; where A i is the i-th answer subset; A i =(A i,1 ,A i,2 ,…,A i,a ,…,A i,b )。
[0034] The preferred scheme inputs the obtained material and question into the large language model again, and generates an answer after detailed thinking and reasoning.
[0035] Preferably, the calculating at least one real-world-based material according to the S obtains at least one question, comprising: obtaining a real-world-based material set R = (R1, R2, …, R g , …, R k ); wherein g = 1, 2, …, k; k is a preset number of obtained materials; R g is the gth obtained material; processing the R through an information validity filter to obtain a high-quality material set G = (G1, G2, …, G v , …, G z ); wherein v = 1, 2, …, z; z is a number of obtained high-quality materials; G v is the vth high-quality material; processing the G through a Few-shot mode according to the S to obtain a question set U = (U1, U2, …, U v , …, U z ) corresponding to the G; wherein U v is a question subset corresponding to the G v ; when any task type question corresponding to the S cannot be obtained according to the G v , U v is marked as 0; when at least one task type question corresponding to the S can be obtained according to the G v , U v = (U v,1 , U v,2 , …, U v,e , …, U v,r ); e = 1, 2, …, r; r is a number of questions contained in the U v ; U v,e is the e th question contained in the U v .
[0036] The above preferred scheme first obtains a batch of real-world-based materials, and then filters and selects the materials to filter out materials with too low information quantity, too simple content or not suitable for generating complex analysis questions, to obtain a high-quality material set. For the selected high-quality materials, questions are generated through a Few-shot mode according to a task quality seed set. In this process, the basic task type can be expanded to more than ten, and higher-level cognitive tasks such as causal relationship analysis, comparative analysis, semantic understanding, counterfactual reasoning, etc. are added to further enrich the diversity of the data set. The prompt instructions for generating questions based on real materials are as follows: As an information extraction and question generation expert, I will provide a segment of material and generate a targeted question based on the content's characteristics.
[0037] I can choose the most suitable question type from the following categories based on the material's content: 1. Time calculation questions: involving the calculation of a certain stay time, the time interval between two events, etc.
[0038] 2. Statistical analysis questions: requiring the number of different categories, groups, events, or states in the material to be counted, compared, or trended.
[0039] 3. Inductive summary questions: requiring the integration of multiple key information in the material to extract lists, points, or common characteristics, such as locations, characters, objects, event attributes, etc.
[0040] 4. Event understanding questions: requiring an overall understanding of the background, development process, causes, results, characteristics, or potential impact of a certain event or phenomenon.
[0041] 5. Information extraction questions: requiring the direct extraction of factual information described in the material, such as the specific behavior, speech, time node, location, quantity of a certain character.
[0042] 6. Causal relationship questions: requiring the identification of the causal chain described in the material, asking about the cause and effect between events or behavior motivation.
[0043] 7. Comparative analysis questions: requiring the analysis of differences or similarities between two or more objects, time periods, or states mentioned in the material.
[0044] 8. Reasoning judgment questions: based on the explicit information in the material, asking questions that require reasonable reasoning to answer, such as relationship judgment.
[0045] 9. Viewpoint attitude questions: asking questions about the position, attitude, or opinion expressed by the characters, organizations, or groups in the material.
[0046] 10. Semantic understanding questions: testing the understanding of certain language, description, or expression in the material, such as metaphors, implications, rhetorical meanings, etc.
[0047] Please follow the following requirements to generate the question: - Based on the content of the material, choose the most suitable question type - The question should be clear, specific, and directly answerable based on the material - Do not introduce information not mentioned in the material or assumptions beyond reasonable extrapolation - Do not add suggestive words or answer content in the question - If the material information is insufficient to support any type of question, return "None" and mark the type as 0 The material content is as follows: {content} Please return the result in JSON format as follows: {{ "question": "You generated the question content", "type": The number corresponding to the type of question (1-10) }} If you cannot determine the appropriate question, return: {{ "question": "None", "type": 0 }} Please give priority to generating: time calculation type questions and statistical analysis type questions.
[0048] Preferably, the inference calculation of the real-world based material and its corresponding question to obtain the answer corresponding to each real-world based material and its corresponding question includes: The inference calculation of G v and U v,e results in the answer E v,e , to obtain the answer set E=(E1,E2,…,E v , …, E z ) corresponding to G; wherein E v is the vth answer subset; when U v is marked as 0, E v is marked as 0; when U v is not marked as 0, E v =(E v,1 , E v,2 , …, E v,e , …, E v,r ).
[0049] The above preferred scheme inputs the material and the generated corresponding question into a large language model again, and generates an answer after detailed thinking and reasoning.
[0050] Preferably, the step S04 includes: Based on different models or different inference paths based on the same model, the S i,a and Q i,aThrough reasoning and calculation, several first-pass answers are obtained; A is determined. i,a Whether the consistency with the aforementioned first verification answers meets the specified threshold; if it does, then A is... i,a This has been identified as the target answer. Based on different models or different inference paths based on the same model, the G v and U v,e By performing reasoning and calculations, several second verification answers are obtained; E is determined. v,e Whether the consistency with the aforementioned second verification answers meets the specified threshold; if it does, then E... v,e This has been identified as the target answer.
[0051] To ensure data quality, the above preferred approach employs multiple different models or uses different inference paths (e.g., Chain-of-Thought, Tree-of-Thought) on the same model to answer the same question. A consistency threshold is set. A data point is only accepted if the consistency of answers from different sources exceeds this threshold. The quality score for a data point is Q. score It can be represented as: Q score (D i = w1·Consistency(A1,A2,...)+w2·Complexity(M,Q); where Consistency is the consistency measure of the answer, Complexity is the complexity assessment of the material and the question, and w1 and w2 are weighting coefficients.
[0052] Preferably, when A is determined i,a E v,e For the target answer, step S05 includes: S i,a Q i,a and A i,a Store the instruction fine-tuning dataset in triplet format; G v U v,e and E v,e Store the instruction fine-tuning dataset in triplet format.
[0053] Accordingly, the present invention provides an embodiment of an instruction fine-tuning dataset generation system, such as... Figure 2 As shown, it includes: Instruction-Driven Generation Module 21: This module is based on a predefined "task instruction seed set," which covers various core capabilities such as time calculation, statistical analysis, inductive summarization, and event understanding. This module adopts a Few-shot learning paradigm to guide the large language model to generate data based on instructions; Dual-path data source processing module 22: This module is used to execute two paths in parallel Path one (synthetic): Directly generate materials containing complex plots and corresponding questions from scratch according to task instructions. This path includes a key material quality improvement step, which actively injects background information and interference information to increase the complexity and authenticity of the materials; Path two (real): Starting from unprocessed real materials, first filter out high-value materials suitable for generating complex questions through an information effectiveness filter, and then generate corresponding instruction questions based on these materials; Answer reasoning and generation module 23: Use large language models to perform in-depth analysis on the generated "material-question" pairs, simulate the thinking process of human experts, and deduce logically rigorous and accurate information answers; Cross-validation and quality control module 24: Perform multi-dimensional and automated cross-validation on the generated "question-answer" data pairs to ensure data accuracy, consistency and logic, and finally filter out high-quality data sets.
[0054] Through the cooperative work of the above modules, the system can produce high-quality and diversified instruction fine-tuning data sets in parallel and on a large scale.
[0055] Figure 2 The specific implementation process of the system embodiment is similar to Figure 1 The method embodiment, so the system embodiment is described relatively simply, and please refer to the foregoing method embodiment for related matters.
[0056] Compared with the prior art, Figure 1 , Figure 2 The embodiments have the following significant advantages and beneficial effects: First, highly automated and scaled, from material generation, question construction to answer reasoning and quality inspection, human intervention is minimized, data sets can be efficiently and massively produced, and the problems of high cost and low efficiency of traditional methods are solved.
[0057] Second, the data quality is high and challenging, through the unique "material quality improvement" step and strict "cross-validation" mechanism, the accuracy, logic and complexity of the generated data are ensured. Especially for synthetic materials, interference information is actively added, and real materials are screened, so that the data set can effectively evaluate and train the robust reasoning ability of the model in a noisy environment, avoiding the model to learn shallow shortcuts.
[0058] Third, the task dimension is rich and controllable. Based on the customizable "task instruction seed set", it can generate questions covering multiple core competencies and can be flexibly extended to more cognitive dimensions. This design makes the generated dataset highly diverse and balanced in terms of task type, and can provide more comprehensive ability training and evaluation for the model.
[0059] Fourth, the generation path is parallel with creativity and authenticity. It creatively combines the two generation paths of synthetic data and real data. The synthetic path ensures that specific, extreme or rare logical challenge scenarios can be created as needed. The real material path ensures that the data closely matches the topics and contexts of the real world. The two paths complement each other, greatly improving the breadth and depth of the final dataset.
[0060] Embodiments of the present application also provide a non-transitory computer readable storage medium, which can be arranged in an electronic device to save at least one instruction or at least one program related to a method in the method embodiments. The at least one instruction or the at least one program is loaded and executed by the processor to implement the method provided by the above embodiments.
[0061] Embodiments of the present application also provide an electronic device comprising a processor and the aforementioned non-transitory computer readable storage medium.
[0062] Embodiments of the present application also provide a computer program product comprising program code for causing an electronic device to perform the steps of the method according to the various exemplary embodiments of the present application described above when the program product is run on the electronic device.
[0063] Although some specific embodiments of the present application have been described in detail above, those skilled in the art should understand that the above examples are only for illustration, not for limiting the scope of the present application. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present application. The scope of the present application is defined by the appended claims.
Claims
1. A method for generating instruction fine-tuning datasets, applied to large language models, characterized in that, The method includes the following steps: Step S01: Obtain the task instruction seed set S = (S1, S2, ..., S...) i S j ), respectively proceed to steps S02 and S03; where i = 1, 2, ..., j; j is the number of instruction templates contained in S; S i This is the i-th instruction template; each instruction template corresponds to a task type; each instruction template is used to generate the materials and / or problems corresponding to its corresponding task type; Step S02: Obtain at least one material and its corresponding question generated based on S; perform reasoning calculations on the material and its corresponding question to obtain the answer corresponding to each material and its corresponding question; Step S03: Calculate at least one real-world material based on S to obtain at least one question; perform reasoning calculations on the real-world material and its corresponding question to obtain the answer corresponding to each real-world material and its corresponding question; Step S04: Cross-validate the answers generated in steps S02 and S03, and determine the answers that pass the validation as the target answers; Step S05: Store each target answer, its corresponding question, and its corresponding material or real-world material into the instruction fine-tuning dataset.
2. The method according to claim 1, characterized in that, The process of obtaining at least one material generated based on S and its corresponding problem includes: By processing S using the Few-shot Prompting technique, we obtain the corresponding material and problem set W = (W1, W2, ..., W...). i ,…,W j ); where W i According to S i The generated materials and problem subset; W i =(W i,1 W i,2 ,…,W i,a ,…,W i,b ); a = 1, 2, ..., b; b is S i Corresponding preset materials and number of questions; W i,a For W i The a-th material and question; W i,a =(C i,a Q i,a ); C i,a For W i The a-th material in the middle; Q i,a C i,a The corresponding question.
3. The method according to claim 2, characterized in that, The method further includes: Obtain the material set C = (C1, C2, ..., C) corresponding to S. i C j ); where C i For the i-th subset of materials; C i =(C i,1 C i,2 C i,a C i,b ); Analysis C i,a The information units contained therein, resulting in C i,a The corresponding set of information units X = (X1, X2, ..., X...) h , ..., X f ); where h = 1, 2, ..., f; f is C i,a The number of information units contained; X h C i,a The corresponding h-th information unit; According to X, C i,a Add at least one interference information to obtain S. i,a To improve C i,a Information entropy H; the interference information is the C i,a Corresponding background information, and / or information similar to but unrelated to any information unit in X; where H = ;P(X) h ) is X h In C i,a The probability of it appearing in the middle.
4. The method according to claim 3, characterized in that, The step of reasoning and calculating the materials and their corresponding questions to obtain the answer for each material and its corresponding question includes: For the S i,a and Q i,a The answer A obtained through reasoning and calculation i,a To obtain the answer set A = (A1, A2, ..., A...) corresponding to S. i A j ); where A i A is a subset of the i-th answer; i = (A i,1 A i,2 A i,a A i,b ).
5. The method according to any one of claims 1 to 4, characterized in that, The calculation based on S on at least one real-world material yields at least one problem, including: Obtain a real-world material set R = (R1, R2, ..., R...) g , ..., R k ); where g = 1, 2, ..., k; k is the preset quantity of materials to be acquired; R g To obtain the g-th material; The R is processed by an information validity filter to obtain a high-quality material set G = (G1, G2, ..., G...). v , ..., G z ); where v = 1, 2, ..., z; z is the quantity of high-quality material obtained; G v It is the vth high-quality material; Based on the given S, G is processed using a Few-shot method to obtain the problem set U = (U1, U2, ..., U...) corresponding to G. v , ..., U z ); where U v For G v The corresponding subset of problems; when according to G v When it is impossible to obtain any task type corresponding to S, U v Marked as 0; when according to G v When U can obtain a problem corresponding to at least one task type of S, v =(U v,1 U v,2 , ..., U v,e , ..., U v,r e = 1, 2, ..., r; r is U v Number of issues included; U v,e For U v The e-th question included.
6. The method according to claim 5, characterized in that, The step of reasoning and calculating based on the real-world materials and their corresponding questions to obtain the answer for each real-world material and its corresponding question includes: For the G v and U v,e The answer E obtained through reasoning and calculation v,e To obtain the answer set E = (E1, E2, ..., E...) corresponding to G. v , ..., E z ); where E v For the v-th answer subset; when U v When marked as 0, E v Marked as 0; when U v When not marked as 0, E v =(E v,1 E v,2 , ..., E v,e , ..., E v,r ).
7. The method according to claim 6, characterized in that, Step S04 includes: Based on different models or different inference paths based on the same model, the S i,a and Q i,a Through reasoning and calculation, several first-pass answers are obtained; A is determined. i,a Whether the consistency with the aforementioned first verification answers meets the specified threshold; if it does, then A is... i,a This has been identified as the target answer. Based on different models or different inference paths based on the same model, the G v and U v,e By performing reasoning and calculations, several second verification answers are obtained; E is determined. v,e Whether the consistency with the aforementioned second verification answers meets the specified threshold; if it does, then E... v,e This has been selected as the target answer.
8. The method according to claim 7, characterized in that, When A is determined i,a E v,e For the target answer, step S05 includes: S i,a Q i,a and A i,a Store the instruction fine-tuning dataset in triplet format; G v U v,e and E v,e Store the instruction fine-tuning dataset in triplet format.
9. A non-transitory computer-readable storage medium storing at least one instruction or at least one program segment, said at least one instruction or said at least one program segment being loaded and executed by a processor to implement the method of any one of claims 1-8.
10. An electronic device, characterized in that, Includes a processor and the non-transitory computer-readable storage medium of claim 9.
Citation Information
Patent Citations
Mathematical big language model fine tuning method, system and equipment with cooperation of data enhancement method and prediction enhancement method and medium
CN118014056A
Sample verification method, electronic equipment, storage medium and program product
CN119514689A
Large language model evaluation method and device in data science field and storage medium
CN119578522A
Intelligent question and answer method and device fusing knowledge graph and large model and storage medium
CN119669428A
Data processing method and device
CN119990333A