Lightweight domain instruction fine-tuning data synthesis method and device based on large model

By building a label-free data set and using a large language model to generate problems, logic and answers, combined with low-rank decomposition adapter and fine-tuning loss function, the problem of high cost and poor generalization capabilities of large language models in vertical field data synthesis is solved, and efficient and low-cost data synthesis is achieved.

CN120430303BActive Publication Date: 2025-09-02HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510938346.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-09-02
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

In the prior art, the data synthesis method of large language models in vertical fields has problems such as high cost, limited performance and poor generalization capabilities, especially in terms of data quality and cross-task generalization.

Method used

Build a label-free data set, including multiple task types and multi-domain label-free text data, generate questions, logic and answers through large language models, use loss functions to train data synthesis models, and fine-tune the model through low-rank decomposition adapter and fine-tune loss functions to generate high-quality domain-specific data.

Benefits of technology

Improves the efficiency of data synthesis and cross-task generalization capabilities, reduces costs, and generates high-quality domain-specific data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430303B_ABST
    Figure CN120430303B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for synthesizing data for lightweight domain instruction fine-tuning based on a large model, relating to the technical field of data synthesis for large language models for question answering. The method comprises: inputting a constructed unlabeled dataset into a large language model to obtain a synthesized dataset; training with the synthesized dataset to obtain a trained data synthesis model; fine-tuning the trained data synthesis model with instructions using a low-rank decomposition adapter to obtain a fine-tuned model; inputting a given task type and its related unlabeled data into the trained data synthesis model to generate data with questions, logic, and answers; merging the unlabeled data with the data with questions, logic, and answers to obtain a new synthesized dataset; inputting the new synthesized dataset into the fine-tuned model for evaluation to obtain an evaluation score; and filtering based on the evaluation score to obtain a high-quality dataset. The present invention can improve the efficiency of data synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language model data synthesis for question answering, and in particular to a method and device for synthesizing lightweight domain instruction fine-tuning data based on a large model. Background Art

[0002] Large language models (LLMs) have become powerful tools for solving a wide range of tasks, from natural language processing to complex reasoning. However, for specialized verticals such as law or medicine, LLMs are relatively limited in their capabilities. Leveraging the inherent capabilities of LLMs for domain data synthesis has become a viable solution. However, current LLM-based domain data synthesis methods face significant challenges in practical application. High-performance LLM data synthesis is prohibitively expensive, while lightweight LLM-based synthesis instructions are of poor quality and limited in the types of tasks for which they can be used.

[0003] In recent years, training vertical expert models (LLMs) by synthesizing domain-specific fine-tuning data has significantly improved the performance of LLMs in vertical domains. Existing data synthesis methods fall into two main categories: those not based on unlabeled data and those based on unlabeled data. Methods that do not rely on unlabeled data primarily leverage the parameterized knowledge of large language models to synthesize domain-specific data. This approach avoids the use of domain-specific unlabeled data, but its efficiency is limited by the domain knowledge of the LLM, and the need to generate high-performance large models with large parameter sizes is costly when generating complex tasks. On the other hand, methods based on unlabeled data synthesize data from unlabeled corpora in specific domains. These methods effectively leverage the lexical, syntactic, and stylistic features contained in the domain corpora to generate high-quality domain-specific data. To balance the quality of synthesized data with the cost of data synthesis, approaches are currently exploring the combination of large models with relatively small parameter sizes, such as around 7 billion parameters, and domain-specific unlabeled data. These approaches attempt to strike a balance between cost and performance, but currently still face challenges such as limited task coverage and overly simple generated data. Summary of the Invention

[0004] To address the existing problems of high cost, limited performance, and poor generalization in large-scale model instruction fine-tuning data synthesis technology based on unlabeled data, especially the technical problem that existing methods are difficult to balance in terms of data quality and cross-task generalization, the present invention provides a method and device for synthesizing lightweight domain instruction fine-tuning data based on large models. The technical solution is as follows:

[0005] On the one hand, a method for synthesizing lightweight domain instruction fine-tuning data based on a large model is provided. The method is implemented by a lightweight domain instruction fine-tuning data synthesis device based on a large model. The method includes:

[0006] S1. Construct an unlabeled dataset; the unlabeled dataset includes: unlabeled text data of various task types and multiple fields;

[0007] S2. Input unlabeled text data of various task types and multiple fields into the large language model to obtain data with questions, logic, and answers.

[0008] S3. Synthesize unlabeled text data of multiple task types and multiple fields and the first data with questions, logic, and answers to obtain a first synthetic dataset; use the first synthetic dataset to train with the constructed loss function to obtain a trained data synthesis model;

[0009] S4. Randomly sample unlabeled text data and task types from the unlabeled dataset and input them into the trained data synthesis model to synthesize second data with questions, logic, and answers; merge the randomly sampled unlabeled text data and task types with the second data with questions, logic, and answers to obtain a second synthesized dataset;

[0010] S5. Evaluate the second synthetic data set using the large language model to obtain a first evaluation score; obtain an evaluation data set based on the first evaluation score; and fine-tune the trained data synthesis model based on the evaluation data set using a low-rank decomposition adapter and the constructed fine-tuning loss function to obtain a fine-tuned data synthesis model.

[0011] S6. Input the given task type and the unlabeled text data related to the given task type into the trained data synthesis model to generate third data with questions, logic, and answers; merge the given task type and the unlabeled text data related to the given task type and the third data with questions, logic, and answers to obtain a third merged data set;

[0012] S7. Input the third merged dataset into the fine-tuned data synthesis model for evaluation to obtain a second evaluation score; perform evaluation filtering based on the second evaluation score to obtain a high-quality domain-specific dataset.

[0013] Optionally, the step S2 inputs unlabeled text data of various task types and domains into a large language model to obtain first data having questions, logic, and answers, including:

[0014] Each task type and the unlabeled text data associated with each task type are input into the large language model to obtain the question, logic, and answer, which are expressed by the following formula (1):

[0015] (1)

[0016] in, Indicates the answer; Indicates a problem; Representation logic; Indicates the task type; Represents unlabeled text data related to the task type; Represents a high-performance large language model for synthesizing corresponding answers, questions, and logic based on input unlabeled text data of any type;

[0017] The obtained questions, answers, each task type, and the unlabeled text data associated with each task type are input into the large language model to calculate the missing logic, which is expressed by the following formula (2):

[0018] (2)

[0019] in, Representation logic; Represents a high-performance large language model for synthesizing corresponding logic based on input unlabeled text data, task type, answer, and question.

[0020] Optionally, the constructed loss function is expressed by the following formula (3):

[0021] (3)

[0022] in, Represents the loss function value; Represents the constructed training dataset; represents an untrained data synthesis model; represents the untrained data synthesis model parameters; Indicates the answer to the j-th data; Indicates the problem of the jth data; The logic representing the jth data; Represents the unlabeled text of the jth data; Indicates the task type of the j-th data; Indicates the data sequence number.

[0023] Optionally, the process of inputting the randomly sampled unlabeled text data and task type into the trained data synthesis model in S4 to synthesize the second data with questions, logic and answers is expressed by the following formula (4):

[0024] (4)

[0025] in, express Synthetic answers; express

[0026] Synthesis issues; express The logic of synthesis; represents the trained data synthesis model; Represents unlabeled text data related to the task type; Indicates the task type.

[0027] Optionally, the process of evaluating the second synthetic dataset using the large language model in S5 to obtain a first evaluation score is represented by the following formula (5):

[0028] (5)

[0029] in, represents the first evaluation score; Represents a high-performance large language model for synthesizing evaluation scores based on input unlabeled text data, task type, answer, question, and logic; express Synthetic answers; express Synthesis issues; Represents unlabeled text data related to the task type; express The logic of synthesis; Indicates the task type.

[0030] Optionally, the constructed fine-tuning loss function is expressed by the following formula (6):

[0031] (6)

[0032] in, Represents the fine-tuning loss function value; Represents the constructed self-assessment training dataset; represents a model with a low-rank factorization adapter; Represents the evaluation score of the jth data in the self-evaluation dataset; Represents the answer to the j-th data in the self-assessment dataset; Represents the problem of the jth data in the self-evaluation dataset; Represents the unlabeled text of the jth data in the self-evaluation dataset; The logic representing the jth data in the self-assessment dataset; Indicates the task type of the j-th data in the self-evaluation dataset.

[0033] Optionally, the process of inputting the third merged data set into the fine-tuned data synthesis model for evaluation in S7 to obtain a second evaluation score is represented by the following formula (7):

[0034] (7)

[0035] in, represents the second assessment score; represents the fine-tuned data synthesis model; express Synthetic answers; express Synthesis issues; Represents unlabeled text data related to the task type; express The logic of synthesis; Indicates the task type.

[0036] On the other hand, a device for synthesizing lightweight domain instruction fine-tuning data based on a large model is provided. The device is applied to a method for synthesizing lightweight domain instruction fine-tuning data based on a large model. The device includes:

[0037] A construction unit is used to construct an unlabeled dataset; wherein the unlabeled dataset includes: unlabeled text data of various task types and multiple fields;

[0038] A first acquisition unit is used to input unlabeled text data of multiple different task types and multiple fields into the large language model to obtain first data with questions, logic and answers;

[0039] The second acquisition unit is configured to synthesize unlabeled text data of multiple different task types and fields and the first data having questions, logic, and answers to obtain a first synthetic data set; and to use the first synthetic data set to perform training using the constructed loss function to obtain a trained data synthesis model;

[0040] A second acquisition unit is configured to randomly sample unlabeled text data and task types from the unlabeled dataset and input them into the trained data synthesis model to synthesize second data with questions, logic, and answers; and merge the randomly sampled unlabeled text data and task types with the second data with questions, logic, and answers to obtain a second synthesized dataset;

[0041] a fourth acquisition unit, configured to evaluate the second synthetic data set using the large language model to obtain a first evaluation score; obtain an evaluation data set based on the first evaluation score; and fine-tune the trained data synthesis model based on the evaluation data set using a low-rank decomposition adapter and the constructed fine-tuning loss function to obtain a fine-tuned data synthesis model;

[0042] a fifth acquisition unit, configured to input the given task type and the unlabeled text data related to the given task type into the trained data synthesis model to generate third data having questions, logic, and answers; and merge the given task type and the unlabeled text data related to the given task type with the third data having questions, logic, and answers to obtain a third merged data set;

[0043] The sixth acquisition unit is used to input the third merged data set into the fine-tuned data synthesis model for evaluation to obtain a second evaluation score; and perform evaluation filtering based on the second evaluation score to obtain a high-quality domain-specific data set.

[0044] Optionally, the first acquiring unit is configured to:

[0045] Each task type and the unlabeled text data associated with each task type are input into the large language model to obtain the question, logic, and answer, which are expressed by the following formula (1):

[0046] (1)

[0047] in, Indicates the answer; Indicates a problem; Representation logic; Indicates the task type; Represents unlabeled text data related to the task type; Represents a high-performance large language model for synthesizing corresponding answers, questions, and logic based on input unlabeled text data of any type;

[0048] The obtained questions, answers, each task type, and the unlabeled text data associated with each task type are input into the large language model to calculate the missing logic, which is expressed by the following formula (2):

[0049] (2)

[0050] in, Representation logic; Represents a high-performance large language model for synthesizing corresponding logic based on input unlabeled text data, task type, answer, and question.

[0051] Optionally, the constructed loss function is expressed by the following formula (3):

[0052] (3)

[0053] in, Represents the loss function value; Represents the constructed training dataset; represents an untrained data synthesis model; Represents the parameters of the untrained data synthesis model; Indicates the answer to the j-th data; Indicates the problem of the jth data; The logic representing the jth data; Represents the unlabeled text of the jth data; Indicates the task type of the j-th data; Indicates the data sequence number.

[0054] Optionally, the process of inputting the randomly sampled unlabeled text data and task type into the trained data synthesis model to synthesize the second data with questions, logic and answers is expressed by the following formula (4):

[0055] (4)

[0056] in, express Synthetic answers; express Synthesis issues; express The logic of synthesis; represents the trained data synthesis model; Represents unlabeled text data related to the task type; Indicates the task type.

[0057] Optionally, the process of evaluating the second synthetic dataset using the large language model to obtain the first evaluation score is represented by the following formula (5):

[0058] (5)

[0059] in, represents the first evaluation score; Represents a high-performance large language model for synthesizing evaluation scores based on input unlabeled text data, task type, answer, question, and logic; express Synthetic answers; express Synthesis issues; Represents unlabeled text data related to the task type; express The logic of synthesis; Indicates the task type.

[0060] Optionally, the constructed fine-tuning loss function is expressed by the following formula (6):

[0061] (6)

[0062] in, Represents the fine-tuning loss function value; Represents the constructed self-assessment training dataset; represents a model with a low-rank factorization adapter; Represents the evaluation score of the jth data in the self-evaluation dataset; Represents the answer to the j-th data in the self-assessment dataset; Represents the problem of the jth data in the self-evaluation dataset; Represents the unlabeled text of the jth data in the self-evaluation dataset; The logic representing the jth data in the self-assessment dataset; Indicates the task type of the j-th data in the self-evaluation dataset.

[0063] Optionally, the process of inputting the third merged data set into the fine-tuned data synthesis model for evaluation to obtain a second evaluation score is represented by the following formula (7):

[0064] (7)

[0065] in, represents the second assessment score; represents the fine-tuned data synthesis model; express Synthetic answers; express Synthesis issues; Represents unlabeled text data related to the task type; express The logic of synthesis; Indicates the task type.

[0066] On the other hand, a lightweight domain instruction fine-tuning data synthesis device based on a large model is provided, and the lightweight domain instruction fine-tuning data synthesis device based on a large model includes: a processor; a memory, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, any one of the above-mentioned lightweight domain instruction fine-tuning data synthesis methods based on a large model is implemented.

[0067] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned large-model-based lightweight domain instruction fine-tuning data synthesis methods.

[0068] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0069] The embodiment of the present invention first constructs an unlabeled dataset; wherein the unlabeled dataset includes: unlabeled text data of various task types and multiple fields; the unlabeled text data of various task types and multiple fields are input into a large language model to obtain first data with questions, logic and answers; the unlabeled text data of various task types and multiple fields and the first data with questions, logic and answers are synthesized to obtain a first synthetic dataset; the first synthetic dataset is used to train through the constructed loss function to obtain a trained data synthesis model; secondly, the randomly sampled unlabeled text data and task types are input into the trained data synthesis model to synthesize second data with questions, logic and answers; the randomly sampled unlabeled text data and task types and the second data with questions, logic and answers are merged to obtain a second A synthetic dataset; a large language model is used to evaluate the second synthetic dataset to obtain a first evaluation score; an evaluation dataset is obtained based on the first evaluation score; based on the evaluation dataset, the trained data synthesis model is fine-tuned through a low-rank decomposition adapter and a constructed fine-tuning loss function to obtain a fine-tuned data synthesis model; a given task type and unlabeled text data related to the given task type are input into the trained data synthesis model to generate a third data with questions, logic and answers; finally, the given task type and the unlabeled text data related to the given task type and the third data with questions, logic and answers are merged to obtain a third merged dataset; the third merged dataset is input into the fine-tuned data synthesis model for evaluation to obtain a second evaluation score; evaluation filtering is performed based on the second evaluation score to obtain a high-quality domain-specific dataset.

[0070] This embodiment of the present invention first generates high-quality data by introducing an intermediate inference process and filtering out duplicate data through prompt engineering and recognition of high-frequency vocabulary. Secondly, the trained data synthesis model acquires self-assessment capabilities to score the quality of the synthesized data. Finally, a lightweight domain instruction data synthesis model based on unlabeled text is constructed based on multi-domain data distillation. This embodiment of the present invention can address the high cost, limited performance, and poor generalization capabilities of existing large-model instruction fine-tuning data synthesis technologies for labeled data. This embodiment of the present invention can improve cross-task generalization capabilities and the efficiency of data synthesis. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0072] Figure 1 This is a flow chart of a method for synthesizing lightweight domain instruction fine-tuning data based on a large model provided by an embodiment of the present invention;

[0073] Figure 2 This is a block diagram of a device for synthesizing lightweight domain instruction fine-tuning data based on a large model provided by an embodiment of the present invention;

[0074] Figure 3 It is a structural diagram of a large-model-based lightweight domain instruction fine-tuning data synthesis device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0075] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0076] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "exemplary" in the present invention should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner. Furthermore, in the embodiments of the present invention, "and / or" can mean both or either of the two.

[0077] In the embodiments of the present invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, when the distinction is not emphasized, the meanings they convey are the same. The terms "of," "corresponding," and "corresponding" may sometimes be used interchangeably. It should be noted that, when the distinction is not emphasized, the meanings they convey are the same.

[0078] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.

[0079] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0080] The embodiment of the present invention provides a method for synthesizing lightweight domain instruction fine-tuning data based on a large model. The method can be implemented by a lightweight domain instruction fine-tuning data synthesis device based on a large model. The lightweight domain instruction fine-tuning data synthesis device based on a large model can be a terminal or a server. Figure 1 The flowchart of the method for synthesizing lightweight domain instruction fine-tuning data based on a large model is shown. The processing flow of the method may include the following steps:

[0081] S1. Construct an unlabeled dataset; the unlabeled dataset includes: unlabeled text data of various task types and multiple fields.

[0082] Among them, the various task types include: extractive question answering, natural language reasoning, multi-choice question answering (single / multiple choices), text generation, text summarization, text classification, and natural language understanding.

[0083] In addition, embodiments of the present invention also need to collect existing task data with questions, answers, unlabeled text data and task types, including extractive question answering, natural language reasoning, multiple-choice question answering (single choice) and summary datasets, to enhance question answering diversity and address the challenges of large language models in generating extractive question answering data.

[0084] To enhance the generalization capability of downstream tasks, the present invention adds two new task types: open-book question answering and closed-book question answering. Since the newly added task types allow for customized questions, the two newly added tasks are not limited to specific categories.

[0085] In a feasible implementation, when synthesizing data for a new task type, an embodiment of the present invention designates the newly added task type as closed-book question answering or open-book question answering based on whether the new task requires unlabeled text data as input.

[0086] In a feasible implementation, when two newly added task types are given, the instructions of the new task type can be added to the question as a prefix of the question, so that the data synthesis model trained by the present invention can synthesize instructions that meet specific needs, which can effectively enhance the generalization ability of data synthesis for new field-specific task types.

[0087] Among them, the embodiment of the present invention adds task types including open question answering and closed question answering, further enhancing the adaptability and generalization ability of the data synthesis model to different tasks.

[0088] Among them, unlabeled text data in multiple fields include: unlabeled data in the news field, unlabeled data in the encyclopedia field, unlabeled data in the legal field, and unlabeled data in the medical field.

[0089] S2. Input unlabeled text data of various task types and multiple fields into the large language model to obtain the first data with questions, logic and answers.

[0090] In one feasible implementation, the intermediate reasoning process of the model, i.e., logic, needs to be introduced during the data synthesis process, which can promote a more structured reasoning process and improve the overall data quality.

[0091] Optionally, S2 inputs unlabeled text data of various task types and domains into a large language model to obtain first data with questions, logic, and answers, including:

[0092] Each task type and the unlabeled text data associated with each task type are input into the large language model to obtain the question, logic, and answer, which are expressed by the following formula (1):

[0093] (1)

[0094] in, Indicates the answer; Indicates a problem; Representation logic; Indicates the task type; Represents unlabeled text data related to the task type; Represents a high-performance large language model used to synthesize corresponding answers, questions, and logic based on the input unlabeled text data and task type;

[0095] in, A high-performance large language model for synthesizing corresponding answers, questions, and logic based on input unlabeled text data and task types can be a DeepSeek-V3 model.

[0096] The obtained questions, answers, each task type, and the unlabeled text data associated with each task type are input into the large language model to calculate the missing logic, which is expressed by the following formula (2):

[0097] (2)

[0098] in, Representation logic; Represents a high-performance large language model for synthesizing corresponding logic based on input unlabeled text data, task type, answer, and question.

[0099] In a feasible implementation, a large language model is used to generate high-quality data that includes a logical reasoning process, and intermediate logical reasoning steps are added to data synthesis to improve the rationality and structure of the data.

[0100] S3. Synthesize unlabeled text data of multiple different task types and multiple fields and the first data with questions, logic and answers to obtain a first synthetic data set; use the first synthetic data set to train through the constructed loss function to obtain a trained data synthesis model.

[0101] In one possible implementation, existing methods often rely heavily on unlabeled text data. To generate Yes, this introduces the risk of synthesizing low-relevance data for tasks that don't rely on unlabeled text data, such as multiple-choice or closed-book questions. This embodiment of the present invention employs a relevance-aware data filtering approach. First, through prompt engineering, the model is explicitly guided to avoid synthesizing highly dependent data. Then, data that doesn't meet these criteria is filtered out by identifying prohibited terms such as "context" and "text." This ensures that the generated questions are applicable to downstream tasks with or without unlabeled text data.

[0102] In one feasible implementation, to mitigate potential bias in large language models, an embodiment of the present invention employs word frequency statistics to address potential bias. Based on the first synthetic dataset obtained, for each task type, word frequency statistics are employed to identify the most frequently occurring words, eliminating stop words to obtain a refined dataset. During the identification of the most frequently occurring words, if a particular word appears in more than 10% of the data, stylistic bias may exist. By eliminating keyword inclusion and reducing the probability of keyword appearance, the resulting refined dataset is ensured to be diverse and unbiased. This embodiment of the present invention employs a refined dataset to train a data synthesis model.

[0103] Optionally, the constructed loss function is expressed by the following formula (3):

[0104] (3)

[0105] in, Represents the loss function value; Represents the constructed training dataset; represents an untrained data synthesis model; represents the untrained data synthesis model parameters; Indicates the answer to the j-th data; Indicates the problem of the jth data; The logic representing the jth data; Represents the unlabeled text of the jth data; Indicates the task type of the j-th data; Indicates the data sequence number.

[0106] Among them, the untrained data synthesis model can be one of Qwen2.5-7B-Base or Llama3-8B-Base, which is not limited in this embodiment of the present invention.

[0107] In one feasible implementation, the data synthesis model constructed by the present invention can generate data from domain-specific text for any task. However, the generated data is of low quality due to factors such as the inherent randomness of the model itself. Therefore, the present invention further trains the trained data synthesis model to improve its self-assessment capabilities.

[0108] Among them, the embodiment of the present invention uses a trained data synthesis model to generate new data, which can ensure that the distribution of the synthesized data is consistent with the final generated data.

[0109] Among them, the embodiment of the present invention uses a fine-tuned data synthesis model to synthesize data for the vertical domain large model.

[0110] S4. Randomly sample unlabeled text data and task types from the unlabeled dataset and input them into the trained data synthesis model to synthesize the second data with questions, logic, and answers; merge the randomly sampled unlabeled text data and task types and the second data with questions, logic, and answers to obtain a second synthetic dataset.

[0111] Optionally, the process of inputting the randomly sampled unlabeled text data and task type into the trained data synthesis model in S4 to synthesize the second data with questions, logic and answers is expressed by the following formula (4):

[0112] (4)

[0113] in, express Synthetic answers; express Synthesis issues; express The logic of synthesis; represents the trained data synthesis model; Represents unlabeled text data related to the task type; Indicates the task type.

[0114] S5. Use the large language model to evaluate the second synthetic dataset to obtain a first evaluation score; based on the first evaluation score, obtain an evaluation dataset; based on the evaluation dataset, fine-tune the trained data synthesis model through a low-rank decomposition adapter and the constructed fine-tuning loss function to obtain a fine-tuned data synthesis model.

[0115] In this embodiment of the present invention, the scoring range is defined as 1-5. The second synthetic dataset is evaluated using a large language model to obtain an evaluated dataset. The evaluated dataset is then downsampled to ensure that the sample data for each scoring score does not exceed 2,000, thereby obtaining a final evaluation dataset.

[0116] Among them, the embodiment of the present invention uses a low-rank decomposer to fine-tune the trained data synthesis model so that the model can fine-tune its own data synthesis model based on Generated dataset Give score.

[0117] Optionally, the process of using the large language model to evaluate the second synthetic dataset in S5 to obtain the first evaluation score is expressed by the following formula (5):

[0118] (5)

[0119] in, represents the first evaluation score; Represents a high-performance large language model for synthesizing evaluation scores based on input unlabeled text data, task type, answer, question, and logic; express Synthetic answers; express Synthesis issues; Represents unlabeled text data related to the task type; express The logic of synthesis; Indicates the task type.

[0120] Optionally, the constructed fine-tuning loss function is expressed by the following formula (6):

[0121] (6)

[0122] in, Represents the fine-tuning loss function value; Represents the constructed self-evaluation training dataset; represents a model with a low-rank factorization adapter; Represents the evaluation score of the jth data in the self-evaluation dataset; Represents the answer to the j-th data in the self-assessment dataset; Represents the problem of the jth data in the self-evaluation dataset; Represents the unlabeled text of the jth data in the self-evaluation dataset; The logic representing the jth data in the self-assessment dataset; Indicates the task type of the j-th data in the self-evaluation dataset.

[0123] S6. Input the given task type and the unlabeled text data related to the given task type into the trained data synthesis model to generate third data with questions, logic and answers; merge the given task type and the unlabeled text data related to the given task type and the third data with questions, logic and answers to obtain a third merged data set.

[0124] The process of generating the third data having questions, logic, and answers is expressed by the following formula (7):

[0125] (7)

[0126] in, express Synthetic answers; express Synthesis issues; express The logic of synthesis; Represents the trained data synthesis model.

[0127] Among them, if the task type t does not appear in the first synthetic dataset, the task type t is designated as closed-book question answering or open-book question answering according to whether it requires unlabeled text data as input; the instructions of the new task are added as a prefix before the question.

[0128] To ensure the quality of the final data generated, the embodiment of the present invention adopts a filtering method based on self-assessment.

[0129] S7. Input the third merged dataset into the fine-tuned data synthesis model for evaluation to obtain a second evaluation score; perform evaluation filtering based on the second evaluation score to obtain a high-quality domain-specific dataset.

[0130] Optionally, the process of inputting the third merged data set into the fine-tuned data synthesis model for evaluation in S7 to obtain a second evaluation score is represented by the following formula (8):

[0131] (8)

[0132] in, represents the second assessment score; represents the fine-tuned data synthesis model; express Synthetic answers; express Synthesis issues; Represents unlabeled text data related to the task type; express The logic of synthesis; Indicates the task type.

[0133] During the evaluation process, the model removes data with an evaluation score below 2. If more than 20% of the data scores are less than or equal to 2 points, the model removes data with an evaluation score of 1.

[0134] Among them, the embodiment of the present invention uses high-quality domain-specific datasets to train the target model, which can improve the performance of the target model on domain-specific tasks.

[0135] In a feasible implementation, the embodiment of the present invention is used to perform vertical domain training and testing on the basis of the Qwen2.5-7B-Instruct model, and to perform vertical domain training and testing on the basis of the Llama3-8B-Instruct model; the embodiment of the present invention is used together with the direct inference method, the adaptive training method, the Bonito model, DeepSeek-V3 with self-guidance, DeepSeek-V3 with unlabeled data, and DeepSeek-V3 with self-guidance and unlabeled data to perform vertical domain training and testing on the basis of the Qwen2.5-7B-Instruct model and the Llama3-8B-Instruct model, respectively, to obtain comparative results, as shown in Tables 1 and 2.

[0136] Table 1

[0137]

[0138] Table 2

[0139]

[0140] In a feasible implementation, as shown in Table 1, the data synthesis model trained by the embodiment of the present invention surpasses most baseline methods in average performance, and is comparable to the best settings of DeepSeek-V3 with self-guidance and unlabeled data methods. In addition, adaptive pre-training that relies solely on unlabeled data does not achieve significant improvements in various tasks, highlighting the value of annotated data synthesis. Extractive question answering mainly tests whether synthetic data can effectively learn domain-specific knowledge from unlabeled data. Since DeepSeek-V3 with self-guidance lacks unlabeled data, it cannot follow the learning of domain-specific knowledge from unlabeled data, and its results are marked as NA. In contrast, the data synthesis model trained by the embodiment of the present invention significantly improves the results on extractive question answering, proving that high-performance large-scale language models perform poorly on extractive question answering tasks.

[0141] Among them, the embodiment of the present invention calculates the dollar cost of different data synthesis and training methods. For the DeepSeek-V3 model, its official API is used, and the cost is calculated based on the total data synthesis expenditure. For other sources, a local NVIDIA 4090 24GB GPU is used, and the cost is calculated by the GPU rental price of Vast AI, and the total cost is calculated based on the number of GPU hours used. Among them, the data synthesis model trained by the embodiment of the present invention is comparable to the DeepSeek-V3 with self-guidance and unlabeled data method in performance, but the cost is only 17% of the latter, and better results are achieved than the DeepSeek-V3 with unlabeled data method at 31% of the cost, reflecting the significant efficiency advantage of the data synthesis model trained by the present invention in data synthesis.

[0142] Among them, the Bonito model cannot generate data for the three tasks because it only supports English tasks that rely on unlabeled data, and the results are marked as NA. In contrast, the data synthesis model trained in the embodiment of the present invention achieves better task generalization by introducing Chinese data and defining more flexible task types. Specifically, the data synthesis model trained in the embodiment of the present invention specifies the task type as closed-book / open-book question answering and uses task requirements as question prefixes. Experimental results show that the data synthesis model trained in the embodiment of the present invention continues to surpass various baseline methods in translation and material question answering tasks, highlighting its strong cross-task generalization ability.

[0143] When there is a lack of labeled instruction fine-tuning data for a certain vertical field, the embodiments of the present invention can be used to synthesize labeled instruction fine-tuning data for the vertical field at a relatively low cost based on unlabeled text in the vertical field.

[0144] The embodiment of the present invention first constructs an unlabeled dataset; wherein the unlabeled dataset includes: unlabeled text data of various task types and multiple fields; the unlabeled text data of various task types and multiple fields are input into a large language model to obtain first data with questions, logic and answers; the unlabeled text data of various task types and multiple fields and the first data with questions, logic and answers are synthesized to obtain a first synthetic dataset; the first synthetic dataset is used to train through the constructed loss function to obtain a trained data synthesis model; secondly, the randomly sampled unlabeled text data and task types are input into the trained data synthesis model to synthesize second data with questions, logic and answers; the randomly sampled unlabeled text data and task types and the second data with questions, logic and answers are merged to obtain a second A synthetic dataset; a large language model is used to evaluate the second synthetic dataset to obtain a first evaluation score; an evaluation dataset is obtained based on the first evaluation score; based on the evaluation dataset, the trained data synthesis model is fine-tuned through a low-rank decomposition adapter and a constructed fine-tuning loss function to obtain a fine-tuned data synthesis model; a given task type and unlabeled text data related to the given task type are input into the trained data synthesis model to generate a third data with questions, logic and answers; finally, the given task type and the unlabeled text data related to the given task type and the third data with questions, logic and answers are merged to obtain a third merged dataset; the third merged dataset is input into the fine-tuned data synthesis model for evaluation to obtain a second evaluation score; evaluation filtering is performed based on the second evaluation score to obtain a high-quality domain-specific dataset.

[0145] This embodiment of the present invention first generates high-quality data by introducing an intermediate inference process and filtering out duplicate data through prompt engineering and recognition of high-frequency vocabulary. Secondly, the trained data synthesis model acquires self-assessment capabilities to score the quality of the synthesized data. Finally, a lightweight domain instruction data synthesis model based on unlabeled text is constructed based on multi-domain data distillation. This embodiment of the present invention can address the high cost, limited performance, and poor generalization capabilities of existing large-model instruction fine-tuning data synthesis technologies for labeled data. This embodiment of the present invention can improve cross-task generalization capabilities and the efficiency of data synthesis.

[0146] Figure 2 This is a block diagram of a device for synthesizing lightweight domain instruction fine-tuning data based on a large model according to an exemplary embodiment. The device is used in a method for synthesizing lightweight domain instruction fine-tuning data based on a large model. Figure 2 The device includes a construction unit 210, a first acquisition unit 220, a second acquisition unit 230, a third acquisition unit 240, a fourth acquisition unit 250, a fifth acquisition unit 260 and a sixth acquisition unit 270.

[0147] A construction unit 210 is configured to construct an unlabeled data set, wherein the unlabeled data set includes: unlabeled text data of various task types and multiple fields;

[0148] A first acquisition unit 220 is configured to input unlabeled text data of various task types and fields into a large language model to obtain first data having questions, logic, and answers;

[0149] The second acquisition unit 230 is configured to synthesize unlabeled text data of multiple different task types and fields and the first data having questions, logic, and answers to obtain a first synthetic data set; and to train the first synthetic data set using the constructed loss function to obtain a trained data synthesis model.

[0150] A second acquisition unit 240 is configured to randomly sample unlabeled text data and task types from the unlabeled dataset and input them into the trained data synthesis model to synthesize second data with questions, logic, and answers; and merge the randomly sampled unlabeled text data and task types with the second data with questions, logic, and answers to obtain a second synthesized dataset.

[0151] The fourth acquisition unit 250 is configured to evaluate the second synthetic data set using the large language model to obtain a first evaluation score; obtain an evaluation data set based on the first evaluation score; and fine-tune the trained data synthesis model based on the evaluation data set using a low-rank decomposition adapter and the constructed fine-tuning loss function to obtain a fine-tuned data synthesis model.

[0152] a fifth acquisition unit 260 configured to input the given task type and the unlabeled text data related to the given task type into the trained data synthesis model to generate third data having questions, logic, and answers; and merge the given task type and the unlabeled text data related to the given task type with the third data having questions, logic, and answers to obtain a third merged dataset;

[0153] The sixth acquisition unit 270 is used to input the third merged dataset into the fine-tuned data synthesis model for evaluation to obtain a second evaluation score; and perform evaluation filtering based on the second evaluation score to obtain a high-quality domain-specific dataset.

[0154] Optionally, the first acquiring unit 220 is configured to:

[0155] Each task type and the unlabeled text data associated with each task type are input into the large language model to obtain the question, logic, and answer, which are expressed by the following formula (1):

[0156] (1)

[0157] in, Indicates the answer; Indicates a problem; Representation logic; Indicates the task type; Represents unlabeled text data related to the task type; Represents a high-performance large language model for synthesizing corresponding answers, questions, and logic based on input unlabeled text data of any type;

[0158] The obtained questions, answers, each task type, and the unlabeled text data associated with each task type are input into the large language model to calculate the missing logic, which is expressed by the following formula (2):

[0159] (2)

[0160] in, Representation logic; Represents a high-performance large language model for synthesizing corresponding logic based on input unlabeled text data, task type, answer, and question.

[0161] Optionally, the constructed loss function is expressed by the following formula (3):

[0162] (3)

[0163] in, Represents the loss function value; Represents the constructed training dataset; represents an untrained data synthesis model; Represents the parameters of the untrained data synthesis model; Indicates the answer to the j-th data; Indicates the problem of the jth data; The logic representing the jth data; Represents the unlabeled text of the jth data; Indicates the task type of the j-th data; Indicates the data sequence number.

[0164] Optionally, the process of inputting the randomly sampled unlabeled text data and task type into the trained data synthesis model to synthesize the second data with questions, logic and answers is expressed by the following formula (4):

[0165] (4)

[0166] in, express Synthetic answers; express Synthesis issues; express The logic of synthesis; represents the trained data synthesis model; Represents unlabeled text data related to the task type; Indicates the task type.

[0167] Optionally, the process of evaluating the second synthetic dataset using the large language model to obtain the first evaluation score is represented by the following formula (5):

[0168] (5)

[0169] in, represents the first evaluation score; Represents a high-performance large language model for synthesizing evaluation scores based on input unlabeled text data, task type, answer, question, and logic; express Synthetic answers; express Synthesis issues; Represents unlabeled text data related to the task type; express The logic of synthesis; Indicates the task type.

[0170] Optionally, the constructed fine-tuning loss function is expressed by the following formula (6):

[0171] (6)

[0172] in, Represents the fine-tuning loss function value; Represents the constructed self-assessment training dataset; represents a model with a low-rank factorization adapter; Represents the evaluation score of the jth data in the self-evaluation dataset; Represents the answer to the j-th data in the self-assessment dataset; Represents the problem of the jth data in the self-evaluation dataset; Represents the unlabeled text of the jth data in the self-evaluation dataset; The logic representing the jth data in the self-assessment dataset; Indicates the task type of the j-th data in the self-evaluation dataset.

[0173] Optionally, the process of inputting the third merged data set into the fine-tuned data synthesis model for evaluation to obtain a second evaluation score is represented by the following formula (7):

[0174] (7)

[0175] in, represents the second assessment score; represents the fine-tuned data synthesis model; express Synthetic answers; express Synthesis issues; Represents unlabeled text data related to the task type; express The logic of synthesis; Indicates the task type.

[0176] The embodiment of the present invention first constructs an unlabeled dataset; wherein the unlabeled dataset includes: unlabeled text data of various task types and multiple fields; the unlabeled text data of various task types and multiple fields are input into a large language model to obtain first data with questions, logic and answers; the unlabeled text data of various task types and multiple fields and the first data with questions, logic and answers are synthesized to obtain a first synthetic dataset; the first synthetic dataset is used to train through the constructed loss function to obtain a trained data synthesis model; secondly, the randomly sampled unlabeled data and task types are input into the trained data synthesis model to synthesize second data with questions, logic and answers; the randomly sampled unlabeled text data and task types and the second data with questions, logic and answers are merged to obtain a second synthetic dataset. into a dataset; use a large language model to evaluate the second synthetic dataset to obtain a first evaluation score; based on the first evaluation score, obtain an evaluation dataset; based on the evaluation dataset, fine-tune the trained data synthesis model through a low-rank decomposition adapter and the constructed fine-tuning loss function to obtain a fine-tuned data synthesis model; input the given task type and the unlabeled text data related to the given task type into the trained data synthesis model to generate a third data with questions, logic and answers; finally, merge the given task type and the unlabeled text data related to the given task type and the third data with questions, logic and answers to obtain a third merged dataset; input the third merged dataset into the fine-tuned data synthesis model for evaluation to obtain a second evaluation score; perform evaluation filtering according to the second evaluation score to obtain a high-quality domain-specific dataset.

[0177] This embodiment of the present invention first generates high-quality data by introducing an intermediate inference process and filtering out duplicate data through prompt engineering and recognition of high-frequency vocabulary. Secondly, the trained data synthesis model acquires self-assessment capabilities to score the quality of the synthesized data. Finally, a lightweight domain instruction data synthesis model based on unlabeled text is constructed based on multi-domain data distillation. This embodiment of the present invention can address the high cost, limited performance, and poor generalization capabilities of existing large-model instruction fine-tuning data synthesis technologies for labeled data. This embodiment of the present invention can improve cross-task generalization capabilities and the efficiency of data synthesis.

[0178] Figure 3 This is a structural diagram of a large-model-based lightweight domain instruction fine-tuning data synthesis device provided by an embodiment of the present invention. Figure 3 As shown, the light-weight domain instruction fine-tuning data synthesis device based on the large model may include the above Figure 2 Optionally, the device 310 for synthesizing lightweight domain instruction fine-tuning data based on a large model may include a first processor 2001 .

[0179] Optionally, the large model-based lightweight domain instruction fine-tuning data synthesis device 310 may further include a memory 2002 and a transceiver 2003 .

[0180] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.

[0181] The following combination Figure 3 The components of the large-model-based lightweight domain instruction fine-tuning data synthesis device 310 are described in detail:

[0182] The first processor 2001 is the control center of the large-model-based lightweight domain instruction fine-tuning data synthesis device 310 and can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs) or one or more field programmable gate arrays (FPGAs).

[0183] Optionally, the first processor 2001 can perform various functions of the large model-based lightweight domain instruction fine-tuning data synthesis device 310 by running or executing a software program stored in the memory 2002 and calling data stored in the memory 2002.

[0184] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 3 CPU0 and CPU1 are shown in FIG.

[0185] In a specific implementation, as an embodiment, the large model-based lightweight domain instruction fine-tuning data synthesis device 310 may also include multiple processors, such as Figure 3 1 and 2. The first processor 2001 and the second processor 2004 are shown in FIG. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). A processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0186] The memory 2002 is used to store the software program for executing the solution of the present invention, and is controlled by the first processor 2001 for execution. The specific implementation method can refer to the above method embodiment and will not be repeated here.

[0187] Alternatively, the memory 2002 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, a random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, an optical disc storage (including a compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and capable of being accessed by a computer, but not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently and be connected to the interface circuit ( Figure 3 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0188] The transceiver 2003 is used to communicate with a network device or a terminal device.

[0189] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 3 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.

[0190] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or may exist independently and fine-tune the interface circuit of the data synthesis device 310 through lightweight domain instructions based on a large model ( Figure 3 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0191] It should be noted that Figure 3 The structure of the large model-based lightweight domain instruction fine-tuning data synthesis device 310 shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0192] In addition, the technical effects of the large-model-based lightweight domain instruction fine-tuning data synthesis device 310 can refer to the technical effects of the large-model-based lightweight domain instruction fine-tuning data synthesis method described in the above method embodiment, and will not be repeated here.

[0193] It should be understood that the first processor 2001 in the embodiment of the present invention may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc.

[0194] It should also be understood that the memory in the embodiments of the present invention may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0195] The above embodiments can be implemented in whole or in part via software, hardware (e.g., circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions or computer programs. When loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are fully or partially performed. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0196] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.

[0197] In this disclosure, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0198] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0199] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0200] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0201] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of the device or unit, which can be electrical, mechanical or other forms.

[0202] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0203] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0204] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical disks.

[0205] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A lightweight domain instruction fine-tuning data synthesis method based on a large model, characterized by: The method comprises: S1. Construct an unlabeled dataset; the unlabeled dataset includes: unlabeled text data of various task types and multiple fields; S2. Input unlabeled text data of various task types and multiple fields into the large language model to obtain data with questions, logic, and answers. S3. Synthesize unlabeled text data of multiple task types and multiple fields and the first data with questions, logic, and answers to obtain a first synthetic dataset; use the first synthetic dataset to train with the constructed loss function to obtain a trained data synthesis model; S4. Randomly sample unlabeled text data and task types from the unlabeled dataset and input them into the trained data synthesis model to synthesize second data with questions, logic, and answers; merge the randomly sampled unlabeled text data and task types with the second data with questions, logic, and answers to obtain a second synthesized dataset; S5. Evaluate the second synthetic data set using the large language model to obtain a first evaluation score; obtain an evaluation data set based on the first evaluation score; and fine-tune the trained data synthesis model based on the evaluation data set using a low-rank decomposition adapter and the constructed fine-tuning loss function to obtain a fine-tuned data synthesis model. S6. Input the given task type and the unlabeled text data related to the given task type into the trained data synthesis model to generate third data with questions, logic, and answers; merge the given task type and the unlabeled text data related to the given task type and the third data with questions, logic, and answers to obtain a third merged data set; S7. Input the third merged dataset into the fine-tuned data synthesis model for evaluation to obtain a second evaluation score; perform evaluation filtering based on the second evaluation score to obtain a high-quality domain-specific dataset.

2. The method for synthesizing lightweight domain instruction fine-tuning data based on a large model according to claim 1 is characterized in that: The S2 inputs unlabeled text data of various task types and domains into the large language model to obtain data with questions, logic, and answers, including: Each task type except the extractive question-answering task and the unlabeled text data associated with each task type are input into the large language model to obtain the question, logic, and answer, which are expressed by the following formula (1): (1) in, Indicates the answer; Indicates a problem; Representation logic; Indicates the task type; Represents unlabeled text data related to the task type; Represents a high-performance large language model for synthesizing corresponding answers, questions, and logic based on input unlabeled text data of any type; The obtained questions, answers, each task type, and the unlabeled text data associated with each task type are input into the large language model to calculate the missing logic, which is expressed by the following formula (2): (2) in, Representation logic; Represents a high-performance large language model for synthesizing corresponding logic based on input unlabeled text data, task type, answer, and question.

3. The method for synthesizing lightweight domain instruction fine-tuning data based on a large model according to claim 1 is characterized in that: The constructed loss function is expressed by the following formula (3): (3) in, Represents the loss function value; Represents the constructed training dataset; represents an untrained data synthesis model; represents the untrained data synthesis model parameters; Indicates the answer to the j-th data; Indicates the problem of the jth data; The logic representing the jth data; Represents the unlabeled text of the jth data; Indicates the task type of the j-th data; Indicates the data sequence number.

4. The method for synthesizing lightweight domain instruction fine-tuning data based on a large model according to claim 1 is characterized in that: The process of inputting the randomly sampled unlabeled text data and task type into the trained data synthesis model in S4 to synthesize the second data with questions, logic and answers is expressed by the following formula (4): (4) in, express Synthetic answers; express Synthesis issues; express The logic of synthesis; represents the trained data synthesis model; Represents unlabeled text data related to the task type; Indicates the task type.

5. The method for synthesizing lightweight domain instruction fine-tuning data based on a large model according to claim 1 is characterized in that: The process of using the large language model to evaluate the second synthetic dataset in S5 to obtain the first evaluation score is expressed by the following formula (5): (5) in, represents the first evaluation score; Represents a high-performance large language model for synthesizing evaluation scores based on input unlabeled text data, task type, answer, question, and logic; express Synthetic answers; express Synthesis issues; Represents unlabeled text data related to the task type; express The logic of synthesis; Indicates the task type.

6. The method for synthesizing lightweight domain instruction fine-tuning data based on a large model according to claim 1, characterized in that: The constructed fine-tuning loss function is expressed by the following formula (6): (6) in, Represents the fine-tuning loss function value; Represents the constructed self-assessment training dataset; represents a model with a low-rank factorization adapter; Represents the evaluation score of the jth data in the self-evaluation dataset; Represents the answer to the j-th data in the self-assessment dataset; Represents the problem of the jth data in the self-evaluation dataset; Represents the unlabeled text of the jth data in the self-evaluation dataset; The logic representing the jth data in the self-assessment dataset; Indicates the task type of the j-th data in the self-evaluation dataset.

7. The method for synthesizing lightweight domain instruction fine-tuning data based on a large model according to claim 1 is characterized in that: The process of inputting the third merged data set into the fine-tuned data synthesis model for evaluation in S7 to obtain a second evaluation score is represented by the following formula (7): (7) in, represents the second assessment score; represents the fine-tuned data synthesis model; express Synthetic answers; express Synthesis issues; Represents unlabeled text data related to the task type; express The logic of synthesis; Indicates the task type.

8. A device for synthesizing lightweight domain instruction fine-tuning data based on a large model, wherein the device is used to implement the method for synthesizing lightweight domain instruction fine-tuning data based on a large model according to any one of claims 1 to 7, characterized in that: The device comprises: A construction unit is used to construct an unlabeled dataset; wherein the unlabeled dataset includes: unlabeled text data of various task types and multiple fields; A first acquisition unit is used to input unlabeled text data of multiple different task types and multiple fields into the large language model to obtain first data with questions, logic and answers; The second acquisition unit is configured to synthesize unlabeled text data of multiple different task types and fields and the first data having questions, logic, and answers to obtain a first synthetic data set; and to use the first synthetic data set to perform training using the constructed loss function to obtain a trained data synthesis model; A second acquisition unit is configured to randomly sample unlabeled text data and task types from the unlabeled dataset and input them into the trained data synthesis model to synthesize second data with questions, logic, and answers; and merge the randomly sampled unlabeled text data and task types with the second data with questions, logic, and answers to obtain a second synthesized dataset; a fourth acquisition unit, configured to evaluate the second synthetic data set using the large language model to obtain a first evaluation score; obtain an evaluation data set based on the first evaluation score; and fine-tune the trained data synthesis model based on the evaluation data set using a low-rank decomposition adapter and the constructed fine-tuning loss function to obtain a fine-tuned data synthesis model; a fifth acquisition unit, configured to input the given task type and the unlabeled text data related to the given task type into the trained data synthesis model to generate third data having questions, logic, and answers; and merge the given task type and the unlabeled text data related to the given task type with the third data having questions, logic, and answers to obtain a third merged data set; The sixth acquisition unit is used to input the third merged data set into the fine-tuned data synthesis model for evaluation to obtain a second evaluation score; and perform evaluation filtering based on the second evaluation score to obtain a high-quality domain-specific data set.

9. A lightweight domain instruction fine-tuning data synthesis device based on a large model, characterized in that: The large model-based lightweight domain instruction fine-tuning data synthesis device includes: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program code, which can be called by a processor to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for carrying out sample screening on large language model for questions and answers

    CN117493890A

  • Synthetic data generation using large language models

    US20250156644A1