Data generation method and device, electronic equipment and storage medium
By using the big model and data generation method, QA instruction data is iteratively generated based on reference text documents and prompt word documents, and data that does not meet the requirements are filtered through the principle of pair pairing, which solves the problems of low efficiency and poor accuracy of manual rewriting in the prior art, and realizes efficient and accurate QA instruction data generation.
Patent Information
- Application Number
- CN202510165557.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, there are problems of low efficiency and poor accuracy in manually rewriting professional technical documents into QA instruction data.
A data generation method is adopted, and a large model is used to iterate the problem and answer instruction data collection based on reference text documents, prompt word documents and seed data documents, and filter data whose similarity does not meet the requirements through the pair pairing principle.
It significantly improves the efficiency and accuracy of QA instruction data generation, reduces manual intervention, and improves data diversity and accuracy.
Smart Images

Figure CN120104738A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a data generation method and device, an electronic device and a storage medium. Background Art
[0002] The acquisition of fine-tuning data is the most important basic work in the application of large models in vertical fields (for example, using fine-tuning data to fine-tune large models to obtain privacy leakage detection models). However, the professional technical documents written by experts in this vertical field cannot be directly applied to the fine-tuning of large models. Therefore, it is necessary to write the professional technical documents into instruction data in the format of question answer (QA) (which can be called QA instruction data, or question instruction data and answer instruction data) to fine-tune large models in vertical fields.
[0003] Currently, professional technical documents are manually rewritten into QA instruction data, but this manual rewriting method has the problems of low efficiency and poor accuracy. Summary of the invention
[0004] In order to overcome the deficiencies of the prior art, the present application provides a data generation method and device, an electronic device and a storage medium to improve the generation efficiency and accuracy of QA instruction data.
[0005] The embodiment of the present application provides a data generation method, which includes: obtaining a reference text document, a first prompt word document, and a seed data document; the reference text document is used to record the knowledge in the specified technical field; the first prompt word document includes a plurality of first prompt words for describing the content generation requirements of the generated question instruction data; the seed data document includes a plurality of reference formats for limiting the generated question instruction data. In the current generation round, the obtained large model is used to generate a question instruction data set under the current generation round based on the reference text document, the first prompt word document, and the seed data document; and the question instruction data under the current generation round is filtered based on the pairwise pairing principle to obtain the reserved question instruction data set under the current generation round. The obtained large model is used to generate an answer instruction data set under the current generation round based on the reference text document, the second prompt word document, and the reserved question instruction data set under the current generation round; wherein the second prompt word document includes a plurality of second prompt words for describing the content generation requirements of the generated answer instruction data. Iteratively execute generation rounds, and determine the retained question instruction data set under all generation rounds when the round iteration stopping condition is met, and the answer instruction data set under all generation rounds as the question instruction data set of the reference text document, and the question instruction data set of the reference text document, respectively.
[0006] The embodiment of the present application also provides a data generation device, which includes: a configuration module, which is used to obtain a reference text document, a first prompt word document, and a seed data document; the reference text document is used to record the knowledge in the specified technical field; the first prompt word document includes a plurality of first prompt words for describing the content generation requirements of the generated question instruction data; the seed data document includes a plurality of reference formats for limiting the generated question instruction data. The first generation module is used to generate a question instruction data set under the current generation round by using the obtained large model, based on the reference text document, the first prompt word document, and the seed data document; and perform filtering operations on the question instruction data under the current generation round based on the pairwise pairing principle to obtain the reserved question instruction data set under the current generation round. The second generation module is used to generate an answer instruction data set under the current generation round by using the obtained large model, based on the reference text document, the second prompt word document, and the reserved question instruction data set under the current generation round; wherein the second prompt word document includes a plurality of second prompt words for describing the content generation requirements of the generated answer instruction data; wherein the second prompt word document includes a plurality of content generation requirements for describing the generated answer instruction data. The determination module, in which the user iteratively executes generation rounds, determines the retained question instruction data set under all generation rounds when the round iteration stopping condition is met, and the answer instruction data set under all generation rounds as the question instruction data set of the reference text document, and the question instruction data set of the reference text document, respectively.
[0007] An embodiment of the present application also provides an electronic device, comprising: a processor and a computer-readable storage medium for storing computer program instructions, wherein the computer program instructions, when executed by the computer-readable storage medium, enable the processor to execute the steps of the above method.
[0008] An embodiment of the present application also provides a machine-readable storage medium, which stores computer program instructions. When the computer program instructions are executed, the steps of the above method can be implemented.
[0009] In this embodiment, after obtaining the reference text document, the first prompt word document, and the seed data document, in the current generation round, the obtained large model is used to generate a question instruction data set for the current generation round based on the reference text document, the first prompt word document, and the seed data document, and a filtering operation is performed on the question instruction data for the current generation round based on the pairing principle to obtain a retained question instruction data set for the current generation round, and the obtained large model is used to generate an answer instruction data set for the current generation round based on the reference text document, the second prompt word document, and the retained question instruction data set for the current generation round. Since this embodiment only needs to use a small amount of seed data documents to automatically generate a large amount of question instruction data and answer instruction data using the large model, that is, the reference text document can be automatically converted into QA instruction data using the large model, which can significantly improve the generation efficiency of QA instruction data compared to manual changes.
[0010] In addition, in this embodiment, after obtaining the question instruction data set under the current generation round, a filtering operation is performed on the question instruction data under the current generation round based on the pairing principle to obtain the retained question instruction data set under the current generation round. Since the question instruction data that do not meet the similarity requirements are filtered out, this provides a favorable basis for the subsequent generation of accurate answer instruction data, that is, this can improve the accuracy of the generation of QA instruction data.
[0011] In addition, compared with conventional filtering methods, this screening method based on the pairing principle can reduce the amount of data transmission during the screening process, thereby improving screening efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a flow chart of the data generation method provided in an embodiment of the present application.
[0013] Figure 2a , Figure 2b , Figure 2c It is a schematic diagram of a reference text document, a first prompt word document, and a seed data document provided in an embodiment of the present application.
[0014] Figure 3 It is a flowchart of the data generation method provided in an embodiment of the present application.
[0015] Figure 4 It is a structural diagram of a data generating device provided in an embodiment of the present application.
[0016] Figure 5 A structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0017] The preferred embodiments of the present application are described in detail below in conjunction with the accompanying drawings so that the advantages and features of the present application can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present application.
[0018] See also Figure 1 , Figure 1 is a flow chart of a data generation method provided in an embodiment of the present application. Optionally, the method is applied in a server. Figure 1 As shown, the detection method includes the following steps.
[0019] S101, obtaining a reference text document, a first prompt word document, and a seed data document.
[0020] In this embodiment, the reference text document is used to record the knowledge in a specified technical field. The reference text document is the professional technical document mentioned in the background technology, which is used to describe the knowledge in a certain vertical field. Figure 2a The reference text document is shown as an example, of course, the reference text document is not limited to Figure 2a The form shown in .
[0021] The first prompt word document includes a plurality of first prompt words for describing the content generation requirements for generating question instruction data. These prompt words require the large model to write out {} different question instruction data line by line around the above reference text document, imitating the reference format, and strictly prohibiting the output of any other explanatory text.
[0022] For example, content generation requirements may include the following:
[0023] 1. Each question instruction data should not have repeated verbs to maximize the diversity of question instruction data.
[0024] 2. The tone of voice used in question-based instructions needs to be diversified. For example, you can combine question-based instructions with imperative instructions.
[0025] 3. The types of question instruction data also need to be diversified, that is, the types of question instruction data should include different types of tasks, such as open generation, classification, editing, etc.
[0026] Figure 2b The first prompt word document is shown as an example. Of course, the content generation requirements in the first prompt word document are not limited to Figure 2b The content generation requirements shown in .
[0027] The seed data document includes multiple reference formats for limiting the generation of question instruction data. The reference format is the object that the large model needs to imitate. The large model needs to imitate the reference format to generate question instruction data. The reference format is mainly used to limit the format of question instruction data, such as being more colloquial or more written. The content of the seed data may be irrelevant to the reference text, and the reference text limits the scope of the content of the question.
[0028] Figure 2c The seed data document is shown as an example. Of course, the reference format in the seed data document is not limited to Figure 2c The content generation requirements shown in .
[0029] It should be noted that the reference text document, the first prompt word document, and the seed data document are all manually written.
[0030] S102, in the current generation round, using the obtained large model, based on the reference text document, the first prompt word document, and the seed data document, a question instruction data set for the current generation round is generated; and based on the pairwise pairing principle, a filtering operation is performed on the question instruction data for the current generation round to obtain a retained question instruction data set for the current generation round.
[0031] It should be noted that this large model includes but is not limited to GPT, etc. The large model can be set in the cloud as long as it can be called, and the embodiment of the present application is not specifically limited.
[0032] S103, using the obtained large model, based on the reference text document, the second prompt word document, and the retained question instruction data set under the current generation round, generate an answer instruction data set under the current generation round.
[0033] The second prompt word document includes a plurality of second prompt words for describing content generation requirements for generating answer instruction data. The second prompt word document is different from the first prompt word document. Although both describe content generation requirements, one is a requirement for generating question instruction data, and the other is a requirement for generating answer instruction data, and the two are not the same.
[0034] S104, iteratively execute generation rounds, and determine the retained question instruction data set under all generation rounds when the round iteration stopping condition is met, and the answer instruction data set under all generation rounds as the question instruction data set of the reference text document, and the question instruction data set of the reference text document, respectively.
[0035] In this embodiment, when the generated question instruction data corresponding to the reference text document reaches a preset number, such as 500 question instruction data, it is determined that the round iteration stop condition is met. Otherwise, the iteration execution continues.
[0036] The specific implementation of the above steps S102 to S104 will be described later and will not be repeated here.
[0037] So far, completed Figure 1 The process shown.
[0038] pass Figure 1 It can be seen from the process shown that in this embodiment, after obtaining the reference text document, the first prompt word document, and the seed data document, in the current generation round, the obtained large model is used to generate a question instruction data set for the current generation round based on the reference text document, the first prompt word document, and the seed data document, and a filtering operation is performed on the question instruction data for the current generation round based on the pairing principle to obtain a reserved question instruction data set for the current generation round, and the obtained large model is used to generate an answer instruction data set for the current generation round based on the reference text document, the second prompt word document, and the reserved question instruction data set for the current generation round. Since this embodiment only needs to use a small amount of seed data documents to automatically generate a large amount of question instruction data and answer instruction data using the large model, that is, the reference text document can be automatically converted into QA instruction data using the large model, which can significantly improve the generation efficiency of QA instruction data compared to manual changes.
[0039] In addition, in this embodiment, after obtaining the question instruction data set under the current generation round, a filtering operation is performed on the question instruction data under the current generation round based on the pairing principle to obtain the retained question instruction data set under the current generation round. Since the question instruction data that do not meet the similarity requirements are filtered out, this provides a favorable basis for the subsequent generation of accurate answer instruction data, that is, this can improve the accuracy of the generation of QA instruction data.
[0040] In addition, compared with conventional filtering methods, this screening method based on the pairing principle can reduce the amount of data transmission during the screening process, thereby improving screening efficiency.
[0041] Combine the following Figure 3 The above steps S102 to S104 are described in detail:
[0042] The specific implementation method of using the large model to generate the answer instruction data set for the current generation round in step S102 is described in detail below:
[0043] In one embodiment, in the current generation round, N question generation tasks are executed in parallel, for example, 20 question generation tasks are executed in parallel. When executing each question generation task, based on the first prompt word document and the seed data document, the first prompt word and the reference format matching the question generation task are determined, and the reference text document, the matching first prompt word and the reference format are input into the obtained large model. For example, as a specific implementation method, the reference text document and the matching first prompt word and the reference format can be spliced into a document in the order of the reference text document, the first prompt word and the reference format, and the generate function is called to ask questions to the large model through the OpenAI application programming interface (Application Programming Interface, API) to obtain the question instruction data set under the large model generated question generation task; N is greater than 1. Among them, the question instruction data set of each of the N question generation tasks generated by the large model is the question instruction data set under the current generation round.
[0044] In the specific implementation, the main program instantiates a queue pipe for communication between the main thread and the worker thread, and then instantiates multiple worker threads (one worker thread is used to execute a problem generation task), and concurrently calls the generate function N times, thereby realizing the parallel execution of N problem generation tasks, which can significantly improve the generation efficiency of problem instruction data.
[0045] Specifically, the following rules need to be followed to determine the first prompt word and reference format that match the question generation task:
[0046] Regardless of the question generation task in which generation round, P first prompt words are randomly selected from the first prompt word file, and the P first prompt words are determined as matching content generation requirements; P is greater than 1.
[0047] If the current generation round is the first generation round, K reference formats are randomly selected from the seed data document, and the K reference formats are determined as matching reference formats; K is greater than 1. If the current generation round is not the first generation round, KL reference formats are randomly selected from the seed data document, L question instruction data are selected from the reserved question instruction data set under the previous historical generation round of the current generation round as L reference formats, and the KL reference formats and the L reference formats are determined as matching reference formats.
[0048] For example, in the first generation round, there is only a manually compiled seed data document. For each problem generation task, K reference formats are randomly selected from the manually compiled seed data document and input into the large model, so that the large model generates according to these K reference formats. Starting from the second generation round, there is no seed data document, and there is a reserved question instruction data set generated in the previous round. Therefore, KL reference formats are randomly selected from the seed data document, and L question instruction data are selected from the reserved question instruction data set under the previous historical generation round of the current generation round as L reference formats, thereby obtaining K reference formats.
[0049] Through the above method, the repetition rate of the retained question instruction data set under the question generation tasks with different contents can be reduced as much as possible, thereby improving the diversity of the question instruction data.
[0050] In this embodiment, through the above method, not only the generation rate of question instruction data can be improved, but also the diversity of question instruction data can be improved.
[0051] The specific implementation method of the above step S102 of using the large model to generate the answer instruction data set for the current generation round is described in detail above.
[0052] The specific implementation of the filtering operation in the above step S103 is described in detail below.
[0053] In one embodiment, if the current generation round is the first generation round, every two problem instruction data in the N problem instruction data sets generated by the current generation round are determined as an instruction pair. For each instruction pair, the similarity value of the two problem instruction data in the instruction pair is calculated, and a similarity matrix is generated based on the similarity values corresponding to each instruction pair. Reference problem instruction data is selected from the N problem instruction data sets, and other reference problem instruction data except the reference problem instruction data are traversed in a set order, and the currently traversed reference problem instruction data is determined as the current reference problem instruction data. The similarity value corresponding to the instruction corresponding to the reference problem instruction data and the current reference problem instruction data is queried from the similarity matrix. If the queried similarity value is greater than the set similarity threshold, the current reference problem instruction data is deleted, and if the queried similarity value is less than or equal to the set similarity threshold, the current reference problem instruction data is retained.
[0054] In another embodiment, if the current generation round is not the first generation round, each two problem instruction data in the N problem instruction data sets generated by the current generation round and the reserved problem instruction data set corresponding to the previous generation round of the current generation round are determined as an instruction pair. For each instruction pair, the similarity value of the two problem instruction data in the instruction pair is calculated, and a similarity matrix is generated based on the similarity values corresponding to each instruction pair. Reference problem instruction data is selected from the N problem instruction data sets, and other reference problem instruction data except the reference problem instruction data are traversed in a set order, and the reference problem instruction data currently traversed is determined as the current reference problem instruction data. The similarity value corresponding to the instruction corresponding to the reference problem instruction data and the current reference problem instruction data is queried from the similarity matrix. If the queried similarity value is greater than the set similarity threshold, the current reference problem instruction data is deleted, and if the queried similarity value is less than or equal to the set similarity threshold, the current reference problem instruction data is retained.
[0055] In the specific implementation, specifically, after the problem instruction data set under the problem generation task is generated from the large model, the format of each problem instruction data is judged to determine whether the format of the generated problem instruction data matches the reference format, the unmatched ones are processed, and the matched problem instruction data are passed to the main thread through the queue pipeline for aggregation, and added to the total instruction text library in list format.
[0056] Furthermore, considering that there is no logical dependency between the similarity calculation and the problem instruction data generation, and there is no problem of mutual occupation of resources, the two tasks are run in parallel. In addition, similarity calculation is the main source of program performance bottlenecks, and the main performance optimization work is carried out here. Similarity calculation mainly involves the following functions: processor, load_model, generate_combinations, judge_text, bert_similarity.
[0057] Based on this, in the initial stage of program operation for generating problem instruction data, the program first instantiates a task queue pipeline and a result queue pipeline for communication between the main thread and the worker thread. Then the main program instantiates multiple worker threads by calling the Processer function multiple times, and instantiates a Bert model in each thread. It waits for the main thread to pass the specific calculation task through an infinite loop, and returns the calculation result to the main thread through the result queue pipeline, avoiding the performance problems caused by repeatedly loading and unloading the Bert model in each round of similarity calculation.
[0058] In each generation round, the main program calls the generate_combinations function to combine all the different indexes in pairs with the indexes of the problem instruction data generated in the previous round in the total library and the indexes of the saved problem instruction data in the total library, and stores them in frozen set format. Frozen sets are a subset of the set data structure. While satisfying the unordered storage of internal elements, they also have the property of immutable internal elements and can be used as dictionary keys.
[0059] In this way, the program avoids repeated operations on the lower triangular matrix and the upper triangular matrix in the similarity matrix, that is, the similarity between text No. 1 and text No. 2 must be equal to the similarity between text No. 2 and text No. 1; avoids the calculation of diagonal elements in the similarity matrix, that is, the similarity between text No. 1 and text No. 1 must be 1; avoids unnecessary similarity calculations between highly similar elements and remaining elements, that is, if the retained elements are 0, 1, 2, and element No. 3 is highly similar to element No. 1, then there is no need to calculate the similarity between element No. 3 and the remaining elements in the subsequent program, which reduces the amount of computational complexity of similarity calculations by about 95% and significantly simplifies the computational requirements.
[0060] After all the indexes are combined, the main program will group and package the remaining problem instruction data according to the set number in the format of a two-dimensional array. Each row of the two-dimensional array is a task package, and each element is an index pair stored in a frozen set format. The main program then passes each package to the queue pipeline to wait for processing by the worker thread.
[0061] After the worker thread receives the package, it processes the task by calling the judge_text function. The judge_text function first unpacks the passed task, reads the actual text through the remove_list function of the tool function module, and performs text preprocessing to remove the inactive text. The processed text is then packaged again and passed to the GPU using the bert_similarity function to calculate the semantic similarity. After obtaining the result of the similarity calculation, the function returns the text pairs and similarities in the frozen set format as the keys and values of the dictionary in the list order to the main thread. In this way, the program significantly reduces the performance loss caused by the communication between the CPU and GPU, and obtains a performance improvement of about 20 times.
[0062] In this embodiment, a filtering operation is performed on the N question instruction data sets generated in the current generation round based on the pairing principle to remove question instruction data whose similarity does not meet the set requirements, and obtain the retained question instruction data set under the current generation round.
[0063] The specific implementation of the filtering operation in the above step S103 is described in detail above.
[0064] The specific implementation method of using the large model to generate the answer instruction data set for the current generation round in the above step S104 is described in detail below:
[0065] In one embodiment, for the reserved question instruction data set under the current generation round, M answer generation tasks are executed in parallel. When executing each answer generation task, based on the second prompt word document, the second prompt word matching the answer generation task is determined, and the reference text document, the matching second prompt word, and a reserved question instruction data are input into the large model to obtain the answer instruction data set of the reserved question instruction data generated by the large model. In specific implementation, the reference text document, the matching second prompt word, and a reserved question instruction data are spliced into a calling function to ask questions to the large model through the OpenAI API to generate the answer instruction data of the question instruction data. And the answer instruction data set corresponding to each reserved question instruction data is determined as the answer instruction data set under the current generation round; M is greater than 1.
[0066] In a further embodiment, when executing each answer generation task, after requesting the execution of the answer generation task to the large model, the index corresponding to the answer generation task is added to the answer generation task index set; the index is determined according to the index coordinates of the reserved question instruction data carried by the answer generation task in the two-dimensional generation matrix; the two-dimensional generation matrix is established according to the reference text document and the reserved question instruction data corresponding to the reference text document;
[0067] For each answer instruction data that retains the question instruction data in the answer generation task obtained from the big model, if the answer instruction data meets the set format requirements, the index coordinate corresponding to the answer generation task is used as the key key, and the answer instruction data is used as the value value to form a kye-value for the answer generation task in the dictionary; wherein the dictionary and the collection user track whether the answer generation task is successfully executed.
[0068] Furthermore, after the current generation round is completed, each index in the collection will be initialized as the key of the dictionary to update the dictionary, and the updated dictionary is used to determine the retained question instruction data carried in the answer generation task in the next generation round.
[0069] In the specific implementation, the above steps are composed of the generate_data function, which is responsible for managing the progress of answer instruction data generation at the current moment. This function instantiates multiple generate tasks through multi-threading according to the set number of threads and concurrently requests the cloud large model API. For this module, since the cloud request may not be successful, and the successful request may not guarantee that the obtained data meets the required format requirements, a dictionary and a collection are used in the module to track the completion effect of the generate task. Specifically, the collection is used to track the task index submitted for the current round of answer generation tasks, and the dictionary is used to track the successful task index among all tasks. That is, in each round of answer generation tasks, the collection is first initialized as the key of the current dictionary, and the index of the current task is added to the collection for each task submitted. After all threads are completed, the process_results function of the text processing module is used to judge one by one. If the format is correct, the corresponding index and value are added to the dictionary respectively, so as to realize the monitoring of the progress of answer instruction data generation at the current moment.
[0070] In one embodiment, considering that the generation of answer instruction data (i.e., the answer generation task) is a complex task that takes a long time, there is a high possibility that the program will be interrupted. Therefore, the idea of saving breakpoints is used to avoid the waste of resources caused by task interruptions. The breakpoint relay module consists of two functions, load_checkpoint and save_checkpoint. The load_checkpoint function is executed at the beginning of the program running. If there is a checkpoint file in the program execution directory, the parameters are initialized according to the checkpoint. If not, they are initialized to zero values. The save_checkpoint function is executed when the program is interrupted or when the time from the last save is greater than the set threshold, and the data that the current program has produced is saved as a local file in json format.
[0071] The specific implementation method of using the large model to generate the answer instruction data set for the current generation round in the above step S104 is described in detail above.
[0072] The method provided in the embodiment of the present application is described above. The device provided in the embodiment of the present application is described below:
[0073] See also Figure 4 , Figure 4 This is a structural diagram of a data generation device provided in an embodiment of the present application. Figure 4 As shown, the device 400 includes: a configuration module 401, a first generation module 402, a second generation module 403, and a determination module.
[0074] Configuration module 401, used to obtain a reference text document, a first prompt word document, and a seed data document; the reference text document is used to record the knowledge in a specified technical field; the first prompt word document includes a plurality of prompt words used to describe the content generation requirements for generating question instruction data; the seed data document includes a plurality of reference formats used to define the generation of question instruction data;
[0075] The first generation module 402 is used to generate a question instruction data set for the current generation round by using the obtained large model, based on the reference text document, the first prompt word document, and the seed data document in the current generation round; and perform a filtering operation on the question instruction data for the current generation round based on a pairing principle to obtain a reserved question instruction data set for the current generation round;
[0076] The second generation module 403 is used to generate the answer instruction data set in the current generation round by using the obtained large model, based on the reference text document, the second prompt word document, and the reserved question instruction data set in the current generation round; wherein the second prompt word document includes a plurality of content generation requirements for describing the generation of answer instruction data;
[0077] Determination module 404, the user iterates and executes generation rounds, and determines the retained question instruction data set under all generation rounds when the round iteration stopping condition is met, and the answer instruction data set under all generation rounds as the question instruction data set of the reference text document, and the question instruction data set of the reference text document, respectively.
[0078] In one embodiment, the first generation module 402 is further used to: execute N question generation tasks in parallel in the current generation round, and when executing each question generation task, determine the first prompt word and reference format matching the question generation task based on the first prompt word document and the seed data document, and input the reference text document, the matching first prompt word and reference format into the obtained large model to obtain the question instruction data set generated by the large model for the question generation task; N is greater than 1; wherein the obtained question instruction data set of each of the N question generation tasks generated by the large model is the question instruction data set under the current generation round;
[0079] Based on the pairwise pairing principle, a filtering operation is performed on the N question instruction data sets generated in the current generation round to remove question instruction data whose similarity does not meet the set requirements, and obtain the retained question instruction data set under the current generation round.
[0080] In one embodiment, the first generation module 403 is further used to: execute M answer generation tasks in parallel for the reserved question instruction data set under the current generation round, and when executing each answer generation task, determine the second prompt word matching the answer generation task based on the second prompt word document, input the reference text document, the matching second prompt word, and a reserved question instruction data into the large model to obtain the answer instruction data set of the reserved question instruction data generated by the large model; and determine the answer instruction data set corresponding to each reserved question instruction data as the answer instruction data set under the current generation round; M is greater than 1.
[0081] In one embodiment, based on the first prompt word document and the seed data document, determining the first prompt word and the reference format matching the question generation task includes:
[0082] Randomly select P first prompt words from the first prompt word file, and determine the P first prompt words as matching content generation requirements; P is greater than 1;
[0083] If the current generation round is the first generation round, K reference formats are randomly selected from the seed data document and the K reference formats are determined as matching reference formats; K is greater than 1;
[0084] If the current generation round is not the first generation round, KL reference formats are randomly selected from the seed data document, L question instruction data are selected from the retained question instruction data set under the previous historical generation round of the current generation round as L reference formats, and the KL reference formats and L reference formats are determined as matching reference formats.
[0085] In one embodiment, based on the pairing principle, a filtering operation is performed on the N question instruction data sets generated in the current generation round to remove the question instruction data whose similarity does not meet the set requirements, and the retained question instruction data set under the current generation round is obtained, including:
[0086] If the current generation round is the first generation round, every two question instruction data in the N question instruction data sets generated in the current generation round are determined as an instruction pair;
[0087] For each instruction pair, calculate the similarity value of the two problem instruction data in the instruction pair, and generate a similarity matrix according to the similarity values corresponding to each instruction pair;
[0088] Select reference question instruction data from the N question instruction data sets, traverse other reference question instruction data except the reference question instruction data in a set order, and determine the currently traversed reference question instruction data as the current reference question instruction data;
[0089] Querying the similarity values corresponding to the instructions corresponding to the reference question instruction data and the current reference question instruction data from the similarity matrix;
[0090] If the similarity value queried is greater than the set similarity threshold, the current reference question instruction data is deleted; if the similarity value queried is less than or equal to the set similarity threshold, the current reference question instruction data is retained.
[0091] In one embodiment, based on the pairing principle, a filtering operation is performed on the N question instruction data sets generated in the current generation round to remove the question instruction data whose similarity does not meet the set requirements, and the retained question instruction data set under the current generation round is obtained, which also includes:
[0092] If the current generation round is not the first generation round, every two problem instruction data in the N problem instruction data sets generated by the current generation round and the reserved problem instruction data set corresponding to the previous generation round of the current generation round are determined as an instruction pair;
[0093] For each instruction pair, calculate the similarity value of the two problem instruction data in the instruction pair, and generate a similarity matrix according to the similarity values corresponding to each instruction pair;
[0094] Select reference question instruction data from the N question instruction data sets, traverse other reference question instruction data except the reference question instruction data in a set order, and determine the currently traversed reference question instruction data as the current reference question instruction data;
[0095] Querying the similarity values corresponding to the instructions corresponding to the reference question instruction data and the current reference question instruction data from the similarity matrix;
[0096] If the similarity value queried is greater than the set similarity threshold, the current reference question instruction data is deleted; if the similarity value queried is less than or equal to the set similarity threshold, the current reference question instruction data is retained.
[0097] In one embodiment, inputting a reference text document, a matching first prompt word and a reference format into the obtained large model includes:
[0098] In the order of reference text document, generation requirements and reference format, the reference text document, the matching first prompt word and the reference format are spliced into one file, and the file is input into the large model.
[0099] In one embodiment, the device further comprises a tracking module, configured to:
[0100] When executing each answer generation task, after requesting the large model to execute the answer generation task, the index corresponding to the answer generation task is added to the answer generation task index set; the index is determined according to the index coordinates of the reserved question instruction data carried by the answer generation task in the two-dimensional generation matrix; the two-dimensional generation matrix is established according to the reference text document and the reserved question instruction data corresponding to the reference text document;
[0101] For each answer instruction data that retains the question instruction data in the answer generation task obtained from the big model, if the answer instruction data meets the set format requirements, the index coordinate corresponding to the answer generation task is used as the key key, and the answer instruction data is used as the value value to form a kye-value for the answer generation task in the dictionary; wherein the dictionary and the collection user track whether the answer generation task is successfully executed.
[0102] In one embodiment, the tracking module is further configured to:
[0103] After the current generation round is completed, each index in the collection will be initialized as the key of the dictionary to update the dictionary. The updated dictionary is used to determine the retained question instruction data carried in the answer generation task in the next generation round.
[0104] So far, completed Figure 4 Structural description of the device shown.
[0105] See also Figure 5 , Figure 5 This is a structural diagram of an electronic device provided in an embodiment of the present application. Figure 5 As shown, the hardware structure may include: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the method disclosed in the above example of this application.
[0106] Based on the same application concept as the above method, an embodiment of the present application also provides a machine-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed by a processor, the method disclosed in the above example of the present application can be implemented.
[0107] Exemplarily, the above-mentioned machine-readable storage medium can be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium can be: RAM (Radom Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as hard disk drive), solid state drive, any type of storage disk (such as CD, DVD, etc.), or similar storage medium, or a combination thereof.
[0108] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A data generation method, characterized in that: The method comprises: Obtaining a reference text document, a first prompt word document, and a seed data document; the reference text document is used to record knowledge in a specified technical field; the first prompt word document includes a plurality of first prompt words for describing content generation requirements for generating question instruction data; the seed data document includes a plurality of reference formats for defining generation question instruction data; In the current generation round, using the obtained large model, based on the reference text document, the first prompt word document, and the seed data document, a question instruction data set for the current generation round is generated; and based on the pairing principle, a filtering operation is performed on the question instruction data for the current generation round to obtain a reserved question instruction data set for the current generation round; Using the obtained large model, based on the reference text document, the second prompt word document, and the reserved question instruction data set under the current generation round, a set of answer instruction data under the current generation round is generated; wherein the second prompt word document includes a plurality of second prompt words for describing content generation requirements for generating answer instruction data; Iteratively execute generation rounds, and determine the retained question instruction data set under all generation rounds when the round iteration stopping condition is met, and the answer instruction data set under all generation rounds as the question instruction data set of the reference text document, and the question instruction data set of the reference text document, respectively.
2. The method according to claim 1, characterized in that In the current generation round, the obtained large model is used to generate a question instruction data set under the current generation round based on the reference text document, the first prompt word document, and the seed data document, including: In the current generation round, N question generation tasks are executed in parallel. When executing each question generation task, based on the first prompt word document and the seed data document, a first prompt word and a reference format matching the question generation task are determined, and the reference text document, as well as the matching first prompt word and reference format are input into the obtained large model to obtain a question instruction data set under which the large model generates the question generation task; N is greater than 1; wherein the obtained question instruction data set of each of the N question generation tasks generated by the large model is the question instruction data set under the current generation round; The filtering operation is performed on the question instruction data under the current generation round based on the pairing principle to obtain a set of retained question instruction data under the current generation round, including: Based on the pairwise pairing principle, a filtering operation is performed on the N question instruction data sets generated in the current generation round to remove the question instruction data whose similarity does not meet the set requirements, and obtain the retained question instruction data set under the current generation round; The step of using the obtained large model to generate the answer instruction data set for the current generation round based on the reference text document, the second prompt word document, and the reserved question instruction data set for the current generation round includes: For the reserved question instruction data set under the current generation round, M answer generation tasks are executed in parallel. When executing each answer generation task, the second prompt word matching the answer generation task is determined based on the second prompt word document, and the reference text document, the matching second prompt word, and a reserved question instruction data are input into the large model to obtain the answer instruction data set of the reserved question instruction data generated by the large model; and the answer instruction data set corresponding to each reserved question instruction data is determined as the answer instruction data set under the current generation round; M is greater than 1.
3. The method according to claim 2, characterized in that The determining, based on the first prompt word document and the seed data document, a first prompt word and a reference format matching the question generation task comprises: Randomly select P first prompt words from the first prompt word file, and determine the P first prompt words as the matching first prompt words; P is greater than 1; If the current generation round is the first generation round, randomly selecting K reference formats from the seed data document, and determining the K reference formats as the matching reference formats; K is greater than 1; If the current generation round is not the first generation round, KL reference formats are randomly selected from the seed data document, L question instruction data are selected from the retained question instruction data set under the previous historical generation round of the current generation round as L reference formats, and the KL reference formats and the L reference formats are determined as the matching reference formats.
4. The method according to claim 2, characterized in that: The filtering operation is performed on the N question instruction data sets generated in the current generation round based on the pairwise pairing principle to remove the question instruction data whose similarity does not meet the set requirements, and the retained question instruction data set under the current generation round is obtained, including: If the current generation round is the first generation round, every two question instruction data in the N question instruction data sets generated in the current generation round are determined as an instruction pair; For each instruction pair, calculate the similarity value of the two problem instruction data in the instruction pair, and generate a similarity matrix according to the similarity values corresponding to each instruction pair; Selecting reference question instruction data from the N question instruction data sets, traversing other reference question instruction data except the reference question instruction data in a set order, and determining the currently traversed reference question instruction data as current reference question instruction data; Querying the similarity values corresponding to the instructions corresponding to the reference question instruction data and the current reference question instruction data from the similarity matrix; If the similarity value found is greater than the set similarity threshold, the current reference question instruction data is deleted; if the similarity value found is less than or equal to the set similarity threshold, the current reference question instruction data is retained.
5. The method according to claim 4, characterized in that The filtering operation is performed on the N question instruction data sets generated in the current generation round based on the pairwise pairing principle to remove the question instruction data whose similarity does not meet the set requirements to obtain the reserved question instruction data set under the current generation round, and also includes: If the current generation round is not the first generation round, every two question instruction data in the N question instruction data sets generated by the current generation round and the reserved question instruction data set corresponding to the previous generation round of the current generation round are determined as an instruction pair; For each instruction pair, calculate the similarity value of the two problem instruction data in the instruction pair, and generate a similarity matrix according to the similarity values corresponding to each instruction pair; Selecting reference question instruction data from the N question instruction data sets, traversing other reference question instruction data except the reference question instruction data in a set order, and determining the currently traversed reference question instruction data as current reference question instruction data; Querying the similarity values corresponding to the instructions corresponding to the reference question instruction data and the current reference question instruction data from the similarity matrix; If the similarity value found is greater than the set similarity threshold, the current reference question instruction data is deleted; if the similarity value found is less than or equal to the set similarity threshold, the current reference question instruction data is retained.
6. The method according to claim 2, characterized in that The step of inputting the reference text document, the matched first prompt word and the reference format into the obtained large model comprises: According to the order of reference text document, generation requirements and reference format, the reference text document, the matching first prompt word and the reference format are spliced into one file, and the file is input into the large model.
7. The method according to claim 2, characterized in that The method further comprises: When executing each answer generation task, after requesting the large model to execute the answer generation task, the index corresponding to the answer generation task is added to the answer generation task index set; the index is determined according to the index coordinates of the reserved question instruction data carried by the answer generation task in the two-dimensional generation matrix; the two-dimensional generation matrix is established according to the reference text document and the reserved question instruction data corresponding to the reference text document; For each answer instruction data that retains the question instruction data in the answer generation task obtained from the large model, if the answer instruction data meets the set format requirements, the index coordinate corresponding to the answer generation task is used as the key key, and the answer instruction data is used as the value value to form a kye-value for the answer generation task in the dictionary; wherein the dictionary and the collection user track whether the answer generation task is successfully executed.
8. The method according to claim 7, characterized in that The method also includes: after the current generation round is completed, each index in the set will be initialized as the key of the dictionary to update the dictionary, and the updated dictionary is used to determine the retained question instruction data carried in the answer generation task in the next generation round.
9. A data generating device, characterized in that: The device comprises: A configuration module is used to obtain a reference text document, a first prompt word document, and a seed data document; the reference text document is used to record knowledge in a specified technical field; the first prompt word document includes a plurality of prompt words for describing content generation requirements for generating question instruction data; the seed data document includes a plurality of reference formats for defining generation question instruction data; A first generation module is used to generate a question instruction data set for the current generation round by using the obtained large model based on the reference text document, the first prompt word document, and the seed data document in the current generation round; and to perform a filtering operation on the question instruction data for the current generation round based on a pairwise pairing principle to obtain a reserved question instruction data set for the current generation round; A second generation module is used to generate an answer instruction data set under the current generation round by using the obtained large model, based on the reference text document, the second prompt word document, and the reserved question instruction data set under the current generation round; wherein the second prompt word document includes a plurality of second prompt words for describing content generation requirements for generating answer instruction data; A determination module, in which a user iterates and executes generation rounds, determines the retained question instruction data set under all generation rounds when the round iteration stopping condition is met, and the answer instruction data set under all generation rounds as the question instruction data set of the reference text document, and the question instruction data set of the reference text document, respectively.
10. An electronic device, characterized in that: The electronic device includes: Processor; and A computer-readable storage medium, wherein computer program instructions are stored in the computer-readable storage medium, and when the computer program instructions are executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 8.
Citation Information
Cited By
Large model data set generation method and device, electronic equipment and storage medium
CN120974180A