Data processing method, device and equipment

Through the standardized data synthesis framework and large language model generation, the problem of poor data quality and correlation in large model instruction data synthesis is solved, and the full utilization of the potential value of data and the improvement of synthesis efficiency is achieved.

CN120408027APending Publication Date: 2025-08-01ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510415057.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In the process of synthesis of large-model instruction data, the data quality and correlation are poor, and the potential value of the data is not effectively utilized, resulting in inconsistent and unpredictable synthetic data.

Method used

Through a standardized data synthesis framework, based on data construction prompt information in preset databases, the synthetic data is generated using a large language model, and the data value-added processing is performed by determining the importance of the synthetic data and covering distribution information to ensure that the synthetic data meets the requirements of the large model instruction data synthesis.

Benefits of technology

It improves the quality and relevance of synthetic data, makes full use of the potential value of data, reduces manual intervention, improves the efficiency and automation of data synthesis, expands the depth and breadth of data, and achieves the dual improvement of the quality and value of large-model instruction data synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408027A_ABST
    Figure CN120408027A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method, device and equipment, and the method comprises the steps: building prompt information through data in a preset database based on a preset large model instruction data synthesis demand; based on the prompt information, generating multiple pieces of synthetic data corresponding to the prompt information through a large language model; determining the importance of each piece of synthetic data, determining the coverage distribution information of the plurality of pieces of synthetic data, and performing data value-added processing on the plurality of pieces of synthetic data based on the importance of each piece of synthetic data and the coverage distribution information of the plurality of pieces of synthetic data to obtain a plurality of pieces of synthetic data after data value-added processing; and determining large model instruction data meeting the large model instruction data synthesis requirements on the basis of the synthesized data after value addition of the multiple pieces of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of computer technology, and in particular, to a data processing method, apparatus, and device. Background Art

[0002] In the rapid development of machine learning and artificial intelligence, data has become a key element in training large models (such as large language models) under the premise of privacy protection. Large models are known for their huge number of parameters and deep learning capabilities, and are highly dependent on data, especially in the pre-training and supervised fine-tuning stages. With the progress of technology, the growing demand for high-quality and diverse data has not only promoted the development of synthetic data technology, but also put forward higher requirements for data management.

[0003] However, current related technologies have several limitations in data synthesis and management. Although synthetic data technology provides a potential solution to the problem of data scarcity, most data synthesis processes are inconsistent and unpredictable. In addition, the data quality and relevance of the obtained large model instruction data are poor, and the potential value of the data has not been effectively utilized. Therefore, it is necessary to provide a better way to synthesize large model instruction data, through a standardized data synthesis framework, to improve the data quality and relevance of large model instruction data, and to make full use of the potential value of the data during the data synthesis process. Summary of the Invention

[0004] The purpose of the embodiments of this specification is to provide a better way to synthesize large model instruction data, through a standardized data synthesis framework, to improve the data quality and relevance of large model instruction data, and to make full use of the potential value of the data during the data synthesis process.

[0005] In order to implement the above technical solution, the embodiments of this specification are implemented as follows: A data processing method provided by the embodiments of this specification, the method includes: constructing prompt information based on the data in a preset database according to the preset large model instruction data synthesis requirements; generating a plurality of synthetic data corresponding to the prompt information through a large language model based on the prompt information; determining the importance of each synthetic data, and determining the coverage distribution information of the plurality of synthetic data, and performing data value-added processing on the plurality of synthetic data based on the importance of each synthetic data and the coverage distribution information of the plurality of synthetic data to obtain a plurality of synthetic data after data value-added; determining large model instruction data that meets the large model instruction data synthesis requirements based on the plurality of synthetic data after data value-added.

[0006] A data processing device provided by an embodiment of this specification, the device includes: a data processing module, which constructs prompt information through data in a preset database based on a preset large model instruction data synthesis requirement; a data synthesis module, which generates a plurality of synthetic data corresponding to the prompt information through a large language model based on the prompt information; a data value-added module, which determines the importance of each synthetic data and determines the coverage distribution information of the plurality of synthetic data, and performs data value-added processing on the plurality of synthetic data based on the importance of each synthetic data and the coverage distribution information of the plurality of synthetic data to obtain a plurality of synthetic data after data value-added; an instruction data determination module, which determines large model instruction data that meets the large model instruction data synthesis requirement based on the plurality of synthetic data after data value-added.

[0007] A data processing device provided by an embodiment of this specification, the data processing device includes: a processor; and a memory arranged to store computer-executable instructions, the executable instructions, when executed, cause the processor to: construct prompt information through data in a preset database based on a preset large model instruction data synthesis requirement; generate a plurality of synthetic data corresponding to the prompt information through a large language model based on the prompt information; determine the importance of each synthetic data and determine the coverage distribution information of the plurality of synthetic data, and perform data value-added processing on the plurality of synthetic data based on the importance of each synthetic data and the coverage distribution information of the plurality of synthetic data to obtain a plurality of synthetic data after data value-added; determine large model instruction data that meets the large model instruction data synthesis requirement based on the plurality of synthetic data after data value-added.

[0008] An embodiment of this specification also provides a storage medium, the storage medium is used to store computer-executable instructions, and the executable instructions, when executed by a processor, implement the following processes: construct prompt information through data in a preset database based on a preset large model instruction data synthesis requirement; generate a plurality of synthetic data corresponding to the prompt information through a large language model based on the prompt information; determine the importance of each synthetic data and determine the coverage distribution information of the plurality of synthetic data, and perform data value-added processing on the plurality of synthetic data based on the importance of each synthetic data and the coverage distribution information of the plurality of synthetic data to obtain a plurality of synthetic data after data value-added; determine large model instruction data that meets the large model instruction data synthesis requirement based on the plurality of synthetic data after data value-added.

[0009] An embodiment of this specification also provides a computer program product, including a computer program, which when executed by a processor implements the following processes: constructing prompt information based on the data in a preset database according to the preset large model instruction data synthesis requirements; generating a plurality of synthetic data corresponding to the prompt information through a large language model based on the prompt information; determining the importance of each synthetic data, and determining the coverage distribution information of the plurality of synthetic data, and performing data value-added processing on the plurality of synthetic data based on the importance of each synthetic data and the coverage distribution information of the plurality of synthetic data to obtain a plurality of synthetic data after data value-added processing; determining large model instruction data that meets the large model instruction data synthesis requirements based on the plurality of synthetic data after data value-added processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Figure 1 Schematic diagram of an embodiment of a data processing method in this specification; Figure 2 Schematic diagram of a data synthesis page in this specification; Figure 3 Schematic diagram of another embodiment of a data processing method in this specification; Figure 4 Schematic diagram of a data processing process in this specification; Figure 5 Schematic diagram of yet another embodiment of a data processing method in this specification; Figure 6 Schematic diagram of another data processing process in this specification; Figure 7 Schematic diagram of yet another embodiment of a data processing method in this specification; Figure 8 Schematic diagram of yet another data processing process in this specification; Figure 9 Schematic diagram of yet another embodiment of a data processing method in this specification; Figure 10 Schematic diagram of yet another embodiment of a data processing method in this specification; Figure 11 Schematic diagram of yet another embodiment of a data processing method in this specification; Figure 12 Schematic diagram of a data processing device in this specification; Figure 13 This is a schematic diagram of a data processing device in this specification. Specific implementation manners

[0011] Embodiments of this specification provide a data processing method, apparatus, and device.

[0012] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only some of the embodiments of this specification, rather than all of them. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.

[0013] Embodiments of this specification provide a domain large model instruction data synthesis mechanism. In the rapid development of machine learning and artificial intelligence, data has become a key element in training large models (such as large language models). Large models are known for their huge number of parameters and deep learning capabilities, and are highly dependent on data, especially in the pre-training and supervised fine-tuning stages. With the progress of technology, the growing demand for high-quality and diverse data not only promotes the development of synthetic data technology, but also poses higher requirements for data management.

[0014] However, current related technologies have several limitations in data synthesis and management. Although synthetic data technology provides a potential solution to the problem of data scarcity, most data synthesis processes are inconsistent and unpredictable. In addition, the data quality and relevance of the obtained large model instruction data are poor, and the potential value of the data is not effectively utilized. Therefore, a better large model instruction data synthesis method is needed to improve the data quality and relevance of large model instruction data through a standardized data synthesis framework, and to fully utilize the potential value of the data during the data synthesis process. The specific processing can refer to the specific content in the following embodiments.

[0015] Such as Figure 1As shown, an embodiment of this specification provides a data processing method. The execution subject of this method can be a terminal device, a server, etc. The terminal device can be a mobile terminal device such as a mobile phone or a tablet computer, or a computer device such as a notebook computer or a desktop computer. Alternatively, it can also be an IoT device (such as a smart watch, a vehicle-mounted device, etc.). The server can be an independent server or a server cluster composed of multiple servers. The server can be a background server in fields such as finance or online shopping, or a background server of a certain application. In this embodiment, the case where the execution subject is a server is taken as an example for detailed description. For the case where the execution subject is a terminal device, reference can be made to the following treatment of the server, which will not be elaborated here. The method can specifically include the following steps: In step S102, based on the preset large model instruction data synthesis requirement, construct a prompt message through the data in the preset database.

[0016] Among them, the large model instruction data is the data used to train the large model, aiming to enable the large model to more accurately understand and execute specified instructions. The large model instruction data usually contains a data pair composed of a clear task instruction (or prompt message) and corresponding answer data to help the large model learn how to execute specific tasks. The large model instruction data synthesis requirement can be used to indicate the requirement for synthesizing large model instruction data of a certain type or field (such as the financial risk field, the medical field, etc.). The large model instruction data synthesis requirement can also include the quantity, attribute information, etc. of the large model instruction data to be synthesized, which can be specifically set according to the actual situation. The preset database can include various different data. For example, the preset database can include physical data, financial data, chemical data, medical field data, etc. The data in the preset database can be obtained from specified basic knowledge data and institutional operation data, or can be crawled from the Internet through web crawlers, etc., or can be obtained from specified external data sources (such as specified open source databases), etc., which can be specifically set according to the actual situation. The prompt message can be a Prompt applied to the large language model. The prompt message is the text input into the large language model, used to guide the large language model to generate an output result that meets the requirements. The prompt message can be a question, an instruction, or a piece of context, helping the large language model better understand and respond to various queries. The prompt message can include explicit instruction type prompt messages, context supplement type prompt messages, and example guidance type prompt messages, etc. The prompt message helps the large language model generate more relevant and accurate output results by providing detailed background information and examples.

[0017] In implementation, when it is necessary to synthesize large model instruction data for a certain large language model, the large model instruction data synthesis requirement can be set. For example Figure 2As shown, in order to synthesize large model instruction data, a data synthesis page can be set up. The data synthesis page may include an input box for the requirements of large model instruction data synthesis, an output box for large model instruction data, an OK button, a Cancel button, etc. When it is necessary to synthesize large model instruction data for a certain large language model, the above data synthesis page can be opened, and the requirements for large model instruction data synthesis can be entered in the input box for the requirements of large model instruction data synthesis on the data synthesis page. After the input is completed, the OK button can be clicked. At this time, the requirements for large model instruction data synthesis entered in the input box for the requirements of large model instruction data synthesis can be obtained, and based on the above requirements for large model instruction data synthesis, using the data in the preset database as the basis, a prompt message Prompt can be generated for each piece of data. Among them, the generated prompt message Prompt can include one or multiple. In addition, the answer data corresponding to the prompt message Prompt can also be generated based on the data in the preset database.

[0018] In step S104, based on the above prompt message, multiple synthetic data corresponding to the prompt message are generated through a large language model.

[0019] Among them, the large language model can be any large language model that needs to be trained or fine-tuned. The large language model can be any large language model, such as GPT-4, ChatGPT, Qwen, KIMI, Deepseek, etc., and can be specifically set according to the actual situation. The synthetic data can be data generated by computer algorithms, used to supplement or enhance the real dataset, especially in the case where real data is difficult to obtain. In this embodiment, the synthetic data can be a data pair composed of question-and-answer data pairs, that is, prompt messages and answer data.

[0020] In implementation, the above constructed prompt message can be input into the large language model. The large language model can output the corresponding synthetic answer data according to the input prompt message. The prompt message and the output synthetic answer data can be used as the synthetic data corresponding to the prompt message. In addition, in actual applications, the large language model can also output different synthetic answer data for the same prompt message according to different fields, etc. Based on this, the prompt message can be input into the large language model, and the large language model can output multiple different synthetic answer data respectively according to the input prompt message. The prompt message and each synthetic answer data can form a data pair, and this data pair is a synthetic data, and thus multiple synthetic data corresponding to the prompt message can be obtained. Through the above method, the synthetic data corresponding to each prompt message can be obtained, thereby generating multiple synthetic data corresponding to the above prompt message.

[0021] In step S106, determine the importance of each synthetic data and determine the coverage distribution information of multiple synthetic data. Based on the importance of each synthetic data and the coverage distribution information of multiple synthetic data, perform data value-added processing on the multiple synthetic data to obtain multiple synthetic data after data value-added processing.

[0022] Among them, the importance of synthetic data can be used to indicate the training gain of this synthetic data for the large language model. The higher the importance of the synthetic data, the greater the training gain of this synthetic data for the large language model, and the lower the importance of the synthetic data, the greater the training gain of this synthetic data for the large language model. The coverage distribution information of multiple synthetic data can be used to indicate the situation of abnormal synthetic data among multiple synthetic data (or the situation of outlier synthetic data among multiple synthetic data). Data value-added processing can be a process of improving the value of data through various methods to give play to the value of data. Data value-added processing can include various processing methods. For example, new data can be generated by replacing specified words in the data with synonyms, or part of the important data can be collected through data sampling, etc. Specifically, it can be set according to the actual situation.

[0023] In implementation, in order to make full use of the potential value of data, data value-added processing can be performed on the above-mentioned synthetic data, so as to identify and extract high-value elements in the synthetic data. Since data value-added processing can be achieved through a variety of different methods, in the embodiments of this specification, starting from aspects such as the importance of synthetic data and the coverage distribution information of multiple synthetic data, the importance and diversity of synthetic data are realized, and the depth of synthetic data is expanded. For the importance of synthetic data, the importance of each synthetic data can be determined through various methods. For example, the correlation between different synthetic data can be calculated. Specifically, the correlation coefficient between different synthetic data can be calculated through algorithms such as the Pearson correlation coefficient algorithm and the Spearman rank correlation coefficient algorithm. The relationship strength between different synthetic data can be judged through the calculated correlation coefficient, and then the importance of each synthetic data can be determined. For another example, the importance of each synthetic data can also be comprehensively evaluated according to aspects such as the credibility, key content, influence range, and time sensitivity of the synthetic data. For another example, corresponding weights can also be assigned to each synthetic data through a weight and scoring model, and multidimensional data analysis can be used to judge its importance. For another example, a specified model can also be trained, and the trained model can automatically identify important features and patterns in each synthetic data, and then determine the importance of each synthetic data.

[0024] For the coverage distribution information of multiple synthetic data, the coverage distribution information of multiple synthetic data can be determined in various ways. For example, the average value of multiple synthetic data can be calculated or a preset average value can be set. Then, the distance between each synthetic data and the average value can be calculated. If the calculated distance is greater than the preset threshold, it is determined that the synthetic data belongs to abnormal data. If the calculated distance is less than the preset threshold, it is determined that the synthetic data does not belong to abnormal data, and then the coverage distribution information of multiple synthetic data is determined. For another example, an outlier factor LOF dependent on neighborhood density can be assigned to each synthetic data, and then it can be judged whether the synthetic data is an outlier synthetic data (i.e., the LOF algorithm), and then the coverage distribution information of multiple synthetic data is determined. For another example, the coverage distribution information of multiple synthetic data can be determined by the DBSCAN algorithm, etc., which can be specifically set according to the actual situation.

[0025] After obtaining the importance of each synthetic data and the coverage distribution information of multiple synthetic data through the above methods, data augmentation processing can be performed on multiple synthetic data. Specifically, for example, homomorphic weighted sampling can be performed on multiple synthetic data through two indicators, namely, the importance of each synthetic data and the coverage distribution information of multiple synthetic data, so as to sample synthetic data with high importance and few distributions, so as to realize data augmentation processing on multiple synthetic data and obtain multiple synthetic data after data augmentation. Through the above data augmentation processing, the quality, diversity and coverage range of the selected data can be improved.

[0026] In step S108, based on multiple synthetic data after data augmentation, the large model instruction data that meets the data synthesis requirements of the large model instruction is determined.

[0027] In implementation, multiple synthetic data after data augmentation can be directly determined as the large model instruction data that meets the data synthesis requirements of the large model instruction. Or, data screening rules can be preset in advance, and multiple synthetic data after data augmentation can be screened through the data screening rules to obtain a certain number of synthetic data after data augmentation. The obtained certain number of synthetic data after data augmentation can be determined as the large model instruction data that meets the data synthesis requirements of the large model instruction. Or, other algorithms (such as similarity algorithms, clustering algorithms, etc.) can be preset, and multiple synthetic data after data augmentation can be processed through the set algorithms (such as calculating the similarity of synthetic data after data augmentation with a similarity greater than the preset threshold through the similarity algorithm or obtaining synthetic data after data augmentation belonging to the same category through clustering processing by the clustering algorithm, etc.). The processed data can be determined as the large model instruction data that meets the data synthesis requirements of the large model instruction, etc., which can be specifically set according to the actual situation.

[0028] An embodiment of this specification provides a data processing method. Based on a preset large model instruction data synthesis requirement, prompt information is constructed through data in a preset database. Then, based on this prompt information, multiple synthetic data corresponding to the prompt information can be generated through a large language model. After that, the importance of each synthetic data can be determined, and the coverage distribution information of the multiple synthetic data can be determined. Based on the importance of each synthetic data and the coverage distribution information of the multiple synthetic data, value-added processing is performed on the multiple synthetic data to obtain multiple synthetic data after value-added processing. Finally, based on the multiple synthetic data after value-added processing, large model instruction data that meets the large model instruction data synthesis requirement can be determined. In this way, a standardized data synthesis framework is provided. Through the standardized data synthesis framework, value-added processing is performed based on the dual dimensions of the importance of the synthetic data and the diversity corresponding to the coverage distribution information, improving the comprehensive value of the synthetic data, enhancing the data quality and relevance of the large model instruction data, and being able to fully utilize the potential value of the data during the data synthesis process. In addition, manual intervention is reduced, the efficiency and automation degree of data synthesis are improved, the depth and breadth of the data are expanded, and an effective dual improvement in the quality and value of the large model instruction data synthesis is achieved.

[0029] In practical applications, the above-mentioned preset database may include a seed data set and an externally collected data set. Based on this, the specific processing method of the above step S102 can be various. Hereinafter, an optional processing method is provided again, as Figure 3 shown, and it may specifically include the processing of the following step S1022 and step S1024.

[0030] In step S1022, based on the preset large model instruction data synthesis requirement, multiple different prompt information examples are constructed through the data in the seed data set.

[0031] Among them, the seed data set may include data in multiple different fields. For example, as Figure 4 shown, the seed data set may include data in fields such as chemistry, physics, and finance, and it can be specifically set according to the actual situation.

[0032] In implementation, for problems such as serious homogenization of the generated synthetic data, during the process of constructing the prompt information, diverse clustering sampling Few-shot examples (i.e., prompt information examples) can be added to improve the diversity of the prompt information Prompt. Specifically, through the above Figure 2The page inputs the large model instruction data synthesis requirements. After the input is completed, the input large model instruction data synthesis requirements can be obtained. In combination with the above large model instruction data synthesis requirements, the data in the seed data set can be used as the basis to collect diverse data for context learning to generate a small number of diverse prompt information Prompt examples, thereby obtaining multiple different prompt information examples.

[0033] In step S1024, based on the preset large model instruction data synthesis requirements, prompt information matching the prompt information example is constructed using data in the seed dataset and / or the externally collected dataset.

[0034] Among them, the externally collected data set may include multiple different types of data. For example, the externally collected data set may include Web data, text data in PDF format, RAG data, etc., and may also include attribute category information of the data, etc., which can be set specifically according to actual conditions.

[0035] In implementation, Figure 4 As shown, a data synthesis template can be set in advance. Through the data synthesis template, combined with the above-mentioned large model instruction data synthesis requirements, the data in the seed data set or the externally collected data set can be used as the basis, and the corresponding prompt information can be generated for each data with reference to the prompt information example, thereby obtaining prompt information that matches the prompt information example. In addition, in order to improve the fine-grained information of the prompt information Prompt, the prompt information that matches the prompt information example can be constructed based on the data synthesis template by adding the externally collected data set (including the attribute category information of the data) and the data in the seed data set. In this way, based on the fine-grained data in the externally collected data set and combined with the diverse data in the seed data set, diverse prompt information Prompt can be generated, thereby increasing the diversity of the prompt information Prompt.

[0036] In practical applications, the specific processing methods of the above step S104 can be varied. An optional processing method is provided below, such as Figure 5 As shown, the processing may specifically include the following steps S10402 to S10410.

[0037] In step S10402, the prompt information is input into a plurality of different large language models respectively to obtain first synthesized data output by each large language model.

[0038] In implementation, Figure 6As shown, multiple different large language models can include multiple ones such as GPT-4, ChatGPT, Qwen, KIMI, DeepSeek, etc. The above-mentioned prompt information can be respectively input into each of the multiple different large language models. Each large language model can output one or more answer data. The above-mentioned prompt information and each output answer data can form a first synthetic data, and thus multiple first synthetic data can be obtained.

[0039] In step S10404, each first synthetic data is evaluated for quality by a large language model to obtain the evaluation result of each first synthetic data.

[0040] In implementation, as Figure 6 shown, in order to achieve control over data quality, a self-reflection function can be set for the large language model. The self-reflection function means that the large language model can evaluate and improve its own output data to improve the quality and accuracy of the output data. Specifically, quality evaluation rules can be preset. For example, the quality evaluation rule is to calculate the similarity between the answer data in the first synthetic data and the benchmark answer data, and the first synthetic data is evaluated for quality based on the calculated similarity. Based on this, the large language model can calculate the similarity between the answer data in the first synthetic data and the benchmark answer data (the answer data corresponding to the prompt information Prompt generated based on the data in the preset database can be used as the benchmark answer data) through this quality evaluation rule to obtain the corresponding score value, and this score value can be used as the evaluation result of this first synthetic data. Through the above method, the evaluation result of each first synthetic data can be obtained.

[0041] In step S10406, each first synthetic data is processed for reflection by a large language model based on the evaluation result of each first synthetic data to obtain the reflection result corresponding to each first synthetic data.

[0042] In implementation, as Figure 6 shown, the large language model can analyze the evaluation result of each first synthetic data to identify the errors or deficiencies in the answer data of each first synthetic data, so as to achieve the reflection processing of each first synthetic data and obtain the reflection result corresponding to each first synthetic data.

[0043] In step S10408, each first synthetic data is processed for correction by a large language model based on the reflection result corresponding to each first synthetic data to obtain the corrected first synthetic data.

[0044] In implementation, the large language model can, according to the reflection results corresponding to each first synthetic data, correct the answer data in each first synthetic data, thereby generating a new and improved answer data, and can combine the prompt information in the first synthetic data with the generated new and improved answer data into the corrected first synthetic data.

[0045] In step S10410, based on the corrected first synthetic data, determine multiple synthetic data corresponding to the above-mentioned prompt information.

[0046] In implementation, the corrected first synthetic data can be directly determined as the multiple synthetic data corresponding to the above-mentioned prompt information, or the corrected first synthetic data can be used as the initial first synthetic data, and the processing of steps S10404 to S10408 is repeated until the answer data generated by the large language model reaches the preset quality standard or the number of iterations reaches the upper limit, and the obtained result is used as the multiple synthetic data corresponding to the above-mentioned prompt information.

[0047] In practical applications, the specific processing method for determining the importance of each synthetic data in step S106 can be various. Hereinafter, an optional processing method is provided again, as Figure 7 shown, and it specifically may include the processing of steps S10602 to S10608.

[0048] In step S10602, fine-tune the large language model based on the multiple synthetic data corresponding to the above-mentioned prompt information to obtain a fine-tuned large language model.

[0049] In step S10604, input the prompt information in the above-mentioned synthetic data into the fine-tuned large language model to obtain the first answer data corresponding to the prompt information in each synthetic data.

[0050] In step S10606, calculate the cross entropy between the first answer data corresponding to the prompt information in each synthetic data and the corresponding synthetic answer data in the above-mentioned synthetic data.

[0051] In implementation, as Figure 8 shown, the calculated cross entropy can be used as the difficulty of the synthetic data. The greater the difficulty of the synthetic data, the greater the possible training gain of the synthetic data for the large language model, and the smaller the difficulty of the synthetic data, the smaller the possible training gain of the synthetic data for the large language model.

[0052] In step S10608, based on the calculated cross entropy, determine the importance of each synthetic data.

[0053] In implementation, the importance of synthetic data represents the gain of synthetic data for model training. The gain situation can be determined by a gain value. The calculated cross-entropy can be directly used as the basis for judging the importance of each synthetic data. Alternatively, the calculated cross-entropy can be further processed (such as mapping, normalization, or regularization, etc.), and the processed result can be used as the basis for judging the importance of each synthetic data, etc. It can be specifically set according to the actual situation, and the embodiments of this specification do not limit this.

[0054] In practical applications, the specific processing methods for determining the coverage distribution information of multiple synthetic data in step S106 above can be diverse. Here is another optional processing method, such as Figure 9 as shown, it can specifically include the processing of the following steps S10610 and S10612.

[0055] In step S10610, based on a preset anomaly detection algorithm, determine the probability of outlier data among multiple synthetic data. The anomaly detection algorithm includes a similarity algorithm.

[0056] In step S10612, based on the probability of outlier data among multiple synthetic data, determine the coverage distribution information of multiple synthetic data.

[0057] In practical applications, in addition to the above data augmentation processing, additional data augmentation processing can also be performed. Specifically, refer to the processing of the following steps A2 and A4.

[0058] In step A2, obtain multiple first instruction data from a preset open-source dataset.

[0059] Among them, the first instruction data can be a type of large model instruction data, which can be data for training a large language model. The first instruction data can include a data pair composed of a clear task instruction (or prompt information) and corresponding answer data. The open-source dataset can include various types, for example, a logical reasoning open-source dataset, a mathematical open-source dataset, a brainstorming open-source dataset, etc. It can be specifically set according to the actual situation.

[0060] In implementation, the preset open-source dataset can include multiple large model instruction data. At this time, multiple large model instruction data can be directly obtained from the preset open-source dataset, and the obtained large model instruction data can be used as the first instruction data. Or, based on the above large model instruction data synthesis requirements, prompt information can be constructed through the data in the open-source dataset, and the answer data corresponding to the prompt information can be generated based on the data in the open-source dataset, thereby obtaining the first instruction data. Multiple first instruction data can be obtained through the above methods, etc. It can be specifically set according to the actual situation.

[0061] In step A4, the importance of each first instruction data is determined, and the coverage distribution information of multiple first instruction data is determined. Based on the importance of each first instruction data and the coverage distribution information of multiple first instruction data, data value-added processing is performed on the multiple first instruction data to obtain multiple first instruction data after data value-added.

[0062] The specific processing of the above step A4 can be found in the above related content and will not be repeated here.

[0063] Based on the processing of the above steps A2 and A4, the specific processing method of the above step S108 may include: determining the large model instruction data that meets the large model instruction data synthesis requirements based on the synthesized data after multiple data values are added and the first instruction data after multiple data values are added.

[0064] In implementation, the synthesized data after multiple data value-added and the first instruction data after multiple data value-added can be directly determined as the large model instruction data that meets the large model instruction data synthesis requirements, or a data screening rule can be pre-set, and the synthesized data after multiple data value-added and the first instruction data after multiple data value-added can be screened by the data screening rule to obtain a certain amount of data, and the obtained certain amount of data can be determined as the large model instruction data that meets the large model instruction data synthesis requirements, or other algorithms (such as similarity algorithms, clustering algorithms, etc.) can be preset, and the synthesized data after multiple data value-added and the first instruction data after multiple data value-added can be processed separately by the set algorithm (such as performing similarity calculation through a similarity algorithm to obtain data with similarity greater than a preset threshold or performing clustering processing through a clustering algorithm to obtain data belonging to the same category, etc.), and the processed data can be determined as the large model instruction data that meets the large model instruction data synthesis requirements, etc., and the specific setting can be based on actual conditions.

[0065] In practical applications, the specific processing methods of the above step S108 can be varied. An optional processing method is provided below, such as Figure 10 As shown, the processing may specifically include the following steps S1082 and S1084.

[0066] In step S1082, data quality assessment is performed on the multiple synthetic data after data value-added based on the preset data quality assessment rules to obtain the data quality assessment results corresponding to each synthetic data after data value-added. The data quality assessment results are used to indicate the degree of gain of the synthetic data after data value-added to the large language model.

[0067] Among them, the data quality assessment rules can include various types. The data quality assessment rules can perform data quality assessment processing on the synthetic data after data value addition through the stability of the data, or can perform data quality assessment processing on the synthetic data after data value addition through the complexity of the data, etc., and can be specifically set according to the actual situation.

[0068] In implementation, data quality assessment processing can be respectively performed on the synthetic data after data value addition of multiple data based on preset data quality assessment rules. Through the data quality assessment rules, information such as the stability degree or complexity of the synthetic data after data value addition can be judged. Through the obtained information such as the stability degree or complexity, the data quality assessment result corresponding to the synthetic data after data value addition is determined, and further, the gain degree of the synthetic data after data value addition to the large language model is determined. Through the above method, the data quality assessment result corresponding to each synthetic data after data value addition can be obtained.

[0069] In step S1084, based on the data quality assessment result corresponding to each synthetic data after data value addition, large model instruction data that meets the data synthesis requirements of the large model instructions is determined.

[0070] In implementation, if the data quality assessment result corresponding to a certain synthetic data after data value addition indicates that the synthetic data after data value addition has a relatively large gain degree to the large language model, then the synthetic data after data value addition can be determined as the large model instruction data that meets the data synthesis requirements of the large model instructions. If the data quality assessment result corresponding to a certain synthetic data after data value addition indicates that the synthetic data after data value addition has a relatively small gain degree to the large language model, then the synthetic data after data value addition can be discarded. Finally, the large model instruction data that meets the data synthesis requirements of the large model instructions can be obtained.

[0071] In practical applications, the specific processing method of the above step S1082 can be various. Hereinafter, another optional processing method is provided, as Figure 11 shown, and specifically, it can include the processing of the following steps S10822 to S10826.

[0072] In step S10822, based on the synthetic data after data value addition of multiple data and the data synthesis requirements of the large model instructions, a data quality assessment index is constructed. The data quality assessment index includes one or more of the distribution difference between the synthetic data after data value addition and the real data, the complexity of the data, and the perplexity of the data.

[0073] In step S10824, based on the data quality assessment index, one or more synthetic data after data value addition that meet the data quality assessment rules corresponding to the data quality assessment index are selected from the synthetic data after data value addition of multiple data.

[0074] In step S10826, based on the selected one or more value-added synthetic data, a data quality assessment result corresponding to the selected one or more value-added synthetic data is determined by the large language model.

[0075] During implementation, the prompt information in the selected synthetic data after data value-added can be input into the large language model to obtain a corresponding output result. The degree of gain of the synthetic data after data value-added to the large language model can be judged based on the output result, so as to determine the data quality assessment result corresponding to the synthetic data after the selected data value-added. In the above manner, the data quality assessment result corresponding to each selected synthetic data after data value-added can be obtained.

[0076] The embodiments of this specification provide a data processing method, which constructs prompt information based on data in a preset database based on a preset large model instruction data synthesis requirement. Then, based on the prompt information, multiple synthetic data corresponding to the prompt information can be generated through a large language model. Thereafter, the importance of each synthetic data can be determined, and the coverage distribution information of the multiple synthetic data can be determined. Based on the importance of each synthetic data and the coverage distribution information of the multiple synthetic data, data value-added processing is performed on the multiple synthetic data to obtain multiple synthetic data after value-added. Finally, based on the multiple synthetic data after value-added, large model instruction data that meets the large model instruction data synthesis requirement can be determined. In this way, a standardized data synthesis framework is provided. Through the standardized data synthesis framework, data value-added processing is performed based on the dual dimensions of diversity corresponding to the importance of the synthetic data and the coverage distribution information, thereby improving the comprehensive value of the synthetic data, improving the data quality and relevance of the large model instruction data, and making full use of the potential value of the data in the process of data synthesis. In addition, manual intervention is reduced, the efficiency and automation of data synthesis are improved, the depth and breadth of the data are expanded, and the effective dual improvement of the quality and value of the large model instruction data synthesis is achieved.

[0077] In addition, during the data value-added processing stage, the importance and diversity of the data are defined and the depth of the data is expanded by calculating cross-entropy and applying unsupervised anomaly data detection algorithms. In addition, the introduction of externally collected data sets and seed data sets, as well as the self-reflection and iteration mechanism of the large language model, improves the quality and fine-grained information of the synthetic data. The data quality assessment indicators and the large language model evaluation mechanism ensure the consistency and stability of the synthetic data with the real data in multiple dimensions. Moreover, high-quality and diverse synthetic data helps to train more robust models and improves the generalization ability of the model in different application scenarios.

[0078] The above is a data processing method provided in the embodiment of this specification. Based on the same idea, the embodiment of this specification also provides a data processing device, such asFigure 12 as shown

[0079] The data processing device includes: a data processing module 1201, a data synthesis module 1202, a data value-added module 1203, and an instruction data determination module 1204, where: The data processing module 1201 constructs a prompt message based on the data in a preset database according to the preset large model instruction data synthesis requirement; The data synthesis module 1202 generates a plurality of synthesized data corresponding to the prompt message through a large language model based on the prompt message; The data value-added module 1203 determines the importance of each synthesized data and determines the coverage distribution information of the plurality of synthesized data, and performs data value-added processing on the plurality of synthesized data based on the importance of each synthesized data and the coverage distribution information of the plurality of synthesized data to obtain a plurality of synthesized data after data value-added; The instruction data determination module 1204 determines the large model instruction data that meets the large model instruction data synthesis requirement based on the plurality of synthesized data after data value-added.

[0080] In the embodiment of this specification, the preset database includes a seed data set and an externally collected data set, and the data processing module 1201 includes: An example construction unit constructs a plurality of different prompt message examples through the data in the seed data set according to the preset large model instruction data synthesis requirement; A prompt message construction unit constructs a prompt message that matches the prompt message example through the data in the seed data set and / or the externally collected data set according to the preset large model instruction data synthesis requirement.

[0081] In the embodiment of this specification, the data synthesis module 1202 includes: A first data synthesis unit inputs the prompt message into a plurality of different large language models respectively to obtain first synthesized data output by each large language model; A quality evaluation unit performs quality evaluation on each first synthesized data through the large language model to obtain an evaluation result of each first synthesized data; A reflection unit performs reflection processing on each first synthesized data through the large language model based on the evaluation result of each first synthesized data to obtain a reflection result corresponding to each first synthesized data; A correction unit performs correction processing on the corresponding first synthesized data through the large language model based on the reflection result corresponding to each first synthesized data to obtain the corrected first synthesized data; The synthetic data determination unit determines multiple synthetic data corresponding to the prompt information based on the corrected first synthetic data.

[0082] In the embodiments of this specification, the data augmentation module 1203 includes: The fine-tuning unit fine-tunes the large language model based on multiple synthetic data corresponding to the prompt information to obtain a fine-tuned large language model; The answer determination unit inputs the prompt information in the synthetic data into the fine-tuned large language model to obtain first answer data corresponding to the prompt information in each synthetic data; The cross-entropy determination unit calculates the cross-entropy between the first answer data corresponding to the prompt information in each synthetic data and the corresponding synthetic answer data in the synthetic data; The importance determination unit determines the importance of each synthetic data based on the calculated cross-entropy.

[0083] In the embodiments of this specification, the data augmentation module 1203 includes: The outlier detection unit determines the outlier data probability in multiple synthetic data based on a preset outlier detection algorithm, and the outlier detection algorithm includes a similarity algorithm; The coverage distribution determination unit determines the coverage distribution information of the multiple synthetic data based on the outlier data probability in the multiple synthetic data.

[0084] In the embodiments of this specification, the device further includes: The first instruction data acquisition module acquires multiple first instruction data from a preset open-source dataset; The additional data augmentation module determines the importance of each first instruction data and determines the coverage distribution information of the multiple first instruction data, and performs data augmentation processing on the multiple first instruction data based on the importance of each first instruction data and the coverage distribution information of the multiple first instruction data to obtain multiple data-augmented first instruction data; The instruction data determination module 1204 determines large model instruction data that meets the large model instruction data synthesis requirements based on the multiple data-augmented synthetic data and the multiple data-augmented first instruction data.

[0085] In the embodiments of this specification, the instruction data determination module 1204 includes: The data quality evaluation unit performs data quality evaluation processing on the multiple data-augmented synthetic data respectively based on preset data quality evaluation rules to obtain a data quality evaluation result corresponding to each data-augmented synthetic data, and the data quality evaluation result is used to indicate the gain degree of the data-augmented synthetic data to the large language model; The large model instruction data determining unit determines the large model instruction data that meets the large model instruction data synthesis requirements based on the data quality assessment results corresponding to the synthesized data after each of the data values are increased.

[0086] In an embodiment of the present specification, the data quality assessment unit constructs a data quality assessment indicator based on the multiple data-valued synthesized data and the large model instruction data synthesis requirements, and the data quality assessment indicator includes one or more of the distribution difference between the data-valued synthesized data and the real data, the complexity of the data, and the perplexity of the data; based on the data quality assessment indicator, one or more data-valued synthesized data that meet the data quality assessment rules corresponding to the data quality assessment indicator are selected from the multiple data-valued synthesized data; based on the selected one or more data-valued synthesized data, the data quality assessment results corresponding to the selected one or more data-valued synthesized data are determined through the large language model.

[0087] An embodiment of the present specification provides a data processing device, which constructs prompt information based on data in a preset database based on a preset large model instruction data synthesis requirement. Then, based on the prompt information, multiple synthetic data corresponding to the prompt information can be generated through a large language model. Thereafter, the importance of each synthetic data can be determined, and the coverage distribution information of the multiple synthetic data can be determined. Based on the importance of each synthetic data and the coverage distribution information of the multiple synthetic data, data value-added processing is performed on the multiple synthetic data to obtain multiple synthetic data after value-added. Finally, based on the multiple synthetic data after value-added, large model instruction data that meets the large model instruction data synthesis requirement can be determined. In this way, a standardized data synthesis framework is provided. Through the standardized data synthesis framework, data value-added processing is performed based on the dual dimensions of diversity corresponding to the importance of the synthetic data and the coverage distribution information, thereby improving the comprehensive value of the synthetic data, improving the data quality and relevance of the large model instruction data, and making full use of the potential value of the data in the process of data synthesis. In addition, manual intervention is reduced, the efficiency and automation of data synthesis are improved, the depth and breadth of the data are expanded, and the effective dual improvement of the quality and value of the large model instruction data synthesis is achieved.

[0088] In addition, in the data value-added processing stage, the importance and diversity of data are defined by calculating cross-entropy and applying unsupervised anomaly data detection algorithms, expanding the depth of the data. In addition, an externally collected dataset, a seed dataset, and a large language model self-reflection and iteration mechanism are introduced to improve the quality and fine-grained information of the synthetic data. The data quality evaluation metrics and the large language model judgment mechanism ensure the consistency and stability of the synthetic data with the real data in multiple dimensions. Moreover, the high-quality and diverse synthetic data helps train a more robust model and improves the generalization ability of the model in different application scenarios.

[0089] The above is the data processing device provided by the embodiments of this specification. Based on the same idea, the embodiments of this specification also provide a data processing device, as Figure 13 shown.

[0090] The data processing device may be the terminal device or server provided in the above embodiments, etc.

[0091] The data processing device may vary greatly due to configuration or performance, and may include one or more processors 1301 and a memory 1302. One or more application programs or data may be stored in the memory 1302. Among them, the memory 1302 may be short-term storage or persistent storage. The application programs stored in the memory 1302 may include one or more modules (not shown in the figure), and each module may include a series of computer-executable instructions in the data processing device. Further, the processor 1301 may be configured to communicate with the memory 1302 and execute a series of computer-executable instructions in the memory 1302 on the data processing device. The data processing device may also include one or more power supplies 1303, one or more wired or wireless network interfaces 1304, one or more input / output interfaces 1305, and one or more keyboards 1306.

[0092] Specifically, in this embodiment, the data processing device includes a memory and one or more programs. One or more of the programs are stored in the memory, and one or more of the programs may include one or more modules, and each module may include a series of computer-executable instructions in the data processing device, and is configured to be executed by one or more processors. The one or more programs include the following computer-executable instructions: Construct prompt information based on the data in the preset database according to the preset large model instruction data synthesis requirements; Generate a plurality of synthetic data corresponding to the prompt information through a large language model based on the prompt information; Determine the importance of each synthetic data, and determine the coverage distribution information of the multiple synthetic data. Based on the importance of each synthetic data and the coverage distribution information of the multiple synthetic data, perform data enhancement processing on the multiple synthetic data to obtain multiple synthetic data after data enhancement; Based on the multiple synthetic data after data enhancement, determine the large model instruction data that meets the requirements of large model instruction data synthesis.

[0093] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the embodiment of the data processing device, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the corresponding part of the method embodiment for the relevant content.

[0094] The embodiments of this specification provide a data processing device. By based on the preset requirements for synthesizing large model instruction data, construct prompt information through the data in the preset database. Then, based on this prompt information, generate multiple synthetic data corresponding to the prompt information through a large language model. After that, determine the importance of each synthetic data, and determine the coverage distribution information of the multiple synthetic data. Based on the importance of each synthetic data and the coverage distribution information of the multiple synthetic data, perform data enhancement processing on the multiple synthetic data to obtain multiple synthetic data after data enhancement. Finally, based on the multiple synthetic data after data enhancement, determine the large model instruction data that meets the requirements of large model instruction data synthesis. In this way, a standardized data synthesis framework is provided. Through the standardized data synthesis framework, perform data enhancement processing based on the dual dimensions of the importance of the synthetic data and the diversity corresponding to the coverage distribution information, improve the comprehensive value of the synthetic data, enhance the data quality and relevance of the large model instruction data, and can fully utilize the potential value of the data during the data synthesis process. In addition, reduce manual intervention, improve the efficiency and automation degree of data synthesis, expand the depth and breadth of the data, and effectively double improve the quality and value of large model instruction data synthesis.

[0095] Further, based on the above Figures 1 to 11 , one or more embodiments of this specification also provide a storage medium for storing computer-executable instruction information. In a specific embodiment, the storage medium can be a USB flash drive, an optical disc, a hard disk, etc. When the computer-executable instruction information stored in the storage medium is executed by a processor, the following process can be realized: Based on the preset requirements for synthesizing large model instruction data, construct prompt information through the data in the preset database; Based on the prompt information, generate multiple synthetic data corresponding to the prompt information through a large language model; Determine the importance of each synthetic data, and determine the coverage distribution information of the multiple synthetic data. Based on the importance of each synthetic data and the coverage distribution information of the multiple synthetic data, perform data enhancement processing on the multiple synthetic data to obtain multiple synthetic data after data enhancement; Based on the multiple synthetic data after data enhancement, determine the large model instruction data that meets the requirements of the large model instruction data synthesis.

[0096] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the above-mentioned storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.

[0097] This specification embodiment provides a storage medium. By constructing prompt information based on the preset requirements for synthesizing large model instruction data through the data in the preset database, then, based on this prompt information, a large language model can be used to generate multiple synthetic data corresponding to this prompt information. After that, the importance of each synthetic data can be determined, and the coverage distribution information of the multiple synthetic data can be determined. Based on the importance of each synthetic data and the coverage distribution information of the multiple synthetic data, perform data enhancement processing on the multiple synthetic data to obtain multiple synthetic data after data enhancement. Finally, based on the multiple synthetic data after data enhancement, determine the large model instruction data that meets the requirements of the large model instruction data synthesis. In this way, a standardized data synthesis framework is provided. Through the standardized data synthesis framework, data enhancement processing is performed based on the dual dimensions of the importance of the synthetic data and the diversity corresponding to the coverage distribution information, improving the comprehensive value of the synthetic data, enhancing the data quality and relevance of the large model instruction data, and being able to fully utilize the potential value of the data during the data synthesis process. In addition, it reduces manual intervention, improves the efficiency and automation degree of data synthesis, expands the depth and breadth of the data, and realizes the effective dual improvement of the quality and value of the large model instruction data synthesis.

[0098] Further, based on the above Figures 1 to 11 , one or more embodiments of this specification also provide a computer program product, including a computer program. When the computer program in this computer program product is executed by a processor, it can implement the following processes: Construct prompt information based on the preset requirements for synthesizing large model instruction data through the data in the preset database; Based on the prompt information, use a large language model to generate multiple synthetic data corresponding to the prompt information; Determine the importance of each synthetic data, and determine the coverage distribution information of the multiple synthetic data. Based on the importance of each synthetic data and the coverage distribution information of the multiple synthetic data, perform data enhancement processing on the multiple synthetic data to obtain multiple synthetic data after data enhancement; Based on the multiple synthetic data after data enhancement, determine the large model instruction data that meets the requirements for synthesizing large model instruction data.

[0099] Each embodiment in this specification is described in a progressive manner. For the same or similar parts between each embodiment, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the above embodiment of a computer program product, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.

[0100] The embodiments of this specification provide a computer program product. By based on the preset requirements for synthesizing large model instruction data, construct prompt information through the data in the preset database. Then, based on this prompt information, generate multiple synthetic data corresponding to this prompt information through a large language model. After that, it is possible to determine the importance of each synthetic data and determine the coverage distribution information of the multiple synthetic data. Based on the importance of each synthetic data and the coverage distribution information of the multiple synthetic data, perform data enhancement processing on the multiple synthetic data to obtain multiple synthetic data after data enhancement. Finally, based on the multiple synthetic data after data enhancement, determine the large model instruction data that meets the requirements for synthesizing large model instruction data. In this way, a standardized data synthesis framework is provided. Through the standardized data synthesis framework, perform data enhancement processing based on the dual dimensions of the importance of synthetic data and the corresponding diversity of coverage distribution information, improve the comprehensive value of synthetic data, enhance the data quality and relevance of large model instruction data, and can make full use of the potential value of data during the data synthesis process. In addition, reduce manual intervention, improve the efficiency and automation degree of data synthesis, expand the depth and breadth of data, and achieve an effective double improvement in the quality and value of large model instruction data synthesis.

[0101] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0102] In the 1990s, an improvement in a technology could be clearly distinguished as either a hardware improvement (e.g., an improvement in the circuit structure of diodes, transistors, switches, etc.) or a software improvement (an improvement in the method flow). However, with the development of technology, many method flow improvements today can be regarded as direct improvements in the hardware circuit structure. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is such an integrated circuit whose logic function is determined by the user's programming of the device. The designer can program by himself to "integrate" a digital system on a piece of PLD without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development writing. The original code before compilation also has to be written in a specific programming language, which is called a hardware description language (HDL). And there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be clear that as long as the method flow is slightly logically programmed with the above-mentioned several hardware description languages and programmed into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0103] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0104] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0105] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0106] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0107] The embodiments of this specification are described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable serial-parallel devices for fraud cases to generate a machine, so that the instructions executed by the processors of the computer or other programmable serial-parallel devices for fraud cases generate means for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0108] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable serial-parallel device for fraud cases to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0109] These computer program instructions can also be loaded onto a computer or other programmable serial-parallel device for fraud cases, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0110] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0111] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0112] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0113] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0114] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Thus, one or more embodiments of this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.

[0115] One or more embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0116] The various embodiments in this specification are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.

[0117] The above is only the embodiments of this specification and is not intended to limit this document. For those skilled in the art, various changes and modifications can be made to this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.

Claims

1. A data processing method, the method comprising: Constructing prompt information based on the data in a preset database according to the preset large model instruction data synthesis requirements; Generating a plurality of synthetic data corresponding to the prompt information through a large language model based on the prompt information; Determining the importance of each synthetic data, and determining the coverage distribution information of the plurality of synthetic data. Based on the importance of each synthetic data and the coverage distribution information of the plurality of synthetic data, performing data value-added processing on the plurality of synthetic data to obtain a plurality of synthetic data after data value-added processing; Determining large model instruction data that meets the large model instruction data synthesis requirements based on the plurality of synthetic data after data value-added processing.

2. The method according to claim 1, wherein the preset database includes a seed data set and an externally collected data set. The constructing prompt information based on the data in the preset database according to the preset large model instruction data synthesis requirements includes: Constructing a plurality of different prompt information examples based on the data in the seed data set according to the preset large model instruction data synthesis requirements; Constructing prompt information that matches the prompt information examples based on the data in the seed data set and / or the externally collected data set according to the preset large model instruction data synthesis requirements.

3. The method according to claim 1 or 2, wherein the generating a plurality of synthetic data corresponding to the prompt information through a large language model based on the prompt information includes: Inputting the prompt information into a plurality of different large language models respectively to obtain first synthetic data output by each large language model; Performing quality evaluation on each first synthetic data through the large language model to obtain an evaluation result of each first synthetic data; Performing reflection processing on each first synthetic data through the large language model based on the evaluation result of each first synthetic data to obtain a reflection result corresponding to each first synthetic data; Performing correction processing on the corresponding first synthetic data through the large language model based on the reflection result corresponding to each first synthetic data to obtain a corrected first synthetic data; Determining a plurality of synthetic data corresponding to the prompt information based on the corrected first synthetic data.

4. The method according to claim 3, wherein the determining the importance of each synthetic data includes: Fine-tuning the large language model based on a plurality of synthetic data corresponding to the prompt information to obtain a fine-tuned large language model; Inputting the prompt information in the synthetic data into the fine-tuned large language model to obtain first answer data corresponding to the prompt information in each synthetic data; Calculating the cross entropy between the first answer data corresponding to the prompt information in each synthetic data and the corresponding synthetic answer data in the synthetic data; Determining the importance of each synthetic data based on the calculated cross entropy.

5. The method according to claim 4, wherein the determining the coverage distribution information of the plurality of synthetic data includes: Determining the probability of outlier data appearing in a plurality of synthetic data based on a preset anomaly detection algorithm, and the anomaly detection algorithm includes a similarity algorithm; Determine the coverage distribution information of the multiple synthetic data based on the outlier data probabilities appearing in the multiple synthetic data.

6. The method according to claim 5, wherein the method further comprises: Obtain multiple first instruction data from a preset open-source dataset; Determine the importance of each first instruction data, and determine the coverage distribution information of the multiple first instruction data. Based on the importance of each first instruction data and the coverage distribution information of the multiple first instruction data, perform data augmentation processing on the multiple first instruction data to obtain multiple first instruction data after data augmentation; The determining the large model instruction data that meets the large model instruction data synthesis requirements based on the multiple synthetic data after data augmentation includes: Based on the multiple synthetic data after data augmentation and the multiple first instruction data after data augmentation, determine the large model instruction data that meets the large model instruction data synthesis requirements.

7. The method according to claim 6, wherein the determining the large model instruction data that meets the large model instruction data synthesis requirements based on the multiple synthetic data after data augmentation includes: Perform data quality assessment processing on the multiple synthetic data after data augmentation respectively based on a preset data quality assessment rule to obtain a data quality assessment result corresponding to each synthetic data after data augmentation, and the data quality assessment result is used to indicate the degree of gain of the synthetic data after data augmentation to the large language model; Based on the data quality assessment result corresponding to each synthetic data after data augmentation, determine the large model instruction data that meets the large model instruction data synthesis requirements.

8. The method according to claim 7, wherein the performing data quality assessment processing on the multiple synthetic data after data augmentation respectively based on a preset data quality assessment rule to obtain a data quality assessment result corresponding to each synthetic data after data augmentation includes: Based on the multiple synthetic data after data augmentation and the large model instruction data synthesis requirements, construct a data quality assessment index, and the data quality assessment index includes one or more of the distribution difference between the synthetic data after data augmentation and the real data, the complexity of the data, and the perplexity of the data; Based on the data quality assessment index, select one or more synthetic data after data augmentation that meet the data quality assessment rules corresponding to the data quality assessment index from the multiple synthetic data after data augmentation; Based on the selected one or more synthetic data after data augmentation, determine the data quality assessment result corresponding to the selected one or more synthetic data after data augmentation through the large language model.

9. A data processing device, the device comprising: A data processing module that constructs a prompt message based on data in a preset database according to a preset large model instruction data synthesis requirement; A data synthesis module that generates multiple synthetic data corresponding to the prompt message through a large language model based on the prompt message; A data value-added module determines the importance of each synthetic data and determines the coverage distribution information of the multiple synthetic data. Based on the importance of each synthetic data and the coverage distribution information of the multiple synthetic data, data value-added processing is performed on the multiple synthetic data to obtain multiple synthetic data after data value-added processing; An instruction data determination module determines large model instruction data that meets the large model instruction data synthesis requirements based on the multiple synthetic data after data value-added processing.

10. A data processing device, the data processing device comprising: A processor; And A memory arranged to store computer-executable instructions, the executable instructions, when executed, cause the processor to: Construct prompt information based on data in a preset database based on preset large model instruction data synthesis requirements; Generate multiple synthetic data corresponding to the prompt information through a large language model based on the prompt information; Determine the importance of each synthetic data and determine the coverage distribution information of the multiple synthetic data. Based on the importance of each synthetic data and the coverage distribution information of the multiple synthetic data, data value-added processing is performed on the multiple synthetic data to obtain multiple synthetic data after data value-added processing; Determine large model instruction data that meets the large model instruction data synthesis requirements based on the multiple synthetic data after data value-added processing.

Citation Information

Cited By

  • Instruction fine tuning data set generation method, electronic equipment and storage medium

    CN121579077A

  • An instruction fine-tuning dataset generation method, an electronic device, and a storage medium

    CN121579077B