Data processing method and device, computer equipment and storage medium
By training LLM to identify and modify various quality defects in AI-generated prompt data, the integrity and quality issues of the prompt dataset were resolved, and high-quality updates of the dataset and improvements in LLM's instruction-following capabilities were achieved.
Patent Information
- Application Number
- CN202410276010.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-11
- Publication Date
- 2025-09-12
AI Technical Summary
In existing technologies, the quality of prompt data generated by AI models is difficult to guarantee. Directly discarding low-quality data will damage data integrity, resulting in a reduced ability of LLM to comply with certain types of instructions. Existing cleaning methods cannot effectively eliminate deep-seated quality problems.
The modification instructions and expected responses are learned through a large language model (LLM). The LLM is trained with multiple training samples to identify and modify various quality defects in the prompt data, generating a high-quality modified prompt dataset.
It improves the data quality and integrity of the prompt data set, ensures LLM's ability to comply with instructions, and is suitable for a variety of scenarios including private domains, with wide applicability.
Smart Images

Figure CN120632441A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a data processing method, apparatus, computer equipment, and storage medium. Background Art
[0002] Large language models (LLMs) have the ability to execute tasks and provide responses based on user instructions. To enable LLMs to respond to user instructions, they can be trained using a large amount of prompt data contained in a prompt dataset. Prompt data consists of an instruction and its corresponding expected response. The instruction is the user's instruction, and the expected response represents the desired output response to the instruction. Therefore, providing a high-quality prompt dataset is crucial to improving LLMs' responsiveness.
[0003] At present, with the development of AI technology, AI models can be used to automatically generate a large amount of diverse prompt data to construct a prompt data set. However, the data quality of the prompt data generated by the AI model is difficult to guarantee. Based on this, the relevant technology can use the existing LLM to perform quality scoring on each prompt data in the prompt data set, and discard a part of the prompt data with a lower quality score to obtain an optimized prompt data set. However, this method of directly discarding prompt data with a low quality score will damage the data integrity of the prompt data set, which may lead to the lack of certain types of instructions in the prompt data set. On this basis, using the optimized prompt data set to train the LLM may reduce the ability of the trained LLM to follow these types of instructions. Summary of the Invention
[0004] The present application provides a data processing method, apparatus, computer device, and storage medium, which can improve the data quality of prompt data in a prompt data set.
[0005] To achieve the above objectives, this application adopts the following technical solutions:
[0006] In a first aspect, a data processing method is provided, the method comprising: obtaining a first modification instruction, the first modification instruction comprising first prompt data to be modified in a prompt data set, the first modification instruction being used to instruct modification of the first prompt data; processing the first modification instruction through a first large language model (LLM) to obtain a first response, the first response comprising second prompt data obtained after modification of the first prompt data, the first LLM being obtained by training through multiple training samples in a first training data set, wherein each training sample comprises a modification instruction sample and a corresponding expected response, the modification instruction sample comprising a prompt data sample to be modified, and the modification instruction sample being used to instruct modification of the prompt data sample, and the expected response comprising an expected modified prompt data sample.
[0007] In the present application, a first modification instruction for indicating that the first prompt data is to be modified is obtained. Afterwards, the first modification instruction is processed by the first LLM to obtain a first response, and the first response includes the second prompt data obtained after the first prompt data is modified. Among them, the first LLM is obtained by training with multiple training samples, and each training sample contains a modification instruction and a corresponding expected response. The modification instruction is used to indicate that the prompt data to be modified is to be modified, and the expected response includes the expected modified prompt data. It can be seen that the first LLM obtained by training with multiple training samples will be able to learn the ability to modify the prompt data to be modified. On this basis, the prompt data in the prompt data set is modified by the first LLM to obtain the modified prompt data set. In this way, not only the data quality of the prompt data set can be improved, but also the data integrity of the prompt data set can be guaranteed.
[0008] Optionally, the method also includes: obtaining the first training data set; obtaining multiple target training samples from the first training data set, wherein the data quality of the multiple target training samples is higher than the data quality of other training samples in the first training data set; and using the multiple target training samples to train the first initial model to obtain the first LLM.
[0009] In the present application, a plurality of target training samples with higher data quality can be screened out from the training samples included in the first training data set to train the first initial model, thereby improving the ability of the first LLM finally trained to follow the user's modification instructions.
[0010] Optionally, the implementation process of obtaining multiple target training samples from the first training data set may include: determining the difference value between the prompt data sample to be modified and the modified prompt data sample included in each training sample in the first training data set; and obtaining the multiple target training samples based on the difference value corresponding to each training sample.
[0011] In the present application, the purpose of training the first initial model using training samples is to enable the first initial model to learn the ability to modify the prompt data sample to be modified in the modification instruction sample into the modified prompt data sample in the expected response. Therefore, the greater the difference between the prompt data sample to be modified in the modification instruction sample in the training sample and the modified prompt data sample in the expected response, the greater the modification ability that the first initial model can learn from the training sample, and accordingly, the stronger the modification ability of the first LLM obtained by the final training to modify the prompt data.
[0012] Optionally, the implementation process of obtaining the multiple target training samples based on the difference value corresponding to each training sample may include: sorting the multiple training samples in the first training data set in descending order of the difference value corresponding to each training sample to obtain a sorting result; determining the first N training samples in the sorting result as the multiple target training samples, where N is a preset value or N is determined based on a preset quality coefficient and the total number of training samples in the first training data set.
[0013] In the present application, the number of target training samples to be obtained can be controlled by the size of N, and the number of target training samples will affect the training effect of the first initial model.
[0014] Optionally, the implementation process of training the first initial model using the multiple target training samples may include: processing the modification instruction sample in each target training sample through the first initial model to obtain the predicted response corresponding to each modification instruction sample; and updating the model parameters of the first initial model based on the predicted response and expected response corresponding to each modification instruction sample.
[0015] Optionally, the multiple prompt data samples to be modified in the multiple training samples include quality defects in multiple dimensions, and the multiple dimensions include readability, feasibility and contextualization of instructions, and readability, correctness, relevance, richness, humanity, safety and comprehensiveness of responses.
[0016] In the present application, the prompt data samples in the multiple training samples used to train the first initial model can contain not only shallow-level quality defects such as language, form, simple logic, etc., but also deep-level quality defects in various semantic aspects such as infeasible instructions, irrelevant responses, etc. In this way, the first initial model can not only learn the ability to modify prompt data with shallow-level quality defects from these training samples, but also learn the ability to modify prompt data with deep-level quality defects. In this way, the subsequent modification of prompt data through the first LLM can better improve the data quality of the prompt data set.
[0017] Optionally, the method further includes: training a second initial model using the second prompt data to obtain a second LLM.
[0018] In the present application, after obtaining the second prompt data, the data processing device may further use the second prompt data to train the second initial model to obtain a second LLM. Since the data quality of the second prompt data is improved, the second LLM obtained by using the second prompt data to train the second initial model will have better instruction-following capabilities.
[0019] In a second aspect, a data processing device is provided, which includes at least one module, and the at least one module is used to execute the data processing method described in the first aspect.
[0020] In a third aspect, a computer device is provided, comprising a processor configured to execute at least one program instruction or code stored in a memory to implement the data processing method described in the first aspect.
[0021] In a fourth aspect, a computer device cluster is provided, wherein the computer device cluster includes at least one computer device, each computing device includes a processor and a memory; the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the data processing method described in the first aspect above.
[0022] In a fifth aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer device, the computer device executes the data processing method described in the first aspect.
[0023] In a sixth aspect, a computer program product comprising instructions is provided, which, when executed on a computer device, enables the computer device to execute the data processing method described in the first aspect.
[0024] The technical effects obtained in the above-mentioned second to sixth aspects are similar to the technical effects obtained by the corresponding technical means in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A flowchart of training a first initial model to obtain a first LLM is provided in an embodiment of the present application;
[0026] Figure 2 A flowchart of a data processing method provided in an embodiment of the present application;
[0027] Figure 3 A distribution diagram of the quality scores of responses in an original alpaca52k dataset and a modified alpaca52k dataset provided in an embodiment of the present application;
[0028] Figure 4 A graph showing the win rates of alpaca-coachLM evaluated by the pandaLM and GPT-4 models at different quality coefficients provided in the embodiments of this application;
[0029] Figure 5 A flowchart of a data processing device provided in an embodiment of the present application;
[0030] Figure 6 A schematic diagram of the structure of a computer device provided in an embodiment of the present application;
[0031] Figure 7 A schematic diagram of a computer device cluster provided in an embodiment of the present application;
[0032] Figure 8 A schematic diagram of another computer device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION
[0033] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0034] Before explaining the embodiments of the present application in detail, the application scenarios involved in the embodiments of the present application are first introduced.
[0035] The rapid development of large language models (LLMs) has had a profound impact on various fields. LLMs have the ability to perform complex tasks and provide responses based on user instructions, and this ability of LLMs can be obtained through three stages of training. Among them, the first stage is the pre-training stage, in which the basic model can be trained with a large corpus to obtain a pre-trained model that can predict the subsequent text of the input text based on the user's input text. The second stage is the prompt instruction fine-tuning stage, in which the pre-trained model can be adjusted with a large amount of prompt data so that the pre-trained model can acquire the ability to respond to user instructions. The third stage is the reinforcement learning (RL) stage, so that the model trained with prompt data can further learn the boundaries of the response to avoid the model generating harmful or sensitive content. Among the three stages mentioned above, the prompt instruction fine-tuning stage is the key process that enables LLM to accurately respond to user instructions and obtain AI generalization capabilities. Based on this, the data quality of the prompt data used to train the pre-trained model in this stage is crucial.
[0036] Typically, prompt data may include instructions and corresponding expected responses. Instructions refer to human instructions, and expected responses refer to the responses expected to be obtained for the human instructions. In related technologies, high-quality prompt data can be manually written to construct a prompt data set. However, since the amount of prompt data required for training is large, and manual writing of prompt data requires the writer to have more comprehensive knowledge, the cost is relatively high. Based on this, AI models can currently be used to automatically generate a large amount of diverse prompt data to construct a prompt data set. However, the data quality of the prompt data generated by the AI model is difficult to guarantee, so the prompt data in the prompt data set needs to be optimized.
[0037] In some related technologies, a pre-trained LLM can be used to perform a quality score on each prompt in a prompt dataset and discard the low-quality prompts to obtain an optimized prompt dataset. However, this method of directly discarding low-quality prompts compromises the integrity of the prompt dataset, potentially resulting in the loss of certain types of instructions in the prompt dataset. Consequently, using the optimized prompt dataset to train the LLM to be trained may reduce the trained LLM's ability to follow these types of instructions.
[0038] Other related technologies use predefined regular expressions to clean prompt data sets, removing low-quality prompt data. However, this method of using regular expressions to clean prompt data often only removes prompt data with formal errors or simple logical errors, but cannot remove prompt data with deeper quality issues such as semantic defects. Therefore, this cleaning method has limited effect on improving the data quality of prompt datasets.
[0039] In response to the various problems mentioned above, an embodiment of the present application provides a data processing method, in which, for prompt data in a prompt data set, such as first prompt data, a data processing device can obtain a first modification instruction for instructing to modify the first prompt data. Afterwards, the first modification instruction is processed by the first LLM to obtain a first response, which includes the second prompt data obtained after modifying the first prompt data. The first LLM is obtained by training with multiple training samples, each training sample contains a modification instruction sample and a corresponding expected response, the modification instruction sample is used to indicate that the prompt data sample to be modified is to be modified, and the expected response includes the expected modified prompt data sample. It can be seen that the first LLM obtained by training with multiple training samples will be able to learn the ability to modify the prompt data to be modified. On this basis, the prompt data in the prompt data set is modified by the first LLM to obtain a modified prompt data set. In this way, not only the data quality of the prompt data set can be improved, but also the data integrity of the prompt data set can be guaranteed.
[0040] In addition, compared to the method of cleaning prompt data through regular expressions, in the embodiment of the present application, training samples containing prompt data with multiple quality defects can be selected as needed to train the first LLM. In this way, the first LLM can learn the ability to identify and modify prompt data with multiple quality defects. In this way, by modifying the prompt data through the first LLM, the data quality of the prompt data set can be better improved.
[0041] Next, the data processing method provided in the embodiment of the present application is explained in detail.
[0042] The data processing method provided in the embodiment of the present application can be applied to a data processing device. In this method, after receiving a modification instruction containing prompt data to be modified, the data processing device processes the modification instruction through the first LLM, thereby outputting a response containing the modified prompt data. The first LLM is obtained by training using multiple training samples in the first training data set. Based on this, the process of obtaining the first LLM is first introduced. Figure 1 , the process may include the following steps:
[0043] Step 101: Obtain a first training data set.
[0044] In an embodiment of the present application, the data processing device may obtain multiple prompt data from the prompt data set, and then obtain a first training data set based on the multiple prompt data.
[0045] The prompt data set can include a large amount of prompt data to be modified using the AI model. In addition, it can also include some manually written prompt data. Each prompt data can include an instruction and a corresponding response. For example, the prompt data can be expressed as (instruction, response).
[0046] The data processing device can extract multiple pieces of prompt data from the prompt data set as prompt data samples to be modified, and sequentially display each prompt data sample to be modified, allowing a professional to modify the corresponding prompt data sample, thereby obtaining a modified prompt data sample. Subsequently, a plurality of training samples are generated based on the plurality of prompt data samples to be modified and the corresponding modified prompt data samples, and a first training data set is obtained based on the plurality of training samples. The first training data set includes the plurality of training samples.
[0047] Exemplarily, the plurality of prompt data samples to be modified extracted by the data processing device from the prompt data set may include quality defects of multiple dimensions. That is, the plurality of prompt data samples to be modified may include quality defects of different types.
[0048] For example, the multiple dimensions may include readability, feasibility, and contextualization of instructions, and at least two of readability, correctness, relevance, comprehensiveness, richness, humanity, and safety of responses.
[0049] Among them, the readability of the instructions mainly refers to whether the instructions comply with language conventions and style specifications. Specifically, the readability of the instructions can be mainly reflected by the grammar, spelling, and punctuation of the instructions. Based on this, in the embodiment of the present application, the quality defect of the readability of the instructions in the prompt data sample to be modified means that the instructions in the prompt data sample do not comply with language conventions or problem specifications and have poor readability. Specifically, the instructions in the prompt data sample may have one or more errors in grammar, spelling, and punctuation.
[0050] The feasibility of an instruction mainly refers to whether the instruction is clear, specific, feasible and easy to understand. Specifically, the feasibility of an instruction is mainly reflected in whether the instruction has ambiguous or vague expressions, whether there are logical errors, and whether the content requested by the instruction exceeds the capabilities of the AI model. Based on this, in an embodiment of the present application, the quality defect of the prompt data sample to be modified in the dimension of instruction feasibility may refer to the problem that the instructions in the prompt data sample are unclear, non-specific, infeasible or not easy to understand. Specifically, it may be that the instruction is ambiguous or vaguely expressed, or there is a logical error, or the instruction is an instruction that the AI model cannot execute, etc.
[0051] The contextualization of instructions mainly refers to whether the instructions contain rich context or effective prompting skills, which can facilitate detailed and accurate responses. Specifically, the contextualization of instructions is mainly reflected in whether the instructions contain scenes, roles, examples or other demand information, and whether they contain prompt information such as thought chains. Based on this, in an embodiment of the present application, the quality defect of the contextualization of instructions in the prompt data sample to be modified may refer to the lack of context or effective prompting skills in the instructions of the prompt data sample. For example, the instruction may lack scene information, examples, role information or other relevant demand information.
[0052] The readability of a response mainly refers to the fact that the language of the response is fluent, concise, the content is correct and the structure is appropriate. Specifically, the readability of a response can be mainly reflected in three aspects: language, content and structure. Among them, in terms of language, the readability of a response mainly refers to the fact that the response uses precise vocabulary and there are no incorrect spellings. In terms of content, the readability of a response mainly refers to the fact that the response contains meaningful content and there is no meaningless or redundant content. In terms of structure, the readability of a response mainly refers to the fact that the information in the response is clear, orderly and logical, and the layout of the information is user-friendly. Based on this, in an embodiment of the present application, the quality defect of the prompt data sample to be modified in the dimension of readability of the response may refer to the poor readability of the response in the prompt data sample. Specifically, it may refer to the fact that the response does not meet the requirements of at least one of the above-mentioned language, content and structure.
[0053] The correctness of the response mainly refers to whether the response is based on real information, common sense and logical reasoning, and can be kept up to date and meet the detailed needs of the user. Specifically, the lack of correctness of the response can be mainly reflected in five aspects. Among them, the first aspect is that the response contains factual errors, which is mainly manifested in that the content of the response does not conform to reality. The second aspect is that the response contains common sense errors, which is mainly manifested in that the content of the response violates human common sense. The third aspect is that the response contains logical errors, which is mainly manifested in that the response contains problems such as concept substitution, self-contradiction, ambiguity and circular reasoning. The fourth aspect is that the response does not follow constraints, where the constraints here include constraints on the number of words, genre and style of the response. The fifth aspect is that the response lacks timeliness, that is, the information contained in the response is not the latest. Based on this, in an embodiment of the present application, the quality defect of the prompt data sample to be modified in the dimension of correctness of the response can mean that the response in the prompt data sample does not meet the above-mentioned correctness requirements. Specifically, it may be that the response has problems described in one or more of the above five aspects.
[0054] The relevance of the response mainly means that the response should be effective and direct, and provide a solution that is consistent with the topic. Specifically, the lack of relevance of the response is mainly reflected in two aspects, among which the first aspect is that the response is irrelevant, which is mainly reflected in the fact that the response misinterprets the user's intention, and the second aspect is that the response is related to the user's topic but deviates from the focus. Based on this, in the embodiment of the present application, the quality defect of the dimension of relevance of the prompt data sample to be modified mainly refers to the lack of relevance of the response in the prompt data sample to the instruction. Specifically, it can be manifested in that the response in the prompt data sample has problems in any of the above two aspects.
[0055] The comprehensiveness of the response mainly refers to the ability of the response to comprehensively cover all necessary angles and information. Specifically, the comprehensiveness of the response can be mainly reflected in two aspects. Among them, the first aspect is that there are no omissions or deficiencies in the response in fully explaining the user's problem. The second aspect is that the response is multi-angle, rich in context and details. Based on this, in an embodiment of the present application, the quality defect of the prompt data sample to be modified in the dimension of comprehensiveness means that the response in the prompt data sample is not comprehensive enough and does not cover the necessary angles and information, which can be specifically reflected in not meeting the requirements of any of the above two aspects.
[0056] The richness of the response means that the information contained in the response should be diverse, useful, creative and extensible. Specifically, the richness of the response can be mainly reflected in two aspects. Among them, the first aspect is that the response provides detailed, diverse information with depth and breadth. The second aspect is that the response contains novel, unique and imaginative content. Based on this, in an embodiment of the present application, the quality defect of the dimension of richness of the prompt data sample to be modified may refer to the content of the response in the prompt data sample being not rich enough, specifically, it may refer to that the response does not meet the requirements of any of the above two aspects.
[0057] The humanity of the response means that the response should be warm, resonant and engaging, and the response is customized based on the user's background and preferences. Specifically, the humanity of the response is mainly reflected in two aspects. The first aspect is emotional perception, which is mainly reflected in the response responding to the user's emotions with empathetic content. The second aspect is humanized language, which is mainly reflected in the response interacting with the user in natural and friendly language, while avoiding robotic language. Based on this, in an embodiment of the present application, the quality defect of the dimension of humanity in the prompt data sample to be modified may refer to the lack of humanity in the response of the prompt data sample, which is specifically manifested in that the response does not meet the requirements of any of the above two aspects.
[0058] The security of the response means that the response should be harmless and non-sensitive. For example, the response should not contain information that is harmful to the user's emotions, body, and property. Specifically, the security of the response can be reflected in the fact that the response does not violate the law, does not contain personal attacks, does not expose user privacy, and does not contain irresponsible suggestions on issues in the fields of medical care and finance. Based on this, in an embodiment of the present application, the quality defect of the prompt data sample to be modified in the dimension of security can mean that the response in the prompt data sample is harmful and / or sensitive. Specifically, it can mean that the response does not meet any of the above requirements.
[0059] It should be noted that, as can be seen from the above introduction to quality defects in various dimensions, the multiple prompt data samples to be processed extracted by the data processing device may include not only prompt data samples with shallow defects such as language, format, and simple logical defects, but also prompt data samples with various deep defects such as infeasible instructions, lack of contextualization, and incorrect, irrelevant, and incomplete responses. In this way, the training samples in the first training data set obtained through the multiple prompt data samples to be processed will contain a rich variety of quality defects. On this basis, the first LLM trained based on the first training data set will be able to better identify and modify prompt data with different types of quality defects.
[0060] After extracting multiple prompt data samples to be modified from the prompt data set, the data processing device can display any prompt data sample to be modified so that professionals can make corresponding modifications and edits based on the quality defects of the prompt data sample, thereby obtaining a modified prompt data sample. Afterwards, the data processing device can generate a corresponding modification instruction sample based on the prompt data sample to be modified, and use the corresponding modified prompt data sample as the expected response corresponding to the modification instruction sample. A training sample is generated based on the modification instruction sample and the corresponding expected response. The training sample includes the modification instruction sample and the corresponding expected response. In this way, multiple training samples can be obtained based on the multiple prompt data samples to be modified and the corresponding modified prompt data samples modified by professionals.
[0061] For example, a data processing device acquires a prompt data sample to be modified, X. After a professional modifies X, Xr is obtained. In this case, the data processing device can generate a modification instruction sample based on X, and generate an expected response corresponding to the modification instruction sample based on Xr. For example, the modification instruction sample can be: improve the instructions and responses in X to make X more specific, detailed, with more logical steps and more accurate syntax, and the expected response is Xr. Accordingly, the training sample Xc generated based on the modification instruction sample and the corresponding expected response can be as follows:
[0062]
[0063] Optionally, in some possible implementations, the first training data set may also be obtained and stored in advance by other devices based on the above method. In this case, the data processing device may directly obtain the first training data set from the device that stores the first training data set.
[0064] Step 102: Train a first initial model using multiple training samples in a first training data set to obtain a first LLM.
[0065] After acquiring the first training data set, the data processing device can use multiple training samples in the first training data set to train the first initial model to obtain a first LLM. The first initial model can be an open source base model obtained by pre-training with a large amount of text data.
[0066] In one possible implementation, the data processing device may process the modification instruction sample in each training sample in the first training dataset using a first initial model to obtain a predicted response corresponding to the modification instruction sample. The model parameters of the first initial model may then be adjusted based on the predicted response and expected response corresponding to the modification instruction sample in each training sample, thereby training the first initial model.
[0067] Exemplarily, the data processing device may input the modification instruction sample in each training sample into the first initial model, thereby outputting the predicted response corresponding to each modification instruction sample through the first initial model, and the predicted response is the actual response after the first initial model executes the corresponding modification instruction sample. Thereafter, the data processing device may obtain the expected response corresponding to each modification instruction sample, and determine the conditional probability corresponding to the training sample containing the corresponding modification instruction sample based on the expected response and the predicted response corresponding to each modification instruction sample. The conditional probability refers to the probability that the predicted response output when the first initial model executes the modification instruction sample is the same as the expected response corresponding to the modification instruction sample. Thereafter, the data processing device may calculate the loss function value corresponding to each training sample based on the conditional probability corresponding to each training sample, and then calculate the average loss function value based on the loss function value corresponding to each training sample, and update the model parameters of the first initial model based on the average loss function value, thereby completing a round of training for the first initial model. At this time, the average loss function value is the average loss function value of this round of training.
[0068] Taking any training sample as an example, the data processing device may calculate the similarity value between the predicted response and the expected response corresponding to the modification instruction sample in the training sample, and use the similarity value as the conditional probability corresponding to the training sample.
[0069] It should be noted that, in some possible cases, the training samples in the first training data set may also be divided into multiple batches. The data processing device may, based on the training samples of each batch, update the model parameters of the first initial model once with reference to the above method. In this way, the model parameters of the first initial model are updated multiple times based on the training samples of multiple batches, thereby completing one round of training for the first initial model. In this case, the data processing device may also calculate the average of the average loss function values corresponding to the training samples of each batch during the training round to obtain the average loss function value of the training round.
[0070] After completing a first round of training of the first initial model using training samples in the first training dataset, the data processing device may refer to the above method to perform a second round of training on the updated first initial model using training samples in the first training dataset, and determine the average loss function value of the second round. If the average loss function value of the second round decreases compared to the average loss function value of the first round, the data processing device may continue with a third round of training until the average loss function value of M consecutive rounds no longer decreases, and use the most recently updated model as the first LLM. Where M can be an integer greater than 1.
[0071] In another possible implementation, the data processing device may obtain multiple target training samples from the first training dataset and use the multiple target training samples to train the first initial model to obtain the first LLM, wherein the data quality of the multiple target training samples is higher than the data quality of other training samples in the first training dataset.
[0072] It should be noted that, in the embodiment of the present application, the use of training samples containing modification instruction samples and corresponding expected responses to train the first initial model is to promote the alignment of the predicted response of the modification instruction samples output by the first initial model based on its own stored knowledge with the expected response. However, low-quality training samples will hinder the first initial model from correctly establishing the connection between its stored knowledge and the user's modification instruction samples. Based on this, in the embodiment of the present application, the data processing device can filter out multiple target training samples with higher data quality from the training samples contained in the first training data set to train the first initial model, so as to improve the ability of the first LLM finally trained to follow the user's modification instruction samples.
[0073] Specifically, in an embodiment of the present application, the first initial model is trained using training samples in order to enable the first initial model to learn the ability to modify the prompt data sample to be modified in the modification instruction sample to the modified prompt data sample in the expected response. Therefore, the greater the difference between the prompt data sample to be modified in the modification instruction sample in the training sample and the modified prompt data sample in the expected response, the more modification ability the first initial model can learn from the training sample; conversely, the smaller the difference between the prompt data sample to be modified in the modification instruction sample in the training sample and the modified prompt data sample in the expected response, the smaller the modification ability the first initial model can learn from the training sample. Based on this, in an embodiment of the present application, the data processing device can determine the difference value between the prompt data sample to be modified and the modified prompt data sample included in each training sample in the first training data set; based on the difference value corresponding to each training sample, multiple target training samples are obtained from the first training data set.
[0074] Taking any training sample as an example, the data processing device can calculate the edit distance between the prompt data sample to be modified included in the modification instruction sample in the training sample and the modified prompt data sample in the expected response. The edit distance refers to the minimum number of single-character edits required to convert one of two character strings into another. It can be seen that the greater the edit distance between the two character strings, the greater the difference between the two character strings. Based on this, in an embodiment of the present application, the edit distance between the prompt data sample to be modified included in the modification instruction sample in the training sample and the modified prompt data sample in the corresponding expected response can be used as the difference value corresponding to the training sample to characterize the difference between the two prompt data samples.
[0075] Of course, in other possible implementations, the data processing device may also use other methods to determine the difference value between the prompt data sample to be modified in the training sample and the modified prompt data sample. For example, a trained AI model may be used to detect the similarity probability between the two, and the dissimilarity probability between the two may be calculated based on the similarity probability, and the dissimilarity probability may be used as the difference value between the two.
[0076] After determining the difference value corresponding to each training sample in the first training data set, the data processing device may sort the training samples in the first training data set in descending order of the difference value to obtain a sorting result. Thereafter, the first N training samples in the sorting result may be determined as the plurality of target training samples.
[0077] Here, N can be a preset fixed value. Alternatively, N can be determined based on a preset quality coefficient and the total number of training samples included in the sorting result. Here, the preset quality coefficient is a value greater than 0 and not greater than 1. For example, if the preset quality coefficient is 0.3 and the sorting result includes 10,000 training samples, then N = 10,000 * 0.3 = 3,000 can be determined, that is, the data processing device can obtain the first 3,000 training samples in the sorting result as target training samples.
[0078] After obtaining multiple target training samples, the data processing device can use the multiple target training samples to train the first initial model. The training process can refer to the process of training the first initial model using training samples introduced above, and will not be repeated here.
[0079] In an embodiment of the present application, a first initial model can be trained using multiple training samples to obtain a first LLM. Each training sample includes a modification instruction sample and a corresponding expected response, wherein the modification instruction sample is used to indicate that the prompt data sample to be modified is to be modified, and the expected response includes the expected modified prompt data sample. It can be seen that the first LLM obtained by training multiple training samples will be able to learn the ability to follow the modification instruction to modify the prompt data to be modified. On this basis, the prompt data in the prompt data set can be modified by the first LLM to obtain a modified prompt data set. In this way, not only the data quality of the prompt data set can be improved, but also the data integrity of the prompt data set can be guaranteed.
[0080] In addition, in an embodiment of the present application, the prompt data samples in the multiple training samples used to train the first initial model can contain not only shallow quality defects such as language, form, simple logic, etc., but also deep quality defects in various semantic aspects such as infeasible instructions, irrelevant responses, etc. In this way, the first initial model can not only learn the ability to modify prompt data with shallow quality defects from these training samples, but also learn the ability to modify prompt data with deep quality defects. In this way, the subsequent modification of prompt data through the first LLM can better improve the data quality of the prompt data set.
[0081] Finally, in an embodiment of the present application, the first initial model can be an open source base model, on this basis, the first initial model can be deployed in any scenario, and the first initial model can be trained by the above method to obtain a first LLM, and then the first LLM is used to optimize the prompt data. It can be seen that the data processing method provided in the embodiment of the present application can be applied to various scenarios, including industrial scenarios, especially in private domains with limited network access rights. It can also be deployed locally, and has a wide range of applications. In addition, for the first initial model deployed in a certain private professional field, the first initial model can also be trained using training samples with the characteristics of the field, so as to obtain a customized first LLM suitable for the field.
[0082] After the first LLM is obtained through training by the above method, the first LLM can be used to modify the prompt data to be modified in the prompt data set. The following takes the use of the first LLM to modify the first prompt data in the prompt data set as an example to introduce the implementation process in detail. Figure 2 , the process includes the following steps:
[0083] Step 201: Obtain a first modification instruction, where the first modification instruction includes first prompt data to be modified.
[0084] In an embodiment of the present application, the data processing device can generate a first modification instruction based on the first prompt data using a preset text format. The first modification instruction is used to instruct to modify the first prompt data, and the first modification instruction includes the first prompt data.
[0085] For example, if the first prompt data is Y, the first modification instruction may be: improve the instructions and responses in Y to make Y more specific, more detailed, have more logical steps and more accurate grammar.
[0086] Optionally, the data processing device may also receive a first modification instruction input by a user.
[0087] Step 202: Process the first modification instruction through the first LLM to obtain a first response, where the first response includes second prompt data obtained by modifying the first prompt data.
[0088] The data processing device may input the first modification instruction into the first LLM, process the first modification instruction through the first LLM, and output a first response.
[0089] Since the first LLM has learned the ability to follow the user's modification instructions to modify the prompt data to be modified, after receiving the first modification instruction, the first LLM can use its own stored knowledge to modify the first prompt data in the first modification instruction to obtain the second prompt data, and output a first response containing the second prompt data.
[0090] For each prompt data in the prompt data set, the data processing device may modify the prompt data using the above method, and replace the prompt data with the modified prompt data, thereby obtaining the modified prompt data set.
[0091] After obtaining the modified prompt dataset, the data processing device can further use the modified prompt dataset to train a second initial model to obtain a second LLM. Since the data quality of the modified prompt dataset is improved, the second LLM obtained by training the second initial model with the modified prompt dataset will have better instruction-following capabilities.
[0092] In an embodiment of the present application, a data processing device can obtain a first modification instruction for instructing to modify the first prompt data. Afterwards, the first modification instruction is processed by the first LLM to obtain a first response, which includes the second prompt data obtained after modifying the first prompt data. Among them, the first LLM is obtained by training with multiple training samples, each training sample contains a modification instruction sample and a corresponding expected response, and the modification instruction sample is used to indicate that the prompt data sample to be modified is modified, and the expected response includes the expected modified prompt data sample. It can be seen that the first LLM obtained by training with multiple training samples will be able to learn the ability to modify the prompt data to be modified. On this basis, the prompt data in the prompt data set is modified by the first LLM to obtain the modified prompt data set. In this way, not only the data quality of the prompt data set can be improved, but also the data integrity of the prompt data set can be guaranteed.
[0093] Next, the technical effects of the data processing method provided in the embodiment of the present application are tested and evaluated from multiple perspectives.
[0094] 1. Data Quality Assessment
[0095] Based on the data processing method provided above, the embodiment of the present application uses the first LLM obtained by training to modify the alpaca52k dataset generated by the machine in the related art, which contains 52,000 (52k) prompt data, to obtain a modified alpaca52k dataset. Thereafter, the modified alpaca52k dataset is processed using regular expressions to filter out invalid prompt data containing invalid characters or repeated character strings. The filtered invalid prompt data accounts for 1.3% of the total prompt data in the alpaca52k dataset. This part of invalid prompt data is replaced with the corresponding original prompt data. Thereafter, for the prompt data in the alpaca52k dataset used to train the first initial model, in order to protect data privacy, this part of prompt data can also be replaced with the corresponding original prompt data. This part of prompt data also accounts for 1.3% of the total prompt data in the alpaca52k dataset. In this way, the final modified alpaca52k dataset is obtained.
[0096] On this basis, the LLM of a third party is used to perform a quality score on the response in each prompt data in the modified alpaca52k dataset, and the LLM of the third party is used to perform a quality score on the response in each prompt data in the original alpaca52k dataset. Among them, the LLM of the third party adopted in the embodiment of the present application is the GPT-3.5-turbo model, and the quality score ranges from 0 to 5 points, wherein GPT is the abbreviation of generative pre-trained transformer. In this case, based on the quality score of each prompt data in the original alpaca52k dataset, the quality average score of the original alpaca52k dataset is calculated to be 3.95 points, while based on the quality score of each prompt data in the modified alpaca52k dataset, the quality average score of the modified alpaca52k dataset is calculated to be 4.31 points. It can be seen that the average quality of the modified alpaca52k dataset is significantly improved.
[0097] Also, see Figure 3 , the left figure shows the distribution of quality scores of responses in the prompt data in the original alpaca52k dataset, and the right figure shows the distribution of quality scores of responses in the prompt data in the modified alpaca52k dataset. It can be seen that in the original alpaca52k dataset, only 17.7% of the prompt data scored more than 4.5 points. However, in the alpaca52k dataset modified by the first LLM, this proportion increased significantly to 78.9%. This result shows that the alpaca52k dataset modified by the first LLM did not refine the alpaca52k dataset by discarding a large number of samples, but was mainly composed of high-quality prompt data, preserving the integrity of the original dataset. Therefore, it can positively influence the instruction tuning of the subsequent second LLM.
[0098] In view of the fact that the third-party LLM can only score the response to the prompt data, the embodiment of the present application also manually scores the prompt data according to the multiple dimensions introduced in the aforementioned embodiment to comprehensively assess the quality of the instructions and responses in the prompt data. To this end, the embodiment of the present application randomly selected 150 prompt data from the modified alpaca52k data set, and invited three independent reviewers (R1, R2 and R3) to score these 150 modified prompt data and the corresponding original prompt data. Among these 150 modified prompt data, 18 prompt data instructions were modified by the first LLM. On this basis, for the 150 modified prompt data and the corresponding original prompt data, the average quality score of the response obtained by each reviewer is shown in Table 1 below. Among them, the average quality scores of the responses in the 150 original prompt data evaluated by R1, R2 and R3 are respectively: 71.1 points, 71.2 points and 71.3 points, and the average value of the three is 71.2 points. The average quality scores of the responses in the 150 modified prompt data evaluated by R1, R2, and R3 were 73.9, 77.2, and 74, respectively, with an average of 75. This shows that after the first LLM modification, the responses in the prompt data received higher average scores in the evaluations of the three reviewers.
[0099] Table 1 Comparison of the scoring results of 150 modified prompt data and the corresponding original prompt data
[0100]
[0101] In addition, for the 18 modified prompts among the 150 modified prompts, the scoring results of each reviewer for these 18 prompts and the corresponding 18 original prompts are shown in Table 2. Among them, the average quality scores of the instructions in the 18 original prompts obtained by R1, R2, and R3 were 76.6, 74.7, and 77.2, respectively, with an average of 76.2; the average quality scores of the responses in the 18 original prompts obtained by R1, R2, and R3 were 67.9, 70, and 68.4, respectively, with an average of 68.8. The average quality scores of the instructions in the 18 modified prompt data evaluated by R1, R2, and R3 were 78.3, 79.6, and 79.1, respectively, with an average of 79 points. The average quality scores of the responses in the 18 modified prompt data evaluated by R1, R2, and R3 were 75.3, 81.8, and 75.6, respectively, with an average of 77.6 points. This shows that after the first LLM modification, the instructions and responses received higher average scores in the evaluations of the three reviewers. Furthermore, it is worth noting that in the 18 prompt data with modified instructions, the improvement in responses was more significant, which shows the importance of practical and accurate instructions in improving the quality of corresponding responses.
[0102] Table 2 Comparison of the scoring results of the modified prompt data and the corresponding original prompt data for 18 instructions
[0103]
[0104] 2. Evaluation of the model capability of the LLM trained using the modified prompt dataset
[0105] Based on the modified alpaca52k dataset, the embodiment of the present application uses the modified alpaca52k dataset to train the second initial model, resulting in alpaca-coachLM. The second initial model corresponding to alpaca-coachLM uses the same settings as the base model corresponding to the alpaca model in the related art. The difference is that the training dataset used to train alpaca-coachLM is the modified alpaca52k dataset, while the original alpaca52k dataset is used to train the alpaca model in the related art.
[0106] In the embodiments of the present application, different test sets are used to test the performance of alpaca-coachLM, each model in the baseline group LLM, and each model in the enhanced group LLM. Among them, each model in the baseline group LLM has the same number of parameters, that is, the number of trainable parameters in each model is 7B (billion). In addition, the amount of training data for each model is similar to LLaMA and has been fine-tuned by prompt instructions. Each model in the enhanced group LLM is more powerful than each model in the baseline group LLM. Among them, the enhanced group LLM can include models that are larger in scale (for example, the number of parameters is 13B) and have been fine-tuned using private prompt data sets. In addition, the enhanced group LLM can also include models that have undergone RL.
[0107] Table 3 compares the test results of alpaca-coachLM and the baseline LLM on multiple different test sets. As shown in Table 3, the baseline LLM includes the vicuna-7b model, the alpaca model, the alpaca-cleaned model, the alpaca-pandaLM model, and the alpaGasus model. The test sets include the coachLM150 test set, the pandaLM170 test set, the vicuna80 test set, and the self-instruct252 test set. The coachLM150 test set includes 150 test samples, the pandaLM170 test set includes 170 test samples, the vicuna80 test set includes 80 test samples, and the self-instruct252 test set includes 252 test samples. For each model in a test set, the pandaLM output scores the predicted response and the expected response for the instruction in the corresponding sample. The quality scores of the predicted and expected responses are compared to determine whether the model wins, loses, or draws on the sample. Next, the model's performance on the test set is determined based on the number of wins, losses, or draws: win rate 1 (WR1), win rate 2 (WR2), and quality score (QS). WR1 = (number of wins + 0.5 * number of draws) / total number of samples, where total number of samples refers to the total number of samples in the test set; WR2 = number of wins / (total number of samples - number of draws); and QS = (number of wins + number of draws) / total number of samples, which measures the proportion of predicted responses that meet the expected response.
[0108] Table 3 Comparison of test results of alpaca-coachLM and baseline group LLM
[0109]
[0110] The comparison in Table 3 shows that alpaca-coachLM is further enhanced after training on the modified alpaca52k dataset. The values of the three performance parameters all exceed the corresponding parameter values of all models in the baseline group LLM. This shows that the performance of alpaca-coachLM exceeds that of the models in the baseline group LLM.
[0111] Table 4 Comparison of test results of alpaca-coachLM and enhanced group LLM
[0112]
[0113] Table 4 is a comparison table of the test results of alpaca-coachLM and the enhanced group LLM on multiple different test sets. As shown in Table 4, the enhanced group LLM includes the LLaMA2-13b-chat model, the vicuna-13b model, the LLaMA2-7b-chat model, the ChatGLM model, and the ChatGLM2 model. Among them, except for the vicuna-13b model, the other models are all models that have undergone RL. In addition, the number of trainable parameters of the LLaMA2-13b-chat model and the vicuna-13b model is 13B, the number of trainable parameters of the LLaMA2-7b-chat model and the alpaca-coachLM trained by the embodiment of the present application is 7B, and the number of trainable parameters of the ChatGLM model and the ChatGLM2 model is 6B. The test set is still the four test sets in the above example. Accordingly, the values of the three performance parameters WR1, WR2 and QS corresponding to each model on each test set are shown in Table 4.
[0114] The test results shown in Table 4 show that the alpaca-coachLM trained in the examples of the present application still achieved impressive results compared to the various models in the enhanced group. Among the 12 performance parameter values compared on the four test sets, 5 performance parameter values (see the bold values in Table 4) were the highest. Moreover, on all four test sets, the performance parameter values of the alpaca-coachLM with 7B training parameters were better than those of the vicuna-13b model with 13B training parameters.
[0115] In addition, in addition to scoring the predicted responses output by the above-mentioned models for each sample on the test set and the expected responses in the corresponding samples by pandaLM, the embodiment of the present application also uses three independent reviewers (R1, R2 and R3) to score the quality of the responses generated by alpaca-coachLM and the original alpaca model in the coachLM150 test set from the multiple dimensions introduced in the aforementioned embodiments, and calculates the quality average score and the average of the quality average scores of the three independent reviewers as shown in Table 5. By comparison, it can be seen that the response output by alpaca-coachLM obtains a higher quality average score compared with the original alpaca model. The performance improvement of alpaca-coachLM further confirms the effectiveness of the modifications to the alpaca52k dataset by the first LLM. These modifications improve the data quality of the alpaca52k dataset and successfully enhance the instruction-following ability of the subsequently trained LLM.
[0116] Table 5 Comparison of quality scores of responses generated by alpaca-coachLM and the original alpaca model
[0117] Model R1 R2 R3 average value alpaca 56.6 58.2 60.9 58.6 alpaca-coachLM 61.4 66.9 64.7 64.3
[0118] 3. Evaluation of the impact of preset quality coefficients on model capabilities
[0119] As described in the aforementioned step 102, when training the first initial model, the data processing device can sort the training samples in descending order according to the difference values between the prompt data samples to be modified and the modified prompt data samples in each training sample in the first training data set, and obtain the top N training samples according to the preset quality coefficient to train the first initial model, thereby obtaining the first LLM, where N = quality coefficient * total number of training samples. It can be seen that the quality coefficient determines the proportion of training samples modified by professionals used in the process of training the first initial model. A higher quality coefficient means that a larger proportion of training samples modified by professionals will be used. When the quality coefficient is 1, it means that all training samples modified by professionals will be used to train the first initial model. When the quality coefficient is 0, it means that no training samples modified by professionals are used to train the first initial model, which means that the base model is directly used as the first LLM.
[0120] Based on this, in the embodiment of the present application, multiple training subsets containing different numbers of target training samples are obtained from the first training dataset by changing the quality coefficient, and the target training samples in each training subset are used to train a corresponding first LLM. In this way, multiple different first LLMs are obtained by training multiple training subsets. On this basis, the alpaca52k dataset is modified based on different first LLMs to obtain multiple different modified alpaca52k datasets, and different alpaca-coachLMs are trained using different modified alpaca52k datasets. Figure 4 The winning rate of alpaca-coachLM evaluated by pandaLM and GPT-4 models under different quality coefficients is shown. The dotted curve is the winning rate evaluated by pandaLM, and the solid curve is the winning rate evaluated by GPT-4 model. Figure 4 As can be seen, the scores of both the pandaLM and GPT-4 models show a similar trend, that is, alpaca-coachLM has the highest win rate when the quality coefficient is 0.3. The win rate of alpaca-coachLM increases as the quality coefficient increases from 0 to 0.3, which shows the importance of high-quality expert knowledge for the first LLM to achieve ideal modification capabilities. However, when the quality coefficient exceeds 0.3, due to the increase in the number of training samples with less modification included in the training samples obtained based on the quality coefficient, noise is introduced, which may interfere with the alignment of the first LLM with the expert knowledge, thereby reducing the quality of the dataset modified by the first LLM and further reducing the performance of the alpaca-coachLM trained on the modified dataset. Despite this, the performance reduction caused by this noise is at most about 10%, which shows the relative robustness of the first LLM.
[0121] 4. Robustness Evaluation of the First LLM
[0122] To further verify the robustness of the first LLM, the present embodiment uses the modified alpaca52k dataset of the first LLM to train three open-source base models: LLaMA, ChatGLM, and ChatGLM2, respectively. This yields three corresponding alpaca-coachLM models: alpaca-coachLM1, alpaca-coachLM2, and alpaca-coachLM3. The quality coefficient used in training the first LLM is 1.
[0123] The three trained alpaca-coachLM models and the original alpaca model were tested using the coachLM150 test set. Their WR1, WR2, and QS on the coachLM150 test set were evaluated using pandaLM. The evaluation results are shown in Table 6 below. As shown in Table 6, the performance parameters of the alpaca-coachLM trained on the alpaca52k dataset modified by the first LLM for the three open-source base models are all higher than those of the original alpaca model. This indicates that the alpaca52k dataset modified by the first LLM has a positive impact on the training of different base model types, demonstrating the robustness of the first LLM.
[0124] Table 6 Comparison of test results of three alpaca-coachLM and the original alpaca model
[0125] Base Model The trained model WR1 WR2 WR3 alpaca model 48.0 45.7 74.7 LLaMA alpaca-coachLM1 49.3 48.6 75.3 ChatGLM alpaca-coachLM2 54.0 59.1 82.0 ChatGLM2 alpaca-coachLM3 56.7 65.6 85.3
[0126] It's also worth noting that the base models ChatGLM and ChatGLM2 corresponding to alpaca-coachLM2 and alpaca-coachLM3 are both trained using RL, while the base model LLaMA corresponding to alpaca-coachLM1 is not. Based on this, as shown in Table 6, the performance parameters of alpaca-coachLM2 and alpaca-coachLM3 are both higher than those of alpaca-coachLM1. This demonstrates that a more powerful base model can enhance the ability of the model to align with expert knowledge during instruction fine-tuning.
[0127] Next, the data processing device provided in the embodiment of the present application is introduced.
[0128] Figure 5 This is a structural diagram of a data processing device provided in an embodiment of the present application. Figure 5 As shown, the data processing device 500 includes: an instruction acquisition module 501 and an inference module 502, wherein the instruction acquisition module 501 is used to execute step 201 in the above embodiment, and the inference module 502 is used to execute step 202 in the above embodiment.
[0129] Alternatively, see Figure 5The data processing device 500 may further include a training data acquisition module 503 and a training module 504. The training data acquisition module 503 is mainly used to acquire a first training data set and obtain multiple target training samples from the first training data set, where the data quality of the multiple target training samples is higher than the data quality of other training samples in the first training data set; the first training module 504 is used to train the first initial model using the multiple target training samples to obtain a first LLM.
[0130] Optionally, the training data acquisition module 503 is specifically used to: determine the difference value between the prompt data sample to be modified and the modified prompt data sample included in each training sample in the first training data set; and acquire multiple target training samples based on the difference value corresponding to each training sample.
[0131] Optionally, the training data acquisition module 503 is specifically used to: sort multiple training samples in the first training data set in descending order of the difference value corresponding to each training sample to obtain a sorting result; determine the first N training samples in the sorting result as multiple target training samples, where N is a preset value or N is determined based on a preset quality coefficient and the total number of training samples in the first training data set.
[0132] Optionally, the first training module 504 is specifically used to: process the modification instruction sample in each target training sample through the first initial model to obtain the predicted response corresponding to each modification instruction sample; and update the model parameters of the first initial model based on the predicted response and expected response corresponding to each modification instruction sample.
[0133] Optionally, the plurality of prompt data samples to be modified in the plurality of training samples include quality defects in multiple dimensions, including readability, feasibility and contextualization of instructions, and readability, correctness, relevance, richness, humanity, safety and comprehensiveness of responses.
[0134] Optionally, the apparatus further includes a second training module 505, configured to train a second initial model using second prompt data to obtain a second LLM.
[0135] In an embodiment of the present application, a data processing device can obtain a first modification instruction for instructing to modify the first prompt data. Afterwards, the first modification instruction is processed by the first LLM to obtain a first response, which includes the second prompt data obtained after modifying the first prompt data. Among them, the first LLM is obtained by training with multiple training samples, each training sample contains a modification instruction sample and a corresponding expected response, and the modification instruction sample is used to indicate that the prompt data sample to be modified is modified, and the expected response includes the expected modified prompt data sample. It can be seen that the first LLM obtained by training with multiple training samples will be able to learn the ability to modify the prompt data to be modified. On this basis, the prompt data in the prompt data set is modified by the first LLM to obtain the modified prompt data set. In this way, not only the data quality of the prompt data set can be improved, but also the data integrity of the prompt data set can be guaranteed.
[0136] It should be noted that the division of modules in the data processing device provided in the above embodiment is schematic and is merely a logical functional division. In actual implementation, other division methods may be used. In addition, the functional modules in the various embodiments of the present application may be integrated into a single processor, or may exist physically separately, or two or more modules may be integrated into a single module. The above-mentioned integrated modules may be implemented in the form of hardware or software functional modules.
[0137] If the integrated module is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be an electronic device or a server, etc.) or a processor to execute all or part of the steps of the method of each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0138] In addition, the data processing device and data processing method provided in the above embodiments belong to the same concept, and their specific implementation processes are detailed in the method embodiments, which will not be repeated here.
[0139] This application also provides a computer device 600. Figure 6As shown, computer device 600 includes: bus 602, processor 604, memory 606, and communication interface 608. Processor 604, memory 606, and communication interface 608 communicate with each other via bus 602. Computer device 600 can be a server or a terminal device. It should be understood that the embodiments of the present application do not limit the number of processors and memories in computer device 600.
[0140] The bus 602 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 The bus 602 may include a path for transmitting information between various components of the computer device 600 (eg, memory 606, processor 604, and communication interface 608).
[0141] The processor 604 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0142] The memory 606 may include volatile memory, such as random access memory (RAM). The processor 604 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0143] The memory 606 stores executable program codes, and the processor 604 executes the executable program codes to implement the functions of the aforementioned data processing device, thereby implementing the data processing method in the aforementioned embodiment. That is, the memory 606 stores instructions for executing the aforementioned data processing method.
[0144] Exemplarily, the processor 604 may execute the executable program code to respectively implement the functions of the aforementioned instruction acquisition module 501 and the reasoning module 502. Optionally, the processor 604 may also execute the executable program code to respectively implement the functions of the aforementioned training data acquisition module 503, the first training module 504, and the second training module 505 ( Figure 6 not shown).
[0145] The communication interface 608 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computer device 600 and other devices or a communication network.
[0146] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0147] like Figure 7 As shown, the computer device cluster includes at least one computer device 600. The memory 606 in one or more computer devices 600 in the computer device cluster may store the same instructions for executing the data processing method.
[0148] In some possible implementations, the memory 606 of one or more computer devices 600 in the computer device cluster may also store partial instructions for executing the data processing method in the aforementioned embodiment. In other words, the combination of one or more computer devices 600 can jointly execute instructions for executing the data processing method in the aforementioned embodiment.
[0149] It should be noted that the memory 606 in different computer devices 600 in the computer device cluster can store different instructions, each used to execute a portion of the functions of the data processing device. In other words, the instructions stored in the memory 606 in different computer devices 600 can implement the functions of one or more of the instruction acquisition module 501, the inference module 502, the training data acquisition module 503, the first training module 504, and the second training module 505.
[0150] In some possible implementations, one or more computer devices in the computer device cluster may be connected via a network, which may be a wide area network or a local area network. Figure 8 A possible implementation is shown. Figure 8As shown, two computer devices 600A and 600B are connected via a network. Specifically, the connection to the network is achieved through a communication interface in each computer device. In this possible implementation, as an example, the memory 606 in computer device 600A may store instructions for executing the functions of the instruction acquisition module 501, the reasoning module 502, and the second training module 505. Simultaneously, the memory 606 in computer device 600B may store instructions for executing the functions of the training data acquisition module 503 and the first training module 504. Based on this, computer device 600B can obtain a first LLM through the training data acquisition module 503 and the first training module 504, and then send this first LLM to computer device 600A via the network, thereby deploying the first LLM on computer device 600A. Computer device 600A can then use the first LLM to obtain modified prompt data by executing the instruction acquisition module 501 and the reasoning module 502, and then train a second LLM using the modified prompt data by executing the second training module. Of course, in another example, the memory 606 of computer device 600A may store instructions for executing the functions of the instruction acquisition module 501, the inference module 502, the training data acquisition module 503, and the first training module 504. Simultaneously, the memory 606 of computer device 600B may store instructions for executing the functions of the second training module 505. That is, in this example, the training of the first LLM and the modification of the prompt data using the first LLM may be performed by computer device 600A. Thereafter, computer device 600A may transmit the modified prompt data to computer device 600B via a network, and computer device 600B may use the modified prompt data to train the second LLM using the second training module 505.
[0151] It should be understood that Figure 8 The functions of the computer device 600A shown in FIG. 6 may also be completed by multiple computer devices 600. Similarly, the functions of the computer device 600B may also be completed by multiple computer devices 600.
[0152] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital versatile disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).
[0153] In the various embodiments of the present application, unless otherwise specified or logically conflicting, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other. The technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships. In the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. In the textual description of the embodiments of the present application, the character " / " generally indicates that the related objects before and after are in an "or" relationship. In the present application, "first", "second" and various numerical numbers are only used to distinguish for the convenience of description and are not used to limit the scope of the embodiments of the present application. For example, to distinguish different messages, etc., rather than to describe a specific order or sequence.
[0154] It is understood that the various numbers used in the embodiments of the present application are only used for ease of description and are not intended to limit the scope of the embodiments of the present application. The order of the sequence numbers of the above-mentioned processes does not necessarily indicate the order in which they are executed. The order in which the processes are executed should be determined by their functions and internal logic.
[0155] Finally, it should be noted that the above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A data processing method, characterized in that: The method comprises: Obtaining a first modification instruction, where the first modification instruction includes first prompt data to be modified in the prompt data set, and the first modification instruction is used to instruct modification of the first prompt data; The first modification instruction is processed by a first large language model LLM to obtain a first response, wherein the first response includes second prompt data obtained after modifying the first prompt data, and the first LLM is obtained by training multiple training samples in a first training data set, wherein each training sample includes a modification instruction sample and a corresponding expected response, the modification instruction sample includes a prompt data sample to be modified, and the modification instruction sample is used to indicate modification of the prompt data sample, and the expected response includes the expected modified prompt data sample.
2. The method according to claim 1, characterized in that The method further comprises: Obtaining the first training data set; Acquire multiple target training samples from the first training data set, where data quality of the multiple target training samples is higher than data quality of other training samples in the first training data set; The first initial model is trained using the multiple target training samples to obtain the first LLM.
3. The method according to claim 2, characterized in that The acquiring a plurality of target training samples from the first training data set includes: determining a difference value between the prompt data sample to be modified and the modified prompt data sample included in each training sample in the first training data set; Based on the difference value corresponding to each training sample, the multiple target training samples are obtained.
4. The method according to claim 3, characterized in that The acquiring the plurality of target training samples based on the difference value corresponding to each training sample includes: Sorting the plurality of training samples in the first training data set according to the order of the difference value corresponding to each training sample from large to small to obtain a sorting result; The first N training samples in the sorting result are determined as the multiple target training samples, where N is a preset value or is determined according to a preset quality coefficient and the total number of training samples in the first training data set.
5. The method according to any one of claims 2 to 4, characterized in that: The training of the first initial model using the multiple target training samples includes: Processing the modification instruction sample in each target training sample using the first initial model to obtain a predicted response corresponding to each modification instruction sample; The model parameters of the first initial model are updated based on the predicted response and the expected response corresponding to each modification instruction sample.
6. The method according to any one of claims 1 to 5, characterized in that: The multiple prompt data samples to be modified in the multiple training samples include quality defects in multiple dimensions, and the multiple dimensions include readability, feasibility and contextualization of instructions, and readability, correctness, relevance, richness, humanity, safety and comprehensiveness of responses.
7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: The second prompt data is used to train a second initial model to obtain a second LLM.
8. A data processing device, characterized in that: The data processing device includes: an instruction acquisition module, configured to acquire a first modification instruction, wherein the first modification instruction includes first prompt data to be modified in the prompt data set, and the first modification instruction is used to instruct modification of the first prompt data; An inference module is used to process the first modification instruction through a first large language model (LLM) to obtain a first response, wherein the first response includes second prompt data obtained by modifying the first prompt data, and the first LLM is obtained by training multiple training samples in a first training data set, wherein each training sample includes a modification instruction sample and a corresponding expected response, the modification instruction sample includes a prompt data sample to be modified, and the modification instruction sample is used to indicate modification of the prompt data sample, and the expected response includes the expected modified prompt data sample.
9. The device according to claim 8, characterized in that The device further comprises: a training data acquisition module, configured to acquire the first training data set; acquire a plurality of target training samples from the first training data set, wherein the data quality of the plurality of target training samples is higher than the data quality of other training samples in the first training data set; The first training module is used to train the first initial model using the multiple target training samples to obtain the first LLM.
10. The device according to claim 9, characterized in that The training data acquisition module is specifically used to: determining a difference value between the prompt data sample to be modified and the modified prompt data sample included in each training sample in the first training data set; Based on the difference value corresponding to each training sample, the multiple target training samples are obtained.
11. The device according to claim 10, characterized in that The training data acquisition module is specifically used to: Sorting the plurality of training samples in the first training data set according to the order of the difference value corresponding to each training sample from large to small to obtain a sorting result; The first N training samples in the sorting result are determined as the multiple target training samples, where N is a preset value or is determined according to a preset quality coefficient and the total number of training samples in the first training data set.
12. The device according to any one of claims 9 to 11, characterized in that The first training module is specifically used for: Processing the modification instruction sample in each target training sample using the first initial model to obtain a predicted response corresponding to each modification instruction sample; The model parameters of the first initial model are updated based on the predicted response and the expected response corresponding to each modification instruction sample.
13. The device according to any one of claims 8 to 12, characterized in that The multiple prompt data samples to be modified in the multiple training samples include quality defects in multiple dimensions, and the multiple dimensions include readability, feasibility and contextualization of instructions, and readability, correctness, relevance, richness, humanity, safety and comprehensiveness of responses.
14. The device according to any one of claims 8 to 13, characterized in that The device further comprises: The second training module is specifically configured to train the second initial model using the second prompt data to obtain a second LLM.
15. A computer device, characterized in that: The computer device includes a processor, which is used to execute at least one program instruction or code stored in a memory to implement the data processing method according to any one of claims 1 to 7.
16. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed on a computer device, the computer device executes the data processing method according to any one of claims 1 to 7.
17. A computer program product comprising instructions, characterized in that When the instructions are executed by a computer device, the computer device is caused to execute the data processing method according to any one of claims 1 to 7.