Data processing method and device, electronic equipment and computer readable storage medium
By obtaining user interaction data and feedback, the role-playing model is fine-tuned to solve the problems of differences in emotional expression, inadequate creativity and insufficient ethical security, and the model output is achieved more in line with human preferences.
Patent Information
- Application Number
- CN202411848264.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-16
AI Technical Summary
In role-playing scenarios, large models have problems such as differences in emotional expression, insufficient ability to generate creative content, and insufficient ethical security.
By obtaining the user's historical interaction data and feedback data, we determine the training data used to fine-tune the role-playing model and use this data to fine-tune the model to better match user preferences.
The emotional expression consistency, creative content generation ability and ethical security of role-playing models are improved, ensuring that the model output is more in line with human preferences and ethical norms.
Smart Images

Figure CN120011803A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, in particular to the technical fields of large models, intelligent dialogue, etc., and can be used in application scenarios such as generative search, intelligent document editing, intelligent assistants, virtual assistants, and intelligent e-commerce. Specifically, the present disclosure relates to a data processing method and device, an electronic device, and a computer-readable storage medium. Background Art
[0002] In recent years, LLM (Large Language Model), also known as "big model", has become a core technology in the field of natural language processing. Large models represented by ChatGPT have a huge number of parameters and can demonstrate excellent natural language understanding, intent recognition, scenario reasoning, multi-language translation and other powerful capabilities.
[0003] As large-scale model products with role-playing functions such as character.AI (chatbots based on artificial intelligence) have gained widespread attention, the role-playing capabilities of large models have become increasingly valued. Users are also increasingly demanding a customizable, highly anthropomorphic, emotional and warm chatbot. Summary of the invention
[0004] The present disclosure provides a data processing method and device, an electronic device, and a computer-readable storage medium.
[0005] According to a first aspect of the present disclosure, a data processing method is provided, the method comprising:
[0006] Acquire historical interaction data and feedback data of the user; the historical interaction data is the interaction data during the interaction between the user and the pre-trained role-playing model, and the feedback data is the feedback of the user on the output data of the role-playing model during the interaction;
[0007] Determining training data for fine-tuning the role-playing model based on the feedback data and the historical interaction data;
[0008] The role-playing model is fine-tuned using the training data.
[0009] According to a second aspect of the present disclosure, there is provided a data processing device, the device comprising:
[0010] A data collection module, used to obtain historical interaction data and feedback data of the user; the historical interaction data is the interaction data during the interaction between the user and the pre-trained role-playing model, and the feedback data is the feedback of the user on the output data of the role-playing model during the interaction;
[0011] A data processing module, configured to determine training data for fine-tuning the role-playing model based on the feedback data and the historical interaction data;
[0012] A model training module is used to fine-tune the role-playing model using the training data.
[0013] According to a third aspect of the present disclosure, an electronic device is provided, the electronic device comprising:
[0014] at least one processor; and
[0015] A memory in communication with the at least one processor; wherein,
[0016] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data processing method.
[0017] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the above-mentioned data processing method.
[0018] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the above data processing method when executed by a processor.
[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0021] Figure 1 is a flowchart of a data processing method provided by an embodiment of the present disclosure;
[0022] Figure 2 It is a process diagram of some steps of another data processing method provided by an embodiment of the present disclosure;
[0023] Figure 3 It is a flowchart of some steps of another data processing method provided by an embodiment of the present disclosure;
[0024] Figure 4 It is a flowchart of some steps of another data processing method provided by an embodiment of the present disclosure;
[0025] Figure 5It is a flowchart of some steps of another data processing method provided by an embodiment of the present disclosure;
[0026] Figure 6 It is a process diagram of some steps of another data processing method provided by an embodiment of the present disclosure;
[0027] Figure 7 is a structural schematic diagram of a data processing device provided by an embodiment of the present disclosure;
[0028] Figure 8 It is a block diagram of an electronic device used to implement the data processing method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0029] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0030] In some related technologies, affected by the training data, the current large models still have the following problems in role-playing scenarios:
[0031] There are differences in emotional expression. Although large models can simulate emotional expressions, the emotional responses they generate may be significantly different from human responses in the same situation. For example, the model may not accurately capture subtle emotional changes or complex emotional mixtures, causing its output to feel too mechanical or unnatural.
[0032] Poor ability to generate creative content. Human players are often able to create creative and highly personalized content in role-playing games. Large models, on the other hand, may rely on existing information in their training data when generating novel content, resulting in output that lacks originality or looks like a repetition of known information.
[0033] Insufficient ethical safety. Humans usually consider moral and social norms to guide their behavior. The training corpus of large models may contain a lot of bad information, which makes them unable to make detailed and reasonable judgments in complex moral and ethical situations like humans.
[0034] These problems are essentially that the inputs to large models are not aligned with human preferences.
[0035] In some related technologies, RLHF (Reinforcement Learning from Human Feedback) can be used to align the output of the big model with human preferences. The basic idea behind RLHF is to use a pre-trained big model and let people rank the results of its output, using the ranking as a signal to guide the big model to "prefer" certain results, thereby inducing a response and making the big model safer and more reliable.
[0036] Although RLHF provides an effective way to adjust the output tendency of large models, especially in complex interactive scenarios such as role-playing, it also has some shortcomings:
[0037] Sensitive to data quality. RLHF relies on feedback data provided by humans, which means that the quality and bias of the data have a huge impact on the large model. If the feedback data is not diverse enough or biased, then these biases may also be absorbed and amplified during the model training process.
[0038] Labeling data is expensive. Collecting high-quality human feedback takes time and money. This can become especially expensive and inefficient for large models that require large amounts of data to train and fine-tune. In addition, as the size of the model grows, the amount of data and computational resources required also increase, which can limit the scalability of RLHF methods.
[0039] The feedback dimension is single. In RLHF, the output results are sorted. However, different people have different preferences for results, and simple sorting may lead to conflicts in the data passed to the large model.
[0040] The data processing method and device, electronic device, and computer-readable storage medium provided by the embodiments of the present disclosure are intended to solve at least one of the above technical problems in the prior art.
[0041] The data processing method provided in the embodiments of the present disclosure may be executed by an electronic device such as a terminal device or a server, and the terminal device may be a vehicle-mounted device, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method may be implemented by a processor calling a computer-readable program instruction stored in a memory. Alternatively, the method may be executed by a server.
[0042] Figure 1 FIG. 1 is a flow chart showing a data processing method provided by an embodiment of the present disclosure. Figure 1As shown in , the data processing method provided by the embodiment of the present disclosure may include step S110, step S120, and step S130.
[0043] In step S110, the user's historical interaction data and feedback data are obtained;
[0044] In step S120, training data for fine-tuning the role-playing model is determined based on the feedback data and the historical interaction data;
[0045] In step S130 , the role-playing model is fine-tuned using the training data.
[0046] For example, the data processing method provided by the embodiments of the present disclosure can be used to process data generated during the interaction between a user and a role-playing model.
[0047] Among them, the role-playing model is a model used by a chat robot that can communicate with users by playing a certain role and provide users with role-playing services.
[0048] The pre-trained role-playing model may be a model obtained by pre-training the LLM using a large amount of data.
[0049] In some possible implementations, in step S110, the user and the role-playing model may conduct multiple rounds of dialogue interactions, and the user's historical interaction data and feedback data may be data obtained during the multiple rounds of dialogue interactions, and the historical interaction data and feedback data of multiple users may be obtained.
[0050] The user's historical interaction data may be interaction data during the user's interaction with the role-playing model, and specifically may be behavioral data related to the interaction during the interaction, such as the number of interaction rounds between the user and the role-playing model and the frequency of interaction between the user and the role-playing model.
[0051] The feedback data may be feedback from the user on output data output by the role-playing model according to the user's input data during the interaction with the role-playing model.
[0052] Feedback data can be collected by setting corresponding functions in the interactive interface where the user interacts with the role-playing model.
[0053] The rewrite function can be set, and the user can rewrite the output data through the rewrite function;
[0054] A regeneration function can be set. When the user uses the regeneration function, the role-playing model will regenerate the output data based on the user's input data;
[0055] An evaluation function can be set up, through which users can rate the output of the role-playing model according to their satisfaction with the output data;
[0056] That is to say, the feedback data may include rewriting data of the output data of the role-playing model during the interaction process by the user; it may also include rewriting data re-output by the role-playing model; it may also include evaluation data of the output data of the role-playing model during the interaction process by the user.
[0057] In some possible implementations, in step S120, whether the user is a malicious user can be determined based on historical interaction data, whether the user is satisfied with the output data output by the role-playing model during the interaction process can be determined based on feedback data, and the output data that satisfies non-malicious users can be selected as training data for fine-tuning the role-playing model.
[0058] In some possible implementations, in step S130, the acquired training data may be used as user preferences, and the role-playing model may be fine-tuned using preference alignment technology.
[0059] For example, RLHF (Reinforcement Learning from Human Feedback) can be used to build a reward model that reflects human preferences based on training data, and use reinforcement learning to fine-tune the role-playing model.
[0060] Other training methods can also be used to fine-tune the role-playing model based on the training data.
[0061] In the data processing method provided in the embodiment of the present disclosure, training data is determined based on the user's historical interaction data and feedback data, and data of some abnormal users can be removed from the training data to improve the quality of the training data, thereby improving the fine-tuning effect of the role-playing model.
[0062] The data processing method provided by the embodiment of the present disclosure is introduced in detail below.
[0063] Figure 2 A schematic diagram of the process of obtaining a user's historical interaction data and feedback data and processing the user's historical interaction data and feedback data is shown.
[0064] like Figure 2As shown, in some possible implementations, data generated during the interaction between multiple users (such as user1, user2, user3, etc.) and the role-playing model (such as the number of interaction rounds between the user and the role-playing model, the input data of the user into the role-playing model, the output data of the role-playing model, the feedback data as described above, etc.) can be collected and stored in a database.
[0065] Obtaining the user's historical interaction data and feedback data may involve extracting data from a database to obtain the user's historical interaction data and feedback data.
[0066] As described above, in some possible implementations, the feedback data may include rewriting data of the output data of the role-playing model during the interaction process by the user.
[0067] The rewritten data is data generated by rewriting the output data when the user is not satisfied with the output data output by the role-playing model. Therefore, the rewritten data is more in line with the user's preference than the output data.
[0068] Therefore, the rewritten data can be used as a positive sample in the fine-tuning process, that is, the positive example data, and the output data can be used as a negative sample in the fine-tuning process, that is, the counterexample data, to form a preferred data pair.
[0069] In some possible implementations, there may be malicious users who maliciously rewrite the data. If such poor quality rewritten data is used as positive data, it may affect the fine-tuning effect of the role-playing model. Figure 2 As shown, after data extraction, data filtering is required to filter the rewritten data to improve the data quality of the rewritten data.
[0070] Figure 3 A flowchart of an implementation method for filtering rewritten data and generating training data is shown, such as Figure 3 As shown, filtering the rewritten data and generating training data may include step S310, step S320, and step S330.
[0071] In step S310, the data quality of the rewritten data is determined based on the historical interaction data;
[0072] In step S320, when the data quality of the rewritten data meets the preset requirements, the similarity between the rewritten data and the output data is calculated;
[0073] In step S330, when the similarity is greater than a preset threshold, the output data is used as counter-example data and the rewritten data is used as positive example data.
[0074] In some possible implementations, in step S310, whether the user is a high-quality user can be judged based on the user's historical interaction data. For example, if the user interacts with the role-playing model for many rounds and with a high interaction frequency, it means that the user often uses the role-playing model and the user is a high-quality user. Obviously, the data quality of the rewritten data rewritten by the high-quality user on the output data is better than the data quality of the rewritten data rewritten by the non-high-quality user on the output data.
[0075] In some possible implementations, in step S320, the preset requirement may be a requirement satisfied by the user's historical interaction data, such as the number of interaction rounds between the user and the role-playing model meeting the requirement or the interaction frequency between the user and the role-playing model meeting the requirement.
[0076] Calculating the similarity between the rewritten data and the output data may be performed using a similarity calculation method. By calculating the similarity, a user behavior of directly copying and pasting the output data as the rewritten data may be identified.
[0077] In some possible implementations, the similarity between the rewritten data and the output data output by the role-playing model in other rounds of interaction may also be calculated to identify the user behavior of copying and pasting other output data as the rewritten data.
[0078] In some possible implementations, the similarity between the rewritten data and the user's input data may also be calculated to identify the user behavior of directly copying and pasting the user's input as the rewritten data.
[0079] The rewritten data generated by these abnormal user behaviors are obviously low-quality rewritten data.
[0080] In some possible implementations, in step S330, the rewritten data generated by the above-mentioned abnormal user behavior is screened out by comparing the similarity with a preset threshold, which can improve the quality of the rewritten data, and also improve the quality of the training data using the output data as counter-example data and the rewritten data as positive example data.
[0081] As described above, in some possible implementations, the feedback data may include rewritten data re-output by the role-playing model.
[0082] The rewritten data is data re-output by the role-playing model when the user is not satisfied with the output data output by the role-playing model. Therefore, the rewritten data is more in line with the user's preference than the output data.
[0083] Therefore, the rewritten data can be used as a positive sample in the fine-tuning process, that is, positive example data, and the output data can be used as a negative sample in the fine-tuning process, that is, counterexample data, to form a preferred data pair.
[0084] In some possible implementations, there may be malicious users who maliciously rewrite the role-playing model or some users who are curious about the rewriting function, which leads to continuous rewriting. If such poor quality rewriting data is used as positive data, it may affect the fine-tuning effect of the role-playing model. Therefore, Figure 2 As shown, after data extraction, data filtering is required, and the rewritten data needs to be filtered to improve the data quality of the rewritten data.
[0085] Figure 4 A flowchart of an implementation method for filtering rewritten data and generating training data is shown, such as Figure 4 As shown, filtering the rewritten data and generating training data may include step S410 and step S420.
[0086] In step S410, the data quality of the rewritten data is determined based on the historical interaction data;
[0087] In step S420, when the data quality of the rewritten data meets the preset requirements, the rewritten data is used as positive example data and the output data is used as negative example data.
[0088] In some possible implementations, in step S410, whether the user is a high-quality user can be judged based on the user's historical interaction data. For example, if the user interacts with the role-playing model for many rounds and with a high frequency, it means that the user often uses the role-playing model, and the user is a high-quality user. Obviously, the preferences of high-quality users for output data and rewriting data are more valuable for reference.
[0089] In some possible implementations, the user's historical interaction data may also include the number of times the user uses the rewrite function during the interaction with the role-playing model, and the user's corresponding regeneration rate is determined based on the ratio of the number of interaction rounds between the user and the role-playing model.
[0090] If the regeneration rate corresponding to a user is too high, it means that the user maliciously causes the role-playing model to be rewritten or the user is curious about the rewriting function, resulting in continuous rewriting. The rewriting data generated in this way cannot represent the user's preferences.
[0091] In some possible implementations, in step S420, the preset requirements may be requirements that the user's historical interaction data meets, such as the number of interaction rounds between the user and the role-playing model meets the requirements, the interaction frequency between the user and the role-playing model, and the regeneration rate of the user meets the requirements. By setting the requirements, the rewritten data that cannot represent the user's preferences can be screened out, and the quality of the rewritten data can be improved, which can also improve the quality of the training data that uses the output data as counter-example data and the rewritten data as positive example data.
[0092] As described above, in some possible implementations, the feedback data may include evaluation data of the output data of the role-playing model during the interaction process by the user. Specifically, the evaluation data may be the user's rating of the output data of the role-playing model during the interaction process.
[0093] Since the evaluation data is the user's evaluation of the output data, the output data that meets the user's preferences and the output data that the user evaluates low and does not meet the user's expectations can be determined through the evaluation data. Among them, the output data that meets the user's preferences can be used as positive samples in the fine-tuning process, that is, positive data, and the output data that the user evaluates low and does not meet the user's expectations can be used as negative samples in the fine-tuning process, that is, counter-example data, to form a preference data pair.
[0094] Through the above method, multiple preference data pairs can be obtained, and these multiple preference data pairs can constitute training data for fine-tuning the role-playing model. In other words, through the above method, not only the training data is obtained, but also the training data is cleaned, and the data with poor quality is eliminated, but the retained data is not completely without problems.
[0095] Some output data, rewritten data and adapted data, especially adapted data such as adapted data written by users themselves, may retain the user's own habits, resulting in over-frequent use or errors of some language symbols (such as emoticons, punctuation marks, etc.).
[0096] In order to avoid the influence of language symbols used too frequently and erroneous language symbols involved in fine-tuning on the performance of the role-playing model, the data needs to be further processed.
[0097] In some possible implementations, frequently used language symbols may be processed by downsampling.
[0098] Specifically, language symbols in the training data are obtained, and when the frequency of occurrence of the language symbols in the training data is higher than a preset frequency, the language symbols are downsampled.
[0099] Acquiring the language symbols in the training data may be to use a pattern recognition method to identify the language pattern of the training data, and to acquire the language pattern used in the training data and the usage frequency of the corresponding language pattern.
[0100] Language patterns are formed through language symbols, such as emoticons are composed of symbols, and bracket literature, which writes words expressing emotions in brackets, is achieved through the use of brackets.
[0101] In other words, language patterns correspond to language symbols, so the frequency of use of a language pattern is the frequency of occurrence of the language symbol corresponding to the language pattern in the training data.
[0102] If the frequency of occurrence of a language symbol in the training data is higher than the preset frequency, it means that the language symbol is used too frequently. The training data is processed using a downsampling method to reduce the frequency of occurrence of the language symbol in the training data.
[0103] Downsampling, also known as down sampling, is a commonly used technology in the field of signal processing. It is mainly used to solve the problem of uneven data distribution (too frequent use of a language symbol can be regarded as an uneven distribution of the language symbol with other language symbols in the training data). Specifically, the downsampling of language symbols can be achieved by deleting language symbols in the training data according to a certain rule.
[0104] It should be emphasized that the embodiments of the present disclosure do not limit the order of cleaning and processing training data. In some specific embodiments, after collecting output data, rewriting data and rewriting data, the method described above can be used to downsample frequently used language symbols, and then the output data, rewriting data and rewriting data can be cleaned to form preference data pairs.
[0105] In some specific embodiments, the training data may be cleaned and processed, and after removing invalid or noisy data, the training data may be manually reviewed to further ensure the quality of the training data.
[0106] Since the role-playing model is a model obtained by pre-training LLM with a large amount of data, fine-tuning the role-playing model also requires a large amount of data. In some possible implementations, based on some training data obtained based on the historical interaction data and feedback data of some users, the In-Context Learning capability of LLM can be used to enable the large model to judge the output data, rewritten data, and rewritten data generated during the interaction between the remaining users and the role-playing model, and generate training data for fine-tuning the role-playing model.
[0107] Figure 5 A flowchart of an implementation method of constructing training data based on LLM is shown, as Figure 5 As shown, constructing training data based on LLM may include step S510, step S520, step S530, and step S540.
[0108] In step S510, a preference score corresponding to the training data is determined based on the historical interaction data and the evaluation data;
[0109] In step S520, the training data, the preference scores corresponding to the training data, and the role features corresponding to the role-playing model are input into the large language model, so that the large language model learns the corresponding relationship between the training data and the preference scores;
[0110] In step S530, the pre-acquired reply data is scored using the large language model to obtain a score for the reply data;
[0111] In step S540, it is determined whether the reply data is positive example data or negative example data according to the score of the reply data.
[0112] In some possible implementations, in step S510, the acquired training data can be manually scored from multiple dimensions (such as user behavior characteristics, quality of data content, frequency of interaction, etc.) to generate a preference score corresponding to the training data.
[0113] It is also possible to automatically generate preference scores corresponding to the training data based on historical interaction data and evaluation data.
[0114] Among them, the historical interaction data may include the number of rounds of interaction between the user and the role-playing model, the frequency of interaction between the user and the role-playing model, the number of times the user uses the rewrite function during the interaction with the role-playing model, other user behaviors (such as the user's clicks, browsing history, etc.), and the interaction data of the user and the role-playing model (i.e., the input data entered by the user into the role-playing model).
[0115] Whether a user is a high-quality user can be determined based on the number of rounds of interaction between the user and the role-playing model, the frequency of interaction between the user and the role-playing model, and other behaviors of the user.
[0116] Specifically, if the number of rounds of interaction between the user and the role-playing model is large and the interaction frequency is high, it means that the user often uses the role-playing model, and the user is a high-quality user. The user's other behaviors can be used to judge whether the user is a real user, and even the user's willingness to interact with the role-playing model can be determined based on the user's clicks and browsing history. Real users with stronger willingness are high-quality users.
[0117] The evaluation of output data by high-quality users is more meaningful for reference. That is, the output data with high evaluation by high-quality users has a higher preference score, and the output data with low evaluation by high-quality users has a lower preference score.
[0118] In some possible implementations, the user's historical interaction data may also include the number of times the user uses the rewrite function during the interaction with the role-playing model, and the user's corresponding regeneration rate is determined based on the ratio of the number of interaction rounds between the user and the role-playing model.
[0119] If the user's corresponding regeneration rate is within a reasonable range, it means that the user is not curious about the rewrite function, which leads to continuous rewriting. The rewrite data generated in this way can represent the user's preference, and the corresponding preference score is higher.
[0120] Whether a user is a high-quality user can also be determined based on the interaction data of the user interacting with the role-playing model.
[0121] Specifically, the similarities of the plurality of interaction data are calculated, and the data quality of the interaction data is determined according to the similarities.
[0122] The higher the similarity of the interaction data, the higher the duplication of the input data input by the user into the role-playing model. The probability that the user may not be a real person or simply interacts with the role-playing model by copying and pasting is very high. The reference value for evaluating the output data is small, and the data quality of the interaction data is lower.
[0123] On the contrary, the lower the similarity of the interaction data, the lower the duplication of the input data of the user input into the role-playing model, and the user's evaluation of the output data is more meaningful. In other words, the higher the data quality of the interaction data, the higher the user is.
[0124] The output data with high evaluation has a higher preference score, and the output data with low evaluation has a lower preference score.
[0125] In some possible implementations, the data quality of the interaction data can also be weighted and summed with the number of interaction rounds between the user and the role-playing model, and the interaction frequency between the user and the role-playing model, so as to judge whether the user is a high-quality user from different dimensions. The output data with high evaluation by high-quality users has a higher preference score, and the output data with low evaluation by high-quality users has a lower preference score.
[0126] In some possible implementations, in step S520, the training data and the preference scores corresponding to the training data are input into the LLM, so that the LLM can learn the data features of the training data and then determine the corresponding relationship between the data features and the preference scores.
[0127] In some possible implementations, the role characteristics (such as domineering, confident, etc.) of the role played by the role-playing model can be input into the LLM as labels together with the training data and the preference scores corresponding to the training data, so as to provide more information for the LLM, so that the LLM can better learn the correspondence between data characteristics and preference scores, and then understand how to perform preference scoring based on the multi-dimensional characteristics of the data.
[0128] Among them, inputting the training data, the preference score corresponding to the training data, and the label into the LLM may be writing the training data, the preference score corresponding to the training data, and the label into a preset Prompt template to generate a Prompt input LLM.
[0129] In some possible implementations, in step S530, the pre-acquired reply data may be the output data, rewritten data, and rewritten data generated during the interaction between the remaining users and the role-playing model. These data are scored using the LLM that has understood how to perform preference scoring based on the multi-dimensional features of the data in step S520, thereby obtaining the preference scores corresponding to the output data, rewritten data, and rewritten data generated during the interaction between the remaining users and the role-playing model.
[0130] In some possible implementations, in step S540, a reply data with a high preference score can be determined as a positive sample for the fine-tuning process, i.e., positive example data, and a reply data with a low preference score can be used as a negative sample for the fine-tuning process, i.e., counterexample data, to form a preference data pair. The process is repeated to determine multiple preference data pairs.
[0131] By letting LLM automatically perform preference scoring, it realizes automatic labeling of data (i.e., determining positive and negative data). Compared with relying entirely on manual labeling, it greatly reduces labor costs and improves the efficiency of processing large-scale data.
[0132] In some possible implementations, the obtained multiple preference data pairs may be used as training data to fine-tune the role-playing model using RLHF.
[0133] RLHF is sensitive to data quality. If the training data is not diverse enough or has biases, these biases may be absorbed and amplified during the model training process. The training data obtained by the data processing method provided in the embodiment of the present disclosure includes both output data and rewritten data rewritten by the user on the output data and rewritten data regenerated by the role-playing model. In other words, the training data is diverse enough.
[0134] The RLHF in the related art determines the training data by sorting the output results. However, different people have different preferences for the results, and simple sorting may cause conflicts in the data passed to the large model. The data processing method provided in the embodiment of the present disclosure obtains training data not only based on the user's evaluation data on the output data, but also based on other dimensions such as the user's historical interaction data, thereby solving the problem of data conflicts.
[0135] At the same time, using LLM to assist in generating training data reduces the time and money required to manually collect high-quality training data.
[0136] In some possible implementations, Figure 6As shown, the role-playing model is fine-tuned using DPO (Direct Preference Optimization) to generate a fine-tuned role-playing model.
[0137] Specifically, multiple preference data pairs are formed based on the training data; the role-playing model is fine-tuned using the preference data pairs so that the similarity between the output of the role-playing model and the positive data of the preference data pairs is greater than the similarity between the output of the role-playing model and the negative data of the preference data pairs.
[0138] Each preference data pair includes a positive data and a negative data. The composition of the preference data pair is as described above and will not be repeated here.
[0139] The core idea of the DPO algorithm is to directly use human preference data to optimize the model, rather than indirectly through training the reward model. Therefore, the preference data in the preference data pair can be directly used as the optimization target, and the output generated by the role-playing model can be made more in line with human preferences by adjusting the model parameters. In other words, the output of the role-playing model can be made closer to the positive data of the preference data pair and further away from the negative data of the preference data pair through the pain model parameters.
[0140] The DPO algorithm omits the training process of the reward model, simplifies the training process, reduces the unstable factors that may occur during training, and also reduces the complexity and cost of training. The DPO algorithm is particularly suitable for processing multi-dimensional scoring data (i.e., the training data obtained by the data processing method provided in the embodiment of the present disclosure), which directly reflects human preferences for model output in many aspects. By fine-tuning the model to adapt to these multi-dimensional preference data, the content generated by the model can be more accurately aligned with the expectations of human users, thereby improving the practicality and user satisfaction of the role-playing model.
[0141] In some possible implementations, after obtaining the fine-tuned role-playing model, the fine-tuned role-playing model is tested, and the test results are evaluated based on a plurality of preset evaluation indicators; when the evaluation results do not meet the preset requirements, the step of obtaining the user's historical interaction data and feedback data is returned to re-fine-tune the role-playing model using the data processing method provided in the embodiment of the present disclosure, until a role-playing model whose splicing result meets the preset requirements is obtained.
[0142] Among them, the evaluation indicators may include repetition detection indicators, word count detection indicators, dialogue content diversity indicators, character stability indicators, and long-context memory ability indicators.
[0143] The repeatability indicator can be an indicator of the repeatability of the output of the role-playing model during the testing process.
[0144] The word count detection indicator may be an indicator of the number of words output by the role-playing model during the testing process.
[0145] The dialogue content diversity index may be an index of the diversity of the content output by the role-playing model during the test process.
[0146] The character stability index may be an index of the degree of fit between the output of the role-playing model and the personality of the character corresponding to the role-playing model during the test process.
[0147] Long-context memory capability may be an indicator of how well the role-playing model can remember historical outputs during testing.
[0148] Through these indicators, user preferences can be mapped to specific indicators, indicating the direction for specific optimization goals.
[0149] Based on these indicators, the fine-tuned role-playing model can be tested using manual testing methods or automatically tested using algorithms.
[0150] Based on Figure 1 The same principle as shown in the method, Figure 7 A schematic diagram of the structure of a data processing device provided by an embodiment of the present disclosure is shown. Figure 7 As shown, the data processing device 70 may include:
[0151] The data collection module 710 is used to obtain the user's historical interaction data and feedback data; the historical interaction data is the interaction data during the interaction between the user and the pre-trained role-playing model, and the feedback data is the user's feedback on the output data of the role-playing model during the interaction process;
[0152] A data processing module 720, for determining training data for fine-tuning the role-playing model based on the feedback data and the historical interaction data;
[0153] The model training module 730 is used to fine-tune the role-playing model using the training data.
[0154] In the data processing device provided in the embodiment of the present disclosure, training data is determined based on the user's historical interaction data and feedback data, and data of some abnormal users can be removed from the training data to improve the quality of the training data, thereby improving the fine-tuning effect of the role-playing model.
[0155] In some possible implementations, the training data includes positive example data and negative example data; and the data processing module is used to determine whether the output data is positive example data or negative example data based on the feedback data and the historical interaction data.
[0156] In some possible implementations, the feedback data includes rewritten data of the output data of the role-playing model during the interaction process by the user; the data processing module is used to: use the output data as counter-example data and use the rewritten data as positive example data.
[0157] In some possible implementations, the data processing module includes: a data quality unit, used to determine the data quality of the rewritten data based on historical interaction data; when the data quality of the rewritten data meets preset requirements, calculate the similarity between the rewritten data and the output data; a data processing unit, used to use the output data as counterexample data and the rewritten data as positive example data when the similarity is greater than a preset threshold.
[0158] In some possible implementations, the feedback data includes rewritten data re-output by the role-playing model when the user is not satisfied with the output data; the data processing module is used to: determine the data quality of the rewritten data based on the historical interaction data; when the data quality of the rewritten data meets the preset requirements, use the rewritten data as positive example data and the output data as negative example data.
[0159] In some possible implementations, the feedback data includes evaluation data of the user on the output data of the role-playing model during the interaction process; the data processing device also includes a large model training module, which is used to: determine the preference score corresponding to the training data based on the historical interaction data and the evaluation data; input the training data, the preference score corresponding to the training data, and the role characteristics corresponding to the role-playing model into the large language model, so that the large language model learns the correspondence between the training data and the preference score; use the large language model to score the pre-acquired reply data to obtain the score of the reply data; and determine whether the reply data is positive example data or negative example data according to the score of the reply data.
[0160] In some possible implementations, the historical interaction data includes the number of interaction rounds and the frequency of interaction between the user and the role-playing model, as well as the interaction data of the user and the role-playing model; the large model training module includes a preference scoring unit for: calculating the similarity of multiple interaction data, and determining the data quality of the interaction data based on the similarity; performing weighted summation of the number of interaction rounds, the frequency of interaction, and the data quality of the interaction data between the user and the role-playing model, to determine the preference score corresponding to the training data.
[0161] In some possible implementations, the large model training module further includes: a downsampling unit, which is used to obtain language symbols in the training data, and downsample the language symbols when the frequency of occurrence of the language symbols in the training data is higher than a preset frequency.
[0162] In some possible implementations, the model training module is used to: form a plurality of preference data pairs based on the training data, each preference data pair including a positive example data and a negative example data; and use the preference data pairs to fine-tune the role-playing model so that the output of the role-playing model has a greater similarity to the positive example data of the preference data pair than to the negative example data of the preference data pair.
[0163] In some possible implementations, the data processing device also includes: a model testing module, which is used to test the fine-tuned role-playing model and evaluate the test results based on multiple preset evaluation indicators; if the evaluation results do not meet the preset requirements, the role-playing model is re-fine-tuned.
[0164] It can be understood that the above modules of the data processing device in the embodiment of the present disclosure have the function of implementing Figure 1 The functions of the corresponding steps of the data processing method in the embodiment shown in . The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. The above modules can be software and / or hardware, and the above modules can be implemented separately or integrated with multiple modules. For the functional description of each module of the above data processing device, please refer to Figure 1 The corresponding description of the data processing method in the embodiment shown in will not be repeated here.
[0165] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.
[0166] In the technical solution of the present disclosure, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.
[0167] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0168] The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data processing method provided in the embodiment of the present disclosure.
[0169] Compared with the existing technology, the electronic device determines training data based on the user's historical interaction data and feedback data, and can remove data of some abnormal users from the training data, thereby improving the quality of the training data and further improving the fine-tuning effect of the role-playing model.
[0170] The readable storage medium is a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute a data processing method as provided in an embodiment of the present disclosure.
[0171] Compared with the prior art, the readable storage medium determines training data based on the user's historical interaction data and feedback data, and can remove data of some abnormal users from the training data, thereby improving the quality of the training data and further improving the fine-tuning effect of the role-playing model.
[0172] The computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the data processing method provided in the embodiment of the present disclosure.
[0173] Compared with the existing technology, this computer program product determines training data based on users' historical interaction data and feedback data, and can remove data of some abnormal users from the training data, thereby improving the quality of the training data and further improving the fine-tuning effect of the role-playing model.
[0174] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0175] like Figure 8 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0176] A number of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0177] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as data processing methods. For example, in some embodiments, the data processing method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the data processing method in any other appropriate manner (e.g., by means of firmware).
[0178] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0179] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0180] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0181] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0182] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0183] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0184] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0185] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A data processing method, comprising: Obtain users’ historical interaction data and feedback data; The historical interaction data is the interaction data during the interaction between the user and the pre-trained role-playing model, and the feedback data is the feedback of the user on the output data of the role-playing model during the interaction process; Determining training data for fine-tuning the role-playing model based on the feedback data and the historical interaction data; The role-playing model is fine-tuned using the training data.
2. The method according to claim 1, wherein: The training data includes positive example data and negative example data; the training data used for fine-tuning the role-playing model determined based on the feedback data and the historical interaction data includes: The output data is determined to be the positive example data or the negative example data based on the feedback data and the historical interaction data.
3. The method according to claim 2, wherein: The feedback data includes rewriting data of the output data of the role-playing model during the interaction process by the user; The determining, based on the feedback data and the historical interaction data, that the output data is the positive example data or the negative example data comprises: The output data is used as the counter-example data, and the rewritten data is used as the positive example data.
4. The method according to claim 3, wherein: The step of using the output data as the counter-example data and the rewritten data as the positive example data includes: determining data quality of the rewritten data based on the historical interaction data; When the data quality of the rewritten data meets the preset requirement, calculating the similarity between the rewritten data and the output data; When the similarity is greater than a preset threshold, the output data is used as the counter-example data, and the rewritten data is used as the positive example data.
5. The method according to claim 2, wherein: The feedback data includes rewritten data re-output by the role-playing model if the user is not satisfied with the output data; The determining, based on the feedback data and the historical interaction data, that the output data is the positive example data or the negative example data comprises: determining data quality of the rewritten data based on the historical interaction data; In the case that the data quality of the rewritten data meets the preset requirement, the rewritten data is used as the positive example data, and the output data is used as the negative example data.
6. The method according to claim 2, wherein: The feedback data includes evaluation data of the user on the output data of the role-playing model during the interaction process; the method further includes: Determine a preference score corresponding to the training data based on the historical interaction data and the evaluation data; Inputting the training data, the preference scores corresponding to the training data, and the role features corresponding to the role-playing model into a large language model, so that the large language model learns the corresponding relationship between the training data and the preference scores; Using the large language model to score the pre-acquired reply data, to obtain a score for the reply data; The reply data is determined to be the positive example data or the negative example data according to the score of the reply data.
7. The method according to claim 6, wherein: The historical interaction data includes the number of interaction rounds and the interaction frequency between the user and the role-playing model, and the interaction data of the user and the role-playing model; The determining the preference score corresponding to the training data based on the historical interaction data and the evaluation data includes: Calculating similarities of a plurality of the interaction data, and determining data quality of the interaction data according to the similarities; The number of interaction rounds and the interaction frequency between the user and the role-playing model and the data quality of the interaction data are weightedly summed to determine a preference score corresponding to the training data.
8. The method according to claim 6, wherein: Before inputting the training data, the preference scores corresponding to the training data, and the role features corresponding to the role-playing model into the large language model, the method further includes: A language symbol in the training data is obtained, and when the appearance frequency of the language symbol in the training data is higher than a preset frequency, the language symbol is downsampled.
9. The method according to claim 2, wherein: The fine-tuning of the role-playing model using the training data comprises: Based on the training data, a plurality of preference data pairs are formed, each preference data pair comprising one positive example data and one negative example data; The role-playing model is fine-tuned using the preference data pair so that the output of the role-playing model has a greater similarity to the positive data of the preference data pair than to the negative data of the preference data pair.
10. The method according to claim 1, further comprising: Test the fine-tuned role-playing model and evaluate the test results based on multiple preset evaluation indicators; If the evaluation result does not meet the preset requirements, return to the step of obtaining the user's historical interaction data and feedback data.
11. A data processing device, comprising: Data collection module, used to obtain users' historical interaction data and feedback data; The historical interaction data is the interaction data during the interaction between the user and the pre-trained role-playing model, and the feedback data is the feedback of the user on the output data of the role-playing model during the interaction process; A data processing module, configured to determine training data for fine-tuning the role-playing model based on the feedback data and the historical interaction data; A model training module is used to fine-tune the role-playing model using the training data.
12. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.
14. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 10.