Data enhancement method and related equipment
By performing multiple rounds of preference alignment fine-tuning of large language models, the problems of semantic incoherence, unreasonableness and low diversity of enhanced text generated by traditional data augmentation methods are solved, and high-quality and diverse data augmentation is achieved, improving the performance of downstream content recognition models.
Patent Information
- Application Number
- CN202510293064.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
The enhanced text generated by traditional data augmentation methods has problems of semantic incoherence, irrationality and low diversity, which affects the generalization ability of downstream models.
By performing multiple rounds of preference alignment fine-tuning of large language models, high-quality and diverse data-enhanced text is generated, suitable for downstream content recognition tasks.
It effectively improves the performance of downstream content recognition models and ensures that the generated enhanced text meets high standards in terms of quality and diversity.
Smart Images

Figure CN120216697A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a data enhancement method and related equipment. Background Art
[0002] In the current industry, there are a lot of demands for building content recognition capabilities in different business scenarios. When building relevant content recognition capabilities, data enhancement is usually used to improve model performance, but traditional data enhancement methods have problems such as stiff enhancement results, confusing semantics, or language style misalignment and cannot be used in actual business problems. For example, the simple data enhancement (EDA) method commonly used in the industry obtains enhanced text by performing synonym replacement, random insertion, random deletion or random exchange operations on the original text. Although the above EDA method is highly transferable and applicable to different scenarios, on the one hand, due to the existence of random operations, the generated enhanced text has semantic incoherence or unreasonable problems, and the text quality is relatively poor. On the other hand, the enhanced text is too similar to the original text in terms of literal similarity and low diversity, which will affect the generalization ability of the downstream model. Therefore, there is a need for a data enhancement scheme suitable for real business scenarios to ensure that the samples have high generation quality and diversity. Summary of the invention
[0003] In view of this, an embodiment of the present disclosure provides a data enhancement method, which can obtain a data enhancement model based on a large language model by performing multiple rounds of preference alignment fine-tuning on a large language model, thereby performing high-quality and diversified data enhancement on the data set of the downstream target content recognition task, thereby effectively improving the performance of the downstream content recognition model.
[0004] The data enhancement method described in the embodiment of the present disclosure may include: obtaining a text sample; generating data enhancement prompt words based on the text sample; inputting the data enhancement prompt words into a large language model that has undergone multiple rounds of preference alignment fine-tuning; and obtaining enhanced text corresponding to the text sample output by the large language model that has undergone multiple rounds of preference alignment fine-tuning.
[0005] In an embodiment of the present disclosure, the above method may further include: constructing a training data set; wherein the training data set includes a plurality of training texts; generating data augmentation prompt words for each training text in the training data set respectively; inputting the data augmentation prompt words into a large language model; obtaining a plurality of test augmented texts respectively corresponding to the plurality of training texts output by the large language model; performing preference annotation on the plurality of test augmented texts based on the plurality of training texts, and respectively determining preference labels reflecting whether the test augmented texts meet the expectations; generating a plurality of preference data pairs based on the plurality of training texts, the plurality of test augmented texts, and the preference labels of the plurality of test augmented texts; wherein the preference data pair includes: a training text, a test augmented text corresponding to the training text, and the preference label of the test augmented text; performing preference alignment fine-tuning on the large language model based on the preference data pairs; returning to the step of constructing the training data set until a preset fine-tuning termination condition is met; and using the large language model after preference alignment fine-tuning as the large language model after multiple rounds of preference alignment.
[0006] In an embodiment of the present disclosure, the above preference annotation on the plurality of test augmented texts includes: performing the following operations on each test augmented text in the plurality of test augmented texts respectively: performing key intent analysis on the training text based on the content label of the training text corresponding to the test augmented text, and determining the key points in the training text that belong to the content label of the training text; evaluating the retention degree of the key points by the test augmented text to obtain a consistency evaluation result; evaluating the diversity of the test augmented text in terms of expression and vocabulary usage compared with the training text to obtain a diversity evaluation result; and determining the preference label of the test augmented text based on the consistency evaluation result and the diversity evaluation result.
[0007] In an embodiment of the present disclosure, the preference annotation for the multiple test enhanced texts includes: performing the following operations on each of the multiple test enhanced texts respectively: generating a preference annotation prompt word based on the test enhanced text and the training text corresponding to the test enhanced text; wherein, the preference annotation prompt word includes: instructing the preference annotation model to perform key intent analysis on the training text based on the content label of the training text corresponding to the test enhanced text, and determining a first indication of the key points belonging to the content label of the training text in the training text; instructing the preference annotation model to evaluate the retention degree of the test enhanced text for the key points to obtain a second indication of the consistency evaluation result; instructing the preference annotation model to evaluate the diversity of the test enhanced text in terms of expression mode and vocabulary usage compared with the training text to obtain a third indication of the diversity evaluation result; and instructing the preference annotation model to determine a fourth indication of the preference label of the test enhanced text based on the consistency evaluation result and the diversity evaluation result; wherein, the preference annotation model is a pre-trained large language model; inputting the preference annotation prompt word into the preference annotation model; and obtaining the preference label of the test enhanced text output by the preference annotation model.
[0008] In an embodiment of the present disclosure, the above-mentioned preference alignment fine-tuning of the large language model based on the preference data includes: using the KTO algorithm to perform preference alignment fine-tuning on the large language model based on the preference data.
[0009] In an embodiment of the present disclosure, the above-mentioned preset fine-tuning termination conditions include: the availability rate of the multiple test enhanced texts reaches a preset first availability rate threshold; or, the number of times of preference alignment fine-tuning of the large language model reaches a preset number of iteration rounds.
[0010] In an embodiment of the present disclosure, when the availability rate of the test enhanced text output by the large language model is lower than a preset second availability rate threshold, the method further includes: rewriting the test enhanced text that does not meet the expectation; modifying the preference label of the test enhanced text to meet the expectation; and generating a preference data pair based on the training text corresponding to the test enhanced text, the rewritten test enhanced text, and the label of the rewritten test enhanced text.
[0011] In an embodiment of the present disclosure, before constructing the training data set, the following steps are further included: constructing a paired supervised data set; wherein, the paired supervised data set includes: a plurality of supervised data pairs; each supervised data pair includes: a first text and a second text; for each supervised data pair in the paired supervised data set, the following operations are respectively performed: generating a model fine-tuning prompt word based on the first text in the supervised data pair; inputting the model fine-tuning prompt word into the large language model; obtaining a third text output by the large language model; and performing supervised fine-tuning on the large language model based on the cross-entropy loss between the second text and the third text, to obtain a large language model that has undergone supervised fine-tuning.
[0012] In an embodiment of the present disclosure, constructing the paired supervised data set includes: extracting a first candidate text and a second candidate text belonging to the same label from a text data set with labels; determining the similarity between the first candidate text and the second candidate text; determining the proportion of co-occurring words between the first candidate text and the second candidate text; and in response to determining that the similarity is greater than a pre-set similarity threshold and the proportion of co-occurring words is greater than a pre-set co-occurring word proportion threshold, combining the first candidate text and the second candidate text in different permutations into two supervised data pairs.
[0013] In an embodiment of the present disclosure, performing supervised fine-tuning on the large language model based on the cross-entropy loss between the second text and the third text includes: performing supervised fine-tuning on the large language model based on the cross-entropy loss between the second text and the third text by means of low-rank adjustment.
[0014] Corresponding to the above data augmentation method, an embodiment of the present disclosure also discloses a data augmentation device, including: a sample acquisition module, configured to acquire text samples; a data augmentation prompt word generation module, configured to generate data augmentation prompt words based on the text samples; a data augmentation module, configured to input the data augmentation prompt words into a large language model that has undergone multiple rounds of preference alignment fine-tuning; and obtain enhanced texts corresponding to the text samples output by the large language model that has undergone multiple rounds of preference alignment fine-tuning.
[0015] In addition, an embodiment of the present disclosure further provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the above data augmentation method is implemented.
[0016] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, where the non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the above data augmentation method.
[0017] Embodiments of the present disclosure also provide a computer program product, including computer program instructions, which when running on a computer, cause the computer to execute the above data augmentation method.
[0018] It can be seen from this that on the one hand, the above data augmentation method and related devices can utilize the semantic understanding and generation capabilities mastered by the large language model during pre-training to alleviate the problem of incoherence or unreasonableness in the generated augmented text; on the other hand, through the preference data constructed based on the training text carrying content tags and the test augmented text generated by the large language model, further multi-round preference alignment fine-tuning of the large language model can be performed to obtain a data augmentation model, and then high-quality and diverse data augmentation can be performed on the dataset of the downstream content recognition task, thereby effectively improving the performance of the downstream content recognition model.
[0019] In addition, the above data augmentation method and related devices construct a paired supervised dataset through a heuristic strategy, and first perform supervised fine-tuning on the large language model. By implicitly adjusting the large language model through the training loss function, the language style of the generated augmented text is made consistent with the language style of the original text, so that the language style of the generated augmented text is consistent with the language style of the text used in the real business scenario, further improving the performance of the downstream content recognition model. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are only embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 Shows the implementation process of the method for multi-round preference alignment fine-tuning of the large language model according to some embodiments of the present disclosure.
[0022] Figure 2 Shows the implementation process of the method for preference annotation of the test augmented text according to some embodiments of the present disclosure.
[0023] Figure 3 Shows the implementation process of the method for text correction according to some embodiments of the present disclosure.
[0024] Figure 4 Shows the implementation process of the method for supervised fine-tuning of the large language model according to some embodiments of the present disclosure.
[0025] Figure 5 Shows the implementation process of the data augmentation method according to some embodiments of the present disclosure.
[0026] Figure 6 shows the internal structure of the data enhancement device according to some embodiments of the present disclosure.
[0027] Figure 7 shows a more specific schematic diagram of the hardware structure of an electronic device according to some embodiments of the present disclosure. Detailed implementation manners
[0028] To make the objectives, technical solutions, and advantages of the present disclosure more clear and understandable, the present disclosure will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0029] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure belongs. The "first", "second", and similar terms used in the embodiments of the present disclosure do not indicate any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or items appearing before this term cover the elements or items listed after this term and their equivalents, without excluding other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0030] It can be understood that, before using the technical solutions of the various embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.
[0031] For example, when responding to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.
[0032] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0033] It can be understood that the above-mentioned notification and the process of obtaining user authorization are only illustrative and do not limit the implementation manner of the present disclosure. Other methods that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0034] For the sake of clarity in description, before describing the specific technical solutions of the embodiments of the present disclosure, several technical terms related to the embodiments of the present disclosure will be described first.
[0035] A large language model (LLM, Large Language Model) is a neural network model used for natural language processing tasks and having a large number of parameters and a complex structure. The LLM is trained by using a vast amount of data and uses deep learning techniques to learn the basic patterns and structures of language. Compared with small models, the LLM can achieve higher performance and usually has better generalization ability in various tasks.
[0036] Data Augmentation refers to a technique of generating more training samples by performing various processes and transformations on existing data, thereby improving the performance and generalization ability of the model. In natural language processing tasks, data augmentation can be equivalent to text augmentation.
[0037] Supervised Fine Tuning (SFT) refers to the process of fine-tuning a pre-trained neural network model using new labeled data to improve its performance.
[0038] Low-Rank Adaptation (LoRA) is an efficient model fine-tuning training method proposed by Microsoft scholars. LoRA significantly reduces the training cost required for fine-tuning large models and does not introduce additional latency during the inference phase.
[0039] The Perfermance Alignment technology aims to enable large models to continuously optimize their behavior by receiving human feedback, so as to better meet human expectations. It can be divided into online preference alignment and offline preference alignment technologies according to whether the human preference feedback is online.
[0040] A Prompt is an injection instruction used to "command" a large model to think about problems and output content according to a preset idea. A Prompt is an instruction or information that guides or triggers a large model to make a response.
[0041] The embodiments of the present disclosure will be described in detail below with specific examples.
[0042] As mentioned above, traditional data augmentation methods have problems such as stiff augmentation results, confusing semantics, or language style misalignment and cannot be used in actual business problems. Therefore, there is a need for a data augmentation solution suitable for real business scenarios to ensure that samples have high generation quality and diversity.
[0043] The current large language model is pre-trained on massive data and has strong semantic understanding and text generation capabilities, as well as high diversity brought by large amounts of data. Therefore, it is possible to consider introducing a large language model as the foundation of the data enhancement model. However, when directly using a large model for data enhancement, the samples generated have poor quality in terms of language style and deviate from the text used in real business scenarios. Direct use has low gains for downstream models.
[0044] To this end, an embodiment of the present disclosure provides a data enhancement method, which can obtain a data enhancement model based on the large language model by performing multiple rounds of preference alignment fine-tuning on the large language model, thereby performing high-quality and diversified data enhancement on the data set of the downstream target content recognition task, thereby effectively improving the performance of the downstream content recognition model.
[0045] Figure 1 The implementation process of the method for performing multiple rounds of preference alignment fine-tuning on a large language model described in some embodiments of the present disclosure is shown. Figure 1 As shown, the specific method for performing multiple rounds of preference alignment fine-tuning on a large language model according to an embodiment of the present disclosure includes the following steps:
[0046] In step 110, a training data set is constructed.
[0047] In an embodiment of the present disclosure, the above training dataset may include multiple training texts. Further, the above training texts may also include content labels corresponding to the training texts. The above content labels are associated with downstream content recognition tasks and are mainly used to label the categories to which the training texts belong in downstream content recognition tasks. For example, if the downstream content recognition task is an evaluation recognition task, the above content labels may generally be labels indicating whether the training text is a positive evaluation or a negative evaluation; and if the downstream content recognition task is a classification task, the above content labels may generally refer to the categories of the training texts in a specific classification dimension. For example, for a training text such as "The overall taste of this northeastern cuisine is quite ordinary. They should have spent a lot on marketing fees", if the downstream is an evaluation recognition task, its content label may be "negative evaluation", and if the downstream is a catering classification task, its content label may be "northeastern cuisine". Therefore, it can be seen that the same training text may correspond to different content labels in different content recognition tasks. It should be noted that in an embodiment of the present disclosure, the content labels of the above training texts will mainly be used to perform preference annotation on the test enhancement texts generated by the large language model.
[0048] In step 120, data augmentation prompt words are generated for each training text in the above training dataset respectively.
[0049] In an embodiment of the present disclosure, the above data augmentation prompt words may be generated based on the training texts and a preset data augmentation prompt word template. The above data augmentation prompt words are mainly used to instruct the large language model to generate test enhancement texts based on the input training texts, and it is required that the generated test enhancement texts should be basically consistent with the input training texts in terms of text content and expression style, but there should be diversity in their expression methods and vocabulary usage.
[0050] In step 130, the above data augmentation prompt words are input into the pre-trained large language model.
[0051] In step 140, multiple test enhancement texts corresponding to the multiple training texts output by the large language model are obtained.
[0052] In step 150, preference annotation is performed on the above multiple test enhancement texts based on the above multiple training texts, and preference labels reflecting whether each test enhancement text meets the expectations are determined respectively.
[0053] In an embodiment of the present disclosure, the preference label may be a binary label. That is, the preference label may only include two values: "compliant" or "non - compliant". For example, in practical applications, "compliant" or "non - compliant" may be represented by 0 or 1 or other numerical values. Among them, "compliant" means that the generated test - enhanced text meets the expectation, that is, the generated test - enhanced text is basically consistent with the input training text in terms of text content and expression style, and there is diversity in the expression mode and vocabulary usage between the two; while "non - compliant" means that the generated test - enhanced text does not meet the expectation, that is, the generated test - enhanced text is inconsistent with the input training text in terms of text content and expression style, or there is basically no diversity in the expression mode and vocabulary usage between the two.
[0054] In step 160, a plurality of preference data pairs are generated based on the plurality of training texts, the plurality of test - enhanced texts, and the preference labels of the plurality of test - enhanced texts.
[0055] In an embodiment of the present disclosure, the preference data pair may include: a training text, the test - enhanced text corresponding to the training text, and the preference label of the test - enhanced text. For example, in an embodiment of the present disclosure, the preference data pair may be in the form of (x, y, label). Wherein, x represents the input training text; y represents the test - enhanced text corresponding to the training text; and label represents the preference label of the test - enhanced text. In an embodiment of the present disclosure, the preference data pair (x, y, label) will be used for subsequent preference alignment fine - tuning of the large - language model.
[0056] In step 170, the large - language model is fine - tuned for preference alignment based on the preference data pairs.
[0057] In some embodiments of the present disclosure, in step 170, the Kahneman - Taversky optimization (KTO) algorithm may be used to fine - tune the large - language model for preference alignment based on the preference data pairs. As an alternative to the above KTO, in some other embodiments of the present disclosure, other offline preference alignment techniques may also be used to fine - tune the large - language model for preference alignment.
[0058] In the above method, after the preference alignment fine - tuning of a certain round is completed, the step of constructing the training data set described in step 110 above may be returned, and the training data set is updated until a pre - set fine - tuning termination condition is met.
[0059] In an embodiment of the present disclosure, the above-mentioned fine-tuning termination condition may include one or a combination of the following multiple conditions: the availability rate of the above-mentioned multiple test enhanced texts reaches a preset first availability rate threshold; and the number of times the large language model performs preference alignment fine-tuning reaches a preset number of iteration rounds. For example, when the availability rate of the above-mentioned multiple test enhanced texts reaches the preset first availability rate threshold or the number of times the large language model performs preference alignment fine-tuning reaches the preset number of iteration rounds, the preset fine-tuning termination condition is satisfied.
[0060] Specifically, the availability rate of the above-mentioned multiple test enhanced texts may be the proportion of the test enhanced texts corresponding to the preference label "conform" among the above-mentioned multiple test enhanced texts in all test enhanced texts. It can be understood that when the proportion of the test enhanced texts corresponding to the preference label "conform" among the above-mentioned multiple test enhanced texts in all test enhanced texts reaches the preset first availability rate threshold (such as 85% or 90%), it can be considered that the test enhanced texts that meet the expectations in the test enhanced texts output by the large language model after this round of preference alignment fine-tuning have reached a certain proportion, that is, the output test enhanced texts basically meet the expectations. Therefore, the preference alignment fine-tuning of the large language model can be ended. In addition, in order to prevent infinite rounds of preference alignment fine-tuning of the large language model or prevent the number of rounds of preference alignment fine-tuning from being too many, resulting in too long fine-tuning time or too high fine-tuning cost, the number of iteration rounds can also be preset, that is, the maximum number of rounds of preference alignment fine-tuning. During the above-mentioned multiple rounds of preference alignment, when the number of times the large language model performs preference alignment fine-tuning has reached the preset number of iteration rounds, the next round of preference alignment fine-tuning can also be terminated.
[0061] In an embodiment of the present disclosure, since the KTO algorithm is an offline preference alignment method, therefore, the embodiment of the present disclosure introduces multiple rounds of preference alignment fine-tuning, and improves the data augmentation quality of the large language model step by step by manually annotating a small number of samples in each round. In addition, it can be understood that after the termination of multiple rounds of preference alignment fine-tuning, the above-mentioned large language model that has undergone multiple rounds of preference alignment can be used as a data augmentation model and applied to subsequent content recognition tasks to perform high-quality and diversified data augmentation on the dataset of the content recognition task, thereby effectively improving the performance of the task recognition model.
[0062] The following will specifically illustrate the specific method for preference annotation of test enhanced texts in the embodiments of the present disclosure with specific examples.
[0063] In some embodiments of the present disclosure, in the above step 150, the specific method for preference annotation of each test enhanced text in the preference annotation of the above-mentioned multiple test enhanced texts can refer to Figure 2 . As Figure 2As shown, the specific method for preference annotation of a certain test-enhanced text may include the following multiple steps:
[0064] In step 210, based on the content labels of the training text corresponding to the test-enhanced text, perform key intent analysis on the training text to determine the key points in the training text that belong to the above content labels.
[0065] As mentioned above, in the embodiments of the present disclosure, the above content labels are mainly used to label the categories to which the training text belongs in the downstream content recognition task. Therefore, in the embodiments of the present disclosure, by performing key intent analysis on the training text corresponding to the test-enhanced text, the key points related to the content labels in the training text can be determined. It can be understood that the key points determined in this way are associated with the downstream content recognition task to improve the execution effect of the content recognition task. For another example, for the training text "The overall taste of this Northeast cuisine is quite average, and the marketing expenses should not be small", if the downstream is an evaluation recognition task, its content label can be "negative evaluation", then its key point can be "the taste is quite average"; and if the downstream is a catering classification task, its content label can be "Northeast cuisine", then its key point can be "Northeast cuisine".
[0066] In step 220, evaluate the retention degree of the above test-enhanced text for the above key points to obtain a consistency evaluation result.
[0067] In the embodiments of the present disclosure, based on the above key points, evaluate the above test-enhanced text to obtain a consistency evaluation result that reflects whether the test-enhanced text is consistent with the training text in terms of text content and expression style. For example, for the aforementioned training text and the test-enhanced text "The taste and texture of this pickled fish restaurant are quite average, maybe the marketing expenses are too much", if the key point "the taste is quite average" determined in the above step 210, it can be evaluated that the two have a high consistency, that is, they are consistent; while if the key point "Northeast cuisine" determined in the above step 210, it can be evaluated that the two have a low consistency, that is, they are inconsistent. It can be seen from this that due to the different content labels corresponding to the training text, the determined key points are different. Therefore, for the same test-enhanced text, different consistency evaluation results may also be obtained.
[0068] In step 230, evaluate the diversity of the above test-enhanced text compared with the above training text in terms of expression and vocabulary usage to obtain a diversity evaluation result.
[0069] In step 240, based on the above consistency evaluation result and the above diversity evaluation result, determine the preference label of the above test-enhanced text.
[0070] In some embodiments of the present disclosure, when the consistency evaluation result is consistent and the diversity evaluation result is diverse, the preference label of the test augmented text can be determined as "conforming"; while when the consistency evaluation result is inconsistent or the diversity evaluation result is non-diverse, the preference label of the test augmented text can be determined as "non-conforming".
[0071] It can be seen from this that in the embodiments of the present disclosure, in the process of annotating the preference label of the test augmented text, the evaluation is mainly carried out from two aspects: the consistency of the content and the diversity of the expression. Therefore, the evaluation result can meet the requirements for both the content and the expression mode in data augmentation.
[0072] In addition, in the process of evaluating the consistency of the content, the influence of the content label of the training text on the content of the training text is mainly considered, so that the annotation of the preference label is more accurate in the downstream content recognition task. As shown in the aforementioned example, for the training text "The overall taste of this northeastern cuisine is quite ordinary, and the marketing expenses should not be small" and the test augmented text "The taste and texture of this pickled fish restaurant are both quite ordinary, probably because of too much marketing expenses", for different content recognition tasks, the consistency evaluation results of the above training text and test augmented text may be very different, and thus the determined preference labels are also very different. It can be seen that for the same training text and test augmented text, there are diametrically opposite judgment results under the evaluation criteria of different content recognition tasks. That is to say, the preference label of the test augmented text should be strongly related to the downstream content recognition task. Based on this, in the embodiments of the present disclosure, in order to ensure the accuracy of the preference annotation of the test augmented text, thereby generating high-quality and diverse augmented texts, and further ensuring the performance of the downstream content recognition model, in the above multi-round preference alignment process, the consistency between the generated test augmented text and the training text is determined by extracting key points from the training text based on the content label and mainly evaluating the retention degree of the above key points by the test augmented text in the process of consistency evaluation, so that the generation constraint conditions can be implicitly added in the preference alignment fine-tuning process, and the content label of the generated test augmented text defined in the downstream content recognition task is kept consistent with the content label of the training text, so as to perform high-quality and diverse data augmentation on the dataset of the downstream target content recognition task, and further improve the performance of the downstream content recognition model.
[0073] In the embodiments of the present disclosure, the above preference annotation can be implemented in various ways. For example, one or more proprietary machine learning models can be used to implement the above Figure 2One or more steps shown. In some other embodiments of the present disclosure, the semantic understanding ability mastered by the large language model during pre-training can also be fully utilized to complete the preference annotation of the test enhanced text. In the above examples, in order to distinguish from the foregoing large language model, the large language model used for preference annotation of the test enhanced text can be referred to as a preference annotation model. It can be understood that the above preference annotation model is the large language model that has been pre-trained.
[0074] Specifically, in the above embodiment of using the preference annotation model to annotate the preference of the test enhanced text, the following operations can be respectively performed on each of the multiple test enhanced texts: First, generate a preference annotation prompt word based on the test enhanced text and the training text corresponding to the test enhanced text. Specifically, the above preference annotation prompt word can include: instructing the preference annotation model to perform key intention analysis on the training text based on the content label of the training text corresponding to the test enhanced text, and determining the first indication of the key points belonging to the content label of the training text in the training text; instructing the preference annotation model to evaluate the retention degree of the test enhanced text for the key points to obtain the second indication of the consistency evaluation result; instructing the preference annotation model to evaluate the diversity of the test enhanced text in terms of expression mode and vocabulary usage compared with the training text to obtain the third indication of the diversity evaluation result; and instructing the preference annotation model to determine the fourth indication of the preference label of the test enhanced text based on the consistency evaluation result and the diversity evaluation result. Then, input the above preference annotation prompt word into the preference annotation model. Finally, obtain the preference label of the test enhanced text output by the preference annotation model.
[0075] It can be seen that in the above embodiment, the preference annotation of the test enhanced text can be completed by using the semantic understanding ability mastered by the large language model during pre-training, and since the large language model is instructed in the preference annotation prompt word to perform consistency evaluation based on the retention degree of the test enhanced text for the key points reflecting the content label of the training text, generation constraint conditions are implicitly added during the preference alignment fine-tuning process, so that the content label of the generated test enhanced text in the downstream content recognition task is kept consistent with the content label of the training text, thereby enabling high-quality and diverse data augmentation for the dataset of the downstream target content recognition task, and further improving the performance of the downstream content recognition model.
[0076] To further solve the problem that the preference alignment fine-tuning effect is not ideal or the fine-tuning cost is too high due to too few positive example samples during the multi-round preference alignment fine-tuning cold start of the above large language model, when the availability rate of the test enhanced text output by the above large language model is lower than a pre-set second availability rate threshold, the following can be further performed Figure 3The text correction method shown, thus effectively increasing the positive example samples in the preference data pairs and improving the efficiency and effect of multi-round preference alignment fine-tuning. Specifically, the above second availability rate can be preset according to the actual situation, for example, it can be set to 20% or other values, etc. As Figure 3 shown, the above text correction method may include the following multiple steps:
[0077] In step 310, rewrite the test augmented text that does not meet the expectations.
[0078] In the embodiments of the present disclosure, the above test augmented text that does not meet the expectations is the test augmented text with a preference label of "does not meet".
[0079] In addition, in the embodiments of the present disclosure, the above rewrite can be implemented by a pre-trained large language model. Specifically, the test augmented text and the corresponding training text can be input into the large language model. In order to guide the large language model to perform text rewriting, the above consistency evaluation result, diversity evaluation result, and the reasons related to the above evaluation results can also be input into the large language model, and the large language model corrects based on the evaluation attribution on the basis of the test augmented text, so that the rewritten test augmented text can meet the expectations.
[0080] In step 320, modify the preference label of the above test augmented text to meet the expectations.
[0081] In step 330, generate a preference data pair based on the training text corresponding to the above test augmented text, the rewritten test augmented text, and the label of the rewritten test augmented text.
[0082] It can be understood that the test augmented text corrected by the above method usually meets the expectations. Therefore, the preference data pair generated based on the training text corresponding to the above test augmented text, the rewritten test augmented text, and the label of the rewritten test augmented text will be a positive example sample, which can shorten the process of the large language model for preference alignment cold start when the positive example samples are too few, thereby improving the efficiency and effect of multi-round preference alignment fine-tuning.
[0083] In some other embodiments of the present disclosure, before performing the above multi-round preference alignment fine-tuning on the large language model, the large language model can be further fine-tuned in a supervised manner to solve the problem of language style misalignment existing in directly using the large language model for data augmentation through prompt engineering. In the embodiments of the present disclosure, the process of the above supervised fine-tuning can refer to Figure 4 . As Figure 4 shown, the method of supervised fine-tuning described in the embodiments of the present disclosure may include the following steps:
[0084] In step 410, a paired supervised dataset is constructed.
[0085] In an embodiment of the present disclosure, the above paired supervised dataset may include: a plurality of supervised data pairs. Specifically, each supervised data pair may include: a first text and a second text.
[0086] In step 420, the following operations are respectively performed for each supervised data pair in the above paired supervised dataset: generating a model fine-tuning prompt based on the first text in the supervised data pair.
[0087] The above model fine-tuning prompt is similar to the above data augmentation prompt and can be generated based on the first text and a preset model fine-tuning prompt template. The above model fine-tuning prompt is mainly used to instruct the large language model to generate a third text based on the input first text, and it is required that the generated third text should be basically consistent with the input first text in terms of expression style.
[0088] In step 430, the model fine-tuning prompt is input into the pre-trained large language model.
[0089] In step 440, the third text output by the large language model is obtained.
[0090] In step 450, the above large language model is supervised and fine-tuned based on the cross-entropy loss between the second text and the third text to obtain a supervised fine-tuned large language model.
[0091] Specifically, in an embodiment of the present disclosure, the construction of the paired supervised dataset described in the above step 410 may include the following multiple steps: First, extract a first candidate text and a second candidate text belonging to the same label from a text dataset with labels; wherein, the above extraction may be random extraction; determine the similarity between the first candidate text and the second candidate text; determine the proportion of co-occurring words between the first candidate text and the second candidate text; and in response to determining that the above similarity is greater than a preset similarity threshold and the above proportion of co-occurring words is greater than a preset co-occurring word proportion threshold, combine the first candidate text and the second candidate text into two supervised data pairs in different permutation orders. For example, assuming that the above first candidate text is S1 and the second candidate text is S2, the supervised data pairs obtained by the above method may be two: (S1, S2) and (S2, S1).
[0092] In some specific examples, the above similarity can be cosine similarity or other methods for determining text similarity, and the present disclosure does not limit this. It can be understood that the above similarity mainly reflects the degree of similarity in content between the first candidate text and the second candidate text, while the co-occurrence word ratio mainly reflects the degree of similarity in word usage between the co-occurrence words in the first candidate text and the second candidate text. Therefore, in the embodiments of the present disclosure, by simultaneously considering the degree of similarity in content and word usage, the constructed supervised data pair can have higher quality.
[0093] In some embodiments of the present disclosure, when performing supervised fine-tuning, the large language model can be supervised and fine-tuned based on the cross-entropy loss between the above second text and the third text by means of low-rank adjustment (LoRA).
[0094] Based on the above large language model that has undergone multiple rounds of preference alignment fine-tuning or the large language model that has first undergone supervised fine-tuning and then multiple rounds of preference alignment fine-tuning, data augmentation can be performed as a data augmentation model. Figure 5 Shows the implementation process of the data augmentation method described in the embodiments of the present disclosure. As Figure 5 shown, the above data augmentation method may include the following multiple steps:
[0095] In step 510, a text sample is obtained.
[0096] In the embodiments of the present disclosure, the above text sample can be a text sample in the dataset of the downstream content recognition task. The main objective of the data augmentation method described in the embodiments of the present disclosure is to perform data augmentation based on these text samples to obtain multiple augmented texts, thereby enriching the dataset of the downstream content recognition task. Thus, when training the content recognition model of the downstream content recognition task subsequently, the goal of improving the performance of the content recognition model can be achieved.
[0097] In step 520, a data augmentation prompt is generated based on the text sample.
[0098] In the embodiments of the present disclosure, the above data augmentation prompt can be generated based on the text sample and a preset data augmentation prompt template. Similar to the aforementioned data augmentation prompt, the above data augmentation prompt is mainly used to instruct the large language model to generate an augmented text based on the input text sample, and it is required that the generated augmented text should be basically consistent with the input text sample in terms of text content and expression style, but there should be diversity in the expression method and vocabulary usage between the two.
[0099] In step 530, the data augmentation prompt is input into the above large language model that has undergone multiple rounds of preference alignment fine-tuning.
[0100] As described above, in the embodiments of the present disclosure, the large language model that has undergone multiple rounds of preference alignment fine-tuning may also be the large language model that has first undergone supervised fine-tuning and then multiple rounds of preference alignment fine-tuning.
[0101] In step 540, obtain the enhanced text corresponding to the text sample output by the large language model that has undergone multiple rounds of preference alignment fine-tuning.
[0102] It can be seen from this that on the one hand, in the embodiments of the present disclosure, the semantic understanding and generation capabilities mastered by the large language model during pre-training can be used to alleviate the problem of incoherence or unreasonableness of the enhanced text; on the other hand, through the preference data pair constructed based on the training text carrying content tags and the test enhanced text generated by the large language model, the large language model can be further fine-tuned for multiple rounds of preference alignment to obtain a data enhancement model, and then the dataset of the downstream content recognition task can be enhanced with high quality and diversity, thereby effectively improving the performance of the downstream content recognition model.
[0103] In addition, the above data enhancement method and related device construct a pairwise supervised dataset through a heuristic strategy, and perform supervised fine-tuning on the large language model. By implicitly adjusting the large language model through the training loss function, the language style of the text generated by the large language model is made to be consistent with the language style of the original sample, so that the generated enhanced text is consistent with the language style of the text used in the real business scenario, further improving the performance of the downstream content recognition model.
[0104] Furthermore, in the application stage, by adjusting the temperature parameter and sampling parameter of the large language model during generation, the diversity of the generated enhanced text can be further increased. In addition, in the training stage, when the downstream content recognition task has a higher demand for the diversity of the enhanced text, the diversity of the enhanced text output by the above large language model can also be improved by adjusting the similarity threshold for constructing the above pairwise supervised dataset and the emphasis on diversity evaluation in preference annotation.
[0105] Corresponding to the above data enhancement method, some embodiments of the present disclosure also disclose a data enhancement device. Figure 6 Shows the internal structure of the data enhancement device described in the embodiments of the present disclosure. As Figure 6 shown, the above data enhancement device may include the following multiple modules:
[0106] A sample acquisition module 610, configured to acquire text samples;
[0107] A data enhancement prompt word generation module 620, configured to generate data enhancement prompt words based on the text samples;
[0108] A data augmentation module 630, configured to input data augmentation prompting words into a large language model that has undergone multi-round preference alignment fine-tuning and obtain augmented text corresponding to a text sample output by the large language model that has undergone multi-round preference alignment fine-tuning.
[0109] It should be noted that for the implementation methods of each module in the above device and the specific technical effects that can be achieved, reference can be made to the implementation methods of each step in the foregoing embodiments, and details will not be repeated here. In addition, for the method of performing multi-round preference alignment fine-tuning on the large language model and the method of supervised fine-tuning, reference can also be made to the specific methods described in the foregoing embodiments, and details will not be repeated here.
[0110] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the data augmentation method described in any of the above embodiments.
[0111] Figure 7 FIG. shows a schematic hardware structure diagram of a more specific electronic device provided in this embodiment. The device may include: a processor 2010, a memory 2020, an input / output interface 2030, a communication interface 2040, and a bus 2050. Among them, the processor 2010, the memory 2020, the input / output interface 2030, and the communication interface 2040 are communicatively connected to each other inside the device through the bus 2050.
[0112] The processor 2010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0113] The memory 2020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 2020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 2020 and are called and executed by the processor 2010.
[0114] The input / output interface 2030 is used to connect input / output devices to achieve information input and output. Among them, the input / output devices can be configured as components in the device or externally connected to the device to provide corresponding functions. The input devices can include microphones, various sensors, etc., and the output devices can include displays, speakers, vibrators, indicator lights, etc.
[0115] The communication interface 2040 is used to connect a communication module (not shown in the figure) to achieve communication interaction between this device and other devices. Among them, the communication module can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0116] The bus 2050 includes a path for transmitting information between various components of the device (such as the processor 2010, the memory 2020, the input / output interface 2030, and the communication interface 2040).
[0117] It should be noted that although the above device only shows the processor 2010, the memory 2020, the input / output interface 2030, the communication interface 2040, and the bus 2050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and do not have to include all the components shown in the figure.
[0118] The electronic device in the above embodiment is used to implement the corresponding data enhancement method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0119] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium, and the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the data enhancement method described in any of the foregoing embodiments.
[0120] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0121] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the task processing method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0122] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; within the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of brevity.
[0123] In addition, for the sake of simplicity of description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. In addition, the devices may be shown in block diagram form in order to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure will be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0124] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0125] Embodiments of the present disclosure also provide a computer program product, including computer program instructions that, when run on a computer, cause the computer to execute the above-described data enhancement method.
[0126] Embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A data enhancement method, comprising: Get a sample of text; Generate data enhancement prompt words based on the text sample; Inputting the data augmentation cue words into a large language model that has been fine-tuned through multiple rounds of preference alignment; as well as Obtain enhanced text corresponding to the text sample output by the large language model that has undergone multiple rounds of preference alignment fine-tuning.
2. The method according to claim 1, further comprising: Constructing a training data set; wherein the training data set includes a plurality of training texts; Generating data enhancement prompt words for each training text in the training data set; Inputting the data augmentation prompt words into a large language model; Acquire a plurality of test enhancement texts output by the large language model and corresponding to the plurality of training texts respectively; Performing preference annotation on the multiple test enhancement texts based on the multiple training texts, and respectively determining preference labels reflecting whether the test enhancement texts meet expectations; Generate a plurality of preference data pairs based on the plurality of training texts, the plurality of test enhancement texts, and the preference labels of the plurality of test enhancement texts; wherein the preference data pairs include: training texts, test enhancement texts corresponding to the training texts, and the preference labels of the test enhancement texts; Performing preference alignment fine-tuning on the large language model based on the preference data; Returning to the step of constructing the training data set until a preset fine-tuning termination condition is met; and The large language model after preference alignment fine-tuning is used as the large language model that has undergone multiple rounds of preference alignment.
3. The method according to claim 2, wherein: Performing preference marking on the multiple test enhancement texts includes: performing the following operations on each of the multiple test enhancement texts: Performing a key intent analysis on the training text based on the content label of the training text corresponding to the test enhanced text, and determining key points in the training text belonging to the content label; Evaluate the degree to which the test enhanced text retains the key points to obtain a consistency evaluation result; evaluating the diversity of expressions and vocabulary usage of the test enhanced text compared with the training text to obtain a diversity evaluation result; and Based on the consistency evaluation result and the diversity evaluation result, a preference label of the test enhanced text is determined.
4. The method according to claim 2, wherein: Performing preference marking on the multiple test enhancement texts includes: performing the following operations on each of the multiple test enhancement texts: Generate preference labeling prompt words based on the test enhanced text and the training text corresponding to the test enhanced text; wherein the preference labeling prompt words include: allowing the preference labeling model to perform key intent analysis on the training text based on the content label of the training text corresponding to the test enhanced text, and determine a first indication of key points in the training text that belong to the content label of the training text; allowing the preference labeling model to evaluate the degree to which the test enhanced text retains the key points, and obtain a second indication of the consistency evaluation result; allowing the preference labeling model to evaluate the diversity of expression and vocabulary usage of the test enhanced text compared with the training text, and obtain a third indication of the diversity evaluation result; and allowing the preference labeling model to determine a fourth indication of the preference label of the test enhanced text based on the consistency evaluation result and the diversity evaluation result; wherein the preference labeling model is a pre-trained large language model; Inputting the preference labeling prompt word into the preference labeling model; and Obtain a preference label of the test enhanced text output by the preference annotation model.
5. The method according to claim 2, wherein: Performing preference alignment fine-tuning on the large language model based on the preference data pair includes: performing preference alignment fine-tuning on the large language model based on the preference data pair using a KTO algorithm.
6. The method according to claim 2, wherein: The preset fine-tuning termination condition includes: the availability rate of the multiple test enhanced texts reaches a preset first availability rate threshold; or the number of preference alignment fine-tuning performed on the large language model reaches a preset number of iteration rounds.
7. The method according to claim 2, wherein: When the availability of the test enhanced text output by the large language model is lower than a preset second availability threshold, the method further includes: Rewrite the test enhancement text that does not meet expectations; Modifying the preference label of the test enhancement text to match the expected value; and A preference data pair is generated based on the training text corresponding to the test enhancement text, the rewritten test enhancement text, and the label of the rewritten test enhancement text.
8. The method according to claim 2, wherein: Before constructing the training data set, the method further includes: Constructing a paired supervised data set; wherein the paired supervised data set includes: a plurality of supervised data pairs; each supervised data pair includes: a first text and a second text; For each supervised data pair in the paired supervised data set, the following operations are performed respectively: Fine-tune the prompt word based on the first text generation model in the supervised data pair; Inputting the model fine-tuning prompt words into the large language model; Obtaining a third text output by the large language model; and The large language model is fine-tuned in a supervised manner based on the cross entropy loss between the second text and the third text to obtain a large language model that has undergone supervised fine-tuning.
9. The method according to claim 8, wherein: Constructing a paired supervision dataset includes: Extracting a first candidate text and a second candidate text belonging to the same label from a text dataset containing labels; Determining a similarity between the first candidate text and the second candidate text; Determining a co-occurring word ratio between the first candidate text and the second candidate text; and In response to determining that the similarity is greater than a preset similarity threshold and the co-occurrence word ratio is greater than a preset co-occurrence word ratio threshold, the first candidate text and the second candidate text are combined into two supervisory data pairs in different arrangement orders.
10. The method according to claim 8, wherein: Performing supervised fine-tuning on the large language model based on the cross entropy loss between the second text and the third text includes: performing supervised fine-tuning on the large language model based on the cross entropy loss between the second text and the third text by means of a low-rank adjustment method.
11. A data enhancement device, comprising: A sample acquisition module is used to acquire text samples; A data enhancement prompt word generation module generates data enhancement prompt words based on the text sample; The data enhancement module is used to input the data enhancement prompt word into a large language model that has undergone multiple rounds of preference alignment and fine-tuning; and obtain the enhanced text corresponding to the text sample output by the large language model that has undergone multiple rounds of preference alignment and fine-tuning.
12. An electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the data enhancement method according to any one of claims 1 to 10 is implemented.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the data enhancement method described in any one of claims 1-10.
14. A computer program product, comprising computer program instructions, which, when executed on a computer, enable the computer to execute the data enhancement method according to any one of claims 1 to 10.
Citation Information
Cited By
Training data generation method, electronic equipment, storage medium and program product
CN120632467A
A training data generation method, electronic equipment, storage medium and program product
CN120632467B