Controllable text generation model training method and related device

Through the method of supervised fine-tuning and reward model fusion, the controllable text generation model is trained, which solves the problem of high computing power consumption, and achieves the effect of resource saving and the generated text in line with human evaluation standards.

CN120337995APending Publication Date: 2025-07-18IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510474317.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the process of training a controllable text generation model, although the classic human feedback reinforcement learning (RLHF) link is important, it leads to high computing resource consumption. How to improve training effect while reducing resource consumption has become a key challenge.

Method used

The exception reward model and multiple constraint reward models are trained in a supervised fine-tuning manner, and fuse them to freeze the head weight of the controllable text generation model prototype, and trained in combination with the fusion reward model to reduce computing power consumption and ensure that the generated text meets human evaluation standards.

Benefits of technology

It effectively reduces the consumption of computing power resources during the training process, while ensuring that the generated text can be aligned with human evaluation standards and maintains the training effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337995A_ABST
    Figure CN120337995A_ABST
Patent Text Reader

Abstract

The invention discloses a controllable text generation model training method and a related device, and relates to the technical field of natural language processing, in the scheme, not only is a supervised fine tuning mode adopted, but also human feedback reinforcement learning is adopted, and when a controllable text generation model prototype is trained in a human feedback reinforcement learning mode, the training efficiency of the controllable text generation model prototype is improved. According to the method, most of parameters in the prototype of the controllable text generation model are frozen, a large reward model is replaced by adopting a mode of fusing a plurality of small reward models, the computing power resource consumption in the training process is greatly reduced, and the value function obtained by training can guarantee that the generated controllable text can be aligned with the human evaluation standard, so that the calculation efficiency of the controllable text generation model is improved. The training method can still guarantee the training effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, and in particular, to a method for training a controllable text generation model and related devices. Background Art

[0002] Controllable text generation means being able to add control over some attributes, styles, key information, etc. of the generated text on the basis of traditional text generation, so that the generated text meets a certain expectation. In the office field, controllable text generation tasks include meeting minutes generation, meeting to-do generation, official document writing outline generation, etc.

[0003] With the rapid development of large language model (hereinafter referred to as large model) technology in recent years, controllable text generation mainly first trains a controllable text generation model, then designs a prompt containing multi-category and multi-level elements or constraints, and then the controllable text generation model generates results that meet the prompt.

[0004] However, in the process of training a controllable text generation model, the training link based on classical human feedback reinforcement learning (RLHF) is an indispensable link. How to improve the training effect of this training link while minimizing the consumption of computing power resources during the inference process is of great significance. Summary of the Invention

[0005] In view of the above problems, this application provides a method for training a controllable text generation model and related devices to achieve the purpose of consuming computing power resources for training the controllable text generation model and ensuring the training effect. The specific solutions are as follows:

[0006] The first aspect of this application provides a method for training a controllable text generation model, including:

[0007] Perform supervised fine-tuning on the base large model to obtain a prototype of the controllable text generation model;

[0008] Use the prototype of the controllable text generation model to train an anomaly reward model and multiple constraint reward models;

[0009] Fuse the anomaly reward model and the multiple constraint reward models to obtain a fused reward model;

[0010] Freeze all weights outside the head of the prototype of the controllable text generation model, and train in combination with the fused reward model to obtain a trained controllable text generation model and a value function.

[0011] In a possible implementation, the performing supervised fine-tuning on the base large model to obtain a prototype of the controllable text generation model includes:

[0012] Obtain a training data set; the training data set includes a task instruction data set, a constraint instruction data set, and an anti-constraint instruction data set;

[0013] Adopt a strategy with the task instruction data set as the main and the constraint instruction data set and the anti-constraint instruction data set as the auxiliary to perform supervised fine-tuning on the base large model to obtain a controllable text generation model prototype.

[0014] In a possible implementation, the construction method of the training data set includes:

[0015] Collect original text data;

[0016] Based on the original text data, construct task instructions, constraint instructions, and anti-constraint instructions;

[0017] Construct output data for the task instructions to obtain a task instruction data set, construct output data for the constraint instructions to obtain a constraint instruction data set, and construct output data for the anti-constraint instructions to obtain an anti-constraint instruction data set.

[0018] In a possible implementation, the training method of the anomaly reward model includes:

[0019] Use the controllable text generation model prototype to construct an anomaly control data set; the anomaly control data set includes multiple groups of anomaly control data, and each group of anomaly control data includes normal generated text data and abnormal generated text data;

[0020] Use the anomaly control data set to train the base language model to obtain an anomaly reward model.

[0021] In a possible implementation, the training method of the multiple constraint reward models includes:

[0022] Use the controllable text generation model prototype to construct a constraint control data set. The constraint control data set includes multiple constraint control data subsets, and each constraint control data subset corresponds to a constraint condition; each constraint control data subset includes multiple groups of constraint control data, and each group of constraint control data includes generated text data that satisfies the constraint condition and generated text data that does not satisfy the constraint condition;

[0023] Use each of the constraint control data sets to train the base language model to obtain the multiple constraint reward models.

[0024] In a possible implementation, the use of the controllable text generation model prototype to construct a constraint control data set includes:

[0025] Based on the controllable text generation model prototype, the constraint instruction dataset, and the anti-constraint instruction dataset, generate constraint comparison data with different constraint conditions for each original text data sample;

[0026] Filter the constraint comparison data of each original text data sample under different constraint conditions to obtain the constraint comparison dataset.

[0027] In a possible implementation, the fusing the anomaly reward model and the multiple constraint reward models to obtain a fused reward model includes:

[0028] Obtain human preference data and the corresponding human preference ranking;

[0029] Use the human preference data and the human preference ranking to determine the fusion weights of the anomaly reward model and the multiple constraint reward models;

[0030] Use the fusion weights to fuse the anomaly reward model and the multiple constraint reward models to obtain a fused reward model.

[0031] The second aspect of this application provides a controllable text generation model training device, including:

[0032] A supervised fine-tuning unit for performing supervised fine-tuning on the base large model to obtain a controllable text generation model prototype;

[0033] A reward model training unit for training an anomaly reward model and multiple constraint reward models by using the controllable text generation model prototype;

[0034] A reward model fusion unit for fusing the anomaly reward model and the multiple constraint reward models to obtain a fused reward model;

[0035] A reinforcement learning training unit for freezing all weights except the head of the controllable text generation model prototype and performing training in combination with the fused reward model to obtain a trained controllable text generation model and a value function.

[0036] In a possible implementation, the supervised fine-tuning unit includes:

[0037] A training dataset acquisition unit for acquiring a training dataset; the training dataset includes a task instruction dataset, a constraint instruction dataset, and an anti-constraint instruction dataset;

[0038] A supervised fine-tuning subunit for performing supervised fine-tuning on the base large model by using the task instruction dataset as the main and the constraint instruction dataset and the anti-constraint instruction dataset as the auxiliary to obtain a controllable text generation model prototype.

[0039] In a possible implementation, the device further includes a training dataset construction unit, and specifically, the training dataset construction unit is configured to:

[0040] Collect original text data;

[0041] Based on the original text data, construct task instructions, constraint instructions, and anti-constraint instructions;

[0042] Construct output data for the task instructions to obtain a task instruction dataset, construct output data for the constraint instructions to obtain a constraint instruction dataset, and construct output data for the anti-constraint instructions to obtain an anti-constraint instruction dataset.

[0043] In a possible implementation, the device further includes an abnormal reward model training unit, and specifically, the abnormal reward model training unit is configured to:

[0044] Use the controllable text generation model prototype to construct an abnormal comparison dataset; the abnormal comparison dataset contains multiple groups of abnormal comparison data, and each group of abnormal comparison data contains normal generated text data and abnormal generated text data;

[0045] Use the abnormal comparison dataset to train a base language model to obtain an abnormal reward model.

[0046] In a possible implementation, the device further includes a constraint reward model training unit, and the constraint reward model training unit includes:

[0047] A constraint comparison dataset construction unit, configured to use the controllable text generation model prototype to construct a constraint comparison dataset, where the constraint comparison dataset contains multiple constraint comparison data subsets, and each constraint comparison data subset corresponds to a constraint condition; each constraint comparison data subset contains multiple groups of constraint comparison data, and each group of constraint comparison data contains generated text data that satisfies the constraint condition and generated text data that does not satisfy the constraint condition;

[0048] A constraint reward model training subunit, configured to use each of the constraint comparison datasets to train a base language model to obtain the multiple constraint reward models.

[0049] In a possible implementation, the constraint comparison dataset construction unit, specific unit:

[0050] Based on the controllable text generation model prototype, the constraint instruction dataset, and the anti-constraint instruction dataset, generate constraint comparison data with different constraint conditions for each original text data sample;

[0051] Filter the constraint comparison data of each of the original text data samples under different constraint conditions to obtain the constraint comparison data set.

[0052] In a possible implementation, the reward model fusion unit is specifically configured to:

[0053] Obtain human preference data and the corresponding human preference ranking;

[0054] Use the human preference data and the human preference ranking to determine the fusion weights of the abnormal reward model and the multiple constraint reward models;

[0055] Use the fusion weights to fuse the abnormal reward model and the multiple constraint reward models to obtain a fused reward model.

[0056] A third aspect of the present application provides a computer program product, including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement the controllable text generation model training method of the first aspect or any implementation manner of the first aspect.

[0057] A fourth aspect of the present application provides an electronic device, including at least one processor and a memory connected to the processor, where:

[0058] The memory is used to store a computer program;

[0059] The processor is used to execute the computer program so that the electronic device can implement the controllable text generation model training method of the first aspect or any implementation manner of the first aspect.

[0060] A fifth aspect of the present application provides a computer-readable storage medium, which carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, can enable the electronic device to implement the controllable text generation model training method of the first aspect or any implementation manner of the first aspect.

[0061] With the above technical solution, a method for training a controllable text generation model and related devices provided by this application first performs supervised fine-tuning on a base large model to obtain a prototype of the controllable text generation model, and then uses the prototype of the controllable text generation model to train an anomaly reward model and multiple constraint reward models, and fuses the anomaly reward model and the multiple constraint reward models to obtain a fused reward model; finally, freeze all weights outside the head of the prototype of the controllable text generation model, and train in combination with the fused reward model to obtain a trained controllable text generation model and a value function. In this solution, not only the method of supervised fine-tuning is adopted, but also human feedback reinforcement learning is adopted. When training the prototype of the controllable text generation model by the method of human feedback reinforcement learning, most of the parameters in the prototype of the controllable text generation model are frozen, and the method of fusing multiple smaller reward models is used to replace the large reward model, which greatly reduces the consumption of computing resources during the training process, and the trained value function can ensure that the generated controllable text can align with the human evaluation standard. Therefore, the training method of this application can still ensure the training effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.

[0063] Figure 1 It is a schematic flowchart of a method for training a controllable text generation model provided by an embodiment of this application;

[0064] Figure 2 It is a schematic structural diagram of a prototype of a controllable text generation model provided by an embodiment of this application;

[0065] Figure 3 It is a schematic diagram of training a controllable text generation model provided by an embodiment of this application;

[0066] Figure 4 It is a schematic diagram of a reward model fusion strategy provided by an embodiment of this application;

[0067] Figure 5 It is a schematic structural diagram of a device for training a controllable text generation model provided by an embodiment of this application;

[0068] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0069] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the embodiments of the present application are only for explaining the specific embodiments of the present application, and are not intended to limit the present application.

[0070] The embodiments of the present application will be described below with reference to the accompanying drawings. Those of ordinary skill in the art will understand that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are equally applicable to similar technical problems.

[0071] The terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device comprising a series of units does not have to be limited to those units, but may include other units not expressly listed or inherent to these processes, methods, products or devices.

[0072] Controllable text generation is to be able to add control over some attributes, styles, key information, etc. of the generated text on the basis of traditional text generation, so that the generated text meets a certain expectation. In the office field, controllable text generation tasks include meeting minutes generation, meeting to-do generation, official document writing outline generation, etc.

[0073] With the rapid development of large language model (hereinafter referred to as large model) technology in recent years, controllable text generation mainly trains a controllable text generation model first, then designs a prompt containing multi-category and multi-level elements or constraints, and then the controllable text generation model generates results that meet the prompt.

[0074] At present, the training methods of controllable text generation models mainly include the following:

[0075] (1) Only use supervised fine-tuning (SFT); only using supervised fine-tuning usually causes the model to be able to follow instructions due to the lack of effective feedback signals during the supervised fine-tuning process, but it cannot be well aligned with human standards.

[0076] (2) Through supervised fine-tuning (SFT) plus classical human feedback reinforcement learning (RLHF): The RLHF method is usually used for further "alignment" work. Here, alignment means making the generated results of the model align with human evaluations and increasing the probability that the model generates text preferred by humans.

[0077] However, with the advancement of large model technology, the size of large models is getting larger and larger, and the number of parameters has rapidly increased from about 100M to about 100B. Their training and inference require a large amount of computing power resources. In this context, the classic Reinforcement Learning from Human Feedback (RLHF) algorithm introduces multiple additional models during training, doubling the training parameters and soaring the consumption of computing power resources. With such a high consumption of computing power resources as the model size increases, the optimization efficiency and iteration difficulty also gradually increase.

[0078] (3) Through methods such as supervised fine-tuning (SFT) combined with diffusion models and contrastive learning: This method can combine the decoding process of the model with diffusion models and contrastive learning methods to achieve controllable text generation. However, there are still some gaps between introducing diffusion models and contrastive learning and aligning with human evaluation criteria compared to Reinforcement Learning from Human Feedback (RLHF).

[0079] Therefore, in the process of training a controllable text generation model, the training link based on the classic Reinforcement Learning from Human Feedback (RLHF) is an essential part. How to improve the training effect of this training link while minimizing the increase in computing power resource consumption during the inference process is of great significance.

[0080] To solve the above problems, the embodiments of the present application provide a method for training a controllable text generation model. The following will introduce the method for training a controllable text generation model according to the embodiments of the present application in detail with reference to the accompanying drawings.

[0081] Refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for training a controllable text generation model provided by the embodiments of the present application. As Figure 1 shown, a method for training a controllable text generation model provided by the embodiments of the present application may include the following steps, and these steps will be described in detail below.

[0082] S101: Perform supervised fine-tuning on the base large model to obtain a prototype of the controllable text generation model;

[0083] In the present application, the base large model can be any large model, and the present application does not make any limitations on this. Assume the base large model is denoted as M base , and after performing supervised fine-tuning on the base large model, the obtained prototype of the controllable text generation model can be denoted as M SFT .

[0084] S102: Use the prototype of the controllable text generation model to train an anomaly reward model and multiple constraint reward models;

[0085] In this application, the controllable text generation model prototype obtained after supervised fine-tuning of the base large model can generate data that meets the requirements to a certain extent, but there will be significant fluctuations. On the one hand, there may be anomalies such as model degradation. On the other hand, it is difficult to follow the instructions or the constraints after splitting the human subjective descriptions. Therefore, in this application, on the one hand, an anomaly reward model is pre-trained, and the anomaly reward model has the ability to determine whether there are anomalies (such as model degradation, etc.) in the controllable text generation model prototype. On the other hand, multiple constraint reward models can also be pre-trained, and each constraint reward model corresponds to a constraint condition. Different constraint conditions refer to the instructions that are difficult for the model after supervised fine-tuning to follow or the constraints after splitting the human subjective descriptions. The multiple constraint reward models have the ability to determine whether the output of the controllable text generation model prototype meets each constraint condition.

[0086] S103: Integrate the anomaly reward model and the multiple constraint reward models to obtain an integrated reward model;

[0087] In this application, the anomaly reward model and the multiple constraint reward models have been trained. If each reward model is independently used for subsequent training, the overall evaluation effect cannot be achieved. Therefore, in this application, it is necessary to integrate the anomaly reward model and the multiple constraint reward models to obtain an integrated reward model for subsequent training to improve the overall evaluation effect. In a possible implementation, the anomaly reward model and the multiple constraint reward models can be integrated based on a fusion strategy guided by human preference data to obtain an integrated reward model.

[0088] S104: Freeze all weights outside the head of the controllable text generation model prototype, and train in combination with the integrated reward model to obtain a trained controllable text generation model and a value function.

[0089] In this application, it is possible to freeze all weights of the trained controllable text generation model prototype M SFT except for the LM head. For model M SFT , its network structure is as Figure 2 shown. During the model training process, first freeze all Transformer Decoder Layers so that they cannot backpropagate gradients during training. Keep the LM head as the trainable part.

[0090] In this application, all weights outside the head of the controllable text generation model prototype are frozen, and the trained controllable text generation model and value function are obtained by training in combination with the fusion reward model. Substantially, the controllable text generation model prototype and value function are trained by using human feedback reinforcement learning in combination with the fusion reward model. In a possible implementation, the TD Lambda algorithm in reinforcement learning can be used to adjust the controllable text generation model prototype and learn the value function:

[0091]

[0092] , where the independent variable of the function is a t * |V| tensor, and the output is a |V|-dimensional vector. V(·) represents the value function; X 1:t represents the sequence from the 1st token to the t-th token of sequence x; t represents the number of tokens; |V| represents the vocabulary size of the large model.

[0093] For ease of understanding, refer to Figure 3 , Figure 3 which is a schematic diagram of training a controllable text generation model provided by an embodiment of this application. As Figure 3 shown, a series of Reward Models are fusion reward models.

[0094] After the training is completed, a well-learned value function and an adjusted controllable text generation model prototype are obtained, which are the trained controllable text generation model and value function.

[0095] During the inference process, the above value function can be used to adjust the output of the controllable text generation model to obtain a controllable text output that meets the human evaluation standard.

[0096] In this application, when all weights outside the head of the controllable text generation model prototype are frozen and the controllable text generation model prototype is trained in combination with the fusion reward model, only the task instruction dataset can be used.

[0097] In this application, first, a base large model is fine-tuned with supervision to obtain a prototype of a controllable text generation model. Then, using the prototype of the controllable text generation model, an anomaly reward model and multiple constraint reward models are trained, and the anomaly reward model and the multiple constraint reward models are fused to obtain a fused reward model. Finally, all weights except the head of the prototype of the controllable text generation model are frozen, and training is performed in combination with the fused reward model to obtain a trained controllable text generation model and a value function. In this solution, not only is the method of supervised fine-tuning adopted, but also human feedback reinforcement learning is used. When training the prototype of the controllable text generation model using the method of human feedback reinforcement learning, most of the parameters in the prototype of the controllable text generation model are frozen, and the method of fusing multiple smaller reward models is used to replace the large reward model, which greatly reduces the consumption of computing resources during the training process. Moreover, the trained value function can ensure that the generated controllable text can align with the human evaluation criteria. Therefore, the training method of this application can still ensure the training effect.

[0098] In another embodiment of this application, the specific implementation manner of fine-tuning the base large model with supervision to obtain a prototype of a controllable text generation model is described. This manner may include the following steps:

[0099] S201: Obtain a training data set; the training data set includes a task instruction data set, a constraint instruction data set, and an anti-constraint instruction data set;

[0100] The task instruction data set includes multiple task instruction data, and each task instruction data includes a task instruction and corresponding output data; the constraint instruction data set includes multiple constraint instruction data, and each constraint instruction data includes a constraint instruction and corresponding output data; the anti-constraint instruction data set includes multiple anti-constraint instruction data, and each anti-constraint instruction data includes an anti-constraint instruction and corresponding output data.

[0101] S202: Using the task instruction data set as the main one and the constraint instruction data set and the anti-constraint instruction data set as the auxiliary ones, fine-tune the base large model with supervision to obtain a prototype of a controllable text generation model.

[0102] In this application, the base large model can be any large model, and this application does not make any limitations on this. In this application, when fine-tuning the base model with supervision using the task instruction data set, the constraint instruction data set, and the anti-constraint instruction data set, the task instruction data set can be used as the main one, and the constraint instruction data set and the anti-constraint instruction data set can be used as the auxiliary ones. In a possible implementation, the proportion of the task instruction data set, the constraint instruction data set, and the anti-constraint instruction data set can be 1:0.1:0.1.

[0103] Specifically, assume that the supervised fine-tuning training dataset is D, and the task instruction dataset is D T , and the constraint instruction dataset is D C , and the anti-constraint instruction dataset is D AC , then the supervised fine-tuning training dataset D can be expressed as:

[0104]

[0105] where sampling is a random sampling function.

[0106] A sampling rate of 1 for the task instruction dataset can ensure that the model follows the task instructions. A sampling rate of 0.1 for the constraint instruction dataset and the anti-constraint instruction dataset can, on the one hand, enable the model to learn the basic ability to follow constraint instructions and anti-constraint instructions, and on the other hand, will not overly increase the burden of supervised fine-tuning.

[0107] Assume that the base large model is denoted as M base , after performing supervised fine-tuning on the base large model, the controllable text generation model prototype can be denoted as M SFT .

[0108] In another embodiment of the present application, the construction method of the training dataset is described, and this method may include the following steps:

[0109] S301: Collect original text data;

[0110] In the present application, the original text data can be data in the original text format, or text data transcribed by means of speech-to-text after collecting original audio data first. The original text format data and the original audio data can be collected from the network or obtained from internal services.

[0111] S302: Based on the original text data, construct task instructions, constraint instructions, and anti-constraint instructions;

[0112] The controllable text generation task involves multiple types of tasks, such as meeting minutes generation, meeting to-do list generation, official document writing outline generation, etc. Each type of task has a corresponding task description. In the present application, task instructions can be constructed based on the original text data and the task description.

[0113] Considering that different controllable text generation tasks may contain different types of constraint information, which is used to restrict the model to follow relevant constraints and better align with human evaluation metrics. For example: length constraint: the generated text must exceed 100 words; style constraint: the style of the generated text must use formal and written language and quote classics, etc. Therefore, in this application, constraint instructions can be constructed based on the original text data, the constraint information, and the task description.

[0114] In addition, in this application, reverse constraint information is also introduced to exclude unwanted content to guide the model to generate more compliant and desired results. In this application, reverse constraint instructions can also be constructed based on the original text data, the reverse constraint information, and the task description.

[0115] In a possible implementation, in the task instructions, the task description is placed in the prefix and suffix parts of the original text data to avoid truncation problems caused by overly long original text data. In the constraint instructions, the task description is placed in the prefix and suffix parts of the original text data, and the constraint information can be placed between the original text data and the prefix task description. In the reverse constraint instructions, the task description is placed in the prefix and suffix parts of the original text data, and the reverse constraint information can be placed between the original text data and the prefix task description.

[0116] It should be noted that the above format is just one implementation. The formats of the task instructions, constraint instructions, and reverse constraint instructions can also be in other forms, as long as they contain the necessary information. This application does not make any limitations in this regard.

[0117] For further understanding, in the embodiments of this application, taking the task of generating an official document writing outline as an example, examples of the corresponding task instructions, constraint instructions, and reverse constraint instructions are provided as follows:

[0118] Task instructions: You are an expert in writing work report outlines, with in-depth understanding of the typical wording and format of work reports, especially rich experience in writing outlines. Please use your professional knowledge to make a summary report based on the content in the <> below and construct a standard work report outline.

[0119] <Work in the past five years>

[0120] Constraint instructions: You are an expert in writing work report outlines, with in-depth understanding of the typical wording and format of work reports, especially rich experience in writing outlines. Please use your professional knowledge to make a summary report based on the content in the <> below and construct a standard work report outline.

[0121] 1. Convert the spoken and informal expressions within <> into formal written language, in a format similar to: "Focus on project promotion to boost the economy to a new level" or "Fully implement systematic management to optimize environmental governance", and describe it in the way of adding the project effect to the project name.

[0122] 2. Hierarchical arrangement of headings: The first-level headings start with "One, Two, Three", etc.; the second-level headings start with "(One), (Two), (Three)", etc.; the third-level headings start with "1. / 2. / 3.", etc.; the fourth-level headings start with "(1), (2), (3)", etc.

[0123] 3. Extract the key information of the work summary from <>, and summarize it into the content under the title of the work summary around the name and achievements of the work; list the specific problems and their locations regarding problem reflection in the problem reflection section. If there are no specific problems, supplement them according to the content of the work summary; extract the details about work planning, arrangement, and goals from the elements of the work plan and list them under the title of Work Plan and Arrangement. Ensure that the hierarchical relationship of the outline content is clear: If the content is extensive and contains three or more levels of headings, a structure of first-level heading followed by second-level heading, second-level heading followed by third-level heading, and third-level heading followed by fourth-level heading should be adopted.

[0124] <Work in the Past Five Years>

[0125] Anti-constraint instruction: You are an expert good at writing the outline of work reports, with in-depth understanding of the typical wording and format of work reports, especially rich experience in writing outlines. Please use your professional knowledge to make a summary report based on the content within <> below and construct a standard outline of the work report.

[0126] 1. Convert the spoken and informal expressions within <> into formal written language, in a format similar to: "Focus on project promotion to boost the economy to a new level" or "Fully implement systematic management to optimize environmental governance", and describe it in the way of adding the project effect to the project name.

[0127] 2. Hierarchical arrangement of headings: The first-level headings start with "One, Two, Three", etc.; the second-level headings start with "(One), (Two), (Three)", etc.; the third-level headings start with "1. / 2. / 3.", etc.; the fourth-level headings start with "(1), (2), (3)", etc.

[0128] 3. Extract the key information of the work summary from <>, and summarize it into the content under the title of the work summary around the name and achievements of the work; list the specific problems and their locations regarding problem reflection in <>, and if there are no specific problems, supplement them according to the content of the work summary; extract the details regarding work planning, arrangement, and goals from the elements of the work plan and list them under the title of Work Plan and Arrangement. Ensure that the hierarchical relationship of the title levels in the outline is clear: if the content is extensive and contains three or more levels of titles, a structure with a first-level title followed by a second-level title, a second-level title followed by a third-level title, and a third-level title followed by a fourth-level title should be adopted.

[0129] Do not make the following mistakes:

[0130] 1. The title content is too long;

[0131] 2. Forced problem reflection;

[0132] 3. Fabricating and inventing multi-level content;

[0133] 4. Generating detailed content for each sub-summary;

[0134] <Work in the past five years>

[0135] S303: Construct output data for the task instruction to obtain a task instruction dataset, construct output data for the constraint instruction to obtain a constraint instruction dataset, and construct output data for the anti-constraint instruction to obtain an anti-constraint instruction dataset.

[0136] In this application, manual annotation and / or AI synthesis can be used to construct output data for the task instruction to obtain a task instruction dataset; in this application, based on the output data of the task instruction, manual annotation and / or AI synthesis can be used to construct output data for the constraint instruction to obtain a constraint instruction dataset; in this application, based on the output data of the task instruction, manual annotation and / or AI synthesis can be used to construct output data for the anti-constraint instruction to obtain an anti-constraint instruction dataset.

[0137] In another embodiment of this application, the training method of the abnormal reward model is described, and this method may include the following steps:

[0138] S401: Use the controllable text generation model prototype to construct an abnormal control dataset; the abnormal control dataset contains multiple groups of abnormal control data, and each group of abnormal control data contains normal generated text data and abnormal generated text data;

[0139] In this application, an abnormal control data set can be constructed by dynamically adjusting the decoding parameters of the controllable text generation model prototype and using the task instruction data set. The decoding parameters of the controllable text generation model prototype can be dynamically adjusted by adding a repetition penalty function to control the generation of text data of different qualities by the controllable text generation model prototype. The repetition penalty function is as follows:

[0140]

[0141] where logits i is the original generation probability of the i-th token, p is the penalty value, and r i is the number of times this token has been repeatedly generated currently. Through repetition penalty, by adjusting the p value, the controllable text generation model prototype can generate abnormal text data after generating a certain number of tokens.

[0142] In a possible implementation, three different decoding parameters can be adopted to generate two types of data based on the task instruction data set D T and the controllable text generation model prototype: normal generated text data and abnormally generated text data. All generated data is collected to construct an abnormal control data set D E .

[0143] In a possible implementation, the abnormal control data set can be expressed as:

[0144]

[0145] where X represents the token sequence of the task instruction, Y1 represents the token sequence of the normally generated text data, Y2 represents the token sequence of the abnormally generated text data, R represents real numbers, L represents the length of the task instruction (the number of tokens), V represents the vocabulary size of the controllable text generation model prototype, L1 represents the sequence length of Y1, and L2 represents the sequence length of Y2.

[0146] S402: Use the abnormal control data set to train the base language model to obtain an abnormal reward model.

[0147] In this application, the abnormal control data set can be used to train the base language model M r to obtain an abnormal reward model.

[0148] The training loss of the abnormal reward model can be:

[0149]

[0150] This loss indicates that the output score of the normally generated text data is expected to be higher than that of the abnormally generated text data.

[0151] In another embodiment of the present application, the training method for multiple constraint reward models is described, and this method may include the following steps:

[0152] S501: Using the controllable text generation model prototype, construct a constraint comparison dataset, where the constraint comparison dataset contains multiple constraint comparison data subsets, and each constraint comparison data subset corresponds to a constraint condition; each constraint comparison data subset contains multiple groups of constraint comparison data, and each group of constraint comparison data contains generated text data that satisfies the constraint condition and generated text data that does not satisfy the constraint condition;

[0153] For ease of understanding, the constraint comparison dataset Dpair can be expressed as Dpair = D CI ∪D C2 ∪…∪D Cn , where D Cn represents the nth constraint comparison data subset, and each constraint comparison data subset corresponds to a constraint condition.

[0154] S502: Using each of the constraint comparison datasets, train the base language model to obtain the multiple constraint reward models.

[0155] For each constraint condition, the base language model can be trained using the constraint comparison dataset corresponding to the constraint condition to obtain the reward model for the constraint condition. The training method can refer to the training method of the anomaly reward model and will not be elaborated here.

[0156] In another embodiment of the present application, the specific implementation method for constructing the constraint comparison dataset using the controllable text generation model prototype is described, and this method may include:

[0157] S601: Based on the controllable text generation model prototype, the constraint instruction dataset, and the anti-constraint instruction dataset, generate constraint comparison data with different constraint conditions for each original text data sample;

[0158] In the present application, using the constraint instructions in the constraint instruction dataset and the anti-constraint instructions in the anti-constraint instruction dataset, different prompts can be constructed according to different constraint conditions and different original text data samples and input into the controllable text generation model prototype to obtain the constraint comparison data for each original text data sample under each constraint condition.

[0159] S602: Filter the constraint comparison data for each original text data sample under different constraint conditions to obtain the constraint comparison dataset.

[0160] In this application, the constraint comparison data that does not meet the requirements among the constraint comparison data of each of the original text data samples under different constraint conditions can be filtered. In a possible implementation, it can be achieved by means of AI feedback. Specifically, each piece of constraint comparison data is judged by AI whether it meets the requirements, and the constraint comparison data that meets the requirements is retained to construct the constraint comparison data set.

[0161] In a possible implementation, a constraint comparison data judgment prompt template can be pre-constructed. The constraint comparison data judgment prompt template includes a constraint condition filling slot and a constraint comparison data filling slot. The constraint condition and the constraint comparison data are filled into the above template to obtain a constraint comparison data judgment prompt. The constraint comparison data judgment prompt is input into a large model (such as GPT), and the judgment result can be obtained. The large model will first judge whether the constraint comparison data complies with the constraint condition x or does not comply with the constraint condition x. If the two outputs of the same sample both follow the instruction, a pair of constraint comparison data regarding the constraint condition x can be formed.

[0162] For ease of understanding, a constraint comparison data judgment prompt template example is provided in this application, which is specifically as follows:

[0163] "You are a rigorous text comparison expert who can accurately judge whether a pair of texts meets the requirements. For texts generated under opposite constraints for the same piece of input, you have very accurate recognition ability. If two texts are generated under opposite constraints and indeed meet the opposite constraint conditions, you will return True as the result. If the above conditions are not met, you will return False.

[0164] <Constraint>Constraint condition 1< / Constraint>

[0165] <Positive>The output text 1 that meets constraint condition 1< / Positive>

[0166] <Negative>The output text 1 that does not meet constraint condition 1< / Negative>".

[0167] In another embodiment of this application, the specific implementation manner of fusing the abnormal reward model and the multiple constraint reward models to obtain a fused reward model is described. This method may include the following steps:

[0168] S701: Obtain human preference data and the corresponding human preference ranking;

[0169] In this application, human preference data can be pre-collected and labeled, and the human preference ranking can be recorded.

[0170] S702: Determine the fusion weights of the abnormal reward model and the multiple constrained reward models by using the human preference data and the human preference ranking.

[0171] In this application, the fusion weights of the abnormal reward model and the multiple constrained reward models can be initialized with an even distribution, and then the human preference data is input into the abnormal reward model and the multiple constrained reward models to obtain model scores and calculate the model preference ranking. Based on the consistency between 1 - the model preference ranking and the human preference ranking as the optimization objective, the initialized fusion weights are updated, and finally the fusion weights of the abnormal reward model and the multiple constrained reward models are obtained. For ease of understanding, refer to Figure 4 , Figure 4 FIG. is a schematic diagram of a reward model fusion strategy provided by an embodiment of this application.

[0172] S703: Use the fusion weights to fuse the abnormal reward model and the multiple constrained reward models to obtain a fused reward model.

[0173] The above introduces a training method for a controllable text generation model provided by an embodiment of this application. The following will introduce the device for executing the above - mentioned training method for the controllable text generation model.

[0174] Please refer to Figure 5 , Figure 5 FIG. is a schematic structural diagram of a controllable text generation model training device provided by an embodiment of this application. As Figure 5 shown, the controllable text generation model training device includes:

[0175] A supervised fine - tuning unit 11, configured to perform supervised fine - tuning on a base large model to obtain a controllable text generation model prototype;

[0176] A reward model training unit 12, configured to train an abnormal reward model and multiple constrained reward models by using the controllable text generation model prototype;

[0177] A reward model fusion unit 13, configured to fuse the abnormal reward model and the multiple constrained reward models to obtain a fused reward model;

[0178] A reinforcement learning training unit 14, configured to freeze all weights except the head of the controllable text generation model prototype and perform training in combination with the fused reward model to obtain a trained controllable text generation model and a value function.

[0179] In a possible implementation, the supervised fine - tuning unit includes:

[0180] A training dataset acquisition unit for acquiring a training dataset; the training dataset includes a task instruction dataset, a constraint instruction dataset, and an anti-constraint instruction dataset;

[0181] A supervised fine-tuning subunit for performing supervised fine-tuning on a base large model by using the task instruction dataset as the main and the constraint instruction dataset and the anti-constraint instruction dataset as the auxiliary to obtain a controllable text generation model prototype.

[0182] In a possible implementation, the device further includes a training dataset construction unit, and the training dataset construction unit is specifically used for:

[0183] Collecting original text data;

[0184] Constructing task instructions, constraint instructions, and anti-constraint instructions based on the original text data;

[0185] Constructing output data for the task instructions to obtain a task instruction dataset, constructing output data for the constraint instructions to obtain a constraint instruction dataset, and constructing output data for the anti-constraint instructions to obtain an anti-constraint instruction dataset.

[0186] In a possible implementation, the device further includes an abnormal reward model training unit, and the abnormal reward model training unit is specifically used for:

[0187] Using the controllable text generation model prototype to construct an abnormal comparison dataset; the abnormal comparison dataset includes multiple groups of abnormal comparison data, and each group of abnormal comparison data includes normal generated text data and abnormal generated text data;

[0188] Using the abnormal comparison dataset to train a base language model to obtain an abnormal reward model.

[0189] In a possible implementation, the device further includes a constraint reward model training unit, and the constraint reward model training unit includes:

[0190] A constraint comparison dataset construction unit for using the controllable text generation model prototype to construct a constraint comparison dataset, the constraint comparison dataset includes multiple constraint comparison data subsets, and each constraint comparison data subset corresponds to a constraint condition; each constraint comparison data subset includes multiple groups of constraint comparison data, and each group of constraint comparison data includes generated text data that satisfies the constraint condition and generated text data that does not satisfy the constraint condition;

[0191] A constraint reward model training subunit for using each of the constraint comparison datasets to train a base language model to obtain the multiple constraint reward models.

[0192] In a possible implementation, the constraint comparison dataset construction unit specifically includes the following units:

[0193] Based on the controllable text generation model prototype, the constraint instruction dataset, and the anti-constraint instruction dataset, generate constraint comparison data with different constraint conditions for each original text data sample;

[0194] Filter the constraint comparison data of each original text data sample under different constraint conditions to obtain the constraint comparison dataset.

[0195] In a possible implementation, the reward model fusion unit is specifically configured to:

[0196] Obtain human preference data and the corresponding human preference rankings;

[0197] Use the human preference data and the human preference rankings to determine the fusion weights of the abnormal reward model and the multiple constraint reward models;

[0198] Use the fusion weights to fuse the abnormal reward model and the multiple constraint reward models to obtain a fused reward model.

[0199] An embodiment of the present application also provides an electronic device. Refer to Figure 6 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiment of the present application. The electronic device in the embodiment of the present application may include, but is not limited to, fixed terminals such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), desktop computers, and the like. Figure 6 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiment of the present application.

[0200] As Figure 6 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0201] Typically, the following devices can be connected to the I / O interface 605: input devices 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 can allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.

[0202] An embodiment of the present application also provides a computer program product including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement any one of the controllable text generation model training methods provided by the embodiments of the present application.

[0203] An embodiment of the present application also provides a computer-readable storage medium carrying one or more computer programs, which, when executed by an electronic device, can enable the electronic device to implement any one of the controllable text generation model training methods provided by the embodiments of the present application.

[0204] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.

[0205] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, software program implementation is a better embodiment in more cases. Based on this understanding, the technical solution of this application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.

[0206] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0207] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

Claims

1. A method for training a controllable text generation model, characterized in that, Including: Performing supervised fine-tuning on a base large model to obtain a prototype of a controllable text generation model; Using the prototype of the controllable text generation model to train an anomaly reward model and multiple constraint reward models; Fusing the anomaly reward model and the multiple constraint reward models to obtain a fused reward model; Freezing all weights except the head of the prototype of the controllable text generation model and training in combination with the fused reward model to obtain a trained controllable text generation model and a value function.

2. The method according to claim 1, wherein The performing supervised fine-tuning on the base large model to obtain a prototype of the controllable text generation model includes: Obtaining a training data set; the training data set contains a task instruction data set, a constraint instruction data set, and an anti-constraint instruction data set; Using the task instruction data set as the main and the constraint instruction data set and the anti-constraint instruction data set as the auxiliary to perform supervised fine-tuning on the base large model to obtain a prototype of the controllable text generation model.

3. The method according to claim 2, wherein The construction method of the training data set includes: Collecting original text data; Constructing task instructions, constraint instructions, and anti-constraint instructions based on the original text data; Constructing output data for the task instructions to obtain a task instruction data set, constructing output data for the constraint instructions to obtain a constraint instruction data set, and constructing output data for the anti-constraint instructions to obtain an anti-constraint instruction data set.

4. The method according to claim 1, wherein The training method of the anomaly reward model includes: Using the prototype of the controllable text generation model to construct an anomaly comparison data set; the anomaly comparison data set contains multiple groups of anomaly comparison data, and each group of anomaly comparison data contains normal generated text data and abnormal generated text data; Using the anomaly comparison data set to train a base language model to obtain an anomaly reward model.

5. The method according to claim 2, wherein The training methods of the multiple constraint reward models include: Using the prototype of the controllable text generation model to construct a constraint comparison data set, the constraint comparison data set contains multiple constraint comparison data subsets, and each constraint comparison data subset corresponds to a constraint condition; each constraint comparison data subset contains multiple groups of constraint comparison data, and each group of constraint comparison data contains generated text data that satisfies the constraint condition and generated text data that does not satisfy the constraint condition; Using each of the constraint comparison data sets to train a base language model to obtain the multiple constraint reward models.

6. The method according to claim 5, wherein The using the prototype of the controllable text generation model to construct a constraint comparison data set includes: Based on the prototype of the controllable text generation model, the constraint instruction data set, and the anti-constraint instruction data set, generating constraint comparison data with different constraint conditions for each original text data sample; Filtering the constraint comparison data of each original text data sample under different constraint conditions to obtain the constraint comparison data set.

7. The method according to claim 1, characterized in that The fusing the anomaly reward model and the multiple constraint reward models to obtain a fused reward model includes: Obtaining human preference data and the corresponding human preference rankings; Using the human preference data and the human preference rankings to determine the fusion weights of the anomaly reward model and the multiple constraint reward models; Fuse the anomaly reward model and the multiple constraint reward models using the fusion weights to obtain a fused reward model.

8. A controllable text generation model training device, characterized in that, It includes: A supervised fine-tuning unit for performing supervised fine-tuning on a base large model to obtain a prototype of a controllable text generation model; A reward model training unit for training an anomaly reward model and multiple constraint reward models using the prototype of the controllable text generation model; A reward model fusion unit for fusing the anomaly reward model and the multiple constraint reward models to obtain a fused reward model; A reinforcement learning training unit for freezing all weights outside the head of the prototype of the controllable text generation model and training in combination with the fused reward model to obtain a trained controllable text generation model and a value function.

9. A computer program product, characterized in that, It includes computer-readable instructions that, when running on an electronic device, cause the electronic device to implement the controllable text generation model training method according to any one of claims 1 to 7.

10. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store a computer program; The processor is used to execute the computer program so that the electronic device can implement the controllable text generation model training method according to any one of claims 1 to 7.

11. A computer-readable storage medium, characterized in that, The storage medium carries one or more computer programs that, when executed by an electronic device, can cause the electronic device to implement the controllable text generation model training method according to any one of claims 1 to 7.