Cboth case generation method and device based on large model, equipment and medium
Through reinforcement learning algorithms, large language models are trained, and combined with rule-driven and model-driven reward function optimization, the problem that LLM cannot meet user expectations when generating marketing copy, achieving better copy generation and improving user experience.
Patent Information
- Application Number
- CN202510735918.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
In the prior art, large-scale language models (LLMs) are difficult to effectively meet the user's expected style requirements and content quality when generating marketing copy. The traditional supervised fine-tuning method cannot fully align the model output with human preferences.
The reinforcement learning algorithm is used to train the large language model, and the weighted combination of the rule-driven first reward function and the model-driven second reward function is formed to form the target reward function, and the base model is fine-grained to optimize the base model to generate copy that meets expectations.
It improves the reliability and quality of copywriting, meets user content requirements, and improves user experience. The generated copywriting performs excellently in terms of creative uniqueness, language fluency and marketing effectiveness.
Smart Images

Figure CN120257948A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a copywriting generation method, device, equipment and medium based on a large model. Background Art
[0002] The copywriting generation ability of LLM (Large Language Model) has shown great potential in the marketing field. However, how to make the model generate marketing copy that meets style requirements and has high-quality content still requires further optimization of the training method. The traditional method is to perform SFT (Supervised Fine-Tuning) on the LLM, that is, to train with a large amount of domain data, but simple SFT often fails to fully align the model output with human preferences.
[0003] Therefore, how to more effectively adjust the training of the LLM so that the trained LLM can generate copy that better meets user expectations and has high-quality content is a technical problem that needs to be solved. Summary of the Invention
[0004] In view of this, the embodiments of this application provide a copywriting generation method, device, equipment and medium based on a large model to solve the problem that the copy automatically generated based on the large model in the prior art cannot well meet user expectations.
[0005] In the first aspect of the embodiments of this application, a copywriting generation method based on a large model is provided, including: Upon receiving a copy generation instruction, obtain a base model, where the base model is a first pre-trained large language model; Determine a first reward function, where the first reward function is a rule-driven reward function, and the rule is determined based on at least one of copy content elements, copy style constraints, copy violation rules, and copy generation process constraints; Use a reinforcement learning algorithm to fine-tune the base model based on the first reward function to obtain the base model after the first training; Call at least one reward model to determine a second reward function; Weight-combine the first reward function and the second reward function to obtain a target reward function, and use a reinforcement learning algorithm to fine-tune the base model after the first training based on the target reward function to obtain the base model after the second training; Generate copy using the base model after the second training based on the copy generation instruction.
[0006] In the second aspect of the embodiments of this application, a copywriting generation device based on a large model is provided, including: An acquisition module, configured to acquire a base model in response to receiving a copywriting generation instruction, where the base model is a first pre-trained large language model; A determination module, configured to determine a first reward function, where the first reward function is a rule-driven reward function, and the rule is determined based on at least one of copywriting content elements, copywriting style constraints, copywriting violation rules, and copywriting generation process constraints; A training module, configured to perform fine-tuning training on the base model based on the first reward function using a reinforcement learning algorithm to obtain the base model after the first training; The determination module is further configured to call at least one reward model to determine a second reward function; The training module is further configured to weight-combine the first reward function and the second reward function to obtain a target reward function, and perform fine-tuning training on the base model after the first training based on the target reward function using a reinforcement learning algorithm to obtain the base model after the second training; A generation module, configured to generate copywriting using the base model after the second training based on the copywriting generation instruction.
[0007] In a third aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0008] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0009] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: The embodiments of the present application use a reinforcement learning algorithm to adjust and train the copywriting generation base model composed of a large language model, so that the trained base model can generate more reliable and high-quality copywriting; among them, when using the reinforcement learning algorithm to perform fine-tuning training on the base model, first perform the first training on the base model based on the rule-driven first reward function, then use the reward model to determine the second reward function, weight-combine the first reward function and the second reward function to obtain the target reward function, and then perform the second training on the base model based on the target reward function, thereby realizing fine-grained optimization of the base model, enabling the trained base model to generate copywriting that meets expectations, and improving the user experience. Description of the Drawings
[0010] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0011] Figure 1 It is a schematic flowchart of a copywriting generation method based on a large model provided by an embodiment of the present application.
[0012] Figure 2 It is a schematic flowchart of a method for determining a first reward function provided by an embodiment of the present application.
[0013] Figure 3 It is a schematic flowchart of a method for determining a rule subtask reward function based on each rule provided by an embodiment of the present application.
[0014] Figure 4 It is a schematic flowchart of a method for calling at least one reward model to determine a second reward function provided by an embodiment of the present application.
[0015] Figure 5 It is a schematic flowchart of another method for calling at least one reward model to determine a second reward function provided by an embodiment of the present application.
[0016] Figure 6 It is a schematic flowchart of a method for periodically statistically calculating the difference between the generated copywriting quality level and the expected quality level provided by an embodiment of the present application.
[0017] Figure 7 It is a schematic flowchart of a method for fine-tuning and training a base model or a base model after the first training using a reinforcement learning algorithm provided by an embodiment of the present application.
[0018] Figure 8 It is a schematic diagram of a copywriting generation device based on a large model provided by an embodiment of the present application.
[0019] Figure 9 It is a schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0020] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0021] The following will describe in detail a method and apparatus for generating copywriting based on a large model according to an embodiment of the present application with reference to the accompanying drawings.
[0022] As mentioned above, the traditional method for fine-tuning and training an LLM is to perform SFT on the LLM, that is, to use a large amount of domain data for training. However, simple SFT often fails to fully align the model output with human preferences.
[0023] In related technologies, a reinforcement learning algorithm can be used to fine-tune and train the LLM to make the trained LLM more in line with user expectations. In some examples, the marketing copy generated by the LLM can be used as training data, and the LLM can be fine-tuned and trained using the reinforcement learning algorithm, and then the trained LLM can be used to generate marketing copy that is more in line with user expectations.
[0024] However, using a reinforcement learning algorithm to fine-tune and train the LLM and using the trained LLM for marketing copy generation still faces the following problems: 1) It is difficult to quantify the copywriting quality target. The quality of marketing copy depends on many factors, such as whether the writing style conforms to the brand tone, whether the format is standardized, whether the content covers product selling points, and whether it can effectively attract the attention of the audience. These targets are difficult to directly quantify through a single loss function or rule, resulting in the generated results of the model still deviating from marketing requirements even after supervised fine-tuning. For example, the model may output off-topic content, inconsistent styles, or lack call-to-action statements.
[0025] 2) How to optimize the design and optimization of the reward function. In the design of the reward function, how to balance different reward items and avoid the model from having an unbalanced reward orientation is a technical difficulty.
[0026] 3) Most current generation models directly output results and lack constraints on intermediate thinking. This results in the model possibly missing key information or having inconsistent logic. How to guide the model to plan before expressing when generating copywriting and assign rewards to the rationality of the thinking process and the quality of the output content respectively is a technical problem that needs to be solved.
[0027] 4) How to improve the efficiency and stability of LLM reinforcement learning training.
[0028] 5) How to achieve integration compatibility with existing copywriting generation and copywriting review processes.
[0029] In view of this, an embodiment of the present application provides a copywriting generation method based on a large model. This method uses a reinforcement learning algorithm to adjust and train the copywriting generation base model composed of a large language model, so that the trained base model can generate more reliable and high-quality copywriting. Among them, when using the reinforcement learning algorithm to fine-tune the base model, first, the base model is trained for the first time based on a rule-driven first reward function, then a reward model is used to determine a second reward function, and the first reward function and the second reward function are weighted and combined to obtain a target reward function. Then, the base model is trained for the second time based on the target reward function, thereby realizing fine-grained optimization of the base model, enabling the trained base model to generate copywriting that meets expectations and improving the user experience. Figure 1 FIG. is a schematic flowchart of a copywriting generation method based on a large model provided by an embodiment of the present application. Figure 1 As shown in the figure, the method includes the following steps: In step S101, in response to receiving a copywriting generation instruction, a base model is obtained.
[0030] Among them, the base model is a first pre-trained large language model.
[0031] In step S102, a first reward function is determined, and the first reward function is a rule-driven reward function.
[0032] Among them, the rule is determined based on at least one of copywriting content elements, copywriting style constraints, copywriting violation rules, and copywriting generation process constraints.
[0033] In step S103, the reinforcement learning algorithm is used to fine-tune the base model based on the first reward function to obtain the base model after the first training.
[0034] In step S104, at least one reward model is called to determine a second reward function.
[0035] In step S105, the first reward function and the second reward function are weighted and combined to obtain a target reward function. The reinforcement learning algorithm is used to fine-tune the base model after the first training based on the target reward function to obtain the base model after the second training.
[0036] In step S106, the base model after the second training is used to generate copywriting based on the copywriting generation instruction.
[0037] In some embodiments of the present application, this method can be executed by a terminal device or a server to automatically generate copywriting according to a copywriting generation instruction input by a user. In one example, the generated copywriting can be a Chinese marketing copywriting or other copywriting.
[0038] In some embodiments of the present application, after receiving a generated copywriting instruction input by a user, a base model can be obtained. The base model can be a first pre-trained LLM model. In one example, the pre-trained LLM model can be an LLM model with Chinese understanding and generation capabilities. If the LLM model is a general model, the general model can also be initially fine-tuned. For example, a certain scale of Chinese marketing copywriting data can be used to perform SFT on the base model so that the pre-trained LLM model after initial fine-tuning can preliminarily learn the corpus style and basic patterns of marketing copywriting, such as how to write relevant copy when given a product or keyword. Among them, the Chinese marketing copywriting data can include product descriptions and corresponding advertising copy, marketing titles, short copywriting examples, etc.
[0039] In addition, if the base model itself already has a certain instruction alignment ability, the SFT step can also be skipped and directly enter the reinforcement learning stage.
[0040] In some embodiments of the present application, a first reward function of the base model can be determined. The first reward function can be a rule-driven reward function. Among them, the rule can be determined based on at least one of copywriting content elements, copywriting style constraints, copywriting violation rules, and copywriting generation process constraints.
[0041] The reinforcement learning algorithm can be used to fine-tune and train the base model based on the first reward function to obtain the base model after the first training. At this time, the base model after the first training can meet the basic rule requirements of the user for the copywriting.
[0042] In some embodiments of the present application, at least one reward model can also be called to determine a second reward function. The second reward function is a model-driven reward function. The first reward function and the second reward function can be weighted and combined to obtain a target reward function.
[0043] In some implementation manners, the first reward function and the second reward function can be weighted and summed, and then normalized processing or truncation processing can be performed to obtain the target reward function. The allocation of the weights can be set according to actual needs. Positive rewards are given for positive indicators that need to be maximized, and negative rewards are given for those that need to be minimized, so as to guide the model to improve in the direction of the comprehensive optimum of multiple objectives. In addition, the weight values can also be adjusted in real time according to the model training situation.
[0044] After determining the target reward function, the reinforcement learning algorithm can be used again to fine-tune and train the base model after the first training based on the target reward function to obtain the base model after the second training. At this time, the base model after the second training can further meet the quality requirements of the user for the copywriting. Among them, the quality requirements include but are not limited to copywriting language fluency, copywriting creativity uniqueness, copywriting marketing effectiveness, etc.
[0045] Finally, the text can be generated using the base model after the second training based on the text generation instruction input by the user.
[0046] According to the technical solution provided by the embodiments of the present application, by using the reinforcement learning algorithm to adjust and train the text generation base model composed of the large language model, the trained base model can generate more reliable and high-quality text; wherein, when using the reinforcement learning algorithm to fine-tune the base model, first, the base model is trained for the first time based on the rule-driven first reward function, then the reward model is used to determine the second reward function, and the first reward function and the second reward function are weighted and combined to obtain the target reward function, and then the base model is trained for the second time based on the target reward function, thereby realizing the fine-grained optimization of the base model, enabling the trained base model to generate text that meets expectations and improving the user experience.
[0047] Figure 2 is a schematic flowchart of the method for determining the first reward function provided by the embodiments of the present application. As Figure 2 shown, the method includes the following steps: In step S201, based on the text generation instruction, determine the text content element rules and the text style constraint rules, and obtain the text violation rules and the text generation process constraint rules.
[0048] In step S202, determine a rule sub-task reward function based on each rule.
[0049] In step S203, weight and combine the rule sub-task reward functions to obtain the first reward function.
[0050] In some embodiments of the present application, when determining the first reward function, the text content element rules and the text style constraint rules can be determined first based on the text generation instruction, and then a rule sub-task reward function is determined for each rule. Finally, the rule sub-task reward functions are weighted and combined to obtain the first reward function.
[0051] Among them, the weighted combination of the rule sub-task reward functions can be to set weights for each rule sub-task reward function respectively, and then perform weighted summation and normalization processing on the rule sub-task reward functions to obtain the first reward function. The weights of the rule sub-task reward functions can be set according to empirical values, or an initial weight can be set and then dynamically adjusted when using reinforcement learning to perform the first fine-tuning training on the base model.
[0052] Figure 3 is a schematic flowchart of the method for determining a rule sub-task reward function based on each rule provided by the embodiments of the present application. As Figure 3As shown, the method comprises the following steps: In step S301, a first rule subtask reward function is determined based on the copy content element rule.
[0053] Among them, the first rule subtask reward function rewards the generated copy that includes content elements, and punishes the generated copy that omits content elements.
[0054] In step S302, a second rule subtask reward function is determined based on the copywriting style constraint rule.
[0055] Among them, the second rule subtask reward function rewards the generated copy that conforms to the copy style, and punishes the generated copy that does not conform to the copy style; the copy style is determined by at least one of the format of the copy, the style type of the copy, and the language style of the copy.
[0056] In step S303, a third rule subtask reward function is determined based on the text violation rule.
[0057] Among them, the third rule subtask reward function rewards the generated copy that complies with the copy violation rules, and punishes the generated copy that does not comply with the copy violation rules.
[0058] In step S304, the fourth rule subtask reward function is determined based on the copywriting generation process constraint rules.
[0059] Among them, the fourth rule subtask reward function rewards the generated copy that meets the copy generation process constraint rules, and rewards the generated copy that does not meet the copy generation process constraint rules; the copy generation process constraint rules include that the copy is generated in a format that thinks about the process first and then outputs the content.
[0060] In some embodiments of the present application, determining a rule subtask reward function based on each rule may be to determine a first rule subtask reward function based on the copy content element rule, wherein the first rule subtask reward function rewards the generated copy that includes the content element and penalizes the generated copy that omits the content element.
[0061] In one example, the content elements may include, for example, a specified product name, a core selling point keyword, etc. If the first rule subtask reward function determines that the generated copy includes the specified product name and the core selling point keyword, a reward may be given, otherwise a penalty may be given.
[0062] In some other embodiments of the present application, determining a rule subtask reward function based on each rule may be to determine a second rule subtask reward function based on the copywriting style constraint rule. The second rule subtask reward function rewards the generated copywriting that conforms to the copywriting style and punishes the generated copywriting that does not conform to the copywriting style; the copywriting style is determined by at least one of the format of the copywriting, the language type of the copywriting, and the language style of the copywriting.
[0063] In one example, the copywriting style may include the format of the copywriting, such as the output length, the language type of the copywriting, and the language style of the copywriting, such as whether it includes a specified sentence pattern. For example, if the second rule subtask reward function determines that the output length of the copywriting is within a preset range, or the end of the copywriting sentence contains a call to action sentence pattern "Experience it now!", then a reward may be given, otherwise a penalty may be given.
[0064] In some other embodiments of the present application, determining a rule subtask reward function based on each rule may also be to determine a third rule subtask reward function based on the copy violation rule. The third rule subtask reward function rewards the generated copy that complies with the copy violation rule, and punishes the generated copy that does not comply with the copy violation rule.
[0065] In one example, the third rule subtask reward function can detect the generated copy based on the sensitive word list and preset violation rules to determine whether it contains any violation, content that denigrates competitors, etc. If not, a reward can be given, otherwise a penalty will be imposed.
[0066] In some other embodiments of the present application, determining a rule subtask reward function based on each rule may also be to determine a fourth rule subtask reward function based on the text generation process constraint rule. Wherein, the fourth rule subtask reward function rewards the generated text that meets the text generation process constraint rule, and rewards the generated text that does not meet the text generation process constraint rule; the text generation process constraint rule includes that the text is generated in a format of thinking process first and then outputting content.
[0067] In order to accurately evaluate the LLM model, the LLM model can be set to explicitly go through the thinking process when generating the copy. In this way, the model's thinking process and the generated copy as a whole can be evaluated, which improves the accuracy of the evaluation.
[0068] In other words, you can set up LLM to separate the thinking process from the generation of the copy. In one example, you can specify that the format of the LLM answer includes two parts:<think(思考过程)> Partial and<answer(生成文案)> .in, <think>The label can list the internal thinking process or reasoning chain within the LLM. <answer>The text inside the label is the finally generated copywriting.
[0069] For example, when the model receives a generation task, it first <think>List the key points, structure, or creative ideas in the paragraph, and then at <answer>The paragraphs give a complete and coherent marketing copy. This format makes the intermediate reasoning of the model explicit and easy to evaluate separately.
[0070] Based on the generation process constraint rules, the fourth rule subtask reward function can be used to detect the content in the copywriting generation process. <think>Part of the reasoning is reasonable. For example, if enough product selling points are enumerated or the pain points of the target users are mentioned, it can be determined that it meets the generation process constraint rules and a reward is given. At the same time, if the reasoning process covers more comprehensive information and has a clearer structure, the higher the reward. On the contrary, if the reasoning process does not meet the generation process constraint rules, a penalty can be given.
[0071] The above rules can be directly calculated according to the pre-specified criteria and have clear interpretability. Using these rules to formulate the corresponding first reward function, and then using the reinforcement learning algorithm to fine-tune the base model based on the first reward function, the base model after the first training that conforms to the rules that must be followed in the copywriting generation process can be obtained first, providing a basis for subsequent training of the base model based on the copywriting quality level.
[0072] Figure 4 It is a schematic flowchart of a method for calling at least one reward model to determine a second reward function provided by an embodiment of the present application. As Figure 4 shown, the method includes the following steps: In step S401, at least one copywriting evaluation index is obtained, and each copywriting evaluation index is used to evaluate a quality level of the generated copywriting.
[0073] Among them, the quality level includes at least the language fluency, creative uniqueness, and marketing effectiveness of the generated copywriting.
[0074] In step S402, the reward models corresponding to the respective evaluation indexes are determined.
[0075] Among them, the reward model is a pre-trained reward model or a second pre-trained large language model, and the second pre-trained large language model is the same as or different from the first pre-trained large language model.
[0076] In step S403, the reward values of the respective reward models are determined as the model sub-task reward functions.
[0077] In step S404, the respective model sub-task reward functions are weighted and combined to obtain the second reward function.
[0078] In some embodiments of the present application, when determining the second reward function, at least one copywriting evaluation index can be obtained first, and each copywriting evaluation index is used to evaluate a quality level of the generated copywriting, and then the reward models corresponding to the respective evaluation indexes are determined. Next, the reward values of the respective reward models are determined as the model sub-task reward functions. Finally, the respective model sub-task reward functions are weighted and combined to obtain the second reward function.
[0079] Among them, the reward model can be, for example, a pre-trained reward model. For example, some marketing copy and the scoring data of humans on its dimensions such as creativity, writing style, and persuasiveness can be collected in advance to train a small model to predict the comprehensive score of the copy. During reinforcement learning, the model output is input into the scoring model to obtain a score as the reward value.
[0080] On the other hand, the reward model can also be a second pre-trained LLM. For example, the LLM can be used as an evaluator to achieve zero-shot scoring. For example, an evaluation prompt can be constructed, and the generated copy together with the requirements is input into the LLM to ask it to give an evaluation or judge whether it meets the expectations. With the strong model knowledge of the LLM and the simulation of human preferences, high-quality feedback signals can also be obtained.
[0081] Different evaluation models can be trained or prompted for different dimensions. For example, one focuses on checking the language fluency of the copy, and the other focuses on marketing effectiveness. Scores are given respectively, and then weighted and aggregated into the final model-based reward, that is, the second reward function.
[0082] It should be noted that bias should be avoided when training the reward model, and it is ensured that the evaluation model is independent of the generation model.
[0083] Figure 5 It is a schematic flowchart of another method provided by the embodiments of the present application for determining the second reward function by invoking at least one reward model. Among them, Figure 5 Steps S501 to S504 in the illustrated embodiment are basically the same as Figure 4 Steps S401 to S404 in the illustrated embodiment, and will not be elaborated here. As Figure 5 shown, the method further includes the following steps: In step S505, when using the reinforcement learning algorithm to fine-tune the base model after the first training based on the target reward function, in response to determining that the base model after the first training has not converged, the difference between the quality level of the generated copy and the expected quality level is periodically counted.
[0084] In step S506, the weights of each reward model are updated based on the difference.
[0085] In step S507, the sub-task reward functions of each model are weighted and combined based on the updated weights to obtain the second reward function.
[0086] In some embodiments of the present application, when using the reinforcement learning algorithm to fine-tune the base model after the first training based on the target reward function, if the base model after the first training does not converge, the difference between the generated copywriting quality level and the expected quality level can be periodically statistically analyzed, and the weights of each reward model can be updated based on this difference. Finally, the second reward function can be obtained by weighted combining the sub-task reward functions of each model based on the updated weights.
[0087] Figure 6 FIG. 4 is a schematic flowchart of a method for periodically statistically analyzing the difference between the generated copywriting quality level and the expected quality level provided by an embodiment of the present application. As Figure 6 shown, the method further includes the following steps: In step S601, a validation data set is periodically obtained.
[0088] Among them, the validation data set includes different types of historical generated copywriting instructions and corresponding historical copywriting.
[0089] In step S602, the base model after the first training that does not converge during the fine-tuning training is used to generate validation copywriting based on the historical generated copywriting instructions.
[0090] In step S603, the difference between the quality level of the validation copywriting and the quality level of the historical copywriting is statistically analyzed to obtain the difference between the generated copywriting quality level and the expected quality level.
[0091] In some embodiments of the present application, periodically statistically analyzing the difference between the generated copywriting quality level and the expected quality level may be to periodically obtain a validation data set, where the validation data set includes different types of historical generated copywriting instructions and corresponding historical copywriting. Then, the base model after the first training that does not converge during the fine-tuning training is used to generate validation copywriting based on the historical generated copywriting instructions. Finally, the difference between the quality level of the validation copywriting and the quality level of the historical copywriting is statistically analyzed to obtain the difference between the generated copywriting quality level and the expected quality level.
[0092] That is to say, the reward model can be iteratively optimized during the fine-tuning training of the base model. In one example, the performance of the model on the validation set (including some unseen marketing copywriting tasks) can be periodically evaluated, including the achievement of rule metrics, the scoring of the auxiliary model, and the subjective evaluation of users when necessary. If the model still has obvious shortcomings, such as a certain style not meeting the standard, the reward weights can be adjusted accordingly or new training samples can be added to continue fine-tuning. When the model meets the expected threshold on all indicators and the training reward tends to be stable, it can be determined that the convergence is achieved and the training is ended.
[0093] In some embodiments, rejection sampling and supervised fine-tuning can also be combined. For example, after using a reinforcement learning algorithm to fine-tune and train the base model to obtain the base model after the second training, high-quality samples generated by the model can be collected and further supervised fine-tuning can be performed to consolidate the model's performance. For copywriting generation, excellent copywriting samples produced by the model can also be collected, reviewed, and added to the training to continuously improve the model's level.
[0094] Figure 7 FIG. is a schematic flowchart of a method for fine-tuning and training a base model or a base model after the first training using a reinforcement learning algorithm provided by an embodiment of the present application. As Figure 7 shown, the method further includes the following steps: In step S701, obtain the currently available resources.
[0095] In step S702, in response to determining that the currently available resources are greater than or equal to a preset resource threshold, use the Proximal Policy Optimization algorithm (PPO) or the Group Relative Policy Optimization algorithm (GRPO) to fine-tune and train the base model or the base model after the first training.
[0096] In step S703, in response to determining that the currently available resources are less than the preset resource threshold, use the Reinforcement Learning with Optimized Objectives algorithm (RLOO) to fine-tune and train the base model or the base model after the first training.
[0097] In some embodiments of the present application, when using a reinforcement learning algorithm to fine-tune and train a base model or a base model after the first training, the currently available resources can be obtained first, and then it can be determined whether the currently available resources are greater than or equal to a preset resource threshold. If so, the Proximal Policy Optimization algorithm (PPO) or the Group Relative Policy Optimization algorithm (GRPO) can be used to fine-tune and train the base model or the base model after the first training. Conversely, if the currently available resources are less than the preset resource threshold, the Reinforcement Learning with Optimized Objectives algorithm (RLOO) can be used to fine-tune and train the base model or the base model after the first training.
[0098] That is to say, in the case of sufficient resources, PPO, GRPO, and their variant algorithms can be selected to fine-tune and train the base model. Conversely, in the case of limited resources, algorithms such as RLOO can be selected to fine-tune and train the base model. In addition, the Kullback-Leibler Divergence (KL divergence, relative entropy) constraint can be introduced, that is, the difference between the base model after fine-tuning training and the initial base model is restricted not to exceed a certain range to avoid the model generation deviating from the human language style.
[0099] By adopting the technical solution provided by the embodiment of the present application, the finally obtained base model after the second fine-tuning training can output marketing copy that meets the requirements for a given Chinese input, such as product descriptions, marketing keywords, etc. In this process, the model is driven by a finely designed reward function and learns strategies to follow rules and pursue copy quality, thus overcoming the deficiencies of the original model in marketing copy generation. The technical solution provided by the embodiment of the present application is optimized for Chinese content and marketing scenarios, integrating a composite reward of rules and model evaluation, as well as an advanced and efficient reinforcement learning algorithm, so it can achieve a significant improvement in effect at a relatively low cost.
[0100] The following gives some typical usage scenarios of the technical solution provided by the embodiment of the present application.
[0101] E-commerce product advertisement generation: On online retail platforms, thousands of products need to have promotional copywritten. The model provided by the embodiment of the present application can be used to automatically generate attractive advertising slogans or product descriptions based on the title, selling points, and parameter descriptions of each product.
[0102] For example, for a newly launched coffee product, after the model understands its flavor characteristics and target consumers, it first lists selling points such as "rich and refreshing taste", "produced in a well-known origin", "limited-time preferential promotion" during the thinking process, and then outputs an attractive advertisement like: " <answer>The first cup of awakening in the early morning: Select Yirgacheffe beans, with a rich and mellow aroma that refreshes the mind. Place an order now to enjoy the limited-time discount and start your day more energized than greeting your boss good morning! < / answer> ". The entire process requires no manual intervention. The generated copy not only contains product highlights but also has a call to action, and meets the platform's requirements for word count and content specifications.
[0103] Social media marketing copy: Brands need to frequently post marketing content on social media such as WeChat and Weibo, such as new product promotion soft articles, holiday promotion tweets, etc. The model provided by the embodiment of the present application can generate creative copy that conforms to the brand tone according to the given theme or event information.
[0104] For example, for a Valentine's Day promotion event, by inputting the theme and preferential information, the model can output a tweet with a romantic atmosphere without losing the brand positioning, and ensure that the copy contains necessary information such as event details and participation methods. Due to the addition of rewards for style and format, the model can control the tone between the lines, making the copy both infectious and maintaining brand consistency. Social platforms have strict restrictions on content compliance and word count. The model has considered these factors under the drive of rewards before generating copy, reducing the risk of content review failure.
[0105] Advertising creative copy generation service: Advertising agencies or copywriting teams can package this model as an AI (Artificial Intelligence) copywriting assistant service.
[0106] The typical usage is that the user provides some keywords (brand names, product functions, target users, etc.) and requirements (lively style, suspenseful, etc.), and the model immediately generates multiple versions of copywriting plans with different wordings for selection. Since the model is optimized through reinforcement learning, it knows how to follow different style requirements (for example, in the reward, the preference of the evaluation model can be adjusted for the "lively style" or "suspenseful style"), so the output plans are diverse in style but all relevant to the topic. This accelerates the creative iteration process. Copywriters can obtain inspiration from the model suggestions or directly adopt them, greatly improving work efficiency.
[0107] At the same time, the service also allows users to customize rules (such as prohibiting the appearance of competitors' names, etc.). These rules will be injected into the model's reward system in real time to ensure that the generated results meet the user's customized needs. This scenario reflects the flexibility and practical value of the technical solution provided by the embodiments of this application in industrial applications.
[0108] Personalized marketing email generation: When enterprises send customized marketing emails to different customer groups, they can use this model to generate corresponding wordings and contents according to the user portraits.
[0109] For example, for VIP (Very Important Person) customers, ordinary users, and potential users, the emphasized selling points and tones have different focuses.
[0110] Traditionally, copywriters need to write multiple sets of templates. However, by using the model provided by the embodiments of this application, it can automatically generate email copy according to the individual situation, while ensuring that all emails follow the unified brand voice and format. The thinking process part of the model can list the key points for this user group (such as VIPs emphasizing a sense of dignity, ordinary users highlighting the preferential intensity, etc.), and the output is the complete email body. This not only saves manpower but also enables the production of truly large-scale personalized marketing content.
[0111] Multimodal content marketing: The technical solution provided by the embodiments of this application can also be extended to content scenarios combining pictures and texts. For example, given a product picture or a promotional poster, the model provided by the embodiments of this application can be combined with a visual model to first analyze the picture to extract key information, and then generate a matching text description or copywriting. After specialized training, the model can integrate image clues into the thinking process, such as identifying the scene of the product in the picture, high-class elements, etc. <think>Partially record this information and then generate corresponding copywriting.
[0112] In typical applications, e-commerce platforms can automatically generate product selling point descriptions for a vast number of product images; the tourism industry can generate promotional texts based on scenic photos, etc. Since the core of the technical solution provided by the embodiments of this application lies in the optimization of text generation, it can be combined with an image processing module to achieve the generation of multi-modal marketing content, further expanding the application scope.
[0113] The above scenarios are just representative examples. The technical solution provided by the embodiments of this application is applicable to almost all fields that require batch generation of high-quality copywriting, including but not limited to: news headline writing, film and television promotional slogans, biopharmaceutical product promotion copywriting, etc. Whether it is on a content creation platform, a digital marketing company, or an in-house marketing department of an enterprise, the technical solution provided by the embodiments of this application can be deployed and used to help automatically generate high-quality Chinese copywriting. By flexibly setting reward rules, it can also adapt to the special needs of different industries and different markets, and is a truly universal and expandable AI copywriting generation solution.
[0114] The following gives some embodiments of automatically generating copywriting using the technical solution provided by the embodiments of this application.
[0115] Embodiment 1: Reinforcement learning optimization of e-commerce category marketing copywriting This embodiment takes the generation of product marketing copywriting on an e-commerce platform as the scenario to demonstrate how to fine-tune and optimize the LLM using the technical solution provided by the embodiments of this application.
[0116] Step 1: Training data preparation and base model fine-tuning. First, collect marketing text data for several product categories on the e-commerce platform, including product titles, descriptions, selling points, and manually written advertising copy, such as short product descriptions, promotional slogans, etc. When constructing the dataset, the basic product information can be used as the input, and the manual copywriting as the target output. In one example, the open-source Chinese base model Qwen2.5-7B can be selected as the initial model. Use the above data to perform SFT fine-tuning on Qwen2.5-7B to enable it to learn the ability to generate corresponding copywriting from product information. After fine-tuning, the model can already output basically qualified copywriting, but the creativity and consistency need to be strengthened.
[0117] Step 2: Design the output format separation idea and content. Before reinforcement learning fine-tuning, the interaction prompt can be modified to make the model follow a specific output format: <think> …< / think> The paragraph contains the model's internal thinking about the input product, such as listing the main selling points of the product, the target consumer group, and the intended tone style; <answer> …< / answer> The paragraph is the final marketing copy presented to the user.
[0118] Taking a smartwatch as an example, when the model receives the product parameters, it is expected that it <think>Partially list content such as "Selling point 1: Health monitoring function, Selling point 2: Long battery life, Target users: Fitness enthusiasts, Style: Inspiring" and then <answer>Generate a promotional text that incorporates these key points. This format requires teaching by adding several example conversations to the training samples and clearly specifying the format in the system prompt. The model will be repeatedly reminded to follow this format during the reinforcement learning process.
[0119] Step 3: Implementation of the reward function. Based on the requirements of this scenario, a composite reward function can be implemented, including: 1) Rule checking - selling point coverage. Extract the list of key selling points provided by the product and match it with the model output. If <answer>If the copywriting covers all the major selling points, a +1 bonus is given; if any one is omitted, a negative bonus is given according to the number of omissions. For example, if a mobile phone has three major selling points: "triple cameras", "fast charging", and "high refresh rate screen", and the copywriting does not mention "fast charging", the bonus for this item is recorded as -1.
[0120] 2) Rule checking - Format and tone. Check <answer>Whether the paragraph meets the platform requirements, such as a length not exceeding 50 characters, the presence of punctuation at the end of the sentence, and whether the overall tone matches the expected style (which can be roughly judged by the presence of exclamation marks, etc.). For each requirement met, a +0.5 bonus is given, and -0.5 for non-compliance. <think>The part must be hidden and not output to the end user, which can be recognized in the environment <think>Label to ensure that only the final result is presented <answer>Content, if it is found that the model violates the format, severe punishment will be imposed (for example, a reward of -2).
[0121] 3) Rule check - prohibited words. A list of prohibited words for e-commerce copywriting can be established (such as legal prohibited words like extreme terms "optimal" and "national level", etc.). Scan <answer>, if there is any forbidden word, immediately impose a huge penalty of -5 and mark the end of this round, so as to ensure that the model tries its best to avoid touching the red line.
[0122] 4) Model scoring - Copywriting quality scoring. A small evaluation model RM1 can be trained using a batch of product copywriting and user feedback (such as click-through rate, highly liked copywriting) collected previously. The scoring range is 0-10, which measures the attractiveness and persuasiveness of the copywriting. Every time the model generates <answer>After that, RM1 scores it and maps it to a reward value in the range of -1 to +1. In one example, a score of 8 or above can be set as +1, a score below 5 as -1, and the rest are linearly mapped.
[0123] 5) Model scoring - style matching degree. At the same time, another evaluation model RM2 can also be trained, or a large evaluation model can be directly prompted. For example, let GPT-4 judge whether the output copywriting meets the preferences of the target users of the product. In one example, for the copywriting of fitness products, another model can be used to judge whether it is full of dynamics and encouragement. RM2 can output yes or no. If the output is yes, for example, it can be recorded as +0.5, otherwise it is recorded as 0. This sub-reward ensures that the copywriting style is close to the expected audience.
[0124] 6) Reward for the thinking process. To encourage the model to correctly use <think>Part, it is possible to formulate if <think>There are more than N selling points listed in the paragraph, and these selling points are later in <answer>If it is mentioned in most cases, the bonus is increased by 1; otherwise, if <think>With <answer>If the content does not match seriously (for example, listed in the thinking but not used in the output), then -1. Where N is a positive integer and can be dynamically set according to the amount of product information. This prompts the model to take the thinking process seriously and not be perfunctory.
[0125] Each of the above sub-rewards will be calculated and accumulated in real time to form the total reward R_total. For example, for a certain output, if the copywriting covers all the selling points (+1), has an appropriate length with an exclamation mark (+0.5), has no prohibited words (+0), has an RM1 score of about 0.8 (+0.8), RM2 is judged as Yes (+0.5), and the thinking process is sufficient (+1), then the total reward ≈ +3.8 points; on the contrary, if a selling point is missed (-1), there are prohibited words (-5), and the quality score is low (-0.7), then the total reward may be negative, and the model will quickly eliminate such strategies.
[0126] Among them, the values of each reward can be adjusted according to actual needs and are not restricted here.
[0127] Step 4: Fine-tuning training with reinforcement learning. The open-source EasyR1 framework can be used to perform fine-tuning using the GRPO algorithm. In one example, the basic model fine-tuned in Step 1 can be used as the initial strategy Load, and prepare a reference model with the same structure but frozen parameters for calculating the KL penalty.
[0128] During training, for each product input, the policy model samples, for example, k = 4 different <think> + <answer>Candidate. Calculate the above-mentioned rewards \(R\) for each candidate i , and then calculate the average reward as the baseline, where \(i\) is a positive integer. For each sample, according to adjust the policy: if is positive, increase the probability that the sample is generated by the policy, otherwise decrease its probability. The gradient approximation can be obtained by weighting the log probability using the policy gradient formula. At the same time, a KL constraint term can also be added to ensure that the new policy does not deviate too much from the original model. For example, a coefficient can be set to control KL within a certain threshold. After each batch is updated, let the parameters take a small step in the direction of improving the reward. After tens of thousands of steps of training, the model gradually learns the output policy that makes the reward function score higher.
[0129] During the training process, the trend of the average reward can be monitored. When it stabilizes on the validation set and approaches the theoretical full score, stop the training. The finally obtained policy model is the optimized copywriting generation model.
[0130] Step 5: Result verification. A group of product information that has not participated in the training can be selected for the model to generate copywriting, and it can be compared with the output of the original SFT model and the artificial copywriting.
[0131] Experiments show that the copywriting output by the technical solution provided in the embodiments of the present application completely covers the product selling points, is fluent and contagious in language, has no illegal vocabulary, and the format also fully meets the requirements. In contrast, the SFT model sometimes misses details or has plain sentences.
[0132] In the blind artificial test, the members of the marketing team generally believe that the copywriting creative quality of the model of this solution is close to that of humans, and some are even considered to be better than some ordinary copywriting written by humans. This shows that reinforcement learning does enable the model to capture the essence of human preferences.
[0133] In the quantitative evaluation, the model output approaches the reference copywriting in coverage metrics such as Rouge-L (Recall-Oriented Understudy for Gisting Evaluation - Longest Common Subsequence) and BLEU (Bilingual Evaluation Understudy), that is, it is consistent with the encouragement to cover the selling points during training, and the user satisfaction is significantly higher than that of the non-RL optimized model.
[0134] Example 2: Copywriting generation with multiple style requirements This embodiment demonstrates how the model uses this solution for training under various style requirements, so as to flexibly adjust the copywriting tone according to the needs of different brands or activities. Suppose there is a copywriting generation service that needs to output two very different advertising copy in "humorous style" and "formal style" for different customers.
[0135] Step 1: Collection of multi-style training corpus. Prepare two sets of copywriting samples: one set is more humorous and sassy, for example, it is required to use Internet buzzwords and more metaphorical gags; the other set is more formal and professional, for example, it is required to use standard words and a serious tone. Each set contains example copywriting on several topics.
[0136] Step 2: Introduce style instructions and format the output. A field can be added to the model input or the required style type can be specified through system prompts. For example, add "[Style: Humorous]" or "[Style: Formal]" to the prompt (0 prompt). The output format still adopts <think> / <answer>Separate, but require <think>Partially consider style elements.
[0137] Step 3: The reward function is extended for style. Based on Example 1, a style matching reward can be added: Use pre-prepared style samples to train two discriminator models to respectively judge whether a given copywriting belongs to a humorous style or a formal style. If the style of the copywriting output by the model matches the input requirement, the corresponding discriminator gives a high score (reward +1), otherwise the reward is -1. In addition, add inspection items for specific styles in the rule check. For example, for the humorous style, emojis or colloquial words are expected to appear, then detect these features; for the formal style, filler words are avoided, then detect whether there are colloquial fillers such as "oh, right". Once the style does not match, immediately impose a penalty.
[0138] Step 4: Adjust the reinforcement learning training process. During training, the style requirement can be randomly specified in the input, allowing the model to alternately optimize among different styles. The reward function evaluates the output in real time according to the required style. This is equivalent to allowing the model to learn two strategies simultaneously, but switching through conditional control under unified parameters. The GRPO algorithm itself remains unchanged, but attention needs to be paid to controlling the proportion of different style samples in training to prevent the model from preferring one of the styles. In practice, a content encouraging diversity can be added to the reward, or a multi-strategy fusion method can be used to ensure balance.
[0139] Step 5: Testing and effects. After training is completed, let the model generate two types of copywriting, a humorous version and a formal version, for the same theme. The results show that the model can flexibly switch the writing style: the humorous version is full of wit and has continuous punchlines, while the formal version is worded rigorously and emphasizes credibility. This shows that the model has successfully learned style as a controllable dimension. When changing the brand or scenario, only need to provide new style samples and re-fine-tune for a small stage, then it can adapt to the new requirements, demonstrating the adaptability of the technical solution provided by the embodiment of the present application in diverse copywriting generation.
[0140] As shown by the above embodiments, the technical solution provided by the embodiment of the present application can be transformed and extended in many aspects according to actual application requirements. For example, as long as a suitable reward function is defined, the technical solution provided by the embodiment of the present application can also be used for the generation and optimization of English marketing copywriting or multilingual copywriting; or apply it to the scenario of a conversational marketing assistant, allowing the model to learn to better interact with users and promote products through reinforcement learning. All these changes are within the scope of the idea of the present application and do not deviate from its spirit and protection scope.
[0141] Adopting the technical solution provided by the embodiments of the present application can comprehensively improve the quality of the generated copywriting. Due to the adoption of composite reward optimization, the fine-tuned model is significantly superior to the model that does not use this method in terms of content accuracy, style consistency, creativity, etc. Through the guidance of the reward function, the model has learned to follow the essential elements of marketing (such as including selling points and calls to action) and avoid low-quality content. The experimental results show that the model fine-tuned by RL is more favored by human reviewers than the model with only SFT, and its generated results have significantly improved in terms of usefulness and effectiveness scores. Especially in the automatic evaluation metrics (such as Rouge, BLEU) and human evaluation in the advertising copywriting scenario, higher scores are obtained.
[0142] At the same time, the generated copywriting meets the requirements of rule compliance. Thanks to the embedded rule rewards, the model output can strictly follow the preset business rules and legal norms. For example, there will be no violation of vocabulary, and it will not deviate from the style guide formulated by the brand; ensuring the usability and security of the copywriting in the real scenario. In contrast, traditional models often need to filter out non-compliant content through rule filtering or even manual review after inference, while the model trained by this solution is inherently built with a compliance tendency, greatly reducing the post-processing cost.
[0143] In addition, the thinking ability of the generated copywriting model is enhanced. Through the reward constraint on the "thinking process", the model has learned to deliberate on the copywriting concept more rigorously, manifested as the output text having a clear structure, strict logic, and being less likely to have contradictions or omissions of key information. This improvement in the internal thinking ability also brings the benefit of generalization performance: when facing new products or marketing themes, the model can also organize the copywriting content in an orderly manner, rather than simply relying on the clichés seen in training. This solves the problem of scattered content or missing the point that sometimes occurs in previous models.
[0144] Moreover, the training efficiency and stability of the base model for generating copywriting have been improved. By adopting optimization algorithms such as GRPO / RLOO, the training resource overhead is reduced, enabling RL fine-tuning of large models under relatively limited resource conditions. While ensuring the effect, the training time is shortened, and the video memory occupancy is reduced by more than half. In addition, the design of multiple rewards plays a role in shaping the rewards, avoiding the situation where the model may go to extremes under a single target (such as over-responding to a certain metric), and the mutual restraint of each metric makes the training converge more smoothly.
[0145] Moreover, this method is flexible, controllable, and scalable. Since the reward function design is modular, users can adjust or expand the metric weights according to specific application requirements to achieve customized generation optimization. For example, a reward item targeting brand tone can be easily added, or localization style evaluation metrics can be incorporated before launching in a new market. This method has high controllability. Instead of being as difficult to intervene as a black box, the model can be "tuned" by adjusting the rewards. In addition, this solution does not rely on a specific model and is applicable to various pre-trained LLMs. It can be migrated as long as the corresponding computing power is provided. This means that whether it is for generating e-commerce product descriptions, social media copywriting, or marketing content in other languages, the technical solutions provided in the embodiments of this application can be used for optimization, and it has good scalability.
[0146] Any combination of the above optional technical solutions can form optional embodiments of this application, which will not be elaborated one by one here.
[0147] The following are embodiments of the device of this application, which can be used to execute the method embodiments of this application. For details not disclosed in the device embodiments of this application, please refer to the method embodiments of this application.
[0148] Figure 8 is a schematic diagram of a copywriting generation device based on a large model provided by an embodiment of this application. As Figure 8 shown, the device includes: An acquisition module 801, configured to acquire a base model in response to receiving a copywriting generation instruction, where the base model is a first pre-trained large language model.
[0149] A determination module 802, configured to determine a first reward function, where the first reward function is a rule-driven reward function, and the rule is determined based on at least one of copywriting content elements, copywriting style constraints, copywriting violation rules, and copywriting generation process constraints.
[0150] A training module 803, configured to perform fine-tuning training on the base model based on the first reward function using a reinforcement learning algorithm to obtain the base model after the first training.
[0151] The determination module 802 is further configured to call at least one reward model to determine a second reward function.
[0152] The training module 803 is further configured to weight-combine the first reward function and the second reward function to obtain a target reward function, and perform fine-tuning training on the base model after the first training based on the target reward function using a reinforcement learning algorithm to obtain the base model after the second training.
[0153] A generation module 804, configured to generate copywriting using the base model after the second training based on the copywriting generation instruction.
[0154] According to the technical solution provided in the embodiment of the present application, the copy generation base model composed of a large language model is adjusted and trained by using a reinforcement learning algorithm, so that the trained base model can generate more reliable and high-quality copy; wherein, when using the reinforcement learning algorithm to fine-tune the base model, the base model is first trained for the first time based on the first reward function of the rule-driven type, and then the second reward function is determined by using the reward model, and the first reward function and the second reward function are weightedly combined to obtain the target reward function, and then the base model is trained for the second time based on the target reward function, thereby achieving fine-grained optimization of the base model, so that the trained base model can generate expected copy and improve the user experience.
[0155] In some implementations, determining the first reward function includes: determining copy content element rules and copy style constraint rules based on copy generation instructions, and obtaining copy violation rules and copy generation process constraint rules; determining a rule subtask reward function based on each rule; and weightedly combining the rule subtask reward functions to obtain the first reward function.
[0156] In some implementations, a rule subtask reward function is determined based on each rule, including: determining a first rule subtask reward function based on the copy content element rule, the first rule subtask reward function rewards the generated copy that includes the content element, and punishes the generated copy that omits the content element; determining a second rule subtask reward function based on the copy style constraint rule, the second rule subtask reward function rewards the generated copy that conforms to the copy style, and punishes the generated copy that does not conform to the copy style; wherein the copy style is determined by the format of the copy, the style of the copy, and the language style of the copy at least one of the following is determined; based on the copy violation rules, a third rule subtask reward function is determined, the third rule subtask reward function rewards the generated copy that complies with the copy violation rules, and punishes the generated copy that does not comply with the copy violation rules; based on the copy generation process constraint rules, a fourth rule subtask reward function is determined, the fourth rule subtask reward function rewards the generated copy that meets the copy generation process constraint rules, and rewards the generated copy that does not meet the copy generation process constraint rules; wherein the copy generation process constraint rules include that the copy is generated in a format of thinking process first and then outputting content.
[0157] In some embodiments, calling at least one reward model to determine a second reward function includes: obtaining at least one copywriting evaluation index, where each copywriting evaluation index is used to evaluate a quality level of a generated copy, and the quality level includes at least the language fluency, creative uniqueness, and marketing effectiveness of the generated copy; determining a reward model corresponding to each evaluation index, where the reward model is a pre-trained reward model or a second pre-trained large language model, and the second pre-trained large language model is the same as or different from the first pre-trained large language model; determining the reward value of each reward model as a model sub-task reward function; and combining the model sub-task reward functions with weights to obtain the second reward function.
[0158] In some embodiments, it further includes: when using the reinforcement learning algorithm to fine-tune the base model after the first training based on the target reward function, in response to determining that the base model after the first training has not converged, periodically statistically calculating the difference between the quality level of the generated copy and the expected quality level; updating the weights of each reward model based on the difference; and combining the model sub-task reward functions with weights based on the updated weights to obtain the second reward function.
[0159] In some embodiments, periodically statistically calculating the difference between the quality level of the generated copy and the expected quality level includes: periodically obtaining a validation data set, where the validation data set includes different types of historical generated copywriting instructions and corresponding historical copies; using the base model after the first training that has not converged during the fine-tuning training to generate validation copies based on the historical generated copywriting instructions; and statistically calculating the difference between the quality level of the validation copies and the quality level of the historical copies to obtain the difference between the quality level of the generated copy and the expected quality level.
[0160] In some embodiments, using the reinforcement learning algorithm to fine-tune the base model or the base model after the first training includes: obtaining the currently available resources; in response to determining that the currently available resources are greater than or equal to the preset resource threshold, using the Proximal Policy Optimization (PPO) algorithm or the Group Relative Policy Optimization (GRPO) algorithm to fine-tune the base model or the base model after the first training; or in response to determining that the currently available resources are less than the preset resource threshold, using the Reinforcement Learning with Out-of-the-Box Optimization (RLOO) algorithm to fine-tune the base model or the base model after the first training.
[0161] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0162] Figure 9 is a schematic diagram of an electronic device provided by an embodiment of the present application. As Figure 9 As shown, the electronic device 9 of this embodiment includes: a processor 901, a memory 902, and a computer program 903 stored in the memory 902 and executable on the processor 901. When the processor 901 executes the computer program 903, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 901 executes the computer program 903, the functions of each module / unit in the above-mentioned device embodiments are implemented.
[0163] The electronic device 9 can be a desktop computer, a notebook, a palm computer, a cloud server, or other electronic devices. The electronic device 9 may include, but is not limited to, a processor 901 and a memory 902. Those skilled in the art can understand that Figure 9 merely examples of the electronic device 9, which do not constitute a limitation on the electronic device 9, and may include more or fewer components than shown in the figure, or different components.
[0164] The processor 901 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0165] The memory 902 can be an internal storage unit of the electronic device 9. For example, the hard disk or memory of the electronic device 9. The memory 902 can also be an external storage device of the electronic device 9. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 9. The memory 902 can also include both an internal storage unit and an external storage device of the electronic device 9. The memory 902 is used to store computer programs and other programs and data required by the electronic device.
[0166] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0167] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.< / think> < / answer> < / think> < / answer> < / think> < / answer> < / think> < / answer> < / think> < / think> < / answer> < / answer> < / answer> < / think> < / think> < / answer> < / answer> < / answer> < / think> < / think> < / think> < / answer> < / think> < / answer> < / think>
Claims
1. A copywriting generation method based on a large model, characterized in that, include: In response to receiving a copywriting generation instruction, obtaining a base model, wherein the base model is a first pre-trained large language model; Determine a first reward function, where the first reward function is a rule-driven reward function, and the rule is determined based on at least one of a text content element, a text style constraint, a text violation rule, and a text generation process constraint; Using a reinforcement learning algorithm to fine-tune the base model based on the first reward function to obtain a base model after the first training; Calling at least one reward model to determine a second reward function; A target reward function is obtained by weighted combination of the first reward function and the second reward function, and a reinforcement learning algorithm is used to fine-tune the base model after the first training based on the target reward function to obtain a base model after the second training; Generate copy based on the copy generation instruction using the base model trained for the second time.
2. The method according to claim 1, characterized in that, The determining of the first reward function comprises: Determine the copy content element rules and copy style constraint rules based on the copy generation instruction, and obtain the copy violation rules and copy generation process constraint rules; Determine a rule subtask reward function based on each rule; The reward functions of each rule subtask are weighted and combined to obtain the first reward function.
3. The method according to claim 2, wherein A rule subtask reward function is determined based on each rule, including: Determining a first rule subtask reward function based on the copy content element rule, wherein the first rule subtask reward function rewards a generated copy that includes the content element and penalizes a generated copy that omits the content element; Determining a second rule subtask reward function based on the copywriting style constraint rule, wherein the second rule subtask reward function rewards the generated copywriting that conforms to the copywriting style and punishes the generated copywriting that does not conform to the copywriting style; wherein the copywriting style is determined by at least one of the format of the copywriting, the language type of the copywriting, and the language style of the copywriting; Determine a third rule subtask reward function based on the copy violation rule, wherein the third rule subtask reward function rewards the generated copy that complies with the copy violation rule, and punishes the generated copy that does not comply with the copy violation rule; The fourth rule subtask reward function is determined based on the copy generation process constraint rules, and the fourth rule subtask reward function rewards the generated copy that meets the copy generation process constraint rules, and rewards the generated copy that does not meet the copy generation process constraint rules; wherein the copy generation process constraint rules include that the copy is generated in a format of thinking process first and then outputting content.
4. The method according to claim 1, wherein Calling at least one reward model to determine a second reward function includes: Obtaining at least one copywriting evaluation indicator, each copywriting evaluation indicator is used to evaluate a quality level of the generated copywriting, wherein the quality level at least includes language fluency, creative uniqueness, and marketing effectiveness of the generated copywriting; Determine a reward model corresponding to each evaluation indicator, where the reward model is a pre-trained reward model or a second pre-trained large language model, where the second pre-trained large language model is the same as or different from the first pre-trained large language model; Determine the reward value of each reward model as the model subtask reward function; Weightedly combine the model subtask reward functions to obtain the second reward function.
5. The method according to claim 4, wherein The method further includes: When using the reinforcement learning algorithm to fine-tune the base model after the first training based on the target reward function, in response to determining that the base model after the first training has not converged, periodically statistically calculate the difference between the generated copywriting quality level and the expected quality level; Update the weights of each reward model based on the difference; Weightedly combine the model subtask reward functions based on the updated weights to obtain the second reward function.
6. The method according to claim 5, characterized in that, The periodically statistically calculating the difference between the generated copywriting quality level and the expected quality level includes: Periodically obtain a validation data set, where the validation data set includes different types of historical generated copywriting instructions and corresponding historical copywriting; Use the base model after the first training that has not converged during the fine-tuning training to generate validation copywriting based on the historical generated copywriting instructions; Statistically calculate the difference between the quality level of the validation copywriting and the quality level of the historical copywriting to obtain the difference between the generated copywriting quality level and the expected quality level.
7. The method according to any one of claims 1 to 6, characterized in that Using the reinforcement learning algorithm to fine-tune the base model or the base model after the first training includes: Obtain the currently available resources; In response to determining that the currently available resources are greater than or equal to the preset resource threshold, use the Proximal Policy Optimization algorithm (PPO) or the Group Relative Policy Optimization algorithm (GRPO) to fine-tune the base model or the base model after the first training; or In response to determining that the currently available resources are less than the preset resource threshold, use the Reinforcement Style Optimization algorithm (RLOO) to fine-tune the base model or the base model after the first training.
8. A copywriting generation device based on a large model, characterized in that, Includes: An acquisition module, configured to obtain a base model in response to receiving a generated copywriting instruction, where the base model is a first pre-trained large language model; A determination module, configured to determine a first reward function, where the first reward function is a rule-driven reward function, and the rule is determined based on at least one of copywriting content elements, copywriting style constraints, copywriting violation rules, and copywriting generation process constraints; A training module, configured to use the reinforcement learning algorithm to fine-tune the base model based on the first reward function to obtain the base model after the first training; The determination module is further configured to call at least one reward model to determine a second reward function; The training module is further configured to weightedly combine the first reward function and the second reward function to obtain a target reward function, and use the reinforcement learning algorithm to fine-tune the base model after the first training based on the target reward function to obtain the base model after the second training; A generation module, configured to generate copywriting using the base model after the second training based on the generated copywriting instruction.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor implements the steps of the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that, The computer program implements the steps of the method according to any one of claims 1 to 7 when executed by the processor.
Citation Information
Patent Citations
Large language model training method for document writing, computer device and storage medium
CN118410776A
Reward model training method and device, electronic equipment and storage medium
CN118656607A
Ensemble learning-oriented question and answer method and device under large model fine tuning
CN118761459A
Man-machine reinforcement learning method based on multi-dimensional human feedback fusion
CN119005287A
Method and device for training language model based on reinforcement learning
CN119558428A
Cited By
Marketing document automatic generation method and device based on reinforcement learning and storage medium
CN120746646A
Report generation method and device based on reward mechanism, equipment and medium
CN120930625A
Methods, apparatus, equipment, and media for generating reports based on reward mechanisms
CN120930625B
Multimodal X-ray image diagnosis report generation method based on reinforcement learning optimization
CN121034518A
Academic paper automatic review method and device based on reinforcement learning
CN121213010A