Marketing document automatic generation method and device based on reinforcement learning and storage medium

By building slotted rewriting instructions and reinforcement learning training, the marketing copy generation model is optimized, the problems of compliance and multiple constraints are solved, and efficient, creative and compliant marketing copy generation is achieved.

CN120746646AActive Publication Date: 2025-10-03北京衔远有限公司 +1

Patent Information

Application Number
CN202511241555.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-10-03
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

When generating marketing copy, existing technologies find it difficult to simultaneously meet compliance, accurately integrate product information, and meet multiple constraint requirements. In addition, the feedback loop is not sound, which makes it easy for the generated content to touch regulatory red lines, distort facts, and lack output stability.

Method used

By obtaining product information and marketing points, slot-based rewriting instructions are constructed, and pre-generated language models are used to fill in information to generate preliminary copy. The copy generation process is optimized by training the model through supervised fine-tuning and reinforcement learning, combined with a regularized evaluation system.

Benefits of technology

It achieves the generation of marketing copy with high compliance and consistent multi-constraints, reduces labor costs, and generates creative marketing copy in batches that meet user needs and platform specifications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120746646A_ABST
    Figure CN120746646A_ABST
Patent Text Reader

Abstract

The invention provides an automatic marketing document generation method and device based on reinforcement learning and a storage medium. The method comprises the following steps: performing semantic matching retrieval on public copywriting data to obtain candidate copywriting; inputting the slot rewriting instruction into a pre-generated language model to generate a first marketing copywriting; performing supervision fine tuning training on a preset basic language model to obtain a first training model; inputting new user product information and promotion requirements to the first training model, generating a second marketing copywriting, scoring the second marketing copywriting and generating evaluation data; constructing a partial sequence training sample according to the evaluation data, taking the partial sequence training sample as a reward signal, and executing reinforcement learning training on the first training model to obtain a second training model; and calling the second training model in the copywriting generation system, and outputting the target marketing copywriting based on the product information and the promotion requirement of the user. According to the method and the device, high-compliance and multi-constraint consistent marketing copywriting batch generation can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, and storage medium for automatically generating marketing copy based on reinforcement learning. Background Art

[0002] With the rapid development of e-commerce, short videos, and social media platforms, companies' demand for product marketing copy has exploded. Compliant, accurate, and creative marketing copy can significantly increase product exposure and conversion rates. Consequently, various automatic copy generation technologies based on natural language processing (NLP) have garnered widespread attention. Large-scale pre-trained language models (LLMs), with their ability to generate general text and understand context, have become a core technology for automated content production.

[0003] Mainstream solutions typically pre-train LLMs using publicly available, large-scale corpora, then generate marketing text using a few-shot or prompt engineering approach. To improve performance in specific verticals, some research has introduced supervised fine-tuning (SFT) or reinforcement learning with human feedback (RLHF) to retrain the models, infusing them with specific industry knowledge and writing styles. Other approaches incorporate product attribute information into the generation process through techniques such as keyword insertion, template replacement, or retrieval-augmented generation (RAG), improving the correspondence between text and products.

[0004] Existing solutions still have the following problems: 1. Low rule adaptability: The existing LLM lacks explicit modeling of advertising laws, platform banned words and industry compliance requirements, and the generated content is prone to touching regulatory red lines. 2. Insufficient coupling of professional information: Structured information such as product parameters and domain terminology cannot be fully and accurately integrated into the generated copy, resulting in factual distortion or missing selling points. 3. Poor consistency of multiple constraints: When users simultaneously put forward multiple requirements such as style, length, selling points, etc., it is difficult for the model to take into account all constraints in a single generation, and the output stability is insufficient. 4. Incomplete feedback loop: Traditional SFT training data is limited in amount and of varying quality. The RLHF stage mostly relies on manual dialogue scoring and lacks a scalable and explainable evaluation rule library, which limits the effectiveness of reinforcement learning. Summary of the Invention

[0005] In view of this, the embodiments of the present application provide a method, device and storage medium for automatically generating marketing copy based on reinforcement learning to solve the problems of low compliance adaptability and difficulty in simultaneously integrating product information and multiple constraints in the existing technology.

[0006] In a first aspect of an embodiment of the present application, a method for automatically generating marketing copy based on reinforcement learning is provided, comprising: obtaining original input including product information, marketing key points and public copy data, performing semantic matching retrieval on the public copy data, and obtaining candidate copy related to the product information; constructing a slotted rewriting instruction including the candidate copy, product information and marketing key points, inputting the slotted rewriting instruction into a pre-generated language model, causing the pre-generated language model to fill in the corresponding product information and marketing key points in the preset slot position, generating a first marketing copy, and arranging the candidate copy, product information and the first marketing copy into a supervised fine-tuning training training samples; using the supervised fine-tuning training samples to perform supervised fine-tuning training on the preset basic language model to obtain a first training model; input new user product information and promotion requirements into the first training model to generate at least one second marketing copy, input the second marketing copy into the regularized evaluation system, score the second marketing copy and generate evaluation data; construct a partial order training sample based on the evaluation data, use the partial order training sample as a reward signal, perform reinforcement learning training on the first training model to obtain a second training model; call the second training model in the copy generation system, and output the target marketing copy based on the user's product information and promotion requirements.

[0007] According to a second aspect of the embodiment of the present application, a device for automatically generating marketing copy based on reinforcement learning is provided, comprising: an acquisition module for acquiring original input including product information, marketing key points and public copy data, performing semantic matching retrieval on the public copy data, and obtaining candidate copy related to the product information; a generation module for constructing slotted rewriting instructions including candidate copy, product information and marketing key points, inputting the slotted rewriting instructions into a pre-generated language model, causing the pre-generated language model to fill in the corresponding product information and marketing key points in the preset slot position, generate a first marketing copy, and organize the candidate copy, product information and the first marketing copy into supervised fine-tuning training samples; a first training module for A block is used to perform supervised fine-tuning training on a preset basic language model using supervised fine-tuning training samples to obtain a first training model; an evaluation module is used to input new user product information and promotion requirements into the first training model, generate at least one second marketing copy, input the second marketing copy into the regularized evaluation system, score the second marketing copy and generate evaluation data; a second training module is used to construct a partial order training sample based on the evaluation data, use the partial order training sample as a reward signal, perform reinforcement learning training on the first training model, and obtain a second training model; an output module is used to call the second training model in the copy generation system, and output the target marketing copy based on the user's product information and promotion requirements.

[0008] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the steps of the above method are implemented when the processor executes the computer program.

[0009] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.

[0010] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: By obtaining original input containing product information, marketing key points, and public copy data, semantic matching retrieval is performed on the public copy data to obtain candidate copy related to the product information; slot-based rewriting instructions containing the candidate copy, product information, and marketing key points are constructed, and the slot-based rewriting instructions are input into a pre-generated language model, so that the pre-generated language model fills the corresponding product information and marketing key points in the preset slot position to generate a first marketing copy, and the candidate copy, product information, and the first marketing copy are organized into supervised fine-tuning training samples; supervised fine-tuning training is performed on the preset basic language model using the supervised fine-tuning training samples to obtain a first training model; new user product information and promotion requirements are input into the first training model to generate at least one second marketing copy, and the second marketing copy is input into a regularized evaluation system to score the second marketing copy and generate evaluation data; partial order training samples are constructed based on the evaluation data, and the partial order training samples are used as reward signals to perform reinforcement learning training on the first training model to obtain a second training model; the second training model is called in the copy generation system to output a target marketing copy based on the user's product information and promotion requirements. This application can achieve batch generation of marketing copy with high compliance and consistent multiple constraints. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Figure 1 This is a flowchart of a method for automatically generating marketing copy based on reinforcement learning provided in an embodiment of the present application; Figure 2 Schematic diagram of the structure of the automatic generation device of marketing copy based on reinforcement learning provided in an embodiment of the present application; Figure 3 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0013] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0014] Product marketing copy not only incorporates numerous communication techniques, such as creating suspense and exaggerating to attract readers, but also has numerous special requirements based on industry regulations and advertising laws. Furthermore, marketing copy often includes a wealth of domain-specific terminology and information about the product's parameters and characteristics. However, existing large-scale text generation models lack an understanding of these relevant domains and special regulations. When given excessive input requirements, these models struggle to fully adhere to these rules and may even produce illusions of fact, failing to accurately represent product information.

[0015] In response to the shortcomings of existing solutions, the technical solution of this application proposes the following solutions: 1. Based on the user's proposed target product and promotional needs, marketing experts rewrite and reconstruct the original content based on publicly available internet text data and an open-source large language model. This rewriting process incorporates user product information, selling points, and relevant marketing considerations to generate creative marketing copy. Based on this generated creative copy data, a small-sample training template for supervised fine-tuning is constructed, fine-tuning training data is generated, and supervised fine-tuning of the small-scale language model is performed to optimize the model's output performance in specific marketing scenarios.

[0016] 2. Apply the fine-tuned language model to new user product needs. Marketing experts will evaluate the quality of the marketing copy automatically generated by the model, referencing the promotion platform's specifications and generating structured evaluation information. This evaluation information will be summarized using an open-source large language model. By integrating user needs with platform standards, a unified, rule-based evaluation system will be established, along with a systematic evaluation rule library.

[0017] 3. Further integrate the initial creative copywriting data with the small model's generation results and evaluation information, and generate the partially ordered training data required for reinforcement learning using the open-source large language model, forming a high-quality feedback dataset for training. Based on this partially ordered training data, reinforcement learning is performed on the fine-tuned language model, ultimately resulting in a creative marketing copywriting generation model with stable output quality and platform adaptability.

[0018] The marketing copy generated through the technical solution of this application is more in line with user needs and platform specifications; it reduces the labor cost of creative copywriting and generates diverse marketing copy in batches.

[0019] The contents of the technical solution of this application are described in detail below with reference to the accompanying drawings and specific embodiments.

[0020] Figure 1 This is a flow chart of the method for automatically generating marketing copy based on reinforcement learning provided in the embodiment of the present application. Figure 1 As shown, the method for automatically generating marketing copy based on reinforcement learning may specifically include: S101, obtaining original input including product information, marketing points, and public copy data, performing semantic matching search on the public copy data, and obtaining candidate copy related to the product information; S102: Constructing a slotted rewriting instruction containing candidate copywriting, product information, and marketing key points, inputting the slotted rewriting instruction into a pre-generated language model, causing the pre-generated language model to fill in the corresponding product information and marketing key points in the preset slot positions to generate a first marketing copywriting, and organizing the candidate copywriting, product information, and the first marketing copywriting into supervised fine-tuning training samples; S103, performing supervised fine-tuning training on a preset basic language model using the supervised fine-tuning training sample to obtain a first training model; S104: Input new user product information and promotion requirements into the first training model to generate at least one second marketing copy. Input the second marketing copy into the rule-based evaluation system to score the second marketing copy and generate evaluation data. S105, constructing partial order training samples based on the evaluation data, using the partial order training samples as reward signals, performing reinforcement learning training on the first training model to obtain a second training model; S106: Calling the second training model in the copy generation system to output target marketing copy based on the user's product information and promotion requirements.

[0021] In some embodiments, semantic matching retrieval is performed on public copy data to obtain candidate copy related to product information, including: Input product information and public copywriting data into the semantic encoding network respectively to obtain the corresponding vector representation; Calculate and sort the similarity of the vector representations of the public copywriting data based on a preset similarity metric; According to a preset number or threshold, the public copywriting with the highest similarity is selected from the sorting results as the candidate copywriting.

[0022] Specifically, in an example of semantic matching retrieval of public copywriting data, the system follows a three-stage process: "text vectorization - similarity calculation and ranking - and high-relevance candidate screening." This process is the front-end data preparation step in the automatic marketing copywriting chain, and its output directly determines the quality and compliance of subsequent rewriting.

[0023] First, the key technical features and implementation process are summarized. The system concatenates the function description, selling point keywords, and target audience description fields in the target product information into a continuous text; at the same time, it reads the public copy data set to be retrieved and performs the same text standardization processing on each copy. All texts are uniformly input into the semantic encoding network to obtain a fixed-length vector representation; then, the correlation between the product vector and the copy vector is calculated using a preset similarity metric (cosine similarity in this example) and a ranking result is generated; finally, highly relevant copy is extracted from the sorted list as candidate materials according to the preset number of candidates and the minimum relevance threshold.

[0024] The following is an explanation of the innovative technical features and important terminology involved in the embodiment. The semantic coding network is a shared-weight Transformer architecture text twin-tower model, which is trained through comparative learning so that semantically similar text pairs are closer in the vector space. Vector representation refers to the fixed-length floating-point vector obtained by mapping the text through the semantic coding network. This embodiment uses a 768-dimensional representation to take into account both expressive power and retrieval efficiency. The similarity metric is used to measure the degree of closeness between text vectors. The cosine similarity value range is 0 to 1. The larger the value, the higher the semantic relevance. The public copywriting dataset refers to the publicly available text materials from social media, e-commerce reviews and industry case libraries, which are used for matching retrieval after desensitization and segmented denoising.

[0025] In its implementation, the system first invokes a semantic encoding network for each product document and each publicly available copywriting text, generating a product vector Vp and a set of copywriting vectors {Vi}. This network is pre-trained on a training set containing millions of industry documents, and the encoder parameters are frozen during training to ensure consistency during inference. The system then uses batched matrix multiplication to parallelize the cosine similarity between Vp and {Vi}, generating a relevance list. To accelerate sorting and subsequent online calls, the relevance list, along with the copywriting vectors, is written into an approximate nearest neighbor index based on the HNSW structure, facilitating reuse of candidate copywriting upon new product requests. The system sets an upper limit of 40 candidates and a minimum relevance threshold of 0.70. The sorted results are iterated, and if cos(Vp,Vi) ≥ 0.70, the copywriting corresponding to Vi is added to the candidate set until the set size reaches 40 or the sorted list is exhausted. Selected copies, along with their similarity scores, copywriting source identifiers, and timestamps, are written to the source buffer for direct access by the slotted rewriting module. To ensure retrieval accuracy in multilingual scenarios, the semantic encoding network also incorporates cross-language alignment loss during the training phase, allowing Chinese and English texts to be compared in the same embedding space.

[0026] Through the above embodiments, the system can locate original materials that are highly relevant to the target product and have a good compliance basis among millions of public documents in milliseconds, significantly reducing the risk of semantic deviation in the rewriting stage and improving the response efficiency and accuracy of the overall generation link.

[0027] In some embodiments, constructing a slotted rewriting instruction containing candidate copywriting, product information, and marketing points includes: Perform text structure analysis on the candidate copy and determine at least one replaceable semantic segment; Writing a slot marker corresponding to product information or marketing points at each replaceable semantic segment to obtain a rewritten template with placeholders; Setting a filling rule for each slot tag, wherein the filling rule specifies a mapping relationship between the slot tag and a product information field or a marketing key point field; The rewriting template and the filling rules are assembled into a unified instruction format to generate slotted rewriting instructions.

[0028] Specifically, before rewriting the candidate text into slots, the system needs to complete four consecutive steps: "fragment recognition - slot labeling - rule binding - instruction encapsulation", so as to convert the original text into rewriting instructions that can be directly parsed by the generative model.

[0029] First, a summary of the key technical features is given. The system receives the candidate copy output by the semantic matching module, and performs text structure analysis on the copy, using technologies such as syntactic dependency, named entity recognition, and rhetorical pattern detection to identify semantic segments such as the title area, product description area, functional appeal area, emotional rendering area, and call to action area. Placeholder-style slot tags are inserted into each semantic segment that is judged to be replaceable. The slot tags correspond one-to-one to product information fields or marketing key points fields. Then, filling rules are set for each slot tag. The rules include field mapping relationships, optional synonym sets, length restrictions, and style constraints, etc., which are used to guide the language model to complete content filling in the decoding stage. Finally, the rewriting template with placeholders and the corresponding rules are assembled into a unified instruction format to provide input for subsequent language model rewriting.

[0030] In this embodiment, replaceable semantic segments refer to text segments that carry product attributes or marketing emotions in candidate copywriting, and can be quickly adapted to different products by replacing key components therein. A slot tag is a symbolic placeholder representation that prompts the language model to replace it with specified content during generation; typical styles include 〈PROD_NAME〉 or 〈BENEFIT_1〉. The filling rule is a set of constraints defined at the slot tag level, which clarifies the mapping between tags and product information fields (such as name, specifications, core functions) or marketing key points fields (such as promotional offers, brand tone), and requires the output to comply with character length, tone, and platform restrictions. The unified instruction format refers to encapsulating the rewritten template body, slot-field mapping table, and generation style parameters into a parseable input sequence to ensure that the model can obtain all constraints in one read.

[0031] The specific implementation process is as follows. The system first performs sentence and paragraph segmentation on the candidate copy, and combines the dependency tree and domain keyword dictionary to locate replaceable semantic fragments. For example, it detects that "XX essence" in the sentence "XX essence makes the skin brighter" is a placeholder for the product name, and "brighter" is a placeholder for the efficacy description. The system then inserts slot tags such as 〈PROD_NAME〉 and 〈BENEFIT〉 at these positions, and retains the original sentence skeleton to generate a rewrite template with placeholders. The system then reads the product information table and the marketing key points table, binds the "product name" field to 〈PROD_NAME〉, and binds the "skin brightening efficacy" field to 〈BENEFIT〉, and limits the corresponding output of 〈BENEFIT〉 to use positive emotional adjectives and a length of no more than ten Chinese characters in the rules. If the copy needs to meet the platform's restrictions on title length, add a character limit constraint to the 〈TITLE〉 slot rule. The system places the rewriting template text at the beginning of the instruction, the slot-field mapping and various rules at the end of the instruction, connects them through delimiters to form the overall input, and finally generates the slotted rewriting instruction.

[0032] Through the above embodiments, the system maintains the overall rhetorical structure of the candidate copy while explicitly embedding product information and marketing points in a slot-rule manner, providing a clear and executable filling guide for subsequent language model rewriting, thereby significantly improving the correspondence between content and product information in the rewriting stage and reducing compliance risks.

[0033] In some embodiments, the slotting rewriting instruction is input into the pre-generated language model, so that the pre-generated language model fills the preset slot position with the corresponding product information and marketing points to generate the first marketing copy, including: Perform text encoding on the slotted rewriting instructions to obtain a semantic vector containing slot tags and mapping relationships; When calling the pre-generated language model for conditional text generation, the slot markers are replaced in real time during the decoding phase according to the mapping relationship, and the corresponding product information or marketing points are written into the generated sequence; At the end of the generation, the first marketing copy containing complete product information and marketing points is output, and identification metadata corresponding to the copy style is attached.

[0034] Specifically, after the slot-based rewriting phase is completed, the system uses the generated rewritten instructions as the entire text input to pre-generate a language model. The model first performs text encoding on the entire instruction, using its internal embedding layer and multi-layer Transformer self-attention structure to map the character sequence into a semantic vector, where the slot marker itself is embedded as a unique placeholder vector, and the "slot-field mapping" explicitly stated at the end of the instruction is parsed into a set of key-value pairs and also converted into vector form. During the encoding phase, the model soft-aligns the placeholder vector with the corresponding field vector, so that the correct fill value can be retrieved during subsequent decoding.

[0035] Furthermore, the generation process adopts a conditional text generation strategy: the model uses the existing context in the rewritten template as the decoding starting point, and the dynamic replacement mechanism is triggered at each slot marker position. The dynamic replacement mechanism reads the mapping relationship through a table lookup, writes the pre-bound product information field or marketing key points field into the current generation sequence, and performs post-processing according to the filling rules, such as applying synonym replacement or emotional color transformation to meet character length and tone requirements. If the slot marker is located in the title or call to action fragment, the model will refer to the domain parameters in the style instructions to adjust the sentence compactness and punctuation. Temperature control and penalty repetition strategies are adopted throughout the generation process to ensure that the output is both diverse and avoids information redundancy.

[0036] Furthermore, when all tokens are replaced and the sentence generation termination condition is triggered, the system outputs the complete first marketing copy and appends metadata. This metadata includes the copy style label, generation timestamp, mapping rule version number used, and the source field for each slot filling, facilitating subsequent quality review and tracking. The copy and metadata are written to a cache area for direct access by the supervised fine-tuning training sample construction module.

[0037] Through the process design of this embodiment, the pre-generated language model accurately embeds product information and marketing points while ensuring sentence coherence, realizes efficient parsing and content filling of slotted rewriting instructions, and provides high-quality training samples with clear structure and consistent semantics for subsequent supervision and fine-tuning.

[0038] In some embodiments, supervised fine-tuning training is performed on a preset basic language model using supervised fine-tuning training samples to obtain a first training model, including: Format the supervised fine-tuning training samples according to the unified prompt template and label specification to generate a dataset that conforms to the input-output pairs; Perform quality screening on the dataset to remove samples that do not meet the semantic completeness threshold or slot filling completeness threshold; Divide the filtered data into a training set and a validation set, and perform parameter update training on the preset basic language model based on the training set; During the training process, the validation set is used to monitor the loss value in real time and the training is terminated according to the early stopping condition to obtain the first training model.

[0039] Specifically, before executing the supervised fine-tuning phase, the system has already obtained raw samples containing slotted rewriting instructions, candidate copy, and the first marketing copy. This phase first formats these samples into a unified template, then performs a double quality screening. Finally, the parameters of the pre-set basic language model are fine-tuned using the qualified data, ultimately resulting in a first training model that can produce output tailored to the marketing scenario.

[0040] In some examples, to ensure that the model can accurately parse the context and execute slot constraints, the system defines a "prompt template" for each sample. The template consists of three sections: the first section is the key-value pair description of the product information and marketing key points fields; the second section is the slot-based rewriting instruction body; and the third section is the generated result placeholder identifier "<GEN>". At the same time, a "label specification" is attached to the sample. The label field contains the style identifier, slot mapping version number, source channel identifier and timestamp, which makes it easier for the model to perceive the domain and the age of the data during training. The formatting process is completed by a pipeline script. The script reads the fields in sequence, concatenates the key-value pairs according to the preset delimiter, and adds the "<END>" marker at the end of the text to ensure that the model can distinguish the input and output boundaries in a single forward process.

[0041] Furthermore, quality screening is divided into two levels: semantic completeness detection and slot filling completeness detection. The semantic completeness detection uses the dual-tower retrieval model to calculate the relevance score between the generated copy and the source instruction. If the score is lower than the threshold of 0.85, the sample is judged to have missing content and is eliminated. The slot filling completeness detection uses regular matching and rule base to confirm one by one whether each slot in the generated copy has been replaced by legal text. If there is a situation where it is not replaced or still contains banned terms after replacement, the sample is also removed. After double screening, the coverage of the retained samples is maintained at an average of about 78% of the total original sample volume, ensuring that the training signal is both diverse and reliable.

[0042] In some examples, qualified samples were randomly split into training and validation sets in a ratio of nine to one. The training set was fed into a pre-set basic language model with approximately seven billion parameters, pre-trained on a general corpus. To reduce GPU memory usage and quickly adapt to specific scenarios, the system inserted two layers of attention weights based on LoRA low-rank adaptation into the model, freezing the original parameters and training only the LoRA weights and output layer. The optimizer selected was AdamW, with an initial learning rate of 3e-5, a linearly decreasing schedule, and a learning rate warmup for the first ten percent of the steps. Due to GPU memory limitations, the training batch size was set to sixteen, and gradient accumulation was used to increase the effective batch size to one hundred and twenty-eight to improve gradient estimation stability.

[0043] In some examples, the average cross-entropy loss, BLEU, and ROUGE metrics are calculated on the validation set after every thousand gradient updates during training. The "slot constraint compliance rate"—the percentage of validation set examples where all slots are correctly filled—is also recorded. The system uses an early stopping strategy: if the validation loss decreases by less than 0.01 for three consecutive times and the slot constraint compliance rate does not significantly improve, training is terminated and rolled back to the model snapshot with the best validation metrics. This snapshot becomes the first trained model. Model weights, optimizer state, vocabulary, and LoRA configuration are archived, and the training set version number and filtering threshold are recorded in the metadata to ensure repeatability of subsequent experiments.

[0044] Through the above embodiment, the preset basic language model explicitly learns the slot filling rules and marketing domain characteristics while maintaining the advantages of grammar and fluency, so that the generated results can accurately embed product information and meet the platform compliance requirements, providing a basic model with high semantic consistency and strong constraint compliance for the subsequent partial order reinforcement learning stage.

[0045] In some embodiments, the second marketing copy is input into a rule-based evaluation system, and the second marketing copy is scored and evaluation data is generated, including: Establish an evaluation rule library that includes compliance rules, content quality rules, and user preference weight templates; Perform text parsing on the second marketing copy to generate corresponding feature representations; Based on the evaluation rule library, the feature representation is matched with rules to calculate the compliance dimension score, content quality dimension score and user preference dimension score respectively; The scores of each dimension are weighted and integrated according to the user preference weight template to obtain the comprehensive score of the second marketing copy; Generate structured evaluation data including the second marketing copy, scores of each dimension, comprehensive scores and trigger rule identifiers.

[0046] Specifically, before the second marketing copy is input into the rule-based evaluation system, the system first pre-builds a three-layer evaluation rule library of "compliance rules-content quality rules-user preference templates". Among them, compliance rules focus on mandatory restrictions such as banned words in advertising laws, medical claims, and numerical exaggeration; content quality rules cover textual indicators such as language fluency, selling point coverage, and sentence diversity; user preference templates record the target audience's relative weights on dimensions such as emotional tone, length, and style labels. The above rules are all stored in the form of callable regular expressions, keyword Trie trees, or interpretable machine learning classifiers, and are accompanied by deduction or addition parameters in the rule nodes for subsequent scoring calculations.

[0047] Furthermore, after receiving the second marketing copy, the system first performs sentence segmentation, part-of-speech tagging, and named entity recognition through the text parsing component, breaking the copy into a title segment, a body segment, and a call-to-action segment. Features such as keywords, sentiment polarity, and numerical quantifiers are then extracted, ultimately constructing a multidimensional feature representation vector. The rule matching engine then iterates through each compliance rule node. If any illegal terms or exaggerated descriptions are detected, the triggering rule flag is recorded and the compliance score is deducted according to the node settings. If no risk rules are triggered, the full score is retained based on the baseline score. The content quality rule component utilizes a pre-trained language quality assessment model to calculate syntactic completeness, repetition rate, and selling point coverage, and then maps these weights to a percentage score. The user preference score assesses the copy's emotional tone, length, and style tags based on the weights within the template. For example, if a user prefers a lively tone but the copy is labeled as formal, the corresponding score is deducted.

[0048] Furthermore, after the scores of the three dimensions are calculated, the score fusion module reads the user preference weight template corresponding to the current scenario, and generates a comprehensive score by weighted average of the compliance score, content quality score and user preference score. For example, the weight of the compliance dimension can be fixed at 0.5 to ensure the bottom line, and the weights of the content quality and user preference dimensions are 0.3 and 0.2 respectively. The weight template can be dynamically adjusted to adapt to situations such as holiday promotions or new product releases. After the comprehensive score is generated, the system encapsulates the second marketing copy, the scores of each dimension, the comprehensive score and the triggered rule ID into a structured JSON and stores it in the evaluation data warehouse to provide a data source for the subsequent partial order sample construction.

[0049] Through the process of this embodiment, the system can automatically and explainably quantify the performance of each marketing copy in terms of compliance, text quality and personalized adaptation, and output it in a unified structured format, so that the subsequent reinforcement learning stage can directly reference these high-confidence scores as partial order labels, thereby effectively improving the model's ability to respond to the linkage between compliance bottom lines and user preferences.

[0050] In some embodiments, constructing partial order training samples based on evaluation data, using the partial order training samples as reward signals, and performing reinforcement learning training on the first training model to obtain a second training model includes: Perform sample selection on the evaluation data and extract at least two candidate copywritings from the second marketing copywriting according to a preset sampling strategy; Compare the comprehensive scores of candidate copywriting, determine the copywriting with higher scores as the preferred copywriting, and the copywriting with lower scores as the second-choice copywriting, and generate partial order labels based on the comparison results; The preferred copy, second-choice copy, partial-order label, and input context corresponding to the same input context are packaged into partial-order training samples; A reinforcement learning environment is constructed using partially ordered training samples, with the first trained model as the policy network, and the degree of consistency between the generated results of the policy network and the partially ordered labels is used as the immediate reward; A reinforcement learning algorithm based on policy optimization is used to update the parameters of the first training model and constrain the distribution shift during the update process to obtain the second training model.

[0051] Specifically, in an embodiment of performing partial order reinforcement learning based on evaluation data, the system first selects samples of the generated second marketing copy and its structured evaluation data, and then constructs partial order training items in a "preferred-secondary" manner. Subsequently, the first training model is used as the policy network to build a reinforcement learning environment, and the model is guided by the reward signal to learn the output sorting preference, and finally the second training model is obtained.

[0052] To ensure sample representativeness, the system uses a stratified sampling strategy during the sample selection phase: all evaluation data is first divided into three tiers: high, medium, and low, based on the comprehensive score range. Within each tier, a number of samples are randomly selected to form a candidate set. For the same input context (i.e., the same product information and marketing points), the system extracts at least two copywriting documents from the candidate set with different scores, with a difference of at least ten points, as comparison objects. The system compares the comprehensive scores of the two copywriting documents, marking the one with the higher score as the preferred copywriting and the one with the lower score as the second-choice copywriting. It also generates a partial order label of "1" to indicate that the preferred copywriting outperforms the second-choice copywriting. If the difference in the comprehensive scores between the two copywriting documents is too small, the pair is discarded to avoid reward signal noise. The system then encapsulates the input context, the preferred copywriting document, the second-choice copywriting document, and the partial order label into a single partial order training sample and writes it into the partial order dataset in JSONL format. Each sample also includes a score version number and a rule base snapshot ID to ensure traceability of subsequent training results.

[0053] When constructing a reinforcement learning environment, the system uses the first trained model as the policy network and a fixed evaluation rule base as the external reward model. Each time the policy network generates two candidate texts based on a given input context, the system calls the rule base to calculate their combined scores and uses this to determine their superiority. A positive reward is given if the model-generated ranking is consistent with the partially ordered labels; otherwise, a negative reward is given. The reward function is designed as r = α·(consistency score)-β·KL, where the consistency score is 1 for consistency and -1 for inconsistency. The KL term measures the divergence of the current policy distribution relative to the distribution used when the first trained model was initialized, and β is a regularization coefficient used to limit excessive deviation from the original language fluency. The system uses the PPO algorithm for policy optimization. Each update round includes multiple partially ordered samples, leveraging the mini-batch advantage to estimate gradients and perform weight updates. Gradient clipping and linear learning rate decay are also enabled to prevent training instability. Immediately after the policy update, the consistency rate and average KL divergence are evaluated on a retained validation partial-order set. If the consistency rate does not improve or the KL divergence exceeds a threshold for three consecutive rounds, training is terminated early and rolled back to the best model snapshot.

[0054] Through this process, the partially ordered training samples provide the model with interpretable relative preference signals. The policy optimization algorithm, driven by both rewards and distributional constraints, guides the model to adjust its generated probability distribution, favoring it to produce marketing copy with higher overall scores while maintaining language quality. Implementation results show that the second-trained model, after undergoing this reinforcement learning phase, achieved an average overall score improvement of approximately 15 points in offline evaluations, and significantly higher click-through rates and compliance pass rates in online A / B tests than the first-trained model.

[0055] This embodiment adopts a multi-stage collaborative optimization method of "fine-tuning + regularized scoring + partial order reinforcement learning". By constructing a fine-tuning template and introducing a regularized evaluation system, the small model has higher generation consistency and adaptability in specific marketing scenarios; using partial order training data for reinforcement learning can effectively optimize the model output sorting ability, thereby giving priority to outputting higher quality candidates when generating multiple copywriting candidates; the multiple rounds of data construction and human-computer collaborative evaluation process in this embodiment enhance the controllability and interpretability of training data, and effectively avoid data noise and overfitting problems that are prone to occur in traditional model fine-tuning; the model training process is compatible with mainstream open source language model architectures, and has strong system scalability and platform deployment friendliness.

[0056] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0057] Figure 2 This is a schematic diagram of the structure of the automatic generation device of marketing copy based on reinforcement learning provided by the embodiment of the present application. Figure 2As shown, the automatic generation device for marketing copy based on reinforcement learning includes: The acquisition module 201 is used to obtain the original input including product information, marketing points and public copy data, perform semantic matching search on the public copy data, and obtain candidate copy related to the product information; Generation module 202 is configured to construct a slotted rewriting instruction containing candidate copy, product information, and marketing key points, input the slotted rewriting instruction into a pre-generated language model, and cause the pre-generated language model to fill in the corresponding product information and marketing key points in the preset slot positions to generate a first marketing copy. The candidate copy, product information, and first marketing copy are then organized into supervised fine-tuning training samples. A first training module 203 is configured to perform supervised fine-tuning training on a preset basic language model using supervised fine-tuning training samples to obtain a first training model; Evaluation module 204, configured to input new user product information and promotion requirements into the first training model, generate at least one second marketing copy, input the second marketing copy into the rule-based evaluation system, score the second marketing copy, and generate evaluation data; The second training module 205 is configured to construct a partial order training sample based on the evaluation data, use the partial order training sample as a reward signal, and perform reinforcement learning training on the first training model to obtain a second training model; The output module 206 is used to call the second training model in the copy generation system to output the target marketing copy based on the user's product information and promotion requirements.

[0058] In some embodiments, Figure 2 The acquisition module 201 inputs the product information and the public copy data into the semantic coding network respectively to obtain the corresponding vector representation; calculates and sorts the vector representation of the public copy data based on a preset similarity measure; and selects the public copy with the highest similarity from the sorting results as the candidate copy according to a preset number or threshold.

[0059] In some embodiments, Figure 2 The generation module 202 performs text structure analysis on the candidate copy to determine at least one replaceable semantic segment; writes a slot tag corresponding to the product information or marketing point at each replaceable semantic segment to obtain a rewriting template with a placeholder; sets a filling rule for each slot tag, and the filling rule specifies the mapping relationship between the slot tag and the product information field or the marketing point field; assembles the rewriting template and the filling rule into a unified instruction format to generate a slotted rewriting instruction.

[0060] In some embodiments, Figure 2The generation module 202 performs text encoding on the slot rewriting instruction to obtain a semantic vector containing slot tags and mapping relationships; when calling the pre-generated language model for conditional text generation, the slot tags are replaced in real time during the decoding stage according to the mapping relationship, and the corresponding product information or marketing points are written into the generation sequence; at the end of the generation, the first marketing copy containing complete product information and marketing points is output, and identification metadata corresponding to the copy style is attached.

[0061] In some embodiments, Figure 2 The first training module 203 formats the supervised fine-tuning training samples according to the unified prompt template and label specification to generate a data set that conforms to the input-output pairs; performs quality screening on the data set to eliminate samples that do not meet the semantic completeness threshold or the slot filling completeness threshold; divides the screened data into a training set and a validation set, and performs parameter update training on the preset basic language model based on the training set; uses the validation set to monitor the loss value in real time during the training process and terminates the training according to the early stopping condition to obtain the first training model.

[0062] In some embodiments, Figure 2 The evaluation module 204 establishes an evaluation rule library including compliance rules, content quality rules and user preference weight templates; performs text parsing on the second marketing copy to generate corresponding feature representations; performs rule matching on the feature representations based on the evaluation rule library, and calculates the compliance dimension score, content quality dimension score and user preference dimension score respectively; performs weighted fusion of the scores of each dimension according to the user preference weight template to obtain a comprehensive score of the second marketing copy; generates structured evaluation data including the second marketing copy, scores of each dimension, comprehensive score and trigger rule identifier.

[0063] In some embodiments, Figure 2 The second training module 205 performs sample selection on the evaluation data, extracts at least two candidate copywritings from the second marketing copywriting according to a preset sampling strategy; compares the comprehensive scores of the candidate copywritings, determines the copywriting with a higher score as the preferred copywriting, and the copywriting with a lower score as the second-selected copywriting, and generates a partial order label based on the comparison result; packages the preferred copywriting, the second-selected copywriting, the partial order label and the input context corresponding to the same input context into a partial order training sample; uses the partial order training sample to construct a reinforcement learning environment, uses the first training model as the policy network, and uses the consistency between the generated result and the partial order label of the policy network as an immediate reward; uses the reinforcement learning algorithm based on policy optimization to update the parameters of the first training model, constrains the distribution offset during the update process, and obtains the second training model.

[0064] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0065] Figure 3 Schematic diagram of the structure of the electronic device 3 provided in the embodiment of the present application. Figure 3 As shown, the electronic device 3 of this embodiment includes: a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, the steps of the above-mentioned method embodiments are implemented. Alternatively, when the processor 301 executes the computer program 303, the functions of the modules / units in the above-mentioned device embodiments are implemented.

[0066] For example, computer program 303 may be divided into one or more modules / units, which are stored in memory 302 and executed by processor 301 to implement the present application. One or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of computer program 303 in electronic device 3.

[0067] The electronic device 3 may be a desktop computer, a notebook, a PDA, a cloud server or other electronic device. The electronic device 3 may include but is not limited to a processor 301 and a memory 302. Those skilled in the art will understand that Figure 3 It is only an example of electronic device 3 and does not constitute a limitation of electronic device 3. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0068] The processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0069] Memory 302 can be an internal storage unit of electronic device 3, such as a hard drive or memory of electronic device 3. Memory 302 can also be an external storage device of electronic device 3, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, memory 302 can include both an internal storage unit of electronic device 3 and an external storage device. Memory 302 is used to store computer programs and other programs and data required by the electronic device. Memory 302 can also be used to temporarily store data that has been output or is about to be output.

[0070] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0071] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0072] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0073] In the embodiments provided in this application, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely schematic. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods. Multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection of the apparatus or unit, which may be electrical, mechanical or other forms.

[0074] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0075] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0076] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application can implement all or part of the processes in the above-mentioned embodiment method by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program may include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. Computer-readable media may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium.

[0077] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the technical solutions of the present application are described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for automatically generating marketing copy based on reinforcement learning, characterized in that: include: Obtaining original input including product information, marketing points, and public copywriting data, performing semantic matching search on the public copywriting data, and obtaining candidate copywriting related to the product information; Constructing a slotted rewriting instruction including the candidate copy, the product information, and the marketing key points, inputting the slotted rewriting instruction into a pre-generated language model, causing the pre-generated language model to fill in the corresponding product information and marketing key points in the preset slot positions to generate a first marketing copy, and organizing the candidate copy, the product information, and the first marketing copy into a supervised fine-tuning training sample; Performing supervised fine-tuning training on a preset basic language model using the supervised fine-tuning training samples to obtain a first training model; Inputting new user product information and promotion requirements into the first training model to generate at least one second marketing copy, inputting the second marketing copy into the rule-based evaluation system, scoring the second marketing copy and generating evaluation data; constructing partial order training samples based on the evaluation data, using the partial order training samples as reward signals, and performing reinforcement learning training on the first training model to obtain a second training model; The second training model is called in the copy generation system to output target marketing copy based on the user's product information and promotion requirements.

2. The method according to claim 1, characterized in that The performing semantic matching search on the public copywriting data to obtain candidate copywriting related to the product information includes: Input the product information and the public copy data into a semantic encoding network respectively to obtain corresponding vector representations; Calculating and sorting the similarity of the vector representations of the public copywriting data based on a preset similarity metric; According to a preset number or threshold, the public copy with the highest similarity is selected from the sorting results as the candidate copy.

3. The method according to claim 1, characterized in that The constructing of slotted rewriting instructions including the candidate copy, the product information, and the marketing points includes: Performing text structure analysis on the candidate copy to determine at least one replaceable semantic segment; Writing a slot mark corresponding to the product information or the marketing point at each replaceable semantic segment to obtain a rewriting template with a placeholder; Setting a filling rule for each slot tag, wherein the filling rule specifies a mapping relationship between the slot tag and a product information field or a marketing key point field; The rewriting template and the filling rule are assembled into a unified instruction format to generate the slotted rewriting instruction.

4. The method according to claim 1, wherein Inputting the slotting rewriting instruction into a pre-generated language model, so that the pre-generated language model fills in the corresponding product information and marketing points in the preset slot position to generate the first marketing copy, includes: Performing text encoding on the slotted rewriting instruction to obtain a semantic vector including slot tags and mapping relationships; When calling the pre-generated language model to generate conditional text, the slot mark is replaced in real time during the decoding phase according to the mapping relationship, and the corresponding product information or marketing points are written into the generated sequence; At the end of the generation, the first marketing copy containing complete product information and marketing points is output, and identification metadata corresponding to the copy style is attached.

5. The method according to claim 1, wherein The method of performing supervised fine-tuning training on a preset basic language model using the supervised fine-tuning training sample to obtain a first training model includes: Formatting the supervised fine-tuning training samples according to a unified prompt template and label specification to generate a data set that conforms to the input-output pairs; Performing quality screening on the dataset to remove samples that do not meet a semantic completeness threshold or a slot filling completeness threshold; Dividing the filtered data into a training set and a validation set, and performing parameter update training on the preset basic language model based on the training set; During the training process, the validation set is used to monitor the loss value in real time and the training is terminated according to the early stopping condition to obtain the first training model.

6. The method according to claim 1, characterized in that The step of inputting the second marketing copy into a regularized evaluation system, scoring the second marketing copy and generating evaluation data includes: Establish an evaluation rule library that includes compliance rules, content quality rules, and user preference weight templates; Performing text parsing on the second marketing copy to generate corresponding feature representations; Perform rule matching on the feature representation based on the evaluation rule library, and calculate compliance dimension scores, content quality dimension scores, and user preference dimension scores respectively; Performing weighted fusion on the scores of each dimension according to the user preference weight template to obtain a comprehensive score for the second marketing copy; Generate structured evaluation data including the second marketing copy, scores of each dimension, comprehensive score and trigger rule identifier.

7. The method according to claim 1, characterized in that The step of constructing a partial order training sample based on the evaluation data, using the partial order training sample as a reward signal, and performing reinforcement learning training on the first training model to obtain a second training model includes: Performing sample selection on the evaluation data, and extracting at least two candidate copywritings from the second marketing copywriting according to a preset sampling strategy; Comparing the comprehensive scores of the candidate documents, determining the document with the higher score as the preferred document and the document with the lower score as the second-selected document, and generating a partial order label based on the comparison results; Packing the preferred copy, the second-selected copy, the partial-order labels and the input context corresponding to the same input context into a partial-order training sample; Using the partially ordered training samples to construct a reinforcement learning environment, using the first training model as a policy network, and using the consistency between the generated results of the policy network and the partially ordered labels as an immediate reward; A reinforcement learning algorithm based on policy optimization is used to update the parameters of the first training model, and the distribution shift during the constraint update process is constrained to obtain the second training model.

8. A device for automatically generating marketing copy based on reinforcement learning, characterized in that: include: An acquisition module is used to obtain original input including product information, marketing points and public copy data, perform semantic matching search on the public copy data, and obtain candidate copy related to the product information; a generation module configured to construct a slotted rewriting instruction including the candidate copy, the product information, and the marketing key points, input the slotted rewriting instruction into a pre-generated language model, cause the pre-generated language model to fill in the corresponding product information and marketing key points in the preset slot positions, generate a first marketing copy, and organize the candidate copy, the product information, and the first marketing copy into a supervised fine-tuning training sample; A first training module is configured to perform supervised fine-tuning training on a preset basic language model using the supervised fine-tuning training samples to obtain a first training model; an evaluation module, configured to input new user product information and promotion requirements into the first training model, generate at least one second marketing copy, input the second marketing copy into a rule-based evaluation system, score the second marketing copy, and generate evaluation data; a second training module, configured to construct a partial order training sample based on the evaluation data, use the partial order training sample as a reward signal, perform reinforcement learning training on the first training model, and obtain a second training model; The output module is used to call the second training model in the copy generation system to output the target marketing copy based on the user's product information and promotion requirements.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Cboth generation method and device, electronic equipment and storage medium

    CN117313670A

  • Method and device for generating recommended copywriting for multimedia object

    CN118410232A

  • Marketing copywriting generation method and system

    CN119203941A

  • Cboth case generation method and device based on large model, equipment and medium

    CN120257948A

  • Copywriting generation method and apparatus, and storage medium

    WO2024022066A1

Cited By

  • Model training method and device, copywriting generation method and device, equipment and medium

    CN121659026A

  • Large model-based sequence generation scoring method, device, equipment and medium

    CN122490484A