Cboth case generation method and device, equipment, medium and program product
By supervising the fine-tuning of training sample pairs and generating multi-turn dialogues on large-scale generative language models, the problems of long processing time and low quality of copywriting generation tools are solved, and efficient and stable copywriting generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2024-11-01
- Publication Date
- 2026-05-08
AI Technical Summary
Existing copywriting generation tools are time-consuming and have unpredictable stability, resulting in low copywriting quality.
Fine-tuning is performed on a large-scale generative language model, and supervised fine-tuning is carried out using training sample pairs to generate a copywriting generation model. Sample copywriting is generated using a multi-turn dialogue approach, and training sample pairs are constructed to improve copywriting quality.
It reduces the reasoning time required for copy generation and improves the quality and stability of copy generation.
Smart Images

Figure CN121998068A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to computer technology, and more particularly to a document generation method, apparatus, device, medium, and program product. Background Technology
[0002] With the development of computer technology, copywriting generation tools are being used by an increasing number of users to generate copy for target audiences, including products.
[0003] Current copywriting generation tools suffer from lengthy reasoning processes when performing copywriting generation tasks, impacting efficiency. Furthermore, the unpredictable stability of these tools leads to low-quality generated copy that fails to meet user expectations. Summary of the Invention
[0004] This disclosure provides a copywriting generation method, apparatus, device, medium, and program product, which can improve copywriting generation efficiency and copywriting quality.
[0005] In a first aspect, embodiments of this disclosure provide a text generation method, including:
[0006] Obtain object information and text prompts for the target object, wherein the object information represents the text information corresponding to the multimodal content of the target object, and the text prompts represent the text description information required by the text generation model to generate the target text;
[0007] The object information and the text prompt text are input into the text generation model to obtain the target text output by the text generation model. The text generation model is obtained by fine-tuning a large generative language model based on training sample pairs. The training sample pairs include sample description text and sample text. The sample description text includes sample object information and sample text prompt text. The sample text is a recommended text generated based on the sample object information and at least two sample text prompt texts after at least two rounds of dialogue.
[0008] Secondly, this disclosure also provides a document generation apparatus, which includes:
[0009] The information acquisition module is used to acquire object information and text prompts of the target object, wherein the object information represents the text information corresponding to the multimodal content of the target object, and the text prompts represent the text description information required by the text generation model to generate the target text.
[0010] The copy generation module is used to input the object information and the copy prompt text into the copy generation model to obtain the target copy output by the copy generation model. The copy generation model is obtained by fine-tuning a large generative language model based on training sample pairs. The training sample pairs include sample description text and sample copy. The sample description text includes sample object information and sample copy prompt text. The sample copy is a recommended copy generated based on the sample object information and at least two sample copy prompt texts through at least two rounds of dialogue.
[0011] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0012] One or more processors;
[0013] Storage device for storing one or more programs.
[0014] When the one or more programs are executed by the one or more processors, the one or more processors implement the text generation method as described in any embodiment of this disclosure.
[0015] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the document generation method as described in any embodiment of this disclosure.
[0016] This disclosure provides a method, apparatus, device, medium, and program product for generating copywriting. By acquiring object information and copywriting prompt text of a target object, and inputting the object information and prompt text into a copywriting generation model, the target copywriting output by the copywriting generation model is obtained. Since sample copywriting is generated through a multi-turn dialogue using a large-scale generative language model based on sample object information and sample copywriting prompt text, the quality of the sample copywriting can be improved, avoiding the problem of unsatisfactory copywriting quality when directly generated by a large-scale generative language model. Sample copywriting and sample descriptive text are used to form training sample pairs. Then, supervised fine-tuning of the large-scale generative language model is performed using the training sample pairs to obtain a smaller-scale copywriting generation model. Using the copywriting generation model to perform the copywriting generation task can reduce inference time and improve the quality of copywriting generation. Attached Figure Description
[0017] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0018] Figure 1A flowchart illustrating a text generation method provided in an embodiment of this disclosure;
[0019] Figure 2 A flowchart illustrating a training method for a copywriting generation model provided in an embodiment of this disclosure;
[0020] Figure 3 A flowchart illustrating another method for training a copywriting generation model provided in this embodiment of the present disclosure;
[0021] Figure 4 This is a schematic diagram of the structure of a document generation device provided in an embodiment of the present disclosure;
[0022] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0023] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0024] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0025] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0026] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0027] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0028] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0029] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0030] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0031] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0032] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0033] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0034] Figure 1 This is a flowchart illustrating a copywriting generation method provided in an embodiment of this disclosure. This embodiment is applicable to situations requiring automatic copywriting generation, such as generating marketing copy. The method can be executed by a copywriting generation device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, PC, or server.
[0035] like Figure 1 As shown, the method includes:
[0036] S110. Obtain the object information and text prompts of the target object.
[0037] Wherein, the object information represents the text information corresponding to the multimodal content of the target object, and the copy prompt text represents the copy description information required by the copy generation model to generate the target copy.
[0038] In this embodiment of the disclosure, the target object can be an object in the marketing industry from which copy is to be generated. For example, the target object includes products such as cosmetics, snacks, electronic products, clothing, or accessories. The object information can be text information corresponding to multimodal content, including object name and object description. The multimodal content includes at least one of image content, text content, and audio content. This embodiment of the disclosure does not specifically limit the meaning of multimodal content; other content not listed that is used to characterize the target object also belongs to the multimodal content corresponding to the target object.
[0039] The accompanying text prompts represent the descriptive information required by the copy generation model to generate the target text. Specifically, the text prompts constrain the operations performed by the copy generation model during the text generation process. The copy generation model generates the target text corresponding to the target object. A large generative language model (LLM) can be used as the base model, and the copy generation model can be obtained through fine-tuning.
[0040] Large-scale generative language models (MLLMs) are computer models capable of processing and generating natural language. They represent a significant advancement in artificial intelligence and hold the promise of transforming the field through learned knowledge. MLLMs can predict the next word or sentence by learning statistical patterns and semantic information from language data; their capabilities increase as the input dataset and parameter space expand. They are used in various application areas, such as robotics, machine learning, machine translation, speech recognition, and image processing, hence the name "multimodal large-scale generative language model" (MLLM).
[0041] Model fine-tuning refers to using an LLM-based model and supervising training based on training samples to enhance the model's understanding of copywriting generation tasks. This involves providing more explicit copywriting generation instructions to enable the model to understand and make correct responses.
[0042] Since the target audience is the object of the copy to be generated within the marketing industry, the target copy can be marketing copy for that target audience. The copy description information can include the generation requirements set for generating the target copy. For example, the copy description information can include at least one of type information, content information, and style information. Type information defines the copy type of the target copy, which can include promotional copy, etc. Content information defines the copy content of the target copy, which can include copy requirements, selling points, target users, and word count, etc. Style information defines the writing style of the target copy, which can include a serious style or a lighthearted style, etc. For a serious style, formal written language can be used to generate the target copy; for a lighthearted style, popular internet slang can be used.
[0043] For example, the network address of the target object is obtained, the multimodal content corresponding to the target object is obtained based on the network address, and the object information is generated based on the multimodal content, wherein the multimodal content includes at least one of image content, text content, and audio content.
[0044] Information input controls can be displayed on the target interactive interface for users to input the target object's network address. The network address of the target object is entered into the input control, and multimodal content corresponding to the target object is obtained based on the network address through a pre-agreed interface with the e-commerce platform. Character information in the image content of the target object is recognized using OCR (Optical Character Recognition) technology. Optionally, image features are mapped to text features using a large-scale multimodal generative language model to obtain image description information corresponding to the image content. Text information corresponding to the audio content of the target object is generated using speech recognition technology. The text information corresponding to the multimodal content is determined based on the character information and / or image description information corresponding to the image content, the text information corresponding to the audio content, and the text content of the target object.
[0045] Obtain a copy generation instruction, and determine the copy prompt text based on the copy description information corresponding to the copy generation instruction, wherein the copy description information includes at least one of type information, content information, and style information.
[0046] For example, the text generation instruction entered in the information input control of the target interactive interface is obtained. Since the text generation instruction includes relevant requirements for generating the target text, such as text type, text content, text word count, target audience, and text style, the text description information can be obtained by parsing the text generation instruction. The text description information includes at least one of type information, content information, and style information. Optionally, key content in the text description information corresponding to the text generation instruction is extracted to generate the text prompt text corresponding to the target object.
[0047] S120. Input the object information and the copy prompt text into the copy generation model to obtain the target copy output by the copy generation model.
[0048] The copy generation model is obtained by fine-tuning a large generative language model based on training sample pairs. The training sample pairs include sample description text and sample copy. The sample description text includes sample object information and sample copy prompt text. The sample copy is a recommended copy generated based on the sample object information and at least two sample copy prompt texts, after at least two rounds of dialogue.
[0049] In this embodiment, the sample prompt text is determined through a multi-round iterative process. First, candidate prompt texts are determined based on business scenarios and historical experience. Then, the candidate prompt texts are modified through manual debugging to finally obtain prompt texts that meet the training requirements.
[0050] Training sample pairs include sample description text and sample copy, where the sample description text describes the training samples. For example, the sample description text includes sample object information and sample copy prompts. The sample object information represents the text information corresponding to the multimodal content of the sample object. Sample object information can be obtained from the internet based on the object type. The object type can refer to the marketing industry category to which the sample object belongs. For example, product details data for beauty products over a past period can be obtained as sample object information. Or, product details data for snack products over a past period can be obtained as sample object information, etc. The sample copy prompt text represents the copy description information required for the large-scale generative language model to generate the sample copy. The sample object information and at least two sample copy prompt texts are input into the large-scale generative language model, and the sample copy is generated through at least two rounds of dialogue. Furthermore, by combining sample description text and sample copy to form training sample pairs, since the copy quality of sample copy generated through at least two rounds of dialogue is significantly improved compared to the quality of sample copy generated by a large generative language model in a single round, supervised fine-tuning of the large generative language model using training sample pairs yields a copy generation model that is smaller in scale and more stable. Using the copy generation model to generate target copy for the target object can improve the quality of the copy.
[0051] For example, the training methods for copywriting generation models include:
[0052] Obtain sample object information and at least two sample text prompts, wherein the sample text prompts are the question texts in the dialogue. Input the sample object information and at least two sample text prompts into the large-scale generative language model, and generate the sample text through at least two rounds of dialogue using the large-scale generative language model. Determine sample description text based on the sample object information and at least two sample text prompts. Construct training sample pairs based on the sample description text and sample text, and train the large-scale generative language model based on the training sample pairs to obtain the text generation model.
[0053] Optionally, the copywriting generation model can be a marketing industry vertical model, and correspondingly, the sample objects can be products or items corresponding to that marketing industry. Specifically, sample object information of sample objects in the Internet is obtained according to the object type corresponding to the marketing industry. For example, sample object information of sample objects can be obtained from e-commerce platforms and other channels based on product type or product description through fuzzy search. Fuzzy search can refer to the search engine using the product type or product description entered by the user as keywords, and searching for products according to keywords and their synonyms. After searching for products from e-commerce platforms, product detail data is collected as sample object information. Product detail data includes, but is not limited to, product name, product description, product image, product reviews, and selling points. By obtaining real product detail information as training sample pairs, the copywriting generation model trained based on the training sample pairs in this embodiment of the disclosure can cover a wide range of product information content.
[0054] Because large generative language models significantly outperform multi-task processing on a single task, directly generating text using such models might lead to the model neglecting certain tasks during multi-task execution, resulting in text quality falling short of expectations. To address this instability and improve text quality, the text generation task can be broken down into at least two sub-tasks using a chain-like algorithm. The large generative language model then executes these sub-tasks sequentially, significantly improving the model's performance and ultimately enhancing text quality.
[0055] Since the copy generation task needs to be broken down into at least two sub-tasks, at least two sample copy prompt texts need to be constructed. The sample copy prompt texts are the question texts in the dialogue of the large generative language model, and the large generative language model generates response texts based on the sample object information and the sample copy prompt texts.
[0056] In this embodiment of the disclosure, the copywriting generation task can be divided into three sub-tasks, including the expansion of selling points, copywriting generation, and copywriting modification.
[0057] Specifically, the sample object information and at least two sample text prompts are input into the large-scale generative language model, and the sample text is generated through at least two rounds of dialogue using the large-scale generative language model, including:
[0058] The sample object information and the first sample copy prompt text are input into the large-scale generative language model to obtain the selling point content output by the large-scale generative language model, wherein the first sample copy prompt text is determined based on the selling point content generation requirements.
[0059] Since the first subtask of the copywriting generation task is the expansion of selling points, a first sample copywriting prompt text can be constructed by combining sample object information and a pre-set prompt text template. The prompt text template can include a template recording the requirements for generating selling point content. For example, the prompt text template includes standard statements corresponding to the selling point content generation requirements, which define the relevant requirements for expanding selling point content based on sample object information. Fields in the prompt text template such as product name and market information are fields to be filled. These fields can be filled based on the product name and market information of the sample object information to obtain the first sample copywriting prompt text. The sample object information and the first sample copywriting prompt text are input into a large-scale generative language model, which generates the selling point content, marketing scenarios, and target users of the sample object. Optionally, the large-scale generative language model can also be used to generate selling point script formulas corresponding to the selling point content. The selling point script formulas include content tags for the sample copy to be generated. These content tags include promotions, etc.
[0060] The selling points and the second sample copy prompt text are input into the large-scale generative language model to obtain the initial copy output by the large-scale generative language model, wherein the second sample copy prompt text is determined based on the copy generation requirements.
[0061] Since the copywriting generation task is broken down into a second subtask, copywriting generation can be used to obtain the marketing scenario and target users output by the large generative language model when executing the first subtask. This is combined with the model output of the first subtask and a pre-set prompt text template to construct the second sample copywriting prompt text. The prompt text template can also include a template recording the copywriting generation requirements. For example, the prompt text template includes standard statements corresponding to the copywriting generation requirements, which include copywriting description information such as product name, product details, product industry, marketing scenario and target users, word count, and copywriting style. The fields corresponding to the above copywriting description information in the prompt text template can be fields to be filled. Based on the model output of the first subtask and the sample object information, these fields are filled to obtain the second sample copywriting prompt text. The selling point content output by the model and the second sample copywriting prompt text are input into the large generative language model, which generates the initial copywriting corresponding to the sample object. Optionally, the selling point content, selling point script formula, and second sample copywriting prompt text output by the model can also be input into the large generative language model to generate the initial copywriting corresponding to the sample object.
[0062] The initial text and the third sample text prompt text are input into the large-scale generative language model to obtain the sample text output by the large-scale generative language model, wherein the third sample text prompt text is determined based on the text modification requirements.
[0063] Because the second subtask includes multiple objectives during execution, such as ensuring the copy is in authentic English, meets the requirements of the target marketing market, is limited to 200 words, and has a lighthearted style, large-scale generative language models may overlook some objectives when processing these multiple objectives, as each objective is equivalent to a model task. This could lead to the generation of copy that does not meet expectations. By inputting the output of the second subtask and the third sample copy prompt text into the large-scale generative language model, the model can modify the initial copy output of the second subtask, resulting in high-quality sample copy. To guide the modification process, the third sample copy prompt text can be a pre-set prompt text template that includes the copy modification requirements. Alternatively, the third sample copy prompt text can be used to allow the large-scale generative language model to provide its own feedback. For example, the third sample copy prompt text could be something like, "As a video creator, please help me see what problems exist with the initial copy."
[0064] After generating the sample copy mentioned above, a sample description text is generated based on the sample object information, the first sample copy prompt text, the second sample copy prompt text, and the third sample copy prompt text. For example, key content is extracted from the sample object information, the first sample copy prompt text, the second sample copy prompt text, and the third sample copy prompt text, and the sample description text is generated based on this key content. This key content refers to the content corresponding to pre-defined fields, including product name, product details, copy word count, copy requirements, selling points, target users, and industry, etc.
[0065] Training sample pairs for a single inference are constructed based on the correspondence between sample description text and sample copy. A large-scale generative language model is then trained based on these training sample pairs. Optionally, some model parameters of the large-scale generative language model can be frozen, and the unfrozen model parameters can be adjusted based on the training sample pairs to achieve fine-tuning of the large-scale generative language model and obtain the copy generation model.
[0066] The technical solution of this disclosure involves acquiring object information and textual prompts for a target object, inputting these information and prompts into a text generation model, and obtaining the target text output by the model. Since the sample text is generated through multiple rounds of dialogue using a large-scale generative language model based on sample object information and sample textual prompts, the quality of the sample text can be improved, avoiding the problem of suboptimal text quality when directly generated by a large-scale generative language model. Sample text and sample descriptive text are used to form training sample pairs. Then, supervised fine-tuning of the large-scale generative language model is performed using these training sample pairs to obtain a smaller-scale text generation model. Using the text generation model to perform the text generation task can reduce inference time and improve the quality of text generation.
[0067] Figure 2 This is a flowchart illustrating a training method for a copywriting generation model provided in this embodiment. Based on the above embodiments, this embodiment specifically defines the acquisition of sample object information and at least two sample copywriting prompt texts.
[0068] like Figure 2 As shown, the method includes:
[0069] S210. Obtain the sample object information of the sample object according to the object type.
[0070] Since products have already been categorized within the marketing industry, data for the corresponding marketing industry vertical can be obtained based on the product type. Sample object information includes product details data from the marketing industry, including but not limited to product name, description, and selling points.
[0071] S220. Decompose the copy generation task of the large-scale generative language model to obtain at least two sub-tasks.
[0072] Since the copywriting generation task is to generate sample copy based on sample object information, it can be divided into two sub-tasks using the thought chain algorithm. The first sub-task is to generate selling point content based on the sample object information, and the second sub-task is to generate sample copy based on the selling point content. Optionally, since the copy obtained by modifying the copy result with prompts from a large-scale generative language model is of higher quality than the copy generated directly, the sample copy generated based on the selling point content and the sample copy prompt text can also be input into the large-scale generative language model to modify the sample copy and obtain a copy result with higher quality. Optionally, the copywriting generation task can also be divided into a first sub-task, a second sub-task, and a sub-task of modifying the sample copy at least once through a large-scale generative language model.
[0073] Taking the modification of sample copy through a large generative language model as an example, the copy generation task needs to be broken down into three sub-tasks. The first sub-task is to generate selling point content based on sample object information, the second sub-task is to generate initial copy based on selling point content, and the third sub-task is to modify the initial copy to obtain sample copy.
[0074] It should be noted that if multiple different sample copy prompt texts are set for the sample copy to modify the copy result multiple times from different angles, the copy generation task can be divided into 4 sub-tasks or more. This disclosure does not specifically limit this.
[0075] S230. Determine at least two types of text description information of the sample text based on the at least two sub-tasks, and determine at least two sample text prompt texts based on the at least two types of text description information.
[0076] The text description information is used to represent the task requirements corresponding to the sub-tasks. For example, the first sub-task is to generate the selling points of the sample object, and the corresponding text description information could be to expand on the selling points of the sample object and generate the corresponding selling point content. The second sub-task is to generate the text for the sample object, and the corresponding text description information could be to generate marketing copy for the xx region targeting the ss audience, with the text length within dd, etc. Optionally, the third sub-task could be to modify the text for the sample object, and the corresponding text description information could be to assume that the large generative language model is an experienced copywriter and ask the large generative language model to modify the text result of the second sub-task. Optionally, the text description information corresponding to the third sub-task may also include copywriting style, etc.
[0077] For example, based on the task requirement information of the at least two sub-tasks, at least two types of text description information for the sample text are determined. A corresponding sample text prompt text is generated based on the first type of text description information, wherein the first type of text description information is used to characterize the task requirement information of the first sub-task with the sample object information as input data. A corresponding sample text prompt text is generated based on the second type of text description information and the task execution result of the first sub-task, wherein the second type of text description information is used to characterize the task requirement information of the second sub-task with the task execution result of the first sub-task as input data. A corresponding sample text prompt text is generated based on the third type of text description information, wherein the third type of text description information is used to characterize the task requirement information of the third sub-task with the task execution result of the second sub-task as input data.
[0078] S240. Input the sample object information and at least two sample text prompts into the large-scale generative language model, and generate the sample text through at least two rounds of dialogue using the large-scale generative language model.
[0079] S250. Determine the sample description text based on the sample object information and at least two sample text prompts.
[0080] S260. Based on the sample description text and sample copy, a training sample pair is formed, and the large-scale generative language model is trained based on the training sample pair to obtain the copy generation model.
[0081] Figure 3 This is a flowchart illustrating another method for training a copywriting generation model provided in an embodiment of this disclosure. Figure 3 As shown, the method includes:
[0082] S310. Obtain sample object information according to the marketing industry.
[0083] For example, the marketing industry includes categories such as electronics, clothing, snacks, and accessories. Based on the marketing industry to which the constructed copywriting generation model is applicable, sample object information is obtained from the corresponding industry. For instance, if the copywriting generation model is a vertical model for generating copy for snacks, then product details information for snack products is obtained. By acquiring real product details data from the internet as sample object information, the copywriting generation model can cover a wide range of product content. Obtaining sample object information based on the marketing industry reduces the amount of sample data. Failing to distinguish between marketing industries could result in a large amount of data, increasing the resources required for model training.
[0084] S320, a data construction scheme based on thought chain.
[0085] For example, the copywriting generation task is broken down into three sub-tasks based on the thought chain algorithm: the first sub-task is selling point extraction, the second is copywriting creation, and the third is copywriting modification. Specifically, firstly, based on product and product detail data combined with sample copywriting prompts, the selling points are expanded and the target users are predicted; then, based on the selling points and target users combined with the sample copywriting prompts, a rough copy is generated; finally, the rough copy is modified based on the sample copywriting prompts to obtain a high-quality sample copy. This approach can solve the problem that large generative language models may forget some tasks when processing multiple tasks, leading to poor copywriting quality.
[0086] For the first subtask, based on the product details data, through debugging and constructing sample prompt text once, a large-scale generative language model is used to generate corresponding selling point content and corresponding selling point script formulas. The selling point content can include expanded content about the product's selling points. The selling point script formula can be tags corresponding to the copywriting content. For example, assuming the product copywriting is about product discounts and promotions, the product copywriting includes promotional tags. Or, assuming the product copywriting includes limited-purchase content, the product copywriting includes scarcity marketing tags. Optionally, the selling point script formula can be used to generate sample prompt text for the second subtask.
[0087] For the second subtask, the selling points output from the first subtask are used as input data, combined with the second sample text prompts, to generate initial copy using a large-scale generative language model. This initial copy is a rough marketing draft.
[0088] For the third subtask, the initial text output from the second subtask is used as the input data for the third subtask. Combined with the prompt text of the third sample text, the initial text is modified through a large-scale generative language model to obtain the sample text.
[0089] Extract the key content from the prompt texts of the first, second, and third samples, and use it as metadata to establish the association between the sample texts and the metadata.
[0090] S330, Store sample text.
[0091] S340, Data Filtering.
[0092] Filter the product detail data. Specifically, filter the product detail data according to data review requirements to remove data that does not meet preset requirements.
[0093] The sample text is filtered to remove data that does not meet the preset requirements. Key content is extracted from the filtered product details data, and the extracted results are added to the metadata to obtain the sample description text.
[0094] S350. Using a large-scale generative language model as the base model, supervised fine-tuning is performed on training sample pairs consisting of sample descriptive text and sample copy to obtain a copy generation model.
[0095] The technical solution of this disclosure decomposes the text generation task of a large-scale generative language model into at least two sub-tasks. Based on these sub-tasks, at least two types of text description information are determined. Sample text prompts are generated based on the different types of text description information. The output of the previous sub-task is used as input data, and the text prompts are used as question texts. The large-scale generative language model generates sample text through multi-turn dialogue, resulting in sample text with higher quality than the output of a single-turn dialogue model. Furthermore, sample description text is determined based on sample object information and at least two sample text prompts. These sample description texts and sample text constitute training sample pairs. Fine-tuning the large-scale generative language model using these training sample pairs yields a text generation model that improves model stability and text generation quality while reducing inference time.
[0096] Figure 4 This is a schematic diagram of a text generation device provided in an embodiment of the present disclosure. The device can be implemented in the form of software and / or hardware, and optionally, it can be implemented in the form of an electronic device, such as a mobile terminal, a PC, or a server.
[0097] like Figure 4 As shown, the device includes an information acquisition module 410 and a document generation module 420.
[0098] The information acquisition module 410 is used to acquire object information and text prompt text of the target object, wherein the object information represents the text information corresponding to the multimodal content of the target object, and the text prompt text represents the text description information required by the text generation model to generate the target text.
[0099] The copy generation module 420 is used to input the object information and the copy prompt text into the copy generation model to obtain the target copy output by the copy generation model. The copy generation model is obtained by fine-tuning a large generative language model based on training sample pairs. The training sample pairs include sample description text and sample copy. The sample description text includes sample object information and sample copy prompt text. The sample copy is a recommended copy generated based on the sample object information and at least two sample copy prompt texts through at least two rounds of dialogue.
[0100] Optionally, the information acquisition module 410 is specifically used for:
[0101] Obtain the network address of the target object, obtain the multimodal content corresponding to the target object based on the network address, and generate the object information based on the multimodal content, wherein the multimodal content includes at least one of image content, text content, and audio content;
[0102] Obtain a copy generation instruction, and determine the copy prompt text based on the copy description information corresponding to the copy generation instruction, wherein the copy description information includes at least one of type information, content information, and style information.
[0103] Optionally, the training method for the copywriting generation model includes:
[0104] Obtain sample object information and at least two sample text prompts, wherein the sample text prompts are the question texts in the dialogue;
[0105] The sample object information and at least two sample text prompts are input into the large-scale generative language model, and the sample text is generated through at least two rounds of dialogue using the large-scale generative language model.
[0106] The sample description text is determined based on the sample object information and at least two sample text prompts.
[0107] The training sample pairs are constructed based on the sample description text and sample copy, and the large-scale generative language model is trained based on the training sample pairs to obtain the copy generation model.
[0108] Furthermore, the acquisition of sample object information and at least two sample text prompts includes:
[0109] Obtain the sample object information of the sample object according to the object type;
[0110] The copy generation task of the large generative language model is decomposed into at least two sub-tasks;
[0111] Based on the at least two sub-tasks, at least two types of text description information of the sample text are determined, and at least two sample text prompt texts are determined based on the at least two types of text description information.
[0112] Further, the step of determining at least two types of text description information of the sample text based on the at least two sub-tasks, and determining at least two sample text prompt texts based on the at least two types of text description information, includes:
[0113] Based on the task requirement information of the at least two sub-tasks, determine at least two types of text description information for the sample text;
[0114] The first type of text description information is used to generate corresponding sample text prompt text, wherein the first type of text description information is used to characterize the task requirement information of the first sub-task with the sample object information as input data.
[0115] Based on the second type of text description information and the task execution result of the first subtask, a corresponding sample text prompt text is generated, wherein the second type of text description information is used to characterize the task requirement information of the second subtask with the task execution result of the first subtask as input data;
[0116] The corresponding sample text prompt text is generated based on the third type of text description information, wherein the third type of text description information is used to characterize the task requirement information of the third subtask with the task execution result of the second subtask as input data.
[0117] Further, the step of inputting the sample object information and at least two sample text prompts into the large-scale generative language model, and generating the sample text through at least two rounds of dialogue using the large-scale generative language model, includes:
[0118] The sample object information and the first sample copy prompt text are input into the large-scale generative language model to obtain the selling point content output by the large-scale generative language model, wherein the first sample copy prompt text is determined based on the selling point content generation requirements.
[0119] The selling point content and the second sample copy prompt text are input into the large generative language model to obtain the initial copy output by the large generative language model, wherein the second sample copy prompt text is determined based on the copy generation requirements.
[0120] The initial text and the third sample text prompt text are input into the large-scale generative language model to obtain the sample text output by the large-scale generative language model, wherein the third sample text prompt text is determined based on the text modification requirements.
[0121] The copy generation apparatus provided in this disclosure can execute the copy generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0122] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0123] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 5 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 5The diagram below shows the structure of the terminal device or server 500. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0124] like Figure 5 As shown, electronic device 500 may include a processing unit (e.g., central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. An edit / output (I / O) interface 505 is also connected to bus 504.
[0125] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0126] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0127] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0128] The electronic device provided in this embodiment and the text generation method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0129] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the document generation method provided in the above embodiments.
[0130] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0131] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0132] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0133] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0134] Obtain object information and text prompts for the target object, wherein the object information represents the text information corresponding to the multimodal content of the target object, and the text prompts represent the text description information required by the text generation model to generate the target text;
[0135] The object information and the text prompt text are input into the text generation model to obtain the target text output by the text generation model. The text generation model is obtained by fine-tuning a large generative language model based on training sample pairs. The training sample pairs include sample description text and sample text. The sample description text includes sample object information and sample text prompt text. The sample text is a recommended text generated based on the sample object information and at least two sample text prompt texts after at least two rounds of dialogue.
[0136] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0137] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0138] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.
[0139] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0140] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0141] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0142] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0143] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A method for generating copy, characterized in that, include: Obtain object information and text prompts for the target object, wherein the object information represents the text information corresponding to the multimodal content of the target object, and the text prompts represent the text description information required by the text generation model to generate the target text; The object information and the text prompt text are input into the text generation model to obtain the target text output by the text generation model. The text generation model is obtained by fine-tuning a large generative language model based on training sample pairs. The training sample pairs include sample description text and sample text. The sample description text includes sample object information and sample text prompt text. The sample text is a recommended text generated based on the sample object information and at least two sample text prompt texts after at least two rounds of dialogue.
2. The method according to claim 1, characterized in that, The acquisition of object information and text prompts for the target object includes: Obtain the network address of the target object, obtain the multimodal content corresponding to the target object based on the network address, and generate the object information based on the multimodal content, wherein the multimodal content includes at least one of image content, text content, and audio content; Obtain a copy generation instruction, and determine the copy prompt text based on the copy description information corresponding to the copy generation instruction, wherein the copy description information includes at least one of type information, content information, and style information.
3. The method according to claim 1, characterized in that, The training methods for the copywriting generation model include: Obtain sample object information and at least two sample text prompts, wherein the sample text prompts are the question texts in the dialogue; The sample object information and at least two sample text prompts are input into the large-scale generative language model, and the sample text is generated through at least two rounds of dialogue using the large-scale generative language model. The sample description text is determined based on the sample object information and at least two sample text prompts. The training sample pairs are constructed based on the sample description text and sample copy, and the large-scale generative language model is trained based on the training sample pairs to obtain the copy generation model.
4. The method according to claim 3, characterized in that, The acquisition of sample object information and at least two sample text prompts includes: Obtain the sample object information of the sample object according to the object type; The copy generation task of the large generative language model is decomposed into at least two sub-tasks; Based on the at least two sub-tasks, at least two types of text description information of the sample text are determined, and at least two sample text prompt texts are determined based on the at least two types of text description information.
5. The method according to claim 4, characterized in that, The step of determining at least two types of text description information of the sample text based on the at least two sub-tasks, and determining at least two sample text prompt texts based on the at least two types of text description information, includes: Based on the task requirement information of the at least two sub-tasks, determine at least two types of text description information for the sample text; The first type of text description information is used to generate corresponding sample text prompt text, wherein the first type of text description information is used to characterize the task requirement information of the first sub-task with the sample object information as input data. Based on the second type of text description information and the task execution result of the first subtask, a corresponding sample text prompt text is generated, wherein the second type of text description information is used to characterize the task requirement information of the second subtask with the task execution result of the first subtask as input data; The corresponding sample text prompt text is generated based on the third type of text description information, wherein the third type of text description information is used to characterize the task requirement information of the third subtask with the task execution result of the second subtask as input data.
6. The method according to claim 3, characterized in that, The step of inputting the sample object information and at least two sample text prompts into the large-scale generative language model, and generating the sample text through at least two rounds of dialogue using the large-scale generative language model, includes: The sample object information and the first sample copy prompt text are input into the large-scale generative language model to obtain the selling point content output by the large-scale generative language model, wherein the first sample copy prompt text is determined based on the selling point content generation requirements. The selling point content and the second sample copy prompt text are input into the large generative language model to obtain the initial copy output by the large generative language model, wherein the second sample copy prompt text is determined based on the copy generation requirements. The initial text and the third sample text prompt text are input into the large-scale generative language model to obtain the sample text output by the large-scale generative language model, wherein the third sample text prompt text is determined based on the text modification requirements.
7. A copywriting generation device, characterized in that, include: The information acquisition module is used to acquire object information and text prompts of the target object, wherein the object information represents the text information corresponding to the multimodal content of the target object, and the text prompts represent the text description information required by the text generation model to generate the target text. The copy generation module is used to input the object information and the copy prompt text into the copy generation model to obtain the target copy output by the copy generation model. The copy generation model is obtained by fine-tuning a large generative language model based on training sample pairs. The training sample pairs include sample description text and sample copy. The sample description text includes sample object information and sample copy prompt text. The sample copy is a recommended copy generated based on the sample object information and at least two sample copy prompt texts through at least two rounds of dialogue.
8. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the copy generation method as described in any one of claims 1-6.
9. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the copy generation method as described in any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the copy generation method as described in any one of claims 1-6.