Copywriting generation method, device, electronic device and storage medium
Through the multimodal copy generation model combining images and text, users support the modification of the first draft of copywriting, solving the problems of insufficient image processing and inflexible editing in existing tools, achieving efficient and flexible copywriting generation and editing, and improving user experience.
Patent Information
- Application Number
- CN202410847601.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2044-06-27
AI Technical Summary
The existing copywriting generation tools lack support for image processing and cannot meet the interaction needs of users for free editing, which increases creative complexity and time cost, and lacks flexibility and editability.
Through a multimodal copy generation model, combining multiple images and/or initial requirements descriptions entered by the user, a first draft of the copy is generated, and the user's modification requirements description of the first draft is supported, further optimization and editing of the copy is realized, and user feedback and iterative modification mechanisms are introduced.
It improves the efficiency and accuracy of copywriting generation, meets users' free editing interaction needs, provides greater flexibility and diversity, and enhances users' satisfaction with the generated copywriting.
Smart Images

Figure CN118673136B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a copywriting generation method, device, electronic device and storage medium. Background Art
[0002] With the rapid development and widespread adoption of artificial intelligence (AI) technology, copywriting generation has become a growing focus. Copywriting generation uses natural language processing (NLP) and machine learning algorithms to enable computers to automatically generate grammatically correct content. Currently, copywriting generation technology is widely used in advertising, marketing, journalism, social media, and other fields, significantly improving content creation efficiency and reducing costs while also providing creators with more inspiration and possibilities.
[0003] However, existing copywriting tools often focus solely on automatic text generation and lack support for image processing. Users must manually match and integrate images with text, significantly increasing the complexity and time cost of copywriting generation. Furthermore, existing copywriting tools lack flexibility and editability, failing to support user interaction and free editing, which negatively impacts the user experience. Summary of the Invention
[0004] The present invention provides a copy generation method, device, electronic device and storage medium, which are used to solve the defects of the related art that copy generation lacks support for image processing and cannot meet the interactive needs of users for free editing.
[0005] The present invention provides a copywriting generation method, comprising:
[0006] Obtain multiple images input by the user and / or an initial requirement description for the copy to be generated;
[0007] Based on the copy generation model, the plurality of images and / or the initial requirement description are used to generate a copy to obtain a first draft of the copy; and a modification requirement description for the first draft of the copy input by the user is obtained;
[0008] Based on the document generation model, the modification requirement description is applied, or the multiple images and the modification requirement description are applied, to modify the document draft to generate a target document.
[0009] According to a method for generating a document provided by the present invention, the method generates a document based on a document generation model by applying the plurality of images and / or the initial requirement description to obtain a first draft of the document, and then further comprises:
[0010] Obtain multiple images currently input by the user and a description of modification requirements for the draft copy;
[0011] Based on the text generation model, the multiple images inputted by the user in the past, the multiple images inputted by the user currently, and the modification requirement description are applied to modify the text draft and generate a target text.
[0012] A copywriting generation method provided by the present invention further includes:
[0013] Obtaining text content for a to-be-generated image copy, wherein the text content is determined based on at least one of a user input, the draft copy, and the target copy;
[0014] Based on the copywriting generation model, applying the text content, a plurality of image prompt texts are generated, wherein the image prompt texts include an image content description and an image style description;
[0015] Based on the image generation model, the multiple image prompt texts are applied to generate target image copy corresponding to the text content, and the target image copy includes multiple target images.
[0016] A copywriting generation method provided by the present invention further includes:
[0017] Acquiring text information for a graphic copy to be generated, wherein the text information is determined based on at least one of user input, the draft copy, and the target copy;
[0018] Based on the copywriting generation model, applying the text information, generating prompt text, the prompt text including text body and image prompt text;
[0019] Based on the image prompt text, a target image is generated, and the text body and the target image are applied to generate a target graphic copy.
[0020] According to a copywriting generation method provided by the present invention, the prompt text includes multiple text sections and multiple image prompt texts, and each image prompt text corresponds to at least one text section;
[0021] The step of generating a target image based on the image prompt text and applying the text body and the target image to generate a target graphic copy includes:
[0022] Based on the image generation model, applying the multiple image prompt texts, generating multiple target images, wherein the target images correspond to the image prompt texts one by one;
[0023] The multiple image prompt texts in the prompt text are replaced with the corresponding multiple target images, and the target image-text copy is obtained based on the replaced prompt text.
[0024] According to a copywriting generation method provided by the present invention, when the text information includes the text input by the user and the draft copywriting, or when the text information includes the text input by the user and the target copywriting, generating a prompt text based on the copywriting generation model and applying the text information includes:
[0025] Based on the text generation model, the prompt text is generated by applying the text information and multiple images input by the user in the past.
[0026] The present invention also provides a document generation device, comprising:
[0027] A first acquisition unit is configured to acquire multiple images input by a user and / or an initial requirement description for a document to be generated;
[0028] A copy generation unit, configured to generate a copy based on a copy generation model and applying the plurality of images and / or the initial requirement description to obtain a first draft of the copy;
[0029] A second acquiring unit is configured to acquire a description of modification requirements for the first draft of the document input by the user;
[0030] A document modification unit is configured to modify the document draft based on the document generation model and apply the modification requirement description, or apply the multiple images and the modification requirement description, to generate a target document.
[0031] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the above-described methods for generating text is implemented.
[0032] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described methods for generating text.
[0033] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned methods for generating text.
[0034] The copy generation method, device, electronic device and storage medium provided by the present invention support users to input multiple images and / or initial requirement descriptions for the copy to be generated, providing users with greater flexibility and diversity. By combining the requirement descriptions in the form of images and text, the user's intentions can be understood more comprehensively, thereby generating a copy that better meets the user's expectations. Moreover, the present invention introduces a mechanism for user feedback and iterative modification. The user can modify the generated draft copy and guide the model to further optimize the copy by inputting the modified requirement description. This interactivity and iterativeness greatly improves the efficiency and accuracy of copy generation, while meeting the user's interactive needs for free editing. In addition, due to the support for multimodal input and interactive modification, users can participate more directly in the copy generation process, and can adjust and optimize the copy according to their own needs and preferences, thereby improving user satisfaction with the final generated copy. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following is a brief introduction to the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0036] Figure 1 It is a flowchart of the copywriting generation method provided by the present invention;
[0037] Figure 2 This is one of the flowcharts of text copy generation and editing provided by the present invention;
[0038] Figure 3 This is the second flowchart of the text copy generation and editing provided by the present invention;
[0039] Figure 4 This is one of the flow charts for generating image text provided by the present invention;
[0040] Figure 5 This is the second flow chart of image copy generation provided by the present invention;
[0041] Figure 6 This is one of the flow charts for generating graphic and text copy provided by the present invention;
[0042] Figure 7 This is the second flow chart of the graphic and text copy generation process provided by the present invention;
[0043] Figure 8 It is a structural diagram of the copywriting generation device provided by the present invention;
[0044] Figure 9It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0045] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0046] With the advent of the digital media age, the combination of images and text has become an indispensable part of copywriting. Currently, copywriting tools often focus solely on automatic text generation and lack support for image processing. As user demand for multimedia content grows, copywriting technology must adapt to this change, supporting image processing and enabling interactive, free-form editing.
[0047] However, existing copywriting generation tools often only handle single text input and cannot directly process multiple images. Users must manually match and integrate images with text, significantly increasing the complexity and time cost of creation. Furthermore, during the copywriting process, users may need to edit and adjust the generated text or images to meet specific needs or styles. However, existing copywriting generation tools often lack flexibility and editability, failing to provide the free editing and interactive features required by users. Users can only passively accept the generated results without being able to modify or optimize them.
[0048] To address this issue, an embodiment of the present invention provides a method for generating a document to overcome the above-mentioned drawbacks. Figure 1 It is a flow chart of the copywriting generation method provided by the present invention, such as Figure 1 As shown, the method includes:
[0049] Step 110 , obtaining multiple images input by the user and / or an initial requirement description for the document to be generated;
[0050] It should be noted that the to-be-generated copy refers to the specific copy content that the user wishes to generate, which can be text copy, image copy, or a combination of text and image copy. For example, the to-be-generated copy can be a story, composition, news, advertising copy, blog, public account, or WeChat Moments copy, and the embodiments of the present invention do not specifically limit this.
[0051] Specifically, users can upload multiple images through the user interface and enter an initial description of the requirements for the copy to be generated in the corresponding text box. The system can then receive these inputs through the back-end code and prepare for subsequent processing. Here, the initial description of requirements refers to the specific requirements provided by the user in the first round of creation to generate a specific copy. It can include the subject matter, theme, word count, etc., so that the copy generation system can generate the desired copy based on these requirements.
[0052] Step 120 , based on a document generation model, applying the multiple images and / or the initial requirement description to generate a document to obtain a first draft of the document;
[0053] It should be noted that the copy generation model refers to a machine learning or deep learning model obtained by training multimodal information such as text and images, which is used to generate corresponding copy based on input information (such as images and / or demand descriptions). In an embodiment of the present invention, the copy generation model can be a multimodal large language model (Large Language Model, LLM, referred to as the large model), which processes large-scale text data and image data during training and has the ability to understand and generate natural language. For example, the copy generation model can include the Spark large model.
[0054] Specifically, after the user provides an image and / or an initial description of their needs, the system can pass these inputs to the copywriting model. The model processes the input internally and outputs a corresponding draft copy. Here, the draft copy refers to the copywriting content generated by the copywriting model based on the user's initial input.
[0055] In one embodiment, when a user has a need for copywriting, he or she can upload multiple images on the user interface, where the requirements for copywriting can also be uploaded in the form of images. After receiving the multiple images uploaded by the user, the copywriting generation model can generate a draft of the copywriting related to the image content by identifying and analyzing the images.
[0056] In another embodiment, when a user has a need for copy generation, he or she can enter an initial requirement description for the copy to be generated in the text box of the user interface. The requirement description may include the user's creation requirements for the copy to be generated and a background introduction to the copy to be generated. After receiving the initial requirement description input by the user, the copy generation model can generate a draft of the copy that meets the initial requirement description by understanding and analyzing the description content.
[0057] In another embodiment, when a user has a need for copywriting generation, he or she can upload multiple images in the user interface and enter an initial requirement description for the copywriting to be generated in the text box. After receiving the content input by the user, the copywriting generation model can jointly create the copywriting based on the image content and the requirement description, thereby generating a first draft of the copywriting that meets the user's requirements.
[0058] In this embodiment of the present invention, by applying a copywriting generation model, it is possible to automatically process multiple images and / or requirement descriptions input by users and generate copywriting based on this information. As the model continues to learn and optimize, it can better understand the user's intentions and needs and generate more accurate copywriting that meets the user's expectations.
[0059] Step 130: Obtain the modification requirement description for the first draft of the document input by the user;
[0060] It should be noted that the initial draft may contain some elements that meet user expectations, but there may also be areas that need to be revised and improved. In this regard, users can further optimize the initial draft by providing a description of their modification requirements to generate a target copy that better meets the requirements.
[0061] Specifically, after generating a draft document, the system first displays it to the user for review. The user can then enter a description of the required changes to the draft document through the user interface. Upon receiving this input, the system backend performs the necessary processing, such as parsing the modification instructions and converting them into a format understandable to the model.
[0062] After processing the modification requirement description entered by the user, the system will pass this information to the copy generation model so that the model can make further modifications and optimizations. It should be understood that the modification requirement description for the first draft of the copy refers to the specific modification opinions or requirements for the first draft proposed by the user after reviewing the first draft of the copy generated by the copy generation model. These modification requirements may include multiple aspects such as the content, structure, style, grammar, and wording of the copy. For example, the user may want to rewrite, continue, or rewrite (the content of a paragraph or a sentence), adjust the logical order of the copy, change the tone or style of the copy, or correct grammatical errors in the copy.
[0063] Step 140 : Based on the document generation model, the modification requirement description is applied, or the plurality of images and the modification requirement description are applied, to modify the document draft to generate a target document.
[0064] Specifically, after receiving the modification requirement description input by the user, the copy generation model can modify the draft copy according to the requirement description to generate a target copy that better meets the user's requirements.
[0065] In one embodiment, the copywriting generation model can modify the initial draft of the copywriting based solely on the modification requirement description. Specifically, the system first parses the modification requirement description entered by the user and converts it into instructions or parameters that the model can understand. This may include natural language processing steps such as word segmentation, part-of-speech tagging, and syntactic analysis. The parsed modification requirements are then passed to the copywriting generation model, which then generates the modified copywriting, i.e., the target copywriting, based on the modification requirement description.
[0066] In another embodiment, the copy generation model can simultaneously apply multiple images input by the user history (i.e., multiple images uploaded by the user during the first round of creation) and the modification requirement description to modify the first draft of the copy. Specifically, for multiple images input by the user history, the system can use image processing technology to extract key features, such as color, texture, shape, etc., or extract high-level semantic information of the image through a deep learning model. The extracted image features are combined with the modification requirement description to form a comprehensive input vector, which contains both the user's text requirements and image information. Subsequently, the comprehensive input vector is passed to the copy generation model, and the model will generate a target copy that is related to the image content and meets the user's modification requirements based on the vector.
[0067] It's understandable that the target copy mentioned above refers to the copy that, after modification and optimization, ultimately meets user expectations. By considering both image and text information during the revision process, we can provide more material and inspiration for copy generation, resulting in more diverse and personalized copy.
[0068] The method provided by the embodiment of the present invention supports users to input multiple images and / or initial requirement descriptions for the copy to be generated, providing users with greater flexibility and diversity. By combining the requirement descriptions in the form of images and text, it is possible to more comprehensively understand the user's intentions, thereby generating a copy that better meets the user's expectations. Moreover, the present invention introduces a mechanism for user feedback and iterative modification. Users can modify the generated draft copy and guide the model to further optimize the copy by inputting and modifying the requirement description. This interactivity and iterativeness greatly improves the efficiency and accuracy of copy generation, while meeting the user's interactive needs for free editing. In addition, due to the support for multimodal input and interactive modification, users can participate more directly in the copy generation process, and can adjust and optimize the copy according to their own needs and preferences, thereby improving user satisfaction with the final generated copy.
[0069] Based on the above embodiment, after step 120, the method further includes:
[0070] Step 150 , obtaining multiple images currently input by the user and a description of modification requirements for the draft document;
[0071] Step 160 , based on the text generation model, the multiple images inputted by the user in the past, the multiple images inputted by the user currently, and the modification requirement description are applied to modify the text draft to generate a target text.
[0072] Specifically, after generating the first draft of the copy, the user can also upload multiple additional images, requesting to continue creating based on the first draft or rewrite the entire copy. In this case, the user can re-upload multiple images through the user interface and enter a description of the modification requirements for the first draft of the copy in the text box. After the system receives the user input through the back-end code, it will parse and process the image and requirement description respectively, and pass the processed information to the copy generation model. After receiving the input information, the model can modify the first draft of the copy based on the multiple images previously input by the user, the multiple images currently input by the user, and the modification requirement description, thereby generating the target copy that meets the user's requirements.
[0073] It is understandable that the multiple images historically input by the user refer to images that the user has previously input during the copy generation process and that were used by the system to generate the first draft of the copy. In step 110, the user may have already input some images, which, together with the initial requirements description, are used to generate the first draft of the copy. These historically input images record the user's original creative intent and style, and therefore still have reference value in subsequent steps. The multiple images currently input by the user refer to new images that the user has input again after the first draft of the copy is generated in order to further improve or modify the copy. In step 150, the user may supplement the input of some new images based on the content of the first draft of the copy or his or her own creative needs. These images contain the user's new creative intent or supplementary information.
[0074] Specifically, after receiving the multiple images currently input by the user and the modification requirement description, the copy generation model can generate the target copy through the following steps: First, the model can re-analyze the user's historical input images to understand the themes, content, emotions and styles represented by these images, and interpret the images currently input by the user to extract new information, themes and styles from these images. At the same time, the model can analyze the modification requirement description provided by the user to clarify how the user wishes to modify the draft copy. Subsequently, the model will fuse the information in the historical images, current images and modification requirement description to generate a comprehensive creative intent, and based on the comprehensive creative intent, modify the draft copy, for example, adding new content, modifying the expression, adjusting the emotional color, etc., to generate the target copy.
[0075] The method provided by the embodiments of the present invention meets the user's need to supplement image creation. Users can add new images at any time as needed, and the model can modify the copy based on these new images, improving creative flexibility. Users do not need to create from scratch; they only need to add new images and modify the required description, and the model can quickly generate new copy, improving creative efficiency. In addition, by combining historical images with current images, the model can better understand the user's creative intent and generate copy that better meets the user's needs.
[0076] Based on any of the above embodiments, the method further includes:
[0077] Step 210: obtaining text content for the image copy to be generated, wherein the text content is determined based on at least one of user input, the draft copy, and the target copy;
[0078] It should be noted that image copywriting (such as story comic strips) can more vividly present the storyline and emotional expression through images, making the content more interesting. Therefore, users' demand for copywriting generation is not limited to text copywriting, but also includes the need to create image copywriting. In this regard, embodiments of the present invention can generate multiple images based on the text content entered by the user or based on the copywriting content created by the user in the first round, thereby meeting the user's demand for creating image copywriting.
[0079] Specifically, the image copy to be generated refers to the set of images that the user wishes to generate based on a given text content. The text content of the image copy to be generated refers to the original or foundational text used to generate the image copy. This text content can be directly entered by the user or obtained from other sources (such as a draft copy or target copy). This text content serves as input to the image generation model to guide the generation of the image copy.
[0080] It is understandable that users can directly enter a piece of text in the user interface, and this text will serve as the text content of the image copy to be generated. User input can be keywords, phrases, sentences or paragraphs to describe the theme, content or style of the image copy they want to generate, and can also include instructions for generating corresponding images. In addition, if the user has conducted a first round of creation, the generated draft copy can also be directly used as the basis for further generating image copy. The user can enter the instruction "Generate corresponding image based on the draft copy". After receiving the instruction, the model can generate related image copy based on the draft copy. Similarly, the target copy obtained after the copy is modified and improved can also be used as the text content for generating image copy.
[0081] Step 220 , based on the text generation model, applying the text content, generating a plurality of image prompt texts, wherein the image prompt texts include image content description and image style description;
[0082] Specifically, after receiving the text content for the image copy to be generated, the copy generation model will analyze and understand the text content. For example, the model can use natural language processing techniques (such as word segmentation, part-of-speech tagging, syntactic analysis, etc.) to extract key information from the text. Then, based on this key information and the language-to-visual mapping relationship learned during training, the model will generate multiple image prompt texts. These image prompt texts are intended to guide the subsequent image generation model on how to generate corresponding images based on the text content.
[0083] Here, image hint text refers to a series of text instructions or descriptions that describe the image content and style. Each image hint text can include two parts: an image content description and an image style description. The image content description is a textual description of the main elements, scenes, or objects in the image. For example, the image content description could be "a turtle and a hare standing at the starting line," which guides the image generation model in generating an image containing these elements. The image content description ensures that the generated image is consistent with the text content in terms of subject matter and objects. The image style description is a textual description of the image's visual style or presentation. For example, the image style description could be "black and white stick figure" or "cartoon style." These style descriptions can help the image generation model understand the user's desired visual effect and adjust the image generation algorithm or parameters accordingly. The image style description ensures that the generated image visually meets the user's expectations and aesthetic preferences. By combining the image content description and the image style description, the image hint text provides comprehensive and specific guidance to the subsequent image generation model, ensuring that the generated image meets both the text content requirements and the user's visual style preferences.
[0084] Step 230 : Based on the image generation model, the plurality of image prompt texts are applied to generate a target image copy corresponding to the text content, wherein the target image copy includes a plurality of target images.
[0085] It should be noted that an image generation model is a machine learning model that can generate corresponding images based on input information (such as text descriptions). For example, an image generation model can be a generative adversarial network (GAN), a variational autoencoder (VAE), or a Transformer model. Image generation models learn the mapping relationship from input to image through training, thereby generating images that match the input information.
[0086] Specifically, the image generation model receives multiple image prompt texts as input and parses these texts to extract key information, such as the main objects, scenes, actions, and desired visual style. Based on the extracted information, the model generates a series of image feature vectors that encode the specific content and style information of the target image. The model then uses a decoder to convert these image feature vectors into pixel values to generate the final image. This process may include multiple steps such as convolution and upsampling to ensure that the generated image conforms to the input description in terms of details and structure. Due to the presence of multiple image prompt texts (which may describe different scenes, angles, or styles), the model can generate a set of images corresponding to the text content, namely the target image copy. Here, the target image copy refers to a set of images generated based on the text content and closely related to the text content. These images describe a scene, character, event, or emotion in the text and are presented in a visual form.
[0087] The method provided by the embodiment of the present invention can meet the needs of users to generate image copy. By generating multiple images corresponding to the text content, it can provide users with rich visual expressions, making the text content more vivid and intuitive. Compared with traditional hand-drawing or photography methods, the use of image generation models can quickly generate a large number of high-quality image copy, improving creation efficiency. In addition, if the text content is a coherent story or a series of events, the model can generate a series of related images to form a story comic strip, providing users with an immersive reading experience.
[0088] Based on any of the above embodiments, the method further includes:
[0089] Step 310: obtaining text information for the graphic copy to be generated, wherein the text information is determined based on at least one of user input, the draft copy, and the target copy;
[0090] It should be noted that graphic copy refers to copy that combines text and images. It combines the advantages of text and images to convey information and emotions more comprehensively and vividly. The text description and image display complement each other, making the content more complete and accurate. The present invention can create a copy according to user requirements and insert images into the copy, thereby meeting the user's demand for creating graphic copy.
[0091] Specifically, the to-be-generated graphic and text copy refers to a copy that the user wishes to generate based on given text information, including both text and images. The to-be-generated text copy's text information refers to the text description or instructions used to guide the generation of the graphic and text copy. This text information can be directly input by the user or can be an existing draft or target copy.
[0092] It is understandable that the user can directly enter a text in the user interface, and this text will serve as the text information of the graphic copy to be generated. It can include the user's specific creative requirements, the theme and content of the graphic copy to be generated, and indicate that the generated copy is a graphic copy. In addition, if the user has conducted a first round of creation, the generated draft copy can also be directly used as the basis for further generation of graphic copy. The user can enter the instruction "Generate graphic copy based on the draft copy" and enter the corresponding creative requirements. After receiving the instruction, the model can generate related graphic copy based on the draft copy. Similarly, the target copy obtained after the copy is modified and improved can also be used as text information to generate graphic copy.
[0093] Step 320: Based on the text generation model, apply the text information to generate prompt text, where the prompt text includes a text body and an image prompt text;
[0094] Specifically, the copy generation model first receives the text information from step 310 as input and parses and processes it, extracting key information and understanding the content and context of the text. The copy generation model then uses its internal mechanisms to generate prompt text. Here, prompt text is generated by the copy generation model and serves as a text instruction or description to guide subsequent image generation or graphic copy generation. It can include two parts: the text body and the image prompt text.
[0095] Among them, the text body is the text part of the prompt text, which directly describes the text content in the graphic copy. This part of the content is closely related to the input text information and is an extraction, summary or expansion of the input text information. The text body can be in the form of a title, paragraph, sentence, etc., which is used to present the main text information in the graphic copy. The image prompt text is the image part of the prompt text, which describes the image requirements or description related to the text content. This part of the content can be used to guide the subsequent image generation process to ensure that the generated image is consistent with the text content. The image prompt text can include an image content description and an image style description. By generating such prompt text, the copy generation model provides clear guidance and requirements for subsequent image generation or graphic copy generation, ensuring that the generated graphic content can accurately convey the text information and meet the needs of users.
[0096] Step 330 : Generate a target image based on the image prompt text, and generate a target graphic copy by applying the text body and the target image.
[0097] Specifically, the image hint text contains detailed information and requirements about the target image. The model will first parse the text and extract key information, such as the main object, scene, action, style, etc. Based on the parsed image hint text, the model will generate corresponding image feature vectors, which describe the detailed information and attributes of the target image at the pixel level. Use a trained image generator (such as a GANs generator) to decode the image feature vector into specific pixel values. After decoding, the model will output one or more target images that match the description of the image hint text. Here, the target image refers to the image generated according to the image hint text and meets the user's needs.
[0098] Subsequently, the generated target image is combined with the text body generated in step 320, and the combined text body and target image are output as the final graphic copy. Alternatively, the image prompt texts in the prompt text can be replaced one by one with the generated target image, while retaining the text body, so that the target graphic copy can be obtained based on the replaced prompt texts.
[0099] The method provided by the embodiments of the present invention can meet the user's need to generate graphic and text copy. By applying a copy generation model, it can automatically convert text information into prompt text. The generated prompt text contains not only the main text but also image prompt text, which makes it possible to generate text-graphic combined copy. The use of image prompt text can accurately describe and convey information that is difficult to express in words, making the generated graphic and text copy more vivid, intuitive, and easy to understand.
[0100] Based on any of the above embodiments, the prompt text includes multiple text segments and multiple image prompt texts, and each image prompt text corresponds to at least one text segment;
[0101] Accordingly, step 330 specifically includes:
[0102] Based on the image generation model, applying the multiple image prompt texts, generating multiple target images, wherein the target images correspond to the image prompt texts one by one;
[0103] The multiple image prompt texts in the prompt text are replaced with the corresponding multiple target images, and the target image-text copy is obtained based on the replaced prompt text.
[0104] Specifically, during the process of generating graphic copy, the prompt text generated by the copy generation model may include multiple paragraphs of text and multiple image prompt texts, where each image prompt text corresponds to at least one paragraph of text. When the system retrieves the prompt text output by the copy generation model and finds that the image prompt text is included in the prompt text, it calls the image generation model and passes the image prompt text to the image generation model, thereby obtaining the target image output by the image generation model.
[0105] Once the target image corresponding to the image prompt text is generated through the image generation model, the image prompt text in the prompt text can be replaced with the target image. Specifically, the prompt text can be traversed to find all the image prompt texts; for each image prompt text, the corresponding target image can be found; then, in the prompt text, the corresponding image prompt text is replaced with the placeholder of the target image (such as a URL, file path, or embedded image code). After the replacement is completed, the result is a graphic copy that contains the original text body and the embedded target image.
[0106] Based on any of the above embodiments, when the text information includes the text input by the user and the draft copy, or when the text information includes the text input by the user and the target copy, step 320 specifically includes: based on the copy generation model, applying the text information and multiple images input by the user in history to generate the prompt text.
[0107] Specifically, when the text information includes user-entered text and a first draft, it indicates that after the user generated the first draft during the first round of creation, they continued to enter corresponding text in the user interface, requesting to continue creating with both text and images to generate a text-and-image copy. When the text information includes user-entered text and a target copy, it indicates that after the target copy was generated, the user continued to enter corresponding text in the user interface, requesting to continue creating with both text and images to generate a text-and-image copy.
[0108] In both scenarios, the copy generation model can simultaneously utilize textual information and multiple images from the user's past input to generate prompt text. Here, textual information provides fundamental semantics and context, while images convey visual information and emotion. Combining the two can generate richer and more diverse prompt text. Furthermore, information such as details and textures in images can supplement the textual content, making the generated prompt text more specific and vivid.
[0109] Based on any of the above embodiments, an embodiment of the present invention provides an interactive copywriting generation method based on a large model, which can be used for copywriting such as stories, essays, news, blogs, public accounts, and social media. The method includes:
[0110] Step S1: Text generation and editing
[0111] Figure 2 This is one of the flow charts of text copy generation and editing provided by the present invention, such as Figure 2 As shown in the figure, images are multiple images uploaded by the user during the first round of creation, and query1 is the creation requirements (i.e., initial requirement description) entered by the user during the first round of creation, such as subject matter, word count, etc. By inputting images and query1 into the multimodal large model, the first round of creation result text1 (i.e., the first draft of the copy) output by the model can be obtained. Query2 is the modification requirement (i.e., modification requirement description) proposed by the user during the second round of interaction. After receiving the modification requirement description entered by the user, the multimodal large model modifies and creates based on the historical input images images, modification requirement description query2, and creation content text1, thereby generating the target copy text2.
[0112] Figure 3 This is the second flow chart of the text copy generation and editing process provided by the present invention, such as Figure 3 As shown, images1 is a plurality of images uploaded by the user during the first round of creation, and query1 is the creation requirements (i.e., the initial requirement description) input by the user during the first round of creation, such as subject matter, word count, etc. By inputting images and query1 into the multimodal large model, the first round of creation result text1 (i.e., the first draft of the copy) output by the model can be obtained. images2 is a plurality of images supplemented and uploaded by the user during the second round of interaction, and query2 is the modification requirements (i.e., the modification requirement description) such as continuation and overall rewriting proposed by the user during the second round of interaction. After receiving the plurality of images images2 uploaded by the user and the modification requirement description query2 input, the multimodal large model will modify and create according to the historical input image images1, the current input image images2, the modification requirement description query2, and the creation content text1, thereby generating the target copy text2.
[0113] Step S2: Image copy generation
[0114] Figure 4 This is one of the flow charts of image text generation provided by the present invention, such as Figure 4 As shown in the figure, the query is the text content entered by the user and the corresponding image is required to be generated. After receiving the query, the multimodal large model will apply the text content to generate text to obtain text. The text contains several image prompts to guide the image generation model to generate correct and coherent images. For example, the text can include the following content:
[0115] Prompt 1: Image content—a tortoise and a rabbit standing at the starting line; Image style—black and white sketch
[0116] Prompt 2: Image content—a rabbit sleeping by a tree; Image style—a black and white sketch
[0117] Prompt 3: Image content—The tortoise overtakes the sleeping hare; Image style—Black and white sketch
[0118] Prompt 4: Image content—The tortoise crosses the finish line before the hare; Image style—Black and white sketch
[0119] When a prompt is found in the retrieved text, the image generation model is called to generate the corresponding target image. Images are a series of content-related images generated based on the text input by the user.
[0120] Figure 5 This is the second flow chart of image text generation provided by the present invention. Figure 5 As shown, query1 is the creation requirement (i.e., the initial demand description) entered by the user during the first round of creation, such as subject matter, word count, etc. By inputting query1 into the multimodal large model, the first round of creation text text1 (i.e., the first draft of the copy) output by the model can be obtained. In this case, the user can continue to request the generation of several images based on the first round of creation content. Query2 is the instruction entered by the user to generate corresponding images based on the first round of creation. After receiving the instruction, the multimodal large model will generate text2 containing several image prompt texts based on the first round of creation text text1 and query2. When prompt is retrieved in text2, the image generation model will be called to generate the corresponding target image through the model. Images is a series of content-related images generated according to the instructions entered by the user and the first round of creation text.
[0121] Step S3: Graphic and text generation
[0122] Figure 6 This is one of the flow charts for generating graphic and text copy provided by the present invention, such as Figure 6 As shown, the query is the user's creative requirement, specifying the generation of text with both text and images. After receiving the query, the multimodal model generates text to produce a text. The text consists of multiple paragraphs of text, and where illustrations are needed, there are corresponding image prompts to guide the image generation model in generating accurate and coherent images. For example, the text might include the following: Text 1: One sunny morning, the hare and the tortoise decided to have a running race. With the attention of all the animals, they arrived at the starting line and waited for the sound of the starting gun.
[0123] Prompt 1: Image content—a tortoise and a rabbit standing at the starting line; Image style—black and white sketch
[0124] Text 2: The rabbit started out fast and ran very fast. It was not long before the tortoise was far behind. The rabbit felt that he had won the race, so he found a place to sleep under the shade of a tree.
[0125] Prompt 2: Image content—a rabbit sleeping by a tree; Image style—a black and white sketch
[0126] Text 3: While the rabbit was sleeping, the tortoise continued to crawl on the track and finally overtook the sleeping rabbit in the afternoon.
[0127] Prompt 3: Image content—The tortoise overtakes the sleeping hare; Image style—Black and white sketch
[0128] Text 4: When the rabbit woke up, it realized the tortoise was no longer behind it, so it raced to catch up. But just as it neared the finish line, it saw the tortoise again. The tortoise, determined and slow, continued its march toward its goal, successfully crossing the finish line just before the rabbit could catch up.
[0129] Prompt 4: Image content—The tortoise crosses the finish line before the hare; Image style—Black and white sketch
[0130] Text 5: This story tells us that no matter how slow we are, as long as we persevere, we can eventually reach our goal.
[0131] When a prompt is found in the retrieved text, the image generation model is invoked to generate the corresponding target image. Images are a series of context-sensitive images generated based on the text input by the user. Once the target image sequence is generated, the image prompts in the prompt text are replaced one by one with the corresponding target images, resulting in the final image-text copy.
[0132] Figure 7 This is the second flow chart of the graphic and text copy generation process provided by the present invention, such as Figure 7As shown, images1 represents the multiple images uploaded by the user during the first round of creation, and query1 represents the creative requirements (i.e., initial description of requirements) entered by the user during the first round of creation, such as subject matter and word count. Inputting images1 and query1 into the multimodal large model yields the model's output, text1 (i.e., the initial draft of the copy). In this case, the user can request continued creation with both text and images based on the first round of creation. query2 represents the user-entered instruction to generate a copy with both images and text based on the first round of creation. Upon receiving this instruction, the multimodal large model generates prompt text2 based on the first round of creation text1, the historical input images images1, and query2. text2 contains multiple paragraphs of text, and corresponding image prompts appear where illustrations are required. When prompts are found in text2, the image generation model is invoked to generate the corresponding target images. images is a series of content-related images generated based on the user's instructions and the first round of creation text. After the target image sequence is generated, the image prompt text prompt in the prompt text text2 can be replaced one by one with the corresponding target image to obtain the final image and text copy.
[0133] Based on any of the above embodiments, Figure 8 Schematic diagram of the structure of the text generation device provided by the present invention. Figure 8 As shown, the device includes:
[0134] The first acquisition unit 810 is configured to acquire multiple images input by a user and / or an initial requirement description for a document to be generated;
[0135] A document generation unit 820 is configured to generate a document based on a document generation model and using the multiple images and / or the initial requirement description to obtain a draft document;
[0136] The second obtaining unit 830 is configured to obtain a description of modification requirements for the first draft document input by the user;
[0137] The document modification unit 840 is configured to modify the document draft based on the document generation model and apply the modification requirement description, or apply the multiple images and the modification requirement description, to generate a target document.
[0138] The device provided by the embodiment of the present invention supports users to input multiple images and / or initial requirement descriptions for the copy to be generated, providing users with greater flexibility and diversity. By combining the requirement descriptions in the form of images and text, it is possible to more comprehensively understand the user's intentions, thereby generating a copy that better meets the user's expectations. Moreover, the present invention introduces a mechanism for user feedback and iterative modification. Users can modify the generated draft copy and guide the model to further optimize the copy by inputting and modifying the requirement description. This interactivity and iterativeness greatly improves the efficiency and accuracy of copy generation, while meeting the user's interactive needs for free editing. In addition, due to the support for multimodal input and interactive modification, users can participate more directly in the copy generation process, and can adjust and optimize the copy according to their own needs and preferences, thereby improving user satisfaction with the final generated copy.
[0139] Based on any of the above embodiments, the device further includes a modification unit, which is configured to:
[0140] Obtain multiple images currently input by the user and a description of modification requirements for the draft copy;
[0141] Based on the text generation model, the multiple images inputted by the user in the past, the multiple images inputted by the user currently, and the modification requirement description are applied to modify the text draft and generate a target text.
[0142] Based on any of the above embodiments, the device further includes an image copy generation unit, the image copy generation unit being configured to: obtain text content for the image copy to be generated, the text content being determined based on at least one of user input, the draft copy, and the target copy;
[0143] Based on the copywriting generation model, applying the text content, a plurality of image prompt texts are generated, wherein the image prompt texts include an image content description and an image style description;
[0144] Based on the image generation model, the multiple image prompt texts are applied to generate target image copy corresponding to the text content, and the target image copy includes multiple target images.
[0145] Based on any of the above embodiments, the device further includes a graphic copy generation unit, the graphic copy generation unit including: an acquisition subunit, configured to acquire text information for the graphic copy to be generated, the text information being determined based on at least one of user input, the draft copy, and the target copy;
[0146] A first generating subunit is configured to generate prompt text based on the text generation model and applying the text information, wherein the prompt text includes a text body and an image prompt text;
[0147] The second generating subunit is configured to generate a target image based on the image prompt text, and generate a target graphic copy by applying the text body and the target image.
[0148] Based on any of the above embodiments, the prompt text includes multiple text segments and multiple image prompt texts, and each image prompt text corresponds to at least one text segment; accordingly, the second generation subunit is specifically configured to: based on the image generation model, apply the multiple image prompt texts to generate multiple target images, and the target images correspond one-to-one to the image prompt texts;
[0149] The multiple image prompt texts in the prompt text are replaced with the corresponding multiple target images, and the target image-text copy is obtained based on the replaced prompt text.
[0150] Based on any of the above embodiments, when the text information includes the text input by the user and the draft copy, or when the text information includes the text input by the user and the target copy, the first generating subunit is specifically configured to:
[0151] Based on the text generation model, the prompt text is generated by applying the text information and multiple images input by the user in the past.
[0152] Figure 9 An example of a physical structure diagram of an electronic device is shown below. Figure 9 As shown, the electronic device may include: a processor 910, a communication interface 920, a memory 930, and a communication bus 940, wherein the processor 910, the communication interface 920, and the memory 930 communicate with each other via the communication bus 940. The processor 910 may call the logic instructions in the memory 930 to execute a text generation method, which includes: obtaining multiple images input by a user and / or an initial requirement description for a text to be generated; based on a text generation model, applying the multiple images and / or the initial requirement description to generate a text to obtain a first draft of the text; obtaining a modification requirement description for the first draft of the text input by the user; based on the text generation model, applying the modification requirement description, or applying the multiple images and the modification requirement description, to modify the first draft of the text to generate a target text.
[0153] In addition, the logic instructions in the above-mentioned memory 930 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the relevant technology, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0154] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the copy generation method provided by the above methods, which includes: obtaining multiple images input by the user and / or an initial requirement description for the copy to be generated; based on a copy generation model, applying the multiple images and / or the initial requirement description to generate the copy to obtain a draft copy; obtaining a modification requirement description for the draft copy input by the user; based on the copy generation model, applying the modification requirement description, or applying the multiple images and the modification requirement description to modify the draft copy to generate a target copy.
[0155] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the text generation method provided by the above-mentioned methods, the method comprising: obtaining multiple images input by a user and / or an initial requirement description for a text to be generated; based on a text generation model, applying the multiple images and / or the initial requirement description to generate a text to obtain a first draft of the text; obtaining a modification requirement description for the first draft of the text input by the user; based on the text generation model, applying the modification requirement description, or applying the multiple images and the modification requirement description, to modify the first draft of the text to generate a target text.
[0156] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0157] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0158] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A copywriting generation method, characterized in that: include: Obtain multiple images input by the user and / or an initial requirement description for the copy to be generated; Based on the copy generation model, the plurality of images and / or the initial requirement description are used to generate the copy to obtain a first draft of the copy; Obtaining a description of modification requirements for the first draft of the document input by the user; Based on the copy generation model, applying the modification requirement description, or applying the multiple images and the modification requirement description, modifying the copy draft to generate a target copy; The method further comprises: Acquiring text information for a graphic copy to be generated, wherein the text information is determined based on at least one of user input, the draft copy, and the target copy; Based on the copywriting generation model, applying the text information, generating prompt text, the prompt text including text body and image prompt text; Based on the image prompt text, a target image is generated, and the text body and the target image are applied to generate a target graphic copy; The prompt text includes multiple text segments and multiple image prompt texts, and each image prompt text corresponds to at least one text segment; The step of generating a target image based on the image prompt text and applying the text body and the target image to generate a target graphic copy includes: Based on the image generation model, applying the multiple image prompt texts to generate multiple target images; The multiple image prompt texts in the prompt text are replaced with the corresponding multiple target images, and the target image-text copy is obtained based on the replaced prompt text.
2. The method for generating a copywriting according to claim 1, wherein: The method further includes: generating a copy based on the copy generation model by applying the plurality of images and / or the initial requirement description to obtain a first draft of the copy; and then: Obtain multiple images currently input by the user and a description of modification requirements for the draft copy; Based on the text generation model, the multiple images inputted by the user in the past, the multiple images inputted by the user currently, and the modification requirement description are applied to modify the text draft and generate a target text.
3. The method for generating a copywriting according to claim 1, wherein: Also includes: Obtaining text content for a to-be-generated image copy, wherein the text content is determined based on at least one of a user input, the draft copy, and the target copy; Based on the copywriting generation model, applying the text content, a plurality of image prompt texts are generated, wherein the image prompt texts include an image content description and an image style description; Based on the image generation model, the multiple image prompt texts are applied to generate target image copy corresponding to the text content, and the target image copy includes multiple target images.
4. The method for generating a copywriting according to any one of claims 1 to 3, wherein: The target image corresponds to the image prompt text one by one.
5. The method for generating a copywriting according to any one of claims 1 to 3, characterized in that: In a case where the text information includes the text input by the user and the draft copy, or in a case where the text information includes the text input by the user and the target copy, generating the prompt text based on the copy generation model and applying the text information includes: Based on the text generation model, the prompt text is generated by applying the text information and multiple images input by the user in the past.
6. A copywriting generation device, characterized in that: include: A first acquisition unit is configured to acquire multiple images input by a user and / or an initial requirement description for a document to be generated; A copy generation unit, configured to generate a copy based on a copy generation model and applying the plurality of images and / or the initial requirement description to obtain a first draft of the copy; A second acquiring unit is configured to acquire a description of modification requirements for the first draft of the document input by the user; A document modification unit, configured to modify the document draft based on the document generation model and applying the modification requirement description, or applying the multiple images and the modification requirement description, to generate a target document; The device further includes a graphic and text copy generation unit, and the graphic and text copy generation unit includes: an acquisition subunit, configured to acquire text information for the graphic copy to be generated, wherein the text information is determined based on at least one of user input, the draft copy, and the target copy; A first generating subunit is configured to generate prompt text based on the text generation model and applying the text information, wherein the prompt text includes a text body and an image prompt text; A second generating subunit is configured to generate a target image based on the image prompt text, and generate a target graphic copy by applying the text body and the target image; The prompt text includes multiple text segments and multiple image prompt texts, each of the image prompt texts corresponds to at least one text segment; accordingly, the second generating subunit is specifically configured to: Based on the image generation model, applying the multiple image prompt texts to generate multiple target images; The multiple image prompt texts in the prompt text are replaced with the corresponding multiple target images, and the target image-text copy is obtained based on the replaced prompt text.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the document generation method according to any one of claims 1 to 5 is implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the document generation method according to any one of claims 1 to 5 is implemented.
9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the document generation method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Cboth generation method and related device, electronic equipment and storage medium
CN116611401A
Image-text content generation method and device, equipment and storage medium
CN117032869A
Article editing automatic illustration method, device and equipment based on AIGC and storage medium
CN117078802A
Cited By
Multi-scene bird image-oriented vertical large model copywriting generation method
CN122113869A
A method for generating text for large vertical models of bird images in multiple scenes
CN122113869B