Poster image generation method and device, equipment, medium and product
By generating independent text element and image element layers during the poster generation process and superimposing them reasonably, the problems of high computing resource consumption and poor effect in generating high-quality poster images in the existing technology are solved, and low-cost and efficient poster image generation is achieved.
Patent Information
- Application Number
- CN202510749467.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies are unable to generate high-quality poster images, and training artificial intelligence models specifically for poster image generation requires a lot of computing resources and complex processes.
By determining the initial prompt words for poster generation, independent layers of text elements and image elements are generated, and they are superimposed in non-overlapping positions to avoid covering phenomena. The poster image is generated using a large language model and a large image generation model.
Generate high-quality, controllable poster images without model training, reduce computing resource costs, ensure reasonable graphic and text layout, improve the aesthetics of image elements and text clarity, expand user range and increase product stickiness.
Smart Images

Figure CN120672907A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, specifically to technical fields such as large models, generative formulas, generative models, Transformers, computer graphics, artificial intelligence, virtual images, image synthesis and cloud computing, and especially to a method, device, equipment, medium and product for generating poster images. Background Art
[0002] In recent years, with the continuous development of deep learning technology, images generated by AI models have significantly improved in terms of visual quality, diversity, and content control, greatly enhancing the inspiration and efficiency of people's daily creations. Among them, poster image generation, as a novel and unique application area, has attracted widespread attention.
[0003] However, poster image generation differs significantly from typical image generation. It often requires the inclusion of eye-catching promotional slogans, and the generated slogans must be properly laid out with the generated image content. Typical AI models are unable to generate high-quality poster images.
[0004] In order to solve the above problems, the existing technology is to train an artificial intelligence model specifically for poster image generation and apply it to the generation of poster images. Summary of the Invention
[0005] The present disclosure provides a method, apparatus, device, medium, and product for generating a poster image, which are used to reduce the cost of computing resources required for generating the poster image.
[0006] According to one aspect of the present disclosure, a method for generating a poster image is provided, comprising:
[0007] Generate initial prompt words based on the obtained poster, and determine the text elements included in the poster image to be generated;
[0008] Generating a text element layer based on the text element and the text element position corresponding to the text element, and generating an image element layer based on the initial prompt word and the image element position generated on the poster; wherein there is no overlapping position between the text element position and the image element position;
[0009] The text element layer and the image element layer are superimposed to obtain the poster image to be generated.
[0010] According to another aspect of the present disclosure, there is provided a device for generating a poster image, comprising:
[0011] A text element determination module is used to generate initial prompt words based on the obtained poster and determine the text elements included in the poster image to be generated;
[0012] an element layer generation module, configured to generate a text element layer based on the text element and the text element position corresponding to the text element, and generate an image element layer based on the initial prompt word and the image element position generated by the poster; wherein the text element position and the image element position do not overlap;
[0013] The poster image generation module is used to perform superposition processing on the text element layer and the image element layer to obtain the poster image to be generated.
[0014] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0015] at least one processor; and
[0016] a memory communicatively connected to at least one processor; wherein,
[0017] The memory stores instructions that can be executed by at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform any method in the present disclosure.
[0018] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute any one of the methods of the present disclosure.
[0019] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program executes any one of the methods of the present disclosure when executed by a processor.
[0020] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0022] Figure 1A is a flowchart of a method for generating some poster images according to an embodiment of the present disclosure;
[0023] Figure 1B It is a schematic diagram of the covering phenomenon generated by some text elements and image elements disclosed in the embodiment of the present disclosure;
[0024] Figure 1C is a schematic diagram of some poster images to be generated according to an embodiment of the present disclosure;
[0025] Figure 2 is a flowchart of another method for generating poster images according to an embodiment of the present disclosure;
[0026] Figure 3A is a flowchart of another method for generating poster images according to an embodiment of the present disclosure;
[0027] Figure 3B It is a schematic diagram of the process of generating some poster images according to the embodiments of the present disclosure;
[0028] Figure 4 is a schematic structural diagram of some poster image generation devices disclosed in an embodiment of the present disclosure;
[0029] Figure 5 It is a block diagram of an electronic device used to implement the method for taking a photo with a virtual image disclosed in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0030] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0031] Typical AI models are unable to generate high-quality poster images. To address this issue, existing approaches involve training an AI model specifically for poster image generation. For example, an encoder is first trained for text understanding, and then a new AI model is retrained to support slogan generation. Rotational position encoding and additional typographic constraints are used to improve control over the slogan's layout.
[0032] However, training an artificial intelligence model specifically for poster image generation requires a large amount of high-quality data and involves complex processes such as multi-stage training and model fine-tuning, which consumes extremely large computing resources.
[0033] Figure 1A This is a flowchart of some poster image generation methods disclosed in embodiments of the present disclosure. This embodiment is applicable to generating poster images without the need for model training. This method can be executed by the poster image generation device disclosed in embodiments of the present disclosure. The device can be implemented using software and / or hardware and can be integrated into any electronic device with computing capabilities, such as a server.
[0034] like Figure 1A As shown, the method for generating a poster image disclosed in this embodiment may include:
[0035] S101: Generate initial prompt words based on the obtained poster, and determine text elements included in the poster image to be generated.
[0036] The initial poster generation prompt refers to the original, unoptimized poster generation prompt entered by the user. The poster generation prompt is the core instruction text the user enters into the mapping system, describing key information such as the desired poster theme, style, elements, or composition. Essentially, it translates human design intent into a machine-understandable parameter combination. For example, the initial poster generation prompt includes, but is not limited to, "Help me generate a poster with a HD map theme."
[0037] The poster image to be generated refers to the final poster image generated by the algorithm after the user inputs the initial poster generation prompt words into the drawing system. The poster image to be generated includes text elements and image elements. Text elements refer to the text content used to convey the core information in the poster image, including the text itself and its visual representation. They are not only information carriers but also key design components that form the visual hierarchy of the poster image and strengthen the theme. For example, text elements can be promotional slogans designed with seal carving effects or calligraphy fonts.
[0038] The user generates initial prompt words for poster generation according to his / her own poster needs, and sends a poster image generation request including the initial prompt words for poster generation to the drawing system.
[0039] In one embodiment, the drawing system parses the poster image generation request to obtain initial prompt words for poster generation, further calls the large language model to perform semantic induction on the initial prompt words for poster generation, and determines the text elements included in the poster image to be generated based on at least one semantic information output by the large language model.
[0040] In another embodiment, the drawing system parses a poster image generation request to obtain initial prompt words for poster generation, and further generates at least one candidate text set based on the initial prompt words. Each candidate text set is evaluated to determine a quality score corresponding to each candidate text set; wherein the evaluation dimensions include at least one of visual harmony, semantic relevance, and aesthetic score. Based on the quality score, a low-scoring text set and a high-scoring text set are determined from each candidate text set. The low-scoring text set is semantically replaced to generate a replacement text set. Furthermore, based on the replacement text set and the high-scoring text set, text elements to be included in the poster image to be generated are determined.
[0041] In another embodiment, the drawing system parses the poster image generation request to obtain initial prompt words for poster generation, further performs keyword extraction on the initial prompt words to obtain festival keywords, brand keywords, or core slogan keywords, and then determines the text elements to be included in the poster image to be generated based on the festival keywords, brand keywords, or core slogan keywords. For example, if the extracted festival keyword is "May Fourth Youth Day" and the core slogan keyword is "Youth Power", then "May Fourth Youth Day" and "Youth Power" are output as the text elements to be included in the poster image to be generated.
[0042] S102: Generate a text element layer according to the text elements and the text element positions corresponding to the text elements, and generate an image element layer according to the initial prompt words and the image element positions generated on the poster.
[0043] The text element position refers to the spatial coordinates and layout of the text elements within the generated poster image. Essentially, it defines the positioning rules for design elements within a two-dimensional plane, directly impacting the visual focus and information delivery efficiency of the generated poster image. A text element layer is a separate layer specifically used to store and edit text elements. By separating text elements from image elements, editing within the text element layer does not affect the content of other image element layers, enabling non-destructive editing.
[0044] The image element position refers to the spatial coordinates and layout of the image elements in the poster image to be generated, which directly affects the information transmission efficiency and visual aesthetics of the poster image to be generated. Its essence is to achieve visual guidance and information priority control through composition rules. The image element layer refers to an independent layer specifically used to store and edit image elements. It separates image elements from text elements, so that when the image element layer edits the image element, it will not affect the content of the text element layer, thus achieving non-destructive editing. There is no overlapping position between the text element position and the image element position, that is, there is no element covering between the text element and the image element.
[0045] Figure 1B is a schematic diagram of some text elements and image elements covering each other according to the embodiment of the present disclosure, such as Figure 1B As shown, 100 represents the poster image to be generated, wherein 101 represents the text element included in the poster image 100, and 102 represents the image element included in the poster image 100. It can be seen that there is an obstruction between the text element 101 and the image element 102, which significantly reduces the aesthetics of the image element 102 and the clarity of the text element 101.
[0046] Therefore, in order to ensure that the layout of the graphics and text of the poster image to be generated is reasonable, it is necessary to avoid overlapping between the text elements and the image elements. Therefore, in this embodiment, the positions of the text elements and the image elements are set to have no overlapping positions.
[0047] In one embodiment of S102, the image position of the text element in the poster image to be generated is determined as the text element position, and a text element layer is generated based on the text element and its corresponding text element position. Furthermore, the image positions in the poster image to be generated other than the text element position are used as image element positions, and an image element layer is generated based on the initial prompt words and image element positions for poster generation.
[0048] Regarding the generation of the text element layer in S102, in one embodiment, the drawing system parses the initial prompt words for poster generation to determine whether the initial prompt words for poster generation contain the expected position information of the text element. If so, the text element position is determined based on the expected position information. For example, assuming that the initial prompt words for poster generation include "text is located directly above the poster", the expected position information of the text element can be determined to be "directly above the poster". Further, based on the expected position information "directly above the poster", an appropriate image position is determined from the poster image to be generated as the text element position. The drawing system generates a blank layer based on the image size of the poster image to be generated, and generates text elements in the blank layer based on the text elements and text element positions to obtain a text element layer.
[0049] Regarding the generation of the text element layer in S102, in another embodiment, the drawing system parses the initial prompt words for poster generation to determine whether the initial prompt words for poster generation contain the expected position information of the text element. If not, based on the number of text elements of the determined text elements, a target image position that matches the number of text elements is determined from pre-set candidate image positions as the text element position. For example, assuming the number of text elements is three, a candidate image position that satisfies the number of three text elements is determined from pre-set candidate image positions as the text element position. The drawing system generates a blank layer based on the image size of the poster image to be generated, and generates text elements in the blank layer based on the text elements and text element positions to obtain a text element layer.
[0050] For example, if the image size of the poster image to be generated is A*B, a blank layer with the same image size of A*B is generated. If the determined text element is "Life is Beautiful" and the text element position is at image position "x1-x2, y1-y2", the text element "Life is Beautiful" is generated at image position "x1-x2, y1-y2" on the blank layer, thus obtaining a text element layer.
[0051] Regarding the generation of the image element layer in S102, in one embodiment, the drawing system generates a blank layer according to the image size of the poster image to be generated, and determines the text element area in the blank layer according to the text element position, and then uses the image position other than the text element position as the image element position, and further determines the image element area in the blank layer according to the image element position.
[0052] Furthermore, the drawing system generates a first prompt word for the text element area and directly uses the initial prompt word generated by the poster as the second prompt word. The drawing system calls the image generation model and inputs the blank layer, the first prompt word, and the second prompt word into the image generation model.
[0053] The image generation model generates a first layer slice that does not contain any elements in the text element area of the blank layer according to the first prompt word. For example, the first prompt word can be "only fill color, do not generate any elements".
[0054] The image generation model generates image elements in the image element area of the blank layer according to the second prompt word to obtain a second layer slice. For example, the second prompt word can be "Generate a poster image according to the XXX prompt word".
[0055] The image generation model obtains and outputs the image element layer according to the first layer slice and the second layer slice. The drawing system obtains the image element layer output by the image generation model.
[0056] S103: Overlay the text element layer and the image element layer to obtain a poster image to be generated.
[0057] In one embodiment, the drawing system overlays the text element layer onto the image element layer to obtain the poster image to be generated, that is, the poster image to be finally output.
[0058] Figure 1C Schematic diagram of some poster images to be generated according to the embodiment of the present disclosure, which are obtained by using the poster image generation method provided by the embodiment of the present disclosure. Figure 1C As shown, 103 represents the poster image to be generated, wherein 104 represents the text element included in the poster image 103 to be generated, and 105 represents the image element included in the poster image 103 to be generated. It can be seen that there is no overlap between the text element 104 and the image element 105, and the layout of the text and graphics is reasonable.
[0059] The present disclosure generates initial prompt words based on an acquired poster to determine the text elements included in the poster image to be generated; generates a text element layer based on the text elements and the text element positions corresponding to the text elements; and generates an image element layer based on the initial prompt words and image element positions generated on the poster; wherein there is no overlap between the text element positions and the image element positions; and superimposes the text element layer and the image element layer to obtain the poster image to be generated. The beneficial effects are:
[0060] First, high-quality and controllable poster images can be generated without model training, which reduces the computing resource cost required for poster generation and achieves the effect of low-cost poster image generation.
[0061] Secondly, by controlling the position of text elements and image elements so that there is no overlapping, the phenomenon of covering between text elements and image elements in the generated poster image can be avoided, thereby ensuring the rationality of the graphic and text layout in the poster image, and further ensuring the beauty of the image elements in the poster image and the clarity of the text elements.
[0062] Thirdly, it has expanded the user range and greatly increased the number of times users use the poster generation service and the retention rate, increased the stickiness of the product, and improved the core competitiveness of the product.
[0063] Figure 2 It is a flowchart of other poster image generation methods disclosed in the embodiments of the present disclosure, which can be used to further optimize and expand the above technical solution and can be combined with the above optional implementation methods.
[0064] like Figure 2 As shown, the method for generating a poster image disclosed in this embodiment may include:
[0065] S201: Generate initial prompt words according to the obtained poster, and determine text elements included in the poster image to be generated.
[0066] S202: Determine the image position of the text element in the poster image to be generated as the text element position, and generate a text element layer according to the text element and the text element position.
[0067] S203: Using the image positions other than the text element positions in the poster image to be generated as image element positions.
[0068] S204: Generate a blank layer according to the image size of the poster image to be generated, and determine the text element area in the blank layer according to the text element position, and determine the image element area in the blank layer according to the image element position.
[0069] Among them, a blank layer refers to an independent layer that can be freely edited and does not contain any pixel content.
[0070] In one embodiment, the drawing system generates a blank layer of the same size as the poster image to be generated based on the image size of the poster image to be generated. Furthermore, the blank layer is divided into regions based on the positions of text elements to obtain text element regions, and the blank layer is divided into regions based on the positions of image elements to obtain image element regions.
[0071] S205 , generating initial prompt words and image element areas according to the poster, generating image elements in a blank layer, and obtaining an image element layer.
[0072] In one embodiment, the drawing system calls the image generation model, inputs the poster generation initial prompt words and the blank layer into the image generation model, and controls the image generation model to generate image elements in the image element area of the blank layer according to the poster generation initial prompt words, and controls the image generation model to only fill the text element area of the blank layer with color without generating any elements, thereby generating an image element layer, and then obtaining the image element layer output by the image generation model.
[0073] By generating a blank layer according to the image size of the poster image to be generated, determining the text element area in the blank layer according to the position of the text element, and determining the image element area in the blank layer according to the position of the image element; generating initial prompt words and image element areas according to the poster, generating image elements in the blank layer, and obtaining an image element layer, the beneficial effects are:
[0074] First, using the image element area as a generation constraint can guide the image generation model to generate image elements within the specified area, ensuring the separation of image elements and text elements, and ensuring the rationality of the poster image and text layout.
[0075] Secondly, the use of independently generated image element layers, combined with the principle of non-destructive editing, allows image elements to be adjusted individually without affecting the content of text elements, thereby improving the flexibility of later modifications.
[0076] Thirdly, the initial prompt words generated by the poster are bound to the image element area to achieve precise semantic-visual mapping and ensure the accuracy and credibility of image element generation.
[0077] Optionally, generating an initial prompt word and an image element area according to the poster, generating image elements in a blank layer, and obtaining an image element layer, including:
[0078] S2051. Add a mask to the text element area as a mask area, and use the image element area as a non-mask area, and determine a mask layer according to the mask area and the non-mask area.
[0079] A mask is a special matrix or binary / grayscale image used to selectively control the scope and intensity of image processing operations. Its core function is to mark the pixel areas that need to be processed or retained, and to block out irrelevant areas.
[0080] In one embodiment, the drawing system generates a mask based on the area size of the text element area and multiplies the mask by the image matrix of the text element area to add the mask to the text element area. The masked text element area is then treated as the masked area. No mask is added to the image element area, which is treated as the non-masked area. Furthermore, a mask layer is determined based on the image stitching result of the masked and non-masked areas.
[0081] S2052: Obtain a first area prompt word corresponding to the masked area and a second area prompt word corresponding to the non-masked area.
[0082] The first region prompt word represents the region prompt word for the masked area; the second region prompt word represents the region prompt word for the unmasked area. Region prompt words are a technical solution used within the image generation model to control image content generation by region. Their core function is to achieve precise composition control by using spatial segmentation and conditional settings to enable different regions of the image content to respond to different text descriptions. The second region prompt word is the initial prompt word for poster generation.
[0083] S2053: Input the mask layer, the first area prompt word, and the second area prompt word into the image generation model, so that the image generation model outputs the image element layer.
[0084] Large image generation models are deep learning models with massive parameter counts, primarily used to generate high-quality images from text descriptions. They typically employ a Transformer-based architecture, efficiently processing and understanding complex language descriptions and converting them into visual information. Typically, large image generation models can have parameters in the hundreds of billions.
[0085] In one embodiment, the drawing system calls the image generation model and inputs the mask layer, the first region prompt word, and the second region prompt word into the image generation model.
[0086] The image generation model generates a first layer slice containing no elements in the masked area based on the first region prompt. Based on the second region prompt, the image generation model generates image elements in the unmasked area to obtain a second layer slice. The image generation model generates and outputs an image element layer based on the first and second layer slices. The drawing system obtains the image element layer output by the image generation model.
[0087] Optionally, the prompt word for the first area may be “only fill color, do not generate any elements”, etc. The second prompt word may be “generate a poster image according to the prompt word XXX”, etc.
[0088] By adding a mask to the text element area as a mask area, and using the image element area as a non-mask area, and determining a mask layer based on the mask area and the non-mask area; obtaining a first area prompt word corresponding to the mask area, and a second area prompt word corresponding to the non-mask area; wherein the second area prompt word is the initial prompt word for poster generation; and inputting the mask layer, the first area prompt word, and the second area prompt word into the image generation model, so that the image generation model outputs the image element layer, the beneficial effects are:
[0089] First, by using different area prompt words for masked areas and non-masked areas, the image content in different areas responds to different text descriptions, achieving precise composition control.
[0090] Secondly, by using a dedicated first-area prompt word in the mask area, it is ensured that the mask area does not contain any elements, thereby reserving the position of the text element and ensuring that the text elements will not be covered when subsequent layers are superimposed; by using a dedicated second-area prompt word in the non-mask area, it is ensured that image elements are generated in the non-mask area. On the premise of ensuring that the image elements do not cover the text elements, image elements that meet user needs are generated based on the second area prompt (initial prompt word generated by the poster).
[0091] Thirdly, masking technology is used to clearly distinguish between text element areas and image element areas, avoiding the problem of cross-contamination between text and image elements in the process of generating image element layers in the large image generation model.
[0092] Fourthly, independent area prompts can be provided without affecting the main body of the poster image, thus reducing the cost of iterative optimization.
[0093] Optionally, when the image generation model is a diffusion model, the image generation model obtains the image element layer in the following manner:
[0094] A. During the image denoising process, the image generation model calculates a first attention representation based on the first text feature of the prompt word in the first region and the first image feature of the first noisy image, and denoises the first noisy image based on the first attention representation to generate a first layer slice.
[0095] Diffusion models are a broad class of image generation models that generate data by simulating the physical diffusion process. Their core concept is to gradually corrupt the original data by adding noise (forward diffusion) and then learn the inverse denoising process to reconstruct high-quality samples (reverse diffusion). In other words, the image generation process of the diffusion model involves both image denoising and image denoising.
[0096] The first text feature refers to the text feature obtained by extracting features from the prompt word in the first region; the first image feature refers to the image feature obtained by extracting features from the first noise image. The first noise image is the noise image corresponding to the masked region, i.e., the noise image obtained by adding noise to the masked region using the diffusion model during the image noise addition process. Image noise addition can be performed by adding Gaussian noise, for example.
[0097] In one embodiment, the drawing system invokes a diffusion model and inputs the mask layer, the first region cue word, and the second region cue word into the diffusion model. During the image denoising process, the diffusion model determines a first text feature based on the first region cue word and a first image feature based on the first noisy image. The model then calculates a first attention representation based on the first text feature and the first image feature. The model then denoises the first noisy image based on the first attention representation to generate a first layer slice.
[0098] B. During the image denoising process, the image generation model calculates a second attention representation based on the second text features of the second region prompt word and the second image features of the second noisy image, and denoises the second noisy image based on the second attention representation to generate a second layer slice.
[0099] The second text feature refers to the text feature obtained by extracting features from the prompt word in the second region; the second image feature refers to the image feature obtained by extracting features from the second noise image. The second noise image is the noise-added image corresponding to the unmasked region, i.e., the noise image obtained by adding noise to the unmasked region using the diffusion model during the image noise addition process. Image noise addition can be performed by adding Gaussian noise, for example.
[0100] In one embodiment, the drawing system invokes a diffusion model and inputs the mask layer, the first region cue word, and the second region cue word into the diffusion model. During the image denoising process, the diffusion model determines second text features based on the second region cue word and second image features based on the second noisy image. The model then calculates a second attention representation based on the second text features and the second image features. The second noisy image is denoised using the second attention representation to generate a second layer slice.
[0101] C. The image generation model obtains the image element layer based on the first layer slice and the second layer slice.
[0102] In the image denoising process, the existing diffusion model often denoises the entire noisy image. For the scene of poster image generation, the problem of overlapping between text elements and image elements often arises. In the embodiment of the present invention, during the image denoising process, the large image generation model calculates a first attention representation based on the first text feature of the prompt word in the first area and the first image feature of the first noise image, and denoises the first noise image based on the first attention representation to generate a first layer slice; during the image denoising process, the large image generation model calculates a second attention representation based on the second text feature of the prompt word in the second area and the second image feature of the second noise image, and denoises the second noise image based on the second attention representation to generate a second layer slice. This can more accurately control the correspondence between the prompt word and the noise image, ensure the accurate generation of the content of each area, avoid interference between irrelevant areas, and ensure the aesthetic layout of the poster image.
[0103] S206: Overlay the text element layer and the image element layer to obtain a poster image to be generated.
[0104] Figure 3A It is a flowchart of other poster image generation methods disclosed in the embodiments of the present disclosure, which can be used to further optimize and expand the above technical solution and can be combined with the above optional implementation methods.
[0105] like Figure 3A As shown, the method for generating a poster image disclosed in this embodiment may include:
[0106] S301: Input the initial prompt words generated by the poster into the large language model, expand the prompt words of the initial prompt words generated by the poster through the large language model, and output the expanded prompt words of the poster generation.
[0107] Among them, a large language model refers to a deep learning model with a huge parameter scale, which can generate natural language text or understand the meaning of language text. The large language model can handle a variety of natural language tasks, such as text classification, question and answer, dialogue, etc. Under normal circumstances, the parameter scale of the large language model can even reach hundreds of billions.
[0108] Prompt word expansion refers to automatically expanding the initial prompt words generated by the poster input by the user into a prompt word combination that contains a complete description of style, details, parameters, etc., which is used to guide the drawing system to generate more accurate and high-quality poster images.
[0109] In one embodiment, the drawing system inputs the initial prompt words and the extended prompt words for poster generation into the large language model. The large language model expands the initial prompt words for poster generation according to the extended prompt words and outputs the extended prompt words for poster generation.
[0110] For example, suppose the initial prompt word for poster generation is "Help me generate a poster with the theme of high-definition maps."
[0111] The large language model generates initial prompt words based on the poster, and the output poster generates extended prompt words such as "A high-definition, highly detailed map poster that shows rich geographical and topographical features. The map accurately marks major cities and famous landmarks. The overall style is both modern and incorporates traditional aesthetic elements, such as ink textures, gold embellishments, or elegant calligraphy fonts to mark province names. The background can use a soft rice paper texture or a deep blue and gold gradient to enhance the visual beauty. Mountains and plateaus use realistic shadows and relief effects to show the undulating terrain. Key landmarks can be embellished with exquisite small illustrations or icons to enhance the cultural charm. The overall layout is clear and the visual hierarchy is distinct, which not only ensures the practicality of the map, but also has artistic appreciation value. It is suitable for education, decoration or reference purposes. There are no redundant elements to ensure that the picture is clean, professional and advanced." etc.
[0112] S302: Determine the text elements included in the poster image to be generated according to the poster generation extended prompt word, and determine the image position of the text elements in the poster image to be generated as the text element position.
[0113] By inputting the initial prompt words for poster generation into a large language model, expanding the initial prompt words for poster generation using the large language model, and outputting the extended prompt words for poster generation; and determining the text elements included in the poster image to be generated based on the extended prompt words for poster generation, the beneficial effects are:
[0114] First, a large language model is used to parse the poster to generate the fuzzy requirements in the initial prompt words, which are automatically expanded into actionable visual element descriptions, thereby improving the semantic consistency between text and images and improving the generation quality of poster images.
[0115] Secondly, compared with manual expansion of prompt words, this solution can improve the efficiency of prompt word expansion and further improve the efficiency of poster image generation. It is suitable for scenarios where standardized poster images need to be produced quickly.
[0116] Optionally, determining text elements included in the poster image to be generated based on the extended prompt words for poster generation includes:
[0117] S3021: Input the poster generation extension prompt words into the large language model, use the large language model to perform semantic induction on the poster generation extension prompt words, and output at least one semantic information.
[0118] Among them, semantic induction refers to the deep semantic abstraction of the extended prompt words generated by the poster, and the extraction of universal semantic features by analyzing their usage patterns in different contexts.
[0119] In one embodiment, the drawing system inputs the poster generation extension prompt words and the semantic induction prompt words into the large language model. The large language model performs semantic induction on the poster generation extension prompt words according to the semantic induction prompt words and outputs at least one semantic information.
[0120] S3022: Determine a poster title corresponding to the poster image to be generated according to each semantic information, and determine text elements according to the poster title.
[0121] The poster title is the core text element used to convey the core theme in the poster image. It has the characteristics of concise text, eye-catching font and concise semantics. Its essence is to quickly attract the audience's attention and convey the core information through the combination of text language and visual design.
[0122] In one embodiment, the drawing system extracts semantic information from each semantic information, obtains a target amount of semantic information as a poster title, and uses the poster title as a text element.
[0123] By inputting extended prompt words for poster generation into a large language model, using the large language model to perform semantic induction on the extended prompt words for poster generation, and outputting at least one semantic information; determining the poster title corresponding to the poster image to be generated based on each semantic information, and determining the text elements based on the poster title, the beneficial effects are:
[0124] First, by using a large language model to perform semantic induction on the extended prompt words generated by the poster, it is possible to eliminate ambiguous information in the extended prompt words generated by the poster and ensure the semantic accuracy of the final text elements.
[0125] Secondly, by using the poster title determined based on semantic information as a text element, the semantic consistency between the text element and the prompt word is guaranteed, ensuring the accuracy of poster generation.
[0126] Optionally, determining a poster title corresponding to the poster image to be generated according to each semantic information includes:
[0127] S30221. Determine the feature information corresponding to each piece of semantic information.
[0128] Feature information includes at least one of text length, keyword density, and sentiment intensity. Text length refers to the number of characters in the semantic information. Keyword density refers to the ratio of the frequency of occurrence of core semantic words and their associated words in the semantic information to the total text volume. Sentiment intensity refers to the intensity of the emotional expression in the semantic information and is a core indicator for measuring the magnitude of emotional tendencies in sentiment analysis.
[0129] S30222. Calculate weight information corresponding to each piece of semantic information based on the feature information, and determine at least one main poster title and at least one subtitle of the poster from each piece of semantic information based on the weight information.
[0130] The poster title consists of a main title and a subtitle. The main title is the core information carrier of the poster image and is usually placed in the most eye-catching position on the page. The subtitle is used to explain the main title, provide additional details, or emphasize the main feature. Weight information is used to quantify the suitability of each semantic information as the main title of the poster.
[0131] It can be understood that when the text length of any semantic information is longer, it is more suitable as a poster subtitle rather than a poster main title, so the weight information of the semantic information is smaller; when the keyword density of any semantic information is higher, it is more suitable as a poster main title rather than a poster subtitle, so the weight information of the semantic information is greater; when the emotional intensity of any semantic information is higher, it is more suitable as a poster main title rather than a poster subtitle, so the weight information of the semantic information is greater.
[0132] In one embodiment, a weighted calculation is performed based on the text length, keyword density and sentiment intensity of each semantic information to determine the weight information corresponding to each semantic information, and based on the sorting results of the weight information corresponding to each semantic information, at least one main title of the poster and at least one subtitle of the poster are determined from each semantic information.
[0133] By determining feature information corresponding to each semantic information, wherein the feature information includes at least one of text length, keyword density, and sentiment intensity; calculating weight information corresponding to each semantic information based on the feature information, and determining at least one main title and at least one subtitle of the poster from each semantic information based on the weight information, the beneficial effects are:
[0134] First, by calculating keyword density, core semantic information is automatically identified to ensure that the main title of the poster directly reflects the content topic that users are most concerned about.
[0135] Secondly, considering the emotional intensity, semantic information with high emotional intensity should be given priority as the main title of the poster to enhance its appeal.
[0136] Thirdly, short text semantic information should be selected as the main title of the poster to highlight the core information and comply with the principle of visual focus.
[0137] Fourthly, the weight information of semantic information is calculated through feature information and used to screen the main title and subtitle of the poster, which can eliminate the influence of personal aesthetic preferences and ensure that the poster title is strictly aligned with the content semantics.
[0138] Optionally, the text element position is determined as follows:
[0139] Determine the number of main titles of the poster main title and the number of subtitles of the poster subtitle; determine the target image position from the candidate image positions based on the number of main titles and the number of subtitles, as the image position of the text element in the poster image to be generated; determine the text element position based on the image position of the text element in the poster image to be generated.
[0140] Among them, based on the number of different main titles and subtitles, and considering the aesthetics of the poster image design, the association relationship between the number of different main titles, the number of subtitles, and the candidate image position is pre-configured. For example, assuming there is one main title and two subtitles, considering the aesthetics of the poster image design, the candidate image position (x0-xn, y0-yn) is set to be associated with them. That is, when there is one main title and two subtitles, setting the poster main title and poster subtitle at the candidate image position (x0-xn, y0-yn) is more aesthetically pleasing.
[0141] In one embodiment, the drawing system determines the number of main titles and subtitles, matches the number of main titles and subtitles with pre-configured associations, and determines a target image position from the candidate image positions based on the matching results, which serves as the image position of the text element in the poster image to be generated. Furthermore, the image position of the text element in the poster image to be generated serves as the text element position corresponding to the text element.
[0142] By determining the number of main titles of the poster and the number of subtitles of the poster's subtitles; determining the target image position from the candidate image positions based on the number of main titles and the number of subtitles, as the image position of the text element in the poster image to be generated; determining the text element position based on the image position of the text element in the poster image to be generated, the aesthetics of the text element in the poster image to be generated can be guaranteed.
[0143] S303: Taking the image positions excluding the text element positions in the poster image to be generated as image element positions, and determining the current poster type corresponding to the poster image to be generated according to the poster generation extended prompt words.
[0144] Among them, the current poster types are mainly classified according to the usage scenarios and content attributes of the poster images, such as e-commerce posters, landscape posters, etc.
[0145] In one embodiment, the drawing system calls the large language model, inputs the poster generation extension prompt word and the poster type determination prompt word into the large language model, and the large language model determines the current poster type corresponding to the poster image to be generated based on the poster type determination prompt word.
[0146] S304: Determine a target poster image style corresponding to the poster image to be generated based on the current poster type and the association between the candidate poster types and the candidate poster image styles.
[0147] The poster image style refers to the overall aesthetic characteristics formed by the combination of visual elements in the poster image design. Based on the aesthetic considerations of the poster image design, the association between candidate poster types and candidate poster image styles is pre-configured.
[0148] In one embodiment, the current poster type is matched with the association relationship between the candidate poster types and the candidate poster image styles, and the target poster image style corresponding to the poster image to be generated is determined from the candidate poster image styles according to the matching result.
[0149] For example, assuming that the current poster type is "e-commerce poster", then the target poster image style can be "3D style"; for another example, assuming that the current poster type is "landscape poster", then the target poster image style can be "illustration style".
[0150] S305 , updating the poster generation extension prompt words according to the target poster image style to generate poster generation update prompt words, and generating an image element layer according to the poster generation update prompt words and the image element positions.
[0151] By determining the current poster type corresponding to the poster image to be generated based on the extended poster generation prompt words, and determining the target poster image style corresponding to the poster image to be generated based on the current poster type and the association between the candidate poster types and the candidate poster image styles; updating the extended poster generation prompt words based on the target poster image style to generate updated poster generation prompt words, and generating an image element layer based on the updated poster generation prompt words and the image element positions, the beneficial effects are:
[0152] First, the extended prompt words generated by the poster are used to automatically determine the current poster type, and the accurate mapping of the poster image style is achieved by combining the association relationship library. Compared with the traditional fixed template solution, the accuracy of poster image style adaptation is improved.
[0153] Secondly, prompt words are automatically added based on the target poster image style, which reduces the number of manual adjustments on the one hand and ensures the quality of poster image generation on the other.
[0154] S306: Determine the current layer color corresponding to the position of the text element in the image element layer, and determine the target text color corresponding to the text element based on the current layer color and the association relationship between the candidate layer colors and the candidate text colors.
[0155] Among them, in order to ensure that the text color of the text element is not close to the background color or too inconsistent with the background color, it is necessary to adaptively determine the target text color corresponding to the text element based on the current layer color corresponding to the text element position in the image element layer. For example, assuming that the text element position is (x1-x10, y1-y10), the current layer color corresponding to the position (x1-x10, y1-y10) in the image element layer is determined, and based on the current layer color corresponding to the position (x1-x10, y1-y10), the target text color corresponding to the text element is adaptively determined so that the target text color is compatible with the current layer color.
[0156] In one embodiment, the drawing system invokes an image understanding model and inputs an image element layer into the model. Based on the current layer color and the relationship between candidate layer colors and candidate text colors, the model determines and outputs the target text color corresponding to the text element. The drawing system then obtains the target text color output by the image understanding model.
[0157] S307 : Generate a text element layer according to the target text color, text element, and text element position, and perform superposition processing on the text element layer and the image element layer to obtain a poster image to be generated.
[0158] In one embodiment, the drawing system generates text elements in a blank layer according to the target text color, text elements, and text element positions to obtain a text element layer.
[0159] By determining the current layer color corresponding to the position of the text element in the image element layer, and based on the current layer color, as well as the correlation between the candidate layer color and the candidate text color, the target text color corresponding to the text element is determined; based on the target text color, the text element and the text element position, a text element layer is generated so that the target text color of the text element is not similar to the background color of the layer, thereby avoiding the problem of unclear text color; the target text color can also be adapted to the background color of the layer, thereby avoiding the problem of the target text color and the background color of the layer being separated, resulting in a problem of insufficient aesthetics of the picture.
[0160] Optionally, also include:
[0161] In response to a modification operation on a text element in a text element layer, a modified text element layer is generated; and the modified text element layer and the image element layer are superimposed to obtain a poster image to be generated.
[0162] In one embodiment, after a text element layer is generated, the user can modify the text elements in the text element layer according to their needs. In response to the modification operation on the text elements in the text element layer, the drawing system generates a modified text element layer and overlays the modified text element layer with the image element layer to obtain the poster image to be generated.
[0163] By responding to the modification operation of the text elements in the text element layer, a modified text element layer is generated; the modified text element layer and the image element layer are superimposed to obtain the poster image to be generated, so that the generation of the poster image supports the editing of text elements, thereby enhancing the flexibility and commercial application value of poster generation.
[0164] Figure 3B This is a schematic diagram of the process of generating some poster images according to the embodiments of the present disclosure, such as Figure 3B As shown:
[0165] The prompt word is expanded based on the poster generated initial prompt word 300 to determine the poster generated extended prompt word 301. The text element 302 is determined based on the poster generated extended prompt word 301. Based on the text element 302 and the determined text element position, the text element is generated in the blank layer 303 to obtain a text element layer 304.
[0166] Based on the poster generated extended prompt word 301 and the determined image element position, the image element is generated in the mask layer 305 to obtain the image element layer 306. Furthermore, based on the current layer color of the text element position in the image element layer 306, the target text color of the text element 302 is determined.
[0167] Finally, the text element layer 304 and the image element layer 306 are superimposed to obtain a poster image 307 to be generated.
[0168] The specific implementation of each step in the above process can refer to the relevant content in the embodiments of the present disclosure, and will not be repeated here.
[0169] Figure 4 This is a schematic diagram of the structure of some poster image generation devices disclosed in embodiments of the present disclosure. This device can be used to generate poster images without the need for model training. This device can be implemented using software and / or hardware and integrated into any electronic device with computing capabilities.
[0170] like Figure 4 As shown, the poster image generation device 40 disclosed in this embodiment may include a text element determination module 41, a text element layer generation module 42, an image element layer generation module 43, and a poster image generation module 44, wherein:
[0171] A text element determination module 41 is configured to generate initial prompt words based on the obtained poster and determine text elements included in the poster image to be generated;
[0172] An element layer generation module 42 is configured to generate a text element layer based on the text element and the text element position corresponding to the text element, and to generate an image element layer based on the initial prompt word and the image element position generated on the poster; wherein the text element position and the image element position do not overlap.
[0173] The poster image generating module 43 is configured to perform a superposition process on the text element layer and the image element layer to obtain the poster image to be generated.
[0174] Optionally, the element layer generation module 42 is specifically configured to:
[0175] generating a blank layer according to the image size of the poster image to be generated, and determining a text element area in the blank layer according to the position of the text element, and determining an image element area in the blank layer according to the position of the image element;
[0176] An initial prompt word and the image element area are generated according to the poster, and image elements are generated in the blank layer to obtain the image element layer.
[0177] Optionally, the element layer generation module 42 is further configured to:
[0178] Adding a mask to the text element area as a mask area, and using the image element area as a non-mask area, and determining a mask layer according to the mask area and the non-mask area;
[0179] Obtaining a first area prompt word corresponding to the masked area and a second area prompt word corresponding to the unmasked area; wherein the second area prompt word is an initial prompt word generated for the poster;
[0180] Inputting the mask layer, the first area prompt word, and the second area prompt word into an image generation model, so that the image generation model outputs the image element layer;
[0181] The image generation model generates a first layer slice that does not contain any elements in the mask area according to the first area prompt word;
[0182] The image generation model generates image elements in the non-masked area according to the second area prompt word to obtain a second layer slice;
[0183] The image generation model obtains the image element layer according to the first layer slice and the second layer slice.
[0184] Optionally, when the image generation model is a diffusion model, the image generation model obtains the image element layer in the following manner:
[0185] During the image denoising process, the image generation model calculates a first attention representation based on the first text feature of the prompt word in the first region and the first image feature of the first noise image, and denoises the first noise image based on the first attention representation to generate the first layer slice; wherein the first noise image is the noise-added image corresponding to the mask region;
[0186] During the image denoising process, the image generation model calculates a second attention representation based on the second text feature of the prompt word in the second area and the second image feature of the second noise image, and denoises the second noise image based on the second attention representation to generate a second layer slice; wherein the second noise image is the noise-added image corresponding to the non-masked area;
[0187] The image generation model obtains the image element layer according to the first layer slice and the second layer slice.
[0188] Optionally, the text element determination module 41 is specifically configured to:
[0189] Inputting the initial prompt words generated by the poster into a large language model, performing prompt word expansion on the initial prompt words generated by the poster using the large language model, and outputting the expanded prompt words generated by the poster;
[0190] The text elements included in the poster image to be generated are determined according to the poster generation extended prompt words.
[0191] Optionally, the text element determination module 41 is further configured to:
[0192] Inputting the poster generation extended prompt word into a large language model, performing semantic induction on the poster generation extended prompt word using the large language model, and outputting at least one semantic information;
[0193] The poster title corresponding to the poster image to be generated is determined according to each of the semantic information, and the text element is determined according to the poster title.
[0194] Optionally, the text element determination module 41 is further configured to:
[0195] Determining feature information corresponding to each of the semantic information; wherein the feature information includes at least one of text length, keyword density, and sentiment intensity;
[0196] Weight information corresponding to each piece of the semantic information is calculated according to the feature information, and at least one main title of the poster and at least one subtitle of the poster are determined from each piece of the semantic information according to the weight information.
[0197] Optionally, the position of the text element is determined by:
[0198] Determine the number of main titles of the poster main title and the number of subtitles of the poster subtitle;
[0199] determining a target image position from candidate image positions according to the number of main titles and the number of subtitles, as the image position of the text element in the poster image to be generated;
[0200] The position of the text element is determined according to the image position of the text element in the poster image to be generated.
[0201] Optionally, the element layer generation module 42 is further configured to:
[0202] Determining a current layer color corresponding to the position of the text element in the image element layer, and determining a target text color corresponding to the text element based on the current layer color and an association relationship between candidate layer colors and candidate text colors;
[0203] The text element layer is generated according to the target text color, the text element, and the text element position.
[0204] Optionally, the element layer generation module 42 is further configured to:
[0205] Inputting the initial prompt words generated by the poster into a large language model, performing prompt word expansion on the initial prompt words generated by the poster using the large language model, and outputting the expanded prompt words generated by the poster;
[0206] Determining a current poster type corresponding to the poster image to be generated based on the poster generation extended prompt word, and determining a target poster image style corresponding to the poster image to be generated based on the current poster type and an association between candidate poster types and candidate poster image styles;
[0207] The poster generation extension prompt word is updated according to the target poster image style to generate a poster generation update prompt word, and the image element layer is generated according to the poster generation update prompt word and the image element position.
[0208] Optionally, the device further includes a text element modification module, specifically configured to:
[0209] In response to a modification operation on the text element in the text element layer, generating a modified text element layer;
[0210] The modified text element layer and the image element layer are superimposed to obtain the poster image to be generated.
[0211] The poster image generation device 40 disclosed in the embodiment of the present disclosure can execute the poster image generation method disclosed in the embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method. For the contents not fully described in this embodiment, please refer to the description of the embodiment of the method disclosed in this disclosure.
[0212] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0213] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0214] Figure 5A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0215] like Figure 5 As shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the device 500 can also be stored in the RAM 503. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0216] Various components in device 500 are connected to I / O interface 505, including: an input unit 506, such as a keyboard, mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, optical disk, etc.; and a communication unit 509, such as a network card, modem, wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0217] The computing unit 501 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the poster image generation method. For example, in some embodiments, the poster image generation method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the poster image generation method described above can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the poster image generation method by any other appropriate means (e.g., by means of firmware).
[0218] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0219] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0220] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0221] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0222] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0223] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0224] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0225] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for generating a poster image, comprising: Generate initial prompt words based on the obtained poster, and determine the text elements included in the poster image to be generated; Generating a text element layer based on the text element and the text element position corresponding to the text element, and generating an image element layer based on the initial prompt word and the image element position generated on the poster; wherein there is no overlapping position between the text element position and the image element position; The text element layer and the image element layer are superimposed to obtain the poster image to be generated.
2. The method according to claim 1, wherein The step of generating an initial prompt word and an image element position according to the poster and generating an image element layer includes: generating a blank layer according to the image size of the poster image to be generated, and determining a text element area in the blank layer according to the position of the text element, and determining an image element area in the blank layer according to the position of the image element; An initial prompt word and the image element area are generated according to the poster, and image elements are generated in the blank layer to obtain the image element layer.
3. The method according to claim 2, wherein: The step of generating the initial prompt word and the image element area according to the poster, generating the image element in the blank layer, and obtaining the image element layer includes: Adding a mask to the text element area as a mask area, and using the image element area as a non-mask area, and determining a mask layer according to the mask area and the non-mask area; Obtaining a first area prompt word corresponding to the masked area and a second area prompt word corresponding to the unmasked area; wherein the second area prompt word is an initial prompt word generated for the poster; Inputting the mask layer, the first area prompt word, and the second area prompt word into an image generation model, so that the image generation model outputs the image element layer; The image generation model generates a first layer slice that does not contain any elements in the mask area according to the first area prompt word; The image generation model generates image elements in the non-masked area according to the second area prompt word to obtain a second layer slice; The image generation model obtains the image element layer according to the first layer slice and the second layer slice.
4. The method according to claim 3, wherein: When the image generation model is a diffusion model, the image generation model obtains the image element layer in the following manner: During the image denoising process, the image generation model calculates a first attention representation based on the first text feature of the prompt word in the first region and the first image feature of the first noise image, and denoises the first noise image based on the first attention representation to generate the first layer slice; wherein the first noise image is the noise-added image corresponding to the mask region; During the image denoising process, the image generation model calculates a second attention representation based on the second text feature of the prompt word in the second area and the second image feature of the second noise image, and denoises the second noise image based on the second attention representation to generate a second layer slice; wherein the second noise image is the noise-added image corresponding to the non-masked area; The image generation model obtains the image element layer according to the first layer slice and the second layer slice.
5. The method according to claim 1, wherein The step of generating initial prompt words based on the obtained poster and determining text elements included in the poster image to be generated includes: Inputting the initial prompt words generated by the poster into a large language model, performing prompt word expansion on the initial prompt words generated by the poster using the large language model, and outputting the expanded prompt words generated by the poster; The text elements included in the poster image to be generated are determined according to the poster generation extended prompt words.
6. The method according to claim 5, wherein: The step of determining the text elements included in the poster image to be generated according to the extended prompt words generated by the poster includes: Inputting the poster generation extended prompt word into a large language model, performing semantic induction on the poster generation extended prompt word using the large language model, and outputting at least one semantic information; The poster title corresponding to the poster image to be generated is determined according to each of the semantic information, and the text element is determined according to the poster title.
7. The method according to claim 6, wherein: The determining the poster title corresponding to the poster image to be generated according to each of the semantic information includes: Determining feature information corresponding to each of the semantic information; wherein the feature information includes at least one of text length, keyword density, and sentiment intensity; Weight information corresponding to each piece of the semantic information is calculated according to the feature information, and at least one main title of the poster and at least one subtitle of the poster are determined from each piece of the semantic information according to the weight information.
8. The method according to claim 7, wherein: The position of the text element is determined as follows: Determine the number of main titles of the poster main title and the number of subtitles of the poster subtitle; determining a target image position from candidate image positions according to the number of main titles and the number of subtitles, as the image position of the text element in the poster image to be generated; The position of the text element is determined according to the image position of the text element in the poster image to be generated.
9. The method according to claim 1, wherein Generating a text element layer according to the text element and the text element position corresponding to the text element includes: Determining a current layer color corresponding to the position of the text element in the image element layer, and determining a target text color corresponding to the text element based on the current layer color and an association relationship between candidate layer colors and candidate text colors; The text element layer is generated according to the target text color, the text element, and the text element position.
10. The method according to claim 1, wherein The step of generating an initial prompt word and an image element position according to the poster and generating an image element layer includes: Inputting the initial prompt words generated by the poster into a large language model, performing prompt word expansion on the initial prompt words generated by the poster using the large language model, and outputting the expanded prompt words generated by the poster; Determining a current poster type corresponding to the poster image to be generated based on the poster generation extended prompt word, and determining a target poster image style corresponding to the poster image to be generated based on the current poster type and an association between candidate poster types and candidate poster image styles; The poster generation extension prompt word is updated according to the target poster image style to generate a poster generation update prompt word, and the image element layer is generated according to the poster generation update prompt word and the image element position.
11. The method according to claim 1 , further comprising: In response to a modification operation on the text element in the text element layer, generating a modified text element layer; The modified text element layer and the image element layer are superimposed to obtain the poster image to be generated.
12. A device for generating a poster image, comprising: A text element determination module is used to generate initial prompt words based on the obtained poster and determine the text elements included in the poster image to be generated; an element layer generation module, configured to generate a text element layer based on the text element and the text element position corresponding to the text element, and generate an image element layer based on the initial prompt word and the image element position generated by the poster; wherein the text element position and the image element position do not overlap; The poster image generation module is used to perform superposition processing on the text element layer and the image element layer to obtain the poster image to be generated.
13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.
15. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 11.
Citation Information
Cited By
Poster generation method and device, electronic equipment and storage medium
CN121639849A