Poster generation method and device and storage medium

By performing feature annotation and text encoding on the target product image, and using a multimodal conditional generation optimization diffusion model to generate posters, the poster distortion problem caused by mask dependence in existing technologies is solved, and high-quality and diversified poster generation is achieved.

CN120953437APending Publication Date: 2025-11-14SHENZHEN DONSON CLOUD TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510945478.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing poster generation technology relies on masks, which can easily lead to size distortion and shape warping of product images, affecting the quality stability and automation of poster generation.

Method used

By annotating the target product image with features, extracting product prompts, and combining object and interactive action prompts with text encoding, a multimodal conditional generation optimization diffusion model is used to generate posters, avoiding reliance on masks.

Benefits of technology

It achieves high-fidelity and diverse poster generation, automatically and accurately generating high-quality product display posters without the need for manual drawing or mask correction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953437A_ABST
    Figure CN120953437A_ABST
Patent Text Reader

Abstract

The invention discloses a poster generation method and device and a storage medium, and relates to the technical field of computer vision, and the method comprises the steps: extracting a feature label of a target product based on a product image corresponding to the target product, and obtaining a product prompt word according to the feature label; obtaining an object prompt word and an interaction action prompt word, and carrying out text coding on the product prompt word, the object prompt word and the interaction action prompt word to obtain a corresponding text embedding vector; carrying out subsurface space coding on the product image to obtain a corresponding image subsurface space feature vector; inputting the image latent space feature vector and the text embedding vector into a multi-modal condition generation optimization diffusion model, and generating a first-stage base map through the multi-modal condition generation optimization diffusion model; and generating a product poster corresponding to the target product according to the first-stage base map, thereby solving the technical problem that poster generation in the prior art seriously depends on a mask, and automatically and accurately generating a high-quality product display poster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to a poster generation method, device, and storage medium. Background Technology

[0002] In the e-commerce sector, with the surge in demand for personalized marketing, efficiently generating high-fidelity product-on-the-go / body-on posters that meet diverse customization requirements has become a key need for merchants' promotion and conversion. Existing poster generation technologies widely employ image restoration or replacement-based methods, typically relying on high-quality source model images as the foundation. Product area masks are then manually drawn or semi-automatically annotated to complete the image synthesis. However, due to the wide variety of products and display scenarios, when suitable source model images are lacking or the masks are inaccurately drawn, the product images in the generated posters are prone to distortions such as size distortion and shape warping, severely affecting the realistic display effect of the product and directly limiting the quality stability and automation level of poster generation.

[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of this application is to provide a poster generation method, device, and storage medium, aiming to solve the technical problem that poster generation in the prior art heavily relies on masks.

[0005] To achieve the above objectives, this application proposes a poster generation method, which includes:

[0006] Based on the product image corresponding to the target product, feature annotations of the target product are extracted, and product prompt words are obtained based on the feature annotations;

[0007] Obtain object prompts and interaction prompts, and perform text encoding on the product prompts, object prompts, and interaction prompts to obtain corresponding text embedding vectors;

[0008] The product image is latent space encoded to obtain the corresponding image latent space feature vector;

[0009] The image latent space feature vector and the text embedding vector are input into a multimodal conditional generation optimization diffusion model, and a first-stage base map is generated through the multimodal conditional generation optimization diffusion model;

[0010] Generate a product poster corresponding to the target product based on the base map of the first stage.

[0011] In some embodiments, the step of extracting feature annotations of the target product based on the product image corresponding to the target product, and obtaining product prompts based on the feature annotations, includes:

[0012] Based on the field lookup table associated with the preset structured template, the product target attributes corresponding to the fields to be filled in the preset structured template are determined.

[0013] A multimodal large model is used to extract the target attributes of the target product in the product image to obtain the target attribute description text corresponding to the target product;

[0014] The target attribute description text is determined as the product prompt word.

[0015] In some embodiments, the step of inputting the image latent space feature vector and the text embedding vector into a multimodal conditional generation optimization diffusion model, and generating a first-stage base map through the multimodal conditional generation optimization diffusion model includes:

[0016] The text embedding vector is spatially expanded by linear projection to obtain the corresponding text spatial feature vector;

[0017] Cross-attention processing is performed based on the latent space feature vector of the image and the feature vector of the text space to obtain the corresponding cross-attention processing result;

[0018] Multi-scale feature extraction is performed on the cross-attention processing results to obtain feature processing results at multiple different scales;

[0019] The latent space feature vector of the image and the feature vector of the text space are concatenated, and the corresponding concatenation result is convolved to obtain the corresponding pixel-wise weight map.

[0020] A multimodal fusion feature vector is obtained by weighting and fusing the pixel-by-pixel weight map, the image latent space feature vector, and the text space feature vector.

[0021] Using the multimodal fusion feature vector and the feature processing result as multimodal conditions, the U-Net network that optimizes the diffusion model is used to predict noise and gradually denoise it to generate the latent vector of the first-stage base map.

[0022] The latent vectors of the first-stage base map are decoded to obtain the first-stage base map.

[0023] In some embodiments, the step of using the multimodal fused feature vector and the feature processing result as multimodal conditions, and using the multimodal conditions as input to optimize the diffusion model to predict noise and progressively denoise, to generate the latent vector of the first-stage base map, includes:

[0024] Obtain a noise map, perform latent space encoding on the noise map, and obtain the corresponding noise map feature vector;

[0025] The noise map feature vector is input into the U-Net network, the feature processing result is input into each downsampling block of the U-Net network, the multimodal fusion feature vector is input into each downsampling block and each upsampling block of the U-Net network, and the noise map feature vector, the multimodal fusion feature vector, and the feature processing result are iteratively diffused based on each downsampling block and each upsampling block to obtain the latent vector of the first stage base map.

[0026] In some embodiments, the step of generating a product poster corresponding to the target product based on the first stage base map includes:

[0027] Obtain the facial image of the target object;

[0028] Based on the first stage base map, determine the facial region of the object in the first stage base map;

[0029] Based on the facial image, a face-swapping algorithm is used to match and fuse the facial region of the object in the first stage base image to obtain the second stage initial base image;

[0030] The product poster corresponding to the target product is generated based on the initial base map of the second stage.

[0031] In some embodiments, the step of generating the product poster corresponding to the target product based on the second-stage initial base map includes:

[0032] The facial region of the object in the initial base map of the second stage is re-textured using a facial texture enhancement algorithm to obtain the base map of the second stage.

[0033] The second-stage base map is determined as the product poster corresponding to the target product;

[0034] Wherein, the product in the product poster is the target product, the interaction between the object in the product poster and the target product satisfies the interaction action prompt words, and the object in the product poster is the target object.

[0035] In some embodiments, the step of generating a product poster corresponding to the target product based on the first stage base map includes:

[0036] The first-stage base map is input into the segmentation model to segment the target product and the object in the first-stage base map, thereby obtaining the segmentation mask corresponding to the first-stage base map.

[0037] A background reference image is obtained, and features are extracted from the background reference image using a style transfer model to obtain the corresponding style condition feature vector;

[0038] Based on the segmentation mask, determine the foreground region and background region in the segmentation mask;

[0039] Based on the style conditional feature vector, the background region in the segmentation mask is diffused to generate the initial base map for the third stage.

[0040] Based on the initial base map of the third stage, the lighting re-rendering algorithm is used to re-render the initial base map of the third stage to obtain the base map of the third stage, and the base map of the third stage is determined as the product poster corresponding to the target product.

[0041] The product in the product poster is the target product, the interaction between the object in the product poster and the target product satisfies the interaction prompt words, and the style conditions of the background area of ​​the product poster are consistent with the background reference image.

[0042] In some embodiments, before the step of determining the second stage base map as the product poster corresponding to the target product, the method further includes:

[0043] The second-stage base map is input into the segmentation model, and the target product and the target object in the second-stage base map are segmented to obtain the second segmentation mask corresponding to the second-stage base map.

[0044] A background reference image is obtained, and features are extracted from the background reference image using a style transfer model to obtain the corresponding style condition feature vector;

[0045] Based on the second segmentation mask, determine the foreground region and background region in the second segmentation mask;

[0046] Based on the style conditional feature vector, the background region in the second segmentation mask is diffused to generate the initial base map of the third stage.

[0047] Based on the initial base map of the third stage, the lighting re-rendering algorithm is used to re-render the initial base map of the third stage to obtain the base map of the third stage, and the base map of the third stage is determined as the product poster corresponding to the target product.

[0048] Wherein, the product in the product poster is the target product, the interaction between the object in the product poster and the target product satisfies the interaction prompt words, and the style conditions of the background area of ​​the product poster are consistent with the background reference image.

[0049] Furthermore, to achieve the above objectives, this application also proposes a poster generation device, the poster generation device comprising:

[0050] The feature annotation module is used to extract feature annotations of the target product based on the product image corresponding to the target product, and to obtain product prompt words based on the feature annotations;

[0051] The text encoding module is used to acquire object prompts and interactive action prompts, and to perform text encoding on the product prompts, object prompts, and interactive action prompts to obtain corresponding text embedding vectors;

[0052] The latent space encoding module is used to perform latent space encoding on the product image to obtain the corresponding image latent space feature vector;

[0053] The diffusion module is used to input the image latent space feature vector and text embedding vector into a multimodal conditional generation optimization diffusion model, and generate a first-stage base map through the multimodal conditional generation optimization diffusion model;

[0054] The generation module is used to generate a product poster corresponding to the target product based on the base map of the first stage.

[0055] In addition, to achieve the above objectives, this application also proposes a poster generation device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the poster generation method as described above.

[0056] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the poster generation method described above.

[0057] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the poster generation method described above.

[0058] The one or more technical solutions proposed in this application have at least the following technical effects: By acquiring product images, object prompts, and interaction prompts describing the interaction between the target product and the object, basic information support is provided for the generation process. Feature annotations of the target product are extracted based on the product image, resulting in product prompts. Key visual features of the product are extracted and transformed into structured semantics. Product prompts, object prompts, and interaction prompts are text-encoded to obtain corresponding text embedding vectors, enabling the model to understand the semantic information of the text. Latent space encoding is performed on the product image to obtain corresponding image latent space feature vectors. The product image is mapped to the latent space feature vectors to extract high-level visual features of the image. Image generation is achieved by inputting the latent space feature vector of the image and the text embedding vector into a multimodal conditional generation optimization diffusion model. A first-stage base map is generated based on this model. The model uses the latent space feature vector (visual features of the product) and the text embedding vector (semantic constraints of the product, object, and interaction) as conditions, and gradually generates a first-stage base map that meets the requirements by fusing multimodal information through a diffusion process. This first-stage base map can be post-processed to obtain the product poster corresponding to the target product. In existing technologies, poster generation heavily relies on masks. If the mask size or shape deviates from the product, the generated poster will be distorted or deformed. This application, however, directly analyzes and encodes the product image and related descriptive text, and then uses the multimodal conditional generation optimization diffusion model to generate the target poster, avoiding reliance on masks. This solves the technical problem of existing poster generation heavily relying on masks, automatically and accurately generating high-quality product display posters without the need for manual mask drawing or correction, thus achieving high-fidelity and diverse poster generation needs. Attached Figure Description

[0059] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0060] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0061] Figure 1 This is a flowchart illustrating the poster generation method of this application in Embodiment 1;

[0062] Figure 2 This application provides a schematic diagram of a structure for performing a poster generation method;

[0063] Figure 3 This application provides another structural diagram for performing a poster generation method;

[0064] Figure 4 A schematic diagram of another structure for performing a poster generation method provided in this application;

[0065] Figure 5 This is a schematic diagram of the module structure of the poster generation device according to an embodiment of this application;

[0066] Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the poster generation method in this application embodiment.

[0067] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0068] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0069] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0070] The main solution of this application embodiment is as follows: extract feature annotations of the target product based on the product image corresponding to the target product, and obtain product prompt words based on the feature annotations; perform text encoding on the product prompt words, object prompt words, and interactive action prompt words to obtain corresponding text embedding vectors; perform latent space encoding on the product image to obtain corresponding image latent space feature vectors; input the image latent space feature vectors and text embedding vectors into a multimodal conditional generation optimization diffusion model, and generate a first-stage base map through the multimodal conditional generation optimization diffusion model; generate a product poster corresponding to the target product based on the first-stage base map.

[0071] In this embodiment, for ease of description, the following description will focus on the poster generation device as the execution subject.

[0072] Existing poster generation technologies widely employ image inpainting or replacement methods, typically relying on high-quality source model images as the foundation. Product area masks are manually drawn or semi-automatically annotated to complete image synthesis. However, due to the wide variety of products and display scenarios, when suitable source model images are lacking or masks are inaccurately drawn, product images in the generated poster are prone to distortions such as size distortion and shape warping. This severely affects the realistic display of the product, directly limiting the quality stability and automation level of poster generation.

[0073] This application provides a solution that directly analyzes and encodes product images and related descriptive text, and then uses a multimodal conditional generation optimization diffusion model to generate target posters. This avoids dependence on masks and solves the technical problem of poster generation in the prior art that heavily relies on masks. It automatically and accurately generates high-quality product display posters without the need for manual drawing or correction of masks, thus achieving high-fidelity and diversified poster generation needs.

[0074] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or a poster generation device capable of performing the above functions. The following description uses a poster generation device as an example to illustrate this embodiment and the subsequent embodiments.

[0075] Based on this, this application provides a poster generation method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the poster generation method of this application.

[0076] In this embodiment, the poster generation method includes steps 101 to 105:

[0077] Step 101: Extract feature annotations of the target product based on the product image corresponding to the target product, and obtain product prompt words based on the feature annotations.

[0078] Specifically, the process involves acquiring product images, object prompts, and interaction prompts describing the interaction between the target product and the object. The product image corresponding to the target product is the original visual data of the core product to be displayed in the poster. It can be a photo or scanned copy of the target product. The product image includes the complete appearance, structure, and key details of the target product, including the product's brand logo, material texture, and color. In this application, the target product can be various products that can be sold online, such as clothing, pants, hats, and bags. The object prompts are textual information describing the subject (i.e., the "object") interacting with the target product, including the object's physical and facial features. Physical features can include gender, age, height, and hair color, while facial features can include facial expression categories, such as a smiling expression. The object prompts provide semantic constraints on the object during the poster generation process, clarifying "who interacts with the product" and "the object's basic characteristics," avoiding discrepancies between the generated object and the requirements, such as incorrectly generating a "male model" instead of a "female model." Interactive action prompts are textual information used to describe specific interactive behaviors between the target product and the object. They typically include action type (such as "holding", "wearing", "displaying"), action details (such as "lifting the hem of the garment with one hand" or "crossing arms"), and action state (such as "bending over", "standing naturally", "turning around", or "squatting"). Interactive action prompts can provide semantic constraints on interactive behaviors during the poster generation process, clarifying "how the object and the product interact", and avoiding a disconnect between the generated actions and the requirements (such as incorrectly generating "holding the product" as "placing the product on the ground").

[0079] Product images can be uploaded by users or retrieved from a product database, which can be the product image library of an e-commerce platform. Product images must be clear, unobstructed, and cover the core display areas of the target product; for example, they should include front, back, and side views of clothing. Both object prompts and interaction prompts can be user-defined or generated based on historical data. By acquiring product images, object prompts, and interaction description text, and through multimodal information input (visual + semantic), clear generation guidance can be provided for generating optimized diffusion models under multimodal conditions, without relying on masked regions.

[0080] Before annotating the target product's features in the product image, preprocessing operations such as cropping and background removal can be performed. The preprocessed image is then fed into a multimodal large model, which can be Qwen Vision-Language Model Version 2.5 (Qwen-VL 2.5). This multimodal large model automatically identifies various attributes in the product image, including color, material, style, and logo, obtaining corresponding attribute description text. For example, the attribute description text for color could be "red," for material "cotton," and for style "dress," etc. This attribute description text is then used as product prompts. By using computer vision technology to extract key visual features (such as shape, color, and details) of the target product from the product image and converting them into structured text descriptions, the core features of the target product are clearly defined. Converting key details of the target product in the product image (such as the brand logo) into text descriptions avoids the loss of details due to masking bias. Furthermore, by replacing the traditional method of "mask-defined regions" with feature annotations of the image itself, the subsequent multimodal conditional generation optimization diffusion model can directly and accurately restore the complete form of the product based on its inherent features, without the need for external mask-assisted positioning.

[0081] Step 102: Obtain object prompts and interactive action prompts, and perform text encoding on the product prompts, object prompts, and interactive action prompts to obtain the corresponding text embedding vectors.

[0082] Specifically, product prompts, object prompts, and interactive action prompts are input into a text encoder for text encoding to obtain the corresponding text embedding vectors.

[0083] For example, the text encoder can be a text encoder based on a contrastive language-image pre-training (CLIP) model. Product prompts, object prompts, and interactive action prompts are input into the text encoder for semantic feature extraction, capturing the semantic information of these prompts. The product prompts, object prompts, and interactive action prompts are then converted into a fixed-length feature vector, i.e., a text embedding vector. These prompts can provide users with their expectations and descriptions of the target poster. The encoded text embedding vector can be used to guide a multimodal conditional generation optimization diffusion model to generate a target poster that meets the user's needs and expectations.

[0084] Step 103: Perform latent space encoding on the product image to obtain the corresponding image latent space feature vector.

[0085] Specifically, the product image is input into a variational autoencoder (VAE) for latent space encoding to obtain the corresponding image latent space feature vector.

[0086] For example, a product image is input into a Visual Image Generator (VAE) for latent space encoding. The VAE can extract local features of the product image through a convolutional neural network, and then compress them into a low-dimensional latent space through a fully connected layer to obtain the image latent space feature vector. The image latent space feature vector is a visual "compressed representation" of the image, containing the core visual information of the product image (while filtering out redundant details). The image latent space feature vector is used as a conditional input to a multimodal conditional generation optimization diffusion model to help guide the generation process and ensure that the generated poster image is consistent with the original product image and has high fidelity in detail.

[0087] Step 104: Input the image latent space feature vector and text embedding vector into the multimodal conditional generation optimization diffusion model, and generate the first-stage base map through the multimodal conditional generation optimization diffusion model.

[0088] In this model, the product in the first-stage base map is the target product, and the interaction actions between the object and the target product in the first-stage base map satisfy the interaction action prompts. The objects in the first-stage base map also satisfy the object prompts. When the multimodal conditional generation optimization diffusion model receives the input image latent space feature vector and text embedding vector, it performs noise prediction and denoising based on these vectors. Based on the prediction results, it progressively removes noise to generate latent vectors that conform to the multimodal constraints (the constraints of the image latent space feature vector and the text embedding vector), thus generating the first-stage base map.

[0089] Specifically, the poster generation method of this application is applied to a trained multimodal conditional generation optimization diffusion model. In this application, the multimodal conditional generation optimization diffusion model can be the FLUX.1 model. The FLUX.1 model is a multimodal conditional generation optimization model based on an improved diffusion model architecture. Its core innovation lies in combining the cross-modal attention mechanism of the Transformer with the progressive generation capability of the diffusion process, achieving high-precision fusion control of multimodal conditions such as text, images, and actions. The image latent space feature vector and text embedding vector are input into the multimodal conditional generation optimization diffusion model for image generation. Through a cross-attention mechanism or conditional projection, the image latent space feature vector and text embedding vector are deeply fused into a multimodal fusion feature vector. The multimodal fusion feature vector simultaneously conveys the visual features of the target product and object-action semantic constraints, clearly defining the generation target, such as "a young woman standing naturally wearing a red dress." In the diffusion generation stage, initialization starts from standard normal distribution noise, followed by inverse denoising iteration. At each time step of denoising, the time step embedding and multimodal condition vector (i.e., the multimodal fusion feature vector) are input to predict and remove noise. By fusing feature vector constraints across multiple modalities and learning the correlations and interactions between input feature vectors, multiple diffusion processes are performed. Each diffusion process adjusts and optimizes based on previous results and the input feature vectors, gradually generating a feature vector that integrates all input feature vectors—the output of the multimodal conditional generation optimization diffusion model. The output feature vector of the multimodal conditional generation optimization diffusion model is then decoded to obtain the first-stage base map. The generated first-stage base map meets requirements such as "the product in the first-stage base map is the target product" and "the interaction between the object and the target product in the first-stage base map satisfies the interaction prompts." Furthermore, super-resolution processing can be applied to the generated first-stage base map to optimize local details. Through the multimodal conditional generation optimization diffusion model, the latent space feature vectors (visual essence) of the image are deeply fused with the text embedding vectors (semantic constraints) to generate a high-fidelity poster that meets all requirements. By conveying the visual essence and globally optimizing multimodal constraints, the poster image distortion problem caused by inaccurate masking is solved.

[0090] Step 105: Generate the product poster corresponding to the target product based on the base map of the first stage.

[0091] Optionally, the first-stage base image can be determined as the product poster corresponding to the target product; or, the first-stage base image can be post-processed, including super-resolution enhancement, background replacement, etc., and the result of the post-processing can be used as the product poster corresponding to the target product.

[0092] Based on the poster generation method proposed in this application, the generation process is supported by acquiring product images, object prompts, and interaction prompts describing the interaction between the target product and the object. Feature annotations of the target product are extracted from the product image to obtain product prompts, and key visual features of the product are extracted and transformed into structured semantics. Product prompts, object prompts, and interaction prompts are text-encoded to obtain corresponding text embedding vectors, enabling the model to understand the semantic information of the text. Latent space encoding of the product image is performed to obtain corresponding image latent space feature vectors. The product image is mapped to the latent space feature vectors to extract high-level visual features of the image. The image latent space feature vectors and text embedding vectors are input into a multimodal conditional generation optimization diffusion model for image generation. A first-stage base map is generated based on the multimodal conditional generation optimization diffusion model. The multimodal conditional generation optimization diffusion model uses the image latent space feature vectors (visual features of the product) and text embedding vectors (semantic constraints of the product, object, and interaction) as conditions, and gradually generates a first-stage base map that meets the requirements by fusing multimodal information through a diffusion process. The first-stage base map can be post-processed to obtain the product poster corresponding to the target product. In existing technologies, poster generation heavily relies on masks. If the mask size or shape deviates from the product, the generated poster will be distorted or deformed. This application, however, directly analyzes and encodes the product image and related descriptive text, and then uses a multimodal conditional generation optimization diffusion model to generate the target poster. This avoids the reliance on masks and solves the technical problem of poster generation in existing technologies that heavily depend on masks. It automatically and accurately generates high-quality product display posters without the need for manual mask drawing or correction, thus meeting the needs for high-fidelity and diverse poster generation.

[0093] In some embodiments, the step of extracting feature annotations of the target product based on the product image corresponding to the target product, and obtaining product prompts based on the feature annotations, includes:

[0094] Based on the field lookup table associated with the preset structured template, determine the product target attributes corresponding to the fields to be filled in the preset structured template;

[0095] A multimodal large model is used to extract the target product attributes from the product image, and the target attribute description text corresponding to the target product is obtained.

[0096] The target attribute description text is identified as the product prompt word.

[0097] Specifically, the preset structured template is a predefined standardized text framework containing four main parts: object appearance features, object expression features, product-object interaction actions, and product features. The product features section contains pre-defined target attributes of the product, such as color, material, style, pattern, and logo. The preset structured template can also be used to standardize the scope and format of attribute extraction. The field lookup table maps each field in the preset structured template to the product attributes to be extracted from the image, clarifying the type of information to be extracted from each field, i.e., the target attributes of the product. The Multimodal Large Language Model (MLLM) is a large language model that integrates multiple modal information such as text, images, video, and speech. It achieves multimodal "understanding-generation-interaction" through a unified framework. By integrating visual (image / video) and auditory (speech) multimodal inputs, MLLM enables the model to "understand" non-textual content and generate cross-modal outputs. For example, multimodal large models can be Qwen-VL 2.5, Large Language and Vision Assistant (LLaVA), Bootstrapped Language-Image Pre-training (BLIP-2), etc.

[0098] Before determining the product target attributes corresponding to the fields to be populated in the preset structured template based on the field lookup table associated with the preset structured template, the process also includes: pre-designing a structured template according to business needs. The template contains attribute fields to be extracted (such as "color," "collar type," "material," and "brand"). The number and type of fields are determined according to the product type; for example, clothing requires "collar type" and "sleeve length," while 3C products require "size" and "interface type," etc. A field lookup table is then constructed to bind specific product attributes, object attributes, or action attributes to each field in the template, forming a mapping relationship table between fields and attributes.

[0099] As an example, a product image is input into a multimodal model. The multimodal model extracts local features (such as color distribution and neckline shape) and global features (such as overall outline and material texture) from the image using an image encoder. Based on a field lookup table, it determines the product target attributes corresponding to the fields to be filled in the preset structured template. The extracted image features are then analyzed based on these product target attributes, including: identifying the target product's color using a color classification head, identifying the target product's material using a texture analysis module, and extracting the brand name and logo using an Optical Character Recognition (OCR) module or a brand logo recognition model. Specifically, the target attribute description text corresponding to the target product is obtained as follows: the attribute description text for the color attribute can be "red," the attribute description text for the material attribute can be "cotton," the attribute description text for the style attribute can be "dress," and so on. The target attribute description text corresponding to the target product is mapped to the preset structured template according to the field lookup table to generate the final product prompt. The product prompt can be in JavaScript Object Notation Schema format.

[0100] Alternatively, object prompts can also be obtained through the following steps:

[0101] Based on the field lookup table associated with the preset structured template, determine the target attributes of the objects corresponding to the fields to be filled in the preset structured template;

[0102] Obtain the object image corresponding to the target object, and use a multimodal large model to extract the target object's target attributes from the object image to obtain the target attribute description text corresponding to the target object;

[0103] The target attribute description text corresponding to the target object is determined as the object prompt word.

[0104] In some embodiments, the step of inputting the image latent space feature vector and the text embedding vector into a multimodal conditional generation optimization diffusion model, and generating a first-stage base map through the multimodal conditional generation optimization diffusion model includes:

[0105] The text embedding vector is spatially expanded by linear projection to obtain the corresponding text space feature vector;

[0106] Cross-attention processing is performed based on image latent space feature vectors and text space feature vectors to obtain the corresponding cross-attention processing results;

[0107] Multi-scale feature extraction is performed on the cross-attention processing results to obtain feature processing results at multiple different scales;

[0108] The latent space feature vectors of the image and the feature vectors of the text space are concatenated, and the concatenation results are convolved to obtain the corresponding pixel-wise weight map.

[0109] A multimodal fusion feature vector is obtained by weighting and fusing the pixel-wise weight map, the latent space feature vector of the image, and the feature vector of the text space.

[0110] Using multimodal fusion feature vectors and feature processing results as multimodal conditions, the U-Net network that optimizes the diffusion model predicts noise and gradually denoises it to generate the latent vector of the first-stage base map.

[0111] The latent vectors of the first-stage base map are decoded to obtain the first-stage base map.

[0112] Specifically, the text embedding vector can be projected into a space with the same dimension as the latent space feature vector of the image through a fully connected layer, and then transformed from a one-dimensional sequence into a two-dimensional spatial feature tensor through a reshape operation to simulate the spatial layout of the image and obtain the text spatial feature vector. This aligns the text features with the image features in the same latent space, which facilitates subsequent cross-modal interaction.

[0113] In some embodiments, cross-attention processing is performed based on image latent space feature vectors and text space feature vectors to obtain the corresponding cross-attention processing result. This includes: using the image latent space feature vector as the key vector and value vector, and the text space feature vector as the query vector; multiplying the key vector and query vector to obtain the corresponding weights, and then multiplying the weights and value vectors to obtain the cross-attention processing result. Multiplying the key vector and query vector can calculate the weights between the image latent space feature vector and the text space feature vector, and multiplying the weights and value vectors to adjust the value vector, i.e., the contribution of the image latent space feature vector, to obtain the features of the product image that match the features of the descriptive text, i.e., the cross-attention processing result. Further, the cross-attention processing result is subjected to alternating convolution and downsampling processing, wherein the number of convolution processing operations is one more than the number of downsampling processing operations, to achieve multi-scale feature extraction of the cross-attention processing result, extracting feature information from multiple levels, thereby more comprehensively understanding the fusion and matching of the features of the descriptive text and the features of the product image. Through multi-scale feature processing, multiple feature processing results at different scales can be obtained, which contain matching information of the features of the prompt words and the features of the product image at different levels.

[0114] The latent space feature vector of the image and the feature vector of the text space are concatenated along the channel dimension to form a joint feature that contains complete visual and semantic information. A 1×1 convolutional layer is then used to extract spatial features from the concatenated result, generating a pixel-wise weight map W. This pixel-wise weight map reflects the importance of each pixel location to multimodal fusion. The multimodal fusion feature vector is obtained by weighted fusion using the pixel-wise weight map, the latent space feature vector of the image, and the feature vector of the text space. The weighted fusion can be performed as follows:

[0115] F fused =W⊙F img +(1-W)⊙F txt

[0116] Among them, F fused For multimodal fusion feature vectors, W is the pixel-wise weight map, and F is the multimodal fusion feature vector. img F is the latent space feature vector of the image. txt The text space feature vector is represented by ⊙, which denotes element-wise multiplication. Weighted fusion balances the contributions of visual and semantic elements, ensuring the generated poster image conforms to both visual features and semantic constraints. The multimodal fusion feature vector and the multi-scale feature processing results are used as multimodal conditions and input into the U-Net network for iterative denoising. The U-Net network contains multiple downsampling blocks and multiple upsampling blocks. In each diffusion process, the downsampling block downsamples the current noisy image to low-resolution features and extracts the global context, while the upsampling block upsamples the low-resolution features to high-resolution features. Combined with skip connections, local details are preserved. With each iterative diffusion process, the generated vector gets closer and closer to the latent space representation of the first-stage base image. After a preset number of diffusion processes, the latent vector of the denoised first-stage base image is generated. The latent vector of the denoised first-stage base image is input into a decoder for decoding, generating a realistic and accurate product display image, i.e., the first-stage base image. The decoder can be a Variational Auto-Decoder (VAE).

[0117] refer to Figure 2 , Figure 2A schematic diagram of a poster generation method is provided for generating a first-stage base map, including a multimodal large model 201, a text encoder 202, a variational autoencoder 203, a multimodal conditional generation optimization diffusion model 204, and a variational autodecoder 205. The product image is input into the multimodal large model 201, which extracts the target product attributes from the image to obtain the corresponding target attribute description text. Product prompts, object prompts, and interactive action prompts are input into the text encoder 202 for text encoding to obtain corresponding text embedding vectors. The product image is input into the variational autoencoder 203 for latent space encoding to obtain corresponding image latent space feature vectors. The image latent space feature vectors and text embedding vectors are input into the multimodal conditional generation optimization diffusion model 204 for image generation, generating the latent vectors of the first-stage base map. The latent vectors of the first-stage base map are then input into the variational autodecoder 205 for decoding to obtain the first-stage base map.

[0118] In some embodiments, the step of using multimodal fusion feature vectors and feature processing results as multimodal conditions, predicting noise and progressively denoising through a U-Net network that optimizes the diffusion model by inputting multimodal conditions, and generating latent vectors for the first-stage base map includes:

[0119] Obtain the noise map, perform latent space encoding on the noise map, and obtain the corresponding noise map feature vector;

[0120] The noise map feature vector is input into the U-Net network, the feature processing result is input into each downsampling block of the U-Net network, and the multimodal fusion feature vector is input into each downsampling block and each upsampling block of the U-Net network. Based on each downsampling block and each upsampling block, the noise map feature vector, the multimodal fusion feature vector, and the feature processing result are subjected to iterative diffusion processing to obtain the latent vector of the first-stage base map.

[0121] Specifically, the noise map is a randomly generated pure noise image (such as Gaussian noise), serving as the initial input for the inverse denoising process of the diffusion model. The noise map can be generated by sampling from a standard normal distribution. The shape and size of the noise map are consistent with the target poster, and the noise map is a collection of random pixel values. The noise map is compressed and encoded using a latent space encoder to extract its low-dimensional latent space features, resulting in a noise map feature vector. This transforms disordered noise into a computable vector representation, providing a foundation for subsequent processing. The multimodal fusion feature vector (a visual-semantic joint representation fusing image latent space features and text embedding vectors) and the feature processing results (outputs of cross-attention or multi-scale feature extraction) are injected into the downsampling and upsampling blocks of the U-Net network, respectively. The process involves several steps: The downsampling block receives the noise map feature vector, the multimodal fusion feature vector, and the feature processing result. It extracts global context information and, through feature fusion and downsampling, fuses the noise map features with the multimodal information and the feature processing result to generate a global feature map. The upsampling block receives the global features and feature processing result output from the downsampling block and, through feature fusion and upsampling, gradually restores the spatial resolution of the image. The U-Net network uses a preset number of iterations for denoising (the specific number of steps is determined by the noise scheduling strategy) to gradually correct random interference in the noise map. In the downsampling stage, each downsampling block performs cross-attention processing on the current noise map features and the multimodal fusion feature vector, and also performs cross-attention processing on the current noise map features and the feature processing result to extract multimodal global features and then performs downsampling. In the upsampling stage, each upsampling block concatenates the global features output from the downsampling block with the output of the previous upsampling block, performs cross-attention processing on the concatenated result with the multimodal fusion feature vector, and then performs upsampling to restore local details, ensuring that the details of the generated poster image are consistent with the overall image. After each upsampling block outputs the final feature, the noise value of the current step is predicted through the convolutional layer, and the noise map feature vector is updated according to the noise scheduling formula of the diffusion model. The latent space representation of the first-stage base map is gradually approximated to obtain the final denoised latent vector, which is the latent vector of the first-stage base map. The latent vector of the first-stage base map has been fused with multimodal conditions (visual features + semantic constraints).

[0122] In some embodiments, the step of generating a product poster corresponding to the target product based on the first-stage base map includes:

[0123] Obtain the facial image of the target object;

[0124] Based on the first-stage base map, determine the facial region of the object described in the first-stage base map;

[0125] Based on facial images, a face-swapping algorithm is used to match and fuse the facial regions of the objects in the first-stage base image to obtain the initial base image for the second stage.

[0126] Generate product posters corresponding to the target products based on the initial base map of the second stage.

[0127] Specifically, the face-swapping algorithm can be a Visual Geometry Group Face Recognition Model (VGG-Face). Pre-trained object detection models, such as Multi-task Cascaded Convolutional Networks (MTCNN) or You Only Look Once version 8 (YOLOv8), can be used to detect faces in the first-stage base image, outputting the bounding box coordinates of the face and clarifying the specific location of the object's face in the poster within the first-stage base image. Based on this, facial landmark detection models can be used to locate facial feature points, such as the center of the eyes, the tip of the nose, and the corners of the mouth, for alignment optimization during subsequent face swapping. A deep learning face-swapping algorithm is then used to replace the facial regions in the first-stage base image. The facial image of the target object and the facial region of the object in the first-stage base image are unified to the same reference coordinate system. Pose and expression information of the facial region image of the object in the first-stage base image are extracted, and the texture features of the target object's facial image are transferred to the facial structure of the object in the first-stage base image. Image inpainting and color matching techniques are used to eliminate edge traces after face swapping, and skin tone and lighting consistency are adjusted. The processed target facial image is seamlessly embedded into the corresponding position in the first-stage base image to form the second-stage initial base image. In the second-stage initial base image, the object's face has been replaced with the target object's face, and it blends naturally with the product and interactive actions. The second-stage initial base image undergoes resolution enhancement and detail enhancement (such as edge sharpening), and finally, a product poster corresponding to the target product is output.

[0128] By using face-swapping fusion, the user-specified character image can be accurately embedded into the poster, meeting the diverse needs of users for generating posters.

[0129] In some embodiments, after obtaining the second-stage initial base map, the method further includes:

[0130] The facial region of the object in the initial base map of the second stage is re-drawn using a facial texture enhancement algorithm to obtain the base map of the second stage.

[0131] In this context, the product in the second-stage base map is the target product, the interaction between the object in the second-stage base map and the target product is the target interaction action, and the object in the second-stage base map is the target object.

[0132] As an example, facial texture enhancement algorithms are deep learning-based facial texture synthesis and optimization techniques, such as texture generation models and diffusion models based on Generative Adversarial Networks (GANs), which can be used to enhance facial details. A facial detection model can be used to locate the facial region of the object in the initial base image of the second stage, ensuring coverage of key areas such as the eyes, nose, and mouth. Then, a pre-trained facial texture encoder extracts texture features of this region, such as skin roughness, pore distribution, and gloss direction, providing a basis for subsequent texture enhancement. A texture enhancement model based on GANs is used to redraw and optimize the facial region. Using the texture features of the target object's facial image as a condition, more realistic and delicate textures are generated. Simultaneously, an attention mechanism focuses on key areas such as the area around the eyes and the nostrils, enhancing local details. The color and contrast of the generated texture can be adjusted to ensure the enhanced facial region matches the overall style of the poster. The texture-enhanced facial region is then merged with other parts of the initial base image of the second stage (such as the body and product) to finally generate the second-stage base image. The facial texture of the second-stage base map is highly consistent with the actual target object and perfectly integrated with the target product and interactive actions to present a highly realistic and personalized effect. Through facial region analysis, texture enhancement processing, and redrawing, the facial texture of the initial base map in the second stage is optimized to a delicate and realistic texture consistent with the actual target object, generating facial textures that conform to real physiological characteristics, making the face of the object in the poster closer to the real object. After completing the facial texture enhancement, the processed second-stage base map is determined as the product poster corresponding to the target product, and the target object, target product, and interactive actions between the target object and target product are all integrated.

[0133] refer to Figure 3 , Figure 3 A schematic diagram of another method for generating posters is provided, which generates a second-stage base map and obtains the facial image of the target object. Based on the first-stage base map, the facial region of the object in the first-stage base map is determined. Based on the facial image, a face-swapping algorithm is used to match and fuse the facial region of the object in the first-stage base map to obtain the second-stage initial base map. A facial texture enhancement algorithm is used to redraw the texture of the facial region of the object in the second-stage initial base map to obtain the second-stage base map. In the second-stage base map, the product is the target product, the interaction between the object and the target product in the second-stage base map satisfies the interaction prompt words, and the object in the second-stage base map is the target object.

[0134] In some embodiments, the step of generating a product poster corresponding to the target product based on the first-stage base map includes:

[0135] The first-stage base map is input into the segmentation model to segment the target products and objects in the first-stage base map, and the segmentation mask corresponding to the first-stage base map is obtained.

[0136] Obtain a background reference image, and extract features from the background reference image using a style transfer model to obtain the corresponding style conditional feature vector;

[0137] Based on the segmentation mask, determine the foreground and background regions in the segmentation mask;

[0138] Based on the style conditional feature vector, the background region in the segmentation mask is diffused to generate the initial base map of the third stage.

[0139] Based on the initial base map of the third stage, the lighting re-rendering algorithm is used to re-render the initial base map of the third stage to obtain the base map of the third stage, and the base map of the third stage is determined as the product poster corresponding to the target product.

[0140] The product in the product poster is the target product, the interaction between the object in the product poster and the target product meets the interaction prompt words, and the style conditions of the background area of ​​the product poster are consistent with the background reference image.

[0141] Specifically, style transfer models can be style-based generative adversarial networks (StyleGAN) or CLIP-based image style transfer (CLIPstyler), used to extract and transfer image style features. Segmentation models are semantic segmentation-based deep learning models, such as fully convolutional networks (FCN) and segment anything models (SAM), used to separate target products / objects from the background in an image and generate segmentation masks. Background reference images are user-provided images for style reference (such as city night scenes, landscape paintings, retro styles, etc.), used to control the visual style of the final poster background. Relighting algorithms are image processing or deep learning-based algorithms that optimize the illumination distribution of an image by adjusting lighting parameters (including brightness, contrast, shadows, and highlights). Relighting algorithms can be RelightingGAN.

[0142] The first-stage base image is input into a pre-trained segmentation model. The segmentation model extracts image features and outputs a segmentation mask. In the segmentation mask, the area containing the target product and object is marked as "1" (foreground), and the remaining areas are marked as "0" (background), clearly distinguishing the foreground to be retained from the background to be optimized. A background reference image is obtained and input into a style transfer model. The style transfer model's feature extraction network extracts style features from the reference image, obtaining a style conditional feature vector containing information such as color, texture, and composition. Based on the segmentation mask, the pixel coordinates of the foreground area are extracted from the first-stage base image, while the remaining pixels in the background area are extracted, ensuring that style optimization is performed only on the background area to avoid affecting foreground details. The style conditional feature vector and the pixel values ​​of the background area in the first-stage base image are input into a diffusion model, such as Stable Diffusion. Through an iterative denoising process, the RGB values, texture details, and lighting of the background area are adjusted so that the style features of the newly generated background area are consistent with the style conditional feature vector. After a preset number of diffusion steps, a background area with a style consistent with the reference image is obtained. The optimized background area is merged with the foreground area of ​​the first-stage base image to generate the third-stage initial base image. The background style of the third-stage initial base image is consistent with the reference image, and the details of the products and objects in the third-stage initial base image are fully preserved. By introducing image segmentation, style transfer, and conditional diffusion, efficient stylized redrawing of the poster background is achieved, allowing for personalized and diverse customization of the background style while preserving the display effect of the products and objects. A re-lighting algorithm is used to re-render the lighting of the third-stage initial base image, further optimizing the lighting effects of the background. For example, natural shadow transitions are generated based on the direction of the light source, and subtle gloss is added to the highlight areas of the background to enhance the realism of the materials, resulting in a more natural and realistic lighting effect in the output third-stage base image. This addresses the unnatural lighting issues that may exist after diffusion processing, making the background closer to a real scene. Finally, the third-stage base image is determined as the product poster corresponding to the target product.

[0143] For details, please refer to Figure 4 , Figure 4 A schematic diagram of another poster generation method is provided for generating a third-stage base map. The first-stage base map is input into a segmentation model to segment the target product and objects, obtaining a segmentation mask. A background reference image is acquired, and features are extracted from it using a style transfer model to obtain a corresponding style conditional feature vector. Based on the segmentation mask, the foreground and background regions are determined. The style conditional feature vector and the segmentation mask are input into a diffusion model to diffuse the background region in the segmentation mask, generating the initial third-stage base map. Based on this initial third-stage base map, a relighting algorithm is used to re-render the initial third-stage base map, resulting in the final third-stage base map.

[0144] In some embodiments, before the step of determining the second-stage base map as the product poster corresponding to the target product, the method further includes:

[0145] The second-stage base map is input into the segmentation model, and the target products and target objects in the second-stage base map are segmented to obtain the second segmentation mask corresponding to the second-stage base map.

[0146] Obtain a background reference image, and extract features from the background reference image using a style transfer model to obtain the corresponding style conditional feature vector;

[0147] Based on the second segmentation mask, determine the foreground region and background region in the second segmentation mask;

[0148] Based on the style conditional feature vector, the background region in the second segmentation mask is diffused to generate the initial base map for the third stage.

[0149] Based on the initial base map of the third stage, the lighting re-rendering algorithm is used to re-render the initial base map of the third stage to obtain the base map of the third stage, and the base map of the third stage is determined as the product poster corresponding to the target product.

[0150] In this context, the product in the product poster is the target product, the interaction between the object in the product poster and the target product satisfies the interaction prompt words, and the style conditions of the background area of ​​the product poster are consistent with the background reference image.

[0151] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the poster generation method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0152] This application also provides a poster generating device, please refer to... Figure 5 The poster generating device includes:

[0153] The feature annotation module 501 is used to extract feature annotations of the target product based on the product image corresponding to the target product, and to obtain product prompt words based on the feature annotations;

[0154] The text encoding module 502 is used to obtain object prompts and interactive action prompts, and to perform text encoding on the product prompts, object prompts and interactive action prompts to obtain the corresponding text embedding vectors;

[0155] The latent space encoding module 503 is used to perform latent space encoding on the product image to obtain the corresponding image latent space feature vector;

[0156] The diffusion module 504 is used to input the image latent space feature vector and the text embedding vector into the multimodal conditional generation optimization diffusion model, and generate the first-stage base map through the multimodal conditional generation optimization diffusion model;

[0157] The generation module 505 is used to generate product posters corresponding to the target product based on the base map of the first stage.

[0158] The poster generation apparatus provided in this application, employing the poster generation method described in the above embodiments, can solve the technical problem of prior art poster generation heavily relying on masks. Compared with the prior art, the beneficial effects of the poster generation apparatus provided in this application are the same as those of the poster generation method provided in the above embodiments, and other technical features in the poster generation apparatus are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0159] This application provides a poster generation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the poster generation method in the first embodiment described above.

[0160] The following is for reference. Figure 6 The diagram illustrates a structural schematic of a poster generation device suitable for implementing embodiments of this application. The poster generation device in these embodiments may include, but is not limited to, mobile terminals such as mobile phones, laptops, tablets (Portable Application Description, PADs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 6 The poster generation device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0161] like Figure 6As shown, the poster generating device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1002 or a program loaded from storage device 1003 into random access memory (RAM) 1004. The random access memory 1004 also stores various programs and data required for the operation of the poster generating device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the poster generating device to communicate wirelessly or wiredly with other devices to exchange data. Although a poster generating device with various systems is shown in the figure, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.

[0162] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0163] The poster generation device provided in this application, employing the poster generation method described in the above embodiments, can solve the technical problem of prior art poster generation heavily relying on masks. Compared with the prior art, the beneficial effects of the poster generation device provided in this application are the same as those of the poster generation method provided in the above embodiments, and other technical features of this poster generation device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0164] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0165] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0166] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the poster generation method in the above embodiments.

[0167] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0168] The aforementioned computer-readable storage medium may be included in the poster generating device; or it may exist independently and not be assembled into the poster generating device.

[0169] The aforementioned computer-readable storage medium carries one or more programs. When the aforementioned one or more programs are executed by the poster generation device, the poster generation device causes the poster generation device to: extract feature annotations of the target product based on the product image corresponding to the target product, and obtain product prompt words based on the feature annotations; perform text encoding on the product prompt words, object prompt words, and interactive action prompt words to obtain corresponding text embedding vectors; perform latent space encoding on the product image to obtain corresponding image latent space feature vectors; input the image latent space feature vectors and text embedding vectors into a multimodal conditional generation optimization diffusion model, and generate a first-stage base map through the multimodal conditional generation optimization diffusion model; and generate a product poster corresponding to the target product based on the first-stage base map.

[0170] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0171] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0172] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0173] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described poster generation method, which can solve the technical problem that poster generation in the prior art heavily relies on masks. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the poster generation method provided in the above embodiments, and will not be repeated here.

[0174] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the poster generation method described above.

[0175] The computer program product provided in this application can solve the technical problem that poster generation in the prior art heavily relies on masks. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the poster generation method provided in the above embodiments, and will not be repeated here.

[0176] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A poster generation method, characterized in that, The poster generation method includes: Based on the product image corresponding to the target product, feature annotations of the target product are extracted, and product prompt words are obtained based on the feature annotations; Obtain object prompts and interaction prompts, and perform text encoding on the product prompts, object prompts, and interaction prompts to obtain corresponding text embedding vectors; The product image is latent space encoded to obtain the corresponding image latent space feature vector; The image latent space feature vector and the text embedding vector are input into a multimodal conditional generation optimization diffusion model, and a first-stage base map is generated through the multimodal conditional generation optimization diffusion model; Generate a product poster corresponding to the target product based on the base map of the first stage.

2. The method as described in claim 1, characterized in that, The step of extracting feature annotations of the target product based on the product image corresponding to the target product, and obtaining product prompt words based on the feature annotations, includes: Based on the field lookup table associated with the preset structured template, the product target attributes corresponding to the fields to be filled in the preset structured template are determined. A multimodal large model is used to extract the target attributes of the target product in the product image to obtain the target attribute description text corresponding to the target product; The target attribute description text is determined as the product prompt word.

3. The method as described in claim 1, characterized in that, The step of inputting the image latent space feature vector and the text embedding vector into a multimodal conditional generation optimization diffusion model, and generating a first-stage base map through the multimodal conditional generation optimization diffusion model includes: The text embedding vector is spatially expanded by linear projection to obtain the corresponding text spatial feature vector; Cross-attention processing is performed based on the latent space feature vector of the image and the feature vector of the text space to obtain the corresponding cross-attention processing result; Multi-scale feature extraction is performed on the cross-attention processing results to obtain feature processing results at multiple different scales; The latent space feature vector of the image and the feature vector of the text space are concatenated, and the corresponding concatenation result is convolved to obtain the corresponding pixel-wise weight map. A multimodal fusion feature vector is obtained by weighting and fusing the pixel-by-pixel weight map, the image latent space feature vector, and the text space feature vector. Using the multimodal fusion feature vector and the feature processing result as multimodal conditions, the U-Net network that optimizes the diffusion model is used to predict noise and gradually denoise it to generate the latent vector of the first-stage base map. The latent vectors of the first-stage base map are decoded to obtain the first-stage base map.

4. The method as described in claim 3, characterized in that, The step of using the multimodal fused feature vector and the feature processing result as multimodal conditions, and using the multimodal conditions as input to optimize the diffusion model to predict noise and gradually denoise, to generate the latent vector of the first-stage base map includes: Obtain a noise map, perform latent space encoding on the noise map, and obtain the corresponding noise map feature vector; The noise map feature vector is input into the U-Net network, the feature processing result is input into each downsampling block of the U-Net network, the multimodal fusion feature vector is input into each downsampling block and each upsampling block of the U-Net network, and iterative diffusion processing is performed on the noise map feature vector, the multimodal fusion feature vector and the feature processing result based on each downsampling block and each upsampling block to obtain the latent vector of the first stage base map.

5. The method as described in claim 1, characterized in that, The step of generating the product poster corresponding to the target product based on the first stage base map includes: Obtain the facial image of the target object; Based on the first stage base map, determine the facial region of the object in the first stage base map; Based on the facial image, a face-swapping algorithm is used to match and fuse the facial region of the object in the first stage base image to obtain the second stage initial base image; The product poster corresponding to the target product is generated based on the initial base map of the second stage.

6. The method as described in claim 5, characterized in that, The step of generating the product poster corresponding to the target product based on the initial base map of the second stage includes: The facial region of the object in the initial base map of the second stage is re-textured using a facial texture enhancement algorithm to obtain the base map of the second stage. The second-stage base map is determined as the product poster corresponding to the target product; Wherein, the product in the product poster is the target product, the interaction between the object in the product poster and the target product satisfies the interaction prompt words, and the object in the product poster is the target object.

7. The method as described in claim 1, characterized in that, The step of generating the product poster corresponding to the target product based on the first stage base map includes: The first-stage base map is input into the segmentation model to segment the target product and the object in the first-stage base map, thereby obtaining the segmentation mask corresponding to the first-stage base map. A background reference image is obtained, and features are extracted from the background reference image using a style transfer model to obtain the corresponding style condition feature vector; Based on the segmentation mask, determine the foreground region and background region in the segmentation mask; Based on the style conditional feature vector, the background region in the segmentation mask is diffused to generate the initial base map for the third stage. Based on the initial base map of the third stage, the lighting re-rendering algorithm is used to re-render the initial base map of the third stage to obtain the base map of the third stage, and the base map of the third stage is determined as the product poster corresponding to the target product. The product in the product poster is the target product, the interaction between the object in the product poster and the target product satisfies the interaction prompt words, and the style conditions of the background area of ​​the product poster are consistent with the background reference image.

8. The method as described in claim 6, characterized in that, Before the step of determining the second-stage base map as the product poster corresponding to the target product, the method further includes: The second-stage base map is input into the segmentation model, and the target product and the target object in the second-stage base map are segmented to obtain the second segmentation mask corresponding to the second-stage base map. A background reference image is obtained, and features are extracted from the background reference image using a style transfer model to obtain the corresponding style condition feature vector; Based on the second segmentation mask, determine the foreground region and background region in the second segmentation mask; Based on the style conditional feature vector, the background region in the second segmentation mask is diffused to generate the initial base map of the third stage. Based on the initial base map of the third stage, the lighting re-rendering algorithm is used to re-render the initial base map of the third stage to obtain the base map of the third stage, and the base map of the third stage is determined as the product poster corresponding to the target product. Wherein, the product in the product poster is the target product, the interaction between the object in the product poster and the target product satisfies the interaction prompt words, and the style conditions of the background area of ​​the product poster are consistent with the background reference image.

9. A poster generation device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the poster generation method as described in any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the poster generation method as described in any one of claims 1 to 8.