Method and device for generating advertisement image and storage medium
By generating text control parameters and image space control parameters, combining noise images, and using AI image generation models to generate advertising images, the problem of insufficient quality and consistency of advertising images in the prior art is solved, and efficient and detailed advertising image generation is achieved.
Patent Information
- Application Number
- CN202311733904.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art is difficult to generate high-quality advertising images, especially when restoring product details and adapting to rare products.
Generate text control parameters based on text descriptions, generate image spatial control parameters using foreground image blocks and single background colors, and combine noise images, generate initial product images using AI image generation models, and then generate high-quality advertising images through replacement and post-processing.
It improves the production efficiency and quality of advertising images, fully reflects the details of the product, and improves the consistency between the product and the actual product in advertising images.
Smart Images

Figure CN120163898A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to image processing, and more particularly, to a computer-implemented method for generating advertisement images, an apparatus for generating advertisement images, and a computer-readable non-transitory storage medium storing a program. Background Art
[0002] With the development of computer science and artificial intelligence, it has become increasingly common and effective to use a computer to run an artificial intelligence (AI) model based on a neural network to implement image processing. Image generation is an important application field of artificial intelligence models.
[0003] Product advertisements containing product images help to display the features and effects of the target product, help to promote the product, help to increase the popularity of the product, and thus increase the sales volume and profit of the product. Traditional advertisement production requires manually creating product images that display the features of the product as advertisement images for the product. A large amount of manpower and time are required in traditional advertisement image production. Even when creating advertisement images based on basic templates, a large amount of manpower and time are still required.
[0004] Currently, methods for generating images based on generative AI models have emerged. Although such AI models can generate creative images, they cannot restore the image details of the product. Moreover, for particularly rare products, existing AI models usually cannot generate advertisement images that are consistent with the products (product image blocks) in the product images of the products.
[0005] Generating high-quality advertisement images of products is desirable. Summary of the Invention
[0006] A brief overview of the present disclosure will be given below to provide a basic understanding of certain aspects of the present disclosure. It should be understood that this overview is not an exhaustive overview of the present disclosure. It is not intended to identify the key or important parts of the present disclosure, nor is it intended to limit the scope of the present disclosure. Its purpose is merely to present certain concepts in a simplified form as a prelude to a more detailed description to be presented later.
[0007] The inventors have studied image generation models and image processing techniques and proposed an improved method for generating advertisement images, which is expected to have advantages in terms of efficiency and image quality.
[0008] According to one aspect of the present disclosure, a computer-implemented method for generating an advertisement image is provided. The method includes: generating text control parameters for an AI image generation model based on a text description of an advertisement image of a product; generating image space control parameters for the AI image generation model based on a product image with an image block indicating the product as the foreground and a single color as the background; generating a noise image for the AI image generation model based on the product image; using the AI image generation model to generate an initial product image based on the text control parameters, the image space control parameters, and the noise image; generating a restored product image by replacing corresponding regions in the initial product image with image blocks; and generating an advertisement image of the product by performing post-processing on the restored product image to fuse a foreground image and a background image.
[0009] According to one aspect of the present disclosure, an apparatus for generating an advertisement image is provided. The apparatus includes: a memory storing instructions thereon; and at least one processor. The at least one processor executes the instructions to implement the aforementioned method for generating an advertisement image.
[0010] According to another aspect of the present disclosure, a computer-readable non-transitory storage medium storing a program is provided. When the program is executed by a computer, the program causes the computer to perform the following operations: generating text control parameters for an AI image generation model based on a text description of an advertisement image of a product; generating image space control parameters for the AI image generation model based on a product image with an image block indicating the product as the foreground and a single color as the background; generating a noise image for the AI image generation model based on the product image; using the AI image generation model to generate an initial product image based on the text control parameters, the image space control parameters, and the noise image; generating a restored product image by replacing corresponding regions in the initial product image with image blocks; and generating an advertisement image of the product by performing post-processing on the restored product image to fuse a foreground image and a background image.
[0011] The beneficial effects of the method, apparatus, and storage medium of the present disclosure include at least one of the following effects: improving the efficiency of producing advertisement images; improving the quality of advertisement images; fully reflecting the details of products; and improving the consistency between the products in advertisement images and the actual products. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Embodiments of the present disclosure will be described below with reference to the accompanying drawings, which will help to more easily understand the above and other objects, features, and advantages of the present disclosure. The drawings are only for showing the principles of the present disclosure. The dimensions and relative positions of the units do not have to be drawn to scale in the drawings. The same reference numerals may represent the same features. In the drawings:
[0013] Figure 1Shows an exemplary flowchart of a method for generating an advertisement image according to an embodiment of the present disclosure;
[0014] Figure 2 Shows an exemplary flowchart of a method for generating a product image according to an embodiment of the present disclosure;
[0015] Figure 3 Shows an exemplary original image of a product according to an embodiment of the present disclosure;
[0016] Figure 4 Shows an exemplary product image according to an embodiment of the present disclosure;
[0017] Figure 5 Shows an exemplary edge image according to an embodiment of the present disclosure;
[0018] Figure 6 Shows an image corresponding to a mask according to an embodiment of the present disclosure;
[0019] Figure 7 Shows an exemplary initial product image according to an embodiment of the present disclosure;
[0020] Figure 8 Shows an exemplary Poisson fusion image according to an embodiment of the present disclosure;
[0021] Figure 9 Shows an exemplary second smoothed image according to an embodiment of the present disclosure;
[0022] Figure 10 Shows an exemplary block diagram of an apparatus for generating an advertisement image according to an embodiment of the present disclosure; and
[0023] Figure 11 Shows an exemplary block diagram of an information processing device according to an embodiment of the present disclosure. Detailed Description of the Invention
[0024] Hereinafter, exemplary embodiments of the present disclosure will be described in conjunction with the accompanying drawings. For clarity and conciseness, not all features of the actual embodiments are described in the specification. However, it should be understood that many embodiment-specific decisions may be made during the development of any such actual embodiment to achieve the specific goals of the developer, and these decisions may vary depending on the embodiment.
[0025] Here, it should also be noted that in order to avoid obscuring the present disclosure with unnecessary details, only the device structures closely related to the solutions according to the present disclosure are shown in the drawings, while other details less related to the present disclosure are omitted.
[0026] It should be understood that the present disclosure is not limited to the described embodiments only by the following description with reference to the drawings. In this article, where feasible, embodiments can be combined with each other, features can be replaced or borrowed between different embodiments, and one or more features can be omitted in one embodiment.
[0027] The computer program code for performing the operations of the various aspects of the embodiments of the present disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, C++, etc., and also including conventional procedural programming languages such as the "C" programming language or similar programming languages.
[0028] The method of the present disclosure can be implemented by a circuit with corresponding functional configurations. The circuit includes a circuit for a processor.
[0029] One aspect of the present disclosure relates to a method for generating an advertisement image. The method can be implemented by a computer. The following refers to Figure 1 An exemplary description of the method is given.
[0030] Figure 1 An exemplary flowchart of a method 100 for generating an advertisement image according to an embodiment of the present disclosure is shown.
[0031] In operation Op101, a text control parameter c_txt for an AI image generation model imM is generated based on a text description txt of an advertisement image of a product prdt. Exemplarily, the text control parameter c_txt can be a vector having a predetermined number of elements.
[0032] For example, the product prdt can be a bottled beverage with specified specifications and packaging. The product has various features involving details. For example, the color, material, shape, height, transparency, pattern, texture, label, pattern, etc. of the container (bottle).
[0033] The text description txt can include a description of the desired background and foreground in the image to be generated, where the description of the foreground is associated with the product prdt. For example, when the product prdt is a bottled beverage, the text description txt can include the noun "bottle".
[0034] An exemplary text description txt for a bottled beverage could be: "A close-up of the crystal-clear bottle of water floating in mid-air against a backdrop of a tranquil waterfall".
[0035] In one example, the text description txt can be text manually input by the user. In one example, the text description txt can be generated by speech recognition software to recognize the user's oral input. In another example, a language model can be used to generate the text description txt. Exemplarily, the language model can be ChatGPT (Chat Generative Pre-trained Transformer, a natural language processing tool driven by artificial intelligence technology developed by OpenAI). ChatGPT is conventional technology and will not be elaborated here.
[0036] Various text encoders can be used to convert the text description txt into a text control parameter c_txt. In one example, a CLIP text encoder can be used to generate the text control parameter c_txt based on the text description txt. For the CLIP text encoder, reference can be made to the article "Learning Transferable Visual Models From Natural Language Supervision" by Alec Radford et al. shown in the following link:
[0037] https: / / arxiv.org / pdf / 2103.00020.pdf.
[0038] The CLIP text encoder is conventional technology and will not be elaborated here.
[0039] The AI image generation model imM can generate an image based on the text control parameter, the image space control parameter, and the noise image. Exemplarily, the AI image generation model imM can be ControlNet. Regarding ControlNet, for example, reference can be made to the article "Adding Conditional Control to Text-to-Image Diffusion Models" by Lvmin Zhang et al. The acquisition link for this article is:
[0040] https: / / arxiv.org / pdf / 2302.05543.pdf.
[0041] ControlNet is a conventional technology and will not be described in detail here. It can be understood that the AI image generation model imM can also be a user-customized image generation model with reference to ControlNet. The customized model can be, for example, a model trained using samples desired by the user.
[0042] In operation Op103, an image space control parameter c_prdt for the AI image generation model imM is generated based on a product image F with an image block Bp indicating a product prdt in the foreground and a single color Cbg in the background. The image block Bp may come from an original product image Imo provided by a user. The image block Bp carries information about the details of the actual product. The image block Bp may be obtained by segmenting an image of a product that the user wants to display in an advertisement. The original product image Imo preferably has a resolution greater than or equal to a predetermined resolution. The original product image Imo may be, for example, an image obtained by photographing an actual product at a predetermined angle, predetermined lighting, and predetermined background. The original product image Imo may be, for example, a hand-drawn image or a product image drawn using drawing software under manual operation. The shape of the image block Bp is determined by the contour of the product prdt in the original product image Imo.
[0043] In operation Op105, a noise image z is generated for the AI image generation model imM based on the product image F. p,r .
[0044] In operation Op107, the AI image generation model imM is used based on the text control parameter c_txt, the image space control parameter c_prdt and the noise image z p,r Generate initial product images ori . When generating the initial product image using ControlNet ori In the case of , the generation operation can be expressed as formula (1).
[0045] gen ori =ControlNet(c _txt ,c_ prdt ,z p,r ;θ) (1)
[0046] Among them, θ represents multiple configuration parameters of the model. The initial product image gen ori The product prdt' in is inconsistent with the product prdt shown in the image block Bp. For example, in the case where the product prdt is a white car, the initial product image gen generated by the model ori The car in the picture is red which is different from white. Therefore, the initial product image is not suitable to be used directly as an advertising image for the product prdt.
[0047] In operation Op109, a restored product image gen is generated by replacing a corresponding region in the initial product image gen with an image patch Bp ori in it. That is, the image patch Bp is pasted into the corresponding region in the image generated by the model to replace or substantially replace the product image patch in the initial product image. prdt In operation Op111, an advertisement image I_final of the product prdt is generated by performing post-processing post_Pr that fuses a foreground image and a background image on the restored product image gen
[0048] prdt Method 100 can be configured to automatically output an advertisement image of a product based on an input text description txt and an original product image Imo. Therefore, the efficiency of producing advertisement images can be improved, the innovativeness of advertisement images can be enhanced, the labor cost can be reduced, and the time cost can be decreased. In method 100, the product image patch in the final advertisement image can be substantially considered to come from the original product image of the product desired by the user, rather than being completely generated by an image model. Therefore, the consistency between the product in the advertisement image and the actual product can be ensured, and the details and features of the actual (real) product can be fully reflected. Due to the post-processing, in the vicinity of the product contour line in the synthesized final advertisement image, the foreground-background transition will be more natural and appear more realistic. That is, method 100 can improve the quality of advertisement images and the effect of advertisements in terms of consistency, details, and realism. Generally, method 100 can ensure the high quality of advertisement images while efficiently generating advertisement images.
[0049] The details of each operation in method 100 are further described below.
[0050]
[0051] In one embodiment, generating text control parameters c_txt for an AI image generation model includes: generating a text description txt using a language model; and generating a feature vector of the text description using a text encoder as the text control parameters c_txt. The language model includes but is not limited to ChatGPT. The text encoder includes but is not limited to a CLIP text encoder.
[0052] Figure 2 FIG. shows an exemplary flowchart of a method 200 for generating a product image F according to an embodiment of the present disclosure. Method 200 includes operations Op201, Op203, and Op205,
[0053] In operation Op201, the foreground image Bp' indicating the product prdt in the original image Imo is extracted by segmenting the original image Imo provided by the user. In one example, a segmentation model such as the model SAM (SegmentAnything Model) is used to extract the foreground image Bp' of the product prdt from the original image Imo. The segmentation model SAM is a conventional technique. For example, reference can be made to the article "SegmentAnything" by Alexander Kirillov et al. provided at the following link: https: / / arxiv.org / pdf / 2304.02643.pdf.
[0054] Figure 3 An exemplary original image Imo of the product is shown. For simplicity, features including patterns, labels, colors, etc. are omitted, and the original image of the product, which is schematically shown in grayscale as a bottled drink, is presented. It can be understood that the original image of the product actually used is usually a color image and includes details such as labels, trademarks, patterns, etc.
[0055] In operation Op203, an image patch Bp of the product is generated by scaling the foreground image Bp' based on a predetermined image size S. The predetermined image size S is, for example, the size of the advertisement image Im_final desired by the user for the product prdt. Generally, in order to display the entire product in the original product image in the advertisement image, the scaling is performed such that the maximum width w of the image patch Bp is less than or equal to the width W of the advertisement image Im_final (e.g., w < W, w ≤ 0.80W, w ≤ 0.62W, w ≤ 0.50W, or w ≤ 0.40W), and the scaling is performed such that the maximum height h of the image patch Bp is less than or equal to the height H of the advertisement image Im_final (h < H, h ≤ 0.80H, h ≤ 0.62H, h ≤ 0.50H, or h ≤ 0.40H). Operation Op203 can be omitted. For example, if the size of the foreground image Bp' is already suitable for pasting into the advertisement image Im_final of a predetermined size for displaying the product prdt, the foreground image Bp' can be directly regarded as the product image patch Bp.
[0056] In operation Op205, a product image F is generated based on the image patch Bp, a single color Cbg, and the predetermined position p(x, y) of the image patch Bp. The predetermined position p(x, y) can be the coordinates of the upper left corner of the bounding box of the image patch Bp in the product image F.
[0057] Figure 4An exemplary product image F according to an embodiment of the present disclosure is shown, which includes an image block Bp as the foreground and a white background (i.e., the single color Cbg is "white"). The image block Bp is a combination of a bottle body and a bottle cap that forms the image block, and it can be obtained by segmenting the original image Imo. It can be understood that the single color Cbg includes but is not limited to "white". Here, for simplicity, the product image F is shown as a grayscale image. It can be understood that in actual application scenarios, the product image F is usually a color image.
[0058] In one embodiment, in method 100, generating the image space control parameter c_prdt for the AI image generation model imM includes: generating an edge image Fe by performing edge detection on the product image F; and encoding the edge image Fe by a neural network to generate the features of the product image F as the image space control parameter c_prdt, where pi is the input of the neural network f (here, it can be Fe), and θ is the configuration parameter of the neural network f. In one example, the Canny edge detection algorithm is used for edge detection. The Canny edge detection algorithm is a conventional technique. For example, the Canny detection function is included in the cross-platform computer vision and machine learning software library OpenCV. Additionally, for example, the description provided in the following article can also be referred to:
[0059] A Computational Approach To Edge Detection;JOHN CANNY,IEEETransactions on Pattern Analysis and Machine Intelligence;Volume:PAMI-8,Issue:6,November 1986;page(s)679-698.
[0060] Figure 5 A schematic edge image Fe according to an embodiment of the present disclosure is shown. It can be seen that the edge image gives a lot of edge information including the product contour line. This is useful for generating product images using the image generation model.
[0061] In one embodiment, in method 100, generating the noise image z p,r includes the following operations: generating a product noise image z by iteratively and progressively adding noise to the product image F p ; generating a random noise image z r ; and combining the product noise image z p and the random noise image z r based on a mask msk indicating the coverage area of the image block Bp in the product image F as the noise image zp,r The mask msk is, for example, a two-dimensional matrix corresponding to the product image F, and the matrix element m i,j is determined according to Equation (2).
[0062]
[0063] where pix i,j is the pixel at (i, j) in the product image F.
[0064] The product noise image z p and the random noise image z r can be combined according to Equation (3) to obtain the noise image z p,r .
[0065]
[0066] where is the pixel at (i, j) in the noise image z p,r , is the pixel at (i, j) in the product noise image z p , is the pixel at (i, j) in the random noise image z r .
[0067] Here, an example of the image corresponding to the mask msk is as shown Figure 6 , where the mask is schematically shown as a black-and-white image. In Figure 6 , the outline of the combined body of the bottle body and the bottle cap that forms the image block (Bp) can be recognized (at the position where the black-and-white colors change). That is, the mask msk can indicate the positions of the pixel points on the outline of the image block Bp in the product image F. The corresponding region in the initial product image can be replaced with the image block Bp based on this mask to generate the restored product image.
[0068] In one embodiment, in method 100, generating the restored product image gen prdt includes: combining the product image F and the initial product image gen ori based on the mask indicating the coverage area of the image block Bp in the product image F as the restored product image. For example, the restored product image gen prdt is synthesized according to Equation (4).
[0069]
[0070] where is the pixel at (i, j) in the restored product image gen prdt , is the pixel at (i, j) in the initial product image gen ori , Is the pixel at (i, j) in the product image F.
[0071] Figure 7 Shows an initial product image according to an embodiment of the present disclosure. Comparing Figure 4 , Figure 7 It can be seen that the product in the initial product image generated by the model is inconsistent with the product in the original product image, and there are differences in details.
[0072] In one embodiment, in method 100, the AI image generation model imM is configured to generate the initial product image gen p,r by progressively removing the noise in the noise image z ori .
[0073] In one embodiment, in method 100, the post-processing may include the following processes: Poisson fusion post-processing, boundary transition smoothing post-processing, and Gaussian filtering smoothing post-processing.
[0074] The Poisson fusion post-processing includes: performing Poisson image fusion on the restored product image and the image patch to obtain the Poisson fusion image Img possion . Poisson fusion is a conventional image processing technique. For example, the article "Poisson Image Editing" by Patrick Perez et al. provided at the following link can be referred to: https: / / www.cs.jhu.edu / ~misha / Fall07 / Papers / Perez03.pdf.
[0075] Figure 8 Shows an exemplary Poisson fusion image according to an embodiment of the present disclosure. Compared with the restored product image, the Poisson fusion image will appear more natural and realistic near the contour line of the product, weakening the synthesis traces. Here, the Poisson fusion image is exemplarily shown as a grayscale image. It can be understood that in practical applications, the Poisson fusion image is usually a color image.
[0076] The boundary transition smoothing post-processing includes: performing boundary transition smoothing on the inner region inside the contour line of the product in the Poisson fusion image Img possion based on the boundary of the image patch Bp in the product image F, so as to obtain the first smoothed image Img1. The first smoothed image Img1 can be determined according to Equation (5).
[0077]
[0078] Where is the pixel at (i, j) in the first smoothed image Img1, the pixel at (i, j) in the product image F, is the Poisson fusion image Imgpossion The pixel at (i,j) in the i,j For image Img possion The minimum distance from the pixel at (i, j) to the contour line, N is the first smoothing threshold. That is, in the post-processing of boundary transition smoothing, the Poisson fusion image Img possion The pixels in the inner area with a width of N inside the contour line of the product are smoothed, where the contour line corresponds to the boundary of the image block Bp in the product image F, and the boundary can be determined by the mask msk.
[0079] The Gaussian filter smoothing post-processing includes: performing Gaussian filter smoothing on the contour area near the contour line of the product in the first smoothed image Img1 to obtain a second smoothed image Img2 as the product's advertising image I_final. The second smoothed image Img2 can be determined according to formula (6).
[0080]
[0081] in, is the pixel at (i, j) in the second smoothed image Img2, is the pixel at (i, j) in the first smoothed image Img1, Gaussian is the Gaussian filter function, d i,j is the minimum distance from the pixel at (i, j) in the image Img1 to the contour line of the product, and M is the second smoothing threshold. That is, in the Gaussian filter smoothing post-processing, the pixels in the contour area with a width of 2M with the contour line of the product as the center line in the first smoothed image Img1 are smoothed to smooth the contour area near the contour line, especially the jagged edges near the contour line, and weaken the image synthesis traces and synthesis defects in the advertising image. It can be understood that the image synthesis traces and synthesis defects will become visible or obvious after the image is enlarged, affecting the quality of the advertising image, and therefore need to be weakened.
[0082] Figure 9 An exemplary second smoothed image is shown. Compared with the restored product image, the second smoothed image will have a smoother and more natural transition near the product contour, thereby improving the quality of the synthesized image. Here, the second smoothed image is exemplarily shown as a grayscale image. It can be understood that in practical applications, the second smoothed image is usually a color image.
[0083] In the above embodiment, although three post-processings are shown to be performed sequentially, it is understandable that only one or two of them may be performed when the image quality requirement is not high. In addition, in the present disclosure, post-processing may also include other edge smoothing processes in the image field.
[0084] One aspect of the present disclosure relates to an apparatus for generating an advertising image. Figure 10An exemplary description of the device is given.
[0085] Figure 10 An exemplary block diagram of a device 1000 for generating an advertisement image according to an embodiment of the present disclosure is shown.
[0086] The device 1000 includes: a memory 1001 on which instructions Inst are stored; and at least one processor 1003, which is connected to the memory and configured to execute the instructions to implement the method 100. Exemplarily, executing the instructions can perform the following operations: generating text control parameters for an AI image generation model based on a text description of an advertisement image of a product; generating image space control parameters for the AI image generation model based on a product image with an image block indicating the product as the foreground and a single color as the background; generating a noise image for the AI image generation model based on the product image; using the AI image generation model to generate an initial product image based on the text control parameters, the image space control parameters, and the noise image; generating a restored product image by replacing corresponding regions in the initial product image with image blocks; and generating an advertisement image of the product by performing post-processing on the restored product image to fuse a foreground image and a background image.
[0087] According to another aspect of the present disclosure, a computer-readable non-transitory storage medium storing a program is provided. When the program is executed by a computer, the program causes the computer to perform the following operations: generating text control parameters for an AI image generation model based on a text description of an advertisement image of a product; generating image space control parameters for the AI image generation model based on a product image with an image block indicating the product as the foreground and a single color as the background; generating a noise image for the AI image generation model based on the product image; using the AI image generation model to generate an initial product image based on the text control parameters, the image space control parameters, and the noise image; generating a restored product image by replacing corresponding regions in the initial product image with image blocks; and generating an advertisement image of the product by performing post-processing on the restored product image to fuse a foreground image and a background image. More details of the program can be referred to the description of the method 100.
[0088] According to one aspect of the present disclosure, an information processing device is also provided.
[0089] Figure 11 is an exemplary block diagram of an information processing device 110 according to an embodiment of the present disclosure. In Figure 11In this case, the central processing unit (CPU) 1101 performs various processes according to a program stored in the read-only memory (ROM) 1102 or a program loaded from the storage device 1108 into the random access memory (RAM) 1103. In the RAM 1103, data and the like required when the CPU 1101 performs various processes are also stored as needed.
[0090] The CPU 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. The input / output interface 1105 is also connected to the bus 1104.
[0091] The following components are connected to the input / output interface 1105: an input device 1106 including a soft keyboard or the like; an output device 1107 including a display such as a liquid crystal display (LCD) or the like and a speaker or the like; a storage device 1108 such as a hard disk; and a communication device 1109 including a network interface card such as a LAN card, a modem, or the like. The communication device 1109 performs communication processing via a network such as the Internet, a local area network, a mobile network, or a combination thereof.
[0092] The drive 1110 is also connected to the input / output interface 1105 as needed. A removable medium 1111 such as a semiconductor memory or the like is mounted on the drive 1110 as needed, so that a program read therefrom is installed in the storage device 1108 as needed.
[0093] The CPU 1101 can run a program for a method applied to generate an advertisement image.
[0094] The beneficial effects of the method, device, and storage medium of the present disclosure include at least one of the following: improving the efficiency of producing an advertisement image; improving the quality of the advertisement image; fully reflecting the details of the product; improving the consistency between the product in the advertisement image and the actual product. In particular, even if the product in the original product image is rare, an advertisement image relatively consistent with the product in the original product image can be generated.
[0095] As described above, according to the present disclosure, the principle of improving the advertisement image generation process has been disclosed. It should be noted that the effects of the solution of the present disclosure are not necessarily limited to the above effects, and any effect shown in this specification or other effects that can be understood from this specification can be achieved in addition to or instead of the effects described in the previous paragraphs.
[0096] Although the present invention has been disclosed above by the description of specific embodiments of the present invention, it should be understood that those skilled in the art can design various modifications (including combinations or substitutions of features between embodiments where applicable), improvements, or equivalents of the present invention within the scope of the appended claims. These modifications, improvements, or equivalents should also be considered to be included within the scope of protection of the present disclosure.
[0097] It should be emphasized that the term "comprising / including" when used herein refers to the presence of features, elements, steps, or components, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0098] In addition, the methods of the embodiments of the present invention are not limited to being executed in the chronological order described in the specification or shown in the drawings, and can also be executed in other chronological orders, in parallel, or independently. Therefore, the execution order of the methods described in this specification does not limit the technical scope of the present invention.
[0099] Supplementary Note
[0100] The present disclosure includes but is not limited to the following solutions.
[0101] 1. A computer-implemented method for generating an advertisement image, characterized by comprising:
[0102] Generating text control parameters for an AI image generation model based on a text description of an advertisement image of a product;
[0103] Generating image space control parameters for the AI image generation model based on a product image with an image block indicating the product as the foreground and a single color as the background;
[0104] Generating a noise image for the AI image generation model based on the product image;
[0105] Using the AI image generation model to generate an initial product image based on the text control parameters, the image space control parameters, and the noise image;
[0106] Generating a restored product image by replacing a corresponding area in the initial product image with the image block; and
[0107] Generating the advertisement image of the product by performing post-processing on the restored product image to fuse a foreground image and a background image.
[0108] 2. The method according to item 1 above, wherein generating text control parameters for an AI image generation model includes:
[0109] Using a language model to generate the text description; and
[0110] Use a text encoder to generate a feature vector of the text description as the text control parameter.
[0111] 3. The method according to appendix 1, wherein the product image is generated by performing the following operations:
[0112] Extract a foreground image indicating the product in the original image by segmenting the original image of the product provided by the user;
[0113] Generate image patches of the product by scaling the foreground image based on a predetermined image size; and
[0114] Generate the product image based on the image patches, the single color, and the predetermined positions of the image patches;
[0115] Wherein the advertisement image has the predetermined image size.
[0116] 4. The method according to appendix 1, wherein generating the image space control parameter for the AI image generation model includes:
[0117] Generate an edge image by performing edge detection on the product image; and
[0118] Generate features of the product image as the image space control parameter by encoding the edge image with a neural network.
[0119] 5. The method according to appendix 1, wherein generating the noise image includes:
[0120] Generate a product noise image by iteratively and progressively adding noise to the product image;
[0121] Generate a random noise image; and
[0122] Combine the product noise image and the random noise image as the noise image based on a mask indicating the coverage area of the image patches in the product image.
[0123] 6. The method according to appendix 1, wherein generating the restored product image includes:
[0124] Combine the product image and the initial product image as the restored product image based on a mask indicating the coverage area of the image patches in the product image.
[0125] 7. The method according to appendix 1, wherein the AI image generation model is configured to generate the initial product image by progressively removing noise from the noise image.
[0126] 8. The method according to Note 1, wherein the post-processing includes:
[0127] Obtaining a Poisson fusion image by performing Poisson image fusion on the restored product image and the image block;
[0128] Obtaining a first smoothed image by performing boundary transition smoothing on the inner region inside the contour line of the product in the Poisson fusion image based on the boundary of the image block in the product image; and
[0129] Obtaining a second smoothed image as the advertisement image of the product by performing Gaussian filtering smoothing on the contour region near the contour line of the product in the first smoothed image.
[0130] 9. The method according to Note 8, wherein the inner region has a width indicated by a first smoothing threshold.
[0131] 10. The method according to Note 8, wherein the contour region has a width twice that indicated by a second smoothing threshold.
[0132] 11. An apparatus for generating an advertisement image, comprising:
[0133] A memory storing instructions thereon; and
[0134] At least one processor configured to execute the instructions to perform the following operations:
[0135] Generating text control parameters for an AI image generation model based on a text description of an advertisement image of a product;
[0136] Generating image space control parameters for the AI image generation model based on a product image with an image block indicating the product as the foreground and a single color as the background;
[0137] Generating a noise image for the AI image generation model based on the product image;
[0138] Using the AI image generation model to generate an initial product image based on the text control parameters, the image space control parameters, and the noise image;
[0139] Generating a restored product image by replacing a corresponding region in the initial product image with the image block; and
[0140] Generating the advertisement image of the product by performing post-processing of fusing a foreground image and a background image on the restored product image.
[0141] 12. The apparatus according to Note 11, wherein generating text control parameters for an AI image generation model includes:
[0142] Generate the text description using a language model; and
[0143] Generate a feature vector of the text description using a text encoder as the text control parameter.
[0144] 13. The apparatus according to claim 11, wherein the product image is generated by performing the following operations:
[0145] Extract a foreground image indicating the product in the original image by segmenting the original image of the product provided by the user;
[0146] Generate image patches of the product by scaling the foreground image based on a predetermined image size; and
[0147] Generate the product image based on the image patches, the single color, and the predetermined positions of the image patches;
[0148] wherein the advertisement image has the predetermined image size.
[0149] 14. The apparatus according to claim 11, wherein generating image space control parameters for the AI image generation model includes:
[0150] Generate an edge image by performing edge detection on the product image; and
[0151] Generate features of the product image as the image space control parameters by encoding the edge image with a neural network.
[0152] 15. The apparatus according to claim 11, wherein generating the noise image includes:
[0153] Generate a product noise image by iteratively and progressively adding noise to the product image;
[0154] Generate a random noise image; and
[0155] Combine the product noise image and the random noise image as the noise image based on a mask indicating a coverage area of the image patches in the product image.
[0156] 16. The apparatus according to claim 11, wherein generating the restored product image includes:
[0157] Combine the product image and the initial product image as the restored product image based on a mask indicating a coverage area of the image patches in the product image.
[0158] 17. The device according to Note 11, wherein the AI image generation model is configured to generate the initial product image by gradually removing noise from the noisy image.
[0159] 18. The device according to Note 11, wherein the post-processing includes:
[0160] Obtaining a Poisson fusion image by performing Poisson image fusion on the restored product image and the image patch;
[0161] Obtaining a first smoothed image by performing boundary transition smoothing on the inner region inside the contour line of the product in the Poisson fusion image based on the boundary of the image patch in the product image; and
[0162] Obtaining a second smoothed image as the advertisement image of the product by performing Gaussian filtering smoothing on the contour region near the contour line of the product in the first smoothed image.
[0163] 19. The device according to Note 18, wherein the inner region has a width indicated by a first smoothing threshold; and the contour region has a width twice that indicated by a second smoothing threshold.
[0164] 20. A computer-readable non-transitory storage medium storing a program, characterized in that when the program is executed by a computer, the program causes the computer to perform the following operations:
[0165] Generating text control parameters for the AI image generation model based on a text description of an advertisement image of a product;
[0166] Generating image space control parameters for the AI image generation model based on a product image with an image patch indicating the product as the foreground and a single color as the background;
[0167] Generating a noisy image for the AI image generation model based on the product image;
[0168] Using the AI image generation model to generate an initial product image based on the text control parameters, the image space control parameters, and the noisy image;
[0169] Generating a restored product image by replacing a corresponding region in the initial product image with the image patch; and
[0170] Generating the advertisement image of the product by performing post-processing on the restored product image to fuse a foreground image and a background image.
Claims
1. A computer-implemented method for generating an advertising image, characterized in that, Comprising: Generating text control parameters for an AI image generation model based on a text description of an advertisement image of a product; Generating image space control parameters for the AI image generation model based on a product image with an image block indicating the product as the foreground and a single color as the background; Generating a noise image for the AI image generation model based on the product image; Using the AI image generation model to generate an initial product image based on the text control parameters, the image space control parameters, and the noise image; Generating a restored product image by replacing a corresponding region in the initial product image with the image block; and Generating the advertisement image of the product by performing post-processing on the restored product image to fuse a foreground image and a background image.
2. The method according to claim 1, wherein, Generating the text control parameters for the AI image generation model includes: Using a language model to generate the text description; and Using a text encoder to generate a feature vector of the text description as the text control parameters.
3. The method according to claim 1, wherein, Generating the product image by performing the following operations: Extracting a foreground image indicating the product in the original image by segmenting the original image of the product provided by a user; Generating the image block of the product by scaling the foreground image based on a predetermined image size; And Generating the product image based on the image block, the single color, and a predetermined position of the image block; Wherein, the advertisement image has the predetermined image size.
4. The method according to claim 1, wherein, Generating the image space control parameters for the AI image generation model includes: Generating an edge image by performing edge detection on the product image; and Encoding the edge image by a neural network to generate features of the product image as the image space control parameters.
5. The method according to claim 1, wherein, Generating the noise image includes: Generating a product noise image by iteratively and progressively adding noise to the product image; Generating a random noise image; and Combining the product noise image and the random noise image based on a mask indicating a coverage area of the image block in the product image as the noise image.
6. The method according to claim 1, wherein, Generating the restored product image includes: Combining the product image and the initial product image based on a mask indicating a coverage area of the image block in the product image as the restored product image.
7. The method according to claim 1, wherein, The AI image generation model is configured to generate the initial product image by progressively removing noise from the noise image.
8. The method according to claim 1, wherein, The post-processing includes: Obtaining a Poisson fusion image by performing Poisson image fusion on the restored product image and the image block; Obtaining a first smoothed image by performing boundary transition smoothing on an inner region inside the contour line of the product in the Poisson fusion image based on a boundary of the image block in the product image; and Obtaining a second smoothed image as the advertisement image of the product by performing Gaussian filtering smoothing on a contour region near the contour line of the product in the first smoothed image.
9. An apparatus for generating an advertising image, characterized in that, Comprising: A memory storing instructions thereon; And At least one processor configured to execute the instructions to implement the method according to any one of claims 1 to 8.
10. A computer-readable non-transitory storage medium storing a program, characterized in that, When the program is executed by a computer, the program causes the computer to perform the following operations: Generate text control parameters for an AI image generation model based on a text description of an advertisement image of a product; Generate image space control parameters for the AI image generation model based on a product image with an image block indicating the product as the foreground and a single color as the background; Generate a noise image for the AI image generation model based on the product image; Use the AI image generation model to generate an initial product image based on the text control parameters, the image space control parameters, and the noise image; Generate a restored product image by replacing a corresponding area in the initial product image with the image block; and Generate an advertisement image of the product by performing post-processing on the restored product image to fuse a foreground image and a background image.
Citation Information
Cited By
Multi-modal collaborative delivery content generation method and device
CN122049120A
A multi-modal collaborative delivery content generation method and device
CN122049120B