Image generation method and apparatus, model training method and apparatus, and electronic device
By using a pre-trained text generation model and a text-guided image generation model, background image description information is automatically generated, solving the problem of the manpower-intensive manual design of prompts in existing technologies, and achieving efficient and high-quality product display image generation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2026-04-02
AI Technical Summary
In existing technologies, text-guided image generation models require manual design of accurate prompts to generate high-quality product display images, resulting in high manpower costs and low matching accuracy, making it difficult to efficiently generate display images with backgrounds that match the products in a variety of products.
Background image descriptions are generated through a pre-trained text generation model, and backgrounds are added to product images using a text-guided image generation model, reducing reliance on manually designed prompts. High-quality display images are generated using machine learning models such as neural networks and generative adversarial networks.
It enables the efficient generation of high-quality display images that match the product, saving labor costs and improving generation efficiency and image quality.
Smart Images

Figure CN2025110398_02042026_PF_FP_ABST
Abstract
Description
Image generation method, model training method, device and electronic equipment
[0001] The present disclosure claims priority to Chinese Patent Application No. 202411390309.6, filed on September 30, 2024, with the Chinese Patent Office, entitled "Image generation method, model training method, device and electronic equipment", the entire contents of which are incorporated herein by reference. TECHNICAL FIELD
[0002] The present disclosure relates to the technical field of computer, in particular to an image generation method, a model training method, a device, an electronic equipment and a computer readable storage medium. BACKGROUND
[0003] In an e-commerce platform, visual content plays a key role in attracting and retaining audience attention, a high-quality, well-designed product display image can quickly capture the attention of consumers and increase customer purchase rate. The product display image includes product graphics (such as product real image) and background image, the product image is usually the product real image provided by the merchant, and the background image is usually generated by the platform or system, therefore, generating a matching background for the product has a great impact on the quality of the product display image.
[0004] In related technologies, a text-guided image generation model is usually used to generate a product display image with a matching background for the product, however, the text-guided image generation model relies on carefully designed accurate prompts, therefore, accurate and delicate prompts need to be designed manually for each product, so that the product display image generated by the text-guided image generation model has a higher matching degree and image quality, which is a great challenge in a large number of diversified products, consumes a lot of manpower, and the manually designed prompts may have low accuracy, resulting in low matching degree and quality of the generated product background. SUMMARY
[0005] The present disclosure provides an image generation method, an image generation model training method, a device, an electronic equipment and a computer readable storage medium, which can more efficiently generate a high-quality product display image with a matching background and product, while saving the manpower cost required for background generation. The specific solutions are as follows:
[0006] In a first aspect, the present disclosure provides an image generation method, the method comprising:
[0007] obtaining product description information and product graphics corresponding to a target product for which a display image is to be generated;
[0008] generating, based on the product description information, background image description information corresponding to the product description information by a pre-trained text generation model;
[0009] generate, based on the product pattern, a display image with a background added to the product pattern by a pre-trained text-guided image generation model and by taking the background image description information as guide text of the text-guided image generation model.
[0010] Optionally, the text generation model is configured to generate corresponding description information according to instruction information.
[0011] The generating, based on the product description information, of the background image description information corresponding to the product description information by the pre-trained text generation model includes:
[0012] The generating, based on the product description information, of the background image description information corresponding to the product description information by the pre-trained text generation model includes:
[0013] Optionally, before the generating, based on the product pattern, of the display image with a background added to the product pattern by the pre-trained text-guided image generation model and by taking the background image description information as guide text of the text-guided image generation model, the method further includes:
[0014] The generating, based on the product description information, of the background image description information corresponding to the product description information by the pre-trained text generation model includes:
[0015] The generating, based on the product pattern, of the display image with a background added to the product pattern by the pre-trained text-guided image generation model and by taking the background image description information as guide text of the text-guided image generation model includes:
[0016] The generating, based on the product pattern, of the display image with a background added to the product pattern by the pre-trained text-guided image generation model and by taking the background image description information as guide text of the text-guided image generation model includes:
[0017] Optionally, the product description information includes at least one of the following: product name, product introduction, and product title.
[0018] The generating, based on the product pattern, of the display image with a background added to the product pattern by the pre-trained text-guided image generation model and by taking the background image description information as guide text of the text-guided image generation model includes:
[0019] determine product region annotation information corresponding to the product pattern, the product region annotation information marking a region of the target product in the product pattern;
[0020] generate a display pattern with a background added to the product pattern based on the product pattern and the product region annotation information, by a pre-trained text-guided image generation model, and taking the background picture description information as a guide text of the text-guided image generation model.
[0021] Optionally, the text-guided image generation model is a stable diffusion model.
[0022] The generating a display pattern with a background added to the product pattern based on the product pattern and the product region annotation information, by a pre-trained text-guided image generation model, and taking the background picture description information as a guide text of the text-guided image generation model, comprises:
[0023] obtaining noise to be added;
[0024] generate a display pattern with a background added to the product pattern based on the product pattern, the product region annotation information and the noise to be added, by a pre-trained text-guided image generation model, and taking the background picture description information as a guide text of the text-guided image generation model.
[0025] Optionally, the generating a display pattern with a background added to the product pattern based on the product pattern, the product region annotation information and the noise to be added, by a pre-trained text-guided image generation model, and taking the background picture description information as a guide text of the text-guided image generation model, comprises:
[0026] inputting the product pattern into a pre-trained dimension reduction encoder to obtain reduced dimension pattern encoding data corresponding to the product pattern;
[0027] inputting the pattern encoding data, the background picture description information, the product region annotation information and the noise to be added into a pre-trained text-guided image generation model to generate a display pattern with a background added to the product pattern.
[0028] In a second aspect, the present disclosure further provides a model training method, the method comprising:
[0029] obtaining a training sample, the training sample comprising sample product description information of a sample product and a pre-designed sample product display pattern with a background;
[0030] obtaining output background picture description information corresponding to the sample product description information by a text generation model to be trained based on the sample product description information.
[0031] According to the sample product display image, a sample background feature corresponding to the sample product display image is determined;
[0032] According to a difference between the output background image description information and the sample background feature, a model parameter of the to-be-trained text generation model is adjusted to obtain a trained text generation model, and the text generation model is used to generate corresponding background image description information according to product description information.
[0033] Optionally, the output background image description information corresponding to the sample product description information is obtained by the to-be-trained text generation model based on the sample product description information, and the output background image description information includes:
[0034] The output background image description information corresponding to the sample product description information is obtained by the to-be-trained text generation model based on the sample product description information and the editing background indication information.
[0035] Optionally, before the model parameter of the to-be-trained text generation model is adjusted according to the difference between the output background image description information and the sample background feature, the method further includes:
[0036] The output display image description information corresponding to the sample product description information is obtained by the to-be-trained text generation model based on the sample product description information and the editing display image indication information.
[0037] According to the sample product display image, a sample background feature corresponding to the sample product display image is determined;
[0038] The model parameter of the to-be-trained text generation model is adjusted according to a difference between the output background image description information and the sample background feature, and the model parameter of the to-be-trained text generation model is adjusted according to a difference between the output display image description information and the sample display image feature.
[0039] The model parameter of the to-be-trained text generation model is adjusted according to a difference between the output background image description information and the sample background feature, and the model parameter of the to-be-trained text generation model is adjusted according to a difference between the output display image description information and the sample display image feature.
[0040] Optionally, the sample background feature corresponding to the sample product display image is determined according to the sample product display image, and the sample background feature includes:
[0041] A sample background region image is cropped from the sample display image.
[0042] The sample background region image is input into a pre-trained image encoder to obtain the sample background feature corresponding to the sample product display image.
[0043] The method further comprises:
[0044] The method further comprises:
[0045] Optionally, the training sample further comprises a sample product pattern corresponding to the sample product, the sample product pattern being a product region image cut from the sample product display image.
[0046] The method further comprises:
[0047] The method further comprises:
[0048] According to the difference between the output display image and the sample display image, adjusting the model parameters of the to-be-trained text guided image generation model and the to-be-trained text generation model to obtain a trained text guided image generation model and a trained text generation model.
[0049] Optionally, the to-be-trained text guided image generation model is a stable diffusion model.
[0050] The method further comprises:
[0051] The method further comprises:
[0052] The method further comprises:
[0053] According to the difference between the predicted noise and the preset sample noise, adjusting the model parameters of the to-be-trained text guided image generation model and the to-be-trained text generation model.
[0054] Optionally, based on the sample product image and the preset sample noise, the output background image description information and the output display image description information are obtained as the guide text of the to-be-trained text guided image generation model, and output display images and predicted noises are obtained by training the to-be-trained text guided image generation model.
[0055] Sample product region annotation information corresponding to the sample product image is determined, and the sample product region annotation information marks a region in which the sample product is located in the sample product image.
[0056] The output background image description information, the output display image description information, the sample product image, the sample product region annotation information, and the preset sample noise are input into a to-be-trained diffusion model to obtain output display images and predicted noises.
[0057] Optionally, the method further comprises:
[0058] A candidate training sample set is obtained.
[0059] A screening prompt is obtained, and the screening prompt is used to instruct a large model to screen out training samples that meet a preset screening rule.
[0060] Each candidate training sample in the candidate training sample set and the screening prompt are input into a pre-trained large model to obtain screened training samples.
[0061] The training samples used for model training are determined according to the screened training samples.
[0062] Optionally, the method further comprises:
[0063] The trained text generation model and the text guided image generation model are verified and evaluated by using a verification sample set to obtain an evaluation result.
[0064] The screening prompt is adjusted and updated according to the evaluation result.
[0065] When the expected training condition is not met, the steps of obtaining the training samples, adjusting and updating the model parameters of the to-be-trained text guided image generation model and the to-be-trained text generation model are continuously performed.
[0066] The method further comprises:
[0067] A screening prompt that is updated most recently is obtained.
[0068] The method further comprises:
[0069] The training sample set after deletion is obtained by deleting the training samples that have been excluded by the large model screening from the initial training sample set, and the training sample set after deletion is determined as a candidate training sample set.
[0070] Optionally, the expected training condition comprises at least one of the following:
[0071] The evaluation result meets a preset result.
[0072] The number of training iterations reaches a preset number.
[0073] The change amount of the loss function corresponding to each training round is less than a set threshold.
[0074] In a third aspect, the present disclosure also provides an image generation device, which comprises:
[0075] An acquisition unit is configured to acquire product description information and a product pattern corresponding to a target product to be generated into a display image.
[0076] A first generation unit is configured to generate background image description information corresponding to the product description information by using a pre-trained text generation model based on the product description information.
[0077] A second generation unit is configured to generate a display image with a background added to the product pattern by using a pre-trained text-guided image generation model and taking the background image description information as a guide text of the text-guided image generation model.
[0078] In a fourth aspect, the present disclosure also provides a model training device, which comprises:
[0079] A sample acquisition unit is configured to acquire training samples, wherein the training samples comprise sample product description information of sample products and pre-designed sample product display images with backgrounds.
[0080] A third generation unit is configured to obtain output background image description information corresponding to the sample product description information by using a to-be-trained text generation model based on the sample product description information.
[0081] A determination unit is configured to determine sample background features corresponding to the sample product display images according to the sample product display images.
[0082] A training unit is configured to adjust model parameters of the to-be-trained text generation model according to a difference between the output background image description information and the sample background features, so as to obtain a trained text generation model, wherein the text generation model is used to generate corresponding background image description information according to product description information.
[0083] In a fifth aspect, the present disclosure provides an electronic device, comprising: a processor, a memory, and computer program instructions stored on the memory and executable on the processor; the processor implements the method of any one of the first aspect or the second aspect when executing the computer program instructions.
[0084] In a sixth aspect, the present disclosure provides a computer readable storage medium, the computer readable storage medium storing computer execution instructions, the computer execution instructions being used for implementing the method of any one of the first aspect or the second aspect when executed by a processor.
[0085] In a sixth aspect, the present disclosure provides a computer program product, comprising a computer program, the computer program being used for implementing the method of any one of the first aspect or the second aspect when executed by a processor.
[0086] Compared with the prior art, the present disclosure has the following advantages:
[0087] The image generation method provided by the embodiments of the present disclosure obtains product description information and product patterns corresponding to a target product to be generated into a display image, and then generates background image description information corresponding to the product description information based on the product description information through a pre-trained text generation model, that is, the text generation model can generate corresponding background image description information through the product description information. Since the text generation model is a machine learning model trained by a large number of well-designed sample product display images as training samples, the text generation model can quickly and accurately generate background image description information corresponding to the product description information and meeting the image aesthetic requirements, that is, the text generation model can generate background image description information corresponding to the target product. The background image description information can be used as guide text (i.e., prompt) for the text-guided image generation model. The product pattern can be understood as a product image without a background. Therefore, the pre-trained text-guided image generation model can generate a display image by adding a background to the product pattern based on the product image and using the above background image description information as guide text, and the generated display image is more matched with the product and has higher quality.
[0088] It can be seen that the scheme provided by the disclosure does not need to manually design the prompt of the text-guided image generation model, but can efficiently generate higher-quality background image description information corresponding to the product description information through the pre-trained text generation model, that is, through the text generation model, the hidden features between the product description information and the image background thereof can be aligned, so that high-quality background image description information strongly related to the product is quickly and efficiently obtained as the prompt of the text-guided image generation model, so that the text-guided image generation model can generate a display image with a higher quality and a more matched product, thereby more efficiently generating a high-quality display image with a matched background and product, while saving the human cost required for background generation.
[0089] The model training method provided by the disclosure obtains a training sample, the training sample includes sample product description information of a sample product and a pre-designed sample product display image with a background, and based on the sample product description information and through a to-be-trained text generation model, output background image description information corresponding to the sample product description information is obtained, and then sample background features corresponding to the sample product display image are determined, and the model parameters of the to-be-trained text generation model are adjusted according to the difference between the output background image description information and the sample background features, so as to obtain a trained text generation model. The disclosure trains the to-be-trained text generation model by comparing the output background image description information and the sample background features, which can enhance the understanding of the text generation model for background elements. Through the model training method provided by the disclosure, a text generation model for generating corresponding background image description information according to product description information can be trained, and the trained text generation model can efficiently and accurately generate high-quality background image description information strongly related to the product, so that the background image description information can be used as the prompt of the text-guided image generation model to efficiently generate a high-quality display image with a matched background and product. BRIEF DESCRIPTION OF DRAWINGS
[0090] FIG. 1 is a schematic diagram of an application scenario of an image generation scheme provided by the disclosure;
[0091] FIG. 2 is a flowchart of an example of an image generation method provided by an embodiment of the disclosure;
[0092] FIG. 3 is a flowchart of another example of an image generation method provided by an embodiment of the disclosure;
[0093] FIG. 4 is a flowchart of an example of a model training method provided by an embodiment of the disclosure;
[0094] FIG. 5 is a flowchart of another example of a model training method provided by an embodiment of the disclosure;
[0095] FIG. 6 is a flowchart of an example of model training through data screening provided by an embodiment of the disclosure;
[0096] FIG. 7 is a structural block diagram of an example of an image generation apparatus according to an embodiment of the present disclosure;
[0097] FIG. 8 is a structural block diagram of an electronic device according to the present disclosure. DETAILED DESCRIPTION
[0098] For those skilled in the art to better understand the technical solutions of the present disclosure, the present disclosure will be described clearly and completely below in conjunction with the drawings in the embodiments of the present disclosure. However, the present disclosure can be implemented in many other ways different from the description below, therefore, based on the embodiments provided by the present disclosure, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present disclosure.
[0099] It should be noted that the terms "first", "source domain", "third" and the like in the claims, description and drawings of the present disclosure are used to distinguish similar objects and do not describe a specific order or sequence. The data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include", "have" and their variants are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0100] In order to facilitate understanding of the embodiments of the present disclosure, the application background of the embodiments is described.
[0101] In the e-commerce platform, visual content plays a key role in attracting and retaining audience attention, a high-quality, well-designed product display image can quickly capture the attention of consumers and increase customer purchase rate. The product display image contains product graphics (such as product real image) and background image, product image is usually the product real image provided by the merchant, and the background image is usually generated by the platform or system, therefore, generating a matching background for the product has a great impact on the quality of the product display image.
[0102] In the related art, a text-guided image generation model is usually used to generate product display images matching the background of products. However, the text-guided image generation model relies on carefully designed accurate prompts, and therefore, accurate and fine prompts need to be manually designed for each product so that the background of the product display image generated by the text-guided image generation model is more matched with the product and the image quality is higher. This is a great challenge in a large number of diversified products, consumes a lot of manpower, and the accuracy of the manually designed prompts may be low, resulting in low matching degree and quality of the generated product background.
[0103] To solve the above problems, the embodiment of the present disclosure provides an image generation method, a model training method, an apparatus, an electronic device and a computer readable storage medium. The purpose is to more efficiently generate high-quality product display images with backgrounds matched with products, while saving the manpower cost required for background generation.
[0104] The image generation method provided by the present disclosure can be used to generate product display images, and specifically used to quickly generate corresponding product display images for products based on obtained product patterns and product description information. The scheme provided by the present disclosure can be used for product display image generation in e-commerce platforms, and can also be used for product display image generation in other scenarios, such as product display image generation in advertising platforms and website design, and the present disclosure does not specifically limit the specific application scenarios of product display image generation.
[0105] In order to facilitate understanding of the method embodiment of the present disclosure, the application scenario thereof is introduced. Please refer to FIG. 1, which is a schematic diagram of the application scenario of the scheme provided by the embodiment of the present disclosure. For ease of description, the embodiment of the present disclosure is described below by taking an e-commerce platform as an example, and this application scenario is a schematic example and does not constitute a specific description of the application scenario. As shown in FIG. 1, the application scenario is provided with a server 102, a merchant terminal 103 and a client terminal 101. In this embodiment, the client terminal 101 and the server 102, and the merchant terminal 103 and the server 102 establish a connection through network communication.
[0106] The merchant terminal 103 can be a mobile phone, a tablet computer (pad), a smart watch, a desktop computer, a smart television, a VR device, a vehicle-mounted device, a wearable device, a notebook computer, etc. having a display function and a data processing function. The merchant terminal 103 is used to upload product patterns to the server 102 and obtain and display the generated product display images from the server 102. The merchant terminal 103 can also be used to upload product information, upload order processing information, upload promotion information, obtain order information from the server 102, etc., but is not limited thereto.
[0107] A specific communication connection needs to be established between the merchant end 103 and the service end 102, so as to perform data transmission.
[0108] The client end 101 can be a mobile phone, a pad, a smart watch, a desktop computer, a smart television, a VR device, a vehicle-mounted device, a wearable device, a notebook computer, or the like electronic device with a display function and a data processing function. The client end 101 is used to acquire and display a product display page of an e-commerce platform from the service end 102, and the product display page contains a product display image generated by the service end 102. The client end 101 can also be used to send order information to the service end 102, send browsing information, add purchase information, acquire and display promotion information of a merchant from the service end 102, and the like, but is not limited thereto.
[0109] A specific communication connection needs to be established between the client end 101 and the service end 102, so as to perform data transmission.
[0110] The service end 102 has high computing capability. The service end 102 can be a server, and the service end 102 has high CPU computing capability, long-time reliable operation, strong I / O external data throughput capability, and better scalability. The service end 102 can be a single server or a server cluster. The service end 102 is used to receive a product image from the merchant end 103, and generate a corresponding product display image according to the product image. The service end 102 can also provide other specific services for the client end 101 and the merchant end 103, such as user information access, website access, application program access, and the like, and the present disclosure is not specifically limited.
[0111] The client 101 and the server 102, and the merchant terminal 103 and the server 102 can communicate by using various communication systems, for example, can be by using a wired communication system or a wireless communication system. The wireless communication system can be, for example, a global system for mobile communications (GSM) system, a code division multiple access (CDMA) system, a wideband code division multiple access (WCDMA) system, a general packet radio service (GPRS), a long term evolution (LTE) system, an LTE frequency division duplex (FDD) system, an LTE time division duplex (TDD), a universal mobile telecommunication system (UMTS), a worldwide interoperability for microwave access (WiMAX) communication system, a future 5th generation (5G) system or new radio (NR), a satellite communication system, and the like.
[0112] Embodiment one
[0113] The first embodiment of the present disclosure provides an image generation method, which is applied to an electronic device, which can be a server, a notebook computer, a tablet computer, a desktop computer, or the like, and has a data processing function.
[0114] The electronic device can be deployed with a pre-trained text generation model and a text-guided image generation model, or can not be deployed with a pre-trained text generation model and a text-guided image generation model. In this case, the electronic device can call a pre-trained text generation model and a text-guided image generation model deployed on other devices to generate images through the models deployed on the other devices.
[0115] As shown in FIG. 2 and FIG. 3, the image generation method provided by the first embodiment of the present disclosure includes the following steps S110-S130.
[0116] Step S110: Obtain product description information and product drawings corresponding to a target product to be generated for a display image.
[0117] The target product can be a product sold on an e-commerce platform, such as clothing, food, dishes, or meals; it can also be a product sold on other platforms, such as virtual equipment or virtual characters in the gaming field; or it can be any non-commodity product that requires a display image.
[0118] The aforementioned product description information can include product name, product introduction, product details, product title, and other information used to describe the product. In the e-commerce field, merchants typically set product description information such as titles, names, and introductions for their products, which electronic devices can directly access. Product names or titles accurately express the product through concise and simple text, making them easy to set. Therefore, many merchants accurately set product names or titles. Thus, in this embodiment, the product description information can be a product name or title, allowing electronic devices to easily and conveniently obtain product description information for each product. Furthermore, the obtained product description information is usually relatively simple, making it easier to obtain product display images applicable to all products. The product description information can be in Chinese, English, German, or other languages; this disclosure does not specifically limit it. For example, as shown in Figure 3, the product description information is the English product name "XX UV Shield Essential Sunscreen Gel SPF 35 PA++".
[0119] The aforementioned product illustration refers to an image containing a photograph of the actual product. The background of the product illustration can be a preset background, such as a background of a preset color or a transparent background; that is, the background of the product illustration is an initial background that has not yet been designed. The photograph included in the product illustration can be an actual product image. If the product is a virtual product, such as virtual equipment or a virtual character, then the photograph included in the product illustration is an image of the virtual product's appearance. The product illustration can contain product images to be generated for display. As shown in Figure 3, the product illustration can be an image containing a photograph of the actual product with a black background.
[0120] In one implementation, step S110 involves obtaining the product image by: acquiring the product image uploaded by the merchant; extracting the product area from the product image; and adding a preset color background to the product area to obtain the product image. Since the product images uploaded by merchants may vary, and some product images may have cluttered backgrounds that hinder the generation of subsequent display images, this implementation, by extracting the product area and adding a preset color background, results in a clearer and more distinguishable boundary between the product area and the background in the obtained product image. This facilitates the subsequent generation of the background, thereby generating a display image with a corresponding background.
[0121] Step S120: based on the product description information, generate the background picture description information corresponding to the product description information through the pre-trained text generation model.
[0122] The text generation model is used to generate the corresponding background picture description information through the product description information. Specifically, the product description information can be input into the pre-trained text generation model to obtain the background picture description information corresponding to the product description information. Alternatively, the product description information can be processed to obtain processed product description information, and then the processed product description information is input into the pre-trained text generation model to obtain the background picture description information corresponding to the product description information. The text processing can be at least one of text expansion, text deduplication, and key information extraction, but is not limited thereto, so that the text description information input into the text generation model is more accurate.
[0123] In one embodiment, the text generation model can generate corresponding description information according to the indication information, which is used to indicate what kind of description information the text generation model generates. In this case, the product description information is the indication information of the text generation model, so that the text generation model can generate description text consistent with the product description information. In order to enable the text generation model to generate description information about the background, the indication information can also include editing background indication information to indicate the text generation model to generate relevant description about the background. In this case, step S120 can be implemented as follows.
[0124] Step S121: through the pre-trained text generation model, taking the product description information and the editing background indication information as the indication information of the text generation model, instructing the text generation model to generate the background picture description information corresponding to the product description information.
[0125] Specifically, the editing background indication information can be a prompt word for editing the background, such as "generate background text", "background" as shown in FIG. 3, or other prompt words that can indicate the editing background. In step S121, the product description information and the editing background indication information can be input into the text generation model to generate the background picture description information corresponding to the product description information; or the product description information and the editing background indication information can be processed, such as splicing processing, encoding processing, etc., to obtain processed information, and then the processed information is input into the text generation model to obtain the background picture description information corresponding to the product description information.
[0126] The embodiment sets the editing background indication information, so that the text generation model can accurately and efficiently generate the background picture description information corresponding to the to-be-generated display product under the indication of the editing background indication information.
[0127] The text generation model is pre-trained and generated, and the specific training process will be described in detail in the subsequent embodiments. The pre-trained text generation model is a machine learning model, which can be trained to be a model that generates accurate background picture description information. The background picture description information is used to accurately describe the background picture of the product in the form of text to guide the text-guided image generation model to generate a display picture with a corresponding background. The text generation model can be a neural network model, a generative adversarial network model, a decision tree model, etc., or other machine learning models.
[0128] Step S130: based on the product pattern, a pre-trained text-guided image generation model is used to generate a display picture with a background added to the product pattern, and the background picture description information is used as the guide text of the text-guided image generation model.
[0129] The display picture with the background added to the product pattern is the generated display picture of the target product.
[0130] The text-guided image generation model can be a Stable Diffusion model, a Midjourney model, a DALL·E model, etc., or other machine learning models that can generate images corresponding to the images described by the text. The pre-trained text-guided image generation model can accurately generate corresponding images according to the text description information, and specifically can generate a display picture with a corresponding background added to the product pattern according to the description of the background picture description information.
[0131] The text-guided image generation model is pre-trained and generated, and the specific training process will be described in detail in the subsequent embodiments. The pre-trained text-guided image generation model is a machine learning model, which can be trained to be a model that generates accurate corresponding images according to the text description. That is, the text-guided image generation model can generate a display picture with a background added to the product pattern under the guidance of the background picture description information, and the background of the generated display picture conforms to the background described in the background picture description information.
[0132] The step S130 can specifically input the product pattern and the background picture description information into the pre-trained text guided image generation model, so that the text guided image generation model generates a display picture with a background added to the product pattern according to the background picture description information as a guided text. Alternatively, the product pattern, the background picture description information and other information for generating the display picture can be input into the text guided image generation model to generate the display picture with a background added to the product pattern. Alternatively, the product pattern or the background picture description information can be processed to obtain processed information, such as encoding processing, text extraction processing, etc., and the processed information is input into the text guided image generation model to generate the display picture with a background added to the product pattern. The specific manner in which the text guided image generation model generates the display picture is not specifically limited by the present disclosure.
[0133] In one specific embodiment, the step S130 can further include the following step S130a before the step S130.
[0134] The step S130a: generating display picture description information corresponding to the product description information by a pre-trained text generation model, with the product description information and the editing display picture indication information as indication information of the text generation model.
[0135] The display picture description information is the description information of the display picture corresponding to the target product as a whole, and the display picture of the product can be described as a whole through the display picture description information.
[0136] The editing display picture indication information can be a prompt word for editing the display picture, such as “generate display picture text”, “eos” as shown in FIG. 3, etc., or other prompt words that can indicate editing of the display picture.
[0137] The step S130a can be executed simultaneously with the step S121, in which case the step S130a and the step S121 can be combined into the following step: generating display picture description information and background picture description information corresponding to the product description information by a pre-trained text generation model, with the product description information, the editing background indication information and the editing display picture indication information as indication information of the text generation model. Specifically, the product description information, the editing background indication information and the editing display picture indication information can be input into the text generation model to generate the display picture description information and the background picture description information, for example, as shown in FIG. 3, the product description information P, the editing background indication information B and the editing display picture indication information D can be input into the text generation model to generate the display picture description information and the background picture description information. <background> <eos>The product name P and the background editing instruction information background are input into the text-to-image generation model to generate the display image description information and the background image description information. In this way, the display image description information and the background image description information can be generated at the same time through one model call, and the efficiency of image generation is improved.
[0138] The display image description information is used to guide the text-to-image generation model to generate the display image description information corresponding to the product description information.
[0139] Correspondingly, the step S130 can be implemented in the following step S131.
[0140] In the step S131, the display image with the background added is generated based on the product image, the background image description information and the display image description information by using the pre-trained text-to-image generation model.
[0141] Specifically, the product image, the background image description information and the display image description information are input into the text-to-image generation model to generate the display image with the background added.
[0142] In the embodiment, the text-to-image generation model is added with the display image description information as the guide text. Since the display image description information can describe the image as a whole, the display image generated by the text-to-image generation model has a better image effect as a whole.
[0143] In an embodiment, the step S130 can be implemented in the following steps S132 and S133.
[0144] In the step S132, the product region annotation information corresponding to the product image is determined, and the product region annotation information marks the region of the target product in the product image.
[0145] The product region annotation information can be an image mask corresponding to the product image. The image mask can specify the product region in the product image that needs to be paid attention to or processed. The image mask can isolate or highlight the product region in the product image. Specifically, the image mask can be a binary image with the same size as the product image, in which the white or non-zero region in the mask represents the product region part that needs to be paid attention to, and the black or zero region represents the part that can be ignored. The product region annotation information can clearly mark the region of the product. The product region annotation information can also be a contour line of the product region, or other information that can mark the product region. The present disclosure is not specifically limited.
[0146] In this step, the electronic device can use image annotation tools such as Labelme, VGG Image Annotator (VIA), COCO Annotator, etc. to determine the corresponding image mask or product contour line and other product region annotation information from the product drawing, and the specific determination method of the product region annotation information is not limited in the present disclosure.
[0147] Step S133: Based on the product drawing and the product region annotation information, a display image with a background added to the product drawing is generated by a pre-trained text-guided image generation model, and the background image description information is used as the guide text of the text-guided image generation model.
[0148] Specifically, the product drawing, the product region annotation information, and the background image description information can be input into the text-guided image generation model, so that the text-guided image generation model generates a display image with a background added to the product drawing based on the background image description information.
[0149] The present embodiment can make the text-guided image generation model accurately distinguish between the background region and the product region by inputting the product region annotation information into the text-guided image generation model to assist the text-guided image generation model in generating the display image, so that the background can be more accurately and appropriately fused with the product region, and the quality of the finally generated display image is higher.
[0150] In one specific embodiment, the text-guided image generation model can be a Stable Diffusion model. The Stable Diffusion model is a generative model that can generate the required image through the processes of forward diffusion and reverse diffusion. Forward diffusion refers to gradually adding noise to the original image until it becomes pure noise, and reverse diffusion starts from pure noise and gradually removes noise through a learned model to finally generate a clear required image. The Stable Diffusion model generates the required image by adding noise, which can be referred to related technologies, and will not be described in detail in the present disclosure.
[0151] Correspondingly, the step S133 can be implemented according to the following steps S133a-S133b.
[0152] Step S133a: obtaining noise to be added.
[0153] Specifically, the preset noise can be determined as the noise to be added, or the noise to be added can be randomly generated based on a Gaussian distribution.
[0154] Step S133b: based on the product pattern, the product region annotation information, and the noise to be added, generating a display image with a background added to the product pattern by a pre-trained text-guided image generation model, and taking the background image description information as the guide text of the text-guided image generation model.
[0155] In this step, the product pattern, the product region annotation information, the noise to be added, and the background image description information can be input into the trained Stable Diffusion model to generate a display image with a background added to the product pattern. Other guide texts and input information can also be added to make the text-guided image generation model generate a display image. In the process of generating a display image, the Stable Diffusion model gradually adds noise to the product pattern by using the noise to be added through a forward diffusion process, and then removes the noise from the image through a reverse diffusion process based on the product region annotation information and the background image description information to obtain a denoised display image that meets the requirements.
[0156] The image generation process of the Stable Diffusion model is relatively stable, and can generate high-quality images with strong controllability. By using the Stable Diffusion model to generate a display image, the quality of the generated display image is more stable and reliable.
[0157] In one specific embodiment, as shown in FIG. 3, step S133b can be implemented in the following steps A-B.
[0158] Step A: inputting the product pattern into a pre-trained dimension reduction encoder to obtain reduced dimension pattern encoding data corresponding to the product pattern.
[0159] The dimension reduction encoder can be any one of a variational autoencoder (VAE), a principal component analysis (PCA) model, and a t-distributed stochastic neighbor embedding (t-SNE) model, but is not limited thereto.
[0160] The dimension reduction encoder can be trained by supervised training, unsupervised training, semi-supervised training, or other training methods, and the specific training method of the dimension reduction encoder is not specifically limited in the present disclosure.
[0161] Step B: input the above pattern coding data, the above background picture description information, the above product region annotation information and the above to-be-added noise into the pre-trained text guided image generation model to generate a display picture with a background added to the product pattern.
[0162] Specifically, as shown in FIG. 3, the pattern coding data output by the dimension reduction encoder, the product region annotation information (for example, the mask) and the to-be-added noise can be spliced to obtain spliced features, and then the spliced features and the background picture description information generated by the text generation model are input into the text guided image generation model to generate a display picture with a background added to the product pattern through multiple noise adding and denoising processes.
[0163] After the product pattern is processed by dimension reduction in this embodiment, the dimension of the image data can be reduced, thereby reducing the calculation and storage requirements, making the image generation process more efficient. In addition, dimension reduction can also remove redundant information, further improving the efficiency and accuracy of image generation.
[0164] The image generation method provided in the embodiments of the present disclosure obtains product description information and a product pattern corresponding to a target product to be generated into a display picture, and then generates background picture description information corresponding to the product description information based on the product description information through a pre-trained text generation model. That is, the text generation model can generate corresponding background picture description information through product description information. Since the text generation model is a machine learning model trained based on a large number of designed product display pictures as training samples, the text generation model can quickly and accurately generate background picture description information corresponding to the product description information and meeting the requirements of image aesthetics, that is, the background picture description information corresponding to the target product can be generated. The background picture description information can be used as the guide text (i.e., the prompt) of the text guided image generation model. The product pattern can be understood as a product pattern without a background. In this way, the pre-trained text guided image generation model can generate a display picture with a background added to the product pattern based on the product pattern and the background picture description information as the guide text, and the generated display picture is more matched with the product and has higher quality.
[0165] It can be seen that the scheme provided by the present disclosure does not need to manually design the prompt of the text-guided image generation model, but can efficiently generate higher-quality background image description information corresponding to the product description information through the pre-trained text generation model, that is, through the text generation model, the hidden features between the product description information and the image background thereof can be aligned, so that high-quality background image description information strongly related to the product is quickly and efficiently obtained as the prompt of the text-guided image generation model, so that the text-guided image generation model can generate a display image with a higher quality and a more matched image, thereby more efficiently generating a high-quality display image with a matched background and product, while saving the human cost required for background generation.
[0166] Embodiment Two
[0167] The second embodiment of the present disclosure also provides a model training method, which is applied to an electronic device, which can be a desktop computer, a notebook computer, a server, a mobile terminal, a gateway, or other electronic devices with data processing capabilities. The present embodiment is at least used to train the text generation model required for image generation in the first embodiment. Other models in the first embodiment (such as the text-guided image generation model and the dimension reduction encoder) can be trained during the training of the text generation model, or can be pre-trained. The model training method of the present embodiment also synchronously trains the training process of the text-guided image generation model and the dimension reduction encoder.
[0168] As shown in FIGS. 4 and 5, the model training method provided by the present embodiment includes the following steps S210-S240.
[0169] Step S210: Obtain training samples.
[0170] The training samples include sample product description information of sample products and pre-designed sample product display images with backgrounds. The sample product display images in the training samples are display images of products with pre-designed backgrounds by designers, and each sample product display image is pre-provided with corresponding sample product description information.
[0171] The sample product description information can include at least one of a sample product name, a sample product introduction, and a sample product title, and can also include other information capable of describing a product.
[0172] Step S220: Based on the sample product description information, obtain output background image description information corresponding to the sample product description information through the text generation model to be trained.
[0173] The to-be-trained text generation model can be a neural network model, a generative adversarial network model, a decision tree model, or other to-be-trained machine learning model. The initial parameters of the to-be-trained text generation model can be randomly generated, or a model that has been trained and used can be used as the to-be-trained text generation model in the present disclosure, so that the parameters of the model that has been trained and used are trained and adjusted, or the initial parameters of the to-be-trained text generation model can be determined in other manners, which are not specifically limited in the present disclosure.
[0174] In step S220, the sample product description information can be input into the to-be-trained text generation model to obtain output background image description information corresponding to the sample product description information. The sample product description information can also be input into the to-be-trained text generation model after being processed by encoding or other processing to obtain the output background image description information corresponding to the sample product description information.
[0175] In an embodiment, step S220 can be implemented in the following step S221.
[0176] Step S221: Based on the sample product description information and the editing background indication information, the to-be-trained text generation model is used to obtain output background image description information corresponding to the sample product description information.
[0177] The editing background indication information in the present embodiment is similar to the editing background indication information in the first embodiment, which will not be described in detail here.
[0178] Specifically, the sample product description information and the editing background indication information can be input into the to-be-trained text generation model to obtain the output background image description information corresponding to the sample product description information. The sample product description information and the editing background indication information can also be processed, for example, spliced, encoded, or processed in other manners to obtain sample processed information, and the sample processed information is input into the to-be-trained text generation model to obtain the output background image description information corresponding to the sample product description information.
[0179] The present embodiment can make the to-be-trained text generation model generate background image description information corresponding to the to-be-generated display product under the indication of the editing background indication information by setting the editing background indication information. In this way, the trained text generation model can accurately and efficiently generate background image description information under the indication of the editing background indication information, and the trained text generation model is not limited to generating background image description information, but can also be used to generate other text information, thereby improving the applicability of the text generation model.
[0180] Step S230: Determine the sample background feature corresponding to the sample product display image according to the sample product display image.
[0181] In one specific embodiment, the sample product display image corresponding sample background feature can be determined according to the following steps S231-S232.
[0182] Step S231: The sample background region image is cropped from the sample display image.
[0183] The sample background region can be cropped from the sample display image by using an image annotation tool such as Labelme, VGG Image Annotator (VIA), COCO Annotator, etc.
[0184] Step S232: The sample background region image is input into a pre-trained image encoder to obtain the sample product display image corresponding sample background feature.
[0185] The pre-trained image encoder is used to encode images to obtain image features. The image encoder can be trained by supervised training, unsupervised training, semi-supervised training or other training methods, and the specific training method of the image encoder is not limited in the present disclosure.
[0186] In this embodiment, the sample background region image is first cropped, so that the corresponding sample background feature can be quickly encoded from the sample background region image by using the image encoder. This method can quickly and accurately obtain the corresponding sample background feature by using the image encoder, and the image encoder does not need to have a background recognition function, but only needs to have a feature encoding function, so that the training of the image encoder is simpler.
[0187] Alternatively, the corresponding sample background feature can also be determined from the sample display image by using a pre-trained background feature extraction model, which can directly extract the background feature from the image. In this embodiment, the sample background feature can be quickly and conveniently extracted by using the model.
[0188] Step S240: According to the difference between the output background image description information and the sample background feature, the model parameters of the to-be-trained text generation model are adjusted to obtain a trained text generation model.
[0189] The text generation model is used to generate the corresponding background image description information according to the product description information.
[0190] Step S240 is to train the text generation model by using a contrast learning method.
[0191] Optionally, the model parameters of the to-be-trained text generation model can be adjusted based on the principle that the difference between the output background graph description information and the sample background feature is less than a second preset threshold, and the to-be-trained text generation model is trained through a large number of training samples until the difference between the output background graph description information and the sample background feature is less than the second preset threshold.
[0192] Alternatively, the model parameters of the to-be-trained text generation model can also be adjusted based on the principle of minimizing the difference between the output background graph description information and the sample background feature, and the to-be-trained text generation model is trained through a large number of training samples until the difference between the output background graph description information and the sample background feature converges in multiple iterations, that is, the difference between the background graph description information corresponding to multiple iterations and the sample background feature is the same or less than a preset difference.
[0193] In one specific embodiment, when the model parameters of the to-be-trained text generation model are adjusted based on the principle of minimizing the difference between the output background graph description information and the sample background feature, the parameter adjustment can be performed through the loss function of formula (1) as follows.
[0194] In formula (1), L cl (b,B) represents the target loss (i.e., the difference between the output background graph description information and the sample background feature), represents the similarity between the current sample b and its positive sample normalized by the temperature parameter τ, represents the sum of the similarities between the current sample b and all other samples (including positive samples and negative samples) normalized by the temperature parameter τ.
[0195] The loss function of formula (1) can minimize the distance between the current sample and its positive sample, and maximize the distance between the current sample and other negative samples. In this way, the model can learn a representation that can distinguish between positive samples and negative samples. During the training process, by minimizing this loss function, the model adjusts its parameters so that the similarity score of the positive sample is high, and the similarity score of the negative sample is low, and a text generation model with adjusted parameters is obtained.
[0196] When other training processes are used, the loss function is also different. Those skilled in the art can set the loss function according to the actual training process, and the present disclosure does not specifically limit it.
[0197] In one specific embodiment, before step S240, the following steps S230a and S230b can also be included.
[0198] Step S230a: based on the sample product description information and the editing display image indication information, the output display image description information corresponding to the sample product description information is obtained through the text generation model to be trained.
[0199] The editing display image indication information in this embodiment is similar to that in the first embodiment, and will not be described in detail here.
[0200] Specifically, the sample product description information and the editing display image indication information can be input into the text generation model to be trained to obtain the output display image description information corresponding to the sample product description information. The sample product description information and the editing display image indication information can also be processed, such as splicing processing, encoding processing, etc., to obtain sample processed information, and the sample processed information is input into the text generation model to be trained to obtain the output display image description information corresponding to the sample product description information.
[0201] Step S230a can be executed simultaneously with step S221, in which case step S230a and step S221 can be combined into the following step: based on the sample product description information, the editing display image indication information, and the editing background image indication information, the output display image description information and the output background image description information corresponding to the sample product description information are obtained through the text generation model to be trained. Specifically, the sample product description information, the editing display image indication information, and the editing background image indication information can be input into the text generation model to be trained to obtain the output display image description information and the output background image description information corresponding to the sample product description information. For example, as shown in FIG. 5, the sample product description information P, the editing display image indication information D, and the editing background image indication information B can be input into the text generation model to be trained to obtain the output display image description information and the output background image description information corresponding to the sample product description information P. <eos>, edit background picture indication information <background>inputting a text to be trained into the model to generate an output display image description information <eos>pooling and outputting context map description information <background>pooling.
[0202] In this way, the output display image description information and the output background image description information can be generated simultaneously through one model call, the efficiency of model training is improved, and the efficiency of data output of the trained text generation model is also improved.
[0203] The output display image description information is used to guide the to-be-trained text generation model to generate description information of a sample product display image corresponding to sample product description information.
[0204] Step S230b: determining sample display image features corresponding to the sample product display image according to the sample product display image.
[0205] Specifically, as shown in FIG. 5, the sample product display image can be input into the image encoder to obtain sample display image features corresponding to the sample product display image.
[0206] Correspondingly, step S240 can adjust the model parameters of the to-be-trained text generation model according to the following step S241.
[0207] Step S241: adjusting the model parameters of the to-be-trained text generation model according to the difference between the output background image description information and the sample background features and the difference between the output display image description information and the sample display image features.
[0208] The trained text generation model is specifically used to generate corresponding background image description information and product display image description information according to product description information.
[0209] The process of adjusting the model parameters of the to-be-trained text generation model according to the difference between the output display image description information and the sample display image features in step S241 is similar to the process of adjusting the parameters according to the difference between the output background image description information and the sample background features in step S240, which will not be described in detail here.
[0210] In step S241, the process of parameter adjustment can comprehensively refer to the difference between the output background image description information and the sample background features and the difference between the output display image description information and the sample display image features to quickly train a text generation model capable of simultaneously outputting background image description information and product display image description information.
[0211] The text generation model trained in this embodiment can generate corresponding background image description information and product display image description information according to the product description information, that is, it can generate a more comprehensive type of description information for the product display image, so as to generate a more comprehensive prompt for generating the product display image, and therefore, the subsequent text-to-image generation model can generate a product display image with higher quality based on more comprehensive information.
[0212] In one specific embodiment, the training sample described above can further include a sample product image corresponding to the sample product, which is a product region image extracted from the sample product display image. The sample product image can also be understood as an image obtained by removing the background region in the sample product display image.
[0213] Correspondingly, the method described above can further include steps S250-S260.
[0214] Step S250: based on the sample product image, the output display image is obtained by the text-to-image generation model to be trained, with the output background image description information and the output display image description information as the guide text of the text-to-image generation model to be trained.
[0215] Specifically, the sample product image, the output background image description information and the output display image description information can be input into the text-to-image generation model to be trained, so that the text-to-image generation model to be trained generates an output display image with the output background image description information and the output display image description information as guide text. Alternatively, the sample product image, the sample background image description information, the sample display image description information and other information used to generate the sample product display image can also be input into the text-to-image generation model to be trained to generate the sample product display image.
[0216] The text-to-image generation model to be trained can be a Stable Diffusion model, a Midjourney model, a DALL·E model, etc., or other machine learning models that can generate images corresponding to the images described by the text through text guidance.
[0217] The initial parameters of the text-to-image generation model to be trained can be randomly generated, or a model that has been trained and used can be used as the text-to-image generation model to be trained in the present disclosure, so as to train and adjust the parameters of the text-to-image generation model that has been trained and used, or the initial parameters of the text-to-image generation model to be trained can be determined in other ways, which are not specifically limited in the present disclosure.
[0218] Step S260: According to the difference between the output display image and the sample display image, the model parameters of the to-be-trained text guided image generation model and the to-be-trained text generation model are adjusted to obtain a trained text guided image generation model and a trained text generation model.
[0219] Specifically, the model parameters of the to-be-trained text generation model and the to-be-trained text guided image generation model can be adjusted based on the principle that the difference between the output display image and the sample display image is less than a third preset threshold. Alternatively, the model parameters of the to-be-trained text generation model and the to-be-trained text guided image generation model can also be adjusted based on the principle of minimizing the difference between the output display image and the sample display image.
[0220] In the embodiments of the present disclosure, both step S260 and step S240 adjust the model parameters of the to-be-trained text generation model, so that the generated text generation model can more accurately generate corresponding background image description information and product display image description information, thereby better guiding the subsequent generation of product display images with better quality.
[0221] The present embodiment can train a text guided image generation model while training a text generation model, thereby improving the training efficiency of multiple models. The trained text guided image generation model can be directly used subsequently, and the text generation model can be further adjusted in parameters based on the output display image output by the to-be-trained text guided image generation model, so that the prediction accuracy of the trained text generation model is higher.
[0222] In one specific embodiment, the to-be-trained text guided image generation model can be a stable diffusion model, and accordingly, step S250 can be implemented as follows.
[0223] Step S251: Based on the sample product image and the preset sample noise, the to-be-trained text guided image generation model is used to obtain an output display image and a predicted noise, with the output background image description information and the output display image description information as the guided text of the to-be-trained text guided image generation model.
[0224] Specifically, the sample product image, the preset sample noise, the output background image description information and the output display image description information can be input into the to-be-trained text guided image generation model to obtain the output display image and the predicted noise. Alternatively, at least one of the sample product image, the preset sample noise, the output background image description information and the output display image description information can be encoded, spliced or processed in other manners before being input into the to-be-trained text guided image generation model to obtain the output display image and the predicted noise.
[0225] Since the Stable Diffusion model can gradually add noise to the original image and then gradually remove the noise, finally generating a clear required image, and the predicted noise can be obtained, the Stable Diffusion model can obtain the output display image and the predicted noise.
[0226] Correspondingly, the model training method can further include the following step S270:
[0227] Step S270: adjusting the model parameters of the text-guided image generation model to be trained and the text generation model to be trained according to the difference between the predicted noise and the preset sample noise.
[0228] Specifically, the model parameters of the text generation model to be trained and the text-guided image generation model to be trained can be adjusted based on the principle that the difference between the predicted noise and the preset sample noise is less than a fourth preset threshold. Alternatively, the model parameters of the text generation model to be trained and the text-guided image generation model to be trained can also be adjusted based on the principle of minimizing the difference between the predicted noise and the preset sample noise.
[0229] In this embodiment, the Stable Diffusion model is used as the text-guided image generation model, and the accuracy of the preset noise and the output product display image is considered during the training process, so that a more accurate text-guided image generation model can be obtained.
[0230] In one specific embodiment, as shown in FIG. 5, step S251 can be implemented in the following steps S251a-S251b.
[0231] Step S251a: determining sample product region annotation information corresponding to the sample product image.
[0232] The sample product region annotation information marks the region where the sample product is located in the sample product image. The related content of the sample product region annotation information is similar to the product region annotation information in the first embodiment, which will not be described in detail here.
[0233] Step S251b: inputting the output background image description information, the output display image description information, the sample product image, the sample product region annotation information, and the preset sample noise into the diffusion model to be trained to obtain the output display image and the predicted noise.
[0234] This embodiment can enable the trained text-guided image generation model to more accurately distinguish between the background region and the product region under the auxiliary guidance of the product region annotation information, so as to more accurately and appropriately fuse the background with the product region, thereby improving the quality of the finally generated display image.
[0235] In one specific embodiment, when adjusting the model parameters of the to-be-trained text generation model and the to-be-trained text guided image generation model based on the principle of minimizing the difference between the predicted noise and the preset sample noise, the difference between the output display image and the sample display image, the parameter adjustment can be performed through the loss function of formula (2) as follows.
[0236] In formula (2), L inpaint represents the target loss (i.e., the comprehensive value of the two differences between the predicted noise and the preset sample noise, and the difference between the output display image and the sample display image), represents the expectation of each condition (including product sample c, product region annotation information m, display image I, product description information p, background image B, time step T, noise level ε t )ε t represents the preset sample noise at the current time step t, ε θ represents the predicted noise, x' represents the sample display image, Γ θ (P,b) represents the output display image. Formula (2) is a commonly used loss function in the training process of the stable diffusion model, and the specific interpretation will not be described in detail in the present disclosure, and other loss functions can also be used for model training by those skilled in the art.
[0237] In one specific embodiment, the embodiments of the present disclosure can adjust the parameters of each model through the following formula (3) comprehensive loss function when adjusting the parameters of each model. L=L inpaint +λL cl (b,B)+λL cl (eos,I) (3)
[0238] wherein λ is a weight coefficient, L cl (eos,I) is the difference between the output display image description information and the sample display image features.
[0239] In one embodiment, step S210 can be implemented in the following steps S211-S214.
[0240] Step S211: Obtain a candidate training sample set.
[0241] The candidate training sample set includes each candidate training sample.
[0242] Step S212: Obtain a screening prompt, the screening prompt being used to instruct the large model to screen out training samples meeting a preset screening rule.
[0243] The screening prompt can be manually pre-designed. The screening prompt can include conditions that the training samples to be screened need to meet, and features of unqualified training samples. Those skilled in the art can flexibly set the screening prompt to screen the training samples meeting the requirements. For example, the screening prompt can include that the image meets the aesthetic features, the image contains the product and the background, etc., but is not limited thereto.
[0244] Step S213: input each candidate training sample in the candidate training sample set and the screening prompt into the pre-trained large model to obtain a screened training sample.
[0245] The large model can be a large language model (LLM) and can be a GPT model or a BERT model, but is not limited thereto.
[0246] Steps S211-S213 are the data screening link in FIG. 6.
[0247] Step S214: determine a training sample for model training according to the screened training sample.
[0248] Specifically, the sample in the screened training sample can be determined as the training sample for model training.
[0249] The present embodiment can screen each candidate training sample in the candidate training sample set by the large language model, so that the quality of the sample for training the model is higher, and thus a model with higher accuracy and higher image quality can be trained.
[0250] In one embodiment, the above model training method can further include steps S280-S2100.
[0251] Step S280: verify and evaluate the trained text generation model and text guided image generation model by using a verification sample set to obtain an evaluation result.
[0252] The evaluation result can include verification pass or verification fail, or the evaluation result can be a verification score, which is not specifically limited in the present disclosure.
[0253] Step S290: update the screening prompt according to the above evaluation result.
[0254] Specifically, when the evaluation result is evaluation pass, it indicates that the accuracy of the trained model is relatively high, and thus it indicates that the training sample used in the current training meets the requirements, and the adjustment and update of the screening prompt can be omitted. When the evaluation result is evaluation fail, the screening prompt is updated so as to screen more qualified training samples in the subsequent training. For example, when the original prompt is a long sentence, the screening prompt can be shortened to a short sentence, or other adjustments can be made, which are not specifically limited in the present disclosure. When the screening prompt is updated, the large language model can be adjusted to improve the efficiency.
[0255] Steps S280-S290 are the model evaluation and feedback improvement prompt link in FIG. 6.
[0256] Step S2100: When the expected training condition is not met, the process of model training in steps S210-S270 is continued.
[0257] The expected training condition can be at least one of the following: the number of training iterations reaches a preset number, the loss function converges, the verification on the verification sample is passed, the evaluation result meets a preset result, and the change amount of the loss function corresponding to each training round is less than a set threshold. The present disclosure is not specifically limited to other training conditions.
[0258] Correspondingly, step S211 can obtain the screening prompt in the following steps: obtaining the screening prompt updated for the last time. That is, the screening prompt obtained each time is the screening prompt updated for the last time.
[0259] The present embodiment updates the screening prompt according to the evaluation result of the model, so that the screened training sample is more in line with the model training requirements, and the model can be accurately and quickly trained in the subsequent iteration training process, so that the model with high accuracy can be more efficiently and quickly trained.
[0260] Optionally, step S210 can be implemented in the following step S211.
[0261] Step S211: deleting the training samples in the initial training sample set that have been excluded by the large model screening to obtain a deleted training sample set, and determining the deleted training sample set as a candidate training sample set.
[0262] Through step S211, the unqualified training samples can be deleted and then screened by the large model, so as to improve the screening efficiency of the large model.
[0263] The model training method provided in the present disclosure obtains training samples, the training samples include sample product description information of a sample product and a sample product display image with a pre-designed background, and based on the sample product description information and by using a to-be-trained text generation model, output background image description information corresponding to the sample product description information is obtained, sample background features corresponding to the sample product display image are determined, and the model parameters of the to-be-trained text generation model are adjusted according to the difference between the output background image description information and the sample background features, so as to obtain a trained text generation model. The present disclosure trains the to-be-trained text generation model by comparing the output background image description information and the sample background features, which can enhance the understanding of the text generation model for background elements. The model training method provided in the present disclosure can train a text generation model for generating corresponding background image description information according to product description information. The trained text generation model can efficiently and accurately generate high-quality background image description information that is strongly related to a product, so that the background image description information can be used as a prompt for a text-to-image generation model to efficiently generate high-quality product display images with a background matching the product.
[0264] The present embodiment introduces the training process of the estimation model from the perspective of model training. The processing process of the to-be-trained model on the training samples in the training process is similar to the processing process of the estimation model on the to-be-estimated information in the estimation process. Therefore, the specific implementation process and beneficial effects of the present embodiment can be referred to the first embodiment. The similar parts will not be described in detail.
[0265] Embodiment Three
[0266] The third embodiment of the present disclosure also provides an image generation device corresponding to the image generation method embodiment provided in the first embodiment. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. The details of the technical features and the effects achieved can be referred to the corresponding description of the image generation method embodiment provided above. The image generation device provided in the present embodiment includes:
[0267] The acquisition unit 301 is configured to acquire product description information and a product pattern corresponding to a target product of a to-be-generated display image.
[0268] The first generation unit 302 is configured to generate background image description information corresponding to the product description information by using a pre-trained text generation model based on the product description information.
[0269] The second generation unit 303 is configured to generate a display image with a background added to the product pattern by using a pre-trained text-to-image generation model and taking the background image description information as the guide text of the text-to-image generation model based on the product pattern.
[0270] Optionally, the text generation model is configured to generate corresponding description information according to the indication information.
[0271] The first generation unit is specifically configured to: generate, by a pre-trained text generation model, background picture description information corresponding to the product description information, by taking the product description information and editing background indication information as indication information of the text generation model.
[0272] Optionally, the first generation unit is further configured to: generate, by a pre-trained text generation model, display picture description information corresponding to the product description information, by taking the product description information and editing display picture indication information as indication information of the text generation model.
[0273] The second generation unit is specifically configured to: generate, by a pre-trained text-guided image generation model, a display picture with a background added to the product picture, by taking the background picture description information and the display picture description information as guide text of the text-guided image generation model, based on the product picture.
[0274] Optionally, the product description information includes at least one of the following: product name, product introduction, and product title.
[0275] Optionally, the second generation unit is specifically configured to: determine product region annotation information corresponding to the product picture, the product region annotation information marking a region of the target product in the product picture; and generate, by a pre-trained text-guided image generation model, a display picture with a background added to the product picture, by taking the background picture description information as guide text of the text-guided image generation model, based on the product picture and the product region annotation information.
[0276] Optionally, the text-guided image generation model is a stable diffusion model.
[0277] The second generation unit is specifically configured to: obtain noise to be added; and generate, by a pre-trained text-guided image generation model, a display picture with a background added to the product picture, by taking the background picture description information as guide text of the text-guided image generation model, based on the product picture, the product region annotation information, and the noise to be added.
[0278] Optionally, the second generation unit is specifically configured to: input the product picture into a pre-trained dimension reduction encoder to obtain reduced dimension picture encoding data corresponding to the product picture; and input the picture encoding data, the background picture description information, the product region annotation information, and the noise to be added into a pre-trained text-guided image generation model to generate a display picture with a background added to the product picture.
[0279] Embodiment Four
[0280] The fourth embodiment of the present disclosure also provides a model training device corresponding to the model training method embodiment provided by the second embodiment. Since the device embodiment is basically similar to the method embodiment, it is described more simply, and the details of the related technical features and the effects achieved can be seen by referring to the corresponding description of the model training method embodiment provided above. The image generation device provided by the present embodiment comprises:
[0281] a sample acquisition unit configured to acquire a training sample, the training sample comprising sample product description information of a sample product and a pre-designed sample product display image with a background;
[0282] a third generation unit configured to obtain output background image description information corresponding to the sample product description information by a to-be-trained text generation model based on the sample product description information;
[0283] a determination unit configured to determine a sample background feature corresponding to the sample product display image according to the sample product display image;
[0284] a training unit configured to adjust model parameters of the to-be-trained text generation model according to a difference between the output background image description information and the sample background feature, to obtain a trained text generation model, the text generation model being configured to generate corresponding background image description information according to product description information.
[0285] The fifth embodiment of the present disclosure also provides an electronic device corresponding to the image generation method provided by the first embodiment, and the following description of the electronic device embodiment is only illustrative. The electronic device embodiment comprises the following:
[0286] The above electronic device can be understood with reference to FIG. 8, which is a schematic diagram of the electronic device. The electronic device provided by the present embodiment comprises a processor 1001, a memory 1002, a communication bus 1003, and a communication interface 1004;
[0287] The memory 1002 is configured to store computer instructions for data processing, and the computer instructions are executed when read by the processor 1001 to perform the following steps:
[0288] acquire product description information and product patterns corresponding to a target product to be generated display image;
[0289] generate background image description information corresponding to the product description information by a pre-trained text generation model based on the product description information;
[0290] Based on the product pattern, a pre-trained text guided image generation model is used to generate a display image with a background added to the product pattern, with the background image description information as the guiding text of the text guided image generation model.
[0291] The sixth embodiment of the present disclosure also provides an electronic device embodiment corresponding to the model training method provided by the second embodiment, which is as follows.
[0292] The electronic device provided in the embodiment includes a processor, a memory, a communication bus, and a communication interface.
[0293] The memory is used to store computer instructions for data processing, and the computer instructions are executed by the processor to perform the following steps:
[0294] Obtain a training sample, which includes sample product description information of a sample product and a pre-designed sample product display image with a background;
[0295] Based on the sample product description information, a text generation model to be trained is used to obtain output background image description information corresponding to the sample product description information;
[0296] According to the sample product display image, determine the sample background feature corresponding to the sample product display image;
[0297] According to the difference between the output background image description information and the sample background feature, adjust the model parameters of the text generation model to be trained to obtain a trained text generation model, which is used to generate corresponding background image description information according to product description information.
[0298] The seventh embodiment of the present disclosure also provides a computer readable storage medium for implementing the method described in the first embodiment. The computer readable storage medium provided by the present disclosure is described relatively simply, and the related parts can be referred to the corresponding description of the above method embodiment. The following described embodiments are only illustrative.
[0299] The computer readable storage medium provided in the embodiment stores computer instructions, which are executed by the processor to implement the following steps:
[0300] Obtain product description information and a product pattern corresponding to a target product to be generated;
[0301] Based on the product description information, a pre-trained text generation model is used to generate background image description information corresponding to the product description information;
[0302] Based on the product pattern, a pre-trained text guided image generation model is used to generate a display image with a background added to the product pattern, with the background image description information as the guiding text of the text guided image generation model.
[0303] The seventh embodiment of the present disclosure also provides a computer readable storage medium for implementing the method of the second embodiment. The computer readable storage medium provided by the present disclosure is relatively simple, and the related parts can be referred to the corresponding description of the above method embodiment. The following described embodiments are only illustrative.
[0304] The computer readable storage medium provided by the present embodiment stores computer instructions, which are executed by a processor to implement the following steps:
[0305] Obtain a training sample, which includes sample product description information of a sample product and a pre-designed sample product display image with a background;
[0306] Based on the sample product description information, a text generation model to be trained is used to obtain output background image description information corresponding to the sample product description information;
[0307] According to the sample product display image, determine the sample background feature corresponding to the sample product display image;
[0308] According to the difference between the output background image description information and the sample background feature, adjust the model parameters of the text generation model to be trained to obtain a trained text generation model, which is used to generate corresponding background image description information according to product description information.
[0309] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0310] The memory can include non-persistent memory in computer readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer readable media.
[0311] 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device. According to the definition in this paper, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carriers.
[0312] 2. Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system or a computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0313] Although the present disclosure is disclosed with the preferred embodiments as above, it is not intended to limit the present disclosure, and any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present disclosure, therefore the protection scope of the present disclosure should be limited by the scope defined by the claims of the present disclosure.< / background> < / eos> < / background> < / eos> < / eos> < / background>
Claims
1. An image generation method, wherein, The method comprises: obtaining product description information and product pattern corresponding to a target product to be generated into a display pattern; based on the product description information, a pre-trained text generation model is used to generate background picture description information corresponding to the product description information; based on the product pattern, a pre-trained text guided image generation model is used to generate a display pattern with a background added to the product pattern, and the background picture description information is used as the guide text of the text guided image generation model.
2. The image generation method of claim 1, wherein, The text generation model is used to generate corresponding description information according to instruction information; The text generation model is used to generate corresponding description information according to instruction information; The text generation model is used to generate corresponding description information according to instruction information; 3. The image generation method of claim 2, wherein, The text generation model is used to generate corresponding description information according to instruction information; Before the step of generating a display pattern with a background added to the product pattern based on the product pattern, a pre-trained text guided image generation model is used, and the background picture description information is used as the guide text of the text guided image generation model, the method further comprises: a pre-trained text generation model is used to generate display pattern description information corresponding to the product description information, and the product description information and the edited display pattern instruction information are used as the instruction information of the text generation model; The text generation model is used to generate corresponding description information according to instruction information; 4. The image generation method according to any one of claims 1 to 3, wherein The product description information comprises at least one of the following: product name, product introduction, product title.
5. The image generation method according to claim 1 or 2, wherein The text generation model is used to generate corresponding description information according to instruction information; The product description information comprises at least one of the following: product name, product introduction, product title. The text generation model is used to generate corresponding description information according to instruction information; 6. The image generation method of claim 5, wherein, The product description information comprises at least one of the following: product name, product introduction, product title. The text generation model is used to generate corresponding description information according to instruction information; The text guided image generation model is a stable diffusion model; The method comprises: obtaining a product sample and product description information of the product sample; obtaining a pre-designed sample product display image with a background corresponding to the product sample; 7. The image generation method of claim 6, wherein, obtaining an output background image description information corresponding to the product description information by using a text generation model; obtaining a sample background feature corresponding to the sample product display image; adjusting model parameters of the text generation model according to a difference between the output background image description information and the sample background feature, to obtain a trained text generation model, wherein the trained text generation model is used to generate corresponding background image description information according to product description information.
8. A model training method in which, The method comprises: obtaining a product sample and product description information of the product sample; obtaining a pre-designed sample product display image with a background corresponding to the product sample; obtaining an output background image description information corresponding to the product description information by using a text generation model; obtaining a sample background feature corresponding to the sample product display image; 9.The model training method of claim 8, wherein, adjusting model parameters of the text generation model according to a difference between the output background image description information and the sample background feature, to obtain a trained text generation model, wherein the trained text generation model is used to generate corresponding background image description information according to product description information. The method comprises: 10.The model training method of claim 9, wherein, obtaining a product sample and product description information of the product sample; obtaining a pre-designed sample product display image with a background corresponding to the product sample; obtaining an output background image description information corresponding to the product description information by using a text generation model; obtaining a sample background feature corresponding to the sample product display image; adjusting model parameters of the text generation model according to a difference between the output background image description information and the sample background feature, to obtain a trained text generation model, wherein the trained text generation model is used to generate corresponding background image description information according to product description information. The method comprises: obtaining a product sample and product description information of the product sample; obtaining a pre-designed sample product display image with a background corresponding to the product sample; obtaining an output background image description information corresponding to the product description information by using a text generation model; obtaining a sample background feature corresponding to the sample product display image; adjusting model parameters of the text generation model according to a difference between the output background image description information and the sample background feature, to obtain a trained text generation model, wherein the trained text generation model is used to generate corresponding background image description information according to product description information. According to the difference between the output background picture description information and the sample background feature, the difference between the output display picture description information and the sample display picture feature, the model parameters of the text generation model to be trained are adjusted, wherein the text generation model is specifically used for generating corresponding background picture description information and product display picture description information according to product description information. 11.The model training method of claim 10, wherein, The sample product display picture is input into the image encoder to obtain the sample display picture feature corresponding to the sample product display picture. The sample product display picture is input into the image encoder to obtain the sample display picture feature corresponding to the sample product display picture. The training sample further includes a sample product pattern corresponding to the sample product, and the sample product pattern is a product region image cut from the sample product display picture. The method further includes: According to the difference between the output display picture and the sample display picture, the model parameters of the text guided image generation model to be trained and the text generation model to be trained are adjusted to obtain a trained text guided image generation model and a trained text generation model. 12.The model training method of claim 10 or 11, wherein, The text guided image generation model to be trained is a stable diffusion model. The method further includes: According to the difference between the output display picture and the sample display picture, the model parameters of the text guided image generation model to be trained and the text generation model to be trained are adjusted to obtain a trained text guided image generation model and a trained text generation model. The text guided image generation model to be trained is a stable diffusion model. 13.The model training method of claim 12, wherein, The method further includes: According to the difference between the output display picture and the sample display picture, the model parameters of the text guided image generation model to be trained and the text generation model to be trained are adjusted to obtain a trained text guided image generation model and a trained text generation model. The method further includes: According to the difference between the output display picture and the sample display picture, the model parameters of the text guided image generation model to be trained and the text generation model to be trained are adjusted to obtain a trained text guided image generation model and a trained text generation model. The method further includes: 14.The model training method of claim 13, wherein, According to the difference between the output display picture and the sample display picture, the model parameters of the text guided image generation model to be trained and the text generation model to be trained are adjusted to obtain a trained text guided image generation model and a trained text generation model. The method further includes: According to the difference between the output display picture and the sample display picture, the model parameters of the text guided image generation model to be trained and the text generation model to be trained are adjusted to obtain a trained text guided image generation model and a trained text generation model. The output background picture description information, the output display picture description information, the sample product picture, the sample product region annotation information, and a preset sample noise are input into a trained diffusion model to obtain an output display picture and predicted noise.
15. The model training method of any one of claims 8 to 14, wherein, The training sample acquisition includes: A candidate training sample set is acquired. A screening prompt is acquired, and the screening prompt is used to instruct a large model to screen out training samples that meet a preset screening rule. Each candidate training sample in the candidate training sample set and the screening prompt are input into a pre-trained large model to obtain screened training samples. The training sample used for model training is determined according to the screened training samples.
16. The model training method of claim 15, wherein, The method further includes: The trained text generation model and the text guided image generation model are verified and evaluated through a verification sample set to obtain an evaluation result. The screening prompt is adjusted and updated according to the evaluation result. When an expected training condition is not met, the steps of acquiring the training sample to adjusting the model parameters of the to-be-trained text guided image generation model and the to-be-trained text generation model are continuously performed. The screening prompt acquisition includes: A screening prompt that is updated most recently is acquired. The candidate training sample acquisition includes: Training samples that have been excluded by the large model in an initial training sample set are deleted to obtain a post-deletion training sample set, and the post-deletion training sample set is determined as a candidate training sample set.
17. The model training method of claim 16, wherein, The expected training condition includes at least one of the following: The evaluation result meets a preset result. The number of training iterations reaches a preset number. The change amount of a loss function corresponding to each training round is less than a set threshold.
18. An image generation apparatus, wherein The device includes: An acquisition unit is configured to acquire product description information and a product picture corresponding to a target product for which a display picture is to be generated. A first generation unit is configured to generate, based on the product description information, background picture description information corresponding to the product description information by using a pre-trained text generation model. A second generation unit is configured to generate, based on the product picture, a display picture with a background added to the product picture by using a pre-trained text guided image generation model and taking the background picture description information as a guide text of the text guided image generation model.
19. A model training apparatus, wherein, The device includes: A sample acquisition unit is configured to acquire training samples, and the training samples include sample product description information of sample products and a pre-designed sample product display picture with a background. A third generation unit is configured to generate, based on the sample product description information, output background picture description information corresponding to the sample product description information by using a to-be-trained text generation model. A determination unit is configured to determine sample background features corresponding to the sample product display picture according to the sample product display picture. A training unit is configured to adjust model parameters of the to-be-trained text generation model according to a difference between the output background picture description information and the sample background features, to obtain a trained text generation model, and the text generation model is used to generate corresponding background picture description information according to product description information.
20. An electronic device, comprising: The device includes: a processor, a memory, and computer program instructions stored on the memory and executable on the processor; the processor implements the method of any one of claims 1-17 when executing the computer program instructions.
21. A computer readable storage medium, wherein, the computer program instructions stored in the computer readable storage medium are executed by the processor to implement the method of any one of claims 1-17.
22. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-17.
Citation Information
Patent Citations
Method for generating commodity main image background and computing equipment
CN117058271A
Commodity display graph generation method and device, and medium
CN117315072A
Image generation method and device, program product and storage medium
CN118154727A
Image generation method and device, program product and storage medium
CN118505337A
Image generation method and device, model training method and device and electronic equipment
CN119540933A