Image generation methods, devices, storage media, and computer program products
By configuring a target low-rank adaptive image generation model, and combining the image to be processed with a reference style image, the problems of low generation efficiency and style inconsistency are solved, and the goal of efficiently generating target images that meet user expectations is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ALIBABA CLOUD COMPUTING CO LTD
- Filing Date
- 2024-11-25
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies suffer from low generation efficiency and image styles that do not meet user expectations when generating images containing specific objects and having a specific style.
By configuring a target low-rank adaptive model to fine-tune the image generation model, and combining the image to be processed, the reference style image, and the prompt words, the generation model can obtain the desired style features and visual features from different dimensions to generate a target image that meets the user's expectations.
It enables the rapid generation of target images that are highly consistent with the user's desired style, improving generation efficiency and reducing the cost of traditional live-action shooting and 3D modeling.
Smart Images

Figure CN122089557A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to an image generation method, apparatus, storage medium, and computer program product. Background Technology
[0002] With the rapid development of computer vision computing, there is a demand in various fields such as e-commerce and film and television production for generating target images containing specific objects and possessing a particular style. For example, generating product images with a photographic style featuring a particular item, or generating poster images with an anime style featuring a film or television character. In practical applications, target images can be generated through methods such as live-action photography and 3D modeling. However, these methods suffer from low generation efficiency and the inability to match user expectations in terms of image style. Summary of the Invention
[0003] This invention provides an image generation method, device, storage medium, and computer program product for efficiently generating images that meet user style requirements.
[0004] In a first aspect, embodiments of the present invention provide an image generation method, the method comprising:
[0005] Obtain the image to be processed, as well as the reference style image and prompt words corresponding to the image to be processed. The image to be processed contains a target object, and the prompt words are used to describe the visual feature information to be presented.
[0006] A target low-rank adaptive model is fine-tuned by a first style feature, wherein the reference style image has a second style feature;
[0007] The image to be processed, the reference style image, and the prompt word are input into the image generation model to generate a first target image, which contains the target object.
[0008] In a second aspect, embodiments of the present invention provide an image generation apparatus, the apparatus comprising:
[0009] The acquisition module is used to acquire the image to be processed, as well as the reference style image and prompt words corresponding to the image to be processed. The image to be processed contains a target object, and the prompt words are used to describe the visual feature information to be presented.
[0010] A generation module is used to determine an image generation model finely tuned by a target low-rank adaptive model corresponding to a first style feature, wherein the reference style image has a second style feature; the image to be processed, the reference style image, and the prompt word are input into the image generation model to generate a first target image, wherein the first target image contains the target object.
[0011] Thirdly, embodiments of the present invention provide an electronic device, including: a memory, a processor, and a communication interface; wherein, the memory stores executable code, and when the executable code is executed by the processor, the processor performs the image generation method as described in the first aspect.
[0012] Fourthly, embodiments of the present invention provide a non-transitory machine-readable storage medium storing executable code, which, when executed by a processor of an electronic device, enables the processor to at least implement the image generation method as described in the first aspect.
[0013] Fifthly, embodiments of the present invention provide a computer program product, the computer program product including a computer program, which, when executed by a processor, can implement the image generation method as described in the first aspect.
[0014] In the image generation method provided by this invention, to generate a target image containing a specific object (i.e., the target object) and having a specific style, firstly, an image to be processed containing the target object, a reference style image corresponding to the image to be processed, and a prompt word are obtained. The prompt word describes the visual feature information expected to be presented, such as the object contained in the expected target image, the object's size, color, lighting, and visual feature information such as image style. Next, an image generation model fine-tuned by a target low-rank adaptation model corresponding to the first style feature is determined. Finally, the image to be processed, the reference style image, and the aforementioned prompt word are input into the image generation model to generate the target image containing the target object. During the process of generating the target image containing the target object through the image generation model, the target low-rank adaptation model can provide information related to the first style feature; the reference style image can provide information related to the second style feature. Based on this, the target image generated by the image generation model can present the first style feature, the second style feature, and the visual features corresponding to the visual feature information in the prompt word.
[0015] In this scheme, by configuring a target low-rank adaptive model corresponding to the desired generated style features for the image generation model, and inputting a reference style image corresponding to the desired generated style features, the image generation model can fully acquire the desired generated style features from different dimensions. As a result, the generated target image can present the visual features corresponding to the visual feature information, and the presented style features can be highly consistent with the user's expectations. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A flowchart of an image generation method provided in an embodiment of the present invention;
[0018] Figure 2 A flowchart of another image generation method provided in an embodiment of the present invention;
[0019] Figure 3 A flowchart illustrating yet another image generation method provided in an embodiment of the present invention;
[0020] Figure 4 A scene diagram illustrating an image generation method provided in an embodiment of the present invention;
[0021] Figure 5 This is a schematic diagram of the structure of an image generation device provided in an embodiment of the present invention;
[0022] Figure 6 This is a schematic diagram of the structure of an electronic device provided in this embodiment. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. In addition, the timing of the steps in the following method embodiments is only an example and not a strict limitation.
[0024] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0025] To facilitate understanding, the relevant concepts involved in the embodiments of the present invention will be explained below.
[0026] Image generation models refer to various large models capable of performing text-to-image and image-to-image tasks, such as the Static Diffusion (SD) model. Among these, image generation models include fast versions that can generate high-quality images in a short time and with a few steps, such as fast versions of the Static Diffusion eXtra Large (SDXL) model, such as SDXL-turbo and SDXL-lighting.
[0027] A style adapter is a component used to introduce style features from a reference style image into the image generation process of an image generation model in order to control the style of the images generated by the image generation model. It includes, but is not limited to, an image-prompt adapter (IP-Adapter) and a style encoder.
[0028] Low-Rank Adaptation (LoRA) models can be viewed as a style feature annotation tool for image generation models. By configuring a LoRA model into an image generation model, the generated images can exhibit the style features corresponding to the LoRA model while maintaining the original image generation performance. In practical applications, the image generation model can be fine-tuned using a small number of images of a specific style to obtain a low-rank matrix corresponding to that style for adjusting the weights of the image generation model, i.e., the LoRA model. Optionally, to improve the image generation quality, images with higher image quality can be used for fine-tuning the image generation model.
[0029] Prompts are descriptive text entered by the user to tell the image generation model the visual features that the generated image should present, such as: which objects should be included in the generated image, the color, size and position of the included objects, and the style of the image.
[0030] In practical applications, there is often a need to generate target images containing specific objects and possessing a particular image style. For example, in e-commerce scenarios, merchants may need to generate product images taken in different usage scenarios, from different shooting angles, or with different models showcasing the product. In film production, advertising design, and other application scenarios, designers may need to generate promotional posters in different styles for the design subject (e.g., film characters), such as anime style, hand-drawn style, or illustration style. Traditionally, target images can be generated by setting up corresponding shooting scenes and taking real-world photos, or through 3D modeling. However, these methods suffer from low generation efficiency and the image style not meeting user expectations.
[0031] To address the issues of low efficiency and inconsistency between the target image generation efficiency and the user's desired style in traditional image generation schemes, this invention provides a novel image generation scheme. In this scheme, the target image is generated through an image generation model. The image generation model is configured with a target low-rank adaptive model corresponding to a first style feature; that is, the image generation model is fine-tuned using the target low-rank adaptive model corresponding to the first style feature. Furthermore, the image generation model can acquire a second style feature from an input reference style image. Based on the style learning capability of this image generation model, a target image with a high degree of consistency with the user's desired style can be generated quickly.
[0032] Optionally, the embodiments of the present invention may employ a fast generation model, such as the fine-tuning model of SDXL-turbo / SDXL-lightning. Taking the product image generation scenario below as an example, fine-tuning here refers to fine-tuning the training through high-quality product image data to better generate target image types (such as product images).
[0033] The following detailed description of some embodiments of the present invention is provided in conjunction with the accompanying drawings. Where there is no conflict between the embodiments, the following embodiments and features thereof can be combined with each other.
[0034] The image generation method provided in this embodiment of the invention can be executed by an electronic device, such as a processor like a GPU in the electronic device. This electronic device can be a terminal device such as a PC, laptop, or smartphone, or it can be a server. The server can be a physical server containing an independent host, a virtual server, or a server or server cluster in the cloud.
[0035] Figure 1 A flowchart of an image generation method provided in an embodiment of the present invention is shown below. Figure 1 As shown, the method may include the following steps:
[0036] 101. Obtain the image to be processed, as well as the corresponding reference style image and prompt words. The image to be processed contains the target object, and the prompt words are used to describe the visual feature information expected to be presented.
[0037] 102. Determine the image generation model finely tuned by the target low-rank adaptation model corresponding to the first style feature, where the reference style image has the second style feature.
[0038] 103. Input the image to be processed, the reference style image, and the prompt words into the image generation model to generate a first target image, which contains the target object.
[0039] In image processing scenarios, an image can generally be divided into two parts: foreground and background. The foreground is usually the object or area of interest to the user and contains more information, such as a person or a product; while the background refers to the part of the image other than the foreground and contains less information, such as the ground, sky, or walls.
[0040] The image generation method provided in this invention can be used to generate a first target image with a target object (such as a product or person) as the foreground and other content to be presented as the background, and the foreground and background have a desired visual effect. The first target image is generated using an image generation model.
[0041] In summary, during the process of generating the first target image using an image generation model, on the one hand, the model is provided with the relevant information needed to generate the first target image by including the image to be processed containing the target object, the reference style image, and the target low-rank adaptation model; on the other hand, prompts describing the visual features that the expected first target image should present are used to guide the image generation model in generating the first target image. Thus, the image generation model can output a first target image that meets the user's expectations based on its configured target low-rank adaptation model and the input image to be processed, the reference style image, and the prompts.
[0042] The following section elaborates on the process of generating the first target image based on the image generation model.
[0043] First, the inputs to the image generation model, namely the image to be processed, the reference style image and prompt words corresponding to the image to be processed, and the target low-rank adaptation model configured by the image generation model are explained.
[0044] The image to be processed refers to any image containing the target object. It is used to provide the contour information of the target object during the generation of the first target image by the image generation model, ensuring a clear boundary between the target object and the background when the target object is used as the foreground in the first target image, resulting in a high-quality first target image. Optionally, the image to be processed can be a real photograph of the target object or a modeled image of the target object. In this embodiment, there is no limitation on other regions in the image to be processed besides the target object; for example, the image to be processed may contain only the target object or other objects besides the target object.
[0045] The reference style image corresponding to the image to be processed is an image with the same or similar style as the image style expected by the user. For example, assuming the target object in the image to be processed is a product, and the first target image to be generated is a product image of that product, and the product image of that product has the same style as other product images already published by the merchant, then the reference style image can be other product images already published by the merchant.
[0046] The prompts corresponding to the image to be processed describe the visual features that the user expects the first target image to present. It's understandable that before generating the first target image through the image generation model, the user has already conceived the visual effect they want the generated first target image to present; the prompts are essentially a textual expression of the user's expected visual effect.
[0047] Optionally, the visual feature information described by the prompt may include at least one of the following visually related information: the image style presented by the first target image, the objects contained in the first target image, and the visual features of the contained objects such as their location, color, size, material, and lighting. The objects contained in the first target image described in the prompt may include only the target object, or they may include other objects besides the target object.
[0048] To facilitate understanding, consider this example: In a scenario where a product image (i.e., the first target image) is generated with a perfume bottle in the foreground, the image to be processed could be a real-life photograph of the perfume bottle. The corresponding prompt could be: "A small perfume bottle is placed on a metal display tray, surrounded by sunlight, a metal wall, and a metal floor, using a large aperture with a blurred background." Here, "perfume bottle" and "display tray" are the objects contained in the first target image. "Large aperture" and "blurred background" reflect that the first target image's style is a real-life photograph with a blurred background. Furthermore, the prompt uses words like "small," "metal," and "placed on..." to define the size, material, and position of the "perfume bottle" and "display tray."
[0049] Based on the above explanation of the prompt words, the image generation model can obtain certain information for generating the first target image through the visual feature information described by the prompt words. This information includes, for example, what objects are contained in the first target image, the visual features of these objects, and the image style that the first target image needs to present. Regarding the image style that the first target image needs to present, it is understandable that the description in the prompt words is often general and lacks detail.
[0050] To ensure that the image style of the first target image generated by the image generation model is consistent with the image style expected by the user, a reference style image with the same or similar image style as the expected image style can be input into the image generation model. This allows the image generation model to obtain the corresponding style features from the reference style image and transfer these style features to the first target image, thereby making the first target image have the same image style as the reference style image, that is, the first target image presents the style features of the reference style image.
[0051] Optionally, the image generation model includes a style adapter, which is used to obtain style features from a reference style image and introduce the style features of the reference style image into the image generation process of the image generation model to control the style features presented by the first target image. The style adapter includes, but is not limited to, an image-prompt adapter (IP-Adapter) and a style encoder.
[0052] In practical applications, during the training phase of an image generation model, reference style image samples with different image styles are typically used for training. This allows the image generation model to extract the style features corresponding to the learned image styles from the reference style images when it is in use.
[0053] However, image styles are diverse, and it cannot be guaranteed that the image generation model has learned all image styles. Therefore, in practical applications, the following situations may occur: The image style corresponding to the reference style image input to the image generation model is one that the model has not learned. The model cannot accurately extract the corresponding style features from this reference style image, resulting in the style features of the generated first target image being inconsistent with the user's expected style. Alternatively, the user's expected image style is a combination of multiple image styles, leading to the input of multiple reference style images with different styles into the image generation model. Since the model has not learned the style features of the combined image styles, similarly, it cannot accurately extract the corresponding style features from this reference style image, resulting in the style features of the generated first target image being inconsistent with the user's expected style.
[0054] To enable the image generation model to accurately acquire style features of various image styles, this solution configures a target low-rank adaptation model within the image generation model. Specifically, the image generation model in this embodiment refers to an image generation model fine-tuned using a target low-rank adaptation model corresponding to a specific style feature. The target low-rank adaptation model provides the image generation model with style features corresponding to image styles that the model did not learn during training and that are desired by the user. By configuring the target low-rank adaptation model into the image generation model, the first target image generated by the model can exhibit the style features corresponding to the target low-rank adaptation model while maintaining its original image generation performance.
[0055] In practical applications, image generation models can be fine-tuned using a small number of images with specific image styles. This yields a low-rank matrix corresponding to that image style, used to adjust the weights of the image generation model—a low-rank adaptive model. Compared to retraining the image generation model to learn style features of new image styles, the low-rank adaptive model provides a faster and more efficient way to supplement the image generation model with missing style features.
[0056] Assuming the image generation model before fine-tuning is called the pre-trained image generation model, then the image generation model fine-tuned by the target low-rank adaptation model corresponding to a certain style feature as described in this embodiment of the invention refers to the image generation model obtained by fine-tuning the pre-trained image generation model by the target low-rank adaptation model corresponding to a certain style feature.
[0057] Optionally, determining an image generation model fine-tuned by a target low-rank adaptation model corresponding to a certain style feature includes: fine-tuning a pre-trained image generation model by using low-rank adaptation models corresponding to different style features to obtain multiple image generation models corresponding to different style features; storing multiple image generation models corresponding to different style features; and obtaining an image generation model corresponding to a certain style feature from the stored image generation models corresponding to multiple different style features for use in image generation.
[0058] Optionally, determining an image generation model fine-tuned by a target low-rank adaptation model corresponding to a certain style feature includes: storing low-rank adaptation models corresponding to different style features; obtaining a target low-rank adaptation model corresponding to a certain style feature from the low-rank adaptation models corresponding to different style features; and fine-tuning the pre-trained image generation model using the target low-rank adaptation model to obtain an image generation model for image generation.
[0059] The above provides two methods for determining the image generation model fine-tuned using a target low-rank adaptive model corresponding to a specific style feature. The first method pre-generates and stores the image generation model fine-tuned using the low-rank adaptive model. This allows for direct use of the fine-tuned model without needing to update the weight parameters of the pre-trained model based on the low-rank adaptive model, improving image generation efficiency. However, since each fine-tuned model requires storing the same number of parameters as the pre-trained model, it consumes significant storage space. The second method uses an un-fine-tuned pre-trained image generation model plus a low-rank adaptive model. This requires updating the weight parameters of the pre-trained model based on the low-rank adaptive model, potentially increasing computational overhead. However, since only the parameters of the pre-trained model and the corresponding low-rank adaptive models need to be stored, storage space consumption is significantly reduced.
[0060] Furthermore, in the second method described above, where the image generation model is fine-tuned using the target low-rank adaptation model, the target low-rank adaptation model and the pre-trained image generation model exhibit a "pluggable" relationship. This means that by switching between low-rank adaptation models corresponding to different style features, the pre-trained image generation model can be fine-tuned to obtain different image generation models. This allows for the generation of images with various style features, thus providing good application flexibility. This will be explained in more detail below.
[0061] In practical applications, the target low-rank adaptation model configured by the image generation model can be flexibly switched according to the user's desired image style. For example, when the user's desired image style is image style 1, the low-rank adaptation model 1 corresponding to image style 1 can be selected for the image generation model; when the user's desired image style is image style 2, the low-rank adaptation model 2 corresponding to image style 2 can be selected for the image generation model. In this embodiment, the low-rank adaptation model configured by the image generation model when generating the first target image is referred to as the target low-rank adaptation model.
[0062] In one optional embodiment, image style can be represented by style features corresponding to the image style. Correspondingly, low-rank adaptation models corresponding to different style features can be pre-generated, that is, low-rank adaptation models corresponding to different image styles. When it is necessary to generate a first target image containing the target object in the image to be processed, multiple style features corresponding to the image to be processed are displayed for the user to select. In response to the user's selection operation of a certain style feature x among the multiple style features, the low-rank adaptation model corresponding to style feature x is obtained, and the low-rank adaptation model corresponding to style feature x is configured in the image generation model. Here, the style feature x selected by the user is a style feature that the user expects the first target image to present.
[0063] In practical applications, different image types typically correspond to different applicable image styles. For example, e-commerce images may be suitable for live-action photography, while film and television production may be suitable for hand-drawn or illustration styles. To address the differences in image types, or the different scenarios in which images are applicable, image styles can be grouped according to image type to establish a mapping relationship between image type and style grouping. For example, image type i corresponds to style group j, where style group j includes style feature A corresponding to image style a, style feature B corresponding to image style b, etc. The style features contained in the style groups corresponding to different image types can be the same or different.
[0064] In another optional embodiment, when it is necessary to generate a first target image containing the target object in the image to be processed, multiple image types that can be generated can be displayed, wherein different image types correspond to different style groups; in response to the user's selection operation of the target image type to which the first target image belongs, the style group corresponding to the target image type is displayed, wherein the style group corresponding to the target image type includes multiple style features, and different style features correspond to different low-rank adaptation models; further, in response to the user's selection operation of a certain style feature y among the multiple style features, the low-rank adaptation model corresponding to style feature y is obtained and configured in the image generation model. Here, the style feature y selected by the user is a style feature that the user expects the first target image to present.
[0065] Optionally, an image type may contain two style groups: a first style group and a second style group. Both the first and second style groups contain multiple style features. The style features in the first style group correspond to the low-rank adaptation model; selecting a style feature from the first style group yields the corresponding low-rank adaptation model. The style features in the second style group correspond to a reference style image; selecting a style feature from the second style group yields the corresponding reference style image. By selecting style features from the first and second style groups, users can flexibly switch the image style presented by the first target image.
[0066] After explaining the input to the image generation model and the target low-rank adaptive model configured in the image generation model, the process of generating the first target image by the image generation model will be explained next.
[0067] For ease of distinction, the style features corresponding to the image style configured in the target low-rank adaptation model of the image generation model are called the first style features, and the style features that the image generation model can obtain from the reference style image are called the second style features. Based on the above introduction, the first style features are the style features that the image generation model has not learned during the training phase; the second style features are the style features that the image generation model has learned during the training phase.
[0068] In the specific implementation process, after the image to be processed, the reference style image, and the prompt words are input into the image generation model, the image generation model obtains the contour information of the target object from the image to be processed, the first style feature from the target low-rank adaptation model, the second style feature from the reference style image, and the visual feature information to be presented from the prompt words. Based on the information obtained, the image generation model outputs the first target image, in which the target object is the foreground and presents the first style feature, the second style feature, and the visual features corresponding to the visual feature information.
[0069] In summary, by configuring a target low-rank adaptive model corresponding to the desired generated style features for the image generation model (i.e., fine-tuning the image generation model through the target low-rank adaptive model), and by inputting a reference style image corresponding to the desired generated style features, an image to be processed containing the target object, and prompts describing the visual feature information to be presented, the image generation model can fully acquire the desired generated style features and related visual features from different dimensions. As a result, the first target image generated can present the visual features corresponding to the visual feature information, and the presented style features can be highly consistent with the user's expectations. In addition, the speed of generating the first target image through the image generation model is much faster than the traditional real-scene shooting and 3D modeling.
[0070] Figure 2 A flowchart of another image generation method provided in an embodiment of the present invention, such as... Figure 2 As shown, the method may include the following steps:
[0071] 201. Obtain the image to be processed, as well as the corresponding reference style image and prompt words. The image to be processed contains the target object, and the prompt words are used to describe the visual feature information expected to be presented.
[0072] 202. Determine the image generation model finely tuned by the target low-rank adaptation model corresponding to the first style feature, and the reference style image has the second style feature.
[0073] 203. Extract the cutout image corresponding to the target object from the image to be processed.
[0074] 204. Perform edge detection on the cutout image to obtain the edge contour image of the target object.
[0075] 205. Input the edge contour image, reference style image and prompt words into the image generation model to generate the first target image through the image generation model.
[0076] The specific implementation process of steps 201 and 202 can be referred to the aforementioned embodiments, and will not be repeated in this embodiment.
[0077] As mentioned in the foregoing embodiments, the image to be processed can be directly input into an image generation model to obtain the contour information of the target object. However, in practical applications, obtaining the contour information of the target object from the image to be processed through an image generation model suffers from low efficiency.
[0078] To solve this technical problem, a cutout image corresponding to the target object can be extracted from the image to be processed first. The cutout image contains only the target object. Then, edge detection is performed on the cutout image to obtain the edge contour image of the target object. The edge contour image is a binary black and white line image. After that, the edge contour image is used to replace the image to be processed and input into the image generation model so that the image generation model can obtain the contour information of the target object from the edge contour.
[0079] Optionally, deep learning models can be used, such as U-shaped Network with Nested U-structures (U2net) to extract the cutout image corresponding to the target object from the image to be processed; edge detection methods such as Canny edge detection and convolutional neural networks can be used to perform edge detection on the image to be processed to obtain the edge contour image corresponding to the image to be processed.
[0080] Understandably, edge contour images contain less information than richly colored images, making it easier for image generation models to learn the contour information of the target object. When computational resources are limited, edge contour images can effectively reduce the computational burden on image generation models. Furthermore, for generating the first target image, since the image to be processed is used as input, the image generation model focuses more on the shape, i.e., the contour information, of the target object. The visual features of the target object, such as color and texture, presented in the first target image are obtained from the prompts. Therefore, processing the image to be processed into an edge contour image and inputting it into the image generation model can effectively reduce excessive interference from visual features such as color in the image to be processed, allowing the image generation model to better understand the contour information of the target object.
[0081] In this embodiment, after acquiring the image to be processed, the corresponding reference style image, and the prompt word, the matted image corresponding to the target object is extracted from the image to be processed. Edge detection is then performed on the matted image to obtain the edge contour image of the target object. Finally, the edge contour image, the reference style image, and the prompt word are input into an image generation model fine-tuned by a target low-rank adaptive model corresponding to the first style feature, so as to generate a first target image. Because the edge contour image contains less information and has fewer interfering factors in acquiring the target object's contour information, the image generation model can quickly and accurately obtain the target object's contour information from the edge contour image. The generated first target image has a clear boundary between the target object and the background, and the overall image quality of the first target image is higher.
[0082] Figure 3 A flowchart of another image generation method provided in an embodiment of the present invention, such as... Figure 3 As shown, the method may include the following steps:
[0083] 301. Obtain the image to be processed, as well as the corresponding reference style image and prompt words. The image to be processed contains the target object, and the prompt words are used to describe the visual feature information expected to be presented.
[0084] 302. Determine the image generation model finely tuned by the target low-rank adaptation model corresponding to the first style feature, and the reference style image has the second style feature.
[0085] 303. Input the image to be processed, the reference style image, and the prompt words into the image generation model to generate the first target image through the image generation model.
[0086] 304. The cutout image corresponding to the target object extracted from the image to be processed is superimposed on the first target image to generate the second target image. The aspect ratio of the cutout image and the first target image is the same.
[0087] The specific implementation process of steps 301 to 303, and the process of obtaining the cutout image corresponding to the target object, can be referred to the aforementioned embodiments, and will not be repeated in this embodiment.
[0088] Since the visual display effect (e.g., color) of the target object in the first target image generated by the image generation model is determined based on the prompt words, in order to make the target object in the first target image closer to the visual features such as color of the target object in the image to be processed, that is, to more realistically restore the visual information of the target object, after the image generation model outputs the first target image, the cutout image corresponding to the target object can be superimposed on the first target image as a layer, so that the target object in the cutout image occludes the target object in the first target image, while other image areas in the first target image are unaffected.
[0089] In this process, the image to be processed and the cutout image of the target object are the same size, and the cutout image and the first target image have the same aspect ratio. When overlaying the cutout image corresponding to the target object with the first target image, the size of the cutout image can be adjusted so that the target object it contains has the same size as the target object in the first target image, thus enabling the overlay.
[0090] In this embodiment, the result of overlaying the cutout image corresponding to the target object with the first target image is called the second target image. Compared with the first target image, the target object in the second target image presents the visual features such as color of the target object in the image to be processed, which can better display the target object.
[0091] Figure 4This is a scene illustration of an image generation method provided in an embodiment of the present invention. The following is in conjunction with... Figure 4 The actual application process of the image generation method provided in the embodiments of the present invention will be described.
[0092] exist Figure 4 In the example, a user wants to generate a product image of a perfume bottle, which is placed on a metal display tray surrounded by sunlight, metal walls, and a metal floor. The product image has the image style of an e-commerce image and presents a photography style with a large aperture and a blurred background.
[0093] Assuming that the image generation model has not learned the first style features corresponding to the photography style of shooting with a large aperture and blurred background during the training phase, but has learned the second style features corresponding to the image style of e-commerce images, then the target low-rank adaptation model matrix corresponding to the first style features is configured for the image generation model. This is to determine the image generation model finely tuned by the target low-rank adaptation model matrix corresponding to the first style features, so that the perfume bottle product image generated by the image generation model can present the first style features.
[0094] like Figure 4 As shown, in the process of generating a product image of a perfume bottle through an image generation model, firstly, the image to be processed containing the perfume bottle (i.e., the target object) is obtained, existing e-commerce images are sampled as reference style images, and prompts describing the visual features of the product image of the perfume bottle (i.e., the first target image) that the user expects to generate are obtained. For example: a small perfume bottle is placed on a metal display tray, surrounded by sunlight, metal walls and metal ground, e-commerce image, using a large aperture, background blur processing.
[0095] The image to be processed may or may not contain elements other than perfume bottles, such as... Figure 4 The image to be processed contains a small flower in addition to the perfume bottle.
[0096] Then, the image to be processed, the reference style image, and the prompt words can be directly input into the image generation model to generate a product image of the perfume bottle. Or, as... Figure 4 As shown, the perfume bottle image is first extracted from the image to be processed. Then, edge detection is performed on the extracted image to obtain the edge contour image of the perfume bottle. After that, the edge contour image of the perfume bottle, the reference style image, and the prompt words are input into the image generation model to generate a product image of the perfume bottle.
[0097] like Figure 4The product image of the perfume bottle shows it placed on a metal display tray, surrounded by sunlight, metal walls, and a metal floor. It also features the image style of e-commerce images, taken with a wide aperture and a blurred background, which is consistent with the visual effect expected by users.
[0098] Furthermore, after generating the product image of the perfume bottle using the image generation model, the cut-out image of the perfume bottle can be overlaid onto the product image of the perfume bottle generated by the image generation model (i.e., the first target image) to generate a second target image, which serves as the final product image. The perfume bottle in the second target image exhibits the same visual features, such as color, as the perfume bottle in the image to be processed, thus better showcasing the perfume bottle.
[0099] The specific implementation process involved in the embodiments of the present invention can be referred to the content of the above embodiments, and will not be repeated here.
[0100] The image generation apparatus of one or more embodiments of the present invention will be described in detail below. Those skilled in the art will understand that these apparatuses can all be configured using commercially available hardware components through the steps taught in this solution.
[0101] Figure 5 This is a schematic diagram of the structure of an image generation device provided in an embodiment of the present invention, such as... Figure 5 As shown, the device includes: an acquisition module 11 and a generation module 12.
[0102] The acquisition module 11 is used to acquire the image to be processed, as well as the reference style image and prompt words corresponding to the image to be processed. The image to be processed contains a target object, and the prompt words are used to describe the visual feature information to be presented.
[0103] The generation module 12 is used to determine an image generation model finely tuned by a target low-rank adaptive model corresponding to a first style feature, wherein the reference style image has a second style feature; and to input the image to be processed, the reference style image, and the prompt word into the image generation model to generate a first target image, wherein the first target image contains the target object.
[0104] Optionally, the image generation model includes a style adapter for obtaining the second style feature from the reference style image.
[0105] Optionally, during the training phase of the image generation model, the image generation model has not learned the first style feature; during the training phase of the image generation model, the image generation model has learned the second style feature.
[0106] Optionally, the generation module 12 is specifically used to extract the cutout image corresponding to the target object from the image to be processed; perform edge detection on the cutout image to obtain the edge contour image of the target object; and input the edge contour image, the reference style image, and the prompt word into the image generation model to generate a first target image through the image generation model.
[0107] Optionally, the generation module 12 is further configured to overlay the cutout image corresponding to the target object extracted from the image to be processed onto the first target image to generate a second target image, wherein the cutout image and the first target image have the same aspect ratio.
[0108] Optionally, the acquisition module 11 is further configured to display multiple style features corresponding to the image to be processed, wherein different style features correspond to different low-rank adaptation models; and in response to the user's selection operation of the first style feature among the multiple style features, to acquire the target low-rank adaptation model corresponding to the first style feature.
[0109] Optionally, the acquisition module 11 is further configured to display multiple image types that can be generated, with different image types corresponding to different style groups; in response to the user's selection operation of the target image type to which the first target image belongs, the style group corresponding to the target image type is displayed, and the style group corresponding to the target image type includes the multiple style features.
[0110] Figure 5 The apparatus shown can perform the steps in the image generation method in the foregoing embodiments. For detailed execution process and technical effects, please refer to the description in the foregoing embodiments, which will not be repeated here.
[0111] This invention also provides an electronic device, such as... Figure 6 As shown, the electronic device may include: a processor 21, a memory 22, and a communication interface 23. The memory 22 stores executable code, which, when executed by the processor 21, enables the processor 21 to implement the image generation method as described in the foregoing embodiments.
[0112] In addition, embodiments of the present invention provide a non-transitory machine-readable storage medium storing executable code, which, when executed by a processor of an electronic device, enables the processor to at least implement the image generation method provided in the foregoing embodiments.
[0113] The invention also provides a computer program product comprising a computer program that, when executed by a processor, enables the processor to at least implement the image generation method provided in the foregoing embodiments.
[0114] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separate. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0115] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of a necessary general-purpose hardware platform, or by a combination of hardware and software. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a computer product. The present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image generation method, characterized in that, include: Obtain the image to be processed, as well as the reference style image and prompt words corresponding to the image to be processed. The image to be processed contains a target object, and the prompt words are used to describe the visual feature information to be presented. A target low-rank adaptive model is fine-tuned by a first style feature, wherein the reference style image has a second style feature; The image to be processed, the reference style image, and the prompt word are input into the image generation model to generate a first target image, which contains the target object.
2. The method according to claim 1, characterized in that, The image generation model includes a style adapter, which is used to obtain the second style feature from the reference style image.
3. The method according to claim 1, characterized in that, During the training phase of the image generation model, the image generation model has not learned the first style feature; during the training phase of the image generation model, the image generation model has learned the second style feature.
4. The method according to claim 1, characterized in that, The step of inputting the image to be processed, the reference style image, and the prompt word into an image generation model to generate a first target image through the image generation model includes: Extract the cutout image corresponding to the target object from the image to be processed; Edge detection is performed on the cut-out image to obtain the edge contour image of the target object; The edge contour image, the reference style image, and the prompt word are input into the image generation model to generate a first target image.
5. The method according to claim 1, characterized in that, The method further includes: The cutout image corresponding to the target object extracted from the image to be processed is superimposed on the first target image to generate a second target image, wherein the aspect ratio of the cutout image and the first target image is the same.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: Displays multiple style features corresponding to the image to be processed, wherein different style features correspond to different low-rank adaptation models; In response to the user's selection of the first style feature among the multiple style features, the target low-rank adaptive model corresponding to the first style feature is obtained.
7. The method according to claim 6, characterized in that, The method further includes: The display shows the various image types that can be generated, with different style groups corresponding to different image types; In response to the user's selection of the target image type to which the first target image belongs, the style grouping corresponding to the target image type is displayed, and the style grouping corresponding to the target image type includes the multiple style features.
8. An electronic device, characterized in that, include: The system includes a memory, a processor, and a communication interface; wherein the memory stores executable code, which, when executed by the processor, causes the processor to perform the image generation method as described in any one of claims 1 to 7.
9. A non-transitory machine-readable storage medium, characterized in that, The non-transitory machine-readable storage medium stores executable code that, when executed by a processor of an electronic device, causes the processor to perform the image generation method as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, include: A computer program, when executed by a processor of an electronic device, causes the processor to perform the image generation method as described in any one of claims 1 to 7.