Method, apparatus, device and program product for adjusting image generation model
By introducing a conditional extraction model and a deep learning method into the image generation model and comparing the first and second conditional images, the problem of the difficulty in accurately controlling the image generation model in the existing technology is solved, and high-quality image generation and adjustment are achieved.
Patent Information
- Application Number
- CN202410346424.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-25
- Publication Date
- 2025-09-26
AI Technical Summary
Existing image generation models have difficulty achieving precise and fine-grained control over images, and have difficulty generating images that are consistent with the input conditions, especially in terms of details that are difficult to meet user expectations.
A first image is generated based on a first conditional image by an image generation model, and a second conditional image is generated by a condition extraction model. The image generation model is adjusted by comparing the first and second conditional images, and image generation and adjustment are performed using deep learning and diffusion methods.
It achieves precise control over image generation, reduces training costs, improves user experience, and generates high-quality images consistent with input conditions.
Smart Images

Figure CN120707654A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to the field of image processing, and more particularly, to methods, apparatuses, electronic devices, and program products for adjusting image generation models. Background Art
[0002] With the continuous advancement and innovation of technology, using image generation models to generate images has become a very active research direction. These technologies can not only generate high-quality images, but also have a wide range of applications and innovations in many fields such as art creation, data enhancement, game development, virtual reality, etc.
[0003] Image generation models can generate realistic images, perform creative transformations and enhancements on existing images, create entirely new visual content, and are capable of performing activities such as automatic content generation, personalized media production, game development, film and video production, advertising, education and training simulations, and more. For example, product or architectural designers can use image generation models to generate concept drawings for product designs, interior design plans, or other works of art. Image generation models can also help create clearer medical images. Summary of the Invention
[0004] Embodiments of the present disclosure provide a method, apparatus, electronic device, and program product for adjusting an image generation model.
[0005] According to a first aspect of the present disclosure, a method for adjusting an image generation model is provided. The method includes generating a first image based on a first conditional image by the image generation model, where the first conditional image is used to control the generation of the first image. The method also includes generating a second conditional image based on the first image by a conditional extraction model. Furthermore, the method includes adjusting the image generation model based on a comparison between the first conditional image and the second conditional image.
[0006] In a second aspect of the present disclosure, a device for adjusting an image generation model is provided. The device includes an image generation module configured to generate a first image based on a first conditional image by the image generation model, where the first conditional image is used to control the generation of the first image. The device also includes a conditional image generation module configured to generate a second conditional image based on the first image by a conditional extraction model. The device also includes a comparison module configured to adjust the image generation model based on a comparison between the first conditional image and the second conditional image.
[0007] In a third aspect of the present disclosure, an electronic device is provided, comprising a processor and a memory coupled to the processor, wherein the memory has instructions stored therein, and when the instructions are executed by the processor, the electronic device executes the method according to the first aspect.
[0008] In a fourth aspect of the present disclosure, a computer program product is provided, wherein a computer-readable storage medium stores computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method according to the first aspect.
[0009] This summary is intended to introduce a selection of concepts in a simplified form that are further described below in the detailed description. It is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0011] Figure 1A A schematic diagram illustrating an example environment in which some embodiments of the present disclosure may be implemented;
[0012] Figure 1B A schematic diagram illustrating another example environment in which some embodiments of the present disclosure may be implemented;
[0013] Figure 1C A schematic diagram illustrating yet another example environment in which some embodiments of the present disclosure may be implemented;
[0014] Figure 2 A flowchart illustrating a method for adjusting an image generation model according to some embodiments of the present disclosure is shown;
[0015] Figure 3 A schematic diagram showing a process for adjusting an image generation model according to some embodiments of the present disclosure is shown;
[0016] Figure 4 Another schematic diagram of a process for adjusting an image generation model according to some embodiments of the present disclosure is shown;
[0017] Figure 5 A block diagram illustrating an apparatus for adjusting an image generation model according to some embodiments of the present disclosure; and
[0018] Figure 6 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.
[0019] Throughout the drawings, the same or similar reference numbers denote the same or similar elements. DETAILED DESCRIPTION
[0020] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0021] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0022] For example, upon receiving a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0023] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0024] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0025] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0026] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. can refer to different or the same objects, unless explicitly stated otherwise. Other explicit and implicit definitions may also be included below.
[0027] Among current image generation methods, some related methods typically generate images based on existing pictures or text. However, these methods have difficulty achieving precise and fine-grained control over images, making it difficult to accurately generate images consistent with the input image conditions and difficult to meet user expectations. For example, in some cases, current image generation methods have difficulty capturing users' specific requirements for details, such as a specific style, color, or the exact location of an object. In some cases, the generated image may not fully match the input conditions in detail because the model loses some details during the generation process.
[0028] To at least address the above and other potential issues, embodiments of the present disclosure provide a method for adjusting an image generation model. The method includes generating a first image based on a first conditional image by the image generation model, where the first conditional image is used to control the generation of the first image. The method also includes generating a second conditional image based on the first image by a conditional extraction model. Furthermore, the method also includes adjusting the image generation model based on a comparison between the first conditional image and the second conditional image. By using this method, the image generation model can achieve precise control over the generated image, reducing the training cost of the image generation model, thereby improving the user experience.
[0029] Figure 1A 1 shows a schematic diagram of an example environment 100A in which some embodiments of the present disclosure may be implemented. Figure 1A As shown, the example environment 100A may include elements such as a conditional image 102, an image generation model 110, and a generated image 104. According to an embodiment of the present disclosure, the image generation model 110 may be deployed in a computing device, and the computing device may be configured as a computing system, a single server, a distributed server, or a cloud-based server, or may be configured as a user terminal, a mobile device, a computer, or a combination of the above devices.
[0030] According to an embodiment of the present disclosure, the conditional image 102 may include an information image associated with real image data. For example, in some embodiments, the conditional image 102 may be an image such as a binary image, a mask image, a sketch image, a depth image, etc. determined based on the real image data. In some embodiments, each pixel value in the binary image can be represented by 0 or 1, thereby characterizing the image pixel as white or black. The binary image can be used to represent the simplified shape or contour information of an object or object in the image, so that the image generation model can depict the boundary of the object or perform object detection. As an example, in the present disclosure, the conditional image 102 may be a mask image associated with a castle building.
[0031] In some embodiments, a mask image can be used to indicate a region of interest (ROI) in an image to define specific areas for processing or analysis, so that the image generation model 110 focuses on these designated areas and ignores other parts. For example, in some embodiments, white areas of the mask image represent regions of interest, while black areas can represent areas to be ignored. In some embodiments, an image can be segmented into regions of interest (foreground) and regions of no interest (background) by setting a threshold. For example, all pixels in the image above the threshold can be identified as foreground, and the remaining pixels are identified as background.
[0032] According to embodiments of the present disclosure, a sketch may include the outline or rough sketch information of an object in an image, such as the shape and main feature lines of the object, but not color or texture details. In some embodiments, the outline of the object in the image may be determined by converting the image into a grayscale image, applying edge detection to identify the outline of the object in the image, and binarizing the result.
[0033] In some examples, the conditional image 102 may also be a depth map, which indicates the distance between each pixel in the image and the observation point, with larger grayscale values in the depth map indicating greater distance. The image generation model 110 may perform three-dimensional reconstruction or virtual reality reconstruction based on the depth map. In some embodiments, the depth information of each pixel may be determined by using two or more cameras to capture the same scene from different angles and then comparing these images (e.g., comparing the pixel differences between the two images).
[0034] Additionally or alternatively, in some embodiments, the conditional image 102 may also be a segmentation map that can assign each pixel in the image to a specific category, thereby segmenting the image into multiple regions or objects, thereby providing detailed scene layout and object shape information to the image generation model 110. For example, in some embodiments, the segmentation map may include a semantic-based segmentation map and an instance-based segmentation map. The instance-based segmentation map can further determine which instance the segmented pixel information comes from based on the semantic segmentation map.
[0035] In some embodiments, conditional image 102 may also be a light map, which can be used to characterize the distribution of lighting in a scene, such as the location, intensity, and area of influence of a light source. For example, image generation model 110 can use the light map to control the lighting conditions in an image, such as for applications such as daylight transitions and shadow generation. In some embodiments, conditional image 102 may also be a normal map, which can characterize the details and texture of an object's surface, thereby helping image generation model 110 generate images with rich textures and details during 3D modeling and augmented reality content creation.
[0036] According to an embodiment of the present disclosure, the image generation model 110 may be a machine learning model that automatically generates images by using machine learning techniques such as deep learning methods, which is capable of learning complex data distributions and generating new images based on the learned information. In some embodiments, the image generation model 110 may generate images based on a random diffusion process.
[0037] For example, during the diffusion process, the image generation model 110 can gradually transform the image into Gaussian noise. Through a series of transformation steps, the image generation model 110 can gradually add noise to the input original image, and ultimately, the image generation model 110 can completely transform the input original image into noise. During the noise addition process, the noisy image at each step can correspond to a distribution. Upon completion of the noise addition process, the original image distribution can be transformed into a complete noise distribution.
[0038] The image generation model 110 can then perform an inverse diffusion operation on the generated noise distribution to restore the original image from the noise. By gradually reducing the noise, the image generation model 110 can generate a new, high-quality image 104. During the training process of the image generation model 110, the image generation model 110 can learn how to reduce noise at each step and gradually restore the details and structure of the image based on a large amount of training data.
[0039] In some embodiments, the image generation model 110 may further extract certain conditions of the generated new, high-quality image 104, such as outlines, depth information, and other conditions, to generate another conditional image 105. According to the methods implemented in the present disclosure, the features or details of the another conditional image 105 may be consistent with or similar to the conditional image 102. Furthermore, the image generation model 110 may compare the conditional image 102 with the another conditional image 105 and adjust parameters of the image generation model 110 based on the comparison results.
[0040] Additionally or alternatively, in some embodiments, the image generation model 110 may further include elements such as an encoder and a decoder. For example, the encoder is configured to encode an input image (or other form of data) into a latent space representation. The image generation model 110 receives the latent space representation input from the encoder and adds noise, and generates a new image based on the noisy latent space representation.
[0041] This allows the image generation model 110 to explore new data representations in the latent space, thereby generating images 104 with new features or styles while maintaining some core properties of the input data. The decoder can evaluate the authenticity of the generated images 104 to ensure the quality and conditional consistency of the generated images, thereby further improving the quality and diversity of the generated images. As an example, the image generation model 110 can generate a castle image 104 based on a mask associated with a castle building.
[0042] Figure 1B 1 is a schematic diagram of another example environment 100B in which some embodiments of the present disclosure may be implemented. Figure 1B As shown, the example environment 100B may include elements such as a conditional image 102 , an image generation model 110 , prompt words 106 , and a generated image 104 .
[0043] According to an embodiment of the present disclosure, prompt word 106 may conflict with conditional image 102. For example, conditional image 102 may be a mask image associated with a castle building, while prompt word 106 may be "delicious cake." The method implemented according to an embodiment of the present disclosure can still effectively generate castle image 108.
[0044] For example, the image generation model 110 can use a reward strategy to learn data that is consistent with the sample or facts, and obtain higher rewards by generating images that are close to the sample. In some embodiments, the reward strategy can also be integrated into the model's training loop to enable the image generation model 110 to prioritize generating images that are consistent with the conditional image. This ensures that even if the prompt word 106 conflicts with the conditional image 102, the desired image can still be generated, ensuring the robustness of the generated image 108.
[0045] According to an embodiment of the present disclosure, the reward score can be determined based on the similarity between the generated image 108 and the sample image associated with the conditional image 102, such as pixel-level differences, structural similarity index (SSIM) and other factors. In some cases, the image generation model 110 can learn to ignore or appropriately handle irrelevant or conflicting information through a reward strategy. For example, when the image generation model 110 generates irrelevant or stylistically different images based on conflicting prompt words 106, the image generation model 110 will receive a lower reward score or penalty. In this way, even if the prompt word 106 is cake, the image 108 generated by the image generation model 110 will not include element details associated with the cake.
[0046] In some embodiments, the image generation model 110 may further extract certain conditions of the generated new, high-quality image 108, such as outlines, depth information, and other conditions, to generate another conditional image 107. According to the method implemented in the present disclosure, the features or details of the another conditional image 107 may be consistent with or similar to the conditional image 102. In addition, the image generation model 110 may also compare the conditional image 102 with the another conditional image 107 and adjust the parameters of the image generation model 110 based on the comparison results.
[0047] Figure 1C 1 is a schematic diagram illustrating another example environment 100C in which some embodiments of the present disclosure may be implemented. Figure 1C As shown, the example environment 100C may include elements such as a conditional image 102 , an image generation model 110 , prompt words 112 , and a generated image 114 .
[0048] According to an embodiment of the present disclosure, prompt word 112 can be a prompt word that further refines conditional image 102. For example, conditional image 102 can be a mask image associated with a castle building, while prompt word 106 can be "house, high quality, extreme detail, 4K." The method implemented according to an embodiment of the present disclosure can effectively generate castle image 108.
[0049] Specifically, prompt word 106 can provide basic subject or category information, while prompt word 112 can provide further detailed requirements regarding image quality, detail level, and resolution (e.g., "high quality, extreme detail, 4k"). Image generation model 110 can comprehensively consider these factors based on a reward strategy to generate an image that not only conforms to mask outline 112 but also meets the detail and quality requirements. In this way, input information from different modalities can be integrated to generate the desired image.
[0050] In some embodiments, the image generation model 110 may further extract some conditions of the generated new, high-quality image 114, such as the outline, depth information, etc., to generate another conditional image 109. According to the method implemented in the present disclosure, the features or details of the another conditional image 109 may be consistent with or similar to the conditional image 102. In addition, the image generation model 110 may also compare the conditional image 102 with the another conditional image 109 and adjust the parameters of the image generation model 110 based on the comparison results.
[0051] It should be understood that the architecture and functions of the example environments 100A to 100C are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure. The embodiments of the present disclosure may also be applied to other environments with different structures and / or functions.
[0052] The following will be combined Figures 2 to 6 The process according to the embodiment of the present disclosure is described in detail. For ease of understanding, the specific data mentioned in the following description are exemplary and are not intended to limit the scope of protection of the present disclosure. It is understood that the embodiments described below may also include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this respect.
[0053] Figure 2 A flowchart of a method 200 for adjusting an image generation model according to some embodiments of the present disclosure is shown. At block 202, the image generation model generates a first image based on a first conditional image, where the first conditional image is used to control the generation of the first image. According to embodiments of the present disclosure, the image generation model may be a deep learning model based on a diffusion method and a generative adversarial network.
[0054] The conditional image can be used to control the generated image. For example, in some embodiments, the conditional image can be a binary image, depth map, or other image of the user-input image. In some embodiments, the conditional image can also be a handwritten sketch or drawing input by the user, which includes the general outline and composition of the image the user wants to generate. The image generation model can fill in the details of the image to be generated based on the guidance of these sketches or drawings, thereby generating a first image that matches the user's intention.
[0055] At block 204, a second conditional image is generated by the conditional extraction model based on the first image. According to an embodiment of the present disclosure, the conditional extraction model may be a reward model based on reward score allocation and used to control the allocation of reward scores to the image generation model. In some embodiments, the conditional extraction model may be part of or a component of the image generation model.
[0056] In some embodiments, the conditional extraction model can be based on one or more image recognition models selected from a holistic nested edge detection (HED) model, a depth model, a Canny model, and / or a segmentation model. The conditional extraction model can identify key features, patterns, textures, or other useful information in the first image and generate another conditional image based on the identified information. In some embodiments, the conditional extraction model can be adjusted or trained based on these multiple types of conditional images.
[0057] For example, in some embodiments, when the conditional extraction model is a HED-based model, it can perform edge detection on the image by learning multi-scale and multi-level features of the image, thereby identifying fine edges and contours in the image and generating a second conditional image including a contour image. In some embodiments, when the conditional extraction model is a depth model, it can determine the depth information of each pixel in the image, thereby generating a second conditional image including a depth map that identifies scene depth information. The depth map can provide information about spatial relationships and object sizes in the image.
[0058] According to an embodiment of the present disclosure, the conditional extraction model can also be a Canny-based model, which can also perform edge detection on the image, thereby identifying elements in the image such as hair, patterns on clothing, details of background trees, textures of animals, etc. These element details help preserve the structural features of the original image when generating another conditional image. In some embodiments, the conditional extraction model can also be a model based on a segmentation method, which can segment the image into multiple parts or regions, each of which can represent different objects or image features, thereby allowing for precise control of the details of each part of the image when generating another conditional image.
[0059] At block 206, the image generation model is adjusted based on the comparison between the first conditional image and the second conditional image. According to an embodiment of the present disclosure, a decoder or discriminator in the conditional extraction model can compare the generated second conditional image with the first conditional image input by the user and assign a reward score. For example, in some embodiments, the conditional extraction model can extract first conditional image features associated with the first conditional image and second conditional image features associated with the second conditional image. These features can include elements such as edge information, texture patterns, color distribution, object shape, etc. that can characterize the content and style of the image.
[0060] The conditional extraction model can then compare the first conditional image feature with the second conditional image feature. In some embodiments, the similarity of the image features can be based on methods such as Euclidean distance, cosine similarity, Manhattan distance, structural similarity index, feature matching, etc.
[0061] In some embodiments, if the difference between the first conditional image feature and the second conditional image feature is lower than a predetermined similarity condition, this indicates that the generated second conditional image is closer in visual features to the first conditional image and maintains a high degree of consistency, and the conditional extraction model can assign a high reward score or the first reward score to the image generation model. In some embodiments, if the difference between the first conditional image feature and the second conditional image feature is higher than a predetermined similarity condition, this indicates that the difference between the two is large, and the conditional extraction model can assign a low reward score or the second reward score to the image generation model.
[0062] Figure 3 FIG. 3 is a schematic diagram of a process 300 for training an image generation model according to some embodiments of the present disclosure. Figure 3 As shown, the conditional image C includes one or more image data such as a binary image, a mask image, a sketch image, a depth image, and a grayscale image. v 302 can be fed into an image generation model 310 implemented based on the present disclosure. As an example, in the present disclosure, the conditional image C v 302 may be a conditional image associated with a “heart-shaped” image, such as a binary image, a contour image, a depth image, etc. of the heart shape.
[0063] Additionally or alternatively, in some embodiments, prompt words C such as "heart shape, mountain, natural image" t It can also be combined with the conditional image C v 302 is fed into the image generation model 310 to assist in generating a new image x'0 304. The new image x'0 304 can be based on the prompt word C t and conditional image C v 302 and the recreated image, which may include the prompt word C t and conditional image C v One or more elements in 302 may also include other unknown elements. In this way, the image generation model 310 can convert an image from one domain to another domain, for example, from a conditional image C v 302 Transition to the newly generated image x'0 304.
[0064] The conditional extraction model 312 can then extract features from the newly generated image x'0 304 to generate another conditional image C' v 306. According to an embodiment of the present disclosure, the conditional extraction model 312 can be a pre-trained image processing reward model. For example, in some embodiments, the conditional extraction model 312 can be one or more models or a combination of a Hed model for image edge processing, a depth model for processing image depth, a Canny model, and a segmentation model for image segmentation. It should be understood that the types of conditional extraction models 312 listed here are only for example, and the conditional extraction model 312 also includes other known or unknown image processing models, which are not limited by the present disclosure, such as an illumination map image processing model, a normal image processing model, etc.
[0065] In some embodiments, the conditional extraction model 312 can be based on the input conditional image C v 302 type to generate another conditional image C' of the same type v306. For example, in some embodiments, if the condition image C v 302 is a depth map, the conditional extraction model 312 can extract the depth information in the newly generated image x'0 304 to generate another conditional image C' of the depth map type. v 306.
[0066] In some embodiments, if the conditional image C v 302 is a binary image, the conditional extraction model 312 can extract the binary information in the newly generated image x'0 304 to generate another conditional image C' of the binary image type. v 306. In some embodiments, if the conditional image C v 302 is a mask image, the conditional extraction model 312 can extract the mask information in the newly generated image x'0 304 to generate another conditional image C' of the mask image type. v 306. Additionally or alternatively, in some embodiments, if the condition image C v 302 is a sketch, the conditional extraction model 312 can extract the sketch contour information in the newly generated image x'0 304 to generate another conditional image C' of the sketch type. v 306.
[0067] According to an embodiment of the present disclosure, the condition extraction model 312 may then convert the condition image C v 302 features and condition image C' v 306, and assigns a reward value to the image generation model 310 according to the difference in the comparison, and adjusts the image generation model 310 based on the reward value. v When 302 is a mask image, the condition extraction model 312 can be used to extract the condition image C' v The mask data in 306 is segmented to determine the consistency between the two. In some embodiments, the mask features may include information such as the edge, outline, shape, etc. of the object in the image.
[0068] For example, in some embodiments, when the conditional image C v Mask features of 302 and conditional image C' v 306 mask features are consistent or similar (for example, below a predetermined similarity condition or C v =C' v ), the conditional extraction model 312 may assign a higher reward value to the image generation model 310. For example, the conditional extraction model 312 may compare the mask value of each pixel of the two images to quantify the difference between them.
[0069] In some embodiments, when the conditional image C v Mask features of 302 and conditional image C' v When the mask features of 306 differ greatly, the conditional extraction model 312 can assign a lower reward value to the image generation model 310. In this way, the cycle consistency loss can be expressed as the conditional image C v Mask features of 302 and conditional image C' v The pixel-by-pixel loss between the mask features of 306. By adjusting the assigned reward value, the image generation model 310 can select a specific image generation path or action with a high reward value and update the corresponding model parameters, and the image generation model 310 can be adjusted through cyclic iterative training.
[0070] The method implemented according to the present disclosure can achieve better performance in pixel-space display optimization. The method implemented according to the present disclosure can also make the image generation process more focused on maintaining consistency with the original condition image, achieve better performance in preserving key visual features and details, and ultimately improve the quality of the generated image 304.
[0071] The methods implemented by this disclosure can also avoid the significant costs associated with image sampling, thereby allowing for more efficient reward fine-tuning. Multiple sampling in some methods can lead to efficiency issues and require storing gradients at each time step, consuming significant time and GPU and memory resources. The methods implemented according to this disclosure can prove that starting with random Gaussian noise sampling is unnecessary.
[0072] Figure 4 FIG. 4 is a schematic diagram of another process 400 for training an image generation model according to some embodiments of the present disclosure. Figure 4 As shown, in some embodiments, sample data 402, such as those based on real images, may be fed into an image generation model 410. An encoder 404 in the image generation model 410 may extract and encode the sample data 402 and generate extracted latent features x0 406, thereby converting the image data into feature data.
[0073] According to an embodiment of the present disclosure, the latent features 406 may then be subjected to a noise addition process by the image generation model 410 to obtain the noisy latent features x t 408. Noised latent feature x t 408 can then be diffused by the diffusion model 420 input into the image generation model 410 to obtain the diffused feature x ’ 0 412.
[0074] In some embodiments, partial Gaussian noise can be gradually added to the feature data, which can be represented as a Markov chain and can eventually transition to a completely random noise state. In some embodiments, the diffused feature x ’ 0 412 can be denoised and decoded by the decoder 414 in the image generation model 410 to obtain another sample image 422 by converting the processed latent features back to the image space.
[0075] In some embodiments, the conditional extraction model may then extract a sample-conditional image feature 416 associated with the sample data 402 and another sample-conditional image feature 418 associated with another sample image 422. According to an embodiment of the present disclosure, the conditional extraction model may compare the sample-conditional image feature 416 with the other sample-conditional image feature 418 using a pixel-space cycle consistency loss. For example, the conditional extraction model may compare the similarity of the two sets of image features to determine pixel-level consistency.
[0076] In some embodiments, the conditional extraction model can also assign a reward value to the image generation model 410 based on the comparison. For example, when the difference between the two is less than a predetermined similarity condition, a higher reward value is assigned to the image generation model 410; and when the difference between the two is greater than the predetermined similarity condition, a lower reward value is assigned to the image generation model 410. The method implemented according to the present disclosure can perform more effective reward fine-tuning by directly adding noise to the training images to disturb their consistency with the input conditional control, and then using a single-step denoised image to restore consistency.
[0077] Figure 5 FIG. 5 is a block diagram of an apparatus 500 for adjusting an image generation model according to some embodiments of the present disclosure. Figure 5 As shown, apparatus 500 includes an image generation module 502 configured to generate a first image based on a first conditional image by an image generation model, wherein the first conditional image is used to control the generation of the first image. Apparatus 500 also includes a conditional image generation module 504 configured to generate a second conditional image based on the first image by a conditional extraction model. Apparatus 500 also includes a comparison module 506 configured to adjust the image generation model based on a comparison between the first conditional image and the second conditional image.
[0078] Figure 6 FIG1 shows a block diagram of an electronic device 600 according to some embodiments of the present disclosure. The device 600 may be a device or apparatus described in the embodiments of the present disclosure. Figure 6As shown, the device 600 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 601, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 602 or computer program instructions loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The CPU / GPU 601, the ROM 602, and the RAM 603 are connected to each other via a bus 608. An input / output (I / O) interface 605 is also connected to the bus 604. Although not shown in FIG. Figure 6 As shown in FIG, device 600 may further include a co-processor.
[0079] Various components in device 600 are connected to I / O interface 605, including: an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0080] The various methods or processes described above may be performed by the CPU / GPU 601. For example, in some embodiments, the methods may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the CPU / GPU 601, one or more steps or actions in the methods or processes described above may be performed.
[0081] In some embodiments, the methods and processes described above may be implemented as a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present disclosure.
[0082] Computer-readable storage medium can be a tangible device that can keep and store the instructions used by the instruction execution device.Computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device or any suitable combination thereof.More specific examples (non-exhaustive list) of computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, for example, a punch card or a convex structure in a groove having instructions stored thereon, and any suitable combination thereof.Computer-readable storage medium used herein is not interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated by waveguides or other transmission media (for example, light pulses by fiber optic cables), or electrical signals transmitted by wires.
[0083] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0084] The computer program instructions for performing the disclosed operation can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data or source code or the object code written in any combination of one or more programming languages, programming languages include object-oriented programming languages, and conventional procedural programming languages.Computer-readable program instructions can be performed completely on a user's computer, partially on a user's computer, performed as an independent software package, partly on a user's computer and partly on a remote computer, or performed completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer by any type of network-including local area network (LAN) or wide area network (WAN), or can be connected to an external computer (such as utilizing an internet service provider to connect by the internet). In certain embodiments, by utilizing the state information of computer-readable program instructions to carry out personalized customization electronic circuits, such as programmable logic circuits, field programmable gate arrays (FPGAs) or programmable logic arrays (PLA), this electronic circuit can perform computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0085] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0086] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0087] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart, can be implemented by a special hardware-based system that performs the prescribed function or action, or can be implemented by a combination of special hardware and computer instructions.
[0088] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, practical applications, or technical improvements to the technology in the market, or to enable other persons skilled in the art to understand the embodiments disclosed herein.
[0089] Some example implementations of the present disclosure are listed below.
[0090] 1. A method for adjusting an image generation model, the method comprising:
[0091] generating a first image based on a first conditional image by an image generation model, wherein the first conditional image is used to control generation of the first image;
[0092] generating, by a conditional extraction model, a second conditional image based on the first image; and
[0093] The image generation model is adjusted based on a comparison between the first condition image and the second condition image.
[0094] Example 2. The method of Example 1, wherein the conditional extraction model is a reward model based on reward score assignment and is used to control assignment of reward scores to the image generation model.
[0095] Example 3. The method of any of Examples 1-2, wherein adjusting the image generation model based on the comparison between the first condition image and the second condition image comprises:
[0096] extracting, by the condition extraction model, a first condition image feature associated with the first condition image and a second condition image feature associated with the second condition image;
[0097] comparing the first conditional image feature with the second conditional image feature;
[0098] In response to a difference between the first conditional image feature and the second conditional image feature being below a predetermined similarity condition, assigning, by the conditional extraction model, a first reward score to the image generation model; and
[0099] In response to a difference between the first conditional image feature and the second conditional image feature being higher than a predetermined similarity condition, the conditional extraction model assigns a second reward score to the image generation model, the first reward score being greater than the second reward score.
[0100] Example 4. The method of any one of Examples 1-3, wherein the first conditional image is based on one type of conditional image among a plurality of types of conditional images; and
[0101] The image generation model and the conditional extraction model are adjusted based on the multiple types of conditional images.
[0102] Example 5. The method of any one of Examples 1-4, further comprising:
[0103] In response to the first conditional image being a depth map, the conditional extraction model extracts depth information from the first image to generate the second conditional image.
[0104] Example 6. The method of any one of Examples 1-5, further comprising:
[0105] In response to the first conditional image being a binary image, the conditional extraction model extracts binary information from the first image to generate the second conditional image.
[0106] Example 7. The method of any one of Examples 1-6, further comprising:
[0107] In response to the first condition image being a mask image, the condition extraction model extracts mask information from the first image to generate the second condition image.
[0108] Example 8. The method of any one of Examples 1-7, further comprising:
[0109] In response to the first condition image being a sketch, the condition extraction model extracts sketch contour information from the first image to generate the second condition image.
[0110] Example 9. The method of any one of Examples 1-8, wherein generating, by an image generation model, the first image based on the first conditional image further comprises:
[0111] extracting the first conditional image features of the first conditional image by an encoder of the image generation model and adding noise to the extracted first conditional image features; and
[0112] The noisy extracted first conditional image features are decoded by a decoder of the image generation model to generate the first image.
[0113] Example 10. The method of any one of Examples 1-9, wherein generating, by an image generation model, the first image based on the first conditional image further comprises:
[0114] The first image is generated by an image generation model based on the first conditional image and text information.
[0115] Example 11. The method of any one of Examples 1-10, further comprising:
[0116] generating, by the image generation model, a second sample image based on the first sample image, wherein the first sample image is based on data of a real image; and
[0117] A first sample condition image feature associated with the first sample image and a second sample condition image feature associated with the second sample image are extracted respectively.
[0118] Example 12. The method of any one of Examples 1-11, further comprising:
[0119] comparing the first sample condition image feature with the second sample condition image feature; and
[0120] The image generation model is adjusted by minimizing the difference between the first sample-condition image feature and the second sample-condition image feature.
[0121] Example 13. The method of any of Examples 1-12, wherein the conditional extraction model is a reward model based on one or more of:
[0122] Holistic Nested Edge Detection (HED) model;
[0123] Deep Model;
[0124] Canny model; or
[0125] Segmentation model.
[0126] Example 14. An apparatus for adjusting an image generation model, comprising:
[0127] an image generation module, configured to generate a first image based on a first condition image by an image generation model, wherein the first condition image is used to control the generation of the first image;
[0128] a conditional image generating module configured to generate a second conditional image based on the first image by a conditional extraction model; and
[0129] A comparison module is configured to adjust the image generation model based on a comparison between the first condition image and the second condition image.
[0130] Example 15. An apparatus according to Example 14, wherein the conditional extraction model is a reward model based on reward score allocation and is used to control the allocation of reward scores to the image generation model.
[0131] Example 16. The apparatus of any of Examples 14-15, wherein the conditional extraction model comprises:
[0132] an extraction module configured to extract a first conditional image feature associated with the first conditional image and a second conditional image feature associated with the second conditional image;
[0133] a comparison module configured to compare the first conditional image feature with the second conditional image feature;
[0134] a first reward score assigning module configured to assign, by the conditional extraction model, a first reward score to the image generation model in response to a difference between the first conditional image feature and the second conditional image feature being lower than a predetermined similarity condition; and
[0135] A second reward score allocation module is configured to allocate a second reward score from the conditional extraction model to the image generation model in response to a difference between the first conditional image feature and the second conditional image feature being higher than a predetermined similarity condition, wherein the first reward score is greater than the second reward score.
[0136] Example 17. An apparatus according to any of Examples 14-16, wherein the first conditional image is based on a type of conditional image among a plurality of types of conditional images; and
[0137] The image generation model and the conditional extraction model are adjusted based on the multiple types of conditional images.
[0138] Example 18. The apparatus of any of Examples 14-17, wherein the conditional extraction model further comprises:
[0139] The depth map generating module is configured to extract depth information from the first image to generate the second condition image in response to the first condition image being a depth map.
[0140] Example 19. The apparatus of any one of Examples 14-18, wherein the conditional extraction model further comprises:
[0141] The binary image generating module is configured to extract binary information from the first image to generate the second condition image in response to the first condition image being a binary image.
[0142] Example 20. The apparatus of any one of Examples 14-19, wherein the conditional extraction model further comprises:
[0143] The mask image generating module is configured to extract mask information from the first image to generate the second condition image in response to the first condition image being a mask image.
[0144] Example 21. The apparatus of any one of Examples 14-20, wherein the conditional extraction model further comprises:
[0145] The sketch image generating module is configured to extract sketch contour information from the first image to generate the second condition image in response to the first condition image being a sketch image.
[0146] Example 22. The apparatus of any of Examples 14-21, wherein the image generation model further comprises:
[0147] an encoder module configured to extract the first conditional image features of the first conditional image and add noise to the extracted first conditional image features; and
[0148] A decoder module is configured to decode the noisy extracted first conditional image features by a decoder of the image generation model to generate the first image.
[0149] Example 23. The apparatus of any of Examples 14-22, wherein the image generation model further comprises:
[0150] The generating module is configured to generate the first image based on the first condition image and text information.
[0151] Example 24. The apparatus of any of Examples 14-23, wherein the image generation model further comprises:
[0152] a sample image generating module configured to generate a second sample image based on a first sample image, wherein the first sample image is based on data of a real image; and
[0153] The feature extraction module is configured to extract a first sample condition image feature associated with the first sample image and a second sample condition image feature associated with the second sample image respectively.
[0154] Example 25. The apparatus of any of Examples 14-24, wherein the image generation model further comprises:
[0155] a sample feature comparison module configured to compare the first sample condition image feature with the second sample condition image feature; and
[0156] An adjustment module is configured to adjust the image generation model by minimizing the difference between the first sample condition image feature and the second sample condition image feature.
[0157] Example 26. The apparatus of any of Examples 14-25, wherein the conditional extraction model is a reward model based on one or more of:
[0158] Holistic Nested Edge Detection (HED) model;
[0159] Deep Model;
[0160] Canny model; or
[0161] Segmentation model.
[0162] 27. An electronic device comprising:
[0163] processor; and
[0164] A memory coupled to the processor, the memory having instructions stored therein, wherein when the instructions are executed by the processor, the electronic device performs actions, the actions comprising:
[0165] generating a first image based on a first conditional image by an image generation model, wherein the first conditional image is used to control generation of the first image;
[0166] generating, by a conditional extraction model, a second conditional image based on the first image; and
[0167] The image generation model is adjusted based on a comparison between the first condition image and the second condition image.
[0168] Example 28. An electronic device according to Example 27, wherein the conditional extraction model is a reward model based on reward score allocation and is used to control the allocation of reward scores to the image generation model.
[0169] Example 29. The electronic device of any of Examples 27-28, wherein adjusting the image generation model based on the comparison between the first condition image and the second condition image comprises:
[0170] extracting, by the condition extraction model, a first condition image feature associated with the first condition image and a second condition image feature associated with the second condition image;
[0171] comparing the first conditional image feature with the second conditional image feature;
[0172] In response to a difference between the first conditional image feature and the second conditional image feature being below a predetermined similarity condition, assigning, by the conditional extraction model, a first reward score to the image generation model; and
[0173] In response to a difference between the first conditional image feature and the second conditional image feature being higher than a predetermined similarity condition, the conditional extraction model assigns a second reward score to the image generation model, the first reward score being greater than the second reward score.
[0174] Example 30. An electronic device according to any of Examples 27-29, wherein the first conditional image is based on a type of conditional image among a plurality of types of conditional images; and
[0175] The image generation model and the conditional extraction model are adjusted based on the multiple types of conditional images.
[0176] Example 31. The electronic device of any of Examples 27-30, further comprising:
[0177] In response to the first conditional image being a depth map, the conditional extraction model extracts depth information from the first image to generate the second conditional image.
[0178] Example 32: The electronic device of any of Examples 27-31, further comprising:
[0179] In response to the first conditional image being a binary image, the conditional extraction model extracts binary information from the first image to generate the second conditional image.
[0180] Example 33. The electronic device of any of Examples 27-32, further comprising:
[0181] In response to the first condition image being a mask image, the condition extraction model extracts mask information from the first image to generate the second condition image.
[0182] Example 34. The electronic device of any of Examples 27-33, further comprising:
[0183] In response to the first condition image being a sketch, the condition extraction model extracts sketch contour information from the first image to generate the second condition image.
[0184] Example 35. The electronic device of any of Examples 27-34, wherein generating, by the image generation model, the first image based on the first conditional image further comprises:
[0185] extracting the first conditional image features of the first conditional image by an encoder of the image generation model and adding noise to the extracted first conditional image features; and
[0186] The noisy extracted first conditional image features are decoded by a decoder of the image generation model to generate the first image.
[0187] Example 36. The electronic device of any of Examples 27-35, wherein generating, by the image generation model, the first image based on the first conditional image further comprises:
[0188] The first image is generated by an image generation model based on the first conditional image and text information.
[0189] Example 37. The electronic device of any of Examples 27-36, further comprising:
[0190] generating, by the image generation model, a second sample image based on the first sample image, wherein the first sample image is based on data of a real image; and
[0191] A first sample condition image feature associated with the first sample image and a second sample condition image feature associated with the second sample image are extracted respectively.
[0192] Example 38. The electronic device of any of Examples 27-37, further comprising:
[0193] comparing the first sample condition image feature with the second sample condition image feature; and
[0194] The image generation model is adjusted by minimizing the difference between the first sample-condition image feature and the second sample-condition image feature.
[0195] Example 39. The electronic device of any of Examples 27-38, wherein the conditional extraction model is a reward model based on one or more of:
[0196] Holistic Nested Edge Detection (HED) model;
[0197] Deep Model;
[0198] Canny model; or
[0199] Segmentation model.
[0200] Example 40. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method according to any one of Examples 1 to 13.
[0201] Example 41. A computer program product tangibly stored on a computer-readable medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method of any one of Examples 1 to 13.
[0202] Although the present disclosure has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for adjusting an image generation model, the method comprising: generating a first image based on a first conditional image by an image generation model, wherein the first conditional image is used to control generation of the first image; generating a second conditional image based on the first image by a conditional extraction model; as well as The image generation model is adjusted based on a comparison between the first condition image and the second condition image. 2 . The method according to claim 1 , wherein the conditional extraction model is a reward model based on reward score allocation and is used to control allocation of reward scores to the image generation model.
3. The method of claim 1 , wherein adjusting the image generation model based on the comparison between the first condition image and the second condition image comprises: extracting, by the condition extraction model, a first condition image feature associated with the first condition image and a second condition image feature associated with the second condition image; comparing the first conditional image feature with the second conditional image feature; In response to a difference between the first conditional image feature and the second conditional image feature being below a predetermined similarity condition, assigning, by the conditional extraction model, a first reward score to the image generation model; as well as In response to a difference between the first conditional image feature and the second conditional image feature being higher than a predetermined similarity condition, the conditional extraction model assigns a second reward score to the image generation model, the first reward score being greater than the second reward score.
4. The method of claim 1 , wherein the first conditional image is based on one type of conditional image among a plurality of types of conditional images; and The image generation model and the conditional extraction model are adjusted based on the multiple types of conditional images.
5. The method according to claim 4, further comprising: In response to the first conditional image being a depth map, the conditional extraction model extracts depth information from the first image to generate the second conditional image.
6. The method according to claim 4, further comprising: In response to the first conditional image being a binary image, the conditional extraction model extracts binary information from the first image to generate the second conditional image.
7. The method according to claim 4, further comprising: In response to the first condition image being a mask image, the condition extraction model extracts mask information from the first image to generate the second condition image.
8. The method according to claim 4, further comprising: In response to the first condition image being a sketch, the condition extraction model extracts sketch contour information from the first image to generate the second condition image.
9. The method according to claim 3, wherein generating the first image based on the first conditional image by an image generation model further comprises: extracting the first conditional image features of the first conditional image by an encoder of the image generation model and adding noise to the extracted first conditional image features; as well as The noisy extracted first conditional image features are decoded by a decoder of the image generation model to generate the first image.
10. The method according to claim 9, wherein generating the first image based on the first conditional image by an image generation model further comprises: The first image is generated by an image generation model based on the first conditional image and text information.
11. The method according to claim 1 , further comprising: generating a second sample image based on the first sample image by the image generation model, wherein the first sample image is based on data of a real image; as well as A first sample condition image feature associated with the first sample image and a second sample condition image feature associated with the second sample image are extracted respectively.
12. The method according to claim 11, further comprising: comparing the first sample condition image feature with the second sample condition image feature; as well as The image generation model is adjusted by minimizing the difference between the first sample-condition image feature and the second sample-condition image feature.
13. The method of claim 3, wherein the conditional extraction model is a reward model based on one or more of the following: Holistic Nested Edge Detection (HED) model; Deep Model; Canny model; or Segmentation model.
14. A device for adjusting an image generation model, comprising: an image generation module, configured to generate a first image based on a first condition image by an image generation model, wherein the first condition image is used to control the generation of the first image; a conditional image generating module, configured to generate a second conditional image based on the first image by a conditional extraction model; as well as A comparison module is configured to adjust the image generation model based on a comparison between the first condition image and the second condition image.
15. An electronic device comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, wherein when the instructions are executed by the processor, the electronic device performs the method according to any one of claims 1 to 13.
16. A computer program product comprising computer executable instructions, wherein the computer executable instructions are executed by a processor to implement the method according to any one of claims 1 to 13.