Image generation method and device, electronic equipment and storage medium
Through multiple iterative noise reduction and area guidance methods, the problems of regional incoherence and resource consumption in image generation are solved, and high-quality image generation with harmonious visual effects are achieved.
Patent Information
- Application Number
- CN202410008339.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-02
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, when generating images, independent noise reduction in each area leads to incoherent image content, dissonant visual effects, and difficult area division, resource consumption and poor generation quality.
By obtaining the reference data and mask images of multiple target objects in the target image, performing multiple iterative noise reduction, and using multiple prompt text and mask images for different areas in the initial noise map for noise reduction to ensure that the overlapping content is consistent, realizing noise reduction in sub-region and guiding unnoised content.
The generated image visual effects are harmonious and coherent, improving the quality of image generation, ensuring differential processing of contents in each area and high credibility in noise reduction results.
Smart Images

Figure CN120259117A_ABST
Abstract
Description
Background Art
[0002] Target image. Currently, an image mask constructed for the whole of each target object expected to be included in the target image and text guidance words corresponding to each target object are usually used to perform noise reduction guidance on different regions divided in the noise map to obtain a target image including multiple target objects; the image mask is the same size as the noise map and is used to describe the overall shape of each target object in the target image.
[0003] For example, using the multidiffusion algorithm, in each step of noise reduction, the following operations are performed: dividing the current noise map into several regions, performing regional noise reduction on the current noise map under the guidance of the text content, and realizing the consistency processing of the overall shape of the content with the help of the mask content in the corresponding region of the image mask to obtain a denoised image after noise reduction.
[0004] However, since the noise reduction for different regions is independent, and the image mask can only perform shape constraints globally, the image content of each region in the finally generated target image may be incoherent and very inharmonious in visual effect; in addition, there are great difficulties in dividing the regions of the noise map. If the regions are divided too small, the noise reduction process will consume a large amount of time and resources, and if the regions are divided too large, affected by the incoherence of the image content, the generated image quality is very poor. Summary of the Invention
[0005] The embodiments of the present application provide an image generation method, device, electronic device and storage medium to improve the generation quality of images.
[0006] In a first aspect, an image generation method is proposed, including:
[0007] Obtaining multiple pieces of reference data configured for multiple target objects included in the target image, where one piece of reference data includes: a prompt text for describing one target object, and a mask image for indicating the shape and position of the one target object in the target image;
[0008] Obtaining an initial noise map;
[0009] Performing multiple rounds of iterative noise reduction on the initial noise map, and using the denoised noise map obtained in the last round of noise reduction as the target image, where in one round of iterative noise reduction, the following operations are performed:
[0010] For a first region in the initial noise image, performing noise reduction guidance using multiple prompt texts respectively to obtain multiple initial denoised results of the first region, and obtaining a final denoised result of the first region according to multiple mask images and the multiple initial denoised results of the first region;
[0011] For the second region in the initial noise image, perform noise reduction guidance using the multiple pieces of prompt text respectively to obtain multiple initial denoising results for the second region, and obtain the final denoising result for the second region according to the multiple mask images and the multiple initial denoising results for the second region, where there is an overlapping part between the first region and the second region, and the content of the overlapping part in the second region is the same as the content of the overlapping part in the final denoising result of the first region;
[0012] Determine a denoised noise map containing the multiple target objects according to the final denoising result of the first region and the part of the final denoising result of the second region that does not overlap with the first region.
[0013] In a second aspect, an image generation device is proposed, including:
[0014] A first acquisition unit, configured to acquire multiple pieces of reference data configured for multiple target objects included in a target image, where one piece of reference data includes: a prompt text for describing one target object, and a mask image for indicating the shape and position of the one target object in the target image;
[0015] A second acquisition unit, configured to acquire an initial noise map;
[0016] A noise reduction unit, configured to perform multiple rounds of iterative noise reduction on the initial noise map, and use the denoised noise map obtained in the last round of noise reduction as the target image, where in a round of iterative noise reduction process, the following operations are performed:
[0017] For the first region in the initial noise image, perform noise reduction guidance using multiple pieces of prompt text respectively to obtain multiple initial denoising results for the first region, and obtain the final denoising result for the first region according to the multiple mask images and the multiple initial denoising results for the first region;
[0018] For the second region in the initial noise image, perform noise reduction guidance using the multiple pieces of prompt text respectively to obtain multiple initial denoising results for the second region, and obtain the final denoising result for the second region according to the multiple mask images and the multiple initial denoising results for the second region, where there is an overlapping part between the first region and the second region, and the content of the overlapping part in the second region is the same as the content of the overlapping part in the final denoising result of the first region;
[0019] Determine a denoised noise map containing the multiple target objects according to the final denoising result of the first region and the part of the final denoising result of the second region that does not overlap with the first region.
[0020] Optionally, after using the denoised noise map obtained from the last round of noise reduction as the target image, the noise reduction unit is further configured to:
[0021] Perform multiple rounds of iterative noise reduction on the target image, and use the denoised image obtained from the last round of noise reduction as the final image. During one round of iteration, the following operations are performed:
[0022] For the target image, perform noise reduction guidance using the multiple prompt texts respectively to obtain multiple intermediate images of the target image, and obtain the denoised image of the target image according to the multiple mask images and the multiple intermediate images of the target image.
[0023] Optionally, after using the denoised image obtained from the last round of noise reduction as the final image, the noise reduction unit is further configured to:
[0024] Extract corresponding comprehensive text features based on the multiple prompt texts, and extract corresponding image features based on the final image;
[0025] Calculate the feature similarity between the comprehensive text features and the image features, and when the feature similarity reaches a set threshold, stop the current image generation process and output the final image.
[0026] Optionally, after calculating the feature similarity between the comprehensive text features and the image features, the noise reduction unit is further configured to:
[0027] When the feature similarity does not reach the set threshold, re-obtain the initial noise map and re-obtain multiple pieces of reference data configured for the target image;
[0028] Based on the re-obtained initial noise map and the multiple pieces of reference data, re-execute the image generation process.
[0029] Optionally, the initial noise map is generated by the second acquisition unit in the following manner:
[0030] Obtain an original noise map;
[0031] Obtain multiple pieces of reference data configured for multiple target objects included in the target image. One piece of reference data includes: a prompt text for describing one target object, and a mask image for indicating the shape and position of the one target object in the target image;
[0032] Perform multiple rounds of iterative noise reduction on the original noise map, and use the denoised noise map obtained from the last round of noise reduction as the initial noise map. During one round of iteration, the following operations are performed:
[0033] Based on the original noise map and multiple mask images, a single-object noise map for each target object is obtained, and for the single-object noise map of each target object, the following operations are performed: Using the prompt text of the target object, noise reduction guidance is performed on the single-object noise map of the target object to obtain a single-object noise-reduced map;
[0034] The contents of multiple single-object noise-reduced maps are superimposed to obtain a noise-reduced map of the original noise map.
[0035] Optionally, the multiple pieces of reference data are generated by the first acquisition unit in the following manner:
[0036] Determine multiple target objects expected to be generated in the target image, and for each target object, perform the following operations:
[0037] Based on the prompt words of the target object, generate prompt text, and in a pre-constructed object mask set, obtain the target object mask corresponding to the target object;
[0038] According to the preset resolution for the target image, the object mask of the target object, and the expected size and expected position of the target object in the target image, obtain the mask image of the target object, and use the prompt text and mask image of the target object as a piece of reference data.
[0039] Optionally, before obtaining the original noise map, the second acquisition unit is further configured to:
[0040] Obtain the preset total number of noise reduction rounds;
[0041] According to the total number of noise reduction rounds and the preset round ratio, configure the first number of noise reduction rounds used to denoise the original noise map to obtain the initial noise map, the second number of noise reduction rounds used to denoise the initial noise map to obtain the target noise map, and the third number of noise reduction rounds used to denoise the target image to obtain the final image.
[0042] Optionally, the regional positions of the first region and the second region in the initial noise map are obtained by the noise reduction unit in the following manner:
[0043] Obtain the preset regional window, sliding direction, sliding step, and the initial position in the initial noise map;
[0044] When the regional window is at the initial position, determine the regional position covered by the regional window in the initial noise map as the regional position corresponding to the first region;
[0045] In the initial noise map, after sliding the sliding step length along the sliding direction starting from the initial position, the area position covered by the current area window in the initial noise map is determined as the area position corresponding to the second area.
[0046] Optionally, when obtaining the final denoising result of the first area according to multiple mask images and multiple initial denoising results of the first area, the denoising unit is configured to:
[0047] For the initial denoising result guided by the prompt text of each target object, perform the following operations: In the mask image of the target object, determine the mask content corresponding to the same area position as the first area, and perform a Hadamard product operation on the initial denoising result guided by the prompt text of the target object and the mask content to obtain the single-object denoising result corresponding to the target object;
[0048] Superimpose the single-object denoising results of the multiple target objects to obtain the final denoising result of the first area.
[0049] In a third aspect, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the above method when executing the computer program.
[0050] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored, and the computer program implements the above method when executed by a processor.
[0051] In a fifth aspect, a computer program product is provided, including a computer program, and the computer program implements the above method when executed by a processor.
[0052] The beneficial effects of this application are as follows:
[0053] In an embodiment of the present application, an image generation method, apparatus, electronic device, and storage medium are provided. It is disclosed to obtain an initial noise map; then obtain multiple pieces of reference data configured for multiple target objects included in the target image. One piece of reference data includes: a prompt text for describing a target object, and a mask image for indicating the shape and position of a target object in the target image. Further, perform multiple rounds of iterative noise reduction on the initial noise map, and use the denoised noise map obtained in the last round of noise reduction as the target image. In the process of one round of iterative noise reduction, the following operations are performed: for the first region in the initial noise image, perform noise reduction guidance using multiple prompt texts respectively to obtain multiple initial denoised results of the first region, and obtain the final denoised result of the first region according to the multiple mask images and the multiple initial denoised results of the first region; for the second region in the initial noise image, perform noise reduction guidance using multiple prompt texts respectively to obtain multiple initial denoised results of the second region, and obtain the final denoised result of the second region according to the multiple mask images and the multiple initial denoised results of the second region. There is an overlapping part between the first region and the second region, and the overlapping part in the second region has the same content as the overlapping part in the final denoised result of the first region; determine the denoised noise map including multiple target objects according to the final denoised result of the first region and the non-overlapping part of the final denoised result of the second region with respect to the first region.
[0054] In this way, in the process of performing multiple rounds of iterative noise reduction on the initial noise map, by performing noise reduction on the first region and the second region in the initial noise map successively, it is equivalent to performing block noise reduction on the initial noise map from the perspective of different regions. Moreover, since there is an overlapping part between the first region and the second region, and the overlapping part in the second region has the same content as the overlapping part in the final denoised result of the first region, it is equivalent to, on the one hand, adopting a region division method suitable for regional noise reduction to divide the initial noise map as a whole, and on the other hand, since the second region includes denoised content and non-denoised content, during the specific noise reduction process, the denoised content can be used to guide the non-denoised content, which helps to generate coherent image content between different regions. Based on this, by means of the iterative regional noise reduction process, it is finally possible to generate image content with a very harmonious visual effect, thus greatly improving the generation quality of the image.
[0055] In addition, when performing noise reduction on either the first region or the second region, the respective prompt texts of each target object are used to perform noise reduction guidance separately, obtaining multiple initial denoising results. Moreover, the final denoising results corresponding to the regions are determined by using the mask images of each target object and the initial denoising results respectively guided by each prompt text. Therefore, it is possible to achieve differential noise reduction processing for the same region under the action of different prompt texts. Furthermore, by means of the positions and shapes of each target object described in each mask image in the target image, it is possible to specifically determine the denoising results corresponding to each target object from the initial denoising results obtained by differential noise reduction, such that the final denoising results obtained by noise reduction include the denoising results differentially generated for each target object. This not only improves the credibility of each final denoising result but also provides a guarantee for the denoising generation effect of the target image. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1A FIG. is a schematic diagram of the process of generating a target image based on a mask image using multidiffusion in the embodiments of the present application;
[0057] Figure 1B FIG. is a schematic diagram of a target image generated based on a mask image using multidiffusion in the embodiments of the present application;
[0058] Figure 2 FIG. is a schematic diagram of a possible application scenario in the embodiments of the present application;
[0059] Figure 3A FIG. is a schematic diagram of the process of image generation in the embodiments of the present application;
[0060] Figure 3B FIG. is a schematic diagram of the process of obtaining each mask image in the embodiments of the present application;
[0061] Figure 4 FIG. is a schematic diagram of the process of obtaining an initial noise map by noise reduction in the embodiments of the present application;
[0062] Figure 5 FIG. is a schematic diagram of a single-object noise map of each target object generated in the embodiments of the present application;
[0063] Figure 6A FIG. is a schematic diagram of the architecture of the SD model used in the embodiments of the present application;
[0064] Figure 6B FIG. is a schematic diagram of the process of performing one round of noise reduction in the embodiments of the present application;
[0065] Figure 7A FIG. is a schematic diagram of the process of obtaining multiple initial denoising results for the first region by noise reduction in the embodiments of the present application;
[0066] Figure 7B Schematic diagram of the process of aggregating to obtain the final denoising result of the first region in the embodiment of the present application;
[0067] Figure 7C Schematic diagram of the regional positions when two regions to be denoised are divided in the embodiment of the present application;
[0068] Figure 7D Schematic diagram of the regional positions when three regions to be denoised are divided in the embodiment of the present application;
[0069] Figure 8 Schematic diagram of a round of denoising process for the target image in the embodiment of the present application;
[0070] Figure 9A Schematic diagram of the overall process of image generation in the embodiment of the present application;
[0071] Figure 9B Schematic diagram for comparing the image generation effects in the embodiment of the present application;
[0072] Figure 9C Schematic diagram for displaying the generated image in the embodiment of the present application;
[0073] Figure 10 Schematic diagram of the logical structure of the image generation device in the embodiment of the present application;
[0074] Figure 11 Schematic diagram of the hardware composition structure of an electronic device applying the embodiment of the present application. Detailed implementation manners
[0075] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the technical solutions of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments recorded in this application document without creative efforts belong to the scope protected by the technical solutions of the present application.
[0076] Terms such as "first", "second", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here.
[0077] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0078] Some of the terms used in the embodiments of the present application are explained below to facilitate understanding by those skilled in the art.
[0079] Stable Diffusion (SD) model: is a potential text-to-image diffusion model that can generate realistic images based on the text content by removing noise given any text input.
[0080] Controllable image generation model: also known as controllable generation with diffusion models, or controllable multi-target generation model, is one of the research directions in stable diffusion models. It mainly improves controllability through various input formats (such as personal images, edge maps, posture maps, key points and bounding boxes). Controllable generation with diffusion models play an important role in personalized image generation and multi-target generation in image generation systems.
[0081] MultiDiffusion algorithm: It is a controllable generative diffusion model that combines multiple diffusion generation processes with shared parameters or layout constraints and uses the least squares optimal solution to coordinate all different diffusion processes.
[0082] The following is a brief introduction to the design concept of the embodiment of the present application:
[0083] In the process of conceiving and implementing image generation, the applicant thought of using a controllable multi-target generation model to generate target objects at corresponding positions based on the image layout; specifically, by utilizing the approximate position of the image layout to generate corresponding target objects, a target image including multiple target objects is obtained.
[0084] However, when there are overlapping layout contents in the generated target image, the generated image will be very inconsistent, which greatly affects the overall look and feel.
[0085] Based on this, the applicant of the present application thought that a controllable multi-object generation model could be adopted to generate a target image based on an overall image mask. In this case, when generating an image with the help of the image mask, the generated target image can better fit the input shape.
[0086] Specifically, the multidiffusion algorithm can be used to perform the following operations in each denoising process: divide the noise map into several regions for multi-region target generation. Moreover, the main task is to make the denoising of each step in the region unified in one direction, that is, for the denoising of each step in multiple regions, the denoising result of the previous step should be as similar as possible to the result of the next step, which corresponds to minimizing the difference between the two; furthermore, overall consistency is achieved through local consistency. In addition to maintaining consistency with the previous step direction and minimizing differences for the generation of each region, the overall mask image is also used for shape constraint.
[0087] For example, refer to Figure 1A As shown, it is a schematic diagram of the process of generating a target image based on a mask image using multidiffusion in an embodiment of the present application. According to Figure 1A As shown in the figure, after obtaining the mask image, the noise map, and each text guidance word, in each step of the noise reduction process, several regions are divided in the noise map, and independent noise reduction guidance is performed on each region, and the mask image is used to perform shape constraint on the noise reduction result; furthermore, after completing the set number of noise reduction steps, the generated target image is obtained.
[0088] However, since different regions are denoised independently and the image mask can only perform shape constraint overall, it is very likely that the image content between regions in the finally generated target image is not coherent, which is very inharmonious in terms of visual effect; in addition, there is great difficulty in dividing the regions of the noise map. If the regions are divided too small, the denoising process will consume a large amount of time and resources, and if the regions are divided too large, affected by the incoherence of the image content, the quality of the generated image is very poor.
[0089] For example, refer to Figure 1B As shown, it is a schematic diagram of the target image generated based on a mask image using multidiffusion in an embodiment of the present application. According to Figure 1B As can be seen from the content shown in the figure, in the generated target image, there is obvious incoherence in the parts framed by the two black lines, so the quality of the target image is very poor and cannot meet the expected image generation requirements.
[0090] In view of this, in the technical solution claimed in this application, an image generation method, apparatus, electronic device, and storage medium are disclosed, and obtaining an initial noise map is disclosed; then obtaining multiple pieces of reference data configured for multiple target objects included in the target image, where one piece of reference data includes: a prompt text for describing one target object, and a mask image for indicating the shape and position of one target object in the target image; further, performing multiple rounds of iterative noise reduction on the initial noise map, and using the denoised noise map obtained in the last round of noise reduction as the target image, where in the process of one round of iterative noise reduction, the following operations are performed: for the first region in the initial noise image, performing noise reduction guidance respectively using multiple prompt texts to obtain multiple initial denoised results of the first region, and obtaining the final denoised result of the first region according to the multiple mask images and the multiple initial denoised results of the first region; for the second region in the initial noise image, performing noise reduction guidance respectively using multiple prompt texts to obtain multiple initial denoised results of the second region, and obtaining the final denoised result of the second region according to the multiple mask images and the multiple initial denoised results of the second region, where there is an overlapping part between the first region and the second region, and the overlapping part in the second region has the same content as the overlapping part in the final denoised result of the first region; determining the denoised noise map including multiple target objects according to the final denoised result of the first region and the non-overlapping part of the final denoised result of the second region with the first region.
[0091] In this way, in the process of performing multiple rounds of iterative noise reduction on the initial noise map, by first performing noise reduction on the first region and the second region in the initial noise map, it is equivalent to performing block noise reduction on the initial noise map from the perspective of different regions; furthermore, since there is an overlapping part between the first region and the second region, and the overlapping part in the second region has the same content as the overlapping part in the final denoised result of the first region, it is equivalent to on the one hand adopting a region division method adapted to block noise reduction to divide the initial noise map as a whole, and on the other hand, since the second region includes denoised content and non-denoised content, during the specific noise reduction process, the denoised content can be used to guide the non-denoised content, which helps to guide the generation of coherent image content between different regions; based on this, with the help of the iterative block noise reduction process, it is finally possible to generate image content that is very harmonious in visual effect, thus greatly improving the generation quality of the image;
[0092] In addition, when performing noise reduction on any one of the first region and the second region, the respective prompt texts of each target object are used to perform noise reduction guidance separately, obtaining multiple initial denoising results, and the initial denoising results obtained by guiding with the mask images of each target object and the respective prompt texts are used to determine the final denoising result corresponding to the region. Therefore, it is possible to achieve differential noise reduction processing for the same region under the action of different prompt texts. Moreover, by means of the positions and shapes of each target object described in each mask image in the target image, it is possible to specifically determine the denoising results corresponding to each target object from the initial denoising results obtained by differential noise reduction, so that the final denoising results obtained by noise reduction include the denoising results differentially generated for each target object. This not only improves the credibility of each final denoising result but also provides a guarantee for the denoising generation effect of the target image.
[0093] The preferred embodiments of the present application will be described below with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only for explaining and illustrating the present application and are not used to limit the present application. And without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0094] Refer to Figure 2 As shown, it is a schematic diagram of a possible application scenario in an embodiment of the present application. In this schematic diagram of the application scenario, a processing device 210 and a client device 220 are included.
[0095] In some feasible embodiments of the present application, the processing device 210 can respond to an image generation request initiated by a target object on the client device 220, obtain the data provided by the client device 220 for describing the expected generation situation of the target image, and process the obtained data to obtain an initial noise map and multiple pieces of reference data configured for multiple target objects included in the target image. Furthermore, the processing device 210 can perform multiple rounds of iterative noise reduction on the initial noise map based on the multiple pieces of reference data to obtain the target image.
[0096] In some other feasible embodiments of the present application, the processing device 210 can respond to an image generation request initiated locally by a target object, obtain the data provided by the target object for describing the expected generation situation of the target image, process the obtained data to obtain an initial noise map and multiple pieces of reference data configured for multiple target objects included in the target image, and perform multiple rounds of iterative noise reduction on the initial noise map based on the multiple pieces of reference data to obtain the target image.
[0097] The processing device 210 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms; or, in a feasible implementation, the processing device 210 is a terminal device such as a tablet computer or a notebook with the required processing capabilities.
[0098] The client device 220 includes, but is not limited to, mobile phones, tablet computers, notebooks, e-book readers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. The embodiments of the present invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.
[0099] It should be noted that in the embodiments of the present application, when the client device 220 initiates an image generation request, it can send a request to the processing device 210 through a target application. Among them, the target application can be a mini-program application, or a client application, or a web application. The present application does not make specific limitations on this.
[0100] In the embodiments of the present application, the processing device 210 and the client device 220 can communicate through a wired network or a wireless network.
[0101] The following describes the image generation process in combination with feasible application scenarios:
[0102] Application scenario: Image generation is performed in a scenario where artificial intelligence can be used to generate content (Artificial Intelligence Generated Content, AIGC).
[0103] Specifically, in specific scenarios such as artificial intelligence (AI) cloud painting and AI image generation, the processing device can process the corresponding initial noise map according to the AI image generation requirements proposed by relevant objects, and configure multiple reference data for multiple target objects included in the target image; then, perform multiple rounds of iterative noise reduction on the initial noise map based on the multiple reference data to obtain the denoised target image. Among them, one reference data includes: a prompt text for describing a target object, and a mask image for indicating the shape and position of a target object in the target image.
[0104] During an iterative noise reduction process, the processing device adopts a serial noise reduction method to sequentially perform noise reduction on each region in the initial noise map, obtaining the respective final denoising results for each region, and then obtaining a denoised noise map based on the final denoising results of each region. Among them, the intersection of each region covers all the content of the initial noise map, and there is an overlapping part between two adjacent regions.
[0105] Taking the determination of two regions (the first region and the second region respectively) from the initial noise map as an example, based on the final denoising result of the first region and the non-overlapping part of the final denoising result of the second region with respect to the first region, a denoised noise map containing multiple target objects is finally determined. Among them, there is an overlapping part between the first region and the second region, and the overlapping part in the second region has the same content as the overlapping part in the final denoising result of the first region.
[0106] Among them, taking the noise reduction process of the first region as an example, multiple hint texts are used to perform noise reduction guidance respectively to obtain multiple initial denoising results of the first region, and based on multiple mask images and the multiple initial denoising results of the first region, the final denoising result of the first region is obtained.
[0107] Based on this, after the image generation is completed, a coherent, harmonious, and natural target image can be obtained.
[0108] In addition, it should be understood that in the specific implementation of the present application, it involves the acquisition and processing of the image generation requirements of relevant objects. When the embodiments described in the present application are applied to specific products or technologies, the permission or consent of relevant objects needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0109] First, the process of the processing device implementing image generation will be described below with reference to the accompanying drawings:
[0110] Refer to Figure 3A As shown, it is a schematic flowchart of image generation in an embodiment of the present application. The process of the processing device implementing image generation will be described below with reference to the accompanying Figure 3A drawings, and the process of the processing device implementing image generation will be described:
[0111] Step 301: The processing device obtains multiple pieces of reference data configured for multiple target objects included in the target image.
[0112] In the embodiments of the present application, in order to implement the generation of the target image, the processing device needs to obtain data that can indicate the generation effect of the target image. Among them, the data that can indicate the generation effect of the target image is specifically multiple pieces of reference data configured for multiple target objects included in the target image; one piece of reference data includes: a hint text for describing a target object, and a mask image for indicating the shape and position of a target object in the target image.
[0113] It should be noted that in the embodiments of the present application, each target object is each object expected to be denoised and generated in the target image. The target objects specifically include foreground objects and background objects. Among them, the present application regards all background contents as one target object; moreover, in the generated target image, the image contents corresponding to each target object cover the entire target image; in addition, there is a corresponding prompt text and mask image for one target object; a mask image refers to an image containing the object mask of a corresponding target object, and is used to describe the expected size and expected position of the corresponding target object in the generated target image.
[0114] The multiple reference data obtained by the processing device can be specifically constructed by itself, or can be obtained from other devices. Taking the processing device constructing multiple reference data by itself as an example, in the process of generating multiple reference data, the processing device determines multiple target objects expected to be generated in the target image, and for each target object, performs the following operations: based on the prompt words of the target object, generate a prompt text, and in the pre-constructed object mask set, obtain the target object mask corresponding to the target object; according to the preset resolution of the target image, the object mask of the target object, and the expected size and expected position of the target object in the target image, obtain the mask image of the target object, and use the prompt text and mask image of the target object as a piece of reference data.
[0115] It should be noted that in the embodiments of the present application, since the image generation is performed by the processing device under the expected requirements of relevant objects, it should be understood that before the image generation, there is a known image concept information set, where the image concept information set includes: the resolution representing the size of the target image, each target object expected to be generated in the target image, and the expected position and expected size of each target object in the target image respectively. Furthermore, the processing device can obtain the generation of multiple reference data based on each image concept information.
[0116] In addition, in the embodiments of the present application, the pre-constructed object mask set includes the object masks of each object, where the object masks can be pre-generated based on the corresponding object images by using a trained mask segmentation model; the selected mask segmentation model can be any segmentation model (Segment Anything Model, SAM). The prompt words of each target object can be obtained from the pre-constructed prompt word set, or can be temporarily determined for the target object.
[0117] Specifically, when determining the prompt words for the target object, if the target object is known, the corresponding prompt words can be directly configured; if the target object is unknown, the trained BLIP model can be used to obtain the prompt words output by the model based on the image of the target object. Here, the training process of the SAM model and the BLIP model is not limited in this application.
[0118] For example, assume that for a target object to be generated in an image, only the image of the target object exists, but the target object cannot be accurately described; at this time, the trained BLIP model can be used to obtain the corresponding prompt words based on the image of the target object.
[0119] Another example, assume that a cat needs to be generated in the generated image, then the corresponding prompt word "cat" can be configured by itself.
[0120] After that, after the processing device obtains the prompt words of each target object, it can generate the corresponding prompt text according to the actual image generation requirements.
[0121] For example, if it is expected to generate a cat, a bear, and a rabbit in the image, then the target objects include: cat, bear, and rabbit; for the target object "cat", the obtained prompt word is "cat", and the generated prompt text is "a cat"; for the target object "bear", the obtained prompt word is "bear", and the generated prompt text is "a bear"; for the target object "rabbit", the obtained prompt word is "rabbit", and the generated prompt text is "a rabbit".
[0122] In the process of obtaining the mask images corresponding to each target object, the processing device can first construct a solid-color image with a specified resolution, and then perform the following operations for the object mask of each target object: determine the object shape shown by the object mask, and determine the area indicated by the object shape as the corresponding target object area, and according to the generation requirements, change the size of the target object in the generated image by adjusting the total number of pixel points occupied by the target object area; in a constructed solid-color image, specify the generation position of a target object, and overlap the center of the target object area with the determined generation position; determine the pixel points of each area covered by the target object area in a solid-color image, and then uniformly adjust the other pixel points in the solid-color image except the pixel points of each area to other pixel values, so as to obtain the mask image corresponding to a target object.
[0123] Specifically, when the target object is the image background, considering that the image background usually does not have a specific shape, it is impossible to directly obtain the corresponding object mask; in this case, after determining the object masks of each foreground object, the background shape can be correspondingly determined, so the background part can be masked to obtain the mask image corresponding to the background.
[0124] For example, a solid color image can be a pure white image. Correspondingly, the pixel values of other pixels except for the pixel points in each region are adjusted from white to black.
[0125] For another example, referring to Figure 3B As shown, it is a schematic diagram of the process of obtaining each mask image in the embodiment of the present application. Suppose the foreground objects expected to be generated in the image include: a bird, a bear, and a rabbit; then, after determining the image size based on the resolution of the image, corresponding mask images can be generated for each foreground object and background object respectively.
[0126] In this way, by generating mask images that can constrain the size and shape of the target object in the target image for each target object to be generated, the image generation requirements can be visualized, so that in the subsequent image generation process, it can better assist in generating images that meet the actual requirements.
[0127] Step 302: The processing device obtains an initial noise map.
[0128] In the embodiment of the present application, in order to reduce noise to generate a target image, the processing device needs to obtain an initial noise map for the noise reduction process. Among them, the initial noise map is a noise image including each target object generated after denoising guidance on the original noise map.
[0129] In some feasible implementation manners, the processing device can obtain the initial noise map by itself based on the original noise reduction; in other feasible implementation manners, the processing device can obtain the initial noise map denoised by other devices based on the same image generation requirements. The present application does not make specific limitations on this; in addition, in the embodiment of the present application, the mask image, the original noise map, the initial noise map, and the target image have the same image size.
[0130] Referring to Figure 4 As shown, it is a schematic diagram of the process of obtaining the initial noise map by noise reduction in the embodiment of the present application. The following will be combined with the attached Figure 4 , taking the processing device obtaining the initial noise map by itself through noise reduction as an example, to illustrate the process of obtaining the initial noise map:
[0131] Step 3021: The processing device obtains the original noise map.
[0132] When executing step 3021, the processing device can use the constructed Gaussian noise map as the obtained original noise map.
[0133] Step 3022: The processing device obtains multiple pieces of reference data configured for multiple target objects included in the target image.
[0134] Specifically, the multiple reference data obtained by the processing device are the same as the multiple reference data involved in step 301. In other words, in the process of denoising the initial noise map and the original noise map, multiple reference data are required. The process of obtaining multiple reference data will not be described in this application. One piece of reference data includes: a prompt text for describing a target object, and a mask image for indicating the shape and position of a target object in the target image.
[0135] Step 3023: The processing device performs multiple rounds of iterative denoising on the original noise map, and uses the denoised noise map obtained in the last round of denoising as the initial noise map.
[0136] In the embodiment of this application, when performing step 3023, the processing device performs multiple rounds of iterative denoising on the original noise map until the preset first number of denoising rounds is reached, and uses the denoised noise map obtained in the last round of denoising as the initial noise map. The value of the first number of denoising rounds is set according to actual processing needs.
[0137] Taking the initial denoising process for the original noise map as an example, the operations performed will be described below:
[0138] The processing device obtains the single-object noise map of each target object according to the original noise map and multiple mask images, and for the single-object noise map of each target object, the following operations are performed: using the prompt text of the target object to perform denoising guidance on the single-object noise map of the target object to obtain a single-object denoised map; then superimposing the contents of multiple single-object denoised maps to obtain the denoised noise map of the original noise map.
[0139] In the embodiment of this application, based on each mask image, the processing device can determine the shape and position of each target object in the target image respectively, and then determine the single-object noise map corresponding to each target object respectively. One single-object noise map corresponds to one target object; only the content area corresponding to the same position as the object mask area of the target object in the mask image in one single-object noise map retains the noise content in the original noise map, so that only the content area corresponding to the target object in the single-object noise map is in a state where denoising can be performed; the object mask area of the target object refers to the area covered by each pixel point used to describe the shape and position of the target object in the mask image; the image sizes of the single-object noise map and the mask image are the same.
[0140] Specifically, when the pixel values corresponding to each pixel point in the target object area in the single-object noise map are all 1 and the pixel values corresponding to each pixel point in other areas except the target object area are all 0, the following operations can be performed for each target object: calculating the Hadamard product between the mask image of the target object and the original noise map to obtain the single-object noise map of the target object.
[0141] For example, referring to Figure 5 as shown, it is a schematic diagram of the single-body noise map of each target object generated in the embodiment of the present application. According to Figure 5 the content shown, after obtaining the mask images of 4 target objects and the original noise map with the same size as the mask images, the following operations are respectively performed according to each mask image: In the original noise map, determine the content area corresponding to the target object area in the corresponding mask image, and only retain the noise content within this content area in the original noise map to obtain a single-body noise map determined based on the original noise map. Similarly, single-body noise maps corresponding to each of the 4 target objects can be obtained.
[0142] Furthermore, in the process of denoising the single-body noise map of each target object, in a feasible implementation, the processing device can use a pre-trained SD model to obtain a denoised single-body denoised map based on the single-body noise map of the target object under the denoising guidance of the prompt text of the target object. Among them, the present application can perform denoising with the help of an open-source pre-trained SD model, and the present application does not explain the training process of the SD model here; in other feasible implementation manners, the processing device can use other denoising models, such as the multidiffusion model, to implement the denoising process.
[0143] It should be noted that in the embodiment of the present application, when using the SD model for denoising processing, one round of denoising process is one step of denoising performed by the SD model. In the embodiment of the present application, the involved denoising processes are all carried out in the latent space.
[0144] For example, taking the denoising processing implemented based on the SD model as an example, referring to Figure 6A as shown, it is a schematic diagram of the architecture of the SD model adopted in the embodiment of the present application. According to Figure 6A the content shown, the main modules in the SD model include: an encoder E, a decoder D, a denoising module implemented by an inverse diffusion network (or a denoising network, denoted as U-Net), and a conditioning module; Figure 6A where Tθ represents extracting the latent feature of the prompt text, and the part composed of multiple QKV represents the multi-head attention layer, and Z T represents the noise map in the latent space obtained by processing when encoding the input picture; D is the corresponding decoder.
[0145] Again, for example, referring to Figure 6BAs shown, it is a schematic diagram of the process of performing one round of noise reduction in an embodiment of the present application. Assuming there are 4 target objects, there are 4 single-body noise maps. Among them, the pixel values of the pixel points in the non-noise part of each single-body noise map can be set to 0, so that the pixel values of the pixel points at the same position can be directly added to the denoised single-body denoised maps to obtain the corresponding denoised noise map in the subsequent process.
[0146] Continuing with reference to the attached Figure 6B For further illustration, when the processing device performs noise reduction on the single-body noise map of a target object, it needs to input the prompt text of the target object in the hidden layer space and the single-body noise map into the SD model together, and perform noise reduction processing by means of the U-net network. Among them, considering that the SD model can perform parallel noise reduction on multiple noise maps, in a feasible implementation, 4 single-body noise maps and the corresponding prompt texts can be input into the SD model for parallel denoising at the same time, and the denoised single-body denoised maps are obtained respectively. Then, the pixel values at the same pixel positions in the 4 denoised single-body denoised maps are superimposed to obtain the denoised noise map after one-step noise reduction, where Figure 6B Only for illustrative purposes, the specific noise reduction process is carried out in the hidden layer space.
[0147] Similarly, after the denoised noise map is obtained in the initial noise reduction, when it is determined that the noise reduction process needs to be continued, the processing device can use the denoised noise map as the original noise map for a new round of noise reduction process, and perform the noise reduction process in the same way.
[0148] In this way, in the process of reducing the original noise map to obtain the initial noise map, by combining the mask image of each target object, the area corresponding to the corresponding target object can be located in the original noise map, so that the corresponding single-body noise map can be obtained for each target object respectively; moreover, since only the content area corresponding to the target object in the single-body noise map has the noise content that needs to be reduced, it is equivalent to restricting the noise position to be reduced in terms of shape; furthermore, the prompt text of the target object is used for noise reduction guidance, and the corresponding target object is generated by reducing the noise in the content area of the single-body noise map of the target object, so that the noise reduction generation of the specific object can be realized at the position of the target object area indicated by the mask image, which is equivalent to performing noise reduction generation from the perspective of each single target object respectively, so that the finally obtained initial noise map is a noise map that includes the noise-reduced generated target objects and needs further noise reduction, which helps to better reduce the noise and generate the image content that meets the expectations.
[0149] Step 303: The processing device performs multiple rounds of iterative noise reduction on the initial noise map, and uses the denoised noise map obtained in the last round of noise reduction as the target image.
[0150] In the embodiment of the present application, when performing step 303, the processing device performs multiple rounds of iterative noise reduction on the initial noise map until the preset second number of noise reduction rounds is reached, and uses the denoised noise map obtained in the last round of noise reduction as the target image, where the value of the second number of noise reduction rounds is set according to actual processing needs.
[0151] Taking the initial noise reduction process for the initial noise map as an example, the operations performed will be described below:
[0152] In one round of iterative noise reduction process, the processing device adopts a serial noise reduction method and sequentially performs noise reduction on each area to be noise-reduced divided from the initial noise map.
[0153] The following combines with the attached Figure 3A Taking the example of dividing two areas from the initial noise map, namely the first area and the second area, the operations performed in the initial round of noise reduction process will be described:
[0154] Step 3031: For the first area in the initial noise image, multiple hint texts are used to perform noise reduction guidance respectively to obtain multiple initial denoising results of the first area, and according to multiple mask images and the multiple initial denoising results of the first area, the final denoising result of the first area is obtained.
[0155] It should be noted that in the embodiment of the present application, the area positions of the first area and the second area in the initial noise map can be divided by the processing device in the following way: obtain the preset area window, sliding direction, sliding step, and the initial position in the initial noise map; when the area window is located at the initial position, the area position covered in the initial noise map is determined as the area position corresponding to the first area; in the initial noise map, after sliding the sliding step along the sliding direction from the initial position, the area position covered by the current area window in the initial noise map is determined as the area position corresponding to the second area.
[0156] In a feasible implementation manner, the processing device can successively determine a set number of areas to be noise-reduced in the initial noise map according to actual processing needs. Specifically, the processing device can preset the area size covered by each area position according to the size of the initial noise map, the set number of areas to be noise-reduced, and the size of the overlapping part between the areas to be noise-reduced obtained twice adjacent; then, determine the size of the area window according to the preset area size, so that an area window can cover each area position during the movement; in addition, since there is an overlapping part between the areas to be noise-reduced obtained twice adjacent, it is necessary to specify the sliding direction and sliding step of the area window, and the initial position of the area window.
[0157] Among them, the regional position is used to indicate different regions to be denoised in the initial noise map; the sliding direction is used to indicate the positional relationship between the successively obtained regions to be denoised; the sliding step is used to define the positional intersection situation between two regions to be denoised (such as the first region and the second region), in other words, it is used to describe the size of the overlapping part between the two successively determined regions to be denoised.
[0158] In this way, by means of the configured size, sliding direction and sliding step of the regional window, as well as the initial position, the regional positions corresponding to the first region and the second region with the same size can be effectively determined.
[0159] In some other feasible embodiments of the present application, the processing device can obtain each region to be denoised by performing a box selection operation on the image according to actual processing needs. Among them, taking the example of successively determining two regions to be denoised (the first region and the second region) from the initial candidate image, there is an overlapping part between the first region and the second region, and this overlapping part in the second region is obtained after denoising the first region.
[0160] Furthermore, after the processing device obtains the first region, in the process of denoising the first region to obtain the final denoised result of the first region, multiple prompt texts of multiple target objects are used to respectively guide the denoising of the first region to obtain multiple initial denoised results of the first region, and then the mask images of multiple target objects and the multiple initial denoised results of the first region are used to obtain the final denoised result of the first region.
[0161] Specifically, after the processing device determines the first region, multiple prompt texts are respectively used to guide the denoising of the first region to obtain the initial denoised results generated under the denoising guidance of each prompt text.
[0162] It should be noted that in the embodiments of the present application, when denoising the first region, a pre-trained SD model can be used to respectively output multiple initial denoised results of the first region based on multiple prompt texts and the first region in the initial noise map; when the SD model allows parallel independent denoising of multiple images, different input data composed of different guiding statements and the first region can be simultaneously input into the SD model; and when the SD model does not allow parallel independent denoising of multiple images, different input data composed of different guiding statements and the first region can be successively input into the SD model for processing. Among them, the denoising process involved in the present application is carried out in the latent space. In the process of guiding denoising according to the prompt text, it also involves the process of injecting the encoded prompt text into the latent space, generating a noise map in the latent space, and injecting the mask image into the latent space. The present application does not make specific descriptions on this.
[0163] For example, refer to Figure 7AAs shown, it is a schematic diagram of the process of obtaining multiple initial denoising results of the first region through denoising in an embodiment of the present application. Assume that there are 4 target objects to be generated in the target image, namely three foreground objects and one background object; then, the processing device respectively uses the prompt texts corresponding to each of the 4 target objects to guide the denoising of the first region, and obtains an initial denoising result of the corresponding first region.
[0164] Continue to describe in combination with the content Figure 7A schematically shown. Assume that the currently determined region to be denoised is the first region, and the SD model is used for denoising processing, and the model allows 4 images to be input in parallel for independent denoising at the same time. Then, as Figure 7A shown, the prompt text of the first region and target object 1, the prompt text of the first region and target object 2, the prompt text of the first region and target object 3, and the prompt text of the first region and target object 4 can be independently used as the input of the SD model, and the U-Net is specifically used for denoising processing to obtain multiple initial denoising results of the first region output by the U-Net.
[0165] In this way, by respectively using the prompt texts of multiple target objects to guide the denoising of the first region, multiple initial denoising results corresponding to the first region obtained by differential denoising under the guidance of different prompt texts can be obtained.
[0166] After that, in the process of the processing device obtaining the final denoising result of the first region according to multiple mask images and multiple initial denoising results of the first region, the processing device performs the following operations on the initial denoising result guided by the prompt text of each target object: in the mask image of the target object, determine the mask content corresponding to the same region position as the first region, and perform a Hadamard product operation on the initial denoising result guided by the prompt text of the target object and the mask content to obtain the single-object denoising result corresponding to the target object; furthermore, superimpose the single-object denoising results of multiple target objects to obtain the final denoising result of the first region.
[0167] In an embodiment of the present application, after denoising the first region to obtain multiple initial denoising results of the first region, the processing device needs to comprehensively obtain the final denoising result of the first region based on the initial denoising results obtained by denoising the first region under the guidance of different prompt statements.
[0168] Specifically, in the process of fusing multiple initial denoising results of the first region to obtain the final denoising result of the first region, the processing device can, with the help of the mask content at the same region position corresponding to the first region in each mask image, determine the single-body denoising result corresponding to the target object region in the corresponding initial denoising result. Then, by superimposing multiple single-body denoising results, the final denoising result of the first region is obtained, where the first region, the initial denoising result, and the final denoising result have the same size; the single-body denoising result refers to the denoising content whose position in the initial denoising result matches the target object region where the target object is located after generating the initial denoising result under the denoising guidance of the prompt text of the target object. In other words, the single-body denoising result refers to the denoising result specifically corresponding to the target object determined in the initial denoising result generated under the guidance of the prompt text of the target object.
[0169] For example, referring to Figure 7B shown, it is a schematic diagram of the process of aggregating the final denoising result of the first region in the embodiment of the present application. According to the content Figure 7B shown, it can be known that assuming there are 4 target objects, then under the guidance of the prompt statements of the 4 target objects, 4 initial denoising results can be guided to be obtained, and the mask regions (or mask content) corresponding to the initial denoising results can be determined in the mask images of the 4 target objects respectively; then, the processing device performs Hadamard product operations on the initial denoising results and the mask content corresponding to the same target object to obtain the single-body denoising results of each target object respectively; then, the single-body denoising results of each target object are superimposed to obtain the final denoising result of the first region, where the region corresponding to the mask content, the region corresponding to the first region, the initial denoising result, and the region corresponding to the final denoising result have the same region size; when calculating the Hadamard product, the mask content of each target object has been pre-normalized in terms of pixel values, so that the pixel values of each pixel point in the object mask part are 1 (appearing white visually), and the pixel values of each pixel point in the non-object mask part are 0 (appearing black visually).
[0170] In this way, under the action of the mask content at the same region position corresponding to the first region in the mask images of each target object, the single-body denoising results corresponding to each target object can be determined from the initial denoising results obtained by guiding with each prompt text respectively. Thus, by synthesizing the single-body denoising results of different target objects, the final denoising result of the first region is obtained, so that in the process of denoising the first region obtained by segmentation, differential denoising guidance can be performed for different target objects.
[0171] Step 3032: For the second region in the initial noisy image, perform noise reduction guidance using multiple hint texts respectively to obtain multiple initial denoising results of the second region, and obtain the final denoising result of the second region according to the multiple mask images and the multiple initial denoising results of the second region.
[0172] In the embodiment of the present application, after the processing device completes noise reduction for the first region in the initial noisy image, it continues to perform noise reduction processing on the second region in the initial noisy image. Moreover, in the specific process of noise reduction for the second region, the method of generating multiple initial denoising results and the final denoising result involved in step 3031 is adopted. Similarly, multiple initial denoising results of the second region are generated, and based on the multiple initial denoising results of the second region, the final denoising result of the second region is obtained.
[0173] It should be noted that in the embodiment of the present application, there is an overlapping part between the first region and the second region determined from the initial noise map, and the overlapping part in the second region is the same as the overlapping part in the final denoising result of the first region; based on this, it can be known that since there is an overlapping part between the two regions to be denoised obtained successively (i.e., the first region and the second region), there must be denoised content and non-denoised content in the second region during the noise reduction process for the second region. At this time, during the noise reduction process in the latent space, the non-denoised content can be guided for noise reduction by means of the denoised content.
[0174] For example, refer to Figure 7C As shown, it is a schematic diagram of the regional positions when two regions to be denoised are divided in the embodiment of the present application. Assume that two regions to be denoised are obtained from the initial noise map, namely the first region and the second region, and the preset overlapping part size is half of the initial noise map. Therefore, the size, initial position, sliding direction, and sliding step of the regional window shown in Figure 7C can be obtained. Furthermore, combined with the content shown in the attached Figure 7C it can be known that according to the position of the first divided region, the first region can be determined in the initial noise map; after completing noise reduction for the first region, the second region can be obtained in the initial noise map according to the position of the second divided region. Among them, the second region includes the denoised content indicated by the shaded part and the non-denoised content.
[0175] It should be understood that in the embodiments of the present application, dividing the initial noise map into two region blocks is only for illustrative purposes. In the feasible implementation manners of the present application, multiple regions to be denoised can be divided from the initial noise map. Among them, in the case of obtaining multiple regions to be denoised, there is an overlapping part with the same size between two adjacent obtained regions to be denoised. Considering that when there are more than two regions to be denoised, the involved processing logic is the same as that when there are two regions to be denoised, so the present application will not expand the description any further.
[0176] For example, according to Figure 7D shown, it is a schematic diagram of the region positions when three regions to be denoised are divided in total in the embodiments of the present application. Combining the content Figure 7D shown, it can be known that assuming that a total of three region positions need to be divided from the initial noise map, that is, there are three regions to be denoised correspondingly, then the region positions can be determined according to the size, initial position, sliding direction, and sliding step of the region window Figure 7D shown. Furthermore, combining the content Figure 7D shown, according to the first divided region position, the first region can be determined in the initial noise map; after denoising the first region, according to the second divided region position, the second region can be obtained in the initial noise map, where the shaded part included in the second region indicates the denoised content; similarly, after denoising the second region, according to the third divided region position, the third region can be obtained in the initial noise map, where the shaded part included in the third region indicates the denoised content.
[0177] In this way, by dividing the initial noise image into the first region and the second region, and sequentially performing denoising processing on the first region and the second region, it is equivalent to realizing region-based denoising; moreover, when denoising the second region, since the second region includes both denoised content and undenoised content, the denoised region content can be used to guide the undenoised region content during the region-based denoising process, thereby improving the denoising effect and helping to generate an image that meets the expected requirements.
[0178] Step 3033: Determine a denoised noise map containing multiple target objects according to the final denoising result of the first region and the part of the final denoising result of the second region that does not overlap with the first region.
[0179] Specifically, since there is an overlapping part between the first region and the second region, the second region includes both denoised content and undenoised content at the same time. At this time, the denoised content obtained after denoising the first region corresponding to the overlapping part will be denoised again along with the denoising process of the second region. In this case, the present application only retains the initial denoising result.
[0180] That is to say, in the feasible embodiments of the present application, after obtaining the final denoising result for the second region, instead of directly updating the image content corresponding to the second region in the initial noise map with the final denoising result obtained by denoising, the overlapping part and the non-overlapping part are determined in the second region, and only the content of the non-overlapping part is used to update the initial noise map; alternatively, the processing device may combine the final denoising result of the first region with the denoising results of other parts except the overlapping part in the final denoising result of the second region according to the position to obtain a denoised noise map.
[0181] In addition, in a feasible implementation manner, according to the actual processing requirements, considering that the overlapping parts of different regions to be denoised are denoised twice, the intermediate results can be retained, and after the denoising of the initial noise image is completed, the retained intermediate results are mapped back to the original regions, so that only the results of the first denoising are retained in different regions to be denoised. Herein, the first region and the second region in the present application both belong to the regions to be denoised.
[0182] In this way, by means of the processing procedures of steps 3031-3033, it is possible to perform regional denoising on the initial noise map in one round of denoising of the initial noise map. Moreover, in the specific denoising process, differential denoising processing of the same region under the action of different prompt texts is realized; in addition, by means of the positions and shapes of the target objects described in each mask image in the target image, it is possible to specifically determine the denoising results corresponding to each target object from the initial denoising results obtained by differential denoising, so that the final denoising results obtained by denoising a single region include the denoising results differentially generated for each target object. This not only improves the credibility of each final denoising result, but also provides a guarantee for the denoising generation effect of the target image.
[0183] Similarly, after the processing device determines that the number of iterative denoising rounds for the initial noise map has not reached the preset second denoising round, the obtained denoised noise map is used as the initial noise map for the new round of training, and the operations shown in the above steps 3031-3033 are similarly executed until the number of iterative denoising rounds performed reaches the preset second denoising round, and the obtained denoised noise map is determined as the target image.
[0184] Optionally, after obtaining the target image, the processing device may continue to perform multiple rounds of iterative denoising on the target image, and finally obtain a final image obtained by denoising the target image.
[0185] Specifically, the processing device performs multi-round iterative noise reduction on the target image and uses the denoised image obtained in the last round of noise reduction as the final image. During one round of iteration, the following operations are performed: for the target image, multiple prompt texts are used to perform noise reduction guidance respectively to obtain multiple intermediate images of the target image, and based on multiple mask images and the multiple intermediate images of the target image, the denoised image of the target image is obtained.
[0186] In the implementation of this application, the processing device can perform multi-round iterative noise reduction on the target image until the preset third number of noise reduction rounds is reached, and use the denoised image obtained in the last round of noise reduction as the final image; moreover, after one round of noise reduction on the target image, when it is determined that the total number of noise reduction rounds iteratively performed on the target image has not reached the third number of noise reduction rounds, the currently obtained denoised image is used as the target image for the next round of noise reduction, and the next round of noise reduction processing is continued.
[0187] It should be noted that in the embodiments of this application, when performing noise reduction on the target image, a pre-trained SD model can be used, and multiple prompt texts are used to perform noise reduction guidance on the target image respectively to output each intermediate image after noise reduction; in the case where the SD model allows parallel independent noise reduction of multiple images, different input data combined by different guiding statements and the target image can be input into the SD model simultaneously; and in the case where the SD model does not allow parallel independent noise reduction of multiple images, different input data combined by different guiding statements and the target image can be input into the SD model for processing successively, where the image noise reduction process involved is performed in the latent space.
[0188] In addition, during the process of the processing device obtaining the denoised image of the target image based on multiple mask images and the multiple intermediate images of the target image, for each target object, the following operations are performed: the Hadamard product result between the mask image of the target object and the intermediate image generated under the guidance of the prompt text based on this target object is calculated to obtain the noise reduction result of the corresponding target object. Then, by superimposing the noise reduction results of multiple target objects, the denoised image of the target image is obtained, where the image sizes of the denoised image, the intermediate image, the target image, and the mask image are the same; when performing the Hadamard product operation, the pixel values of the pixel points at the corresponding same positions in the intermediate image and the mask image are multiplied, and finally the Hadamard product result composed of the respective product results corresponding to each pixel point is obtained; the pixel values of each pixel point in the mask image have been normalized.
[0189] For example, refer to Figure 8 As shown, it is a schematic diagram of one round of noise reduction process for the target image in the embodiments of this application. According to Figure 8As can be seen from the content shown, the processing device uses the prompt texts corresponding to 4 target objects respectively to perform noise reduction guidance on the target image to obtain corresponding intermediate images; then, the Hadamard product operation is performed on the intermediate image of the target object and the mask image of the target image to obtain the corresponding Hadamard product result; after that, the pixel values of the pixel points at the same position in each Hadamard product result are superimposed, and finally the denoised image obtained by combining each Hadamard product result is obtained.
[0190] In this way, in the process of performing multiple rounds of iterative noise reduction for the target object, it is equivalent to performing noise reduction from the perspective of the overall image. The target image is denoised for the overall picture according to the prompt texts of each target object (including foreground objects and background objects), and the noise reduction results of each target object as a whole are determined respectively according to the mask images of each target object. Then, by superimposing the noise reduction results of each target object, a personalized image including multiple target objects as a whole can be finally obtained.
[0191] Optionally, in the embodiment of the present application, after the processing device obtains the final image through noise reduction, it can quantitatively evaluate the generation quality of the image. Specifically, the processing device can extract the corresponding comprehensive text features based on multiple prompt texts, and extract the corresponding image features based on the final image; calculate the feature similarity between the comprehensive text features and the image features, and when the feature similarity reaches the set threshold, stop the current image generation process and output the final image.
[0192] It should be noted that the processing device can use the CLIP model to calculate the feature similarity between the image features extracted based on the generated final image and the comprehensive text content extracted based on multiple prompt texts; at the same time, the value of the set threshold can be configured according to actual processing needs.
[0193] For example, the CLIP model can be used to obtain the clip-score index for measuring the feature similarity based on the final image and the text content generated based on each prompt text. Assuming that the value of the set threshold is 0.25, if the value of the clip-score index is higher than 0.25, it is considered that the generated final image meets the expectations.
[0194] In this way, by calculating the feature similarity based on the generated final image and the text content, the generation effect of the image can be quantified, and it can be evaluated whether the image includes the object content expected to be generated, providing a reference basis for considering the generation accuracy of the image.
[0195] Specifically, when the feature similarity fails to reach the set threshold, the processing device re-obtains the initial noise map, re-obtains the initial noise map, and re-obtains multiple pieces of reference data configured for the target image; and then, based on the re-obtained initial noise map and multiple pieces of reference data, the image generation process is re-executed.
[0196] Specifically, when it is determined that the feature similarity fails to reach the set threshold, it can be determined that the generated final image fails to meet the expected image generation requirements. At this time, it is necessary to re-obtain the noise map and re-perform noise reduction and generation of the image based on the noise map.
[0197] In this way, after it is determined that the generated image does not meet the expectations, it is possible to trigger re-performing noise reduction and generation of the image, so as to restart the image generation process and re-generate an image that meets the expectations.
[0198] Generally speaking, in the feasible image generation process of this application, the following three noise reduction processes are involved: the process of performing object-based noise reduction on the original noise map to generate the initial noise map; the process of performing block-based noise reduction on the initial noise map to generate the target image; the process of performing overall noise reduction on the target image to generate the final image. Based on this, in the related image generation process, it may only include one noise reduction process, or include two noise reduction processes, or include three noise reduction processes, where each noise reduction process corresponds to multiple rounds of iterative noise reduction processing.
[0199] Based on this, in different feasible embodiments, when the processing device evaluates the generation quality of the image, it can be the target image or the final image. Among them, the process of evaluating the target image can be the same as the above processing process. Based on multiple prompt texts, the corresponding comprehensive text features are extracted, and based on the target image, the corresponding image features are extracted; then, the feature similarity between the comprehensive text features and the image features of the target image is calculated, and when the feature similarity reaches the set threshold, the current image generation process is stopped and the target image is output.
[0200] In addition, the processing device can, according to actual processing needs, determine the involved noise reduction processes before performing image generation, and then determine the number of iterative rounds involved in each noise reduction process.
[0201] Assume that in the image noise reduction process, three noise reduction processes are involved and the SD model is used for noise reduction processing. When determining the number of noise reduction rounds in each noise reduction process, the processing device obtains the preset total number of denoising rounds; and then, according to the total number of denoising rounds and the preset round ratio, configures the first number of noise reduction rounds used when denoising the original noise map to obtain the initial noise map, the second number of noise reduction rounds used when denoising the initial noise map to obtain the target noise map, and the third number of noise reduction rounds used when denoising the target image to obtain the final image.
[0202] Specifically, the processing device can set the ratio of the number of rounds for different noise reduction processes. For example, configure the ratios 0.2, 0.5, 0.3, and configure the total number of rounds (or total number of steps) passed through by the noise reduction, and then respectively determine the total number of rounds of noise reduction corresponding to each noise reduction process.
[0203] In this way, the noise reduction process can be configured as a whole, with different noise reduction processes being differentially emphasized according to actual processing needs. Combining specific processing experience can effectively coordinate different noise reduction processes and help generate higher-quality images.
[0204] The following describes the image generation algorithm abstracted from the overall image generation process:
[0205] In the image generation algorithm disclosed in this application, after obtaining the noise map to be denoised, as well as each prompt text and mask image, through a series of image latent space transformations, cross-attention mechanism incorporation of conditions, and image latent space inverse transformations, the finally generated image is output. The following only takes the example that two target objects are expected to be included in the generated image, and explains the formulas based on in the process of denoising the original noise map to generate the initial noise map and in the process of denoising the target image to generate the final image:
[0206]
[0207]
[0208]
[0209]
[0210] Among them, it is expected to generate two target objects in the image, namely a foreground object and a background object. I0 and I1 are respectively the prompt words of the two input target objects (foreground prompt word and background prompt word), and respectively represent the noise images at the t-th moment (i.e., the first step of noise reduction, or the first round of noise reduction process) and the 0-th moment (i.e., the last step of noise reduction, or the last round of noise reduction process). F t→0 represents the foreground image (i.e., an intermediate image) obtained by denoising and transforming the noise image Nt initialized at the t-th moment according to the foreground prompt word I0; F t→1The noise image Nt initialized at time t is denoised according to the background prompt I1 to obtain a background image (i.e., an intermediate image); w is the denoising operation. In the denoising process of this application, it is necessary to rely on the prompt text of the target object, the object mask, and the noise image guided by the prompt text of the target object for denoising; M is the mask image of the foreground object, used to extract the corresponding foreground, 1 - M represents the mask image corresponding to the background object, and R is the residual, used to enhance the details in the fused image; Represents the foreground image denoised under the guidance of the foreground prompt word; Represents the background image denoised under the guidance of the background prompt word; Represents the final image generated after denoising; Is the initially obtained noise image; M⊙N t Characterizes the calculation of the Hadamard product between the noise image and the mask image.
[0211] Next, taking a specific application scenario as an example, the overall process of image generation will be described:
[0212] Refer to Figure 9A As shown, it is a schematic diagram of the overall process of image generation in an embodiment of this application. Next, in combination with the attached Figure 9A , the related image generation process will be described:
[0213] Step 901: The processing device prepares text prompt words and an image mask.
[0214] Specifically, when the processing device performs text generation, it can obtain existing prompt text, or obtain prompt text according to the text-to-image model.
[0215] For example, the (BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation) model is used to generate prompt text.
[0216] When generating the mask, the processing device can obtain the object mask through a mask generator.
[0217] For example, (SAM: Segment Anything) can be used to generate the object mask.
[0218] Step 902: The processing device determines the resolution of the generated image and the overall number of denoising steps.
[0219] It should be noted that when using the SD model for processing, the concept of time steps in the SD model is followed, so the denoising steps for the whole can be set. Among them, the denoising steps refer to the same content as the total number of denoising rounds described above. The denoising steps are also called the total number of denoising time steps. One denoising time step can be understood as one round of noise reduction.
[0220] Step 903: The processing device injects the encoded text prompt and the scaled image mask into the latent space generation network.
[0221] It should be noted that in the embodiments of the present application, the image denoising process involved is carried out in the latent space. Among them, the latent space generation network is used to realize denoising generation in the latent space. In the case of using the SD model for processing, it specifically refers to the U-Net part.
[0222] Step 904: The processing device initializes the normal distribution noise in the latent space generation network.
[0223] Step 905: The processing device performs multi-stage noise removal in the latent space according to the determined denoising steps.
[0224] Specifically, the multi-stage noise removal includes different denoising processes. Taking the example of generating two target objects in the image as expected, the processing involved in performing noise reduction using a three-stage noise reduction process will be described below:
[0225] The processing device selects a foreground prompt I0 for a foreground object and a background prompt I1 for a background object from the text prompt set according to the need to generate the target object, and obtains the corresponding mask image M. Among them, the dimension of the initialized original Gaussian noise map is (1, 4, h, w); the value of the 0th dimension of the Gaussian noise map is 1, indicating that the number of images is 1, the value of the 1st dimension is 4, indicating that there are 4 channels, the value of the 2nd dimension represents the height of the Gaussian noise map, and the value of the 3rd dimension represents the width of the Gaussian noise map.
[0226] After that, in the first noise reduction process, when performing each step of noise reduction, the following operations are carried out: When the processing device performs noise reduction to obtain the denoised noise map mid1_M based on the input original Gaussian noise map and mask image, as well as the prompt texts I0 and I1, where the dimension of mid1_M is (b, 4, h, w); b is the total number of target objects, which takes the value of 2 in this case; then add up all the elements in the 0th dimension of the image of mid1_M to obtain a matrix mid2_M with the dimension of (1, 4, h, w), and this matrix contains the foreground image content and background image content after noise reduction; it should be noted that when the preset number of noise reduction steps is not reached, the obtained matrix mid2_M is used as the denoised noise map mid1_M in the next noise reduction process, and continue to perform noise reduction on the newly determined mid1_M. After completing the first noise reduction process, output the mid2_M obtained in the last step of noise reduction.
[0227] Furthermore, in the second noise reduction process, when performing each step of noise reduction, the following operations are carried out: Repeat mid2_M to the same number as the total number of target objects, with the dimension of (b, 4, h, w); then divide it into several large regions, and use the already denoised regions to guide the non-denoised regions. Similarly, obtain the denoised noise map mid3_M output by the generation algorithm according to M, I0, and I1, where the dimension of mid3_M is (b, 4, h, w); b is the total number of target objects, which takes the value of 2 in this case; then add up all the elements in the 0th dimension of the image of mid3_M to obtain a matrix mid4_M with the dimension of (1, 4, h, w); it should be noted that when the preset number of noise reduction steps is not reached, the obtained matrix mid4_M is used as the denoised noise map mid3_M in the next noise reduction process, and continue to perform noise reduction on the newly determined mid3_M. After completing the second noise reduction process, output the mid4_M obtained in the last step of noise reduction.
[0228] Then, in the third noise reduction process, when performing each step of noise reduction, the following operations are carried out: Repeat the obtained mid4_M to the same number as the total number of target objects, with the dimension of (b, 4, h, w); perform denoising on the overall picture respectively according to M, I0, and I1, and then add up the noise reduction results (foreground pictures and background pictures) of all obtained target objects to obtain mid5_M with the dimension of (1, 4, h, w); it should be noted that when the preset number of noise reduction steps is not reached, the obtained matrix mid5_M is used as the denoised noise map mid4_M in the next noise reduction process, and continue to perform noise reduction on the newly determined mid4_M. After completing the third noise reduction process, output the mid5_M obtained in the last step of noise reduction.
[0229] It should be noted that the specific implementation methods of the three noise reduction processes have been described in detail in the above process, and this application will not expand on this.
[0230] Step 906: The processing device decodes the denoised picture in the latent space using a variational autoencoder to obtain the final picture from the latent space.
[0231] Specifically, the processing device obtains the final image through an inverse transformation of the latent space on the obtained noise reduction result.
[0232] Step 907: The processing device determines whether the generated picture meets the expectations. If so, it executes Step 908; otherwise, it returns to execute Step 901.
[0233] Specifically, the processing device can use the CLIP model to evaluate the content consistency between the text content and the image content to determine whether the generated image meets the expectations.
[0234] Step 908: The processing device outputs the generated picture.
[0235] Further, referring to Figure 9B As shown, it is a comparison schematic diagram of the image generation effect in the embodiment of the present application. Given a masked image and the prompt text of each target object to be generated, the image generated by the existing Multidiffusion model has a large difference from the image generated by the present application. The image generated by the existing Multidiffusion model is extremely chaotic, and the coherence of the image content is very poor, and it is impossible to generate the expected image; while adopting the technical solution proposed by the present application, in the processing process, it is necessary to construct a masked image and prompt text for each foreground object and the background object corresponding to the background as a whole, and then through a series of denoising operations, it is possible to generate a coherent, natural, and harmonious image, which can meet the personalized generation requirements.
[0236] Further, referring to Figure 9C As shown, it is a display schematic diagram of the generated image in the embodiment of the present application. Before specifically generating an image, after obtaining the overall masked image, foreground prompt words, and background prompt words on which the image generation task is based under the relevant technology, for each image generation task, the following operations are performed: The objects involved in the foreground prompt words and the objects involved in the background prompt words are respectively used as target objects, and the prompt text for each of the multiple target objects is constructed based on multiple prompt words; at the same time, for the determined multiple target objects, corresponding masked images are respectively constructed; then, based on the masked images of the multiple target objects and the prompt text of the multiple target objects obtained, the original noise map can be denoised to generate a high-quality image that meets the requirements.
[0237] In summary, in the technical solution proposed in this application, a controllable multi-stage diffusion generation algorithm is proposed for the characteristics of the controllable generation diffusion model. Compared with the denoising methods of existing solutions, this application adopts a multi-stage denoising method. By successively performing denoising in specific regions, block denoising, and overall denoising generation, the generated multi-object images are more natural and coherent. Moreover, this application proposes a controllable image generation method for desensitizing the divided regions. Through a multi-stage denoising process of denoising the masked region, using the denoised region to guide the denoising of the non-denoised region, and overall denoising, the final generated image result has no obvious defects as a whole and can meet the expected processing requirements.
[0238] Based on the same inventive concept, refer to Figure 10 As shown, it is a schematic logical structure diagram of an image generation device in an embodiment of this application. The image generation device 1000 includes a first acquisition unit 1001, a second acquisition unit 1002, and a noise reduction unit 1003. Among them,
[0239] The first acquisition unit 1001 is configured to acquire multiple pieces of reference data configured for multiple target objects included in a target image. Among them, one piece of reference data includes: a prompt text for describing one target object, and a mask image for indicating the shape and position of one target object in the target image;
[0240] The second acquisition unit 1002 is configured to acquire an initial noise map;
[0241] The noise reduction unit 1003 is configured to perform multi-round iterative noise reduction on the initial noise map, and use the denoised noise map obtained in the last round of noise reduction as the target image. Among them, in one round of iterative noise reduction process, the following operations are performed:
[0242] For the first region in the initial noise image, multiple prompt texts are respectively used for noise reduction guidance to obtain multiple initial denoising results of the first region, and according to multiple mask images and multiple initial denoising results of the first region, the final denoising result of the first region is obtained;
[0243] For the second region in the initial noise image, multiple prompt texts are respectively used for noise reduction guidance to obtain multiple initial denoising results of the second region, and according to multiple mask images and multiple initial denoising results of the second region, the final denoising result of the second region is obtained. Among them, there is an overlapping part between the first region and the second region, and the overlapping part in the second region has the same content as the overlapping part in the final denoising result of the first region;
[0244] According to the final denoising result of the first region and the non-overlapping part of the final denoising result of the second region with respect to the first region, a denoised noise map including multiple target objects is determined.
[0245] Optionally, after using the denoised noise map obtained in the last round of noise reduction as the target image, the noise reduction unit 1003 is further configured to:
[0246] Perform multiple rounds of iterative noise reduction on the target image, and use the denoised image obtained in the last round of noise reduction as the final image. During one round of iteration, the following operations are performed:
[0247] For the target image, perform noise reduction guidance using multiple prompt texts respectively to obtain multiple intermediate images of the target image, and obtain the denoised image of the target image according to the multiple mask images and the multiple intermediate images of the target image.
[0248] Optionally, after using the denoised image obtained in the last round of noise reduction as the final image, the noise reduction unit 1003 is further configured to:
[0249] Extract the corresponding comprehensive text features based on the multiple prompt texts, and extract the corresponding image features based on the final image;
[0250] Calculate the feature similarity between the comprehensive text features and the image features, and when the feature similarity reaches a set threshold, stop the current image generation process and output the final image.
[0251] Optionally, after calculating the feature similarity between the comprehensive text features and the image features, the noise reduction unit 1003 is further configured to:
[0252] When the feature similarity does not reach the set threshold, re-obtain the initial noise map and re-obtain multiple pieces of reference data configured for the target image;
[0253] Based on the re-obtained initial noise map and multiple pieces of reference data, re-execute the image generation process.
[0254] Optionally, the initial noise map is generated by the second acquisition unit 1002 in the following manner:
[0255] Obtain the original noise map;
[0256] Obtain multiple pieces of reference data configured for multiple target objects included in the target image. One piece of reference data includes: a prompt text for describing a target object, and a mask image for indicating the shape and position of a target object in the target image;
[0257] Perform multiple rounds of iterative noise reduction on the original noise map, and use the denoised noise map obtained in the last round of noise reduction as the initial noise map. During one round of iteration, the following operations are performed:
[0258] Based on the original noise map and multiple mask images, a single-object noise map for each target object is obtained, and for the single-object noise map of each target object, the following operations are performed: Using the prompt text of the target object, noise reduction guidance is performed on the single-object noise map of the target object to obtain a single-object noise-reduced map;
[0259] The contents of multiple single-object noise-reduced maps are superimposed to obtain a denoised noise map of the original noise map.
[0260] Optionally, multiple pieces of reference data are generated by the first acquisition unit 1001 in the following manner:
[0261] Determine multiple target objects expected to be generated in the target image, and for each target object, perform the following operations:
[0262] Based on the prompt words of the target object, generate prompt text, and in the pre-constructed object mask set, obtain the target object mask corresponding to the target object;
[0263] According to the preset resolution for the target image, the object mask of the target object, and the expected size and expected position of the target object in the target image, obtain the mask image of the target object, and use the prompt text and mask image of the target object as a piece of reference data.
[0264] Optionally, before obtaining the original noise map, the second acquisition unit 1002 is further used for:
[0265] Obtain the preset total number of denoising rounds;
[0266] According to the total number of denoising rounds and the preset round ratio, configure the first denoising round number used when denoising the original noise map to obtain the initial noise map, the second denoising round number used when denoising the initial noise map to obtain the target noise map, and the third denoising round number used when denoising the target image to obtain the final image.
[0267] Optionally, the regional positions of the first region and the second region in the initial noise map are obtained by the denoising unit 1003 in the following manner:
[0268] Obtain the preset regional window, sliding direction, sliding step size, and the initial position in the initial noise map;
[0269] When the regional window is at the initial position, determine the regional position covered by the regional window in the initial noise map as the regional position corresponding to the first region;
[0270] In the initial noise map, after sliding the sliding step size along the sliding direction from the initial position, determine the regional position covered by the current regional window in the initial noise map as the regional position corresponding to the second region.
[0271] Optionally, when obtaining the final denoising result of the first region based on multiple masked images and multiple initial denoising results of the first region, the denoising unit 1003 is configured to:
[0272] Perform the following operations on the initial denoising result guided by the prompt text of each target object: In the masked image of the target object, determine the masked content corresponding to the same region position as the first region, and perform a Hadamard product operation on the initial denoising result guided by the prompt text of the target object and the masked content to obtain the single-object denoising result corresponding to the target object;
[0273] Overlay the single-object denoising results of multiple target objects to obtain the final denoising result of the first region.
[0274] After introducing the image generation method and apparatus according to the exemplary embodiments of the present application, next, an electronic device according to another exemplary embodiment of the present application will be introduced.
[0275] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here.
[0276] Based on the same inventive concept as the above method embodiment, an electronic device is further provided in an embodiment of the present application. Refer to Figure 11 As shown, it is a schematic diagram of a hardware composition structure of an electronic device applying the embodiment of the present application. The electronic device 1100 may at least include a processor 1101 and a memory 1102. Among them, the memory 702 stores program code. When the program code is executed by the processor 1101, the processor 1101 is caused to execute the steps of any one of the above image generations.
[0277] In some possible implementation manners, the computing device according to the present application may at least include at least one processor and at least one memory. Among them, the memory stores program code. When the program code is executed by the processor, the processor is caused to execute the steps of the image generation according to various exemplary embodiments of the present application described above in this specification. For example, the processor may execute the steps as shown in Figure 3A shown.
[0278] In some possible embodiments, the electronic device according to the present application may include at least one processor and at least one memory. Among them, the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of image generation according to various exemplary embodiments of the present application described above in this specification. For example, the processor may execute as Figure 3A shown in the steps.
[0279] Based on the same inventive concept as the above method embodiments, various aspects of the image generation provided by the present application can also be implemented in the form of a program product, which includes program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps in the image generation according to various exemplary embodiments of the present application described above in this specification. For example, the electronic device may execute as Figure 3A shown in the steps.
[0280] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0281] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept. Therefore, the appended claims are intended to be interpreted to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.
[0282] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.
Claims
1. An image generation method, characterized in that, Including: Obtain multiple pieces of reference data configured for multiple target objects included in a target image, where one piece of reference data includes: a prompt text for describing one target object, and a mask image for indicating the shape and position of the one target object in the target image; Obtain an initial noise map; Perform multi-round iterative noise reduction on the initial noise map, and use the denoised noise map obtained in the last round of noise reduction as the target image. In the process of one-round iterative noise reduction, perform the following operations: For a first region in the initial noise image, perform noise reduction guidance using multiple prompt texts respectively to obtain multiple initial denoised results of the first region, and obtain a final denoised result of the first region according to multiple mask images and the multiple initial denoised results of the first region; For a second region in the initial noise image, perform noise reduction guidance using the multiple prompt texts respectively to obtain multiple initial denoised results of the second region, and obtain a final denoised result of the second region according to the multiple mask images and the multiple initial denoised results of the second region, where there is an overlapping part between the first region and the second region, and the overlapping part in the second region has the same content as the overlapping part in the final denoised result of the first region; Determine a denoised noise map containing the multiple target objects according to the final denoised result of the first region and the non-overlapping part of the final denoised result of the second region with respect to the first region.
2. The method according to claim 1, characterized in that After using the denoised noise map obtained in the last round of noise reduction as the target image, further include: Perform multi-round iterative noise reduction on the target image, and use the denoised image obtained in the last round of noise reduction as the final state image. In the process of one-round iteration, perform the following operations: For the target image, perform noise reduction guidance using the multiple prompt texts respectively to obtain multiple intermediate images of the target image, and obtain a denoised image of the target image according to the multiple mask images and the multiple intermediate images of the target image.
3. The method according to claim 2, wherein After using the denoised image obtained in the last round of noise reduction as the final state image, further include: Extract corresponding comprehensive text features based on the multiple prompt texts, and extract corresponding image features based on the final state image; Calculate the feature similarity between the comprehensive text features and the image features, and when the feature similarity reaches a set threshold, stop the current image generation process and output the final state image.
4. The method according to claim 3, wherein After calculating the feature similarity between the comprehensive text features and the image features, further include: When the feature similarity does not reach the set threshold, re-obtain the initial noise map and re-obtain multiple pieces of reference data configured for the target image; Based on the re-obtained initial noise map and the multiple pieces of reference data, re-execute the image generation process.
5. The method according to any one of claims 1 to 4, characterized in that, The initial noise map is generated in the following manner: Obtain an original noise map; Obtain multiple pieces of reference data configured for multiple target objects included in a target image. One piece of reference data includes: a prompt text for describing one target object, and a mask image for indicating the shape and position of the one target object in the target image; Perform multi-round iterative noise reduction on the original noise map, and use the denoised noise map obtained in the last round of noise reduction as the initial noise map. In one round of iteration, perform the following operations: According to the original noise map and multiple mask images, obtain the single-object noise map of each target object, and for the single-object noise map of each target object, perform the following operations: use the prompt text of the target object to perform noise reduction guidance on the single-object noise map of the target object to obtain a single-object denoised map; Overlay the contents of multiple single-object denoised maps to obtain the denoised noise map of the original noise map.
6. The method according to claim 5, wherein The multiple pieces of reference data are generated in the following manner: Determine multiple target objects expected to be generated in the target image, and for each target object, perform the following operations: Generate a prompt text based on the prompt word of the target object, and obtain the target object mask corresponding to the target object in a pre-constructed object mask set; According to the preset resolution of the target image, the object mask of the target object, and the expected size and expected position of the target object in the target image, obtain the mask image of the target object, and use the prompt text and mask image of the target object as one piece of reference data.
7. The method according to claim 5, characterized in that Before obtaining the original noise map, it further includes: Obtain the preset total number of noise reduction rounds; According to the total number of noise reduction rounds and the preset round ratio, configure the first number of noise reduction rounds used to denoise the original noise map to obtain the initial noise map, the second number of noise reduction rounds used to denoise the initial noise map to obtain the target noise map, and the third number of noise reduction rounds used to denoise the target image to obtain the final image.
8. The method according to any one of claims 1 to 4, characterized in that The regional positions of the first region and the second region in the initial noise map are obtained by the following method: Obtain the preset regional window, sliding direction, sliding step, and the initial position in the initial noise map; When the regional window is located at the initial position, determine the regional position covered by the regional window in the initial noise map as the regional position corresponding to the first region; In the initial noise map, after sliding the sliding step along the sliding direction from the initial position, determine the regional position covered by the current regional window in the initial noise map as the regional position corresponding to the second region.
9. The method according to any one of claims 1 to 4, characterized in that, The obtaining of the final denoising result of the first region according to the multiple mask images and the multiple initial denoising results of the first region includes: For the initial denoising result guided by the prompt text of each target object, perform the following operations: determine the mask content corresponding to the same regional position as the first region in the mask image of the target object, and perform a Hadamard product operation on the initial denoising result guided by the prompt text of the target object and the mask content to obtain the single-object denoising result corresponding to the target object; Overlay the individual denoising results of the multiple target objects to obtain the final denoising result of the first region.
10. An image generation device, characterized in that, It includes: A first acquisition unit for acquiring multiple pieces of reference data configured for multiple target objects included in a target image, where one piece of reference data includes: a prompt text for describing one target object, and a mask image for indicating the shape and position of the one target object in the target image; A second acquisition unit for acquiring an initial noise map; A noise reduction unit for performing multiple rounds of iterative noise reduction on the initial noise map and using the denoised noise map obtained in the last round of noise reduction as the target image. In one round of iterative noise reduction, the following operations are performed: For the first region in the initial noise image, perform noise reduction guidance using multiple prompt texts respectively to obtain multiple initial denoising results of the first region, and obtain the final denoising result of the first region according to multiple mask images and the multiple initial denoising results of the first region; For the second region in the initial noise image, perform noise reduction guidance using the multiple prompt texts respectively to obtain multiple initial denoising results of the second region, and obtain the final denoising result of the second region according to the multiple mask images and the multiple initial denoising results of the second region, where there is an overlapping part between the first region and the second region, and the overlapping part in the second region has the same content as the overlapping part in the final denoising result of the first region; Determine the denoised noise map containing the multiple target objects according to the final denoising result of the first region and the non-overlapping part of the final denoising result of the second region with respect to the first region.
11. The device according to claim 10, characterized in that, After using the denoised noise map obtained in the last round of noise reduction as the target image, the noise reduction unit is further configured to: Perform multiple rounds of iterative noise reduction on the target image and use the denoised image obtained in the last round of noise reduction as the final state image. In one round of iteration, the following operations are performed: For the target image, perform noise reduction guidance using the multiple prompt texts respectively to obtain multiple intermediate images of the target image, and obtain the denoised image of the target image according to the multiple mask images and the multiple intermediate images of the target image.
12. The device according to claim 10 or 11, characterized in that, The initial noise map is generated by the acquisition unit in the following manner: Acquire an original noise map; Acquire multiple pieces of reference data configured for multiple target objects included in a target image, where one piece of reference data includes: a prompt text for describing one target object, and a mask image for indicating the shape and position of the one target object in the target image; Perform multiple rounds of iterative noise reduction on the original noise map and use the denoised noise map obtained in the last round of noise reduction as the initial noise map. In one round of iteration, the following operations are performed: Based on the original noise map and multiple mask images, a single-object noise map of each target object is obtained, and for the single-object noise map of each target object, the following operations are performed: Using the prompt text of the target object, noise reduction guidance is performed on the single-object noise map of the target object to obtain a single-object noise reduction map; The contents of multiple single-object noise reduction maps are superimposed to obtain a denoised noise map of the original noise map.
13. The device according to claim 10 or 11, characterized in that, The regional positions of the first region and the second region in the initial noise map are obtained by the noise reduction unit in the following manner: Obtain a preset regional window, sliding direction, sliding step, and an initial position in the initial noise map; When the regional window is located at the initial position, determine the regional position covered by the regional window in the initial noise map as the regional position corresponding to the first region; In the initial noise map, after sliding the sliding step along the sliding direction from the initial position, determine the regional position covered by the current regional window in the initial noise map as the regional position corresponding to the second region.
14. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method described in any one of claims 1-9 is implemented.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, the method described in any one of claims 1-9 is implemented.
16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1-9 is implemented.