Image generation method and electronic equipment
By obtaining constraint information for iterative denoising processing and information enhancement, the problem of image mismatch in graffiti generation is solved, and the controllable generation and high matching degree of the target image are achieved.
Patent Information
- Application Number
- CN202510572161.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-18
AI Technical Summary
In the graffiti generation scenario, the images generated by the existing diffusion model do not match the graffiti images input by the user, lack controllability, and it is difficult to meet user expectations.
By obtaining constraint information, including constraint text and constraint images, iterative denoising processing is performed, and the intermediate image is enhanced using compensation information and constraint text to generate a target image, so that the content of the target object matches the description content of the constraint text, and the outline matches the object outline of the constraint image.
The controllability and refined control capabilities of image generation are improved, and the generated target images are more consistent with user expectations, and the outline and content match are high.
Smart Images

Figure CN120339111A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to an image generation method and an electronic device. Background Art
[0002] Diffusion models are one of the important technologies in the field of image generation in recent years. Diffusion models gradually transform the original image into a pure noise state by injecting Gaussian noise into the original image data during the training process, and then train a neural network to learn how to reverse this process to restore the original image.
[0003] Diffusion models have been widely applied in multiple fields such as image generation, restoration, and super-resolution. However, in some scenarios, for example, in the scenario of graffiti generation, there may be a situation where the generated image does not match the graffiti image input by the user. Therefore, how to improve the controllability of image generation during the generation process so that the generated target image meets the user's expectations is a direction worthy of exploration. Summary of the Invention
[0004] In view of this, the present disclosure provides an image generation method and an electronic device.
[0005] In one aspect of the present disclosure, an image generation method is provided, including: obtaining constraint information, where the constraint information includes constraint text and a constraint image; performing iterative denoising processing on a noise image based on the constraint information to generate a target image of a target object; the object content of the target object matches the content described in the constraint text, and the object contour of the target object matches the object contour of the constraint image; where performing one denoising process includes: enhancing the information of the constraint information and the intermediate image of the previous denoising process to obtain compensation information, and the intermediate image is obtained by performing at least one denoising process on the noise image; and performing denoising processing on the intermediate image based on the compensation information and the constraint text.
[0006] According to an embodiment of the present disclosure, performing denoising processing on the intermediate image based on the compensation information and the constraint text includes: performing multiple feature restorations on multi-modal features based on the compensation information; the multi-modal features are obtained by processing the constraint text and the intermediate image.
[0007] According to an embodiment of the present disclosure, one denoising process includes I feature restoration processes; performing one feature restoration process includes: using a target compensation feature to perform feature restoration on the i-th multi-modal feature to obtain the (i + 1)-th multi-modal feature; the target compensation feature is a compensation feature in the compensation information that has a mapping relationship with the i-th multi-modal feature, and i is an integer greater than or equal to 1 and less than I.
[0008] According to an embodiment of the present disclosure, feature restoration is performed on the i-th multimodal feature using a target compensation feature to obtain the (i + 1)-th multimodal feature, including: performing information compensation on the i-th multimodal feature using the target compensation feature to obtain the i-th multimodal enhanced feature; performing multimodal feature processing on the i-th multimodal enhanced feature to obtain the (i + 1)-th multimodal feature.
[0009] According to an embodiment of the present disclosure, information enhancement is performed on the constraint information and the intermediate image of the previous denoising process to obtain compensation information, including: performing information fusion on the constraint information and the intermediate image of the previous denoising process to obtain fusion information; performing feature extraction of different feature scales on the fusion information to obtain compensation information composed of multiple compensation features.
[0010] According to an embodiment of the present disclosure, the above method further includes: extracting contour information from the original image to obtain a constraint image.
[0011] On the other hand, the present disclosure provides an image generation method, including: obtaining constraint information, where the constraint information includes a constraint text and a constraint image; using the denoising module of the image generation model to perform iterative denoising processing on the noise image based on the constraint information to generate a target image of a target object; the object content of the target object matches the content described in the constraint text, and the contour of the target object matches the object contour of the constraint image; where performing one denoising process includes: using the compensation module of the image generation model to perform information enhancement on the constraint information and the intermediate image of the previous denoising process to obtain compensation information, and the intermediate image is obtained by performing at least one denoising process on the noise image; using the denoising unit in the denoising module to perform denoising processing on the intermediate image based on the compensation information and the constraint text.
[0012] According to an embodiment of the present disclosure, the above method further includes: obtaining a sample target image and a sample constraint text, where the content described in the sample constraint text matches the object content of the sample object in the sample target image; extracting the contour information of the sample target image to obtain a sample constraint image; using the sample noise image, the sample constraint information formed by the sample constraint text and the sample constraint image, and the sample target image to train an initial image generation model to obtain an image generation model; the sample noise image is obtained by adding noise to the sample target image, or the sample noise image is a random noise image.
[0013] According to an embodiment of the present disclosure, using the sample noise image, the sample constraint information formed by the sample constraint text and the sample constraint image, and the sample target image to train an initial image generation model to obtain an image generation model, including: using the denoising module of the initial image generation model to perform iterative denoising processing on the sample noise image based on the sample constraint information to generate a sample restored image of a sample object;
[0014] Among them, performing one denoising process includes: using the initial compensation module of the initial image generation model to enhance the information of the sample constraint information and the sample intermediate image of the previous denoising process to obtain sample compensation information, where the sample intermediate image is obtained by performing at least one denoising process on the sample noise image; using the denoising unit in the denoising module to perform denoising processing on the sample intermediate image based on the sample compensation information and the sample constraint text; freezing the parameters of the denoising module, and adjusting the parameters of the initial compensation module based on the sample restored image and the sample target image to obtain an image generation model.
[0015] On the other hand, the present disclosure provides an electronic device, including: one or more processors; a memory for storing one or more programs, where when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the steps of the above method.
[0016] On the other hand, the present disclosure further provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the above control method.
[0017] On the other hand, the present disclosure further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above control method is implemented.
[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:
[0020] Figure 1 Schematically shows an exemplary system architecture to which the image generation method and apparatus according to the embodiments of the present disclosure can be applied;
[0021] Figure 2 Schematically shows a flowchart of the image generation method according to the embodiments of the present disclosure;
[0022] Figure 3 Schematically shows a data flow diagram of a single denoising process according to the embodiments of the present disclosure;
[0023] Figure 4 Schematically shows a flowchart of the image generation method according to another embodiment of the present disclosure;
[0024] Figure 5Schematically shows a flowchart of a single denoising process according to an embodiment of the present disclosure;
[0025] Figure 6 Schematically shows a network structure diagram of an image generation method according to an embodiment of the present disclosure;
[0026] Figure 7 Schematically shows an effect diagram of a target image generated by an image generation method according to an embodiment of the present disclosure;
[0027] Figure 8 Schematically shows a comparison diagram of the effects of the target images generated by this embodiment and related embodiments; and
[0028] Figure 9 Schematically shows a block diagram of an electronic device suitable for implementing the above-described method according to an embodiment of the present disclosure. Detailed implementation manners
[0029] To make the objectives, technical solutions, and advantages of the present disclosure more clear and understandable, the present disclosure will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0030] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0031] All terms used herein, including technical and scientific terms, have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0032] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art. For example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C. In the case of using expressions such as "at least one of A, B, or C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art. For example, "a system having at least one of A, B, or C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C.
[0033] It should also be noted that the directional terms mentioned in the embodiments, such as "upper", "lower", "front", "rear", "left", "right", etc., are only references to the directions in the accompanying drawings and are not used to limit the protection scope of the present disclosure. Throughout the accompanying drawings, the same elements are represented by the same or similar reference numerals. When it may cause confusion in the understanding of the present disclosure, the conventional structures or configurations will be omitted.
[0034] Figure 1 An exemplary system architecture to which the image generation method and apparatus according to embodiments of the present disclosure can be applied is schematically shown.
[0035] It should be noted that Figure 1 What is shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios. For example, in another embodiment, the exemplary system architecture to which the image generation method and apparatus can be applied may include a terminal device, but the terminal device can implement the image generation method and the image generation apparatus provided by the embodiments of the present disclosure without interacting with the server.
[0036] As Figure 1 shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0037] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples).
[0038] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0039] The server 105 may be a server that provides various services, such as a background management server (only for example) that supports the content browsed by users using the terminal devices 101, 102, and 103. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0040] It should be noted that the image generation method provided by the embodiments of the present disclosure can generally be executed by the first terminal device 101, the second terminal device 102, and the third terminal device 103. Correspondingly, the image generation device provided by the embodiments of the present disclosure can also be arranged in the first terminal device 101, the second terminal device 102, and the third terminal device 103.
[0041] Alternatively, the image generation method provided by the embodiments of the present disclosure can generally also be executed by the server 105. Correspondingly, the image generation device provided by the embodiments of the present disclosure can generally be arranged in the server 105. The image generation method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the image generation device provided by the embodiments of the present disclosure can also be arranged in a server or a server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0042] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the servers in
[0043] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.
[0044] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, disclosure, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good customs.
[0045] It should be noted that the serial numbers of the respective operations in the following methods are only used as representations of the operations for description, and should not be regarded as indicating the execution order of the respective operations. Unless explicitly stated, the method does not need to be executed completely in the order shown.
[0046] Figure 2 Schematically shows a flowchart of the image generation method according to an embodiment of the present disclosure.
[0047] As Figure 2 described above, the method includes operations S210 to S222.
[0048] In operation S210, constraint information is obtained, where the constraint information includes constraint text and a constraint image.
[0049] The constraint text refers to a text description of the target object to be generated through natural language. The target object can be any object, living being, landscape, etc. The constraint text can be used to constrain information such as the theme, attributes, and scene of the generated content. For example, the constraint text can be "Generate a cute cat" or "Generate a red apple".
[0050] The constraint image refers to a reference image in the process of image generation, and the constraint image can be used to constrain the structure, contour, local details, etc. of the generated target object. The constraint image can be a contour map, a scribble map, etc. of the target object.
[0051] In operation S220, iterative denoising processing is performed on the noise image based on the constraint information to generate a target image of the target object. The object content of the target object matches the content described in the constraint text, and the contour of the target object matches the object contour of the constraint image.
[0052] The noise image can be a random tensor generated by pure Gaussian noise. The present disclosure is not limited thereto, and the noise image can also be generated based on salt-and-pepper noise, Poisson noise, etc.
[0053] The denoising processing refers to directionally removing random noise from the noise image to generate image features highly correlated with the constraint information. The iterative denoising processing refers to continuously performing the denoising processing a specified number of times starting from the noise image to obtain the target image of the target object after the denoising processing. For example, the denoising processing is performed on the noise image T times, where the t-th denoising processing is performed on the intermediate image obtained after the (t - 1)-th denoising processing; where T≥1 and t≥1.
[0054] According to an embodiment of the present disclosure, for operation S220, performing one denoising processing includes: operations S221 to S222.
[0055] In operation S221, information enhancement is performed on the constraint information and the intermediate image of the previous denoising processing to obtain compensation information, and the intermediate image is obtained by performing at least one denoising processing on the noise image.
[0056] In operation S222, denoising processing is performed on the intermediate image based on the compensation information and the constraint text.
[0057] Information enhancement refers to, in each denoising process, by analyzing the relationship between the intermediate image and the constraint information, dynamically strengthening or correcting local or detailed features in the generation process to obtain compensation information for enhanced local features.
[0058] Compensation information for enhancing the contour consistency between the intermediate image and the constraint image can be generated based on the differences between the intermediate image and the constraint image. Compensation information for enhancing the semantic consistency between the intermediate image and the constraint text can also be obtained according to the differences between the intermediate image and the text constraint.
[0059] In the first denoising process, information enhancement can be performed on the constraint information and the noisy image to obtain compensation information. The noisy image is denoised using the compensation information and the constraint text to obtain the intermediate image after the first denoising process.
[0060] In the t-th denoising process, information enhancement can be performed on the intermediate image after the (t - 1)-th denoising process and the constraint information to obtain compensation information. The intermediate image after the (t - 1)-th denoising process is denoised based on the compensation information and the constraint text to obtain the intermediate image after the t-th denoising process.
[0061] According to an embodiment of the present disclosure, in each denoising process, information enhancement is performed on the constraint information and the intermediate image of the previous denoising process to obtain compensation information. The intermediate image is denoised using the compensation information and the constraint text. Since the intermediate image of each denoising process is updated, dynamic compensation information can be generated. Denoising the intermediate image again using the dynamic compensation information and the constraint text can correct the intermediate image in each denoising process, thereby enhancing the controllability of image generation and improving the degree of fine control. And since the constraint information includes the constraint text and the constraint image, the generated target image matches both the content described in the constraint text and the object contour of the constraint image, so that the generated target image meets the user's expectations.
[0062] According to an embodiment of the present disclosure, before performing operation S210, contour information extraction may also be included on the original image to obtain a constraint image.
[0063] The original image refers to the initial input image that has not been processed. The original image can be an image containing complete visual information, such as color, texture, light and shade, etc. The original image can be a photographed photo, a hand-drawn graffiti, a scanned image, or other forms of digital images.
[0064] The contour information of the original image is extracted to obtain a constraint image. For example, the contour information of the original image can be extracted through a feature extraction network, which can be a convolutional neural network, a residual network, or other dedicated edge detection networks, etc. The constraint image can be a binary image or a contour image to highlight the contour information of the target object.
[0065] By extracting the contour information of the original image to obtain a constraint image, the contour features of the constraint image can be strengthened, so that the edges of the generated image follow the contour constraints, improving the controllability of image generation.
[0066] According to an embodiment of the present disclosure, information enhancement of the constraint information and the intermediate image of the previous denoising process to obtain compensation information may include: performing information fusion on the constraint information and the intermediate image of the previous denoising process to obtain fusion information; performing feature extraction on the fusion information at different feature scales to obtain compensation information composed of multiple compensation features.
[0067] Performing information fusion on the constraint information and the intermediate image of the previous denoising process to obtain fusion information may be to perform fusion processing on the constraint image and the intermediate image of the previous denoising process to obtain image fusion information, and then fuse the image fusion information and the constraint text to obtain fusion information.
[0068] Specifically, the constraint image can be encoded to obtain a first encoded feature, the intermediate image can be encoded to obtain a second encoded feature, and after adding the first encoded feature and the second encoded feature, an image fusion feature is obtained. Feature extraction is performed on the constraint text to obtain text features; the image fusion feature and the text features are fused to obtain fusion information.
[0069] In some embodiments, an image encoder can be used to encode the constraint image to obtain a first encoded feature of the constraint image. A text encoder can be used to encode the constraint text to obtain text features. The present disclosure does not limit the types of the image encoder and the text encoder.
[0070] In some embodiments, multiple text encoders can also be used to encode the constraint text to obtain multiple text features. The first text feature and the second text feature are determined based on the multiple text features. The first text feature can be obtained by splicing multiple text features. The second text feature can be obtained by pooling one or more text features. The information granularity of the first text feature is higher than that of the second text feature.
[0071] The time step feature can also be obtained, and the time step feature and the second text feature are fused with the second text feature to obtain a third text feature. The image fusion feature, the first text feature, and the third text feature are fused to obtain fusion information.
[0072] Based on the attention mechanism, feature extraction can be performed on the fused information at different feature scales to obtain compensation information composed of multiple compensation features.
[0073] In some embodiments, feature extraction can be performed on the fused information by using different attention parameters respectively, so as to obtain multiple compensation features of different scales. The attention parameters can include the number of attention heads, the number of attention layers, attention weights, etc.
[0074] According to the embodiments of the present disclosure, information with different spatial resolutions or semantic hierarchical representations can be extracted from the fused information through feature extraction at different scales, so as to capture diverse information from local details to global structures. Thereby avoiding the loss of details or the blurring of structures caused by single-scale features.
[0075] According to the embodiments of the present disclosure, denoising the intermediate image based on the compensation information and the constraint text may include: performing multiple feature recoveries on the multimodal features based on the compensation information; the multimodal features are obtained by processing the constraint text and the intermediate image.
[0076] The multimodal features can be obtained by fusing and processing the constraint text and the intermediate image. Specifically, the constraint text can be used as the Key and Value, and the intermediate image can be used as the Query, and through the attention mechanism, fusion and adaptive normalization processing are performed to obtain the multimodal features.
[0077] In the case of performing the first denoising process, the multimodal features can also be obtained by fusing and processing the constraint text and the noisy image. The constraint text can be used as the Key and Value, and the noisy image can be used as the Query, and through the attention mechanism, fusion and adaptive normalization processing are performed to obtain the multimodal features.
[0078] Feature recovery refers to using the compensation features to iteratively correct the multimodal features to gradually approximate the true feature distribution of the target image. For example, through feature recovery processing, contour deviation, semantic loss, etc. can be corrected.
[0079] Performing multiple feature recoveries on the multimodal features based on the compensation information can be performing multiple feature recoveries on the multimodal features based on one compensation feature. The present disclosure is not limited thereto, and multiple compensation features of different scales can also be used to perform multiple feature recoveries on the multimodal features according to a preset mapping relationship.
[0080] According to the embodiments of the present disclosure, by performing multiple feature recoveries on the multimodal features based on the compensation information, the generation direction can be dynamically corrected through the compensation information, avoiding the deviation of the intermediate image from the constraint, and gradually generating a high-quality image that meets the constraint.
[0081] Figure 3 A flowchart of performing single denoising processing according to an embodiment of the present disclosure is schematically shown.
[0082] As Figure 3 shown, the constraint text 310 and the intermediate image 320 of the (t - 1)-th removal process are fused to obtain the multi-modal feature 340. The constraint information 330 and the intermediate image 320 of the (t - 1)-th removal process are fused to obtain the compensation information 330. Based on the compensation information 350, the multi-modal feature 340 is subjected to multiple feature restorations to obtain the intermediate image 360 of the t-th denoising process. Thus, one denoising process is completed.
[0083] According to an embodiment of the present disclosure, one denoising process may include I feature restoration processes; performing one feature restoration process may include: using the target compensation feature to perform feature restoration on the i-th multi-modal feature to obtain the (i + 1)-th multi-modal feature; the target compensation feature is a compensation feature among the multiple compensation features included in the compensation information that has a mapping relationship with the i-th multi-modal feature, and i is an integer greater than or equal to 1 and less than I.
[0084] Before using the target compensation feature to perform feature restoration on the i-th multi-modal feature, the dimension of the target compensation feature may also be converted so that the dimension of the target compensation feature is the same as that of the i-th multi-modal feature.
[0085] The number of compensation features may be less than or equal to I, and one compensation feature may map multiple consecutive multi-modal features. For example, one denoising process may include 24 feature restoration processes, the number of compensation features may be 6, and each compensation feature maps 4 consecutive multi-modal features respectively. The present disclosure is not limited thereto, and the number of compensation features may be set according to requirements.
[0086] According to an embodiment of the present disclosure, by dividing one denoising process into multiple restoration processes and combining the multi-modal feature with the compensation feature in each restoration process, the cumulative error can be reduced. And the number of compensation features is adjustable, which can adapt to different scenario requirements and has strong scalability.
[0087] According to an embodiment of the present disclosure, using the target compensation feature to perform feature restoration on the i-th multi-modal feature to obtain the (i + 1)-th multi-modal feature includes: using the target compensation feature to perform information compensation on the i-th multi-modal feature to obtain the i-th multi-modal enhanced feature; performing multi-modal feature processing on the i-th multi-modal enhanced feature to obtain the (i + 1)-th multi-modal feature.
[0088] Information compensation means fusing the target compensation feature with the i-th multi-modal feature to enhance or correct the deficiencies of the i-th multi-modal feature.
[0089] The information of the $i$-th multimodal feature is compensated by using the target compensation feature to obtain the $i$-th multimodal enhanced feature. For example, the target compensation feature can be concatenated to the $i$-th multimodal feature to obtain the $i$-th multimodal enhanced feature. This disclosure is not limited thereto. The target compensation feature and the $i$-th multimodal feature can also be fused by element-wise addition or attention weighting to obtain the $i$-th multimodal enhanced feature.
[0090] The multimodal feature processing of the $i$-th multimodal enhanced feature can be performed by cross-attention and self-attention to further extract features from the multimodal enhanced feature, obtaining the $(i + 1)$-th multimodal feature.
[0091] Figure 4 A flowchart of an image generation method according to another embodiment of the present disclosure is schematically shown.
[0092] The method includes operations S410 to S422.
[0093] In operation S410, constraint information is obtained, where the constraint information includes constraint text and constraint images.
[0094] In operation S420, using the denoising module of the image generation model, iterative denoising processing is performed on the noise image based on the constraint information to generate a target image of the target object; the object content of the target object matches the content described in the constraint text, and the object contour of the target object matches the object contour of the constraint image.
[0095] For operation S410, performing one denoising process includes: operations S421 to S422
[0096] In operation S421, using the compensation module of the image generation model, information enhancement is performed on the constraint information and the intermediate image of the previous denoising process to obtain compensation information, where the intermediate image is obtained by performing at least one denoising process on the noise image.
[0097] In operation S422, using the denoising unit in the denoising module, denoising processing is performed on the intermediate image based on the compensation information and the constraint text.
[0098] In an embodiment of the present disclosure, the image generation model includes a denoising module and a compensation module, and the output of the compensation module is used as the input of the denoising module. Among them, the denoising module can be constructed with the DiT architecture as the backbone. The compensation module can be a module constructed based on the transform architecture. This disclosure is not limited thereto. The compensation module can also be constructed based on a convolutional neural network, a generative adversarial network, etc.
[0099] In the early stage, most diffusion models adopted the convolutional-based UNet (U-shaped Network) architecture. Due to its strong local feature extraction ability, the convolutional-based UNet architecture has been widely used. However, with the development of the Transformer architecture, based on the advantages of Transformer in long-range dependence modeling and global feature capture, the DiT (Diffusion Transformer) architecture based on Transformer has gradually replaced the UNet architecture and achieved significant performance improvements in various generation tasks. However, the DiT architecture captures global dependence relationships based on the self-attention mechanism of Transformer and lacks the local feature extraction ability provided by the convolutional layer in UNet, which reduces the controllability of the DiT architecture in the image generation process.
[0100] According to an embodiment of the present disclosure, in each denoising process, by using a compensation module to enhance the information of the constraint information and the intermediate image of the previous denoising process, compensation information is obtained, and then the compensation information is input into the denoising module to denoise the intermediate image by using the compensation information and the constraint text. Since the intermediate image of each denoising process will be updated, dynamic compensation information can be generated. Using the dynamic compensation information and the constraint text for denoising can correct the intermediate image in each denoising process, thereby enhancing the controllability of image generation. Applying the compensation module to the DiT architecture can improve the fine control of the DiT architecture for image generation.
[0101] Figure 5 Schematically shows a flowchart of a single denoising process according to an embodiment of the present disclosure.
[0102] As Figure 5 shown, the constraint information 530 and the intermediate image 520 of the (t - 1)-th denoising process are input into the compensation module M542 of the image generation model M540, and the compensation information is output; the constraint text 510, the noisy image 520, and the compensation information are input into the denoising module M541 of the image generation model M540, and the noisy image is denoised based on the compensation information and the constraint text to obtain the intermediate image 550 of the t-th denoising process.
[0103] According to an embodiment of the present disclosure, using the compensation module of the image generation model to enhance the information of the constraint information and the intermediate image to obtain the compensation information may include: performing information fusion on the constraint information and the intermediate image to obtain fusion information; using multiple compensation units in the compensation module to respectively perform feature extraction on the fusion information at different feature scales to obtain the compensation information composed of multiple compensation features.
[0104] Multiple compensation units can be configured with different attention parameters, so that feature extraction of different feature scales can be performed on the fused information.
[0105] According to an embodiment of the present disclosure, using the denoising unit in the denoising module, the intermediate image is denoised based on the compensation information and the constraint text, including at least one of the following denoising processes: information compensation is performed based on the target compensation feature output by the compensation module and the multimodal feature output by the i-th denoising unit to obtain the i-th multimodal feature. The target compensation feature is output by the target compensation unit mapped to the i-th denoising unit; the input data of the i-th denoising unit is obtained by processing the constraint text and the intermediate image, and i is an integer greater than or equal to 1.
[0106] The denoising module includes multiple denoising units, and the number of denoising units is greater than or equal to the number of compensation units. There is a mapping relationship between the compensation unit and the denoising unit. For example, the number of denoising units can be set to 24, and the number of compensation units can be set to 6. Each compensation unit can correspond to 4 consecutive denoising units.
[0107] According to an embodiment of the present disclosure, feature restoration of the i-th multimodal feature using the target compensation feature output by the compensation module to obtain the (i + 1)-th multimodal feature may include: performing information compensation on the i-th multimodal feature using the target compensation feature output by the compensation module to obtain the i-th multimodal enhanced feature; using the (i + 1)-th denoising unit in the denoising module to perform multimodal feature processing on the i-th multimodal enhanced feature to obtain the (i + 1)-th multimodal feature.
[0108] Figure 6 Schematically shows the network structure diagram of the image generation method according to an embodiment of the present disclosure.
[0109] As Figure 6 shown, when performing the first denoising process, after fusing the noise image 620 and the constraint image 630, they are input into the compensation module M642 of the image generation model M640 together with the constraint text 610. The compensation module M642 includes multiple compensation units, such as the first supplementary unit, the second supplementary unit, the n-th compensation unit, and multiple compensation features of different scales are output using the multiple compensation units.
[0110] The constrained text 610 and the noisy image 620 are input into the denoising module M641 of the image generation model M640. The denoising module M641 includes multiple denoising units, such as the first denoising unit, the second denoising unit, and the i-th denoising unit. There is a mapping relationship between the multiple denoising units and the multiple compensation units. The first denoising unit is used to process the constrained text 610 and the noisy image 620 to obtain the first multi-modal feature. The first compensation feature output by the first compensation unit that has a mapping relationship with the first denoising unit is input into the first denoising unit, and the first multi-modal feature is compensated with the first supplementary feature to obtain the first multi-modal enhanced feature. The first multi-modal enhanced feature is input into the second denoising unit to obtain the second multi-modal feature.
[0111] The operations of compensating the current multi-modal feature with the compensation feature output by the compensation unit that has a mapping relationship with the current denoising unit to obtain the multi-modal enhanced feature, and inputting the multi-modal enhanced feature into the next denoising unit for processing are repeatedly executed until a denoising process is completed, and the intermediate image 650 is output.
[0112] The noisy image is replaced with the intermediate image 650, and the above steps are repeatedly executed and multiple iterative denoising processes are performed. Based on the intermediate image of the last denoising process, the target image 660 is obtained.
[0113] Figure 7 The effect diagram of the target image generated by the image generation method according to the present disclosure is schematically shown.
[0114] As Figure 7 shown, in Examples 1 to 5, the generated target images all have a high degree of matching with the outline of the graffiti image, and the generated image content also has a high degree of matching with the description content of the constrained text. It shows that the image generation method of the embodiments of the present disclosure can realize the controllable generation of the target image based on the constrained text and the constrained image, and obtain the target image that meets the expectations.
[0115] Figure 8 The effect comparison diagram of the target images generated by this embodiment and related embodiments is schematically shown.
[0116] As Figure 8 shown, the effect diagrams of the related embodiments are generated without the function of the compensation module of the embodiments of the present disclosure. Compared with the generation effect of the related embodiments, the image outline of the target image of this embodiment is more fitting to the outline of the graffiti image, while there is a large deviation between the image outline of the target image of the related embodiments and the graffiti image. It shows that the image generation method of the embodiments of the present disclosure has better fine control ability.
[0117] According to an embodiment of the present disclosure, an image generation model can be trained through the following method: obtaining a sample target image and a sample constraint text, where the content described in the sample constraint text matches the object content of the sample object in the sample target image. Extracting the contour information of the sample target image to obtain a sample constraint image. Training an initial image generation model using the sample noise image, the sample constraint information formed by the sample constraint text and the sample constraint image, and the sample target image to obtain an image generation model; the sample noise image is obtained by adding noise to the sample target image, or the sample noise image is a random noise image.
[0118] The sample target image refers to the real image in the training data and can be used as a label to guide the image generation model to generate an image consistent with the sample target image.
[0119] The sample constraint text refers to the natural language text describing the content of the sample target image and is used to provide semantic conditions to guide the model to generate an image that conforms to the text description.
[0120] The sample constraint image is an image generated by extracting the contour information of the sample target image and can be, for example, a binary image or a contour image. It is used to provide structural and contour constraints to ensure that the geometric shape of the generated image is consistent with the sample constraint image.
[0121] Training the initial image generation model using the sample noise image, the sample constraint information formed by the sample constraint text and the sample constraint image, and the sample target image to obtain an image generation model can include: training the denoising module of the initial image generation model using the sample constraint text and the sample target image.
[0122] The text encoder of a pre-trained model can be used to map the sample constraint text to a sample text feature vector, and the image encoder can be used to map the sample target image to a sample image feature. Inputting the sample text feature and the sample image feature into the denoising module. In the forward diffusion process, noise is gradually added to the sample image feature to obtain a sample noise image. In the reverse process, based on the attention mechanism, the sample text feature and the sample noise image are fused, and the noise is predicted. Training is performed with the mean squared error between the predicted noise and the real noise as the loss function, and the parameters of the denoising module are adjusted with the goal of minimizing the mean squared error to complete the training of the denoising module.
[0123] After completing the training of the denoising module, the compensation module of the initial image generation model is then trained using the sample noise image, the sample constraint information formed by the sample constraint text and the sample constraint image, and the sample target image.
[0124] According to an embodiment of the present disclosure, an initial image generation model is trained using a sample noise image, sample constraint information formed by sample constraint text and a sample constraint image, and a sample target image to obtain an image generation model, including: using a denoising module of the initial image generation model to perform iterative denoising processing on the sample noise image based on the sample constraint information to generate a sample restored image of a sample object. Wherein, performing one denoising process includes: using an initial compensation module of the initial image generation model to perform information enhancement on the sample constraint information and a sample intermediate image of the previous denoising process to obtain sample compensation information, and the sample intermediate image is obtained by performing at least one denoising process on the sample noise image; using a denoising unit in the denoising module to perform denoising processing on the sample intermediate image based on the sample compensation information and the sample constraint text; freezing the parameters of the denoising module, and based on the sample restored image and the sample target image, adjusting the parameters of the initial compensation module to obtain the image generation model.
[0125] Using an initial compensation module of the initial image generation model to perform information enhancement on the sample constraint information and a sample intermediate image of the previous denoising process to obtain sample compensation information may include: performing information fusion on the sample constraint image and the sample intermediate image of the previous denoising process to obtain sample fusion information; inputting the sample fusion information into a plurality of initial compensation units of the initial compensation module, and using the plurality of initial compensation units to perform feature extraction on the sample fusion information at different scales to obtain a plurality of sample compensation features as the sample compensation information.
[0126] Inputting the sample text feature, the sample image feature, and the sample compensation information into the denoising module. In the forward diffusion process, gradually adding noise to the sample image feature to obtain a sample noise image. In the reverse process, based on the attention mechanism, fusing the sample text feature, the sample compensation information, and the sample noise image, and predicting the noise. Using the mean square error between the predicted noise and the real noise as the loss function for training, aiming to minimize the mean square error, freezing the parameters of the denoising module, and adjusting the parameters of the initial compensation module to complete the training and obtain the image generation model.
[0127] According to an embodiment of the present disclosure, through a phased training strategy, that is, first completing the training of the denoising module, then freezing the denoising module, and training the compensation module, it is possible to avoid mutual interference between modules and joint overfitting, and improve the training efficiency.
[0128] Figure 9 A block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present disclosure is schematically shown. Figure 9 The electronic device shown is only an example and should not bring any limitations to the functions and usage scopes of the embodiments of the present disclosure.
[0129] The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, personal digital processors, cellular telephones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only by way of example and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0130] As Figure 9 shown, the device 900 includes a computing unit 901 which can perform various appropriate actions and processes in accordance with a computer program stored in a read only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0131] A plurality of components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0132] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the image generation method. For example, in some embodiments, the image generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the image generation method described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the image generation method by any other suitable means (e.g., by means of firmware).
[0133] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0134] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable image generation devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0135] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0136] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0137] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0138] A computer system can include a client and a server. The client and the server are generally far apart from each other and typically interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.
[0139] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. An image generation method, comprising: Obtaining constraint information, where the constraint information includes constraint text and a constraint image; Performing iterative denoising processing on a noise image based on the constraint information to generate a target image of a target object; the object content of the target object matches the content described in the constraint text, and the contour of the target object matches the object contour of the constraint image; Wherein, performing one denoising process includes: Enhancing information of the constraint information and the intermediate image of the previous denoising process to obtain compensation information, where the intermediate image is obtained by performing at least one denoising process on the noise image; Performing denoising processing on the intermediate image based on the compensation information and the constraint text.
2. The method according to claim 1, wherein, The performing denoising processing on the intermediate image based on the compensation information and the constraint text includes: Performing multiple feature restorations on multimodal features based on the compensation information; the multimodal features are obtained by processing the constraint text and the intermediate image.
3. The method according to claim 2, wherein, One denoising process includes I feature restoration processes; Performing one feature restoration process includes: Using a target compensation feature to perform feature restoration on the i-th multimodal feature to obtain the (i + 1)-th multimodal feature; The target compensation feature is a compensation feature in the multiple compensation features included in the compensation information that has a mapping relationship with the i-th multimodal feature, and i is an integer greater than or equal to 1 and less than I.
4. The method according to claim 3, wherein, The using a target compensation feature to perform feature restoration on the i-th multimodal feature to obtain the (i + 1)-th multimodal feature includes: Using the target compensation feature to perform information compensation on the i-th multimodal feature to obtain the i-th multimodal enhanced feature; Performing multimodal feature processing on the i-th multimodal enhanced feature to obtain the (i + 1)-th multimodal feature.
5. The method according to any one of claims 1 to 4, wherein The enhancing information of the constraint information and the intermediate image of the previous denoising process to obtain compensation information includes: Performing information fusion on the constraint information and the intermediate image of the previous denoising process to obtain fusion information; Performing feature extraction of different feature scales on the fusion information to obtain the compensation information composed of multiple compensation features.
6. The method according to any one of claims 1 to 4, further comprising: Extracting contour information from an original image to obtain the constraint image.
7. An image generation method, comprising: Obtaining constraint information, where the constraint information includes constraint text and a constraint image; Using a denoising module of an image generation model to perform iterative denoising processing on a noise image based on the constraint information to generate a target image of a target object; the object content of the target object matches the content described in the constraint text, and the contour of the target object matches the object contour of the constraint image; Wherein, performing one denoising process includes: Using a compensation module of the image generation model to enhance information of the constraint information and the intermediate image of the previous denoising process to obtain compensation information, where the intermediate image is obtained by performing at least one denoising process on the noise image; Using a denoising unit in the denoising module to perform denoising processing on the intermediate image based on the compensation information and the constraint text.
8. The method according to claim 7 further includes: Obtaining a sample target image and a sample constraint text, where the content described in the sample constraint text matches the object content of the sample object in the sample target image; Extracting the contour information of the sample target image to obtain a sample constraint image; Training an initial image generation model by using a sample noise image, sample constraint information formed by the sample constraint text and the sample constraint image, and the sample target image to obtain the image generation model; the sample noise image is obtained by adding noise to the sample target image, or the sample noise image is a random noise image.
9. The method according to claim 8, wherein The step of training the initial image generation model by using the sample noise image, the sample constraint information formed by the sample constraint text and the sample constraint image, and the sample target image to obtain the image generation model includes: Using the denoising module of the initial image generation model to perform iterative denoising processing on the sample noise image based on the sample constraint information to generate a sample restoration image of the sample object; Wherein, performing one denoising process includes: Using the initial compensation module of the initial image generation model to enhance the information of the sample constraint information and the sample intermediate image of the previous denoising process to obtain sample compensation information, and the sample intermediate image is obtained by performing at least one denoising process on the sample noise image; Using the denoising unit in the denoising module to perform denoising processing on the sample intermediate image based on the sample compensation information and the sample constraint text; Freezing the parameters of the denoising module, and adjusting the parameters of the initial compensation module based on the sample restoration image and the sample target image to obtain the image generation model.
10. An electronic device includes: One or more processors; A memory for storing one or more programs, Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 9.