Image denoising method and related equipment
By optimizing the resolution layer design and using a 32-channel variational encoder and attention weight normalization in the image denoising model, the problem of excessive parameter and computational complexity of the diffusion network deployed on mobile devices is solved, and efficient image denoising processing is achieved on mobile phones and other devices.
Patent Information
- Application Number
- CN202411216134.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2025-09-26
AI Technical Summary
The existing diffusion network-based portrait enhancement model has too many parameters and computational complexity, making it difficult to deploy on mobile devices such as mobile phones.
A resolution layer design is adopted to reduce the number of residual blocks in the high-resolution layer and increase the number of residual blocks in the low-resolution layer. A 32-channel variational encoder and attention weight normalization are used to reduce the number of parameters and computation.
While ensuring processing quality, the number of parameters and computational complexity are significantly reduced, allowing the portrait enhancement model to be deployed on mobile devices such as mobile phones.
Smart Images

Figure CN120707413A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an image denoising method and related equipment. Background Art
[0002] Portrait enhancement models primarily enhance image clarity and portraits by denoising them. Current diffusion network-based portrait enhancement models for denoising are too parameter-heavy and computationally intensive to deploy on mobile devices. Summary of the Invention
[0003] The present application provides an image denoising method and related equipment, which can reduce the number of parameters and the amount of calculation.
[0004] In a first aspect, some embodiments of the present application provide an image denoising method. The image denoising method may include:
[0005] The first image is input into the first model for processing to obtain a second image. The first model is used to denoise the image. The noise in the second image is less than that in the first image. The first model includes multiple resolution layers. The resolution layers are used to predict the noise in the image. The number of residual blocks in the first resolution layer is less than the number of residual blocks in the second resolution layer. The resolution of the first resolution layer is greater than the resolution of the second resolution layer. The residual blocks are used to maintain the image gradient of the features propagated between resolution layers.
[0006] Through the above method, the same processing requires more computation at higher resolutions and less computation at lower resolutions. Therefore, the number of residual blocks at lower resolutions is greater than that at higher resolutions, ensuring processing quality while reducing the number of parameters and computation.
[0007] In a possible implementation, the number of residual blocks in the first resolution layer is less than the number of residual blocks in the second resolution layer, including: the number of residual blocks in the first resolution layer is 1, and the number of residual blocks in the second resolution layer is multiple.
[0008] In the above manner, the residual blocks in the resolution layer with high resolution are removed, and the residual blocks in the resolution layer with low resolution are added, thereby reducing the amount of calculation corresponding to the residual blocks in the resolution layer with high resolution.
[0009] In a possible implementation, the second resolution layer is a resolution layer with the smallest resolution among multiple resolution layers.
[0010] Through the above method, due to the same processing, the smaller the resolution layer, the smaller the corresponding calculation amount will be, so adding the residual blocks to the resolution layer with the smallest resolution can minimize the calculation amount while ensuring the quality as much as possible.
[0011] In a possible implementation, the first resolution layer is a resolution layer with the highest resolution among multiple resolution layers.
[0012] In the above manner, for the same processing, a resolution layer with a larger resolution will require a larger amount of calculation, so the residual blocks in the resolution layer with the largest resolution are removed to reduce the amount of calculation.
[0013] In one possible implementation, the number of first resolution layers is the number of residual blocks in the second resolution layer plus one.
[0014] In the above manner, the quality of processing can be ensured.
[0015] In one possible implementation, the first image is input into the first model for processing, including: inputting the first image into the first model; encoding the first image through the 32-channel variational encoder in the first model to obtain a first feature; performing convolution processing on the first feature in the third resolution layer to obtain an attention weight, and the third resolution layer is the resolution layer with the largest resolution among multiple resolution layers; determining a second feature output by the third resolution layer based on the attention weight, and the second feature is used as the input of the next resolution layer.
[0016] In this way, by using a variational encoder with up to 32 channels, the first model can have a higher capacity to learn and characterize the complex features of the data, capturing more data details so that the subsequent decoding stage can reconstruct the original data more accurately.
[0017] In one possible implementation, determining the second feature of the output of the third resolution layer based on the attention weight includes: normalizing the attention weight to obtain the standardized attention weight, wherein the normalization is used to limit the value of the attention weight to a preset range; and determining the second feature of the output of the third resolution layer based on the standardized attention weight.
[0018] In the above manner, the weights are limited to a range through standardization processing, thereby avoiding the problem of excessive calculation results.
[0019] In one possible implementation, an encoder is trained based on the first labeled image to obtain a 32-channel variational encoder; a model is trained based on the second labeled image to obtain a diffusion network; and a model is trained based on the sample image to obtain a conditionally controlled generative network; wherein the noise of both the first labeled image and the second labeled image is greater than that of the sample image.
[0020] Through the above method, model training can be accurately performed to obtain the first model.
[0021] In a second aspect, the present application provides an image denoising device, which may be an electronic device, a device within an electronic device, or a device capable of being used in conjunction with an electronic device; wherein the image denoising device may also be a chip system, and the image denoising device may execute the method executed by the electronic device in the first aspect. The functions of the image denoising device may be implemented by hardware, or by hardware executing corresponding software implementations. The hardware or software includes one or more units corresponding to the above functions. The units may be software and / or hardware. The operations and beneficial effects performed by the image denoising device may refer to the methods and beneficial effects described in the first aspect above, and any repetitions will not be repeated.
[0022] In a third aspect, the present application provides an electronic device comprising one or more processors and one or more memories. The one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code, wherein the computer program code comprises computer instructions. When the one or more processors execute the computer instructions, the electronic device performs the image denoising method according to any possible implementation of the first aspect.
[0023] In a fourth aspect, the present application provides a chip system comprising a processor and an interface, wherein the processor and the interface are coupled; the interface is used to receive or output signals, and the processor is used to execute code instructions to execute the image denoising method in any possible implementation of the first aspect above.
[0024] In a fifth aspect, the present application provides a computer-readable storage medium, which stores a computer program / instructions. When the computer program product runs on a computer, it enables the computer to execute the image denoising method in any possible implementation of the first aspect above.
[0025] In a sixth aspect, the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the image denoising method in any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1A A schematic diagram of a Stable Diffusion UNet architecture provided in an embodiment of the present application;
[0027] Figure 1B A schematic diagram of the structure of a residual block provided in an embodiment of the present application;
[0028] Figure 1C A schematic diagram of a space conversion structure provided in an embodiment of the present application;
[0029] Figure 1DA schematic diagram of the architecture of a conditional control generation network provided in an embodiment of the present application;
[0030] Figure 2 A flowchart of an image denoising method provided in an embodiment of the present application;
[0031] Figure 3A A schematic diagram of the structure of an attention block provided in an embodiment of the present application;
[0032] Figure 3B A schematic diagram of a UNet network architecture provided in an embodiment of the present application;
[0033] Figure 3C A schematic diagram of a training process provided in an embodiment of the present application;
[0034] Figure 4 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;
[0035] Figure 5 A schematic diagram of the software structure of an electronic device provided in an embodiment of the present application;
[0036] Figure 6 A schematic diagram of an image denoising application process provided in an embodiment of the present application;
[0037] Figure 7 A schematic structural diagram of an image denoising device provided in an embodiment of the present application;
[0038] Figure 8 A schematic diagram of the structure of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION
[0039] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in the text is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.
[0040] It should be understood that the terms "first," "second," and the like in the specification, claims, and drawings of this application are used to distinguish between different objects, rather than to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.
[0041] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.
[0042] 1. Diffusion Model
[0043] The diffusion model is mainly divided into two processes: the noise addition process and the denoising process. The noise addition process means gradually adding Gaussian noise to the real image in the dataset, while the denoising process means gradually removing the noise from the noisy image to restore the real image.
[0044] 2. Stable Diffusion UNet Architecture
[0045] To further improve the generation quality and training stability of the diffusion model, Stable Diffusion was introduced. Stable Diffusion is an improved diffusion model that significantly improves the quality and consistency of image generation by introducing the U-Net architecture and other techniques.
[0046] The following is a further introduction to the UNet structure of Stable Diffusion.
[0047] Stable Diffusion consists of the following three core parts:
[0048] A text encoder (using CLIP's ViT-L / 14 text encoder in Stable Diffusion) is used to convert user-entered text prompts into text embeddings.
[0049] Image Auto Encoder-Decoder, used to encode an image into a latent vector z, or restore an image from the latent vector z;
[0050] UNET structure, uses UNET for iterative denoising, performs multiple rounds of prediction under text guidance, and converts random Gaussian noise zt into the image implicit vector z0.
[0051] These three components are independent of each other, with the most important being the UNET structure. The UNET is the primary component for generating images from noise. During the prediction process, the UNET is repeatedly called to remove the noise slices from the original noise, resulting in a progressively denoised image representation. The UNET of the Stable Diffusion Model contains approximately 860M parameters.
[0052] A UNet can include multiple modules, specifically including but not limited to the following four modules: ResnetBlock, Spatial Transformer Block, DownSample and UpSample. It should be noted that a UNet can include multiple residual blocks and multiple spatial transformer blocks. In practical applications, such as Figure 1A (For ease of description Figure 1A Only the architecture of the downsampling part is shown) The UNet network architecture diagram of StableDiffusion is shown, taking the resolution h*w as an example (the resolution is Figure 1A The UNet with the highest resolution in the StableDiffusion UNet structure shown in the figure, or the resolution layer with the highest resolution, includes two residual blocks and two spatial transformation blocks. For each resolution layer, the output of the previous layer is used as the input of the next residual block A, the output of residual block A is used as the input of spatial transformation block A, the output of spatial transformation block A is used as the input of residual block B, and the output of residual block B is used as the input of spatial change block B. Finally, spatial transformation block B is downsampled and input to the next resolution layer.
[0053] in, Figure 1A The resolution of h*w is higher than that of h / 2*w / 2. For example, h*w can be 64*64, and h / 2*w / 2 can be 32*32. The values of h and w can be the same or different.
[0054] The following further Figure 1A The residual block, spatial transformation block and downsampling in are introduced separately:
[0055] Residual block: The structure of the residual block is as follows Figure 1BAs shown in the figure, the residual block accepts two inputs: image features (also called latent vectors) and a time embedding. After convolution, the latent vector is summed with the fully connected projected time embedding. This sum is then added to the original latent vector after skip connections. This is then fed into another convolutional layer to produce the latent output after encoding and transformation by the residual block.
[0056] Spatial Transformation Block: The spatial transformation block also has two inputs, namely the output (latent) of the previous module (i.e. residual block) and the context embedding (the output of the text prompt after CLIP encoding). In this module, the image feature (latent vector) corresponds to the image token, and performs a crossattention operation with the context embedding (such as Figure 1C As shown in Figure 3, this attention mechanism injects semantic information from the context embedding into the corresponding image token, effectively fusing image and text information. In general, the size of the latent output by the spatial transformation block remains consistent with the size of the input latent, but semantic information is incorporated at the corresponding locations.
[0057] The main function of downsampling is to reduce the length and width of the image features (latent vectors) by 2 times, which is achieved through a two-dimensional ordinary convolution with a kernel size of 3 and a stride of 2. Figure 1A Upsampling (the omitted part of the figure) increases the length and width of the latent vector by a factor of 2, achieved through interpolation. In general, downsampling and upsampling simply change the size of the latent vector, maintaining the number of channels unchanged.
[0058] 3. Conditional Control Generation Network - ControlNet
[0059] Although the above diffusion model works well in generating images, it still has certain limitations in some complex scenes and fine control. In order to further enhance the ability of conditional control generation, the ControlNet method was proposed to improve the quality and consistency of the generated image by introducing additional control signals. Compared with Stable Diffusion, ControlNet provides more control and flexibility, enabling it to achieve some functions that are difficult to achieve with Stable Diffusion. The structure of ControlNet is as follows Figure 1DAs shown in the figure, the left half is the Stable Diffusion branch, whose parameters are frozen and not trained; the right half is the ControlNet branch, whose parameters are trainable.
[0060] Each SD encoding module can also be called a resolution layer, or a UNet. A resolution layer corresponds to a resolution. For example, the resolution corresponding to SD encoding module_1 is 64x64, and the resolution corresponding to SD encoding module_2 is 32x32. A resolution layer can include multiple modules, such as Figure 1A The medium resolution h*w layer includes two residual blocks and two spatial transformation blocks.
[0061] The above-mentioned diffusion network / Stable Diffusion's UNet / Conditional Control Generation Network - ControlNet can be used for portrait enhancement, that is, as a portrait enhancement model. The focus of the portrait enhancement model is to denoise the image. Although the diffusion model includes denoising processing, the portrait enhancement model based on the extended model has a very large number of parameters, and the amount of calculation during use is also very large, and it cannot be deployed to mobile devices such as mobile phones. The present application provides an image denoising method. The image denoising method provided by the present application can reduce the number of parameters and the amount of calculation, thereby realizing the deployment of the portrait enhancement model based on the diffusion model to mobile devices such as mobile phones.
[0062] The following is a further introduction to the image denoising method provided in the embodiment of the present application: Figure 2 As shown, Figure 2 This is a flow chart of an image denoising method provided in an embodiment of the present application. The image denoising method includes the following step 201. Figure 2 The method execution subject shown can be an electronic device or a first model in an electronic device. Or Figure 2 The execution subject of the method shown can be a chip or chip system in an electronic device, which is not limited in the embodiments of the present application. Figure 2 The following description is made by taking an electronic device as the execution subject of the method as an example.
[0063] 201. An electronic device inputs a first image into a first model for processing to obtain a second image, wherein the first model is used to denoise the image, the noise in the second image is smaller than that in the first image, the first model includes multiple resolution layers, the resolution layers are used to predict the noise in the image, the number of residual blocks in the first resolution layer is less than the number of residual blocks in the second resolution layer, the resolution of the first resolution layer is greater than the resolution of the second resolution layer, and the residual blocks are used to maintain the image gradient of features propagated between resolution layers.
[0064] Optionally, the first model can be used for portrait enhancement, or the first model can be said to be a portrait enhancement model. The first model can also be used in other scenarios, for example, the first model can also be used for image generation, etc., which is not limited in this application.
[0065] Optionally, the first image is a noisy image, and the second image is a denoised image. The noise in the first image may be similar to Gaussian noise. The first image may be an image captured by a camera of an electronic device, an image stored in an album of the electronic device, or an image downloaded from the Internet. The first image may include a human face.
[0066] Optionally, the resolution layer may include residual blocks and / or spatial transform blocks as well as downsampling / upsampling. If the resolution includes downsampling, then the resolution layer no longer includes upsampling; similarly, if the resolution includes upsampling, then the resolution layer no longer includes downsampling.
[0067] Optionally, the resolution layer can be the UNet mentioned above.
[0068] Optionally, different resolution layers correspond to different resolutions, and the resolution can be referred to as h*w above.
[0069] Optionally, the resolution corresponding to the first resolution layer is different from the resolution corresponding to the second resolution layer. For example, the resolution of the first resolution layer is 64*64, and the resolution of the second resolution layer is 32*32. The first resolution layer and the second resolution layer may be adjacent resolution layers or non-adjacent resolution layers.
[0070] Optionally, the residual block can refer to the introduction above, and this application will not elaborate on it.
[0071] In a possible embodiment, the number of residual blocks in the first resolution layer is less than the number of residual blocks in the second resolution layer, including: the number of residual blocks in the first resolution layer is 1, and the number of residual blocks in the second resolution layer is multiple.
[0072] Optionally, there may be multiple first resolution layers and multiple second resolution layers. That is, the number of residual blocks in multiple resolution layers in the first model is 1, or there are multiple resolution layers without residual blocks. Similarly, multiple resolution layers in the first model contain multiple residual blocks.
[0073] In such Figure 1AIn the Stable Diffusion UNet architecture shown, each resolution layer includes two residual blocks, while in the first model, the number of residual blocks in the first resolution layer is one, and the number of residual blocks in the second resolution layer is multiple. For example, the number of residual blocks in the resolution layer with a resolution of h*w is one, and the number of residual blocks in the resolution layer with a resolution of h / 4*w / 4 is four.
[0074] In a possible embodiment, the second resolution layer is a resolution layer with the smallest resolution among multiple resolution layers.
[0075] Exemplarily, the first model includes the following three resolution layers: resolution layer 1 (resolution h*w), resolution layer 2 (resolution h / 2*w / 2), and resolution layer 3 (resolution h / 4*w / 4), where resolution layer 1 is the resolution layer with the highest resolution and resolution layer 3 is the resolution layer with the lowest resolution. The first resolution layer is resolution layer 1 and resolution layer 2, and each of resolution layer 1 and resolution layer 2 includes only one residual block. The second resolution layer is resolution layer 3, and resolution layer 3 includes three residual blocks.
[0076] In a possible embodiment, the first resolution layer is a resolution layer with the highest resolution among multiple resolution layers.
[0077] For example, the first model includes the following three resolution layers: resolution layer 1 (resolution of h*w), resolution layer 2 (resolution of h / 2*w / 2), and resolution layer 3 (resolution of h / 4*w / 4), where resolution layer 1 is the resolution layer with the highest resolution and resolution layer 3 is the resolution layer with the lowest resolution. The first resolution layer is resolution layer 1, and the second resolution layer is resolution layer 3.
[0078] In a possible embodiment, the number of first resolution layers is the number of residual blocks in the second resolution layer plus one.
[0079] Exemplarily, the first model includes the following three resolution layers: resolution layer 1 (resolution h*w), resolution layer 2 (resolution h / 2*w / 2), and resolution layer 3 (resolution h / 4*w / 4), where resolution layer 1 is the resolution layer with the highest resolution and resolution layer 3 is the resolution layer with the lowest resolution. The first resolution layer is resolution layer 1 and resolution layer 2, and the second resolution layer is resolution layer 3. The number of residual blocks in resolution layer 3 is 1+2=3.
[0080] Optionally, the structure of the residual block can be found in Figure 1B shown.
[0081] Optionally, all residual blocks in the first model may be improved residual blocks, specifically, Figure 1B The PaddedConv2D convolution layer in the residual block structure shown is changed to a Depthwise Separable Convlution convolution layer. The computational complexity of the Depthwise Separable Convlution convolution layer is smaller than that of the PaddedConv2D convolution layer.
[0082] Optional, such as Figure 1B As shown, a residual block includes two PaddedConv2D convolutional layers. Both PaddedConv2D convolutional layers can be changed to Depthwise Separable Convlution convolutional layers, or only one of the PaddedConv2D convolutional layers can be changed to Depthwise Separable Convlution convolutional layers. This application does not impose any restrictions on this.
[0083] In a possible embodiment, the first image is input into the first model for processing, including: inputting the first image into the first model; encoding the first image through the 32-channel variational encoder in the first model to obtain a first feature; performing convolution processing on the first feature in a third resolution layer to obtain an attention weight, and the third resolution layer is the resolution layer with the largest resolution among multiple resolution layers; determining a second feature output by the third resolution layer based on the attention weight, and the second feature is used as the input of the next resolution layer.
[0084] Optionally, the third resolution layer is the resolution layer with the highest resolution, that is, the third resolution layer is processed first among all resolution layers. A second feature of the output of the third resolution layer is determined based on the attention weight, and the second feature is used as the input of the next resolution layer.
[0085] Optionally, the attention block in the third resolution layer performs convolution processing on the first feature to obtain an attention weight; the attention block in the third resolution layer determines the second feature output by the third resolution layer based on the attention weight.
[0086] Optionally, the next resolution layer also includes an attention block, and the next resolution layer also performs the same processing as the third resolution layer.
[0087] Optionally, inputting the first image into the first model for processing includes: inputting the first image into the first model; encoding the first image through a 32-channel variational encoder in the first model to obtain a first feature; performing zero convolution on the first feature, and combining the result of the zero convolution with the first convolution layer (the first convolution layer is as shown in FIG. Figure 1D The output of the convolutional layer in the left network is added (the parameters of the left network are frozen); the result of the addition (Latent vector) is input to the third resolution layer (such as Figure 1A SD encoding module_1 in the right network).
[0088] Among them, the existing model usually adopts a 4-channel variational encoder. Using a 32-channel variational encoder can allow the first model to retain more image features.
[0089] Optionally, the third resolution layer includes a residual block and an attention block, and the attention block is different from Figure 1C The spatial transformation block shown. That is, with Figure 1A Unlike the example where each resolution layer includes a residual block and a spatial transformation block, the resolution layer in the first model of this application includes a residual block and an attention block. This means that the spatial transformation block is replaced by an attention block. Alternatively, the spatial transformation block is improved to obtain a new attention block.
[0090] because Figure 1C The purpose of the demonstrated spatial attention block is to integrate semantic information at the corresponding position to achieve image generation based on text; while the first model in this application is used for denoising to achieve portrait enhancement. Therefore, removing the calculation of the text / semantic part can further reduce model parameters and the amount of computation required when applying the model.
[0091] In a possible embodiment, determining the second feature of the output of the third resolution layer based on the attention weight includes: normalizing the attention weight to obtain the standardized attention weight, wherein the normalization is used to limit the value of the attention weight to a preset range; and determining the second feature of the output of the third resolution layer based on the standardized attention weight.
[0092] Optionally, the attention block in the third resolution layer normalizes the attention weight to obtain the standardized attention weight; the attention block in the third resolution layer determines the second feature of the third resolution layer output based on the standardized attention weight.
[0093] Optionally, the attention block includes normalization processing, which can be performed through the softmax function. That is, the data is input into the softmax function for normalization. However, the softmax function is an exponential function. If the value input to the softmax function is too large, it will cause the exponential function result to overflow. Therefore, normalization is performed on the data so that the model can train effective output without loading pre-trained weights during the training phase, and the model is more stable during the application phase.
[0094] For example, Figure 3A As shown, Figure 3A This is a schematic diagram of the structure of an attention block provided in an embodiment of the present application. The attention block uses image features (latent vectors) as input. After convolution processing on the image features, a query matrix, a key matrix, and a value matrix are obtained. The query matrix and key matrix are normalized separately, and matrix multiplication is performed on the two normalized matrices. Finally, the attention width matrix is obtained through scaling and normalization. The attention width matrix is then matrix multiplied with the value matrix, and the output is obtained through a convolution layer, which is also a latent vector.
[0095] The following is a further introduction to the architecture of the first model provided in the embodiment of the present application. Figure 3B As shown, Figure 3B A schematic diagram of the UNet network architecture in a first model provided in an embodiment of the present application.
[0096] Optional, compared to Figure 1A The UNet network structure diagram shown, Figure 3B In the UNet network architecture shown in the figure, each resolution layer except the lowest resolution layer has only one residual block (that is, in Figure 1A The UNet network shown in the figure has one residual block removed except for the lowest resolution layer. The spatial transformation blocks in each layer are replaced by attention blocks, and each resolution layer except the lowest resolution layer has only one attention block. Figure 3B More residual blocks and more attention blocks are added to the lowest resolution layer in the UNet network architecture diagram shown.
[0097] The above describes the application phase of the first model. The following describes the training phase of the first model:
[0098] In one possible embodiment, the electronic device performs encoder training based on the first label image to obtain a 32-channel variational encoder; performs model training based on the second label image to obtain a diffusion network; and performs model training based on the sample image to obtain a conditionally controlled generative network; wherein the noise of the first label image and the second label image is greater than that of the sample image.
[0099] Optionally, the first model includes two networks: a diffusion network and a conditional control generation network. The architecture of the first network can be found in Figure 1D As shown, Figure 1D The middle left part (the part with high-definition image input) is a diffusion network, which can also be called a stable diffusion network. Figure 1DThe right part (the part with low-definition image input) is the conditional control generation network. This conditional control generation network is used to provide more control and flexibility, enabling the entire first model to achieve some functions that are difficult to achieve with the StableDiffusion network.
[0100] Optionally, the electronic device first performs encoder training based on the first label image to obtain a 32-channel variational encoder; after the 32-channel variational encoder training is completed, the model is trained based on the second label image to obtain a diffusion network; after the diffusion network training is completed, the model is trained based on the sample image to obtain a conditional control generation network.
[0101] Among them, the first label image and the second label image can be the same image or different images. The first label image and the second label image can be high-definition images, that is, the noise in the images is small; the sample image can be a low-definition image, that is, the noise in the image is large.
[0102] For example, the training process can be found in Figure 3C As shown, the autoencoder is first trained. After the autoencoder training is completed, the parameters in the autoencoder are locked, and the light UNet network (diffusion network) is trained; after the UNet network training is completed, the parameters of the autoencoder and the light UNet network are locked, and the conditional control generation network is trained.
[0103] The following is an introduction to the hardware structure of electronic equipment:
[0104] See also Figure 4 , Figure 4 1 is a schematic diagram of the hardware structure of the electronic device 100 provided in an embodiment of the present application. The electronic device 100 may be an electronic device corresponding to the training phase or an electronic device corresponding to the application phase.
[0105] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0106] It should be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0107] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0108] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0109] The processor 110 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by the processor 110. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves system efficiency. The processor 110 calls the instructions or data stored in the memory, causing the electronic device 100 to execute the shooting method performed by the electronic device in the following method embodiment.
[0110] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.
[0111] The charging management module 140 is configured to receive charging input from a charger, which may be a wireless charger or a wired charger.
[0112] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to provide power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, and the wireless communication module 160. In some other embodiments, the power management module 141 can also be set in the processor 110.
[0113] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0114] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0115] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.
[0116] The modem processor includes a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a medium- or high-frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is passed to the application processor.
[0117] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as Wi-Fi networks), Bluetooth (BT), BLE broadcasting, global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. applied to the electronic device 100. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.
[0118] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150 , and antenna 2 is coupled to wireless communication module 160 , so that electronic device 100 can communicate with the network and other devices through wireless communication technology.
[0119] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0120] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. In some embodiments, electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1. Display screen 194 may include an OLED screen.
[0121] Optionally, the display screen 194 may further include: an OLED glass layer, an OLED light-emitting unit, a fingerprint recognition sensor, a microlens array, etc. The display screen 194 supports optical under-screen fingerprint recognition.
[0122] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor. The ISP is used to process data fed back by the camera 193. The camera 193 is used to capture still images or videos. The camera 193 may include a front camera and a rear camera, the front camera is located in the display area of the screen, and the rear camera is located in the back area of the screen. The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. The video codec is used to compress or decompress digital videos. The electronic device 100 may support one or more video codecs.
[0123] NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission mode between neurons in the human brain, it can quickly process input information and can also continuously self-learn.
[0124] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement a data storage function.
[0125] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound playback function), etc. The data storage area can store data (such as audio data) created during the use of the electronic device 100, etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as a flash memory device, etc.
[0126] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0127] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.
[0128] Speaker 170A, also known as a "horn," is used to convert audio electrical signals into sound signals. Receiver 170B, also known as an "earpiece," is used to convert audio electrical signals into sound signals. Microphone 170C, also known as a "microphone" or "microphone," is used to convert sound signals into electrical signals. Headphone jack 170D is used to connect wired headphones. Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A may be provided on display screen 194. Gyroscope sensor 180B may be used to determine the motion posture of electronic device 100. Air pressure sensor 180C is used to measure air pressure. Magnetic sensor 180D includes a Hall sensor. Acceleration sensor 180E may detect the magnitude of acceleration of electronic device 100 in various directions (generally three axes). Distance sensor 180F is used to measure distance. Proximity light sensor 180G may include, for example, a light-emitting diode (LED) and a light detector. Ambient light sensor 180L is used to sense ambient light brightness. Fingerprint sensor 180H is used to collect fingerprints. The temperature sensor 180J is used to detect the temperature. The touch sensor 180K is also called a "touch panel". The touch sensor 180K can be set on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also called a "touch screen". The touch sensor 180K is used to detect touch operations acting on or near it. The bone conduction sensor 180M can obtain vibration signals. The buttons 190 include a power button, a volume button, etc. The motor 191 can generate vibration prompts. The indicator 192 can be an indicator light, which can be used to indicate the charging status, power changes, messages, missed calls, notifications, etc. The SIM card interface 195 is used to connect a SIM card.
[0129] In addition, an operating system runs on top of the above components. For example, operating systems such as iOS and Android. The operating system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a micro-core architecture, a microservice architecture, or a cloud architecture. The embodiment of the present application takes the Android system with a layered architecture as an example to illustrate the software structure of the electronic device 100. It should be noted that although the embodiment of the present application is described using the Android system as an example, its basic principles are also applicable to electronic devices with other operating systems.
[0130] The software structure of the electronic device 100 is introduced below:
[0131] Figure 5Schematic diagram of the software structure of an electronic device 100 provided in an embodiment of the present application. The software structure adopts a layered architecture, which divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In an embodiment of the present application, the operating system (taking the Android system, where the Android system runs on an AP as an example) can be divided into four layers, from top to bottom: the application layer (application, APP), the application framework layer (framework, FWK), the hardware abstraction layer (hardware abstraction layer, HAL) and the kernel layer (Kernal).
[0132] The application layer can include a series of application packages. Figure 5 As shown, the application package may include applications such as a camera and a gallery. In the embodiment of the present application, the camera refers to a camera application. The camera application may include a camera interface module (which may be called a CameraApi2Module), etc. The gallery refers to a gallery application, which is used to store images and videos taken by electronic devices. The gallery application is also used to provide users with a playback function, and users can view historical images and videos in the gallery application.
[0133] The application framework layer provides application developers with an application programming interface (API) framework and various services and management tools for accessing core functions, including interface management, data access, application layer messaging, application package management, phone management, location management, and other functions. The application framework layer includes some predefined functions. Figure 5 As shown, the application framework layer may include but is not limited to the camera service CameraService.
[0134] The CameraService is responsible for scheduling the camera application's startup process, creating and managing processes, and creating and managing windows. For this application, the portrait enhancement model can be built into the camera service or be a module independent of the camera service. For ease of description, we'll take the example of a portrait enhancement model built into the camera service.
[0135] The hardware abstraction layer is an interface layer between the operating system kernel and the hardware circuit. Its purpose is to abstract the hardware. It hides the hardware interface details of a specific platform and can provide a virtual hardware platform for the operating system. Figure 5As shown, the hardware abstraction layer may include camera resource service (CameraResourceService), camera provider service (CameraProvider), etc. In addition, the hardware abstraction layer may also include: camera device session interface (CameraDeviceSession), camera preview screen interface (PreviewFlowImpl), etc. Among them, the camera resource service is used to interact with the memory bar in the hardware; CameraProvider is used to enumerate individual devices and manage their status, and can open and close physical camera devices (such as rear cameras); CameraDeviceSession is used to create camera device sessions and store the properties and configuration information required for camera device sessions; PreviewFlowImpl is responsible for notifying the APP that the first frame of the preview screen has been displayed after the first frame of the preview screen is displayed.
[0136] The kernel layer is the layer between hardware and software. The kernel layer includes at least display drivers, camera drivers, etc.
[0137] Based on the above Figure 4 and Figure 5 The following is a further introduction to the application phase of the first model. The application phase of the first model mainly involves the following four parts: sensor (hardware), ISP (hardware), camera service (application framework layer) and gallery. Among them, the first model is used in the camera service as an example. Figure 6 As shown:
[0138] First, after the sensor captures the original image, it sends the acquired original image to the ISP (for electronic devices such as mobile phones, the original image is usually processed by the ISP).
[0139] The ISP then performs ISP processing on the original image to produce a first image. This ISP processing includes, but is not limited to, the following: gamma correction, which adjusts the image's contrast and brightness to better reflect the human eye's visual characteristics; color correction, which adjusts and optimizes the image's color performance; and LSC, which addresses issues such as uneven image brightness and color caused by lens optical characteristics.
[0140] Then, the first model embedded in the camera service processes the first image to obtain a second image. Finally, the camera service sends the second image to the image library, and the image library saves the second image.
[0141] See Figure 7 , Figure 7 Schematic diagram of the structure of an image denoising device 700 provided in an embodiment of the present application. Figure 7The image denoising device shown may be an electronic device, or a device in an electronic device, or a device that can be used in conjunction with an electronic device. Figure 7 The image denoising apparatus shown may include a processing unit 701. In which:
[0142] Processing unit 701 is used to input the first image into the first model for processing to obtain a second image, the first model is used to denoise the image, the noise in the second image is less than that in the first image, the first model includes multiple resolution layers, the resolution layers are used to predict the noise in the image, the number of residual blocks in the first resolution layer is less than the number of residual blocks in the second resolution layer, the resolution of the first resolution layer is greater than the resolution of the second resolution layer, and the residual blocks are used to maintain the image gradient of the features propagated between resolution layers.
[0143] In a possible implementation, the number of residual blocks in the first resolution layer is less than the number of residual blocks in the second resolution layer, including: the number of residual blocks in the first resolution layer is 1, and the number of residual blocks in the second resolution layer is multiple.
[0144] In a possible implementation, the second resolution layer is a resolution layer with the smallest resolution among multiple resolution layers.
[0145] In a possible implementation, the first resolution layer is a resolution layer with the highest resolution among multiple resolution layers.
[0146] In one possible implementation, the number of first resolution layers is the number of residual blocks in the second resolution layer plus one.
[0147] In one possible implementation, the processing unit 701 is further used to input the first image into the first model; encode the first image through the 32-channel variational encoder in the first model to obtain a first feature; perform convolution processing on the first feature in the third resolution layer to obtain an attention weight, and the third resolution layer is the resolution layer with the largest resolution among multiple resolution layers; determine the second feature output by the third resolution layer based on the attention weight, and the second feature is used as the input of the next resolution layer.
[0148] In one possible implementation, the processing unit 701 is also used to normalize the attention weight to obtain the standardized attention weight, and the normalization is used to limit the value of the attention weight to a preset range; and determine the second feature of the output of the third resolution layer based on the standardized attention weight.
[0149] In one possible implementation, the processing unit 701 is further used to perform encoder training based on the first label image to obtain a 32-channel variational encoder; perform model training based on the second label image to obtain a diffusion network; and perform model training based on the sample image to obtain a conditionally controlled generative network; wherein the noise of the first label image and the second label image is greater than that of the sample image.
[0150] For the case where the image denoising device can be a chip or a chip system, see Figure 8 Schematic diagram of the chip structure shown. Figure 8 The chip 800 shown includes a processor 801 and an interface 802. Optionally, it may also include a memory 803. The number of the processors 801 may be one or more, and the number of the interfaces 802 may be multiple.
[0151] For the case where the chip is used to implement the electronic device in the embodiment of the present application:
[0152] The interface 802 is used to receive or output signals;
[0153] The processor 801 is configured to execute data processing operations of the electronic device.
[0154] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
[0155] It is understood that some optional features in the embodiments of the present application may, in certain scenarios, be implemented independently of other features, such as the solution on which they are currently based, to solve corresponding technical problems and achieve corresponding effects. They may also be combined with other features in certain scenarios as needed. Accordingly, the audio data archiving device provided in the embodiments of the present application may also implement these features or functions accordingly, which will not be described in detail here.
[0156] It should be understood that the processor in the embodiment of the present application can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method embodiment can be completed by an integrated logic circuit of hardware in the processor or instructions in the form of software. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component.
[0157] It is understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0158] The present application also provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed on an electronic device, the functions of any of the above method embodiments are implemented.
[0159] The present application also provides a computer program product, which, when executed on a computer, enables the computer to implement the functions of any of the above method embodiments.
[0160] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a high-density digital video disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).
[0161] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. An image denoising method, characterized in that: The method comprises: A first image is input into a first model for processing to obtain a second image, the first model is used to denoise the image, the noise in the second image is smaller than that in the first image, the first model includes multiple resolution layers, the resolution layers are used to predict the noise in the image, the number of residual blocks in the first resolution layer is less than the number of residual blocks in the second resolution layer, the resolution of the first resolution layer is greater than the resolution of the second resolution layer, and the residual blocks are used to maintain the image gradient of features propagated between resolution layers.
2. The method according to claim 1, characterized in that The number of residual blocks in the first resolution layer is less than the number of residual blocks in the second resolution layer, comprising: The number of residual blocks in the first resolution layer is 1, and the number of residual blocks in the second resolution layer is multiple.
3. The method according to claim 2, characterized in that The second resolution layer is a resolution layer with the smallest resolution among the multiple resolution layers.
4. The method according to claim 3, characterized in that The first resolution layer is a resolution layer with the highest resolution among the multiple resolution layers.
5. The method according to claim 3, characterized in that The number of the first resolution layers is the number of residual blocks in the second resolution layer plus one.
6. The method according to any one of claims 1 to 5, characterized in that Inputting the first image into the first model for processing includes: inputting the first image into the first model; Encoding the first image using a 32-channel variational encoder in the first model to obtain the first feature; In a third resolution layer, convolution processing is performed on the first feature to obtain an attention weight, and the third resolution layer is a resolution layer with the highest resolution among the multiple resolution layers; A second feature of the output of the third resolution layer is determined based on the attention weight, and the second feature is used as an input of the next resolution layer.
7. The method according to claim 6, characterized in that The determining, based on the attention weight, a second feature of the output of the third resolution layer, comprises: performing a normalization process on the attention weight to obtain a normalized attention weight, wherein the normalization process is used to limit a value of the attention weight to a preset range; Determine a second feature of the third resolution layer output based on the normalized attention weight.
8. The method according to claim 6 or 7, characterized in that The first model includes a 32-channel variational encoder, a diffusion network, and a conditional control generation network. The method further includes: Performing encoder training based on the first labeled image to obtain the 32-channel variational encoder; Performing model training based on the second label image to obtain the diffusion network; Performing model training based on sample images to obtain the conditionally controlled generative network; The noise of the first label image and the second label image is greater than that of the sample image.
9. An electronic device comprising one or more memories and one or more processors, characterized in that: The memory is used to store a computer program; the processor is used to call the computer program, so that the electronic device executes the method according to any one of claims 1 to 8.
10. A chip system, applied to electronic equipment, characterized in that: The chip system includes at least one processor and an interface, wherein the interface is used to receive instructions and transmit them to the at least one processor; the at least one processor executes the instructions so that the electronic device executes the method according to any one of claims 1 to 8.
11. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
12. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
An image super-resolution reconstruction method and device
CN109447900A
Mine image super-resolution reconstruction method and system based on multi-scale residual network
CN113592718A
Infrared image deblurring algorithm based on attention mechanism residual network model
CN115345791A
Digital watermark attack method based on conditional diffusion model
CN116645260A
Method for semantic image synthesis using condition diffusion and apparatus for same
KR1020240134643A