Image denoising method and related device

CN120707413BActive Publication Date: 2026-09-25HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411216134.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2026-09-25
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

当前用于去噪的基于扩散网络的人像增强模型参数量和计算量过大,很难部署在手机登移动设备上

Benefits of technology

[0017]通过上述方式,通过标准化处理,将权重限定在范围内,从而避免了计算结果过大灯问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707413B_ABST
    Figure CN120707413B_ABST
Patent Text Reader

Abstract

The application discloses an image denoising method and related equipment. The method comprises the following steps: inputting a first image into a first model for processing to obtain a second image, the first model is used for denoising the image, the noise in the second image is smaller than that in the first image, the first model comprises a plurality of resolution layers, the resolution layers are used for predicting the noise in the image, the number of residual blocks in a first resolution layer is less than that in a second resolution layer, the resolution of the first resolution layer is greater than that of the second resolution layer, and the residual blocks are used for keeping the image gradient of the features propagated between the resolution layers. The method can reduce the parameter quantity and the calculation quantity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to image denoising methods and related equipment. Background Technology

[0002] Image enhancement models primarily improve image clarity by denoising the image, thereby enhancing the human figure within it. However, current diffusion network-based image enhancement models for denoising have excessively large parameter and computational requirements, making them difficult to deploy on mobile devices. Summary of the Invention

[0003] This application provides an image denoising method and related equipment, which can reduce the number of parameters and computational load.

[0004] Firstly, some embodiments of this application provide an image denoising method. This image denoising method may include: The first image is input into the first model for processing to obtain the second image. The first model is used to denoise the image. The noise in the second image is less than that in the first image. The first model includes multiple resolution layers. The resolution layers are used to predict the noise in the image. The number of residual blocks in the first resolution layer is less than the number of residual blocks in the second resolution layer. The resolution of the first resolution layer is greater than the resolution of the second resolution layer. The residual blocks are used to maintain the image gradient of the features propagated between the resolution layers.

[0005] Using the above method, the same processing requires more computation at higher resolution layers and less computation at lower resolution layers. Therefore, the number of residual blocks at lower resolution is greater than the number of residual blocks at higher resolution, which can ensure processing quality while reducing the number of parameters and computational load.

[0006] In one possible implementation, the number of residual blocks in the first resolution layer is less than the number of residual blocks in the second resolution layer, including: the number of residual blocks in the first resolution layer is 1, and the number of residual blocks in the second resolution layer is multiple.

[0007] By using the above method, residual blocks in high-resolution layers are removed and residual blocks in low-resolution layers are added, thereby reducing the computational load corresponding to residual blocks in high-resolution layers.

[0008] In one possible implementation, the second resolution layer is the resolution layer with the lowest resolution among multiple resolution layers.

[0009] By using the above method, since the same processing is performed, the computational load will be smaller for the smaller resolution layer. Therefore, by adding all residual blocks to the smallest resolution layer, the computational load can be reduced to the greatest extent while ensuring quality as much as possible.

[0010] In one possible implementation, the first resolution layer is the resolution layer with the highest resolution among multiple resolution layers.

[0011] Using the above method, since the same processing is applied, the computational load will be greater for higher resolution layers. Therefore, the residual blocks in the highest resolution layer are removed to reduce the computational load.

[0012] In one possible implementation, the number of the first resolution layer is one less than the number of residual blocks in the second resolution layer.

[0013] By using the methods described above, the quality of the processing can be ensured.

[0014] In one possible implementation, the first image is input into the first model for processing, including: inputting the first image into the first model; encoding the first image through a 32-channel variational encoder in the first model to obtain a first feature; performing convolution processing on the first feature in a third resolution layer to obtain attention weights, wherein the third resolution layer is the resolution layer with the highest resolution among multiple resolution layers; determining a second feature output by the third resolution layer based on the attention weights, wherein the second feature is used as the input to the next resolution layer.

[0015] By employing a variational encoder with up to 32 channels, the first model can have a higher capacity to learn and represent the complex features of the data, capturing more data details so that the original data can be reconstructed more accurately in the subsequent decoding stage.

[0016] In one possible implementation, determining the second feature of the third resolution layer output based on the attention weights includes: standardizing the attention weights to obtain standardized attention weights, wherein the standardization process is used to limit the value of the attention weights within a preset range; and determining the second feature of the third resolution layer output based on the standardized attention weights.

[0017] By using the above method and standardizing the process, the weights are limited to a certain range, thus avoiding the problem of overestimating the calculated results.

[0018] In one possible implementation, an encoder is trained based on a first labeled image to obtain a 32-channel variational encoder; a model is trained based on a second labeled image to obtain a diffusion network; and a model is trained based on a sample image to obtain a conditional control generator network; wherein the noise of both the first and second labeled images is greater than that of the sample image.

[0019] The above method allows for accurate model training to obtain the first model.

[0020] Secondly, this application provides an image denoising apparatus. This apparatus can be an electronic device, a device within an electronic device, or a device compatible with an electronic device. The image denoising apparatus can also be a chip system, capable of executing the methods performed by the electronic device in the first aspect. The functions of the image denoising apparatus can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more units corresponding to the aforementioned functions. These units can be software and / or hardware. The operations and beneficial effects performed by the image denoising apparatus are described in the first aspect, and their repetitions will not be repeated.

[0021] Thirdly, this application provides an electronic device including one or more processors and one or more memories. The one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code, including computer instructions, which, when executed by the one or more processors, cause the electronic device to perform the image denoising method in any possible implementation of the first aspect described above.

[0022] Fourthly, this application provides a chip system including a processor and an interface, the processor and the interface being coupled; the interface is used to receive or output signals, and the processor is used to execute code instructions to perform the image denoising method in any possible implementation of the first aspect above.

[0023] Fifthly, this application provides a computer-readable storage medium storing a computer program / instructions that, when the computer program product is run on a computer, cause the computer to perform the image denoising method in any possible implementation of the first aspect described above.

[0024] Sixthly, this application provides a computer program product that, when run on a computer, causes the computer to perform the image denoising method in any possible implementation of the first aspect described above. Attached Figure Description

[0025] Figure 1A A schematic diagram of a Stable Diffusion UNet architecture provided in this application embodiment; Figure 1B This is a schematic diagram of the structure of a residual block provided in an embodiment of this application; Figure 1C This is a schematic diagram of a spatial transformation structure provided in an embodiment of this application; Figure 1D A schematic diagram of the architecture of a condition-controlled generation network provided in an embodiment of this application; Figure 2 A schematic flowchart illustrating an image denoising method provided in an embodiment of this application; Figure 3A A schematic diagram of an attention block structure provided in an embodiment of this application; Figure 3B A schematic diagram of a UNet network architecture provided in an embodiment of this application; Figure 3C A schematic diagram of a training process provided in an embodiment of this application; Figure 4 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application; Figure 5 A schematic diagram of the software structure of an electronic device provided in an embodiment of this application; Figure 6 This is a schematic diagram of an image denoising application process provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an image denoising device provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of a chip provided in an embodiment of this application. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.

[0027] It should be understood that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0028] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.

[0029] I. Diffusion Model The diffusion model mainly consists of two processes: noise addition and denoising. Noise addition means gradually adding Gaussian noise to the real images in the dataset, while denoising means gradually removing noise from the images to restore the real images.

[0030] II. Stable Diffusion's UNet Architecture To further improve the generation quality and training stability of the diffusion model, Stable Diffusion was introduced. Stable Diffusion is an improved diffusion model that significantly enhances the quality and consistency of image generation by incorporating the U-Net architecture and other techniques.

[0031] The UNet architecture of Stable Diffusion will be further introduced below.

[0032] Stable Diffusion comprises the following three core components: A text encoder (using the ViT-L / 14 text encoder of CLIP in Stable Diffusion) is used to convert the user-input text prompt into a text embedding. Image Auto Encoder-Decoder is used to encode an image into a latent vector z, or to reconstruct the image from the latent vector z. The UNET structure is used for iterative noise reduction. Multiple rounds of prediction are performed under text guidance to transform random Gaussian noise zt into image latent vector z0.

[0033] These three parts are independent of each other, with the UNET structure being the most important. UNET is the main component for generating images from noise. During the prediction process, by repeatedly calling UNET, the noise slice of the UNET prediction output is removed from the original noise, resulting in a progressively denoised image representation. The UNET of the Stable Diffusion Model contains approximately 860M parameters.

[0034] A UNet can include multiple modules, specifically including but not limited to the following four modules: ResNetBlock (residual block), Spatial Transformer Block (spatial transformation block), DownSample (downsampling block), and UpSample (upsampling block). It should be noted that a UNet can include multiple residual blocks and multiple spatial transformation blocks. In practical applications, such as... Figure 1A (For ease of description) Figure 1A The StableDiffusion UNet network architecture diagram shown below (only the downsampling part of the architecture is displayed) is an example with a resolution of h*w (this resolution is...). Figure 1A The StableDiffusion UNet structure shown is the highest-resolution UNet (or the highest-resolution layer). This h*w resolution layer includes two residual blocks and two spatial transformation blocks. For each resolution layer, the output of the previous layer is used as the input of the next residual block A, the output of residual block A is used as the input of spatial transformation block A, the output of spatial transformation block A is used as the input of residual block B, the output of residual block B is used as the input of spatial transformation block B, and finally, spatial transformation block B is downsampled and input to the next resolution layer.

[0035] in, Figure 1A The resolution of h*w is higher than that of h / 2*w / 2. For example, h*w can be 64*64, while h / 2*w / 2 can be 32*32. The values ​​of h and w can be the same or different.

[0036] The above will be further discussed below. Figure 1A The residual block, spatial transformation block, and downsampling in the code will be introduced separately: Residual Block: The structure of the residual block is as follows Figure 1B As shown, the residual block accepts two inputs: image features (also known as Latent vectors) and Time Embedding. The Latent vector is transformed by convolution and summed with the Time Embedding after fully connected projection, and then summed with the original Latent vector after skip connections. This summation is then fed into another convolutional layer to obtain the Latent output after residual block encoding transformation.

[0037] Spatial Transformation Block: The spatial transformation block also has two inputs: the output (Latent) of the previous module (i.e., the residual block) and the context embedding (the output of the text prompt after CLIP encoding). In this module, the image features (Latent vector) correspond to the image token, and cross-attention is performed with the context embedding (e.g., ...). Figure 1C As shown in the diagram, this attention mechanism injects semantic information from the Context Embedding into the corresponding image token, effectively fusing image and text information. In summary, the output latent size of the spatial transformation block maintains the same size as the input latent size, but semantic information is fused at the corresponding positions.

[0038] Downsampling primarily reduces the length and width of image features (Latent vectors) by a factor of two, achieved through a standard 2D convolution with a kernel size of 3 and a stride of 2. Correspondingly, upsampling ( Figure 1A The omitted parts (including upsampling) involve doubling the length and width of the latent vector using an interpolation algorithm. In summary, downsampling and upsampling only change the size of the latent vector while maintaining the same number of channels.

[0039] III. Conditional Control Generative Network – ControlNet While the aforementioned diffusion models perform well in generating images, they still have limitations in certain complex scenes and fine-grained control. To further enhance the capabilities of conditionally controlled generation, the ControlNet method was proposed, which improves the quality and consistency of generated images by introducing additional control signals. Compared to Stable Diffusion, ControlNet offers more control and flexibility, enabling it to achieve some functions that Stable Diffusion struggles to accomplish. The structure of ControlNet is as follows... Figure 1D As shown, the left half represents Stable Diffusion, whose parameters are frozen and not trained; the right half represents the ControlNet branch, whose parameters are trainable.

[0040] Each SD encoding module can also be called a resolution layer or a UNet. One resolution layer corresponds to one resolution; for example, SD encoding module_1 corresponds to a resolution of 64x64, and SD encoding module_2 corresponds to a resolution of 32x32. A resolution layer can include multiple modules, for example... Figure 1AThe resolution layer with a medium resolution of h*w includes two residual blocks and two spatial transformation blocks.

[0041] The aforementioned diffusion network / Stable Diffusion UNet / Conditional Controlled Generative Network (ControlNet) can be used for portrait enhancement, i.e., as a portrait enhancement model. The focus of portrait enhancement models is image denoising. While diffusion models include denoising processing, portrait enhancement models based on diffusion models have a very large number of parameters and are computationally intensive, making them unsuitable for deployment on mobile devices such as smartphones. This application provides an image denoising method that reduces the number of parameters and computational cost, thereby enabling the deployment of diffusion-based portrait enhancement models on mobile devices such as smartphones.

[0042] The image denoising method provided in the embodiments of this application is further described below: Please refer to Figure 2 As shown, Figure 2 This is a flowchart illustrating an image denoising method provided in an embodiment of this application. The image denoising method includes the following step 201. Figure 2 The method shown can be executed by an electronic device / the first model within an electronic device. Or... Figure 2 The subject of the method shown can be a chip or chip system in an electronic device, but this application does not limit it. Figure 2 The method will be explained using an electronic device as the executing entity. Specifically: 201. An electronic device inputs a first image into a first model for processing to obtain a second image. The first model is used to denoise the image. The noise in the second image is less than that in the first image. The first model includes multiple resolution layers. The resolution layers are used to predict the noise in the image. The number of residual blocks in the first resolution layer is less than the number of residual blocks in the second resolution layer. The resolution of the first resolution layer is greater than the resolution of the second resolution layer. The residual blocks are used to maintain the image gradient of the features propagated between the resolution layers.

[0043] Optionally, the first model can be used for portrait enhancement, or it can be said that the first model is a portrait enhancement model. The first model can also be used in other scenarios, such as image generation, etc., and this application does not limit it in this regard.

[0044] Optionally, the first image is a noisy image, and the second image is a denoised image. The noise in the first image can be similar to Gaussian noise. The first image can be an image captured by the camera of an electronic device, an image stored in the photo album of an electronic device, or an image downloaded from the Internet, etc. The first image includes a human face.

[0045] Optionally, the resolution layer may include residual blocks and / or spatial transform blocks, as well as downsampling / upsampling. If the resolution layer includes downsampling, then it will not include upsampling; similarly, if the resolution layer includes upsampling, then it will not include downsampling.

[0046] Optionally, the resolution layer can be the UNet described above.

[0047] Optionally, different resolution layers correspond to different resolutions, which can be found in h*w above.

[0048] Optionally, the resolution of the first resolution layer and the resolution of the second resolution layer are different. For example, the resolution of the first resolution layer is 64*64, and the resolution of the second resolution layer is 32*32. The first resolution layer and the second resolution layer can be adjacent resolution layers or non-adjacent resolution layers.

[0049] Optionally, the residual block can be referred to in the above description, and will not be repeated here.

[0050] In one possible embodiment, the number of residual blocks in the first resolution layer is less than the number of residual blocks in the second resolution layer, including: the number of residual blocks in the first resolution layer is 1, and the number of residual blocks in the second resolution layer is multiple.

[0051] Optionally, there can be multiple first resolution layers, and there can also be multiple second resolution layers. That is, the first model has multiple resolution layers with a residual block count of 1, or multiple resolution layers with no residual blocks. Similarly, the first model has multiple resolution layers that each contain multiple residual blocks.

[0052] In such Figure 1A In the Stable Diffusion UNet architecture shown, each resolution layer includes two residual blocks. In the first model, the number of residual blocks in the first resolution layer is 1, while the number of residual blocks in the second resolution layer is multiple. For example, the number of residual blocks in the resolution layer with resolution h*w is 1, and the number of residual blocks in the resolution layer with resolution h / 4*w / 4 is 4.

[0053] In one possible embodiment, the second resolution layer is the resolution layer with the lowest resolution among a plurality of resolution layers.

[0054] For example, the first model includes three resolution layers: resolution layer 1 (resolution is h*w), resolution layer 2 (resolution is h / 2*w / 2), and resolution layer 3 (resolution is h / 4*w / 4), where resolution layer 1 is the resolution layer with the highest resolution, and resolution layer 3 is the resolution layer with the lowest resolution. The first resolution layer consists of resolution layer 1 and resolution layer 2, each containing only one residual block. The second resolution layer is resolution layer 3, which contains three residual blocks.

[0055] In one possible embodiment, the first resolution layer is the resolution layer with the highest resolution among multiple resolution layers.

[0056] For example, the first model includes three resolution layers: resolution layer 1 (resolution is h*w), resolution layer 2 (resolution is h / 2*w / 2), and resolution layer 3 (resolution is h / 4*w / 4), where resolution layer 1 is the resolution layer with the highest resolution, and resolution layer 3 is the resolution layer with the lowest resolution. The first resolution layer is resolution layer 1, and the second resolution layer is resolution layer 3.

[0057] In one possible embodiment, the number of the first resolution layer is one less than the number of residual blocks in the second resolution layer.

[0058] For example, the first model includes three resolution layers: resolution layer 1 (resolution h*w), resolution layer 2 (resolution h / 2*w / 2), and resolution layer 3 (resolution h / 4*w / 4), where resolution layer 1 is the resolution layer with the highest resolution, and resolution layer 3 is the resolution layer with the lowest resolution. The first resolution layer consists of resolution layer 1 and resolution layer 2, and the second resolution layer is resolution layer 3. The number of residual blocks in resolution layer 3 is 1+2=3.

[0059] Optionally, the structure diagram of the residual block can be found in [reference needed]. Figure 1B As shown.

[0060] Optionally, all residual blocks in the first model can be improved residual blocks, specifically, they can be... Figure 1B The PaddedConv2D convolutional layer in the residual block structure shown is replaced with a Depthwise Separable Convolutional layer. The computational cost of the Depthwise Separable Convolutional layer is less than that of the PaddedConv2D convolutional layer.

[0061] Optional, such as Figure 1BAs shown, a residual block includes two PaddedConv2D convolutional layers. Either both PaddedConv2D convolutional layers can be replaced with Depthwise Separable Convolutional layers, or only one of the PaddedConv2D convolutional layers can be replaced with a Depthwise Separable Convolutional layer. This application does not impose any restrictions on this.

[0062] In one possible embodiment, the first image is input into a first model for processing, including: inputting the first image into the first model; encoding the first image using a 32-channel variational encoder in the first model to obtain a first feature; performing convolution processing on the first feature in a third resolution layer to obtain attention weights, wherein the third resolution layer is the resolution layer with the highest resolution among multiple resolution layers; determining a second feature output by the third resolution layer based on the attention weights, wherein the second feature is used as the input to the next resolution layer.

[0063] Optionally, this third resolution layer is the highest resolution layer, meaning it is processed first among all resolution layers. A second feature from the output of the third resolution layer is determined based on attention weights, and this second feature serves as the input to the next resolution layer.

[0064] Optionally, the attention block in the third resolution layer performs convolution processing on the first feature to obtain attention weights; the attention block in the third resolution layer determines the second feature output by the third resolution layer based on the attention weights.

[0065] Optionally, the next resolution layer also includes attention blocks, and the next resolution layer will perform the same processing as the third resolution layer.

[0066] Optionally, the first image is input into the first model for processing, including: inputting the first image into the first model; encoding the first image using a 32-channel variational encoder in the first model to obtain a first feature; performing a zero convolution on the first feature; and combining the result of the zero convolution with a first convolutional layer (the first convolutional layer is as follows...). Figure 1D The outputs of the convolutional layers in the left-hand network are summed (the parameters of the left-hand network are frozen); the summed result (Latent vector) is then input into the third resolution layer (e.g., ...). Figure 1A The SD encoding module_1 in the right-hand network of the network.

[0067] The existing models typically use a 4-channel variational encoder. Using a 32-channel variational encoder allows the first model to retain more image features.

[0068] Optionally, the third resolution layer includes residual blocks and attention blocks, the attention blocks being different from... Figure 1C The space transformation block shown. That is, with... Figure 1A The resolution layers shown differ in their residual blocks and spatial transformation blocks. In the first model of this application, the resolution layers include residual blocks and attention blocks. That is, the spatial transformation block is replaced with an attention block. Or, the spatial transformation block is improved to obtain a new attention block.

[0069] because Figure 1C The purpose of the spatial attention blocks shown is to fuse semantic information at corresponding locations to achieve text-based image generation; while the first model in this application is used for denoising to achieve portrait enhancement. Therefore, removing the computation of the text / semantic parts can further reduce the model parameters and the computational load when applying the model.

[0070] In one possible embodiment, determining the second feature of the third resolution layer output based on the attention weights includes: standardizing the attention weights to obtain standardized attention weights, wherein the standardization process is used to limit the value of the attention weights within a preset range; and determining the second feature of the third resolution layer output based on the standardized attention weights.

[0071] Optionally, the attention block in the third resolution layer performs a normalization process on the attention weights to obtain normalized attention weights; the attention block in the third resolution layer determines the second feature output by the third resolution layer based on the normalized attention weights.

[0072] Optionally, the attention block includes a normalization process, which can be performed using the softmax function. This involves inputting the data into the softmax function for normalization. However, the softmax function is an exponential function; if the input value is too large, it can cause the exponential function to overflow. Therefore, standardization is performed to ensure that the model can be trained to produce effective outputs without loading pre-trained weights during the training phase, and to make the model more stable in application.

[0073] For example, such as Figure 3A As shown, Figure 3AThis is a schematic diagram of an attention block structure provided in an embodiment of this application. Image features (Latent vectors) are used as input to the attention block. After convolution processing of the image features, a query matrix, a key matrix, and a value matrix are obtained. The query matrix and the key matrix are standardized respectively. The two standardized matrices are then multiplied, and finally, the attention width matrix is ​​obtained through scaling and normalization. The attention width matrix and the value matrix are then multiplied, and the output is obtained through a convolutional layer. This output is also a Latent vector.

[0074] The architecture of the first model provided in the embodiments of this application will be further described below. For example... Figure 3B As shown, Figure 3B This is a schematic diagram of a UNet network architecture in a first model provided in an embodiment of this application.

[0075] Optional, compared to Figure 1A The UNet network structure diagram shown is as follows: Figure 3B In the UNet network architecture diagram shown, each resolution layer except the lowest resolution layer has only one residual block (that is, in...). Figure 1A The UNet network shown here has one residual block removed from the lowest resolution layer. The spatial transformation block in each layer is replaced with an attention block, and each resolution layer except the lowest resolution layer has only one attention block. Figure 3B The UNet network architecture diagram shown adds more residual blocks and more attention blocks to the lowest resolution layer.

[0076] The application phase of the first model has been described above. The training phase of the first model will be described below: In one possible embodiment, the electronic device trains an encoder based on a first labeled image to obtain a 32-channel variational encoder; trains a model based on a second labeled image to obtain a diffusion network; and trains a model based on a sample image to obtain a conditional control generative network; wherein the noise of both the first labeled image and the second labeled image is greater than that of the sample image.

[0077] Optionally, the first model includes two networks: a diffusion network and a conditional control generation network. The architecture of the first network can be found in [reference needed]. Figure 1D As shown, Figure 1D The middle left section (the part where the high-resolution image is input) is a diffusion network, also known as a Stable Diffusion network. Figure 1DThe right part (the low-resolution image input portion) is the Conditional Controlled Generation Network. This network provides greater control and flexibility, enabling the first model to achieve some functionalities that are difficult for the StableDiffusion network to accomplish.

[0078] Optionally, the electronic device first trains the encoder based on the first labeled image to obtain a 32-channel variational encoder; after the 32-channel variational encoder is trained, the model is trained based on the second labeled image to obtain a diffusion network; after the diffusion network is trained, the model is trained based on the sample images to obtain a conditional control generative network.

[0079] The first label image and the second label image can be the same image or different images. The first label image and the second label image can be high-definition images, that is, images with low noise; the sample image can be a low-definition image, that is, images with high noise.

[0080] For example, the training process can be found in [reference needed]. Figure 3C As shown, the autoencoder is first trained. After the autoencoder is trained, the parameters in the autoencoder are locked, and the lightweight UNet network (diffusion network) is trained. After the UNet network is trained, the parameters of the autoencoder and the lightweight UNet network are locked, and the conditional control generator network is trained.

[0081] The hardware structure of electronic devices is described below: Please see Figure 4 , Figure 4 This is a schematic diagram of the hardware structure of the electronic device 100 provided in this application embodiment. The electronic device 100 can be an electronic device corresponding to the training phase or an electronic device corresponding to the application phase.

[0082] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0083] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0084] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0085] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.

[0086] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system. The processor 110 retrieves the instructions or data stored in the memory, causing the electronic device 100 to execute the imaging method performed by the electronic device in the following method embodiments.

[0087] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0088] The charging management module 140 is used to receive charging input from the charger. The charger can be a wireless charger or a wired charger.

[0089] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, internal memory 121, external memory, display 194, camera 193, and wireless communication module 160, etc. In some other embodiments, the power management module 141 may also be located in the processor 110.

[0090] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0091] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.

[0092] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0093] A modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor.

[0094] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLAN) (such as Wi-Fi), Bluetooth (BT), BLE broadcasting, global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR). The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signal, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0095] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, so that electronic device 100 can communicate with networks and other devices through wireless communication technology.

[0096] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0097] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. In some embodiments, electronic device 100 may include one or N displays screens 194, where N is a positive integer greater than 1. Display screen 194 may include OLED screens.

[0098] Optionally, the display 194 may further include: an OLED glass layer, an OLED light-emitting unit, a fingerprint recognition sensor, a microlens array, etc. The display 194 supports optical in-display fingerprint recognition.

[0099] Electronic device 100 can achieve shooting functions through an ISP, camera 193, video codec, GPU, display screen 194, and application processor. The ISP processes data fed back by the camera 193. The camera 193 captures still images or videos. The camera 193 may include a front-facing camera and a rear-facing camera; the front-facing camera is located on the display area of ​​the screen, and the rear-facing camera is located on the back area of ​​the screen. The digital signal processor processes digital signals, including digital image signals and other digital signals. The video codec is used to compress or decompress digital video. Electronic device 100 may support one or more video codecs.

[0100] NPU stands for Neural Network (NN) Computing Processor. By drawing inspiration from the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can quickly process input information and continuously learn on its own.

[0101] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to perform data storage functions.

[0102] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as a sound playback function), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data), etc. Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as flash memory devices.

[0103] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0104] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.

[0105] A speaker 170A, also called a "loudspeaker," is used to convert audio electrical signals into sound signals. A receiver 170B, also called a "handpiece," is used to convert audio electrical signals into sound signals. A microphone 170C, also called a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. A headphone jack 170D is used to connect wired headphones. A pressure sensor 180A is used to sense pressure signals and can convert the pressure signals into electrical signals. In some embodiments, the pressure sensor 180A may be located on the display screen 194. A gyroscope sensor 180B can be used to determine the motion posture of the electronic device 100. A barometric pressure sensor 180C is used to measure barometric pressure. A magnetic sensor 180D includes a Hall effect sensor. An accelerometer 180E can detect the magnitude of acceleration of the electronic device 100 in various directions (generally three axes). A distance sensor 180F is used to measure distance. A proximity sensor 180G may include, for example, a light-emitting diode (LED) and a photosensor. An ambient light sensor 180L is used to sense ambient light intensity. A fingerprint sensor 180H is used to collect fingerprints. Temperature sensor 180J is used to detect temperature. Touch sensor 180K, also known as a "touch panel," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touch screen." Touch sensor 180K detects touch operations applied to or near it. Bone conduction sensor 180M can acquire vibration signals. Buttons 190 include power button, volume buttons, etc. Motor 191 can generate vibration prompts. Indicator 192 can be an indicator light, used to indicate charging status, battery level changes, messages, missed calls, notifications, etc. SIM card interface 195 is used to connect a SIM card.

[0106] Furthermore, an operating system runs on top of the aforementioned components. Examples include iOS and Android. The operating system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of electronic device 100. It should be noted that although this application embodiment uses the Android system as an example for illustration, its basic principles are equally applicable to electronic devices with other operating systems.

[0107] The software structure of electronic device 100 is described below: Figure 5This is a schematic diagram of the software structure of an electronic device 100 provided in an embodiment of this application. The software structure adopts a layered architecture, which divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In this embodiment, the operating system (taking Android as an example, with Android running on an AP) can be divided into four layers, from top to bottom: the application layer (APP), the application framework layer (FWK), the hardware abstraction layer (HAL), and the kernel layer.

[0108] The application layer can include a series of application packages. For example... Figure 5 As shown, the application package may include applications such as a camera and a gallery. In this embodiment, "camera" refers to a camera application. A camera application may include a camera interface module (which may be called a CameraApi2Module), etc. "Gallery" refers to a gallery application, which is used to store images and videos taken by electronic devices. This gallery application also provides users with playback functionality, allowing users to view historically captured images and videos within the gallery application.

[0109] The application framework layer provides application developers with an application programming interface (API) framework and various services and management tools to access core functionalities, including interface management, data access, application-layer messaging, application package management, telephony management, and location management. The application framework layer includes some predefined functions. For example... Figure 5 As shown, the application framework layer may include, but is not limited to, the camera service CameraService.

[0110] CameraService is responsible for scheduling the startup process of the camera application, creating and managing processes, and creating and managing windows. In this application, the portrait enhancement model can be built into the camera service or be a module independent of the camera service. For ease of description, we will take the example of the portrait enhancement model being built into the camera service.

[0111] The Hardware Abstraction Layer (HAL) is an interface layer located between the operating system kernel and the hardware circuitry. Its purpose is to abstract the hardware. It hides the hardware interface details of a specific platform, providing the operating system with a virtual hardware platform. For example... Figure 5As shown, the hardware abstraction layer can include Camera Resource Service and Camera Provider. In addition, this hardware abstraction layer can also include Camera Device Session interface and PreviewFlowImpl interface. Specifically, Camera Resource Service interacts with the memory modules in the hardware; Camera Provider enumerates individual devices and manages their states, enabling the opening and closing of physical camera devices (such as rear cameras); Camera Device Session creates camera device sessions and stores the attributes and configuration information required for those sessions; and PreviewFlowImpl is responsible for informing the app that the first frame of the preview has been displayed.

[0112] The kernel layer is the layer between hardware and software. The kernel layer includes at least display drivers, camera drivers, etc.

[0113] Based on the above Figure 4 and Figure 5 The following section further describes the application phases of the first model, as described in the previous section. The application phases of this first model mainly involve the following four parts: sensor (hardware), ISP (hardware), camera service (application framework layer), and image library. The example used is the camera service. Figure 6 As shown: First, after the sensor captures the raw image, it sends the acquired raw image to the ISP (for electronic devices such as mobile phones, the raw image is usually processed by the ISP).

[0114] Then, the ISP processes the original image to obtain the first image. This ISP processing includes, but is not limited to, the following: Gamma processing for gamma correction to adjust the image's contrast and brightness, making the image more consistent with human visual characteristics; CCM processing for color correction to adjust and optimize the image's color performance; and LSC processing to address issues of uneven image brightness and color caused by lens optical characteristics.

[0115] Next, the first model embedded in the camera service processes the first image to obtain the second image. Finally, the camera service sends the second image to the image library, which saves the second image.

[0116] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of an image denoising device 700 provided in an embodiment of this application. Figure 7The image denoising device shown can be an electronic device, a device within an electronic device, or a device that can be used in conjunction with an electronic device. Figure 7 The image denoising apparatus shown may include a processing unit 701. Wherein: The processing unit 701 is used to input the first image into the first model for processing to obtain the second image. The first model is used to denoise the image. The noise in the second image is less than that in the first image. The first model includes multiple resolution layers. The resolution layers are used to predict the noise in the image. The number of residual blocks in the first resolution layer is less than the number of residual blocks in the second resolution layer. The resolution of the first resolution layer is greater than the resolution of the second resolution layer. The residual blocks are used to maintain the image gradient of the features propagated between the resolution layers.

[0117] In one possible implementation, the number of residual blocks in the first resolution layer is less than the number of residual blocks in the second resolution layer, including: the number of residual blocks in the first resolution layer is 1, and the number of residual blocks in the second resolution layer is multiple.

[0118] In one possible implementation, the second resolution layer is the resolution layer with the lowest resolution among multiple resolution layers.

[0119] In one possible implementation, the first resolution layer is the resolution layer with the highest resolution among multiple resolution layers.

[0120] In one possible implementation, the number of the first resolution layer is one less than the number of residual blocks in the second resolution layer.

[0121] In one possible implementation, the processing unit 701 is further configured to input the first image into the first model; encode the first image using a 32-channel variational encoder in the first model to obtain a first feature; perform convolution processing on the first feature in the third resolution layer to obtain attention weights, wherein the third resolution layer is the resolution layer with the highest resolution among multiple resolution layers; determine the second feature output by the third resolution layer based on the attention weights, and use the second feature as the input to the next resolution layer.

[0122] In one possible implementation, the processing unit 701 is further configured to perform a standardization process on the attention weights to obtain standardized attention weights. The standardization process is used to limit the values ​​of the attention weights to a preset range. Based on the standardized attention weights, the second feature of the output of the third resolution layer is determined.

[0123] In one possible implementation, the processing unit 701 is further configured to train an encoder based on the first label image to obtain a 32-channel variational encoder; train a model based on the second label image to obtain a diffusion network; and train a model based on the sample image to obtain a conditional control generator network; wherein the noise of both the first label image and the second label image is greater than that of the sample image.

[0124] For cases where the image denoising device can be a chip or a chip system, please refer to [link / reference]. Figure 8 The diagram shows the structure of the chip. Figure 8 The chip 800 shown includes a processor 801 and an interface 802. Optionally, it may also include a memory 803. The number of processors 801 can be one or more, and the number of interfaces 802 can be multiple.

[0125] For cases where the chip is used to implement the electronic device in the embodiments of this application: The interface 802 is used to receive or output signals; The processor 801 is used to perform data processing operations of the electronic device.

[0126] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0127] It is understood that some optional features in the embodiments of this application can be implemented independently in certain scenarios without relying on other features, such as the current solution on which they are based, to solve the corresponding technical problems and achieve the corresponding effects. Alternatively, they can be combined with other features as needed in certain scenarios. Accordingly, the audio data archiving device given in the embodiments of this application can also implement these features or functions, which will not be elaborated here.

[0128] It should be understood that the processor in the embodiments of this application can be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor described above can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0129] It is understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0130] This application also provides a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed on an electronic device, implement the functions of any of the above method embodiments.

[0131] This application also provides a computer program product that, when run on a computer, enables the computer to perform the functions of any of the above method embodiments.

[0132] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).

[0133] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An image denoising method, characterized in that, The method includes: A first image is input into a first model for processing to obtain a second image. The first model is used to denoise the image. The noise in the second image is less than that in the first image. The first model includes multiple resolution layers. The resolution layers are used to predict noise in the image. The number of residual blocks in the first resolution layer is less than the number of residual blocks in the second resolution layer. The resolution of the first resolution layer is greater than the resolution of the second resolution layer. The residual blocks are used to maintain the image gradient of the features propagated between resolution layers. The first resolution layer includes residual blocks and attention blocks. The output of the residual blocks in the first resolution layer is the input of the attention blocks in the first resolution layer. The output of the attention blocks in the first resolution layer is the input of the residual blocks in the next resolution layer. The step of inputting the first image into the first model for processing includes: The first image is encoded using a 32-channel variational encoder in the first model to obtain the first feature; In the third resolution layer, the first feature is convolved to obtain attention weights. The third resolution layer is the resolution layer with the highest resolution among the plurality of resolution layers. The attention weights are normalized to obtain normalized attention weights. The normalized attention weights are standardized to obtain standardized attention weights, and the standardization process is used to limit the values ​​of the attention weights to a preset range. The second feature of the third resolution layer output is determined based on the attention weights after the standardization process, and the second feature is used as the input of the next resolution layer.

2. The method according to claim 1, characterized in that, The number of residual blocks in the first resolution layer is less than the number of residual blocks in the second resolution layer, including: The number of residual blocks in the first resolution layer is 1, and the number of residual blocks in the second resolution layer is multiple.

3. The method according to claim 2, characterized in that, The second resolution layer is the resolution layer with the lowest resolution among the plurality of resolution layers.

4. The method according to claim 3, characterized in that, The number of the first resolution layer is one less than the number of residual blocks in the second resolution layer.

5. The method according to any one of claims 1-4, characterized in that, The first model includes a 32-channel variational encoder, a diffusion network, and a conditional control generation network; the method further includes: The encoder is trained based on the first labeled image to obtain the 32-channel variational encoder. The diffusion network is obtained by training the model based on the second labeled image; The conditional control generative network is obtained by training the model based on the sample images; The noise levels in both the first and second labeled images are greater than those in the sample image.

6. An electronic device comprising one or more memories and one or more processors, characterized in that, The memory is used to store a computer program; the processor is used to invoke the computer program to cause the electronic device to perform the method of any one of claims 1-5.

7. A chip system for use in electronic devices, characterized in that, The chip system includes at least one processor and an interface for receiving instructions and transmitting them to the at least one processor; the at least one processor executes the instructions to cause the electronic device to perform the method as described in any one of claims 1-5.

8. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1-5.

9. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Infrared image deblurring algorithm based on attention mechanism residual network model

    CN115345791A

  • Digital watermark attack method based on conditional diffusion model

    CN116645260A