Neural network training method and device and image processing method and device
Through neural network training methods, using the enhancement network and degraded network to simultaneously learn image enhancement and degradation, the problem of inability to effectively deal with complex degraded cores in the prior art is solved, and a higher quality image enhancement effect is achieved.
Patent Information
- Application Number
- CN202280100381.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-27
- Publication Date
- 2025-06-03
AI Technical Summary
Existing image enhancement methods cannot effectively deal with image degradation caused by complex degraded nuclei, and it is assumed that the degraded nuclei belongs to the anisotropic Gaussian fuzzy nuclei family, which is inconsistent with the complex degraded nuclei in the real world.
Through the neural network training method, multiple training image pairs are acquired, each of which includes high-quality images and low-quality images. The enhanced network and degraded network are used to learn image enhancement and image degradation simultaneously to generate high-quality output images.
Improve the quality of the output image, and better understand and handle complex degradation processes, thereby achieving better results in image enhancement.
Smart Images

Figure CN120092246A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of artificial intelligence, and more specifically, to neural network training methods and devices, and image processing methods and devices. Background Art
[0002] With the development of terminals, it is very common to take photos using terminals (especially handheld terminals equipped with cameras). However, the quality of the photos taken (especially the detail quality) is usually affected by various degradation sources, as well as prominent degradation sources in real life such as low-quality camera optics and motion blur. Therefore, the task of image enhancement is one of the most important aspects in the field of image processing, especially for images taken by the cameras of handheld terminals (such as mobile phones).
[0003] Image enhancement methods include single image super resolution (SISR) and single image deblurring. Single image super resolution aims to improve the image resolution while restoring the influence of the point spread function (PSF) of the camera lens on the image details. Single image deblurring aims to restore the image details blurred due to camera motion and / or scene object motion.
[0004] Most image enhancement methods are based on convolutional neural networks (CNNs). However, existing image enhancement methods either do not involve any degradation information about the degradation process or assume that the degradation kernel belongs to the anisotropic Gaussian blur kernel family, that is, the convolution of classical downsampling filters, which is inconsistent with the complex degradation kernels in the real world. Summary of the Invention
[0005] Embodiments of the present application provide neural network training methods and devices, and image processing methods and devices, which can simultaneously learn image enhancement and image degradation, thereby improving the quality of the output image.
[0006] According to a first aspect, an embodiment of the present application provides a neural network training method, and the neural network training method includes: obtaining a plurality of training image pairs, where each pair of the plurality of training image pairs includes a high-quality image and a low-quality image, and the high-quality image and the low-quality image are images of the same scene with different qualities; inputting the low-quality image into an enhancement network to generate an enhanced image; inputting the high-quality image and at least two intermediate tensors into a degradation network to generate a degraded image, where the at least two intermediate tensors are determined based on the enhancement network; determining a first loss function based on the high-quality image and the enhanced image; determining a second loss function based on the low-quality image and the degraded image; and updating the neural network based on the first loss function and the second loss function, where the neural network includes the enhancement network and the degradation network.
[0007] According to the above technical solution, the training process involves learning a loss function that simultaneously learns the enhancement loss and the degradation loss. Therefore, by means of the degradation network, information about the degradation process is extracted from the low-quality image, so as to improve the quality of the output image when using the enhancement network to process the image after the training is completed.
[0008] In an alternative implementation, the enhancement network includes a first encoder and a first decoder; the at least two intermediate tensors include a first intermediate tensor and a second intermediate tensor, where the first intermediate tensor is determined based on the bottleneck tensor in the first encoder, and the second intermediate tensor is determined based on at least one network layer in the network layer in the first decoder.
[0009] According to the above technical solution, the first intermediate tensor is extracted from the bottleneck tensor in the first encoder in the enhancement network, and thus includes global information about the degradation of the low-quality image. The second intermediate tensor is extracted from several network layers in the first decoder and can be processed into a local component tensor, and the local component tensor includes local information about the degradation of the low-quality image. By inputting this information into the degradation network, the degradation process can be better understood, so as to improve the image quality in the enhancement network after the training is completed.
[0010] The enhancement network is applicable to many CNN-based networks including an encoder and a decoder (the first encoder and the first decoder in the present application), so that the hardware and software requirements can be highly flexibly met.
[0011] In an alternative implementation, the degradation network includes a first sub-network, and inputting the high-quality image and at least two intermediate tensors into the degradation network to generate a degraded image includes: inputting the first intermediate tensor and the second intermediate tensor into the first sub-network to generate a per-pixel kernel map, wherein the convolutional weights of multiple network layers of the first sub-network are determined based on the first intermediate tensor; inputting the high-quality image into the degradation network; convolving the high-quality image with the per-pixel kernel to generate the degraded image, wherein the per-pixel kernel is determined from the per-pixel kernel map.
[0012] According to the above technical solution, the per-pixel kernel map is estimated by combining the global information and local information about the degradation process. It is a spatially-varying degradation kernel map estimated by the degradation network. The decoupling of the global information and local information about the degradation process greatly reduces the internal dimension of the task of estimating the spatially-varying degradation kernel map (per-pixel kernel map), thereby greatly reducing its ill-posedness and making the per-pixel kernel map more consistent.
[0013] In an alternative implementation, the height×width size of the first intermediate tensor is 1×1 or 2×2 or 3×3; or the height×width size of the first intermediate tensor is the same as the height×width size of the bottleneck tensor, and the number of channels of the first intermediate tensor is 1 / 8 or 1 / 4 or 1 / 2 of the number of channels of the bottleneck tensor.
[0014] According to the above technical solution, the first intermediate tensor is directly extracted based on a part of the bottleneck tensor, or processed into a tensor with a higher dimension but a smaller size.
[0015] In an alternative implementation, the method further includes: injecting the processing result of the first intermediate tensor into at least one network layer in the first decoder, wherein the spatial size of the processing result of the first intermediate tensor is 1×1.
[0016] According to the above technical solution, the first intermediate tensor including the global information about the degradation process is processed to be injected into the first decoder of the enhancement network without much calculation. The first intermediate tensor (global component tensor) includes information about image degradation specific to the entire image (e.g., information about the camera lens, camera motion, etc.). Therefore, if this information is provided to the first decoder of the enhancement network, the first decoder can better recover the details of the high-quality image. On the other hand, since the global component tensor has a small spatial size (i.e., height×width), its estimation (and the injection of its compressed variant) does not require much calculation. Therefore, the enhancement image quality is improved at the cost of a moderately increased computational complexity.
[0017] In an alternative implementation, the first sub-network includes a pointwise convolution weight estimation sub-network and a kernel map estimation sub-network. The step of inputting the first intermediate tensor and the second intermediate tensor into the first sub-network to generate a per-pixel kernel map includes: inputting the first intermediate tensor into the pointwise convolution weight estimation sub-network to generate a plurality of dynamic pointwise convolution weights; inputting the second intermediate tensor into the first sub-network and determining a local component tensor based on the second intermediate tensor; and inputting the local component tensor and the plurality of dynamic pointwise convolution weights into the kernel map estimation sub-network to generate the per-pixel kernel map, wherein the convolution weights of multiple network layers of the kernel map estimation sub-network are the plurality of dynamic pointwise convolution weights.
[0018] The kernel map estimation sub-network is constructed to include several pointwise-convolution-based MLPs. Since the space for estimating the degradation process is richer, the quality of the output enhanced image is better after training.
[0019] The pointwise convolution weight estimation sub-network generates the weights for dynamic convolution. Since it can better adapt to each specific input image, it can improve the quality of the output enhanced image after training.
[0020] In an alternative implementation, the first sub-network further includes a kernel embedding map estimation sub-network. The step of inputting the second intermediate tensor into the first sub-network and determining a local component tensor based on the second intermediate tensor includes: inputting the second intermediate tensor into the kernel embedding map estimation sub-network to generate the local component tensor.
[0021] According to the above technical solution, the degradation network includes three sub-networks: a pointwise convolution weight estimation sub-network, a kernel map estimation sub-network, and a kernel embedding map estimation sub-network. The first intermediate tensor is input into the pointwise convolution weight estimation sub-network to generate a plurality of dynamic pointwise convolution weights. The second intermediate tensor is input into the kernel embedding map estimation sub-network to generate the local component tensor. The local component tensor is then input into the kernel map estimation sub-network to generate the per-pixel kernel map, wherein the dynamic pointwise convolution weights are also input into the kernel map estimation sub-network to be used as convolution weights. The per-pixel kernel map is processed into a per-pixel kernel, which is further used to convolve with the high-quality image to generate a degraded image. The degradation information learned from the degradation process will improve the performance of the enhancement network.
[0022] In an alternative implementation, the enhancement network further includes a kernel embedding graph estimation sub-network, and the local component tensor is the second intermediate tensor. The method further includes: inputting at least one tensor in the network layer tensors in the first decoder into the kernel embedding graph estimation sub-network to generate the second intermediate tensor; concatenating the second intermediate tensor with one network layer tensor in the network layer tensors in the first decoder.
[0023] According to the above technical solution, the kernel embedding graph estimation sub-network is included in the enhancement network, and the second intermediate tensor input into the degradation network is the local component tensor used in the kernel graph estimation network. For example, the kernel embedding graph estimation sub-network is not very complex, thereby reducing the computational cost during the training process.
[0024] In an alternative implementation, the height×width size of the local component tensor is the same as the height×width size of the high-quality image, and the number of channels of the local component tensor is 4 or 8 or 16 or 24 or 32.
[0025] In an alternative implementation, a skip connection is provided between the first encoder and the first decoder.
[0026] According to the above technical solution, a skip connection is provided between the first encoder and the first decoder to reduce the problems of network degradation and gradient vanishing and explosion, thereby improving the performance of the neural network.
[0027] In an alternative implementation, the enhancement network adopts a U-net design.
[0028] According to a second aspect, an embodiment of the present application provides an image processing method, including: obtaining an image to be processed; processing the image to be processed using an enhancement network, where the enhancement network is obtained by training a neural network including the enhancement network and a degradation network based on a plurality of training image pairs, and each pair in the plurality of training image pairs includes a high-quality image and a low-quality image. The neural network is trained based on a first loss function and a second loss function, where the first loss function is determined based on the high-quality image and the enhanced image, and the second loss function is determined based on the low-quality image and the degraded image. The enhanced image is generated by the enhancement network using the low-quality image as an input, and the degraded image is generated by the degradation network using the high-quality image and at least two intermediate tensors as inputs, where the at least two intermediate tensors are determined based on the enhancement network.
[0029] In an alternative implementation, the enhancement network includes a first encoder and a first decoder; the at least two intermediate tensors include a first intermediate tensor and a second intermediate tensor, wherein the first intermediate tensor is determined based on a bottleneck tensor in the first encoder, and the second intermediate tensor is determined based on at least one network layer in the network layers of the first decoder.
[0030] In an alternative implementation, the degradation network includes a first sub-network, and the degraded image is obtained by convolving the high-quality image with a per-pixel kernel, the per-pixel kernel is determined from the per-pixel kernel map, and the per-pixel kernel map is generated by the first sub-network using the first intermediate tensor and the second intermediate tensor as inputs, wherein the convolutional weights of multiple network layers of the first sub-network are determined based on the first intermediate tensor.
[0031] In an alternative implementation, the height×width dimension of the first intermediate tensor is 1×1 or 2×2 or 3×3; or the height×width dimension of the first intermediate tensor is the same as the height×width dimension of the bottleneck tensor, and the number of channels of the first intermediate tensor is 1 / 8 or 1 / 4 or 1 / 2 of the number of channels of the bottleneck tensor.
[0032] In an alternative implementation, the processing result of the first intermediate tensor is also injected into at least one network layer in the first decoder, wherein the spatial dimension of the processing result of the first intermediate tensor is 1×1.
[0033] In an alternative implementation, the first sub-network includes a pointwise convolutional weight estimation sub-network and a kernel map estimation sub-network, and the per-pixel kernel map is generated by the kernel map estimation sub-network using a local component tensor and multiple dynamic pointwise convolutional weights as inputs, wherein the convolutional weights of multiple network layers of the kernel map estimation sub-network are the multiple dynamic pointwise convolutional weights; the multiple dynamic pointwise convolutional weights are generated by the pointwise convolutional weight estimation sub-network using the first intermediate tensor as an input, and the local component tensor is determined based on the second intermediate tensor.
[0034] In an alternative implementation, the first sub-network further includes a kernel embedding map estimation sub-network, and the local component tensor is determined by the kernel embedding map estimation sub-network using the second intermediate tensor as an input.
[0035] In an alternative implementation, the enhancement network further includes a kernel embedding graph estimation sub-network, the local component tensor is the second intermediate tensor, and the second intermediate tensor is generated by the kernel embedding graph estimation sub-network using at least one network layer tensor in the network layer tensors in the first decoder as input; the second intermediate tensor is concatenated with one network layer tensor in the network layer tensors in the first decoder.
[0036] In an alternative implementation, the height×width size of the local component tensor is the same as the height×width size of the high-quality image, and the number of channels of the local component tensor is 4 or 8 or 16 or 24 or 32.
[0037] In an alternative implementation, a skip connection is provided between the first encoder and the first decoder.
[0038] In an alternative implementation, the enhancement network adopts a U-net design.
[0039] According to a third aspect, an embodiment of the present application provides a neural network training device, including: a memory storing computer program instructions; a processor coupled to the memory and, when executing the computer program instructions, configured to: obtain a plurality of training image pairs, where each pair of the plurality of training image pairs includes a high-quality image and a low-quality image, and the high-quality image and the low-quality image are images of the same scene with different qualities; input the low-quality image into the enhancement network to generate an enhanced image; input the high-quality image and at least two intermediate tensors into the degradation network to generate a degraded image, where the at least two intermediate tensors are determined based on the enhancement network; determine a first loss function based on the high-quality image and the enhanced image; determine a second loss function based on the low-quality image and the degraded image; and update the neural network based on the first loss function and the second loss function, where the neural network includes the enhancement network and the degradation network.
[0040] In an alternative implementation, the enhancement network includes a first encoder and a first decoder; the at least two intermediate tensors include a first intermediate tensor and a second intermediate tensor, where the first intermediate tensor is determined based on the bottleneck tensor in the first encoder, and the second intermediate tensor is determined based on at least one network layer in the first decoder.
[0041] In an alternative implementation, the degradation network includes a first sub-network. When the processor executes the computer program instructions, it is specifically configured to: input the first intermediate tensor and the second intermediate tensor into the first sub-network to generate a per-pixel kernel map, where the convolution weights of multiple network layers of the first sub-network are determined based on the first intermediate tensor; input the high-quality image into the degradation network; perform convolution on the high-quality image with the per-pixel kernel to generate the degraded image, where the per-pixel kernel is determined from the per-pixel kernel map.
[0042] In an alternative implementation, the height×width dimension of the first intermediate tensor is 1×1 or 2×2 or 3×3; or the height×width dimension of the first intermediate tensor is the same as the height×width dimension of the bottleneck tensor, and the number of channels of the first intermediate tensor is 1 / 8 or 1 / 4 or 1 / 2 of the number of channels of the bottleneck tensor.
[0043] In an alternative implementation, when the processor executes the computer program instructions, it is further configured to: inject the processing result of the first intermediate tensor into at least one network layer in the first decoder, where the spatial dimension of the processing result of the first intermediate tensor is 1×1.
[0044] In an alternative implementation, the first sub-network includes a pointwise convolution weight estimation sub-network and a kernel map estimation sub-network. When the processor executes the computer program instructions, it is specifically configured to: input the first intermediate tensor into the pointwise convolution weight estimation sub-network to generate multiple dynamic pointwise convolution weights; input the second intermediate tensor into the first sub-network and determine a local component tensor based on the second intermediate tensor; input the local component tensor and the multiple dynamic pointwise convolution weights into the kernel map estimation sub-network to generate the per-pixel kernel map, where the convolution weights of multiple network layers of the kernel map estimation sub-network are the multiple dynamic pointwise convolution weights.
[0045] In an alternative implementation, the first sub-network further includes a kernel embedding map estimation sub-network. When the processor executes the computer program instructions, it is specifically configured to: input the second intermediate tensor into the kernel embedding map estimation sub-network to generate the local component tensor.
[0046] In an alternative implementation, the enhancement network further includes a kernel embedding graph estimation sub-network, and the local component tensor is the second intermediate tensor. When executing the computer program instructions, the processor is further configured to: input at least one tensor in the network layer tensors in the first decoder into the kernel embedding graph estimation sub-network to generate the second intermediate tensor; concatenate the second intermediate tensor with one of the network layer tensors in the network layer tensors in the first decoder.
[0047] In an alternative implementation, the height × width size of the local component tensor is the same as the height × width size of the high-quality image, and the number of channels of the local component tensor is 4 or 8 or 16 or 24 or 32.
[0048] In an alternative implementation, a skip connection is provided between the first encoder and the first decoder.
[0049] In an alternative implementation, the enhancement network adopts a U-net design.
[0050] According to a fourth aspect, an embodiment of the present application provides an image processing apparatus, including: a memory storing computer program instructions; a processor coupled to the memory and configured to, when executing the computer program instructions: obtain an image to be processed; process the image to be processed using an enhancement network, where the enhancement network is obtained by training a neural network including the enhancement network and a degradation network based on a plurality of training image pairs, and each pair of the plurality of training image pairs includes a high-quality image and a low-quality image. The neural network is trained based on a first loss function and a second loss function, where the first loss function is determined based on the high-quality image and an enhanced image, and the second loss function is determined based on the low-quality image and a degraded image. The enhanced image is generated by the enhancement network using the low-quality image as an input, and the degraded image is generated by the degradation network using the high-quality image and at least two intermediate tensors as inputs, where the at least two intermediate tensors are determined based on the enhancement network.
[0051] In an alternative implementation, the enhancement network includes a first encoder and a first decoder; the at least two intermediate tensors include a first intermediate tensor and a second intermediate tensor, where the first intermediate tensor is determined based on a bottleneck tensor in the first encoder, and the second intermediate tensor is determined based on at least one network layer in the first decoder.
[0052] In an alternative implementation, the degradation network includes a first sub-network. The degraded image is obtained by convolving the high-quality image with a per-pixel kernel, and the per-pixel kernel is determined from a per-pixel kernel map. The per-pixel kernel map is generated by the first sub-network using the first intermediate tensor and the second intermediate tensor as inputs. Among them, the convolution weights of multiple network layers of the first sub-network are determined based on the first intermediate tensor.
[0053] In an alternative implementation, the height × width dimension of the first intermediate tensor is 1×1 or 2×2 or 3×3; or the height × width dimension of the first intermediate tensor is the same as the height × width dimension of the bottleneck tensor, and the number of channels of the first intermediate tensor is 1 / 8 or 1 / 4 or 1 / 2 of the number of channels of the bottleneck tensor.
[0054] In an alternative implementation, the processing result of the first intermediate tensor is also injected into at least one network layer in the first decoder. Among them, the spatial dimension of the processing result of the first intermediate tensor is 1×1.
[0055] In an alternative implementation, the first sub-network includes a pointwise convolution weight estimation sub-network and a kernel map estimation sub-network. The per-pixel kernel map is generated by the kernel map estimation sub-network using the local component tensor and multiple dynamic pointwise convolution weights as inputs. Among them, the convolution weights of multiple network layers of the kernel map estimation sub-network are the multiple dynamic pointwise convolution weights; the multiple dynamic pointwise convolution weights are generated by the pointwise convolution weight estimation sub-network using the first intermediate tensor as an input, and the local component tensor is determined based on the second intermediate tensor.
[0056] In an alternative implementation, the first sub-network further includes a kernel embedding map estimation sub-network, and the local component tensor is determined by the kernel embedding map estimation sub-network using the second intermediate tensor as an input.
[0057] In an alternative implementation, the enhancement network further includes a kernel embedding map estimation sub-network. The local component tensor is the second intermediate tensor, and the second intermediate tensor is generated by the kernel embedding map estimation sub-network using at least one network layer tensor among the network layer tensors in the first decoder as an input; the second intermediate tensor is concatenated with one network layer tensor among the network layer tensors in the first decoder.
[0058] In an alternative implementation, the height × width dimension of the local component tensor is the same as the height × width dimension of the high-quality image, and the number of channels of the local component tensor is 4 or 8 or 16 or 24 or 32.
[0059] In an alternative implementation, a skip connection is provided between the first encoder and the first decoder.
[0060] In an alternative implementation, the enhancement network adopts a U-net design.
[0061] According to a fifth aspect, an embodiment of the present application provides a computer-readable storage medium including instructions. When the instructions are run on a computer, the computer is caused to execute the methods in the above-mentioned first aspect or any alternative implementation of the first aspect and the methods in the second aspect or any alternative implementation of the second aspect.
[0062] According to a sixth aspect, an electronic device is provided. The electronic device includes a processor and a memory. The processor is connected to the memory. The memory is used to store instructions, and the processor is used to execute the instructions. When the processor executes the instructions stored in the memory, the processor is caused to execute the methods in the above-mentioned first aspect or any alternative implementation of the first aspect and the methods in the second aspect or any alternative implementation of the second aspect.
[0063] According to a seventh aspect, a chip system is provided. The chip system includes a memory and a processor. The memory is used to store a computer program, and the processor is used to call the computer program from the memory and run the computer program, so that the electronic device where the chip system is located executes the methods in the above-mentioned first aspect or any alternative implementation of the first aspect and the methods in the second aspect or any alternative implementation of the second aspect.
[0064] According to an eighth aspect, a computer program product is provided. When the computer program product runs on an electronic device, the electronic device is caused to execute the methods in the above-mentioned first aspect or any alternative implementation of the first aspect and the methods in the second aspect or any alternative implementation of the second aspect. Description of the Drawings
[0065] Figure 1 is a schematic diagram of the structure of the system architecture provided by an embodiment of the present application.
[0066] Figure 2 is a schematic diagram of a convolutional neural network.
[0067] Figure 3 is a schematic diagram of the hardware structure of the chip provided by an embodiment of the present application.
[0068] Figure 4 is a schematic flowchart of the neural network training method 400 provided by an embodiment of the present application.
[0069] Figure 5 It is a schematic flowchart of the image processing method 500 provided by an embodiment of the present application.
[0070] Figure 6 It is a schematic diagram of the neural network of the training process used in the method 400 provided by an embodiment of the present application.
[0071] Figure 7 It is related to Figure 6 The network of the inference process corresponding to the neural network of the training process in.
[0072] Figure 8 It is a schematic diagram of another neural network of the training process used in the method 400 provided by an embodiment of the present application.
[0073] Figure 9 It is related to Figure 8 The network of the inference process corresponding to the neural network of the training process in.
[0074] Figure 10 It is a schematic diagram of another neural network of the training process used in the method 400 provided by an embodiment of the present application.
[0075] Figure 11 It is related to Figure 10 The network of the inference process corresponding to the neural network of the training process in.
[0076] Figure 12 It is a schematic block diagram of the neural network training device 1200.
[0077] Figure 13 It is a schematic block diagram of the image processing device 1300.
[0078] Figure 14 It is a schematic diagram of the hardware structure of the neural network training device 1400.
[0079] Figure 15 It is a schematic diagram of the hardware structure of the image processing device 1500. Detailed implementation manners
[0080] The technical solutions in the present application will be described below with reference to the accompanying drawings.
[0081] Definition:
[0082] Unless otherwise specified or otherwise implied by the context, the following terms and phrases have the meanings provided below.
[0083] Embodiments of the present application relate to large-scale applications of convolutional neural networks. For ease of understanding, relevant terms and related concepts such as convolutional neural networks in embodiments of the present application will be introduced below.
[0084] (1) Neural network
[0085] A neural network may include neurons, and a neuron may be an operation unit with an input of x s and an intercept of 1. The output of the operation unit may be as shown in formula (1):
[0086]
[0087] Here, s = 1, 2,..., n, where n is a natural number greater than 1, Ws represents the weight of x s , b represents the bias of the neuron, and f represents the activation function of the neuron. Among them, the activation function is used to introduce non-linear characteristics into the neural network to convert the input signal in the neuron into an output signal. The output signal of the activation function can be used as the input of the next convolutional layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting many single neurons together. Specifically, the output of one neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area including several neurons.
[0088] (2) Deep neural network
[0089] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with multiple hidden layers. DNN is divided based on the positions of different layers. The neural network inside DNN can be classified into three types: input layer, hidden layer, and output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the middle layers are the hidden layers. The layers are fully connected. Specifically, any neuron in the i-th layer is definitely connected to any neuron in the (i + 1)-th layer.
[0090] Although DNN seems complex, the operations of each layer are actually not complex and are only represented by the following linear relationship expression: represents the input vector, represents the output vector, represents the bias vector, W represents the weight matrix (also known as the coefficient), and α() represents the activation function. In each layer, only this simple operation needs to be performed on the input vector to obtain the output vector Since DNN has more layers, the number of coefficients W and bias vectors is also large. These parameters are defined as follows in DNN. Taking the coefficient W as an example, assume that in a three-layer DNN, the linear coefficient from the fourth neuron in the second layer to the second neuron in the third layer is defined as The superscript 3 represents the layer number where the coefficient W is located, and the subscript corresponds to the index 2 of the third output layer and the index 4 of the second input layer.
[0091] In summary, the coefficient from the k-th neuron in the (L–1)-th layer to the j-th neuron in the L-th layer is defined as
[0092] It should be noted that the input layer has no parameter W. In a deep neural network, more hidden layers make the network more capable of describing complex situations in the real world. Theoretically, the more parameters the model has, the higher the complexity and the greater the "capacity". This indicates that the model can complete more complex learning tasks. Training a deep neural network is a process of learning the weight matrix, and the ultimate goal of training is to obtain the weight matrix of all layers of the trained deep neural network (the weight matrix composed of vectors W of multiple layers).
[0093] (3) Convolutional Neural Network
[0094] A convolutional neural network (CNN) is a deep convolutional neural network with a convolutional structure. A CNN includes a feature extractor composed of convolutional layers and subsampling layers. The feature extractor can be regarded as a filter. The convolution process can be considered as performing convolution on the input image or convolutional feature plane using a trainable filter to output a convolutional feature plane. The convolutional feature plane can also be called a feature map. A convolutional layer is a layer of neurons that performs convolution processing on the input signal in the convolutional neural network. In the convolutional layer of a convolutional neural network, a neuron can only be connected to some neurons in the adjacent layer. A convolutional layer usually includes several feature planes, and each feature plane can include some neurons arranged in a rectangle. Neurons in the same feature plane share weights, and the weight matrix corresponding to the shared weights here is the convolutional kernel. Weight sharing can be understood as the way of extracting image information being independent of position. An implicit principle here is that the statistical information of a part of the image is the same as that of other parts. This means that the image information learned in one part can also be used in another part. Therefore, the same learned image information can be used for all positions in the image. In the same convolutional layer, multiple convolutional kernels can be used to extract different image information. Generally, the more convolutional kernels there are, the richer the image information reflected by the convolution operation.
[0095] The convolutional kernel can be initialized in the form of a random-size matrix. During the process of training a convolutional neural network, the convolutional kernel can obtain appropriate weights through learning. In addition, a direct benefit brought by weight sharing is to reduce the connections between layers of the convolutional neural network and reduce the risk of overfitting.
[0096] (4) Residual Network
[0097] The Residual Network is a deep convolutional network proposed in 2015. Compared with traditional convolutional neural networks, the Residual Network is easier to optimize and can improve accuracy by increasing a considerable depth. The core of the Residual Network is to solve the side effects (degradation problems) brought about by increasing depth, and the network performance can be improved only by increasing the network depth. The Residual Network usually includes many sub-modules with the same structure.
[0098] (5) Pixel value
[0099] The pixel values of an image can be red-green-blue (RGB) color values, and the pixel values can be long integers representing colors. For example, the pixel value is 256×Red + 100×Green + 76×Blue, where Blue represents the blue component, Green represents the green component, and Red represents the red component. In each color component, the smaller the value, the lower the indicated brightness, and the larger the value, the higher the indicated brightness. For grayscale images, the pixel values can be grayscale values.
[0100] (6) Loss function
[0101] During the process of training a convolutional neural network, since it is expected that the output of the convolutional neural network is as close as possible to the actual expected predicted value, the current predicted value of the network can be compared with the actual expected target value, and then the weight vector of each layer of the convolutional neural network can be updated based on the difference between the predicted value and the target value (of course, there is usually an initialization process before the first update. Specifically, all layers of the convolutional neural network are pre-configured with parameters). For example, if the predicted value of the network is high, the weight vector is adjusted to reduce the predicted value, and the adjustment continues until the convolutional neural network can predict the actual expected target value or a value very close to the actual expected target value. Therefore, "how to obtain the difference between the predicted value and the target value through comparison" needs to be predefined. This is the loss function or the objective function. The loss function and the objective function are important equations used to measure the difference between the predicted value and the target value. Taking the loss function as an example. The higher the output value (loss) of the loss function, the greater the indicated difference. Therefore, the training of a convolutional neural network is a process of minimizing the loss as much as possible.
[0102] (7) Backpropagation algorithm
[0103] A convolutional neural network can correct the parameter values in the convolutional neural network during the training process by using the back propagation (BP) algorithm, so that the error loss between the predicted value output by the convolutional neural network and the actual expected target value becomes smaller and smaller. Specifically, the input signal is propagated forward until an error loss occurs at the output end, and the parameter in the initial convolutional neural network is updated by using the back-propagated error loss information to converge the error loss. The back propagation algorithm is a backward propagation movement led by the error loss, and its purpose is to obtain the optimal parameters of the convolutional neural network, such as the weight matrix, that is, the convolutional kernel of the convolutional layer.
[0104] As Figure 1 shown, an embodiment of the present application provides a system architecture 100. In Figure 1 it, the data acquisition device 160 is used to acquire training data. For the image processing method in the embodiment of the present application, the training data may include pairs of aligned low-quality images and high-quality images depicting the same scene.
[0105] After the data acquisition device 160 acquires the training data, it stores the training data in the database 130. The training device 120 obtains the target model / rule 101 by training based on the training data maintained in the database 130.
[0106] The following describes how the training device 120 obtains the target model / rule 101 based on the training data.
[0107] The target model / rule 101 can be used to implement the image processing method of the embodiment of the present application. Specifically, after performing relevant preprocessing on the image to be processed, the image to be processed obtained after the relevant preprocessing is input into the target model / rule 101 to obtain the processed high-quality result of the image. The target model / rule 101 in the embodiment of the present application can specifically be a neural network. It should be noted that in actual applications, the training data maintained in the database 130 may not necessarily be all acquired by the data acquisition device 160, and may also be received and obtained from other devices. It should also be noted that the training device 120 does not necessarily train the target model / rule 101 completely based on the training data maintained in the database 130, and can also obtain training data from the cloud or other places for model training. The above description should not be construed as a limitation on the embodiments of the present application.
[0108] The target model / rule 101 obtained by the training device 120 through training can be applied to different systems or devices, such as Figure 1The execution device 110 shown. The execution device 110 can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) / virtual reality (VR) terminal, or a vehicle-mounted terminal, or can also be a server, a cloud device, etc. In Figure 1 In it, the input / output (I / O) interface 112 is used for the execution device 110 to exchange data with external devices. A user can input data to the I / O interface 112 by using the client device 140. In the embodiments of the present application, the input data may include the image to be processed input by the client device.
[0109] The preprocessing module 113 and the preprocessing module 114 are used to perform preprocessing based on the input data (such as the image to be processed) received by the I / O interface 112. In this embodiment of the present application, there may be no preprocessing module 113, or there may be no preprocessing module 114 (or there may be only one preprocessing module), and the input data is directly processed by the computing module 111.
[0110] During the process of the execution device 110 performing preprocessing on the input data or the computing module 111 of the execution device 110 performing calculations and other related processing, the execution device 110 can call the data, code, etc. in the data storage system 150 for corresponding processing, or can also store the data, instructions, etc. obtained through the corresponding processing into the data storage system 150.
[0111] Finally, the I / O interface 112 returns the processing result (such as the high-quality result obtained from the image to be processed) to the client device 140 to provide the processing result to the user.
[0112] It should be noted that the training device 120 can generate corresponding target models / rules 101 based on different training data for different targets or tasks. The corresponding target models / rules 101 can be used to achieve the above targets or complete the above tasks to provide the required results for users.
[0113] In Figure 1In the case shown, the user can manually provide input data by operating the interface provided by the I / O interface 112. In another case, the client device 140 can automatically send input data to the I / O interface 112. If obtaining the user's authorization is required to request the client device 140 to automatically send input data, the user can set the corresponding permissions on the client device 140. The user can view the results output by the execution device 110 on the client device 140, and the specific presentation form can be specific ways such as display, sound, action, etc. The client device 140 can also be used as a data collection end to collect the input data input to the I / O interface 112 and the output results output from the I / O interface 112 shown in the figure as new sample data, and store the new sample data in the database 130. Of course, the client device 140 can also not perform the collection, but the I / O interface 112 directly stores the input data input to the I / O interface 112 and the output results output from the I / O interface 112 shown in the figure as new sample data in the database 130.
[0114] It should be noted that Figure 1 is only a schematic diagram of the system architecture provided by the embodiments of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in Figure 1 , the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 can also be set in the execution device 110.
[0115] As Figure 1 shown, the target model / rule 101 is obtained by training by the training device 120. In this embodiment of the present application, the target model / rule 101 is the enhanced network to be discussed later and is a part of the model / rule to be trained. During the training process, the model / rule to be trained includes a degradation network and an enhanced network. The enhanced network can be a neural network in the present application. Specifically, the neural network provided by the embodiments of the present application can be a CNN, a deep convolutional neural network (DCNN), a recurrent neural network (RNN), etc. The degradation network can be a neural network in the present application. Specifically, the neural network provided by the embodiments of the present application can be a CNN, a deep convolutional neural network (DCNN), etc.
[0116] As Figure 2As shown, a convolutional neural network (CNN) 200 may include an input layer 210, a convolutional / pooling layer 220 (where the pooling layer is optional), and a convolutional neural network layer 230.
[0117] Convolutional / pooling layer 220:
[0118] Convolutional layer:
[0119] As Figure 2 shown, the convolutional / pooling layer 220 may include layers 221 to 226. For example, in one implementation, layer 221 is a convolutional layer, layer 222 is a pooling layer, layer 223 is a convolutional layer, layer 224 is a pooling layer, layer 225 is a convolutional layer, and layer 226 is a pooling layer. In another implementation, layers 221 and 222 are convolutional layers, layer 223 is a pooling layer, layers 224 and 225 are convolutional layers, and layer 226 is a pooling layer. Specifically, the output of the convolutional layer can be used as the input of the subsequent pooling layer or as the input of other convolutional layers to continue the convolution operation.
[0120] Taking the convolutional layer 221 as an example, the internal working principle of a convolutional layer will be described below.
[0121] The convolutional layer 221 may include multiple convolutional operators. A convolutional operator is also referred to as a convolution kernel. In image processing, a convolutional operator is equivalent to a filter for extracting specific information from an input image matrix. Essentially, a convolutional operator can be a weight matrix. The weight matrix is usually predefined and depends on the stride value during the convolution operation on the image. The weight matrix usually processes pixels at a granularity level of one pixel or two pixels in the horizontal direction of the input image to extract specific features from the image. The size of the weight matrix needs to be related to the size of the image. It should be noted that the depth dimension of the weight matrix is the same as the depth dimension of the input image. During the convolution operation, the weight matrix extends over the entire depth of the input image. The depth dimension is the channel dimension, corresponding to the number of channels. Therefore, the convolution using a single weight matrix generates a convolution output with a single depth dimension. However, in most cases, instead of using a single weight matrix, multiple weight matrices of the same size (row × column), i.e., multiple matrices of the same type, are applied. The outputs of the weight matrices are stacked to form the depth dimension of the convolutional image. Here, the dimension can be understood as determined based on the above "multiple". Different weight matrices can be used to extract different features from the image. For example, one weight matrix is used to extract edge information of the image, another weight matrix is used to extract specific colors of the image, and another weight matrix is used to blur unnecessary noise in the image. The multiple weight matrices have the same size (row × column). The size of the feature maps extracted by using multiple weight matrices of the same size is also the same, and then the multiple extracted feature maps of the same size are combined to form the output of the convolution operation.
[0122] The weight values in these weight matrices need to be obtained through a large amount of training in actual applications. Each weight matrix formed by using the weight values obtained through training can be used to extract information from the input image so that the convolutional neural network 200 can make correct predictions.
[0123] When the convolutional neural network 200 has multiple convolutional layers, generally more general features are extracted in the initial convolutional layer (e.g., 221). General features can also be called low-level features, corresponding to high-resolution feature maps. As the depth of the convolutional neural network 200 increases, the features extracted in the subsequent convolutional layers (e.g., 226) become more complex, such as high-level semantic features, and correspond to low-resolution feature maps. Features with high-level semantics are more suitable for the problems to be solved.
[0124] Pooling layer:
[0125] The number of training parameters often needs to be reduced. Therefore, it is usually necessary to periodically introduce a pooling layer after the convolutional layer. For Figure 2Among the layers 221 to 226 in 220, a convolutional layer can be followed by a pooling layer, or multiple convolutional layers can be followed by one or more pooling layers. During image processing, the pooling layer is only used to reduce the spatial dimensions of the image. The pooling layer can include an average pooling operator and / or a max pooling operator, and can be used to sample the input image to obtain a smaller image, and can also be used to sample the feature map input to the convolutional layer to obtain a smaller feature map. The average pooling operator can be used to calculate the pixel values in a specific range of the image to generate an average value. The average value is used as the average pooling result. The max pooling operator can be used to select the pixel with the maximum value in a specific range as the max pooling result. In addition, similar to the fact that the size of the weight matrix of the convolutional layer needs to be related to the size of the image, the operators of the pooling layer also need to be related to the size of the image. The size of the processed image output from the pooling layer can be smaller than the size of the image input to the pooling layer. Each pixel in the image output from the pooling layer represents the average value or the maximum value of the corresponding sub-region of the image input to the pooling layer.
[0126] Convolutional neural network layer 230:
[0127] After being processed by the convolutional / pooling layer 220, the convolutional neural network 200 is not yet ready to output the required output information. As described above, at the convolutional / pooling layer 220, only features are extracted and the parameters generated by the input image are reduced. However, in order to generate the final output information, the convolutional neural network 200 needs to generate a processed image result through the convolutional neural network layer 230. Therefore, the convolutional neural network layer 230 can include multiple hidden layers (such as Figure 2 231, 232,..., 23n shown) and an output layer 240. The parameters included in the multiple hidden layers can be obtained through pre-training based on relevant training data of a specific task type. For example, the task type can include image semantic segmentation, image classification, and super-resolution image reconstruction. The hidden layer can perform a series of processes on the feature map output from the convolutional / pooling layer 220 to obtain an image segmentation result. The process of obtaining a processed image result based on the feature map output from the convolutional / pooling layer 220 will be described in detail later and will not be elaborated here.
[0128] In the convolutional neural network layer 230, the multiple hidden layers are followed by the output layer 240, which is the last layer of the entire convolutional neural network 200. The output layer 240 has a loss function similar to categorical cross-entropy, and the loss function is specifically used to calculate the prediction error. Once the forward propagation of the entire convolutional neural network 200 (propagating in the direction from 210 to 240, as Figure 2 shown) is completed, then the backpropagation (propagating in the direction from 240 to 210, as Figure 2as shown in the figure), to update each of the above layer weight values and biases to reduce the loss of the convolutional neural network 200, as well as the error between the result output by the convolutional neural network 200 through the output layer (i.e., the above image processing result) and the ideal result.
[0129] It should be noted that Figure 2 the convolutional neural network 200 shown is merely an exemplary convolutional neural network. In specific applications, the convolutional neural network can also exist in the form of other network models.
[0130] Figure 3 shows the hardware structure of the chip provided by the embodiment of the present application. The chip includes a neural network processing unit 30. The chip can be set in Figure 1 the execution device 110 shown to complete the computing work of the computing module 111. The chip can also be set in Figure 1 the training device 120 shown to complete the training work of the training device 120 and output the target model / rule 101. Figure 2 All algorithms of each layer in the convolutional neural network shown can be implemented on Figure 3 the chip shown.
[0131] The neural network processing unit (NPU) 30 is mounted on the main CPU as a coprocessor, and the main CPU assigns tasks to the NPU 30. The core part of the NPU is the arithmetic circuit 303, and the controller 304 controls the arithmetic circuit 303 to extract data from the memory (weight memory or input memory) and perform operations.
[0132] In some implementation manners, the arithmetic circuit 303 includes a plurality of processing engines (PEs). In some implementation manners, the arithmetic circuit 303 is a two-dimensional systolic array. The arithmetic circuit 303 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementation manners, the arithmetic circuit 303 is a general matrix processor.
[0133] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit 303 extracts the data corresponding to the matrix B from the weight memory 302 and caches the data in each PE of the arithmetic circuit 303. The arithmetic circuit 303 extracts the data of the matrix A from the input memory 301 to perform matrix operations with the matrix B, thereby obtaining partial results or final results of the matrix and storing the results in the accumulator 308.
[0134] The vector calculation unit 307 can perform further processing on the output of the arithmetic circuit 303, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. For example, the vector calculation unit 307 can be used for network calculations such as pooling, batch normalization, or local response normalization in non-convolutional / non-FC layers in a neural network.
[0135] In some implementations, the vector calculation unit 307 can store the processed output vector in the unified buffer 306. For example, the vector calculation unit 307 can apply a non-linear function to the output of the arithmetic circuit 303, such as to the vector of accumulated values, to generate activation values. In some implementations, the vector calculation unit 307 generates normalized values, combined values, or both. In some implementations, the processed output vector can be used as the activation input of the arithmetic circuit 303. For example, the processed output vector is used in subsequent layers in a neural network.
[0136] The unified memory 306 is used to store input data and output data.
[0137] For weight data, the direct memory access controller (DMAC) 305 moves the input data in the external memory to the input memory 301 and / or the unified memory 306, stores the weight data in the external memory in the weight memory 302, and stores the data in the unified memory 306 in the external memory.
[0138] The bus interface unit (BIU) 310 is used to implement the interaction between the main CPU, the DMAC, and the instruction fetch buffer 309 through the bus.
[0139] The instruction fetch buffer 309 connected to the controller 304 is used to store the instructions used by the controller 304.
[0140] The controller 304 is used to call the instructions cached in the instruction fetch buffer 309 to control the working process of the arithmetic accelerator.
[0141] Generally, the unified memory 306, the input memory 301, the weight memory 302, and the instruction fetch buffer 309 are all on-chip memories. The external memory is the memory outside the NPU, which can be a double data rate synchronous dynamic random access memory (DDR SDRAM), a high bandwidth memory (HBM), or other readable and writable memories.
[0142] Figure 2 The operations of each layer in the convolutional neural network shown can be executed by the operation circuit 303 or the vector calculation unit 307.
[0143] In the above text, in combination with Figures 1 to 3 the basic content of the neural network, related devices and models in the embodiments of the present application are described in detail. Next, in combination with Figure 4 the method of the embodiments of the present application will be described in detail.
[0144] Figure 4 FIG. is a schematic flowchart of a neural network training method 400 provided by an embodiment of the present application. The method can be executed by a device with strong operation capabilities such as a computer device, a server device, or an operating device. Figure 4 The method 400 shown includes steps S410 to S440. Each step will be described in detail separately below.
[0145] S410, obtain a plurality of training image pairs, where each pair in the plurality of training image pairs includes a high-quality image and a low-quality image.
[0146] S420, input the low-quality image into an enhancement network to generate an enhanced image; input the high-quality image and at least two intermediate tensors into a degradation network to generate a degraded image, where the at least two intermediate tensors are determined based on the enhancement network.
[0147] S430, determine a first loss function based on the high-quality image and the enhanced image; determine a second loss function based on the low-quality image and the degraded image.
[0148] S440, update the neural network based on the first loss function and the second loss function, where the neural network includes an enhancement network and a degradation network.
[0149] In the present application, when training a neural network, not only the loss of the enhancement process is calculated, but also the loss of the degradation process is calculated. The information about the degradation process extracted from the input low-quality image (due to the image degradation process) can be further used in the enhancement network to improve the quality of the output enhanced image. Learning image enhancement and image degradation simultaneously is beneficial to better perform image enhancement. This method can be applied to fields including but not limited to SISR tasks or image deblurring tasks, etc.
[0150] In step S410, the high-quality image and the low-quality image are images depicting the same scene. In some implementations, these images are all real images captured by a camera. In other implementations, these images are artificial training data sets. In other words, the high-quality image can be a real image captured by a camera, while the low-quality image is a degraded image obtained from the high-quality image using another neural network (e.g., in a non-linear manner). The high-quality image and the low-quality image can have the same size. Alternatively, they can have different sizes. The low-quality image can be a low-resolution image for an SISR task or a blurred image for a deblurring task.
[0151] The low-quality image can be degraded by applying a non-uniform, spatially-varying degradation kernel with different sources. The sources of degradation include (but are not limited to) the PSF of an optical system, a complex scene depth map, camera motion, and scene object motion. Therefore, using these pairs of training images to train a neural network, the trained neural network will be applicable to real-life images.
[0152] In addition, the high-quality image and the low-quality image can be tensors of the same size processed from the original high-quality image and the original low-quality image. For example, the original high-quality image and the original low-quality image are processed through several filters to generate several feature maps or image tensors.
[0153] Different camera parameters can be used to obtain the high-quality image and the low-quality image. Some pairs of training images can include moving objects, while some pairs of training images can be obtained using a moving camera. Some pairs of training images can be captured under low-light conditions. Some pairs of training images can be landscape images, including objects far from the camera. Some pairs of training images can be captured by a camera with a poor optical system, such as the front camera of a smartphone.
[0154] Compared with the images degraded by a neural network, using real low-quality images to train a neural network will be more applicable to real-life images, which involves degradation parameters including but not limited to low resolution, blur, etc.
[0155] In S420, for the training phase, the neural network includes an enhancement network and a degradation network. The enhancement network receives the low-quality image as input, and the output of the enhancement network is an enhanced image. In addition, the intermediate tensor determined by the enhancement network is used as the input of the degradation network. The degradation network receives the high-quality image and at least two intermediate tensors as input, and the output of the degradation network is a degraded image.
[0156] In this application, the enhanced network includes a first encoder and a first decoder; at least two intermediate tensors include a first intermediate tensor and a second intermediate tensor, where the first intermediate tensor is determined based on the bottleneck tensor in the first encoder, and the second intermediate tensor is determined based on at least one network layer in the network layer in the first decoder.
[0157] The enhanced network includes an encoder and a decoder. For example, the enhanced network can follow the UNet design. In addition, the enhanced network can have other structures, such as a residual in residual dense block (RRDB) or a residual channel attention network (RCAN), as long as it has an encoder and a decoder.
[0158] The bottleneck tensor is the output of the encoder of a UNet-like network. The spatial dimensions (height × width) of this tensor are the smallest among all the tensors involved in the UNet-like network.
[0159] Optionally, for the SISR task, the enhanced network can include additional image upscaling operations at the beginning or end of the image enhancement network.
[0160] The first intermediate tensor can be a part of the bottleneck tensor in the first encoder. The first intermediate tensor can have the same height × width dimensions as the bottleneck tensor, but only includes some of the channels in the bottleneck tensor. The proportion of the bottleneck tensor extracted as the first intermediate tensor can be 1 / 8. In addition, other dimensions can also be used, such as 1 / 2 or 1 / 4. That is, the height × width dimensions of the first intermediate tensor are the same as those of the bottleneck tensor, while the number of channels of the first intermediate tensor is 1 / 8 or 1 / 4 or 1 / 2 of the number of channels of the bottleneck tensor.
[0161] The first intermediate tensor can be the processed result of the bottleneck tensor in the first encoder. For example, the bottleneck tensor is processed by several convolutional kernels, and then a reduction method is used to generate a global component tensor. In addition, the global component tensor can be used as the input of the degradation network. In this implementation, the height × width dimensions of the global component tensor as the first intermediate tensor are 1×1 or 2×2 or 3×3. In this case, the enhanced network includes a global component tensor estimation branch (converting a part of the bottleneck tensor into a global component tensor)
[0162] In one example, 1 / 8 of the bottleneck tensor is extracted to produce a global component tensor. The global component tensor estimation branch consists of a series of convolutions, each of which doubles the number of channels of the tensor being processed while keeping its width and height the same. The number of such convolutions is selected such that the output of the last convolution has the same number of channels as the original bottleneck tensor (if 1 / 8 of the bottleneck tensor is extracted, the number is 3). Then a downscaling operation is applied to produce a global component tensor with a fixed spatial size S g ×S g , where the value of S g can be 2 or other values, such as 2 or 3. For the downscaling operation, any suitable downscaling method can be used, such as bilinear, bicubic, or area downscaling.
[0163] Then the global component tensor can be further processed and injected into the first decoder of the enhancement network. In other words, the method also includes: injecting the processing result of the first intermediate tensor into at least one network layer of the first decoder, where the spatial size of the processing result of the first intermediate tensor is 1×1. Similarly, the first intermediate tensor can be injected back into the bottleneck tensor.
[0164] In the above implementation, the global component tensor will be processed and injected into at least one network layer of the first decoder. The injection method can be but is not limited to concatenation or affine injection.
[0165] If the injection method is concatenation, it may include the following process: First, the global component tensor is compressed into a compressed global component tensor with a spatial size of 1×1 through a global average pooling operation (the pooling operation does not involve a convolution kernel); then, for each network layer of the decoder, the compressed global component tensor is processed through at least two fully connected layers to generate a tensor with a spatial size still of 1×1 and the number of channels equal to the number of channels of the decoder tensor of that network layer; finally, the tensor is replicated along the spatial dimension to match the spatial size of the decoder tensor, so that the replicated tensor and the decoder tensor can ultimately be concatenated along the channel axis.
[0166] If it is affine injection, first the global component tensor is still compressed into a spatial size of 1×1 through a pooling operation, and then the compressed global component tensor is processed through multiple fully connected layers to generate a scaling tensor and a bias tensor, and the affine transformation of the decoder tensor is performed by multiplying the decoder tensor by the scaling tensor and adding the bias tensor.
[0167] The injection of the global component tensor (i.e., the global information decoupling the degradation information) will improve the performance of the enhancement network during the inference phase without excessive computation. In addition, the per-pixel kernel map determined by the first intermediate tensor and the second intermediate tensor (discussed later) will not be injected into the enhancement network to avoid a huge computational cost.
[0168] It should be noted that in the case where a part of the bottleneck tensor in the extraction encoder is directly used as the first intermediate tensor input to the degradation network, a global component tensor is also obtained in the degradation network, but the global component tensor is not injected into the enhancement network. Therefore, the global component tensor estimation branch is included in the degradation network as part of the pointwise convolution weight estimation sub-network, rather than in the enhancement network. Pointwise convolution is a convolution with a kernel size of 1×1.
[0169] In the present application, the degradation network includes a first sub-network. Inputting a high-quality image and at least two intermediate tensors into the degradation network to generate a degraded image includes: inputting the first intermediate tensor and the second intermediate tensor into the first sub-network to generate a per-pixel kernel map, where the convolution weights of multiple network layers of the first sub-network are determined based on the first intermediate tensor; inputting the high-quality image into the degradation network; convolving the high-quality image with the per-pixel kernel to generate a degraded image, where the per-pixel kernel is determined from the per-pixel kernel map.
[0170] The first sub-network is used to generate a per-pixel kernel map, and then process it into a per-pixel kernel to convolve with the input high-quality image to generate a degraded image. Similarly, in the first sub-network, the first intermediate tensor determines the weights in the first sub-network to convert the second intermediate tensor into a per-pixel kernel map. In other words, the first sub-network takes the second intermediate tensor (the second intermediate tensor is a local component tensor, or the second intermediate tensor can be processed to obtain a local component tensor) as input, and its parameters (convolution weights) are estimated from the first intermediate tensor.
[0171] This application does not assume a single degradation kernel for each image, but estimates a per-pixel kernel map for the degradation process. Most actual image degradation scenarios (such as the PSF of an optical system and motion blur) imply non-uniform spatially-varying degradation kernels. Therefore, by estimating the per-pixel kernel map, the degradation network of this method can better explore the underlying degradation process, thus providing more information clues for the enhancement network.
[0172] In this application, the first sub-network includes a pointwise convolution weight estimation sub-network and a kernel map estimation sub-network. Inputting the first intermediate tensor and the second intermediate tensor into the first sub-network to generate a per-pixel kernel map includes: inputting the first intermediate tensor into the pointwise convolution weight estimation sub-network to generate multiple dynamic pointwise convolution weights; inputting the second intermediate tensor into the first sub-network and determining a local component tensor based on the second intermediate tensor; inputting the local component tensor and the multiple dynamic pointwise convolution weights into the kernel map estimation sub-network to generate a per-pixel kernel map, where the convolution weights of multiple network layers of the kernel map estimation sub-network are the multiple dynamic pointwise convolution weights.
[0173] The first intermediate tensor determined by the bottleneck tensor can indicate the global information of the degradation information. The second intermediate tensor determines the local component tensor, which indicates the local information about the degradation process. Instead of directly estimating the per-pixel kernel map, this technical solution first decouples the degradation into a tensor about global information and a tensor about local information. The tensor about global information is high-dimensional (e.g., the first intermediate tensor having the same dimension as the bottleneck tensor), but is general for the entire image, and the tensor about local information is low-dimensional (e.g., the local component tensor with dimensions of 8, 4,...), but is specific for each pixel. Then, the method combines the global component and the local component to estimate the per-pixel kernel map. This decoupling greatly reduces the internal dimension of the spatially varying degradation estimation task, thereby greatly reducing its ill-posedness and making the per-pixel kernel map more consistent.
[0174] That is to say, the first sub-network includes two sub-networks: (1) a pointwise convolution weight estimation sub-network that takes the first intermediate tensor as input, and the output is a plurality of dynamic pointwise convolution weights; (2) a kernel map estimation sub-network that takes the local component tensor as input, and the output is the per-pixel kernel map.
[0175] The local component tensor is a tensor with the same height×width size as that of the high-quality image, and the number of channels of the local component tensor is 4 or 8 or 16 or 24 or 32.
[0176] The local component tensor has the same size as the high-quality image. And after passing through several pointwise convolution layers, it still maintains the same size and can be divided into m (height×width) n-dimensional vectors, where each of the m vectors corresponds to each pixel in the high-quality image.
[0177] In some implementations, the local component tensor is the second intermediate tensor. The second intermediate tensor is determined based on at least one network layer in the first decoder. Specifically, the processing result of at least one network layer in the first decoder is used as the second intermediate tensor and input into the first sub-network (kernel map estimation sub-network).
[0178] In this case, the enhancement network further includes a kernel embedding map estimation sub-network, the local component tensor is the second intermediate tensor, and the method further includes: inputting at least one of the network layer tensors in the first decoder into the kernel embedding map estimation sub-network to generate the second intermediate tensor; cascading the second intermediate tensor with one of the network layers in the first decoder.
[0179] The second intermediate tensor (local information about the degradation process or local component tensor) is injected into the enhancement network to improve its performance. Since the local component tensor is low-dimensional, the injection does not incur too much computational cost.
[0180] The nuclear embedding graph estimation sub-network is constructed to be similar to the first decoder and includes a second decoder. Tensors of several network layers in the first decoder are concatenated with tensors of several network layers of the nuclear embedding graph estimation sub-network.
[0181] In other implementations, the tensor of at least one network layer in the first decoder is directly used as a second intermediate tensor and input into the first sub-network.
[0182] In this case, the first sub-network further includes a nuclear embedding graph estimation sub-network, which means that the first network includes a total of three sub-networks: a pointwise convolution weight estimation sub-network, a nuclear graph estimation sub-network, and a nuclear embedding graph estimation sub-network. Then, the second intermediate tensor is input into the first sub-network and the local component tensor is determined based on the second intermediate tensor, including: inputting the second intermediate tensor into the nuclear embedding graph estimation sub-network to generate the local component tensor.
[0183] In addition, the nuclear embedding graph estimation sub-network includes a second decoder. Tensors of several network layers in the first decoder are concatenated with tensors of several network layers of the nuclear embedding graph estimation sub-network. The output of the nuclear embedding graph estimation sub-network is also a local component tensor, but it is directly used as the input of the nuclear graph estimation sub-network and will not be injected into the enhancement network again.
[0184] It should be noted that the difference between these two cases (the nuclear embedding graph estimation sub-network belongs to the enhancement network and the nuclear embedding graph estimation sub-network belongs to the degradation network) is whether the output (local component tensor) of the nuclear embedding graph estimation sub-network is injected back into the enhancement network through concatenation. In addition, in the case where the local component tensor is injected back into the enhancement network, the nuclear embedding graph estimation sub-network can be simplified to reduce the computational cost.
[0185] Through the above introduction, the first intermediate tensor, the second intermediate tensor, and the main structures of the enhancement network and the degradation network are explained.
[0186] In S430, for the enhancement network, a first loss function is determined, and for the degradation network, a second loss function is determined.
[0187] During the training phase, the first loss function is composed of the enhanced image and the high-quality image. The first loss function can be the mean absolute error (MAE) loss or L 1, structural similarity index measure (SSIM), mean squared error loss (MSE loss), perceptual loss, adversarial loss, or any other loss used for SISR and image deblurring tasks. The first loss function quantifies the difference between the enhanced image and the high-quality image, and minimizes the difference through the training process so that the image enhancement network outputs an enhanced image close to the high-quality image, for reference.
[0188] During the training phase, the degraded image and the low-quality image are used to form a degradation loss, denoted as the second loss function, which can be a combination of L 1 , SSIM, MSE, perceptual loss, adversarial loss, or any other loss used for SISR and image deblurring tasks. Experiments have found that the L 1 difference is sufficient for the degradation loss. The degradation loss quantifies the difference between the degraded image and the low-quality image, and minimizes the difference through the training process so that the image degradation network outputs a degraded image close to the low-quality image, thereby indirectly promoting the extraction of indicative global component tensors and local component tensors, which include important features of the underlying degradation process specific to the input low-quality image.
[0189] Taking the first loss function as an example, for each training image pair, the MSE loss is the square of the difference between the pixel values of the enhanced image and the pixel values of the high-quality image. For multiple training image pairs, the MSE loss can be the average of the MSE losses of multiple training images.
[0190] In S440, the neural network is updated using the first loss function and the second loss function.
[0191] Next, the structures of the enhancement network and the degradation network will be described in detail to explain the update process of S440.
[0192] As mentioned above, the main body of the enhancement network (the first encoder and the first decoder) and the kernel embedding graph estimation sub-network follow the same network design, both including an encoder and a decoder. In addition, for the enhancement network, skip connections are provided between the first encoder and the first decoder, which will enable better training results and avoid the problems of gradient vanishing or gradient explosion.
[0193] The first encoder, the first decoder, and the kernel embedding graph estimation sub-network all include multiple convolutional layers. For each convolutional layer, the convolutional kernel needs to be trained. In addition, the weights involved in the convolutional layers of the global component tensor estimation branch and the fully connected layers of the global component tensor injection process need to be trained.
[0194] The per-point convolution weight estimation sub-network takes the first intermediate tensor as input and outputs multiple per-point convolution weights. In the case where the first intermediate tensor is part of the bottleneck tensor in the first encoder, the first intermediate tensor is first processed through several convolutional layers and a downsampling layer to become a global component tensor, and then the global component tensor is processed into multiple dynamic per-point convolution weights. In the case where the first intermediate tensor is the global component tensor, the global component tensor is directly processed into multiple dynamic per-point convolution weights. Since the details of processing the first intermediate tensor (part of the bottleneck tensor in the first encoder) into the global component tensor (global component tensor estimation branch) are described above (the weights involved in the convolutional layers need to be trained), here, how to transform the global component tensor into multiple dynamic per-point convolution weights is introduced.
[0195] The global component tensor is first flattened, for example, reshaping the height×width×channels from S g ×S g ×C g to 1×1×C g S g 2 . Subsequently, the flattened global component tensor is used to estimate multiple dynamic per-point convolution weights using at least two fully connected layers. The multiple dynamic per-point convolution weights are then used in several network layers of the kernel map estimation sub-network, so the number of multiple dynamic per-point convolution weights is determined by the number of several layers using the dynamic per-point convolution weights. In this application, the number can be 5. Additionally, other numbers can also be used, such as 3, 4, 6, 7, 8. During the training process of the network, the weights of at least two fully connected layers need to be trained.
[0196] The kernel map estimation sub-network takes the local component tensor as input and outputs the per-pixel kernel map. The kernel map estimation sub-network is constructed as a multi-layer perceptron (MLP), taking the C e -dimensional local component tensor (H×W, the same as the height×width of the high-quality image) as input and generating a K 2 -dimensional vector. The process of converting the per-pixel kernel map to the per-pixel kernel is as follows: Reshape the K 2 -dimensional vector of the per-pixel kernel map into a K×K degenerate kernel impulse response and normalize it to sum to one to obtain the H×W per-pixel kernels for the H×W pixels in the high-quality image, where K is the degenerate kernel width / height and can be an odd value, such as 25, but other values such as 21, 23, 27, 29 can also be used. In fact, the MLP is implemented as a sequence of per-point convolutions operating on a tensor with the same spatial dimensions as the high-quality image and having the number of channels C e . The final normalization is done by dividing each K×K kernel by the sum of its elements.
[0197] To adapt the conversion of the local component tensor to the per-pixel kernel map to the input image content, only a few of the first pointwise convolutions in the kernel map estimation sub-network are static convolutions, and the rest are dynamic convolutions, whose weights are estimated by the pointwise convolution weight estimation sub-network. Static convolution is a traditional convolution whose weights are independent of the input data. These weights are updated according to the regular training process during the training phase but are fixed during the inference phase. Dynamic convolution is a convolution whose weights (dynamic pointwise convolution weights in this application) depend on the input data, and these weights are usually estimated from the input data through a dedicated part of the network.
[0198] The parameter count of the pointwise convolution in the kernel map estimation sub-network is determined by the number of input and output channels of the network layer. For example, for a convolutional layer, both the number of input channels and the number of output channels are C e , and the parameter count is C e ×C e . The numbers of static convolution and dynamic convolution can be 2 and 5 respectively, but other values can also be used. For example, the number of static convolutions can be 1, 3, 4, 5, 6, etc., and the number of dynamic convolutions can be 2, 3, 4, 6, 7, etc. During the training process, the static convolution weights are updated multiple times. Only the weights of the static convolution need to be trained.
[0199] This technical solution derives the degradation kernel for each pixel (per-pixel kernel) as the output of a very general non-linear transformation operating on the kernel latent space vector in a configurable dimension, rather than restricting the per-pixel kernel (determined by the per-pixel kernel map) to belong to some simple low-parameter kernel families, such as anisotropic Gaussian kernels. This method uses an MLP as the non-linear transformation from the kernel latent space vector to the kernel impulse response. In this way, this method can estimate the per-pixel kernel map including the kernel from the rich kernel space provided by the MLP, and the MLP is known for its universal approximation ability.
[0200] In addition, instead of using a static MLP (whose weights are the same for all input images), this technical solution uses an MLP with dynamic layers, whose weights are estimated by another MLP from the decoupled global component tensor of the degradation. This way can make the method better adapt to each input image, thus better understanding its underlying degradation.
[0201] Then, the high-quality image is convolved with the per-pixel kernel determined according to the per-pixel kernel map to generate the degraded image.
[0202] It should be noted that non-linear activation is usually performed after the convolutional layer and the fully connected layer. For example, the rectified linear unit (ReLU) or the leaky ReLU. Exceptions include the last convolution, which produces the output enhanced image, the per-pixel kernel map, the local component tensor, and the fully connected layer that produces the pointwise convolution weights.
[0203] In summary, the weights to be trained include: the weights in the first encoder and the first decoder, the weights of the fully connected layers in the global component tensor estimation branch and the pointwise convolution weight estimation sub-network, the weights of the kernel embedding graph estimation sub-network, and the weights of the static convolution in the kernel graph estimation sub-network.
[0204] Before training the neural network, initialize the above weights first. Then, after the first training, calculate the first loss function and the second loss function. If the total loss function does not meet the standard, update the weights for the neural network and start the second training, and so on. The training ends until the total loss function reaches the standard.
[0205] The total loss used in the training process to train the entire network is the sum of the enhancement loss and the degradation loss. Therefore, the training process minimizes both of these losses simultaneously. The standard can include that the sum of the first loss function and the second loss function is less than a preset threshold.
[0206] After the training ends, only the enhancement network is used in the inference phase. Therefore, a properly trained image enhancement network should be integrated into the imaging device firmware, for example, the mobile phone camera software. The image degradation network is only used during the training phase, and its purpose is to help the encoder of the image enhancement network learn to efficiently estimate the informative global component tensor and local component tensor of the degradation process.
[0207] It should be noted that the present technical solution can be applied to various CNN-based image enhancement backbone networks to improve their quality. The only requirement for the applicability of this application is that the backbone architecture should be divisible into "encoder" and "decoder" parts. For example, this application can be used to improve the output image quality of many CNN-based lightweight image enhancement networks without significantly reducing their computational performance, thus providing high flexibility for target hardware and software requirements.
[0208] It should also be noted that the process of convolving the per-pixel kernel with the high-quality image to generate the degraded image can be replaced by, for example, some non-linear image filtering.
[0209] It should also be noted that the kernel graph estimation sub-network using the MLP can be replaced by some other transformations. Although the solution implemented in this application well balances the spatial richness of modeling the degradation kernel and the complexity of estimating the transformation parameters.
[0210] It should also be noted that the pointwise convolution weight estimation sub-network can be modified.
[0211] Figure 5 It is a schematic flowchart of the image processing method 500 provided by an embodiment of the present application. This method can be executed by devices such as handheld devices (such as smartphones). Figure 5The method 500 shown includes steps S510 to S520. Each step is described in detail below.
[0212] S510, obtain an image to be processed;
[0213] S520, process the image to be processed using an enhancement network, where the enhancement network is obtained by training a neural network including an enhancement network and a degradation network based on multiple pairs of training images, and each pair of the multiple pairs of training images includes a high-quality image and a low-quality image. The neural network is trained based on a first loss function and a second loss function, where the first loss function is determined based on the high-quality image and the enhanced image, and the second loss function is determined based on the low-quality image and the degraded image. The enhanced image is generated by the enhancement network using the low-quality image as an input, and the degraded image is generated by the degradation network using the high-quality image and at least two intermediate tensors as inputs, where the at least two intermediate tensors are determined based on the enhancement network.
[0214] In S510, the image to be processed is degraded by applying a non-uniform, spatially-varying degradation kernel with different sources. The sources of degradation include (but are not limited to) the PSF of an optical system, a complex scene depth map, camera motion, and scene object motion.
[0215] For example, the image to be processed is a selfie photo taken by the front camera of a smartphone. Generally, the optical system of the front camera is poor, resulting in the degradation of the details of the photo image taken by this camera.
[0216] For another example, the image to be processed is a landscape photo taken by the main camera of a smartphone. Landscape photos usually include objects far from the camera. Due to the blurring of the PSF of the main camera, these objects lack fine-grained details.
[0217] For another example, the image to be processed is a photo taken under low-light conditions. Generally, under such conditions, the photo is taken using a long exposure, resulting in motion blur due to camera motion or moving scene objects, thus causing the image to be blurred.
[0218] In S520, the image to be processed is processed using the enhancement network. The enhancement network used in method 500 is extracted from the neural network trained by the method provided in method 400.
[0219] Details of method 500 can be referred to the description in method 400. To avoid repetition, it will not be elaborated here.
[0220] Figure 6 is an exemplary neural network and can be the neural network used for training in method 400. Next, Figure 6The neural network provided in [reference] is further used to illustrate the process of method 400.
[0221] As Figure 6 shown, the neural network includes an enhancement network and a degradation network. The enhancement network receives the low-quality image 10 as input to generate an enhanced image 19. The degradation network receives the high-quality image 88 as input. The enhancement network determines at least two intermediate tensors: tensor 22, tensor 14, tensor 16, and tensor 18 to generate a degraded image 87, where tensor 22 is the first intermediate tensor, and the second intermediate tensor includes tensor 14, tensor 16, and tensor 18.
[0222] In the enhancement network, the first encoder includes tensor 10, tensor 11, tensor 12, and tensor 13, where tensor 131 is the bottleneck tensor in the first encoder. The first decoder includes tensor 14, tensor 15, tensor 16, tensor 17, tensor 18, and tensor 19. The first decoder uses skip connections from the corresponding layers in the first encoder to combine the extracted features. The skip connections (S29, S30) of tensor 11 and tensor 12 propagate low-level image details to higher levels (tensor 15 and tensor 17) to improve the local component tensor estimation during training.
[0223] In the first encoder, the input low-quality image 10 is processed through several convolutional operations to become the bottleneck tensor 131. Figure 6 Three convolutional operations are shown: S10, S11, and S12.
[0224] The enhancement network also includes a global component tensor estimation branch, which takes a part of the bottleneck tensor 131 (such as 1 / 8 or 1 / 4 or 1 / 2 of the bottleneck tensor 131) as input and outputs a global component tensor 22, which is used as the first intermediate tensor input of the degradation network.
[0225] The global component tensor estimation branch includes a series of convolutional operations (S19 and S20), each of which doubles the number of channels of the tensor being processed while keeping its height and width dimensions the same, so that the output of the last convolutional operation (tensor 21) has the same number of channels as the bottleneck tensor 131, which means that for obtaining 1 / 8, 1 / 4, and 1 / 2 parts of the bottleneck tensor, the number of convolutional operations is 3, 2, and 1 respectively. For simplicity, Figure 6 only two convolutional operations are depicted. Then a downsampling operation (S21) is applied to tensor 21 to produce a global component tensor 22 with a fixed spatial size S g ×S g where S g is the height / width of the global component tensor. The number can be 2, 1, or 3.
[0226] The global component tensor 22 is further injected into the first decoder through concatenation. In Figure 6 , the global component is concatenated with tensors 151 and 152 to form tensor 15, and concatenated with tensors 171 and 172 to form tensor 17. At the same time, the global component tensor is concatenated with the bottleneck tensor 131 to form tensor 13. Here, the global component tensor needs to be processed for concatenation with the corresponding tensors.
[0227] In the example of processing the global component tensor for concatenation with the bottleneck tensor 131, first, the global component tensor 22 is compressed through the global average pooling operation S22 to generate the compressed global component tensor 23, and then at least two fully connected operations S23 and S24 are performed to generate a tensor with a spatial size of 1×1 and the same number of channels as the bottleneck tensor 131. The tensor is replicated along the spatial dimension to match the size of the bottleneck tensor 131. The bottleneck tensor 131 and the injection tensor 132 are concatenated to form tensor 13, which serves as the input to the convolution operation S13.
[0228] In the first decoder, tensor 13 is also processed through several convolution operations (S13, S14, S15, S16, S17, and S18) to generate the enhanced image 19. It should be noted that for operation S15, the input tensor includes three parts: the output result of S14 (tensor 151), the features from tensor 12 through the skip connection (tensor 152), and the injected global component tensor 153. For the convolution operation S17, the input tensor is similar to S15, and for the sake of simplicity, it will not be elaborated here.
[0229] It should be noted that after convolution operations and fully connected operations, there is usually a non-linear activation (e.g., ReLU or leaky ReLU), which is not described in Figure 6 for the sake of simplicity.
[0230] The degradation network includes a first sub-network that takes the second intermediate tensors (tensors 14, 16, and 18) and the first intermediate tensor 22 as inputs and then outputs the per-pixel kernel map 85, where the first intermediate tensor 22 is used to determine the weights of some convolution operations (S82 and S83) in the first sub-network. The first sub-network includes layer tensors: tensors 41 to 45, tensors 80 to 88, as Figure 6 shown. The first sub-network also includes three sub-networks: the pointwise convolution weight estimation sub-network, the kernel map estimation sub-network, and the kernel embedding map estimation sub-network.
[0231] The kernel embedding graph estimation sub-network takes the second intermediate tensor as input and outputs the local component tensor 80 through a number of convolutional operations including S41 to S45. The kernel embedding graph estimation sub-network follows a structure similar to that of the enhancement network. For the convolutional operation S42, the input includes the output result of the convolutional operation S41 (tensor 421) and a copy of the second intermediate tensor 16 (tensor 422). The situation is similar for the convolutional operation S44, and for simplicity, it will not be elaborated here.
[0232] The kernel graph estimation sub-network takes the local component tensor 80 as input and outputs the per-pixel kernel graph 85. The local component tensor 80 is first processed through a number of static pointwise convolutional operations with a kernel size of 1×1 (S80 and S81), and then processed through a number of dynamic pointwise convolutional operations (S82 and S83), whose weights are determined by the pointwise convolutional weight estimation sub-network (discussed later). Tensors 80 to 84 can have the same height×width size as the high-quality image 88, and the number of channels of tensors 80 to 84 is the same, for example, the number of channels is 8 (other numbers of channels such as 4, 16, 24, 32,... can also be used). The output of the convolutional operation S83 is tensor 85, with the same height×width size and the number of channels being K 2 , where K can be an odd number, such as 25, 21, 23, 19, 17, etc. By reshaping tensor 85 into a number of K×K per-pixel kernels (the number of per-pixel kernels is height×width, corresponding to the number of pixels in the high-quality image 88) and normalizing to sum to one, tensor 85 is converted into a per-pixel kernel (tensor 86) to generate the per-pixel kernel 86.
[0233] In the kernel graph estimation sub-network, the number of static pointwise convolutional operations (such as S80, S81) can be 2 (other numbers such as 1, 3, 4, 5, etc. can also be used); the number of dynamic pointwise convolutional operations (such as S82, S83) can be 5 (other numbers such as 3, 4, 5, 6, 7, 8, etc. can also be used).
[0234] The weights of the dynamic pointwise convolutional operations are estimated by the pointwise convolutional weight estimation sub-network, which takes the global component tensor 22 as input and outputs the dynamic pointwise convolutional weights 61 and 62. First, the global component tensor 22 is reshaped from S g ×S g ×C g into 1×1×C g S g 2 to form tensor 60. Then, for each dynamic pointwise convolutional weight, tensor 60 is processed through at least 2 fully connected operations to form the corresponding dynamic pointwise convolutional weight. For the operation S82, a dynamic pointwise convolutional weight of size C e ×C e is required. For the operation S83, a dynamic pointwise convolutional weight of size K is required.2 The dynamic per-point convolution weights of size ×Ce, where C e is the number of channels of tensors 82 and 84.
[0235] Then, the per-pixel kernel 86 is convolved with the high-quality image 88 to generate the degraded image 87. After generating the enhanced image 19 and the degraded image 87, a first loss function that describes the difference between the enhanced image 19 and the high-quality image 88 is calculated, and a second loss function that describes the difference between the degraded image 87 and the low-quality image 10 is calculated. By updating the weights involved in the neural network, the total loss function (the sum of the first loss function and the second loss function) is minimized.
[0236] Once the total loss function converges, the enhanced network can be extracted to implement the image enhancement operation. Figure 7 is the network for the inference process and can be used to implement the image processing method 500. Figure 7 The enhanced network provided in Figure 6 corresponds to the neural network in the training process in
[0237] Figure 8 is another example of a neural network and can be the neural network used for training in method 400.
[0238] As Figure 8 shown, Figure 8 the neural network depicted in Figure 6 differs from the neural network in Figure 6 in that there is no global component tensor injection in the image enhancement network. Accordingly, the global component estimation branch is excluded from the image enhancement network and moved to the image degradation network. In all other respects, it is the same as that shown in
[0239] Figure 9 shows the image enhancement network used in the Figure 8 corresponding inference stage. Since this image enhancement network lacks global component tensor estimation and injection, it is the same as the traditional backbone network widely used for SISR or image deblurring tasks. However, since the training stage includes the image degradation network, the image enhancement network is trained to produce enhanced images of better quality compared to the traditional method of training this network without the aid of the image degradation network.
[0240] Figure 10 is another example of a neural network and can be the neural network used for training in method 400.
[0241] As Figure 10 shown, Figure 10 the neural network depicted in Figure 6The neural network herein is different in that it greatly simplifies the kernel embedding graph estimation sub-network, which only takes the last tensor in the decoder of the image enhancement network as input and processes it with two convolutions to generate the kernel embedding graph.
[0242] Another difference is that the obtained kernel embedding graph is injected back into the decoder of the image enhancement network through concatenation.
[0243] Figure 11 It shows in Figure 10 the image enhancement network used in the corresponding inference stage. In this network, both the global component and the local component of the degradation information are injected into the decoder.
[0244] It should be noted that the neural network training method provided in this application can also be used in other aspects such as image dehazing, image contrast / brightness / color enhancement, and image coloring.
[0245] For the image dehazing application, the image degradation network should be replaced by a certain image blurring network that generates a plausible blurred image from a clean image. This requires inventing an image blurring network that takes local component data as input and whose parameters (such as convolution weights) are estimated from the global component data. And the training images should be replaced by blurred image and dehazed image pairs. In addition, the global component tensor is determined by the encoder of the enhancement network, and the local component tensor is determined by the decoder of the enhancement network.
[0246] Due to the injection of the local component and the global component, Figure 11 the enhancement network described in Figure 9 provides better visual quality. The enhancement network in Figure 7 has a smaller computational cost because no global component tensor and local component tensor are injected into the first decoder of the enhancement network. Due to the injection of the global component tensor into the first decoder of the enhancement network, Figure 11 the enhancement network in Figure 9 has less computational cost than the enhancement network in
[0247] Figure 12 is a schematic block diagram of the neural network training device 1200 provided in the embodiment of this application. As Figure 12As shown, the neural network training device 1200 includes: an acquisition module 1210, configured to acquire a plurality of training image pairs, where each pair of the plurality of training image pairs includes a high-quality image and a low-quality image, and the high-quality image and the low-quality image are images of the same scene with different qualities; input the low-quality image into an enhancement network to generate an enhanced image; a training module 1220, configured to input the high-quality image and at least two intermediate tensors into a degradation network to generate a degraded image, where the at least two intermediate tensors are determined based on the enhancement network; determine a first loss function based on the high-quality image and the enhanced image; determine a second loss function based on the low-quality image and the degraded image; update the neural network based on the first loss function and the second loss function, where the neural network includes an enhancement network and a degradation network.
[0248] In an alternative implementation, the enhancement network includes a first encoder and a first decoder; the at least two intermediate tensors include a first intermediate tensor and a second intermediate tensor, where the first intermediate tensor is determined based on a bottleneck tensor in the first encoder, and the second intermediate tensor is determined based on at least one network layer in the network layer of the first decoder.
[0249] In an alternative implementation, the degradation network includes a first sub-network, and the training module 1220 is specifically configured to: input the first intermediate tensor and the second intermediate tensor into the first sub-network to generate a per-pixel kernel map, where the convolutional weights of multiple network layers of the first sub-network are determined based on the first intermediate tensor; input the high-quality image into the degradation network; convolve the high-quality image with the per-pixel kernel to generate the degraded image, where the per-pixel kernel is determined from the per-pixel kernel map.
[0250] In an alternative implementation, the height × width size of the first intermediate tensor is 1×1 or 2×2 or 3×3; or the height × width size of the first intermediate tensor is the same as the height × width size of the bottleneck tensor, and the number of channels of the first intermediate tensor is the number of channels of the bottleneck tensor.
[0251] In an alternative implementation, the training module 1220 is further configured to: inject the processing result of the first intermediate tensor into at least one network layer in the first decoder, where the spatial size of the processing result of the first intermediate tensor is 1×1.
[0252] In an alternative implementation, the first sub-network includes a pointwise convolution weight estimation sub-network and a kernel map estimation sub-network. The training module 1220 is specifically configured to: input the first intermediate tensor into the pointwise convolution weight estimation sub-network to generate a plurality of dynamic pointwise convolution weights; input the second intermediate tensor into the first sub-network and determine a local component tensor based on the second intermediate tensor; input the local component tensor and the plurality of dynamic pointwise convolution weights into the kernel map estimation sub-network to generate the per-pixel kernel map, where the convolution weights of multiple network layers of the kernel map estimation sub-network are the plurality of dynamic pointwise convolution weights.
[0253] In an alternative implementation, the first sub-network further includes a kernel embedding map estimation sub-network. The training module 1220 is specifically configured to: input the second intermediate tensor into the kernel embedding map estimation sub-network to generate the local component tensor.
[0254] In an alternative implementation, the enhancement network further includes a kernel embedding map estimation sub-network. The local component tensor is the second intermediate tensor. The training module 1220 is further configured to: input at least one tensor in the network layer tensors in the first decoder into the kernel embedding map estimation sub-network to generate the second intermediate tensor; concatenate the second intermediate tensor with one of the network layer tensors in the network layer tensors in the first decoder.
[0255] In an alternative implementation, the height × width size of the local component tensor is the same as the height × width size of the high-quality image, and the number of channels of the local component tensor is 4 or 8 or 16 or 24 or 32.
[0256] In an alternative implementation, a skip connection is provided between the first encoder and the first decoder.
[0257] In an alternative implementation, the enhancement network adopts a U-net design.
[0258] Figure 13 is a schematic block diagram of an image processing device 1300 provided by an embodiment of the present application. As Figure 13As shown, the image processing apparatus 1300 includes: an acquisition module 1310 for acquiring an image to be processed; a processing module 1320 for processing the image to be processed using an enhancement network, where the enhancement network is obtained by training a neural network including an enhancement network and a degradation network based on a plurality of training image pairs, and each pair of the plurality of training image pairs includes a high-quality image and a low-quality image. The neural network is trained based on a first loss function and a second loss function, where the first loss function is determined based on the high-quality image and the enhanced image, and the second loss function is determined based on the low-quality image and the degraded image. The enhanced image is generated by the enhancement network using the low-quality image as an input, and the degraded image is generated by the degradation network using the high-quality image and at least two intermediate tensors as inputs, where the at least two intermediate tensors are determined based on the enhancement network.
[0259] In an alternative implementation, the enhancement network includes a first encoder and a first decoder; the at least two intermediate tensors include a first intermediate tensor and a second intermediate tensor, where the first intermediate tensor is determined based on a bottleneck tensor in the first encoder, and the second intermediate tensor is determined based on at least one network layer in the network layer of the first decoder.
[0260] In an alternative implementation, the degradation network includes a first sub-network, and the degraded image is obtained by convolving the high-quality image with a per-pixel kernel, where the per-pixel kernel is determined from a per-pixel kernel map generated by the first sub-network using the first intermediate tensor and the second intermediate tensor as inputs, and the convolution weights of the multiple network layers of the first sub-network are determined based on the first intermediate tensor.
[0261] In an alternative implementation, the height × width size of the first intermediate tensor is 1×1 or 2×2 or 3×3; or the height × width size of the first intermediate tensor is the same as the height × width size of the bottleneck tensor, and the number of channels of the first intermediate tensor is 1 / 8 or 1 / 4 or 1 / 2 of the number of channels of the bottleneck tensor.
[0262] In an alternative implementation, the processing result of the first intermediate tensor is further injected into at least one network layer in the first decoder, where the spatial size of the processing result of the first intermediate tensor is 1×1.
[0263] In an alternative implementation, the first sub-network includes a pointwise convolution weight estimation sub-network and a kernel map estimation sub-network. The per-pixel kernel map is generated by the kernel map estimation sub-network using the local component tensor and a plurality of dynamic pointwise convolution weights as inputs, where the convolution weights of multiple network layers of the kernel map estimation sub-network are the plurality of dynamic pointwise convolution weights; the plurality of dynamic pointwise convolution weights are generated by the pointwise convolution weight estimation sub-network using the first intermediate tensor as an input, and the local component tensor is determined based on the second intermediate tensor.
[0264] In an alternative implementation, the first sub-network further includes a kernel embedding map estimation sub-network, and the local component tensor is determined by the kernel embedding map estimation sub-network using the second intermediate tensor as an input.
[0265] In an alternative implementation, the enhancement network further includes a kernel embedding map estimation sub-network, the local component tensor is the second intermediate tensor, and the second intermediate tensor is generated by the kernel embedding map estimation sub-network using at least one network layer tensor among the network layer tensors in the first decoder as an input; the second intermediate tensor is concatenated with one network layer tensor among the network layer tensors in the first decoder.
[0266] In an alternative implementation, the height × width dimension of the local component tensor is the same as the height × width dimension of the high-quality image, and the number of channels of the local component tensor is 4 or 8 or 16 or 24 or 32.
[0267] In an alternative implementation, a skip connection is provided between the first encoder and the first decoder.
[0268] In an alternative implementation, the enhancement network adopts a U-net design.
[0269] Figure 14 It is a schematic diagram of the hardware structure of the neural network training device provided in the embodiments of the present application. Figure 14 The illustrated neural network training device 1400 includes a memory 1410, a processor 1420, a communication interface 1430, and a bus 1440. The memory 1410, the processor 1420, and the communication interface 1430 are communicatively connected to each other through the bus 1440.
[0270] The memory 1410 may store a program. When the program stored in the memory 1410 is executed by the processor 1420, the processor 1420 is used to execute the steps of the neural network training method 400 in the embodiments of the present application.
[0271] The processor 1420 may use a general - purpose CPU, microprocessor, ASIC, GPU, or one or more integrated circuits, and is used to execute relevant programs to perform the neural - network training method 400 in the embodiments of the present application.
[0272] The processor 1420 may also be an integrated - circuit chip with signal - processing capabilities. In the implementation process, the steps of the neural - network training method in the embodiments of the present application can be completed by using the integrated logic circuit in the hardware of the processor 1420 or instructions in software form.
[0273] It should be understood that Figure 14 The shown neural - network training device 1400 is used to train a neural network. A part of the neural network (enhanced network) obtained through training can be used to perform the image - processing method 500 in the embodiments of the present application. Specifically, Figure 4 The neural network in the shown method 400 can be obtained by training a neural network using the device 1400.
[0274] Specifically, Figure 14 The shown device can obtain training image pairs and the neural network to be trained from the outside through the communication interface 1430, and then the processor trains the neural network based on the training image pairs.
[0275] Figure 15 is a schematic diagram of the hardware structure of the image - processing device 1500 provided by the embodiments of the present application. Figure 15 The shown image - processing device 1500 includes a memory 1510, a processor 1520, a communication interface 1530, and a bus 1540. The memory 1510, the processor 1520, and the communication interface 1530 are communicatively connected to each other through the bus 1540.
[0276] The memory 1510 can store programs. When the programs stored in the memory 1510 are executed by the processor 1520, the processor 1520 is used to execute the steps of the image - processing method 500 in the embodiments of the present application.
[0277] The processor 1520 may use a general - purpose CPU, microprocessor, ASIC, GPU, or one or more integrated circuits, and is used to execute relevant programs to perform the image - processing method 500 in the embodiments of the present application.
[0278] The processor 1520 may also be an integrated - circuit chip with signal - processing capabilities. In the implementation process, each step of the image - processing method 500 in the embodiments of the present application can be completed by using the integrated logic circuit in the hardware of the processor 1520 or instructions in software form.
[0279] Specifically, Figure 15The device shown can obtain the image to be processed from the outside through the communication interface 1530, and then the processor executes an image processing method based on the image to be processed to generate a processed image.
[0280] It should be noted that although devices 1400 and 1500 only show a memory, a processor, and a communication interface, in the specific implementation process, those skilled in the art should understand that devices 1400 and 1500 may also include other components necessary for normal operation. Additionally, based on specific needs, those skilled in the art should understand that devices 1400 and 1500 may also include hardware components for implementing other additional functions. Additionally, those skilled in the art should understand that devices 1400 and 1500 may only include the components required to implement the embodiments of the present application, and need to include Figure 14 and Figure 15 all the components shown in.
[0281] Those of ordinary skill in the art can realize that, combining the units and algorithm steps described in the exemplary descriptions of the embodiments disclosed herein, the embodiments of the present application can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether the function is executed by hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described function for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0282] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above systems, devices, and units can refer to the corresponding processes in the above method embodiments, and will not be elaborated herein.
[0283] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the unit division is only a logical function division, and there may be other divisions in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the mutually coupled or directly coupled or communication connections shown or discussed can be implemented through some interfaces. The indirect coupling or communication connection between devices or units can be implemented in electronic, mechanical, or other forms.
[0284] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units. They can be located in one place, or can also be distributed on multiple network units. Some or all of the units can be selected based on actual needs to achieve the purpose of the solution of this embodiment.
[0285] In addition, in each embodiment of the present application, each functional unit can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0286] When these functions are implemented in the form of software functional units and sold or used as independent products, these functions can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be implemented in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for instructing a computer device (which can be a personal computer, a server, a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The above storage medium includes: any medium that can store program code, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disc.
[0287] The above are only some specific implementation manners of the present application and are not used to limit the protection scope of the present application. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application.
Claims
1. A neural network training method, characterized in that, comprising: Obtaining a plurality of training image pairs, wherein each pair of the plurality of training image pairs includes a high-quality image and a low-quality image, and wherein the high-quality image and the low-quality image are images of the same scene with different qualities; Inputting the low-quality image into an enhancement network to generate an enhanced image; inputting the high-quality image and at least two intermediate tensors into a degradation network to generate a degraded image, wherein the at least two intermediate tensors are determined based on the enhancement network; Determining a first loss function based on the high-quality image and the enhanced image; determining a second loss function based on the low-quality image and the degraded image; Updating the neural network based on the first loss function and the second loss function, wherein the neural network includes the enhancement network and the degradation network.
2. The method according to claim 1, characterized in that, The enhancement network includes a first encoder and a first decoder; the at least two intermediate tensors include a first intermediate tensor and a second intermediate tensor, wherein the first intermediate tensor is determined based on the bottleneck tensor in the first encoder, and the second intermediate tensor is determined based on at least one network layer in the network layer of the first decoder.
3. The method according to claim 2, characterized in that, The degradation network includes a first sub-network, and the inputting the high-quality image and at least two intermediate tensors into the degradation network to generate a degraded image includes: Inputting the first intermediate tensor and the second intermediate tensor into the first sub-network to generate a per-pixel kernel map, wherein the convolution weights of the multiple network layers of the first sub-network are determined based on the first intermediate tensor; Inputting the high-quality image into the degradation network; Convolving the high-quality image with the per-pixel kernel to generate the degraded image, wherein the per-pixel kernel is determined from the per-pixel kernel map.
4. The method according to claim 2 or 3, characterized in that, The height × width dimension of the first intermediate tensor is 1×1 or 2×2 or 3×3; or the height × width dimension of the first intermediate tensor is the same as the height × width dimension of the bottleneck tensor, and the number of channels of the first intermediate tensor is 1 / 8 or 1 / 4 or 1 / 2 of the number of channels of the bottleneck tensor.
5. The method according to any one of claims 2 to 4, characterized in that, The method further includes: Injecting the processing result of the first intermediate tensor into at least one network layer in the first decoder, wherein the spatial dimension of the processing result of the first intermediate tensor is 1×1.
6. The method according to claim 3, characterized in that, The first sub-network includes a pointwise convolution weight estimation sub-network and a kernel map estimation sub-network, and the inputting the first intermediate tensor and the second intermediate tensor into the first sub-network to generate a per-pixel kernel map includes: Inputting the first intermediate tensor into the pointwise convolution weight estimation sub-network to generate a plurality of dynamic pointwise convolution weights; Input the second intermediate tensor into the first sub-network and determine a local component tensor based on the second intermediate tensor; Input the local component tensor and the multiple dynamic pointwise convolution weights into the kernel map estimation sub-network to generate the per-pixel kernel map, wherein the convolution weights of multiple network layers of the kernel map estimation sub-network are the multiple dynamic pointwise convolution weights.
7. The method according to claim 6, wherein, the first sub-network further includes a kernel embedding map estimation sub-network, and the inputting the second intermediate tensor into the first sub-network and determining a local component tensor based on the second intermediate tensor includes: Input the second intermediate tensor into the kernel embedding map estimation sub-network to generate the local component tensor.
8. The method according to claim 6, wherein, the enhancement network further includes a kernel embedding map estimation sub-network, the local component tensor is the second intermediate tensor, and the method further includes: Input at least one network layer tensor in the network layer tensors in the first decoder into the kernel embedding map estimation sub-network to generate the second intermediate tensor; Concatenate the second intermediate tensor with one network layer tensor in the network layer tensors in the first decoder.
9. The method according to any one of claims 6 to 8, wherein, the height×width size of the local component tensor is the same as the height×width size of the high-quality image, and the number of channels of the local component tensor is 4 or 8 or 16 or 24 or 32.
10. The method according to any one of claims 2 to 9, wherein, A skip connection is provided between the first encoder and the first decoder.
11. The method according to any one of claims 1 to 10, wherein, The enhancement network adopts a U-net design.
12. An image processing method, wherein, comprising: Processing the image to be processed using an enhancement network, wherein the enhancement network is obtained by training a neural network including the enhancement network and a degradation network based on multiple training image pairs, wherein each pair of the multiple training image pairs includes a high-quality image and a low-quality image; the neural network is trained based on a first loss function and a second loss function, wherein the first loss function is determined based on the high-quality image and the enhanced image, the second loss function is determined based on the low-quality image and the degraded image; the enhanced image is generated by the enhancement network using the low-quality image as an input, and the degraded image is generated by the degradation network using the high-quality image and at least two intermediate tensors as inputs, wherein the at least two intermediate tensors are determined based on the enhancement network.
13. The method according to claim 12, wherein, The enhanced network includes a first encoder and a first decoder; the at least two intermediate tensors include a first intermediate tensor and a second intermediate tensor, wherein the first intermediate tensor is determined based on the bottleneck tensor in the first encoder, and the second intermediate tensor is determined based on at least one network layer in the network layer of the first decoder.
14. The method according to claim 13, wherein, the degradation network includes a first sub-network, the degraded image is obtained by convolving the high-quality image with a per-pixel kernel, the per-pixel kernel is determined from a per-pixel kernel map, and the per-pixel kernel map is generated by the first sub-network using the first intermediate tensor and the second intermediate tensor as inputs, wherein the convolutional weights of multiple network layers of the first sub-network are determined based on the first intermediate tensor.
15. The method according to claim 13 or 14, wherein, the height×width dimension of the first intermediate tensor is 1×1 or 2×2 or 3×3; or the height×width dimension of the first intermediate tensor is the same as the height×width dimension of the bottleneck tensor, and the number of channels of the first intermediate tensor is 1 / 8 or 1 / 4 or 1 / 2 of the number of channels of the bottleneck tensor.
16. The method according to any one of claims 13 to 15, wherein, the processing result of the first intermediate tensor is further injected into at least one network layer in the first decoder, wherein the spatial dimension of the processing result of the first intermediate tensor is 1×1.
17. The method according to claim 14, wherein, the first sub-network includes a pointwise convolutional weight estimation sub-network and a kernel map estimation sub-network, the per-pixel kernel map is generated by the kernel map estimation sub-network using a local component tensor and multiple dynamic pointwise convolutional weights as inputs, wherein the convolutional weights of multiple network layers of the kernel map estimation sub-network are the multiple dynamic pointwise convolutional weights; the multiple dynamic pointwise convolutional weights are generated by the pointwise convolutional weight estimation sub-network using the first intermediate tensor as an input, and the local component tensor is determined based on the second intermediate tensor.
18. The method according to claim 17, wherein, the first sub-network further includes a kernel embedding map estimation sub-network, and the local component tensor is determined by the kernel embedding map estimation sub-network using the second intermediate tensor as an input.
19. The method according to claim 17, wherein, the enhanced network further includes a kernel embedding map estimation sub-network, the local component tensor is the second intermediate tensor, and the second intermediate tensor is generated by the kernel embedding map estimation sub-network using at least one network layer tensor in the network layer tensors of the first decoder as an input; the second intermediate tensor is concatenated with one network layer tensor in the network layer tensors of the first decoder.
20. The method according to any one of claims 17 to 19, wherein, The height × width dimension of the local component tensor is the same as the height × width dimension of the high-quality image, and the number of channels of the local component tensor is 4 or 8 or 16 or 24 or 32.
21. The method according to any one of claims 13 to 20, wherein, a skip connection is provided between the first encoder and the first decoder.
22. The method according to any one of claims 12 to 21, wherein, the enhancement network adopts a U-net design.
23. A neural network training device, wherein, comprising: a memory for storing computer program instructions; a processor coupled to the memory and, when executing the computer program instructions, configured to: obtain a plurality of training image pairs, wherein each pair of the plurality of training image pairs includes a high-quality image and a low-quality image, and wherein the high-quality image and the low-quality image are images of the same scene with different qualities; input the low-quality image into an enhancement network to generate an enhanced image; input the high-quality image and at least two intermediate tensors into a degradation network to generate a degraded image, wherein the at least two intermediate tensors are determined based on the enhancement network; determine a first loss function based on the high-quality image and the enhanced image; determine a second loss function based on the low-quality image and the degraded image; update the neural network based on the first loss function and the second loss function, wherein the neural network includes the enhancement network and the degradation network.
24. The neural network training device according to claim 23, wherein, the enhancement network includes a first encoder and a first decoder; the at least two intermediate tensors include a first intermediate tensor and a second intermediate tensor, wherein the first intermediate tensor is determined based on a bottleneck tensor in the first encoder, and the second intermediate tensor is determined based on at least one network layer in the network layer of the first decoder.
25. The neural network training device according to claim 24, wherein, the degradation network includes a first sub-network, and the processor, when executing the computer program instructions, is configured to: input the first intermediate tensor and the second intermediate tensor into the first sub-network to generate a per-pixel kernel map, wherein the convolution weights of multiple network layers of the first sub-network are determined based on the first intermediate tensor; input the high-quality image into the degradation network; convolve the high-quality image with the per-pixel kernel to generate the degraded image, wherein the per-pixel kernel is determined from the per-pixel kernel map.
26. The neural network training device according to claim 24 or 25, wherein, the height × width dimension of the first intermediate tensor is 1×1 or 2×2 or 3×3; or the height × width dimension of the first intermediate tensor is the same as the height × width dimension of the bottleneck tensor, and the number of channels of the first intermediate tensor is 1 / 8 or 1 / 4 or 1 / 2 of the number of channels of the bottleneck tensor.
27. The neural network training device according to any one of claims 24 to 26, Characterized in that, when the processor executes the computer program instructions, it is further configured to inject the processing result of the first intermediate tensor into at least one network layer in the first decoder, wherein the spatial dimension of the processing result of the first intermediate tensor is 1×1.
28. The neural network training device according to claim 25, Characterized in that, the first sub-network includes a pointwise convolution weight estimation sub-network and a kernel map estimation sub-network, and when the processor executes the computer program instructions, it is configured to: input the first intermediate tensor into the pointwise convolution weight estimation sub-network to generate a plurality of dynamic pointwise convolution weights; input the second intermediate tensor into the first sub-network and determine a local component tensor based on the second intermediate tensor; input the local component tensor and the plurality of dynamic pointwise convolution weights into the kernel map estimation sub-network to generate the per-pixel kernel map, wherein the convolution weights of multiple network layers of the kernel map estimation sub-network are the plurality of dynamic pointwise convolution weights.
29. The neural network training device according to claim 28, Characterized in that, the first sub-network further includes a kernel embedding map estimation sub-network, and when the processor executes the computer program instructions, it is configured to: input the second intermediate tensor into the kernel embedding map estimation sub-network to generate the local component tensor.
30. The neural network training device according to claim 28, Characterized in that, the enhancement network further includes a kernel embedding map estimation sub-network, the local component tensor is the second intermediate tensor, and when the processor executes the computer program instructions, it is further configured to: input at least one network layer tensor in the network layer tensors in the first decoder into the kernel embedding map estimation sub-network to generate the second intermediate tensor; cascade the second intermediate tensor with one network layer tensor in the network layer tensors in the first decoder.
31. The neural network training device according to any one of claims 28 to 30, Characterized in that, the height×width dimension of the local component tensor is the same as the height×width dimension of the high-quality image, and the number of channels of the local component tensor is 4 or 8 or 16 or 24 or 32.
32. The neural network training device according to any one of claims 24 to 31, Characterized in that, a skip connection is provided between the first encoder and the first decoder.
33. The neural network training device according to any one of claims 23 to 32, Characterized in that, the enhancement network adopts a U-net design.
34. An image processing device, Characterized in that, comprising: a memory for storing computer program instructions; a processor coupled to the memory and, when executing the computer program instructions, configured to: acquire an image to be processed; Process the image to be processed using an enhancement network, where the enhancement network is obtained by training a neural network including the enhancement network and a degradation network based on a plurality of training image pairs, and each pair of the plurality of training image pairs includes a high-quality image and a low-quality image; the neural network is trained based on a first loss function and a second loss function, where the first loss function is determined based on the high-quality image and the enhanced image, and the second loss function is determined based on the low-quality image and the degraded image; the enhanced image is generated by the enhancement network using the low-quality image as an input, and the degraded image is generated by the degradation network using the high-quality image and at least two intermediate tensors as inputs, where the at least two intermediate tensors are determined based on the enhancement network.
35. The image processing apparatus according to claim 34, wherein, the enhancement network includes a first encoder and a first decoder; the at least two intermediate tensors include a first intermediate tensor and a second intermediate tensor, where the first intermediate tensor is determined based on a bottleneck tensor in the first encoder, and the second intermediate tensor is determined based on at least one network layer in the network layers in the first decoder.
36. The image processing apparatus according to claim 35, wherein, the degradation network includes a first sub-network, and the degraded image is obtained by convolving the high-quality image with a per-pixel kernel, and the per-pixel kernel is determined from a per-pixel kernel map generated by the first sub-network using the first intermediate tensor and the second intermediate tensor as inputs, where the convolution weights of the multiple network layers of the first sub-network are determined based on the first intermediate tensor.
37. The image processing apparatus according to claim 35 or 36, wherein, the height × width size of the first intermediate tensor is 1×1 or 2×2 or 3×3; or the height × width size of the first intermediate tensor is the same as the height × width size of the bottleneck tensor, and the number of channels of the first intermediate tensor is 1 / 8 or 1 / 4 or 1 / 2 of the number of channels of the bottleneck tensor.
38. The image processing apparatus according to any one of claims 35 to 37, wherein, the processing result of the first intermediate tensor is further injected into at least one network layer in the first decoder, where the spatial size of the processing result of the first intermediate tensor is 1×1.
39. The image processing apparatus according to claim 36, wherein, The first sub-network includes a pointwise convolution weight estimation sub-network and a kernel map estimation sub-network. The per-pixel kernel map is generated by the kernel map estimation sub-network using the local component tensor and a plurality of dynamic pointwise convolution weights as inputs. Among them, the convolution weights of multiple network layers of the kernel map estimation sub-network are the plurality of dynamic pointwise convolution weights; the plurality of dynamic pointwise convolution weights are generated by the pointwise convolution weight estimation sub-network using the first intermediate tensor as an input, and the local component tensor is determined based on the second intermediate tensor.
40. The image processing apparatus according to claim 39, wherein, the first sub-network further includes a kernel embedding map estimation sub-network, and the local component tensor is determined by the kernel embedding map estimation sub-network using the second intermediate tensor as an input.
41. The image processing apparatus according to claim 39, wherein, the enhancement network further includes a kernel embedding map estimation sub-network, the local component tensor is the second intermediate tensor, and the second intermediate tensor is generated by the kernel embedding map estimation sub-network using at least one network layer tensor among the network layer tensors in the first decoder as an input; the second intermediate tensor is concatenated with one network layer tensor among the network layer tensors in the first decoder.
42. The image processing apparatus according to any one of claims 39 to 41, wherein, the height × width dimension of the local component tensor is the same as the height × width dimension of the high-quality image, and the number of channels of the local component tensor is 4 or 8 or 16 or 24 or 32.
43. The image processing apparatus according to any one of claims 35 to 42, wherein, a skip connection is provided between the first encoder and the first decoder.
44. The image processing apparatus according to any one of claims 34 to 43, wherein, the enhancement network adopts a U-net design.
45. A computer-readable storage medium, wherein, the computer-readable storage medium stores instructions that, when the instructions are run on a device, cause the device to execute the method according to any one of claims 1 to 11 and 12 to 22.
46. A computer program product, wherein, when the computer program product is run on a device, cause the device to execute the method according to any one of claims 1 to 11 and 12 to 22.
47. A chip system, wherein, includes a memory and a processor. Among them, the memory is used to store a computer program, and the processor is used to call the computer program from the memory and run the computer program, so that the device where the chip system is located executes the method according to any one of claims 1 to 11 and 12 to 22.