Image Enhancement Method and Device

The image is decomposed and processed through wavelet transformation and feature extraction modules, and the application problem of lightweight image enhancement network on devices with insufficient computing power is solved, and high-quality image enhancement is achieved.

CN111951195BActive Publication Date: 2025-06-17HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010654234.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-08
Publication Date
2025-06-17
Estimated Expiration
2040-07-08

AI Technical Summary

Technical Problem

The existing lightweight image enhancement network reduces the computational volume while reducing image enhancement quality, making it difficult to effectively apply on devices with lower computing power.

Method used

The image to be processed is decomposed into multiple feature sub-maps with different frequencies through wavelet transformation, and the feature extraction module is used to extract features of these sub-maps separately, and the frequency differences between different feature sub-maps are processed separately, and the enhanced image is finally generated through inverse wavelet transformation.

Benefits of technology

While reducing the amount of image processing calculation, the image enhancement quality is improved and is suitable for devices with lower computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111951195B_ABST
    Figure CN111951195B_ABST
Patent Text Reader

Abstract

The present application provides an image enhancement method and apparatus. It relates to the field of artificial intelligence, specifically to the field of image enhancement. The method includes: performing wavelet transform on the input feature map of the image to be processed to obtain N feature submaps with different frequencies, where N is an integer greater than 1; respectively performing feature extraction on the N feature submaps to obtain N intermediate results; performing inverse wavelet transform on the N intermediate results to obtain a first feature map; and generating an enhanced image according to the first feature map. The present application performs feature extraction based on low-resolution feature submaps, and considering the frequency differences between different feature submaps, separately processes multiple feature submaps. In this way, on the one hand, the computational complexity in the image processing process can be reduced, and on the other hand, the image enhancement quality can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and particularly to an image enhancement method and apparatus. Background Art

[0002] Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. The research in the field of artificial intelligence includes robots, natural language processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, AI basic theory, etc.

[0003] Image enhancement technology is one of the key technologies in the field of image processing. This technology is used to improve and enhance the quality of the original image, and even reveal the hidden information in the original image, making it more suitable for observation by the human visual system or processing by other subsequent functional modules. Image enhancement technology has important application values in fields such as high-definition television, high-definition conversion of old movies, monitoring devices, satellite images, and medical images.

[0004] With the rapid development of artificial intelligence and the continuous improvement of computing power, the image enhancement effect of neural networks with image enhancement functions (referred to as image enhancement networks for short) has been greatly improved. However, while the effect has been improved, the structure of the image enhancement network has become more complex, and the amount of computation has also increased accordingly. This greatly limits the application of image enhancement networks in some devices with low computing power, such as mobile phones, cameras, and smart homes.

[0005] In order to promote the application of image enhancement networks on devices with low computing power, many current efforts are dedicated to constructing lightweight image enhancement networks. However, existing lightweight image enhancement networks reduce the amount of computation while also reducing the image enhancement quality. Summary of the Invention

[0006] This application provides an image enhancement method and apparatus that can provide high image enhancement quality on the premise of reducing the computation amount of the image enhancement network.

[0007] First aspect, there is provided an image enhancement method, including: obtaining an input feature map of an image to be processed; performing wavelet transform on the input feature map to obtain N feature submaps with different frequencies, where N is an integer greater than 1, and the value of N is associated with the level of the wavelet transform, and the higher the level of the wavelet transform, the larger the value of N; using N feature extraction modules to respectively perform feature extraction on the N feature submaps to obtain N intermediate results; performing inverse wavelet transform on the N intermediate results to obtain a first feature map; generating an enhanced image of the image to be processed according to the first feature map.

[0008] Based on the feature submap with low resolution, feature extraction is performed, and considering the frequency difference between different feature submaps, multiple feature submaps are processed separately. On the one hand, this can reduce the computational complexity in the image processing process, and on the other hand, it can improve the image enhancement quality.

[0009] Combined with the first aspect, in some embodiments of the first aspect, the N feature submaps include a first feature submap and a second feature submap, the first feature submap and the second feature submap are any two feature submaps with different frequencies among the N feature submaps, the N feature extraction modules include a first feature extraction module corresponding to the first feature submap and a second feature extraction module corresponding to the second feature submap, and the step of using N feature extraction modules to respectively perform feature extraction on the N feature submaps includes: using the first feature extraction module to perform feature extraction on the first feature submap to obtain a first intermediate result; performing a splicing operation on the first intermediate result and the second feature submap to obtain a second feature submap fused with the first intermediate result; using the second feature extraction module to perform feature extraction on the second feature submap fused with the first intermediate result to obtain a second intermediate result.

[0010] Before using the second feature extraction module to perform feature extraction on the second feature submap, fusing the first intermediate result obtained by the first feature extraction module with the second feature submap is equivalent to using the information in the submap of one frequency to guide and constrain the feature extraction process of the submap of another frequency, which can ensure the accuracy of feature extraction and thus improve the image enhancement quality.

[0011] Combined with the first aspect, in some embodiments of the first aspect, the frequency of the second feature submap is lower than the frequency of the first feature submap.

[0012] Since the frequency of the second feature sub - graph is lower than that of the first feature sub - graph, the first feature sub - graph contains high - frequency information and the second feature sub - graph contains low - frequency information. High - frequency information usually corresponds to the detailed information of the image, and low - frequency information usually corresponds to the structural information of the image. The two are correlated. Using high - frequency information to guide the feature extraction process of low - frequency information can improve the efficiency of feature extraction.

[0013] Combined with the first aspect, in some embodiments of the first aspect, the frequency of the second feature sub - graph is higher than that of the first feature sub - graph.

[0014] Combined with the first aspect, in some embodiments of the first aspect, the N feature sub - graphs include a third feature sub - graph, the third feature sub - graph is any one of the N feature sub - graphs, the N feature sub - graphs further include M feature sub - graphs with frequencies greater than that of the third feature sub - graph, the N feature extraction modules include a third feature extraction module corresponding to the third feature sub - graph and M feature extraction modules corresponding to the M feature sub - graphs, M is an integer greater than 1. The step of using the N feature extraction modules to extract features from the N feature sub - graphs respectively includes: using the M feature extraction modules to extract features from the M feature sub - graphs respectively to obtain M intermediate results; performing a splicing operation on the M intermediate results and the third feature sub - graph to obtain a third feature sub - graph integrating the M intermediate results; using the third feature extraction module to extract features from the third feature sub - graph integrating the M intermediate results to obtain a third intermediate result.

[0015] Fusing the intermediate results corresponding to the M feature sub - graphs with frequencies greater than that of the third feature sub - graph into the third feature sub - graph can make the high - frequency information in the third feature sub - graph richer. High - frequency information represents detailed information, and the full utilization of detailed information can improve the quality of image enhancement.

[0016] Combined with the first aspect, in some embodiments of the first aspect, the step of generating the enhanced image of the to - be - processed image according to the first feature map includes: generating the enhanced image of the to - be - processed image according to the first feature map and the input feature map.

[0017] Compared with the resolution of the feature sub - graph obtained by wavelet transform, the input feature map is a high - resolution feature map. Retaining the input feature map is equivalent to retaining the high - resolution feature information in the to - be - processed image, which can improve the quality of image enhancement.

[0018] Combined with the first aspect, in some embodiments of the first aspect, the step of generating the enhanced image of the to - be - processed image according to the first feature map includes: performing super - resolution processing on the to - be - processed image according to the first feature map to obtain the enhanced image of the to - be - processed image.

[0019] In combination with the first aspect, in some embodiments of the first aspect, generating the enhanced image of the image to be processed based on the first feature map includes: performing noise reduction or haze removal on the image to be processed based on the first feature map to obtain the enhanced image of the image to be processed.

[0020] In combination with the first aspect, in some embodiments of the first aspect, the wavelet transform is a Haar wavelet transform.

[0021] A second aspect provides a training method for an image enhancement network, including: obtaining the input feature map of a sample image; performing wavelet transform on the input feature map to obtain N feature submaps with different frequencies, where N is an integer greater than 1, and the value of N is associated with the level of the wavelet transform, and the higher the level of the wavelet transform, the larger the value of N; respectively using N feature extraction modules to perform feature extraction on the N feature submaps to obtain N intermediate results; performing inverse wavelet transform on the N intermediate results to obtain the first feature map; generating the enhanced image of the sample image based on the first feature map; and training the image enhancement network according to the difference between the enhanced image and the supervised image of the sample image.

[0022] In combination with the second aspect, in some embodiments of the second aspect, the N feature submaps include a first feature submap and a second feature submap, the first feature submap and the second feature submap are any two feature submaps with different frequencies among the N feature submaps, the N feature extraction modules include a first feature extraction module corresponding to the first feature submap and a second feature extraction module corresponding to the second feature submap, and respectively using N feature extraction modules to perform feature extraction on the N feature submaps includes: using the first feature extraction module to perform feature extraction on the first feature submap to obtain a first intermediate result; performing a splicing operation on the first intermediate result and the second feature submap to obtain a second feature submap fused with the first intermediate result; and using the second feature extraction module to perform feature extraction on the second feature submap fused with the first intermediate result to obtain a second intermediate result.

[0023] In combination with the second aspect, in some embodiments of the second aspect, the frequency of the second feature submap is lower than the frequency of the first feature submap.

[0024] In combination with the second aspect, in certain embodiments of the second aspect, the N feature sub - graphs include a third feature sub - graph, where the third feature sub - graph is any one of the N feature sub - graphs, the N feature sub - graphs further include M feature sub - graphs with a frequency greater than that of the third feature sub - graph, the N feature extraction modules include a third feature extraction module corresponding to the third feature sub - graph and M feature extraction modules corresponding to the M feature sub - graphs, M is an integer greater than 1, and the step of using the N feature extraction modules to respectively extract features from the N feature sub - graphs includes: using the M feature extraction modules to respectively extract features from the M feature sub - graphs to obtain M intermediate results; performing a splicing operation on the M intermediate results and the third feature sub - graph to obtain a third feature sub - graph fused with the M intermediate results; using the third feature extraction module to extract features from the third feature sub - graph fused with the M intermediate results to obtain a third intermediate result.

[0025] In combination with the second aspect, in certain embodiments of the second aspect, the step of generating the enhanced image of the sample image according to the first feature map includes: generating the enhanced image of the sample image according to the first feature map and the input feature map.

[0026] In combination with the second aspect, in certain embodiments of the second aspect, the step of generating the enhanced image of the sample image according to the first feature map includes: performing super - resolution processing on the sample image according to the first feature map to obtain the enhanced image of the sample image.

[0027] In combination with the second aspect, in certain embodiments of the second aspect, the step of generating the enhanced image of the sample image according to the first feature map includes: performing noise reduction or defogging processing on the sample image according to the first feature map to obtain the enhanced image of the sample image.

[0028] In combination with the second aspect, in certain embodiments of the second aspect, the wavelet transform is a Haar wavelet transform.

[0029] In a third aspect, there is provided an image enhancement device, including a module for performing the method described in the first aspect.

[0030] In a fourth aspect, there is provided a training device for an image enhancement network, including a module for performing the method described in the second aspect.

[0031] In a fifth aspect, there is provided an image enhancement device, including a memory for storing a program; a processor for executing the program stored in the memory, and when the program stored in the memory is executed, the processor is used to execute the method described in the first aspect.

[0032] In a sixth aspect, there is provided a training device for an image enhancement network, including a memory for storing programs; and a processor for executing the programs stored in the memory. When the programs stored in the memory are executed, the processor is configured to execute the method described in the second aspect.

[0033] In a seventh aspect, there is provided a computer-readable storage medium storing instructions which, when run on a computer, cause the computer to execute the method described in the first aspect or the second aspect.

[0034] In an eighth aspect, there is provided a computer program product containing instructions which, when run on a computer, cause the computer to execute the method described in the first aspect or the second aspect.

[0035] In a ninth aspect, there is provided a chip including a processor and a data interface. The processor reads instructions stored on a memory through the data interface and executes the method in the first aspect or the second aspect.

[0036] Optionally, as an implementation, the chip may further include a memory storing instructions, and the processor is configured to execute the instructions stored on the memory. When the instructions are executed, the processor is configured to execute the method in the first aspect or the second aspect.

[0037] In a tenth aspect, there is provided an electronic device including the device in any one of the third aspect to the sixth aspect described above. Description of the Drawings

[0038] Figure 1 It is a schematic diagram of an application scenario of an image enhancement method provided in an embodiment of the present application.

[0039] Figure 2 It is a schematic diagram of an application scenario of an image enhancement method provided in another embodiment of the present application.

[0040] Figure 3 It is a schematic diagram of two-dimensional wavelet transform of an image.

[0041] Figure 4 It is a schematic diagram of inverse wavelet transform of an image.

[0042] Figure 5 It is a schematic diagram of a system architecture provided in an embodiment of the present application.

[0043] Figure 6 It is a schematic diagram of the structure of an image enhancement network provided in an embodiment of the present application.

[0044] Figure 7 It is a schematic diagram of the chip hardware structure provided in an embodiment of the present application.

[0045] Figure 8 Schematic flowchart of an image enhancement method provided by an embodiment of the present application.

[0046] Figure 9 Schematic diagram of a processing method for preprocessing using a preprocessing module provided by an embodiment of the present application.

[0047] Figure 10 Schematic diagram of a processing method for image enhancement processing using an image enhancement module provided by an embodiment of the present application.

[0048] Figure 11 Schematic flowchart of an image enhancement method provided by another embodiment of the present application.

[0049] Figure 12 Schematic diagram of a processing method for image enhancement processing using an image enhancement module provided by another embodiment of the present application.

[0050] Figure 13 Schematic diagram of a processing method for image enhancement processing using an image enhancement module provided by another embodiment of the present application.

[0051] Figure 14 Schematic diagram of a processing method for image enhancement processing using an image enhancement module provided by another embodiment of the present application.

[0052] Figure 15 Schematic diagram of a processing method for image enhancement processing using an image enhancement module provided by another embodiment of the present application.

[0053] Figure 16 Schematic diagram of a processing method for image enhancement processing using an image enhancement module provided by another embodiment of the present application.

[0054] Figure 17 Schematic diagram of a processing method for super-resolution processing using a post-processing module provided by an embodiment of the present application.

[0055] Figure 18 Schematic diagram of a processing method for noise reduction or defogging processing using a post-processing module provided by an embodiment of the present application.

[0056] Figure 19 Schematic flowchart of an image enhancement method provided by another embodiment of the present application.

[0057] Figure 20 Schematic diagram of a processing method for image enhancement processing using an image enhancement network provided by an embodiment of the present application.

[0058] Figure 21 Schematic diagram of the structure of a convolutional module provided by an embodiment of the present application.

[0059] Figure 22 The structural schematic diagram of the residual block provided by an embodiment of the present application.

[0060] Figure 23 The structural schematic diagram of the image enhancement device provided by an embodiment of the present application.

[0061] Figure 24 The hardware structural schematic diagram of the image enhancement device provided by an embodiment of the present application. Detailed implementation manners

[0062] Next, the technical solutions in the present application will be described with reference to the accompanying drawings.

[0063] The image enhancement method and device provided by the embodiments of the present application can be applied in intelligent vehicles for assisted driving and autonomous driving, and can also be applied in fields that require image enhancement in the computer vision field such as safe cities and intelligent terminals. Specifically, the technical solutions of the present application can be applied in video stream transmission scenarios and video surveillance scenarios. Next, with reference to Figure 1 and Figure 2 a simple introduction to the video stream transmission scenario and the video surveillance scenario will be given.

[0064] Video stream transmission scenario:

[0065] For example, when playing a video on the client of an intelligent terminal (such as a mobile phone), in order to reduce the bandwidth requirement of the video stream, the server can transmit a downsampled and low-quality video stream with a lower resolution to the client through the network. Then the client can use the image enhancement method provided by the embodiments of the present application to enhance the images in the low-quality video stream. For example, perform operations such as super-resolution and noise reduction on the images in the video, and finally present high-quality images to the user. This method can significantly reduce the network bandwidth requirement of the video stream without significantly reducing the quality of high-definition images.

[0066] Video surveillance scenario:

[0067] In the field of security, due to adverse conditions such as the installation location of surveillance cameras and limited storage space, the image quality of some video surveillance is poor, which will affect the accuracy of people or recognition algorithms in recognizing targets. Therefore, the image enhancement method provided by the embodiments of the present application can be used to convert low-quality video surveillance images into high-quality high-definition images, so as to effectively restore a large number of details in the surveillance images and provide more effective and richer information for subsequent target recognition tasks.

[0068] Since the embodiments of the present application involve applications in image enhancement and neural networks, for the convenience of understanding, the following will first give a simple introduction to the relevant terms and related concepts such as neural networks that may be involved in the embodiments of the present application.

[0069] (1) Image Enhancement

[0070] Image enhancement is an image processing technique whose main goal is to enhance an image from a lower quality to a higher quality. Image enhancement techniques have a wide range of applications in the field of image processing. Common image enhancement methods include super-resolution, noise reduction, defogging, etc.

[0071] (2) Super-Resolution

[0072] Super-resolution is a type of image enhancement technique. Given a single or a set of low-resolution images, it can restore the high-frequency detail information of the image and generate a higher-resolution image by one or more of the following means: learning the prior knowledge of the image, utilizing the similarity of the image itself, and complementing information from multiple frames of images. According to whether the input low-resolution image is a single-frame image or a set of video images, super-resolution can be divided into single-frame image super-resolution and video super-resolution.

[0073] (3) Noise Reduction

[0074] Images are often affected by imaging devices and the external environment during the digitization and transmission processes, resulting in the images containing noise. The process of reducing the noise in an image is called image noise reduction, and sometimes it can also be called image denoising.

[0075] (4) Defogging

[0076] Some images are affected by haze, fog, etc. in the real world and contain fog, resulting in low image quality. Image defogging refers to removing the fog from the image and restoring a fog-free image.

[0077] (5) Image Features

[0078] Image features mainly include color features, texture features, shape features, and spatial relationship features of the image, etc.

[0079] Color features are a type of global feature that describes the surface properties of the scene corresponding to the image or image region; generally, color features are pixel-based features, and at this time, all pixels belonging to the image or image region contribute individually. Since color is insensitive to changes in the direction, size, etc. of the image or image region, color features cannot well capture the local features of the objects in the image.

[0080] Texture features are also a type of global feature that describes the surface properties of the scene corresponding to the image or image region; however, since texture is only a characteristic of the object's surface and cannot fully reflect the essential attributes of the object, only using texture features cannot obtain high-level image content. Different from color features, texture features are not pixel-based features, and they need to be statistically calculated in a region containing multiple pixels.

[0081] There are two types of representation methods for shape features. One is contour features, and the other is region features. The contour features of an image mainly target the outer boundary of an object, while the region features of an image are related to the entire shape region.

[0082] Spatial relationship features refer to the mutual spatial positions or relative direction relationships between multiple objects segmented in an image. These relationships can also be divided into connection / adjacency relationships, overlap / overlay relationships, and inclusion / containment relationships, etc. Generally, spatial position information can be divided into two categories: relative spatial position information and absolute spatial position information. The former relationship emphasizes the relative situation between objects, such as up-down, left-right relationships, etc., and the latter relationship emphasizes the distance and orientation between objects.

[0083] It should be noted that the above-listed image features can be used as some examples of the features existing in an image. An image can also have other features, such as higher-level features: semantic features, which will not be elaborated here.

[0084] (6) Wavelet transform

[0085] Wavelet transform can also be called wavelet analysis, which is a signal analysis method that uses the oscillatory waveform of a "mother wavelet" with finite length or rapid decay to represent a signal. Wavelet transform is divided into two major categories: discrete wavelet transform (DWT) and continuous wavelet transform (CWT). Among them, discrete wavelet transform is commonly used for image decomposition processing. An image can be decomposed into multiple sub-images with separated high and low frequencies through discrete wavelet transform. Haar wavelet transform is the simplest one among wavelet transforms and is also the earliest proposed wavelet transform.

[0086] Figure 3 is a schematic diagram of two-dimensional discrete wavelet transform of an image. Refer to Figure 3 , in the first-level wavelet transform of the input image 300, the low-pass filter 301 and the high-pass filter 302 are respectively used to perform row filtering on the input image 300, and then downsampling 303, 304 of column data is performed on the filtered data; then the low-pass filters 305, 307 and the high-pass filters 306, 308 are respectively used to perform column filtering on the downsampled data, and downsampling 309, 310, 311, 312 of row data is performed on the filtered data. In this way, the first-level wavelet transform of the input image is completed. The image after the first-level wavelet transform of the input image is decomposed into multiple sub-images with different frequencies, namely the LL sub-image 313, the LH sub-image 314, the HL sub-image 315, and the HH sub-image 316.

[0087] The LL sub - graph 313 corresponds to the low - frequency information of the input image. The LH sub - graph 314 corresponds to the changes along the columns, that is, the high - frequency information in the vertical direction. The HL sub - graph 315 corresponds to the changes along the rows, that is, the high - frequency information in the horizontal direction. The HH sub - graph 316 corresponds to the changes in the diagonal direction, that is, the high - frequency information in the diagonal direction. Among them, the low - frequency information corresponds to the structural information of the image, and the high - frequency information corresponds to the detailed information of the image, such as the edges of the image.

[0088] (7) Inverse wavelet transform

[0089] The inverse wavelet transform is the inverse process of the wavelet transform, which is used to reconstruct the input image using the sub - graphs after wavelet transform. Figure 4 It is a schematic diagram of the inverse wavelet transform of the image. The LL sub - graph, LH sub - graph, HL sub - graph, and HH sub - graph respectively perform up - sampling operations along the rows and columns using two high - pass and low - pass filters to obtain the reconstructed input image.

[0090] (8) Neural network

[0091] A neural network can be composed of neural units. A neural unit can refer to an operation unit with x s and intercept 1 as inputs. The output of this operation unit can be:

[0092]

[0093] where s = 1, 2,......n, n is a natural number greater than 1, W s is the weight of x s , b is the bias of the neural unit. f is the activation function of the neural unit (activation functions), which is used to introduce non - linear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting many such single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neural units.

[0094] (9) Deep neural network

[0095] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with many hidden layers. Here, "many" does not have a specific measurement standard. Dividing the DNN according to the positions of different layers, the neural network inside the DNN can be divided into three categories: the input layer, the hidden layer, and the output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the (i + 1)-th layer. Although the DNN looks very complex, in terms of the work of each layer, it is actually not complex. Simply put, it is the following linear relationship expression: Among them, is the input vector. is the output vector, b is the bias vector, W is the weight matrix (also known as the coefficient), and α(.) is the activation function. Each layer simply performs the following simple operation on the input vector to obtain the output vector Since the DNN has many layers, the number of coefficients W and bias vectors b is also large. The definitions of these parameters in the DNN are as follows: Taking the coefficient W as an example: Suppose in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer where the coefficient W is located, and the subscripts correspond to the index 2 of the output third layer and the index 4 of the input second layer. In summary: The coefficient from the k-th neuron in the (L - 1)-th layer to the j-th neuron in the L-th layer is defined as It should be noted that the input layer does not have the W parameter. In a deep neural network, more hidden layers enable the network to better depict complex situations in the real world. Theoretically, the more parameters a model has, the higher its complexity and the larger its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is also the process of learning the weight matrix, and its ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (the weight matrix formed by vectors W of many layers).

[0096] (10) Convolutional neural network

[0097] A convolutional neural network (CNN) is a deep neural network with a convolutional structure. A convolutional neural network contains a feature extractor composed of convolutional layers and subsampling layers. This feature extractor can be regarded as a filter, and the convolution process can be regarded as convolving a trainable filter with an input image or a convolutional feature map. A convolutional layer refers to the neuron layer in a convolutional neural network that performs convolution processing on the input signal. In the convolutional layer of a convolutional neural network, a neuron can be connected to only some adjacent layer neurons. In a convolutional layer, there are usually several feature maps, and each feature map can be composed of some neurons arranged in a rectangle. The neurons in the same feature map share weights, and the shared weights are the convolutional kernels. Sharing weights can be understood as a way of extracting image information that is independent of position. The underlying principle here is that the statistical information of a certain part of an image is the same as that of other parts. That is to say, the image information learned in a certain part can also be used in another part. Therefore, the same learned image information can be used for all positions on the image. In the same convolutional layer, multiple convolutional kernels can be used to extract different image information. Generally, the more convolutional kernels there are, the richer the image information reflected by the convolution operation.

[0098] The convolutional kernel can be initialized in the form of a matrix of random size, and during the training process of the convolutional neural network, the convolutional kernel can learn to obtain reasonable weights. Additionally, the direct benefit brought by sharing weights is to reduce the connections between the layers of the convolutional neural network while reducing the risk of overfitting.

[0099] (11) Residual network

[0100] A residual network can be a network in which, in addition to being connected layer by layer between multiple hidden layers of a neural network, there is also a direct connection branch. For example, a neural network includes 4 hidden layers. Among them, the first hidden layer is connected to the second hidden layer, the second hidden layer is connected to the third hidden layer, and the third hidden layer is connected to the fourth hidden layer. These 4 hidden layers represent an operation path of the data in the neural network. In addition to the above-mentioned operation path connected layer by layer, a residual network can also be set with an additional direct connection branch. This direct connection branch can directly connect from the first hidden layer to the fourth hidden layer, thereby skipping the second and third hidden layers and directly transmitting the data of the first hidden layer to the fourth hidden layer for operation.

[0101] (12) Loss function

[0102] During the process of training a deep neural network, since we hope that the output of the deep neural network is as close as possible to the value we really want to predict, we can compare the predicted value of the current network with the target value we really want, and then update the weight vector of each layer of the neural network according to the difference between the two. (Of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the deep neural network). For example, if the predicted value of the network is too high, we adjust the weight vector to make it predict lower, and keep adjusting until the deep neural network can predict the target value we really want or a value very close to the target value we really want. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function. They are important equations used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the deep neural network becomes a process of minimizing this loss as much as possible.

[0103] (13) Backpropagation algorithm

[0104] Convolutional neural networks can use the backpropagation (BP) algorithm to correct the parameters in the initial neural network model during the training process, so that the reconstruction error loss of the neural network model becomes smaller and smaller. Specifically, forward-propagating the input signal until the output generates an error loss, and updating the parameters in the initial neural network model by backpropagating the error loss information, so that the error loss converges. The backpropagation algorithm is a backpropagation movement dominated by the error loss, aiming to obtain the optimal parameters of the neural network model, such as the weight matrix.

[0105] The following combines Figure 5 to introduce the system architecture provided by the embodiments of the present application in detail. Figure 5 is a schematic diagram of the system architecture provided by an embodiment of the present application. As Figure 5 shown, the system architecture 500 includes an execution device 510, a training device 520, a database 530, a client device 540, a data storage system 550, and a data acquisition system 560.

[0106] The execution device 510 includes a computing module 511, an I / O interface 512, a preprocessing module 513, and a preprocessing module 514. The computing module 511 may include a target model / rule 501, and the preprocessing module 513 and the preprocessing module 514 are optional.

[0107] The data acquisition device 560 is used to acquire training data. The training data in the embodiments of this application includes the sample images to be enhanced and the supervision images. Among them, the sample images are low-quality images, and the supervision images are high-quality images corresponding to the sample images obtained in advance before model training. For example, the sample images can be low-resolution images, and the supervision images are high-resolution images; or, the sample images can be images containing fog or noise, and the supervision images are images with fog or noise removed. After acquiring the training data, the data acquisition device 560 stores these training data in the database 530, and the training device 520 trains to obtain the target model / rule 501 based on the training data maintained in the database 530.

[0108] The above-mentioned target model / rule 501 can be used to implement the image enhancement method in the embodiments of this application, that is, by inputting the image to be processed into the target model / rule 501, the processed enhanced image can be obtained. It should be noted that in practical applications, the training data maintained in the database 530 may not all come from the acquisition of the data acquisition device 560, and it may also be received from other devices. Additionally, it should be noted that the training device 520 may not necessarily train the target model / rule 501 entirely based on the training data maintained in the database 530, and it may also obtain training data from the cloud or other places for model training. The above description should not be regarded as a limitation to the embodiments of this application.

[0109] The target model / rule 501 trained according to the training device 520 can be applied to different systems or devices, such as being applied to Figure 5 the execution device 510 shown. The execution device 510 can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) / virtual reality (VR) device, a vehicle-mounted terminal, etc., or it can also be a server or the cloud, etc. In Figure 5 , the execution device 510 is configured with an input / output (I / O) interface 512 for data interaction with external devices. The user can input data to the I / O interface 512 through the client device 540. The input data in the embodiments of this application can include: the image to be processed input by the client device.

[0110] The preprocessing module 513 and the preprocessing module 514 are used to perform preprocessing according to the input data (such as the image to be processed) received by the I / O interface 512. In the embodiments of this application, there may be no preprocessing module 513 and preprocessing module 514, or there may be only one preprocessing module. When the preprocessing module 513 and the preprocessing module 514 do not exist, the computing module 511 can directly process the input data.

[0111] When the execution device 510 preprocesses the input data, or during the relevant processing such as the calculation module 511 of the execution device 510 performing calculations, the execution device 510 can call the data, code, etc. in the data storage system 550 for corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing into the data storage system 550.

[0112] Finally, the I / O interface 512 presents the processing result, such as the enhanced image obtained after processing, to the client device 540, and thus provides it to the user.

[0113] It should be noted that the training device 520 can generate the corresponding target model / rule 501 based on different targets or different tasks and different training data. The corresponding target model / rule 501 can be used to achieve the above targets or complete the above tasks, so as to provide the required results for the user.

[0114] In Figure 5 In the shown case, the user can manually give the input data (the input data can be the image to be processed), and this "manually giving the input data" can be operated through the interface provided by the I / O interface 512. In another case, the client device 540 can automatically send the input data to the I / O interface 512. If the client device 540 is required to automatically send the input data and user authorization is needed, the user can set the corresponding permissions in the client device 540. The user can view the results output by the execution device 510 in the client device 540, and the specific presentation forms can be display, sound, action and other specific ways. The client device 540 can also be used as a data acquisition end to collect the input data input to the I / O interface 512 and the output result of the output I / O interface 512 shown in the figure as new sample data, and store them in the database 530. Of course, it can also be collected without going through the client device 540, but the I / O interface 512 directly stores the input data input to the I / O interface 512 and the output result of the output I / O interface 512 shown in the figure as new sample data into the database 530.

[0115] It should be noted that Figure 5 This is only a schematic diagram of a system architecture provided by the embodiments of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in Figure 5 in, the data storage system 550 is an external memory relative to the execution device 510. In other cases, the data storage system 550 can also be placed in the execution device 510.

[0116] Such as Figure 5 shown, the target model / rule 501 trained according to the training device 520 can be the image enhancement network in the embodiments of the present application. Specifically, asFigure 6 As shown in Figure 6 , the image enhancement network 600 provided by an embodiment of the present application may include a preprocessing module 610, an image enhancement module 620, and a postprocessing module 630.

[0117] The preprocessing module 610 is configured to perform preliminary feature extraction on an input image to be processed, including features such as color features, texture features, shape features, and spatial relationship features, to obtain a feature map of the image to be processed.

[0118] The image enhancement module 620 is configured to use the feature map obtained by preprocessing the image to be processed as an input feature map, and perform image enhancement processing on it to obtain a processed feature map.

[0119] The postprocessing module 630 is configured to perform different processing on the processed feature map according to different image enhancement tasks (such as super-resolution, noise reduction, or haze removal, etc.) to obtain an enhanced image of the image to be processed.

[0120] In some embodiments, in order to achieve a better image enhancement effect, K (K>1) image enhancement modules 620 may be cascaded. Specifically, the output feature map of the (n-1)th image enhancement module 620 may be used as the input feature map of the nth image enhancement module 620, or the output feature maps of the first n-1 image enhancement modules 620 may be concatenated and used as the input feature map of the nth image enhancement module 620. The present application does not limit this.

[0121] Optionally, in some embodiments, the target model / rule 501 trained according to the training device 520 may also include the above-mentioned image enhancement module 620 and postprocessing module 630. The function of the preprocessing module 610 may be integrated in the preprocessing module 513 and / or preprocessing module 514.

[0122] Next, a chip hardware structure provided by an embodiment of the present application will be introduced.

[0123] Figure 7 FIG. Figure 7 is a chip hardware structure diagram provided by an embodiment of the present application. The chip includes a neural network processor 700. The chip may be disposed in an execution device 510 as shown in Figure 5 to complete the computing work of the computing module 511. The chip may also be disposed in a training device 520 as shown in Figure 5 to complete the training work of the training device 520 and output the target model / rule 501. The algorithms of each layer in the image enhancement network shown in Figure 6 can all be implemented in the chip shown in Figure 7 as shown.

[0124] The neural processing unit (NPU) 700 is mounted on the host central processing unit (host CPU) as a coprocessor and tasks are assigned by the host CPU. The core part of the NPU is the arithmetic circuit 703, and the controller 704 controls the arithmetic circuit 703 to extract data from the memory (weight memory 702 or input memory 701) and perform operations.

[0125] In some implementations, the arithmetic circuit 703 includes multiple processing engines (PEs) inside. In some implementations, the arithmetic circuit 703 is a two-dimensional systolic array. The arithmetic circuit 703 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 703 is a general matrix processor.

[0126] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit 703 fetches the corresponding data of matrix B from the weight memory 702 and caches it on each PE in the arithmetic circuit 703. The arithmetic circuit 703 fetches the data of matrix A from the input memory 701 and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are stored in the accumulator 708.

[0127] The vector calculation unit 707 can further process the output of the arithmetic circuit 703, such as vector multiplication, vector addition, exponential operation, logarithmic operation, magnitude comparison, etc. For example, the vector calculation unit 707 can be used for network calculations in non-convolutional / non-FC layers in a neural network, such as pooling, batch normalization, local response normalization, etc.

[0128] In some implementations, the vector calculation unit 707 can store the processed output vector into the unified memory 706. For example, the vector calculation unit 707 can apply a non-linear function to the output of the arithmetic circuit 703, such as a vector of accumulated values, to generate activation values. In some implementations, the vector calculation unit 707 generates normalized values, combined values, or both. In some implementations, the processed output vector can be used as an activation input to the arithmetic circuit 703, for example, for use in subsequent layers in a neural network.

[0129] The unified memory 706 is used to store input data and output data.

[0130] The weight data directly transfers the input data in the external memory to the input memory 701 and / or the unified memory 706 through the direct memory access controller (DMAC) 705, stores the weight data in the external memory into the weight memory 702, and stores the data in the unified memory 706 into the external memory.

[0131] The bus interface unit (BIU) 710 is used to interact between the main CPU, the DMAC, and the instruction fetch buffer 709 through the bus.

[0132] The instruction fetch buffer 709 connected to the controller 704 is used to store the instructions used by the controller 704.

[0133] The controller 704 is used to call the instructions cached in the instruction fetch buffer 709 to control the working process of the arithmetic accelerator.

[0134] Generally, the unified memory 706, the input memory 701, the weight memory 702, and the instruction fetch buffer 709 are all on-chip memories, and the external memory is the memory outside the NPU. The external memory can be a double data rate synchronous dynamic random access memory (DDR SDRAM), a high bandwidth memory (HBM), or other readable and writable memories.

[0135] Next, in conjunction with Figure 8 The image enhancement method of the embodiments of the present application will be introduced in detail.

[0136] Figure 8 It is a schematic flowchart of the image enhancement method provided by an embodiment of the present application. Figure 8 The method can specifically be executed by the computing module 511 in the execution device 510 shown in Figure 5 . Optionally, Figure 8 The method can be processed by a central processing unit (CPU), or jointly processed by a CPU and a graphics processing unit (GPU), or without using a GPU but using other processors suitable for neural network computing. The present application does not make any restrictions. Figure 8 The method includes the following contents.

[0137] S810: Obtain the input feature map of the image to be processed.

[0138] The input feature map can be a multi-channel feature map. For example, the initial features of the image to be processed can be extracted, so as to convert the image to be processed from a three-channel image (such as an RGB image) into a multi-channel feature map.

[0139] The extraction of the initial features can be completed by using Figure 6 the preprocessing module 610 shown in the figure. In some embodiments, as Figure 9 shown in the figure, when receiving the image to be processed, the preprocessing module 610 can sequentially process the image to be processed by using the convolution operation 611 and the activation function 612, so as to obtain a multi-channel feature map. The convolution operation can be, for example, a 3×3 convolution operation, and the activation function can be, for example, a sigmoid function.

[0140] S820: Perform wavelet transform on the input feature map to obtain N feature sub-maps with different frequencies, where N is an integer greater than 1, and the value of N is associated with the level of the wavelet transform. The higher the level of the wavelet transform, the larger the value of N.

[0141] The wavelet transform can be a Haar wavelet transform, a Daubechies wavelet transform, etc., as long as multiple feature sub-maps with different frequencies can be obtained. In addition, the number of feature sub-maps obtained after wavelet transform in the embodiments of the present application is not limited either, and can be 2 feature sub-maps, 4 feature sub-maps or more feature sub-maps.

[0142] Taking Figure 10 as an example, the discrete wavelet transform (such as Haar wavelet transform) can be performed on the input feature map, so as to convert the input feature map into 4 feature sub-maps, namely the LL sub-map, the LH sub-map, the HL sub-map and the HH sub-map. Among them, the LL sub-map corresponds to the low-frequency information of the input feature map, the LH sub-map corresponds to the high-frequency information in the vertical direction, the HL sub-map corresponds to the high-frequency information in the horizontal direction, and the HH sub-map corresponds to the high-frequency information in the diagonal direction.

[0143] It should be understood that the resolution of the feature sub-map obtained after wavelet transform is smaller than the resolution of the input feature map. The resolution of the feature map is closely related to the overall computational amount of the neural network. For example, if the resolution of the feature map doubles, the computational amount may increase by 4 times. Therefore, in the embodiments of the present application, the input feature map is decomposed into multiple low-resolution feature sub-maps, and feature extraction is performed based on the low-resolution feature sub-maps, which can greatly reduce the computational amount.

[0144] S830: Use N feature extraction modules to respectively extract features from the N feature sub-maps to obtain N intermediate results.

[0145] In other words, the feature extraction processes of the N feature sub - graphs are carried out separately. The N feature sub - graphs can correspond one - to - one with N processing branches, and each processing branch is used to extract features from the feature sub - graph corresponding to this branch. Such a design mainly takes into account the differences between features of different frequencies, and separate processing can better extract the feature information of each frequency.

[0146] For example, as Figure 10 shown, the feature extraction module 1010, the feature extraction module 1020, the feature extraction module 1030, and the feature extraction module 1040 can be used to extract features from the HH sub - graph, the HL sub - graph, the LH sub - graph, and the LL sub - graph respectively, obtaining the first intermediate result, the second intermediate result, the third intermediate result, and the fourth intermediate result. The feature extraction module 1010, the feature extraction module 1020, the feature extraction module 1030, and the feature extraction module 1040 can be convolution modules.

[0147] In some embodiments, step S830 can be replaced by: using N sub - branches to parallelly extract features from the N feature sub - graphs, obtaining N intermediate results.

[0148] For example, as Figure 10 shown, the HH sub - branch can be used to extract features from the HH sub - graph, the HL sub - branch can be used to extract features from the HL sub - graph, the LH sub - branch can be used to extract features from the LH sub - graph, and the LL sub - branch can be used to extract features from the LL sub - graph.

[0149] S840: Perform inverse wavelet transform on the N intermediate results to obtain the first feature map.

[0150] For example, as Figure 10 shown, the first intermediate result, the second intermediate result, the third intermediate result, and the fourth intermediate result can be aggregated together, and then inverse wavelet transform (such as inverse Haar wavelet transform) is performed to obtain the first feature map. The resolution of the first feature map can be the same as that of the input feature map.

[0151] S850: Generate the enhanced image of the image to be processed according to the first feature map.

[0152] The processing method of step S850 is related to the type of the image to be processed. Taking the image to be processed as a low - resolution image to be super - resolved as an example, upsampling and convolution operations can be performed on the first feature map to obtain the super - resolved high - resolution image. Taking the image to be processed as a foggy or noisy image as an example, convolution operations can be performed on the first feature map to obtain a defogged or denoised image. The specific implementation method of step S850 can be referred to in the following Figure 17 and Figure 18 and related descriptions.

[0153] In the embodiments of the present application, feature extraction is performed based on low-resolution feature subgraphs, and the frequency differences between different feature subgraphs are considered, and multiple feature subgraphs are processed separately. On the one hand, this can reduce the computational complexity in the image processing process, and on the other hand, it can improve the image enhancement quality.

[0154] See Figures 11 - 12 . The N feature subgraphs include a first feature subgraph and a second feature subgraph. The N feature extraction modules include a first feature extraction module corresponding to the first feature subgraph and a second feature extraction module corresponding to the second feature subgraph. As Figure 11 shown, Figure 8 step S830 in

[0155] S8310: Use the first feature extraction module to perform feature extraction on the first feature subgraph to obtain a first intermediate result;

[0156] S8320: Perform a splicing operation on the first intermediate result and the second feature subgraph to obtain a second feature subgraph fused with the first intermediate result;

[0157] S8330: Use the second feature extraction module to perform feature extraction on the second feature subgraph fused with the first intermediate result to obtain a second intermediate result.

[0158] In some embodiments, before fusing with the first intermediate result, the second feature subgraph can be preprocessed. For example, a convolution operation can be first performed on the second feature subgraph, and then the convolved second feature subgraph is spliced with the first intermediate result.

[0159] In the embodiments of the present application, before using the second feature extraction module to perform feature extraction on the second feature subgraph, the first intermediate result obtained by the first feature extraction module is fused with the second feature subgraph, which is equivalent to using the information in a subgraph of one frequency to guide and constrain the feature extraction process of a subgraph of another frequency. This can ensure the accuracy of feature extraction and thus improve the image enhancement quality.

[0160] The embodiments of the present application do not specifically limit the relationship between the frequencies of the first feature subgraph and the second feature subgraph.

[0161] In some embodiments, the frequency of the second feature subgraph can be lower than the frequency of the first feature subgraph. In other words, the first feature subgraph contains high-frequency information, and the second feature subgraph contains low-frequency information. High-frequency information usually corresponds to the detailed information of the image, and low-frequency information usually corresponds to the structural information of the image. The two are correlated. Using high-frequency information to guide the feature extraction process of low-frequency information can improve the efficiency of feature extraction. In addition, in the image super-resolution task, integrating high-frequency information into low-frequency information can obtain a super-resolution image with richer details.

[0162] Alternatively, in some other embodiments, the frequency of the second feature sub - graph may also be higher than that of the first feature sub - graph. That is to say, sometimes, the feature extraction process of using low - frequency information to guide high - frequency information can also be adopted. For example, noise belongs to high - frequency information. When performing noise reduction processing on the image to be processed, adopting the method of using low - frequency information to guide high - frequency information can achieve a better noise reduction effect.

[0163] The first feature sub - graph and the second feature sub - graph can be any two feature sub - graphs with different frequencies among the multiple feature sub - graphs obtained after wavelet transform. For example, the first feature sub - graph and the second feature sub - graph can be two adjacent - frequency feature sub - graphs among the multiple feature sub - graphs. Figure 13 For example, the first feature sub - graph can be the HH sub - graph, and the second feature sub - graph can be the HL sub - graph; or, the first feature sub - graph can be the HL sub - graph, and the second feature sub - graph can be the LH sub - graph; or, the first feature sub - graph can be the LH sub - graph, and the second feature sub - graph can be the LL sub - graph.

[0164] Optionally, in some embodiments, the N feature sub - graphs include a third feature sub - graph, the third feature sub - graph is any one of the N feature sub - graphs, the N feature sub - graphs further include M feature sub - graphs with frequencies greater than that of the third feature sub - graph, the N feature extraction modules include a third feature extraction module corresponding to the third feature sub - graph and M feature extraction modules corresponding to the M feature sub - graphs, and M is an integer greater than 1. Figure 8 The step S830 in

[0165] S8311: Use the M feature extraction modules to extract features from the M feature sub - graphs respectively to obtain M intermediate results;

[0166] S8321: Perform a splicing operation on the M intermediate results and the third feature sub - graph to obtain a third feature sub - graph fused with the M intermediate results;

[0167] S8331: Use the third feature extraction module to extract features from the third feature sub - graph fused with the M intermediate results to obtain a third intermediate result.

[0168] Fusing the intermediate results corresponding to the M feature sub - graphs with frequencies greater than that of the third feature sub - graph into the third feature sub - graph can make the high - frequency information in the third feature sub - graph richer. High - frequency information represents detailed information, and making full use of the detailed information can improve the quality of image enhancement.

[0169] For Figure 14For example, if the third feature sub - graph is the LL sub - graph, then the M feature sub - graphs may include the HH sub - graph, the HL sub - graph, and the LH sub - graph. Or, if the third feature sub - graph is the LH sub - graph, then the M feature sub - graphs may include the HH sub - graph and the HL sub - graph.

[0170] Of course, the fusion method between frequency information is not limited to the method described above, and it can also be determined after comprehensively considering the computational complexity and the fusion effect. For example, the fusion of different frequency information can also be carried out in the way as Figure 15 shown, that is, the first intermediate result output by the HH sub - branch is fused in the HL sub - graph; the first intermediate result output by the HH sub - branch is fused in the LH sub - graph; the first intermediate result output by the HH sub - branch, the second intermediate result output by the HL sub - branch, and the third intermediate result output by the LH sub - branch are fused in the LL sub - graph. That is to say, the HH high - frequency information, the HL high - frequency information, and the LH high - frequency information can all guide the LL low - frequency information. In addition, the HH high - frequency information can also guide the other two high - frequency information (the HL high - frequency information and the LH high - frequency information).

[0171] The above has introduced in detail the feature fusion method of high and low frequencies. In order to further improve the quality of image enhancement, the features of high and low resolutions in the image to be processed can also be fused. The following will describe this in detail in combination with specific embodiments.

[0172] Compared with the resolution of the feature sub - graph obtained by wavelet transform, the input feature map is a high - resolution feature map. Therefore, the high - resolution feature information is retained in the input feature map.

[0173] Therefore, in some embodiments, the enhanced image of the image to be processed can be generated according to the first feature map and the input feature map. For example, the first feature map and the input feature map can be fused (or superimposed) to obtain a fused feature map; then, based on the fused feature map, the enhanced image of the image to be processed is generated. Take Figure 16 as an example. First, after discrete wavelet transform, the HH sub - graph, the HL sub - graph, the HL sub - graph, and the LH sub - graph are obtained. Then, feature extraction is performed on these 4 feature sub - graphs, and the result after feature extraction is subjected to inverse wavelet transform to obtain the first feature map. Next, the input feature map and the first feature map are fused to obtain a fused feature map. Subsequently, image enhancement can be performed based on the fused feature map.

[0174] Retaining the input feature map is equivalent to retaining the high - resolution feature information in the image to be processed. There is a large amount of detailed information in the high - resolution feature information, which can significantly improve the quality of image enhancement.

[0175] As mentioned above, the image to be processed can be super-resolved based on the first feature map, or denoised and dehazed based on the first feature map. The following will, in conjunction with Figure 17 and Figure 18 , give examples of these two processing methods respectively.

[0176] Figure 17 is an example diagram of super-resolution processing based on the first feature map (the first feature map can be the first feature map after fusing the input feature map). Refer to Figure 17 . The upsampling module 1710 and the convolution module 1720 can be used to perform upsampling operations and convolution operations on the first feature map respectively; the upsampling module 1730 is used to perform an upsampling operation on the image to be processed (a low-resolution RGB image); then the processing results of the two processing branches are fused (superimposed) to obtain a high-resolution RGB image.

[0177] Figure 18 is an example diagram of denoising or dehazing processing based on the first feature map (the first feature map can be the first feature map after fusing the input feature map). Refer to Figure 18 . The convolution module 1810 can be used to directly perform a convolution operation on the first feature map to obtain a denoised or dehazed RGB image.

[0178] The following will, in conjunction with Figure 19 and Figure 20 , take the discrete Haar wavelet transform as an example to give a more specific example. It should be understood that Figure 19 's processing flow is based on Figure 20 's provided model structure, and the two correspond to each other. The following will, in conjunction with Figure 20 describe each step in Figure 19 in detail.

[0179] S1910: The color three-channel image to be processed is converted into a multi-channel input feature map through a 3×3 convolution operation and an activation function processing by the preprocessing module 610.

[0180] The resolution of the multi-channel input feature map is the same as that of the color three-channel (RGB) image to be processed, both being H×W.

[0181] S1920: The multi-channel input feature map is subjected to discrete Haar wavelet transform by the wavelet transform (DWT) module 621 to obtain 4 feature submaps with the same resolution and different frequencies, namely HH, HL, LH, and LL submaps, corresponding to 4 sub-branches respectively.

[0182] The resolutions of the HH, HL, LH, and LL submaps are all H / 2×W / 2.

[0183] S1930: Perform a convolution operation on the HH sub - graph using the convolution module 622 to obtain a first intermediate result, and output the first intermediate result to the HL sub - branch and the inverse wavelet transform (IWT) module 626.

[0184] S1940: After performing a 3×3 convolution operation and activation function processing on the HL sub - graph, concatenate it with the first intermediate result to form a first multi - frequency feature. Perform a convolution operation on the first multi - frequency feature using the convolution module 623 to obtain a second intermediate result, and output the second intermediate result to the LH sub - branch and the inverse wavelet transform module 626.

[0185] S1950: After performing a 3×3 convolution operation and activation function processing on the LH sub - graph, concatenate it with the second intermediate result to form a second multi - frequency feature. Perform a convolution operation on the second multi - frequency feature using the convolution module 624 to obtain a third intermediate result, and output the third intermediate result to the LL sub - branch and the inverse wavelet transform module 626.

[0186] S1960: After performing a 3×3 convolution operation and activation function processing on the LL sub - graph, concatenate it with the third intermediate result to form a third multi - frequency feature. Perform a convolution operation on the third multi - frequency feature using the convolution module 625 to obtain a fourth intermediate result, and output the fourth intermediate result to the inverse wavelet transform module 626.

[0187] S1970: Use the inverse wavelet transform module 626 to perform an inverse wavelet transform on the first intermediate result, second intermediate result, third intermediate result, and fourth intermediate result output from the HH, HL, LH, and LL sub - branches to obtain a multi - channel first feature map with a resolution of H×W.

[0188] S1980: After performing a 3×3 convolution operation and activation function processing on the multi - channel input feature map, superimpose it on the multi - channel first feature map to obtain a multi - channel superimposed feature map. Perform a convolution operation on the multi - channel superimposed feature map using the convolution module 627 to obtain a multi - channel output feature map with a resolution of H×W.

[0189] S1990: According to different image enhancement task requirements, use the post - processing module 630 to convert the multi - channel output feature map into a color three - channel enhanced image, that is, obtain the enhanced image of the image to be processed.

[0190] It should be noted that the structures of the above - mentioned convolution modules 622, 623, 624, 625, and 627 can be as Figure 21 shown, and respectively include a 3×3 convolution layer 2110, an activation function 2120, a residual block 2130, and a 3×3 convolution layer 2140. Among them, as Figure 22 shown, the residual block 2130 can include two 3×3 convolution layers 2131, 2133, and an activation function 2132.

[0191] Based on the low-resolution HH, HL, LH, and LL sub-images, this technical solution extracts features and considers the frequency differences between different feature sub-images, separately processing the HH, HL, LH, and LL sub-images. On the one hand, this can reduce the computational complexity in the image processing process, and on the other hand, it can improve the image enhancement quality. Additionally, the intermediate results corresponding to the higher-frequency feature sub-images are successively fused into the lower-frequency feature sub-images, making the high-frequency information in the lower-frequency feature sub-images more abundant. High-frequency information represents detailed information, and the full utilization of detailed information can enhance the image enhancement quality. Meanwhile, retaining the high-resolution input feature map is equivalent to retaining the high-resolution feature information in the image to be processed, which can further improve the image enhancement quality.

[0192] The device embodiments of the present application will be described below. Since the device embodiments can execute the above methods, the parts not described in detail can be referred to the previous method embodiments.

[0193] Figure 23 The following is a schematic structural diagram of an image enhancement device provided by an embodiment of the present application. The image enhancement device 2300 includes: an acquisition module 2310, a wavelet transform module 2320, N feature extraction modules 2330, an inverse wavelet transform module 2340, and a generation module 2350. The image enhancement device 2300 can be used to execute Figure 8 the method described in Figure 8 For example, the acquisition module 2310, the wavelet transform module 2320, the N feature extraction modules 2330, the inverse wavelet transform module 2340, and the generation module 2350 can be respectively used to execute

[0194] The acquisition module 2310 is used to acquire the input feature map of the image to be processed.

[0195] The wavelet transform module 2320 is used to perform wavelet transform on the input feature map to obtain N feature sub-images with different frequencies. N is an integer greater than 1, and the value of N is associated with the level of the wavelet transform. The higher the level of the wavelet transform, the larger the value of N.

[0196] The N feature extraction modules 2330 are respectively used to extract features from the N feature sub-images to obtain N intermediate results.

[0197] The inverse wavelet transform module 2340 is used to perform inverse wavelet transform on the N intermediate results to obtain the first feature map.

[0198] The generation module 2350 is configured to generate an enhanced image of the image to be processed according to the first feature map.

[0199] Feature extraction is performed based on the low-resolution feature submaps, and considering the frequency differences between different feature submaps, multiple feature submaps are processed separately. On the one hand, this can reduce the computational complexity in the image processing process, and on the other hand, it can improve the image enhancement quality.

[0200] Optionally, in some embodiments, the N feature submaps include a first feature submap and a second feature submap. The first feature submap and the second feature submap are any two feature submaps with different frequencies among the N feature submaps. The N feature extraction modules include a first feature extraction module corresponding to the first feature submap and a second feature extraction module corresponding to the second feature submap. The image enhancement device further includes a splicing module 2360. The first feature extraction module is configured to perform feature extraction on the first feature submap to obtain a first intermediate result. The splicing module 2360 is configured to perform a splicing operation on the first intermediate result and the second feature submap to obtain a second feature submap fused with the first intermediate result. The second feature extraction module is configured to perform feature extraction on the second feature submap fused with the first intermediate result to obtain a second intermediate result.

[0201] Optionally, in some embodiments, the frequency of the second feature submap is lower than the frequency of the first feature submap.

[0202] Optionally, in some embodiments, the N feature submaps include a third feature submap. The third feature submap is any one of the N feature submaps. The N feature submaps further include M feature submaps with frequencies greater than the third feature submap. The N feature extraction modules include a third feature extraction module corresponding to the third feature submap and M feature extraction modules corresponding to the M feature submaps. M is an integer greater than 1. The image enhancement device further includes a splicing module 2360. The M feature extraction modules are respectively configured to perform feature extraction on the M feature submaps to obtain M intermediate results. The splicing module 2360 is configured to perform a splicing operation on the M intermediate results and the third feature submap to obtain a third feature submap fused with the M intermediate results. The third feature extraction module is configured to perform feature extraction on the third feature submap fused with the M intermediate results to obtain a third intermediate result.

[0203] Optionally, in some embodiments, the generation module 2350 is further configured to generate an enhanced image of the image to be processed according to the first feature map and the input feature map.

[0204] Figure 24 Schematic diagram of the hardware structure of an image enhancement device provided in an embodiment of the present application. Figure 24 The shown image enhancement device 2400 (the device 2400 may specifically be a computer device) includes a memory 2410, a processor 2420, a communication interface 2430, and a bus 2440. Among them, the memory 2410, the processor 2420, and the communication interface 2430 are communicatively connected to each other through the bus 2440.

[0205] The memory 2410 may be a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 2410 may store a program. When the program stored in the memory 2410 is executed by the processor 2420, the processor 2420 and the communication interface 2430 are used to execute each step of the image enhancement method of the embodiment of the present application.

[0206] The processor 2420 may adopt a general central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU), or one or more integrated circuits, and is used to execute relevant programs to implement the functions required to be executed by the units in the image enhancement device of the embodiment of the present application, or to execute the image enhancement method of the method embodiment of the present application.

[0207] The processor 2420 can also be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the image enhancement method of the present application can be completed by the integrated logic circuit of the hardware in the processor 2420 or the instructions in the form of software. The above-mentioned processor 2420 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 2410, and the processor 2420 reads the information in the memory 2410 and combines its hardware to complete the functions required to be executed by the units included in the image enhancement device in the embodiments of the present application, or execute the image enhancement method in the method embodiments of the present application.

[0208] The communication interface 2430 uses a transceiver device such as, but not limited to, a transceiver to implement the communication between the device 2400 and other devices or communication networks. For example, the to-be-processed image described in the embodiments of the present application can be obtained through the communication interface 2430.

[0209] The bus 2440 can include a path for transmitting information between various components of the device 2400 (for example, the memory 2410, the processor 2420, the communication interface 2430).

[0210] It should be understood that the acquisition module 2310 in the image enhancement device 2300 can be equivalent to the communication interface 2430. The wavelet transform module 2320, the N feature extraction modules 2330, the inverse wavelet transform module 2340, the generation module 2350, and the splicing module 2360 in the image enhancement device 2300 can be equivalent to the processor 2420.

[0211] It should be noted that although Figure 24The illustrated device 2400 only shows a memory, a processor, and a communication interface. However, in specific implementations, those skilled in the art should understand that device 2400 also includes other components necessary for normal operation. At the same time, according to specific needs, those skilled in the art should understand that device 2400 may also include hardware components for implementing other additional functions. In addition, those skilled in the art should understand that device 2400 may also only include the components necessary for implementing the embodiments of the present application, and not necessarily include Figure 24 all the components shown in

[0212] It can be understood that device 2400 is equivalent to Figure 5 the execution device 510 described therein. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0213] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0214] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0215] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0216] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0217] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0218] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. An image enhancement method, characterized in that, Including: Obtain an input feature map of the image to be processed; Perform wavelet transform on the input feature map to obtain N feature submaps with different frequencies, where N is an integer greater than 1, and the value of N is associated with the level of the wavelet transform. The higher the level of the wavelet transform, the larger the value of N; Use N feature extraction modules to respectively extract features from the N feature submaps to obtain N intermediate results; Perform inverse wavelet transform on the N intermediate results to obtain a first feature map; Generate an enhanced image of the image to be processed according to the first feature map; Among them, the N feature submaps include a third feature submap, and the third feature submap is any one of the N feature submaps. The N feature submaps further include M feature submaps with frequencies higher than the third feature submap. The N feature extraction modules include a third feature extraction module corresponding to the third feature submap and M feature extraction modules corresponding to the M feature submaps, where M is an integer greater than 1. The step of using N feature extraction modules to respectively extract features from the N feature submaps includes: Use the M feature extraction modules to respectively extract features from the M feature submaps to obtain M intermediate results; Perform a splicing operation on the M intermediate results and the third feature submap to obtain a third feature submap fused with the M intermediate results; Use the third feature extraction module to extract features from the third feature submap fused with the M intermediate results to obtain a third intermediate result.

2. The image enhancement method according to claim 1, characterized in that, The N feature submaps include a first feature submap and a second feature submap, and the first feature submap and the second feature submap are any two feature submaps with different frequencies among the N feature submaps. The N feature extraction modules include a first feature extraction module corresponding to the first feature submap and a second feature extraction module corresponding to the second feature submap. The step of using N feature extraction modules to respectively extract features from the N feature submaps includes: Use the first feature extraction module to extract features from the first feature submap to obtain a first intermediate result; Perform a splicing operation on the first intermediate result and the second feature submap to obtain a second feature submap fused with the first intermediate result; Use the second feature extraction module to extract features from the second feature submap fused with the first intermediate result to obtain a second intermediate result.

3. The image enhancement method according to claim 2, characterized in that, The frequency of the second feature submap is lower than the frequency of the first feature submap.

4. The image enhancement method according to any one of claims 1-3, characterized in that, The step of generating an enhanced image of the image to be processed according to the first feature map includes: Generate an enhanced image of the image to be processed according to the first feature map and the input feature map.

5. The image enhancement method according to any one of claims 1-4, characterized in that, The step of generating an enhanced image of the image to be processed according to the first feature map includes: Perform super-resolution processing on the image to be processed according to the first feature map to obtain an enhanced image of the image to be processed.

6. The image enhancement method according to any one of claims 1-4, characterized in that, The step of generating an enhanced image of the image to be processed according to the first feature map includes: Perform noise reduction or defogging processing on the image to be processed according to the first feature map to obtain an enhanced image of the image to be processed.

7. The image enhancement method according to any one of claims 1-6, characterized in that, The wavelet transform is a Haar wavelet transform.

8. An image enhancement device, characterized in that, It includes: An acquisition module, configured to acquire an input feature map of an image to be processed; A wavelet transform module, configured to perform wavelet transform on the input feature map to obtain N feature submaps with different frequencies, where N is an integer greater than 1, and the value of N is associated with the level of the wavelet transform. The higher the level of the wavelet transform, the larger the value of N; N feature extraction modules, respectively configured to perform feature extraction on the N feature submaps to obtain N intermediate results; An inverse wavelet transform module, configured to perform inverse wavelet transform on the N intermediate results to obtain a first feature map; A generation module, configured to generate an enhanced image of the image to be processed according to the first feature map; The N feature submaps include a third feature submap, where the third feature submap is any one of the N feature submaps, and the N feature submaps further include M feature submaps with frequencies higher than that of the third feature submap. The N feature extraction modules include a third feature extraction module corresponding to the third feature submap and M feature extraction modules corresponding to the M feature submaps, where M is an integer greater than 1. The image enhancement device further includes a splicing module, The M feature extraction modules are respectively configured to perform feature extraction on the M feature submaps to obtain M intermediate results; The splicing module is configured to perform a splicing operation on the M intermediate results and the third feature submap to obtain a third feature submap fused with the M intermediate results; The third feature extraction module is configured to perform feature extraction on the third feature submap fused with the M intermediate results to obtain a third intermediate result.

9. The image enhancement device according to claim 8, wherein, The N feature submaps include a first feature submap and a second feature submap, where the first feature submap and the second feature submap are any two feature submaps with different frequencies among the N feature submaps. The N feature extraction modules include a first feature extraction module corresponding to the first feature submap and a second feature extraction module corresponding to the second feature submap. The image enhancement device further includes a splicing module, The first feature extraction module is configured to perform feature extraction on the first feature submap to obtain a first intermediate result; The splicing module is configured to perform a splicing operation on the first intermediate result and the second feature submap to obtain a second feature submap fused with the first intermediate result; The second feature extraction module is configured to perform feature extraction on the second feature submap fused with the first intermediate result to obtain a second intermediate result.

10. The image enhancement device according to claim 9, wherein, The frequency of the second feature submap is lower than that of the first feature submap.

11. The image enhancement device according to any one of claims 8-10, wherein, The generation module is further configured to generate an enhanced image of the image to be processed according to the first feature map and the input feature map.

12. The image enhancement device according to any one of claims 8-11, wherein, The generation module is further configured to perform super-resolution processing on the image to be processed according to the first feature map to obtain an enhanced image of the image to be processed.

13. The image enhancement device according to any one of claims 8-11, wherein, The generation module is further configured to perform noise reduction or defogging processing on the image to be processed according to the first feature map to obtain an enhanced image of the image to be processed.

14. The image enhancement device according to any one of claims 8-13, wherein, The wavelet transform is a Haar wavelet transform.

15. An image enhancement device, wherein, It includes: A memory, configured to store programs; A processor for executing a program stored in the memory, and when the program stored in the memory is executed, the processor is used to execute the image enhancement method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Image denoising method and device

    CN109993707A

  • Image enhancement method and device and terminal equipment

    CN111047512A